Interpretability
Sparse features are readable genomic concepts
A sparse autoencoder decomposes one layer of a frozen Botanic1-M into 8,192 features, fitted with no genomic annotation at all. Single features nonetheless detect splice donors and acceptors at MCC 0.871 and 0.873, intron bodies, and each of the three positions in a codon. Two of them read zero where TAIR10 and Araport11 annotate a splice donor — and in 2026 TAIR12 moved that donor to where they fire.
The browser below reads that dictionary out base by base over 21 precomputed windows. Every value drawn is a measurement; the annotation rows are a separate, paler register and never an input to a number on this page.
Choose a window
Each window is 8,191 bp, the model's native context, and each was chosen because it shows something. Ten are windows where the features track the annotation closely: over 384 windows scanned, the splice features fire at a median 95% of the annotated donors and acceptors, and these are the clearest of them. Eight are disagreements — five where a feature marks a splice site the canonical transcript drops but another annotated isoform uses, three where it marks a site no annotated transcript uses at all. Each of those eight also carries a verdict from a register independent of the annotation, and one of the three is a site an RNA-seq atlas measures in use. Two are loci where the annotation itself moved between releases. The last is the AT1G65170 locus the report's annotation-correction result is drawn from.
A released annotation says which splice sites someone drew a transcript through. It does not say which sites are used, and the two questions come apart exactly where a feature disagrees with the annotation. So each of the eight disagreements carries a second register, shown under the browser: for Arabidopsis, AtRTD3's 169,503 long-read transcripts, the 2026 TAIR12 reannotation, and PastDB's usage measurements over 202 RNA-seq samples, all three on the same TAIR10 coordinates the windows are cut from.
The three verdicts are different claims, and only one of them is a discovery. Used means an independent register records the site in use: chr1:21,735,155, which no transcript model carries, is an alternative 3′ splice site in PastDB running at up to 7.5% against the dominant site's 99.8%, and both Arabidopsis alternative-isoform sites are ones TAIR12 keeps while dropping the canonical site the feature reads near zero at. Untestable means the registers cover the locus and still cannot answer: chr5:5,427,495 sits 23 bp upstream of the start codon of a gene with one transcript, no annotated 5′ UTR and a median expression of 0.00 cRPKM, so no read can confirm or refute it. Untested means no register was consulted at all — the rice and maize disagreements, where no long-read transcriptome or splicing atlas on those assemblies was checked. A feature firing where nothing is annotated is not evidence of a discovery until a register says so, and it is not presented as one here.
Window
A genomic window, feature by feature
How to read this browser
Hover a base to inspect it; click or press Enter to pin it, Escape to unpin. Drag to pan, pinch or + and − to zoom, arrow keys to step base by base (hold Shift to stride), Home for the whole window. Top features opens a movable panel ranking the strongest features at the pinned base.
Each window opens on one lane per genomic category — splice donor, splice acceptor, intron, coding sequence, the three codon positions, and whichever other annotated element the dictionary marks best here. Each lane is the feature that matches that category best in this window, searched over every feature the window stores rather than a shortlist, scored as Matthews correlation against the annotated category. The threshold is the best of six percentiles of the feature's own values in the window, so the score describes the picture in front of you and is not an out-of-sample estimate. Codon positions are derived from the canonical transcripts' coding sequence in transcription order, because the label tracks carry region and site classes but not reading frame. Top features ranks the strongest features at the pinned base out of all 8,192 and any of them can be given a lane: this dictionary keeps 64 features per base and zeroes the rest, and each window carries a complete track for every feature that reaches a base's top eight anywhere in it.
Select a window to read its features base by base.
How to read the tracks
- Top strip. Every base takes the colour of the feature that fires hardest on it, and stays grey where nothing fires. Base letters appear once the view is narrow enough to hold them.
- Annotation rows. Region bands, single-base site ticks and gene extents, one set per strand, in the paler register. They come from the species' released annotation, never from the model.
- Feature lanes. One lane per shown feature. A stem's height is the activation at that base as a fraction of that feature's peak in this window, so lanes are scaled independently; the gutter carries each lane's maximum in view. A blank base is a zero — the feature did not fire.
- Dashed outline. The anchor: the position the window sampler selected on. It marks where we looked, not an annotation of that base.
- Readout. Under the plot: the coordinate, the base, the annotation on both strands, every shown feature firing there, and the strongest feature of all 8,192.
The twelve named features
The dictionary carries no names: nothing upstream assigns a feature a meaning. These
twelve have one because the report measures what they detect. The two splice features
and the intron feature are the best detector of their annotation class in every one of the
four annotated genomes, at MCC 0.821 to 0.937 for the splice classes and 0.650 to 0.741
for intron, so the naming does not rest on one genome. f4131 and f2634 are also the pair a
two-level decision tree selects in all five chromosome folds, sorting a candidate
GT into decoy, intermediate-usage or constitutive donor at a balanced accuracy
of 0.841 [0.827, 0.855].
Two features the report names are deliberately absent here. Asked for a transcription start site or a termination site, the best feature in the dictionary reaches MCC 0.035 and 0.029 — consistent with neither site having a sharp sequence consensus in plants. They are a real negative result, and a lane for either would mark nothing, so neither is offered.
Annotation sources
The annotation register is read from each species' released annotation, and the browser names the assembly of every window it draws.
The report's interpretability section carries the methodology, the full figure and the annotation history of the locus: Botanic1 technical report ↗