BOTANIC?
Living Models

Interpretability

Sparse features are readable genomic concepts

A sparse autoencoder decomposes one layer of a frozen Botanic1-M into 8,192 features, fitted with no genomic annotation at all. Single features nonetheless detect splice donors and acceptors at MCC 0.871 and 0.873, intron bodies, and each of the three positions in a codon. Two of them read zero where TAIR10 and Araport11 annotate a splice donor — and in 2026 TAIR12 moved that donor to where they fire.

The browser below reads that dictionary out base by base over 21 precomputed windows. Every value drawn is a measurement; the annotation rows are a separate, paler register and never an input to a number on this page.

Choose a window

Each window is 8,191 bp, the model's native context, and each was chosen because it shows something. Ten are windows where the features track the annotation closely: over 384 windows scanned, the splice features fire at a median 95% of the annotated donors and acceptors, and these are the clearest of them. Eight are disagreements — five where a feature marks a splice site the canonical transcript drops but another annotated isoform uses, three where it marks a site no annotated transcript uses at all. Each of those eight also carries a verdict from a register independent of the annotation, and one of the three is a site an RNA-seq atlas measures in use. Two are loci where the annotation itself moved between releases. The last is the AT1G65170 locus the report's annotation-correction result is drawn from.

A released annotation says which splice sites someone drew a transcript through. It does not say which sites are used, and the two questions come apart exactly where a feature disagrees with the annotation. So each of the eight disagreements carries a second register, shown under the browser: for Arabidopsis, AtRTD3's 169,503 long-read transcripts, the 2026 TAIR12 reannotation, and PastDB's usage measurements over 202 RNA-seq samples, all three on the same TAIR10 coordinates the windows are cut from.

The three verdicts are different claims, and only one of them is a discovery. Used means an independent register records the site in use: chr1:21,735,155, which no transcript model carries, is an alternative 3′ splice site in PastDB running at up to 7.5% against the dominant site's 99.8%, and both Arabidopsis alternative-isoform sites are ones TAIR12 keeps while dropping the canonical site the feature reads near zero at. Untestable means the registers cover the locus and still cannot answer: chr5:5,427,495 sits 23 bp upstream of the start codon of a gene with one transcript, no annotated 5′ UTR and a median expression of 0.00 cRPKM, so no read can confirm or refute it. Untested means no register was consulted at all — the rice and maize disagreements, where no long-read transcriptome or splicing atlas on those assemblies was checked. A feature firing where nothing is annotated is not evidence of a discovery until a register says so, and it is not presented as one here.

Window

A genomic window, feature by feature

Select a window to read its features base by base.

Activations:

How to read the tracks

CDS 5′ UTR 3′ UTR intron non-coding intergenic TSS start codon donor acceptor stop TTS

The twelve named features

The dictionary carries no names: nothing upstream assigns a feature a meaning. These twelve have one because the report measures what they detect. The two splice features and the intron feature are the best detector of their annotation class in every one of the four annotated genomes, at MCC 0.821 to 0.937 for the splice classes and 0.650 to 0.741 for intron, so the naming does not rest on one genome. f4131 and f2634 are also the pair a two-level decision tree selects in all five chromosome folds, sorting a candidate GT into decoy, intermediate-usage or constitutive donor at a balanced accuracy of 0.841 [0.827, 0.855].

Two features the report names are deliberately absent here. Asked for a transcription start site or a termination site, the best feature in the dictionary reaches MCC 0.035 and 0.029 — consistent with neither site having a sharp sequence consensus in plants. They are a real negative result, and a lane for either would mark nothing, so neither is offered.

Evidence as published. Bold marks the three features that are the best detector of their annotation class in every one of the four annotated genomes.

Annotation sources

The annotation register is read from each species' released annotation, and the browser names the assembly of every window it draws.

The report's interpretability section carries the methodology, the full figure and the annotation history of the locus: Botanic1 technical report ↗