Botanic1 is a state-of-the-art encoder of plant genomes
The ranking of a large panel of both generalist gLMs (gLMs trained on diverse organisms' genomes) and plant specialist gLMs trained exclusively on plants, with the per-family values reported below. Every value on this page is read from the same tables as the report's Figure 1d and the per-category scorecard.
The aggregate score
We measure over 34 per-(task, species) scalars, of which 22 enter Sbal. Averaging the 22 metrics directly would give artificially more importance to well-represented tasks, in particular because the number of species available per task depends on how the benchmark was created and should be irrelevant: chromatin accessibility alone contributes seven of the 22 tasks, and A. thaliana recurs in six families as this species is an emblematic test species in plants which has been annotated extensively. We therefore report Sbal, which first averages species within each of the nine task families and then averages the nine families, so that every specific biological capability enters with equal weight:
Sbal = (1/9) Ī£f=1..9 (1/|š®f|) Ī£sāš®f mf,s
where š®f is the set of species evaluated for family f and mf,s the corresponding metric. Seven are supervised probes trained on frozen representations; the remaining two are zero-shot and score variants directly from the model's likelihood. Every leaderboard value and model comparison in the article uses Sbaltest, which scores splicing on centre-token embeddings with the prevalence-corrected stratified AUPRC and, for auto-regressive models, performs two forward passes, one on the normal strand and one on its reverse complement.
Model ranking
All four Botanic1 models rank above the strongest competitor, PlantCAD2-L (aggregate score Sbaltest = 0.756), and their performance scales monotonically with model size, from Botanic1-S (0.758) through Botanic1-M (0.765), Botanic1-L (0.768) to Botanic1-XL (0.769). The aggregate-score axis starts at 0.5, so bar length is not proportional to the score. Task views use the full score range.
Hover a bar for the nine family means. Whiskers are 95% confidence intervals from the per-sample paired bootstrap of the technical report (20,000 replicates), centred on the reported score. Botanic0-L is the previous generation, shown in olive. Parameter counts are trainable parameters as declared by each model's configuration and code; PlantBiMoE has 116M trainable parameters and 64M active per token.
The nine task families
Zooming in per task shows that Botanic1 leads overall without dominating every capability: the flagship Botanic1-XL tops the benchmark and two task families, genomic-region classification (0.527 versus 0.525 for PlantCAD2-L) and causal-variant discovery (recall AUC 0.752), Botanic1-L leads three more (conservation, translation initiation and PRO-seq) and Botanic1-M leads splicing (donor and acceptor mean 0.984 versus 0.978 for NTv3-650M-post, which uses test annotations during post-training). The autoregressive Evo 2 leads variant-effect prediction, where its 7B model reaches an LLR AUROC of 0.724 against 0.721 for Botanic1-XL (its 40B model reaches 0.719); PlantCAD2-L leads on chromatin accessibility (0.483 versus 0.481 for Botanic1-L); and GPN leads translation termination (0.888 versus 0.887). Botanic1-XL is the only model that stays competitive across all nine families.
Size and training tokens
Plotting the aggregated score against model size highlights the efficiency of this frontier: the Botanic family leads at every scale, and even Botanic1-S (318M parameters) outperforms baselines with more than an order of magnitude larger, including Carbon-8B (8.3B, 0.727) and Evo 2 (7B, 0.695), both margins significant under both bootstrap tests; the most parameter-efficient baseline, GPN (66M, 0.731), still trails on performance. The right panel replots the same models against the total number of genomic base pairs seen during training: the tokens seen, reconstructed for every model from its publication or model card. The Botanic scores are attained after only 314.6B base pairs, where the strongest competitor, PlantCAD2-L, had to process 4.03T base-pairs and NTv3 12.1T. Interestingly, the Evo 2 family does not scale well with capacity: it peaks at 7B (0.695, from 0.670 at 1B) and declines through 20B (0.678) to 40B (0.629).
Vertical bars are 95% confidence intervals from the per-sample bootstrap. Tokens are converted to base pairs using approximations for 6-mer models. Models without a documented token count are omitted from the right panel.
Are leaderboard differences significant?
We assess leaderboard differences with two paired bootstrap tests, each using 20,000 replicates: one resampling the nine task families, and one resampling test examples within each benchmark cell. Together, they test whether observed margins are robust to variation across task families and to sampling noise in the benchmark. The technical report (v2) reports both tests: the table below reproduces Supplementary Table S7, and the per-sample intervals are drawn as error bars in Figure 1d and as paired differences to PlantCAD2-L in Figure 1e.
The key result is that Botanic1-S performs on par with the much larger PlantCAD2-L: its score is slightly higher (+0.002), but the confidence intervals include zero, so we cannot resolve a significant difference between the two. In contrast, Botanic1-L and Botanic1-XL significantly outperform PlantCAD2-L under both tests. Botanic1-M (+0.010) is significant under the per-sample test but not under the family-level test, and shows a significant gain over Botanic1-S (+0.007) under both. Botanic1-S also significantly outperforms Carbon-8B and Evo 2-7B under both tests.
Overall, the results show that Botanic1 is already highly competitive at small scale, with larger variants delivering statistically supported gains over strong baselines.
| Model | Reference | Ī Sbal | 95% CI, families | 95% CI, examples |
|---|---|---|---|---|
| Botanic1-S | PlantCAD2-L | +0.002 | [ā0.005, +0.010] | [ā0.007, +0.013] |
| Botanic1-M | PlantCAD2-L | +0.010 | [ā0.001, +0.024] | [+0.000, +0.021]* |
| Botanic1-L | PlantCAD2-L | +0.012 | [+0.001, +0.028]* | [+0.003, +0.024]* |
| Botanic1-XL | PlantCAD2-L | +0.014 | [+0.002, +0.031]* | [+0.006, +0.024]* |
| Botanic1-XL | PlantCaduceus-l32 | +0.017 | [+0.010, +0.026]* | [+0.011, +0.024]* |
| Botanic1-M | Botanic1-S | +0.007 | [+0.002, +0.015]* | [+0.002, +0.013]* |
| Botanic1-L | Botanic1-M | +0.003 | [ā0.001, +0.006] | [ā0.002, +0.007] |
| Botanic1-XL | Botanic1-L | +0.002 | [ā0.003, +0.006] | [ā0.005, +0.008] |
What each family of tasks evaluates
| Family | Task description | Species | Metric |
|---|---|---|---|
| Supervised probes | |||
| Chromatin accessibility | open versus closed chromatin | 7 | macro-AUPRC |
| Genomic region | genic compartment of a window | 4 | balanced acc. |
| Conservation | evolutionarily constrained positions | 2 | AUPRC |
| Splicing | donor and acceptor site recognition | 1 | strat. AUPRC (Sbaltest) |
| TIS | translation initiation site | 1 | AUPRC |
| TTS | translation termination site | 1 | AUPRC |
| PRO-seq | nascent transcription | 1 | AUPRC |
| Zero-shot | |||
| LLR mutation effect | deleterious versus benign variants | 3 | AUROC |
| Causal variant discovery | experimentally supported variant recovery (29 studies) | 4 | recall-AUC |