BOTANIC?
Living Models

Botanic1 is a state-of-the-art encoder of plant genomes

The ranking of a large panel of both generalist gLMs (gLMs trained on diverse organisms' genomes) and plant specialist gLMs trained exclusively on plants, with the per-family values reported below. Every value on this page is read from the same tables as the report's Figure 1d and the per-category scorecard.

The aggregate score

We measure over 34 per-(task, species) scalars, of which 22 enter Sbal. Averaging the 22 metrics directly would give artificially more importance to well-represented tasks, in particular because the number of species available per task depends on how the benchmark was created and should be irrelevant: chromatin accessibility alone contributes seven of the 22 tasks, and A. thaliana recurs in six families as this species is an emblematic test species in plants which has been annotated extensively. We therefore report Sbal, which first averages species within each of the nine task families and then averages the nine families, so that every specific biological capability enters with equal weight:

Sbal = (1/9) Ī£f=1..9 (1/|š’®f|) Ī£sāˆˆš’®f mf,s

where š’®f is the set of species evaluated for family f and mf,s the corresponding metric. Seven are supervised probes trained on frozen representations; the remaining two are zero-shot and score variants directly from the model's likelihood. Every leaderboard value and model comparison in the article uses Sbaltest, which scores splicing on centre-token embeddings with the prevalence-corrected stratified AUPRC and, for auto-regressive models, performs two forward passes, one on the normal strand and one on its reverse complement.

Model ranking

All four Botanic1 models rank above the strongest competitor, PlantCAD2-L (aggregate score Sbaltest = 0.756), and their performance scales monotonically with model size, from Botanic1-S (0.758) through Botanic1-M (0.765), Botanic1-L (0.768) to Botanic1-XL (0.769). The aggregate-score axis starts at 0.5, so bar length is not proportional to the score. Task views use the full score range.

Show
Sort by

Hover a bar for the nine family means. Whiskers are 95% confidence intervals from the per-sample paired bootstrap of the technical report (20,000 replicates), centred on the reported score. Botanic0-L is the previous generation, shown in olive. Parameter counts are trainable parameters as declared by each model's configuration and code; PlantBiMoE has 116M trainable parameters and 64M active per token.

The nine task families

Zooming in per task shows that Botanic1 leads overall without dominating every capability: the flagship Botanic1-XL tops the benchmark and two task families, genomic-region classification (0.527 versus 0.525 for PlantCAD2-L) and causal-variant discovery (recall AUC 0.752), Botanic1-L leads three more (conservation, translation initiation and PRO-seq) and Botanic1-M leads splicing (donor and acceptor mean 0.984 versus 0.978 for NTv3-650M-post, which uses test annotations during post-training). The autoregressive Evo 2 leads variant-effect prediction, where its 7B model reaches an LLR AUROC of 0.724 against 0.721 for Botanic1-XL (its 40B model reaches 0.719); PlantCAD2-L leads on chromatin accessibility (0.483 versus 0.481 for Botanic1-L); and GPN leads translation termination (0.888 versus 0.887). Botanic1-XL is the only model that stays competitive across all nine families.

Columns
Click a column header to sort by it. Shading is relative within each column. Hover a cell for its 95% confidence interval.

Size and training tokens

Plotting the aggregated score against model size highlights the efficiency of this frontier: the Botanic family leads at every scale, and even Botanic1-S (318M parameters) outperforms baselines with more than an order of magnitude larger, including Carbon-8B (8.3B, 0.727) and Evo 2 (7B, 0.695), both margins significant under both bootstrap tests; the most parameter-efficient baseline, GPN (66M, 0.731), still trails on performance. The right panel replots the same models against the total number of genomic base pairs seen during training: the tokens seen, reconstructed for every model from its publication or model card. The Botanic scores are attained after only 314.6B base pairs, where the strongest competitor, PlantCAD2-L, had to process 4.03T base-pairs and NTv3 12.1T. Interestingly, the Evo 2 family does not scale well with capacity: it peaks at 7B (0.695, from 0.670 at 1B) and declines through 20B (0.678) to 40B (0.629).

Vertical bars are 95% confidence intervals from the per-sample bootstrap. Tokens are converted to base pairs using approximations for 6-mer models. Models without a documented token count are omitted from the right panel.

Are leaderboard differences significant?

We assess leaderboard differences with two paired bootstrap tests, each using 20,000 replicates: one resampling the nine task families, and one resampling test examples within each benchmark cell. Together, they test whether observed margins are robust to variation across task families and to sampling noise in the benchmark. The technical report (v2) reports both tests: the table below reproduces Supplementary Table S7, and the per-sample intervals are drawn as error bars in Figure 1d and as paired differences to PlantCAD2-L in Figure 1e.

The key result is that Botanic1-S performs on par with the much larger PlantCAD2-L: its score is slightly higher (+0.002), but the confidence intervals include zero, so we cannot resolve a significant difference between the two. In contrast, Botanic1-L and Botanic1-XL significantly outperform PlantCAD2-L under both tests. Botanic1-M (+0.010) is significant under the per-sample test but not under the family-level test, and shows a significant gain over Botanic1-S (+0.007) under both. Botanic1-S also significantly outperforms Carbon-8B and Evo 2-7B under both tests.

Overall, the results show that Botanic1 is already highly competitive at small scale, with larger variants delivering statistically supported gains over strong baselines.

ModelReferenceĪ” Sbal95% CI, families95% CI, examples
Botanic1-SPlantCAD2-L+0.002[āˆ’0.005, +0.010][āˆ’0.007, +0.013]
Botanic1-MPlantCAD2-L+0.010[āˆ’0.001, +0.024][+0.000, +0.021]*
Botanic1-LPlantCAD2-L+0.012[+0.001, +0.028]*[+0.003, +0.024]*
Botanic1-XLPlantCAD2-L+0.014[+0.002, +0.031]*[+0.006, +0.024]*
Botanic1-XLPlantCaduceus-l32+0.017[+0.010, +0.026]*[+0.011, +0.024]*
Botanic1-MBotanic1-S+0.007[+0.002, +0.015]*[+0.002, +0.013]*
Botanic1-LBotanic1-M+0.003[āˆ’0.001, +0.006][āˆ’0.002, +0.007]
Botanic1-XLBotanic1-L+0.002[āˆ’0.003, +0.006][āˆ’0.005, +0.008]
Paired differences in reported Sbaltest between the model and the reference. Both intervals are 95% percentile intervals over 20,000 paired replicates: "families" resamples the nine task families, "examples" resamples the test examples within each of the 22 tasks and species, and is the test that speaks to sampling noise in the benchmark. An asterisk marks an interval that excludes zero.

What each family of tasks evaluates

FamilyTask descriptionSpeciesMetric
Supervised probes
Chromatin accessibilityopen versus closed chromatin7macro-AUPRC
Genomic regiongenic compartment of a window4balanced acc.
Conservationevolutionarily constrained positions2AUPRC
Splicingdonor and acceptor site recognition1strat. AUPRC (Sbaltest)
TIStranslation initiation site1AUPRC
TTStranslation termination site1AUPRC
PRO-seqnascent transcription1AUPRC
Zero-shot
LLR mutation effectdeleterious versus benign variants3AUROC
Causal variant discoveryexperimentally supported variant recovery (29 studies)4recall-AUC
Each family contributes one ninth of the score irrespective of how many species it is evaluated on. Supervised probes (XGBoost) are trained on frozen representations; the two zero-shot families score variants using the model's likelihood directly. Causal variant discovery here is scored on a fixed subset of 29 curated studies from four species; the full 545-study benchmark is reported on its own page.