BOTANIC?
Living Models

Zero-shot causal variant prioritisation

A new and harder zero-shot causal-variant discovery benchmark based on experimentally validated variants, totalling 545 loci across 14 species. Each task ranks one documented causal SNP among population SNPs in its surrounding 100 kbp locus.

ABI5 locus

Arabidopsis thaliana · TAIR10 · Botanic1-XL

Loading locus…

Hover or tap a SNP to inspect. Drag to pan; + / − to zoom. Arrow keys select SNPs.Measured SNP scores · 4,096 bp context · Araport11 gene models ↗

Our newly introduced causal-variant benchmark uses experimentally validated functional variants as a source of ground truth, instead of relying on indirect proxies such as allele frequency or variant class. Every language model is scored at min(4,096, trained context) (PlantCaduceus-l32 and GPN are fixed at 512 bp) for each LLR, putting the SNV in the middle of the window. All gLMs are evaluated on the full dataset from the reference genome sequences only, while bioinformatics baselines require either precise annotations for Ensembl VEP or MSA support for conservation scores. As a result, PhyloP and PhastCons, which are built on PlantRegMap, are only scored on a subset of the studies.

Although it should not necessarily be ranked first, we expect a validated causal variant to rank near the top of its locus, within the top 1% (about 60 variants) or even 0.1% (about 6 variants), which is why we choose to assess the performance on this task using the recall at these thresholds. The recalls are tie-aware: if the candidate is within a tie block in a study and the considered top fraction has to cut within the block, then the contribution of that study to the recall is neither 0 nor 1 but the proportion of the tie block that can fit in the shortlist.

Recall against shortlist size

Fraction of studies whose causal variant falls inside a shortlist of the given size. The dotted line is the random baseline (shuffled ranking). Restrict the studies to one species, or add and remove models, and the curves and the table are recomputed from the per-study ranks.

Species

Hover the plot to read every selected model's recall at that shortlist fraction. Baselines are dashed. Every language model is scored on the same studies; PhyloP and PhastCons cover 471 of the 545 and Ensembl VEP 540.

Recall at the top 0.1% and 1%

R(0.001) corresponds to the recall at the top 0.1% of all scored variants in the 100 kbp locus centred around the documented causal SNP. On average, R(0.001) corresponds to a shortlist of about 6 SNPs, and R(0.01) of about 60. Rows are sorted by the R(0.01) value. The report's Table gives 95% bootstrap confidence intervals and paired differences against Botanic1-XL; on the full cohort, Botanic1-XL's lead over every non-Botanic model excludes zero at both endpoints.

How the baselines fail

In this benchmark, non-gLM baselines fail in somewhat unexpected ways. Ensembl VEP is a categorical score, and usually scores the causal variant as "high impact" along with many other variants in the locus (median of 10.4% of all SNVs), so selecting the top 1% or lower is impossible on the whole dataset. Conservation scores also have a ceiling on their possible recall performance, due to the depth of the original MSA and each score's formula. PhyloP is the most sensitive to this effect, assigning a tied score to a median number of 39 top variants in 89 studies, reproducing the weakness of VEP to a smaller extent. PhastCons uses contextual information and is more continuous, and does not saturate in practice.

gLMs present several advantages over the tested baselines. First, they only require a sequence from the reference genome around the variant that one wants to score. Neither annotation nor multiple sequence alignment is needed. Secondly, since all present models are pre-trained on reference genomes only, their scores are not subject to linkage disequilibrium or influenced by the variant's frequency in arbitrary populations.

Context length

All models for which we performed a context sweep gradually improve performance until 4,096 bp, and plateau after. R(0.01) per scoring window, each model on its own matched set of studies:

Simplifications

This new benchmark also uses some simplifications: the search is limited to a 100 kbp causal region and to a single basic type of mutation, SNPs. Besides, the task relies on the assumption that the documented causal SNP should be ranked among the candidate variants assigned the lowest likelihood by the model, an assumption that could be challenged since the surrounding unlabelled variants may include other functional or phenotypically relevant mutations. This limitation affects the interpretation of the absolute benchmark scores, but not the relative comparison between methods, since all methods are evaluated against the same set of unlabelled variants.

The 545 studies

Below are 20 example studies from the benchmark. The full set of 545 studies, candidate variants, and curation audit will be published later.

Each row gives the validated causal SNP and its rank within the locus (1 is best). The recall curves above use all 545 studies.

Rank columns