Biological benchmarks
Explore evaluation suites and challenges, what they measure, and the models tested against them.
39 benchmark records in release 2026-09-29-06401fd5b220.
Broad tasks, specific protocols and evaluators have their own records. Follow each benchmark to its procedures and results.
ATOM3D
Molecular interactions
ATOM3D provides molecular-structure datasets and utilities for task-specific evaluation.
BEACON
RNA
BEACON compares RNA representations across structural and functional downstream tasks.
BEELINE
Biological networks
BEELINE compares inferred gene-regulatory edge rankings with reference networks.
BEND
Genomics
BEND evaluates DNA representations using task-specific genomic annotations and explicit split membership.
CAFA
Protein function
CAFA evaluates prospective protein-function predictions against annotations that become available after prediction submission.
CAMI
Microbiome
CAMI is a community benchmark program for metagenomic computational methods.
CAPRI
Protein structure
CAPRI assesses blind predictions of protein-complex structures supplied before experimental publication.
CASP
Protein structure
CASP assesses structure-prediction methods through blind predictions and category-specific evaluation.
DART-Eval
Genomics
DART-Eval measures human regulatory-DNA representations under zero-shot, probing and fine-tuning regimes.
ESMFold2 Runs N’ Poses comparison
Proteins and complexes
Bounded complete Figure 2C Runs N’ Poses subpanel: all eleven printed labels, separated into single-sequence and MSA comparisons.
FLIP
Protein function
FLIP evaluates protein-sequence representations using multiple deliberately defined train/test splits.
FLIP2
Protein function
FLIP2 evaluates protein fitness prediction under deliberately shifted training and test distributions.
Gene-MTEB
Genomics
Genomic embedding evaluations covering sequence classification and clustering. Accuracy and V-measure remain separate metrics and each source task retains its own comparison.
GENEB
Genomics
GENEB compares frozen genomic representations across classification tasks and label-budget regimes.
Genie 3 unconditional short monomer comparison
Proteins and complexes
Bounded complete Table 3 comparison across 13 method/version rows. Five directional metrics are included; secondary-structure composition is retained separately without ranking.
Genomic Benchmarks
Genomics
Genomic Benchmarks packages genomic sequence-classification datasets with explicit versions and supplied train/test folders.
GlycanML
Glycomics
GlycanML evaluates glycan learning across multiple classification and interaction tasks.
GUE
Genomics
GUE evaluates genome understanding across multiple datasets, task types and species.
HEST-Benchmark
Spatial omics
HEST-Benchmark tests prediction of gene expression from histological image representations.
MassSpecGym
Metabolomics
MassSpecGym separates spectrum-to-structure generation, candidate retrieval and structure-to-spectrum simulation.
MIST CANOPUS molecular retrieval
Metabolomics
Molecular candidate retrieval from mass spectra, as evaluated in DreaMS Extended Data Table 1. Inputs and extra training data differ across the reported methods.
mRNABench
RNA
mRNABench assesses genomic-model embeddings on transcript-specific expression, stability and regulatory tasks.
MSAlign molecular retrieval evaluations
Metabolomics
Spectrum-to-molecule retrieval comparisons on NPLIB1, Spectraverse and two MassSpecGym splits, reported in MSAlign Table 3. Formula availability and candidate pools remain separate.
NABench
RNA
NABench compares nucleotide foundation models on measured DNA/RNA sequence effects under multiple adaptation settings.
Open Problems
Single cell
Open Problems is an extensible platform hosting benchmark tasks and their datasets.
PerturBench
Single cell
PerturBench evaluates predicted single-cell perturbation responses with explicit aggregation and metric choices.
PEtab benchmark collection
Mechanistic biology
The PEtab collection supports evaluation of computational methods for fitting mathematical models to observations.
PFMBench
Protein function
PFMBench is a configurable suite of protein-model downstream evaluations.
Plant Genomic Benchmark (PGB)
Genomics
Plant genomic evaluations introduced with AgroNT. This release contains the complete promoter and terminator comparison tables from Figure 3e and 3f; other tasks remain unextracted.
PLINDER
Molecular interactions
PLINDER supplies annotated protein–ligand systems and evaluation resources for docking.
ProteinBench
Protein structure
ProteinBench assesses multiple protein-model tasks using quality, novelty, diversity and robustness dimensions.
ProteinGym
Protein function
ProteinGym separates experimental variant-effect and clinical annotation tasks under supervised and zero-shot regimes.
RhoFold+ CASP15 natural RNA comparison
RNA and transcriptomics
Complete Figure 2h source table: ten configurations, six natural CASP15 RNA targets, RMSD and cumulative GDT-TS/TM Z-score. Retrospective best-of-five comparison; eight unavailable entries retained in extraction receipts.
scFoundation cell-type annotation comparison
Cells spatial multiomics
Source-specific reported comparison. Protocol details, fitted configurations and limitations are retained separately from other evaluations.
scIB
Single cell
scIB is an atlas-level single-cell integration benchmark study covering RNA, chromatin-accessibility and simulated datasets. It compares batch removal with preservation of biological variation. The scib Python package and scib-pipeline implement its evaluation workflow.
SegmentNT human genome annotation
Genomics
Fourteen genomic-element annotation tasks from the SegmentNT study, with complete per-nucleotide MCC and auPRC comparisons from Supplementary Tables 2 and 3. Each element and metric remains a separate comparison.
TAPE
Protein function
TAPE evaluates protein representations through five supervised downstream tasks.
TDC molecular tasks
Molecular interactions
TDC organizes molecular prediction tasks into datasets and benchmark groups with explicit splitting and evaluation interfaces.
Virtual Cell Challenge 2026
Single cell
The 2026 Virtual Cell Challenge evaluates perturbation-response prediction in unseen cellular contexts.
This index reflects a dated catalogue, not an exhaustive census. Source checking does not mean independent reproduction; compare results only under compatible protocols, datasets and metrics.