rewire.itbenchmarks

Biological benchmarks

Explore evaluation suites and challenges, what they measure, and the models tested against them.

39 benchmark records in release 2026-09-29-06401fd5b220.

Broad tasks, specific protocols and evaluators have their own records. Follow each benchmark to its procedures and results.

  • ATOM3D

    Molecular interactions

    ATOM3D provides molecular-structure datasets and utilities for task-specific evaluation.

  • BEACON

    RNA

    BEACON compares RNA representations across structural and functional downstream tasks.

  • BEELINE

    Biological networks

    BEELINE compares inferred gene-regulatory edge rankings with reference networks.

  • BEND

    Genomics

    BEND evaluates DNA representations using task-specific genomic annotations and explicit split membership.

  • CAFA

    Protein function

    CAFA evaluates prospective protein-function predictions against annotations that become available after prediction submission.

  • CAMI

    Microbiome

    CAMI is a community benchmark program for metagenomic computational methods.

  • CAPRI

    Protein structure

    CAPRI assesses blind predictions of protein-complex structures supplied before experimental publication.

  • CASP

    Protein structure

    CASP assesses structure-prediction methods through blind predictions and category-specific evaluation.

  • DART-Eval

    Genomics

    DART-Eval measures human regulatory-DNA representations under zero-shot, probing and fine-tuning regimes.

  • ESMFold2 Runs N’ Poses comparison

    Proteins and complexes

    Bounded complete Figure 2C Runs N’ Poses subpanel: all eleven printed labels, separated into single-sequence and MSA comparisons.

  • FLIP

    Protein function

    FLIP evaluates protein-sequence representations using multiple deliberately defined train/test splits.

  • FLIP2

    Protein function

    FLIP2 evaluates protein fitness prediction under deliberately shifted training and test distributions.

  • Gene-MTEB

    Genomics

    Genomic embedding evaluations covering sequence classification and clustering. Accuracy and V-measure remain separate metrics and each source task retains its own comparison.

  • GENEB

    Genomics

    GENEB compares frozen genomic representations across classification tasks and label-budget regimes.

  • Genie 3 unconditional short monomer comparison

    Proteins and complexes

    Bounded complete Table 3 comparison across 13 method/version rows. Five directional metrics are included; secondary-structure composition is retained separately without ranking.

  • Genomic Benchmarks

    Genomics

    Genomic Benchmarks packages genomic sequence-classification datasets with explicit versions and supplied train/test folders.

  • GlycanML

    Glycomics

    GlycanML evaluates glycan learning across multiple classification and interaction tasks.

  • GUE

    Genomics

    GUE evaluates genome understanding across multiple datasets, task types and species.

  • HEST-Benchmark

    Spatial omics

    HEST-Benchmark tests prediction of gene expression from histological image representations.

  • MassSpecGym

    Metabolomics

    MassSpecGym separates spectrum-to-structure generation, candidate retrieval and structure-to-spectrum simulation.

  • MIST CANOPUS molecular retrieval

    Metabolomics

    Molecular candidate retrieval from mass spectra, as evaluated in DreaMS Extended Data Table 1. Inputs and extra training data differ across the reported methods.

  • mRNABench

    RNA

    mRNABench assesses genomic-model embeddings on transcript-specific expression, stability and regulatory tasks.

  • MSAlign molecular retrieval evaluations

    Metabolomics

    Spectrum-to-molecule retrieval comparisons on NPLIB1, Spectraverse and two MassSpecGym splits, reported in MSAlign Table 3. Formula availability and candidate pools remain separate.

  • NABench

    RNA

    NABench compares nucleotide foundation models on measured DNA/RNA sequence effects under multiple adaptation settings.

  • Open Problems

    Single cell

    Open Problems is an extensible platform hosting benchmark tasks and their datasets.

  • PerturBench

    Single cell

    PerturBench evaluates predicted single-cell perturbation responses with explicit aggregation and metric choices.

  • PEtab benchmark collection

    Mechanistic biology

    The PEtab collection supports evaluation of computational methods for fitting mathematical models to observations.

  • PFMBench

    Protein function

    PFMBench is a configurable suite of protein-model downstream evaluations.

  • Plant Genomic Benchmark (PGB)

    Genomics

    Plant genomic evaluations introduced with AgroNT. This release contains the complete promoter and terminator comparison tables from Figure 3e and 3f; other tasks remain unextracted.

  • PLINDER

    Molecular interactions

    PLINDER supplies annotated protein–ligand systems and evaluation resources for docking.

  • ProteinBench

    Protein structure

    ProteinBench assesses multiple protein-model tasks using quality, novelty, diversity and robustness dimensions.

  • ProteinGym

    Protein function

    ProteinGym separates experimental variant-effect and clinical annotation tasks under supervised and zero-shot regimes.

  • RhoFold+ CASP15 natural RNA comparison

    RNA and transcriptomics

    Complete Figure 2h source table: ten configurations, six natural CASP15 RNA targets, RMSD and cumulative GDT-TS/TM Z-score. Retrospective best-of-five comparison; eight unavailable entries retained in extraction receipts.

  • scFoundation cell-type annotation comparison

    Cells spatial multiomics

    Source-specific reported comparison. Protocol details, fitted configurations and limitations are retained separately from other evaluations.

  • scIB

    Single cell

    scIB is an atlas-level single-cell integration benchmark study covering RNA, chromatin-accessibility and simulated datasets. It compares batch removal with preservation of biological variation. The scib Python package and scib-pipeline implement its evaluation workflow.

  • SegmentNT human genome annotation

    Genomics

    Fourteen genomic-element annotation tasks from the SegmentNT study, with complete per-nucleotide MCC and auPRC comparisons from Supplementary Tables 2 and 3. Each element and metric remains a separate comparison.

  • TAPE

    Protein function

    TAPE evaluates protein representations through five supervised downstream tasks.

  • TDC molecular tasks

    Molecular interactions

    TDC organizes molecular prediction tasks into datasets and benchmark groups with explicit splitting and evaluation interfaces.

  • Virtual Cell Challenge 2026

    Single cell

    The 2026 Virtual Cell Challenge evaluates perturbation-response prediction in unseen cellular contexts.

This index reflects a dated catalogue, not an exhaustive census. Source checking does not mean independent reproduction; compare results only under compatible protocols, datasets and metrics.