Datasets
High-throughput assay collections spanning DNA and RNA families.
NABench compares nucleotide foundation models on measured DNA/RNA sequence effects under multiple adaptation settings.
High-throughput assay collections spanning DNA and RNA families.
Zero-shot evaluation reports Spearman correlation, NDCG, AUROC and MCC. Supervised and few-shot DMS evaluation emphasizes Spearman correlation, whereas SELEX evaluation uses AUROC.
DNA/RNA sequences and task-specific measured effects.
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
limited source coverage · Automated source review, 2026-09-16. All specifications and missing details
Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.
spearman (correlation) · Higher values are better.
NABench CCV-CORR-APTAMER: Fitness prediction on aptamer assays, supervised, contiguous cross validation · NABench aptamer assays (NABench split)
Evidence origin: Author-reported evaluation.
NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 11, column(aptamer)Every method NABench reports on Fitness prediction on aptamer assays, supervised, contiguous cross validation, scored with Spearman ρ on NABench aptamer assays.
Automated source review: 2026-09-18. Numerical source review does not establish independent reproduction.
Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.
Showing 12 of 25 matching rows.
NABench compares nucleic-acid fitness prediction across DMS and SELEX assays. Zero-shot sequence scores, supervised ridge probes and low-label settings are separate evaluation regimes. Random and contiguous-position folds distinguish interpolation from transfer to unseen mutational regions.
Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.
These source-backed links do not make different protocols or scores interchangeable.
Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.
No concrete protocols are explicitly linked to this suite. Protocol identification and baseline selection are outstanding.
Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums
Coverage is derived from release 2026-09-29-06401fd5b220. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.
Create the environment, produce embeddings for a model, and run the benchmark's own evaluation over the scored assays.
Generate predictions and evaluate them. This recipe does not establish reproduction of a particular published score.
Source reviewed; these instructions have not been executed by rewire.
# (Recommended) Create environment with conda
conda create -n nabench python=3.9
conda activate nabench
# Install dependencies with pip
pip install -r requirements.txtNABench: repository README · README.md at 99c8681e, Usage and Reproducibility, lines 99-104Source reviewed; these instructions have not been executed by rewire.
# Example for DNABERT
bash scripts/dnabert/seq_emb.sh path/to/input/data.csv path/to/output/embeddings.ptNABench: repository README · README.md at 99c8681e, Usage and Reproducibility, lines 114-115Source reviewed; these instructions have not been executed by rewire.
# Example command (specific script to be provided by you)
python evaluate.py --scores_dir path/to/scores --output_dir benchmarks/NABench: repository README · README.md at 99c8681e, Usage and Reproducibility, lines 122-123Run your model locally and return predictions keyed by the input IDs. The evaluator supplies biological inputs without test labels and owns scoring. This interface is not a sandbox for model code.
Pass your existing prediction function into this adapter. Its output direction must match the selected protocol.
class MyModelAdapter:
def __init__(self, score):
self.score = score
def predict(self, inputs):
return {row["id"]: float(self.score(row)) for row in inputs}
# adapter = MyModelAdapter(your_prediction_function)
# report = rewirebench.run(prepared, adapter, output="runs/my-model")Alternatively, generate a keyed prediction file in your existing model environment and use the score-only recipe. Your model code and weights do not need to be shared.
NABench: repository README · README.md at 99c8681eContribute a result for review. The library can submit an exported evaluation for private review when intake is open. Check the contribution page for access and sign-in.
The official README provides environment and embedding-generation templates. It states that SELEX processed data are still being organized, and leaves the performance-evaluation example as a script to be supplied by the user. A complete benchmark-run guide cannot be inferred from these placeholders.
A maintained rewire runner has not been verified for this benchmark. Check data access, weights, licences, dependencies and hardware in the linked official documentation; requirements have not been fully extracted.
mrzzmrzz/NABench / README.md · README.md lines 68–70 and 94–124 (Resources and Usage and Reproducibility)Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.
Stable record: discovery-benchmark-nabenchExplanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Datasets | High-throughput assay collections spanning DNA and RNA families.Sourcesmrzzmrzz/NABench official source · Pinned README: Introduction; Evaluation settings; evaluation scripts |
| Splits | Zero-shot, few-shot, supervised and transfer-learning settings are separate benchmark regimes.Sourcesmrzzmrzz/NABench official source · Pinned README: Introduction; Evaluation settings; evaluation scripts |
| Metrics | Zero-shot evaluation reports Spearman correlation, NDCG, AUROC and MCC. Supervised and few-shot DMS evaluation emphasizes Spearman correlation, whereas SELEX evaluation uses AUROC.Sourcesnabench primary benchmark evidence · Sections on evaluation settings and metrics; dataset appendix |
| Baselines | BERT-like, GPT-like, Hyena and LLaMA-based model families are included.Sourcesmrzzmrzz/NABench official source · Pinned README: Introduction; Evaluation settings; evaluation scripts |
| Leakage controls | Supervised probes use five-fold random or contiguous-position partitions. The contiguous version holds out variants mutated in a region of the wild-type sequence to test transfer across mutation positions; it is not a global homology or pretraining-overlap audit.Sourcesnabench primary benchmark evidence · Sections on evaluation settings and metrics; dataset appendix |
| Uncertainty | The inspected evaluation sections specify five-fold cross-validation and aggregate assay results, but do not define a uniform seed-based or bootstrap confidence interval for the suite. · Not reported in inspected sourcesSourcesnabench primary benchmark evidence · Sections on evaluation settings and metrics; dataset appendix |
| Entity type | Nucleic-acid variant-effect benchmark suite.Sourcesmrzzmrzz/NABench official source · Pinned README: Introduction; Evaluation settings; evaluation scripts |
| Organisms | Multiple natural DNA/RNA families and synthetic SELEX libraries, with assay-dependent experimental contexts. A synthetic selected sequence need not have a unique organism of origin.Sourcesnabench primary benchmark evidence · Sections on evaluation settings and metrics; dataset appendix |
| Assays | DMS and SELEX-derived nucleic-acid measurements; the README distinguishes released DMS data from pending processed SELEX data.Sourcesmrzzmrzz/NABench official source · Pinned README: Introduction; Evaluation settings; evaluation scripts |
| Allowed inputs | DNA/RNA sequences and task-specific measured effects.Sourcesmrzzmrzz/NABench official source · Pinned README: Introduction; Evaluation settings; evaluation scripts |
| Adaptation | Zero-shot, few-shot, supervised and transfer settings are evaluated separately.Sourcesmrzzmrzz/NABench official source · Pinned README: Introduction; Evaluation settings; evaluation scripts |
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Last literature check: 2026-09-17. Dated primary-source discovery and protocol/table screening. Source checking does not mean experimental reproduction. Only separately extracted and independently reviewed numeric batches are publishable.
| Paper or primary resource | Version | Reference |
|---|---|---|
| NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction | Primary full-text snapshot retrieved 2026-09-17; exact bytes pinned by SHA-256 | Read source |
The catalogue now holds 1,101 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.
primary protocol reviewed
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
29 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions. Individual claims | nabench primary benchmark evidence Pinned README: Introduction; Evaluation settings; evaluation scripts; Sections on evaluation settings and metrics; dataset appendix Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 2511.02888v1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram caption Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions. Individual claims | mrzzmrzz/NABench official source Pinned README: Introduction; Evaluation settings; evaluation scripts; Sections on evaluation settings and metrics; dataset appendix Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 99c8681ec1eab706e10ff90a5c329dcf184cc1d1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| nabench primary benchmark evidence Pinned README: Introduction; Evaluation settings; evaluation scripts; Sections on evaluation settings and metrics; dataset appendix Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 2511.02888v1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| mrzzmrzz/NABench official source Pinned README: Introduction; Evaluation settings; evaluation scripts; Sections on evaluation settings and metrics; dataset appendix Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 99c8681ec1eab706e10ff90a5c329dcf184cc1d1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Evaluation procedure Individual claims | nabench primary benchmark evidence Pinned README: Introduction; Evaluation settings; evaluation scripts; Sections on evaluation settings and metrics; dataset appendix Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 2511.02888v1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Evaluation procedure Individual claims | mrzzmrzz/NABench official source Pinned README: Introduction; Evaluation settings; evaluation scripts; Sections on evaluation settings and metrics; dataset appendix Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 99c8681ec1eab706e10ff90a5c329dcf184cc1d1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Datasets High-throughput assay collections spanning DNA and RNA families. Individual claims | mrzzmrzz/NABench official source Pinned README: Introduction; Evaluation settings; evaluation scripts Version: 99c8681ec1eab706e10ff90a5c329dcf184cc1d1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Splits Zero-shot, few-shot, supervised and transfer-learning settings are separate benchmark regimes. Individual claims | mrzzmrzz/NABench official source Pinned README: Introduction; Evaluation settings; evaluation scripts Version: 99c8681ec1eab706e10ff90a5c329dcf184cc1d1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Adaptation Zero-shot, few-shot, supervised and transfer settings are evaluated separately. Individual claims | mrzzmrzz/NABench official source Pinned README: Introduction; Evaluation settings; evaluation scripts Version: 99c8681ec1eab706e10ff90a5c329dcf184cc1d1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Metrics Zero-shot evaluation reports Spearman correlation, NDCG, AUROC and MCC. Supervised and few-shot DMS evaluation emphasizes Spearman correlation, whereas SELEX evaluation uses AUROC. Individual claims | nabench primary benchmark evidence Sections on evaluation settings and metrics; dataset appendix Version: 2511.02888v1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
View linked audit checks and correction history
Release 2026-09-29-06401fd5b220 · Record review: discovered
Stable ID: discovery-benchmark-nabench