rewire.itbenchmarks
Benchmark

NABench

NABench compares nucleotide foundation models on measured DNA/RNA sequence effects under multiple adaptation settings.

Sourcesmrzzmrzz/NABench official source · Pinned README: Introduction; Evaluation settings; evaluation scripts

1,101 evaluations · 1,101 results

Overview

Datasets

High-throughput assay collections spanning DNA and RNA families.

Sourcesmrzzmrzz/NABench official source · Pinned README: Introduction; Evaluation settings; evaluation scripts

Metrics

Zero-shot evaluation reports Spearman correlation, NDCG, AUROC and MCC. Supervised and few-shot DMS evaluation emphasizes Spearman correlation, whereas SELEX evaluation uses AUROC.

Sourcesnabench primary benchmark evidence · Sections on evaluation settings and metrics; dataset appendix

Allowed inputs

DNA/RNA sequences and task-specific measured effects.

Sourcesmrzzmrzz/NABench official source · Pinned README: Introduction; Evaluation settings; evaluation scripts
Evaluation procedure diagram
How it worksEvaluation procedure
Evaluation procedure1. Allowed inputs: DNA/RNA sequences and task-specific measured effects.. Then: 2. Splits: Zero-shot, few-shot, supervised and transfer-learning settings are separate benchmark regimes.. Then: 3. Metrics: Zero-shot evaluation reports Spearman correlation, NDCG, AUROC and MCC. Supervised and few-shot DMS evaluation emphasizes Spearman correlation, whereas SELEX evaluation uses AUROC.Evaluation procedure1. Allowed inputs: DNA/RNA sequences and task-specific measured effects.. Then: 2. Splits: Zero-shot, few-shot, supervised and transfer-learning settings are separate benchmark regimes.. Then: 3. Metrics: Zero-shot evaluation reports Spearman correlation, NDCG, AUROC and MCC. Supervised and few-shot DMS evaluation emphasizes Spearman correlation, whereas SELEX evaluation uses AUROC.Evaluation procedure1. Allowed inputs: DNA/RNA sequences and task-specific measured effects.. Then: 2. Splits: Zero-shot, few-shot, supervised and transfer-learning settings are separate benchmark regimes.. Then: 3. Metrics: Zero-shot evaluation reports Spearman correlation, NDCG, AUROC and MCC. Supervised and few-shot DMS evaluation emphasizes Spearman correlation, whereas SELEX evaluation uses AUROC.

Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.

Sources (2)mrzzmrzz/NABench official source; nabench primary benchmark evidence · Pinned README: Introduction; Evaluation settings; evaluation scripts; Sections on evaluation settings and metrics; dataset appendix

limited source coverage · Automated source review, 2026-09-16. All specifications and missing details

Results

Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.

NABench CCV-CORR-APTAMER: Fitness prediction on aptamer assays, supervised, contiguous cross validation

spearman (correlation) · Higher values are better.

NABench CCV-CORR-APTAMER: Fitness prediction on aptamer assays, supervised, contiguous cross validation · NABench aptamer assays (NABench split)

Evidence origin: Author-reported evaluation.

NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 11, column(aptamer)
  • Zero-shot, few-shot and supervised cross validation figures come from different protocols and are not comparable to each other.
  • The per-nucleotide-type figures and the overall figures describe the same runs at different resolutions, so they must not be combined.
Comparison details and limitations

Every method NABench reports on Fitness prediction on aptamer assays, supervised, contiguous cross validation, scored with Spearman ρ on NABench aptamer assays.

  • Author-reported numbers, source checked but not independently reproduced.

Automated source review: 2026-09-18. Numerical source review does not establish independent reproduction.

Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.

Showing 12 of 25 matching rows.

Methods and evaluation design

Procedure, tasks and evaluated configurations

How it works

Evaluation methodology

NABench compares nucleic-acid fitness prediction across DMS and SELEX assays. Zero-shot sequence scores, supervised ridge probes and low-label settings are separate evaluation regimes. Random and contiguous-position folds distinguish interpolation from transfer to unseen mutational regions.

Sourcesnabench primary benchmark evidence · Sections on evaluation settings and metrics; dataset appendix

Evaluation design

Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.

Tasks

These source-backed links do not make different protocols or scores interchangeable.

Baseline coverage

Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.

No concrete protocols are explicitly linked to this suite. Protocol identification and baseline selection are outstanding.

Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums

Coverage is derived from release 2026-09-29-06401fd5b220. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.

Run this benchmark

Embed and score a nucleotide model on NABench

Create the environment, produce embeddings for a model, and run the benchmark's own evaluation over the scored assays.

Generate predictions and evaluate them. This recipe does not establish reproduction of a particular published score.

Dataset access
Downloaded as the README's data section describes.
Model and weights
A published nucleotide model checkpoint.
Licences
Project licence: see repository. Upstream data licences are separate and unreported here.
Software
Python with the repository's conda environment and requirements file.
Hardware
Not stated in the cited section. Several of these steps expect a GPU.
Required inputs and expected outputs

Inputs

  • A nucleotide foundation model and the assay data the repository downloads.

Outputs

  • Per-assay scores aggregated the way the benchmark defines.

Execution steps

  1. 1. Create the environment (Command line)

    Source reviewed; these instructions have not been executed by rewire.

    # (Recommended) Create environment with conda
    conda create -n nabench python=3.9
    conda activate nabench
    
    # Install dependencies with pip
    pip install -r requirements.txt
    NABench: repository README · README.md at 99c8681e, Usage and Reproducibility, lines 99-104
  2. 2. Produce embeddings (Command line)

    Source reviewed; these instructions have not been executed by rewire.

    # Example for DNABERT
    bash scripts/dnabert/seq_emb.sh path/to/input/data.csv path/to/output/embeddings.pt
    NABench: repository README · README.md at 99c8681e, Usage and Reproducibility, lines 114-115
  3. 3. Evaluate (Command line)

    Source reviewed; these instructions have not been executed by rewire.

    # Example command (specific script to be provided by you)
    python evaluate.py --scores_dir path/to/scores --output_dir benchmarks/
    NABench: repository README · README.md at 99c8681e, Usage and Reproducibility, lines 122-123

Use your own model

Run your model locally and return predictions keyed by the input IDs. The evaluator supplies biological inputs without test labels and owns scoring. This interface is not a sandbox for model code.

Pass your existing prediction function into this adapter. Its output direction must match the selected protocol.

class MyModelAdapter:
    def __init__(self, score):
        self.score = score

    def predict(self, inputs):
        return {row["id"]: float(self.score(row)) for row in inputs}

# adapter = MyModelAdapter(your_prediction_function)
# report = rewirebench.run(prepared, adapter, output="runs/my-model")

Alternatively, generate a keyed prediction file in your existing model environment and use the score-only recipe. Your model code and weights do not need to be shared.

NABench: repository README · README.md at 99c8681e
Scope and limitations
  • Quoted from the project's README and not executed by rewire, so the commands are evidence of what the project documents rather than a verified run.
  • The project may have changed since the pinned commit.
  • Zero-shot, few-shot and cross-validation figures come from different protocols and are not comparable.
  • The per-type and overall figures on this page describe the same runs at different resolutions.

Contribute a result for review. The library can submit an exported evaluation for private review when intake is open. Check the contribution page for access and sign-in.

Original repository instructions

Run this benchmark

The official README provides environment and embedding-generation templates. It states that SELEX processed data are still being organized, and leaves the performance-evaluation example as a script to be supplied by the user. A complete benchmark-run guide cannot be inferred from these placeholders.

A maintained rewire runner has not been verified for this benchmark. Check data access, weights, licences, dependencies and hardware in the linked official documentation; requirements have not been fully extracted.

mrzzmrzz/NABench / README.md · README.md lines 68–70 and 94–124 (Resources and Usage and Reproducibility)
Strengths, limitations and unresolved questions

Strengths and limitations

Strengths and considerations

  • Separating label-access regimes exposes when supervised adaptation changes model comparisons.
    Sourcesmrzzmrzz/NABench official source · Pinned README: Introduction; Evaluation settings; evaluation scripts

Limitations and conditions

  • Fitness labels arise from different selection and reporter assays. The paper’s generalization controls do not establish that every underlying sequence was absent from model pretraining.
    Sourcesnabench primary benchmark evidence · Sections on evaluation settings and metrics; dataset appendix
Profile review details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Stable record: discovery-benchmark-nabench

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetsHigh-throughput assay collections spanning DNA and RNA families.
Sourcesmrzzmrzz/NABench official source · Pinned README: Introduction; Evaluation settings; evaluation scripts
SplitsZero-shot, few-shot, supervised and transfer-learning settings are separate benchmark regimes.
Sourcesmrzzmrzz/NABench official source · Pinned README: Introduction; Evaluation settings; evaluation scripts
MetricsZero-shot evaluation reports Spearman correlation, NDCG, AUROC and MCC. Supervised and few-shot DMS evaluation emphasizes Spearman correlation, whereas SELEX evaluation uses AUROC.
Sourcesnabench primary benchmark evidence · Sections on evaluation settings and metrics; dataset appendix
BaselinesBERT-like, GPT-like, Hyena and LLaMA-based model families are included.
Sourcesmrzzmrzz/NABench official source · Pinned README: Introduction; Evaluation settings; evaluation scripts
Leakage controlsSupervised probes use five-fold random or contiguous-position partitions. The contiguous version holds out variants mutated in a region of the wild-type sequence to test transfer across mutation positions; it is not a global homology or pretraining-overlap audit.
Sourcesnabench primary benchmark evidence · Sections on evaluation settings and metrics; dataset appendix
UncertaintyThe inspected evaluation sections specify five-fold cross-validation and aggregate assay results, but do not define a uniform seed-based or bootstrap confidence interval for the suite. · Not reported in inspected sources
Sourcesnabench primary benchmark evidence · Sections on evaluation settings and metrics; dataset appendix
Entity typeNucleic-acid variant-effect benchmark suite.
Sourcesmrzzmrzz/NABench official source · Pinned README: Introduction; Evaluation settings; evaluation scripts
OrganismsMultiple natural DNA/RNA families and synthetic SELEX libraries, with assay-dependent experimental contexts. A synthetic selected sequence need not have a unique organism of origin.
Sourcesnabench primary benchmark evidence · Sections on evaluation settings and metrics; dataset appendix
AssaysDMS and SELEX-derived nucleic-acid measurements; the README distinguishes released DMS data from pending processed SELEX data.
Sourcesmrzzmrzz/NABench official source · Pinned README: Introduction; Evaluation settings; evaluation scripts
Allowed inputsDNA/RNA sequences and task-specific measured effects.
Sourcesmrzzmrzz/NABench official source · Pinned README: Introduction; Evaluation settings; evaluation scripts
AdaptationZero-shot, few-shot, supervised and transfer settings are evaluated separately.
Sourcesmrzzmrzz/NABench official source · Pinned README: Introduction; Evaluation settings; evaluation scripts

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Dated primary-source discovery and protocol/table screening. Source checking does not mean experimental reproduction. Only separately extracted and independently reviewed numeric batches are publishable.

Paper or primary resourceVersionReference
NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness PredictionPrimary full-text snapshot retrieved 2026-09-17; exact bytes pinned by SHA-256Read source
Historical gaps recorded on 2026-09-17

The catalogue now holds 1,101 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.

  • complete comparable numeric result batch: Specific candidate tables and protocol boundaries are documented; no graph values or incomplete winner-only selection are converted into publishable rows.
Search and extraction details

primary protocol reviewed

Searches

  • NABench nucleic acid benchmark 2511.02888

Evidence locations

  • v1 Tables 6–10 and dataset/assay tables

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

29 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Individual claims
nabench primary benchmark evidence

Original source ↗

Pinned README: Introduction; Evaluation settings; evaluation scripts; Sections on evaluation settings and metrics; dataset appendix

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2511.02888v1
Retrieved: 2026-09-16T21:06:29.579095+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: fefd48d53b1a7eadf9e14db96adacc8e646304c1b592562d9f136f4508941350

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram caption
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Individual claims
mrzzmrzz/NABench official source

Original source ↗

Pinned README: Introduction; Evaluation settings; evaluation scripts; Sections on evaluation settings and metrics; dataset appendix

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 99c8681ec1eab706e10ff90a5c329dcf184cc1d1
Retrieved: 2026-09-16T10:30:21.898594+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: e9c0b76d743af53198b0197bfa58305bf26cff822538658d0366880fcf56a8a9

Hash scope: Hash scope not separately documented; inspect source record

Diagram steps
  • Allowed inputs: DNA/RNA sequences and task-specific measured effects.
  • Splits: Zero-shot, few-shot, supervised and transfer-learning settings are separate benchmark regimes.
  • Metrics: Zero-shot evaluation reports Spearman correlation, NDCG, AUROC and MCC. Supervised and few-shot DMS evaluation emphasizes Spearman correlation, whereas SELEX evaluation uses AUROC.
Individual claims
nabench primary benchmark evidence

Original source ↗

Pinned README: Introduction; Evaluation settings; evaluation scripts; Sections on evaluation settings and metrics; dataset appendix

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2511.02888v1
Retrieved: 2026-09-16T21:06:29.579095+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: fefd48d53b1a7eadf9e14db96adacc8e646304c1b592562d9f136f4508941350

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps
  • Allowed inputs: DNA/RNA sequences and task-specific measured effects.
  • Splits: Zero-shot, few-shot, supervised and transfer-learning settings are separate benchmark regimes.
  • Metrics: Zero-shot evaluation reports Spearman correlation, NDCG, AUROC and MCC. Supervised and few-shot DMS evaluation emphasizes Spearman correlation, whereas SELEX evaluation uses AUROC.
Individual claims
mrzzmrzz/NABench official source

Original source ↗

Pinned README: Introduction; Evaluation settings; evaluation scripts; Sections on evaluation settings and metrics; dataset appendix

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 99c8681ec1eab706e10ff90a5c329dcf184cc1d1
Retrieved: 2026-09-16T10:30:21.898594+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: e9c0b76d743af53198b0197bfa58305bf26cff822538658d0366880fcf56a8a9

Hash scope: Hash scope not separately documented; inspect source record

Diagram title
Evaluation procedure
Individual claims
nabench primary benchmark evidence

Original source ↗

Pinned README: Introduction; Evaluation settings; evaluation scripts; Sections on evaluation settings and metrics; dataset appendix

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2511.02888v1
Retrieved: 2026-09-16T21:06:29.579095+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: fefd48d53b1a7eadf9e14db96adacc8e646304c1b592562d9f136f4508941350

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title
Evaluation procedure
Individual claims
mrzzmrzz/NABench official source

Original source ↗

Pinned README: Introduction; Evaluation settings; evaluation scripts; Sections on evaluation settings and metrics; dataset appendix

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 99c8681ec1eab706e10ff90a5c329dcf184cc1d1
Retrieved: 2026-09-16T10:30:21.898594+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: e9c0b76d743af53198b0197bfa58305bf26cff822538658d0366880fcf56a8a9

Hash scope: Hash scope not separately documented; inspect source record

Datasets
High-throughput assay collections spanning DNA and RNA families.
Individual claims
mrzzmrzz/NABench official source

Original source ↗

Pinned README: Introduction; Evaluation settings; evaluation scripts

Version: 99c8681ec1eab706e10ff90a5c329dcf184cc1d1
Retrieved: 2026-09-16T10:30:21.898594+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: e9c0b76d743af53198b0197bfa58305bf26cff822538658d0366880fcf56a8a9

Hash scope: Hash scope not separately documented; inspect source record

Splits
Zero-shot, few-shot, supervised and transfer-learning settings are separate benchmark regimes.
Individual claims
mrzzmrzz/NABench official source

Original source ↗

Pinned README: Introduction; Evaluation settings; evaluation scripts

Version: 99c8681ec1eab706e10ff90a5c329dcf184cc1d1
Retrieved: 2026-09-16T10:30:21.898594+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: e9c0b76d743af53198b0197bfa58305bf26cff822538658d0366880fcf56a8a9

Hash scope: Hash scope not separately documented; inspect source record

Adaptation
Zero-shot, few-shot, supervised and transfer settings are evaluated separately.
Individual claims
mrzzmrzz/NABench official source

Original source ↗

Pinned README: Introduction; Evaluation settings; evaluation scripts

Version: 99c8681ec1eab706e10ff90a5c329dcf184cc1d1
Retrieved: 2026-09-16T10:30:21.898594+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: e9c0b76d743af53198b0197bfa58305bf26cff822538658d0366880fcf56a8a9

Hash scope: Hash scope not separately documented; inspect source record

Metrics
Zero-shot evaluation reports Spearman correlation, NDCG, AUROC and MCC. Supervised and few-shot DMS evaluation emphasizes Spearman correlation, whereas SELEX evaluation uses AUROC.
Individual claims
nabench primary benchmark evidence

Original source ↗

Sections on evaluation settings and metrics; dataset appendix

Version: 2511.02888v1
Retrieved: 2026-09-16T21:06:29.579095+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: fefd48d53b1a7eadf9e14db96adacc8e646304c1b592562d9f136f4508941350

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: discovered

5 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: discovery-benchmark-nabench

areas
rna
entity level
suite
scope note
Specialist molecular or omics evaluation; protocol details require review before numerical comparison.
task
DNA and RNA fitness prediction
version
Not reported
benchmark research
review date: 2026-09-17; status: primary_protocol_reviewed; primary sources: evidence-expansion-nabench-fefd48d5; inspected locators: v1 Tables 6–10 and dataset/assay tables; searched queries: NABench nucleic acid benchmark 2511.02888; gaps: complete comparable numeric result batch: Specific candidate tables and protocol boundaries are documented; no graph values or incomplete winner-only selection are converted into publishable rows.; claim scope: Dated primary-source discovery and protocol/table screening. Source checking does not mean experimental reproduction. Only separately extracted and independently reviewed numeric batches are publishable.
historical missing metadata
dataset release: unextracted; metric implementation: unextracted; split manifest: unextracted; version: unextracted
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
entity classification
review date: 2026-09-17; rationale: The cited profile describes a collection of evaluation tasks or protocols; retain it as the top-level benchmark suite. Its datasets and individual protocols remain separate records.; source ids: src-discovery-mrzzmrzz-nabench; source locator: Pinned README: Introduction; Evaluation settings; evaluation scripts; ambiguities: None recorded
run documentation
record id: discovery-benchmark-nabench; source ids: run-doc-nabench-readme-md-99c8681e; status: official_documentation_linked; summary: The official README provides environment and embedding-generation templates. It states that SELEX processed data are still being organized, and leaves the performance-evaluation example as a script to be supplied by the user. A complete benchmark-run guide cannot be inferred from these placeholders.; source locator: README.md lines 68–70 and 94–124 (Resources and Usage and Reproducibility)
run recipes
id: nabench-official; protocol id: discovery-benchmark-nabench; version: 99c8681ec1eab706e10ff90a5c329dcf184cc1d1; title: Embed and score a nucleotide model on NABench; purpose: generate_and_evaluate; summary: Create the environment, produce embeddings for a model, and run the benchmark's own evaluation over the scored assays.; inputs: A nucleotide foundation model and the assay data the repository downloads.; outputs: Per-assay scores aggregated the way the benchmark defines.; requirements: data: Downloaded as the README's data section describes.; weights: A published nucleotide model checkpoint.; licence: Project licence: see repository. Upstream data licences are separate and unreported here.; software: Python with the repository's conda environment and requirements file.; hardware: Not stated in the cited section. Several of these steps expect a GPU.; instructions: runtime: command_line; title: Create the environment; code: # (Recommended) Create environment with conda conda create -n nabench python=3.9 conda activate nabench # Install dependencies with pip pip install -r requirements.txt; status: source_reviewed_not_executed; source ids: project-recipe-nabench-99c8681e; source locator: README.md at 99c8681e, Usage and Reproducibility, lines 99-104; runtime: command_line; title: Produce embeddings; code: # Example for DNABERT bash scripts/dnabert/seq_emb.sh path/to/input/data.csv path/to/output/embeddings.pt; status: source_reviewed_not_executed; source ids: project-recipe-nabench-99c8681e; source locator: README.md at 99c8681e, Usage and Reproducibility, lines 114-115; runtime: command_line; title: Evaluate; code: # Example command (specific script to be provided by you) python evaluate.py --scores_dir path/to/scores --output_dir benchmarks/; status: source_reviewed_not_executed; source ids: project-recipe-nabench-99c8681e; source locator: README.md at 99c8681e, Usage and Reproducibility, lines 122-123; limitations: Quoted from the project's README and not executed by rewire, so the commands are evidence of what the project documents rather than a verified run.; The project may have changed since the pinned commit.; Zero-shot, few-shot and cross-validation figures come from different protocols and are not comparable.; The per-type and overall figures on this page describe the same runs at different resolutions.; source ids: project-recipe-nabench-99c8681e; source locator: README.md at 99c8681e
Related records

Suggest a correction