rewire.itbenchmarks
Benchmark

BEND

BEND evaluates DNA representations using task-specific genomic annotations and explicit split membership.

Sourcesfrederikkemarin/BEND official source · Pinned README: Data format; dataset use; Citation Guidelines

105 evaluations · 105 results

Overview

Datasets

Genomic-coordinate tables paired with labels, including HDF5 labels for complex outputs; reference genome supplies input sequences.

Sourcesfrederikkemarin/BEND official source · Pinned README: Data format; dataset use; Citation Guidelines

Metrics

Gene finding uses multiclass MCC; enhancer annotation uses AUPRC; chromatin accessibility, histone modification and methylation use label-wise AUROC; variant-effect tasks use AUROC.

Sourcesbend primary benchmark evidence · Sections 3–4, Table 1, Table 3, Appendix A.1 and A.6

Allowed inputs

DNA sequences extracted from genomic coordinates and a reference genome; complex labels align by index to HDF5 records.

Sourcesfrederikkemarin/BEND official source · Pinned README: Data format; dataset use; Citation Guidelines
Evaluation procedure diagram
How it worksEvaluation procedure
Evaluation procedure1. Allowed inputs: DNA sequences extracted from genomic coordinates and a reference genome; complex labels align by index to HDF5 records.. Then: 2. Splits: BED task files contain an explicit split column, and paired label files share the same row index.. Then: 3. Metrics: Gene finding uses multiclass MCC; enhancer annotation uses AUPRC; chromatin accessibility, histone modification and methylation use label-wise AUROC; variant-effect tasks use AUROC.Evaluation procedure1. Allowed inputs: DNA sequences extracted from genomic coordinates and a reference genome; complex labels align by index to HDF5 records.. Then: 2. Splits: BED task files contain an explicit split column, and paired label files share the same row index.. Then: 3. Metrics: Gene finding uses multiclass MCC; enhancer annotation uses AUPRC; chromatin accessibility, histone modification and methylation use label-wise AUROC; variant-effect tasks use AUROC.Evaluation procedure1. Allowed inputs: DNA sequences extracted from genomic coordinates and a reference genome; complex labels align by index to HDF5 records.. Then: 2. Splits: BED task files contain an explicit split column, and paired label files share the same row index.. Then: 3. Metrics: Gene finding uses multiclass MCC; enhancer annotation uses AUPRC; chromatin accessibility, histone modification and methylation use label-wise AUROC; variant-effect tasks use AUROC.

Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.

Sources (2)frederikkemarin/BEND official source; bend primary benchmark evidence · Pinned README: Data format; dataset use; Citation Guidelines; Sections 3–4, Table 1, Table 3, Appendix A.1 and A.6

limited source coverage · Automated source review, 2026-09-16. All specifications and missing details

Results

Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.

BEND CHROMATIN: Chromatin accessibility

auroc (fraction) · Higher values are better.

BEND CHROMATIN: Chromatin accessibility · ENCODE chromatin accessibility (BEND split)

Evidence origin: Author-reported evaluation.

BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 1, row(Chromatin accessibility)
  • The expert entries are specialist published models, each compared on one task only.
  • The metric differs by task, taken from Table 1, so these figures cannot be averaged into one score.
Comparison details and limitations

Every method BEND reports on Chromatin accessibility, scored with AUROC on ENCODE chromatin accessibility.

  • Author-reported numbers, source checked but not independently reproduced.

Automated source review: 2026-09-18. Numerical source review does not establish independent reproduction.

Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.

Showing 12 of 15 matching rows.

Methods and evaluation design

Procedure, tasks and evaluated configurations

How it works

Evaluation methodology

BEND asks whether DNA representations support realistic human-genome annotation across short and long contexts. Frozen embeddings feed a small CNN for supervised tasks, while variant-effect prediction compares reference and alternate embeddings without fitting a task predictor. Whole-chromosome or sequence-identity partitions and specialist baselines make the tested capability explicit.

Sourcesbend primary benchmark evidence · Sections 3–4, Table 1, Table 3, Appendix A.1 and A.6

Evaluation design

Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.

These source-backed links do not make different protocols or scores interchangeable.

Baseline coverage

Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.

No concrete protocols are explicitly linked to this suite. Protocol identification and baseline selection are outstanding.

Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums

Coverage is derived from release 2026-09-29-06401fd5b220. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.

Run this benchmark

Embed, train and evaluate on a BEND task

Precompute embeddings for a DNA language model, then train and score the downstream head on one of the seven tasks.

Generate predictions and evaluate them. This recipe does not establish reproduction of a particular published score.

Dataset access
Downloaded by the repository's scripts.
Model and weights
A published DNA language model checkpoint.
Licences
Project licence: BSD-3-Clause. Upstream data licences are separate and unreported here.
Software
Python with the repository's environment and hydra configs.
Hardware
Not stated in the cited section. Several of these steps expect a GPU.
Required inputs and expected outputs

Inputs

  • A supported embedder, or your own.
  • The task data the repository downloads.

Outputs

  • Precomputed embeddings and a scored downstream model.

Execution steps

  1. 1. Precompute embeddings (Command line)

    Source reviewed; these instructions have not been executed by rewire.

    python scripts/precompute_embeddings.py model=resnetlm,awdlstm task=gene_finding,enhancer_annotation
    BEND: repository README · README.md at ac6e80c7, 3. Computing embeddings, lines 58-58
  2. 2. Train and evaluate on a task (Command line)

    Source reviewed; these instructions have not been executed by rewire.

    python scripts/train_on_task.py --config-name gene_finding embedder=resnetlm,awdlstm
    BEND: repository README · README.md at ac6e80c7, Training and evaluating supervised models, lines 117-117
  3. 3. Score variant effects (Command line)

    Source reviewed; these instructions have not been executed by rewire.

    python3 scripts/predict_variant_effects.py {variant_file_name}.bed {output_file_name}.csv {model_type} {path_to_checkpoint} {path_to_reference_genome_fasta} --embedding_idx {position_of_embedding}
    BEND: repository README · README.md at ac6e80c7, Unsupervised tasks, lines 162-162

Use your own model

Run your model locally and return predictions keyed by the input IDs. The evaluator supplies biological inputs without test labels and owns scoring. This interface is not a sandbox for model code.

Pass your existing prediction function into this adapter. Its output direction must match the selected protocol.

class MyModelAdapter:
    def __init__(self, score):
        self.score = score

    def predict(self, inputs):
        return {row["id"]: float(self.score(row)) for row in inputs}

# adapter = MyModelAdapter(your_prediction_function)
# report = rewirebench.run(prepared, adapter, output="runs/my-model")

Alternatively, generate a keyed prediction file in your existing model environment and use the score-only recipe. Your model code and weights do not need to be shared.

BEND: repository README · README.md at ac6e80c7
Scope and limitations
  • Quoted from the project's README and not executed by rewire, so the commands are evidence of what the project documents rather than a verified run.
  • The project may have changed since the pinned commit.
  • The metric differs by task, so the figures on this page cannot be averaged.
  • The expert entries on this page are specialist models, each compared on one task only.

Contribute a result for review. The library can submit an exported evaluation for private review when intake is open. Check the contribution page for access and sign-in.

Original repository instructions

Official run instructions

Run the official gene-finding downstream example using precomputed ResNetLM and AWD-LSTM embeddings. This is one BEND task, not the entire suite.

Checked against the official instructions on 2026-09-17. These commands have not been executed by rewire. Running them does not automatically reproduce the published scores.

Before you start

  • Git and a conda environment with Python 3.10; the README recommends this Python version.
  • Internet access to the official data host and the model resources selected by the configuration. Preserve the downloaded data directory structure.
  1. 1. Check out the reviewed code

    The clone command comes from the README; the detached checkout is added to select the exact revision reviewed here.

    git clone https://github.com/frederikkemarin/BEND.git
    cd BEND
    git checkout --detach ac6e80c75e09d83cf47a7b4bcf0e44599c5706cf
    frederikkemarin/BEND / README.md · README.md lines 36–43; pinned repository revision
  2. 2. Install and download the benchmark data

    Run within the prepared Python 3.10 environment. This downloads data; it is not a small smoke test.

    pip install -r requirements.txt
    pip install -e .
    python scripts/download_bend.py
    frederikkemarin/BEND / README.md · README.md lines 36–43
  3. 3. Precompute the documented embeddings

    This is the README example verbatim; it computes two models and two tasks. Embeddings are stored as Webdataset tar.gz shards; the next step uses the gene-finding subset.

    python scripts/precompute_embeddings.py model=resnetlm,awdlstm task=gene_finding,enhancer_annotation
    frederikkemarin/BEND / README.md · README.md lines 45–60
  4. 4. Train and evaluate gene finding

    Run after gene-finding embeddings are present in data/gene_finding/<embedder>/. Task-specific Hydra configuration supplies the data, model and evaluation settings.

    python scripts/train_on_task.py --config-name gene_finding embedder=resnetlm,awdlstm
    frederikkemarin/BEND / README.md · README.md lines 101–117, 148–155

Expected outputs

  • Embedding shards under data/<task_name>/<embedder>/*.tar.gz.
  • Downstream run results under downstream_tasks/gene_finding/<embedder>/.

Scope and limitations

  • No hardware minimum, runtime or monetary cost is stated in these README sections; no cost estimate is inferred.
  • Checkpoint downloads, dataset bytes and transitive dependencies are not checksum-pinned by this guide. Pinning the repository does not pin those external resources.
  • Changing the data, embedder or Hydra settings changes the evaluation configuration; this example does not establish reproduction of a published score.
Strengths, limitations and unresolved questions

Strengths and limitations

Strengths and considerations

  • Explicit split columns and aligned label indices make task membership inspectable.
    Sourcesfrederikkemarin/BEND official source · Pinned README: Data format; dataset use; Citation Guidelines

Limitations and conditions

  • BEND measures different biological tasks with different metrics. Its zero-shot variant tests, frozen probes and literature specialist results require separate interpretation; pretraining on the reference genome is not equivalent to supervised label leakage.
    Sourcesbend primary benchmark evidence · Sections 3–4, Table 1, Table 3, Appendix A.1 and A.6
Profile review details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Independent automated spot review corrected interval terminology and sequence-identity scope against the original figure captions and methods.

Stable record: discovery-benchmark-bend

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetsGenomic-coordinate tables paired with labels, including HDF5 labels for complex outputs; reference genome supplies input sequences.
Sourcesfrederikkemarin/BEND official source · Pinned README: Data format; dataset use; Citation Guidelines
SplitsBED task files contain an explicit split column, and paired label files share the same row index.
Sourcesfrederikkemarin/BEND official source · Pinned README: Data format; dataset use; Citation Guidelines
MetricsGene finding uses multiclass MCC; enhancer annotation uses AUPRC; chromatin accessibility, histone modification and methylation use label-wise AUROC; variant-effect tasks use AUROC.
Sourcesbend primary benchmark evidence · Sections 3–4, Table 1, Table 3, Appendix A.1 and A.6
BaselinesA one-hot two-layer CNN is matched to frozen-embedding probes. Task specialists include AUGUSTUS, Enformer, Basset and DeepSEA, and the study includes supervised ResNet and simple pretrained language-model controls.
Sourcesbend primary benchmark evidence · Sections 3–4, Table 1, Table 3, Appendix A.1 and A.6
Leakage controlsSupervised tasks hold out chromosomes, except gene finding, whose cross-partition pairs share no more than 80% identity of the mature protein. Enhancer evaluation uses ten chromosome-based folds. Variant effects are zero-shot tests; these partitions do not by themselves remove unsupervised genome-pretraining exposure.
Sourcesbend primary benchmark evidence · Appendix A.1.1 gene-finding split: mature-protein identity; other supervised task and enhancer split descriptions in Appendix A.1
UncertaintyEnhancer evaluation uses ten-fold cross-validation because the dataset is small. Table 3 does not provide a uniform repeated-seed confidence-interval protocol for all task/model scores; its Enformer ± entry must not be generalized to every row. · Not reported in inspected sources
Sourcesbend primary benchmark evidence · Sections 3–4, Table 1, Table 3, Appendix A.1 and A.6
Entity typeHuman genomic task benchmark suite.
Sourcesfrederikkemarin/BEND official source · Pinned README: Data format; dataset use; Citation Guidelines
OrganismsHuman genomic evaluation; the README distinguishes models trained on other organisms from the evaluated set.
Sourcesfrederikkemarin/BEND official source · Pinned README: Data format; dataset use; Citation Guidelines
AssaysTask-dependent regulatory and annotation datasets, including ENCODE-derived resources.
Sourcesfrederikkemarin/BEND official source · Pinned README: Data format; dataset use; Citation Guidelines
Allowed inputsDNA sequences extracted from genomic coordinates and a reference genome; complex labels align by index to HDF5 records.
Sourcesfrederikkemarin/BEND official source · Pinned README: Data format; dataset use; Citation Guidelines
AdaptationDownstream supervised models use precomputed embeddings; unsupervised variant-effect scoring is a separate evaluation route.
Sourcesfrederikkemarin/BEND official source · Pinned README: Data format; dataset use; Citation Guidelines

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Dated primary-source discovery and protocol/table screening. Source checking does not mean experimental reproduction. Only separately extracted and independently reviewed numeric batches are publishable.

Paper or primary resourceVersionReference
BEND: Benchmarking DNA Language Models on Biologically Meaningful TasksPrimary full-text snapshot retrieved 2026-09-17; exact bytes pinned by SHA-256Read source
Historical gaps recorded on 2026-09-17

The catalogue now holds 105 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.

  • complete comparable numeric result batch: Specific candidate tables and protocol boundaries are documented; no graph values or incomplete winner-only selection are converted into publishable rows.
Search and extraction details

primary protocol reviewed

Searches

  • BEND benchmark DNA language models ICLR 2024 results

Evidence locations

  • Final ICLR 2024 Tables 1–3; task-specific Appendix

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

33 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Individual claims
bend primary benchmark evidence

Original source ↗

Pinned README: Data format; dataset use; Citation Guidelines; Sections 3–4, Table 1, Table 3, Appendix A.1 and A.6

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2311.12570v1
Retrieved: 2026-09-16T21:04:56.127991+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Independent automated spot review corrected interval terminology and sequence-identity scope against the original figure captions and methods.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 6ff9f6dc19fc831200e241e3279a06c44d0fcd0aac566da81058e9db7fb70334

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram caption
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Individual claims
frederikkemarin/BEND official source

Original source ↗

Pinned README: Data format; dataset use; Citation Guidelines; Sections 3–4, Table 1, Table 3, Appendix A.1 and A.6

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: ac6e80c75e09d83cf47a7b4bcf0e44599c5706cf
Retrieved: 2026-09-16T10:30:20.985691+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Independent automated spot review corrected interval terminology and sequence-identity scope against the original figure captions and methods.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: a40358504726ee4e086f0623348fb2206c24ab8166d04a83d944579e9c62bdc8

Hash scope: Hash scope not separately documented; inspect source record

Diagram steps
  • Allowed inputs: DNA sequences extracted from genomic coordinates and a reference genome; complex labels align by index to HDF5 records.
  • Splits: BED task files contain an explicit split column, and paired label files share the same row index.
  • Metrics: Gene finding uses multiclass MCC; enhancer annotation uses AUPRC; chromatin accessibility, histone modification and methylation use label-wise AUROC; variant-effect tasks use AUROC.
Individual claims
bend primary benchmark evidence

Original source ↗

Pinned README: Data format; dataset use; Citation Guidelines; Sections 3–4, Table 1, Table 3, Appendix A.1 and A.6

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2311.12570v1
Retrieved: 2026-09-16T21:04:56.127991+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Independent automated spot review corrected interval terminology and sequence-identity scope against the original figure captions and methods.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 6ff9f6dc19fc831200e241e3279a06c44d0fcd0aac566da81058e9db7fb70334

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps
  • Allowed inputs: DNA sequences extracted from genomic coordinates and a reference genome; complex labels align by index to HDF5 records.
  • Splits: BED task files contain an explicit split column, and paired label files share the same row index.
  • Metrics: Gene finding uses multiclass MCC; enhancer annotation uses AUPRC; chromatin accessibility, histone modification and methylation use label-wise AUROC; variant-effect tasks use AUROC.
Individual claims
frederikkemarin/BEND official source

Original source ↗

Pinned README: Data format; dataset use; Citation Guidelines; Sections 3–4, Table 1, Table 3, Appendix A.1 and A.6

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: ac6e80c75e09d83cf47a7b4bcf0e44599c5706cf
Retrieved: 2026-09-16T10:30:20.985691+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Independent automated spot review corrected interval terminology and sequence-identity scope against the original figure captions and methods.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: a40358504726ee4e086f0623348fb2206c24ab8166d04a83d944579e9c62bdc8

Hash scope: Hash scope not separately documented; inspect source record

Diagram title
Evaluation procedure
Individual claims
bend primary benchmark evidence

Original source ↗

Pinned README: Data format; dataset use; Citation Guidelines; Sections 3–4, Table 1, Table 3, Appendix A.1 and A.6

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2311.12570v1
Retrieved: 2026-09-16T21:04:56.127991+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Independent automated spot review corrected interval terminology and sequence-identity scope against the original figure captions and methods.

Field: attributes.profile.diagram.title

Source artifact SHA-256: 6ff9f6dc19fc831200e241e3279a06c44d0fcd0aac566da81058e9db7fb70334

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title
Evaluation procedure
Individual claims
frederikkemarin/BEND official source

Original source ↗

Pinned README: Data format; dataset use; Citation Guidelines; Sections 3–4, Table 1, Table 3, Appendix A.1 and A.6

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: ac6e80c75e09d83cf47a7b4bcf0e44599c5706cf
Retrieved: 2026-09-16T10:30:20.985691+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Independent automated spot review corrected interval terminology and sequence-identity scope against the original figure captions and methods.

Field: attributes.profile.diagram.title

Source artifact SHA-256: a40358504726ee4e086f0623348fb2206c24ab8166d04a83d944579e9c62bdc8

Hash scope: Hash scope not separately documented; inspect source record

Datasets
Genomic-coordinate tables paired with labels, including HDF5 labels for complex outputs; reference genome supplies input sequences.
Individual claims
frederikkemarin/BEND official source

Original source ↗

Pinned README: Data format; dataset use; Citation Guidelines

Version: ac6e80c75e09d83cf47a7b4bcf0e44599c5706cf
Retrieved: 2026-09-16T10:30:20.985691+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Independent automated spot review corrected interval terminology and sequence-identity scope against the original figure captions and methods.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: a40358504726ee4e086f0623348fb2206c24ab8166d04a83d944579e9c62bdc8

Hash scope: Hash scope not separately documented; inspect source record

Splits
BED task files contain an explicit split column, and paired label files share the same row index.
Individual claims
frederikkemarin/BEND official source

Original source ↗

Pinned README: Data format; dataset use; Citation Guidelines

Version: ac6e80c75e09d83cf47a7b4bcf0e44599c5706cf
Retrieved: 2026-09-16T10:30:20.985691+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Independent automated spot review corrected interval terminology and sequence-identity scope against the original figure captions and methods.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: a40358504726ee4e086f0623348fb2206c24ab8166d04a83d944579e9c62bdc8

Hash scope: Hash scope not separately documented; inspect source record

Adaptation
Downstream supervised models use precomputed embeddings; unsupervised variant-effect scoring is a separate evaluation route.
Individual claims
frederikkemarin/BEND official source

Original source ↗

Pinned README: Data format; dataset use; Citation Guidelines

Version: ac6e80c75e09d83cf47a7b4bcf0e44599c5706cf
Retrieved: 2026-09-16T10:30:20.985691+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Independent automated spot review corrected interval terminology and sequence-identity scope against the original figure captions and methods.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: a40358504726ee4e086f0623348fb2206c24ab8166d04a83d944579e9c62bdc8

Hash scope: Hash scope not separately documented; inspect source record

Metrics
Gene finding uses multiclass MCC; enhancer annotation uses AUPRC; chromatin accessibility, histone modification and methylation use label-wise AUROC; variant-effect tasks use AUROC.
Individual claims
bend primary benchmark evidence

Original source ↗

Sections 3–4, Table 1, Table 3, Appendix A.1 and A.6

Version: 2311.12570v1
Retrieved: 2026-09-16T21:04:56.127991+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Independent automated spot review corrected interval terminology and sequence-identity scope against the original figure captions and methods.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: 6ff9f6dc19fc831200e241e3279a06c44d0fcd0aac566da81058e9db7fb70334

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: discovered

5 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: discovery-benchmark-bend

areas
genomics
entity level
suite
scope note
Specialist molecular or omics evaluation; protocol details require review before numerical comparison.
task
DNA representations on biological downstream tasks
version
Not reported
benchmark research
review date: 2026-09-17; status: primary_protocol_reviewed; primary sources: evidence-expansion-bend-final-f709b6be; inspected locators: Final ICLR 2024 Tables 1–3; task-specific Appendix; searched queries: BEND benchmark DNA language models ICLR 2024 results; gaps: complete comparable numeric result batch: Specific candidate tables and protocol boundaries are documented; no graph values or incomplete winner-only selection are converted into publishable rows.; claim scope: Dated primary-source discovery and protocol/table screening. Source checking does not mean experimental reproduction. Only separately extracted and independently reviewed numeric batches are publishable.
historical missing metadata
dataset release: unextracted; metric implementation: unextracted; split manifest: unextracted; version: unextracted
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
entity classification
review date: 2026-09-17; rationale: The cited profile describes a collection of evaluation tasks or protocols; retain it as the top-level benchmark suite. Its datasets and individual protocols remain separate records.; source ids: src-discovery-frederikkemarin-bend; source locator: Pinned README: Data format; dataset use; Citation Guidelines; ambiguities: None recorded
run guide
record id: discovery-benchmark-bend; summary: Run the official gene-finding downstream example using precomputed ResNetLM and AWD-LSTM embeddings. This is one BEND task, not the entire suite.; status: source_reviewed_not_executed; prerequisites: Git and a conda environment with Python 3.10; the README recommends this Python version.; Internet access to the official data host and the model resources selected by the configuration. Preserve the downloaded data directory structure.; steps: title: Check out the reviewed code; shell: git clone https://github.com/frederikkemarin/BEND.git cd BEND git checkout --detach ac6e80c75e09d83cf47a7b4bcf0e44599c5706cf; explanation: The clone command comes from the README; the detached checkout is added to select the exact revision reviewed here.; source ids: run-doc-bend-readme-md-ac6e80c7; source locator: README.md lines 36–43; pinned repository revision; title: Install and download the benchmark data; shell: pip install -r requirements.txt pip install -e . python scripts/download_bend.py; explanation: Run within the prepared Python 3.10 environment. This downloads data; it is not a small smoke test.; source ids: run-doc-bend-readme-md-ac6e80c7; source locator: README.md lines 36–43; title: Precompute the documented embeddings; shell: python scripts/precompute_embeddings.py model=resnetlm,awdlstm task=gene_finding,enhancer_annotation; explanation: This is the README example verbatim; it computes two models and two tasks. Embeddings are stored as Webdataset tar.gz shards; the next step uses the gene-finding subset.; source ids: run-doc-bend-readme-md-ac6e80c7; source locator: README.md lines 45–60; title: Train and evaluate gene finding; shell: python scripts/train_on_task.py --config-name gene_finding embedder=resnetlm,awdlstm; explanation: Run after gene-finding embeddings are present in data/gene_finding/<embedder>/. Task-specific Hydra configuration supplies the data, model and evaluation settings.; source ids: run-doc-bend-readme-md-ac6e80c7; source locator: README.md lines 101–117, 148–155; outputs: Embedding shards under data/<task_name>/<embedder>/*.tar.gz.; Downstream run results under downstream_tasks/gene_finding/<embedder>/.; limitations: No hardware minimum, runtime or monetary cost is stated in these README sections; no cost estimate is inferred.; Checkpoint downloads, dataset bytes and transitive dependencies are not checksum-pinned by this guide. Pinning the repository does not pin those external resources.; Changing the data, embedder or Hydra settings changes the evaluation configuration; this example does not establish reproduction of a published score.; source ids: run-doc-bend-readme-md-ac6e80c7; review: method: official_repository_review; date: 2026-09-17
run documentation
record id: discovery-benchmark-bend; source ids: run-doc-bend-readme-md-ac6e80c7; status: source_reviewed_not_executed; summary: Run the official gene-finding downstream example using precomputed ResNetLM and AWD-LSTM embeddings. This is one BEND task, not the entire suite. Commands were source-reviewed only. No hardware minimum, runtime or monetary cost is stated in these README sections; no cost estimate is inferred.; source locator: README.md lines 36–43; pinned repository revision; README.md lines 36–43; README.md lines 45–60; README.md lines 101–117, 148–155
run recipes
id: bend-official; protocol id: discovery-benchmark-bend; version: ac6e80c75e09d83cf47a7b4bcf0e44599c5706cf; title: Embed, train and evaluate on a BEND task; purpose: generate_and_evaluate; summary: Precompute embeddings for a DNA language model, then train and score the downstream head on one of the seven tasks.; inputs: A supported embedder, or your own.; The task data the repository downloads.; outputs: Precomputed embeddings and a scored downstream model.; requirements: data: Downloaded by the repository's scripts.; weights: A published DNA language model checkpoint.; licence: Project licence: BSD-3-Clause. Upstream data licences are separate and unreported here.; software: Python with the repository's environment and hydra configs.; hardware: Not stated in the cited section. Several of these steps expect a GPU.; instructions: runtime: command_line; title: Precompute embeddings; code: python scripts/precompute_embeddings.py model=resnetlm,awdlstm task=gene_finding,enhancer_annotation; status: source_reviewed_not_executed; source ids: project-recipe-bend-ac6e80c7; source locator: README.md at ac6e80c7, 3. Computing embeddings, lines 58-58; runtime: command_line; title: Train and evaluate on a task; code: python scripts/train_on_task.py --config-name gene_finding embedder=resnetlm,awdlstm; status: source_reviewed_not_executed; source ids: project-recipe-bend-ac6e80c7; source locator: README.md at ac6e80c7, Training and evaluating supervised models, lines 117-117; runtime: command_line; title: Score variant effects; code: python3 scripts/predict_variant_effects.py {variant_file_name}.bed {output_file_name}.csv {model_type} {path_to_checkpoint} {path_to_reference_genome_fasta} --embedding_idx {position_of_embedding}; status: source_reviewed_not_executed; source ids: project-recipe-bend-ac6e80c7; source locator: README.md at ac6e80c7, Unsupervised tasks, lines 162-162; limitations: Quoted from the project's README and not executed by rewire, so the commands are evidence of what the project documents rather than a verified run.; The project may have changed since the pinned commit.; The metric differs by task, so the figures on this page cannot be averaged.; The expert entries on this page are specialist models, each compared on one task only.; source ids: project-recipe-bend-ac6e80c7; source locator: README.md at ac6e80c7
Related records

Suggest a correction