Datasets
Genomic-coordinate tables paired with labels, including HDF5 labels for complex outputs; reference genome supplies input sequences.
BEND evaluates DNA representations using task-specific genomic annotations and explicit split membership.
Genomic-coordinate tables paired with labels, including HDF5 labels for complex outputs; reference genome supplies input sequences.
Gene finding uses multiclass MCC; enhancer annotation uses AUPRC; chromatin accessibility, histone modification and methylation use label-wise AUROC; variant-effect tasks use AUROC.
DNA sequences extracted from genomic coordinates and a reference genome; complex labels align by index to HDF5 records.
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
limited source coverage · Automated source review, 2026-09-16. All specifications and missing details
Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.
auroc (fraction) · Higher values are better.
BEND CHROMATIN: Chromatin accessibility · ENCODE chromatin accessibility (BEND split)
Evidence origin: Author-reported evaluation.
BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 1, row(Chromatin accessibility)Every method BEND reports on Chromatin accessibility, scored with AUROC on ENCODE chromatin accessibility.
Automated source review: 2026-09-18. Numerical source review does not establish independent reproduction.
Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.
Showing 12 of 15 matching rows.
BEND asks whether DNA representations support realistic human-genome annotation across short and long contexts. Frozen embeddings feed a small CNN for supervised tasks, while variant-effect prediction compares reference and alternate embeddings without fitting a task predictor. Whole-chromosome or sequence-identity partitions and specialist baselines make the tested capability explicit.
Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.
These source-backed links do not make different protocols or scores interchangeable.
Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.
No concrete protocols are explicitly linked to this suite. Protocol identification and baseline selection are outstanding.
Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums
Coverage is derived from release 2026-09-29-06401fd5b220. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.
Precompute embeddings for a DNA language model, then train and score the downstream head on one of the seven tasks.
Generate predictions and evaluate them. This recipe does not establish reproduction of a particular published score.
Source reviewed; these instructions have not been executed by rewire.
python scripts/precompute_embeddings.py model=resnetlm,awdlstm task=gene_finding,enhancer_annotationBEND: repository README · README.md at ac6e80c7, 3. Computing embeddings, lines 58-58Source reviewed; these instructions have not been executed by rewire.
python scripts/train_on_task.py --config-name gene_finding embedder=resnetlm,awdlstmBEND: repository README · README.md at ac6e80c7, Training and evaluating supervised models, lines 117-117Source reviewed; these instructions have not been executed by rewire.
python3 scripts/predict_variant_effects.py {variant_file_name}.bed {output_file_name}.csv {model_type} {path_to_checkpoint} {path_to_reference_genome_fasta} --embedding_idx {position_of_embedding}BEND: repository README · README.md at ac6e80c7, Unsupervised tasks, lines 162-162Run your model locally and return predictions keyed by the input IDs. The evaluator supplies biological inputs without test labels and owns scoring. This interface is not a sandbox for model code.
Pass your existing prediction function into this adapter. Its output direction must match the selected protocol.
class MyModelAdapter:
def __init__(self, score):
self.score = score
def predict(self, inputs):
return {row["id"]: float(self.score(row)) for row in inputs}
# adapter = MyModelAdapter(your_prediction_function)
# report = rewirebench.run(prepared, adapter, output="runs/my-model")Alternatively, generate a keyed prediction file in your existing model environment and use the score-only recipe. Your model code and weights do not need to be shared.
BEND: repository README · README.md at ac6e80c7Contribute a result for review. The library can submit an exported evaluation for private review when intake is open. Check the contribution page for access and sign-in.
Run the official gene-finding downstream example using precomputed ResNetLM and AWD-LSTM embeddings. This is one BEND task, not the entire suite.
Checked against the official instructions on 2026-09-17. These commands have not been executed by rewire. Running them does not automatically reproduce the published scores.
The clone command comes from the README; the detached checkout is added to select the exact revision reviewed here.
git clone https://github.com/frederikkemarin/BEND.git
cd BEND
git checkout --detach ac6e80c75e09d83cf47a7b4bcf0e44599c5706cffrederikkemarin/BEND / README.md · README.md lines 36–43; pinned repository revisionRun within the prepared Python 3.10 environment. This downloads data; it is not a small smoke test.
pip install -r requirements.txt
pip install -e .
python scripts/download_bend.pyfrederikkemarin/BEND / README.md · README.md lines 36–43This is the README example verbatim; it computes two models and two tasks. Embeddings are stored as Webdataset tar.gz shards; the next step uses the gene-finding subset.
python scripts/precompute_embeddings.py model=resnetlm,awdlstm task=gene_finding,enhancer_annotationfrederikkemarin/BEND / README.md · README.md lines 45–60Run after gene-finding embeddings are present in data/gene_finding/<embedder>/. Task-specific Hydra configuration supplies the data, model and evaluation settings.
python scripts/train_on_task.py --config-name gene_finding embedder=resnetlm,awdlstmfrederikkemarin/BEND / README.md · README.md lines 101–117, 148–155Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Independent automated spot review corrected interval terminology and sequence-identity scope against the original figure captions and methods.
Stable record: discovery-benchmark-bendExplanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Datasets | Genomic-coordinate tables paired with labels, including HDF5 labels for complex outputs; reference genome supplies input sequences.Sourcesfrederikkemarin/BEND official source · Pinned README: Data format; dataset use; Citation Guidelines |
| Splits | BED task files contain an explicit split column, and paired label files share the same row index.Sourcesfrederikkemarin/BEND official source · Pinned README: Data format; dataset use; Citation Guidelines |
| Metrics | Gene finding uses multiclass MCC; enhancer annotation uses AUPRC; chromatin accessibility, histone modification and methylation use label-wise AUROC; variant-effect tasks use AUROC.Sourcesbend primary benchmark evidence · Sections 3–4, Table 1, Table 3, Appendix A.1 and A.6 |
| Baselines | A one-hot two-layer CNN is matched to frozen-embedding probes. Task specialists include AUGUSTUS, Enformer, Basset and DeepSEA, and the study includes supervised ResNet and simple pretrained language-model controls.Sourcesbend primary benchmark evidence · Sections 3–4, Table 1, Table 3, Appendix A.1 and A.6 |
| Leakage controls | Supervised tasks hold out chromosomes, except gene finding, whose cross-partition pairs share no more than 80% identity of the mature protein. Enhancer evaluation uses ten chromosome-based folds. Variant effects are zero-shot tests; these partitions do not by themselves remove unsupervised genome-pretraining exposure.Sourcesbend primary benchmark evidence · Appendix A.1.1 gene-finding split: mature-protein identity; other supervised task and enhancer split descriptions in Appendix A.1 |
| Uncertainty | Enhancer evaluation uses ten-fold cross-validation because the dataset is small. Table 3 does not provide a uniform repeated-seed confidence-interval protocol for all task/model scores; its Enformer ± entry must not be generalized to every row. · Not reported in inspected sourcesSourcesbend primary benchmark evidence · Sections 3–4, Table 1, Table 3, Appendix A.1 and A.6 |
| Entity type | Human genomic task benchmark suite.Sourcesfrederikkemarin/BEND official source · Pinned README: Data format; dataset use; Citation Guidelines |
| Organisms | Human genomic evaluation; the README distinguishes models trained on other organisms from the evaluated set.Sourcesfrederikkemarin/BEND official source · Pinned README: Data format; dataset use; Citation Guidelines |
| Assays | Task-dependent regulatory and annotation datasets, including ENCODE-derived resources.Sourcesfrederikkemarin/BEND official source · Pinned README: Data format; dataset use; Citation Guidelines |
| Allowed inputs | DNA sequences extracted from genomic coordinates and a reference genome; complex labels align by index to HDF5 records.Sourcesfrederikkemarin/BEND official source · Pinned README: Data format; dataset use; Citation Guidelines |
| Adaptation | Downstream supervised models use precomputed embeddings; unsupervised variant-effect scoring is a separate evaluation route.Sourcesfrederikkemarin/BEND official source · Pinned README: Data format; dataset use; Citation Guidelines |
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Last literature check: 2026-09-17. Dated primary-source discovery and protocol/table screening. Source checking does not mean experimental reproduction. Only separately extracted and independently reviewed numeric batches are publishable.
| Paper or primary resource | Version | Reference |
|---|---|---|
| BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks | Primary full-text snapshot retrieved 2026-09-17; exact bytes pinned by SHA-256 | Read source |
The catalogue now holds 105 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.
primary protocol reviewed
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
33 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions. Individual claims | bend primary benchmark evidence Pinned README: Data format; dataset use; Citation Guidelines; Sections 3–4, Table 1, Table 3, Appendix A.1 and A.6 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 2311.12570v1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Independent automated spot review corrected interval terminology and sequence-identity scope against the original figure captions and methods. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram caption Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions. Individual claims | frederikkemarin/BEND official source Pinned README: Data format; dataset use; Citation Guidelines; Sections 3–4, Table 1, Table 3, Appendix A.1 and A.6 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: ac6e80c75e09d83cf47a7b4bcf0e44599c5706cf | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Independent automated spot review corrected interval terminology and sequence-identity scope against the original figure captions and methods. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| bend primary benchmark evidence Pinned README: Data format; dataset use; Citation Guidelines; Sections 3–4, Table 1, Table 3, Appendix A.1 and A.6 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 2311.12570v1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Independent automated spot review corrected interval terminology and sequence-identity scope against the original figure captions and methods. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| frederikkemarin/BEND official source Pinned README: Data format; dataset use; Citation Guidelines; Sections 3–4, Table 1, Table 3, Appendix A.1 and A.6 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: ac6e80c75e09d83cf47a7b4bcf0e44599c5706cf | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Independent automated spot review corrected interval terminology and sequence-identity scope against the original figure captions and methods. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Evaluation procedure Individual claims | bend primary benchmark evidence Pinned README: Data format; dataset use; Citation Guidelines; Sections 3–4, Table 1, Table 3, Appendix A.1 and A.6 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 2311.12570v1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Independent automated spot review corrected interval terminology and sequence-identity scope against the original figure captions and methods. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Evaluation procedure Individual claims | frederikkemarin/BEND official source Pinned README: Data format; dataset use; Citation Guidelines; Sections 3–4, Table 1, Table 3, Appendix A.1 and A.6 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: ac6e80c75e09d83cf47a7b4bcf0e44599c5706cf | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Independent automated spot review corrected interval terminology and sequence-identity scope against the original figure captions and methods. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Datasets Genomic-coordinate tables paired with labels, including HDF5 labels for complex outputs; reference genome supplies input sequences. Individual claims | frederikkemarin/BEND official source Pinned README: Data format; dataset use; Citation Guidelines Version: ac6e80c75e09d83cf47a7b4bcf0e44599c5706cf | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Independent automated spot review corrected interval terminology and sequence-identity scope against the original figure captions and methods. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Splits BED task files contain an explicit split column, and paired label files share the same row index. Individual claims | frederikkemarin/BEND official source Pinned README: Data format; dataset use; Citation Guidelines Version: ac6e80c75e09d83cf47a7b4bcf0e44599c5706cf | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Independent automated spot review corrected interval terminology and sequence-identity scope against the original figure captions and methods. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Adaptation Downstream supervised models use precomputed embeddings; unsupervised variant-effect scoring is a separate evaluation route. Individual claims | frederikkemarin/BEND official source Pinned README: Data format; dataset use; Citation Guidelines Version: ac6e80c75e09d83cf47a7b4bcf0e44599c5706cf | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Independent automated spot review corrected interval terminology and sequence-identity scope against the original figure captions and methods. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Metrics Gene finding uses multiclass MCC; enhancer annotation uses AUPRC; chromatin accessibility, histone modification and methylation use label-wise AUROC; variant-effect tasks use AUROC. Individual claims | bend primary benchmark evidence Sections 3–4, Table 1, Table 3, Appendix A.1 and A.6 Version: 2311.12570v1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Independent automated spot review corrected interval terminology and sequence-identity scope against the original figure captions and methods. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
View linked audit checks and correction history
Release 2026-09-29-06401fd5b220 · Record review: discovered
Stable ID: discovery-benchmark-bend