rewire.itbenchmarks
Task

regulatory element identification

Regulatory-element identification tests whether DNA models distinguish ENCODE regulatory sequences from composition-matched controls.

SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1

1 evaluation · 1 result

Overview

Datasets

ENCODE candidate cis-regulatory elements paired with synthetic dinucleotide-shuffled controls.

Metrics

Zero-shot accuracy measures which member of a matched pair receives higher sequence likelihood. Supervised absolute classification accuracy and paired ranking accuracy are separate quantities.

Allowed inputs

DNA sequence, with matched controls preserving dinucleotide composition.

SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1
Evaluation procedure diagram
How it worksComputational evaluation flow
Computational evaluation flow1. Input: DNA sequence, with matched controls preserving dinucleotide composition.. Then: 2. Evaluation: Zero-shot likelihood comparison, frozen-model probing, LoRA fine-tuning and supervised CNN training are distinct regimes.. Then: 3. Readout: Zero-shot accuracy measures which member of a matched pair receives higher sequence likelihood. Supervised absolute classification accuracy and paired ranking accuracy are separate quantities.Computational evaluation flow1. Input: DNA sequence, with matched controls preserving dinucleotide composition.. Then: 2. Evaluation: Zero-shot likelihood comparison, frozen-model probing, LoRA fine-tuning and supervised CNN training are distinct regimes.. Then: 3. Readout: Zero-shot accuracy measures which member of a matched pair receives higher sequence likelihood. Supervised absolute classification accuracy and paired ranking accuracy are separate quantities.Computational evaluation flow1. Input: DNA sequence, with matched controls preserving dinucleotide composition.. Then: 2. Evaluation: Zero-shot likelihood comparison, frozen-model probing, LoRA fine-tuning and supervised CNN training are distinct regimes.. Then: 3. Readout: Zero-shot accuracy measures which member of a matched pair receives higher sequence likelihood. Supervised absolute classification accuracy and paired ranking accuracy are separate quantities.

Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.

SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1

limited source coverage · Automated source review, 2026-09-16. All specifications and missing details

Results

Results are available, but no reviewed comparison panel is linked in this release.

All evaluations

1 evaluation · 1 result. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: DNABERT-2Task: regulatory element identification
Dataset: DART-Eval cCREs versus matched shuffled controls
0.876 accuracy
fraction · unknown

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

DNABERT-2: regulatory element identification

zero-shot likelihood ranking: higher likelihood for cCRE than matched control

Aggregation: Not reported

DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 3 (PDF page 5), DNABERT-2 row, Zero-Shot Accuracy column

Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.

Methods and evaluation design

Procedure, tasks and evaluated configurations

How it works

Evaluation methodology

ENCODE candidate cis-regulatory elements paired with synthetic dinucleotide-shuffled controls. Test chromosomes are 5, 10, 14, 18, 20 and 22; validation uses 6 and 21; other chromosomes form training data. Supervised checkpoints are selected by validation loss. Zero-shot accuracy measures which member of a matched pair receives higher sequence likelihood. Supervised absolute classification accuracy and paired ranking accuracy are separate quantities. Six DNA language-model families are compared in zero-shot, frozen-embedding probing and LoRA fine-tuning settings; a supervised CNN supplies an ab-initio reference. Chromosome holdout separates supervised fitting from testing. Synthetic negatives preserve dinucleotide composition; pretraining sequence overlap is not resolved by chromosome splitting. The appendix reports a one-sided Wilcoxon rank-sum test for regulatory versus control likelihoods.

SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1

Recorded evaluations

Each evaluation records what was tested and under which conditions.

Run instructions

No runnable recipe has been reviewed for this task. Dataset access, model requirements, licences and compute requirements must be checked against its sources before execution.

A task describes a biological question. Choose a linked protocol to obtain concrete split and scoring instructions.

Strengths, limitations and unresolved questions

Strengths and limitations

Profile review details

Task-specific computational methodology and field context checked in the cited primary-source artifact. Source-backed fields, inapplicable evaluator dimensions and unresolved details are distinguished. Numerical results were not reproduced.

Stable record: reported-task-cdbee1c9285568

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetsENCODE candidate cis-regulatory elements paired with synthetic dinucleotide-shuffled controls.
SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1
SplitsTest chromosomes are 5, 10, 14, 18, 20 and 22; validation uses 6 and 21; other chromosomes form training data. Supervised checkpoints are selected by validation loss.
SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1
MetricsZero-shot accuracy measures which member of a matched pair receives higher sequence likelihood. Supervised absolute classification accuracy and paired ranking accuracy are separate quantities.
SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1
BaselinesSix DNA language-model families are compared in zero-shot, frozen-embedding probing and LoRA fine-tuning settings; a supervised CNN supplies an ab-initio reference.
SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1
Leakage controlsChromosome holdout separates supervised fitting from testing. Synthetic negatives preserve dinucleotide composition; pretraining sequence overlap is not resolved by chromosome splitting.
SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1
UncertaintyThe appendix reports a one-sided Wilcoxon rank-sum test for regulatory versus control likelihoods.
SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1
Entity typePaper-specific computational evaluation protocol.
SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1
OrganismsHuman.
SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1
AssaysENCODE regulatory-element annotations with synthetic sequence controls.
SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1
Allowed inputsDNA sequence, with matched controls preserving dinucleotide composition.
SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1
AdaptationZero-shot likelihood comparison, frozen-model probing, LoRA fine-tuning and supervised CNN training are distinct regimes.
SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Dated primary-source discovery and protocol/table screening. Source checking does not mean experimental reproduction. Only separately extracted and independently reviewed numeric batches are publishable.

Paper or primary resourceVersionReference
DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNANeurIPS 2024 Datasets and Benchmarks Track proceedingsRead source
Historical gaps recorded on 2026-09-17

The catalogue now holds 1 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.

  • complete numerical transcription and independent cell review: Full primary artifact and table inventory preserved; no new numeric row is published from this audit alone.
  • exact checkpoint hashes and per-method scored denominators: Table labels alone do not establish these fields; do not infer checkpoint or scored count from model name or dataset size.
Search and extraction details

primary comparison tables located

Searches

  • DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA 10.52202/079017-1981

Evidence locations

  • NeurIPS 2024 main Results and supplementary task-level tables

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

18 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.
Individual claims
DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA

Original source ↗

Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1

Version: NeurIPS 2024 Datasets and Benchmarks Track proceedings
Retrieved: 2026-09-16T10:38:57.558203+00:00

source checked

automated source review · 2026-09-16

Audit details

Task-specific computational methodology and field context checked in the cited primary-source artifact. Source-backed fields, inapplicable evaluator dimensions and unresolved details are distinguished. Numerical results were not reproduced.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: e5aee5b1f7cc6fd961b1d2a131d02cf243b79e091d5e418fbabee7fde9b39b22

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps
  • Input: DNA sequence, with matched controls preserving dinucleotide composition.
  • Evaluation: Zero-shot likelihood comparison, frozen-model probing, LoRA fine-tuning and supervised CNN training are distinct regimes.
  • Readout: Zero-shot accuracy measures which member of a matched pair receives higher sequence likelihood. Supervised absolute classification accuracy and paired ranking accuracy are separate quantities.
Individual claims
DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA

Original source ↗

Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1

Version: NeurIPS 2024 Datasets and Benchmarks Track proceedings
Retrieved: 2026-09-16T10:38:57.558203+00:00

source checked

automated source review · 2026-09-16

Audit details

Task-specific computational methodology and field context checked in the cited primary-source artifact. Source-backed fields, inapplicable evaluator dimensions and unresolved details are distinguished. Numerical results were not reproduced.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: e5aee5b1f7cc6fd961b1d2a131d02cf243b79e091d5e418fbabee7fde9b39b22

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title
Computational evaluation flow
Individual claims
DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA

Original source ↗

Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1

Version: NeurIPS 2024 Datasets and Benchmarks Track proceedings
Retrieved: 2026-09-16T10:38:57.558203+00:00

source checked

automated source review · 2026-09-16

Audit details

Task-specific computational methodology and field context checked in the cited primary-source artifact. Source-backed fields, inapplicable evaluator dimensions and unresolved details are distinguished. Numerical results were not reproduced.

Field: attributes.profile.diagram.title

Source artifact SHA-256: e5aee5b1f7cc6fd961b1d2a131d02cf243b79e091d5e418fbabee7fde9b39b22

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Datasets
ENCODE candidate cis-regulatory elements paired with synthetic dinucleotide-shuffled controls.
Individual claims
DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA

Original source ↗

Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1

Version: NeurIPS 2024 Datasets and Benchmarks Track proceedings
Retrieved: 2026-09-16T10:38:57.558203+00:00

source checked

automated source review · 2026-09-16

Audit details

Task-specific computational methodology and field context checked in the cited primary-source artifact. Source-backed fields, inapplicable evaluator dimensions and unresolved details are distinguished. Numerical results were not reproduced.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: e5aee5b1f7cc6fd961b1d2a131d02cf243b79e091d5e418fbabee7fde9b39b22

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Splits
Test chromosomes are 5, 10, 14, 18, 20 and 22; validation uses 6 and 21; other chromosomes form training data. Supervised checkpoints are selected by validation loss.
Individual claims
DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA

Original source ↗

Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1

Version: NeurIPS 2024 Datasets and Benchmarks Track proceedings
Retrieved: 2026-09-16T10:38:57.558203+00:00

source checked

automated source review · 2026-09-16

Audit details

Task-specific computational methodology and field context checked in the cited primary-source artifact. Source-backed fields, inapplicable evaluator dimensions and unresolved details are distinguished. Numerical results were not reproduced.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: e5aee5b1f7cc6fd961b1d2a131d02cf243b79e091d5e418fbabee7fde9b39b22

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Adaptation
Zero-shot likelihood comparison, frozen-model probing, LoRA fine-tuning and supervised CNN training are distinct regimes.
Individual claims
DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA

Original source ↗

Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1

Version: NeurIPS 2024 Datasets and Benchmarks Track proceedings
Retrieved: 2026-09-16T10:38:57.558203+00:00

source checked

automated source review · 2026-09-16

Audit details

Task-specific computational methodology and field context checked in the cited primary-source artifact. Source-backed fields, inapplicable evaluator dimensions and unresolved details are distinguished. Numerical results were not reproduced.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: e5aee5b1f7cc6fd961b1d2a131d02cf243b79e091d5e418fbabee7fde9b39b22

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Metrics
Zero-shot accuracy measures which member of a matched pair receives higher sequence likelihood. Supervised absolute classification accuracy and paired ranking accuracy are separate quantities.
Individual claims
DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA

Original source ↗

Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1

Version: NeurIPS 2024 Datasets and Benchmarks Track proceedings
Retrieved: 2026-09-16T10:38:57.558203+00:00

source checked

automated source review · 2026-09-16

Audit details

Task-specific computational methodology and field context checked in the cited primary-source artifact. Source-backed fields, inapplicable evaluator dimensions and unresolved details are distinguished. Numerical results were not reproduced.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: e5aee5b1f7cc6fd961b1d2a131d02cf243b79e091d5e418fbabee7fde9b39b22

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Baselines
Six DNA language-model families are compared in zero-shot, frozen-embedding probing and LoRA fine-tuning settings; a supervised CNN supplies an ab-initio reference.
Individual claims
DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA

Original source ↗

Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1

Version: NeurIPS 2024 Datasets and Benchmarks Track proceedings
Retrieved: 2026-09-16T10:38:57.558203+00:00

source checked

automated source review · 2026-09-16

Audit details

Task-specific computational methodology and field context checked in the cited primary-source artifact. Source-backed fields, inapplicable evaluator dimensions and unresolved details are distinguished. Numerical results were not reproduced.

Field: attributes.profile.facts.3.value

Source artifact SHA-256: e5aee5b1f7cc6fd961b1d2a131d02cf243b79e091d5e418fbabee7fde9b39b22

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Leakage controls
Chromosome holdout separates supervised fitting from testing. Synthetic negatives preserve dinucleotide composition; pretraining sequence overlap is not resolved by chromosome splitting.
Individual claims
DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA

Original source ↗

Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1

Version: NeurIPS 2024 Datasets and Benchmarks Track proceedings
Retrieved: 2026-09-16T10:38:57.558203+00:00

source checked

automated source review · 2026-09-16

Audit details

Task-specific computational methodology and field context checked in the cited primary-source artifact. Source-backed fields, inapplicable evaluator dimensions and unresolved details are distinguished. Numerical results were not reproduced.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: e5aee5b1f7cc6fd961b1d2a131d02cf243b79e091d5e418fbabee7fde9b39b22

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Uncertainty
The appendix reports a one-sided Wilcoxon rank-sum test for regulatory versus control likelihoods.
Individual claims
DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA

Original source ↗

Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1

Version: NeurIPS 2024 Datasets and Benchmarks Track proceedings
Retrieved: 2026-09-16T10:38:57.558203+00:00

source checked

automated source review · 2026-09-16

Audit details

Task-specific computational methodology and field context checked in the cited primary-source artifact. Source-backed fields, inapplicable evaluator dimensions and unresolved details are distinguished. Numerical results were not reproduced.

Field: attributes.profile.facts.5.value

Source artifact SHA-256: e5aee5b1f7cc6fd961b1d2a131d02cf243b79e091d5e418fbabee7fde9b39b22

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: needs review

2 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: reported-task-cdbee1c9285568

areas
dna-genomes
tasks
regulatory element identification
entity level
task
version
Not reported
task
regulatory element identification
scope note
Paper-specific evaluation task; protocol completeness requires further extraction.
benchmark research
review date: 2026-09-17; status: primary_comparison_tables_located; primary sources: evidence-expansion-dart-eval-regulatory-2024-e5aee5b1; inspected locators: NeurIPS 2024 main Results and supplementary task-level tables; searched queries: DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA 10.52202/079017-1981; gaps: complete numerical transcription and independent cell review: Full primary artifact and table inventory preserved; no new numeric row is published from this audit alone.; exact checkpoint hashes and per-method scored denominators: Table labels alone do not establish these fields; do not infer checkpoint or scored count from model name or dataset size.; claim scope: Dated primary-source discovery and protocol/table screening. Source checking does not mean experimental reproduction. Only separately extracted and independently reviewed numeric batches are publishable.
historical missing metadata
protocol version: not_reported_in_legacy_extract
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
legacy kinds
benchmark
entity classification
review date: 2026-09-17; rationale: This source-scoped record identifies the biological prediction task and holds its paper context. Preserve the existing task identity; exact split, model adaptation and scoring remain in linked evaluations or separate protocol records.; source ids: dart-eval-regulatory-2024; source locator: Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1; ambiguities: A paper- or suite-specific task may constrain some inputs or metrics; that alone does not make it interchangeable with a complete versioned protocol. No protocol equivalence is inferred.; Some legacy profile Entity type facts use the generic phrase computational evaluation protocol. That boilerplate is not sufficient to establish a single fixed protocol identity or to merge this task with another protocol record.
Related records

Suggest a correction