Datasets
ENCODE candidate cis-regulatory elements paired with synthetic dinucleotide-shuffled controls.
Regulatory-element identification tests whether DNA models distinguish ENCODE regulatory sequences from composition-matched controls.
ENCODE candidate cis-regulatory elements paired with synthetic dinucleotide-shuffled controls.
Zero-shot accuracy measures which member of a matched pair receives higher sequence likelihood. Supervised absolute classification accuracy and paired ranking accuracy are separate quantities.
DNA sequence, with matched controls preserving dinucleotide composition.
Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.
limited source coverage · Automated source review, 2026-09-16. All specifications and missing details
Results are available, but no reviewed comparison panel is linked in this release.
1 evaluation · 1 result. Different protocols are not a single leaderboard.
Applied filters: All linked evaluations
| Tested configuration | Protocol and dataset | Finding | Evidence and details |
|---|---|---|---|
| Configuration: DNABERT-2 | Task: regulatory element identification Dataset: DART-Eval cCREs versus matched shuffled controls | 0.876 accuracy fraction · unknown Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceDNABERT-2: regulatory element identification zero-shot likelihood ranking: higher likelihood for cCRE than matched control Aggregation: Not reported DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 3 (PDF page 5), DNABERT-2 row, Zero-Shot Accuracy column |
Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.
ENCODE candidate cis-regulatory elements paired with synthetic dinucleotide-shuffled controls. Test chromosomes are 5, 10, 14, 18, 20 and 22; validation uses 6 and 21; other chromosomes form training data. Supervised checkpoints are selected by validation loss. Zero-shot accuracy measures which member of a matched pair receives higher sequence likelihood. Supervised absolute classification accuracy and paired ranking accuracy are separate quantities. Six DNA language-model families are compared in zero-shot, frozen-embedding probing and LoRA fine-tuning settings; a supervised CNN supplies an ab-initio reference. Chromosome holdout separates supervised fitting from testing. Synthetic negatives preserve dinucleotide composition; pretraining sequence overlap is not resolved by chromosome splitting. The appendix reports a one-sided Wilcoxon rank-sum test for regulatory versus control likelihoods.
Each evaluation records what was tested and under which conditions.
No runnable recipe has been reviewed for this task. Dataset access, model requirements, licences and compute requirements must be checked against its sources before execution.
A task describes a biological question. Choose a linked protocol to obtain concrete split and scoring instructions.
Task-specific computational methodology and field context checked in the cited primary-source artifact. Source-backed fields, inapplicable evaluator dimensions and unresolved details are distinguished. Numerical results were not reproduced.
Stable record: reported-task-cdbee1c9285568Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Datasets | ENCODE candidate cis-regulatory elements paired with synthetic dinucleotide-shuffled controls.SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1 |
| Splits | Test chromosomes are 5, 10, 14, 18, 20 and 22; validation uses 6 and 21; other chromosomes form training data. Supervised checkpoints are selected by validation loss.SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1 |
| Metrics | Zero-shot accuracy measures which member of a matched pair receives higher sequence likelihood. Supervised absolute classification accuracy and paired ranking accuracy are separate quantities.SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1 |
| Baselines | Six DNA language-model families are compared in zero-shot, frozen-embedding probing and LoRA fine-tuning settings; a supervised CNN supplies an ab-initio reference.SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1 |
| Leakage controls | Chromosome holdout separates supervised fitting from testing. Synthetic negatives preserve dinucleotide composition; pretraining sequence overlap is not resolved by chromosome splitting.SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1 |
| Uncertainty | The appendix reports a one-sided Wilcoxon rank-sum test for regulatory versus control likelihoods.SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1 |
| Entity type | Paper-specific computational evaluation protocol.SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1 |
| Organisms | Human.SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1 |
| Assays | ENCODE regulatory-element annotations with synthetic sequence controls.SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1 |
| Allowed inputs | DNA sequence, with matched controls preserving dinucleotide composition.SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1 |
| Adaptation | Zero-shot likelihood comparison, frozen-model probing, LoRA fine-tuning and supervised CNN training are distinct regimes.SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1 |
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Last literature check: 2026-09-17. Dated primary-source discovery and protocol/table screening. Source checking does not mean experimental reproduction. Only separately extracted and independently reviewed numeric batches are publishable.
| Paper or primary resource | Version | Reference |
|---|---|---|
| DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA | NeurIPS 2024 Datasets and Benchmarks Track proceedings | Read source |
The catalogue now holds 1 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.
primary comparison tables located
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
18 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol. Individual claims | DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1 Version: NeurIPS 2024 Datasets and Benchmarks Track proceedings | source checked automated source review · 2026-09-16 Audit detailsTask-specific computational methodology and field context checked in the cited primary-source artifact. Source-backed fields, inapplicable evaluator dimensions and unresolved details are distinguished. Numerical results were not reproduced. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1 Version: NeurIPS 2024 Datasets and Benchmarks Track proceedings | source checked automated source review · 2026-09-16 Audit detailsTask-specific computational methodology and field context checked in the cited primary-source artifact. Source-backed fields, inapplicable evaluator dimensions and unresolved details are distinguished. Numerical results were not reproduced. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Computational evaluation flow Individual claims | DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1 Version: NeurIPS 2024 Datasets and Benchmarks Track proceedings | source checked automated source review · 2026-09-16 Audit detailsTask-specific computational methodology and field context checked in the cited primary-source artifact. Source-backed fields, inapplicable evaluator dimensions and unresolved details are distinguished. Numerical results were not reproduced. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Datasets ENCODE candidate cis-regulatory elements paired with synthetic dinucleotide-shuffled controls. Individual claims | DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1 Version: NeurIPS 2024 Datasets and Benchmarks Track proceedings | source checked automated source review · 2026-09-16 Audit detailsTask-specific computational methodology and field context checked in the cited primary-source artifact. Source-backed fields, inapplicable evaluator dimensions and unresolved details are distinguished. Numerical results were not reproduced. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Splits Test chromosomes are 5, 10, 14, 18, 20 and 22; validation uses 6 and 21; other chromosomes form training data. Supervised checkpoints are selected by validation loss. Individual claims | DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1 Version: NeurIPS 2024 Datasets and Benchmarks Track proceedings | source checked automated source review · 2026-09-16 Audit detailsTask-specific computational methodology and field context checked in the cited primary-source artifact. Source-backed fields, inapplicable evaluator dimensions and unresolved details are distinguished. Numerical results were not reproduced. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Adaptation Zero-shot likelihood comparison, frozen-model probing, LoRA fine-tuning and supervised CNN training are distinct regimes. Individual claims | DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1 Version: NeurIPS 2024 Datasets and Benchmarks Track proceedings | source checked automated source review · 2026-09-16 Audit detailsTask-specific computational methodology and field context checked in the cited primary-source artifact. Source-backed fields, inapplicable evaluator dimensions and unresolved details are distinguished. Numerical results were not reproduced. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Metrics Zero-shot accuracy measures which member of a matched pair receives higher sequence likelihood. Supervised absolute classification accuracy and paired ranking accuracy are separate quantities. Individual claims | DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1 Version: NeurIPS 2024 Datasets and Benchmarks Track proceedings | source checked automated source review · 2026-09-16 Audit detailsTask-specific computational methodology and field context checked in the cited primary-source artifact. Source-backed fields, inapplicable evaluator dimensions and unresolved details are distinguished. Numerical results were not reproduced. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Baselines Six DNA language-model families are compared in zero-shot, frozen-embedding probing and LoRA fine-tuning settings; a supervised CNN supplies an ab-initio reference. Individual claims | DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1 Version: NeurIPS 2024 Datasets and Benchmarks Track proceedings | source checked automated source review · 2026-09-16 Audit detailsTask-specific computational methodology and field context checked in the cited primary-source artifact. Source-backed fields, inapplicable evaluator dimensions and unresolved details are distinguished. Numerical results were not reproduced. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Leakage controls Chromosome holdout separates supervised fitting from testing. Synthetic negatives preserve dinucleotide composition; pretraining sequence overlap is not resolved by chromosome splitting. Individual claims | DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1 Version: NeurIPS 2024 Datasets and Benchmarks Track proceedings | source checked automated source review · 2026-09-16 Audit detailsTask-specific computational methodology and field context checked in the cited primary-source artifact. Source-backed fields, inapplicable evaluator dimensions and unresolved details are distinguished. Numerical results were not reproduced. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Uncertainty The appendix reports a one-sided Wilcoxon rank-sum test for regulatory versus control likelihoods. Individual claims | DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA Paper §3, §4.1, Table 3; Appendices D.2–D.3 and E.1 Version: NeurIPS 2024 Datasets and Benchmarks Track proceedings | source checked automated source review · 2026-09-16 Audit detailsTask-specific computational methodology and field context checked in the cited primary-source artifact. Source-backed fields, inapplicable evaluator dimensions and unresolved details are distinguished. Numerical results were not reproduced. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
View linked audit checks and correction history
Release 2026-09-29-06401fd5b220 · Record review: needs review
Stable ID: reported-task-cdbee1c9285568