Datasets
Task-specific HDF5 inputs/outputs with raw data and evaluated model outputs organized in a Synapse project.
DART-Eval measures human regulatory-DNA representations under zero-shot, probing and fine-tuning regimes.
Task-specific HDF5 inputs/outputs with raw data and evaluated model outputs organized in a Synapse project.
Regulatory and variant classification use AUROC/AUPRC; quantitative accessibility uses Pearson and Spearman correlations on peaks alone and peaks plus background. Cell-type clustering uses adjusted mutual information. Motif sensitivity is a separate paired-sequence evaluation.
Task-specific DNA sequences and HDF5 targets; raw data and model outputs are linked through Synapse.
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Source reviewed · Automated source review, 2026-09-16. All specifications and missing details
Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.
auroc (fraction) · Higher values are better.
DART-Eval CA-AUROC-GM12878: Chromatin activity prediction, GM12878, positives against negatives · ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split)
Evidence origin: Author-reported evaluation.
DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 5, column(AUROC GM12878)Every method DART-Eval reports on Chromatin activity prediction, GM12878, positives against negatives, scored with AUROC on ENCODE chromatin accessibility peaks in five cell lines.
Automated source review: 2026-09-19. Numerical source review does not establish independent reproduction.
Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.
Showing 12 of 13 matching rows.
DART-Eval separates regulatory sequence discrimination, motif sensitivity, cell-type specificity, quantitative accessibility and variant effects. It uses matched controls and consistent chromosome partitions where models are fitted. A result therefore needs its precise task and adaptation setting, rather than a single regulatory-intelligence score.
Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.
These source-backed links do not make different protocols or scores interchangeable.
Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.
No concrete protocols are explicitly linked to this suite. Protocol identification and baseline selection are outstanding.
Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums
Coverage is derived from release 2026-09-29-06401fd5b220. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.
Score regulatory sequences against their existing paired shuffled controls without model fitting. Strict element-score greater-than-control accuracy; paired score-difference summaries and pinned SciPy1.12 one-sided Wilcoxon procedure.
Recompute metrics from supplied predictions. This recipe does not establish reproduction of a particular published score.
Choose one way to run this recipe. These instruction formats are alternatives.
Source reviewed; these instructions have not been executed by rewire.
import rewirebench
prepared = rewirebench.prepare(
'dart-eval-task1-zero-shot-v1', source='demo', output='prepared-dart-task1'
)
report = rewirebench.evaluate(
prepared, 'predictions.json', output='results-dart-task1-rescore',
model={'name': 'My private model', 'training_overlap': 'Unreported'},
)rewirebench 0.4 library interface; rewirebench container and HPC instructions; DART-Eval Task1 zero-shot paired controls runner guide; DART-Eval Task1 zero-shot paired controls protocol implementation; DART-Eval Task1 zero-shot paired controls source manifest; DART-Eval Task1 zero-shot paired controls automated validation receipt · docs/dart-eval.md and protocol implementation; linked validation receipt concerns its stated tests, not execution of this exact snippetRun your model locally and return predictions keyed by the input IDs. The evaluator supplies biological inputs without test labels and owns scoring. This interface is not a sandbox for model code.
Pass your existing prediction function into this adapter. Its output direction must match the selected protocol.
class MyModelAdapter:
def __init__(self, score):
self.score = score
def predict(self, inputs):
return {row["id"]: float(self.score(row)) for row in inputs}
# adapter = MyModelAdapter(your_prediction_function)
# report = rewirebench.run(prepared, adapter, output="runs/my-model")Alternatively, generate a keyed prediction file in your existing model environment and use the score-only recipe. Your model code and weights do not need to be shared.
rewirebench 0.4 library interface; rewirebench container and HPC instructions; DART-Eval Task1 zero-shot paired controls runner guide; DART-Eval Task1 zero-shot paired controls protocol implementation; DART-Eval Task1 zero-shot paired controls source manifest; DART-Eval Task1 zero-shot paired controls automated validation receipt · docs/dart-eval.md; source manifest; protocol score function; automated receipt scopeContribute a result for review. The library can submit an exported evaluation for private review when intake is open. Check the contribution page for access and sign-in.
The official README provides task-specific Python module commands, model identifiers, DART_WORK_DIR layout and Synapse resource IDs. These inspected sections do not document dependency installation or a complete environment lock. The Task 1 link label and target differ; resolve the exact Synapse resource before assembling a copyable recipe.
A maintained rewire runner has not been verified for this benchmark. Check data access, weights, licences, dependencies and hardware in the linked official documentation; requirements have not been fully extracted.
kundajelab/DART-Eval / README.md · README.md lines 8–26 and 28–110 (Data, Preliminaries and Task 1 commands)Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.
Stable record: discovery-benchmark-dart-evalExplanatory profile: source reviewed · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Datasets | Task-specific HDF5 inputs/outputs with raw data and evaluated model outputs organized in a Synapse project.Sourceskundajelab/DART-Eval official source · Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections |
| Splits | Training uses chromosomes other than the held-out sets; validation is chromosomes 6 and 21; test is chromosomes 5, 10, 14, 18, 20 and 22. Fitted checkpoints are chosen using validation loss.Sourcesdart primary benchmark evidence · Appendix: common train/validation/test split |
| Metrics | Regulatory and variant classification use AUROC/AUPRC; quantitative accessibility uses Pearson and Spearman correlations on peaks alone and peaks plus background. Cell-type clustering uses adjusted mutual information. Motif sensitivity is a separate paired-sequence evaluation.Sourcesdart primary benchmark evidence · Appendix: training/test splits, clustering and supervised evaluation; reproducibility checklist |
| Baselines | Probing-head-like ab initio models are documented alongside pretrained-model probing and fine-tuning.Sourceskundajelab/DART-Eval official source · Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections |
| Leakage controls | The common genomic partition holds out chromosomes 5, 10, 14, 18, 20 and 22 for test, with chromosomes 6 and 21 for validation. Lowest validation loss selects checkpoints. Dinucleotide-shuffled negatives preserve composition; this control does not eliminate all pretrained-genome exposure.Sourcesdart primary benchmark evidence · Appendix: training/test splits, clustering and supervised evaluation; reproducibility checklist |
| Uncertainty | Clustering is repeated 100 times and reports a 95% interval across clustering runs. Motif-sensitivity intervals are provided in the linked artifacts; the paper does not report repeated-training intervals for every other task.Sourcesdart primary benchmark evidence · Appendix: training/test splits, clustering and supervised evaluation; reproducibility checklist |
| Entity type | DNA regulatory evaluation suite.Sourceskundajelab/DART-Eval official source · Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections |
| Organisms | Human regulatory datasets.Sourceskundajelab/DART-Eval official source · Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections |
| Assays | Regulatory-element, chromatin-accessibility and variant-effect measurements.Sourceskundajelab/DART-Eval official source · Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections |
| Allowed inputs | Task-specific DNA sequences and HDF5 targets; raw data and model outputs are linked through Synapse.Sourceskundajelab/DART-Eval official source · Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections |
| Adaptation | Zero-shot, probing, fine-tuning and ab-initio comparison regimes are distinct.Sourceskundajelab/DART-Eval official source · Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections |
Applicability is distinct from a completed evaluation.
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Last literature check: 2026-09-17. Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.
| Paper or primary resource | Version | Reference |
|---|---|---|
| DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA | 2412.05430v1 | Read source |
The catalogue now holds 316 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.
source found structured extraction pending
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
127 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions. Individual claims | dart primary benchmark evidence Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections; Appendix: common train/validation/test split; Appendix: training/test splits, clustering and supervised evaluation; reproducibility checklist Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 2412.05430v1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram caption Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions. Individual claims | kundajelab/DART-Eval official source Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections; Appendix: common train/validation/test split; Appendix: training/test splits, clustering and supervised evaluation; reproducibility checklist Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: af2a86d666c35304257c2fa7e15180e1fbcabb01 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| dart primary benchmark evidence Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections; Appendix: common train/validation/test split; Appendix: training/test splits, clustering and supervised evaluation; reproducibility checklist Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 2412.05430v1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| kundajelab/DART-Eval official source Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections; Appendix: common train/validation/test split; Appendix: training/test splits, clustering and supervised evaluation; reproducibility checklist Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: af2a86d666c35304257c2fa7e15180e1fbcabb01 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Evaluation procedure Individual claims | dart primary benchmark evidence Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections; Appendix: common train/validation/test split; Appendix: training/test splits, clustering and supervised evaluation; reproducibility checklist Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 2412.05430v1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Evaluation procedure Individual claims | kundajelab/DART-Eval official source Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections; Appendix: common train/validation/test split; Appendix: training/test splits, clustering and supervised evaluation; reproducibility checklist Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: af2a86d666c35304257c2fa7e15180e1fbcabb01 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Datasets Task-specific HDF5 inputs/outputs with raw data and evaluated model outputs organized in a Synapse project. Individual claims | kundajelab/DART-Eval official source Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections Version: af2a86d666c35304257c2fa7e15180e1fbcabb01 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Splits Training uses chromosomes other than the held-out sets; validation is chromosomes 6 and 21; test is chromosomes 5, 10, 14, 18, 20 and 22. Fitted checkpoints are chosen using validation loss. Individual claims | dart primary benchmark evidence Appendix: common train/validation/test split Version: 2412.05430v1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Adaptation Zero-shot, probing, fine-tuning and ab-initio comparison regimes are distinct. Individual claims | kundajelab/DART-Eval official source Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections Version: af2a86d666c35304257c2fa7e15180e1fbcabb01 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Metrics Regulatory and variant classification use AUROC/AUPRC; quantitative accessibility uses Pearson and Spearman correlations on peaks alone and peaks plus background. Cell-type clustering uses adjusted mutual information. Motif sensitivity is a separate paired-sequence evaluation. Individual claims | dart primary benchmark evidence Appendix: training/test splits, clustering and supervised evaluation; reproducibility checklist Version: 2412.05430v1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
View linked audit checks and correction history
Release 2026-09-29-06401fd5b220 · Record review: discovered
Stable ID: discovery-benchmark-dart-eval