rewire.itbenchmarks
Benchmark

DART-Eval

DART-Eval measures human regulatory-DNA representations under zero-shot, probing and fine-tuning regimes.

Sourceskundajelab/DART-Eval official source · Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections

316 evaluations · 316 results

Overview

Datasets

Task-specific HDF5 inputs/outputs with raw data and evaluated model outputs organized in a Synapse project.

Sourceskundajelab/DART-Eval official source · Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections

Metrics

Regulatory and variant classification use AUROC/AUPRC; quantitative accessibility uses Pearson and Spearman correlations on peaks alone and peaks plus background. Cell-type clustering uses adjusted mutual information. Motif sensitivity is a separate paired-sequence evaluation.

Sourcesdart primary benchmark evidence · Appendix: training/test splits, clustering and supervised evaluation; reproducibility checklist

Allowed inputs

Task-specific DNA sequences and HDF5 targets; raw data and model outputs are linked through Synapse.

Sourceskundajelab/DART-Eval official source · Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections
Evaluation procedure diagram
How it worksEvaluation procedure
Evaluation procedure1. Allowed inputs: Task-specific DNA sequences and HDF5 targets; raw data and model outputs are linked through Synapse.. Then: 2. Splits: Training uses chromosomes other than the held-out sets; validation is chromosomes 6 and 21; test is chromosomes 5, 10, 14, 18, 20 and 22. Fitted checkpoints are chosen using validation loss.. Then: 3. Metrics: Regulatory and variant classification use AUROC/AUPRC; quantitative accessibility uses Pearson and Spearman correlations on peaks alone and peaks plus background. Cell-type clustering uses adjusted mutual information. Motif sensitivity is a separate paired-sequence evaluation.Evaluation procedure1. Allowed inputs: Task-specific DNA sequences and HDF5 targets; raw data and model outputs are linked through Synapse.. Then: 2. Splits: Training uses chromosomes other than the held-out sets; validation is chromosomes 6 and 21; test is chromosomes 5, 10, 14, 18, 20 and 22. Fitted checkpoints are chosen using validation loss.. Then: 3. Metrics: Regulatory and variant classification use AUROC/AUPRC; quantitative accessibility uses Pearson and Spearman correlations on peaks alone and peaks plus background. Cell-type clustering uses adjusted mutual information. Motif sensitivity is a separate paired-sequence evaluation.Evaluation procedure1. Allowed inputs: Task-specific DNA sequences and HDF5 targets; raw data and model outputs are linked through Synapse.. Then: 2. Splits: Training uses chromosomes other than the held-out sets; validation is chromosomes 6 and 21; test is chromosomes 5, 10, 14, 18, 20 and 22. Fitted checkpoints are chosen using validation loss.. Then: 3. Metrics: Regulatory and variant classification use AUROC/AUPRC; quantitative accessibility uses Pearson and Spearman correlations on peaks alone and peaks plus background. Cell-type clustering uses adjusted mutual information. Motif sensitivity is a separate paired-sequence evaluation.

Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.

Sources (2)kundajelab/DART-Eval official source; dart primary benchmark evidence · Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections; Appendix: common train/validation/test split; Appendix: training/test splits, clustering and supervised evaluation; reproducibility checklist

Source reviewed · Automated source review, 2026-09-16. All specifications and missing details

Results

Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.

DART-Eval CA-AUROC-GM12878: Chromatin activity prediction, GM12878, positives against negatives

auroc (fraction) · Higher values are better.

DART-Eval CA-AUROC-GM12878: Chromatin activity prediction, GM12878, positives against negatives · ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split)

Evidence origin: Author-reported evaluation.

DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 5, column(AUROC GM12878)
  • The evaluation setting is part of the method name: a zero-shot, probed and fine-tuned run of the same model are different entries.
  • Metrics and datasets differ between tasks, so these figures cannot be averaged into one score.
Comparison details and limitations

Every method DART-Eval reports on Chromatin activity prediction, GM12878, positives against negatives, scored with AUROC on ENCODE chromatin accessibility peaks in five cell lines.

  • Author-reported numbers, source checked but not independently reproduced.

Automated source review: 2026-09-19. Numerical source review does not establish independent reproduction.

Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.

Showing 12 of 13 matching rows.

Methods and evaluation design

Procedure, tasks and evaluated configurations

How it works

Evaluation methodology

DART-Eval separates regulatory sequence discrimination, motif sensitivity, cell-type specificity, quantitative accessibility and variant effects. It uses matched controls and consistent chromosome partitions where models are fitted. A result therefore needs its precise task and adaptation setting, rather than a single regulatory-intelligence score.

Sourcesdart primary benchmark evidence · Appendix: training/test splits, clustering and supervised evaluation; reproducibility checklist

Evaluation design

Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.

Tasks

These source-backed links do not make different protocols or scores interchangeable.

Baseline coverage

Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.

No concrete protocols are explicitly linked to this suite. Protocol identification and baseline selection are outstanding.

Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums

Coverage is derived from release 2026-09-29-06401fd5b220. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.

Run this benchmark

DART-Eval Task1 zero-shot paired controls: score supplied predictions

Score regulatory sequences against their existing paired shuffled controls without model fitting. Strict element-score greater-than-control accuracy; paired score-difference summaries and pinned SciPy1.12 one-sided Wilcoxon procedure.

Recompute metrics from supplied predictions. This recipe does not establish reproduction of a particular published score.

Dataset access
Official one-hot paired input is Synapse syn64314109 version1; authenticated access is required. Anonymous downloads returned403 and official bytes/complete denominator remain unverified. The local demo is12 synthetic pairs.
Model and weights
No weights needed for rescoring or the supplied toy/composition control. Private-model weights remain local; their access requirements depend on the model.
Licences
Runner MIT. Dataset reuse terms are not established by the software licence; review Synapse access and original ENCODE terms.
Software
Python3.11, pinned rewirebench0.4 environment; sequence extra for parquet and HDF5. Podman or Apptainer is optional.
Hardware
CPU scoring and small controls; memory/accelerator needs for real private-model inference are model-dependent and unreported here.
Required inputs and expected outputs

Inputs

  • Prepared protocol inputs with evaluator-owned labels and opaque IDs.
  • Keyed finite scalar scores; missing predictions must be explicitly permitted and remain in coverage.

Outputs

  • Local report.json with metrics, source/version identity and coverage.
  • Local predictions.json and unscored.json; nothing submitted automatically.

Choose one way to run this recipe. These instruction formats are alternatives.

Execution steps

  1. 1. Score supplied keyed predictions (Python)

    Source reviewed; these instructions have not been executed by rewire.

    import rewirebench
    prepared = rewirebench.prepare(
        'dart-eval-task1-zero-shot-v1', source='demo', output='prepared-dart-task1'
    )
    report = rewirebench.evaluate(
        prepared, 'predictions.json', output='results-dart-task1-rescore',
        model={'name': 'My private model', 'training_overlap': 'Unreported'},
    )
    rewirebench 0.4 library interface; rewirebench container and HPC instructions; DART-Eval Task1 zero-shot paired controls runner guide; DART-Eval Task1 zero-shot paired controls protocol implementation; DART-Eval Task1 zero-shot paired controls source manifest; DART-Eval Task1 zero-shot paired controls automated validation receipt · docs/dart-eval.md and protocol implementation; linked validation receipt concerns its stated tests, not execution of this exact snippet

Use your own model

Run your model locally and return predictions keyed by the input IDs. The evaluator supplies biological inputs without test labels and owns scoring. This interface is not a sandbox for model code.

Pass your existing prediction function into this adapter. Its output direction must match the selected protocol.

class MyModelAdapter:
    def __init__(self, score):
        self.score = score

    def predict(self, inputs):
        return {row["id"]: float(self.score(row)) for row in inputs}

# adapter = MyModelAdapter(your_prediction_function)
# report = rewirebench.run(prepared, adapter, output="runs/my-model")

Alternatively, generate a keyed prediction file in your existing model environment and use the score-only recipe. Your model code and weights do not need to be shared.

rewirebench 0.4 library interface; rewirebench container and HPC instructions; DART-Eval Task1 zero-shot paired controls runner guide; DART-Eval Task1 zero-shot paired controls protocol implementation; DART-Eval Task1 zero-shot paired controls source manifest; DART-Eval Task1 zero-shot paired controls automated validation receipt · docs/dart-eval.md; source manifest; protocol score function; automated receipt scope
Scope and limitations
  • Only zero-shot Task1 is supported, not the full DART-Eval suite; training and embedding probes are forbidden.
  • Official biological inputs and DNABERT-2 prediction artifacts require authenticated access. User-supplied local hashes prove integrity, not source authenticity or complete chromosome coverage.
  • The demo contains synthetic random DNA with shuffled controls. Its toy score checks software only and is not a biological baseline.
  • Missing either pair member excludes that pair from metrics; sequence and complete-pair coverage stay separate. Ties are incorrect for accuracy.
  • Scoring is tested on synthetic inputs against the pinned upstream evaluator and historical SciPy rules. No official biological inference or published evaluation reproduction is claimed.
  • These snippets are source-reviewed instructions. The linked automated receipt supports only the tests it records; it does not certify the exact snippet or published-model reproduction.
  • Smoke examples and partial predictions retain original denominators and cannot establish a complete suite. Prepared execution is offline; exports and submissions are explicit.

Contribute a result for review. The library can submit an exported evaluation for private review when intake is open. Check the contribution page for access and sign-in.

Original repository instructions

Run this benchmark

The official README provides task-specific Python module commands, model identifiers, DART_WORK_DIR layout and Synapse resource IDs. These inspected sections do not document dependency installation or a complete environment lock. The Task 1 link label and target differ; resolve the exact Synapse resource before assembling a copyable recipe.

A maintained rewire runner has not been verified for this benchmark. Check data access, weights, licences, dependencies and hardware in the linked official documentation; requirements have not been fully extracted.

kundajelab/DART-Eval / README.md · README.md lines 8–26 and 28–110 (Data, Preliminaries and Task 1 commands)
Strengths, limitations and unresolved questions

Strengths and limitations

Strengths supported by sources

  • Tasks separate generic regulatory recognition from cell-specific activity and variant effects.
    Sourceskundajelab/DART-Eval official source · Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections

Limitations and conditions

  • The assessed tasks use local regulatory contexts. A chromosome-held-out supervised test does not establish that those sequences were absent from unsupervised pretraining.
    Sourcesdart primary benchmark evidence · Appendix: training/test splits, clustering and supervised evaluation; reproducibility checklist
Profile review details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Stable record: discovery-benchmark-dart-eval

Specifications

Inputs, training, access and other details

Explanatory profile: source reviewed · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetsTask-specific HDF5 inputs/outputs with raw data and evaluated model outputs organized in a Synapse project.
Sourceskundajelab/DART-Eval official source · Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections
SplitsTraining uses chromosomes other than the held-out sets; validation is chromosomes 6 and 21; test is chromosomes 5, 10, 14, 18, 20 and 22. Fitted checkpoints are chosen using validation loss.
Sourcesdart primary benchmark evidence · Appendix: common train/validation/test split
MetricsRegulatory and variant classification use AUROC/AUPRC; quantitative accessibility uses Pearson and Spearman correlations on peaks alone and peaks plus background. Cell-type clustering uses adjusted mutual information. Motif sensitivity is a separate paired-sequence evaluation.
Sourcesdart primary benchmark evidence · Appendix: training/test splits, clustering and supervised evaluation; reproducibility checklist
BaselinesProbing-head-like ab initio models are documented alongside pretrained-model probing and fine-tuning.
Sourceskundajelab/DART-Eval official source · Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections
Leakage controlsThe common genomic partition holds out chromosomes 5, 10, 14, 18, 20 and 22 for test, with chromosomes 6 and 21 for validation. Lowest validation loss selects checkpoints. Dinucleotide-shuffled negatives preserve composition; this control does not eliminate all pretrained-genome exposure.
Sourcesdart primary benchmark evidence · Appendix: training/test splits, clustering and supervised evaluation; reproducibility checklist
UncertaintyClustering is repeated 100 times and reports a 95% interval across clustering runs. Motif-sensitivity intervals are provided in the linked artifacts; the paper does not report repeated-training intervals for every other task.
Sourcesdart primary benchmark evidence · Appendix: training/test splits, clustering and supervised evaluation; reproducibility checklist
Entity typeDNA regulatory evaluation suite.
Sourceskundajelab/DART-Eval official source · Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections
OrganismsHuman regulatory datasets.
Sourceskundajelab/DART-Eval official source · Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections
AssaysRegulatory-element, chromatin-accessibility and variant-effect measurements.
Sourceskundajelab/DART-Eval official source · Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections
Allowed inputsTask-specific DNA sequences and HDF5 targets; raw data and model outputs are linked through Synapse.
Sourceskundajelab/DART-Eval official source · Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections
AdaptationZero-shot, probing, fine-tuning and ab-initio comparison regimes are distinct.
Sourceskundajelab/DART-Eval official source · Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections
Applicable tests and references

Applicability is distinct from a completed evaluation.

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.

Historical gaps recorded on 2026-09-17

The catalogue now holds 316 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.

  • Regulatory sequence detection, TF motifs, cell-type activity and variant effects are distinct protocols. Enhancer broad-task record must link concrete tasks, not inherit every DART result. Existing DART observations require source-cell identity reuse.
Search and extraction details

source found structured extraction pending

Searches

  • DART-Eval primary paper benchmark results

Evidence locations

  • arXiv v1 main benchmark procedures and results tables

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

127 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Individual claims
dart primary benchmark evidence

Original source ↗

Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections; Appendix: common train/validation/test split; Appendix: training/test splits, clustering and supervised evaluation; reproducibility checklist

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2412.05430v1
Retrieved: 2026-09-16T21:06:29.715960+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 4194b137ba55c9a2c269d119a9afec6ae1bb0feaf17d91433ae483c41221a56b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram caption
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Individual claims
kundajelab/DART-Eval official source

Original source ↗

Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections; Appendix: common train/validation/test split; Appendix: training/test splits, clustering and supervised evaluation; reproducibility checklist

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: af2a86d666c35304257c2fa7e15180e1fbcabb01
Retrieved: 2026-09-16T10:30:21.153480+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: d21ee86b56c9b794f5b58f3c39c3e27c51d027a3b280d848457abf53f521f052

Hash scope: Hash scope not separately documented; inspect source record

Diagram steps
  • Allowed inputs: Task-specific DNA sequences and HDF5 targets; raw data and model outputs are linked through Synapse.
  • Splits: Training uses chromosomes other than the held-out sets; validation is chromosomes 6 and 21; test is chromosomes 5, 10, 14, 18, 20 and 22. Fitted checkpoints are chosen using validation loss.
  • Metrics: Regulatory and variant classification use AUROC/AUPRC; quantitative accessibility uses Pearson and Spearman correlations on peaks alone and peaks plus background. Cell-type clustering uses adjusted mutual information. Motif sensitivity is a separate paired-sequence evaluation.
Individual claims
dart primary benchmark evidence

Original source ↗

Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections; Appendix: common train/validation/test split; Appendix: training/test splits, clustering and supervised evaluation; reproducibility checklist

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2412.05430v1
Retrieved: 2026-09-16T21:06:29.715960+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 4194b137ba55c9a2c269d119a9afec6ae1bb0feaf17d91433ae483c41221a56b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps
  • Allowed inputs: Task-specific DNA sequences and HDF5 targets; raw data and model outputs are linked through Synapse.
  • Splits: Training uses chromosomes other than the held-out sets; validation is chromosomes 6 and 21; test is chromosomes 5, 10, 14, 18, 20 and 22. Fitted checkpoints are chosen using validation loss.
  • Metrics: Regulatory and variant classification use AUROC/AUPRC; quantitative accessibility uses Pearson and Spearman correlations on peaks alone and peaks plus background. Cell-type clustering uses adjusted mutual information. Motif sensitivity is a separate paired-sequence evaluation.
Individual claims
kundajelab/DART-Eval official source

Original source ↗

Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections; Appendix: common train/validation/test split; Appendix: training/test splits, clustering and supervised evaluation; reproducibility checklist

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: af2a86d666c35304257c2fa7e15180e1fbcabb01
Retrieved: 2026-09-16T10:30:21.153480+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: d21ee86b56c9b794f5b58f3c39c3e27c51d027a3b280d848457abf53f521f052

Hash scope: Hash scope not separately documented; inspect source record

Diagram title
Evaluation procedure
Individual claims
dart primary benchmark evidence

Original source ↗

Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections; Appendix: common train/validation/test split; Appendix: training/test splits, clustering and supervised evaluation; reproducibility checklist

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2412.05430v1
Retrieved: 2026-09-16T21:06:29.715960+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: 4194b137ba55c9a2c269d119a9afec6ae1bb0feaf17d91433ae483c41221a56b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title
Evaluation procedure
Individual claims
kundajelab/DART-Eval official source

Original source ↗

Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections; Appendix: common train/validation/test split; Appendix: training/test splits, clustering and supervised evaluation; reproducibility checklist

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: af2a86d666c35304257c2fa7e15180e1fbcabb01
Retrieved: 2026-09-16T10:30:21.153480+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: d21ee86b56c9b794f5b58f3c39c3e27c51d027a3b280d848457abf53f521f052

Hash scope: Hash scope not separately documented; inspect source record

Datasets
Task-specific HDF5 inputs/outputs with raw data and evaluated model outputs organized in a Synapse project.
Individual claims
kundajelab/DART-Eval official source

Original source ↗

Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections

Version: af2a86d666c35304257c2fa7e15180e1fbcabb01
Retrieved: 2026-09-16T10:30:21.153480+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: d21ee86b56c9b794f5b58f3c39c3e27c51d027a3b280d848457abf53f521f052

Hash scope: Hash scope not separately documented; inspect source record

Splits
Training uses chromosomes other than the held-out sets; validation is chromosomes 6 and 21; test is chromosomes 5, 10, 14, 18, 20 and 22. Fitted checkpoints are chosen using validation loss.
Individual claims
dart primary benchmark evidence

Original source ↗

Appendix: common train/validation/test split

Version: 2412.05430v1
Retrieved: 2026-09-16T21:06:29.715960+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: 4194b137ba55c9a2c269d119a9afec6ae1bb0feaf17d91433ae483c41221a56b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Adaptation
Zero-shot, probing, fine-tuning and ab-initio comparison regimes are distinct.
Individual claims
kundajelab/DART-Eval official source

Original source ↗

Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections

Version: af2a86d666c35304257c2fa7e15180e1fbcabb01
Retrieved: 2026-09-16T10:30:21.153480+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: d21ee86b56c9b794f5b58f3c39c3e27c51d027a3b280d848457abf53f521f052

Hash scope: Hash scope not separately documented; inspect source record

Metrics
Regulatory and variant classification use AUROC/AUPRC; quantitative accessibility uses Pearson and Spearman correlations on peaks alone and peaks plus background. Cell-type clustering uses adjusted mutual information. Motif sensitivity is a separate paired-sequence evaluation.
Individual claims
dart primary benchmark evidence

Original source ↗

Appendix: training/test splits, clustering and supervised evaluation; reproducibility checklist

Version: 2412.05430v1
Retrieved: 2026-09-16T21:06:29.715960+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: 4194b137ba55c9a2c269d119a9afec6ae1bb0feaf17d91433ae483c41221a56b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: discovered

11 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: discovery-benchmark-dart-eval

areas
genomics
entity level
suite
scope note
Specialist molecular or omics evaluation; protocol details require review before numerical comparison.
task
Human regulatory DNA representation evaluation
version
Not reported
benchmark research
review date: 2026-09-17; status: source_found_structured_extraction_pending; primary sources: evidence-expansion-p2-evidence-discovery-final-dart-4194b137ba55; inspected locators: arXiv v1 main benchmark procedures and results tables; searched queries: DART-Eval primary paper benchmark results; gaps: Regulatory sequence detection, TF motifs, cell-type activity and variant effects are distinct protocols. Enhancer broad-task record must link concrete tasks, not inherit every DART result. Existing DART observations require source-cell identity reuse.; claim scope: Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.
historical missing metadata
dataset release: unextracted; metric implementation: unextracted; split manifest: unextracted; version: unextracted
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
entity classification
review date: 2026-09-17; rationale: The cited profile describes a collection of evaluation tasks or protocols; retain it as the top-level benchmark suite. Its datasets and individual protocols remain separate records.; source ids: src-discovery-kundajelab-dart-eval; source locator: Pinned README: Overview; Data download; task-organized outputs; probing/fine-tuning sections; ambiguities: None recorded
run documentation
record id: discovery-benchmark-dart-eval; source ids: run-doc-dart-eval-readme-md-af2a86d6; status: official_documentation_linked; summary: The official README provides task-specific Python module commands, model identifiers, DART_WORK_DIR layout and Synapse resource IDs. These inspected sections do not document dependency installation or a complete environment lock. The Task 1 link label and target differ; resolve the exact Synapse resource before assembling a copyable recipe.; source locator: README.md lines 8–26 and 28–110 (Data, Preliminaries and Task 1 commands)
run recipes
id: dart-task1-runner-rescore-v1; protocol id: dart-eval-task1-zero-shot-v1; version: fcccbbcdbe3d5cd64a1a312d536615273320f7b4; title: DART-Eval Task1 zero-shot paired controls: score supplied predictions; purpose: rescore_predictions; summary: Score regulatory sequences against their existing paired shuffled controls without model fitting. Strict element-score greater-than-control accuracy; paired score-difference summaries and pinned SciPy1.12 one-sided Wilcoxon procedure.; inputs: Prepared protocol inputs with evaluator-owned labels and opaque IDs.; Keyed finite scalar scores; missing predictions must be explicitly permitted and remain in coverage.; outputs: Local report.json with metrics, source/version identity and coverage.; Local predictions.json and unscored.json; nothing submitted automatically.; requirements: data: Official one-hot paired input is Synapse syn64314109 version1; authenticated access is required. Anonymous downloads returned403 and official bytes/complete denominator remain unverified. The local demo is12 synthetic pairs.; weights: No weights needed for rescoring or the supplied toy/composition control. Private-model weights remain local; their access requirements depend on the model.; licence: Runner MIT. Dataset reuse terms are not established by the software licence; review Synapse access and original ENCODE terms.; software: Python3.11, pinned rewirebench0.4 environment; sequence extra for parquet and HDF5. Podman or Apptainer is optional.; hardware: CPU scoring and small controls; memory/accelerator needs for real private-model inference are model-dependent and unreported here.; instructions: runtime: python; title: Score supplied keyed predictions; code: import rewirebench prepared = rewirebench.prepare( 'dart-eval-task1-zero-shot-v1', source='demo', output='prepared-dart-task1' ) report = rewirebench.evaluate( prepared, 'predictions.json', output='results-dart-task1-rescore', model={'name': 'My private model', 'training_overlap': 'Unreported'}, ); status: source_reviewed_not_executed; source ids: runner-04-docs-sdk-md; runner-04-docs-hpc-md; runner-04-docs-dart-eval-md; runner-04-packages-rewirebench-src-rewirebench-protocols-dart-eval-py; runner-04-packages-rewirebench-src-rewirebench-resources-dart-eval-sources-json; runner-04-packages-rewirebench-src-rewirebench-resources-dart-eval-statistical-reference-json; source locator: docs/dart-eval.md and protocol implementation; linked validation receipt concerns its stated tests, not execution of this exact snippet; runtime: command_line; title: Prepare inputs and score supplied predictions; code: git clone https://github.com/rewire-bio/rewire-benchmarks.git cd rewire-benchmarks git checkout --detach fcccbbcdbe3d5cd64a1a312d536615273320f7b4 uv sync --locked --package rewirebench --extra sequence --python 3.11 uv run --package rewirebench rewirebench prepare dart-eval-task1-zero-shot-v1 \ --source demo --output prepared-dart-task1 \ --options '{}' uv run --package rewirebench rewirebench evaluate \ --prepared prepared-dart-task1 --predictions predictions.json \ --output results-dart-task1-rescore --model-name 'My private model'; status: source_reviewed_not_executed; source ids: runner-04-docs-sdk-md; runner-04-docs-hpc-md; runner-04-docs-dart-eval-md; runner-04-packages-rewirebench-src-rewirebench-protocols-dart-eval-py; runner-04-packages-rewirebench-src-rewirebench-resources-dart-eval-sources-json; runner-04-packages-rewirebench-src-rewirebench-resources-dart-eval-statistical-reference-json; source locator: docs/dart-eval.md and protocol implementation; linked validation receipt concerns its stated tests, not execution of this exact snippet; runtime: podman; title: Score prepared predictions offline; code: # Obtain/build the pinned core OCI image following the linked HPC guide. mkdir -p runs podman run --rm --userns=keep-id --network=none \ -v "$PWD/prepared-dart-task1:/prepared:ro" \ -v "$PWD/predictions:/predictions:ro" -v "$PWD/runs:/outputs:rw" \ localhost/rewirebench:core evaluate --prepared /prepared \ --predictions /predictions/predictions.json --output /outputs/dart-task1-rescore; status: source_reviewed_not_executed; source ids: runner-04-docs-sdk-md; runner-04-docs-hpc-md; runner-04-docs-dart-eval-md; runner-04-packages-rewirebench-src-rewirebench-protocols-dart-eval-py; runner-04-packages-rewirebench-src-rewirebench-resources-dart-eval-sources-json; runner-04-packages-rewirebench-src-rewirebench-resources-dart-eval-statistical-reference-json; source locator: docs/dart-eval.md and protocol implementation; linked validation receipt concerns its stated tests, not execution of this exact snippet; runtime: apptainer; title: Score prepared predictions offline; code: # Obtain/build the pinned SIF following the linked HPC guide. mkdir -p runs apptainer run --cleanenv --containall \ --bind "$PWD/prepared-dart-task1:/prepared:ro" \ --bind "$PWD/predictions:/predictions:ro" --bind "$PWD/runs:/outputs:rw" \ rewirebench-core.sif evaluate --prepared /prepared \ --predictions /predictions/predictions.json --output /outputs/dart-task1-rescore; status: source_reviewed_not_executed; source ids: runner-04-docs-sdk-md; runner-04-docs-hpc-md; runner-04-docs-dart-eval-md; runner-04-packages-rewirebench-src-rewirebench-protocols-dart-eval-py; runner-04-packages-rewirebench-src-rewirebench-resources-dart-eval-sources-json; runner-04-packages-rewirebench-src-rewirebench-resources-dart-eval-statistical-reference-json; source locator: docs/dart-eval.md and protocol implementation; linked validation receipt concerns its stated tests, not execution of this exact snippet; runtime: slurm; title: Schedule prepared scoring with site-specific resources; code: #!/bin/bash set -euo pipefail # Set these for your cluster; no performance or resource estimate is implied. : "${REWIRE_ACCOUNT:?Set your Slurm account}" : "${REWIRE_PARTITION:?Set your Slurm partition}" : "${REWIRE_CPUS:?Set the requested CPU count}" : "${REWIRE_MEMORY:?Set the requested memory}" : "${REWIRE_WALLTIME:?Set the requested time limit}" : "${REWIRE_JOB_ROOT:?Set a shared absolute directory with prepared data and SIF}" export REWIRE_JOB_ROOT sbatch --account="$REWIRE_ACCOUNT" --partition="$REWIRE_PARTITION" \ --cpus-per-task="$REWIRE_CPUS" --mem="$REWIRE_MEMORY" \ --time="$REWIRE_WALLTIME" --export=ALL <<'REWIRE_JOB' #!/bin/bash set -euo pipefail cd "$REWIRE_JOB_ROOT" mkdir -p runs apptainer run --cleanenv --containall \ --bind "$PWD/prepared-dart-task1:/prepared:ro" \ --bind "$PWD/predictions:/predictions:ro" \ --bind "$PWD/runs:/outputs:rw" \ rewirebench-core.sif evaluate \ --prepared /prepared --predictions /predictions/predictions.json \ --output /outputs/dart-task1-rescore REWIRE_JOB; status: source_reviewed_not_executed; source ids: runner-04-docs-sdk-md; runner-04-docs-hpc-md; runner-04-docs-dart-eval-md; runner-04-packages-rewirebench-src-rewirebench-protocols-dart-eval-py; runner-04-packages-rewirebench-src-rewirebench-resources-dart-eval-sources-json; runner-04-packages-rewirebench-src-rewirebench-resources-dart-eval-statistical-reference-json; source locator: docs/hpc.md: Slurm template, caller-supplied account/partition/resources; protocol guide: evaluate prepared predictions; limitations: Only zero-shot Task1 is supported, not the full DART-Eval suite; training and embedding probes are forbidden.; Official biological inputs and DNABERT-2 prediction artifacts require authenticated access. User-supplied local hashes prove integrity, not source authenticity or complete chromosome coverage.; The demo contains synthetic random DNA with shuffled controls. Its toy score checks software only and is not a biological baseline.; Missing either pair member excludes that pair from metrics; sequence and complete-pair coverage stay separate. Ties are incorrect for accuracy.; Scoring is tested on synthetic inputs against the pinned upstream evaluator and historical SciPy rules. No official biological inference or published evaluation reproduction is claimed.; These snippets are source-reviewed instructions. The linked automated receipt supports only the tests it records; it does not certify the exact snippet or published-model reproduction.; Smoke examples and partial predictions retain original denominators and cannot establish a complete suite. Prepared execution is offline; exports and submissions are explicit.; source ids: runner-04-docs-sdk-md; runner-04-docs-hpc-md; runner-04-docs-dart-eval-md; runner-04-packages-rewirebench-src-rewirebench-protocols-dart-eval-py; runner-04-packages-rewirebench-src-rewirebench-resources-dart-eval-sources-json; runner-04-packages-rewirebench-src-rewirebench-resources-dart-eval-statistical-reference-json; source locator: docs/dart-eval.md; source manifest; protocol score function; automated receipt scope; id: dart-task1-runner-control-v1; protocol id: dart-eval-task1-zero-shot-v1; version: fcccbbcdbe3d5cd64a1a312d536615273320f7b4; title: DART-Eval Task1 zero-shot paired controls: local model or software control; purpose: generate_and_evaluate; summary: Score regulatory sequences against their existing paired shuffled controls without model fitting. Strict element-score greater-than-control accuracy; paired score-difference summaries and pinned SciPy1.12 one-sided Wilcoxon procedure.; inputs: Same prepared biological inputs; no evaluation labels are exposed to adapters.; A private local adapter or the documented lightweight software control.; outputs: Local report.json with metrics, source/version identity and coverage.; Local predictions.json and unscored.json; nothing submitted automatically.; requirements: data: Official one-hot paired input is Synapse syn64314109 version1; authenticated access is required. Anonymous downloads returned403 and official bytes/complete denominator remain unverified. The local demo is12 synthetic pairs.; weights: No weights needed for rescoring or the supplied toy/composition control. Private-model weights remain local; their access requirements depend on the model.; licence: Runner MIT. Dataset reuse terms are not established by the software licence; review Synapse access and original ENCODE terms.; software: Python3.11, pinned rewirebench0.4 environment; sequence extra for parquet and HDF5. Podman or Apptainer is optional.; hardware: CPU scoring and small controls; memory/accelerator needs for real private-model inference are model-dependent and unreported here.; instructions: runtime: python; title: Run a small local software control; code: import rewirebench class ToySequenceScore: def predict(self, inputs): return {row["id"]: float(row["sequence"].count("AAA")) for row in inputs} prepared = rewirebench.prepare( 'dart-eval-task1-zero-shot-v1', source='demo', output='prepared-dart-task1' ) report = rewirebench.run( prepared, ToySequenceScore(), output='results-dart-task1-control', model={'name': 'Local software control; not a published baseline'}, ); status: source_reviewed_not_executed; source ids: runner-04-docs-sdk-md; runner-04-docs-hpc-md; runner-04-docs-dart-eval-md; runner-04-packages-rewirebench-src-rewirebench-protocols-dart-eval-py; runner-04-packages-rewirebench-src-rewirebench-resources-dart-eval-sources-json; runner-04-packages-rewirebench-src-rewirebench-resources-dart-eval-statistical-reference-json; source locator: docs/dart-eval.md and protocol implementation; linked validation receipt concerns its stated tests, not execution of this exact snippet; limitations: Only zero-shot Task1 is supported, not the full DART-Eval suite; training and embedding probes are forbidden.; Official biological inputs and DNABERT-2 prediction artifacts require authenticated access. User-supplied local hashes prove integrity, not source authenticity or complete chromosome coverage.; The demo contains synthetic random DNA with shuffled controls. Its toy score checks software only and is not a biological baseline.; Missing either pair member excludes that pair from metrics; sequence and complete-pair coverage stay separate. Ties are incorrect for accuracy.; Scoring is tested on synthetic inputs against the pinned upstream evaluator and historical SciPy rules. No official biological inference or published evaluation reproduction is claimed.; These snippets are source-reviewed instructions. The linked automated receipt supports only the tests it records; it does not certify the exact snippet or published-model reproduction.; Smoke examples and partial predictions retain original denominators and cannot establish a complete suite. Prepared execution is offline; exports and submissions are explicit.; source ids: runner-04-docs-sdk-md; runner-04-docs-hpc-md; runner-04-docs-dart-eval-md; runner-04-packages-rewirebench-src-rewirebench-protocols-dart-eval-py; runner-04-packages-rewirebench-src-rewirebench-resources-dart-eval-sources-json; runner-04-packages-rewirebench-src-rewirebench-resources-dart-eval-statistical-reference-json; source locator: docs/dart-eval.md; source manifest; protocol score function; automated receipt scope; id: dart-eval-official; protocol id: discovery-benchmark-dart-eval; version: af2a86d666c35304257c2fa7e15180e1fbcabb01; title: Run the regulatory element identification task; purpose: generate_and_evaluate; summary: Generate the task's dataset, then score a model zero-shot, probed or fine-tuned, which are the three settings this page separates.; inputs: A DNA language model supported by the repository.; outputs: Per-setting scores for the chosen task.; requirements: data: Generated by the repository's dataset generators from ENCODE inputs.; weights: A published DNA language model checkpoint.; licence: Project licence: see repository. Upstream data licences are separate and unreported here.; software: Python with the dnalm_bench package from the repository.; hardware: Not stated in the cited section. Several of these steps expect a GPU.; instructions: runtime: command_line; title: Generate the dataset; code: python -m dnalm_bench.task_1_paired_control.dataset_generators.encode_ccre --ccre_bed $DART_WORK_DIR/task_1_ccre/input_data/ENCFF420VPZ.bed --output_file $DART_WORK_DIR/task_1_ccre/processed_inputs/ENCFF420VPZ_processed.tsv; status: source_reviewed_not_executed; source ids: project-recipe-dart-eval-af2a86d6; source locator: README.md at af2a86d6, Dataset Generation, lines 39-39; runtime: command_line; title: Score zero-shot; code: python -m dnalm_bench.task_1_paired_control.zero_shot.encode_ccre.$MODEL; status: source_reviewed_not_executed; source ids: project-recipe-dart-eval-af2a86d6; source locator: README.md at af2a86d6, Zero-shot likelihood analyses, lines 56-56; runtime: command_line; title: Extract embeddings; code: python -m dnalm_bench.task_1_paired_control.supervised.encode_ccre.extract_embeddings.$MODEL; status: source_reviewed_not_executed; source ids: project-recipe-dart-eval-af2a86d6; source locator: README.md at af2a86d6, Probing models, lines 84-84; runtime: command_line; title: Train a probing model; code: python -m dnalm_bench.task_1_paired_control.supervised.encode_ccre.train_classifiers.$MODEL; status: source_reviewed_not_executed; source ids: project-recipe-dart-eval-af2a86d6; source locator: README.md at af2a86d6, Probing models, lines 90-90; runtime: command_line; title: Evaluate a probing model; code: python -m dnalm_bench.task_1_paired_control.supervised.encode_ccre.eval_probing.$MODEL; status: source_reviewed_not_executed; source ids: project-recipe-dart-eval-af2a86d6; source locator: README.md at af2a86d6, Probing models, lines 96-96; limitations: Quoted from the project's README and not executed by rewire, so the commands are evidence of what the project documents rather than a verified run.; The project may have changed since the pinned commit.; The evaluation setting is part of the result: zero-shot, probed and fine-tuned runs of one model are different entries on this page.; The commands take a model name through an environment variable, so each one runs a single model.; source ids: project-recipe-dart-eval-af2a86d6; source locator: README.md at af2a86d6
Related records

Suggest a correction