Datasets
Genomic Benchmarks includes human enhancer datasets; Nucleotide Transformer tasks include enhancer/non-enhancer and enhancer-strength labels.
Enhancer classification is evaluated within two published genomic sequence benchmark collections.
Genomic Benchmarks includes human enhancer datasets; Nucleotide Transformer tasks include enhancer/non-enhancer and enhancer-strength labels.
Classification accuracy is described for the benchmark comparison.
Genomic sequence.
Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.
limited source coverage · Automated source review, 2026-09-16. All specifications and missing details
Results are available, but no reviewed comparison panel is linked in this release.
2 evaluations · 2 results. Different protocols are not a single leaderboard.
Applied filters: All linked evaluations
| Tested configuration | Protocol and dataset | Finding | Evidence and details |
|---|---|---|---|
| Configuration: ENBED | Task: Enhancer classification Dataset: Genomic Benchmarks Mouse Enhancers | 90.3 Accuracy % · unknown Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceENBED: Enhancer classification Reported Genomic Benchmarks classification accuracy. Aggregation: Not reported Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Table 2, Mouse Enhancers row, ENBED column |
| Configuration: ENBED (GRCh38) | Task: Enhancer classification Dataset: Genomic Benchmarks Mouse Enhancers | 81.1 Accuracy % · unknown Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceENBED (GRCh38): Enhancer classification ENBED trained on GRCh38; reported Genomic Benchmarks classification accuracy. Aggregation: Not reported Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Table 2, Mouse Enhancers row, ENBED (GRCh38) column |
Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.
Genomic Benchmarks includes human enhancer datasets; Nucleotide Transformer tasks include enhancer/non-enhancer and enhancer-strength labels. For the Nucleotide Transformer enhancer tasks, Supplementary Table 1 lists 14,968 training sequences and 400 test sequences of length 200. This set combines the original strong/weak/non-enhancer collection with 6,000 synthetic enhancers and 6,000 synthetic non-enhancers. Those counts do not describe the separate Genomic Benchmarks enhancer datasets. Classification accuracy is described for the benchmark comparison. Enformer, DNABERT-2, Nucleotide Transformer and HyenaDNA in the enhancer-containing NT task comparison. The enhancer-specific methods and supplement do not specify a chromosome holdout, homology threshold or synthetic-parent separation across these splits. The clade split and 95% similarity threshold described later in the supplement apply to mutation generation, not this enhancer evaluation.
Each evaluation records what was tested and under which conditions.
No runnable recipe has been reviewed for this task. Dataset access, model requirements, licences and compute requirements must be checked against its sources before execution.
A task describes a biological question. Choose a linked protocol to obtain concrete split and scoring instructions.
No source-reviewed explanatory claims are recorded here yet.
Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.
Stable record: reported-task-132da895d4c381Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Datasets | Genomic Benchmarks includes human enhancer datasets; Nucleotide Transformer tasks include enhancer/non-enhancer and enhancer-strength labels.SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions |
| Splits | For the Nucleotide Transformer enhancer tasks, Supplementary Table 1 lists 14,968 training sequences and 400 test sequences of length 200. This set combines the original strong/weak/non-enhancer collection with 6,000 synthetic enhancers and 6,000 synthetic non-enhancers. Those counts do not describe the separate Genomic Benchmarks enhancer datasets.Sources (2)Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision; enbed-2024__vbae117_supplementary_data.pdf · Supplement: Nucleotide Transformer benchmark dataset description and Table 1 |
| Metrics | Classification accuracy is described for the benchmark comparison.SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions |
| Baselines | Enformer, DNABERT-2, Nucleotide Transformer and HyenaDNA in the enhancer-containing NT task comparison.SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions |
| Leakage controls | The enhancer-specific methods and supplement do not specify a chromosome holdout, homology threshold or synthetic-parent separation across these splits. The clade split and 95% similarity threshold described later in the supplement apply to mutation generation, not this enhancer evaluation. · Not reported in inspected sourcesSources (2)Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision; enbed-2024__vbae117_supplementary_data.pdf · Supplement: Nucleotide Transformer benchmark description and Table 1; separate Mutation generation section |
| Uncertainty | The cited text-accessible evaluation sections give no confidence-interval, resampling or repeat-run error-bar specification. Image-only tables and uninspected supplements are outside this absence claim. · Not reported in inspected sourcesSourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions |
| Entity type | Paper-specific computational evaluation protocol.SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions |
| Organisms | Human enhancer datasets within Genomic Benchmarks and Nucleotide Transformer tasks.SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions |
| Assays | Enhancer identity/strength annotations.SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions |
| Allowed inputs | Genomic sequence.SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions |
| Adaptation | Task-level enhancer classification; exact task configuration must be selected from the evaluated benchmark collection.SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions |
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Last literature check: 2026-09-17. Dated primary-source discovery and protocol/table screening. Source checking does not mean experimental reproduction. Only separately extracted and independently reviewed numeric batches are publishable.
| Paper or primary resource | Version | Reference |
|---|---|---|
| Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision | version of record | Read source |
The catalogue now holds 2 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.
primary comparison tables located
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
24 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol. Individual claims | Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions; Supplement: Nucleotide Transformer benchmark dataset description and Table 1 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: version of record | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram caption Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol. Individual claims | enbed-2024__vbae117_supplementary_data.pdf Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions; Supplement: Nucleotide Transformer benchmark dataset description and Table 1 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: Retrieved 2026-09-16; sha256:a169bbebba2298e9c98c33c28053c1a0b42d0c3e88e5a8795934725c4408cc32 | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record Archive member: vbae117_supplementary_data.pdf |
Diagram steps
| Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions; Supplement: Nucleotide Transformer benchmark dataset description and Table 1 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: version of record | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| enbed-2024__vbae117_supplementary_data.pdf Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions; Supplement: Nucleotide Transformer benchmark dataset description and Table 1 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: Retrieved 2026-09-16; sha256:a169bbebba2298e9c98c33c28053c1a0b42d0c3e88e5a8795934725c4408cc32 | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record Archive member: vbae117_supplementary_data.pdf |
| Diagram title Computational evaluation flow Individual claims | Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions; Supplement: Nucleotide Transformer benchmark dataset description and Table 1 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: version of record | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Computational evaluation flow Individual claims | enbed-2024__vbae117_supplementary_data.pdf Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions; Supplement: Nucleotide Transformer benchmark dataset description and Table 1 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: Retrieved 2026-09-16; sha256:a169bbebba2298e9c98c33c28053c1a0b42d0c3e88e5a8795934725c4408cc32 | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record Archive member: vbae117_supplementary_data.pdf |
| Datasets Genomic Benchmarks includes human enhancer datasets; Nucleotide Transformer tasks include enhancer/non-enhancer and enhancer-strength labels. Individual claims | Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions Version: version of record | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Splits For the Nucleotide Transformer enhancer tasks, Supplementary Table 1 lists 14,968 training sequences and 400 test sequences of length 200. This set combines the original strong/weak/non-enhancer collection with 6,000 synthetic enhancers and 6,000 synthetic non-enhancers. Those counts do not describe the separate Genomic Benchmarks enhancer datasets. Individual claims | Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision Supplement: Nucleotide Transformer benchmark dataset description and Table 1 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: version of record | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Splits For the Nucleotide Transformer enhancer tasks, Supplementary Table 1 lists 14,968 training sequences and 400 test sequences of length 200. This set combines the original strong/weak/non-enhancer collection with 6,000 synthetic enhancers and 6,000 synthetic non-enhancers. Those counts do not describe the separate Genomic Benchmarks enhancer datasets. Individual claims | enbed-2024__vbae117_supplementary_data.pdf Supplement: Nucleotide Transformer benchmark dataset description and Table 1 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: Retrieved 2026-09-16; sha256:a169bbebba2298e9c98c33c28053c1a0b42d0c3e88e5a8795934725c4408cc32 | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record Archive member: vbae117_supplementary_data.pdf |
| Adaptation Task-level enhancer classification; exact task configuration must be selected from the evaluated benchmark collection. Individual claims | Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions Version: version of record | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
View linked audit checks and correction history
Release 2026-09-29-06401fd5b220 · Record review: needs review
Stable ID: reported-task-132da895d4c381