Datasets
Nucleotide Transformer tasks, Genomic Benchmarks and GUE.
Regulatory-sequence classification compares genomic models and tokenizers across established benchmark collections.
Nucleotide Transformer tasks, Genomic Benchmarks and GUE.
Mean accuracy and Matthews correlation coefficient are reported per fine-tuning task; multiclass tasks use macro averaging. Results remain task-specific across the three benchmark collections.
Tokenized DNA sequence.
Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.
limited source coverage · Automated source review, 2026-09-16. All specifications and missing details
Results are available, but no reviewed comparison panel is linked in this release.
1 evaluation · 1 result. Different protocols are not a single leaderboard.
Applied filters: All linked evaluations
| Tested configuration | Protocol and dataset | Finding | Evidence and details |
|---|---|---|---|
| Configuration: Caduceus (character tokens) | Task: regulatory sequence classification Dataset: genomic benchmark categories | 0.778 MCC unitless · unknown Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceCaduceus (character tokens): regulatory sequence classification task-category MCC across benchmark datasets Aggregation: Not reported The impact of tokenizer selection in genomic language models · Table 2, Regulatory row, Caduceus (char) MCC column |
Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.
Nucleotide Transformer tasks, Genomic Benchmarks and GUE. Mean accuracy and Matthews correlation coefficient are reported per fine-tuning task; multiclass tasks use macro averaging. Results remain task-specific across the three benchmark collections. Attention-based and state-space genomic language models with different tokenizer choices.
Each evaluation records what was tested and under which conditions.
No runnable recipe has been reviewed for this task. Dataset access, model requirements, licences and compute requirements must be checked against its sources before execution.
A task describes a biological question. Choose a linked protocol to obtain concrete split and scoring instructions.
No source-reviewed explanatory claims are recorded here yet.
Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.
Stable record: reported-task-cd127e56fb1f04Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Datasets | Nucleotide Transformer tasks, Genomic Benchmarks and GUE.SourcesThe impact of tokenizer selection in genomic language models · Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44 |
| Splits | The experiment reuses three published fine-tuning benchmark collections and repeats every task at least ten times. The broad catalogue label spans multiple task partitions; no single regulatory-element split is defined by that label.SourcesThe impact of tokenizer selection in genomic language models · §2.2 Benchmarks |
| Metrics | Mean accuracy and Matthews correlation coefficient are reported per fine-tuning task; multiclass tasks use macro averaging. Results remain task-specific across the three benchmark collections.SourcesThe impact of tokenizer selection in genomic language models · §2.2–2.3 Benchmarks and Metrics |
| Baselines | Attention-based and state-space genomic language models with different tokenizer choices.SourcesThe impact of tokenizer selection in genomic language models · Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44 |
| Leakage controls | The benchmark reuses GB, NTTv2 and GUE task datasets and compares models pretrained on different corpora, including the human reference genome. Its main Methods do not report a unified overlap-removal audit between all task examples and every model’s pretraining corpus. · Not reported in inspected sourcesSourcesThe impact of tokenizer selection in genomic language models · §2 Materials and methods, model/pretraining and task-dataset descriptions; Table 1 |
| Uncertainty | Each fine-tuning task is replicated at least ten times; hyperparameter search differs by model family.SourcesThe impact of tokenizer selection in genomic language models · Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44 |
| Entity type | Paper-specific computational evaluation protocol.SourcesThe impact of tokenizer selection in genomic language models · Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44 |
| Organisms | Dataset-specific organisms in Nucleotide Transformer, Genomic Benchmarks and GUE.SourcesThe impact of tokenizer selection in genomic language models · Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44 |
| Assays | Regulatory and other genomic classification labels.SourcesThe impact of tokenizer selection in genomic language models · Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44 |
| Allowed inputs | Tokenized DNA sequence.SourcesThe impact of tokenizer selection in genomic language models · Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44 |
| Adaptation | Repeated task fine-tuning compares tokenizer/model configurations.SourcesThe impact of tokenizer selection in genomic language models · Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44 |
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Last literature check: 2026-09-17. Primary-paper discovery and source inspection. Source-checked results are not independently reproduced experiments.
| Paper or primary resource | Version | Reference |
|---|---|---|
| The impact of tokenizer selection in genomic language models | journal full text in PMC | Read source DOI: 10.1093/bioinformatics/btaf456 |
The catalogue now holds 1 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.
primary comparison table screened
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
17 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol. Individual claims | The impact of tokenizer selection in genomic language models Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44; Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44; §2.2–2.3 Benchmarks and Metrics Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| The impact of tokenizer selection in genomic language models Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44; Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44; §2.2–2.3 Benchmarks and Metrics Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Computational evaluation flow Individual claims | The impact of tokenizer selection in genomic language models Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44; Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44; §2.2–2.3 Benchmarks and Metrics Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Datasets Nucleotide Transformer tasks, Genomic Benchmarks and GUE. Individual claims | The impact of tokenizer selection in genomic language models Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44 Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Splits The experiment reuses three published fine-tuning benchmark collections and repeats every task at least ten times. The broad catalogue label spans multiple task partitions; no single regulatory-element split is defined by that label. Individual claims | The impact of tokenizer selection in genomic language models §2.2 Benchmarks Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Adaptation Repeated task fine-tuning compares tokenizer/model configurations. Individual claims | The impact of tokenizer selection in genomic language models Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44 Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Metrics Mean accuracy and Matthews correlation coefficient are reported per fine-tuning task; multiclass tasks use macro averaging. Results remain task-specific across the three benchmark collections. Individual claims | The impact of tokenizer selection in genomic language models §2.2–2.3 Benchmarks and Metrics Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Baselines Attention-based and state-space genomic language models with different tokenizer choices. Individual claims | The impact of tokenizer selection in genomic language models Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44 Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Leakage controls The benchmark reuses GB, NTTv2 and GUE task datasets and compares models pretrained on different corpora, including the human reference genome. Its main Methods do not report a unified overlap-removal audit between all task examples and every model’s pretraining corpus. Individual claims | The impact of tokenizer selection in genomic language models §2 Materials and methods, model/pretraining and task-dataset descriptions; Table 1 Version: journal full text in PMC | unreported automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Uncertainty Each fine-tuning task is replicated at least ten times; hyperparameter search differs by model family. Individual claims | The impact of tokenizer selection in genomic language models Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44 Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
View linked audit checks and correction history
Release 2026-09-29-06401fd5b220 · Record review: needs review
Stable ID: reported-task-cd127e56fb1f04