rewire.itbenchmarks
Task

regulatory sequence classification

Regulatory-sequence classification compares genomic models and tokenizers across established benchmark collections.

SourcesThe impact of tokenizer selection in genomic language models · Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44

1 evaluation · 1 result

Overview

Metrics

Mean accuracy and Matthews correlation coefficient are reported per fine-tuning task; multiclass tasks use macro averaging. Results remain task-specific across the three benchmark collections.

SourcesThe impact of tokenizer selection in genomic language models · §2.2–2.3 Benchmarks and Metrics
Evaluation procedure diagram
How it worksComputational evaluation flow
Computational evaluation flow1. Allowed inputs: Tokenized DNA sequence.. Then: 2. Datasets: Nucleotide Transformer tasks, Genomic Benchmarks and GUE.. Then: 3. Metrics: Mean accuracy and Matthews correlation coefficient are reported per fine-tuning task; multiclass tasks use macro averaging. Results remain task-specific across the three benchmark collections.Computational evaluation flow1. Allowed inputs: Tokenized DNA sequence.. Then: 2. Datasets: Nucleotide Transformer tasks, Genomic Benchmarks and GUE.. Then: 3. Metrics: Mean accuracy and Matthews correlation coefficient are reported per fine-tuning task; multiclass tasks use macro averaging. Results remain task-specific across the three benchmark collections.Computational evaluation flow1. Allowed inputs: Tokenized DNA sequence.. Then: 2. Datasets: Nucleotide Transformer tasks, Genomic Benchmarks and GUE.. Then: 3. Metrics: Mean accuracy and Matthews correlation coefficient are reported per fine-tuning task; multiclass tasks use macro averaging. Results remain task-specific across the three benchmark collections.

Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.

SourcesThe impact of tokenizer selection in genomic language models · Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44; Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44; §2.2–2.3 Benchmarks and Metrics

limited source coverage · Automated source review, 2026-09-16. All specifications and missing details

Results

Results are available, but no reviewed comparison panel is linked in this release.

All evaluations

1 evaluation · 1 result. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: Caduceus (character tokens)Task: regulatory sequence classification
Dataset: genomic benchmark categories
0.778 MCC
unitless · unknown

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

Caduceus (character tokens): regulatory sequence classification

task-category MCC across benchmark datasets

Aggregation: Not reported

The impact of tokenizer selection in genomic language models · Table 2, Regulatory row, Caduceus (char) MCC column

Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.

Methods and evaluation design

Procedure, tasks and evaluated configurations

How it works

Evaluation methodology

Nucleotide Transformer tasks, Genomic Benchmarks and GUE. Mean accuracy and Matthews correlation coefficient are reported per fine-tuning task; multiclass tasks use macro averaging. Results remain task-specific across the three benchmark collections. Attention-based and state-space genomic language models with different tokenizer choices.

SourcesThe impact of tokenizer selection in genomic language models · Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44; §2.2–2.3 Benchmarks and Metrics; Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44

Recorded evaluations

Each evaluation records what was tested and under which conditions.

Run instructions

No runnable recipe has been reviewed for this task. Dataset access, model requirements, licences and compute requirements must be checked against its sources before execution.

A task describes a biological question. Choose a linked protocol to obtain concrete split and scoring instructions.

Strengths, limitations and unresolved questions

Strengths and limitations

Strengths and considerations

No source-reviewed explanatory claims are recorded here yet.

Limitations and conditions

  • Different hyperparameter-tuning procedures can affect the comparison. The broad catalogue label does not uniquely identify one dataset or executable protocol.
    SourcesThe impact of tokenizer selection in genomic language models · Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44
Profile review details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Stable record: reported-task-cd127e56fb1f04

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetsNucleotide Transformer tasks, Genomic Benchmarks and GUE.
SourcesThe impact of tokenizer selection in genomic language models · Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44
SplitsThe experiment reuses three published fine-tuning benchmark collections and repeats every task at least ten times. The broad catalogue label spans multiple task partitions; no single regulatory-element split is defined by that label.
SourcesThe impact of tokenizer selection in genomic language models · §2.2 Benchmarks
MetricsMean accuracy and Matthews correlation coefficient are reported per fine-tuning task; multiclass tasks use macro averaging. Results remain task-specific across the three benchmark collections.
SourcesThe impact of tokenizer selection in genomic language models · §2.2–2.3 Benchmarks and Metrics
BaselinesAttention-based and state-space genomic language models with different tokenizer choices.
SourcesThe impact of tokenizer selection in genomic language models · Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44
Leakage controlsThe benchmark reuses GB, NTTv2 and GUE task datasets and compares models pretrained on different corpora, including the human reference genome. Its main Methods do not report a unified overlap-removal audit between all task examples and every model’s pretraining corpus. · Not reported in inspected sources
SourcesThe impact of tokenizer selection in genomic language models · §2 Materials and methods, model/pretraining and task-dataset descriptions; Table 1
UncertaintyEach fine-tuning task is replicated at least ten times; hyperparameter search differs by model family.
SourcesThe impact of tokenizer selection in genomic language models · Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44
Entity typePaper-specific computational evaluation protocol.
SourcesThe impact of tokenizer selection in genomic language models · Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44
OrganismsDataset-specific organisms in Nucleotide Transformer, Genomic Benchmarks and GUE.
SourcesThe impact of tokenizer selection in genomic language models · Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44
AssaysRegulatory and other genomic classification labels.
SourcesThe impact of tokenizer selection in genomic language models · Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44
Allowed inputsTokenized DNA sequence.
SourcesThe impact of tokenizer selection in genomic language models · Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44
AdaptationRepeated task fine-tuning compares tokenizer/model configurations.
SourcesThe impact of tokenizer selection in genomic language models · Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Primary-paper discovery and source inspection. Source-checked results are not independently reproduced experiments.

Paper or primary resourceVersionReference
The impact of tokenizer selection in genomic language modelsjournal full text in PMCRead source
DOI: 10.1093/bioinformatics/btaf456
Historical gaps recorded on 2026-09-17

The catalogue now holds 1 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.

  • Models differ in architecture, size and pretraining as well as tokenizer, so Table 2 alone does not isolate tokenizer causality.
  • Table 3 offers paired tokenizer comparisons; do not relabel whole-family comparisons as controlled tokenizer ablations.
  • No per-cell uncertainty in Table 2.
Search and extraction details

primary comparison table screened

Searches

  • "PMC12453675"

Evidence locations

  • Tables 1–3
  • Benchmark experimental setup

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

17 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.
Individual claims
The impact of tokenizer selection in genomic language models

Original source ↗

Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44; Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44; §2.2–2.3 Benchmarks and Metrics

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558210+00:00

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 0a01c36fdd63f3f6db509777e61c3f87e8a298c810f8aef7974915aaa0655342

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps
  • Allowed inputs: Tokenized DNA sequence.
  • Datasets: Nucleotide Transformer tasks, Genomic Benchmarks and GUE.
  • Metrics: Mean accuracy and Matthews correlation coefficient are reported per fine-tuning task; multiclass tasks use macro averaging. Results remain task-specific across the three benchmark collections.
Individual claims
The impact of tokenizer selection in genomic language models

Original source ↗

Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44; Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44; §2.2–2.3 Benchmarks and Metrics

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558210+00:00

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 0a01c36fdd63f3f6db509777e61c3f87e8a298c810f8aef7974915aaa0655342

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title
Computational evaluation flow
Individual claims
The impact of tokenizer selection in genomic language models

Original source ↗

Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44; Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44; §2.2–2.3 Benchmarks and Metrics

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558210+00:00

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.diagram.title

Source artifact SHA-256: 0a01c36fdd63f3f6db509777e61c3f87e8a298c810f8aef7974915aaa0655342

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Datasets
Nucleotide Transformer tasks, Genomic Benchmarks and GUE.
Individual claims
The impact of tokenizer selection in genomic language models

Original source ↗

Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558210+00:00

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: 0a01c36fdd63f3f6db509777e61c3f87e8a298c810f8aef7974915aaa0655342

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Splits
The experiment reuses three published fine-tuning benchmark collections and repeats every task at least ten times. The broad catalogue label spans multiple task partitions; no single regulatory-element split is defined by that label.
Individual claims
The impact of tokenizer selection in genomic language models

Original source ↗

§2.2 Benchmarks

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558210+00:00

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: 0a01c36fdd63f3f6db509777e61c3f87e8a298c810f8aef7974915aaa0655342

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Adaptation
Repeated task fine-tuning compares tokenizer/model configurations.
Individual claims
The impact of tokenizer selection in genomic language models

Original source ↗

Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558210+00:00

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: 0a01c36fdd63f3f6db509777e61c3f87e8a298c810f8aef7974915aaa0655342

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Metrics
Mean accuracy and Matthews correlation coefficient are reported per fine-tuning task; multiclass tasks use macro averaging. Results remain task-specific across the three benchmark collections.
Individual claims
The impact of tokenizer selection in genomic language models

Original source ↗

§2.2–2.3 Benchmarks and Metrics

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558210+00:00

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: 0a01c36fdd63f3f6db509777e61c3f87e8a298c810f8aef7974915aaa0655342

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Baselines
Attention-based and state-space genomic language models with different tokenizer choices.
Individual claims
The impact of tokenizer selection in genomic language models

Original source ↗

Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558210+00:00

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.facts.3.value

Source artifact SHA-256: 0a01c36fdd63f3f6db509777e61c3f87e8a298c810f8aef7974915aaa0655342

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Leakage controls
The benchmark reuses GB, NTTv2 and GUE task datasets and compares models pretrained on different corpora, including the human reference genome. Its main Methods do not report a unified overlap-removal audit between all task examples and every model’s pretraining corpus.
Individual claims
The impact of tokenizer selection in genomic language models

Original source ↗

§2 Materials and methods, model/pretraining and task-dataset descriptions; Table 1

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558210+00:00

unreported

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: 0a01c36fdd63f3f6db509777e61c3f87e8a298c810f8aef7974915aaa0655342

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Uncertainty
Each fine-tuning task is replicated at least ten times; hyperparameter search differs by model family.
Individual claims
The impact of tokenizer selection in genomic language models

Original source ↗

Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558210+00:00

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.facts.5.value

Source artifact SHA-256: 0a01c36fdd63f3f6db509777e61c3f87e8a298c810f8aef7974915aaa0655342

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: needs review

2 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: reported-task-cd127e56fb1f04

areas
dna-genomes
tasks
regulatory sequence classification
entity level
task
version
Not reported
task
regulatory sequence classification
scope note
Paper-specific evaluation task; protocol completeness requires further extraction.
benchmark research
review date: 2026-09-17; status: primary_comparison_table_screened; primary sources: expansion-p3-genomic-tokenizer-selection-2025; inspected locators: Tables 1–3; Benchmark experimental setup; searched queries: "PMC12453675"; gaps: Models differ in architecture, size and pretraining as well as tokenizer, so Table 2 alone does not isolate tokenizer causality.; Table 3 offers paired tokenizer comparisons; do not relabel whole-family comparisons as controlled tokenizer ablations.; No per-cell uncertainty in Table 2.; claim scope: Primary-paper discovery and source inspection. Source-checked results are not independently reproduced experiments.
historical missing metadata
protocol version: not_reported_in_legacy_extract
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
legacy kinds
benchmark
entity classification
review date: 2026-09-17; rationale: This source-scoped record identifies the biological prediction task and holds its paper context. Preserve the existing task identity; exact split, model adaptation and scoring remain in linked evaluations or separate protocol records.; source ids: genomic-tokenizer-selection-2025; source locator: Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44; ambiguities: A paper- or suite-specific task may constrain some inputs or metrics; that alone does not make it interchangeable with a complete versioned protocol. No protocol equivalence is inferred.; Some legacy profile Entity type facts use the generic phrase computational evaluation protocol. That boilerplate is not sufficient to establish a single fixed protocol identity or to merge this task with another protocol record.
Related records

Suggest a correction