0.778 MCC
Caduceus (character tokens) · MCC · genomic benchmark categories
- Tested configuration
- Caduceus (character tokens)
- Task
- regulatory sequence classification
- Dataset
- genomic benchmark categories
- Procedure
- task-category MCC across benchmark datasets
- Evaluation
- Caduceus (character tokens): regulatory sequence classification
- Coverage
- scored: unreported; eligible: unreported
- Uncertainty
- Not reported
- Evidence
- Independent external evaluation · source checkedThe impact of tokenizer selection in genomic language models · Table 2, Regulatory row, Caduceus (char) MCC column
A source-checked result verifies the numerical transcription, not every model or protocol detail. Evaluation metadata: needs review. Source checked does not mean independently reproduced.
Reproduction
- Split
- paper benchmark summary
- Adaptation
- Not reported
- Scoring implementation
- Not reported
No execution recipe has been verified for this exact configuration and evaluation. A benchmark's general instructions may use different inputs, splits or model settings.
Reproducing this published result requires matching its model configuration, data, split and scorer. Source checking or a successful smoke test does not establish score reproduction.
Evidence
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Evidence table
Inspect claims, sources and review details
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
1 evidence row matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| attributes.printed_value 0.778 Individual claims | The impact of tokenizer selection in genomic language models Table 2, Regulatory row, Caduceus (char) MCC column Version: journal full text in PMC | source checked independent ai table review · 2026-09-16T10:38:57.558210+00:00 independent paper Audit detailsMatched Regulatory row with Caduceus (char) column, 0.778. Caption establishes these as MCC summaries by category; model-size row identifies 3.9M parameters. This is an aggregated category result, not a single unspecified split. Field: Claim: claim-b2-genomic-tokenizer-selection-2025 Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record Extraction artifact SHA-256: |
Sources and history
View linked audit checks and correction history
Release 2026-09-29-06401fd5b220 · Record review: source checked
1 source records and release history
- The impact of tokenizer selection in genomic language models · Original source · journal full text in PMC
Technical metadata and extraction receipts
Stable ID: b2-genomic-tokenizer-selection-2025
- areas
- dna-genomes
- tasks
- regulatory sequence classification
- printed value
- 0.778
- numeric value
- 0.778
- metric
- MCC
- metric direction
- unknown
- unit
- unitless
- uncertainty
- Not reported
- source locator
- Table 2, Regulatory row, Caduceus (char) MCC column
- review
- method: independent_ai_table_review; reviewer: Codex secondary table review; reviewed at: 2026-09-16T10:38:57.558210+00:00; notes: Matched Regulatory row with Caduceus (char) column, 0.778. Caption establishes these as MCC summaries by category; model-size row identifies 3.9M parameters. This is an aggregated category result, not a single unspecified split.; evidence: Original PMC XML table headers, row groups and caption inspected; printed value 0.778.; artifact sha256: 0a01c36fdd63f3f6db509777e61c3f87e8a298c810f8aef7974915aaa0655342; retrieval url: https://www.ebi.ac.uk/europepmc/webservices/rest/PMC12453675/fullTextXML
- legacy id
- b2-genomic-tokenizer-selection-2025
- legacy row
- id: b2-genomic-tokenizer-selection-2025; paper id: genomic-tokenizer-selection-2025; domain id: dna-genomes; task: regulatory sequence classification; model: Caduceus (character tokens); model version: 3.9M parameter variant; dataset: genomic benchmark categories; dataset version: Not reported; split: paper benchmark summary; metric: MCC; value: 0.778; unit: unitless; uncertainty: Not reported; protocol: task-category MCC across benchmark datasets; source locator: Table 2, Regulatory row, Caduceus (char) MCC column; source url: https://pmc.ncbi.nlm.nih.gov/articles/PMC12453675/; evaluation origin: independent_paper; reviewed utc: 2026-09-15T23:33:26Z
- missing metadata
- dataset version: not_reported_in_legacy_extract; uncertainty: not_reported_in_legacy_extract