Datasets
The benchmark archive contains separate sequence-classification datasets; model pretraining data are distributed separately.
GUE evaluates genome understanding across multiple datasets, task types and species.
The benchmark archive contains separate sequence-classification datasets; model pretraining data are distributed separately.
GUE uses Matthews correlation coefficient for the regulatory classification tasks and F1 for COVID variant classification; the exact task metric is tabulated in the dataset appendix.
DNA sequences from the benchmark archive, separate from model-pretraining data.
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Source reviewed · Automated source review, 2026-09-16. All specifications and missing details
Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.
mcc (percent) · Higher values are better.
GUE CORE-PROMOTER-DETECTION-ALL: Core promoter detection, dataset all · GUE Core promoter detection, all (GUE split)
Evidence origin: Author-reported evaluation.
DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 12, row(Core promoter detection)Every method GUE reports on Core promoter detection, dataset all, scored with MCC on GUE Core promoter detection, all.
Automated source review: 2026-09-18. Numerical source review does not establish independent reproduction.
Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.
Showing 10 of 10 matching rows.
GUE combines genome sequence classification tasks with supplied partitions and task-specific metrics. Models are fine-tuned on labelled training examples, selected with validation data and scored on test data. Random partitions occur in the suite, so a high score does not automatically demonstrate transfer to unrelated genomes.
Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.
These source-backed links do not make different protocols or scores interchangeable.
Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.
No concrete protocols are explicitly linked to this suite. Protocol identification and baseline selection are outstanding.
Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums
Coverage is derived from release 2026-09-29-06401fd5b220. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.
Fine-tune and score a model across the GUE datasets, using the authors' own evaluation script.
Generate predictions and evaluate them. This recipe does not establish reproduction of a particular published score.
Source reviewed; these instructions have not been executed by rewire.
import torch
from transformers import AutoTokenizer, AutoModel
tokenizer = AutoTokenizer.from_pretrained("zhihan1996/DNABERT-2-117M", trust_remote_code=True)
model = AutoModel.from_pretrained("zhihan1996/DNABERT-2-117M", trust_remote_code=True)DNABERT-2 and GUE: repository README · README.md at f25bed9e, 4. Quick Start, lines 84-88Source reviewed; these instructions have not been executed by rewire.
export DATA_PATH=/path/to/GUE #(e.g., /home/user)
cd finetune
# Evaluate DNABERT-2 on GUE
sh scripts/run_dnabert2.sh DATA_PATH
# Evaluate DNABERT (e.g., DNABERT with 3-mer) on GUE
# 3 for 3-mer, 4 for 4-mer, 5 for 5-mer, 6 for 6-mer
sh scripts/run_dnabert1.sh DATA_PATH 3
# Evaluate Nucleotide Transformers on GUE
# 0 for 500m-1000g, 1 for 500m-human-ref, 2 for 2.5b-1000g, 3 for 2.5b-multi-species
sh scripts/run_nt.sh DATA_PATH 0DNABERT-2 and GUE: repository README · README.md at f25bed9e, 6.1 Evaluate models on GUE, lines 138-151Run your model locally and return predictions keyed by the input IDs. The evaluator supplies biological inputs without test labels and owns scoring. This interface is not a sandbox for model code.
Pass your existing prediction function into this adapter. Its output direction must match the selected protocol.
class MyModelAdapter:
def __init__(self, score):
self.score = score
def predict(self, inputs):
return {row["id"]: float(self.score(row)) for row in inputs}
# adapter = MyModelAdapter(your_prediction_function)
# report = rewirebench.run(prepared, adapter, output="runs/my-model")Alternatively, generate a keyed prediction file in your existing model environment and use the score-only recipe. Your model code and weights do not need to be shared.
DNABERT-2 and GUE: repository README · README.md at f25bed9eContribute a result for review. The library can submit an exported evaluation for private review when intake is open. Check the contribution page for access and sign-in.
The DNABERT-2 repository supplies GUE fine-tuning scripts and data links. Its documented configuration uses four GPUs and a global batch size of 32; other hardware requires an explicit batch configuration. The shell example declares DATA_PATH but passes the literal token DATA_PATH, so inspect the script argument handling before copying it.
A maintained rewire runner has not been verified for this benchmark. Check data access, weights, licences, dependencies and hardware in the linked official documentation; requirements have not been fully extracted.
MAGICS-LAB/DNABERT_2 / README.md · README.md lines 129–150 (Evaluate models on GUE)Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.
Stable record: discovery-benchmark-gueExplanatory profile: source reviewed · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Datasets | The benchmark archive contains separate sequence-classification datasets; model pretraining data are distributed separately.SourcesMAGICS-LAB/DNABERT_2 official source · Pinned README: GUE section; data download and evaluation scripts |
| Splits | Separate train/validation/test files and counts are defined per dataset. The yeast epigenetic tasks use an 8:1:1 random split; source-derived regulatory and viral tasks retain their own documented partitions. There is no universal chromosome holdout across GUE.Sourcesgue primary benchmark evidence · Appendix C: task construction; Table 9 |
| Metrics | GUE uses Matthews correlation coefficient for the regulatory classification tasks and F1 for COVID variant classification; the exact task metric is tabulated in the dataset appendix.Sourcesgue primary benchmark evidence · Section 5; Appendix C and Table 9: GUE task datasets |
| Baselines | DNABERT-2 and other genomic representation models are compared in the associated benchmark.SourcesMAGICS-LAB/DNABERT_2 official source · Pinned README: GUE section; data download and evaluation scripts |
| Leakage controls | Datasets have explicit train/validation/test partitions, but some use random splits, including yeast epigenetic marks at 8:1:1. The paper’s masked-token leakage discussion concerns tokenization and must not be confused with proof of train/test genome independence.Sourcesgue primary benchmark evidence · Section 5; Appendix C and Table 9: GUE task datasets |
| Uncertainty | Models are fine-tuned with three different random seeds and the mean test result is reported. The stated protocol does not define a uniform confidence interval for every dataset.Sourcesgue primary benchmark evidence · Section 5; Appendix C and Table 9: GUE task datasets |
| Entity type | Genome Understanding Evaluation benchmark suite.SourcesMAGICS-LAB/DNABERT_2 official source · Pinned README: GUE section; data download and evaluation scripts |
| Organisms | Multiple species; the README describes four-species coverage.SourcesMAGICS-LAB/DNABERT_2 official source · Pinned README: GUE section; data download and evaluation scripts |
| Assays | Task-specific genomic classification labels.SourcesMAGICS-LAB/DNABERT_2 official source · Pinned README: GUE section; data download and evaluation scripts |
| Allowed inputs | DNA sequences from the benchmark archive, separate from model-pretraining data.SourcesMAGICS-LAB/DNABERT_2 official source · Pinned README: GUE section; data download and evaluation scripts |
| Adaptation | Supervised fine-tuning; scripts include model-specific training examples.SourcesMAGICS-LAB/DNABERT_2 official source · Pinned README: GUE section; data download and evaluation scripts |
Applicability is distinct from a completed evaluation.
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Last literature check: 2026-09-17. Dated primary-source discovery and protocol/table screening. Source checking does not mean experimental reproduction. Only separately extracted and independently reviewed numeric batches are publishable.
| Paper or primary resource | Version | Reference |
|---|---|---|
| DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes | Primary full-text snapshot retrieved 2026-09-17; exact bytes pinned by SHA-256 | Read source |
The catalogue now holds 280 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.
primary protocol reviewed
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
28 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions. Individual claims | gue primary benchmark evidence Pinned README: GUE section; data download and evaluation scripts; Appendix C: task construction; Table 9; Section 5; Appendix C and Table 9: GUE task datasets Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 2306.15006 retrieved PDF | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram caption Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions. Individual claims | MAGICS-LAB/DNABERT_2 official source Pinned README: GUE section; data download and evaluation scripts; Appendix C: task construction; Table 9; Section 5; Appendix C and Table 9: GUE task datasets Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: f25bed9ee20db966dff39e5c1571249d04e36404 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| gue primary benchmark evidence Pinned README: GUE section; data download and evaluation scripts; Appendix C: task construction; Table 9; Section 5; Appendix C and Table 9: GUE task datasets Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 2306.15006 retrieved PDF | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| MAGICS-LAB/DNABERT_2 official source Pinned README: GUE section; data download and evaluation scripts; Appendix C: task construction; Table 9; Section 5; Appendix C and Table 9: GUE task datasets Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: f25bed9ee20db966dff39e5c1571249d04e36404 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Evaluation procedure Individual claims | gue primary benchmark evidence Pinned README: GUE section; data download and evaluation scripts; Appendix C: task construction; Table 9; Section 5; Appendix C and Table 9: GUE task datasets Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 2306.15006 retrieved PDF | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Evaluation procedure Individual claims | MAGICS-LAB/DNABERT_2 official source Pinned README: GUE section; data download and evaluation scripts; Appendix C: task construction; Table 9; Section 5; Appendix C and Table 9: GUE task datasets Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: f25bed9ee20db966dff39e5c1571249d04e36404 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Datasets The benchmark archive contains separate sequence-classification datasets; model pretraining data are distributed separately. Individual claims | MAGICS-LAB/DNABERT_2 official source Pinned README: GUE section; data download and evaluation scripts Version: f25bed9ee20db966dff39e5c1571249d04e36404 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Splits Separate train/validation/test files and counts are defined per dataset. The yeast epigenetic tasks use an 8:1:1 random split; source-derived regulatory and viral tasks retain their own documented partitions. There is no universal chromosome holdout across GUE. Individual claims | gue primary benchmark evidence Appendix C: task construction; Table 9 Version: 2306.15006 retrieved PDF | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Adaptation Supervised fine-tuning; scripts include model-specific training examples. Individual claims | MAGICS-LAB/DNABERT_2 official source Pinned README: GUE section; data download and evaluation scripts Version: f25bed9ee20db966dff39e5c1571249d04e36404 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Metrics GUE uses Matthews correlation coefficient for the regulatory classification tasks and F1 for COVID variant classification; the exact task metric is tabulated in the dataset appendix. Individual claims | gue primary benchmark evidence Section 5; Appendix C and Table 9: GUE task datasets Version: 2306.15006 retrieved PDF | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
View linked audit checks and correction history
Release 2026-09-29-06401fd5b220 · Record review: discovered
Stable ID: discovery-benchmark-gue