rewire.itbenchmarks
Benchmark

GUE

GUE evaluates genome understanding across multiple datasets, task types and species.

SourcesMAGICS-LAB/DNABERT_2 official source · Pinned README: GUE section; data download and evaluation scripts

280 evaluations · 280 results

Overview

Datasets

The benchmark archive contains separate sequence-classification datasets; model pretraining data are distributed separately.

SourcesMAGICS-LAB/DNABERT_2 official source · Pinned README: GUE section; data download and evaluation scripts

Metrics

GUE uses Matthews correlation coefficient for the regulatory classification tasks and F1 for COVID variant classification; the exact task metric is tabulated in the dataset appendix.

Sourcesgue primary benchmark evidence · Section 5; Appendix C and Table 9: GUE task datasets

Allowed inputs

DNA sequences from the benchmark archive, separate from model-pretraining data.

SourcesMAGICS-LAB/DNABERT_2 official source · Pinned README: GUE section; data download and evaluation scripts
Evaluation procedure diagram
How it worksEvaluation procedure
Evaluation procedure1. Allowed inputs: DNA sequences from the benchmark archive, separate from model-pretraining data.. Then: 2. Splits: Separate train/validation/test files and counts are defined per dataset. The yeast epigenetic tasks use an 8:1:1 random split; source-derived regulatory and viral tasks retain their own documented partitions. There is no universal chromosome holdout across GUE.. Then: 3. Metrics: GUE uses Matthews correlation coefficient for the regulatory classification tasks and F1 for COVID variant classification; the exact task metric is tabulated in the dataset appendix.Evaluation procedure1. Allowed inputs: DNA sequences from the benchmark archive, separate from model-pretraining data.. Then: 2. Splits: Separate train/validation/test files and counts are defined per dataset. The yeast epigenetic tasks use an 8:1:1 random split; source-derived regulatory and viral tasks retain their own documented partitions. There is no universal chromosome holdout across GUE.. Then: 3. Metrics: GUE uses Matthews correlation coefficient for the regulatory classification tasks and F1 for COVID variant classification; the exact task metric is tabulated in the dataset appendix.Evaluation procedure1. Allowed inputs: DNA sequences from the benchmark archive, separate from model-pretraining data.. Then: 2. Splits: Separate train/validation/test files and counts are defined per dataset. The yeast epigenetic tasks use an 8:1:1 random split; source-derived regulatory and viral tasks retain their own documented partitions. There is no universal chromosome holdout across GUE.. Then: 3. Metrics: GUE uses Matthews correlation coefficient for the regulatory classification tasks and F1 for COVID variant classification; the exact task metric is tabulated in the dataset appendix.

Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.

Sources (2)MAGICS-LAB/DNABERT_2 official source; gue primary benchmark evidence · Pinned README: GUE section; data download and evaluation scripts; Appendix C: task construction; Table 9; Section 5; Appendix C and Table 9: GUE task datasets

Source reviewed · Automated source review, 2026-09-16. All specifications and missing details

Results

Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.

GUE CORE-PROMOTER-DETECTION-ALL: Core promoter detection, dataset all

mcc (percent) · Higher values are better.

GUE CORE-PROMOTER-DETECTION-ALL: Core promoter detection, dataset all · GUE Core promoter detection, all (GUE split)

Evidence origin: Author-reported evaluation.

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 12, row(Core promoter detection)
  • Scores are MCC, except Covid variant classification which is F1, both on a 0 to 100 scale.
  • The diamond entry is DNABERT-2 with further pre-training on the GUE training sets, so it is not directly comparable to the others.
Comparison details and limitations

Every method GUE reports on Core promoter detection, dataset all, scored with MCC on GUE Core promoter detection, all.

  • Author-reported numbers, source checked but not independently reproduced.

Automated source review: 2026-09-18. Numerical source review does not establish independent reproduction.

Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.

Showing 10 of 10 matching rows.

Methods and evaluation design

Procedure, tasks and evaluated configurations

How it works

Evaluation methodology

GUE combines genome sequence classification tasks with supplied partitions and task-specific metrics. Models are fine-tuned on labelled training examples, selected with validation data and scored on test data. Random partitions occur in the suite, so a high score does not automatically demonstrate transfer to unrelated genomes.

Sourcesgue primary benchmark evidence · Section 5; Appendix C and Table 9: GUE task datasets

Evaluation design

Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.

Tasks

These source-backed links do not make different protocols or scores interchangeable.

Baseline coverage

Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.

No concrete protocols are explicitly linked to this suite. Protocol identification and baseline selection are outstanding.

Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums

Coverage is derived from release 2026-09-29-06401fd5b220. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.

Run this benchmark

Evaluate a model on the GUE benchmark

Fine-tune and score a model across the GUE datasets, using the authors' own evaluation script.

Generate predictions and evaluate them. This recipe does not establish reproduction of a particular published score.

Dataset access
The GUE archive, downloaded separately as the README describes.
Model and weights
A published checkpoint, or your own.
Licences
Project licence: Apache-2.0. Upstream data licences are separate and unreported here.
Software
Python with transformers and the repository's finetune scripts.
Hardware
Not stated in the cited section. Several of these steps expect a GPU.
Required inputs and expected outputs

Inputs

  • A tokeniser and model checkpoint loadable by transformers.
  • The GUE dataset directory.

Outputs

  • Per-dataset scores under the split the benchmark fixes.

Execution steps

  1. 1. Load the model (Python)

    Source reviewed; these instructions have not been executed by rewire.

    import torch
    from transformers import AutoTokenizer, AutoModel
    
    tokenizer = AutoTokenizer.from_pretrained("zhihan1996/DNABERT-2-117M", trust_remote_code=True)
    model = AutoModel.from_pretrained("zhihan1996/DNABERT-2-117M", trust_remote_code=True)
    DNABERT-2 and GUE: repository README · README.md at f25bed9e, 4. Quick Start, lines 84-88
  2. 2. Evaluate on GUE (Command line)

    Source reviewed; these instructions have not been executed by rewire.

    export DATA_PATH=/path/to/GUE #(e.g., /home/user)
    cd finetune
    
    # Evaluate DNABERT-2 on GUE
    sh scripts/run_dnabert2.sh DATA_PATH
    
    # Evaluate DNABERT (e.g., DNABERT with 3-mer) on GUE
    # 3 for 3-mer, 4 for 4-mer, 5 for 5-mer, 6 for 6-mer
    sh scripts/run_dnabert1.sh DATA_PATH 3
    
    # Evaluate Nucleotide Transformers on GUE
    # 0 for 500m-1000g, 1 for 500m-human-ref, 2 for 2.5b-1000g, 3 for 2.5b-multi-species
    sh scripts/run_nt.sh DATA_PATH 0
    DNABERT-2 and GUE: repository README · README.md at f25bed9e, 6.1 Evaluate models on GUE, lines 138-151

Use your own model

Run your model locally and return predictions keyed by the input IDs. The evaluator supplies biological inputs without test labels and owns scoring. This interface is not a sandbox for model code.

Pass your existing prediction function into this adapter. Its output direction must match the selected protocol.

class MyModelAdapter:
    def __init__(self, score):
        self.score = score

    def predict(self, inputs):
        return {row["id"]: float(self.score(row)) for row in inputs}

# adapter = MyModelAdapter(your_prediction_function)
# report = rewirebench.run(prepared, adapter, output="runs/my-model")

Alternatively, generate a keyed prediction file in your existing model environment and use the score-only recipe. Your model code and weights do not need to be shared.

DNABERT-2 and GUE: repository README · README.md at f25bed9e
Scope and limitations
  • Quoted from the project's README and not executed by rewire, so the commands are evidence of what the project documents rather than a verified run.
  • The project may have changed since the pinned commit.
  • Scores are MCC, except Covid variant classification which is F1.
  • The diamond entry on this page is DNABERT-2 with further pre-training on the GUE training sets.

Contribute a result for review. The library can submit an exported evaluation for private review when intake is open. Check the contribution page for access and sign-in.

Original repository instructions

Run this benchmark

The DNABERT-2 repository supplies GUE fine-tuning scripts and data links. Its documented configuration uses four GPUs and a global batch size of 32; other hardware requires an explicit batch configuration. The shell example declares DATA_PATH but passes the literal token DATA_PATH, so inspect the script argument handling before copying it.

A maintained rewire runner has not been verified for this benchmark. Check data access, weights, licences, dependencies and hardware in the linked official documentation; requirements have not been fully extracted.

MAGICS-LAB/DNABERT_2 / README.md · README.md lines 129–150 (Evaluate models on GUE)
Strengths, limitations and unresolved questions

Strengths and limitations

Strengths supported by sources

Limitations and conditions

  • Dataset partitions and metrics are task-specific. Non-overlapping tokenization addresses masked-token information leakage, which is a different issue from biological homology or pretraining overlap.
    Sourcesgue primary benchmark evidence · Section 5; Appendix C and Table 9: GUE task datasets
Profile review details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Stable record: discovery-benchmark-gue

Specifications

Inputs, training, access and other details

Explanatory profile: source reviewed · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetsThe benchmark archive contains separate sequence-classification datasets; model pretraining data are distributed separately.
SourcesMAGICS-LAB/DNABERT_2 official source · Pinned README: GUE section; data download and evaluation scripts
SplitsSeparate train/validation/test files and counts are defined per dataset. The yeast epigenetic tasks use an 8:1:1 random split; source-derived regulatory and viral tasks retain their own documented partitions. There is no universal chromosome holdout across GUE.
Sourcesgue primary benchmark evidence · Appendix C: task construction; Table 9
MetricsGUE uses Matthews correlation coefficient for the regulatory classification tasks and F1 for COVID variant classification; the exact task metric is tabulated in the dataset appendix.
Sourcesgue primary benchmark evidence · Section 5; Appendix C and Table 9: GUE task datasets
BaselinesDNABERT-2 and other genomic representation models are compared in the associated benchmark.
SourcesMAGICS-LAB/DNABERT_2 official source · Pinned README: GUE section; data download and evaluation scripts
Leakage controlsDatasets have explicit train/validation/test partitions, but some use random splits, including yeast epigenetic marks at 8:1:1. The paper’s masked-token leakage discussion concerns tokenization and must not be confused with proof of train/test genome independence.
Sourcesgue primary benchmark evidence · Section 5; Appendix C and Table 9: GUE task datasets
UncertaintyModels are fine-tuned with three different random seeds and the mean test result is reported. The stated protocol does not define a uniform confidence interval for every dataset.
Sourcesgue primary benchmark evidence · Section 5; Appendix C and Table 9: GUE task datasets
Entity typeGenome Understanding Evaluation benchmark suite.
SourcesMAGICS-LAB/DNABERT_2 official source · Pinned README: GUE section; data download and evaluation scripts
OrganismsMultiple species; the README describes four-species coverage.
SourcesMAGICS-LAB/DNABERT_2 official source · Pinned README: GUE section; data download and evaluation scripts
AssaysTask-specific genomic classification labels.
SourcesMAGICS-LAB/DNABERT_2 official source · Pinned README: GUE section; data download and evaluation scripts
Allowed inputsDNA sequences from the benchmark archive, separate from model-pretraining data.
SourcesMAGICS-LAB/DNABERT_2 official source · Pinned README: GUE section; data download and evaluation scripts
AdaptationSupervised fine-tuning; scripts include model-specific training examples.
SourcesMAGICS-LAB/DNABERT_2 official source · Pinned README: GUE section; data download and evaluation scripts
Applicable tests and references

Applicability is distinct from a completed evaluation.

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Dated primary-source discovery and protocol/table screening. Source checking does not mean experimental reproduction. Only separately extracted and independently reviewed numeric batches are publishable.

Paper or primary resourceVersionReference
DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species GenomesPrimary full-text snapshot retrieved 2026-09-17; exact bytes pinned by SHA-256Read source
Historical gaps recorded on 2026-09-17

The catalogue now holds 280 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.

  • complete comparable numeric result batch: Specific candidate tables and protocol boundaries are documented; no graph values or incomplete winner-only selection are converted into publishable rows.
Search and extraction details

primary protocol reviewed

Searches

  • DNABERT2 GUE benchmark 2306.15006

Evidence locations

  • ICLR 2024 paper Tables 4–6; Appendix A.2

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

28 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Individual claims
gue primary benchmark evidence

Original source ↗

Pinned README: GUE section; data download and evaluation scripts; Appendix C: task construction; Table 9; Section 5; Appendix C and Table 9: GUE task datasets

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2306.15006 retrieved PDF
Retrieved: 2026-09-16T20:00:00+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 49300acee3e4afd44bebc3de9893c3bc310d331bd4805374e0952fdfbf366f06

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram caption
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Individual claims
MAGICS-LAB/DNABERT_2 official source

Original source ↗

Pinned README: GUE section; data download and evaluation scripts; Appendix C: task construction; Table 9; Section 5; Appendix C and Table 9: GUE task datasets

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: f25bed9ee20db966dff39e5c1571249d04e36404
Retrieved: 2026-09-16T10:30:20.936532+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 734a8cec5f667d74d421bf3b273ad7e256216109636da45aa7ceba21cd34de16

Hash scope: Hash scope not separately documented; inspect source record

Diagram steps
  • Allowed inputs: DNA sequences from the benchmark archive, separate from model-pretraining data.
  • Splits: Separate train/validation/test files and counts are defined per dataset. The yeast epigenetic tasks use an 8:1:1 random split; source-derived regulatory and viral tasks retain their own documented partitions. There is no universal chromosome holdout across GUE.
  • Metrics: GUE uses Matthews correlation coefficient for the regulatory classification tasks and F1 for COVID variant classification; the exact task metric is tabulated in the dataset appendix.
Individual claims
gue primary benchmark evidence

Original source ↗

Pinned README: GUE section; data download and evaluation scripts; Appendix C: task construction; Table 9; Section 5; Appendix C and Table 9: GUE task datasets

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2306.15006 retrieved PDF
Retrieved: 2026-09-16T20:00:00+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 49300acee3e4afd44bebc3de9893c3bc310d331bd4805374e0952fdfbf366f06

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps
  • Allowed inputs: DNA sequences from the benchmark archive, separate from model-pretraining data.
  • Splits: Separate train/validation/test files and counts are defined per dataset. The yeast epigenetic tasks use an 8:1:1 random split; source-derived regulatory and viral tasks retain their own documented partitions. There is no universal chromosome holdout across GUE.
  • Metrics: GUE uses Matthews correlation coefficient for the regulatory classification tasks and F1 for COVID variant classification; the exact task metric is tabulated in the dataset appendix.
Individual claims
MAGICS-LAB/DNABERT_2 official source

Original source ↗

Pinned README: GUE section; data download and evaluation scripts; Appendix C: task construction; Table 9; Section 5; Appendix C and Table 9: GUE task datasets

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: f25bed9ee20db966dff39e5c1571249d04e36404
Retrieved: 2026-09-16T10:30:20.936532+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 734a8cec5f667d74d421bf3b273ad7e256216109636da45aa7ceba21cd34de16

Hash scope: Hash scope not separately documented; inspect source record

Diagram title
Evaluation procedure
Individual claims
gue primary benchmark evidence

Original source ↗

Pinned README: GUE section; data download and evaluation scripts; Appendix C: task construction; Table 9; Section 5; Appendix C and Table 9: GUE task datasets

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2306.15006 retrieved PDF
Retrieved: 2026-09-16T20:00:00+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: 49300acee3e4afd44bebc3de9893c3bc310d331bd4805374e0952fdfbf366f06

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title
Evaluation procedure
Individual claims
MAGICS-LAB/DNABERT_2 official source

Original source ↗

Pinned README: GUE section; data download and evaluation scripts; Appendix C: task construction; Table 9; Section 5; Appendix C and Table 9: GUE task datasets

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: f25bed9ee20db966dff39e5c1571249d04e36404
Retrieved: 2026-09-16T10:30:20.936532+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: 734a8cec5f667d74d421bf3b273ad7e256216109636da45aa7ceba21cd34de16

Hash scope: Hash scope not separately documented; inspect source record

Datasets
The benchmark archive contains separate sequence-classification datasets; model pretraining data are distributed separately.
Individual claims
MAGICS-LAB/DNABERT_2 official source

Original source ↗

Pinned README: GUE section; data download and evaluation scripts

Version: f25bed9ee20db966dff39e5c1571249d04e36404
Retrieved: 2026-09-16T10:30:20.936532+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: 734a8cec5f667d74d421bf3b273ad7e256216109636da45aa7ceba21cd34de16

Hash scope: Hash scope not separately documented; inspect source record

Splits
Separate train/validation/test files and counts are defined per dataset. The yeast epigenetic tasks use an 8:1:1 random split; source-derived regulatory and viral tasks retain their own documented partitions. There is no universal chromosome holdout across GUE.
Individual claims
gue primary benchmark evidence

Original source ↗

Appendix C: task construction; Table 9

Version: 2306.15006 retrieved PDF
Retrieved: 2026-09-16T20:00:00+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: 49300acee3e4afd44bebc3de9893c3bc310d331bd4805374e0952fdfbf366f06

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Adaptation
Supervised fine-tuning; scripts include model-specific training examples.
Individual claims
MAGICS-LAB/DNABERT_2 official source

Original source ↗

Pinned README: GUE section; data download and evaluation scripts

Version: f25bed9ee20db966dff39e5c1571249d04e36404
Retrieved: 2026-09-16T10:30:20.936532+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: 734a8cec5f667d74d421bf3b273ad7e256216109636da45aa7ceba21cd34de16

Hash scope: Hash scope not separately documented; inspect source record

Metrics
GUE uses Matthews correlation coefficient for the regulatory classification tasks and F1 for COVID variant classification; the exact task metric is tabulated in the dataset appendix.
Individual claims
gue primary benchmark evidence

Original source ↗

Section 5; Appendix C and Table 9: GUE task datasets

Version: 2306.15006 retrieved PDF
Retrieved: 2026-09-16T20:00:00+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: 49300acee3e4afd44bebc3de9893c3bc310d331bd4805374e0952fdfbf366f06

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: discovered

5 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: discovery-benchmark-gue

areas
genomics
entity level
suite
scope note
Specialist molecular or omics evaluation; protocol details require review before numerical comparison.
task
Multi-species genome understanding tasks
version
Not reported
benchmark research
review date: 2026-09-17; status: primary_protocol_reviewed; primary sources: evidence-expansion-gue-49300ace; inspected locators: ICLR 2024 paper Tables 4–6; Appendix A.2; searched queries: DNABERT2 GUE benchmark 2306.15006; gaps: complete comparable numeric result batch: Specific candidate tables and protocol boundaries are documented; no graph values or incomplete winner-only selection are converted into publishable rows.; claim scope: Dated primary-source discovery and protocol/table screening. Source checking does not mean experimental reproduction. Only separately extracted and independently reviewed numeric batches are publishable.
historical missing metadata
dataset release: unextracted; metric implementation: unextracted; split manifest: unextracted; version: unextracted
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
entity classification
review date: 2026-09-17; rationale: The cited profile describes a collection of evaluation tasks or protocols; retain it as the top-level benchmark suite. Its datasets and individual protocols remain separate records.; source ids: src-discovery-magics-lab-dnabert-2; source locator: Pinned README: GUE section; data download and evaluation scripts; ambiguities: None recorded
run documentation
record id: discovery-benchmark-gue; source ids: run-doc-gue-readme-md-f25bed9e; status: official_documentation_linked; summary: The DNABERT-2 repository supplies GUE fine-tuning scripts and data links. Its documented configuration uses four GPUs and a global batch size of 32; other hardware requires an explicit batch configuration. The shell example declares DATA_PATH but passes the literal token DATA_PATH, so inspect the script argument handling before copying it.; source locator: README.md lines 129–150 (Evaluate models on GUE)
run recipes
id: gue-official; protocol id: discovery-benchmark-gue; version: f25bed9ee20db966dff39e5c1571249d04e36404; title: Evaluate a model on the GUE benchmark; purpose: generate_and_evaluate; summary: Fine-tune and score a model across the GUE datasets, using the authors' own evaluation script.; inputs: A tokeniser and model checkpoint loadable by transformers.; The GUE dataset directory.; outputs: Per-dataset scores under the split the benchmark fixes.; requirements: data: The GUE archive, downloaded separately as the README describes.; weights: A published checkpoint, or your own.; licence: Project licence: Apache-2.0. Upstream data licences are separate and unreported here.; software: Python with transformers and the repository's finetune scripts.; hardware: Not stated in the cited section. Several of these steps expect a GPU.; instructions: runtime: python; title: Load the model; code: import torch from transformers import AutoTokenizer, AutoModel tokenizer = AutoTokenizer.from_pretrained("zhihan1996/DNABERT-2-117M", trust_remote_code=True) model = AutoModel.from_pretrained("zhihan1996/DNABERT-2-117M", trust_remote_code=True); status: source_reviewed_not_executed; source ids: project-recipe-gue-f25bed9e; source locator: README.md at f25bed9e, 4. Quick Start, lines 84-88; runtime: command_line; title: Evaluate on GUE; code: export DATA_PATH=/path/to/GUE #(e.g., /home/user) cd finetune # Evaluate DNABERT-2 on GUE sh scripts/run_dnabert2.sh DATA_PATH # Evaluate DNABERT (e.g., DNABERT with 3-mer) on GUE # 3 for 3-mer, 4 for 4-mer, 5 for 5-mer, 6 for 6-mer sh scripts/run_dnabert1.sh DATA_PATH 3 # Evaluate Nucleotide Transformers on GUE # 0 for 500m-1000g, 1 for 500m-human-ref, 2 for 2.5b-1000g, 3 for 2.5b-multi-species sh scripts/run_nt.sh DATA_PATH 0; status: source_reviewed_not_executed; source ids: project-recipe-gue-f25bed9e; source locator: README.md at f25bed9e, 6.1 Evaluate models on GUE, lines 138-151; limitations: Quoted from the project's README and not executed by rewire, so the commands are evidence of what the project documents rather than a verified run.; The project may have changed since the pinned commit.; Scores are MCC, except Covid variant classification which is F1.; The diamond entry on this page is DNABERT-2 with further pre-training on the GUE training sets.; source ids: project-recipe-gue-f25bed9e; source locator: README.md at f25bed9e
Related records

Suggest a correction