rewire.itbenchmarks
Configuration

DNABERT-2 (further pre-trained on GUE)

DNABERT-2 learns DNA representations that can be adapted to genomic prediction tasks.

Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0

28 evaluations · 28 results

How it worksDNABERT-2 workflow
DNABERT-2 workflow1. DNA sequence. Then: 2. BPE tokens. Then: 3. Transformer encoder. Then: 4. Representations. Then: 5. Specified task headDNABERT-2 workflow1. DNA sequence. Then: 2. BPE tokens. Then: 3. Transformer encoder. Then: 4. Representations. Then: 5. Specified task headDNABERT-2 workflow1. DNA sequence. Then: 2. BPE tokens. Then: 3. Transformer encoder. Then: 4. Representations. Then: 5. Specified task head

Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.

Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0

Overview

Model type

Masked-token DNA transformer encoder

Inputs

DNA sequence tokenized with the supplied tokenizer.

Outputs

Token representations and, after a specified adaptation, task predictions.

Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0

limited source coverage · Automated source review, 2026-09-23. All specifications and missing details

Evaluations and results

28 evaluations · 28 results. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: DNABERT-2 (further pre-trained on GUE)Task: GUE CORE-PROMOTER-DETECTION-ALL: Core promoter detection, dataset all
Dataset subset: GUE Core promoter detection, all (GUE split)
67.5% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (further pre-trained on GUE) on GUE CORE-PROMOTER-DETECTION-ALL: Core promoter detection, dataset all

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Core promoter detection all)
Configuration: DNABERT-2 (further pre-trained on GUE)Task: GUE CORE-PROMOTER-DETECTION-NOTATA: Core promoter detection, dataset notata
Dataset subset: GUE Core promoter detection, notata (GUE split)
69.5% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (further pre-trained on GUE) on GUE CORE-PROMOTER-DETECTION-NOTATA: Core promoter detection, dataset notata

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Core promoter detection notata)
Configuration: DNABERT-2 (further pre-trained on GUE)Task: GUE CORE-PROMOTER-DETECTION-TATA: Core promoter detection, dataset tata
Dataset subset: GUE Core promoter detection, tata (GUE split)
76.2% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (further pre-trained on GUE) on GUE CORE-PROMOTER-DETECTION-TATA: Core promoter detection, dataset tata

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Core promoter detection tata)
Configuration: DNABERT-2 (further pre-trained on GUE)Task: GUE COVID-VARIANT-CLASSIFICATION-COVID: Covid variant classification, dataset Covid
Dataset subset: GUE Covid variant classification, Covid (GUE split)
68.5% f1
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (further pre-trained on GUE) on GUE COVID-VARIANT-CLASSIFICATION-COVID: Covid variant classification, dataset Covid

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Covid variant classification Covid)
Configuration: DNABERT-2 (further pre-trained on GUE)Task: GUE EPIGENETIC-MARKS-PREDICTION-H3: Epigenetic marks prediction, dataset H3
Dataset subset: GUE Epigenetic marks prediction, H3 (GUE split)
80.2% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (further pre-trained on GUE) on GUE EPIGENETIC-MARKS-PREDICTION-H3: Epigenetic marks prediction, dataset H3

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Epigenetic marks prediction H3)
Configuration: DNABERT-2 (further pre-trained on GUE)Task: GUE EPIGENETIC-MARKS-PREDICTION-H3K14AC: Epigenetic marks prediction, dataset H3K14ac
Dataset subset: GUE Epigenetic marks prediction, H3K14ac (GUE split)
57.4% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (further pre-trained on GUE) on GUE EPIGENETIC-MARKS-PREDICTION-H3K14AC: Epigenetic marks prediction, dataset H3K14ac

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Epigenetic marks prediction H3K14ac)
Configuration: DNABERT-2 (further pre-trained on GUE)Task: GUE EPIGENETIC-MARKS-PREDICTION-H3K36ME3: Epigenetic marks prediction, dataset H3K36me3
Dataset subset: GUE Epigenetic marks prediction, H3K36me3 (GUE split)
61.9% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (further pre-trained on GUE) on GUE EPIGENETIC-MARKS-PREDICTION-H3K36ME3: Epigenetic marks prediction, dataset H3K36me3

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Epigenetic marks prediction H3K36me3)
Configuration: DNABERT-2 (further pre-trained on GUE)Task: GUE EPIGENETIC-MARKS-PREDICTION-H3K4ME1: Epigenetic marks prediction, dataset H3K4me1
Dataset subset: GUE Epigenetic marks prediction, H3K4me1 (GUE split)
53% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (further pre-trained on GUE) on GUE EPIGENETIC-MARKS-PREDICTION-H3K4ME1: Epigenetic marks prediction, dataset H3K4me1

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Epigenetic marks prediction H3K4me1)
Configuration: DNABERT-2 (further pre-trained on GUE)Task: GUE EPIGENETIC-MARKS-PREDICTION-H3K4ME2: Epigenetic marks prediction, dataset H3K4me2
Dataset subset: GUE Epigenetic marks prediction, H3K4me2 (GUE split)
39.9% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (further pre-trained on GUE) on GUE EPIGENETIC-MARKS-PREDICTION-H3K4ME2: Epigenetic marks prediction, dataset H3K4me2

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Epigenetic marks prediction H3K4me2)
Configuration: DNABERT-2 (further pre-trained on GUE)Task: GUE EPIGENETIC-MARKS-PREDICTION-H3K4ME3: Epigenetic marks prediction, dataset H3K4me3
Dataset subset: GUE Epigenetic marks prediction, H3K4me3 (GUE split)
41.2% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (further pre-trained on GUE) on GUE EPIGENETIC-MARKS-PREDICTION-H3K4ME3: Epigenetic marks prediction, dataset H3K4me3

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Epigenetic marks prediction H3K4me3)
Configuration: DNABERT-2 (further pre-trained on GUE)Task: GUE EPIGENETIC-MARKS-PREDICTION-H3K79ME3: Epigenetic marks prediction, dataset H3K79me3
Dataset subset: GUE Epigenetic marks prediction, H3K79me3 (GUE split)
65.5% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (further pre-trained on GUE) on GUE EPIGENETIC-MARKS-PREDICTION-H3K79ME3: Epigenetic marks prediction, dataset H3K79me3

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Epigenetic marks prediction H3K79me3)
Configuration: DNABERT-2 (further pre-trained on GUE)Task: GUE EPIGENETIC-MARKS-PREDICTION-H3K9AC: Epigenetic marks prediction, dataset H3K9ac
Dataset subset: GUE Epigenetic marks prediction, H3K9ac (GUE split)
57.1% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (further pre-trained on GUE) on GUE EPIGENETIC-MARKS-PREDICTION-H3K9AC: Epigenetic marks prediction, dataset H3K9ac

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Epigenetic marks prediction H3K9ac)
Configuration: DNABERT-2 (further pre-trained on GUE)Task: GUE EPIGENETIC-MARKS-PREDICTION-H4: Epigenetic marks prediction, dataset H4
Dataset subset: GUE Epigenetic marks prediction, H4 (GUE split)
81.9% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (further pre-trained on GUE) on GUE EPIGENETIC-MARKS-PREDICTION-H4: Epigenetic marks prediction, dataset H4

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Epigenetic marks prediction H4)
Configuration: DNABERT-2 (further pre-trained on GUE)Task: GUE EPIGENETIC-MARKS-PREDICTION-H4AC: Epigenetic marks prediction, dataset H4ac
Dataset subset: GUE Epigenetic marks prediction, H4ac (GUE split)
50.4% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (further pre-trained on GUE) on GUE EPIGENETIC-MARKS-PREDICTION-H4AC: Epigenetic marks prediction, dataset H4ac

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Epigenetic marks prediction H4ac)
Configuration: DNABERT-2 (further pre-trained on GUE)Task: GUE PROMOTER-DETECTION-ALL: Promoter detection, dataset all
Dataset subset: GUE Promoter detection, all (GUE split)
88.3% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (further pre-trained on GUE) on GUE PROMOTER-DETECTION-ALL: Promoter detection, dataset all

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Promoter detection all)
Configuration: DNABERT-2 (further pre-trained on GUE)Task: GUE PROMOTER-DETECTION-NOTATA: Promoter detection, dataset notata
Dataset subset: GUE Promoter detection, notata (GUE split)
94.3% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (further pre-trained on GUE) on GUE PROMOTER-DETECTION-NOTATA: Promoter detection, dataset notata

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Promoter detection notata)
Configuration: DNABERT-2 (further pre-trained on GUE)Task: GUE PROMOTER-DETECTION-TATA: Promoter detection, dataset tata
Dataset subset: GUE Promoter detection, tata (GUE split)
68.8% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (further pre-trained on GUE) on GUE PROMOTER-DETECTION-TATA: Promoter detection, dataset tata

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Promoter detection tata)
Configuration: DNABERT-2 (further pre-trained on GUE)Task: GUE SPLICE-SITE-PREDICTION-RECONSTRUCT: Splice site prediction, dataset Reconstruct
Dataset subset: GUE Splice site prediction, Reconstruct (GUE split)
85.9% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (further pre-trained on GUE) on GUE SPLICE-SITE-PREDICTION-RECONSTRUCT: Splice site prediction, dataset Reconstruct

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Splice site prediction Reconstruct)
Configuration: DNABERT-2 (further pre-trained on GUE)Task: GUE TRANSCRIPTION-FACTOR-PREDICTION-HUMAN-0: Transcription factor prediction (human), dataset 0
Dataset subset: GUE Transcription factor prediction (human), 0 (GUE split)
69.1% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (further pre-trained on GUE) on GUE TRANSCRIPTION-FACTOR-PREDICTION-HUMAN-0: Transcription factor prediction (human), dataset 0

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Transcription factor prediction (human) 0)
Configuration: DNABERT-2 (further pre-trained on GUE)Task: GUE TRANSCRIPTION-FACTOR-PREDICTION-HUMAN-1: Transcription factor prediction (human), dataset 1
Dataset subset: GUE Transcription factor prediction (human), 1 (GUE split)
71.9% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (further pre-trained on GUE) on GUE TRANSCRIPTION-FACTOR-PREDICTION-HUMAN-1: Transcription factor prediction (human), dataset 1

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Transcription factor prediction (human) 1)
Configuration: DNABERT-2 (further pre-trained on GUE)Task: GUE TRANSCRIPTION-FACTOR-PREDICTION-HUMAN-2: Transcription factor prediction (human), dataset 2
Dataset subset: GUE Transcription factor prediction (human), 2 (GUE split)
63% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (further pre-trained on GUE) on GUE TRANSCRIPTION-FACTOR-PREDICTION-HUMAN-2: Transcription factor prediction (human), dataset 2

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Transcription factor prediction (human) 2)
Configuration: DNABERT-2 (further pre-trained on GUE)Task: GUE TRANSCRIPTION-FACTOR-PREDICTION-HUMAN-3: Transcription factor prediction (human), dataset 3
Dataset subset: GUE Transcription factor prediction (human), 3 (GUE split)
55.4% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (further pre-trained on GUE) on GUE TRANSCRIPTION-FACTOR-PREDICTION-HUMAN-3: Transcription factor prediction (human), dataset 3

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Transcription factor prediction (human) 3)
Configuration: DNABERT-2 (further pre-trained on GUE)Task: GUE TRANSCRIPTION-FACTOR-PREDICTION-HUMAN-4: Transcription factor prediction (human), dataset 4
Dataset subset: GUE Transcription factor prediction (human), 4 (GUE split)
74.9% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (further pre-trained on GUE) on GUE TRANSCRIPTION-FACTOR-PREDICTION-HUMAN-4: Transcription factor prediction (human), dataset 4

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Transcription factor prediction (human) 4)
Configuration: DNABERT-2 (further pre-trained on GUE)Task: GUE TRANSCRIPTION-FACTOR-PREDICTION-MOUSE-0: Transcription factor prediction (mouse), dataset 0
Dataset subset: GUE Transcription factor prediction (mouse), 0 (GUE split)
64.2% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (further pre-trained on GUE) on GUE TRANSCRIPTION-FACTOR-PREDICTION-MOUSE-0: Transcription factor prediction (mouse), dataset 0

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Transcription factor prediction (mouse) 0)
Configuration: DNABERT-2 (further pre-trained on GUE)Task: GUE TRANSCRIPTION-FACTOR-PREDICTION-MOUSE-1: Transcription factor prediction (mouse), dataset 1
Dataset subset: GUE Transcription factor prediction (mouse), 1 (GUE split)
86.3% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (further pre-trained on GUE) on GUE TRANSCRIPTION-FACTOR-PREDICTION-MOUSE-1: Transcription factor prediction (mouse), dataset 1

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Transcription factor prediction (mouse) 1)

Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.

Use this model

How it works, versions and access

Related profile: DNABERT-2. This page retains the exact record and its evaluation context.

This configuration

Genome language model fine-tuned on each GUE dataset by the DNABERT-2 authors.

record
DNABERT-2 (further pre-trained on GUE)
configuration
Not reported
entity type
Configuration

How it works

How it works

DNABERT-2 merges recurring DNA substrings into byte-pair tokens, then processes those tokens with a masked-language-model transformer. ALiBi supplies distance-dependent attention biases, while FlashAttention changes how attention is computed. The resulting contextual embeddings need an explicit pooling rule and prediction head for a downstream task.

Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
Versions and reproducibility

DNABERT-2-117M model card and official DNABERT_2 implementation. ALiBi permits inference beyond the pretraining sequence length, subject to attention/memory cost; this does not establish unlimited biological context or validated accuracy at arbitrary lengths.

Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
Strengths, limitations and unresolved questions

Strengths and limitations

Limitations and conditions

  • An embedding model alone is not the same evaluated pipeline as frozen embeddings followed by logistic regression. Tokenization and pooling choices must be preserved.
    Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
Profile review details

Follow-up review of Reference checkpoint, Context limits, Training cutoff, Further pretraining. Source locations and before/after decisions are recorded in the 23 September profile-evidence audit. Other explanatory content retains its earlier source scope. No human scientific review or independent reproduction is implied.

Stable record: discovery-model-dnabert-2

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-23. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Inputs, outputs and configuration
PropertyDescription and evidence
Model typeMasked-token DNA transformer encoder
Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
ArchitectureBERT-style DNA encoder with byte-pair tokenization, ALiBi relative attention biases and FlashAttention; task heads and pooling are separately configured.
Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
InputsDNA sequence tokenized with the supplied tokenizer.
Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
OutputsToken representations and, after a specified adaptation, task predictions.
Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
Parameters117 million for DNABERT-2-117M; family names do not establish a particular checkpoint.
Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
Known versionsDNABERT-2-117M model card and official DNABERT_2 implementation.
Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
Training dataThe paper describes a 32.49-billion-base corpus covering 135 species in six groups, alongside a 2.75-billion-base human corpus. Further GUE-domain pretraining is a separately reported model variant.
Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
Training cutoffSection4.1 identifies the human and 135-species corpora but does not establish a latest-sequence deposition date. The model repository revision is not a training-data cutoff. · Not reported in inspected sources
Sourcesdnabert2 paper v2: primary artifact · Section4.1; Table11
Context limitsThe paper describes pretraining on 700-base sequences and evaluates 5,000–10,000-base GUE+ inputs after fine-tuning. This does not establish frozen-model accuracy at those lengths; tokenisation, truncation and adaptation must be specified.
Sourcesdnabert2 paper v2: primary artifact · Section5.4 Results on GUE+, printed p10; Table2 input lengths
Weights licenceThe official zhihan1996/DNABERT-2-117M checkpoint repository carries Apache-2.0 in its pinned LICENSE. This does not assign terms to a separately fitted downstream predictor.
Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
AccessOfficial project documentation and implementation: https://github.com/MAGICS-LAB/DNABERT_2
Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
Code licenceApache-2.0
SourcesMAGICS-LAB/DNABERT_2: LICENSE · LICENSE: licence text
Reference checkpointDNABERT-2-117M: Hugging Face revision 7bce263b15377fc15361f52cfab88f8b586abda0. Its pytorch_model.bin has registry-reported SHA-256 7ff39ec77a484dd01070a41bfd6e95cdd7247bec80fe357ab43a4be33687aeba. The weight file was not downloaded for this review.
Sourcesdnabert2 release: primary artifact · sha; siblings[pytorch_model.bin].lfs.sha256
Further pretrainingThe diamond-marked DNABERT-2 variant receives additional masked-language-model training on GUE training sets. It must be distinguished from the base pretrained model in comparisons.
Sourcesdnabert2 paper v2: primary artifact · Section5.2FurtherPreTraining; Table3caption

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

2 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Relationship: family
discovery-model-dnabert-2
Individual claims
DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes

Original source ↗

DNABERT-2 paper GUE result tables; further pre-training setting retained; official README GUE benchmark; source-labelled configuration DNABERT-2 (further pre-trained on GUE) | Existing reviewed locator: Table 6, row(DNABERT-2♦)

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: Primary full-text snapshot retrieved 2026-09-17; exact bytes pinned by SHA-256
Retrieved: 2026-09-17T08:06:28.183387+00:00

source checked

automated source review · 2026-09-23

Audit details

Source review establishes this relationship only. Exact evaluated configurations and original numerical review status remain unchanged. Both configurations are explicitly DNABERT-2; retain further pretraining as a distinct configuration.

Field: links:family:discovery-model-dnabert-2

Claim: model-evaluation-identity-642226226b4179cf9b07

Source artifact SHA-256: 49300acee3e4afd44bebc3de9893c3bc310d331bd4805374e0952fdfbf366f06

Hash scope: Exact retrieved primary paper artifact bytes.

Inspected artifact

Relationship: family
discovery-model-dnabert-2
Individual claims
MAGICS-LAB/DNABERT_2: README.md

Original source ↗

DNABERT-2 paper GUE result tables; further pre-training setting retained; official README GUE benchmark; source-labelled configuration DNABERT-2 (further pre-trained on GUE) | Existing reviewed locator: Table 6, row(DNABERT-2♦)

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: f25bed9ee20db966dff39e5c1571249d04e36404
Retrieved: 2026-09-16T19:46:17.892989+00:00

source checked

automated source review · 2026-09-23

Audit details

Source review establishes this relationship only. Exact evaluated configurations and original numerical review status remain unchanged. Both configurations are explicitly DNABERT-2; retain further pretraining as a distinct configuration.

Field: links:family:discovery-model-dnabert-2

Claim: model-evaluation-identity-642226226b4179cf9b07

Source artifact SHA-256: 734a8cec5f667d74d421bf3b273ad7e256216109636da45aa7ceba21cd34de16

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: source checked

1 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: gue-method-dnabert-2-further-pre-trained-on-gue

areas
dna-genomes
source locator
Table 6, row(DNABERT-2♦)
missing metadata
checkpoint revision: unreported; parameters: unextracted
Related records

Suggest a correction