rewire.itbenchmarks
Model

DNABERT-2

DNABERT-2 learns DNA representations that can be adapted to genomic prediction tasks.

Sources (5)zhihan1996/DNABERT-2-117M: README.md; zhihan1996/DNABERT-2-117M: config.json; MAGICS-LAB/DNABERT_2: README.md; zhihan1996/DNABERT-2-117M: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0

125 evaluations · 128 results · 1 evaluated configuration using this model

How it worksDNABERT-2 workflow
DNABERT-2 workflow1. DNA sequence. Then: 2. BPE tokens. Then: 3. Transformer encoder. Then: 4. Representations. Then: 5. Specified task headDNABERT-2 workflow1. DNA sequence. Then: 2. BPE tokens. Then: 3. Transformer encoder. Then: 4. Representations. Then: 5. Specified task headDNABERT-2 workflow1. DNA sequence. Then: 2. BPE tokens. Then: 3. Transformer encoder. Then: 4. Representations. Then: 5. Specified task head

Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.

Sources (5)zhihan1996/DNABERT-2-117M: README.md; zhihan1996/DNABERT-2-117M: config.json; MAGICS-LAB/DNABERT_2: README.md; zhihan1996/DNABERT-2-117M: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0

Overview

Model type

Masked-token DNA transformer encoder

Inputs

DNA sequence tokenized with the supplied tokenizer.

Outputs

Token representations and, after a specified adaptation, task predictions.

Sources (5)zhihan1996/DNABERT-2-117M: README.md; zhihan1996/DNABERT-2-117M: config.json; MAGICS-LAB/DNABERT_2: README.md; zhihan1996/DNABERT-2-117M: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0

limited source coverage · Automated source review, 2026-09-23. All specifications and missing details

Evaluations and results

125 evaluations · 128 results. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: DNABERT-2Task: regulatory element identification
Dataset: DART-Eval cCREs versus matched shuffled controls
0.876 accuracy
fraction · unknown

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

DNABERT-2: regulatory element identification

zero-shot likelihood ranking: higher likelihood for cCRE than matched control

Aggregation: Not reported

DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 3 (PDF page 5), DNABERT-2 row, Zero-Shot Accuracy column
Configuration: DNABERT-2Task: human core-promoter classification
Dataset: GUE H-CPD
70.5% MCC
percent · unknown

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Result quoted from another source · Source checked
Methods, coverage and source

DNABERT-2: human core-promoter classification

DNABERT-2 comparator in consolidated H-CPD table; rerun provenance not explicit

Aggregation: Not reported

EDEN: multiscale expected density of nucleotide encoding for enhanced DNA sequence classification with hybrid deep learning · Table 5, DNABERT-2 row, H-CPD (MCC) column
Configuration: DNABERT-2Task: BEND CHROMATIN: Chromatin accessibility
Dataset subset: ENCODE chromatin accessibility (BEND split)
0.81 auroc
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 on BEND CHROMATIN: Chromatin accessibility

A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1.

Aggregation: Not reported

BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(DNABERT-2), column(Chromatin accessibility)
Configuration: DNABERT-2Task: BEND CPG: CpG methylation
Dataset subset: ENCODE CpG methylation (BEND split)
0.9 auroc
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 on BEND CPG: CpG methylation

A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1.

Aggregation: Not reported

BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(DNABERT-2), column(CpG methylation)
Configuration: DNABERT-2Task: BEND ENHANCER: Enhancer annotation
Dataset subset: Fulco 2019, Gasperini 2019 and Enformer enhancer set (BEND split)
0.03 auprc
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 on BEND ENHANCER: Enhancer annotation

A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1.

Aggregation: Not reported

BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(DNABERT-2), column(Enhancer annotation)
Configuration: DNABERT-2Task: BEND GENE-FINDING: Gene finding
Dataset subset: GENCODE (BEND split)
0.43 mcc
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 on BEND GENE-FINDING: Gene finding

A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1.

Aggregation: Not reported

BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(DNABERT-2), column(Gene finding)
Configuration: DNABERT-2Task: BEND HISTONE: Histone modification
Dataset subset: ENCODE histone modification (BEND split)
0.78 auroc
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 on BEND HISTONE: Histone modification

A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1.

Aggregation: Not reported

BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(DNABERT-2), column(Histone modification)
Configuration: DNABERT-2Task: BEND VARIANT-DISEASE: Noncoding variant effects on disease
Dataset subset: ClinVar disease variants (BEND split)
0.51 auroc
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 on BEND VARIANT-DISEASE: Noncoding variant effects on disease

A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1.

Aggregation: Not reported

BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(DNABERT-2), column(Noncoding variant effects on disease)
Configuration: DNABERT-2Task: BEND VARIANT-EXPRESSION: Noncoding variant effects on expression
Dataset subset: DeepSEA expression variants (BEND split)
0.49 auroc
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 on BEND VARIANT-EXPRESSION: Noncoding variant effects on expression

A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1.

Aggregation: Not reported

BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(DNABERT-2), column(Noncoding variant effects on expression)
Configuration: DNABERT-2 (fine-tuned)Task: DART-Eval CA-AUROC-GM12878: Chromatin activity prediction, GM12878, positives against negatives
Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split)
0.916 auroc
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (fine-tuned) on DART-Eval CA-AUROC-GM12878: Chromatin activity prediction, GM12878, positives against negatives

Separating positive GM12878 peaks from matched negatives.

Aggregation: Not reported

DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 5, row(Fine-Tuned DNABERT-2), column(CA-AUROC-GM12878)
Configuration: DNABERT-2 (fine-tuned)Task: DART-Eval CA-AUROC-H1ESC: Chromatin activity prediction, H1ESC, positives against negatives
Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split)
0.94 auroc
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (fine-tuned) on DART-Eval CA-AUROC-H1ESC: Chromatin activity prediction, H1ESC, positives against negatives

Separating positive H1ESC peaks from matched negatives.

Aggregation: Not reported

DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 5, row(Fine-Tuned DNABERT-2), column(CA-AUROC-H1ESC)
Configuration: DNABERT-2 (fine-tuned)Task: DART-Eval CA-AUROC-HEPG2: Chromatin activity prediction, HEPG2, positives against negatives
Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split)
0.893 auroc
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (fine-tuned) on DART-Eval CA-AUROC-HEPG2: Chromatin activity prediction, HEPG2, positives against negatives

Separating positive HEPG2 peaks from matched negatives.

Aggregation: Not reported

DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 5, row(Fine-Tuned DNABERT-2), column(CA-AUROC-HEPG2)
Configuration: DNABERT-2 (fine-tuned)Task: DART-Eval CA-AUROC-IMR90: Chromatin activity prediction, IMR90, positives against negatives
Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split)
0.963 auroc
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (fine-tuned) on DART-Eval CA-AUROC-IMR90: Chromatin activity prediction, IMR90, positives against negatives

Separating positive IMR90 peaks from matched negatives.

Aggregation: Not reported

DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 5, row(Fine-Tuned DNABERT-2), column(CA-AUROC-IMR90)
Configuration: DNABERT-2 (fine-tuned)Task: DART-Eval CA-AUROC-K562: Chromatin activity prediction, K562, positives against negatives
Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split)
0.917 auroc
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (fine-tuned) on DART-Eval CA-AUROC-K562: Chromatin activity prediction, K562, positives against negatives

Separating positive K562 peaks from matched negatives.

Aggregation: Not reported

DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 5, row(Fine-Tuned DNABERT-2), column(CA-AUROC-K562)
Configuration: DNABERT-2 (fine-tuned)Task: DART-Eval CA-SPEARMAN-GM12878: Chromatin activity prediction, GM12878, positives only
Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split)
0.489 spearman_r
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (fine-tuned) on DART-Eval CA-SPEARMAN-GM12878: Chromatin activity prediction, GM12878, positives only

Rank correlation with measured accessibility among positive GM12878 peaks.

Aggregation: Not reported

DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 5, row(Fine-Tuned DNABERT-2), column(CA-SPEARMAN-GM12878)
Configuration: DNABERT-2 (fine-tuned)Task: DART-Eval CA-SPEARMAN-H1ESC: Chromatin activity prediction, H1ESC, positives only
Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split)
0.717 spearman_r
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (fine-tuned) on DART-Eval CA-SPEARMAN-H1ESC: Chromatin activity prediction, H1ESC, positives only

Rank correlation with measured accessibility among positive H1ESC peaks.

Aggregation: Not reported

DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 5, row(Fine-Tuned DNABERT-2), column(CA-SPEARMAN-H1ESC)
Configuration: DNABERT-2 (fine-tuned)Task: DART-Eval CA-SPEARMAN-HEPG2: Chromatin activity prediction, HEPG2, positives only
Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split)
0.472 spearman_r
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (fine-tuned) on DART-Eval CA-SPEARMAN-HEPG2: Chromatin activity prediction, HEPG2, positives only

Rank correlation with measured accessibility among positive HEPG2 peaks.

Aggregation: Not reported

DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 5, row(Fine-Tuned DNABERT-2), column(CA-SPEARMAN-HEPG2)
Configuration: DNABERT-2 (fine-tuned)Task: DART-Eval CA-SPEARMAN-IMR90: Chromatin activity prediction, IMR90, positives only
Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split)
0.47 spearman_r
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (fine-tuned) on DART-Eval CA-SPEARMAN-IMR90: Chromatin activity prediction, IMR90, positives only

Rank correlation with measured accessibility among positive IMR90 peaks.

Aggregation: Not reported

DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 5, row(Fine-Tuned DNABERT-2), column(CA-SPEARMAN-IMR90)
Configuration: DNABERT-2 (fine-tuned)Task: DART-Eval CA-SPEARMAN-K562: Chromatin activity prediction, K562, positives only
Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split)
0.529 spearman_r
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (fine-tuned) on DART-Eval CA-SPEARMAN-K562: Chromatin activity prediction, K562, positives only

Rank correlation with measured accessibility among positive K562 peaks.

Aggregation: Not reported

DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 5, row(Fine-Tuned DNABERT-2), column(CA-SPEARMAN-K562)
Configuration: DNABERT-2 (fine-tuned)Task: DART-Eval CTS-ACC: Cell-type-specific element classification, overall accuracy
Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split)
0.65 accuracy
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (fine-tuned) on DART-Eval CTS-ACC: Cell-type-specific element classification, overall accuracy

Classify which of five cell lines a accessible element belongs to.

Aggregation: Not reported

DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 4, row(Fine-Tuned DNABERT-2), column(CTS-ACC)
Configuration: DNABERT-2 (fine-tuned)Task: DART-Eval CTS-GM12878: Cell-type-specific element classification, GM12878
Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split)
0.894 auroc
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (fine-tuned) on DART-Eval CTS-GM12878: Cell-type-specific element classification, GM12878

One-against-rest AUROC for GM12878 accessible elements.

Aggregation: Not reported

DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 4, row(Fine-Tuned DNABERT-2), column(CTS-GM12878)
Configuration: DNABERT-2 (fine-tuned)Task: DART-Eval CTS-H1ESC: Cell-type-specific element classification, H1ESC
Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split)
0.93 auroc
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (fine-tuned) on DART-Eval CTS-H1ESC: Cell-type-specific element classification, H1ESC

One-against-rest AUROC for H1ESC accessible elements.

Aggregation: Not reported

DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 4, row(Fine-Tuned DNABERT-2), column(CTS-H1ESC)
Configuration: DNABERT-2 (fine-tuned)Task: DART-Eval CTS-HEPG2: Cell-type-specific element classification, HEPG2
Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split)
0.891 auroc
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (fine-tuned) on DART-Eval CTS-HEPG2: Cell-type-specific element classification, HEPG2

One-against-rest AUROC for HEPG2 accessible elements.

Aggregation: Not reported

DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 4, row(Fine-Tuned DNABERT-2), column(CTS-HEPG2)
Configuration: DNABERT-2 (fine-tuned)Task: DART-Eval CTS-IMR90: Cell-type-specific element classification, IMR90
Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split)
0.922 auroc
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (fine-tuned) on DART-Eval CTS-IMR90: Cell-type-specific element classification, IMR90

One-against-rest AUROC for IMR90 accessible elements.

Aggregation: Not reported

DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 4, row(Fine-Tuned DNABERT-2), column(CTS-IMR90)
Configuration: DNABERT-2 (fine-tuned)Task: DART-Eval CTS-K562: Cell-type-specific element classification, K562
Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split)
0.871 auroc
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (fine-tuned) on DART-Eval CTS-K562: Cell-type-specific element classification, K562

One-against-rest AUROC for K562 accessible elements.

Aggregation: Not reported

DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 4, row(Fine-Tuned DNABERT-2), column(CTS-K562)

Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.

Related configurations, pipelines and services

These configurations, services and pipelines use this model within their own configurations. Their results, where available, are not assigned to the underlying model.

Use this model

How it works, versions and access

Related profile: DNABERT-2. This page retains the exact record and its evaluation context.

How it works

How it works

DNABERT-2 merges recurring DNA substrings into byte-pair tokens, then processes those tokens with a masked-language-model transformer. ALiBi supplies distance-dependent attention biases, while FlashAttention changes how attention is computed. The resulting contextual embeddings need an explicit pooling rule and prediction head for a downstream task.

Sources (5)zhihan1996/DNABERT-2-117M: README.md; zhihan1996/DNABERT-2-117M: config.json; MAGICS-LAB/DNABERT_2: README.md; zhihan1996/DNABERT-2-117M: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
Versions and reproducibility

DNABERT-2-117M model card and official DNABERT_2 implementation. ALiBi permits inference beyond the pretraining sequence length, subject to attention/memory cost; this does not establish unlimited biological context or validated accuracy at arbitrary lengths.

Sources (5)zhihan1996/DNABERT-2-117M: README.md; zhihan1996/DNABERT-2-117M: config.json; MAGICS-LAB/DNABERT_2: README.md; zhihan1996/DNABERT-2-117M: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
Strengths, limitations and unresolved questions

Strengths and limitations

Strengths and considerations

Limitations and conditions

Profile review details

Follow-up review of Reference checkpoint, Context limits, Training cutoff, Further pretraining. Source locations and before/after decisions are recorded in the 23 September profile-evidence audit. Other explanatory content retains its earlier source scope. No human scientific review or independent reproduction is implied.

Stable record: catalog-model-dnabert-2

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-23. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Inputs, outputs and configuration
PropertyDescription and evidence
Model typeMasked-token DNA transformer encoder
Sources (5)zhihan1996/DNABERT-2-117M: README.md; zhihan1996/DNABERT-2-117M: config.json; MAGICS-LAB/DNABERT_2: README.md; zhihan1996/DNABERT-2-117M: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
ArchitectureBERT-style DNA encoder with byte-pair tokenization, ALiBi relative attention biases and FlashAttention; task heads and pooling are separately configured.
Sources (5)zhihan1996/DNABERT-2-117M: README.md; zhihan1996/DNABERT-2-117M: config.json; MAGICS-LAB/DNABERT_2: README.md; zhihan1996/DNABERT-2-117M: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
InputsDNA sequence tokenized with the supplied tokenizer.
Sources (5)zhihan1996/DNABERT-2-117M: README.md; zhihan1996/DNABERT-2-117M: config.json; MAGICS-LAB/DNABERT_2: README.md; zhihan1996/DNABERT-2-117M: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
OutputsToken representations and, after a specified adaptation, task predictions.
Sources (5)zhihan1996/DNABERT-2-117M: README.md; zhihan1996/DNABERT-2-117M: config.json; MAGICS-LAB/DNABERT_2: README.md; zhihan1996/DNABERT-2-117M: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
Parameters117 million for DNABERT-2-117M; family names do not establish a particular checkpoint.
Sources (5)zhihan1996/DNABERT-2-117M: README.md; zhihan1996/DNABERT-2-117M: config.json; MAGICS-LAB/DNABERT_2: README.md; zhihan1996/DNABERT-2-117M: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
Known versionsDNABERT-2-117M model card and official DNABERT_2 implementation.
Sources (5)zhihan1996/DNABERT-2-117M: README.md; zhihan1996/DNABERT-2-117M: config.json; MAGICS-LAB/DNABERT_2: README.md; zhihan1996/DNABERT-2-117M: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
Training dataThe paper describes a 32.49-billion-base corpus covering 135 species in six groups, alongside a 2.75-billion-base human corpus. Further GUE-domain pretraining is a separately reported model variant.
Sources (5)zhihan1996/DNABERT-2-117M: README.md; zhihan1996/DNABERT-2-117M: config.json; MAGICS-LAB/DNABERT_2: README.md; zhihan1996/DNABERT-2-117M: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
Training cutoffSection4.1 identifies the human and 135-species corpora but does not establish a latest-sequence deposition date. The model repository revision is not a training-data cutoff. · Not reported in inspected sources
Sourcesdnabert2 paper v2: primary artifact · Section4.1; Table11
Context limitsThe paper describes pretraining on 700-base sequences and evaluates 5,000–10,000-base GUE+ inputs after fine-tuning. This does not establish frozen-model accuracy at those lengths; tokenisation, truncation and adaptation must be specified.
Sourcesdnabert2 paper v2: primary artifact · Section5.4 Results on GUE+, printed p10; Table2 input lengths
Weights licenceThe official zhihan1996/DNABERT-2-117M checkpoint repository carries Apache-2.0 in its pinned LICENSE. This does not assign terms to a separately fitted downstream predictor.
Sources (5)zhihan1996/DNABERT-2-117M: README.md; zhihan1996/DNABERT-2-117M: config.json; MAGICS-LAB/DNABERT_2: README.md; zhihan1996/DNABERT-2-117M: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
AccessOfficial downloadable model/card and usage examples: https://huggingface.co/zhihan1996/DNABERT-2-117M
Sources (5)zhihan1996/DNABERT-2-117M: README.md; zhihan1996/DNABERT-2-117M: config.json; MAGICS-LAB/DNABERT_2: README.md; zhihan1996/DNABERT-2-117M: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
Code licenceApache-2.0
SourcesMAGICS-LAB/DNABERT_2: LICENSE · LICENSE: licence text
Reference checkpointDNABERT-2-117M: Hugging Face revision 7bce263b15377fc15361f52cfab88f8b586abda0. Its pytorch_model.bin has registry-reported SHA-256 7ff39ec77a484dd01070a41bfd6e95cdd7247bec80fe357ab43a4be33687aeba. The weight file was not downloaded for this review.
Sourcesdnabert2 release: primary artifact · sha; siblings[pytorch_model.bin].lfs.sha256
Further pretrainingThe diamond-marked DNABERT-2 variant receives additional masked-language-model training on GUE training sets. It must be distinguished from the base pretrained model in comparisons.
Sourcesdnabert2 paper v2: primary artifact · Section5.2FurtherPreTraining; Table3caption

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

92 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.
Individual claims
zhihan1996/DNABERT-2-117M: config.json

Original source ↗

DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 7bce263b15377fc15361f52cfab88f8b586abda0
Retrieved: 2026-09-16T19:46:20.874956+00:00

source checked

automated source review · 2026-09-23

Audit details

Follow-up review of Reference checkpoint, Context limits, Training cutoff, Further pretraining. Source locations and before/after decisions are recorded in the 23 September profile-evidence audit. Other explanatory content retains its earlier source scope. No human scientific review or independent reproduction is implied.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: ba9bdafaff0cc3e30556927474d4a179519a9864012bed2628e9f1bc23c84bfd

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram caption
Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.
Individual claims
dnabert2: Primary paper PDF

Original source ↗

DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2306.15006v2
Retrieved: 2026-09-16T20:04:04.777189+00:00

source checked

automated source review · 2026-09-23

Audit details

Follow-up review of Reference checkpoint, Context limits, Training cutoff, Further pretraining. Source locations and before/after decisions are recorded in the 23 September profile-evidence audit. Other explanatory content retains its earlier source scope. No human scientific review or independent reproduction is implied.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 49300acee3e4afd44bebc3de9893c3bc310d331bd4805374e0952fdfbf366f06

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram caption
Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.
Individual claims
zhihan1996/DNABERT-2-117M: README.md

Original source ↗

DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 7bce263b15377fc15361f52cfab88f8b586abda0
Retrieved: 2026-09-16T19:46:20.874956+00:00

source checked

automated source review · 2026-09-23

Audit details

Follow-up review of Reference checkpoint, Context limits, Training cutoff, Further pretraining. Source locations and before/after decisions are recorded in the 23 September profile-evidence audit. Other explanatory content retains its earlier source scope. No human scientific review or independent reproduction is implied.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 48b18abd051eb4952e0c0a50a0740387a1cea145be06517b67ab73a29431a3bc

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram caption
Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.
Individual claims
MAGICS-LAB/DNABERT_2: README.md

Original source ↗

DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: f25bed9ee20db966dff39e5c1571249d04e36404
Retrieved: 2026-09-16T19:46:17.892989+00:00

source checked

automated source review · 2026-09-23

Audit details

Follow-up review of Reference checkpoint, Context limits, Training cutoff, Further pretraining. Source locations and before/after decisions are recorded in the 23 September profile-evidence audit. Other explanatory content retains its earlier source scope. No human scientific review or independent reproduction is implied.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 734a8cec5f667d74d421bf3b273ad7e256216109636da45aa7ceba21cd34de16

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram caption
Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.
Individual claims
zhihan1996/DNABERT-2-117M: LICENSE

Original source ↗

DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 7bce263b15377fc15361f52cfab88f8b586abda0
Retrieved: 2026-09-16T19:46:20.874956+00:00

source checked

automated source review · 2026-09-23

Audit details

Follow-up review of Reference checkpoint, Context limits, Training cutoff, Further pretraining. Source locations and before/after decisions are recorded in the 23 September profile-evidence audit. Other explanatory content retains its earlier source scope. No human scientific review or independent reproduction is implied.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: c71d239df91726fc519c6eb72d318ec65820627232b2f796219e87dcf35d0ab4

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram steps
  • DNA sequence
  • BPE tokens
  • Transformer encoder
  • Representations
  • Specified task head
Individual claims
zhihan1996/DNABERT-2-117M: config.json

Original source ↗

DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 7bce263b15377fc15361f52cfab88f8b586abda0
Retrieved: 2026-09-16T19:46:20.874956+00:00

source checked

automated source review · 2026-09-23

Audit details

Follow-up review of Reference checkpoint, Context limits, Training cutoff, Further pretraining. Source locations and before/after decisions are recorded in the 23 September profile-evidence audit. Other explanatory content retains its earlier source scope. No human scientific review or independent reproduction is implied.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: ba9bdafaff0cc3e30556927474d4a179519a9864012bed2628e9f1bc23c84bfd

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram steps
  • DNA sequence
  • BPE tokens
  • Transformer encoder
  • Representations
  • Specified task head
Individual claims
dnabert2: Primary paper PDF

Original source ↗

DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2306.15006v2
Retrieved: 2026-09-16T20:04:04.777189+00:00

source checked

automated source review · 2026-09-23

Audit details

Follow-up review of Reference checkpoint, Context limits, Training cutoff, Further pretraining. Source locations and before/after decisions are recorded in the 23 September profile-evidence audit. Other explanatory content retains its earlier source scope. No human scientific review or independent reproduction is implied.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 49300acee3e4afd44bebc3de9893c3bc310d331bd4805374e0952fdfbf366f06

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram steps
  • DNA sequence
  • BPE tokens
  • Transformer encoder
  • Representations
  • Specified task head
Individual claims
zhihan1996/DNABERT-2-117M: README.md

Original source ↗

DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 7bce263b15377fc15361f52cfab88f8b586abda0
Retrieved: 2026-09-16T19:46:20.874956+00:00

source checked

automated source review · 2026-09-23

Audit details

Follow-up review of Reference checkpoint, Context limits, Training cutoff, Further pretraining. Source locations and before/after decisions are recorded in the 23 September profile-evidence audit. Other explanatory content retains its earlier source scope. No human scientific review or independent reproduction is implied.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 48b18abd051eb4952e0c0a50a0740387a1cea145be06517b67ab73a29431a3bc

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram steps
  • DNA sequence
  • BPE tokens
  • Transformer encoder
  • Representations
  • Specified task head
Individual claims
MAGICS-LAB/DNABERT_2: README.md

Original source ↗

DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: f25bed9ee20db966dff39e5c1571249d04e36404
Retrieved: 2026-09-16T19:46:17.892989+00:00

source checked

automated source review · 2026-09-23

Audit details

Follow-up review of Reference checkpoint, Context limits, Training cutoff, Further pretraining. Source locations and before/after decisions are recorded in the 23 September profile-evidence audit. Other explanatory content retains its earlier source scope. No human scientific review or independent reproduction is implied.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 734a8cec5f667d74d421bf3b273ad7e256216109636da45aa7ceba21cd34de16

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram steps
  • DNA sequence
  • BPE tokens
  • Transformer encoder
  • Representations
  • Specified task head
Individual claims
zhihan1996/DNABERT-2-117M: LICENSE

Original source ↗

DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 7bce263b15377fc15361f52cfab88f8b586abda0
Retrieved: 2026-09-16T19:46:20.874956+00:00

source checked

automated source review · 2026-09-23

Audit details

Follow-up review of Reference checkpoint, Context limits, Training cutoff, Further pretraining. Source locations and before/after decisions are recorded in the 23 September profile-evidence audit. Other explanatory content retains its earlier source scope. No human scientific review or independent reproduction is implied.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: c71d239df91726fc519c6eb72d318ec65820627232b2f796219e87dcf35d0ab4

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: discovered

9 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: catalog-model-dnabert-2

areas
dna-genomes
method types
foundation model
entity level
family
version
117M
reported name
DNABERT-2
access
Public checkpoint; remote model code needs review before local use.
method type
foundation model
historical missing metadata
checkpoint revision: not_yet_extracted; training data: not_yet_extracted; licence: not_yet_extracted
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
entity classification
review date: 2026-09-17; rationale: The cited profile describes a named learned biological predictor or representation model/family. Preserve this identity separately from task-specific fitting, individual checkpoints, pipelines and hosted access.; source ids: evidence-official-7267eabca7c4a7945878; evidence-official-2cf0a41fec83ee9c5cc9; evidence-official-96b5c3a31a50f7c259d6; evidence-official-a48ae27aa000bdab7442; evidence-official-3aa749d3205c0a768b22; source locator: DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0; ambiguities: None recorded
Related records

Suggest a correction