rewire.itbenchmarks
Configuration

DNABERT-2 (zero-shot)

DNABERT-2 learns DNA representations that can be adapted to genomic prediction tasks.

Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0

1 evaluation · 1 result

How it worksDNABERT-2 workflow
DNABERT-2 workflow1. DNA sequence. Then: 2. BPE tokens. Then: 3. Transformer encoder. Then: 4. Representations. Then: 5. Specified task headDNABERT-2 workflow1. DNA sequence. Then: 2. BPE tokens. Then: 3. Transformer encoder. Then: 4. Representations. Then: 5. Specified task headDNABERT-2 workflow1. DNA sequence. Then: 2. BPE tokens. Then: 3. Transformer encoder. Then: 4. Representations. Then: 5. Specified task head

Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.

Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0

Overview

Model type

Masked-token DNA transformer encoder

Inputs

DNA sequence tokenized with the supplied tokenizer.

Outputs

Token representations and, after a specified adaptation, task predictions.

Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0

limited source coverage · Automated source review, 2026-09-23. All specifications and missing details

Evaluations and results

1 evaluation · 1 result. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: DNABERT-2 (zero-shot)Task: DART-Eval REI-ACC: Regulatory element identification, zero-shot accuracy
Dataset subset: ENCODE cCREs against dinucleotide-shuffled backgrounds (DART-Eval split)
0.876 accuracy
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

DNABERT-2 (zero-shot) on DART-Eval REI-ACC: Regulatory element identification, zero-shot accuracy

Distinguish ENCODE cCREs from dinucleotide-shuffled background sequences, scored zero-shot from model likelihood.

Aggregation: Not reported

DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 3, row(DNABERT-2), column(zero-shot accuracy)

Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.

Use this model

How it works, versions and access

Related profile: DNABERT-2. This page retains the exact record and its evaluation context.

This configuration

DNA language model evaluated by the DART-Eval authors in the zero-shot setting.

record
DNABERT-2 (zero-shot)
configuration
Not reported
entity type
Configuration

How it works

How it works

DNABERT-2 merges recurring DNA substrings into byte-pair tokens, then processes those tokens with a masked-language-model transformer. ALiBi supplies distance-dependent attention biases, while FlashAttention changes how attention is computed. The resulting contextual embeddings need an explicit pooling rule and prediction head for a downstream task.

Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
Versions and reproducibility

DNABERT-2-117M model card and official DNABERT_2 implementation. ALiBi permits inference beyond the pretraining sequence length, subject to attention/memory cost; this does not establish unlimited biological context or validated accuracy at arbitrary lengths.

Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
Strengths, limitations and unresolved questions

Strengths and limitations

Limitations and conditions

  • An embedding model alone is not the same evaluated pipeline as frozen embeddings followed by logistic regression. Tokenization and pooling choices must be preserved.
    Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
Profile review details

Follow-up review of Reference checkpoint, Context limits, Training cutoff, Further pretraining. Source locations and before/after decisions are recorded in the 23 September profile-evidence audit. Other explanatory content retains its earlier source scope. No human scientific review or independent reproduction is implied.

Stable record: discovery-model-dnabert-2

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-23. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Inputs, outputs and configuration
PropertyDescription and evidence
Model typeMasked-token DNA transformer encoder
Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
ArchitectureBERT-style DNA encoder with byte-pair tokenization, ALiBi relative attention biases and FlashAttention; task heads and pooling are separately configured.
Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
InputsDNA sequence tokenized with the supplied tokenizer.
Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
OutputsToken representations and, after a specified adaptation, task predictions.
Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
Parameters117 million for DNABERT-2-117M; family names do not establish a particular checkpoint.
Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
Known versionsDNABERT-2-117M model card and official DNABERT_2 implementation.
Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
Training dataThe paper describes a 32.49-billion-base corpus covering 135 species in six groups, alongside a 2.75-billion-base human corpus. Further GUE-domain pretraining is a separately reported model variant.
Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
Training cutoffSection4.1 identifies the human and 135-species corpora but does not establish a latest-sequence deposition date. The model repository revision is not a training-data cutoff. · Not reported in inspected sources
Sourcesdnabert2 paper v2: primary artifact · Section4.1; Table11
Context limitsThe paper describes pretraining on 700-base sequences and evaluates 5,000–10,000-base GUE+ inputs after fine-tuning. This does not establish frozen-model accuracy at those lengths; tokenisation, truncation and adaptation must be specified.
Sourcesdnabert2 paper v2: primary artifact · Section5.4 Results on GUE+, printed p10; Table2 input lengths
Weights licenceThe official zhihan1996/DNABERT-2-117M checkpoint repository carries Apache-2.0 in its pinned LICENSE. This does not assign terms to a separately fitted downstream predictor.
Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
AccessOfficial project documentation and implementation: https://github.com/MAGICS-LAB/DNABERT_2
Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0
Code licenceApache-2.0
SourcesMAGICS-LAB/DNABERT_2: LICENSE · LICENSE: licence text
Reference checkpointDNABERT-2-117M: Hugging Face revision 7bce263b15377fc15361f52cfab88f8b586abda0. Its pytorch_model.bin has registry-reported SHA-256 7ff39ec77a484dd01070a41bfd6e95cdd7247bec80fe357ab43a4be33687aeba. The weight file was not downloaded for this review.
Sourcesdnabert2 release: primary artifact · sha; siblings[pytorch_model.bin].lfs.sha256
Further pretrainingThe diamond-marked DNABERT-2 variant receives additional masked-language-model training on GUE training sets. It must be distinguished from the base pretrained model in comparisons.
Sourcesdnabert2 paper v2: primary artifact · Section5.2FurtherPreTraining; Table3caption

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

2 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Relationship: family
discovery-model-dnabert-2
Individual claims
DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA

Original source ↗

DART-Eval Table 1 (DNABERT-2, 117M) and adaptation-specific result tables; source-labelled configuration DNABERT-2 (zero-shot) | Existing reviewed locator: Table 1, row(DNABERT-2)

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2412.05430v1
Retrieved: 2026-09-17T07:56:09.182117+00:00

source checked

automated source review · 2026-09-23

Audit details

Source review establishes this relationship only. Exact evaluated configurations and original numerical review status remain unchanged. Official paper fixes backbone identity and separates fine-tuning/probing/zero-shot conditions. Family link does not pool scores.

Field: links:family:discovery-model-dnabert-2

Claim: model-evaluation-identity-44809af7f62ab9cbf95b

Source artifact SHA-256: 4194b137ba55c9a2c269d119a9afec6ae1bb0feaf17d91433ae483c41221a56b

Hash scope: Exact retrieved primary paper artifact bytes.

Inspected artifact

Relationship: family
discovery-model-dnabert-2
Individual claims
MAGICS-LAB/DNABERT_2: README.md

Original source ↗

DART-Eval Table 1 (DNABERT-2, 117M) and adaptation-specific result tables; source-labelled configuration DNABERT-2 (zero-shot) | Existing reviewed locator: Table 1, row(DNABERT-2)

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: f25bed9ee20db966dff39e5c1571249d04e36404
Retrieved: 2026-09-16T19:46:17.892989+00:00

source checked

automated source review · 2026-09-23

Audit details

Source review establishes this relationship only. Exact evaluated configurations and original numerical review status remain unchanged. Official paper fixes backbone identity and separates fine-tuning/probing/zero-shot conditions. Family link does not pool scores.

Field: links:family:discovery-model-dnabert-2

Claim: model-evaluation-identity-44809af7f62ab9cbf95b

Source artifact SHA-256: 734a8cec5f667d74d421bf3b273ad7e256216109636da45aa7ceba21cd34de16

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: source checked

1 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: dart-eval-method-dnabert-2-zero-shot

areas
dna-genomes
source locator
Table 1, row(DNABERT-2)
missing metadata
checkpoint revision: unreported; parameters: unextracted
Related records

Suggest a correction