Model type
Masked-token DNA transformer encoder
DNABERT-2 learns DNA representations that can be adapted to genomic prediction tasks.
Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.
Masked-token DNA transformer encoder
DNA sequence tokenized with the supplied tokenizer.
Token representations and, after a specified adaptation, task predictions.
Official project documentation and implementation: https://github.com/MAGICS-LAB/DNABERT_2
limited source coverage · Automated source review, 2026-09-23. All specifications and missing details
28 evaluations · 28 results. Different protocols are not a single leaderboard.
Applied filters: All linked evaluations
| Tested configuration | Protocol and dataset | Finding | Evidence and details |
|---|---|---|---|
| Configuration: DNABERT-2 (further pre-trained on GUE) | Task: GUE CORE-PROMOTER-DETECTION-ALL: Core promoter detection, dataset all Dataset subset: GUE Core promoter detection, all (GUE split) | 67.5% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Core promoter detection all) |
| Configuration: DNABERT-2 (further pre-trained on GUE) | Task: GUE CORE-PROMOTER-DETECTION-NOTATA: Core promoter detection, dataset notata Dataset subset: GUE Core promoter detection, notata (GUE split) | 69.5% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Core promoter detection notata) |
| Configuration: DNABERT-2 (further pre-trained on GUE) | Task: GUE CORE-PROMOTER-DETECTION-TATA: Core promoter detection, dataset tata Dataset subset: GUE Core promoter detection, tata (GUE split) | 76.2% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Core promoter detection tata) |
| Configuration: DNABERT-2 (further pre-trained on GUE) | Task: GUE COVID-VARIANT-CLASSIFICATION-COVID: Covid variant classification, dataset Covid Dataset subset: GUE Covid variant classification, Covid (GUE split) | 68.5% f1 percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Covid variant classification Covid) |
| Configuration: DNABERT-2 (further pre-trained on GUE) | Task: GUE EPIGENETIC-MARKS-PREDICTION-H3: Epigenetic marks prediction, dataset H3 Dataset subset: GUE Epigenetic marks prediction, H3 (GUE split) | 80.2% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Epigenetic marks prediction H3) |
| Configuration: DNABERT-2 (further pre-trained on GUE) | Task: GUE EPIGENETIC-MARKS-PREDICTION-H3K14AC: Epigenetic marks prediction, dataset H3K14ac Dataset subset: GUE Epigenetic marks prediction, H3K14ac (GUE split) | 57.4% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Epigenetic marks prediction H3K14ac) |
| Configuration: DNABERT-2 (further pre-trained on GUE) | Task: GUE EPIGENETIC-MARKS-PREDICTION-H3K36ME3: Epigenetic marks prediction, dataset H3K36me3 Dataset subset: GUE Epigenetic marks prediction, H3K36me3 (GUE split) | 61.9% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Epigenetic marks prediction H3K36me3) |
| Configuration: DNABERT-2 (further pre-trained on GUE) | Task: GUE EPIGENETIC-MARKS-PREDICTION-H3K4ME1: Epigenetic marks prediction, dataset H3K4me1 Dataset subset: GUE Epigenetic marks prediction, H3K4me1 (GUE split) | 53% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Epigenetic marks prediction H3K4me1) |
| Configuration: DNABERT-2 (further pre-trained on GUE) | Task: GUE EPIGENETIC-MARKS-PREDICTION-H3K4ME2: Epigenetic marks prediction, dataset H3K4me2 Dataset subset: GUE Epigenetic marks prediction, H3K4me2 (GUE split) | 39.9% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Epigenetic marks prediction H3K4me2) |
| Configuration: DNABERT-2 (further pre-trained on GUE) | Task: GUE EPIGENETIC-MARKS-PREDICTION-H3K4ME3: Epigenetic marks prediction, dataset H3K4me3 Dataset subset: GUE Epigenetic marks prediction, H3K4me3 (GUE split) | 41.2% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Epigenetic marks prediction H3K4me3) |
| Configuration: DNABERT-2 (further pre-trained on GUE) | Task: GUE EPIGENETIC-MARKS-PREDICTION-H3K79ME3: Epigenetic marks prediction, dataset H3K79me3 Dataset subset: GUE Epigenetic marks prediction, H3K79me3 (GUE split) | 65.5% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Epigenetic marks prediction H3K79me3) |
| Configuration: DNABERT-2 (further pre-trained on GUE) | Task: GUE EPIGENETIC-MARKS-PREDICTION-H3K9AC: Epigenetic marks prediction, dataset H3K9ac Dataset subset: GUE Epigenetic marks prediction, H3K9ac (GUE split) | 57.1% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Epigenetic marks prediction H3K9ac) |
| Configuration: DNABERT-2 (further pre-trained on GUE) | Task: GUE EPIGENETIC-MARKS-PREDICTION-H4: Epigenetic marks prediction, dataset H4 Dataset subset: GUE Epigenetic marks prediction, H4 (GUE split) | 81.9% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Epigenetic marks prediction H4) |
| Configuration: DNABERT-2 (further pre-trained on GUE) | Task: GUE EPIGENETIC-MARKS-PREDICTION-H4AC: Epigenetic marks prediction, dataset H4ac Dataset subset: GUE Epigenetic marks prediction, H4ac (GUE split) | 50.4% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Epigenetic marks prediction H4ac) |
| Configuration: DNABERT-2 (further pre-trained on GUE) | Task: GUE PROMOTER-DETECTION-ALL: Promoter detection, dataset all Dataset subset: GUE Promoter detection, all (GUE split) | 88.3% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Promoter detection all) |
| Configuration: DNABERT-2 (further pre-trained on GUE) | Task: GUE PROMOTER-DETECTION-NOTATA: Promoter detection, dataset notata Dataset subset: GUE Promoter detection, notata (GUE split) | 94.3% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Promoter detection notata) |
| Configuration: DNABERT-2 (further pre-trained on GUE) | Task: GUE PROMOTER-DETECTION-TATA: Promoter detection, dataset tata Dataset subset: GUE Promoter detection, tata (GUE split) | 68.8% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Promoter detection tata) |
| Configuration: DNABERT-2 (further pre-trained on GUE) | Task: GUE SPLICE-SITE-PREDICTION-RECONSTRUCT: Splice site prediction, dataset Reconstruct Dataset subset: GUE Splice site prediction, Reconstruct (GUE split) | 85.9% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Splice site prediction Reconstruct) |
| Configuration: DNABERT-2 (further pre-trained on GUE) | Task: GUE TRANSCRIPTION-FACTOR-PREDICTION-HUMAN-0: Transcription factor prediction (human), dataset 0 Dataset subset: GUE Transcription factor prediction (human), 0 (GUE split) | 69.1% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Transcription factor prediction (human) 0) |
| Configuration: DNABERT-2 (further pre-trained on GUE) | Task: GUE TRANSCRIPTION-FACTOR-PREDICTION-HUMAN-1: Transcription factor prediction (human), dataset 1 Dataset subset: GUE Transcription factor prediction (human), 1 (GUE split) | 71.9% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Transcription factor prediction (human) 1) |
| Configuration: DNABERT-2 (further pre-trained on GUE) | Task: GUE TRANSCRIPTION-FACTOR-PREDICTION-HUMAN-2: Transcription factor prediction (human), dataset 2 Dataset subset: GUE Transcription factor prediction (human), 2 (GUE split) | 63% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Transcription factor prediction (human) 2) |
| Configuration: DNABERT-2 (further pre-trained on GUE) | Task: GUE TRANSCRIPTION-FACTOR-PREDICTION-HUMAN-3: Transcription factor prediction (human), dataset 3 Dataset subset: GUE Transcription factor prediction (human), 3 (GUE split) | 55.4% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Transcription factor prediction (human) 3) |
| Configuration: DNABERT-2 (further pre-trained on GUE) | Task: GUE TRANSCRIPTION-FACTOR-PREDICTION-HUMAN-4: Transcription factor prediction (human), dataset 4 Dataset subset: GUE Transcription factor prediction (human), 4 (GUE split) | 74.9% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Transcription factor prediction (human) 4) |
| Configuration: DNABERT-2 (further pre-trained on GUE) | Task: GUE TRANSCRIPTION-FACTOR-PREDICTION-MOUSE-0: Transcription factor prediction (mouse), dataset 0 Dataset subset: GUE Transcription factor prediction (mouse), 0 (GUE split) | 64.2% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Transcription factor prediction (mouse) 0) |
| Configuration: DNABERT-2 (further pre-trained on GUE) | Task: GUE TRANSCRIPTION-FACTOR-PREDICTION-MOUSE-1: Transcription factor prediction (mouse), dataset 1 Dataset subset: GUE Transcription factor prediction (mouse), 1 (GUE split) | 86.3% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(DNABERT-2♦), column(Transcription factor prediction (mouse) 1) |
Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.
Related profile: DNABERT-2. This page retains the exact record and its evaluation context.
Genome language model fine-tuned on each GUE dataset by the DNABERT-2 authors.
DNABERT-2 merges recurring DNA substrings into byte-pair tokens, then processes those tokens with a masked-language-model transformer. ALiBi supplies distance-dependent attention biases, while FlashAttention changes how attention is computed. The resulting contextual embeddings need an explicit pooling rule and prediction head for a downstream task.
DNABERT-2-117M model card and official DNABERT_2 implementation. ALiBi permits inference beyond the pretraining sequence length, subject to attention/memory cost; this does not establish unlimited biological context or validated accuracy at arbitrary lengths.
Follow-up review of Reference checkpoint, Context limits, Training cutoff, Further pretraining. Source locations and before/after decisions are recorded in the 23 September profile-evidence audit. Other explanatory content retains its earlier source scope. No human scientific review or independent reproduction is implied.
Stable record: discovery-model-dnabert-2Explanatory profile: limited source coverage · Automated source review, 2026-09-23. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Model type | Masked-token DNA transformer encoderSources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0 |
| Architecture | BERT-style DNA encoder with byte-pair tokenization, ALiBi relative attention biases and FlashAttention; task heads and pooling are separately configured.Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0 |
| Inputs | DNA sequence tokenized with the supplied tokenizer.Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0 |
| Outputs | Token representations and, after a specified adaptation, task predictions.Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0 |
| Parameters | 117 million for DNABERT-2-117M; family names do not establish a particular checkpoint.Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0 |
| Known versions | DNABERT-2-117M model card and official DNABERT_2 implementation.Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0 |
| Training data | The paper describes a 32.49-billion-base corpus covering 135 species in six groups, alongside a 2.75-billion-base human corpus. Further GUE-domain pretraining is a separately reported model variant.Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0 |
| Training cutoff | Section4.1 identifies the human and 135-species corpora but does not establish a latest-sequence deposition date. The model repository revision is not a training-data cutoff. · Not reported in inspected sourcesSourcesdnabert2 paper v2: primary artifact · Section4.1; Table11 |
| Context limits | The paper describes pretraining on 700-base sequences and evaluates 5,000–10,000-base GUE+ inputs after fine-tuning. This does not establish frozen-model accuracy at those lengths; tokenisation, truncation and adaptation must be specified.Sourcesdnabert2 paper v2: primary artifact · Section5.4 Results on GUE+, printed p10; Table2 input lengths |
| Weights licence | The official zhihan1996/DNABERT-2-117M checkpoint repository carries Apache-2.0 in its pinned LICENSE. This does not assign terms to a separately fitted downstream predictor.Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0 |
| Access | Official project documentation and implementation: https://github.com/MAGICS-LAB/DNABERT_2Sources (3)MAGICS-LAB/DNABERT_2: README.md; MAGICS-LAB/DNABERT_2: LICENSE; dnabert2: Primary paper PDF · DNABERT-2 paper Sections 3.2, 4.1 and 5.2 (Further Pre-Training); official model card and implementation README; official DNABERT-2-117M checkpoint LICENSE at revision 7bce263b15377fc15361f52cfab88f8b586abda0 |
| Code licence | Apache-2.0SourcesMAGICS-LAB/DNABERT_2: LICENSE · LICENSE: licence text |
| Reference checkpoint | DNABERT-2-117M: Hugging Face revision 7bce263b15377fc15361f52cfab88f8b586abda0. Its pytorch_model.bin has registry-reported SHA-256 7ff39ec77a484dd01070a41bfd6e95cdd7247bec80fe357ab43a4be33687aeba. The weight file was not downloaded for this review.Sourcesdnabert2 release: primary artifact · sha; siblings[pytorch_model.bin].lfs.sha256 |
| Further pretraining | The diamond-marked DNABERT-2 variant receives additional masked-language-model training on GUE training sets. It must be distinguished from the base pretrained model in comparisons.Sourcesdnabert2 paper v2: primary artifact · Section5.2FurtherPreTraining; Table3caption |
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
2 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Relationship: family discovery-model-dnabert-2 Individual claims | DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes DNABERT-2 paper GUE result tables; further pre-training setting retained; official README GUE benchmark; source-labelled configuration DNABERT-2 (further pre-trained on GUE) | Existing reviewed locator: Table 6, row(DNABERT-2♦) Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: Primary full-text snapshot retrieved 2026-09-17; exact bytes pinned by SHA-256 | source checked automated source review · 2026-09-23 Audit detailsSource review establishes this relationship only. Exact evaluated configurations and original numerical review status remain unchanged. Both configurations are explicitly DNABERT-2; retain further pretraining as a distinct configuration. Field: Claim: model-evaluation-identity-642226226b4179cf9b07 Source artifact SHA-256: Hash scope: Exact retrieved primary paper artifact bytes. |
| Relationship: family discovery-model-dnabert-2 Individual claims | MAGICS-LAB/DNABERT_2: README.md DNABERT-2 paper GUE result tables; further pre-training setting retained; official README GUE benchmark; source-labelled configuration DNABERT-2 (further pre-trained on GUE) | Existing reviewed locator: Table 6, row(DNABERT-2♦) Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: f25bed9ee20db966dff39e5c1571249d04e36404 | source checked automated source review · 2026-09-23 Audit detailsSource review establishes this relationship only. Exact evaluated configurations and original numerical review status remain unchanged. Both configurations are explicitly DNABERT-2; retain further pretraining as a distinct configuration. Field: Claim: model-evaluation-identity-642226226b4179cf9b07 Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
View linked audit checks and correction history
Release 2026-09-29-06401fd5b220 · Record review: source checked
Stable ID: gue-method-dnabert-2-further-pre-trained-on-gue