rewire.itbenchmarks
Configuration

DNABERT-2

This genomic model is adapted to predict G-quadruplex-associated sequence regions.

SourcesBenchmarking DNA large language models on quadruplexes · Discussion and conclusions (paragraph 6); Results/LLM performance at the genome-wide level (paragraph 4)

1 evaluation · 1 result

How it worksEvaluated procedure (conceptual)
Evaluated procedure (conceptual)1. DNA sequences for G-quadruplex classification and genomic scanning. Then: 2. DNABERT-2. Then: 3. G4-associated sequence predictionsEvaluated procedure (conceptual)1. DNA sequences for G-quadruplex classification and genomic scanning. Then: 2. DNABERT-2. Then: 3. G4-associated sequence predictionsEvaluated procedure (conceptual)1. DNA sequences for G-quadruplex classification and genomic scanning. Then: 2. DNABERT-2. Then: 3. G4-associated sequence predictions

Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings.

SourcesBenchmarking DNA large language models on quadruplexes · Materials and methods/Data preparation (paragraph 2); Results/Interpretation of LLMs (paragraph 2)

Overview

Model type

DNA sequence transformer; this record is the paper-specific evaluated configuration.

SourcesMAGICS-LAB/DNABERT_2 README.md · README.md model description

limited source coverage · Automated source review, 2026-09-16. All specifications and missing details

Evaluations and results

1 evaluation · 1 result. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: DNABERT-2Task: G-quadruplex classification
Dataset: KEx
97 Accuracy
% · unknown

Uncertainty: ± 0.5

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

DNABERT-2: G-quadruplex classification

Pretrained model evaluated on KEx as reported in Table 5.

Aggregation: Not reported

Benchmarking DNA large language models on quadruplexes · Table 5, DNABERT-2 (117 M) row, Accuracy column

Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.

Use this model

How it works, versions and access

Related profile: DNABERT-2. This page retains the exact record and its evaluation context.

How it works

How the evaluated method works

The study fine-tunes genomic sequence models for G4 classification and applies them to whole-genome annotation, retaining each model’s native tokenisation.

SourcesBenchmarking DNA large language models on quadruplexes · Materials and methods/Data preparation (paragraph 2); Results/Interpretation of LLMs (paragraph 2)
Underlying method and version boundaries

DNABERT-2 replaces overlapping k-mer tokens with byte-pair encoding and uses ALiBi positional biases. The official 117M model produces 768-dimensional token representations; downstream classifiers and pooling choices are separate configuration details.

SourcesMAGICS-LAB/DNABERT_2 README.md · README.md; introduction, model description, pretrained-model and usage sections at pinned revision
What was evaluated

The linked evaluation record identifies DNABERT-2: G-quadruplex classification. Its dataset, split, adaptation and evidence origin remain attached to the reported results.

SourcesBenchmarking DNA large language models on quadruplexes · The named evaluation’s methods and comparison table; exact preserved evaluation IDs: evaluation-lit-005
Strengths, limitations and unresolved questions

Strengths and limitations

Limitations and conditions

Profile review details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Stable record: reported-model-ade36035f58f27

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Inputs, outputs and configuration
PropertyDescription and evidence
Model typeDNA sequence transformer; this record is the paper-specific evaluated configuration.
SourcesMAGICS-LAB/DNABERT_2 README.md · README.md model description
Architecture / procedureThe study fine-tunes genomic sequence models for G4 classification and applies them to whole-genome annotation, retaining each model’s native tokenisation.
SourcesBenchmarking DNA large language models on quadruplexes · Materials and methods/Data preparation (paragraph 2); Results/Interpretation of LLMs (paragraph 2)
Biological inputsDNA sequences for G-quadruplex classification and genomic scanning
SourcesBenchmarking DNA large language models on quadruplexes · Results/LLM comparison on quadruplex datasets (paragraph 1); Materials and methods/Data preparation (paragraph 1)
OutputsG4-associated sequence predictions
SourcesBenchmarking DNA large language models on quadruplexes · Introduction (paragraph 3); CRediT authorship contribution statement (paragraph 1)
Parameters117 million parameters, as identified for this row
SourcesBenchmarking DNA large language models on quadruplexes · Materials and methods/Low Rank Adaptation (LoRA) (paragraph 1); Discussion and conclusions (paragraph 3)
Known versions / configuration117M
SourcesBenchmarking DNA large language models on quadruplexes · Table tbl0030 (paragraph 1); Table tbl0025 (paragraph 1)
Training data / fittingTask-specific fine-tuning for approximately four to ten epochs, depending on model performance, with gradient accumulation.
SourcesBenchmarking DNA large language models on quadruplexes · Results/LLM comparison on quadruplex datasets (paragraph 5); Materials and methods/Fine-tuning (paragraph 1)
Context limitsA maximum input/context length for this exact evaluated configuration is not established by the inspected sources. · Not reported in inspected sources
Sources (2)Benchmarking DNA large language models on quadruplexes; MAGICS-LAB/DNABERT_2 README.md · Materials and methods/Data preparation; Materials and methods/Tokenization for G4s; Materials and methods/Metrics of evaluation; Materials and methods/Fine-tuning; Materials and methods/Low Rank Adaptation (LoRA); inspected for explicit maximum input length (dataset lengths and family-wide limits are not substituted); README.md at pinned repository revision
AccessOfficial upstream implementation and usage documentation: https://github.com/MAGICS-LAB/DNABERT_2/blob/f25bed9ee20db966dff39e5c1571249d04e36404/README.md. This pinned documentation revision is not automatically the evaluated weight revision.
SourcesMAGICS-LAB/DNABERT_2 README.md · README.md; installation, model download and usage instructions
Code licenceApache 2.0 (upstream repository code at the cited revision; this does not establish every dependency or historical checkpoint licence).
SourcesMAGICS-LAB/DNABERT_2 LICENSE · LICENSE; complete licence text
Weights licenceThe inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately. · Not reported in inspected sources
SourcesMAGICS-LAB/DNABERT_2 README.md · README.md; checkpoint/access documentation and licence scope

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

24 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings.
Individual claims
Benchmarking DNA large language models on quadruplexes

Original source ↗

Materials and methods/Data preparation (paragraph 2); Results/Interpretation of LLMs (paragraph 2)

Version: version of record
Retrieved: 2026-09-16T10:33:35.379Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: c3d7c6d068d3c11a9c8255a932197ece3804d94e8b2d4f858bea373a1b6eb32f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps
  • DNA sequences for G-quadruplex classification and genomic scanning
  • DNABERT-2
  • G4-associated sequence predictions
Individual claims
Benchmarking DNA large language models on quadruplexes

Original source ↗

Materials and methods/Data preparation (paragraph 2); Results/Interpretation of LLMs (paragraph 2)

Version: version of record
Retrieved: 2026-09-16T10:33:35.379Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: c3d7c6d068d3c11a9c8255a932197ece3804d94e8b2d4f858bea373a1b6eb32f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title
Evaluated procedure (conceptual)
Individual claims
Benchmarking DNA large language models on quadruplexes

Original source ↗

Materials and methods/Data preparation (paragraph 2); Results/Interpretation of LLMs (paragraph 2)

Version: version of record
Retrieved: 2026-09-16T10:33:35.379Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: c3d7c6d068d3c11a9c8255a932197ece3804d94e8b2d4f858bea373a1b6eb32f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Model type
DNA sequence transformer; this record is the paper-specific evaluated configuration.
Individual claims
MAGICS-LAB/DNABERT_2 README.md

Original source ↗

README.md model description

Version: f25bed9ee20db966dff39e5c1571249d04e36404
Retrieved: 2026-09-16T20:00:02.624261+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: 734a8cec5f667d74d421bf3b273ad7e256216109636da45aa7ceba21cd34de16

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Architecture / procedure
The study fine-tunes genomic sequence models for G4 classification and applies them to whole-genome annotation, retaining each model’s native tokenisation.
Individual claims
Benchmarking DNA large language models on quadruplexes

Original source ↗

Materials and methods/Data preparation (paragraph 2); Results/Interpretation of LLMs (paragraph 2)

Version: version of record
Retrieved: 2026-09-16T10:33:35.379Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: c3d7c6d068d3c11a9c8255a932197ece3804d94e8b2d4f858bea373a1b6eb32f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Weights licence
The inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately.
Individual claims
MAGICS-LAB/DNABERT_2 README.md

Original source ↗

README.md; checkpoint/access documentation and licence scope

Version: f25bed9ee20db966dff39e5c1571249d04e36404
Retrieved: 2026-09-16T20:00:02.624261+00:00

unreported

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: 734a8cec5f667d74d421bf3b273ad7e256216109636da45aa7ceba21cd34de16

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Biological inputs
DNA sequences for G-quadruplex classification and genomic scanning
Individual claims
Benchmarking DNA large language models on quadruplexes

Original source ↗

Results/LLM comparison on quadruplex datasets (paragraph 1); Materials and methods/Data preparation (paragraph 1)

Version: version of record
Retrieved: 2026-09-16T10:33:35.379Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: c3d7c6d068d3c11a9c8255a932197ece3804d94e8b2d4f858bea373a1b6eb32f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Outputs
G4-associated sequence predictions
Individual claims
Benchmarking DNA large language models on quadruplexes

Original source ↗

Introduction (paragraph 3); CRediT authorship contribution statement (paragraph 1)

Version: version of record
Retrieved: 2026-09-16T10:33:35.379Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.3.value

Source artifact SHA-256: c3d7c6d068d3c11a9c8255a932197ece3804d94e8b2d4f858bea373a1b6eb32f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Parameters
117 million parameters, as identified for this row
Individual claims
Benchmarking DNA large language models on quadruplexes

Original source ↗

Materials and methods/Low Rank Adaptation (LoRA) (paragraph 1); Discussion and conclusions (paragraph 3)

Version: version of record
Retrieved: 2026-09-16T10:33:35.379Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: c3d7c6d068d3c11a9c8255a932197ece3804d94e8b2d4f858bea373a1b6eb32f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Known versions / configuration
117M
Individual claims
Benchmarking DNA large language models on quadruplexes

Original source ↗

Table tbl0030 (paragraph 1); Table tbl0025 (paragraph 1)

Version: version of record
Retrieved: 2026-09-16T10:33:35.379Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.5.value

Source artifact SHA-256: c3d7c6d068d3c11a9c8255a932197ece3804d94e8b2d4f858bea373a1b6eb32f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: needs review

3 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: reported-model-ade36035f58f27

areas
dna-genomes
entity level
method
version
117M
reported name
DNABERT-2
historical missing metadata
checkpoint revision: not_reported_in_legacy_extract; training data: not_reported_in_legacy_extract; licence: not_reported_in_legacy_extract
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
legacy kinds
model
entity classification
review date: 2026-09-17; rationale: This source-scoped entry preserves the method/configuration actually named in an evaluation. It is neither a global family identity nor proof of an immutable checkpoint; the linked evaluation retains adaptation, fitting and scoring details.; source ids: quadruplex-llm-benchmark-2025; evidence-reported-base-dnabert2-readme-md; source locator: Materials and methods/Data preparation (paragraph 2); Results/Interpretation of LLMs (paragraph 2) | README.md model description | Discussion and conclusions (paragraph 6); Results/LLM performance at the genome-wide level (paragraph 4); ambiguities: Configuration means the source-labelled evaluated identity. It does not establish missing checkpoint hashes, default settings or equivalence to same-named records in other papers.
Related records

Suggest a correction