rewire.itbenchmarks
Task

G-quadruplex classification

G-quadruplex classification evaluates balanced positives and sampled genomic-background negatives from several annotation assays.

SourcesBenchmarking DNA large language models on quadruplexes · Methods: Data preparation; Metrics of evaluation; cached text lines 13–16, 22–23; comparative evaluation and ablation passages

2 evaluations · 2 results

Overview

Datasets

KEx, G4 ChIP-seq, G4-seq and G4 CUT&Tag annotation collections.

Metrics

Accuracy, ROC-AUC, F1 and MCC.

Allowed inputs

DNA sequences around candidate G-quadruplex regions.

SourcesBenchmarking DNA large language models on quadruplexes · Methods: Data preparation; Metrics of evaluation; cached text lines 13–16, 22–23; comparative evaluation and ablation passages
Evaluation procedure diagram
How it worksComputational evaluation flow
Computational evaluation flow1. Input: DNA sequences around candidate G-quadruplex regions.. Then: 2. Evaluation: Supervised classification with five-fold cross-validation.. Then: 3. Readout: Accuracy, ROC-AUC, F1 and MCC.Computational evaluation flow1. Input: DNA sequences around candidate G-quadruplex regions.. Then: 2. Evaluation: Supervised classification with five-fold cross-validation.. Then: 3. Readout: Accuracy, ROC-AUC, F1 and MCC.Computational evaluation flow1. Input: DNA sequences around candidate G-quadruplex regions.. Then: 2. Evaluation: Supervised classification with five-fold cross-validation.. Then: 3. Readout: Accuracy, ROC-AUC, F1 and MCC.

Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.

SourcesBenchmarking DNA large language models on quadruplexes · Methods: Data preparation; Metrics of evaluation; cached text lines 13–16, 22–23; comparative evaluation and ablation passages

limited source coverage · Automated source review, 2026-09-16. All specifications and missing details

Results

Results are available, but no reviewed comparison panel is linked in this release.

All evaluations

2 evaluations · 2 results. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: DNABERT-2Task: G-quadruplex classification
Dataset: KEx
97 Accuracy
% · unknown

Uncertainty: ± 0.5

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

DNABERT-2: G-quadruplex classification

Pretrained model evaluated on KEx as reported in Table 5.

Aggregation: Not reported

Benchmarking DNA large language models on quadruplexes · Table 5, DNABERT-2 (117 M) row, Accuracy column
Configuration: CaduceusTask: G-quadruplex classification
Dataset: KEx
95 Accuracy
% · unknown

Uncertainty: ± 0.5

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

Caduceus: G-quadruplex classification

Pretrained model evaluated on KEx as reported in Table 5.

Aggregation: Not reported

Benchmarking DNA large language models on quadruplexes · Table 5, Caduceus (8 M) row, Accuracy column

Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.

Methods and evaluation design

Procedure, tasks and evaluated configurations

How it works

Evaluation methodology

KEx, G4 ChIP-seq, G4-seq and G4 CUT&Tag annotation collections. Accuracy, ROC-AUC, F1 and MCC. DNABERT, DNABERT-2, GENA-LM, HyenaDNA and Caduceus; tested context lengths differ.

SourcesBenchmarking DNA large language models on quadruplexes · Methods: Data preparation; Metrics of evaluation; cached text lines 13–16, 22–23; comparative evaluation and ablation passages; Methods: Data preparation; Metrics of evaluation; cached text lines 13–16, 22–23; comparative evaluation and ablation passages; Methods: Data preparation; Metrics of evaluation; cached text lines 13–16, 22–23; comparative evaluation and ablation passages

Recorded evaluations

Each evaluation records what was tested and under which conditions.

Run instructions

No runnable recipe has been reviewed for this task. Dataset access, model requirements, licences and compute requirements must be checked against its sources before execution.

A task describes a biological question. Choose a linked protocol to obtain concrete split and scoring instructions.

Strengths, limitations and unresolved questions

Strengths and limitations

Strengths and considerations

No source-reviewed explanatory claims are recorded here yet.

Limitations and conditions

  • Balanced sampled negatives do not reproduce genome-wide prevalence. Fold variability is not the same as an independent cohort confidence interval.
    SourcesBenchmarking DNA large language models on quadruplexes · Methods: Data preparation; Metrics of evaluation; cached text lines 13–16, 22–23; comparative evaluation and ablation passages
Profile review details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Stable record: reported-task-c9d2a6435979e9

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetsKEx, G4 ChIP-seq, G4-seq and G4 CUT&Tag annotation collections.
SourcesBenchmarking DNA large language models on quadruplexes · Methods: Data preparation; Metrics of evaluation; cached text lines 13–16, 22–23; comparative evaluation and ablation passages
SplitsFive-fold cross-validation.
SourcesBenchmarking DNA large language models on quadruplexes · Methods: Data preparation; Metrics of evaluation; cached text lines 13–16, 22–23; comparative evaluation and ablation passages
MetricsAccuracy, ROC-AUC, F1 and MCC.
SourcesBenchmarking DNA large language models on quadruplexes · Methods: Data preparation; Metrics of evaluation; cached text lines 13–16, 22–23; comparative evaluation and ablation passages
BaselinesDNABERT, DNABERT-2, GENA-LM, HyenaDNA and Caduceus; tested context lengths differ.
SourcesBenchmarking DNA large language models on quadruplexes · Methods: Data preparation; Metrics of evaluation; cached text lines 13–16, 22–23; comparative evaluation and ablation passages
Leakage controlsSampled negative regions exclude annotated G-quadruplex positives. Data preparation does not specify a chromosome-disjoint or homology-disjoint partition, so non-overlapping positive/negative labels alone do not establish genomic independence.
SourcesBenchmarking DNA large language models on quadruplexes · Materials and methods: Data preparation and Metrics of evaluation
UncertaintyMean and standard deviation across the five folds.
SourcesBenchmarking DNA large language models on quadruplexes · Methods: Data preparation; Metrics of evaluation; cached text lines 13–16, 22–23; comparative evaluation and ablation passages
Entity typePaper-specific computational evaluation protocol.
SourcesBenchmarking DNA large language models on quadruplexes · Methods: Data preparation; Metrics of evaluation; cached text lines 13–16, 22–23; comparative evaluation and ablation passages
OrganismsThe paper reuses KEx, G4 ChIP-seq, G4-seq and G4 CUT&Tag collections. Data preparation lists assays and sample counts without dataset-by-dataset organism/assembly identifiers. Its references include both mammalian and multispecies studies, so a single genome cannot be assigned to all four from this description. · Not reported in inspected sources
SourcesBenchmarking DNA large language models on quadruplexes · Materials and methods: Data preparation, Table 1; primary dataset references 14–17,24
AssaysG4 ChIP-seq, G4-seq, CUT&Tag and KEx annotations.
SourcesBenchmarking DNA large language models on quadruplexes · Methods: Data preparation; Metrics of evaluation; cached text lines 13–16, 22–23; comparative evaluation and ablation passages
Allowed inputsDNA sequences around candidate G-quadruplex regions.
SourcesBenchmarking DNA large language models on quadruplexes · Methods: Data preparation; Metrics of evaluation; cached text lines 13–16, 22–23; comparative evaluation and ablation passages
AdaptationSupervised classification with five-fold cross-validation.
SourcesBenchmarking DNA large language models on quadruplexes · Methods: Data preparation; Metrics of evaluation; cached text lines 13–16, 22–23; comparative evaluation and ablation passages

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.

Paper or primary resourceVersionReference
Benchmarking DNA large language models on quadruplexesversion of recordRead source
DOI: 10.1016/j.csbj.2025.03.007
Historical gaps recorded on 2026-09-17

The catalogue now holds 2 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.

  • Complete raw tables acquired. Positive/negative construction, sequence lengths and train/test splits remain part of each G4 protocol; existing observations preserved. Structured extraction pending.
Search and extraction details

source found structured extraction pending

Searches

  • Benchmarking DNA large language models on quadruplexes primary paper benchmark results

Evidence locations

  • Primary results tables and G4 classification protocol

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

17 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.
Individual claims
Benchmarking DNA large language models on quadruplexes

Original source ↗

Methods: Data preparation; Metrics of evaluation; cached text lines 13–16, 22–23; comparative evaluation and ablation passages

Version: version of record
Retrieved: 2026-09-16T10:33:35.379Z

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: c3d7c6d068d3c11a9c8255a932197ece3804d94e8b2d4f858bea373a1b6eb32f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps
  • Input: DNA sequences around candidate G-quadruplex regions.
  • Evaluation: Supervised classification with five-fold cross-validation.
  • Readout: Accuracy, ROC-AUC, F1 and MCC.
Individual claims
Benchmarking DNA large language models on quadruplexes

Original source ↗

Methods: Data preparation; Metrics of evaluation; cached text lines 13–16, 22–23; comparative evaluation and ablation passages

Version: version of record
Retrieved: 2026-09-16T10:33:35.379Z

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: c3d7c6d068d3c11a9c8255a932197ece3804d94e8b2d4f858bea373a1b6eb32f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title
Computational evaluation flow
Individual claims
Benchmarking DNA large language models on quadruplexes

Original source ↗

Methods: Data preparation; Metrics of evaluation; cached text lines 13–16, 22–23; comparative evaluation and ablation passages

Version: version of record
Retrieved: 2026-09-16T10:33:35.379Z

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.diagram.title

Source artifact SHA-256: c3d7c6d068d3c11a9c8255a932197ece3804d94e8b2d4f858bea373a1b6eb32f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Datasets
KEx, G4 ChIP-seq, G4-seq and G4 CUT&Tag annotation collections.
Individual claims
Benchmarking DNA large language models on quadruplexes

Original source ↗

Methods: Data preparation; Metrics of evaluation; cached text lines 13–16, 22–23; comparative evaluation and ablation passages

Version: version of record
Retrieved: 2026-09-16T10:33:35.379Z

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: c3d7c6d068d3c11a9c8255a932197ece3804d94e8b2d4f858bea373a1b6eb32f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Splits
Five-fold cross-validation.
Individual claims
Benchmarking DNA large language models on quadruplexes

Original source ↗

Methods: Data preparation; Metrics of evaluation; cached text lines 13–16, 22–23; comparative evaluation and ablation passages

Version: version of record
Retrieved: 2026-09-16T10:33:35.379Z

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: c3d7c6d068d3c11a9c8255a932197ece3804d94e8b2d4f858bea373a1b6eb32f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Adaptation
Supervised classification with five-fold cross-validation.
Individual claims
Benchmarking DNA large language models on quadruplexes

Original source ↗

Methods: Data preparation; Metrics of evaluation; cached text lines 13–16, 22–23; comparative evaluation and ablation passages

Version: version of record
Retrieved: 2026-09-16T10:33:35.379Z

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: c3d7c6d068d3c11a9c8255a932197ece3804d94e8b2d4f858bea373a1b6eb32f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Metrics
Accuracy, ROC-AUC, F1 and MCC.
Individual claims
Benchmarking DNA large language models on quadruplexes

Original source ↗

Methods: Data preparation; Metrics of evaluation; cached text lines 13–16, 22–23; comparative evaluation and ablation passages

Version: version of record
Retrieved: 2026-09-16T10:33:35.379Z

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: c3d7c6d068d3c11a9c8255a932197ece3804d94e8b2d4f858bea373a1b6eb32f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Baselines
DNABERT, DNABERT-2, GENA-LM, HyenaDNA and Caduceus; tested context lengths differ.
Individual claims
Benchmarking DNA large language models on quadruplexes

Original source ↗

Methods: Data preparation; Metrics of evaluation; cached text lines 13–16, 22–23; comparative evaluation and ablation passages

Version: version of record
Retrieved: 2026-09-16T10:33:35.379Z

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.facts.3.value

Source artifact SHA-256: c3d7c6d068d3c11a9c8255a932197ece3804d94e8b2d4f858bea373a1b6eb32f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Leakage controls
Sampled negative regions exclude annotated G-quadruplex positives. Data preparation does not specify a chromosome-disjoint or homology-disjoint partition, so non-overlapping positive/negative labels alone do not establish genomic independence.
Individual claims
Benchmarking DNA large language models on quadruplexes

Original source ↗

Materials and methods: Data preparation and Metrics of evaluation

Version: version of record
Retrieved: 2026-09-16T10:33:35.379Z

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: c3d7c6d068d3c11a9c8255a932197ece3804d94e8b2d4f858bea373a1b6eb32f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Uncertainty
Mean and standard deviation across the five folds.
Individual claims
Benchmarking DNA large language models on quadruplexes

Original source ↗

Methods: Data preparation; Metrics of evaluation; cached text lines 13–16, 22–23; comparative evaluation and ablation passages

Version: version of record
Retrieved: 2026-09-16T10:33:35.379Z

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.facts.5.value

Source artifact SHA-256: c3d7c6d068d3c11a9c8255a932197ece3804d94e8b2d4f858bea373a1b6eb32f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: needs review

2 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: reported-task-c9d2a6435979e9

areas
dna-genomes
tasks
G-quadruplex classification
entity level
task
version
Not reported
task
G-quadruplex classification
scope note
Paper-specific evaluation task; protocol completeness requires further extraction.
benchmark research
review date: 2026-09-17; status: source_found_structured_extraction_pending; primary sources: evidence-expansion-p2-quadruplex-llm-benchmark-2025-c3d7c6d068d3; inspected locators: Primary results tables and G4 classification protocol; searched queries: Benchmarking DNA large language models on quadruplexes primary paper benchmark results; gaps: Complete raw tables acquired. Positive/negative construction, sequence lengths and train/test splits remain part of each G4 protocol; existing observations preserved. Structured extraction pending.; claim scope: Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.
historical missing metadata
protocol version: not_reported_in_legacy_extract; split: not_reported_in_legacy_extract
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
legacy kinds
benchmark
entity classification
review date: 2026-09-17; rationale: This source-scoped record identifies the biological prediction task and holds its paper context. Preserve the existing task identity; exact split, model adaptation and scoring remain in linked evaluations or separate protocol records.; source ids: quadruplex-llm-benchmark-2025; source locator: Methods: Data preparation; Metrics of evaluation; cached text lines 13–16, 22–23; comparative evaluation and ablation passages; ambiguities: A paper- or suite-specific task may constrain some inputs or metrics; that alone does not make it interchangeable with a complete versioned protocol. No protocol equivalence is inferred.; Some legacy profile Entity type facts use the generic phrase computational evaluation protocol. That boilerplate is not sufficient to establish a single fixed protocol identity or to merge this task with another protocol record.
Related records

Suggest a correction