rewire.itbenchmarks
Task

polyadenylation site detection

Polyadenylation-site classification compares two background definitions and two adaptation regimes.

SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51

1 evaluation · 1 result

Overview

Datasets

GENCODE v38 human annotations with balanced Gene-Gene and Gene-Intergene datasets; a separate mouse transfer evaluation.

Metrics

Accuracy, precision, recall, F1 and AUC, averaged across folds.

Allowed inputs

DNA/RNA sequence context for candidate polyadenylation sites.

SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51
Evaluation procedure diagram
How it worksComputational evaluation flow
Computational evaluation flow1. Input: DNA/RNA sequence context for candidate polyadenylation sites.. Then: 2. Evaluation: Five-fold evaluation with 60:20:20 training/validation/test proportions per fold.. Then: 3. Readout: Accuracy, precision, recall, F1 and AUC, averaged across folds.Computational evaluation flow1. Input: DNA/RNA sequence context for candidate polyadenylation sites.. Then: 2. Evaluation: Five-fold evaluation with 60:20:20 training/validation/test proportions per fold.. Then: 3. Readout: Accuracy, precision, recall, F1 and AUC, averaged across folds.Computational evaluation flow1. Input: DNA/RNA sequence context for candidate polyadenylation sites.. Then: 2. Evaluation: Five-fold evaluation with 60:20:20 training/validation/test proportions per fold.. Then: 3. Readout: Accuracy, precision, recall, F1 and AUC, averaged across folds.

Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.

SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51

Source reviewed · Automated source review, 2026-09-16. All specifications and missing details

Results

Results are available, but no reviewed comparison panel is linked in this release.

All evaluations

1 evaluation · 1 result. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: HyenaDNATask: polyadenylation site detection
Dataset: poly(A) Gene-Gene
0.751 AUC
fraction · unknown

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

HyenaDNA: polyadenylation site detection

few-shot Gene-Gene negative-set comparison

Aggregation: Not reported

PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Table 1, Few-shot HyenaDNA row, G-G AUC column

Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.

Methods and evaluation design

Procedure, tasks and evaluated configurations

How it works

Evaluation methodology

GENCODE v38 human annotations with balanced Gene-Gene and Gene-Intergene datasets; a separate mouse transfer evaluation. Five-fold evaluation with 60:20:20 training/validation/test proportions per fold. Accuracy, precision, recall, F1 and AUC, averaged across folds. DNABERT-2, Nucleotide Transformer and HyenaDNA under few-shot and fine-tuned conditions. Negative examples exclude genic regions overlapping GENCODE v38 polyadenylation annotations. This is a label-construction control; the evaluation description does not establish chromosome or homologous-sequence separation between the five-fold training and test partitions.

SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51; Methods: dataset construction and evaluation framework

Recorded evaluations

Each evaluation records what was tested and under which conditions.

Run instructions

No runnable recipe has been reviewed for this task. Dataset access, model requirements, licences and compute requirements must be checked against its sources before execution.

A task describes a biological question. Choose a linked protocol to obtain concrete split and scoring instructions.

Strengths, limitations and unresolved questions

Strengths and limitations

Strengths supported by sources

No source-reviewed explanatory claims are recorded here yet.

Profile review details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Stable record: reported-task-13dfe6b33e71ed

Specifications

Inputs, training, access and other details

Explanatory profile: source reviewed · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetsGENCODE v38 human annotations with balanced Gene-Gene and Gene-Intergene datasets; a separate mouse transfer evaluation.
SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51
SplitsFive-fold evaluation with 60:20:20 training/validation/test proportions per fold.
SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51
MetricsAccuracy, precision, recall, F1 and AUC, averaged across folds.
SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51
BaselinesDNABERT-2, Nucleotide Transformer and HyenaDNA under few-shot and fine-tuned conditions.
SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51
Leakage controlsNegative examples exclude genic regions overlapping GENCODE v38 polyadenylation annotations. This is a label-construction control; the evaluation description does not establish chromosome or homologous-sequence separation between the five-fold training and test partitions.
SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation framework
UncertaintyTable 1 reports averages over five folds. For few-shot predictions, ten random prototype pairs per fold are averaged. No confidence interval or standard-error estimator for the table’s aggregate metrics is specified there.
SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: few-shot evaluation; Table 1 caption
Entity typePaper-specific computational evaluation protocol.
SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51
OrganismsHuman; separate transfer evaluation in mouse.
SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51
AssaysGENCODE polyadenylation annotations.
SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51
Allowed inputsDNA/RNA sequence context for candidate polyadenylation sites.
SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51
AdaptationFew-shot and task fine-tuning settings are compared separately.
SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.

Paper or primary resourceVersionReference
PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language modelsjournal full text in PMCRead source
DOI: 10.1016/j.csbj.2025.12.011
Historical gaps recorded on 2026-09-17

The catalogue now holds 1 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.

  • Raw complete comparison acquired. Few-shot versus fine-tuning and Gene-Gene versus Intergenic-Gene conditions cannot share a chart group; NT100M and 500M are different configurations. Structured extraction pending.
Search and extraction details

source found structured extraction pending

Searches

  • PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models primary paper benchmark results

Evidence locations

  • Table1 two negative-set conditions and few-shot/fine-tuning row groups

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

17 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.
Individual claims
PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models

Original source ↗

Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558220+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: e9ebd53d88837ad8d457881ffee918d2734dcae87d3c5cd03135947b6cf5dbde

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps
  • Input: DNA/RNA sequence context for candidate polyadenylation sites.
  • Evaluation: Five-fold evaluation with 60:20:20 training/validation/test proportions per fold.
  • Readout: Accuracy, precision, recall, F1 and AUC, averaged across folds.
Individual claims
PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models

Original source ↗

Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558220+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: e9ebd53d88837ad8d457881ffee918d2734dcae87d3c5cd03135947b6cf5dbde

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title
Computational evaluation flow
Individual claims
PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models

Original source ↗

Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558220+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.diagram.title

Source artifact SHA-256: e9ebd53d88837ad8d457881ffee918d2734dcae87d3c5cd03135947b6cf5dbde

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Datasets
GENCODE v38 human annotations with balanced Gene-Gene and Gene-Intergene datasets; a separate mouse transfer evaluation.
Individual claims
PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models

Original source ↗

Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558220+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: e9ebd53d88837ad8d457881ffee918d2734dcae87d3c5cd03135947b6cf5dbde

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Splits
Five-fold evaluation with 60:20:20 training/validation/test proportions per fold.
Individual claims
PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models

Original source ↗

Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558220+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: e9ebd53d88837ad8d457881ffee918d2734dcae87d3c5cd03135947b6cf5dbde

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Adaptation
Few-shot and task fine-tuning settings are compared separately.
Individual claims
PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models

Original source ↗

Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558220+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: e9ebd53d88837ad8d457881ffee918d2734dcae87d3c5cd03135947b6cf5dbde

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Metrics
Accuracy, precision, recall, F1 and AUC, averaged across folds.
Individual claims
PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models

Original source ↗

Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558220+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: e9ebd53d88837ad8d457881ffee918d2734dcae87d3c5cd03135947b6cf5dbde

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Baselines
DNABERT-2, Nucleotide Transformer and HyenaDNA under few-shot and fine-tuned conditions.
Individual claims
PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models

Original source ↗

Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558220+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.3.value

Source artifact SHA-256: e9ebd53d88837ad8d457881ffee918d2734dcae87d3c5cd03135947b6cf5dbde

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Leakage controls
Negative examples exclude genic regions overlapping GENCODE v38 polyadenylation annotations. This is a label-construction control; the evaluation description does not establish chromosome or homologous-sequence separation between the five-fold training and test partitions.
Individual claims
PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models

Original source ↗

Methods: dataset construction and evaluation framework

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558220+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: e9ebd53d88837ad8d457881ffee918d2734dcae87d3c5cd03135947b6cf5dbde

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Uncertainty
Table 1 reports averages over five folds. For few-shot predictions, ten random prototype pairs per fold are averaged. No confidence interval or standard-error estimator for the table’s aggregate metrics is specified there.
Individual claims
PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models

Original source ↗

Methods: few-shot evaluation; Table 1 caption

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558220+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.5.value

Source artifact SHA-256: e9ebd53d88837ad8d457881ffee918d2734dcae87d3c5cd03135947b6cf5dbde

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: needs review

2 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: reported-task-13dfe6b33e71ed

areas
dna-genomes
tasks
polyadenylation site detection
entity level
task
version
Not reported
task
polyadenylation site detection
scope note
Paper-specific evaluation task; protocol completeness requires further extraction.
benchmark research
review date: 2026-09-17; status: source_found_structured_extraction_pending; primary sources: evidence-expansion-p2-polya-glm-2025-e9ebd53d8883; inspected locators: Table1 two negative-set conditions and few-shot/fine-tuning row groups; searched queries: PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models primary paper benchmark results; gaps: Raw complete comparison acquired. Few-shot versus fine-tuning and Gene-Gene versus Intergenic-Gene conditions cannot share a chart group; NT100M and 500M are different configurations. Structured extraction pending.; claim scope: Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.
historical missing metadata
protocol version: not_reported_in_legacy_extract
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
legacy kinds
benchmark
entity classification
review date: 2026-09-17; rationale: This source-scoped record identifies the biological prediction task and holds its paper context. Preserve the existing task identity; exact split, model adaptation and scoring remain in linked evaluations or separate protocol records.; source ids: polya-glm-2025; source locator: Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51; ambiguities: A paper- or suite-specific task may constrain some inputs or metrics; that alone does not make it interchangeable with a complete versioned protocol. No protocol equivalence is inferred.; Some legacy profile Entity type facts use the generic phrase computational evaluation protocol. That boilerplate is not sufficient to establish a single fixed protocol identity or to merge this task with another protocol record.
Related records

Suggest a correction