rewire.itbenchmarks
Task

gene fusion breakpoint classification

Fusion-breakpoint classification evaluates curated sequence labels rather than locating breakpoints in raw sequencing data.

SourcesBenchmarking genomic foundation models for binary classification of gene fusion breakpoints from DNA sequences · Methods: Evaluation metrics; implementation; Discussion; cached text lines 34–35, 45, 88–90

1 evaluation · 1 result

Overview

Datasets

The curated FusionAI benchmark.

Metrics

Accuracy, class-weighted precision/recall/F1 and ROC-AUC.

Allowed inputs

Genomic breakpoint sequence represented by foundation-model embeddings.

SourcesBenchmarking genomic foundation models for binary classification of gene fusion breakpoints from DNA sequences · Methods: Evaluation metrics; implementation; Discussion; cached text lines 34–35, 45, 88–90
Evaluation procedure diagram
How it worksComputational evaluation flow
Computational evaluation flow1. Input: Genomic breakpoint sequence represented by foundation-model embeddings.. Then: 2. Evaluation: Supervised prediction heads are compared with FusionAI.. Then: 3. Readout: Accuracy, class-weighted precision/recall/F1 and ROC-AUC.Computational evaluation flow1. Input: Genomic breakpoint sequence represented by foundation-model embeddings.. Then: 2. Evaluation: Supervised prediction heads are compared with FusionAI.. Then: 3. Readout: Accuracy, class-weighted precision/recall/F1 and ROC-AUC.Computational evaluation flow1. Input: Genomic breakpoint sequence represented by foundation-model embeddings.. Then: 2. Evaluation: Supervised prediction heads are compared with FusionAI.. Then: 3. Readout: Accuracy, class-weighted precision/recall/F1 and ROC-AUC.

Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.

SourcesBenchmarking genomic foundation models for binary classification of gene fusion breakpoints from DNA sequences · Methods: Evaluation metrics; implementation; Discussion; cached text lines 34–35, 45, 88–90

limited source coverage · Automated source review, 2026-09-16. All specifications and missing details

Results

Results are available, but no reviewed comparison panel is linked in this release.

All evaluations

1 evaluation · 1 result. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Pipeline: Nucleotide Transformer + NN (middle)Task: gene fusion breakpoint classification
Dataset: gene fusion breakpoint DNA sequences
0.994 ROC AUC
fraction · unknown

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

Nucleotide Transformer + NN (middle): gene fusion breakpoint classification

middle embedding with neural-network classifier

Aggregation: Not reported

Benchmarking genomic foundation models for binary classification of gene fusion breakpoints from DNA sequences · Table 2, NT / NN (middle) row, ROC AUC column

Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.

Methods and evaluation design

Procedure, tasks and evaluated configurations

How it works

Evaluation methodology

The curated FusionAI benchmark. Accuracy, class-weighted precision/recall/F1 and ROC-AUC. Foundation-model embeddings with supervised heads are compared with FusionAI.

SourcesBenchmarking genomic foundation models for binary classification of gene fusion breakpoints from DNA sequences · Methods: Evaluation metrics; implementation; Discussion; cached text lines 34–35, 45, 88–90; Methods: Evaluation metrics; implementation; Discussion; cached text lines 34–35, 45, 88–90; Methods: Evaluation metrics; implementation; Discussion; cached text lines 34–35, 45, 88–90

Recorded evaluations

Each evaluation records what was tested and under which conditions.

Run instructions

No runnable recipe has been reviewed for this task. Dataset access, model requirements, licences and compute requirements must be checked against its sources before execution.

A task describes a biological question. Choose a linked protocol to obtain concrete split and scoring instructions.

Strengths, limitations and unresolved questions

Strengths and limitations

Strengths and considerations

No source-reviewed explanatory claims are recorded here yet.

Profile review details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Stable record: reported-task-ee34721cf55590

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetsThe curated FusionAI benchmark.
SourcesBenchmarking genomic foundation models for binary classification of gene fusion breakpoints from DNA sequences · Methods: Evaluation metrics; implementation; Discussion; cached text lines 34–35, 45, 88–90
SplitsSampling and partitioning use a fixed random seed so models share the same examples. The reported split is not described as holding out entire genes or breakpoint families.
SourcesBenchmarking genomic foundation models for binary classification of gene fusion breakpoints from DNA sequences · Methods: implementation/reproducibility, random seed 42; Dataset and Classification
MetricsAccuracy, class-weighted precision/recall/F1 and ROC-AUC.
SourcesBenchmarking genomic foundation models for binary classification of gene fusion breakpoints from DNA sequences · Methods: Evaluation metrics; implementation; Discussion; cached text lines 34–35, 45, 88–90
BaselinesFoundation-model embeddings with supervised heads are compared with FusionAI.
SourcesBenchmarking genomic foundation models for binary classification of gene fusion breakpoints from DNA sequences · Methods: Evaluation metrics; implementation; Discussion; cached text lines 34–35, 45, 88–90
Leakage controlsAll model comparisons use the same fixed-seed partitions. This supports matched comparisons but does not by itself prevent related genomic loci crossing the split.
SourcesBenchmarking genomic foundation models for binary classification of gene fusion breakpoints from DNA sequences · Methods: Evaluation metrics; implementation; Discussion; cached text lines 34–35, 45, 88–90
UncertaintyThe paper reports final-epoch neural results and single-run SVM results; repeated-run uncertainty is not established in the reviewed passage.
SourcesBenchmarking genomic foundation models for binary classification of gene fusion breakpoints from DNA sequences · Methods: Evaluation metrics; implementation; Discussion; cached text lines 34–35, 45, 88–90
Entity typePaper-specific computational evaluation protocol.
SourcesBenchmarking genomic foundation models for binary classification of gene fusion breakpoints from DNA sequences · Methods: Evaluation metrics; implementation; Discussion; cached text lines 34–35, 45, 88–90
OrganismsHuman gene fusions from the reused FusionAI classification dataset; the foundation models’ multispecies pretraining coverage is a different property.
SourcesBenchmarking genomic foundation models for binary classification of gene fusion breakpoints from DNA sequences · Dataset/methods description; Discussion of human gene-fusion classification; cached paragraphs 77,84
AssaysFusion breakpoint labels.
SourcesBenchmarking genomic foundation models for binary classification of gene fusion breakpoints from DNA sequences · Methods: Evaluation metrics; implementation; Discussion; cached text lines 34–35, 45, 88–90
Allowed inputsGenomic breakpoint sequence represented by foundation-model embeddings.
SourcesBenchmarking genomic foundation models for binary classification of gene fusion breakpoints from DNA sequences · Methods: Evaluation metrics; implementation; Discussion; cached text lines 34–35, 45, 88–90
AdaptationSupervised prediction heads are compared with FusionAI.
SourcesBenchmarking genomic foundation models for binary classification of gene fusion breakpoints from DNA sequences · Methods: Evaluation metrics; implementation; Discussion; cached text lines 34–35, 45, 88–90

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Primary-paper discovery and source inspection. Source-checked results are not independently reproduced experiments.

Paper or primary resourceVersionReference
Benchmarking genomic foundation models for binary classification of gene fusion breakpoints from DNA sequencesjournal full text in PMCRead source
DOI: 10.1186/s13040-026-00553-1
Historical gaps recorded on 2026-09-17

The catalogue now holds 1 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.

  • Counts are approximate (~36k train, ~8k validation, ~8k test) and must remain approximate.
  • NN values are final-epoch performance; SVM is a single run. No replicate confidence bounds.
  • Do not assign embedding-plus-classifier scores to bare foundation-model checkpoints.
Search and extraction details

primary comparison table screened

Searches

  • "PMC13182013"

Evidence locations

  • Table 2
  • Methods; Evaluation metrics; Implementation details

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

17 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.
Individual claims
Benchmarking genomic foundation models for binary classification of gene fusion breakpoints from DNA sequences

Original source ↗

Methods: Evaluation metrics; implementation; Discussion; cached text lines 34–35, 45, 88–90

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558209+00:00

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 0f4d9de77f1e39cfd2164a20653d86370767da684dc22d17e09f589761abeb5f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps
  • Input: Genomic breakpoint sequence represented by foundation-model embeddings.
  • Evaluation: Supervised prediction heads are compared with FusionAI.
  • Readout: Accuracy, class-weighted precision/recall/F1 and ROC-AUC.
Individual claims
Benchmarking genomic foundation models for binary classification of gene fusion breakpoints from DNA sequences

Original source ↗

Methods: Evaluation metrics; implementation; Discussion; cached text lines 34–35, 45, 88–90

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558209+00:00

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 0f4d9de77f1e39cfd2164a20653d86370767da684dc22d17e09f589761abeb5f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title
Computational evaluation flow
Individual claims
Benchmarking genomic foundation models for binary classification of gene fusion breakpoints from DNA sequences

Original source ↗

Methods: Evaluation metrics; implementation; Discussion; cached text lines 34–35, 45, 88–90

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558209+00:00

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.diagram.title

Source artifact SHA-256: 0f4d9de77f1e39cfd2164a20653d86370767da684dc22d17e09f589761abeb5f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Datasets
The curated FusionAI benchmark.
Individual claims
Benchmarking genomic foundation models for binary classification of gene fusion breakpoints from DNA sequences

Original source ↗

Methods: Evaluation metrics; implementation; Discussion; cached text lines 34–35, 45, 88–90

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558209+00:00

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: 0f4d9de77f1e39cfd2164a20653d86370767da684dc22d17e09f589761abeb5f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Splits
Sampling and partitioning use a fixed random seed so models share the same examples. The reported split is not described as holding out entire genes or breakpoint families.
Individual claims
Benchmarking genomic foundation models for binary classification of gene fusion breakpoints from DNA sequences

Original source ↗

Methods: implementation/reproducibility, random seed 42; Dataset and Classification

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558209+00:00

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: 0f4d9de77f1e39cfd2164a20653d86370767da684dc22d17e09f589761abeb5f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Adaptation
Supervised prediction heads are compared with FusionAI.
Individual claims
Benchmarking genomic foundation models for binary classification of gene fusion breakpoints from DNA sequences

Original source ↗

Methods: Evaluation metrics; implementation; Discussion; cached text lines 34–35, 45, 88–90

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558209+00:00

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: 0f4d9de77f1e39cfd2164a20653d86370767da684dc22d17e09f589761abeb5f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Metrics
Accuracy, class-weighted precision/recall/F1 and ROC-AUC.
Individual claims
Benchmarking genomic foundation models for binary classification of gene fusion breakpoints from DNA sequences

Original source ↗

Methods: Evaluation metrics; implementation; Discussion; cached text lines 34–35, 45, 88–90

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558209+00:00

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: 0f4d9de77f1e39cfd2164a20653d86370767da684dc22d17e09f589761abeb5f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Baselines
Foundation-model embeddings with supervised heads are compared with FusionAI.
Individual claims
Benchmarking genomic foundation models for binary classification of gene fusion breakpoints from DNA sequences

Original source ↗

Methods: Evaluation metrics; implementation; Discussion; cached text lines 34–35, 45, 88–90

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558209+00:00

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.facts.3.value

Source artifact SHA-256: 0f4d9de77f1e39cfd2164a20653d86370767da684dc22d17e09f589761abeb5f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Leakage controls
All model comparisons use the same fixed-seed partitions. This supports matched comparisons but does not by itself prevent related genomic loci crossing the split.
Individual claims
Benchmarking genomic foundation models for binary classification of gene fusion breakpoints from DNA sequences

Original source ↗

Methods: Evaluation metrics; implementation; Discussion; cached text lines 34–35, 45, 88–90

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558209+00:00

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: 0f4d9de77f1e39cfd2164a20653d86370767da684dc22d17e09f589761abeb5f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Uncertainty
The paper reports final-epoch neural results and single-run SVM results; repeated-run uncertainty is not established in the reviewed passage.
Individual claims
Benchmarking genomic foundation models for binary classification of gene fusion breakpoints from DNA sequences

Original source ↗

Methods: Evaluation metrics; implementation; Discussion; cached text lines 34–35, 45, 88–90

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558209+00:00

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.facts.5.value

Source artifact SHA-256: 0f4d9de77f1e39cfd2164a20653d86370767da684dc22d17e09f589761abeb5f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: needs review

2 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: reported-task-ee34721cf55590

areas
dna-genomes
tasks
gene fusion breakpoint classification
entity level
task
version
Not reported
task
gene fusion breakpoint classification
scope note
Paper-specific evaluation task; protocol completeness requires further extraction.
benchmark research
review date: 2026-09-17; status: primary_comparison_table_screened; primary sources: expansion-p3-fusion-breakpoint-foundation-models-2026; inspected locators: Table 2; Methods; Evaluation metrics; Implementation details; searched queries: "PMC13182013"; gaps: Counts are approximate (~36k train, ~8k validation, ~8k test) and must remain approximate.; NN values are final-epoch performance; SVM is a single run. No replicate confidence bounds.; Do not assign embedding-plus-classifier scores to bare foundation-model checkpoints.; claim scope: Primary-paper discovery and source inspection. Source-checked results are not independently reproduced experiments.
historical missing metadata
protocol version: not_reported_in_legacy_extract
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
legacy kinds
benchmark
entity classification
review date: 2026-09-17; rationale: This source-scoped record identifies the biological prediction task and holds its paper context. Preserve the existing task identity; exact split, model adaptation and scoring remain in linked evaluations or separate protocol records.; source ids: fusion-breakpoint-foundation-models-2026; source locator: Methods: Evaluation metrics; implementation; Discussion; cached text lines 34–35, 45, 88–90; ambiguities: A paper- or suite-specific task may constrain some inputs or metrics; that alone does not make it interchangeable with a complete versioned protocol. No protocol equivalence is inferred.; Some legacy profile Entity type facts use the generic phrase computational evaluation protocol. That boilerplate is not sufficient to establish a single fixed protocol identity or to merge this task with another protocol record.
Related records

Suggest a correction