rewirebio.iobenchmarks

Use caseDNA and genomesResearch and clinical research

Select a tumour RNA fusion-detection workflow

Which RNA-sequencing workflow detects and prioritises tumour gene fusions with useful sensitivity and a manageable false-positive review burden?

In this release 2 evaluated endpoints · 6 tested configurations · Direct evidence for the stated endpoint · 3 recorded evidence gaps

Question and applicability

Inspect the matched synthetic Seraseq reference-standard and NCH clinical-ascertainment evidence for Arriba, STAR-Fusion and the EnFusion ensemble before selecting a workflow and review burden for the intended specimen type.

Who this is for
  • Clinical researchers scoping the available diagnostic-genomics evidence
  • Computational researchers comparing exact evaluated configurations
Research setting

Research and clinical research. See Clinical research scope for what the evidence does not establish.

Biological setting
A 14-fusion synthetic Seraseq reference standard (undiluted, duplicate libraries) and a 229-sample paediatric cancer/haematologic cohort (67 ensemble-ascertained clinically relevant fusions) measure fusion-calling sensitivity and precision for Arriba, STAR-Fusion and the EnFusion ensemble (with/without filtering and a known-fusion list).

Outside this use case

  • Novel-fusion discovery performance (the optimised known-fusion rescue measures recovery of 14 known synthetic fusions, not novel-fusion discovery)
  • Adult solid-tumour cohorts outside the transcribed NCH cohort
  • Non-Seraseq background fusions, all counted false positive in this intake though some may be real endogenous fusions

Inputs and expected output

Inputs you need

  • Paired-end tumour RNA-seq reads
  • The intended specimen type (frozen, FFPE or other) and acceptable false-positive review burden

Expected output

A sourced set of exact evaluated sensitivity/precision values on a synthetic reference standard and a real clinical cohort, with the ascertainment-bias caveat explicit; no patient-level fusion classification.

Clinical research scope

Clinical applicability is bounded: the NCH clinical-cohort sensitivity figures are ascertained against the optimised EnFusion ensemble's own calls, not an independent exhaustive truth set, so a false-negative rate for fusions the ensemble itself missed cannot be derived from this evidence.

Evaluated evidence

Evidence is grouped by its protocol. Relevance refers to the stated endpoint and context; it is separate from clinical validation and from the review method. Limits specific to each evaluation are listed with it.

NCH clinically relevant fusion ascertainment

Current source-reviewed mapping

Direct evidence for the stated endpoint

A real clinical cohort with ensemble-ascertained clinically relevant fusions directly measures fusion-detection sensitivity in a clinical population, the declared endpoint; the ascertainment denominator is not an independent exhaustive truth set.

Assessed endpoint
Retrospective sensitivity among 67 ensemble-ascertained clinically relevant fusions in a 229-sample paediatric cancer/haematologic cohort, for STAR-Fusion and Arriba
Evaluation protocol
NCH clinically relevant fusion ascertainment
Computational task
A reviewed task relationship is not recorded for this protocol.
Input and population constraints
  • Inspect every linked evaluation's source locator and preserved conflicts before citing a result.
  • Do not combine this mapping's evaluations with any other protocol's results.

Limits on interpretation

  • Ascertainment is from the optimized EnFusion pipeline; denominator is not an independent exhaustive truth set.
  • Only a subset orthogonally confirmed; the endpoint does not quantify novel fusion generalization or patient outcomes.
  • Do not combine with duplicate Seraseq measurements.

Automated source review · 2026-10-07 · Claude Sonnet AMP-integration worker, bounded transcription of Codex-checked primary values; independently reviewed by Codex (workbench/amp-supervision/primary-review.md, integration-review-corrections.md)

Bounded primary-source transcription, independently Codex-checked. No new model execution, independent experimental reproduction, qualified human scientific review or clinical validation.

Evaluated configurations

Each configuration below belongs to this protocol. Inspect its inputs, population and scoring conditions before comparing it with another evaluation.

STAR-Fusion

Independent external evaluation · Source checked

STAR-Fusion NCH clinical-ascertainment evaluation; bounded primary-source AMP candidate.

Inspect results, conditions and reproduction (2 recorded results)
Population and split
67 clinically relevant fusions in 229 pediatric cancer/hematologic samples · Not reported
Inputs and adaptation
paired-end tumor RNA-seq · Not reported
Evaluation budget
Not reported
Runtime and memory

Runtime and memory measurements are not reported in this evaluation. A study budget is not a runtime or memory measurement.

Recorded results for this configuration
MetricValueCoverageUncertainty and source
clinically relevant fusions called63

fusions · higher

Not reported scored / Not reported eligible

Not reported

Result provenance
sensitivity among ensemble-ascertained clinically relevant fusions94%

% · higher

Not reported scored / Not reported eligible

Not reported

Result provenance

Evaluation methods, evidence and reproduction

No execution recipe has been verified for this exact configuration and evaluation. Inspect its methods and original run documentation before attempting reproduction.

Arriba

Independent external evaluation · Source checked

Arriba NCH clinical-ascertainment evaluation; bounded primary-source AMP candidate.

Inspect results, conditions and reproduction (2 recorded results)
Population and split
67 clinically relevant fusions in 229 pediatric cancer/hematologic samples · Not reported
Inputs and adaptation
paired-end tumor RNA-seq · Not reported
Evaluation budget
Not reported
Runtime and memory

Runtime and memory measurements are not reported in this evaluation. A study budget is not a runtime or memory measurement.

Recorded results for this configuration
MetricValueCoverageUncertainty and source
clinically relevant fusions called59

fusions · higher

Not reported scored / Not reported eligible

Not reported

Result provenance
sensitivity among ensemble-ascertained clinically relevant fusions88.1%

% · higher

Not reported scored / Not reported eligible

Not reported

Result provenance

Evaluation methods, evidence and reproduction

No execution recipe has been verified for this exact configuration and evaluation. Inspect its methods and original run documentation before attempting reproduction.

Open the protocol's results and comparison checks →

Mapping sources and review metadata

Mapping use-case-mapping-amp-20261007-issue11-enfusion-nch-clinical · revision 1

Add Codex-checked primary-source protocol evidence from the bounded AMP intake (rewire.it#365).

Reviewed evidence fingerprint fa64eef53c2b9588535b6c09e919f81b7d6e8c9a0b19bf47cf5e211339dce30b

EnFusion undiluted Seraseq duplicate benchmark

Current source-reviewed mapping

Direct evidence for the stated endpoint

A synthetic reference standard with a known fusion set directly measures fusion-calling sensitivity and precision, the declared benchmark endpoint.

Assessed endpoint
Synthetic 14-fusion Seraseq reference-standard sensitivity, precision and total/true fusion counts for Arriba, STAR-Fusion and EnFusion ensemble configurations (undiluted duplicate libraries)
Evaluation protocol
EnFusion undiluted Seraseq duplicate benchmark
Computational task
A reviewed task relationship is not recorded for this protocol.
Input and population constraints
  • Inspect every linked evaluation's source locator and preserved conflicts before citing a result.
  • Do not combine this mapping's evaluations with any other protocol's results.

Limits on interpretation

  • The optimized known-fusion rescue uses prior knowledge of true reference fusions; 100% precision/sensitivity is an optimization-control observation, not unbiased generalization.
  • Table2 caption v3 versus optimization narrative v2 is unresolved; dataset_version=null and automatic comparison blocked.
  • STAR-Fusion precision 43.6% in Table2 vs 43.8% Par31; numeric candidate retains printed table value but requires disputed-result warning or omission pending resolution.
  • All non-Seraseq fusions counted false positives although endogenous background fusions may exist.
  • No CI or dispersion printed for Table2 rows.
  • Clinical Table1 sensitivity is relative to 67 fusions ascertained through optimized ensemble, not independently exhaustive truth or treatment benefit.

Automated source review · 2026-10-07 · Claude Sonnet AMP-integration worker, bounded transcription of Codex-checked primary values; independently reviewed by Codex (workbench/amp-supervision/primary-review.md, integration-review-corrections.md)

Bounded primary-source transcription, independently Codex-checked. No new model execution, independent experimental reproduction, qualified human scientific review or clinical validation.

Evaluated configurations

Each configuration below belongs to this protocol. Inspect its inputs, population and scoring conditions before comparing it with another evaluation.

Arriba

Independent external evaluation · Source checked

Arriba evaluation; bounded primary-source AMP candidate.

Inspect results, conditions and reproduction (4 recorded results)
Population and split
14 synthetic fusions; duplicate experiments · optimization/reference control
Inputs and adaptation
paired-end RNA-seq reads · Not reported
Evaluation budget
Not reported
Runtime and memory

Runtime and memory measurements are not reported in this evaluation. A study budget is not a runtime or memory measurement.

Recorded results for this configuration
MetricValueCoverageUncertainty and source
mean Seraseq fusions identified13

fusions · higher

Not reported scored / Not reported eligible

Not reported

Result provenance
mean total fusions identified23.5

fusions · lower

Not reported scored / Not reported eligible

Not reported

Result provenance
precision55.3%

% · higher

Not reported scored / Not reported eligible

Not reported

Result provenance
sensitivity92.9%

% · higher

Not reported scored / Not reported eligible

Not reported

Result provenance

Evaluation methods, evidence and reproduction

No execution recipe has been verified for this exact configuration and evaluation. Inspect its methods and original run documentation before attempting reproduction.

STAR-Fusion

Independent external evaluation · Source checked

STAR-Fusion evaluation; bounded primary-source AMP candidate.

Inspect results, conditions and reproduction (4 recorded results)
Population and split
14 synthetic fusions; duplicate experiments · optimization/reference control
Inputs and adaptation
paired-end RNA-seq reads · Not reported
Evaluation budget
Not reported
Runtime and memory

Runtime and memory measurements are not reported in this evaluation. A study budget is not a runtime or memory measurement.

Recorded results for this configuration
MetricValueCoverageUncertainty and source
mean Seraseq fusions identified14

fusions · higher

Not reported scored / Not reported eligible

Not reported

Result provenance
mean total fusions identified32

fusions · lower

Not reported scored / Not reported eligible

Not reported

Result provenance
precision43.6%

% · higher

Not reported scored / Not reported eligible

Not reported

Result provenance
sensitivity100%

% · higher

Not reported scored / Not reported eligible

Not reported

Result provenance

Evaluation methods, evidence and reproduction

No execution recipe has been verified for this exact configuration and evaluation. Inspect its methods and original run documentation before attempting reproduction.

EnFusion 3 callers

Author-reported evaluation · Source checked

EnFusion 3 callers evaluation; bounded primary-source AMP candidate.

Inspect results, conditions and reproduction (4 recorded results)
Population and split
14 synthetic fusions; duplicate experiments · optimization/reference control
Inputs and adaptation
paired-end RNA-seq reads · Not reported
Evaluation budget
Not reported
Runtime and memory

Runtime and memory measurements are not reported in this evaluation. A study budget is not a runtime or memory measurement.

Recorded results for this configuration
MetricValueCoverageUncertainty and source
mean Seraseq fusions identified14

fusions · higher

Not reported scored / Not reported eligible

Not reported

Result provenance
mean total fusions identified15.5

fusions · lower

Not reported scored / Not reported eligible

Not reported

Result provenance
precision90.3%

% · higher

Not reported scored / Not reported eligible

Not reported

Result provenance
sensitivity100%

% · higher

Not reported scored / Not reported eligible

Not reported

Result provenance

Evaluation methods, evidence and reproduction

No execution recipe has been verified for this exact configuration and evaluation. Inspect its methods and original run documentation before attempting reproduction.

EnFusion 3 callers + filter + known fusion list

Author-reported evaluation · Source checked

EnFusion 3 callers + filter + known fusion list evaluation; bounded primary-source AMP candidate.

Inspect results, conditions and reproduction (4 recorded results)
Population and split
14 synthetic fusions; duplicate experiments · optimization/reference control
Inputs and adaptation
paired-end RNA-seq reads · Not reported
Evaluation budget
Not reported
Runtime and memory

Runtime and memory measurements are not reported in this evaluation. A study budget is not a runtime or memory measurement.

Recorded results for this configuration
MetricValueCoverageUncertainty and source
mean Seraseq fusions identified14

fusions · higher

Not reported scored / Not reported eligible

Not reported

Result provenance
mean total fusions identified14

fusions · lower

Not reported scored / Not reported eligible

Not reported

Result provenance
precision100%

% · higher

Not reported scored / Not reported eligible

Not reported

Result provenance
sensitivity100%

% · higher

Not reported scored / Not reported eligible

Not reported

Result provenance

Evaluation methods, evidence and reproduction

No execution recipe has been verified for this exact configuration and evaluation. Inspect its methods and original run documentation before attempting reproduction.

Open the protocol's results and comparison checks →

Mapping sources and review metadata

Mapping use-case-mapping-amp-20261007-issue11-enfusion-seraseq · revision 1

Add Codex-checked primary-source protocol evidence from the bounded AMP intake (rewire.it#365).

Reviewed evidence fingerprint 46998424368e98a97385dea2498684ed51e5fa11037ea2666ed678ed62884995

Limitations and missing evidence

These gaps apply to the question as a whole. Absence of evidence is not a zero score.

  • Table 2 caption (v3) vs the optimisation narrative (v2) is an unresolved source version conflict.
  • STAR-Fusion precision is printed as both 43.6% (Table 2) and 43.8% (paragraph) in the source; both are preserved, neither is silently reconciled.
  • No confidence interval or dispersion is printed for the Seraseq Table 2 rows.

Planned work

These plans do not contribute measured results or evaluated winners above.

Contribute evidence or propose a correction

Evidence collection plan

Collecting evidence

Mapped evidence already covers 2 evaluated endpoints, in the evaluated evidence above. The status above describes only the specific comparison in this plan, which remains open; it does not mean no evidence has been collected. The plan defines a comparison to investigate; it does not establish model performance or suitability.

Comparison question

Which RNA-sequencing workflow detects and prioritises tumour gene fusions with useful sensitivity and a manageable false-positive review burden?

Baselines, outcomes and validation requirements

Baselines to include

  • The conventional/author-introduced workflow measured in the linked primary source(s).

Outcomes to measure

  • The declared endpoint in this use case's active mapping(s); see evidence_gaps for what remains open.

Validation requirements

  • Independent held-out population matched to the intended clinical setting.
  • Qualified human scientific review before any clinical-validation claim.

Next collection task

A dedicated rewire-benchmarks protocol/run task with an independent fusion truth set not ascertained by the same ensemble being evaluated.

Sources and review

Automated source review · 2026-10-07 · Claude Sonnet AMP-integration worker, bounded transcription of Codex-checked primary values; independently reviewed by Codex (workbench/amp-supervision/primary-review.md, integration-review-corrections.md)

Bounded primary-source transcription, independently Codex-checked. No new model execution, independent experimental reproduction, qualified human scientific review or clinical validation.

Release provenance and downloads

Release 2026-10-07-1448159e6a81

Use-case input digest d60fd7f669bfec7bd34ec6e5080d8e1cb4e8186f8286ead60888848f1e20e001

Download questions, collection plans and review metadata (JSON) · Verify release checksums

Question use-case-tumour-rna-fusion-detection. Any numerical results on this page come from this release's existing evaluation records.