rewirebio.iobenchmarks

Use caseDNA and genomesResearch and clinical research

Select an execution workflow for large-scale diagnostic genomics

Which eligible model configuration and execution workflow can meet a diagnostic genomics team's resource, throughput and reproducibility constraints?

In this release 1 evaluated endpoint · 1 tested configuration · Proxy evidence only · 3 recorded evidence gaps

Question and applicability

Inspect the DRAGEN 4.2 single-sample runtime baseline before scoping a foundation-model execution workflow's resource/throughput budget; this conventional baseline is a proxy, not a model-execution benchmark.

Who this is for
  • Clinical researchers scoping the available diagnostic-genomics evidence
  • Computational researchers comparing exact evaluated configurations
Research setting

Research and clinical research. See Clinical research scope for what the evidence does not establish.

Biological setting
DRAGEN 4.2 Phase 4 total single-sample wall-clock runtime for HG002 is 1,838.5 seconds on one hardware configuration and 5,521.21 seconds on an AWS HG002 configuration (f1.4xlarge, Xeon E5-2686v4, 16 threads); these are hardware comparators, not model comparators, and stage times cannot be summed due to concurrency.

Outside this use case

  • Any foundation-model inference runtime or throughput benchmark (none is ingested in this bounded intake)
  • Peak memory usage (host RAM capacity in the hardware comparator is not measured peak usage)
  • Clinical-reporting turnaround time

Inputs and expected output

Inputs you need

  • The diagnostic genomics team's available hardware configuration
  • The required per-sample turnaround and batch throughput

Expected output

A sourced single-sample total runtime baseline on two hardware configurations; no foundation-model execution benchmark is ingested in this bounded intake.

Clinical research scope

Clinical applicability is not established: this is a conventional variant-calling pipeline's operational runtime baseline, a proxy for the resource/throughput constraints a foundation-model execution workflow would need to meet, not itself a model-execution benchmark.

Evaluated evidence

Evidence is grouped by its protocol. Relevance refers to the stated endpoint and context; it is separate from clinical validation and from the review method. Limits specific to each evaluation are listed with it.

HG002 DRAGEN full single-genome runtime on Phase4

Current source-reviewed mapping

Proxy evidence: transfer to this question is limited

A conventional variant-calling pipeline's measured single-sample runtime is a proxy baseline for the throughput/resource constraints a foundation-model execution workflow would need to meet; it is not itself a model-execution benchmark.

Assessed endpoint
Total single-sample wall-clock runtime (seconds) for DRAGEN 4.2 Phase 4 processing of HG002, by hardware configuration
Evaluation protocol
HG002 DRAGEN full single-genome runtime on Phase4
Computational task
A reviewed task relationship is not recorded for this protocol.
Input and population constraints
  • Inspect every linked evaluation's source locator and preserved conflicts before citing a result.
  • Do not combine this mapping's evaluations with any other protocol's results.

Limits on interpretation

  • No foundation-model inference benchmark
  • No peak memory
  • No clinical reporting turnaround
  • No runtime repeatCI

Automated source review · 2026-10-07 · Claude Sonnet AMP-integration worker, bounded transcription of Codex-checked primary values; independently reviewed by Codex (workbench/amp-supervision/primary-review.md, integration-review-corrections.md)

Bounded primary-source transcription, independently Codex-checked. No new model execution, independent experimental reproduction, qualified human scientific review or clinical validation.

Evaluated configurations

Each configuration below belongs to this protocol. Inspect its inputs, population and scoring conditions before comparing it with another evaluation.

DRAGEN 4.2 on Phase4 SKY-6200

Author-reported evaluation · Source checked

Scoped author-reported evidence candidate; no clinical recommendation.

Inspect results, conditions and reproduction (1 recorded result)
Population and split
OneHG002 row of seven GIABsamples listed per hardware · Single HG002 row; supplement also lists HG001–HG007
Inputs and adaptation
35× WGS raw reads · Not reported
Evaluation budget
SKY-6200 Phase4,64threads
Runtime and memory

Runtime and memory measurements are not reported in this evaluation. A study budget is not a runtime or memory measurement.

Recorded results for this configuration
MetricValueCoverageUncertainty and source
total_pipeline_runtime1840

seconds · lower

Not reported scored / Not reported eligible

Not reported

Result provenance

Uncertainty: Unreported

Evaluation methods, evidence and reproduction

No execution recipe has been verified for this exact configuration and evaluation. Inspect its methods and original run documentation before attempting reproduction.

Open the protocol's results and comparison checks →

Mapping sources and review metadata

Mapping use-case-mapping-amp-20261007-issue18 · revision 1

Add Codex-checked primary-source protocol evidence from the bounded AMP intake (rewire.it#365).

Reviewed evidence fingerprint 8992490eb9aba66f496e6f9d8950a8b85405062638ca44d2bfed0dbe80e2dd59

Limitations and missing evidence

These gaps apply to the question as a whole. Absence of evidence is not a zero score.

  • No peak-memory endpoint is reported.
  • No runtime repeat/variance (CI) is reported.
  • A separate ~2-hour joint-aggregation figure over 3,202 genomes at concurrency 200 is a different operation and is not assigned to this single-sample runtime.

Planned work

These plans do not contribute measured results or evaluated winners above.

Contribute evidence or propose a correction

Evidence collection plan

Collecting evidence

Mapped evidence already covers 1 evaluated endpoint, in the evaluated evidence above. The status above describes only the specific comparison in this plan, which remains open; it does not mean no evidence has been collected. The plan defines a comparison to investigate; it does not establish model performance or suitability.

Comparison question

Which eligible model configuration and execution workflow can meet a diagnostic genomics team's resource, throughput and reproducibility constraints?

Baselines, outcomes and validation requirements

Baselines to include

  • The conventional/author-introduced workflow measured in the linked primary source(s).

Outcomes to measure

  • The declared endpoint in this use case's active mapping(s); see evidence_gaps for what remains open.

Validation requirements

  • Independent held-out population matched to the intended clinical setting.
  • Qualified human scientific review before any clinical-validation claim.

Next collection task

A dedicated rewire-benchmarks protocol/run task that actually executes an eligible foundation-model configuration and measures its resource/throughput budget.

Sources and review

Automated source review · 2026-10-07 · Claude Sonnet AMP-integration worker, bounded transcription of Codex-checked primary values; independently reviewed by Codex (workbench/amp-supervision/primary-review.md, integration-review-corrections.md)

Bounded primary-source transcription, independently Codex-checked. No new model execution, independent experimental reproduction, qualified human scientific review or clinical validation.

Release provenance and downloads

Release 2026-10-07-1448159e6a81

Use-case input digest d60fd7f669bfec7bd34ec6e5080d8e1cb4e8186f8286ead60888848f1e20e001

Download questions, collection plans and review metadata (JSON) · Verify release checksums

Question use-case-diagnostic-genomics-model-execution. Any numerical results on this page come from this release's existing evaluation records.