rewirebio.iobenchmarks
Dataset

FoundationOne CDx report variants, 612 patients (Lin et al. 2025)

Real-world variants labelled clinically relevant or VUS by the FoundationOne CDx report section they appear in.

Research readiness

These checks assess whether the evidence supports a reproducible investigation. A source-checked score alone does not meet these requirements.

Release 2026-10-09-ba02f2f4a36e · Evidence verified: Not verified

Evidence incomplete

Replay metrics

Exact outcomes, predictions, identifiers and evaluator are connected.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • metric replay: verification is missing

Verified: Not verified

Evidence incomplete

Investigate discrepancies

Replay evidence includes annotations and an assessment of dependence. Unknown independence permits descriptive analysis only.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • metric replay: verification is missing
  • annotations: verification is missing
  • dependence: verification is missing

Verified: Not verified

Evidence incomplete

Run locally

A pinned recipe describes the inputs, environment and resource requirements.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • recipe pinned: verification is missing
  • resource estimate: verification is missing

Verified: Not verified

Evidence incomplete

Validate independently

Separate data and exposure records support an independent test.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • independent validation: verification is missing
  • overlap checked: verification is missing

Verified: Not verified

Readiness describes the evidence in this release. Availability on your computer is checked separately when an investigation runs. Existing data exposure can prevent independent validation even when files are available.

Artifacts and reproduction

No verified artifact manifest is connected to this record yet. The gaps above identify what is needed before analysis can begin.

Read reviewed discrepancy investigations

Evaluation results

6 evaluations · 6 results. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: GPT-4o, basic prompt, temperature 1.0 (default) (Lin et al. 2025)Protocol: Lin et al. 2025 FoundationOne variants: clinically relevant vs VUS
Dataset: FoundationOne CDx report variants, 612 patients (Lin et al. 2025)
0.732 top-1-accuracy
fraction · higher

Uncertainty: 95% CI 0.7307 to 0.7329

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

gpt-4o-basic on foundationone (Lin et al. 2025)

egfrnsclc-20261009-protocol-lin2025-foundationone-relevant-vs-vus

Aggregation: Not reported

Benchmarking large language models GPT-4o, llama 3.1, and qwen 2.5 for cancer genetic variant classification · Table 1 (Tab1), row 'Foundationone', column 'GPT-4o' Mean accuracy; 95% CI in the next column
Configuration: Llama 3.1 70B, basic prompt, temperature 0.8 (default) (Lin et al. 2025)Protocol: Lin et al. 2025 FoundationOne variants: clinically relevant vs VUS
Dataset: FoundationOne CDx report variants, 612 patients (Lin et al. 2025)
0.498 top-1-accuracy
fraction · higher

Uncertainty: 95% CI 0.4974 to 0.4978

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

llama-basic-t0-8 on foundationone (Lin et al. 2025)

egfrnsclc-20261009-protocol-lin2025-foundationone-relevant-vs-vus

Aggregation: Not reported

Benchmarking large language models GPT-4o, llama 3.1, and qwen 2.5 for cancer genetic variant classification · Table 1 (Tab1), row 'Foundationone', column 'Llama 3' Mean accuracy; 95% CI in the next column
Configuration: Qwen 2.5 72B, basic prompt, temperature 0.8 (default) (Lin et al. 2025)Protocol: Lin et al. 2025 FoundationOne variants: clinically relevant vs VUS
Dataset: FoundationOne CDx report variants, 612 patients (Lin et al. 2025)
0.573 top-1-accuracy
fraction · higher

Uncertainty: 95% CI 0.5725 to 0.5736

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

qwen-basic on foundationone (Lin et al. 2025)

egfrnsclc-20261009-protocol-lin2025-foundationone-relevant-vs-vus

Aggregation: Not reported

Benchmarking large language models GPT-4o, llama 3.1, and qwen 2.5 for cancer genetic variant classification · Table 1 (Tab1), row 'Foundationone', column 'Qwen 2.5' Mean accuracy; 95% CI in the next column
Configuration: Qwen 2.5 72B, binary class prompt, temperature 0.8 (default) (Lin et al. 2025)Protocol: Lin et al. 2025 FoundationOne variants: clinically relevant vs VUS
Dataset: FoundationOne CDx report variants, 612 patients (Lin et al. 2025)
0.612 top-1-accuracy
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

qwen-binary on foundationone (Lin et al. 2025)

egfrnsclc-20261009-protocol-lin2025-foundationone-relevant-vs-vus

Aggregation: Not reported

Benchmarking large language models GPT-4o, llama 3.1, and qwen 2.5 for cancer genetic variant classification · Table 2 (Tab2), row 3 ('Qwen2.5', 'Binary class prompt + Default temperature (0.8)', 'FoundationOne'), column 'Accuracy'
Configuration: Qwen 2.5 72B, RAG with basic prompt, temperature 0.8 (default) (Lin et al. 2025)Protocol: Lin et al. 2025 FoundationOne variants: clinically relevant vs VUS
Dataset: FoundationOne CDx report variants, 612 patients (Lin et al. 2025)
0.662 top-1-accuracy
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

qwen-rag on foundationone (Lin et al. 2025)

egfrnsclc-20261009-protocol-lin2025-foundationone-relevant-vs-vus

Aggregation: Not reported

Benchmarking large language models GPT-4o, llama 3.1, and qwen 2.5 for cancer genetic variant classification · Table 2 (Tab2), row 4 ('Qwen2.5', 'RAG + Default temperature (0.8)', 'FoundationOne'), column 'Accuracy'
Configuration: Qwen 2.5 72B, refined prompt, temperature 0.8 (default) (Lin et al. 2025)Protocol: Lin et al. 2025 FoundationOne variants: clinically relevant vs VUS
Dataset: FoundationOne CDx report variants, 612 patients (Lin et al. 2025)
0.725 top-1-accuracy
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

qwen-refined on foundationone (Lin et al. 2025)

egfrnsclc-20261009-protocol-lin2025-foundationone-relevant-vs-vus

Aggregation: Not reported

Benchmarking large language models GPT-4o, llama 3.1, and qwen 2.5 for cancer genetic variant classification · Table 2 (Tab2), row 2 ('Qwen2.5', 'Refined prompt + Default temperature (0.8)', 'FoundationOne'), column 'Accuracy'

Source checking is not independent reproduction. Release 2026-10-09-ba02f2f4a36e.

Dataset and evaluation context

A dataset supplies biological observations. The evaluation protocol defines how those observations are split, used and scored.

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

10 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-10-09-ba02f2f4a36e
Property and statementOriginal source and locationReview and provenance
attributes.access
Hospital data; not published
Context-only references
Benchmarking large language models GPT-4o, llama 3.1, and qwen 2.5 for cancer genetic variant classification

Original source ↗

Lin et al. 2025 Methods 'Dataset' (P35-P38), 'Model selection' (P39-P40), 'System prompts' (P41-P43), 'Testing framework design' (P44-P47)

Version: npj Precision Oncology 9:141, published 2025-05-15; PMC12078457 full-text XML
Retrieved: 2026-10-09T20:43:36Z

not individually reviewed

No individual claim review recorded

Audit details

Field: attributes.access

Source artifact SHA-256: 09c67fcaf7b367d74500db5fd015968389371dcb9ee5a37ccc83241aa63c0f80

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

attributes.negatives
5266
Context-only references
Benchmarking large language models GPT-4o, llama 3.1, and qwen 2.5 for cancer genetic variant classification

Original source ↗

Lin et al. 2025 Methods 'Dataset' (P35-P38), 'Model selection' (P39-P40), 'System prompts' (P41-P43), 'Testing framework design' (P44-P47)

Version: npj Precision Oncology 9:141, published 2025-05-15; PMC12078457 full-text XML
Retrieved: 2026-10-09T20:43:36Z

not individually reviewed

No individual claim review recorded

Audit details

Field: attributes.negatives

Source artifact SHA-256: 09c67fcaf7b367d74500db5fd015968389371dcb9ee5a37ccc83241aa63c0f80

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

attributes.population
10,506 genetic alterations from 612 patients: 5,240 clinically relevant (Genomic Findings) and 5,266 VUS (appendix)
Context-only references
Benchmarking large language models GPT-4o, llama 3.1, and qwen 2.5 for cancer genetic variant classification

Original source ↗

Lin et al. 2025 Methods 'Dataset' (P35-P38), 'Model selection' (P39-P40), 'System prompts' (P41-P43), 'Testing framework design' (P44-P47)

Version: npj Precision Oncology 9:141, published 2025-05-15; PMC12078457 full-text XML
Retrieved: 2026-10-09T20:43:36Z

not individually reviewed

No individual claim review recorded

Audit details

Field: attributes.population

Source artifact SHA-256: 09c67fcaf7b367d74500db5fd015968389371dcb9ee5a37ccc83241aa63c0f80

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

attributes.positives
5240
Context-only references
Benchmarking large language models GPT-4o, llama 3.1, and qwen 2.5 for cancer genetic variant classification

Original source ↗

Lin et al. 2025 Methods 'Dataset' (P35-P38), 'Model selection' (P39-P40), 'System prompts' (P41-P43), 'Testing framework design' (P44-P47)

Version: npj Precision Oncology 9:141, published 2025-05-15; PMC12078457 full-text XML
Retrieved: 2026-10-09T20:43:36Z

not individually reviewed

No individual claim review recorded

Audit details

Field: attributes.positives

Source artifact SHA-256: 09c67fcaf7b367d74500db5fd015968389371dcb9ee5a37ccc83241aa63c0f80

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

attributes.private_data_not_included
true
Context-only references
Benchmarking large language models GPT-4o, llama 3.1, and qwen 2.5 for cancer genetic variant classification

Original source ↗

Lin et al. 2025 Methods 'Dataset' (P35-P38), 'Model selection' (P39-P40), 'System prompts' (P41-P43), 'Testing framework design' (P44-P47)

Version: npj Precision Oncology 9:141, published 2025-05-15; PMC12078457 full-text XML
Retrieved: 2026-10-09T20:43:36Z

not individually reviewed

No individual claim review recorded

Audit details

Field: attributes.private_data_not_included

Source artifact SHA-256: 09c67fcaf7b367d74500db5fd015968389371dcb9ee5a37ccc83241aa63c0f80

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

attributes.source_locator
Lin et al. 2025 Methods 'Dataset' (P35-P38), 'Model selection' (P39-P40), 'System prompts' (P41-P43), 'Testing framework design' (P44-P47)
Context-only references
Benchmarking large language models GPT-4o, llama 3.1, and qwen 2.5 for cancer genetic variant classification

Original source ↗

Lin et al. 2025 Methods 'Dataset' (P35-P38), 'Model selection' (P39-P40), 'System prompts' (P41-P43), 'Testing framework design' (P44-P47)

Version: npj Precision Oncology 9:141, published 2025-05-15; PMC12078457 full-text XML
Retrieved: 2026-10-09T20:43:36Z

not individually reviewed

No individual claim review recorded

Audit details

Field: attributes.source_locator

Source artifact SHA-256: 09c67fcaf7b367d74500db5fd015968389371dcb9ee5a37ccc83241aa63c0f80

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

attributes.total
10506
Context-only references
Benchmarking large language models GPT-4o, llama 3.1, and qwen 2.5 for cancer genetic variant classification

Original source ↗

Lin et al. 2025 Methods 'Dataset' (P35-P38), 'Model selection' (P39-P40), 'System prompts' (P41-P43), 'Testing framework design' (P44-P47)

Version: npj Precision Oncology 9:141, published 2025-05-15; PMC12078457 full-text XML
Retrieved: 2026-10-09T20:43:36Z

not individually reviewed

No individual claim review recorded

Audit details

Field: attributes.total

Source artifact SHA-256: 09c67fcaf7b367d74500db5fd015968389371dcb9ee5a37ccc83241aa63c0f80

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

attributes.version
Single-hospital FoundationOne CDx reports as used by Lin et al. 2025
Context-only references
Benchmarking large language models GPT-4o, llama 3.1, and qwen 2.5 for cancer genetic variant classification

Original source ↗

Lin et al. 2025 Methods 'Dataset' (P35-P38), 'Model selection' (P39-P40), 'System prompts' (P41-P43), 'Testing framework design' (P44-P47)

Version: npj Precision Oncology 9:141, published 2025-05-15; PMC12078457 full-text XML
Retrieved: 2026-10-09T20:43:36Z

not individually reviewed

No individual claim review recorded

Audit details

Field: attributes.version

Source artifact SHA-256: 09c67fcaf7b367d74500db5fd015968389371dcb9ee5a37ccc83241aa63c0f80

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

description
Real-world variants labelled clinically relevant or VUS by the FoundationOne CDx report section they appear in.
Context-only references
Benchmarking large language models GPT-4o, llama 3.1, and qwen 2.5 for cancer genetic variant classification

Original source ↗

Lin et al. 2025 Methods 'Dataset' (P35-P38), 'Model selection' (P39-P40), 'System prompts' (P41-P43), 'Testing framework design' (P44-P47)

Version: npj Precision Oncology 9:141, published 2025-05-15; PMC12078457 full-text XML
Retrieved: 2026-10-09T20:43:36Z

not individually reviewed

No individual claim review recorded

Audit details

Field: description

Source artifact SHA-256: 09c67fcaf7b367d74500db5fd015968389371dcb9ee5a37ccc83241aa63c0f80

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

name
FoundationOne CDx report variants, 612 patients (Lin et al. 2025)
Context-only references
Benchmarking large language models GPT-4o, llama 3.1, and qwen 2.5 for cancer genetic variant classification

Original source ↗

Lin et al. 2025 Methods 'Dataset' (P35-P38), 'Model selection' (P39-P40), 'System prompts' (P41-P43), 'Testing framework design' (P44-P47)

Version: npj Precision Oncology 9:141, published 2025-05-15; PMC12078457 full-text XML
Retrieved: 2026-10-09T20:43:36Z

not individually reviewed

No individual claim review recorded

Audit details

Field: name

Source artifact SHA-256: 09c67fcaf7b367d74500db5fd015968389371dcb9ee5a37ccc83241aa63c0f80

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

Release 2026-10-09-ba02f2f4a36e · Record review: source checked

1 source records and release historyDownload this release (gzip)
Technical metadata and extraction receipts

Stable ID: egfrnsclc-20261009-data-lin2025-foundationone-variants

areas
dna-genomes
contexts
clinical_research
version
Single-hospital FoundationOne CDx reports as used by Lin et al. 2025
total
10506
positives
5240
negatives
5266
population
10,506 genetic alterations from 612 patients: 5,240 clinically relevant (Genomic Findings) and 5,266 VUS (appendix)
access
Hospital data; not published
private data not included
true
source locator
Lin et al. 2025 Methods 'Dataset' (P35-P38), 'Model selection' (P39-P40), 'System prompts' (P41-P43), 'Testing framework design' (P44-P47)
Related records

Suggest a correction