rewirebio.iobenchmarks
Protocol

Lin et al. 2025 OncoKB level-of-evidence assignment

LLM evidence-level classification task from Lin et al. 2025.

6 evaluations · 6 results

Overview

LLM evidence-level classification task from Lin et al. 2025.

Consult the linked sources for architecture or protocol details. Missing evidence is not evidence of a missing capability.

6 recorded evaluations, 6 metric rows. A comparison chart has not yet been validated for these results. The table retains the individual findings and their sources.

View coverage and remaining gaps across all benchmarks

Results

Results are available, but no reviewed comparison panel is linked in this release.

All evaluations

6 evaluations · 6 results. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: GPT-4o, basic prompt, temperature 1.0 (default) (Lin et al. 2025)Protocol: Lin et al. 2025 OncoKB level-of-evidence assignment
Dataset: OncoKB actionable-genes table, 625 variant-cancer-drug associations (accessed 2024-11-20)
0.339 top-1-accuracy
fraction · higher

Uncertainty: 95% CI 0.3369 to 0.3417

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

gpt-4o-basic on oncokb (Lin et al. 2025)

egfrnsclc-20261009-protocol-lin2025-oncokb-level-assignment

Aggregation: Not reported

Benchmarking large language models GPT-4o, llama 3.1, and qwen 2.5 for cancer genetic variant classification · Table 1 (Tab1), row 'OncoKB', column 'GPT-4o' Mean accuracy; 95% CI in the next column
Configuration: Llama 3.1 70B, basic prompt, temperature 0.4 (Lin et al. 2025)Protocol: Lin et al. 2025 OncoKB level-of-evidence assignment
Dataset: OncoKB actionable-genes table, 625 variant-cancer-drug associations (accessed 2024-11-20)
0.318 top-1-accuracy
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

llama-basic-t0-4 on oncokb (Lin et al. 2025)

egfrnsclc-20261009-protocol-lin2025-oncokb-level-assignment

Aggregation: Not reported

Benchmarking large language models GPT-4o, llama 3.1, and qwen 2.5 for cancer genetic variant classification · Table 2 (Tab2), row 10 ('Llama 3.1', 'Basic prompt + temperature (0.4)', 'OncoKB'), column 'Accuracy'
Configuration: Llama 3.1 70B, basic prompt, temperature 0.8 (default) (Lin et al. 2025)Protocol: Lin et al. 2025 OncoKB level-of-evidence assignment
Dataset: OncoKB actionable-genes table, 625 variant-cancer-drug associations (accessed 2024-11-20)
0.307 top-1-accuracy
fraction · higher

Uncertainty: 95% CI 0.3041 to 0.309

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

llama-basic-t0-8 on oncokb (Lin et al. 2025)

egfrnsclc-20261009-protocol-lin2025-oncokb-level-assignment

Aggregation: Not reported

Benchmarking large language models GPT-4o, llama 3.1, and qwen 2.5 for cancer genetic variant classification · Table 1 (Tab1), row 'OncoKB', column 'Llama 3' Mean accuracy; 95% CI in the next column
Configuration: Llama 3.1 70B, basic prompt, temperature 0 (Lin et al. 2025)Protocol: Lin et al. 2025 OncoKB level-of-evidence assignment
Dataset: OncoKB actionable-genes table, 625 variant-cancer-drug associations (accessed 2024-11-20)
0.331 top-1-accuracy
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

llama-basic-t0 on oncokb (Lin et al. 2025)

egfrnsclc-20261009-protocol-lin2025-oncokb-level-assignment

Aggregation: Not reported

Benchmarking large language models GPT-4o, llama 3.1, and qwen 2.5 for cancer genetic variant classification · Table 2 (Tab2), row 11 ('Llama 3.1', 'Basic prompt + temperature (0)', 'OncoKB'), column 'Accuracy'
Configuration: Qwen 2.5 72B, basic prompt, temperature 0.8 (default) (Lin et al. 2025)Protocol: Lin et al. 2025 OncoKB level-of-evidence assignment
Dataset: OncoKB actionable-genes table, 625 variant-cancer-drug associations (accessed 2024-11-20)
0.333 top-1-accuracy
fraction · higher

Uncertainty: 95% CI 0.3316 to 0.334

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

qwen-basic on oncokb (Lin et al. 2025)

egfrnsclc-20261009-protocol-lin2025-oncokb-level-assignment

Aggregation: Not reported

Benchmarking large language models GPT-4o, llama 3.1, and qwen 2.5 for cancer genetic variant classification · Table 1 (Tab1), row 'OncoKB', column 'Qwen 2.5' Mean accuracy; 95% CI in the next column
Configuration: Qwen 2.5 72B, refined prompt, temperature 0.8 (default) (Lin et al. 2025)Protocol: Lin et al. 2025 OncoKB level-of-evidence assignment
Dataset: OncoKB actionable-genes table, 625 variant-cancer-drug associations (accessed 2024-11-20)
0.299 top-1-accuracy
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

qwen-refined on oncokb (Lin et al. 2025)

egfrnsclc-20261009-protocol-lin2025-oncokb-level-assignment

Aggregation: Not reported

Benchmarking large language models GPT-4o, llama 3.1, and qwen 2.5 for cancer genetic variant classification · Table 2 (Tab2), row 6 ('Qwen2.5', 'Refined prompt + Default temperature (0.8)', 'OncoKB'), column 'Accuracy'

Source checking is not independent reproduction. Release 2026-10-09-ba02f2f4a36e.

Methods and evaluation design

Procedure, tasks and evaluated configurations

Recorded evaluations

Each evaluation records what was tested and under which conditions.

Baseline coverage

Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.

0 of 2 active baseline roles have published Rewire measurements in this release. Measurements on a selected protocol do not establish coverage of an entire suite.

No execution recipe linked to this protocol. Recipe availability does not establish a completed evaluation.

External evaluations
6

Literature evidence is not a Rewire measurement. Executed but unpublished runs and private review status are not included.

Null control

Proposed control: requires review

Training-set class prior where supervised fitting is permitted

Protocol-specific applicability, permitted inputs, access, split, evaluator and execution requirements need review before implementation or execution.

This is a suggested selection rule, not a validated method or a measured score.

Conventional reference

Proposed control: requires review

Regularised classifier on simple permitted features, or protocol's conventional reference

Protocol-specific applicability, permitted inputs, access, split, evaluator and execution requirements need review before implementation or execution.

This is a suggested selection rule, not a validated method or a measured score.

Protocol coverage CSV (gzip) · Model evaluation matrix (gzip) · Source table (gzip) · Release and checksums (gzip)

Coverage is derived from release 2026-10-09-ba02f2f4a36e. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.

Run instructions

No runnable recipe has been reviewed for this protocol. Dataset access, model requirements, licences and compute requirements must be checked against its sources before execution.

Strengths, limitations and unresolved questions

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

0 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-10-09-ba02f2f4a36e
Property and statementOriginal source and locationReview and provenance

No evidence rows match these filters. Choose another scope or clear the search.

Sources and history

Release 2026-10-09-ba02f2f4a36e · Record review: source checked

1 source records and release historyDownload this release (gzip)
Technical metadata and extraction receipts

Stable ID: egfrnsclc-20261009-protocol-lin2025-oncokb-level-assignment

areas
dna-genomes
contexts
clinical_research
protocol
Ask each LLM to assign the OncoKB level (1, 2, 3, 4, R1, R2 or VUS) for each gene, alteration and cancer type of the OncoKB actionable-genes table. Reference: OncoKB level in the actionable-genes table. Score the top-1 answer (up to three levels may be returned); repeat each query 100 times for basic prompts and 10 times for other settings; report mean accuracy.
version
Lin et al. 2025 Methods
metric
top-1-accuracy
metric direction
higher
limitations
Pan-cancer variant set; no EGFR- or NSCLC-specific score is reported.; The model assigns an evidence level from gene, alteration and tumour type alone; no prior therapy, line or evidence date is given and no source is cited.; Public OncoKB and CIViC tables may be in the models' training data (Discussion P33).; Queries name the gene, alteration and tumour type but no drug, while the reference table assigns levels per drug or evidence item; the paper does not say how a query with several reference levels was scored, so sensitivity and resistance levels for the same variant cannot be told apart.; The 95% confidence intervals span 0.0004 to 0.0049 in width and appear to describe variation between repeated runs of the same queries, not uncertainty from the choice of variants.
source locator
Lin et al. 2025 Methods 'Dataset' (P35-P38), 'Model selection' (P39-P40), 'System prompts' (P41-P43), 'Testing framework design' (P44-P47); Tables 1-2
Related records

Suggest a correction