rewirebio.iobenchmarks
Result

0.339 top-1-accuracy

lin2025-gpt-4o-basic-oncokb top-1-accuracy (OncoKB level of evidence; mean over iterations)

Tested configuration
GPT-4o, basic prompt, temperature 1.0 (default) (Lin et al. 2025)
Protocol
Lin et al. 2025 OncoKB level-of-evidence assignment
Dataset
OncoKB actionable-genes table, 625 variant-cancer-drug associations (accessed 2024-11-20)
Procedure
egfrnsclc-20261009-protocol-lin2025-oncokb-level-assignment
Evaluation
gpt-4o-basic on oncokb (Lin et al. 2025)
Coverage
Not reported scored / Not reported eligible
Uncertainty
95% CI 0.3369 to 0.3417
Evidence
Independent external evaluation · source checkedBenchmarking large language models GPT-4o, llama 3.1, and qwen 2.5 for cancer genetic variant classification · Table 1 (Tab1), row 'OncoKB', column 'GPT-4o' Mean accuracy; 95% CI in the next column

A source-checked result verifies the numerical transcription, not every model or protocol detail. Evaluation metadata: source checked. Source checked does not mean independently reproduced.

Reproduction

Split
Whole table; no training split (prompted models)
Adaptation
Prompting only
Scoring implementation
Top-1 answer compared with the reference level

No execution recipe has been verified for this exact configuration and evaluation. A benchmark's general instructions may use different inputs, splits or model settings.

Reproducing this published result requires matching its model configuration, data, split and scorer. Source checking or a successful smoke test does not establish score reproduction.

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

1 evidence row matching the loaded filters

Claims, original sources and review scope · Release 2026-10-09-ba02f2f4a36e
Property and statementOriginal source and locationReview and provenance
Reported result
0.3393
Individual claims
Benchmarking large language models GPT-4o, llama 3.1, and qwen 2.5 for cancer genetic variant classification

Original source ↗

Table 1 (Tab1), row 'OncoKB', column 'GPT-4o' Mean accuracy; 95% CI in the next column

Version: npj Precision Oncology 9:141, published 2025-05-15; PMC12078457 full-text XML
Retrieved: 2026-10-09T20:43:36Z

source checked

["source-hash-verification","deterministic-table-parse","independent-cell-check"] · 2026-10-09

independent paper

Audit details

Extracted by deterministic parse of the article JATS XML tables Tab1 and Tab2 (extract/extract_egfr_nsclc.py) with header and row labels asserted. Independent review 2026-10-09: value and identity match the source.

Field: attributes.printed_value

Source artifact SHA-256: 09c67fcaf7b367d74500db5fd015968389371dcb9ee5a37ccc83241aa63c0f80

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Extraction artifact SHA-256: 09c67fcaf7b367d74500db5fd015968389371dcb9ee5a37ccc83241aa63c0f80

Extraction artifact

Sources and history

Release 2026-10-09-ba02f2f4a36e · Record review: source checked

1 source records and release historyDownload this release (gzip)
Technical metadata and extraction receipts

Stable ID: egfrnsclc-20261009-result-lin2025-gpt-4o-basic-oncokb-top1

metric
top-1-accuracy
metric qualifier
OncoKB level of evidence; mean over 100 iterations
metric direction
higher
unit
fraction
printed value
0.3393
numeric value
0.3393
source locator
Table 1 (Tab1), row 'OncoKB', column 'GPT-4o' Mean accuracy; 95% CI in the next column
review
method: source-hash-verification; deterministic-table-parse; independent-cell-check; method note: Read Table 1 and Table 2 from the article XML with a separate parser; checked value, 95% CI, p value, metric, qualifier, unit, direction and the linked configuration and protocol against the row and column headers.; reviewer: claude; reviewer note: Separate Claude review agent, independent of the extractor; no human review claimed; date: 2026-10-09; artifact sha256: 09c67fcaf7b367d74500db5fd015968389371dcb9ee5a37ccc83241aa63c0f80; retrieval url: https://www.ebi.ac.uk/europepmc/webservices/rest/PMC12078457/fullTextXML; note: Extracted by deterministic parse of the article JATS XML tables Tab1 and Tab2 (extract/extract_egfr_nsclc.py) with header and row labels asserted. Independent review 2026-10-09: value and identity match the source.
uncertainty
type: confidence_interval; lower: 0.3369; upper: 0.3417; level: 0.95; printed: 0.3369–0.3417
reported p value
<0.001 (ANOVA across the three models, Methods P52; printed in the row's 'p value' column)
Related records

Suggest a correction