0.248 top-1-accuracy
lin2025-qwen-basic-civic top-1-accuracy (CIViC evidence level; mean over iterations)
- Tested configuration
- Qwen 2.5 72B, basic prompt, temperature 0.8 (default) (Lin et al. 2025)
- Protocol
- Lin et al. 2025 CIViC evidence-level assignment
- Dataset
- CIViC clinical evidence summary, 4,426 variant-disease associations (accessed 2024-11-20)
- Procedure
- egfrnsclc-20261009-protocol-lin2025-civic-level-assignment
- Evaluation
- qwen-basic on civic (Lin et al. 2025)
- Coverage
- Not reported scored / Not reported eligible
- Uncertainty
- 95% CI 0.2477 to 0.2492
- Evidence
- Independent external evaluation · source checkedBenchmarking large language models GPT-4o, llama 3.1, and qwen 2.5 for cancer genetic variant classification · Table 1 (Tab1), row 'CIViC', column 'Qwen 2.5' Mean accuracy; 95% CI in the next column
A source-checked result verifies the numerical transcription, not every model or protocol detail. Evaluation metadata: source checked. Source checked does not mean independently reproduced.
Reproduction
- Split
- Whole table; no training split (prompted models)
- Adaptation
- Prompting only
- Scoring implementation
- Top-1 answer compared with the reference level
No execution recipe has been verified for this exact configuration and evaluation. A benchmark's general instructions may use different inputs, splits or model settings.
Reproducing this published result requires matching its model configuration, data, split and scorer. Source checking or a successful smoke test does not establish score reproduction.
Evidence
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Evidence table
Inspect claims, sources and review details
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
1 evidence row matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Reported result 0.2485 Individual claims | Benchmarking large language models GPT-4o, llama 3.1, and qwen 2.5 for cancer genetic variant classification Table 1 (Tab1), row 'CIViC', column 'Qwen 2.5' Mean accuracy; 95% CI in the next column Version: npj Precision Oncology 9:141, published 2025-05-15; PMC12078457 full-text XML | source checked ["source-hash-verification","deterministic-table-parse","independent-cell-check"] · 2026-10-09 independent paper Audit detailsExtracted by deterministic parse of the article JATS XML tables Tab1 and Tab2 (extract/extract_egfr_nsclc.py) with header and row labels asserted. Independent review 2026-10-09: value and identity match the source. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record Extraction artifact SHA-256: |
Sources and history
Release 2026-10-09-ba02f2f4a36e · Record review: source checked
1 source records and release history
- Benchmarking large language models GPT-4o, llama 3.1, and qwen 2.5 for cancer genetic variant classification · Original source · npj Precision Oncology 9:141, published 2025-05-15; PMC12078457 full-text XML
Technical metadata and extraction receipts
Stable ID: egfrnsclc-20261009-result-lin2025-qwen-basic-civic-top1
- metric
- top-1-accuracy
- metric qualifier
- CIViC evidence level; mean over 100 iterations
- metric direction
- higher
- unit
- fraction
- printed value
- 0.2485
- numeric value
- 0.2485
- source locator
- Table 1 (Tab1), row 'CIViC', column 'Qwen 2.5' Mean accuracy; 95% CI in the next column
- review
- method: source-hash-verification; deterministic-table-parse; independent-cell-check; method note: Read Table 1 and Table 2 from the article XML with a separate parser; checked value, 95% CI, p value, metric, qualifier, unit, direction and the linked configuration and protocol against the row and column headers.; reviewer: claude; reviewer note: Separate Claude review agent, independent of the extractor; no human review claimed; date: 2026-10-09; artifact sha256: 09c67fcaf7b367d74500db5fd015968389371dcb9ee5a37ccc83241aa63c0f80; retrieval url: https://www.ebi.ac.uk/europepmc/webservices/rest/PMC12078457/fullTextXML; note: Extracted by deterministic parse of the article JATS XML tables Tab1 and Tab2 (extract/extract_egfr_nsclc.py) with header and row labels asserted. Independent review 2026-10-09: value and identity match the source.
- uncertainty
- type: confidence_interval; lower: 0.2477; upper: 0.2492; level: 0.95; printed: 0.2477–0.2492
- reported p value
- <0.001 (ANOVA across the three models, Methods P52; printed in the row's 'p value' column)
Related records
- evaluation: qwen-basic on civic (Lin et al. 2025)