37/40 (92.5%) appropriate_retrieval_rate_postcutoff
CIViC-Fact v3 post-cutoff retrieval: fine-tuned MedCPT cross-encoder, appropriate-content rate
- Tested configuration
- Fine-tuned MedCPT cross-encoder (passage retrieval)
- Protocol
- CIViC-Fact v3 within-linked-publication passage retrieval, post-cutoff cohort
- Dataset
- CIViC-Fact v3 post-cutoff temporal-evaluation cohort, 2026-03-03 to 2026-06-09
- Procedure
- Not reported
- Evaluation
- CIViC-Fact v3 post-cutoff within-linked-publication retrieval: fine-tuned MedCPT cross-encoder
- Coverage
- scored: unreported; eligible: unreported
- Uncertainty
- Not reported
- Evidence
- Author-reported evaluation · source checkedCIViC-Fact: a proof-of-concept framework for AI-assisted verification of cancer variant interpretations (bioRxiv v3) · Results, 'Passage Retrieval Models Perform Well without Fine-Tuning' subsection (printed page 14 / PDF page 15): '...both the fine-tuned MedCPT model and the 8B Qwen 3 reranker performed similarly retrieving appropriate content for 37 of the remaining 40 entries (92.5%).'
A source-checked result verifies the numerical transcription, not every model or protocol detail. Evaluation metadata: source checked. Source checked does not mean independently reproduced.
Reproduction
- Split
- Not reported
- Adaptation
- Not reported
- Scoring implementation
- Manual review judgment of 'appropriate' (sufficient) retrieved content per entry; source-defined, not an automated or recomputed metric. Reviewer count, blinding and replicate/seed structure for this judgment are UNREPORTED; the source's separately described curator/group-consensus process (Methods) covers reference-label assignment for the cohort, not this retrieval judgment, and neither is assumed here.
No execution recipe has been verified for this exact configuration and evaluation. A benchmark's general instructions may use different inputs, splits or model settings.
Reproducing this published result requires matching its model configuration, data, split and scorer. Source checking or a successful smoke test does not establish score reproduction.
Evidence
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Evidence table
Inspect claims, sources and review details
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
1 evidence row matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Reported result 37/40 (92.5%) Individual claims | CIViC-Fact: a proof-of-concept framework for AI-assisted verification of cancer variant interpretations (bioRxiv v3) Results, 'Passage Retrieval Models Perform Well without Fine-Tuning' subsection (printed page 14 / PDF page 15): '...both the fine-tuned MedCPT model and the 8B Qwen 3 reranker performed similarly retrieving appropriate content for 37 of the remaining 40 entries (92.5%).' Version: bioRxiv preprint doi:10.1101/2025.09.10.675443, Version 3, posted 2026-08-29 (per in-document header and bioRxiv API version metadata); published: NA. | source checked automated source review · 2026-10-07 author reported Audit detailsSource-backed literature curation with independent automated transcription and scope checks, independently reviewed by Codex across two review cycles before ingestion. No new model execution, independent experimental replication, qualified human scientific review or clinical validation. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Sources and history
View linked audit checks and correction history
Release 2026-10-07-1448159e6a81 · Record review: source checked
1 source records and release history
- CIViC-Fact: a proof-of-concept framework for AI-assisted verification of cancer variant interpretations (bioRxiv v3) · Original source · bioRxiv preprint doi:10.1101/2025.09.10.675443, Version 3, posted 2026-08-29 (per in-document header and bioRxiv API version metadata); published: NA.
Technical metadata and extraction receipts
Stable ID: ucc-clinical-egfr-result-medcpt-ft-postcutoff-appropriate
- areas
- dna-genomes
- contexts
- clinical_research
- metric
- appropriate_retrieval_rate_postcutoff
- unit
- percent
- printed value
- 37/40 (92.5%)
- numeric value
- 92.5
- metric direction
- higher
- uncertainty
- Not reported
- numerator
- 37
- denominator
- 40
- source locator
- Results, 'Passage Retrieval Models Perform Well without Fine-Tuning' subsection (printed page 14 / PDF page 15): '...both the fine-tuned MedCPT model and the 8B Qwen 3 reranker performed similarly retrieving appropriate content for 37 of the remaining 40 entries (92.5%).'
- review
- method: automated_source_review; reviewer: Claude Sonnet EGFR-evidence-research worker; date: 2026-10-07; note: Source-backed literature curation with independent automated transcription and scope checks, independently reviewed by Codex across two review cycles before ingestion. No new model execution, independent experimental replication, qualified human scientific review or clinical validation.; source id: ucc-clinical-egfr-source-civicfact-v3
- source discrepancy
- Not reported
- missing metadata
- uncertainty: Not reported for this endpoint; no confidence interval is recomputed here.; reviewer count: Unreported for this manual retrieval-success judgment.; blinding: Unreported for this manual retrieval-success judgment.; replicates or seeds: Unreported; retrieval is a single reported pass per configuration on this cohort.
- scope note
- Both the fine-tuned MedCPT and the Qwen3-Reranker-8B configurations are reported as achieving the identical 37/40 (92.5%) figure on the same 40-entry post-cutoff cohort. This is NOT an EGFR- or NSCLC-specific score. This is NOT a clinical efficacy result. This is NOT the static stance-classification accuracy (89%/85% under the v2 SUPPORTS/REFUTES/NEI schema, or any v2 BGE-retriever figure); those are separate, version-specific figures and are not ingested here.