rewirebio.iobenchmarks
Result

37/40 (92.5%) appropriate_retrieval_rate_postcutoff

CIViC-Fact v3 post-cutoff retrieval: fine-tuned MedCPT cross-encoder, appropriate-content rate

Tested configuration
Fine-tuned MedCPT cross-encoder (passage retrieval)
Protocol
CIViC-Fact v3 within-linked-publication passage retrieval, post-cutoff cohort
Dataset
CIViC-Fact v3 post-cutoff temporal-evaluation cohort, 2026-03-03 to 2026-06-09
Procedure
Not reported
Evaluation
CIViC-Fact v3 post-cutoff within-linked-publication retrieval: fine-tuned MedCPT cross-encoder
Coverage
scored: unreported; eligible: unreported
Uncertainty
Not reported
Evidence
Author-reported evaluation · source checkedCIViC-Fact: a proof-of-concept framework for AI-assisted verification of cancer variant interpretations (bioRxiv v3) · Results, 'Passage Retrieval Models Perform Well without Fine-Tuning' subsection (printed page 14 / PDF page 15): '...both the fine-tuned MedCPT model and the 8B Qwen 3 reranker performed similarly retrieving appropriate content for 37 of the remaining 40 entries (92.5%).'

A source-checked result verifies the numerical transcription, not every model or protocol detail. Evaluation metadata: source checked. Source checked does not mean independently reproduced.

Reproduction

Split
Not reported
Adaptation
Not reported
Scoring implementation
Manual review judgment of 'appropriate' (sufficient) retrieved content per entry; source-defined, not an automated or recomputed metric. Reviewer count, blinding and replicate/seed structure for this judgment are UNREPORTED; the source's separately described curator/group-consensus process (Methods) covers reference-label assignment for the cohort, not this retrieval judgment, and neither is assumed here.

No execution recipe has been verified for this exact configuration and evaluation. A benchmark's general instructions may use different inputs, splits or model settings.

Reproducing this published result requires matching its model configuration, data, split and scorer. Source checking or a successful smoke test does not establish score reproduction.

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

1 evidence row matching the loaded filters

Claims, original sources and review scope · Release 2026-10-07-1448159e6a81
Property and statementOriginal source and locationReview and provenance
Reported result
37/40 (92.5%)
Individual claims
CIViC-Fact: a proof-of-concept framework for AI-assisted verification of cancer variant interpretations (bioRxiv v3)

Original source ↗

Results, 'Passage Retrieval Models Perform Well without Fine-Tuning' subsection (printed page 14 / PDF page 15): '...both the fine-tuned MedCPT model and the 8B Qwen 3 reranker performed similarly retrieving appropriate content for 37 of the remaining 40 entries (92.5%).'

Version: bioRxiv preprint doi:10.1101/2025.09.10.675443, Version 3, posted 2026-08-29 (per in-document header and bioRxiv API version metadata); published: NA.
Retrieved: 2026-10-07T09:28:49Z

source checked

automated source review · 2026-10-07

author reported

Audit details

Source-backed literature curation with independent automated transcription and scope checks, independently reviewed by Codex across two review cycles before ingestion. No new model execution, independent experimental replication, qualified human scientific review or clinical validation.

Field: attributes.printed_value

Source artifact SHA-256: 6beb79ede82f7263a06eafc844bcb4ae41538323e3c9a9a7fa5420f08a8add55

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-10-07-1448159e6a81 · Record review: source checked

1 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: ucc-clinical-egfr-result-medcpt-ft-postcutoff-appropriate

areas
dna-genomes
contexts
clinical_research
metric
appropriate_retrieval_rate_postcutoff
unit
percent
printed value
37/40 (92.5%)
numeric value
92.5
metric direction
higher
uncertainty
Not reported
numerator
37
denominator
40
source locator
Results, 'Passage Retrieval Models Perform Well without Fine-Tuning' subsection (printed page 14 / PDF page 15): '...both the fine-tuned MedCPT model and the 8B Qwen 3 reranker performed similarly retrieving appropriate content for 37 of the remaining 40 entries (92.5%).'
review
method: automated_source_review; reviewer: Claude Sonnet EGFR-evidence-research worker; date: 2026-10-07; note: Source-backed literature curation with independent automated transcription and scope checks, independently reviewed by Codex across two review cycles before ingestion. No new model execution, independent experimental replication, qualified human scientific review or clinical validation.; source id: ucc-clinical-egfr-source-civicfact-v3
source discrepancy
Not reported
missing metadata
uncertainty: Not reported for this endpoint; no confidence interval is recomputed here.; reviewer count: Unreported for this manual retrieval-success judgment.; blinding: Unreported for this manual retrieval-success judgment.; replicates or seeds: Unreported; retrieval is a single reported pass per configuration on this cohort.
scope note
Both the fine-tuned MedCPT and the Qwen3-Reranker-8B configurations are reported as achieving the identical 37/40 (92.5%) figure on the same 40-entry post-cutoff cohort. This is NOT an EGFR- or NSCLC-specific score. This is NOT a clinical efficacy result. This is NOT the static stance-classification accuracy (89%/85% under the v2 SUPPORTS/REFUTES/NEI schema, or any v2 BGE-retriever figure); those are separate, version-specific figures and are not ingested here.
Related records

Suggest a correction