CIViC evidence-direction retrieval, March 2026
Compare expected CIViC evidence-direction labels for exact molecular-profile, disease and therapy triplets. Exclude No Evidence from performance aggregation; distinguish support, non-support, contradictory and absent evidence. Micro-average significance labels within evidence types.
Overview
Compare expected CIViC evidence-direction labels for exact molecular-profile, disease and therapy triplets. Exclude No Evidence from performance aggregation; distinguish support, non-support, contradictory and absent evidence. Micro-average significance labels within evidence types.
Consult the linked sources for architecture or protocol details. Missing evidence is not evidence of a missing capability.
3 recorded evaluations, 162 metric rows. A comparison chart has not yet been validated for these results. The table retains the individual findings and their sources.
Results
Results are available, but no reviewed comparison panel is linked in this release.
All evaluations
3 evaluations · 162 results. Different protocols are not a single leaderboard.
Filter evaluations
Applied filters: All linked evaluations
| Tested configuration | Protocol and dataset | Finding | Evidence and details |
|---|---|---|---|
| Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluation | Protocol: CIViC evidence-direction retrieval, March 2026 Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026 | 0.83 overall_accuracy fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNot reported Aggregation: Not reported CIViC MCP: integrating large language models with the Clinical Interpretations of Variants in Cancer · Fig1c Overall Accuracy; Section 3 |
| Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluation | Protocol: CIViC evidence-direction retrieval, March 2026 Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026 | 0.84 diagnostic_micro_f1 F1 score · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNot reported Aggregation: Not reported CIViC MCP: integrating large language models with the Clinical Interpretations of Variants in Cancer · Figure 1c, diagnostic row, agent column |
| Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluation | Protocol: CIViC evidence-direction retrieval, March 2026 Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026 | 0 diagnostic_negative_f1 F1 score · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNot reported Aggregation: Not reported CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Diagnostic / Negative, f1 |
| Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluation | Protocol: CIViC evidence-direction retrieval, March 2026 Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026 | 0 diagnostic_negative_precision fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNot reported Aggregation: Not reported CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Diagnostic / Negative, precision |
| Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluation | Protocol: CIViC evidence-direction retrieval, March 2026 Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026 | 0 diagnostic_negative_recall fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNot reported Aggregation: Not reported CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Diagnostic / Negative, recall |
| Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluation | Protocol: CIViC evidence-direction retrieval, March 2026 Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026 | 0.94 diagnostic_positive_f1 F1 score · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNot reported Aggregation: Not reported CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Diagnostic / Positive, f1 |
| Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluation | Protocol: CIViC evidence-direction retrieval, March 2026 Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026 | 0.89 diagnostic_positive_precision fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNot reported Aggregation: Not reported CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Diagnostic / Positive, precision |
| Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluation | Protocol: CIViC evidence-direction retrieval, March 2026 Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026 | 1 diagnostic_positive_recall fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNot reported Aggregation: Not reported CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Diagnostic / Positive, recall |
| Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluation | Protocol: CIViC evidence-direction retrieval, March 2026 Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026 | 1 functional_dominant_negative_f1 F1 score · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNot reported Aggregation: Not reported CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Dominant Negative, f1 |
| Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluation | Protocol: CIViC evidence-direction retrieval, March 2026 Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026 | 1 functional_dominant_negative_precision fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNot reported Aggregation: Not reported CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Dominant Negative, precision |
| Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluation | Protocol: CIViC evidence-direction retrieval, March 2026 Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026 | 1 functional_dominant_negative_recall fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNot reported Aggregation: Not reported CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Dominant Negative, recall |
| Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluation | Protocol: CIViC evidence-direction retrieval, March 2026 Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026 | 0.88 functional_micro_f1 F1 score · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNot reported Aggregation: Not reported CIViC MCP: integrating large language models with the Clinical Interpretations of Variants in Cancer · Figure 1c, functional row, agent column |
| Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluation | Protocol: CIViC evidence-direction retrieval, March 2026 Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026 | 0.67 functional_gain_of_function_f1 F1 score · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNot reported Aggregation: Not reported CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Gain of Function, f1 |
| Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluation | Protocol: CIViC evidence-direction retrieval, March 2026 Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026 | 0.75 functional_gain_of_function_precision fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNot reported Aggregation: Not reported CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Gain of Function, precision |
| Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluation | Protocol: CIViC evidence-direction retrieval, March 2026 Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026 | 0.6 functional_gain_of_function_recall fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNot reported Aggregation: Not reported CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Gain of Function, recall |
| Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluation | Protocol: CIViC evidence-direction retrieval, March 2026 Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026 | 0.92 functional_loss_of_function_f1 F1 score · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNot reported Aggregation: Not reported CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Loss of Function, f1 |
| Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluation | Protocol: CIViC evidence-direction retrieval, March 2026 Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026 | 0.86 functional_loss_of_function_precision fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNot reported Aggregation: Not reported CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Loss of Function, precision |
| Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluation | Protocol: CIViC evidence-direction retrieval, March 2026 Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026 | 1 functional_loss_of_function_recall fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNot reported Aggregation: Not reported CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Loss of Function, recall |
| Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluation | Protocol: CIViC evidence-direction retrieval, March 2026 Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026 | 1 functional_unaltered_function_f1 F1 score · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNot reported Aggregation: Not reported CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Unaltered Function, f1 |
| Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluation | Protocol: CIViC evidence-direction retrieval, March 2026 Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026 | 1 functional_unaltered_function_precision fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNot reported Aggregation: Not reported CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Unaltered Function, precision |
| Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluation | Protocol: CIViC evidence-direction retrieval, March 2026 Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026 | 1 functional_unaltered_function_recall fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNot reported Aggregation: Not reported CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Unaltered Function, recall |
| Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluation | Protocol: CIViC evidence-direction retrieval, March 2026 Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026 | 0.91 micro_f1 F1 score · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNot reported Aggregation: Not reported CIViC MCP: integrating large language models with the Clinical Interpretations of Variants in Cancer · Fig1c Micro F1; Section 3 |
| Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluation | Protocol: CIViC evidence-direction retrieval, March 2026 Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026 | 0.96 oncogenic_micro_f1 F1 score · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNot reported Aggregation: Not reported CIViC MCP: integrating large language models with the Clinical Interpretations of Variants in Cancer · Figure 1c, oncogenic row, agent column |
| Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluation | Protocol: CIViC evidence-direction retrieval, March 2026 Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026 | 0.96 oncogenic_oncogenicity_f1 F1 score · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNot reported Aggregation: Not reported CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Oncogenic / Oncogenicity, f1 |
| Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluation | Protocol: CIViC evidence-direction retrieval, March 2026 Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026 | 0.92 oncogenic_oncogenicity_precision fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNot reported Aggregation: Not reported CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Oncogenic / Oncogenicity, precision |
Source checking is not independent reproduction. Release 2026-09-30-e37e3ab1284d.
Methods and evaluation design
Procedure, tasks and evaluated configurations
Evaluation design
Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.
These source-backed links do not make different protocols or scores interchangeable.
Recorded evaluations
Each evaluation records what was tested and under which conditions.
Baseline coverage
Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.
0 of 2 active baseline roles have published Rewire measurements in this release. Measurements on a selected protocol do not establish coverage of an entire suite.
No execution recipe linked to this protocol. Recipe availability does not establish a completed evaluation.
- Author-reported evaluations
- 3
Literature evidence is not a Rewire measurement. Executed but unpublished runs and private review status are not included.
Null control
Proposed control: requires review
Select a task-valid null control after reviewing inputs and metric
Protocol-specific applicability, permitted inputs, access, split, evaluator and execution requirements need review before implementation or execution.
This is a suggested selection rule, not a validated method or a measured score.
Conventional reference
Proposed control: requires review
Select an upstream conventional reference after reviewing the full protocol
Protocol-specific applicability, permitted inputs, access, split, evaluator and execution requirements need review before implementation or execution.
This is a suggested selection rule, not a validated method or a measured score.
Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums
Coverage is derived from release 2026-09-30-e37e3ab1284d. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.
Run instructions
No runnable recipe has been reviewed for this protocol. Dataset access, model requirements, licences and compute requirements must be checked against its sources before execution.
Strengths, limitations and unresolved questions
Applicable tests and references
Applicability is distinct from a completed evaluation.
- Browser Agent Mode retrieval baseline · source_supported
- Unaided GPT-5 retrieval baseline · source_supported
- Equal-effort EGFR manual lookup baseline (proposed) · proposed
Evidence
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Evidence table
Inspect claims, sources and review details
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
2 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Relationship: part of uc-clinical-20260930-benchmark-civic-study Individual claims | CIViC MCP: integrating large language models with the Clinical Interpretations of Variants in Cancer Figure1c; Supplementary Tables1–3 and implementation methods Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: Final article corrected and typeset 12 August 2026 | source checked automated source review · 2026-09-30 Audit detailsTranscription from inspected primary source. No execution, independent replication or human scientific review. Field: Claim: uc-clinical-20260930-membership-civic-retrieval-protocol Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Relationship: part of uc-clinical-20260930-benchmark-civic-study Individual claims | CIViC MCP supplementary methods and complete Tables 1–3 Figure1c; Supplementary Tables1–3 and implementation methods Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: vbag209_supplementary_data.docx, final article supplement | source checked automated source review · 2026-09-30 Audit detailsTranscription from inspected primary source. No execution, independent replication or human scientific review. Field: Claim: uc-clinical-20260930-membership-civic-retrieval-protocol Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Sources and history
View linked audit checks and correction history
Release 2026-09-30-e37e3ab1284d · Record review: source checked
1 source records and release history
- CIViC MCP supplementary methods and complete Tables 1–3 · Original source · vbag209_supplementary_data.docx, final article supplement
Technical metadata and extraction receipts
Stable ID: uc-clinical-20260930-civic-retrieval-protocol
- areas
- dna-genomes
- contexts
- clinical_research
- protocol
- Compare expected CIViC evidence-direction labels for exact molecular-profile, disease and therapy triplets. Exclude No Evidence from performance aggregation; distinguish support, non-support, contradictory and absent evidence. Micro-average significance labels within evidence types.
- review
- method: automated_source_review; reviewer: Codex clinical coverage worker; date: 2026-09-30; note: Transcription from inspected primary source. No execution, independent replication or human scientific review.; source id: uc-clinical-20260930-source-civic-supp
- limitations
- Pan-cancer retrieval task; no measured EGFR-mutant advanced NSCLC subgroup.; No treatment-line, prior-therapy, jurisdiction or complete primary-literature recall benchmark.; Live March 2026 CIViC source; current reruns may differ.; Supplement describes four outcomes although body says three.; UI Agent Mode model/version and concurrent timings differ from API configurations.
Related records
- part of: CIViC MCP 2026 evidence-retrieval evaluation
- applicable to: Browser Agent Mode retrieval baseline
- applicable to: Unaided GPT-5 retrieval baseline
- applicable to: Equal-effort EGFR manual lookup baseline (proposed)
- protocol: CIViC agent retrieval
- protocol: CIViC alone retrieval
- protocol: CIViC mcp retrieval
- subject: Published study membership: CIViC evidence-direction retrieval, March 2026