rewire.itbenchmarks
Protocol

CIViC evidence-direction retrieval, March 2026

Compare expected CIViC evidence-direction labels for exact molecular-profile, disease and therapy triplets. Exclude No Evidence from performance aggregation; distinguish support, non-support, contradictory and absent evidence. Micro-average significance labels within evidence types.

3 evaluations · 162 results

Overview

Compare expected CIViC evidence-direction labels for exact molecular-profile, disease and therapy triplets. Exclude No Evidence from performance aggregation; distinguish support, non-support, contradictory and absent evidence. Micro-average significance labels within evidence types.

Consult the linked sources for architecture or protocol details. Missing evidence is not evidence of a missing capability.

3 recorded evaluations, 162 metric rows. A comparison chart has not yet been validated for these results. The table retains the individual findings and their sources.

View coverage and remaining gaps across all benchmarks

Results

Results are available, but no reviewed comparison panel is linked in this release.

All evaluations

3 evaluations · 162 results. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
0.83 overall_accuracy
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP: integrating large language models with the Clinical Interpretations of Variants in Cancer · Fig1c Overall Accuracy; Section 3
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
0.84 diagnostic_micro_f1
F1 score · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP: integrating large language models with the Clinical Interpretations of Variants in Cancer · Figure 1c, diagnostic row, agent column
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
0 diagnostic_negative_f1
F1 score · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Diagnostic / Negative, f1
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
0 diagnostic_negative_precision
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Diagnostic / Negative, precision
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
0 diagnostic_negative_recall
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Diagnostic / Negative, recall
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
0.94 diagnostic_positive_f1
F1 score · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Diagnostic / Positive, f1
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
0.89 diagnostic_positive_precision
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Diagnostic / Positive, precision
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
1 diagnostic_positive_recall
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Diagnostic / Positive, recall
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
1 functional_dominant_negative_f1
F1 score · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Dominant Negative, f1
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
1 functional_dominant_negative_precision
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Dominant Negative, precision
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
1 functional_dominant_negative_recall
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Dominant Negative, recall
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
0.88 functional_micro_f1
F1 score · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP: integrating large language models with the Clinical Interpretations of Variants in Cancer · Figure 1c, functional row, agent column
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
0.67 functional_gain_of_function_f1
F1 score · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Gain of Function, f1
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
0.75 functional_gain_of_function_precision
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Gain of Function, precision
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
0.6 functional_gain_of_function_recall
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Gain of Function, recall
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
0.92 functional_loss_of_function_f1
F1 score · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Loss of Function, f1
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
0.86 functional_loss_of_function_precision
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Loss of Function, precision
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
1 functional_loss_of_function_recall
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Loss of Function, recall
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
1 functional_unaltered_function_f1
F1 score · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Unaltered Function, f1
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
1 functional_unaltered_function_precision
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Unaltered Function, precision
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
1 functional_unaltered_function_recall
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Unaltered Function, recall
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
0.91 micro_f1
F1 score · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP: integrating large language models with the Clinical Interpretations of Variants in Cancer · Fig1c Micro F1; Section 3
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
0.96 oncogenic_micro_f1
F1 score · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP: integrating large language models with the Clinical Interpretations of Variants in Cancer · Figure 1c, oncogenic row, agent column
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
0.96 oncogenic_oncogenicity_f1
F1 score · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Oncogenic / Oncogenicity, f1
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
0.92 oncogenic_oncogenicity_precision
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Oncogenic / Oncogenicity, precision

Source checking is not independent reproduction. Release 2026-09-30-e37e3ab1284d.

Methods and evaluation design

Procedure, tasks and evaluated configurations

Evaluation design

Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.

These source-backed links do not make different protocols or scores interchangeable.

Recorded evaluations

Each evaluation records what was tested and under which conditions.

Baseline coverage

Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.

0 of 2 active baseline roles have published Rewire measurements in this release. Measurements on a selected protocol do not establish coverage of an entire suite.

No execution recipe linked to this protocol. Recipe availability does not establish a completed evaluation.

Author-reported evaluations
3

Literature evidence is not a Rewire measurement. Executed but unpublished runs and private review status are not included.

Null control

Proposed control: requires review

Select a task-valid null control after reviewing inputs and metric

Protocol-specific applicability, permitted inputs, access, split, evaluator and execution requirements need review before implementation or execution.

This is a suggested selection rule, not a validated method or a measured score.

Conventional reference

Proposed control: requires review

Select an upstream conventional reference after reviewing the full protocol

Protocol-specific applicability, permitted inputs, access, split, evaluator and execution requirements need review before implementation or execution.

This is a suggested selection rule, not a validated method or a measured score.

Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums

Coverage is derived from release 2026-09-30-e37e3ab1284d. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.

Run instructions

No runnable recipe has been reviewed for this protocol. Dataset access, model requirements, licences and compute requirements must be checked against its sources before execution.

Strengths, limitations and unresolved questions
Applicable tests and references

Applicability is distinct from a completed evaluation.

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

2 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-30-e37e3ab1284d
Property and statementOriginal source and locationReview and provenance
Relationship: part of
uc-clinical-20260930-benchmark-civic-study
Individual claims
CIViC MCP: integrating large language models with the Clinical Interpretations of Variants in Cancer

Original source ↗

Figure1c; Supplementary Tables1–3 and implementation methods

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: Final article corrected and typeset 12 August 2026
Retrieved: 2026-09-30T21:24:15Z

source checked

automated source review · 2026-09-30

Audit details

Transcription from inspected primary source. No execution, independent replication or human scientific review.

Field: links:part_of:uc-clinical-20260930-benchmark-civic-study

Claim: uc-clinical-20260930-membership-civic-retrieval-protocol

Source artifact SHA-256: 1bebf491323f252d27539a73da04fab4dacf7e3f83b531ee9262645f79e40bad

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Relationship: part of
uc-clinical-20260930-benchmark-civic-study
Individual claims
CIViC MCP supplementary methods and complete Tables 1–3

Original source ↗

Figure1c; Supplementary Tables1–3 and implementation methods

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: vbag209_supplementary_data.docx, final article supplement
Retrieved: 2026-09-30T21:25:48Z

source checked

automated source review · 2026-09-30

Audit details

Transcription from inspected primary source. No execution, independent replication or human scientific review.

Field: links:part_of:uc-clinical-20260930-benchmark-civic-study

Claim: uc-clinical-20260930-membership-civic-retrieval-protocol

Source artifact SHA-256: 857b272f9f29b9cb35d4fba7f424ef31d664d8acf8a23475c7c626a452b2e9e6

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-30-e37e3ab1284d · Record review: source checked

1 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: uc-clinical-20260930-civic-retrieval-protocol

areas
dna-genomes
contexts
clinical_research
protocol
Compare expected CIViC evidence-direction labels for exact molecular-profile, disease and therapy triplets. Exclude No Evidence from performance aggregation; distinguish support, non-support, contradictory and absent evidence. Micro-average significance labels within evidence types.
review
method: automated_source_review; reviewer: Codex clinical coverage worker; date: 2026-09-30; note: Transcription from inspected primary source. No execution, independent replication or human scientific review.; source id: uc-clinical-20260930-source-civic-supp
limitations
Pan-cancer retrieval task; no measured EGFR-mutant advanced NSCLC subgroup.; No treatment-line, prior-therapy, jurisdiction or complete primary-literature recall benchmark.; Live March 2026 CIViC source; current reruns may differ.; Supplement describes four outcomes although body says three.; UI Agent Mode model/version and concurrent timings differ from API configurations.
Related records

Suggest a correction