rewire.itbenchmarks
Benchmark

CIViC MCP 2026 evidence-retrieval evaluation

Published-study grouping of the catalogued evaluation protocols. This grouping does not claim an executable suite, full-paper extraction or independent clinical validation.

3 evaluations · 162 results

Overview

Published-study grouping of the catalogued evaluation protocols. This grouping does not claim an executable suite, full-paper extraction or independent clinical validation.

Consult the linked sources for architecture or protocol details. Missing evidence is not evidence of a missing capability.

3 recorded evaluations, 162 metric rows. A comparison chart has not yet been validated for these results. The table retains the individual findings and their sources.

View coverage and remaining gaps across all benchmarks

Results

Results are available, but no reviewed comparison panel is linked in this release.

All evaluations

3 evaluations · 162 results. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
0.83 overall_accuracy
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP: integrating large language models with the Clinical Interpretations of Variants in Cancer · Fig1c Overall Accuracy; Section 3
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
0.84 diagnostic_micro_f1
F1 score · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP: integrating large language models with the Clinical Interpretations of Variants in Cancer · Figure 1c, diagnostic row, agent column
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
0 diagnostic_negative_f1
F1 score · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Diagnostic / Negative, f1
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
0 diagnostic_negative_precision
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Diagnostic / Negative, precision
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
0 diagnostic_negative_recall
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Diagnostic / Negative, recall
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
0.94 diagnostic_positive_f1
F1 score · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Diagnostic / Positive, f1
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
0.89 diagnostic_positive_precision
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Diagnostic / Positive, precision
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
1 diagnostic_positive_recall
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Diagnostic / Positive, recall
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
1 functional_dominant_negative_f1
F1 score · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Dominant Negative, f1
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
1 functional_dominant_negative_precision
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Dominant Negative, precision
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
1 functional_dominant_negative_recall
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Dominant Negative, recall
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
0.88 functional_micro_f1
F1 score · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP: integrating large language models with the Clinical Interpretations of Variants in Cancer · Figure 1c, functional row, agent column
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
0.67 functional_gain_of_function_f1
F1 score · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Gain of Function, f1
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
0.75 functional_gain_of_function_precision
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Gain of Function, precision
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
0.6 functional_gain_of_function_recall
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Gain of Function, recall
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
0.92 functional_loss_of_function_f1
F1 score · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Loss of Function, f1
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
0.86 functional_loss_of_function_precision
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Loss of Function, precision
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
1 functional_loss_of_function_recall
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Loss of Function, recall
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
1 functional_unaltered_function_f1
F1 score · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Unaltered Function, f1
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
1 functional_unaltered_function_precision
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Unaltered Function, precision
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
1 functional_unaltered_function_recall
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Functional / Unaltered Function, recall
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
0.91 micro_f1
F1 score · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP: integrating large language models with the Clinical Interpretations of Variants in Cancer · Fig1c Micro F1; Section 3
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
0.96 oncogenic_micro_f1
F1 score · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP: integrating large language models with the Clinical Interpretations of Variants in Cancer · Figure 1c, oncogenic row, agent column
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
0.96 oncogenic_oncogenicity_f1
F1 score · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Oncogenic / Oncogenicity, f1
Configuration: ChatGPT Agent Mode, 2–7 March 2026 CIViC evaluationProtocol: CIViC evidence-direction retrieval, March 2026
Dataset: CIViC 100 molecular-profile/disease/therapy triplets, March 2026
0.92 oncogenic_oncogenicity_precision
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CIViC agent retrieval

Not reported

Aggregation: Not reported

CIViC MCP supplementary methods and complete Tables 1–3 · Supplementary Table 3, Oncogenic / Oncogenicity, precision

Source checking is not independent reproduction. Release 2026-09-30-e37e3ab1284d.

Methods and evaluation design

Procedure, tasks and evaluated configurations

Evaluation design

Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.

These source-backed links do not make different protocols or scores interchangeable.

Baseline coverage

Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.

0 of 2 active baseline roles have published Rewire measurements in this release. Measurements on a selected protocol do not establish coverage of an entire suite.

Baseline status by linked protocol

Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums

Coverage is derived from release 2026-09-30-e37e3ab1284d. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.

Run this benchmark

Choose a concrete protocol before running an evaluation. Its inputs, split and scoring rules determine which results can be compared.

Run instructions

No runnable recipe has been reviewed for this benchmark. Dataset access, model requirements, licences and compute requirements must be checked against its sources before execution.

Strengths, limitations and unresolved questions

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

0 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-30-e37e3ab1284d
Property and statementOriginal source and locationReview and provenance

No evidence rows match these filters. Choose another scope or clear the search.

Sources and history

View linked audit checks and correction history

Release 2026-09-30-e37e3ab1284d · Record review: source checked

2 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: uc-clinical-20260930-benchmark-civic-study

areas
dna-genomes
contexts
clinical_research
entity level
suite
grouping type
published_study_evaluations
executable suite
false
source locator
Figure1c; Supplementary Tables1–3 and implementation methods
review
method: automated_source_review; reviewer: Codex clinical coverage worker; date: 2026-09-30; note: Transcription from inspected primary source. No execution, independent replication or human scientific review.; source id: uc-clinical-20260930-source-civic
Related records

Suggest a correction