0.981 auroc
kafkas2025-gpcards-gpt-4-one-shot-size75-auroc auroc
- Tested configuration
- GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)
- Protocol
- Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, GPCards (Kafkas et al. 2025 Table 3)
- Dataset
- GPCards gene-phenotype cases (free-text phenotypes) (Kafkas et al. 2025)
- Procedure
- rare-ranking-20261009-protocol-kafkas2025-gpcards-gene-sets
- Evaluation
- GPT-4 One shot on GP-Cards candidate gene sets
- Coverage
- Not reported scored / Not reported eligible
- Uncertainty
- Not reported by the source
- Evidence
- Author-reported evaluation · source checkedThe application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 75, 'ROC AUC', column 'GP-Cards' 'GPT-4' 'One shot'
A source-checked result verifies the numerical transcription, not every model or protocol detail. Evaluation metadata: source checked. Source checked does not mean independently reproduced.
Reproduction
- Split
- No split
- Adaptation
- One shot
- Scoring implementation
- Hits@k, ROC AUC and AUPR over candidate sets
No execution recipe has been verified for this exact configuration and evaluation. A benchmark's general instructions may use different inputs, splits or model settings.
Reproducing this published result requires matching its model configuration, data, split and scorer. Source checking or a successful smoke test does not establish score reproduction.
Evidence
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Evidence table
Inspect claims, sources and review details
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
1 evidence row matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Reported result 0.981 Individual claims | The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients Table 3 row Size 75, 'ROC AUC', column 'GP-Cards' 'GPT-4' 'One shot' Version: Scientific Reports 15, published 2025-04-29; PMC12041562 full-text XML | source checked ["source-hash-verification","deterministic-table-parse","independent-cell-check"] · 2026-10-09 author reported Audit detailsDeterministic parse of the pinned article XML table (extract/extract_rare_ranking.py) with caption, column headers, size, metric and column labels asserted; row-spanning cells resolved by position. Independent review 2026-10-09: matches Table 3 (table-wrap id 'Tab3'), with the three header rows resolved by span. Hits percentages equal a whole number of cases out of the cohort size (50 GPCards, 100 ClinVar or 500 PAVS). Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record Extraction artifact SHA-256: |
Sources and history
Release 2026-10-10-6e93f504adfc · Record review: source checked
1 source records and release history
- The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Original source · Scientific Reports 15, published 2025-04-29; PMC12041562 full-text XML
Technical metadata and extraction receipts
Stable ID: rare-ranking-20261009-result-kafkas2025-gpcards-gpt-4-one-shot-size75-auroc
- metric
- auroc
- metric direction
- higher
- unit
- fraction
- printed value
- 0.981
- numeric value
- 0.981
- source locator
- Table 3 row Size 75, 'ROC AUC', column 'GP-Cards' 'GPT-4' 'One shot'
- missing metadata
- uncertainty: reason: unreported
- metric qualifier
- candidate gene set of 75 genes including the causative gene
- review
- method: source-hash-verification; deterministic-table-parse; independent-cell-check; method note: Re-downloaded the artifact and matched its SHA-256. Read the cell with a separate parser written for this review; the extractor script was not imported or run. Checked printed and numeric value, locator, metric, qualifier, unit, direction and the linked evaluation, configuration, protocol and dataset.; reviewer: claude; reviewer note: Separate Claude review agent, independent of the extractor; no human review claimed; date: 2026-10-09; artifact sha256: 75851f55d59fb0f194ca8bbcf10f47b6f593e8658239dc0fb85f5d2844236a3b; retrieval url: https://www.ebi.ac.uk/europepmc/webservices/rest/PMC12041562/fullTextXML; note: Deterministic parse of the pinned article XML table (extract/extract_rare_ranking.py) with caption, column headers, size, metric and column labels asserted; row-spanning cells resolved by position. Independent review 2026-10-09: matches Table 3 (table-wrap id 'Tab3'), with the three header rows resolved by span. Hits percentages equal a whole number of cases out of the cohort size (50 GPCards, 100 ClinVar or 500 PAVS).
Related records
- evaluation: GPT-4 One shot on GP-Cards candidate gene sets