rewirebio.iobenchmarks
Evaluation

GPT-4 Zero shot on ClinVar candidate gene sets

Published comparison; transcribed, not reproduced.

Research readiness

These checks assess whether the evidence supports a reproducible investigation. A source-checked score alone does not meet these requirements.

Release 2026-10-10-6e93f504adfc · Evidence verified: Not verified

Evidence incomplete

Replay metrics

Exact outcomes, predictions, identifiers and evaluator are connected.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • metric replay: verification is missing

Verified: Not verified

Evidence incomplete

Investigate discrepancies

Replay evidence includes annotations and an assessment of dependence. Unknown independence permits descriptive analysis only.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • metric replay: verification is missing
  • annotations: verification is missing
  • dependence: verification is missing

Verified: Not verified

Evidence incomplete

Run locally

A pinned recipe describes the inputs, environment and resource requirements.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • recipe pinned: verification is missing
  • resource estimate: verification is missing

Verified: Not verified

Evidence incomplete

Validate independently

Separate data and exposure records support an independent test.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • independent validation: verification is missing
  • overlap checked: verification is missing

Verified: Not verified

Readiness describes the evidence in this release. Availability on your computer is checked separately when an investigation runs. Existing data exposure can prevent independent validation even when files are available.

Artifacts and reproduction

No verified artifact manifest is connected to this record yet. The gaps above identify what is needed before analysis can begin.

Read reviewed discrepancy investigations

Evaluation results

1 evaluation · 20 results. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: GPT-4 (gpt-4-1106-preview), zero-shot prompt (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, ClinVar (Kafkas et al. 2025 Table 3)
Dataset: ClinVar variants added 2 July to 7 October 2023, 100 genes (Kafkas et al. 2025)
0.622 auprc
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 Zero shot on ClinVar candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-clinvar-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 100, 'AUPR', column 'ClinVar' 'GPT-4' 'Zero shot'
Configuration: GPT-4 (gpt-4-1106-preview), zero-shot prompt (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, ClinVar (Kafkas et al. 2025 Table 3)
Dataset: ClinVar variants added 2 July to 7 October 2023, 100 genes (Kafkas et al. 2025)
0.911 auroc
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 Zero shot on ClinVar candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-clinvar-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 100, 'ROC AUC', column 'ClinVar' 'GPT-4' 'Zero shot'
Configuration: GPT-4 (gpt-4-1106-preview), zero-shot prompt (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, ClinVar (Kafkas et al. 2025 Table 3)
Dataset: ClinVar variants added 2 July to 7 October 2023, 100 genes (Kafkas et al. 2025)
68% top-1-accuracy
percent · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 Zero shot on ClinVar candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-clinvar-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 100, 'Hits@1 (%)', column 'ClinVar' 'GPT-4' 'Zero shot'
Configuration: GPT-4 (gpt-4-1106-preview), zero-shot prompt (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, ClinVar (Kafkas et al. 2025 Table 3)
Dataset: ClinVar variants added 2 July to 7 October 2023, 100 genes (Kafkas et al. 2025)
81% top-10-accuracy
percent · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 Zero shot on ClinVar candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-clinvar-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 100, 'Hits@10 (%)', column 'ClinVar' 'GPT-4' 'Zero shot'
Configuration: GPT-4 (gpt-4-1106-preview), zero-shot prompt (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, ClinVar (Kafkas et al. 2025 Table 3)
Dataset: ClinVar variants added 2 July to 7 October 2023, 100 genes (Kafkas et al. 2025)
0.802 auprc
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 Zero shot on ClinVar candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-clinvar-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 25, 'AUPR', column 'ClinVar' 'GPT-4' 'Zero shot'
Configuration: GPT-4 (gpt-4-1106-preview), zero-shot prompt (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, ClinVar (Kafkas et al. 2025 Table 3)
Dataset: ClinVar variants added 2 July to 7 October 2023, 100 genes (Kafkas et al. 2025)
0.964 auroc
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 Zero shot on ClinVar candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-clinvar-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 25, 'ROC AUC', column 'ClinVar' 'GPT-4' 'Zero shot'
Configuration: GPT-4 (gpt-4-1106-preview), zero-shot prompt (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, ClinVar (Kafkas et al. 2025 Table 3)
Dataset: ClinVar variants added 2 July to 7 October 2023, 100 genes (Kafkas et al. 2025)
82% top-1-accuracy
percent · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 Zero shot on ClinVar candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-clinvar-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 25, 'Hits@1 (%)', column 'ClinVar' 'GPT-4' 'Zero shot'
Configuration: GPT-4 (gpt-4-1106-preview), zero-shot prompt (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, ClinVar (Kafkas et al. 2025 Table 3)
Dataset: ClinVar variants added 2 July to 7 October 2023, 100 genes (Kafkas et al. 2025)
94% top-10-accuracy
percent · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 Zero shot on ClinVar candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-clinvar-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 25, 'Hits@10 (%)', column 'ClinVar' 'GPT-4' 'Zero shot'
Configuration: GPT-4 (gpt-4-1106-preview), zero-shot prompt (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, ClinVar (Kafkas et al. 2025 Table 3)
Dataset: ClinVar variants added 2 July to 7 October 2023, 100 genes (Kafkas et al. 2025)
0.944 auprc
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 Zero shot on ClinVar candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-clinvar-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 5, 'AUPR', column 'ClinVar' 'GPT-4' 'Zero shot'
Configuration: GPT-4 (gpt-4-1106-preview), zero-shot prompt (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, ClinVar (Kafkas et al. 2025 Table 3)
Dataset: ClinVar variants added 2 July to 7 October 2023, 100 genes (Kafkas et al. 2025)
0.991 auroc
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 Zero shot on ClinVar candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-clinvar-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 5, 'ROC AUC', column 'ClinVar' 'GPT-4' 'Zero shot'
Configuration: GPT-4 (gpt-4-1106-preview), zero-shot prompt (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, ClinVar (Kafkas et al. 2025 Table 3)
Dataset: ClinVar variants added 2 July to 7 October 2023, 100 genes (Kafkas et al. 2025)
93% top-1-accuracy
percent · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 Zero shot on ClinVar candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-clinvar-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 5, 'Hits@1 (%)', column 'ClinVar' 'GPT-4' 'Zero shot'
Configuration: GPT-4 (gpt-4-1106-preview), zero-shot prompt (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, ClinVar (Kafkas et al. 2025 Table 3)
Dataset: ClinVar variants added 2 July to 7 October 2023, 100 genes (Kafkas et al. 2025)
100% top-10-accuracy
percent · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 Zero shot on ClinVar candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-clinvar-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 5, 'Hits@10 (%)', column 'ClinVar' 'GPT-4' 'Zero shot'
Configuration: GPT-4 (gpt-4-1106-preview), zero-shot prompt (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, ClinVar (Kafkas et al. 2025 Table 3)
Dataset: ClinVar variants added 2 July to 7 October 2023, 100 genes (Kafkas et al. 2025)
0.806 auprc
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 Zero shot on ClinVar candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-clinvar-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 50, 'AUPR', column 'ClinVar' 'GPT-4' 'Zero shot'
Configuration: GPT-4 (gpt-4-1106-preview), zero-shot prompt (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, ClinVar (Kafkas et al. 2025 Table 3)
Dataset: ClinVar variants added 2 July to 7 October 2023, 100 genes (Kafkas et al. 2025)
0.971 auroc
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 Zero shot on ClinVar candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-clinvar-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 50, 'ROC AUC', column 'ClinVar' 'GPT-4' 'Zero shot'
Configuration: GPT-4 (gpt-4-1106-preview), zero-shot prompt (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, ClinVar (Kafkas et al. 2025 Table 3)
Dataset: ClinVar variants added 2 July to 7 October 2023, 100 genes (Kafkas et al. 2025)
81% top-1-accuracy
percent · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 Zero shot on ClinVar candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-clinvar-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 50, 'Hits@1 (%)', column 'ClinVar' 'GPT-4' 'Zero shot'
Configuration: GPT-4 (gpt-4-1106-preview), zero-shot prompt (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, ClinVar (Kafkas et al. 2025 Table 3)
Dataset: ClinVar variants added 2 July to 7 October 2023, 100 genes (Kafkas et al. 2025)
96% top-10-accuracy
percent · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 Zero shot on ClinVar candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-clinvar-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 50, 'Hits@10 (%)', column 'ClinVar' 'GPT-4' 'Zero shot'
Configuration: GPT-4 (gpt-4-1106-preview), zero-shot prompt (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, ClinVar (Kafkas et al. 2025 Table 3)
Dataset: ClinVar variants added 2 July to 7 October 2023, 100 genes (Kafkas et al. 2025)
0.663 auprc
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 Zero shot on ClinVar candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-clinvar-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 75, 'AUPR', column 'ClinVar' 'GPT-4' 'Zero shot'
Configuration: GPT-4 (gpt-4-1106-preview), zero-shot prompt (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, ClinVar (Kafkas et al. 2025 Table 3)
Dataset: ClinVar variants added 2 July to 7 October 2023, 100 genes (Kafkas et al. 2025)
0.925 auroc
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 Zero shot on ClinVar candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-clinvar-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 75, 'ROC AUC', column 'ClinVar' 'GPT-4' 'Zero shot'
Configuration: GPT-4 (gpt-4-1106-preview), zero-shot prompt (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, ClinVar (Kafkas et al. 2025 Table 3)
Dataset: ClinVar variants added 2 July to 7 October 2023, 100 genes (Kafkas et al. 2025)
71% top-1-accuracy
percent · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 Zero shot on ClinVar candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-clinvar-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 75, 'Hits@1 (%)', column 'ClinVar' 'GPT-4' 'Zero shot'
Configuration: GPT-4 (gpt-4-1106-preview), zero-shot prompt (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, ClinVar (Kafkas et al. 2025 Table 3)
Dataset: ClinVar variants added 2 July to 7 October 2023, 100 genes (Kafkas et al. 2025)
82% top-10-accuracy
percent · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 Zero shot on ClinVar candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-clinvar-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 75, 'Hits@10 (%)', column 'ClinVar' 'GPT-4' 'Zero shot'

Source checking is not independent reproduction. Release 2026-10-10-6e93f504adfc.

Evaluation procedure

rare-ranking-20261009-protocol-kafkas2025-clinvar-gene-sets

Configuration
GPT-4 (gpt-4-1106-preview), zero-shot prompt (Kafkas et al. 2025)
Protocol
Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, ClinVar (Kafkas et al. 2025 Table 3)
Dataset
ClinVar variants added 2 July to 7 October 2023, 100 genes (Kafkas et al. 2025)
origin
Author-reported evaluation
configuration
Primary source as retrieved 2026-10-09
dataset version
As published
split
No split
population
ClinVar genotype-phenotype pairs
inputs
HPO-coded phenotypes and a candidate gene list
adaptation
Zero shot
metric implementation
Hits@k, ROC AUC and AUPR over candidate sets
aggregation
Per candidate-set size
budget
Candidate sets of 5, 25, 50, 75 and 100 genes
protocol id
rare-ranking-20261009-protocol-kafkas2025-clinvar-gene-sets

Metadata review: source checked. Unreported conditions prevent automatic comparisons.

Reproduction

Split
No split
Adaptation
Zero shot
Scoring implementation
Hits@k, ROC AUC and AUPR over candidate sets

No execution recipe has been verified for this exact configuration and evaluation. A benchmark's general instructions may use different inputs, splits or model settings.

Reproducing this published result requires matching its model configuration, data, split and scorer. Source checking or a successful smoke test does not establish score reproduction.

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

19 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-10-10-6e93f504adfc
Property and statementOriginal source and locationReview and provenance
attributes.comparison.adaptation
Zero shot
Context-only references
The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients

Original source ↗

Table 3 column 'ClinVar', 'GPT-4' 'Zero shot'

Version: Scientific Reports 15, published 2025-04-29; PMC12041562 full-text XML
Retrieved: 2026-10-09T21:20:27Z

not individually reviewed

No individual claim review recorded

author reported

Audit details

Field: attributes.comparison.adaptation

Source artifact SHA-256: 75851f55d59fb0f194ca8bbcf10f47b6f593e8658239dc0fb85f5d2844236a3b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

attributes.comparison.aggregation
Per candidate-set size
Context-only references
The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients

Original source ↗

Table 3 column 'ClinVar', 'GPT-4' 'Zero shot'

Version: Scientific Reports 15, published 2025-04-29; PMC12041562 full-text XML
Retrieved: 2026-10-09T21:20:27Z

not individually reviewed

No individual claim review recorded

author reported

Audit details

Field: attributes.comparison.aggregation

Source artifact SHA-256: 75851f55d59fb0f194ca8bbcf10f47b6f593e8658239dc0fb85f5d2844236a3b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

attributes.comparison.budget
Candidate sets of 5, 25, 50, 75 and 100 genes
Context-only references
The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients

Original source ↗

Table 3 column 'ClinVar', 'GPT-4' 'Zero shot'

Version: Scientific Reports 15, published 2025-04-29; PMC12041562 full-text XML
Retrieved: 2026-10-09T21:20:27Z

not individually reviewed

No individual claim review recorded

author reported

Audit details

Field: attributes.comparison.budget

Source artifact SHA-256: 75851f55d59fb0f194ca8bbcf10f47b6f593e8658239dc0fb85f5d2844236a3b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

attributes.comparison.dataset_version
As published
Context-only references
The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients

Original source ↗

Table 3 column 'ClinVar', 'GPT-4' 'Zero shot'

Version: Scientific Reports 15, published 2025-04-29; PMC12041562 full-text XML
Retrieved: 2026-10-09T21:20:27Z

not individually reviewed

No individual claim review recorded

author reported

Audit details

Field: attributes.comparison.dataset_version

Source artifact SHA-256: 75851f55d59fb0f194ca8bbcf10f47b6f593e8658239dc0fb85f5d2844236a3b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

attributes.comparison.inputs
HPO-coded phenotypes and a candidate gene list
Context-only references
The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients

Original source ↗

Table 3 column 'ClinVar', 'GPT-4' 'Zero shot'

Version: Scientific Reports 15, published 2025-04-29; PMC12041562 full-text XML
Retrieved: 2026-10-09T21:20:27Z

not individually reviewed

No individual claim review recorded

author reported

Audit details

Field: attributes.comparison.inputs

Source artifact SHA-256: 75851f55d59fb0f194ca8bbcf10f47b6f593e8658239dc0fb85f5d2844236a3b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

attributes.comparison.metric_implementation
Hits@k, ROC AUC and AUPR over candidate sets
Context-only references
The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients

Original source ↗

Table 3 column 'ClinVar', 'GPT-4' 'Zero shot'

Version: Scientific Reports 15, published 2025-04-29; PMC12041562 full-text XML
Retrieved: 2026-10-09T21:20:27Z

not individually reviewed

No individual claim review recorded

author reported

Audit details

Field: attributes.comparison.metric_implementation

Source artifact SHA-256: 75851f55d59fb0f194ca8bbcf10f47b6f593e8658239dc0fb85f5d2844236a3b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

attributes.comparison.population
ClinVar genotype-phenotype pairs
Context-only references
The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients

Original source ↗

Table 3 column 'ClinVar', 'GPT-4' 'Zero shot'

Version: Scientific Reports 15, published 2025-04-29; PMC12041562 full-text XML
Retrieved: 2026-10-09T21:20:27Z

not individually reviewed

No individual claim review recorded

author reported

Audit details

Field: attributes.comparison.population

Source artifact SHA-256: 75851f55d59fb0f194ca8bbcf10f47b6f593e8658239dc0fb85f5d2844236a3b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

attributes.comparison.protocol_id
rare-ranking-20261009-protocol-kafkas2025-clinvar-gene-sets
Context-only references
The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients

Original source ↗

Table 3 column 'ClinVar', 'GPT-4' 'Zero shot'

Version: Scientific Reports 15, published 2025-04-29; PMC12041562 full-text XML
Retrieved: 2026-10-09T21:20:27Z

not individually reviewed

No individual claim review recorded

author reported

Audit details

Field: attributes.comparison.protocol_id

Source artifact SHA-256: 75851f55d59fb0f194ca8bbcf10f47b6f593e8658239dc0fb85f5d2844236a3b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

attributes.comparison.split
No split
Context-only references
The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients

Original source ↗

Table 3 column 'ClinVar', 'GPT-4' 'Zero shot'

Version: Scientific Reports 15, published 2025-04-29; PMC12041562 full-text XML
Retrieved: 2026-10-09T21:20:27Z

not individually reviewed

No individual claim review recorded

author reported

Audit details

Field: attributes.comparison.split

Source artifact SHA-256: 75851f55d59fb0f194ca8bbcf10f47b6f593e8658239dc0fb85f5d2844236a3b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

attributes.limitations
1 values
  • The authors designed the GPT-4 prompts, including the one-shot chain-of-thought prompt Q4 (Table 1; Methods 'Prompt engineering'; Author contributions), so these rows are author_reported.
Context-only references
The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients

Original source ↗

Table 3 column 'ClinVar', 'GPT-4' 'Zero shot'

Version: Scientific Reports 15, published 2025-04-29; PMC12041562 full-text XML
Retrieved: 2026-10-09T21:20:27Z

not individually reviewed

No individual claim review recorded

author reported

Audit details

Field: attributes.limitations

Source artifact SHA-256: 75851f55d59fb0f194ca8bbcf10f47b6f593e8658239dc0fb85f5d2844236a3b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

Release 2026-10-10-6e93f504adfc · Record review: source checked

1 source records and release historyDownload this release (gzip)
Technical metadata and extraction receipts

Stable ID: rare-ranking-20261009-eval-kafkas2025-clinvar-gpt-4-zero-shot

areas
dna-genomes
contexts
clinical_research
origin
author_reported
protocol
rare-ranking-20261009-protocol-kafkas2025-clinvar-gene-sets
version
Primary source as retrieved 2026-10-09
comparison
dataset version: As published; split: No split; population: ClinVar genotype-phenotype pairs; inputs: HPO-coded phenotypes and a candidate gene list; adaptation: Zero shot; metric implementation: Hits@k, ROC AUC and AUPR over candidate sets; aggregation: Per candidate-set size; budget: Candidate sets of 5, 25, 50, 75 and 100 genes; protocol id: rare-ranking-20261009-protocol-kafkas2025-clinvar-gene-sets
source locator
Table 3 column 'ClinVar', 'GPT-4' 'Zero shot'
limitations
The authors designed the GPT-4 prompts, including the one-shot chain-of-thought prompt Q4 (Table 1; Methods 'Prompt engineering'; Author contributions), so these rows are author_reported.
Related records

Suggest a correction