rewirebio.iobenchmarks
Evaluation

GPT-4 One shot on PAVS candidate gene sets

Published comparison; transcribed, not reproduced.

Research readiness

These checks assess whether the evidence supports a reproducible investigation. A source-checked score alone does not meet these requirements.

Release 2026-10-10-6e93f504adfc · Evidence verified: Not verified

Evidence incomplete

Replay metrics

Exact outcomes, predictions, identifiers and evaluator are connected.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • metric replay: verification is missing

Verified: Not verified

Evidence incomplete

Investigate discrepancies

Replay evidence includes annotations and an assessment of dependence. Unknown independence permits descriptive analysis only.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • metric replay: verification is missing
  • annotations: verification is missing
  • dependence: verification is missing

Verified: Not verified

Evidence incomplete

Run locally

A pinned recipe describes the inputs, environment and resource requirements.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • recipe pinned: verification is missing
  • resource estimate: verification is missing

Verified: Not verified

Evidence incomplete

Validate independently

Separate data and exposure records support an independent test.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • independent validation: verification is missing
  • overlap checked: verification is missing

Verified: Not verified

Readiness describes the evidence in this release. Availability on your computer is checked separately when an investigation runs. Existing data exposure can prevent independent validation even when files are available.

Artifacts and reproduction

No verified artifact manifest is connected to this record yet. The gaps above identify what is needed before analysis can begin.

Read reviewed discrepancy investigations

Evaluation results

1 evaluation · 20 results. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, PAVS (Kafkas et al. 2025 Table 3)
Dataset: PAVS Saudi phenotype-associated variants, 500 genes (Kafkas et al. 2025)
0.539 auprc
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on PAVS candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-pavs-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 100, 'AUPR', column 'PAVS' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, PAVS (Kafkas et al. 2025 Table 3)
Dataset: PAVS Saudi phenotype-associated variants, 500 genes (Kafkas et al. 2025)
0.937 auroc
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on PAVS candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-pavs-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 100, 'ROC AUC', column 'PAVS' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, PAVS (Kafkas et al. 2025 Table 3)
Dataset: PAVS Saudi phenotype-associated variants, 500 genes (Kafkas et al. 2025)
57.8% top-1-accuracy
percent · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on PAVS candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-pavs-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 100, 'Hits@1 (%)', column 'PAVS' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, PAVS (Kafkas et al. 2025 Table 3)
Dataset: PAVS Saudi phenotype-associated variants, 500 genes (Kafkas et al. 2025)
84.8% top-10-accuracy
percent · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on PAVS candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-pavs-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 100, 'Hits@10 (%)', column 'PAVS' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, PAVS (Kafkas et al. 2025 Table 3)
Dataset: PAVS Saudi phenotype-associated variants, 500 genes (Kafkas et al. 2025)
0.784 auprc
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on PAVS candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-pavs-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 25, 'AUPR', column 'PAVS' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, PAVS (Kafkas et al. 2025 Table 3)
Dataset: PAVS Saudi phenotype-associated variants, 500 genes (Kafkas et al. 2025)
0.979 auroc
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on PAVS candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-pavs-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 25, 'ROC AUC', column 'PAVS' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, PAVS (Kafkas et al. 2025 Table 3)
Dataset: PAVS Saudi phenotype-associated variants, 500 genes (Kafkas et al. 2025)
77.8% top-1-accuracy
percent · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on PAVS candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-pavs-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 25, 'Hits@1 (%)', column 'PAVS' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, PAVS (Kafkas et al. 2025 Table 3)
Dataset: PAVS Saudi phenotype-associated variants, 500 genes (Kafkas et al. 2025)
99% top-10-accuracy
percent · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on PAVS candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-pavs-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 25, 'Hits@10 (%)', column 'PAVS' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, PAVS (Kafkas et al. 2025 Table 3)
Dataset: PAVS Saudi phenotype-associated variants, 500 genes (Kafkas et al. 2025)
0.956 auprc
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on PAVS candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-pavs-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 5, 'AUPR', column 'PAVS' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, PAVS (Kafkas et al. 2025 Table 3)
Dataset: PAVS Saudi phenotype-associated variants, 500 genes (Kafkas et al. 2025)
0.991 auroc
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on PAVS candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-pavs-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 5, 'ROC AUC', column 'PAVS' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, PAVS (Kafkas et al. 2025 Table 3)
Dataset: PAVS Saudi phenotype-associated variants, 500 genes (Kafkas et al. 2025)
94.8% top-1-accuracy
percent · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on PAVS candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-pavs-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 5, 'Hits@1 (%)', column 'PAVS' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, PAVS (Kafkas et al. 2025 Table 3)
Dataset: PAVS Saudi phenotype-associated variants, 500 genes (Kafkas et al. 2025)
100% top-10-accuracy
percent · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on PAVS candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-pavs-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 5, 'Hits@10 (%)', column 'PAVS' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, PAVS (Kafkas et al. 2025 Table 3)
Dataset: PAVS Saudi phenotype-associated variants, 500 genes (Kafkas et al. 2025)
0.677 auprc
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on PAVS candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-pavs-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 50, 'AUPR', column 'PAVS' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, PAVS (Kafkas et al. 2025 Table 3)
Dataset: PAVS Saudi phenotype-associated variants, 500 genes (Kafkas et al. 2025)
0.973 auroc
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on PAVS candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-pavs-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 50, 'ROC AUC', column 'PAVS' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, PAVS (Kafkas et al. 2025 Table 3)
Dataset: PAVS Saudi phenotype-associated variants, 500 genes (Kafkas et al. 2025)
68.4% top-1-accuracy
percent · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on PAVS candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-pavs-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 50, 'Hits@1 (%)', column 'PAVS' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, PAVS (Kafkas et al. 2025 Table 3)
Dataset: PAVS Saudi phenotype-associated variants, 500 genes (Kafkas et al. 2025)
95.2% top-10-accuracy
percent · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on PAVS candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-pavs-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 50, 'Hits@10 (%)', column 'PAVS' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, PAVS (Kafkas et al. 2025 Table 3)
Dataset: PAVS Saudi phenotype-associated variants, 500 genes (Kafkas et al. 2025)
0.595 auprc
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on PAVS candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-pavs-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 75, 'AUPR', column 'PAVS' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, PAVS (Kafkas et al. 2025 Table 3)
Dataset: PAVS Saudi phenotype-associated variants, 500 genes (Kafkas et al. 2025)
0.96 auroc
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on PAVS candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-pavs-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 75, 'ROC AUC', column 'PAVS' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, PAVS (Kafkas et al. 2025 Table 3)
Dataset: PAVS Saudi phenotype-associated variants, 500 genes (Kafkas et al. 2025)
61.8% top-1-accuracy
percent · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on PAVS candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-pavs-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 75, 'Hits@1 (%)', column 'PAVS' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, PAVS (Kafkas et al. 2025 Table 3)
Dataset: PAVS Saudi phenotype-associated variants, 500 genes (Kafkas et al. 2025)
91.4% top-10-accuracy
percent · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on PAVS candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-pavs-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 75, 'Hits@10 (%)', column 'PAVS' 'GPT-4' 'One shot'

Source checking is not independent reproduction. Release 2026-10-10-6e93f504adfc.

Evaluation procedure

rare-ranking-20261009-protocol-kafkas2025-pavs-gene-sets

Configuration
GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)
Protocol
Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, PAVS (Kafkas et al. 2025 Table 3)
Dataset
PAVS Saudi phenotype-associated variants, 500 genes (Kafkas et al. 2025)
origin
Author-reported evaluation
configuration
Primary source as retrieved 2026-10-09
dataset version
As published
split
No split
population
PAVS genotype-phenotype pairs
inputs
HPO-coded phenotypes and a candidate gene list
adaptation
One shot
metric implementation
Hits@k, ROC AUC and AUPR over candidate sets
aggregation
Per candidate-set size
budget
Candidate sets of 5, 25, 50, 75 and 100 genes
protocol id
rare-ranking-20261009-protocol-kafkas2025-pavs-gene-sets

Metadata review: source checked. Unreported conditions prevent automatic comparisons.

Reproduction

Split
No split
Adaptation
One shot
Scoring implementation
Hits@k, ROC AUC and AUPR over candidate sets

No execution recipe has been verified for this exact configuration and evaluation. A benchmark's general instructions may use different inputs, splits or model settings.

Reproducing this published result requires matching its model configuration, data, split and scorer. Source checking or a successful smoke test does not establish score reproduction.

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

19 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-10-10-6e93f504adfc
Property and statementOriginal source and locationReview and provenance
attributes.comparison.adaptation
One shot
Context-only references
The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients

Original source ↗

Table 3 column 'PAVS', 'GPT-4' 'One shot'

Version: Scientific Reports 15, published 2025-04-29; PMC12041562 full-text XML
Retrieved: 2026-10-09T21:20:27Z

not individually reviewed

No individual claim review recorded

author reported

Audit details

Field: attributes.comparison.adaptation

Source artifact SHA-256: 75851f55d59fb0f194ca8bbcf10f47b6f593e8658239dc0fb85f5d2844236a3b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

attributes.comparison.aggregation
Per candidate-set size
Context-only references
The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients

Original source ↗

Table 3 column 'PAVS', 'GPT-4' 'One shot'

Version: Scientific Reports 15, published 2025-04-29; PMC12041562 full-text XML
Retrieved: 2026-10-09T21:20:27Z

not individually reviewed

No individual claim review recorded

author reported

Audit details

Field: attributes.comparison.aggregation

Source artifact SHA-256: 75851f55d59fb0f194ca8bbcf10f47b6f593e8658239dc0fb85f5d2844236a3b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

attributes.comparison.budget
Candidate sets of 5, 25, 50, 75 and 100 genes
Context-only references
The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients

Original source ↗

Table 3 column 'PAVS', 'GPT-4' 'One shot'

Version: Scientific Reports 15, published 2025-04-29; PMC12041562 full-text XML
Retrieved: 2026-10-09T21:20:27Z

not individually reviewed

No individual claim review recorded

author reported

Audit details

Field: attributes.comparison.budget

Source artifact SHA-256: 75851f55d59fb0f194ca8bbcf10f47b6f593e8658239dc0fb85f5d2844236a3b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

attributes.comparison.dataset_version
As published
Context-only references
The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients

Original source ↗

Table 3 column 'PAVS', 'GPT-4' 'One shot'

Version: Scientific Reports 15, published 2025-04-29; PMC12041562 full-text XML
Retrieved: 2026-10-09T21:20:27Z

not individually reviewed

No individual claim review recorded

author reported

Audit details

Field: attributes.comparison.dataset_version

Source artifact SHA-256: 75851f55d59fb0f194ca8bbcf10f47b6f593e8658239dc0fb85f5d2844236a3b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

attributes.comparison.inputs
HPO-coded phenotypes and a candidate gene list
Context-only references
The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients

Original source ↗

Table 3 column 'PAVS', 'GPT-4' 'One shot'

Version: Scientific Reports 15, published 2025-04-29; PMC12041562 full-text XML
Retrieved: 2026-10-09T21:20:27Z

not individually reviewed

No individual claim review recorded

author reported

Audit details

Field: attributes.comparison.inputs

Source artifact SHA-256: 75851f55d59fb0f194ca8bbcf10f47b6f593e8658239dc0fb85f5d2844236a3b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

attributes.comparison.metric_implementation
Hits@k, ROC AUC and AUPR over candidate sets
Context-only references
The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients

Original source ↗

Table 3 column 'PAVS', 'GPT-4' 'One shot'

Version: Scientific Reports 15, published 2025-04-29; PMC12041562 full-text XML
Retrieved: 2026-10-09T21:20:27Z

not individually reviewed

No individual claim review recorded

author reported

Audit details

Field: attributes.comparison.metric_implementation

Source artifact SHA-256: 75851f55d59fb0f194ca8bbcf10f47b6f593e8658239dc0fb85f5d2844236a3b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

attributes.comparison.population
PAVS genotype-phenotype pairs
Context-only references
The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients

Original source ↗

Table 3 column 'PAVS', 'GPT-4' 'One shot'

Version: Scientific Reports 15, published 2025-04-29; PMC12041562 full-text XML
Retrieved: 2026-10-09T21:20:27Z

not individually reviewed

No individual claim review recorded

author reported

Audit details

Field: attributes.comparison.population

Source artifact SHA-256: 75851f55d59fb0f194ca8bbcf10f47b6f593e8658239dc0fb85f5d2844236a3b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

attributes.comparison.protocol_id
rare-ranking-20261009-protocol-kafkas2025-pavs-gene-sets
Context-only references
The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients

Original source ↗

Table 3 column 'PAVS', 'GPT-4' 'One shot'

Version: Scientific Reports 15, published 2025-04-29; PMC12041562 full-text XML
Retrieved: 2026-10-09T21:20:27Z

not individually reviewed

No individual claim review recorded

author reported

Audit details

Field: attributes.comparison.protocol_id

Source artifact SHA-256: 75851f55d59fb0f194ca8bbcf10f47b6f593e8658239dc0fb85f5d2844236a3b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

attributes.comparison.split
No split
Context-only references
The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients

Original source ↗

Table 3 column 'PAVS', 'GPT-4' 'One shot'

Version: Scientific Reports 15, published 2025-04-29; PMC12041562 full-text XML
Retrieved: 2026-10-09T21:20:27Z

not individually reviewed

No individual claim review recorded

author reported

Audit details

Field: attributes.comparison.split

Source artifact SHA-256: 75851f55d59fb0f194ca8bbcf10f47b6f593e8658239dc0fb85f5d2844236a3b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

attributes.limitations
1 values
  • The authors designed the GPT-4 prompts, including the one-shot chain-of-thought prompt Q4 (Table 1; Methods 'Prompt engineering'; Author contributions), so these rows are author_reported.
Context-only references
The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients

Original source ↗

Table 3 column 'PAVS', 'GPT-4' 'One shot'

Version: Scientific Reports 15, published 2025-04-29; PMC12041562 full-text XML
Retrieved: 2026-10-09T21:20:27Z

not individually reviewed

No individual claim review recorded

author reported

Audit details

Field: attributes.limitations

Source artifact SHA-256: 75851f55d59fb0f194ca8bbcf10f47b6f593e8658239dc0fb85f5d2844236a3b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

Release 2026-10-10-6e93f504adfc · Record review: source checked

1 source records and release historyDownload this release (gzip)
Technical metadata and extraction receipts

Stable ID: rare-ranking-20261009-eval-kafkas2025-pavs-gpt-4-one-shot

areas
dna-genomes
contexts
clinical_research
origin
author_reported
protocol
rare-ranking-20261009-protocol-kafkas2025-pavs-gene-sets
version
Primary source as retrieved 2026-10-09
comparison
dataset version: As published; split: No split; population: PAVS genotype-phenotype pairs; inputs: HPO-coded phenotypes and a candidate gene list; adaptation: One shot; metric implementation: Hits@k, ROC AUC and AUPR over candidate sets; aggregation: Per candidate-set size; budget: Candidate sets of 5, 25, 50, 75 and 100 genes; protocol id: rare-ranking-20261009-protocol-kafkas2025-pavs-gene-sets
source locator
Table 3 column 'PAVS', 'GPT-4' 'One shot'
limitations
The authors designed the GPT-4 prompts, including the one-shot chain-of-thought prompt Q4 (Table 1; Methods 'Prompt engineering'; Author contributions), so these rows are author_reported.
Related records

Suggest a correction