rewirebio.iobenchmarks
Protocol

Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, GPCards (Kafkas et al. 2025 Table 3)

Position of the causative gene when a fixed candidate set is ranked from phenotypes.

2 evaluations · 40 results

Overview

Position of the causative gene when a fixed candidate set is ranked from phenotypes.

Consult the linked sources for architecture or protocol details. Missing evidence is not evidence of a missing capability.

2 recorded evaluations, 40 metric rows. A comparison chart has not yet been validated for these results. The table retains the individual findings and their sources.

View coverage and remaining gaps across all benchmarks

Results

Results are available, but no reviewed comparison panel is linked in this release.

All evaluations

2 evaluations · 40 results. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, GPCards (Kafkas et al. 2025 Table 3)
Dataset: GPCards gene-phenotype cases (free-text phenotypes) (Kafkas et al. 2025)
0.559 auprc
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on GP-Cards candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-gpcards-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 100, 'AUPR', column 'GP-Cards' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, GPCards (Kafkas et al. 2025 Table 3)
Dataset: GPCards gene-phenotype cases (free-text phenotypes) (Kafkas et al. 2025)
0.927 auroc
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on GP-Cards candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-gpcards-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 100, 'ROC AUC', column 'GP-Cards' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, GPCards (Kafkas et al. 2025 Table 3)
Dataset: GPCards gene-phenotype cases (free-text phenotypes) (Kafkas et al. 2025)
60% top-1-accuracy
percent · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on GP-Cards candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-gpcards-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 100, 'Hits@1 (%)', column 'GP-Cards' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, GPCards (Kafkas et al. 2025 Table 3)
Dataset: GPCards gene-phenotype cases (free-text phenotypes) (Kafkas et al. 2025)
82% top-10-accuracy
percent · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on GP-Cards candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-gpcards-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 100, 'Hits@10 (%)', column 'GP-Cards' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, GPCards (Kafkas et al. 2025 Table 3)
Dataset: GPCards gene-phenotype cases (free-text phenotypes) (Kafkas et al. 2025)
0.756 auprc
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on GP-Cards candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-gpcards-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 25, 'AUPR', column 'GP-Cards' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, GPCards (Kafkas et al. 2025 Table 3)
Dataset: GPCards gene-phenotype cases (free-text phenotypes) (Kafkas et al. 2025)
0.964 auroc
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on GP-Cards candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-gpcards-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 25, 'ROC AUC', column 'GP-Cards' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, GPCards (Kafkas et al. 2025 Table 3)
Dataset: GPCards gene-phenotype cases (free-text phenotypes) (Kafkas et al. 2025)
78% top-1-accuracy
percent · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on GP-Cards candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-gpcards-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 25, 'Hits@1 (%)', column 'GP-Cards' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, GPCards (Kafkas et al. 2025 Table 3)
Dataset: GPCards gene-phenotype cases (free-text phenotypes) (Kafkas et al. 2025)
98% top-10-accuracy
percent · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on GP-Cards candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-gpcards-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 25, 'Hits@10 (%)', column 'GP-Cards' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, GPCards (Kafkas et al. 2025 Table 3)
Dataset: GPCards gene-phenotype cases (free-text phenotypes) (Kafkas et al. 2025)
0.949 auprc
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on GP-Cards candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-gpcards-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 5, 'AUPR', column 'GP-Cards' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, GPCards (Kafkas et al. 2025 Table 3)
Dataset: GPCards gene-phenotype cases (free-text phenotypes) (Kafkas et al. 2025)
0.99 auroc
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on GP-Cards candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-gpcards-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 5, 'ROC AUC', column 'GP-Cards' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, GPCards (Kafkas et al. 2025 Table 3)
Dataset: GPCards gene-phenotype cases (free-text phenotypes) (Kafkas et al. 2025)
94% top-1-accuracy
percent · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on GP-Cards candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-gpcards-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 5, 'Hits@1 (%)', column 'GP-Cards' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, GPCards (Kafkas et al. 2025 Table 3)
Dataset: GPCards gene-phenotype cases (free-text phenotypes) (Kafkas et al. 2025)
100% top-10-accuracy
percent · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on GP-Cards candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-gpcards-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 5, 'Hits@10 (%)', column 'GP-Cards' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, GPCards (Kafkas et al. 2025 Table 3)
Dataset: GPCards gene-phenotype cases (free-text phenotypes) (Kafkas et al. 2025)
0.68 auprc
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on GP-Cards candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-gpcards-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 50, 'AUPR', column 'GP-Cards' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, GPCards (Kafkas et al. 2025 Table 3)
Dataset: GPCards gene-phenotype cases (free-text phenotypes) (Kafkas et al. 2025)
0.961 auroc
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on GP-Cards candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-gpcards-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 50, 'ROC AUC', column 'GP-Cards' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, GPCards (Kafkas et al. 2025 Table 3)
Dataset: GPCards gene-phenotype cases (free-text phenotypes) (Kafkas et al. 2025)
70% top-1-accuracy
percent · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on GP-Cards candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-gpcards-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 50, 'Hits@1 (%)', column 'GP-Cards' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, GPCards (Kafkas et al. 2025 Table 3)
Dataset: GPCards gene-phenotype cases (free-text phenotypes) (Kafkas et al. 2025)
94% top-10-accuracy
percent · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on GP-Cards candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-gpcards-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 50, 'Hits@10 (%)', column 'GP-Cards' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, GPCards (Kafkas et al. 2025 Table 3)
Dataset: GPCards gene-phenotype cases (free-text phenotypes) (Kafkas et al. 2025)
0.681 auprc
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on GP-Cards candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-gpcards-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 75, 'AUPR', column 'GP-Cards' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, GPCards (Kafkas et al. 2025 Table 3)
Dataset: GPCards gene-phenotype cases (free-text phenotypes) (Kafkas et al. 2025)
0.981 auroc
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on GP-Cards candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-gpcards-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 75, 'ROC AUC', column 'GP-Cards' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, GPCards (Kafkas et al. 2025 Table 3)
Dataset: GPCards gene-phenotype cases (free-text phenotypes) (Kafkas et al. 2025)
72% top-1-accuracy
percent · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on GP-Cards candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-gpcards-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 75, 'Hits@1 (%)', column 'GP-Cards' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), one-shot chain-of-thought prompt Q4 (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, GPCards (Kafkas et al. 2025 Table 3)
Dataset: GPCards gene-phenotype cases (free-text phenotypes) (Kafkas et al. 2025)
94% top-10-accuracy
percent · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 One shot on GP-Cards candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-gpcards-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 75, 'Hits@10 (%)', column 'GP-Cards' 'GPT-4' 'One shot'
Configuration: GPT-4 (gpt-4-1106-preview), zero-shot prompt (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, GPCards (Kafkas et al. 2025 Table 3)
Dataset: GPCards gene-phenotype cases (free-text phenotypes) (Kafkas et al. 2025)
0.343 auprc
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 Zero shot on GP-Cards candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-gpcards-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 100, 'AUPR', column 'GP-Cards' 'GPT-4' 'Zero shot'
Configuration: GPT-4 (gpt-4-1106-preview), zero-shot prompt (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, GPCards (Kafkas et al. 2025 Table 3)
Dataset: GPCards gene-phenotype cases (free-text phenotypes) (Kafkas et al. 2025)
0.928 auroc
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 Zero shot on GP-Cards candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-gpcards-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 100, 'ROC AUC', column 'GP-Cards' 'GPT-4' 'Zero shot'
Configuration: GPT-4 (gpt-4-1106-preview), zero-shot prompt (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, GPCards (Kafkas et al. 2025 Table 3)
Dataset: GPCards gene-phenotype cases (free-text phenotypes) (Kafkas et al. 2025)
60% top-1-accuracy
percent · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 Zero shot on GP-Cards candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-gpcards-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 100, 'Hits@1 (%)', column 'GP-Cards' 'GPT-4' 'Zero shot'
Configuration: GPT-4 (gpt-4-1106-preview), zero-shot prompt (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, GPCards (Kafkas et al. 2025 Table 3)
Dataset: GPCards gene-phenotype cases (free-text phenotypes) (Kafkas et al. 2025)
86% top-10-accuracy
percent · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 Zero shot on GP-Cards candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-gpcards-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 100, 'Hits@10 (%)', column 'GP-Cards' 'GPT-4' 'Zero shot'
Configuration: GPT-4 (gpt-4-1106-preview), zero-shot prompt (Kafkas et al. 2025)Protocol: Ranking the causative gene within synthetic candidate sets of 5 to 100 genes, GPCards (Kafkas et al. 2025 Table 3)
Dataset: GPCards gene-phenotype cases (free-text phenotypes) (Kafkas et al. 2025)
0.77 auprc
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GPT-4 Zero shot on GP-Cards candidate gene sets

rare-ranking-20261009-protocol-kafkas2025-gpcards-gene-sets

Aggregation: Not reported

The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients · Table 3 row Size 25, 'AUPR', column 'GP-Cards' 'GPT-4' 'Zero shot'

Source checking is not independent reproduction. Release 2026-10-10-6e93f504adfc.

Methods and evaluation design

Procedure, tasks and evaluated configurations

Recorded evaluations

Each evaluation records what was tested and under which conditions.

Baseline coverage

Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.

0 of 2 active baseline roles have published Rewire measurements in this release. Measurements on a selected protocol do not establish coverage of an entire suite.

No execution recipe linked to this protocol. Recipe availability does not establish a completed evaluation.

Author-reported evaluations
2

Literature evidence is not a Rewire measurement. Executed but unpublished runs and private review status are not included.

Null control

Proposed control: requires review

Select a task-valid null control after reviewing inputs and metric

Protocol-specific applicability, permitted inputs, access, split, evaluator and execution requirements need review before implementation or execution.

This is a suggested selection rule, not a validated method or a measured score.

Conventional reference

Proposed control: requires review

Select an upstream conventional reference after reviewing the full protocol

Protocol-specific applicability, permitted inputs, access, split, evaluator and execution requirements need review before implementation or execution.

This is a suggested selection rule, not a validated method or a measured score.

Protocol coverage CSV (gzip) · Model evaluation matrix (gzip) · Source table (gzip) · Release and checksums (gzip)

Coverage is derived from release 2026-10-10-6e93f504adfc. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.

Run instructions

No runnable recipe has been reviewed for this protocol. Dataset access, model requirements, licences and compute requirements must be checked against its sources before execution.

Strengths, limitations and unresolved questions

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

0 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-10-10-6e93f504adfc
Property and statementOriginal source and locationReview and provenance

No evidence rows match these filters. Choose another scope or clear the search.

Sources and history

Release 2026-10-10-6e93f504adfc · Record review: source checked

1 source records and release historyDownload this release (gzip)
Technical metadata and extraction receipts

Stable ID: rare-ranking-20261009-protocol-kafkas2025-gpcards-gene-sets

areas
dna-genomes
contexts
clinical_research
protocol
For each genotype-phenotype pair, build candidate sets of 5, 25, 50, 75 or 100 genes containing the causative gene; rank with each method; report Hits@1, Hits@10, ROC AUC and AUPR per set size.
version
Table 3 column 'GPCards'
source locator
Table 3 column 'GPCards'; Methods 'Datasets used' and 'Baseline methods'
limitations
Synthetic candidate sets: the causative gene plus randomly chosen genes, not the filtered variant list of a real exome.; ClinVar phenotypes are OMIM disease annotations from the HPO database, which Exomiser also uses, so the comparison may favour Exomiser (authors' note).; Phenotypes are database annotations, not observed in a patient workup (PAVS is closer to clinical reports).; Exomiser was given one random variant per gene and its variant scores were ignored; only gene scores were compared.; The GPT-4 prompts were designed by the authors, so the GPT-4 rows are author_reported; Exomiser was not developed by them. Two authors (P. N. Schofield, R. Hoehndorf) developed DeepPVP, a competing tool that is not in Table 3.; No comparator: Exomiser is not applicable because phenotypes are free text (table footnote).; Which zero-shot prompt (Q1, Q2, Q3 or Q2Q3) produced the zero-shot columns is not stated; on GPCards with GPT-3.5 the prompts ranged from 30% to 80% Hits@1 (Table 2).; The random candidate genes are drawn from all human genes or from genes with a genotype in the benchmark set (Methods 'Datasets used' paragraph 4); 'Evaluation procedure' paragraph 2 says all human genes. Which pool was used for Table 3 is not stated.; The prompts were chosen on GPCards (with GPT-3.5-turbo, Table 2), so the GPCards GPT-4 values are in-sample for prompt choice.
Related records

Suggest a correction