rewirebio.iobenchmarks
Dataset

OncoKB-annotated somatic missense mutations (oncogenic, likely oncogenic and likely neutral)

Benchmark 2 truth set in Chen et al. 2020.

Research readiness

These checks assess whether the evidence supports a reproducible investigation. A source-checked score alone does not meet these requirements.

Release 2026-10-10-7fcc3e48a123 · Evidence verified: Not verified

Evidence incomplete

Replay metrics

Exact outcomes, predictions, identifiers and evaluator are connected.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • metric replay: verification is missing

Verified: Not verified

Evidence incomplete

Investigate discrepancies

Replay evidence includes annotations and an assessment of dependence. Unknown independence permits descriptive analysis only.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • metric replay: verification is missing
  • annotations: verification is missing
  • dependence: verification is missing

Verified: Not verified

Evidence incomplete

Run locally

A pinned recipe describes the inputs, environment and resource requirements.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • recipe pinned: verification is missing
  • resource estimate: verification is missing

Verified: Not verified

Evidence incomplete

Validate independently

Separate data and exposure records support an independent test.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • independent validation: verification is missing
  • overlap checked: verification is missing

Verified: Not verified

Readiness describes the evidence in this release. Availability on your computer is checked separately when an investigation runs. Existing data exposure can prevent independent validation even when files are available.

Artifacts and reproduction

No verified artifact manifest is connected to this record yet. The gaps above identify what is needed before analysis can begin.

Read reviewed discrepancy investigations

Evaluation results

50 evaluations · 250 results. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: CADD (Chen et al. 2020)Protocol: OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020 Additional file 9)
Dataset: OncoKB-annotated somatic missense mutations (oncogenic, likely oncogenic and likely neutral)
0.64 accuracy
fraction · higher

Uncertainty: Not yet extracted: Printed as 'mean (lower-upper)' under the column header '(±2σ)': per Methods, the mean and two standard deviations over 100 random draws. That is not a confidence interval, and the schema's standard_deviation type would need a standard deviation derived from rounded bounds, so no structured uncertainty is recorded. The printed range is kept in printed_source_cell; every range is symmetric about the mean to rounding.

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

CADD on OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020)

somatic-oncogenicity-20261009-protocol-chen2020-oncokb-median

Aggregation: Not reported

Comprehensive assessment of computational algorithms in predicting cancer driver mutations; Chen et al. 2020, Additional file 9: Performance metrics of 33 algorithms, median-score threshold, benchmark 2 (OncoKB) · Additional file 9 (sheet 'Additional_file_9'), B16; Algorithm 'CADD'; column 'Accuracy (±2σ)'
Configuration: CADD (Chen et al. 2020)Protocol: OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020 Additional file 9)
Dataset: OncoKB-annotated somatic missense mutations (oncogenic, likely oncogenic and likely neutral)
0.64 negative-predictive-value
fraction · higher

Uncertainty: Not yet extracted: Printed as 'mean (lower-upper)' under the column header '(±2σ)': per Methods, the mean and two standard deviations over 100 random draws. That is not a confidence interval, and the schema's standard_deviation type would need a standard deviation derived from rounded bounds, so no structured uncertainty is recorded. The printed range is kept in printed_source_cell; every range is symmetric about the mean to rounding.

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

CADD on OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020)

somatic-oncogenicity-20261009-protocol-chen2020-oncokb-median

Aggregation: Not reported

Comprehensive assessment of computational algorithms in predicting cancer driver mutations; Chen et al. 2020, Additional file 9: Performance metrics of 33 algorithms, median-score threshold, benchmark 2 (OncoKB) · Additional file 9 (sheet 'Additional_file_9'), F16; Algorithm 'CADD'; column 'NPV (±2σ)'
Configuration: CADD (Chen et al. 2020)Protocol: OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020 Additional file 9)
Dataset: OncoKB-annotated somatic missense mutations (oncogenic, likely oncogenic and likely neutral)
0.64 precision
fraction · higher

Uncertainty: Not yet extracted: Printed as 'mean (lower-upper)' under the column header '(±2σ)': per Methods, the mean and two standard deviations over 100 random draws. That is not a confidence interval, and the schema's standard_deviation type would need a standard deviation derived from rounded bounds, so no structured uncertainty is recorded. The printed range is kept in printed_source_cell; every range is symmetric about the mean to rounding.

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

CADD on OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020)

somatic-oncogenicity-20261009-protocol-chen2020-oncokb-median

Aggregation: Not reported

Comprehensive assessment of computational algorithms in predicting cancer driver mutations; Chen et al. 2020, Additional file 9: Performance metrics of 33 algorithms, median-score threshold, benchmark 2 (OncoKB) · Additional file 9 (sheet 'Additional_file_9'), E16; Algorithm 'CADD'; column 'PPV (±2σ)'
Configuration: CADD (Chen et al. 2020)Protocol: OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020 Additional file 9)
Dataset: OncoKB-annotated somatic missense mutations (oncogenic, likely oncogenic and likely neutral)
0.64 recall
fraction · higher

Uncertainty: Not yet extracted: Printed as 'mean (lower-upper)' under the column header '(±2σ)': per Methods, the mean and two standard deviations over 100 random draws. That is not a confidence interval, and the schema's standard_deviation type would need a standard deviation derived from rounded bounds, so no structured uncertainty is recorded. The printed range is kept in printed_source_cell; every range is symmetric about the mean to rounding.

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

CADD on OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020)

somatic-oncogenicity-20261009-protocol-chen2020-oncokb-median

Aggregation: Not reported

Comprehensive assessment of computational algorithms in predicting cancer driver mutations; Chen et al. 2020, Additional file 9: Performance metrics of 33 algorithms, median-score threshold, benchmark 2 (OncoKB) · Additional file 9 (sheet 'Additional_file_9'), C16; Algorithm 'CADD'; column 'Sensitivity (±2σ)'
Configuration: CADD (Chen et al. 2020)Protocol: OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020 Additional file 9)
Dataset: OncoKB-annotated somatic missense mutations (oncogenic, likely oncogenic and likely neutral)
0.64 specificity
fraction · higher

Uncertainty: Not yet extracted: Printed as 'mean (lower-upper)' under the column header '(±2σ)': per Methods, the mean and two standard deviations over 100 random draws. That is not a confidence interval, and the schema's standard_deviation type would need a standard deviation derived from rounded bounds, so no structured uncertainty is recorded. The printed range is kept in printed_source_cell; every range is symmetric about the mean to rounding.

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

CADD on OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020)

somatic-oncogenicity-20261009-protocol-chen2020-oncokb-median

Aggregation: Not reported

Comprehensive assessment of computational algorithms in predicting cancer driver mutations; Chen et al. 2020, Additional file 9: Performance metrics of 33 algorithms, median-score threshold, benchmark 2 (OncoKB) · Additional file 9 (sheet 'Additional_file_9'), D16; Algorithm 'CADD'; column 'Specificity (±2σ)'
Configuration: CanDrA (Chen et al. 2020)Protocol: OncoKB benchmark (positive set not stated), default categorical calls (Chen et al. 2020 Additional file 10)
Dataset: OncoKB-annotated somatic missense mutations (oncogenic, likely oncogenic and likely neutral)
0.49 accuracy
fraction · higher

Uncertainty: Not yet extracted: Printed as 'mean (lower-upper)' under the column header '(±2σ)': per Methods, the mean and two standard deviations over 100 random draws. That is not a confidence interval, and the schema's standard_deviation type would need a standard deviation derived from rounded bounds, so no structured uncertainty is recorded. The printed range is kept in printed_source_cell; every range is symmetric about the mean to rounding.

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CanDrA on OncoKB benchmark (positive set not stated), default categorical calls (Chen et al. 2020)

somatic-oncogenicity-20261009-protocol-chen2020-oncokb-default

Aggregation: Not reported

Comprehensive assessment of computational algorithms in predicting cancer driver mutations; Chen et al. 2020, Additional file 10: Performance metrics of 17 algorithms, default categories, benchmark 2 (OncoKB) · Additional file 10 (sheet 'Additional_file_10'), B18; Algorithm 'CanDrA'; column 'Accuracy (±2σ)'
Configuration: CanDrA (Chen et al. 2020)Protocol: OncoKB benchmark (positive set not stated), default categorical calls (Chen et al. 2020 Additional file 10)
Dataset: OncoKB-annotated somatic missense mutations (oncogenic, likely oncogenic and likely neutral)
0.46 negative-predictive-value
fraction · higher

Uncertainty: Not yet extracted: Printed as 'mean (lower-upper)' under the column header '(±2σ)': per Methods, the mean and two standard deviations over 100 random draws. That is not a confidence interval, and the schema's standard_deviation type would need a standard deviation derived from rounded bounds, so no structured uncertainty is recorded. The printed range is kept in printed_source_cell; every range is symmetric about the mean to rounding.

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CanDrA on OncoKB benchmark (positive set not stated), default categorical calls (Chen et al. 2020)

somatic-oncogenicity-20261009-protocol-chen2020-oncokb-default

Aggregation: Not reported

Comprehensive assessment of computational algorithms in predicting cancer driver mutations; Chen et al. 2020, Additional file 10: Performance metrics of 17 algorithms, default categories, benchmark 2 (OncoKB) · Additional file 10 (sheet 'Additional_file_10'), F18; Algorithm 'CanDrA'; column 'NPV (±2σ)'
Configuration: CanDrA (Chen et al. 2020)Protocol: OncoKB benchmark (positive set not stated), default categorical calls (Chen et al. 2020 Additional file 10)
Dataset: OncoKB-annotated somatic missense mutations (oncogenic, likely oncogenic and likely neutral)
0.49 precision
fraction · higher

Uncertainty: Not yet extracted: Printed as 'mean (lower-upper)' under the column header '(±2σ)': per Methods, the mean and two standard deviations over 100 random draws. That is not a confidence interval, and the schema's standard_deviation type would need a standard deviation derived from rounded bounds, so no structured uncertainty is recorded. The printed range is kept in printed_source_cell; every range is symmetric about the mean to rounding.

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CanDrA on OncoKB benchmark (positive set not stated), default categorical calls (Chen et al. 2020)

somatic-oncogenicity-20261009-protocol-chen2020-oncokb-default

Aggregation: Not reported

Comprehensive assessment of computational algorithms in predicting cancer driver mutations; Chen et al. 2020, Additional file 10: Performance metrics of 17 algorithms, default categories, benchmark 2 (OncoKB) · Additional file 10 (sheet 'Additional_file_10'), E18; Algorithm 'CanDrA'; column 'PPV (±2σ)'
Configuration: CanDrA (Chen et al. 2020)Protocol: OncoKB benchmark (positive set not stated), default categorical calls (Chen et al. 2020 Additional file 10)
Dataset: OncoKB-annotated somatic missense mutations (oncogenic, likely oncogenic and likely neutral)
0.8 recall
fraction · higher

Uncertainty: Not yet extracted: Printed as 'mean (lower-upper)' under the column header '(±2σ)': per Methods, the mean and two standard deviations over 100 random draws. That is not a confidence interval, and the schema's standard_deviation type would need a standard deviation derived from rounded bounds, so no structured uncertainty is recorded. The printed range is kept in printed_source_cell; every range is symmetric about the mean to rounding.

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CanDrA on OncoKB benchmark (positive set not stated), default categorical calls (Chen et al. 2020)

somatic-oncogenicity-20261009-protocol-chen2020-oncokb-default

Aggregation: Not reported

Comprehensive assessment of computational algorithms in predicting cancer driver mutations; Chen et al. 2020, Additional file 10: Performance metrics of 17 algorithms, default categories, benchmark 2 (OncoKB) · Additional file 10 (sheet 'Additional_file_10'), C18; Algorithm 'CanDrA'; column 'Sensitivity (±2σ)'
Configuration: CanDrA (Chen et al. 2020)Protocol: OncoKB benchmark (positive set not stated), default categorical calls (Chen et al. 2020 Additional file 10)
Dataset: OncoKB-annotated somatic missense mutations (oncogenic, likely oncogenic and likely neutral)
0.17 specificity
fraction · higher

Uncertainty: Not yet extracted: Printed as 'mean (lower-upper)' under the column header '(±2σ)': per Methods, the mean and two standard deviations over 100 random draws. That is not a confidence interval, and the schema's standard_deviation type would need a standard deviation derived from rounded bounds, so no structured uncertainty is recorded. The printed range is kept in printed_source_cell; every range is symmetric about the mean to rounding.

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CanDrA on OncoKB benchmark (positive set not stated), default categorical calls (Chen et al. 2020)

somatic-oncogenicity-20261009-protocol-chen2020-oncokb-default

Aggregation: Not reported

Comprehensive assessment of computational algorithms in predicting cancer driver mutations; Chen et al. 2020, Additional file 10: Performance metrics of 17 algorithms, default categories, benchmark 2 (OncoKB) · Additional file 10 (sheet 'Additional_file_10'), D18; Algorithm 'CanDrA'; column 'Specificity (±2σ)'
Configuration: CanDrA (Chen et al. 2020)Protocol: OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020 Additional file 9)
Dataset: OncoKB-annotated somatic missense mutations (oncogenic, likely oncogenic and likely neutral)
0.59 accuracy
fraction · higher

Uncertainty: Not yet extracted: Printed as 'mean (lower-upper)' under the column header '(±2σ)': per Methods, the mean and two standard deviations over 100 random draws. That is not a confidence interval, and the schema's standard_deviation type would need a standard deviation derived from rounded bounds, so no structured uncertainty is recorded. The printed range is kept in printed_source_cell; every range is symmetric about the mean to rounding.

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CanDrA on OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020)

somatic-oncogenicity-20261009-protocol-chen2020-oncokb-median

Aggregation: Not reported

Comprehensive assessment of computational algorithms in predicting cancer driver mutations; Chen et al. 2020, Additional file 9: Performance metrics of 33 algorithms, median-score threshold, benchmark 2 (OncoKB) · Additional file 9 (sheet 'Additional_file_9'), B26; Algorithm 'CanDrA'; column 'Accuracy (±2σ)'
Configuration: CanDrA (Chen et al. 2020)Protocol: OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020 Additional file 9)
Dataset: OncoKB-annotated somatic missense mutations (oncogenic, likely oncogenic and likely neutral)
0.59 negative-predictive-value
fraction · higher

Uncertainty: Not yet extracted: Printed as 'mean (lower-upper)' under the column header '(±2σ)': per Methods, the mean and two standard deviations over 100 random draws. That is not a confidence interval, and the schema's standard_deviation type would need a standard deviation derived from rounded bounds, so no structured uncertainty is recorded. The printed range is kept in printed_source_cell; every range is symmetric about the mean to rounding.

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CanDrA on OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020)

somatic-oncogenicity-20261009-protocol-chen2020-oncokb-median

Aggregation: Not reported

Comprehensive assessment of computational algorithms in predicting cancer driver mutations; Chen et al. 2020, Additional file 9: Performance metrics of 33 algorithms, median-score threshold, benchmark 2 (OncoKB) · Additional file 9 (sheet 'Additional_file_9'), F26; Algorithm 'CanDrA'; column 'NPV (±2σ)'
Configuration: CanDrA (Chen et al. 2020)Protocol: OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020 Additional file 9)
Dataset: OncoKB-annotated somatic missense mutations (oncogenic, likely oncogenic and likely neutral)
0.59 precision
fraction · higher

Uncertainty: Not yet extracted: Printed as 'mean (lower-upper)' under the column header '(±2σ)': per Methods, the mean and two standard deviations over 100 random draws. That is not a confidence interval, and the schema's standard_deviation type would need a standard deviation derived from rounded bounds, so no structured uncertainty is recorded. The printed range is kept in printed_source_cell; every range is symmetric about the mean to rounding.

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CanDrA on OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020)

somatic-oncogenicity-20261009-protocol-chen2020-oncokb-median

Aggregation: Not reported

Comprehensive assessment of computational algorithms in predicting cancer driver mutations; Chen et al. 2020, Additional file 9: Performance metrics of 33 algorithms, median-score threshold, benchmark 2 (OncoKB) · Additional file 9 (sheet 'Additional_file_9'), E26; Algorithm 'CanDrA'; column 'PPV (±2σ)'
Configuration: CanDrA (Chen et al. 2020)Protocol: OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020 Additional file 9)
Dataset: OncoKB-annotated somatic missense mutations (oncogenic, likely oncogenic and likely neutral)
0.59 recall
fraction · higher

Uncertainty: Not yet extracted: Printed as 'mean (lower-upper)' under the column header '(±2σ)': per Methods, the mean and two standard deviations over 100 random draws. That is not a confidence interval, and the schema's standard_deviation type would need a standard deviation derived from rounded bounds, so no structured uncertainty is recorded. The printed range is kept in printed_source_cell; every range is symmetric about the mean to rounding.

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CanDrA on OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020)

somatic-oncogenicity-20261009-protocol-chen2020-oncokb-median

Aggregation: Not reported

Comprehensive assessment of computational algorithms in predicting cancer driver mutations; Chen et al. 2020, Additional file 9: Performance metrics of 33 algorithms, median-score threshold, benchmark 2 (OncoKB) · Additional file 9 (sheet 'Additional_file_9'), C26; Algorithm 'CanDrA'; column 'Sensitivity (±2σ)'
Configuration: CanDrA (Chen et al. 2020)Protocol: OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020 Additional file 9)
Dataset: OncoKB-annotated somatic missense mutations (oncogenic, likely oncogenic and likely neutral)
0.59 specificity
fraction · higher

Uncertainty: Not yet extracted: Printed as 'mean (lower-upper)' under the column header '(±2σ)': per Methods, the mean and two standard deviations over 100 random draws. That is not a confidence interval, and the schema's standard_deviation type would need a standard deviation derived from rounded bounds, so no structured uncertainty is recorded. The printed range is kept in printed_source_cell; every range is symmetric about the mean to rounding.

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CanDrA on OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020)

somatic-oncogenicity-20261009-protocol-chen2020-oncokb-median

Aggregation: Not reported

Comprehensive assessment of computational algorithms in predicting cancer driver mutations; Chen et al. 2020, Additional file 9: Performance metrics of 33 algorithms, median-score threshold, benchmark 2 (OncoKB) · Additional file 9 (sheet 'Additional_file_9'), D26; Algorithm 'CanDrA'; column 'Specificity (±2σ)'
Configuration: CHASM (Chen et al. 2020)Protocol: OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020 Additional file 9)
Dataset: OncoKB-annotated somatic missense mutations (oncogenic, likely oncogenic and likely neutral)
0.65 accuracy
fraction · higher

Uncertainty: Not yet extracted: Printed as 'mean (lower-upper)' under the column header '(±2σ)': per Methods, the mean and two standard deviations over 100 random draws. That is not a confidence interval, and the schema's standard_deviation type would need a standard deviation derived from rounded bounds, so no structured uncertainty is recorded. The printed range is kept in printed_source_cell; every range is symmetric about the mean to rounding.

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

CHASM on OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020)

somatic-oncogenicity-20261009-protocol-chen2020-oncokb-median

Aggregation: Not reported

Comprehensive assessment of computational algorithms in predicting cancer driver mutations; Chen et al. 2020, Additional file 9: Performance metrics of 33 algorithms, median-score threshold, benchmark 2 (OncoKB) · Additional file 9 (sheet 'Additional_file_9'), B15; Algorithm 'CHASM'; column 'Accuracy (±2σ)'
Configuration: CHASM (Chen et al. 2020)Protocol: OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020 Additional file 9)
Dataset: OncoKB-annotated somatic missense mutations (oncogenic, likely oncogenic and likely neutral)
0.65 negative-predictive-value
fraction · higher

Uncertainty: Not yet extracted: Printed as 'mean (lower-upper)' under the column header '(±2σ)': per Methods, the mean and two standard deviations over 100 random draws. That is not a confidence interval, and the schema's standard_deviation type would need a standard deviation derived from rounded bounds, so no structured uncertainty is recorded. The printed range is kept in printed_source_cell; every range is symmetric about the mean to rounding.

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

CHASM on OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020)

somatic-oncogenicity-20261009-protocol-chen2020-oncokb-median

Aggregation: Not reported

Comprehensive assessment of computational algorithms in predicting cancer driver mutations; Chen et al. 2020, Additional file 9: Performance metrics of 33 algorithms, median-score threshold, benchmark 2 (OncoKB) · Additional file 9 (sheet 'Additional_file_9'), F15; Algorithm 'CHASM'; column 'NPV (±2σ)'
Configuration: CHASM (Chen et al. 2020)Protocol: OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020 Additional file 9)
Dataset: OncoKB-annotated somatic missense mutations (oncogenic, likely oncogenic and likely neutral)
0.65 precision
fraction · higher

Uncertainty: Not yet extracted: Printed as 'mean (lower-upper)' under the column header '(±2σ)': per Methods, the mean and two standard deviations over 100 random draws. That is not a confidence interval, and the schema's standard_deviation type would need a standard deviation derived from rounded bounds, so no structured uncertainty is recorded. The printed range is kept in printed_source_cell; every range is symmetric about the mean to rounding.

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

CHASM on OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020)

somatic-oncogenicity-20261009-protocol-chen2020-oncokb-median

Aggregation: Not reported

Comprehensive assessment of computational algorithms in predicting cancer driver mutations; Chen et al. 2020, Additional file 9: Performance metrics of 33 algorithms, median-score threshold, benchmark 2 (OncoKB) · Additional file 9 (sheet 'Additional_file_9'), E15; Algorithm 'CHASM'; column 'PPV (±2σ)'
Configuration: CHASM (Chen et al. 2020)Protocol: OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020 Additional file 9)
Dataset: OncoKB-annotated somatic missense mutations (oncogenic, likely oncogenic and likely neutral)
0.65 recall
fraction · higher

Uncertainty: Not yet extracted: Printed as 'mean (lower-upper)' under the column header '(±2σ)': per Methods, the mean and two standard deviations over 100 random draws. That is not a confidence interval, and the schema's standard_deviation type would need a standard deviation derived from rounded bounds, so no structured uncertainty is recorded. The printed range is kept in printed_source_cell; every range is symmetric about the mean to rounding.

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

CHASM on OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020)

somatic-oncogenicity-20261009-protocol-chen2020-oncokb-median

Aggregation: Not reported

Comprehensive assessment of computational algorithms in predicting cancer driver mutations; Chen et al. 2020, Additional file 9: Performance metrics of 33 algorithms, median-score threshold, benchmark 2 (OncoKB) · Additional file 9 (sheet 'Additional_file_9'), C15; Algorithm 'CHASM'; column 'Sensitivity (±2σ)'
Configuration: CHASM (Chen et al. 2020)Protocol: OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020 Additional file 9)
Dataset: OncoKB-annotated somatic missense mutations (oncogenic, likely oncogenic and likely neutral)
0.64 specificity
fraction · higher

Uncertainty: Not yet extracted: Printed as 'mean (lower-upper)' under the column header '(±2σ)': per Methods, the mean and two standard deviations over 100 random draws. That is not a confidence interval, and the schema's standard_deviation type would need a standard deviation derived from rounded bounds, so no structured uncertainty is recorded. The printed range is kept in printed_source_cell; every range is symmetric about the mean to rounding.

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

CHASM on OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020)

somatic-oncogenicity-20261009-protocol-chen2020-oncokb-median

Aggregation: Not reported

Comprehensive assessment of computational algorithms in predicting cancer driver mutations; Chen et al. 2020, Additional file 9: Performance metrics of 33 algorithms, median-score threshold, benchmark 2 (OncoKB) · Additional file 9 (sheet 'Additional_file_9'), D15; Algorithm 'CHASM'; column 'Specificity (±2σ)'
Configuration: CTAT-cancer (Chen et al. 2020)Protocol: OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020 Additional file 9)
Dataset: OncoKB-annotated somatic missense mutations (oncogenic, likely oncogenic and likely neutral)
0.63 accuracy
fraction · higher

Uncertainty: Not yet extracted: Printed as 'mean (lower-upper)' under the column header '(±2σ)': per Methods, the mean and two standard deviations over 100 random draws. That is not a confidence interval, and the schema's standard_deviation type would need a standard deviation derived from rounded bounds, so no structured uncertainty is recorded. The printed range is kept in printed_source_cell; every range is symmetric about the mean to rounding.

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

CTAT-cancer on OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020)

somatic-oncogenicity-20261009-protocol-chen2020-oncokb-median

Aggregation: Not reported

Comprehensive assessment of computational algorithms in predicting cancer driver mutations; Chen et al. 2020, Additional file 9: Performance metrics of 33 algorithms, median-score threshold, benchmark 2 (OncoKB) · Additional file 9 (sheet 'Additional_file_9'), B18; Algorithm 'CTAT-cancer'; column 'Accuracy (±2σ)'
Configuration: CTAT-cancer (Chen et al. 2020)Protocol: OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020 Additional file 9)
Dataset: OncoKB-annotated somatic missense mutations (oncogenic, likely oncogenic and likely neutral)
0.63 negative-predictive-value
fraction · higher

Uncertainty: Not yet extracted: Printed as 'mean (lower-upper)' under the column header '(±2σ)': per Methods, the mean and two standard deviations over 100 random draws. That is not a confidence interval, and the schema's standard_deviation type would need a standard deviation derived from rounded bounds, so no structured uncertainty is recorded. The printed range is kept in printed_source_cell; every range is symmetric about the mean to rounding.

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

CTAT-cancer on OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020)

somatic-oncogenicity-20261009-protocol-chen2020-oncokb-median

Aggregation: Not reported

Comprehensive assessment of computational algorithms in predicting cancer driver mutations; Chen et al. 2020, Additional file 9: Performance metrics of 33 algorithms, median-score threshold, benchmark 2 (OncoKB) · Additional file 9 (sheet 'Additional_file_9'), F18; Algorithm 'CTAT-cancer'; column 'NPV (±2σ)'
Configuration: CTAT-cancer (Chen et al. 2020)Protocol: OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020 Additional file 9)
Dataset: OncoKB-annotated somatic missense mutations (oncogenic, likely oncogenic and likely neutral)
0.63 precision
fraction · higher

Uncertainty: Not yet extracted: Printed as 'mean (lower-upper)' under the column header '(±2σ)': per Methods, the mean and two standard deviations over 100 random draws. That is not a confidence interval, and the schema's standard_deviation type would need a standard deviation derived from rounded bounds, so no structured uncertainty is recorded. The printed range is kept in printed_source_cell; every range is symmetric about the mean to rounding.

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

CTAT-cancer on OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020)

somatic-oncogenicity-20261009-protocol-chen2020-oncokb-median

Aggregation: Not reported

Comprehensive assessment of computational algorithms in predicting cancer driver mutations; Chen et al. 2020, Additional file 9: Performance metrics of 33 algorithms, median-score threshold, benchmark 2 (OncoKB) · Additional file 9 (sheet 'Additional_file_9'), E18; Algorithm 'CTAT-cancer'; column 'PPV (±2σ)'
Configuration: CTAT-cancer (Chen et al. 2020)Protocol: OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020 Additional file 9)
Dataset: OncoKB-annotated somatic missense mutations (oncogenic, likely oncogenic and likely neutral)
0.63 recall
fraction · higher

Uncertainty: Not yet extracted: Printed as 'mean (lower-upper)' under the column header '(±2σ)': per Methods, the mean and two standard deviations over 100 random draws. That is not a confidence interval, and the schema's standard_deviation type would need a standard deviation derived from rounded bounds, so no structured uncertainty is recorded. The printed range is kept in printed_source_cell; every range is symmetric about the mean to rounding.

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

CTAT-cancer on OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020)

somatic-oncogenicity-20261009-protocol-chen2020-oncokb-median

Aggregation: Not reported

Comprehensive assessment of computational algorithms in predicting cancer driver mutations; Chen et al. 2020, Additional file 9: Performance metrics of 33 algorithms, median-score threshold, benchmark 2 (OncoKB) · Additional file 9 (sheet 'Additional_file_9'), C18; Algorithm 'CTAT-cancer'; column 'Sensitivity (±2σ)'
Configuration: CTAT-cancer (Chen et al. 2020)Protocol: OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020 Additional file 9)
Dataset: OncoKB-annotated somatic missense mutations (oncogenic, likely oncogenic and likely neutral)
0.63 specificity
fraction · higher

Uncertainty: Not yet extracted: Printed as 'mean (lower-upper)' under the column header '(±2σ)': per Methods, the mean and two standard deviations over 100 random draws. That is not a confidence interval, and the schema's standard_deviation type would need a standard deviation derived from rounded bounds, so no structured uncertainty is recorded. The printed range is kept in printed_source_cell; every range is symmetric about the mean to rounding.

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

CTAT-cancer on OncoKB benchmark (positive set not stated), median-score threshold (Chen et al. 2020)

somatic-oncogenicity-20261009-protocol-chen2020-oncokb-median

Aggregation: Not reported

Comprehensive assessment of computational algorithms in predicting cancer driver mutations; Chen et al. 2020, Additional file 9: Performance metrics of 33 algorithms, median-score threshold, benchmark 2 (OncoKB) · Additional file 9 (sheet 'Additional_file_9'), D18; Algorithm 'CTAT-cancer'; column 'Specificity (±2σ)'

Source checking is not independent reproduction. Release 2026-10-10-7fcc3e48a123.

Dataset and evaluation context

A dataset supplies biological observations. The evaluation protocol defines how those observations are split, used and scored.

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

5 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-10-10-7fcc3e48a123
Property and statementOriginal source and locationReview and provenance
attributes.population
Methods: 816 oncogenic, 1,384 likely oncogenic and 421 likely neutral mutations (271 inconclusive excluded); Results: 773 oncogenic and 497 likely neutral used in the main comparison
Context-only references
Comprehensive assessment of computational algorithms in predicting cancer driver mutations

Original source ↗

Results 'Benchmark 2'; Methods 'OncoKB annotation benchmark' and 'Calculation of five evaluation metrics'

Version: Genome Biology 21:43, published 2020-02-20; PMC7033911 full-text XML
Retrieved: 2026-10-09T20:45:50Z

not individually reviewed

No individual claim review recorded

Audit details

Field: attributes.population

Source artifact SHA-256: fef52f70c3a0ff3f82902f87080933c220e12274787de6309b41ee283daf8b8e

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

attributes.source_locator
Results 'Benchmark 2'; Methods 'OncoKB annotation benchmark' and 'Calculation of five evaluation metrics'
Context-only references
Comprehensive assessment of computational algorithms in predicting cancer driver mutations

Original source ↗

Results 'Benchmark 2'; Methods 'OncoKB annotation benchmark' and 'Calculation of five evaluation metrics'

Version: Genome Biology 21:43, published 2020-02-20; PMC7033911 full-text XML
Retrieved: 2026-10-09T20:45:50Z

not individually reviewed

No individual claim review recorded

Audit details

Field: attributes.source_locator

Source artifact SHA-256: fef52f70c3a0ff3f82902f87080933c220e12274787de6309b41ee283daf8b8e

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

attributes.split
400 positives and 400 negatives drawn at random, 100 times, for the binary metrics
Context-only references
Comprehensive assessment of computational algorithms in predicting cancer driver mutations

Original source ↗

Results 'Benchmark 2'; Methods 'OncoKB annotation benchmark' and 'Calculation of five evaluation metrics'

Version: Genome Biology 21:43, published 2020-02-20; PMC7033911 full-text XML
Retrieved: 2026-10-09T20:45:50Z

not individually reviewed

No individual claim review recorded

Audit details

Field: attributes.split

Source artifact SHA-256: fef52f70c3a0ff3f82902f87080933c220e12274787de6309b41ee283daf8b8e

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

description
Benchmark 2 truth set in Chen et al. 2020.
Context-only references
Comprehensive assessment of computational algorithms in predicting cancer driver mutations

Original source ↗

Results 'Benchmark 2'; Methods 'OncoKB annotation benchmark' and 'Calculation of five evaluation metrics'

Version: Genome Biology 21:43, published 2020-02-20; PMC7033911 full-text XML
Retrieved: 2026-10-09T20:45:50Z

not individually reviewed

No individual claim review recorded

Audit details

Field: description

Source artifact SHA-256: fef52f70c3a0ff3f82902f87080933c220e12274787de6309b41ee283daf8b8e

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

name
OncoKB-annotated somatic missense mutations (oncogenic, likely oncogenic and likely neutral)
Context-only references
Comprehensive assessment of computational algorithms in predicting cancer driver mutations

Original source ↗

Results 'Benchmark 2'; Methods 'OncoKB annotation benchmark' and 'Calculation of five evaluation metrics'

Version: Genome Biology 21:43, published 2020-02-20; PMC7033911 full-text XML
Retrieved: 2026-10-09T20:45:50Z

not individually reviewed

No individual claim review recorded

Audit details

Field: name

Source artifact SHA-256: fef52f70c3a0ff3f82902f87080933c220e12274787de6309b41ee283daf8b8e

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

Release 2026-10-10-7fcc3e48a123 · Record review: source checked

1 source records and release historyDownload this release (gzip)
Technical metadata and extraction receipts

Stable ID: somatic-oncogenicity-20261009-data-chen2020-oncokb

areas
dna-genomes
contexts
clinical_research
population
Methods: 816 oncogenic, 1,384 likely oncogenic and 421 likely neutral mutations (271 inconclusive excluded); Results: 773 oncogenic and 497 likely neutral used in the main comparison
split
400 positives and 400 negatives drawn at random, 100 times, for the binary metrics
source locator
Results 'Benchmark 2'; Methods 'OncoKB annotation benchmark' and 'Calculation of five evaluation metrics'
missing metadata
population detail: reason: conflicting; note: Results give 773 oncogenic and 497 likely neutral (and 2,327 oncogenic or likely oncogenic); Methods give 816 oncogenic, 1,384 likely oncogenic and 421 likely neutral. Additional files 9 and 10 do not say which positive set the binary metrics used.; version: reason: unreported; note: OncoKB download date or release not stated
Related records

Suggest a correction