PertEval-scFM Norman single-gene (2,000 HVGs) AUSPC across SPECTRA sparsification splits
Area Under the SPECTRA Performance Curve (AUSPC) of mean squared error (MSE) scoring the perturbation effect delta := P - Xc (Eq. 5; P = perturbed expression, Xc = control expression), described by the source as the 'log fold change perturbation effect' -- not raw post-perturbation expression itself -- integrated by the trapezoidal rule across seven SPECTRA sparsification-probability train-test splits (s in {0.1,0.2,...,0.7}); lower AUSPC is better.
Overview
Area Under the SPECTRA Performance Curve (AUSPC) of mean squared error (MSE) scoring the perturbation effect delta := P - Xc (Eq. 5; P = perturbed expression, Xc = control expression), described by the source as the 'log fold change perturbation effect' -- not raw post-perturbation expression itself -- integrated by the trapezoidal rule across seven SPECTRA sparsification-probability train-test splits (s in {0.1,0.2,...,0.7}); lower AUSPC is better.
Consult the linked sources for architecture or protocol details. Missing evidence is not evidence of a missing capability.
3 recorded evaluations, 3 metric rows. A comparison chart has not yet been validated for these results. The table retains the individual findings and their sources.
Results
Results are available, but no reviewed comparison panel is linked in this release.
All evaluations
3 evaluations · 3 results. Different protocols are not a single leaderboard.
Filter evaluations
Applied filters: All linked evaluations
| Tested configuration | Protocol and dataset | Finding | Evidence and details |
|---|---|---|---|
| Configuration: GEARS (trained from scratch, no pretrained weights) | Protocol: PertEval-scFM Norman single-gene (2,000 HVGs) AUSPC across SPECTRA sparsification splits Dataset subset: Norman et al. 2019 single-gene perturbations, K562, top 2,000 HVGs (PertEval-scFM Table 1) | 0.815 ± 0.039 AUSPC 10^-2 (printed column header units) · lower Uncertainty: type: author_reported_propagated_standard_error; reported spread: 0.039; note: The source describes this quantity as a standard error (main-text Figure 2 caption: 'Average AUSPC (down-arrow) across sparsification probabilities for each model with standard error bars') and separately gives its own propagation formula (Appendix F.2, Eqs. F3-F5): AUSPC's uncertainty is derived from each split's own MSE uncertainty via the trapezoidal integral's partial derivatives (sigma^2 = sum_i (d/2)^2 * sigma_phi_i^2, where d=0.1 is the fixed sparsification step size). These two author statements describe the same quantity and are not in conflict: a propagated quantity can correctly be reported as a standard error. This is recorded as the author's own reported, propagated uncertainty; the F3-F5 derivation is the authors' own formula and its mathematical correctness has not been independently validated here. It must not be read as an independently resampled model-seed standard deviation or a confidence interval. The underlying per-split uncertainty is attributed by the source to triplicate experiments per model (Appendix I, Figure I1 caption: 'Experiments were carried out in triplicate for each model'), not to the main-text Figure 2 region. Figure I1's own caption separately states '8 train-test splits of increasing difficulty' for this same Norman single-gene evaluation, while Table 1 prints seven S-columns (S0.1-S0.7) and Appendix F.2 describes the sparsification probabilities as spanning 0.1 to 0.7. This 7-vs-8 discrepancy between the Figure I1 caption and the Table 1 / F.2 grid is preserved exactly as printed, not resolved; it must not be read as establishing an eighth Table 1 column or a confirmed n_runs=7, and no significance claim is made from any overlapping error bars. Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceGEARS on PertEval-scFM Norman single-gene (2,000 HVGs) AUSPC Trapezoidal-rule AUSPC of MSE, scored on the perturbation-effect delta=P-Xc (Eq. 5), across seven SPECTRA sparsification splits (s=0.1..0.7), Norman single-gene, 2,000 HVGs. Aggregation: Not reported PertEval-scFM (Wenteler et al., ICML 2025), full text · Table 1, Norman single-gene section, row GEARS, column AUSPC (10^-2). |
| Configuration: Mean baseline (context mean, no perturbation-specific effect) | Protocol: PertEval-scFM Norman single-gene (2,000 HVGs) AUSPC across SPECTRA sparsification splits Dataset subset: Norman et al. 2019 single-gene perturbations, K562, top 2,000 HVGs (PertEval-scFM Table 1) | 4.612 ± 0.317 AUSPC 10^-2 (printed column header units) · lower Uncertainty: type: author_reported_propagated_standard_error; reported spread: 0.317; note: The source describes this quantity as a standard error (main-text Figure 2 caption: 'Average AUSPC (down-arrow) across sparsification probabilities for each model with standard error bars') and separately gives its own propagation formula (Appendix F.2, Eqs. F3-F5): AUSPC's uncertainty is derived from each split's own MSE uncertainty via the trapezoidal integral's partial derivatives (sigma^2 = sum_i (d/2)^2 * sigma_phi_i^2, where d=0.1 is the fixed sparsification step size). These two author statements describe the same quantity and are not in conflict: a propagated quantity can correctly be reported as a standard error. This is recorded as the author's own reported, propagated uncertainty; the F3-F5 derivation is the authors' own formula and its mathematical correctness has not been independently validated here. It must not be read as an independently resampled model-seed standard deviation or a confidence interval. The underlying per-split uncertainty is attributed by the source to triplicate experiments per model (Appendix I, Figure I1 caption: 'Experiments were carried out in triplicate for each model'), not to the main-text Figure 2 region. Figure I1's own caption separately states '8 train-test splits of increasing difficulty' for this same Norman single-gene evaluation, while Table 1 prints seven S-columns (S0.1-S0.7) and Appendix F.2 describes the sparsification probabilities as spanning 0.1 to 0.7. This 7-vs-8 discrepancy between the Figure I1 caption and the Table 1 / F.2 grid is preserved exactly as printed, not resolved; it must not be read as establishing an eighth Table 1 column or a confirmed n_runs=7, and no significance claim is made from any overlapping error bars. Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceMean baseline on PertEval-scFM Norman single-gene (2,000 HVGs) AUSPC Trapezoidal-rule AUSPC of MSE, scored on the perturbation-effect delta=P-Xc (Eq. 5), across seven SPECTRA sparsification splits (s=0.1..0.7), Norman single-gene, 2,000 HVGs. Aggregation: Not reported PertEval-scFM (Wenteler et al., ICML 2025), full text · Table 1, Norman single-gene section, row Mean baseline, column AUSPC (10^-2). |
| Configuration: MLP baseline (raw control expression + gene co-expression input, Eq. 3) | Protocol: PertEval-scFM Norman single-gene (2,000 HVGs) AUSPC across SPECTRA sparsification splits Dataset subset: Norman et al. 2019 single-gene perturbations, K562, top 2,000 HVGs (PertEval-scFM Table 1) | 4.484 ± 0.299 AUSPC 10^-2 (printed column header units) · lower Uncertainty: type: author_reported_propagated_standard_error; reported spread: 0.299; note: The source describes this quantity as a standard error (main-text Figure 2 caption: 'Average AUSPC (down-arrow) across sparsification probabilities for each model with standard error bars') and separately gives its own propagation formula (Appendix F.2, Eqs. F3-F5): AUSPC's uncertainty is derived from each split's own MSE uncertainty via the trapezoidal integral's partial derivatives (sigma^2 = sum_i (d/2)^2 * sigma_phi_i^2, where d=0.1 is the fixed sparsification step size). These two author statements describe the same quantity and are not in conflict: a propagated quantity can correctly be reported as a standard error. This is recorded as the author's own reported, propagated uncertainty; the F3-F5 derivation is the authors' own formula and its mathematical correctness has not been independently validated here. It must not be read as an independently resampled model-seed standard deviation or a confidence interval. The underlying per-split uncertainty is attributed by the source to triplicate experiments per model (Appendix I, Figure I1 caption: 'Experiments were carried out in triplicate for each model'), not to the main-text Figure 2 region. Figure I1's own caption separately states '8 train-test splits of increasing difficulty' for this same Norman single-gene evaluation, while Table 1 prints seven S-columns (S0.1-S0.7) and Appendix F.2 describes the sparsification probabilities as spanning 0.1 to 0.7. This 7-vs-8 discrepancy between the Figure I1 caption and the Table 1 / F.2 grid is preserved exactly as printed, not resolved; it must not be read as establishing an eighth Table 1 column or a confirmed n_runs=7, and no significance claim is made from any overlapping error bars. Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceMLP baseline on PertEval-scFM Norman single-gene (2,000 HVGs) AUSPC Trapezoidal-rule AUSPC of MSE, scored on the perturbation-effect delta=P-Xc (Eq. 5), across seven SPECTRA sparsification splits (s=0.1..0.7), Norman single-gene, 2,000 HVGs. Aggregation: Not reported PertEval-scFM (Wenteler et al., ICML 2025), full text · Table 1, Norman single-gene section, row MLP baseline, column AUSPC (10^-2). |
Source checking is not independent reproduction. Release 2026-10-07-1448159e6a81.
Methods and evaluation design
Procedure, tasks and evaluated configurations
Recorded evaluations
Each evaluation records what was tested and under which conditions.
Baseline coverage
Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.
0 of 2 active baseline roles have published Rewire measurements in this release. Measurements on a selected protocol do not establish coverage of an entire suite.
No execution recipe linked to this protocol. Recipe availability does not establish a completed evaluation.
- Author-reported evaluations
- 3
Literature evidence is not a Rewire measurement. Executed but unpublished runs and private review status are not included.
Null control
Proposed control: requires review
No-change prediction under matched control conditions
Protocol-specific applicability, permitted inputs, access, split, evaluator and execution requirements need review before implementation or execution.
This is a suggested selection rule, not a validated method or a measured score.
Conventional reference
Proposed control: requires review
Training-only mean-effect or linear prediction
Protocol-specific applicability, permitted inputs, access, split, evaluator and execution requirements need review before implementation or execution.
This is a suggested selection rule, not a validated method or a measured score.
Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums
Coverage is derived from release 2026-10-07-1448159e6a81. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.
Run instructions
No runnable recipe has been reviewed for this protocol. Dataset access, model requirements, licences and compute requirements must be checked against its sources before execution.
Strengths, limitations and unresolved questions
Evidence
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Evidence table
Inspect claims, sources and review details
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
0 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|
No evidence rows match these filters. Choose another scope or clear the search.
Sources and history
View linked audit checks and correction history
Release 2026-10-07-1448159e6a81 · Record review: source checked
1 source records and release history
- PertEval-scFM (Wenteler et al., ICML 2025), full text · Original source · PMLR v267 wenteler25a (as served by the PMLR-affiliated mlresearch/v267 GitHub mirror; ETag "ed9f0fe44cf6edc939ee6950d024f65dcd8dd6f16414bd9f5297cda9395d6e58" at retrieval)
Technical metadata and extraction receipts
Stable ID: perteval-scfm-2025-protocol-norman-single-2000hvg-auspc
- areas
- cells-spatial-multiomics
- tasks
- Norman single-gene perturbation effect prediction, 2,000 HVGs, SPECTRA distribution shift
- metric
- AUSPC (Area Under the SPECTRA Performance Curve of MSE)
- metric direction
- lower
- unit
- printed column units: MSE (10^-2) per split; AUSPC (10^-2)
- protocol
- SPECTRA (Ektefaie et al., 2024) generates seven train-test splits of increasing sparsification probability s=0.1..0.7 (step 0.1), each a controlled-overlap distribution-shift condition for unseen perturbations (Fig. F1) -- this protocol concerns generalisation across perturbations under increasing distribution shift, not a claim about forecasting unseen cells. MSE is computed at each split on the delta=P-Xc perturbation-effect target (Eq. 5), not on raw post-perturbation expression (Table 1 columns S0.1-S0.7); AUSPC is np.trapz(phi, s) over those seven points (Appendix F.2, Eq. F2, Algorithm 1) -- not a ROC-AUC, accuracy or rank-based score. This delta/expression-change MSE is a different metric construction from GEARS Supplementary Table 6's own Pearson-delta-DE metric and must not be conflated with it.
- source locator
- Table 1 (Norman single-gene section); Appendix A.1 (dataset); Section 2.1 ('Raw expression data', Eq. 3, MLP baseline input); Section 2.2 ('MLP baseline' Eq. 4, 'GEARS baseline' Eq. 5, 'Mean baseline'); Appendix F.2 and Algorithm 1 (AUSPC/uncertainty propagation definitions, Eqs. F2-F5); Section 2.1.1 (HVG selection); Section 2.3.2 (SPECTRA splits); main-text Figure 2 caption ('standard error bars'); Appendix I, Figure I1 caption (triplicate experiments, '8 train-test splits').
- configurations in this intake
- GEARS; MLP baseline; Mean baseline
- configurations in same table not intaken
- Geneformer; scBERT; scFoundation; scGPT; UCE
- missing metadata
- per split scored sample counts: unreported; see dataset record; random seeds: unreported beyond the word 'triplicate'; exact seed values not printed; hyperparameters: GEARS: official implementation defaults, trained from scratch without pretrained weights (paper's own Methods wording, Section 2.2 region: 'we train GEARS from scratch without using pretrained weights'); MLP baseline / Mean baseline: architecture per Eqs. 4-6 and 'Mean baseline' text in Section 2.2, exact training hyperparameters not separately tabulated for these two baselines in the main text
- caveats
- This is a different dataset preprocessing (2,000 HVGs, not GEARS Supplementary Table 6's own unstated gene subset), a different split mechanism (SPECTRA sparsification, not Table 6's unstated split) and a different metric construction (AUSPC, a trapezoidal integral of MSE over seven distribution-shift splits, not Table 6's single-split MSE/Pearson-DeltaExpression) from the existing GEARS Supplementary Table 6 catalogue entries for this same use case (protocols gears-2023-supp-table6-task-mse / -task-pearson-de). Do not merge, average or otherwise combine these values with Table 6's.; GEARS in this protocol is trained from scratch by this paper's authors (an independent execution using the official GEARS implementation's default hyperparameters apart from the SPECTRA split), not the original GEARS authors' own reported Table 6 checkpoint/run. This is a second, independent GEARS evaluation, not a reproduction or replication claim for Table 6.; The same Table 1's double-gene section prints a GEARS AUSPC of 0.808 and a printed Mean-baseline-relative delta of 4.254, which is not exactly reproducible from the printed Mean-baseline AUSPC of 4.255 (4.255-0.808=3.447, not 4.254); this inconsistency is specific to the double-gene row, is not resolved or derived here, and is explicitly out of scope for this intake, which covers the single-gene row only.; Table 2 (Replogle RPE1) separately shows a running-text/table-row labeling conflict (text attributes an AUSPC of 0.1251 to the 'Mean baseline', while the table's own row data assigns 0.1251 to Geneformer and 0.1341 to the Mean baseline). This conflict concerns a different dataset (RPE1) not covered by this intake and is noted in the accompanying dossier, not resolved here.; Input populations differ across the three intaken configurations and are not an identical input budget: GEARS uses a graph-based architecture over raw expression and prior-knowledge gene relationships (official implementation); the MLP baseline's input ZGE = Xc (+) Gc concatenates raw control expression with a gene-co-expression matrix for the perturbed gene(s) (Eq. 3), not raw expression alone; the Mean baseline uses no learned input at all (see its own configuration record for the paper's exact, population-ambiguous definition). None of the three should be assumed to have received comparable information.; This protocol concerns generalisation to unseen perturbations under SPECTRA's increasing-sparsification distribution shift. It does not test, and does not establish, forecasting for unseen cells, donors, or cell-line/context transfer.
- uncertainty definition
- The source describes this quantity as a standard error (main-text Figure 2 caption: 'Average AUSPC (down-arrow) across sparsification probabilities for each model with standard error bars') and separately gives its own propagation formula (Appendix F.2, Eqs. F3-F5): AUSPC's uncertainty is derived from each split's own MSE uncertainty via the trapezoidal integral's partial derivatives (sigma^2 = sum_i (d/2)^2 * sigma_phi_i^2, where d=0.1 is the fixed sparsification step size). These two author statements describe the same quantity and are not in conflict: a propagated quantity can correctly be reported as a standard error. This is recorded as the author's own reported, propagated uncertainty; the F3-F5 derivation is the authors' own formula and its mathematical correctness has not been independently validated here. It must not be read as an independently resampled model-seed standard deviation or a confidence interval. The underlying per-split uncertainty is attributed by the source to triplicate experiments per model (Appendix I, Figure I1 caption: 'Experiments were carried out in triplicate for each model'), not to the main-text Figure 2 region. Figure I1's own caption separately states '8 train-test splits of increasing difficulty' for this same Norman single-gene evaluation, while Table 1 prints seven S-columns (S0.1-S0.7) and Appendix F.2 describes the sparsification probabilities as spanning 0.1 to 0.7. This 7-vs-8 discrepancy between the Figure I1 caption and the Table 1 / F.2 grid is preserved exactly as printed, not resolved; it must not be read as establishing an eighth Table 1 column or a confirmed n_runs=7, and no significance claim is made from any overlapping error bars.