rewirebio.iobenchmarks
Protocol

PertEval-scFM Norman single-gene (2,000 HVGs) AUSPC across SPECTRA sparsification splits

Area Under the SPECTRA Performance Curve (AUSPC) of mean squared error (MSE) scoring the perturbation effect delta := P - Xc (Eq. 5; P = perturbed expression, Xc = control expression), described by the source as the 'log fold change perturbation effect' -- not raw post-perturbation expression itself -- integrated by the trapezoidal rule across seven SPECTRA sparsification-probability train-test splits (s in {0.1,0.2,...,0.7}); lower AUSPC is better.

3 evaluations · 3 results

Overview

Area Under the SPECTRA Performance Curve (AUSPC) of mean squared error (MSE) scoring the perturbation effect delta := P - Xc (Eq. 5; P = perturbed expression, Xc = control expression), described by the source as the 'log fold change perturbation effect' -- not raw post-perturbation expression itself -- integrated by the trapezoidal rule across seven SPECTRA sparsification-probability train-test splits (s in {0.1,0.2,...,0.7}); lower AUSPC is better.

Consult the linked sources for architecture or protocol details. Missing evidence is not evidence of a missing capability.

3 recorded evaluations, 3 metric rows. A comparison chart has not yet been validated for these results. The table retains the individual findings and their sources.

View coverage and remaining gaps across all benchmarks

Results

Results are available, but no reviewed comparison panel is linked in this release.

All evaluations

3 evaluations · 3 results. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: GEARS (trained from scratch, no pretrained weights)Protocol: PertEval-scFM Norman single-gene (2,000 HVGs) AUSPC across SPECTRA sparsification splits
Dataset subset: Norman et al. 2019 single-gene perturbations, K562, top 2,000 HVGs (PertEval-scFM Table 1)
0.815 ± 0.039 AUSPC
10^-2 (printed column header units) · lower

Uncertainty: type: author_reported_propagated_standard_error; reported spread: 0.039; note: The source describes this quantity as a standard error (main-text Figure 2 caption: 'Average AUSPC (down-arrow) across sparsification probabilities for each model with standard error bars') and separately gives its own propagation formula (Appendix F.2, Eqs. F3-F5): AUSPC's uncertainty is derived from each split's own MSE uncertainty via the trapezoidal integral's partial derivatives (sigma^2 = sum_i (d/2)^2 * sigma_phi_i^2, where d=0.1 is the fixed sparsification step size). These two author statements describe the same quantity and are not in conflict: a propagated quantity can correctly be reported as a standard error. This is recorded as the author's own reported, propagated uncertainty; the F3-F5 derivation is the authors' own formula and its mathematical correctness has not been independently validated here. It must not be read as an independently resampled model-seed standard deviation or a confidence interval. The underlying per-split uncertainty is attributed by the source to triplicate experiments per model (Appendix I, Figure I1 caption: 'Experiments were carried out in triplicate for each model'), not to the main-text Figure 2 region. Figure I1's own caption separately states '8 train-test splits of increasing difficulty' for this same Norman single-gene evaluation, while Table 1 prints seven S-columns (S0.1-S0.7) and Appendix F.2 describes the sparsification probabilities as spanning 0.1 to 0.7. This 7-vs-8 discrepancy between the Figure I1 caption and the Table 1 / F.2 grid is preserved exactly as printed, not resolved; it must not be read as establishing an eighth Table 1 column or a confirmed n_runs=7, and no significance claim is made from any overlapping error bars.

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GEARS on PertEval-scFM Norman single-gene (2,000 HVGs) AUSPC

Trapezoidal-rule AUSPC of MSE, scored on the perturbation-effect delta=P-Xc (Eq. 5), across seven SPECTRA sparsification splits (s=0.1..0.7), Norman single-gene, 2,000 HVGs.

Aggregation: Not reported

PertEval-scFM (Wenteler et al., ICML 2025), full text · Table 1, Norman single-gene section, row GEARS, column AUSPC (10^-2).
Configuration: Mean baseline (context mean, no perturbation-specific effect)Protocol: PertEval-scFM Norman single-gene (2,000 HVGs) AUSPC across SPECTRA sparsification splits
Dataset subset: Norman et al. 2019 single-gene perturbations, K562, top 2,000 HVGs (PertEval-scFM Table 1)
4.612 ± 0.317 AUSPC
10^-2 (printed column header units) · lower

Uncertainty: type: author_reported_propagated_standard_error; reported spread: 0.317; note: The source describes this quantity as a standard error (main-text Figure 2 caption: 'Average AUSPC (down-arrow) across sparsification probabilities for each model with standard error bars') and separately gives its own propagation formula (Appendix F.2, Eqs. F3-F5): AUSPC's uncertainty is derived from each split's own MSE uncertainty via the trapezoidal integral's partial derivatives (sigma^2 = sum_i (d/2)^2 * sigma_phi_i^2, where d=0.1 is the fixed sparsification step size). These two author statements describe the same quantity and are not in conflict: a propagated quantity can correctly be reported as a standard error. This is recorded as the author's own reported, propagated uncertainty; the F3-F5 derivation is the authors' own formula and its mathematical correctness has not been independently validated here. It must not be read as an independently resampled model-seed standard deviation or a confidence interval. The underlying per-split uncertainty is attributed by the source to triplicate experiments per model (Appendix I, Figure I1 caption: 'Experiments were carried out in triplicate for each model'), not to the main-text Figure 2 region. Figure I1's own caption separately states '8 train-test splits of increasing difficulty' for this same Norman single-gene evaluation, while Table 1 prints seven S-columns (S0.1-S0.7) and Appendix F.2 describes the sparsification probabilities as spanning 0.1 to 0.7. This 7-vs-8 discrepancy between the Figure I1 caption and the Table 1 / F.2 grid is preserved exactly as printed, not resolved; it must not be read as establishing an eighth Table 1 column or a confirmed n_runs=7, and no significance claim is made from any overlapping error bars.

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

Mean baseline on PertEval-scFM Norman single-gene (2,000 HVGs) AUSPC

Trapezoidal-rule AUSPC of MSE, scored on the perturbation-effect delta=P-Xc (Eq. 5), across seven SPECTRA sparsification splits (s=0.1..0.7), Norman single-gene, 2,000 HVGs.

Aggregation: Not reported

PertEval-scFM (Wenteler et al., ICML 2025), full text · Table 1, Norman single-gene section, row Mean baseline, column AUSPC (10^-2).
Configuration: MLP baseline (raw control expression + gene co-expression input, Eq. 3)Protocol: PertEval-scFM Norman single-gene (2,000 HVGs) AUSPC across SPECTRA sparsification splits
Dataset subset: Norman et al. 2019 single-gene perturbations, K562, top 2,000 HVGs (PertEval-scFM Table 1)
4.484 ± 0.299 AUSPC
10^-2 (printed column header units) · lower

Uncertainty: type: author_reported_propagated_standard_error; reported spread: 0.299; note: The source describes this quantity as a standard error (main-text Figure 2 caption: 'Average AUSPC (down-arrow) across sparsification probabilities for each model with standard error bars') and separately gives its own propagation formula (Appendix F.2, Eqs. F3-F5): AUSPC's uncertainty is derived from each split's own MSE uncertainty via the trapezoidal integral's partial derivatives (sigma^2 = sum_i (d/2)^2 * sigma_phi_i^2, where d=0.1 is the fixed sparsification step size). These two author statements describe the same quantity and are not in conflict: a propagated quantity can correctly be reported as a standard error. This is recorded as the author's own reported, propagated uncertainty; the F3-F5 derivation is the authors' own formula and its mathematical correctness has not been independently validated here. It must not be read as an independently resampled model-seed standard deviation or a confidence interval. The underlying per-split uncertainty is attributed by the source to triplicate experiments per model (Appendix I, Figure I1 caption: 'Experiments were carried out in triplicate for each model'), not to the main-text Figure 2 region. Figure I1's own caption separately states '8 train-test splits of increasing difficulty' for this same Norman single-gene evaluation, while Table 1 prints seven S-columns (S0.1-S0.7) and Appendix F.2 describes the sparsification probabilities as spanning 0.1 to 0.7. This 7-vs-8 discrepancy between the Figure I1 caption and the Table 1 / F.2 grid is preserved exactly as printed, not resolved; it must not be read as establishing an eighth Table 1 column or a confirmed n_runs=7, and no significance claim is made from any overlapping error bars.

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

MLP baseline on PertEval-scFM Norman single-gene (2,000 HVGs) AUSPC

Trapezoidal-rule AUSPC of MSE, scored on the perturbation-effect delta=P-Xc (Eq. 5), across seven SPECTRA sparsification splits (s=0.1..0.7), Norman single-gene, 2,000 HVGs.

Aggregation: Not reported

PertEval-scFM (Wenteler et al., ICML 2025), full text · Table 1, Norman single-gene section, row MLP baseline, column AUSPC (10^-2).

Source checking is not independent reproduction. Release 2026-10-07-1448159e6a81.

Methods and evaluation design

Procedure, tasks and evaluated configurations

Recorded evaluations

Each evaluation records what was tested and under which conditions.

Baseline coverage

Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.

0 of 2 active baseline roles have published Rewire measurements in this release. Measurements on a selected protocol do not establish coverage of an entire suite.

No execution recipe linked to this protocol. Recipe availability does not establish a completed evaluation.

Author-reported evaluations
3

Literature evidence is not a Rewire measurement. Executed but unpublished runs and private review status are not included.

Null control

Proposed control: requires review

No-change prediction under matched control conditions

Protocol-specific applicability, permitted inputs, access, split, evaluator and execution requirements need review before implementation or execution.

This is a suggested selection rule, not a validated method or a measured score.

Conventional reference

Proposed control: requires review

Training-only mean-effect or linear prediction

Protocol-specific applicability, permitted inputs, access, split, evaluator and execution requirements need review before implementation or execution.

This is a suggested selection rule, not a validated method or a measured score.

Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums

Coverage is derived from release 2026-10-07-1448159e6a81. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.

Run instructions

No runnable recipe has been reviewed for this protocol. Dataset access, model requirements, licences and compute requirements must be checked against its sources before execution.

Strengths, limitations and unresolved questions

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

0 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-10-07-1448159e6a81
Property and statementOriginal source and locationReview and provenance

No evidence rows match these filters. Choose another scope or clear the search.

Sources and history

View linked audit checks and correction history

Release 2026-10-07-1448159e6a81 · Record review: source checked

1 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: perteval-scfm-2025-protocol-norman-single-2000hvg-auspc

areas
cells-spatial-multiomics
tasks
Norman single-gene perturbation effect prediction, 2,000 HVGs, SPECTRA distribution shift
metric
AUSPC (Area Under the SPECTRA Performance Curve of MSE)
metric direction
lower
unit
printed column units: MSE (10^-2) per split; AUSPC (10^-2)
protocol
SPECTRA (Ektefaie et al., 2024) generates seven train-test splits of increasing sparsification probability s=0.1..0.7 (step 0.1), each a controlled-overlap distribution-shift condition for unseen perturbations (Fig. F1) -- this protocol concerns generalisation across perturbations under increasing distribution shift, not a claim about forecasting unseen cells. MSE is computed at each split on the delta=P-Xc perturbation-effect target (Eq. 5), not on raw post-perturbation expression (Table 1 columns S0.1-S0.7); AUSPC is np.trapz(phi, s) over those seven points (Appendix F.2, Eq. F2, Algorithm 1) -- not a ROC-AUC, accuracy or rank-based score. This delta/expression-change MSE is a different metric construction from GEARS Supplementary Table 6's own Pearson-delta-DE metric and must not be conflated with it.
source locator
Table 1 (Norman single-gene section); Appendix A.1 (dataset); Section 2.1 ('Raw expression data', Eq. 3, MLP baseline input); Section 2.2 ('MLP baseline' Eq. 4, 'GEARS baseline' Eq. 5, 'Mean baseline'); Appendix F.2 and Algorithm 1 (AUSPC/uncertainty propagation definitions, Eqs. F2-F5); Section 2.1.1 (HVG selection); Section 2.3.2 (SPECTRA splits); main-text Figure 2 caption ('standard error bars'); Appendix I, Figure I1 caption (triplicate experiments, '8 train-test splits').
configurations in this intake
GEARS; MLP baseline; Mean baseline
configurations in same table not intaken
Geneformer; scBERT; scFoundation; scGPT; UCE
missing metadata
per split scored sample counts: unreported; see dataset record; random seeds: unreported beyond the word 'triplicate'; exact seed values not printed; hyperparameters: GEARS: official implementation defaults, trained from scratch without pretrained weights (paper's own Methods wording, Section 2.2 region: 'we train GEARS from scratch without using pretrained weights'); MLP baseline / Mean baseline: architecture per Eqs. 4-6 and 'Mean baseline' text in Section 2.2, exact training hyperparameters not separately tabulated for these two baselines in the main text
caveats
This is a different dataset preprocessing (2,000 HVGs, not GEARS Supplementary Table 6's own unstated gene subset), a different split mechanism (SPECTRA sparsification, not Table 6's unstated split) and a different metric construction (AUSPC, a trapezoidal integral of MSE over seven distribution-shift splits, not Table 6's single-split MSE/Pearson-DeltaExpression) from the existing GEARS Supplementary Table 6 catalogue entries for this same use case (protocols gears-2023-supp-table6-task-mse / -task-pearson-de). Do not merge, average or otherwise combine these values with Table 6's.; GEARS in this protocol is trained from scratch by this paper's authors (an independent execution using the official GEARS implementation's default hyperparameters apart from the SPECTRA split), not the original GEARS authors' own reported Table 6 checkpoint/run. This is a second, independent GEARS evaluation, not a reproduction or replication claim for Table 6.; The same Table 1's double-gene section prints a GEARS AUSPC of 0.808 and a printed Mean-baseline-relative delta of 4.254, which is not exactly reproducible from the printed Mean-baseline AUSPC of 4.255 (4.255-0.808=3.447, not 4.254); this inconsistency is specific to the double-gene row, is not resolved or derived here, and is explicitly out of scope for this intake, which covers the single-gene row only.; Table 2 (Replogle RPE1) separately shows a running-text/table-row labeling conflict (text attributes an AUSPC of 0.1251 to the 'Mean baseline', while the table's own row data assigns 0.1251 to Geneformer and 0.1341 to the Mean baseline). This conflict concerns a different dataset (RPE1) not covered by this intake and is noted in the accompanying dossier, not resolved here.; Input populations differ across the three intaken configurations and are not an identical input budget: GEARS uses a graph-based architecture over raw expression and prior-knowledge gene relationships (official implementation); the MLP baseline's input ZGE = Xc (+) Gc concatenates raw control expression with a gene-co-expression matrix for the perturbed gene(s) (Eq. 3), not raw expression alone; the Mean baseline uses no learned input at all (see its own configuration record for the paper's exact, population-ambiguous definition). None of the three should be assumed to have received comparable information.; This protocol concerns generalisation to unseen perturbations under SPECTRA's increasing-sparsification distribution shift. It does not test, and does not establish, forecasting for unseen cells, donors, or cell-line/context transfer.
uncertainty definition
The source describes this quantity as a standard error (main-text Figure 2 caption: 'Average AUSPC (down-arrow) across sparsification probabilities for each model with standard error bars') and separately gives its own propagation formula (Appendix F.2, Eqs. F3-F5): AUSPC's uncertainty is derived from each split's own MSE uncertainty via the trapezoidal integral's partial derivatives (sigma^2 = sum_i (d/2)^2 * sigma_phi_i^2, where d=0.1 is the fixed sparsification step size). These two author statements describe the same quantity and are not in conflict: a propagated quantity can correctly be reported as a standard error. This is recorded as the author's own reported, propagated uncertainty; the F3-F5 derivation is the authors' own formula and its mathematical correctness has not been independently validated here. It must not be read as an independently resampled model-seed standard deviation or a confidence interval. The underlying per-split uncertainty is attributed by the source to triplicate experiments per model (Appendix I, Figure I1 caption: 'Experiments were carried out in triplicate for each model'), not to the main-text Figure 2 region. Figure I1's own caption separately states '8 train-test splits of increasing difficulty' for this same Norman single-gene evaluation, while Table 1 prints seven S-columns (S0.1-S0.7) and Appendix F.2 describes the sparsification probabilities as spanning 0.1 to 0.7. This 7-vs-8 discrepancy between the Figure I1 caption and the Table 1 / F.2 grid is preserved exactly as printed, not resolved; it must not be read as establishing an eighth Table 1 column or a confirmed n_runs=7, and no significance claim is made from any overlapping error bars.
Related records

Suggest a correction