rewirebio.iobenchmarks
Result

4.484 ± 0.299 AUSPC

MLP baseline · PertEval-scFM Norman single-gene (2,000 HVGs) AUSPC · Area Under the SPECTRA Performance Curve

Tested configuration
MLP baseline (raw control expression + gene co-expression input, Eq. 3)
Protocol
PertEval-scFM Norman single-gene (2,000 HVGs) AUSPC across SPECTRA sparsification splits
Dataset subset
Norman et al. 2019 single-gene perturbations, K562, top 2,000 HVGs (PertEval-scFM Table 1)
Procedure
Trapezoidal-rule AUSPC of MSE, scored on the perturbation-effect delta=P-Xc (Eq. 5), across seven SPECTRA sparsification splits (s=0.1..0.7), Norman single-gene, 2,000 HVGs.
Evaluation
MLP baseline on PertEval-scFM Norman single-gene (2,000 HVGs) AUSPC
Coverage
scored: unreported; eligible: unreported
Uncertainty
type: author_reported_propagated_standard_error; reported spread: 0.299; note: The source describes this quantity as a standard error (main-text Figure 2 caption: 'Average AUSPC (down-arrow) across sparsification probabilities for each model with standard error bars') and separately gives its own propagation formula (Appendix F.2, Eqs. F3-F5): AUSPC's uncertainty is derived from each split's own MSE uncertainty via the trapezoidal integral's partial derivatives (sigma^2 = sum_i (d/2)^2 * sigma_phi_i^2, where d=0.1 is the fixed sparsification step size). These two author statements describe the same quantity and are not in conflict: a propagated quantity can correctly be reported as a standard error. This is recorded as the author's own reported, propagated uncertainty; the F3-F5 derivation is the authors' own formula and its mathematical correctness has not been independently validated here. It must not be read as an independently resampled model-seed standard deviation or a confidence interval. The underlying per-split uncertainty is attributed by the source to triplicate experiments per model (Appendix I, Figure I1 caption: 'Experiments were carried out in triplicate for each model'), not to the main-text Figure 2 region. Figure I1's own caption separately states '8 train-test splits of increasing difficulty' for this same Norman single-gene evaluation, while Table 1 prints seven S-columns (S0.1-S0.7) and Appendix F.2 describes the sparsification probabilities as spanning 0.1 to 0.7. This 7-vs-8 discrepancy between the Figure I1 caption and the Table 1 / F.2 grid is preserved exactly as printed, not resolved; it must not be read as establishing an eighth Table 1 column or a confirmed n_runs=7, and no significance claim is made from any overlapping error bars.
Evidence
Author-reported evaluation · source checkedPertEval-scFM (Wenteler et al., ICML 2025), full text · Table 1, Norman single-gene section, row MLP baseline, column AUSPC (10^-2).

A source-checked result verifies the numerical transcription, not every model or protocol detail. Evaluation metadata: source checked. Source checked does not mean independently reproduced.

Reproduction

Split
Not reported
Adaptation
Not reported
Scoring implementation
AUSPC (trapezoidal-rule integral of MSE across seven SPECTRA sparsification-probability train-test splits, s=0.1..0.7; Appendix F.2 Eq. F2)

No execution recipe has been verified for this exact configuration and evaluation. A benchmark's general instructions may use different inputs, splits or model settings.

Reproducing this published result requires matching its model configuration, data, split and scorer. Source checking or a successful smoke test does not establish score reproduction.

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

1 evidence row matching the loaded filters

Claims, original sources and review scope · Release 2026-10-07-1448159e6a81
Property and statementOriginal source and locationReview and provenance
Reported result
4.484 ± 0.299
Individual claims
PertEval-scFM (Wenteler et al., ICML 2025), full text

Original source ↗

Table 1, Norman single-gene section, row MLP baseline, column AUSPC (10^-2).

Version: PMLR v267 wenteler25a (as served by the PMLR-affiliated mlresearch/v267 GitHub mirror; ETag "ed9f0fe44cf6edc939ee6950d024f65dcd8dd6f16414bd9f5297cda9395d6e58" at retrieval)
Retrieved: 2026-10-07T11:39:51Z

source checked

Exact printed value transcribed from the primary PDF (pdftotext -layout extraction), independently checked against the rendered table text twice in this session. Not model execution. · 2026-10-07

author reported

Audit details

Source checked, not reproduced. Rank 2nd of eight Table 1 configurations by this metric in the source's own printed rank column; only three of eight rows are intaken.

Field: attributes.printed_value

Source artifact SHA-256: c116a153872645d5c91b9ee836df3945cd8ffd02a6e736f58f854c4b70978bcc

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Extraction artifact SHA-256: c116a153872645d5c91b9ee836df3945cd8ffd02a6e736f58f854c4b70978bcc

Extraction artifact

Sources and history

View linked audit checks and correction history

Release 2026-10-07-1448159e6a81 · Record review: source checked

1 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: perteval-scfm-2025-result-mlp-baseline-auspc

areas
cells-spatial-multiomics
tasks
Norman single-gene perturbation effect prediction, 2,000 HVGs, SPECTRA distribution shift
metric
AUSPC
metric direction
lower
unit
10^-2 (printed column header units)
printed value
4.484 ± 0.299
numeric value
4.484
derived normalized value
0.04484
derived normalized value note
4.484 x 10^-2, computed here for cross-scale comparability only; not a separate printed source value.
uncertainty
type: author_reported_propagated_standard_error; reported spread: 0.299; note: The source describes this quantity as a standard error (main-text Figure 2 caption: 'Average AUSPC (down-arrow) across sparsification probabilities for each model with standard error bars') and separately gives its own propagation formula (Appendix F.2, Eqs. F3-F5): AUSPC's uncertainty is derived from each split's own MSE uncertainty via the trapezoidal integral's partial derivatives (sigma^2 = sum_i (d/2)^2 * sigma_phi_i^2, where d=0.1 is the fixed sparsification step size). These two author statements describe the same quantity and are not in conflict: a propagated quantity can correctly be reported as a standard error. This is recorded as the author's own reported, propagated uncertainty; the F3-F5 derivation is the authors' own formula and its mathematical correctness has not been independently validated here. It must not be read as an independently resampled model-seed standard deviation or a confidence interval. The underlying per-split uncertainty is attributed by the source to triplicate experiments per model (Appendix I, Figure I1 caption: 'Experiments were carried out in triplicate for each model'), not to the main-text Figure 2 region. Figure I1's own caption separately states '8 train-test splits of increasing difficulty' for this same Norman single-gene evaluation, while Table 1 prints seven S-columns (S0.1-S0.7) and Appendix F.2 describes the sparsification probabilities as spanning 0.1 to 0.7. This 7-vs-8 discrepancy between the Figure I1 caption and the Table 1 / F.2 grid is preserved exactly as printed, not resolved; it must not be read as establishing an eighth Table 1 column or a confirmed n_runs=7, and no significance claim is made from any overlapping error bars.
source locator
Table 1, Norman single-gene section, row MLP baseline, column AUSPC (10^-2).
review
method: Exact printed value transcribed from the primary PDF (pdftotext -layout extraction), independently checked against the rendered table text twice in this session. Not model execution.; reviewer: Claude (local extraction, this session), with corrections independently identified by a separate Codex review pass before this intake; date: 2026-10-07; artifact sha256: c116a153872645d5c91b9ee836df3945cd8ffd02a6e736f58f854c4b70978bcc; retrieval url: https://raw.githubusercontent.com/mlresearch/v267/main/assets/wenteler25a/wenteler25a.pdf; notes: Source checked, not reproduced. Rank 2nd of eight Table 1 configurations by this metric in the source's own printed rank column; only three of eight rows are intaken.
evidence overlap
Independent of the GEARS Supplementary Table 6 catalogue entries for this use case.
Related records

Suggest a correction