Datasets
Predicted and observed perturbation responses; the associated study evaluates protocols across public datasets.
scPertEval evaluates and calibrates scoring protocols for single-cell perturbation predictions.
No reviewed evaluations are linked here in this release. See the sources and separately identified configurations below.
Predicted and observed perturbation responses; the associated study evaluates protocols across public datasets.
Protocol scoring is separated from calibration against empirical positive/negative controls using DRF and BDS.
Predicted/observed responses plus a protocol specifying representation, metric, transformation and reporting.
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Source reviewed · Automated source review, 2026-09-16. All specifications and missing details
0 evaluations · 0 results. Different protocols are not a single leaderboard.
Applied filters: All linked evaluations
No evaluations linked in this release.
Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.
scPertEval supplies explicit scoring protocols for perturbation predictions. A metric can compare one perturbation or operate across the complete perturbation panel, depending on its declared scope. Ground-truth references and context are passed to the evaluator, so protocol and preprocessing choices remain part of the result.
No runnable recipe has been reviewed for this evaluator. Dataset access, model requirements, licences and compute requirements must be checked against its sources before execution.
Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.
Stable record: discovery-benchmark-scpertevalExplanatory profile: source reviewed · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Datasets | Predicted and observed perturbation responses; the associated study evaluates protocols across public datasets.SourcesVirtual-Cell-Research-Community/scPertEval official source · Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy |
| Splits | The evaluator scores supplied outputs; predictor train/test partitions belong to the dataset/run being evaluated. · Not applicableSourcesVirtual-Cell-Research-Community/scPertEval official source · Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy |
| Metrics | Protocol scoring is separated from calibration against empirical positive/negative controls using DRF and BDS.SourcesVirtual-Cell-Research-Community/scPertEval official source · Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy |
| Baselines | Empirical controls calibrate how well a protocol distinguishes expected response quality.SourcesVirtual-Cell-Research-Community/scPertEval official source · Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy |
| Leakage controls | The metric implementation receives ground truth, predictions and an evaluation context; it does not construct model-training partitions or audit training data. Predictor leakage controls belong to the protocol and dataset used to produce the submitted predictions. · Not applicableSourcesscperteval0 primary benchmark evidence · Pinned src/scperteval/protocols/metrics.py: metric input contract and context |
| Uncertainty | Uncertainty across samples, datasets or training runs must be defined by the evaluation study; this evaluator entry does not fix one experiment. · Not applicableSourcesVirtual-Cell-Research-Community/scPertEval official source · Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy |
| Entity type | Perturbation scoring-protocol calibration toolkit.SourcesVirtual-Cell-Research-Community/scPertEval official source · Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy |
| Organisms | Organism scope belongs to the selected perturbation dataset. · Not applicableSourcesVirtual-Cell-Research-Community/scPertEval official source · Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy |
| Assays | Observed single-cell perturbation responses and empirical positive/negative controls.SourcesVirtual-Cell-Research-Community/scPertEval official source · Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy |
| Allowed inputs | Predicted/observed responses plus a protocol specifying representation, metric, transformation and reporting.SourcesVirtual-Cell-Research-Community/scPertEval official source · Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy |
| Adaptation | The toolkit scores and calibrates evaluation protocols; it does not impose predictor fine-tuning. · Not applicableSourcesVirtual-Cell-Research-Community/scPertEval official source · Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy |
| Implementation | Not extracted or verified for this record. |
Applicability is distinct from a completed evaluation.
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Last literature check: 2026-09-17. Primary-paper discovery and source inspection. Source-checked results are not independently reproduced experiments.
| Paper or primary resource | Version | Reference |
|---|---|---|
| Towards Principled Evaluation of Single-Cell Perturbation Prediction Models | bioRxiv 2026.07.23.740433v1 | Read source DOI: 10.64898/2026.07.23.740433 |
| scPertEval documentation | Documentation snapshot at retrieval; byte-pinned by SHA-256 | Read source |
primary protocol screened
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
18 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions. Individual claims | Virtual-Cell-Research-Community/scPertEval official source Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy Version: 4685f11927e887745737600170da7a655b727553 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| Virtual-Cell-Research-Community/scPertEval official source Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy Version: 4685f11927e887745737600170da7a655b727553 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Evaluation procedure Individual claims | Virtual-Cell-Research-Community/scPertEval official source Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy Version: 4685f11927e887745737600170da7a655b727553 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Datasets Predicted and observed perturbation responses; the associated study evaluates protocols across public datasets. Individual claims | Virtual-Cell-Research-Community/scPertEval official source Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy Version: 4685f11927e887745737600170da7a655b727553 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Splits The evaluator scores supplied outputs; predictor train/test partitions belong to the dataset/run being evaluated. Individual claims | Virtual-Cell-Research-Community/scPertEval official source Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy Version: 4685f11927e887745737600170da7a655b727553 | inapplicable automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Adaptation The toolkit scores and calibrates evaluation protocols; it does not impose predictor fine-tuning. Individual claims | Virtual-Cell-Research-Community/scPertEval official source Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy Version: 4685f11927e887745737600170da7a655b727553 | inapplicable automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Metrics Protocol scoring is separated from calibration against empirical positive/negative controls using DRF and BDS. Individual claims | Virtual-Cell-Research-Community/scPertEval official source Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy Version: 4685f11927e887745737600170da7a655b727553 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Baselines Empirical controls calibrate how well a protocol distinguishes expected response quality. Individual claims | Virtual-Cell-Research-Community/scPertEval official source Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy Version: 4685f11927e887745737600170da7a655b727553 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Leakage controls The metric implementation receives ground truth, predictions and an evaluation context; it does not construct model-training partitions or audit training data. Predictor leakage controls belong to the protocol and dataset used to produce the submitted predictions. Individual claims | scperteval0 primary benchmark evidence Pinned src/scperteval/protocols/metrics.py: metric input contract and context Version: 4685f11927e887745737600170da7a655b727553:src/scperteval/protocols/metrics.py | inapplicable automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Uncertainty Uncertainty across samples, datasets or training runs must be defined by the evaluation study; this evaluator entry does not fix one experiment. Individual claims | Virtual-Cell-Research-Community/scPertEval official source Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy Version: 4685f11927e887745737600170da7a655b727553 | inapplicable automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
View linked audit checks and correction history
Release 2026-09-29-06401fd5b220 · Record review: discovered
Stable ID: discovery-benchmark-scperteval