Datasets
Processed AnnData datasets with perturbation/covariate metadata and configurable feature selections.
PerturBench evaluates predicted single-cell perturbation responses with explicit aggregation and metric choices.
Processed AnnData datasets with perturbation/covariate metadata and configurable feature selections.
Expression/change aggregation precedes metrics such as cosine, Pearson, RMSE, MSE, MAE and R-squared; optional rank metrics form another view.
Predicted/observed expression and perturbation/covariate metadata in AnnData.
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Source reviewed · Automated source review, 2026-09-16. All specifications and missing details
Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.
cosine_logfc (fraction) · Higher values are better.
PerturBench CB-COSINE: combination prediction on Norman19, Cosine similarity of log fold change · Norman19 (PerturBench split)
Evidence origin: Author-reported evaluation.
PerturBench: Benchmarking Machine Learning Models for Cellular Perturbation Analysis · Table 3, column(Cosine, log fold change (LogFC))Every method PerturBench reports on combination prediction on Norman19, Cosine similarity of log fold change, scored with Cosine similarity of log fold change on Norman19.
Automated source review: 2026-09-18. Numerical source review does not establish independent reproduction.
Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.
Showing 9 of 9 matching rows.
PerturBench predicts single-cell responses in held-out perturbation–context combinations. It implements cross-covariate, combinatorial and inverse-combinatorial partitions, comparing learned models with simple controls. Rank-based metrics complement expression-error metrics to reveal models that fail to distinguish perturbations.
Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.
These source-backed links do not make different protocols or scores interchangeable.
Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.
No concrete protocols are explicitly linked to this suite. Protocol identification and baseline selection are outstanding.
Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums
Coverage is derived from release 2026-09-29-06401fd5b220. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.
Install the environment, load a benchmark dataset with its split, and train a model through the repository's configuration system.
Generate predictions and evaluate them. This recipe does not establish reproduction of a particular published score.
Source reviewed; these instructions have not been executed by rewire.
conda create -n [env-name] python=3.11
conda activate [env-name]
cd [/path/to/PerturBench/]
pip3 install -e .
# or
pip3 install -e .[cli]PerturBench: repository README · README.md at c84038bc, Install PerturBench, lines 17-22Source reviewed; these instructions have not been executed by rewire.
from perturbench.data.accessors.srivatsan20 import Sciplex3
srivatsan20_accessor = Sciplex3()
adata = srivatsan20_accessor.get_anndata() ## Get the preprocessed anndata object
torch_dataset = srivatsan20_accessor.get_dataset() ## Get a PyTorch DatasetPerturBench: repository README · README.md at c84038bc, Dataset Access, lines 35-39Source reviewed; these instructions have not been executed by rewire.
from perturbench.data.accessors.jiang24 import Jiang24
jiang24_accessor = Jiang24()
split = jiang24_accessor.get_split()PerturBench: repository README · README.md at c84038bc, Data Splitting, lines 49-52Source reviewed; these instructions have not been executed by rewire.
python <path-to-repo-folder>/src/perturbench/modelcore/train.py <config-options>PerturBench: repository README · README.md at c84038bc, Hydra Training Script, lines 66-66Run your model locally and return predictions keyed by the input IDs. The evaluator supplies biological inputs without test labels and owns scoring. This interface is not a sandbox for model code.
Pass your existing prediction function into this adapter. Its output direction must match the selected protocol.
class MyModelAdapter:
def __init__(self, score):
self.score = score
def predict(self, inputs):
return {row["id"]: float(self.score(row)) for row in inputs}
# adapter = MyModelAdapter(your_prediction_function)
# report = rewirebench.run(prepared, adapter, output="runs/my-model")Alternatively, generate a keyed prediction file in your existing model environment and use the score-only recipe. Your model code and weights do not need to be shared.
PerturBench: repository README · README.md at c84038bcContribute a result for review. The library can submit an exported evaluation for private review when intake is open. Check the contribution page for access and sign-in.
Official installation, processed-data download, dataset accessors and Hydra/evaluator instructions are available. Local cache paths and task-specific split files must be configured; several datasets have manual split-generation notebooks. A default invocation must not be presented as reproducing every study split.
A maintained rewire runner has not been verified for this benchmark. Check data access, weights, licences, dependencies and hardware in the linked official documentation; requirements have not been fully extracted.
altoslabs/perturbench / README.md · README.md lines 15–63 (Install, datasets, splits and usage)Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.
Stable record: discovery-benchmark-perturbenchExplanatory profile: source reviewed · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Datasets | Processed AnnData datasets with perturbation/covariate metadata and configurable feature selections.Sourcesaltoslabs/perturbench official source · Pinned README: Data; Model evaluation; custom dataset configuration |
| Splits | Cross-cell-type and combination-prediction splits are supported, along with explicit custom split files.Sourcesaltoslabs/perturbench official source · Pinned README: Data; Model evaluation; custom dataset configuration |
| Metrics | Expression/change aggregation precedes metrics such as cosine, Pearson, RMSE, MSE, MAE and R-squared; optional rank metrics form another view.Sourcesaltoslabs/perturbench official source · Pinned README: Data; Model evaluation; custom dataset configuration |
| Baselines | Reproduction configurations include linear reference models.Sourcesaltoslabs/perturbench official source · Pinned README: Data; Model evaluation; custom dataset configuration |
| Leakage controls | Split and covariate definitions are configuration inputs; exact evaluation files must be pinned.Sourcesaltoslabs/perturbench official source · Pinned README: Data; Model evaluation; custom dataset configuration |
| Uncertainty | For the best hyperparameter configuration the authors run four additional training seeds, yielding five runs. Error bars represent standard deviation of model performance across those runs.Sourcesperturbench primary benchmark evidence · Experimental setup; Appendix datasets and data splitting |
| Entity type | Perturbation-response evaluation framework.Sourcesaltoslabs/perturbench official source · Pinned README: Data; Model evaluation; custom dataset configuration |
| Organisms | The evaluated datasets use human cell-line perturbation systems, including McFaline-Figueroa’s glioblastoma cell contexts and Srivatsan’s chemical perturbation cell lines. The framework itself is not restricted to those organisms.Sourcesperturbench primary benchmark evidence · Experimental setup; Appendix datasets and data splitting |
| Assays | Single-cell perturbation response measurements.Sourcesaltoslabs/perturbench official source · Pinned README: Data; Model evaluation; custom dataset configuration |
| Allowed inputs | Predicted/observed expression and perturbation/covariate metadata in AnnData.Sourcesaltoslabs/perturbench official source · Pinned README: Data; Model evaluation; custom dataset configuration |
| Adaptation | Supports supervised response prediction with explicit cell-type and combination holdouts.Sourcesaltoslabs/perturbench official source · Pinned README: Data; Model evaluation; custom dataset configuration |
Applicability is distinct from a completed evaluation.
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Last literature check: 2026-09-17. Dated primary-source discovery and protocol/table screening. Source checking does not mean experimental reproduction. Only separately extracted and independently reviewed numeric batches are publishable.
| Paper or primary resource | Version | Reference |
|---|---|---|
| PerturBench: Benchmarking Machine Learning Models for Cellular Perturbation Analysis | Primary full-text snapshot retrieved 2026-09-17; exact bytes pinned by SHA-256 | Read source |
The catalogue now holds 76 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.
primary protocol reviewed
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
27 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions. Individual claims | altoslabs/perturbench official source Pinned README: Data; Model evaluation; custom dataset configuration Version: c84038bc1ea409aa54f3832cfa6f34f5059adf0c | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| altoslabs/perturbench official source Pinned README: Data; Model evaluation; custom dataset configuration Version: c84038bc1ea409aa54f3832cfa6f34f5059adf0c | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Evaluation procedure Individual claims | altoslabs/perturbench official source Pinned README: Data; Model evaluation; custom dataset configuration Version: c84038bc1ea409aa54f3832cfa6f34f5059adf0c | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Datasets Processed AnnData datasets with perturbation/covariate metadata and configurable feature selections. Individual claims | altoslabs/perturbench official source Pinned README: Data; Model evaluation; custom dataset configuration Version: c84038bc1ea409aa54f3832cfa6f34f5059adf0c | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Splits Cross-cell-type and combination-prediction splits are supported, along with explicit custom split files. Individual claims | altoslabs/perturbench official source Pinned README: Data; Model evaluation; custom dataset configuration Version: c84038bc1ea409aa54f3832cfa6f34f5059adf0c | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Adaptation Supports supervised response prediction with explicit cell-type and combination holdouts. Individual claims | altoslabs/perturbench official source Pinned README: Data; Model evaluation; custom dataset configuration Version: c84038bc1ea409aa54f3832cfa6f34f5059adf0c | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Metrics Expression/change aggregation precedes metrics such as cosine, Pearson, RMSE, MSE, MAE and R-squared; optional rank metrics form another view. Individual claims | altoslabs/perturbench official source Pinned README: Data; Model evaluation; custom dataset configuration Version: c84038bc1ea409aa54f3832cfa6f34f5059adf0c | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Baselines Reproduction configurations include linear reference models. Individual claims | altoslabs/perturbench official source Pinned README: Data; Model evaluation; custom dataset configuration Version: c84038bc1ea409aa54f3832cfa6f34f5059adf0c | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Leakage controls Split and covariate definitions are configuration inputs; exact evaluation files must be pinned. Individual claims | altoslabs/perturbench official source Pinned README: Data; Model evaluation; custom dataset configuration Version: c84038bc1ea409aa54f3832cfa6f34f5059adf0c | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Uncertainty For the best hyperparameter configuration the authors run four additional training seeds, yielding five runs. Error bars represent standard deviation of model performance across those runs. Individual claims | perturbench primary benchmark evidence Experimental setup; Appendix datasets and data splitting Version: 2408.10609v1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
View linked audit checks and correction history
Release 2026-09-29-06401fd5b220 · Record review: discovered
Stable ID: discovery-benchmark-perturbench