Datasets
Paired spatial-transcriptomic measurements and histology images from HEST resources.
HEST-Benchmark tests prediction of gene expression from histological image representations.
Paired spatial-transcriptomic measurements and histology images from HEST resources.
Pearson correlation between predicted and measured log1p gene expression, using the 50 genes with highest normalized variance.
Histology patches for prediction; spatial expression supplies evaluation labels.
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Source reviewed · Automated source review, 2026-09-16. All specifications and missing details
Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.
pearson_r (correlation) · Higher values are better.
HEST-Benchmark CCRCC: Gene expression prediction from histology, Clear cell renal cell carcinoma · HEST-Benchmark CCRCC (HEST-Benchmark split)
Evidence origin: Author-reported evaluation.
HEST-1k: A Dataset for Spatial Transcriptomics and Histology Image Analysis · Table 1, row(CCRCC)Every method HEST-Benchmark reports on Gene expression prediction from histology, Clear cell renal cell carcinoma, scored with Pearson correlation on HEST-Benchmark CCRCC.
Automated source review: 2026-09-18. Numerical source review does not establish independent reproduction.
Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.
Showing 10 of 10 matching rows.
HEST-Benchmark connects histology patches to spatially measured gene expression. A patch encoder supplies features to a regression model, which predicts highly variable genes. Patient-stratified folds evaluate transfer between individuals, and correlation is summarized across those folds.
Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.
These source-backed links do not make different protocols or scores interchangeable.
Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.
No concrete protocols are explicitly linked to this suite. Protocol identification and baseline selection are outstanding.
Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums
Coverage is derived from release 2026-09-29-06401fd5b220. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.
Install the library with its benchmark extras, then evaluate a patch encoder across the ten cohorts on this page.
Generate predictions and evaluate them. This recipe does not establish reproduction of a particular published score.
Source reviewed; these instructions have not been executed by rewire.
git clone https://github.com/mahmoodlab/HEST.git
cd HEST
conda create -n "hest" python=3.11
conda activate hest
pip install -e .HEST: repository README · README.md at 3ddb5eaf, HEST-Library installation, lines 46-50Source reviewed; these instructions have not been executed by rewire.
pip install -e ".[benchmark]"HEST: repository README · README.md at 3ddb5eaf, Additional dependencies (HEST-Benchmark), lines 56-56Source reviewed; these instructions have not been executed by rewire.
from hest import iter_hest
for st in iter_hest('../hest_data', id_list=['TENX95']):
print(st)HEST: repository README · README.md at 3ddb5eaf, Inspect HEST-1k with HEST-Library, lines 80-83Run your model locally and return predictions keyed by the input IDs. The evaluator supplies biological inputs without test labels and owns scoring. This interface is not a sandbox for model code.
Pass your existing prediction function into this adapter. Its output direction must match the selected protocol.
class MyModelAdapter:
def __init__(self, score):
self.score = score
def predict(self, inputs):
return {row["id"]: float(self.score(row)) for row in inputs}
# adapter = MyModelAdapter(your_prediction_function)
# report = rewirebench.run(prepared, adapter, output="runs/my-model")Alternatively, generate a keyed prediction file in your existing model environment and use the score-only recipe. Your model code and weights do not need to be shared.
HEST: repository README · README.md at 3ddb5eafContribute a result for review. The library can submit an exported evaluation for private review when intake is open. Check the contribution page for access and sign-in.
Official installation, benchmark extras and a dedicated HEST-Benchmark notebook are linked. The complete HEST-1k collection is described as exceeding 1 TB; select the benchmark subset and patch encoder deliberately. This pass does not promote dataset download or WSI preparation into a model evaluation command.
A maintained rewire runner has not been verified for this benchmark. Check data access, weights, licences, dependencies and hardware in the linked official documentation; requirements have not been fully extracted.
mahmoodlab/HEST / README.md · README.md lines 36–73 and 137–139 (Data, installation and benchmarking your own model)Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.
Stable record: discovery-benchmark-hest-benchmarkExplanatory profile: source reviewed · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Datasets | Paired spatial-transcriptomic measurements and histology images from HEST resources.Sourcesmahmoodlab/HEST official source · Pinned README: HEST-Benchmark overview; evaluation notes |
| Splits | Patient-stratified cross-validation: one fold per patient, except ccRCC uses half as many folds because of its larger patient cohort.Sourceshest primary benchmark evidence · Sections 5.1–5.2; Appendix Table A11 |
| Metrics | Pearson correlation between predicted and measured log1p gene expression, using the 50 genes with highest normalized variance.Sourceshest primary benchmark evidence · Sections 5.1–5.2; Appendix Table A11 |
| Baselines | Ridge regression on PCA-reduced embeddings is the reported comparison setup.Sourcesmahmoodlab/HEST official source · Pinned README: HEST-Benchmark overview; evaluation notes |
| Leakage controls | All samples from a patient stay within their fold, preventing patch-level mixing of the same patient between training and test. Histology-encoder pretraining overlap is a separate concern.Sourceshest primary benchmark evidence · Sections 5.1–5.2; Appendix Table A11 |
| Uncertainty | The paper reports the mean and standard deviation across folds or patients, not a universal retraining-seed interval.Sourceshest primary benchmark evidence · Sections 5.1–5.2; Appendix Table A11 |
| Entity type | Spatial-transcriptomics prediction benchmark.Sourcesmahmoodlab/HEST official source · Pinned README: HEST-Benchmark overview; evaluation notes |
| Organisms | Multiple species selectable in the HEST metadata.Sourcesmahmoodlab/HEST official source · Pinned README: HEST-Benchmark overview; evaluation notes |
| Assays | Paired tissue histology and spatial gene-expression measurements.Sourcesmahmoodlab/HEST official source · Pinned README: HEST-Benchmark overview; evaluation notes |
| Allowed inputs | Histology patches for prediction; spatial expression supplies evaluation labels.Sourcesmahmoodlab/HEST official source · Pinned README: HEST-Benchmark overview; evaluation notes |
| Adaptation | Patch embeddings are evaluated through downstream expression prediction.Sourcesmahmoodlab/HEST official source · Pinned README: HEST-Benchmark overview; evaluation notes |
Applicability is distinct from a completed evaluation.
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Last literature check: 2026-09-17. Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.
| Paper or primary resource | Version | Reference |
|---|---|---|
| HEST-1k: A Dataset for Spatial Transcriptomics and Histology Image Analysis | Version pinned by URL and artifact SHA256 when available | Read source |
The catalogue now holds 100 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.
source found structured extraction pending
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
29 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions. Individual claims | hest primary benchmark evidence Pinned README: HEST-Benchmark overview; evaluation notes; Sections 5.1–5.2; Appendix Table A11 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 2406.16192v1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram caption Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions. Individual claims | mahmoodlab/HEST official source Pinned README: HEST-Benchmark overview; evaluation notes; Sections 5.1–5.2; Appendix Table A11 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 3ddb5eaf5bd2a8133e0c0e8015816489a3d99dc3 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| hest primary benchmark evidence Pinned README: HEST-Benchmark overview; evaluation notes; Sections 5.1–5.2; Appendix Table A11 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 2406.16192v1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| mahmoodlab/HEST official source Pinned README: HEST-Benchmark overview; evaluation notes; Sections 5.1–5.2; Appendix Table A11 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 3ddb5eaf5bd2a8133e0c0e8015816489a3d99dc3 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Evaluation procedure Individual claims | hest primary benchmark evidence Pinned README: HEST-Benchmark overview; evaluation notes; Sections 5.1–5.2; Appendix Table A11 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 2406.16192v1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Evaluation procedure Individual claims | mahmoodlab/HEST official source Pinned README: HEST-Benchmark overview; evaluation notes; Sections 5.1–5.2; Appendix Table A11 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 3ddb5eaf5bd2a8133e0c0e8015816489a3d99dc3 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Datasets Paired spatial-transcriptomic measurements and histology images from HEST resources. Individual claims | mahmoodlab/HEST official source Pinned README: HEST-Benchmark overview; evaluation notes Version: 3ddb5eaf5bd2a8133e0c0e8015816489a3d99dc3 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Splits Patient-stratified cross-validation: one fold per patient, except ccRCC uses half as many folds because of its larger patient cohort. Individual claims | hest primary benchmark evidence Sections 5.1–5.2; Appendix Table A11 Version: 2406.16192v1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Adaptation Patch embeddings are evaluated through downstream expression prediction. Individual claims | mahmoodlab/HEST official source Pinned README: HEST-Benchmark overview; evaluation notes Version: 3ddb5eaf5bd2a8133e0c0e8015816489a3d99dc3 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Metrics Pearson correlation between predicted and measured log1p gene expression, using the 50 genes with highest normalized variance. Individual claims | hest primary benchmark evidence Sections 5.1–5.2; Appendix Table A11 Version: 2406.16192v1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
View linked audit checks and correction history
Release 2026-09-29-06401fd5b220 · Record review: discovered
Stable ID: discovery-benchmark-hest-benchmark