Datasets
The pinned reference contains 217 assays and 2,465,767 assay–variant records, including multiple substitutions. These are reference counts, not a claim that this review obtained or verified every assay file.
A version-pinned zero-shot protocol for scoring amino-acid substitutions against ProteinGym v1.3 deep mutational scanning measurements. It defines track-wide aggregation, but a selected assay or partial prediction file does not establish complete-track performance.
No reviewed evaluations are linked here in this release. See the sources and separately identified configurations below.
The pinned reference contains 217 assays and 2,465,767 assay–variant records, including multiple substitutions. These are reference counts, not a claim that this review obtained or verified every assay file.
Per-assay Spearman rank correlation, ROC AUC, MCC, NDCG and top-10% recall. MCC thresholds predictions at the assay median; NDCG uses the pinned top-floor(n×0.1) implementation; top recall uses inclusive 90th-percentile thresholds, which can retain more than 10% with ties. Non-finite outputs become null, and the runner omits null values from subsequent means.
limited source coverage · Automated source review, 2026-09-23. All specifications and missing details
0 evaluations · 0 results. Different protocols are not a single leaderboard.
Applied filters: All linked evaluations
No evaluations linked in this release.
Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.
Models supply one scalar score per assay–mutant pair, with higher values predicting better experimental function. The evaluator retains the supplied DMS_score and DMS_score_bin labels. This track includes single and multiple amino-acid substitutions; indels, supervised learning and clinical tracks require different protocols.
Metrics are first rounded within each assay, then averaged over assays belonging to the same UniProt protein and selection type, then over proteins within selection type, then equally over selection types. The runner exposes track metrics only when all 217 assays are fully scored. Category summaries from selected complete assays describe that selection only.
Pinned code and documentation specify the procedure. Their synthetic metric-comparison tests are source-reported implementation checks, not reproduction of a published ProteinGym model row. The separate AMFR report documents one later local evaluation and must not be generalized to this complete track.
Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.
These source-backed links do not make different protocols or scores interchangeable.
Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.
0 of 2 active baseline roles have published Rewire measurements in this release. Measurements on a selected protocol do not establish coverage of an entire suite.
2 execution recipes Recipe availability does not establish a completed evaluation.
No reviewed evaluations with results linked in this release.
Literature evidence is not a Rewire measurement. Executed but unpublished runs and private review status are not included.
Proposed control: requires review
Protocol-valid seeded random ranking; no fitting on assay labels
Protocol-specific applicability, permitted inputs, access, split, evaluator and execution requirements need review before implementation or execution.
This is a suggested selection rule, not a validated method or a measured score.
Proposed control: requires review
Established label-free reference with explicitly permitted sequence/MSA inputs
Protocol-specific applicability, permitted inputs, access, split, evaluator and execution requirements need review before implementation or execution.
This is a suggested selection rule, not a validated method or a measured score.
Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums
Coverage is derived from release 2026-09-29-06401fd5b220. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.
Score supplied zero-shot substitution predictions using the pinned ProteinGym evaluator definitions and report explicit per-assay coverage.
Recompute metrics from supplied predictions. This recipe does not establish reproduction of a particular published score.
Choose one way to run this recipe. These instruction formats are alternatives.
Source reviewed; these instructions have not been executed by rewire.
import rewirebench
prepared = rewirebench.prepare(
"proteingym-v1.3-dms-substitutions", source="data/DMS_ProteinGym_substitutions",
output="prepared", assay_ids=["AMFR_HUMAN_Tsuboyama_2023_4G3O"],
)
report = rewirebench.evaluate(
prepared, "predictions/predictions.tsv", output="runs/proteingym-rescore",
model={"name": "My protein model", "training_overlap": "Unreported"},
)
# Omit assay_ids only when preparing all assays for an explicitly requested full-track run.rewirebench: Library guide; rewirebench: ProteinGym guide; rewirebench: ProteinGym scoring code; rewirebench: ESM-2 adapter code · docs/proteingym.md: preparation, selected assays and evaluationRun your model locally and return predictions keyed by the input IDs. The evaluator supplies biological inputs without test labels and owns scoring. This interface is not a sandbox for model code.
Pass your existing prediction function into this adapter. Its output direction must match the selected protocol.
class MyModelAdapter:
def __init__(self, score):
self.score = score
def predict(self, inputs):
return {row["id"]: float(self.score(row)) for row in inputs}
# adapter = MyModelAdapter(your_prediction_function)
# report = rewirebench.run(prepared, adapter, output="runs/my-model")Alternatively, generate a keyed prediction file in your existing model environment and use the score-only recipe. Your model code and weights do not need to be shared.
rewirebench: Library guide; rewirebench: ProteinGym guide; rewirebench: ProteinGym scoring code; rewirebench: ESM-2 adapter code · Pinned runner documentation, protocol implementation and adapter implementationContribute a result for review. The library can submit an exported evaluation for private review when intake is open. Check the contribution page for access and sign-in.
Prepare official v1.3 substitution assays, generate or supply model scores, and evaluate the selected track. Full-suite scores require complete track coverage.
A maintained rewire runner has not been verified for this benchmark. Check data access, weights, licences, dependencies and hardware in the linked official documentation; requirements have not been fully extracted.
rewirebench: Library guide; rewirebench: ProteinGym guide; rewirebench: ProteinGym scoring code; rewirebench: ESM-2 adapter code · docs/proteingym.md and pinned protocols/proteingym.pyBounded review of pinned protocol sources and an existing public report. No model execution, numerical-result re-review, independent reproduction or named human scientific review was performed.
Stable record: rewire-proteingym-v1-3-dms-substitutionsExplanatory profile: limited source coverage · Automated source review, 2026-09-23. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Version and reference | ProteinGym v1.3 substitutions; upstream revision 144fe22b07dfaeec2b366f2346203a9838a55b4c; reference CSV SHA-256 a8f498011532a74aa9fe556a50555a75e928c5837d19c06a87592ae04049b308. The runner recipe is pinned separately to f80cef7f818bec33e51b7f43ad499eb5078c8d87; this does not assert that v1.3 is the latest release.Sources (2)rewirebench: ProteinGym guide; rewirebench: ProteinGym scoring code · docs/proteingym.md lines 7–11; protocols/proteingym.py lines 21–24; source records: version and artifact_sha256 |
| Datasets | The pinned reference contains 217 assays and 2,465,767 assay–variant records, including multiple substitutions. These are reference counts, not a claim that this review obtained or verified every assay file.SourcesProtocol metadata evidence: proteingym-reference_files-DMS_substitutions.csv · DMS_substitutions.csv: 217 data rows; sum of DMS_total_number_mutants; includes_multiple_mutants |
| Inputs and labels | Inputs contain assay ID, wild-type sequence, mutated sequence and substitution notation. Predictions use globally unique <DMS_id>::<mutant> IDs. Higher scores mean better function; official DMS_score and DMS_score_bin labels remain with the evaluator.Sources (2)rewirebench: ProteinGym guide; rewirebench: ProteinGym scoring code · docs/proteingym.md lines 3, 25–33; protocols/proteingym.py lines 126–153 |
| Splits | This zero-shot track has no assay-label fitting or cross-validation training split. An adapter exposing fit is rejected. This restriction does not establish that pretraining data exclude benchmark proteins.Sourcesrewirebench: ProteinGym guide · docs/proteingym.md line 33 |
| Metrics | Per-assay Spearman rank correlation, ROC AUC, MCC, NDCG and top-10% recall. MCC thresholds predictions at the assay median; NDCG uses the pinned top-floor(n×0.1) implementation; top recall uses inclusive 90th-percentile thresholds, which can retain more than 10% with ties. Non-finite outputs become null, and the runner omits null values from subsequent means.Sources (2)rewirebench: ProteinGym scoring code; Protocol metadata evidence: proteingym-proteingym-performance_DMS_benchmarks.py · protocols/proteingym.py lines 157–189; performance_DMS_benchmarks.py lines 14–78, 212–226 |
| Aggregation | Round assay metrics to three decimals; average assays within each UniProt ID and selection type; average proteins within each selection type; average selection types equally and round the final runner summary to three decimals. This is not an assay-weighted or variant-weighted mean.Sources (2)rewirebench: ProteinGym scoring code; Protocol metadata evidence: proteingym-proteingym-performance_DMS_benchmarks.py · protocols/proteingym.py lines 183–189, 221–239; performance_DMS_benchmarks.py lines 269–315 |
| Coverage | Only a full-scope, fully scored 217-assay track exposes aggregate metrics. A selected assay is subset scope, and any limit is smoke scope. Incomplete assay metrics retain the original assay denominator and are excluded from category aggregation.Sources (2)rewirebench: ProteinGym guide; rewirebench: ProteinGym scoring code · docs/proteingym.md line 31; protocols/proteingym.py lines 104–108, 141, 208–249 |
| Uncertainty | The runner estimates no uncertainty interval. Upstream zero-shot bootstrap standard errors concern differences from the best aggregate model, resampling within selection categories for 10,000 repetitions by default; they are not absolute model confidence intervals or variation across inference seeds.Sources (2)rewirebench: ProteinGym scoring code; Protocol metadata evidence: proteingym-proteingym-performance_DMS_benchmarks.py · protocols/proteingym.py line 246; performance_DMS_benchmarks.py lines 95–111, 310–315 |
| Data verification | The reference CSV is checked against its pinned SHA-256. Observed local assay hashes and optional expected_hashes detect byte changes but are explicitly not independently verified official archive checksums.Sources (2)rewirebench: ProteinGym guide; rewirebench: ProteinGym scoring code · docs/proteingym.md line 11; protocols/proteingym.py lines 69–71, 95–98, 145–148 |
| Checkpoint identity | The protocol accepts supplied predictions or a compatible model adapter and therefore defines no universal checkpoint. Its ESM-2 example is a separate configuration; it does not identify weights for historical ProteinGym result rows. · Not applicableSourcesrewirebench: ProteinGym guide · docs/proteingym.md lines 3, 25–33, 35–53 |
| Reuse conditions | The upstream software MIT notice does not establish permission for every assay dataset. Underlying study conditions and checkpoint terms must be checked separately; weights are not redistributed by the runner.Sourcesrewirebench: ProteinGym guide · docs/proteingym.md lines 7–9 |
| Organisms | Not extracted or verified for this record. |
| Assays | Not extracted or verified for this record. |
| Adaptation | Not extracted or verified for this record. |
| Baselines | Not extracted or verified for this record. |
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
93 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Version and reference ProteinGym v1.3 substitutions; upstream revision 144fe22b07dfaeec2b366f2346203a9838a55b4c; reference CSV SHA-256 a8f498011532a74aa9fe556a50555a75e928c5837d19c06a87592ae04049b308. The runner recipe is pinned separately to f80cef7f818bec33e51b7f43ad499eb5078c8d87; this does not assert that v1.3 is the latest release. Individual claims | rewirebench: ProteinGym guide docs/proteingym.md lines 7–11; protocols/proteingym.py lines 21–24; source records: version and artifact_sha256 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: f80cef7f818bec33e51b7f43ad499eb5078c8d87 | source checked automated source review · 2026-09-23 Audit detailsBounded review of pinned protocol sources and an existing public report. No model execution, numerical-result re-review, independent reproduction or named human scientific review was performed. Field: Source artifact SHA-256: Hash scope: complete file bytes Format: text |
| Version and reference ProteinGym v1.3 substitutions; upstream revision 144fe22b07dfaeec2b366f2346203a9838a55b4c; reference CSV SHA-256 a8f498011532a74aa9fe556a50555a75e928c5837d19c06a87592ae04049b308. The runner recipe is pinned separately to f80cef7f818bec33e51b7f43ad499eb5078c8d87; this does not assert that v1.3 is the latest release. Individual claims | rewirebench: ProteinGym scoring code docs/proteingym.md lines 7–11; protocols/proteingym.py lines 21–24; source records: version and artifact_sha256 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: f80cef7f818bec33e51b7f43ad499eb5078c8d87 | source checked automated source review · 2026-09-23 Audit detailsBounded review of pinned protocol sources and an existing public report. No model execution, numerical-result re-review, independent reproduction or named human scientific review was performed. Field: Source artifact SHA-256: Hash scope: complete file bytes Format: text |
| Datasets The pinned reference contains 217 assays and 2,465,767 assay–variant records, including multiple substitutions. These are reference counts, not a claim that this review obtained or verified every assay file. Individual claims | Protocol metadata evidence: proteingym-reference_files-DMS_substitutions.csv DMS_substitutions.csv: 217 data rows; sum of DMS_total_number_mutants; includes_multiple_mutants Version: 144fe22b07dfaeec2b366f2346203a9838a55b4c | source checked automated source review · 2026-09-23 Audit detailsBounded review of pinned protocol sources and an existing public report. No model execution, numerical-result re-review, independent reproduction or named human scientific review was performed. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Reuse conditions The upstream software MIT notice does not establish permission for every assay dataset. Underlying study conditions and checkpoint terms must be checked separately; weights are not redistributed by the runner. Individual claims | rewirebench: ProteinGym guide docs/proteingym.md lines 7–9 Version: f80cef7f818bec33e51b7f43ad499eb5078c8d87 | source checked automated source review · 2026-09-23 Audit detailsBounded review of pinned protocol sources and an existing public report. No model execution, numerical-result re-review, independent reproduction or named human scientific review was performed. Field: Source artifact SHA-256: Hash scope: complete file bytes Format: text |
| Inputs and labels Inputs contain assay ID, wild-type sequence, mutated sequence and substitution notation. Predictions use globally unique <DMS_id>::<mutant> IDs. Higher scores mean better function; official DMS_score and DMS_score_bin labels remain with the evaluator. Individual claims | rewirebench: ProteinGym guide docs/proteingym.md lines 3, 25–33; protocols/proteingym.py lines 126–153 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: f80cef7f818bec33e51b7f43ad499eb5078c8d87 | source checked automated source review · 2026-09-23 Audit detailsBounded review of pinned protocol sources and an existing public report. No model execution, numerical-result re-review, independent reproduction or named human scientific review was performed. Field: Source artifact SHA-256: Hash scope: complete file bytes Format: text |
| Inputs and labels Inputs contain assay ID, wild-type sequence, mutated sequence and substitution notation. Predictions use globally unique <DMS_id>::<mutant> IDs. Higher scores mean better function; official DMS_score and DMS_score_bin labels remain with the evaluator. Individual claims | rewirebench: ProteinGym scoring code docs/proteingym.md lines 3, 25–33; protocols/proteingym.py lines 126–153 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: f80cef7f818bec33e51b7f43ad499eb5078c8d87 | source checked automated source review · 2026-09-23 Audit detailsBounded review of pinned protocol sources and an existing public report. No model execution, numerical-result re-review, independent reproduction or named human scientific review was performed. Field: Source artifact SHA-256: Hash scope: complete file bytes Format: text |
| Splits This zero-shot track has no assay-label fitting or cross-validation training split. An adapter exposing fit is rejected. This restriction does not establish that pretraining data exclude benchmark proteins. Individual claims | rewirebench: ProteinGym guide docs/proteingym.md line 33 Version: f80cef7f818bec33e51b7f43ad499eb5078c8d87 | source checked automated source review · 2026-09-23 Audit detailsBounded review of pinned protocol sources and an existing public report. No model execution, numerical-result re-review, independent reproduction or named human scientific review was performed. Field: Source artifact SHA-256: Hash scope: complete file bytes Format: text |
| Metrics Per-assay Spearman rank correlation, ROC AUC, MCC, NDCG and top-10% recall. MCC thresholds predictions at the assay median; NDCG uses the pinned top-floor(n×0.1) implementation; top recall uses inclusive 90th-percentile thresholds, which can retain more than 10% with ties. Non-finite outputs become null, and the runner omits null values from subsequent means. Individual claims | Protocol metadata evidence: proteingym-proteingym-performance_DMS_benchmarks.py protocols/proteingym.py lines 157–189; performance_DMS_benchmarks.py lines 14–78, 212–226 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 144fe22b07dfaeec2b366f2346203a9838a55b4c | source checked automated source review · 2026-09-23 Audit detailsBounded review of pinned protocol sources and an existing public report. No model execution, numerical-result re-review, independent reproduction or named human scientific review was performed. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Metrics Per-assay Spearman rank correlation, ROC AUC, MCC, NDCG and top-10% recall. MCC thresholds predictions at the assay median; NDCG uses the pinned top-floor(n×0.1) implementation; top recall uses inclusive 90th-percentile thresholds, which can retain more than 10% with ties. Non-finite outputs become null, and the runner omits null values from subsequent means. Individual claims | rewirebench: ProteinGym scoring code protocols/proteingym.py lines 157–189; performance_DMS_benchmarks.py lines 14–78, 212–226 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: f80cef7f818bec33e51b7f43ad499eb5078c8d87 | source checked automated source review · 2026-09-23 Audit detailsBounded review of pinned protocol sources and an existing public report. No model execution, numerical-result re-review, independent reproduction or named human scientific review was performed. Field: Source artifact SHA-256: Hash scope: complete file bytes Format: text |
| Aggregation Round assay metrics to three decimals; average assays within each UniProt ID and selection type; average proteins within each selection type; average selection types equally and round the final runner summary to three decimals. This is not an assay-weighted or variant-weighted mean. Individual claims | Protocol metadata evidence: proteingym-proteingym-performance_DMS_benchmarks.py protocols/proteingym.py lines 183–189, 221–239; performance_DMS_benchmarks.py lines 269–315 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 144fe22b07dfaeec2b366f2346203a9838a55b4c | source checked automated source review · 2026-09-23 Audit detailsBounded review of pinned protocol sources and an existing public report. No model execution, numerical-result re-review, independent reproduction or named human scientific review was performed. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
View linked audit checks and correction history
Release 2026-09-29-06401fd5b220 · Record review: source checked
Stable ID: rewire-proteingym-v1-3-dms-substitutions