rewire.itbenchmarks
Protocol

ProteinGym v1.3 zero-shot DMS substitutions

A version-pinned zero-shot protocol for scoring amino-acid substitutions against ProteinGym v1.3 deep mutational scanning measurements. It defines track-wide aggregation, but a selected assay or partial prediction file does not establish complete-track performance.

Sources (2)rewirebench: ProteinGym guide; rewirebench: ProteinGym scoring code · docs/proteingym.md lines 3–11, 31–33, 59–63; protocols/proteingym.py lines 232–249

No reviewed evaluations are linked here in this release. See the sources and separately identified configurations below.

0 evaluations · 0 results

Overview

Datasets

The pinned reference contains 217 assays and 2,465,767 assay–variant records, including multiple substitutions. These are reference counts, not a claim that this review obtained or verified every assay file.

SourcesProtocol metadata evidence: proteingym-reference_files-DMS_substitutions.csv · DMS_substitutions.csv: 217 data rows; sum of DMS_total_number_mutants; includes_multiple_mutants

Metrics

Per-assay Spearman rank correlation, ROC AUC, MCC, NDCG and top-10% recall. MCC thresholds predictions at the assay median; NDCG uses the pinned top-floor(n×0.1) implementation; top recall uses inclusive 90th-percentile thresholds, which can retain more than 10% with ties. Non-finite outputs become null, and the runner omits null values from subsequent means.

Sources (2)rewirebench: ProteinGym scoring code; Protocol metadata evidence: proteingym-proteingym-performance_DMS_benchmarks.py · protocols/proteingym.py lines 157–189; performance_DMS_benchmarks.py lines 14–78, 212–226

limited source coverage · Automated source review, 2026-09-23. All specifications and missing details

Results

All evaluations

0 evaluations · 0 results. Different protocols are not a single leaderboard.

Applied filters: All linked evaluations

No evaluations linked in this release.

Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.

Methods and evaluation design

Procedure, tasks and evaluated configurations

How it works

What the protocol evaluates

Models supply one scalar score per assay–mutant pair, with higher values predicting better experimental function. The evaluator retains the supplied DMS_score and DMS_score_bin labels. This track includes single and multiple amino-acid substitutions; indels, supervised learning and clinical tracks require different protocols.

Sources (2)rewirebench: ProteinGym guide; rewirebench: ProteinGym scoring code · docs/proteingym.md lines 3, 25–33; protocols/proteingym.py lines 41–55, 99–153
How to interpret a track summary

Metrics are first rounded within each assay, then averaged over assays belonging to the same UniProt protein and selection type, then over proteins within selection type, then equally over selection types. The runner exposes track metrics only when all 217 assays are fully scored. Category summaries from selected complete assays describe that selection only.

Sources (2)rewirebench: ProteinGym scoring code; Protocol metadata evidence: proteingym-proteingym-performance_DMS_benchmarks.py · protocols/proteingym.py lines 173–189, 216–244; performance_DMS_benchmarks.py lines 269–315
Source review and execution are separate

Pinned code and documentation specify the procedure. Their synthetic metric-comparison tests are source-reported implementation checks, not reproduction of a published ProteinGym model row. The separate AMFR report documents one later local evaluation and must not be generalized to this complete track.

Sources (2)rewirebench: ProteinGym guide; ESM-2 8M masked-marginal scoring: local execution report (20 September 2026) · docs/proteingym.md lines 65–70; report.json: /scope, /protocol_results/complete_assays, /protocol_results/total_assays, /independently_reproduced

Evaluation design

Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.

Benchmarks

These source-backed links do not make different protocols or scores interchangeable.

Baseline coverage

Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.

0 of 2 active baseline roles have published Rewire measurements in this release. Measurements on a selected protocol do not establish coverage of an entire suite.

2 execution recipes Recipe availability does not establish a completed evaluation.

No reviewed evaluations with results linked in this release.

Literature evidence is not a Rewire measurement. Executed but unpublished runs and private review status are not included.

Null control

Proposed control: requires review

Protocol-valid seeded random ranking; no fitting on assay labels

Protocol-specific applicability, permitted inputs, access, split, evaluator and execution requirements need review before implementation or execution.

This is a suggested selection rule, not a validated method or a measured score.

Conventional reference

Proposed control: requires review

Established label-free reference with explicitly permitted sequence/MSA inputs

Protocol-specific applicability, permitted inputs, access, split, evaluator and execution requirements need review before implementation or execution.

This is a suggested selection rule, not a validated method or a measured score.

Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums

Coverage is derived from release 2026-09-29-06401fd5b220. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.

Run this benchmark

ProteinGym v1.3: score mutation predictions

Score supplied zero-shot substitution predictions using the pinned ProteinGym evaluator definitions and report explicit per-assay coverage.

Recompute metrics from supplied predictions. This recipe does not establish reproduction of a particular published score.

Dataset access
Download and extract the official ProteinGym v1.3 DMS substitutions archive. The packaged reference metadata are pinned to upstream 144fe22b07dfaeec2b366f2346203a9838a55b4c; input assay bytes are hashed during preparation.
Model and weights
No model or weights are needed to score existing mutation predictions.
Licences
Runner code is MIT and retains upstream evaluator notices. Dataset and checkpoint reuse terms must be checked at their original sources.
Software
Python 3.11 with the pinned core environment; no GPU framework required for scoring.
Hardware
CPU scoring. Full-track storage, memory and runtime must be assessed locally; no universal resource estimate is established.
Required inputs and expected outputs

Inputs

  • Extracted official v1.3 DMS substitution assay CSV files.
  • Keyed prediction files; use the IDs from prepared input rather than constructing ambiguous mutation IDs.

Outputs

  • Local report.json with metrics, coverage, protocol and artifact hashes.
  • Local predictions.json and unscored.json; neither is submitted automatically.

Choose one way to run this recipe. These instruction formats are alternatives.

Execution steps

  1. 1. Prepare a selected assay and score predictions (Python)

    Source reviewed; these instructions have not been executed by rewire.

    import rewirebench
    prepared = rewirebench.prepare(
        "proteingym-v1.3-dms-substitutions", source="data/DMS_ProteinGym_substitutions",
        output="prepared", assay_ids=["AMFR_HUMAN_Tsuboyama_2023_4G3O"],
    )
    report = rewirebench.evaluate(
        prepared, "predictions/predictions.tsv", output="runs/proteingym-rescore",
        model={"name": "My protein model", "training_overlap": "Unreported"},
    )
    # Omit assay_ids only when preparing all assays for an explicitly requested full-track run.
    rewirebench: Library guide; rewirebench: ProteinGym guide; rewirebench: ProteinGym scoring code; rewirebench: ESM-2 adapter code · docs/proteingym.md: preparation, selected assays and evaluation

Use your own model

Run your model locally and return predictions keyed by the input IDs. The evaluator supplies biological inputs without test labels and owns scoring. This interface is not a sandbox for model code.

Pass your existing prediction function into this adapter. Its output direction must match the selected protocol.

class MyModelAdapter:
    def __init__(self, score):
        self.score = score

    def predict(self, inputs):
        return {row["id"]: float(self.score(row)) for row in inputs}

# adapter = MyModelAdapter(your_prediction_function)
# report = rewirebench.run(prepared, adapter, output="runs/my-model")

Alternatively, generate a keyed prediction file in your existing model environment and use the score-only recipe. Your model code and weights do not need to be shared.

rewirebench: Library guide; rewirebench: ProteinGym guide; rewirebench: ProteinGym scoring code; rewirebench: ESM-2 adapter code · Pinned runner documentation, protocol implementation and adapter implementation
Scope and limitations
  • Only zero-shot DMS substitutions are supported. Indels, supervised tracks and clinical tasks are excluded.
  • A selected assay or incomplete coverage is not a full-suite score.
  • Preserve mutation identities and upstream score direction and aggregation.
  • Input hashes record local bytes; they do not independently establish source authenticity.

Contribute a result for review. The library can submit an exported evaluation for private review when intake is open. Check the contribution page for access and sign-in.

Original repository instructions

Run this benchmark

Prepare official v1.3 substitution assays, generate or supply model scores, and evaluate the selected track. Full-suite scores require complete track coverage.

A maintained rewire runner has not been verified for this benchmark. Check data access, weights, licences, dependencies and hardware in the linked official documentation; requirements have not been fully extracted.

rewirebench: Library guide; rewirebench: ProteinGym guide; rewirebench: ProteinGym scoring code; rewirebench: ESM-2 adapter code · docs/proteingym.md and pinned protocols/proteingym.py
Strengths, limitations and unresolved questions

Strengths and limitations

Strengths and considerations

  • Versioned reference metadata, mutation validation and explicit partial-coverage handling make the intended comparison scope inspectable.
    Sourcesrewirebench: ProteinGym scoring code · protocols/proteingym.py lines 69–80, 104–125, 216–249

Limitations and conditions

  • A model can satisfy the scoring interface while having pretraining overlap. The example adapter leaves proteins longer than 1,022 residues unscored and cannot guarantee complete-track inference.
    Sourcesrewirebench: ProteinGym guide · docs/proteingym.md lines 33, 53
  • No uncertainty intervals, independent assay-archive authentication or complete-track model run are established by this source review.
    Sources (2)rewirebench: ProteinGym guide; rewirebench: ProteinGym scoring code · docs/proteingym.md lines 11, 63–70; protocols/proteingym.py lines 245–246
Profile review details

Bounded review of pinned protocol sources and an existing public report. No model execution, numerical-result re-review, independent reproduction or named human scientific review was performed.

Stable record: rewire-proteingym-v1-3-dms-substitutions

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-23. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
Version and referenceProteinGym v1.3 substitutions; upstream revision 144fe22b07dfaeec2b366f2346203a9838a55b4c; reference CSV SHA-256 a8f498011532a74aa9fe556a50555a75e928c5837d19c06a87592ae04049b308. The runner recipe is pinned separately to f80cef7f818bec33e51b7f43ad499eb5078c8d87; this does not assert that v1.3 is the latest release.
Sources (2)rewirebench: ProteinGym guide; rewirebench: ProteinGym scoring code · docs/proteingym.md lines 7–11; protocols/proteingym.py lines 21–24; source records: version and artifact_sha256
DatasetsThe pinned reference contains 217 assays and 2,465,767 assay–variant records, including multiple substitutions. These are reference counts, not a claim that this review obtained or verified every assay file.
SourcesProtocol metadata evidence: proteingym-reference_files-DMS_substitutions.csv · DMS_substitutions.csv: 217 data rows; sum of DMS_total_number_mutants; includes_multiple_mutants
Inputs and labelsInputs contain assay ID, wild-type sequence, mutated sequence and substitution notation. Predictions use globally unique <DMS_id>::<mutant> IDs. Higher scores mean better function; official DMS_score and DMS_score_bin labels remain with the evaluator.
Sources (2)rewirebench: ProteinGym guide; rewirebench: ProteinGym scoring code · docs/proteingym.md lines 3, 25–33; protocols/proteingym.py lines 126–153
SplitsThis zero-shot track has no assay-label fitting or cross-validation training split. An adapter exposing fit is rejected. This restriction does not establish that pretraining data exclude benchmark proteins.
Sourcesrewirebench: ProteinGym guide · docs/proteingym.md line 33
MetricsPer-assay Spearman rank correlation, ROC AUC, MCC, NDCG and top-10% recall. MCC thresholds predictions at the assay median; NDCG uses the pinned top-floor(n×0.1) implementation; top recall uses inclusive 90th-percentile thresholds, which can retain more than 10% with ties. Non-finite outputs become null, and the runner omits null values from subsequent means.
Sources (2)rewirebench: ProteinGym scoring code; Protocol metadata evidence: proteingym-proteingym-performance_DMS_benchmarks.py · protocols/proteingym.py lines 157–189; performance_DMS_benchmarks.py lines 14–78, 212–226
AggregationRound assay metrics to three decimals; average assays within each UniProt ID and selection type; average proteins within each selection type; average selection types equally and round the final runner summary to three decimals. This is not an assay-weighted or variant-weighted mean.
Sources (2)rewirebench: ProteinGym scoring code; Protocol metadata evidence: proteingym-proteingym-performance_DMS_benchmarks.py · protocols/proteingym.py lines 183–189, 221–239; performance_DMS_benchmarks.py lines 269–315
CoverageOnly a full-scope, fully scored 217-assay track exposes aggregate metrics. A selected assay is subset scope, and any limit is smoke scope. Incomplete assay metrics retain the original assay denominator and are excluded from category aggregation.
Sources (2)rewirebench: ProteinGym guide; rewirebench: ProteinGym scoring code · docs/proteingym.md line 31; protocols/proteingym.py lines 104–108, 141, 208–249
UncertaintyThe runner estimates no uncertainty interval. Upstream zero-shot bootstrap standard errors concern differences from the best aggregate model, resampling within selection categories for 10,000 repetitions by default; they are not absolute model confidence intervals or variation across inference seeds.
Sources (2)rewirebench: ProteinGym scoring code; Protocol metadata evidence: proteingym-proteingym-performance_DMS_benchmarks.py · protocols/proteingym.py line 246; performance_DMS_benchmarks.py lines 95–111, 310–315
Data verificationThe reference CSV is checked against its pinned SHA-256. Observed local assay hashes and optional expected_hashes detect byte changes but are explicitly not independently verified official archive checksums.
Sources (2)rewirebench: ProteinGym guide; rewirebench: ProteinGym scoring code · docs/proteingym.md line 11; protocols/proteingym.py lines 69–71, 95–98, 145–148
Checkpoint identityThe protocol accepts supplied predictions or a compatible model adapter and therefore defines no universal checkpoint. Its ESM-2 example is a separate configuration; it does not identify weights for historical ProteinGym result rows. · Not applicable
Sourcesrewirebench: ProteinGym guide · docs/proteingym.md lines 3, 25–33, 35–53
Reuse conditionsThe upstream software MIT notice does not establish permission for every assay dataset. Underlying study conditions and checkpoint terms must be checked separately; weights are not redistributed by the runner.
Sourcesrewirebench: ProteinGym guide · docs/proteingym.md lines 7–9
OrganismsNot extracted or verified for this record.
AssaysNot extracted or verified for this record.
AdaptationNot extracted or verified for this record.
BaselinesNot extracted or verified for this record.

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

93 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Version and reference
ProteinGym v1.3 substitutions; upstream revision 144fe22b07dfaeec2b366f2346203a9838a55b4c; reference CSV SHA-256 a8f498011532a74aa9fe556a50555a75e928c5837d19c06a87592ae04049b308. The runner recipe is pinned separately to f80cef7f818bec33e51b7f43ad499eb5078c8d87; this does not assert that v1.3 is the latest release.
Individual claims
rewirebench: ProteinGym guide

Original source ↗

docs/proteingym.md lines 7–11; protocols/proteingym.py lines 21–24; source records: version and artifact_sha256

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: f80cef7f818bec33e51b7f43ad499eb5078c8d87
Retrieved: 2026-09-17

source checked

automated source review · 2026-09-23

Audit details

Bounded review of pinned protocol sources and an existing public report. No model execution, numerical-result re-review, independent reproduction or named human scientific review was performed.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: d2b1fc009acc590b82ca125d886f3d52b07c2f36bbe0e9cb0d30776618166b5f

Hash scope: complete file bytes

Format: text

Version and reference
ProteinGym v1.3 substitutions; upstream revision 144fe22b07dfaeec2b366f2346203a9838a55b4c; reference CSV SHA-256 a8f498011532a74aa9fe556a50555a75e928c5837d19c06a87592ae04049b308. The runner recipe is pinned separately to f80cef7f818bec33e51b7f43ad499eb5078c8d87; this does not assert that v1.3 is the latest release.
Individual claims
rewirebench: ProteinGym scoring code

Original source ↗

docs/proteingym.md lines 7–11; protocols/proteingym.py lines 21–24; source records: version and artifact_sha256

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: f80cef7f818bec33e51b7f43ad499eb5078c8d87
Retrieved: 2026-09-17

source checked

automated source review · 2026-09-23

Audit details

Bounded review of pinned protocol sources and an existing public report. No model execution, numerical-result re-review, independent reproduction or named human scientific review was performed.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: 2d62959db5611b1c7c1c7e8e44056e96c90b3f52975b2080bc36da576f0ac871

Hash scope: complete file bytes

Format: text

Datasets
The pinned reference contains 217 assays and 2,465,767 assay–variant records, including multiple substitutions. These are reference counts, not a claim that this review obtained or verified every assay file.
Individual claims
Protocol metadata evidence: proteingym-reference_files-DMS_substitutions.csv

Original source ↗

DMS_substitutions.csv: 217 data rows; sum of DMS_total_number_mutants; includes_multiple_mutants

Version: 144fe22b07dfaeec2b366f2346203a9838a55b4c
Retrieved: 2026-09-23T18:39:35.604049+00:00

source checked

automated source review · 2026-09-23

Audit details

Bounded review of pinned protocol sources and an existing public report. No model execution, numerical-result re-review, independent reproduction or named human scientific review was performed.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: a8f498011532a74aa9fe556a50555a75e928c5837d19c06a87592ae04049b308

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Reuse conditions
The upstream software MIT notice does not establish permission for every assay dataset. Underlying study conditions and checkpoint terms must be checked separately; weights are not redistributed by the runner.
Individual claims
rewirebench: ProteinGym guide

Original source ↗

docs/proteingym.md lines 7–9

Version: f80cef7f818bec33e51b7f43ad499eb5078c8d87
Retrieved: 2026-09-17

source checked

automated source review · 2026-09-23

Audit details

Bounded review of pinned protocol sources and an existing public report. No model execution, numerical-result re-review, independent reproduction or named human scientific review was performed.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: d2b1fc009acc590b82ca125d886f3d52b07c2f36bbe0e9cb0d30776618166b5f

Hash scope: complete file bytes

Format: text

Inputs and labels
Inputs contain assay ID, wild-type sequence, mutated sequence and substitution notation. Predictions use globally unique <DMS_id>::<mutant> IDs. Higher scores mean better function; official DMS_score and DMS_score_bin labels remain with the evaluator.
Individual claims
rewirebench: ProteinGym guide

Original source ↗

docs/proteingym.md lines 3, 25–33; protocols/proteingym.py lines 126–153

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: f80cef7f818bec33e51b7f43ad499eb5078c8d87
Retrieved: 2026-09-17

source checked

automated source review · 2026-09-23

Audit details

Bounded review of pinned protocol sources and an existing public report. No model execution, numerical-result re-review, independent reproduction or named human scientific review was performed.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: d2b1fc009acc590b82ca125d886f3d52b07c2f36bbe0e9cb0d30776618166b5f

Hash scope: complete file bytes

Format: text

Inputs and labels
Inputs contain assay ID, wild-type sequence, mutated sequence and substitution notation. Predictions use globally unique <DMS_id>::<mutant> IDs. Higher scores mean better function; official DMS_score and DMS_score_bin labels remain with the evaluator.
Individual claims
rewirebench: ProteinGym scoring code

Original source ↗

docs/proteingym.md lines 3, 25–33; protocols/proteingym.py lines 126–153

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: f80cef7f818bec33e51b7f43ad499eb5078c8d87
Retrieved: 2026-09-17

source checked

automated source review · 2026-09-23

Audit details

Bounded review of pinned protocol sources and an existing public report. No model execution, numerical-result re-review, independent reproduction or named human scientific review was performed.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: 2d62959db5611b1c7c1c7e8e44056e96c90b3f52975b2080bc36da576f0ac871

Hash scope: complete file bytes

Format: text

Splits
This zero-shot track has no assay-label fitting or cross-validation training split. An adapter exposing fit is rejected. This restriction does not establish that pretraining data exclude benchmark proteins.
Individual claims
rewirebench: ProteinGym guide

Original source ↗

docs/proteingym.md line 33

Version: f80cef7f818bec33e51b7f43ad499eb5078c8d87
Retrieved: 2026-09-17

source checked

automated source review · 2026-09-23

Audit details

Bounded review of pinned protocol sources and an existing public report. No model execution, numerical-result re-review, independent reproduction or named human scientific review was performed.

Field: attributes.profile.facts.3.value

Source artifact SHA-256: d2b1fc009acc590b82ca125d886f3d52b07c2f36bbe0e9cb0d30776618166b5f

Hash scope: complete file bytes

Format: text

Metrics
Per-assay Spearman rank correlation, ROC AUC, MCC, NDCG and top-10% recall. MCC thresholds predictions at the assay median; NDCG uses the pinned top-floor(n×0.1) implementation; top recall uses inclusive 90th-percentile thresholds, which can retain more than 10% with ties. Non-finite outputs become null, and the runner omits null values from subsequent means.
Individual claims
Protocol metadata evidence: proteingym-proteingym-performance_DMS_benchmarks.py

Original source ↗

protocols/proteingym.py lines 157–189; performance_DMS_benchmarks.py lines 14–78, 212–226

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 144fe22b07dfaeec2b366f2346203a9838a55b4c
Retrieved: 2026-09-23T18:39:35.572257+00:00

source checked

automated source review · 2026-09-23

Audit details

Bounded review of pinned protocol sources and an existing public report. No model execution, numerical-result re-review, independent reproduction or named human scientific review was performed.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: 0002bc1b031c02dc2d67fde53da092c8fd301415645e8e83b425b4183570eba2

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Metrics
Per-assay Spearman rank correlation, ROC AUC, MCC, NDCG and top-10% recall. MCC thresholds predictions at the assay median; NDCG uses the pinned top-floor(n×0.1) implementation; top recall uses inclusive 90th-percentile thresholds, which can retain more than 10% with ties. Non-finite outputs become null, and the runner omits null values from subsequent means.
Individual claims
rewirebench: ProteinGym scoring code

Original source ↗

protocols/proteingym.py lines 157–189; performance_DMS_benchmarks.py lines 14–78, 212–226

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: f80cef7f818bec33e51b7f43ad499eb5078c8d87
Retrieved: 2026-09-17

source checked

automated source review · 2026-09-23

Audit details

Bounded review of pinned protocol sources and an existing public report. No model execution, numerical-result re-review, independent reproduction or named human scientific review was performed.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: 2d62959db5611b1c7c1c7e8e44056e96c90b3f52975b2080bc36da576f0ac871

Hash scope: complete file bytes

Format: text

Aggregation
Round assay metrics to three decimals; average assays within each UniProt ID and selection type; average proteins within each selection type; average selection types equally and round the final runner summary to three decimals. This is not an assay-weighted or variant-weighted mean.
Individual claims
Protocol metadata evidence: proteingym-proteingym-performance_DMS_benchmarks.py

Original source ↗

protocols/proteingym.py lines 183–189, 221–239; performance_DMS_benchmarks.py lines 269–315

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 144fe22b07dfaeec2b366f2346203a9838a55b4c
Retrieved: 2026-09-23T18:39:35.572257+00:00

source checked

automated source review · 2026-09-23

Audit details

Bounded review of pinned protocol sources and an existing public report. No model execution, numerical-result re-review, independent reproduction or named human scientific review was performed.

Field: attributes.profile.facts.5.value

Source artifact SHA-256: 0002bc1b031c02dc2d67fde53da092c8fd301415645e8e83b425b4183570eba2

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: source checked

8 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: rewire-proteingym-v1-3-dms-substitutions

domain
proteomics
entity level
protocol
protocol version
v1.3
upstream revision
144fe22b07dfaeec2b366f2346203a9838a55b4c
run documentation
record id: rewire-proteingym-v1-3-dms-substitutions; status: source_reviewed_not_executed; summary: Prepare official v1.3 substitution assays, generate or supply model scores, and evaluate the selected track. Full-suite scores require complete track coverage.; source ids: runner-recipe-docs-sdk-md-f80cef7f; runner-recipe-docs-proteingym-md-f80cef7f; runner-recipe-packages-rewirebench-src-rewirebench-protocols-proteingym-py-f80cef7f; runner-recipe-packages-rewirebench-src-rewirebench-adapters-esm-py-f80cef7f; source locator: docs/proteingym.md and pinned protocols/proteingym.py
run recipes
id: proteingym-v1-3-rescore; protocol id: proteingym-v1.3-dms-substitutions; version: f80cef7f818bec33e51b7f43ad499eb5078c8d87; title: ProteinGym v1.3: score mutation predictions; purpose: rescore_predictions; summary: Score supplied zero-shot substitution predictions using the pinned ProteinGym evaluator definitions and report explicit per-assay coverage.; inputs: Extracted official v1.3 DMS substitution assay CSV files.; Keyed prediction files; use the IDs from prepared input rather than constructing ambiguous mutation IDs.; outputs: Local report.json with metrics, coverage, protocol and artifact hashes.; Local predictions.json and unscored.json; neither is submitted automatically.; requirements: data: Download and extract the official ProteinGym v1.3 DMS substitutions archive. The packaged reference metadata are pinned to upstream 144fe22b07dfaeec2b366f2346203a9838a55b4c; input assay bytes are hashed during preparation.; weights: No model or weights are needed to score existing mutation predictions.; licence: Runner code is MIT and retains upstream evaluator notices. Dataset and checkpoint reuse terms must be checked at their original sources.; software: Python 3.11 with the pinned core environment; no GPU framework required for scoring.; hardware: CPU scoring. Full-track storage, memory and runtime must be assessed locally; no universal resource estimate is established.; instructions: runtime: python; title: Prepare a selected assay and score predictions; code: import rewirebench prepared = rewirebench.prepare( "proteingym-v1.3-dms-substitutions", source="data/DMS_ProteinGym_substitutions", output="prepared", assay_ids=["AMFR_HUMAN_Tsuboyama_2023_4G3O"], ) report = rewirebench.evaluate( prepared, "predictions/predictions.tsv", output="runs/proteingym-rescore", model={"name": "My protein model", "training_overlap": "Unreported"}, ) # Omit assay_ids only when preparing all assays for an explicitly requested full-track run.; status: source_reviewed_not_executed; source ids: runner-recipe-docs-sdk-md-f80cef7f; runner-recipe-docs-proteingym-md-f80cef7f; runner-recipe-packages-rewirebench-src-rewirebench-protocols-proteingym-py-f80cef7f; runner-recipe-packages-rewirebench-src-rewirebench-adapters-esm-py-f80cef7f; source locator: docs/proteingym.md: preparation, selected assays and evaluation; runtime: command_line; title: Prepare a selected assay and score predictions; code: git clone https://github.com/rewire-bio/rewire-benchmarks.git cd rewire-benchmarks git checkout --detach f80cef7f818bec33e51b7f43ad499eb5078c8d87 uv sync --locked --package rewirebench --python 3.11 uv run --package rewirebench rewirebench prepare proteingym-v1.3-dms-substitutions \ --source data/DMS_ProteinGym_substitutions --output prepared \ --options '{"assay_ids":["AMFR_HUMAN_Tsuboyama_2023_4G3O"]}' uv run --package rewirebench rewirebench evaluate --prepared prepared \ --predictions predictions/predictions.tsv --output runs/proteingym-rescore \ --model-name "My protein model"; status: source_reviewed_not_executed; source ids: runner-recipe-docs-sdk-md-f80cef7f; runner-recipe-docs-proteingym-md-f80cef7f; runner-recipe-packages-rewirebench-src-rewirebench-protocols-proteingym-py-f80cef7f; runner-recipe-packages-rewirebench-src-rewirebench-adapters-esm-py-f80cef7f; source locator: docs/proteingym.md: CLI preparation and scoring; runtime: podman; title: Score in a local Podman image; code: # Build the pinned core image as described in the linked HPC guide first. mkdir -p runs podman run --rm --userns=keep-id --network=none \ -v "$PWD/prepared:/prepared:ro" \ -v "$PWD/predictions:/predictions:ro" \ -v "$PWD/runs:/outputs:rw" \ localhost/rewirebench:core evaluate \ --prepared /prepared --predictions /predictions/predictions.tsv \ --output /outputs/proteingym-v1-3-rescore; status: source_reviewed_not_executed; source ids: runner-recipe-docs-hpc-md-f80cef7f; runner-recipe-docs-sdk-md-f80cef7f; source locator: docs/hpc.md: Podman build, read-only input mounts and offline execution; docs/sdk.md: evaluate; runtime: apptainer; title: Score in an Apptainer image; code: # Build or obtain the pinned SIF as described in the linked HPC guide first. mkdir -p runs apptainer run --cleanenv --containall \ --bind "$PWD/prepared:/prepared:ro" \ --bind "$PWD/predictions:/predictions:ro" \ --bind "$PWD/runs:/outputs:rw" \ rewirebench-core.sif evaluate \ --prepared /prepared --predictions /predictions/predictions.tsv \ --output /outputs/proteingym-v1-3-rescore; status: source_reviewed_not_executed; source ids: runner-recipe-docs-hpc-md-f80cef7f; runner-recipe-docs-sdk-md-f80cef7f; source locator: docs/hpc.md: Apptainer execution and source image identity; docs/sdk.md: evaluate; runtime: slurm; title: Submit the prepared scoring job to Slurm; code: #!/bin/bash set -euo pipefail # Set these for your cluster; no performance or resource estimate is implied. : "${REWIRE_ACCOUNT:?Set your Slurm account}" : "${REWIRE_PARTITION:?Set your Slurm partition}" : "${REWIRE_CPUS:?Set the requested CPU count}" : "${REWIRE_MEMORY:?Set the requested memory}" : "${REWIRE_WALLTIME:?Set the requested time limit}" : "${REWIRE_JOB_ROOT:?Set a shared absolute directory with prepared data and SIF}" export REWIRE_JOB_ROOT sbatch --account="$REWIRE_ACCOUNT" --partition="$REWIRE_PARTITION" \ --cpus-per-task="$REWIRE_CPUS" --mem="$REWIRE_MEMORY" \ --time="$REWIRE_WALLTIME" --export=ALL <<'REWIRE_JOB' #!/bin/bash set -euo pipefail cd "$REWIRE_JOB_ROOT" mkdir -p runs apptainer run --cleanenv --containall \ --bind "$PWD/prepared:/prepared:ro" \ --bind "$PWD/predictions:/predictions:ro" \ --bind "$PWD/runs:/outputs:rw" \ rewirebench-core.sif evaluate \ --prepared /prepared --predictions /predictions/predictions.tsv \ --output /outputs/proteingym-v1-3-rescore REWIRE_JOB; status: source_reviewed_not_executed; source ids: runner-recipe-docs-hpc-md-f80cef7f; runner-recipe-docs-sdk-md-f80cef7f; source locator: docs/hpc.md: Slurm template, preparation outside compute nodes and site-specific resources; limitations: Only zero-shot DMS substitutions are supported. Indels, supervised tracks and clinical tasks are excluded.; A selected assay or incomplete coverage is not a full-suite score.; Preserve mutation identities and upstream score direction and aggregation.; Input hashes record local bytes; they do not independently establish source authenticity.; source ids: runner-recipe-docs-sdk-md-f80cef7f; runner-recipe-docs-proteingym-md-f80cef7f; runner-recipe-packages-rewirebench-src-rewirebench-protocols-proteingym-py-f80cef7f; runner-recipe-packages-rewirebench-src-rewirebench-adapters-esm-py-f80cef7f; source locator: Pinned runner documentation, protocol implementation and adapter implementation; id: proteingym-v1-3-esm2; protocol id: proteingym-v1.3-dms-substitutions; version: f80cef7f818bec33e51b7f43ad499eb5078c8d87; title: ProteinGym v1.3: ESM-2 8M example; purpose: generate_and_evaluate; summary: Generate zero-shot substitution scores with the public ESM-2 8M checkpoint and score a limited assay smoke input.; inputs: Extracted ProteinGym v1.3 assay files.; The local official esm2_t6_8M_UR50D.pt checkpoint; its SHA-256 is verified before loading.; outputs: Local report.json with metrics, coverage, protocol and artifact hashes.; Local predictions.json and unscored.json; neither is submitted automatically.; requirements: data: Download and extract the official ProteinGym v1.3 DMS substitutions archive. The packaged reference metadata are pinned to upstream 144fe22b07dfaeec2b366f2346203a9838a55b4c; input assay bytes are hashed during preparation.; weights: ESM-2 esm2_t6_8M_UR50D checkpoint from the official fair-esm distribution; download before the offline job.; licence: Runner code is MIT and retains upstream evaluator notices. Dataset and checkpoint reuse terms must be checked at their original sources.; software: Python 3.11 with rewirebench[esm], fair-esm 2.0.0 and PyTorch.; hardware: CPU example. Sequences longer than 1,022 residues are explicitly unscored and remain in coverage; this example cannot produce complete predictions for tracks containing those inputs.; instructions: runtime: python; title: Run a limited ESM-2 example; code: import rewirebench from rewirebench.adapters.esm import ESM2Adapter prepared = rewirebench.prepare( "proteingym-v1.3-dms-substitutions", source="data/DMS_ProteinGym_substitutions", output="prepared-smoke", assay_ids=["AMFR_HUMAN_Tsuboyama_2023_4G3O"], limit=10, ) report = rewirebench.run( prepared, ESM2Adapter(checkpoint="weights/esm2_t6_8M_UR50D.pt"), output="runs/esm2-smoke", model={"name": "ESM-2 8M masked marginals", "training_overlap": "Unreported"}, ); status: source_reviewed_not_executed; source ids: runner-recipe-docs-sdk-md-f80cef7f; runner-recipe-docs-proteingym-md-f80cef7f; runner-recipe-packages-rewirebench-src-rewirebench-protocols-proteingym-py-f80cef7f; runner-recipe-packages-rewirebench-src-rewirebench-adapters-esm-py-f80cef7f; source locator: docs/proteingym.md: public ESM-2 example; adapters/esm.py: ESM2Adapter; runtime: command_line; title: Run a limited ESM-2 example; code: git clone https://github.com/rewire-bio/rewire-benchmarks.git cd rewire-benchmarks git checkout --detach f80cef7f818bec33e51b7f43ad499eb5078c8d87 uv sync --locked --package rewirebench --extra esm --python 3.11 uv run --package rewirebench --extra esm rewirebench prepare proteingym-v1.3-dms-substitutions --source data/DMS_ProteinGym_substitutions --output prepared-smoke --options '{"assay_ids":["AMFR_HUMAN_Tsuboyama_2023_4G3O"],"limit":10}' uv run --package rewirebench --extra esm rewirebench run --prepared prepared-smoke --adapter rewirebench.adapters.esm:ESM2Adapter --adapter-options '{"checkpoint":"weights/esm2_t6_8M_UR50D.pt"}' --output runs/esm2-smoke --model-name "ESM-2 8M masked marginals"; status: source_reviewed_not_executed; source ids: runner-recipe-docs-sdk-md-f80cef7f; runner-recipe-docs-proteingym-md-f80cef7f; runner-recipe-packages-rewirebench-src-rewirebench-protocols-proteingym-py-f80cef7f; runner-recipe-packages-rewirebench-src-rewirebench-adapters-esm-py-f80cef7f; source locator: docs/sdk.md: run; docs/proteingym.md: public ESM example; limitations: Only zero-shot DMS substitutions are supported. Indels, supervised tracks and clinical tasks are excluded.; A selected assay or incomplete coverage is not a full-suite score.; Preserve mutation identities and upstream score direction and aggregation.; Input hashes record local bytes; they do not independently establish source authenticity.; Masked-marginal log odds are calculated in wild-type context and summed over substitutions.; This is an example adapter, not reproduction of a published ESM ProteinGym result.; source ids: runner-recipe-docs-sdk-md-f80cef7f; runner-recipe-docs-proteingym-md-f80cef7f; runner-recipe-packages-rewirebench-src-rewirebench-protocols-proteingym-py-f80cef7f; runner-recipe-packages-rewirebench-src-rewirebench-adapters-esm-py-f80cef7f; source locator: Pinned runner documentation, protocol implementation and adapter implementation
Related records

Suggest a correction