rewire.itbenchmarks
Benchmark

PFMBench

PFMBench is a configurable suite of protein-model downstream evaluations.

Sourcesbiomap-research/PFMBench official source · Pinned README: Overview; Features; repository architecture

141 evaluations · 141 results

Overview

Datasets

Tasks span structure, function, localization, interactions and other protein properties.

Sourcesbiomap-research/PFMBench official source · Pinned README: Overview; Features; repository architecture

Metrics

Task-specific metrics include AUROC for several binary function and interaction tasks, accuracy for categorical tasks, and Spearman correlation for continuous fitness, affinity and enzyme properties; the task appendix defines the metric for each dataset.

Sourcespfmbench primary benchmark evidence · Benchmark construction and evaluation setup; Appendix task definitions
Evaluation procedure diagram
How it worksEvaluation procedure
Evaluation procedure1. Allowed inputs: Protein task datasets via configurable loaders and prediction heads.. Then: 2. Splits: Most datasets are split 8:1:1 using a 30% protein sequence-similarity threshold. Mutation datasets are explicitly exempt and preserve their original train/validation/test partitions.. Then: 3. Metrics: Task-specific metrics include AUROC for several binary function and interaction tasks, accuracy for categorical tasks, and Spearman correlation for continuous fitness, affinity and enzyme properties; the task appendix defines the metric for each dataset.Evaluation procedure1. Allowed inputs: Protein task datasets via configurable loaders and prediction heads.. Then: 2. Splits: Most datasets are split 8:1:1 using a 30% protein sequence-similarity threshold. Mutation datasets are explicitly exempt and preserve their original train/validation/test partitions.. Then: 3. Metrics: Task-specific metrics include AUROC for several binary function and interaction tasks, accuracy for categorical tasks, and Spearman correlation for continuous fitness, affinity and enzyme properties; the task appendix defines the metric for each dataset.Evaluation procedure1. Allowed inputs: Protein task datasets via configurable loaders and prediction heads.. Then: 2. Splits: Most datasets are split 8:1:1 using a 30% protein sequence-similarity threshold. Mutation datasets are explicitly exempt and preserve their original train/validation/test partitions.. Then: 3. Metrics: Task-specific metrics include AUROC for several binary function and interaction tasks, accuracy for categorical tasks, and Spearman correlation for continuous fitness, affinity and enzyme properties; the task appendix defines the metric for each dataset.

Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.

Sources (2)biomap-research/PFMBench official source; pfmbench primary benchmark evidence · Pinned README: Overview; Features; repository architecture; Benchmark construction; Benchmark construction and evaluation setup; Appendix task definitions

Source reviewed · Automated source review, 2026-09-16. All specifications and missing details

Results

Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.

PFMBench ANTI-RES: Antibiotic resistance

accuracy (fraction) · Higher values are better.

PFMBench ANTI-RES: Antibiotic resistance · Antibiotic resistance (PFMBench split)

Evidence origin: Author-reported evaluation.

pfmbench primary benchmark evidence · Table 1, row(Antibiotic resistance)
  • The metric differs by task, taken from Table 1, so these figures cannot be averaged into one score.
  • Table 3 scores come from adapter fine-tuning; the ProteinGym figure is zero-shot and is not comparable to them.
Comparison details and limitations

Every method PFMBench reports on Antibiotic resistance, scored with Accuracy on Antibiotic resistance.

  • Author-reported numbers, source checked but not independently reproduced.

Automated source review: 2026-09-18. Numerical source review does not establish independent reproduction.

Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.

Showing 12 of 12 matching rows.

Methods and evaluation design

Procedure, tasks and evaluated configurations

How it works

Evaluation methodology

PFMBench evaluates protein representations across structural, functional, interaction and engineering tasks. Most datasets use sequence-similarity partitions, while mutation datasets retain their original assay splits. A stability screen identifies a core task subset, which must remain distinguishable from the full collection.

Sourcespfmbench primary benchmark evidence · Benchmark construction and evaluation setup; Appendix task definitions

Evaluation design

Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.

These source-backed links do not make different protocols or scores interchangeable.

Baseline coverage

Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.

No concrete protocols are explicitly linked to this suite. Protocol identification and baseline selection are outstanding.

Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums

Coverage is derived from release 2026-09-29-06401fd5b220. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.

Run this benchmark

Fine-tune and score a protein foundation model

Clone the benchmark, create its environment, and run either a fine-tuning task or the zero-shot evaluation.

Generate predictions and evaluate them. This recipe does not establish reproduction of a particular published score.

Dataset access
Prepared by the repository as its README describes.
Model and weights
A published protein foundation model checkpoint.
Licences
Project licence: see repository. Upstream data licences are separate and unreported here.
Software
Python with the repository's conda environment.
Hardware
Not stated in the cited section. Several of these steps expect a GPU.
Required inputs and expected outputs

Inputs

  • A protein foundation model supported by the repository.

Outputs

  • Per-task scores under the benchmark's own splits.

Execution steps

  1. 1. Clone and create the environment (Command line)

    Source reviewed; these instructions have not been executed by rewire.

    # Clone the repo
    git clone https://github.com/biomap-research/PFMBench.git
    cd PFMBench
    
    # Install Python dependencies
    conda env create -f environment.yml
    
    # Or you can use our Docker image via: docker pull whwendell/pfmbench:latest
    PFMBench: repository README · README.md at 53758ffc, Installation, lines 29-36
  2. 2. Fine-tune a single task (Command line)

    Source reviewed; these instructions have not been executed by rewire.

    # Example: run fine-tuning with specific GPU and configs
    env CUDA_VISIBLE_DEVICES=0 \
        python tasks/main.py \
        --config_name binding_db \
        --pretrain_model_name esm2_35m \
        --offline 0
    PFMBench: repository README · README.md at 53758ffc, Fine-tuning a single task, lines 104-109
  3. 3. Zero-shot evaluation (Command line)

    Source reviewed; these instructions have not been executed by rewire.

    # Example: run zero-shot MSA KL-div scoring
    env CUDA_VISIBLE_DEVICES=0 \
        python zeroshot/msa_kl_light.py \
        --config_name zero_msa_kl \
        --pretrain_model_name esm2_35m \
        --offline 0
    PFMBench: repository README · README.md at 53758ffc, Zero-shot evaluation, lines 115-120

Use your own model

Run your model locally and return predictions keyed by the input IDs. The evaluator supplies biological inputs without test labels and owns scoring. This interface is not a sandbox for model code.

Pass your existing prediction function into this adapter. Its output direction must match the selected protocol.

class MyModelAdapter:
    def __init__(self, score):
        self.score = score

    def predict(self, inputs):
        return {row["id"]: float(self.score(row)) for row in inputs}

# adapter = MyModelAdapter(your_prediction_function)
# report = rewirebench.run(prepared, adapter, output="runs/my-model")

Alternatively, generate a keyed prediction file in your existing model environment and use the score-only recipe. Your model code and weights do not need to be shared.

PFMBench: repository README · README.md at 53758ffc
Scope and limitations
  • Quoted from the project's README and not executed by rewire, so the commands are evidence of what the project documents rather than a verified run.
  • The project may have changed since the pinned commit.
  • The metric differs by task, so the figures on this page cannot be averaged.
  • Table 3 scores are adapter fine-tuning; the ProteinGym figure is zero-shot and not comparable to them.

Contribute a result for review. The library can submit an exported evaluation for private review when intake is open. Check the contribution page for access and sign-in.

Original repository instructions

Run this benchmark

Official environment creation and fine-tuning/zero-shot examples are provided. Data download, model_zoom weights, task YAML and GPU selection must be supplied; the suggested Docker latest tag and model URLs are not immutable execution pins. This pass links the recipe without claiming a verified environment or resource budget.

A maintained rewire runner has not been verified for this benchmark. Check data access, weights, licences, dependencies and hardware in the linked official documentation; requirements have not been fully extracted.

biomap-research/PFMBench / readme.md · readme.md lines 26–56 and 99–124 (Installation and Quick Start)
Strengths, limitations and unresolved questions

Strengths and limitations

Strengths supported by sources

Limitations and conditions

  • Core-task selection based on one adapter’s repeated runs does not guarantee equal reliability for every model. Functional-label pretraining and mutation-specific splits require separate overlap checks.
    Sourcespfmbench primary benchmark evidence · Benchmark construction and evaluation setup; Appendix task definitions
Profile review details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Stable record: discovery-benchmark-pfmbench

Specifications

Inputs, training, access and other details

Explanatory profile: source reviewed · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetsTasks span structure, function, localization, interactions and other protein properties.
Sourcesbiomap-research/PFMBench official source · Pinned README: Overview; Features; repository architecture
SplitsMost datasets are split 8:1:1 using a 30% protein sequence-similarity threshold. Mutation datasets are explicitly exempt and preserve their original train/validation/test partitions.
Sourcespfmbench primary benchmark evidence · Benchmark construction
MetricsTask-specific metrics include AUROC for several binary function and interaction tasks, accuracy for categorical tasks, and Spearman correlation for continuous fitness, affinity and enzyme properties; the task appendix defines the metric for each dataset.
Sourcespfmbench primary benchmark evidence · Benchmark construction and evaluation setup; Appendix task definitions
BaselinesThe framework supports both fine-tuning on labels and zero-shot evaluations.
Sourcesbiomap-research/PFMBench official source · Pinned README: Overview; Features; repository architecture
Leakage controlsMost datasets use an 8:1:1 split with a 30% sequence-similarity threshold. Mutation datasets retain their original partitions. The paper also flags possible functional-label overlap for annotation-aware pretrained models.
Sourcespfmbench primary benchmark evidence · Benchmark construction and evaluation setup; Appendix task definitions
UncertaintyESM2-Adapter is evaluated over three runs to screen task stability. Its reported bias is the best-to-worst spread divided by mean performance; this is not a confidence interval or a rule proven for all models.
Sourcespfmbench primary benchmark evidence · Benchmark construction and evaluation setup; Appendix task definitions
Entity typeConfigurable protein foundation-model evaluation suite.
Sourcesbiomap-research/PFMBench official source · Pinned README: Overview; Features; repository architecture
OrganismsDataset dependent: named tasks include human and yeast protein interactions, broader protein-property collections, molecular binding data and mutation assays. Species is a property of each source dataset, not one suite-wide organism.
Sourcespfmbench primary benchmark evidence · Benchmark construction and evaluation setup; Appendix task definitions
AssaysTask-specific structure, function, localization and interaction labels.
Sourcesbiomap-research/PFMBench official source · Pinned README: Overview; Features; repository architecture
Allowed inputsProtein task datasets via configurable loaders and prediction heads.
Sourcesbiomap-research/PFMBench official source · Pinned README: Overview; Features; repository architecture
AdaptationBoth fine-tuning and zero-shot evaluation are supported; configurations define the adaptation budget.
Sourcesbiomap-research/PFMBench official source · Pinned README: Overview; Features; repository architecture

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Primary-paper discovery and source inspection. Source-checked results are not independently reproduced experiments.

Paper or primary resourceVersionReference
pfmbench primary benchmark evidence2506.14796v1Read source
DOI: 10.48550/arXiv.2506.14796
Historical gaps recorded on 2026-09-17

The catalogue now holds 141 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.

  • Core-model filtering and excluded poorly performing tasks make Table 3 a selected subset, not all 17 models across every task.
  • Task metrics and 30% sequence-identity splits vary; Table 1 maps each.
  • Full values staged only after task-specific extraction; no overall family ranking inferred.
Search and extraction details

primary protocol screened

Searches

  • Genomic Benchmarks collection genomic sequence classification PMC10150520
  • mRNABench PMC12265608
  • PFMBench 2506.14796
  • ProteinBench 2409.06744

Evidence locations

  • Section 3.3
  • Tables 1–3

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

29 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Individual claims
pfmbench primary benchmark evidence

Original source ↗

Pinned README: Overview; Features; repository architecture; Benchmark construction; Benchmark construction and evaluation setup; Appendix task definitions

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2506.14796v1
Retrieved: 2026-09-16T21:06:29.716004+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 59c7bbb888e8e91f33c1e2cabfde062c32381d6ca727d23b6655c977aabf97a2

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram caption
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Individual claims
biomap-research/PFMBench official source

Original source ↗

Pinned README: Overview; Features; repository architecture; Benchmark construction; Benchmark construction and evaluation setup; Appendix task definitions

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 53758ffcbdf1d79b5d125383e4dd52d6fd59d2a1
Retrieved: 2026-09-16T10:30:21.987781+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 2736a9546e94b367e1bb3d1052e22460bb2188229d432b71eb9b01fe6b2a9b1a

Hash scope: Hash scope not separately documented; inspect source record

Diagram steps
  • Allowed inputs: Protein task datasets via configurable loaders and prediction heads.
  • Splits: Most datasets are split 8:1:1 using a 30% protein sequence-similarity threshold. Mutation datasets are explicitly exempt and preserve their original train/validation/test partitions.
  • Metrics: Task-specific metrics include AUROC for several binary function and interaction tasks, accuracy for categorical tasks, and Spearman correlation for continuous fitness, affinity and enzyme properties; the task appendix defines the metric for each dataset.
Individual claims
pfmbench primary benchmark evidence

Original source ↗

Pinned README: Overview; Features; repository architecture; Benchmark construction; Benchmark construction and evaluation setup; Appendix task definitions

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2506.14796v1
Retrieved: 2026-09-16T21:06:29.716004+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 59c7bbb888e8e91f33c1e2cabfde062c32381d6ca727d23b6655c977aabf97a2

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps
  • Allowed inputs: Protein task datasets via configurable loaders and prediction heads.
  • Splits: Most datasets are split 8:1:1 using a 30% protein sequence-similarity threshold. Mutation datasets are explicitly exempt and preserve their original train/validation/test partitions.
  • Metrics: Task-specific metrics include AUROC for several binary function and interaction tasks, accuracy for categorical tasks, and Spearman correlation for continuous fitness, affinity and enzyme properties; the task appendix defines the metric for each dataset.
Individual claims
biomap-research/PFMBench official source

Original source ↗

Pinned README: Overview; Features; repository architecture; Benchmark construction; Benchmark construction and evaluation setup; Appendix task definitions

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 53758ffcbdf1d79b5d125383e4dd52d6fd59d2a1
Retrieved: 2026-09-16T10:30:21.987781+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 2736a9546e94b367e1bb3d1052e22460bb2188229d432b71eb9b01fe6b2a9b1a

Hash scope: Hash scope not separately documented; inspect source record

Diagram title
Evaluation procedure
Individual claims
pfmbench primary benchmark evidence

Original source ↗

Pinned README: Overview; Features; repository architecture; Benchmark construction; Benchmark construction and evaluation setup; Appendix task definitions

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2506.14796v1
Retrieved: 2026-09-16T21:06:29.716004+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: 59c7bbb888e8e91f33c1e2cabfde062c32381d6ca727d23b6655c977aabf97a2

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title
Evaluation procedure
Individual claims
biomap-research/PFMBench official source

Original source ↗

Pinned README: Overview; Features; repository architecture; Benchmark construction; Benchmark construction and evaluation setup; Appendix task definitions

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 53758ffcbdf1d79b5d125383e4dd52d6fd59d2a1
Retrieved: 2026-09-16T10:30:21.987781+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: 2736a9546e94b367e1bb3d1052e22460bb2188229d432b71eb9b01fe6b2a9b1a

Hash scope: Hash scope not separately documented; inspect source record

Datasets
Tasks span structure, function, localization, interactions and other protein properties.
Individual claims
biomap-research/PFMBench official source

Original source ↗

Pinned README: Overview; Features; repository architecture

Version: 53758ffcbdf1d79b5d125383e4dd52d6fd59d2a1
Retrieved: 2026-09-16T10:30:21.987781+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: 2736a9546e94b367e1bb3d1052e22460bb2188229d432b71eb9b01fe6b2a9b1a

Hash scope: Hash scope not separately documented; inspect source record

Splits
Most datasets are split 8:1:1 using a 30% protein sequence-similarity threshold. Mutation datasets are explicitly exempt and preserve their original train/validation/test partitions.
Individual claims
pfmbench primary benchmark evidence

Original source ↗

Benchmark construction

Version: 2506.14796v1
Retrieved: 2026-09-16T21:06:29.716004+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: 59c7bbb888e8e91f33c1e2cabfde062c32381d6ca727d23b6655c977aabf97a2

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Adaptation
Both fine-tuning and zero-shot evaluation are supported; configurations define the adaptation budget.
Individual claims
biomap-research/PFMBench official source

Original source ↗

Pinned README: Overview; Features; repository architecture

Version: 53758ffcbdf1d79b5d125383e4dd52d6fd59d2a1
Retrieved: 2026-09-16T10:30:21.987781+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: 2736a9546e94b367e1bb3d1052e22460bb2188229d432b71eb9b01fe6b2a9b1a

Hash scope: Hash scope not separately documented; inspect source record

Metrics
Task-specific metrics include AUROC for several binary function and interaction tasks, accuracy for categorical tasks, and Spearman correlation for continuous fitness, affinity and enzyme properties; the task appendix defines the metric for each dataset.
Individual claims
pfmbench primary benchmark evidence

Original source ↗

Benchmark construction and evaluation setup; Appendix task definitions

Version: 2506.14796v1
Retrieved: 2026-09-16T21:06:29.716004+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: 59c7bbb888e8e91f33c1e2cabfde062c32381d6ca727d23b6655c977aabf97a2

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: discovered

5 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: discovery-benchmark-pfmbench

areas
protein-function
entity level
suite
scope note
Specialist molecular or omics evaluation; protocol details require review before numerical comparison.
task
Protein representation evaluation
version
Not reported
benchmark research
review date: 2026-09-17; status: primary_protocol_screened; primary sources: expansion-p3-pfmbench; inspected locators: Section 3.3; Tables 1–3; searched queries: Genomic Benchmarks collection genomic sequence classification PMC10150520; mRNABench PMC12265608; PFMBench 2506.14796; ProteinBench 2409.06744; gaps: Core-model filtering and excluded poorly performing tasks make Table 3 a selected subset, not all 17 models across every task.; Task metrics and 30% sequence-identity splits vary; Table 1 maps each.; Full values staged only after task-specific extraction; no overall family ranking inferred.; claim scope: Primary-paper discovery and source inspection. Source-checked results are not independently reproduced experiments.
historical missing metadata
dataset release: unextracted; metric implementation: unextracted; split manifest: unextracted; version: unextracted
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
entity classification
review date: 2026-09-17; rationale: The cited profile describes a collection of evaluation tasks or protocols; retain it as the top-level benchmark suite. Its datasets and individual protocols remain separate records.; source ids: src-discovery-biomap-research-pfmbench; source locator: Pinned README: Overview; Features; repository architecture; ambiguities: None recorded
run documentation
record id: discovery-benchmark-pfmbench; source ids: run-doc-pfmbench-readme-md-53758ffc; status: official_documentation_linked; summary: Official environment creation and fine-tuning/zero-shot examples are provided. Data download, model_zoom weights, task YAML and GPU selection must be supplied; the suggested Docker latest tag and model URLs are not immutable execution pins. This pass links the recipe without claiming a verified environment or resource budget.; source locator: readme.md lines 26–56 and 99–124 (Installation and Quick Start)
run recipes
id: pfmbench-official; protocol id: discovery-benchmark-pfmbench; version: 53758ffcbdf1d79b5d125383e4dd52d6fd59d2a1; title: Fine-tune and score a protein foundation model; purpose: generate_and_evaluate; summary: Clone the benchmark, create its environment, and run either a fine-tuning task or the zero-shot evaluation.; inputs: A protein foundation model supported by the repository.; outputs: Per-task scores under the benchmark's own splits.; requirements: data: Prepared by the repository as its README describes.; weights: A published protein foundation model checkpoint.; licence: Project licence: see repository. Upstream data licences are separate and unreported here.; software: Python with the repository's conda environment.; hardware: Not stated in the cited section. Several of these steps expect a GPU.; instructions: runtime: command_line; title: Clone and create the environment; code: # Clone the repo git clone https://github.com/biomap-research/PFMBench.git cd PFMBench # Install Python dependencies conda env create -f environment.yml # Or you can use our Docker image via: docker pull whwendell/pfmbench:latest; status: source_reviewed_not_executed; source ids: project-recipe-pfmbench-53758ffc; source locator: README.md at 53758ffc, Installation, lines 29-36; runtime: command_line; title: Fine-tune a single task; code: # Example: run fine-tuning with specific GPU and configs env CUDA_VISIBLE_DEVICES=0 \ python tasks/main.py \ --config_name binding_db \ --pretrain_model_name esm2_35m \ --offline 0; status: source_reviewed_not_executed; source ids: project-recipe-pfmbench-53758ffc; source locator: README.md at 53758ffc, Fine-tuning a single task, lines 104-109; runtime: command_line; title: Zero-shot evaluation; code: # Example: run zero-shot MSA KL-div scoring env CUDA_VISIBLE_DEVICES=0 \ python zeroshot/msa_kl_light.py \ --config_name zero_msa_kl \ --pretrain_model_name esm2_35m \ --offline 0; status: source_reviewed_not_executed; source ids: project-recipe-pfmbench-53758ffc; source locator: README.md at 53758ffc, Zero-shot evaluation, lines 115-120; limitations: Quoted from the project's README and not executed by rewire, so the commands are evidence of what the project documents rather than a verified run.; The project may have changed since the pinned commit.; The metric differs by task, so the figures on this page cannot be averaged.; Table 3 scores are adapter fine-tuning; the ProteinGym figure is zero-shot and not comparable to them.; source ids: project-recipe-pfmbench-53758ffc; source locator: README.md at 53758ffc
Related records

Suggest a correction