Datasets
Tasks span structure, function, localization, interactions and other protein properties.
PFMBench is a configurable suite of protein-model downstream evaluations.
Tasks span structure, function, localization, interactions and other protein properties.
Task-specific metrics include AUROC for several binary function and interaction tasks, accuracy for categorical tasks, and Spearman correlation for continuous fitness, affinity and enzyme properties; the task appendix defines the metric for each dataset.
Protein task datasets via configurable loaders and prediction heads.
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Source reviewed · Automated source review, 2026-09-16. All specifications and missing details
Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.
accuracy (fraction) · Higher values are better.
PFMBench ANTI-RES: Antibiotic resistance · Antibiotic resistance (PFMBench split)
Evidence origin: Author-reported evaluation.
pfmbench primary benchmark evidence · Table 1, row(Antibiotic resistance)Every method PFMBench reports on Antibiotic resistance, scored with Accuracy on Antibiotic resistance.
Automated source review: 2026-09-18. Numerical source review does not establish independent reproduction.
Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.
Showing 12 of 12 matching rows.
PFMBench evaluates protein representations across structural, functional, interaction and engineering tasks. Most datasets use sequence-similarity partitions, while mutation datasets retain their original assay splits. A stability screen identifies a core task subset, which must remain distinguishable from the full collection.
Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.
These source-backed links do not make different protocols or scores interchangeable.
Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.
No concrete protocols are explicitly linked to this suite. Protocol identification and baseline selection are outstanding.
Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums
Coverage is derived from release 2026-09-29-06401fd5b220. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.
Clone the benchmark, create its environment, and run either a fine-tuning task or the zero-shot evaluation.
Generate predictions and evaluate them. This recipe does not establish reproduction of a particular published score.
Source reviewed; these instructions have not been executed by rewire.
# Clone the repo
git clone https://github.com/biomap-research/PFMBench.git
cd PFMBench
# Install Python dependencies
conda env create -f environment.yml
# Or you can use our Docker image via: docker pull whwendell/pfmbench:latestPFMBench: repository README · README.md at 53758ffc, Installation, lines 29-36Source reviewed; these instructions have not been executed by rewire.
# Example: run fine-tuning with specific GPU and configs
env CUDA_VISIBLE_DEVICES=0 \
python tasks/main.py \
--config_name binding_db \
--pretrain_model_name esm2_35m \
--offline 0PFMBench: repository README · README.md at 53758ffc, Fine-tuning a single task, lines 104-109Source reviewed; these instructions have not been executed by rewire.
# Example: run zero-shot MSA KL-div scoring
env CUDA_VISIBLE_DEVICES=0 \
python zeroshot/msa_kl_light.py \
--config_name zero_msa_kl \
--pretrain_model_name esm2_35m \
--offline 0PFMBench: repository README · README.md at 53758ffc, Zero-shot evaluation, lines 115-120Run your model locally and return predictions keyed by the input IDs. The evaluator supplies biological inputs without test labels and owns scoring. This interface is not a sandbox for model code.
Pass your existing prediction function into this adapter. Its output direction must match the selected protocol.
class MyModelAdapter:
def __init__(self, score):
self.score = score
def predict(self, inputs):
return {row["id"]: float(self.score(row)) for row in inputs}
# adapter = MyModelAdapter(your_prediction_function)
# report = rewirebench.run(prepared, adapter, output="runs/my-model")Alternatively, generate a keyed prediction file in your existing model environment and use the score-only recipe. Your model code and weights do not need to be shared.
PFMBench: repository README · README.md at 53758ffcContribute a result for review. The library can submit an exported evaluation for private review when intake is open. Check the contribution page for access and sign-in.
Official environment creation and fine-tuning/zero-shot examples are provided. Data download, model_zoom weights, task YAML and GPU selection must be supplied; the suggested Docker latest tag and model URLs are not immutable execution pins. This pass links the recipe without claiming a verified environment or resource budget.
A maintained rewire runner has not been verified for this benchmark. Check data access, weights, licences, dependencies and hardware in the linked official documentation; requirements have not been fully extracted.
biomap-research/PFMBench / readme.md · readme.md lines 26–56 and 99–124 (Installation and Quick Start)Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.
Stable record: discovery-benchmark-pfmbenchExplanatory profile: source reviewed · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Datasets | Tasks span structure, function, localization, interactions and other protein properties.Sourcesbiomap-research/PFMBench official source · Pinned README: Overview; Features; repository architecture |
| Splits | Most datasets are split 8:1:1 using a 30% protein sequence-similarity threshold. Mutation datasets are explicitly exempt and preserve their original train/validation/test partitions.Sourcespfmbench primary benchmark evidence · Benchmark construction |
| Metrics | Task-specific metrics include AUROC for several binary function and interaction tasks, accuracy for categorical tasks, and Spearman correlation for continuous fitness, affinity and enzyme properties; the task appendix defines the metric for each dataset.Sourcespfmbench primary benchmark evidence · Benchmark construction and evaluation setup; Appendix task definitions |
| Baselines | The framework supports both fine-tuning on labels and zero-shot evaluations.Sourcesbiomap-research/PFMBench official source · Pinned README: Overview; Features; repository architecture |
| Leakage controls | Most datasets use an 8:1:1 split with a 30% sequence-similarity threshold. Mutation datasets retain their original partitions. The paper also flags possible functional-label overlap for annotation-aware pretrained models.Sourcespfmbench primary benchmark evidence · Benchmark construction and evaluation setup; Appendix task definitions |
| Uncertainty | ESM2-Adapter is evaluated over three runs to screen task stability. Its reported bias is the best-to-worst spread divided by mean performance; this is not a confidence interval or a rule proven for all models.Sourcespfmbench primary benchmark evidence · Benchmark construction and evaluation setup; Appendix task definitions |
| Entity type | Configurable protein foundation-model evaluation suite.Sourcesbiomap-research/PFMBench official source · Pinned README: Overview; Features; repository architecture |
| Organisms | Dataset dependent: named tasks include human and yeast protein interactions, broader protein-property collections, molecular binding data and mutation assays. Species is a property of each source dataset, not one suite-wide organism.Sourcespfmbench primary benchmark evidence · Benchmark construction and evaluation setup; Appendix task definitions |
| Assays | Task-specific structure, function, localization and interaction labels.Sourcesbiomap-research/PFMBench official source · Pinned README: Overview; Features; repository architecture |
| Allowed inputs | Protein task datasets via configurable loaders and prediction heads.Sourcesbiomap-research/PFMBench official source · Pinned README: Overview; Features; repository architecture |
| Adaptation | Both fine-tuning and zero-shot evaluation are supported; configurations define the adaptation budget.Sourcesbiomap-research/PFMBench official source · Pinned README: Overview; Features; repository architecture |
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Last literature check: 2026-09-17. Primary-paper discovery and source inspection. Source-checked results are not independently reproduced experiments.
| Paper or primary resource | Version | Reference |
|---|---|---|
| pfmbench primary benchmark evidence | 2506.14796v1 | Read source DOI: 10.48550/arXiv.2506.14796 |
The catalogue now holds 141 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.
primary protocol screened
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
29 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions. Individual claims | pfmbench primary benchmark evidence Pinned README: Overview; Features; repository architecture; Benchmark construction; Benchmark construction and evaluation setup; Appendix task definitions Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 2506.14796v1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram caption Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions. Individual claims | biomap-research/PFMBench official source Pinned README: Overview; Features; repository architecture; Benchmark construction; Benchmark construction and evaluation setup; Appendix task definitions Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 53758ffcbdf1d79b5d125383e4dd52d6fd59d2a1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| pfmbench primary benchmark evidence Pinned README: Overview; Features; repository architecture; Benchmark construction; Benchmark construction and evaluation setup; Appendix task definitions Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 2506.14796v1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| biomap-research/PFMBench official source Pinned README: Overview; Features; repository architecture; Benchmark construction; Benchmark construction and evaluation setup; Appendix task definitions Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 53758ffcbdf1d79b5d125383e4dd52d6fd59d2a1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Evaluation procedure Individual claims | pfmbench primary benchmark evidence Pinned README: Overview; Features; repository architecture; Benchmark construction; Benchmark construction and evaluation setup; Appendix task definitions Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 2506.14796v1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Evaluation procedure Individual claims | biomap-research/PFMBench official source Pinned README: Overview; Features; repository architecture; Benchmark construction; Benchmark construction and evaluation setup; Appendix task definitions Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 53758ffcbdf1d79b5d125383e4dd52d6fd59d2a1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Datasets Tasks span structure, function, localization, interactions and other protein properties. Individual claims | biomap-research/PFMBench official source Pinned README: Overview; Features; repository architecture Version: 53758ffcbdf1d79b5d125383e4dd52d6fd59d2a1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Splits Most datasets are split 8:1:1 using a 30% protein sequence-similarity threshold. Mutation datasets are explicitly exempt and preserve their original train/validation/test partitions. Individual claims | pfmbench primary benchmark evidence Benchmark construction Version: 2506.14796v1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Adaptation Both fine-tuning and zero-shot evaluation are supported; configurations define the adaptation budget. Individual claims | biomap-research/PFMBench official source Pinned README: Overview; Features; repository architecture Version: 53758ffcbdf1d79b5d125383e4dd52d6fd59d2a1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Metrics Task-specific metrics include AUROC for several binary function and interaction tasks, accuracy for categorical tasks, and Spearman correlation for continuous fitness, affinity and enzyme properties; the task appendix defines the metric for each dataset. Individual claims | pfmbench primary benchmark evidence Benchmark construction and evaluation setup; Appendix task definitions Version: 2506.14796v1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
View linked audit checks and correction history
Release 2026-09-29-06401fd5b220 · Record review: discovered
Stable ID: discovery-benchmark-pfmbench