rewire.itbenchmarks
Benchmark

ProteinBench

ProteinBench assesses multiple protein-model tasks using quality, novelty, diversity and robustness dimensions.

Sourcesproteinbench official source · ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions

556 evaluations · 556 results

Overview

Datasets

Task-specific structure, sequence and complex evaluation collections.

Metrics

Metric sets vary by task and include structure-predictor confidence, structural similarity and diversity measures.

Allowed inputs

Task-specific sequences, structures or complexes.

Sourcesproteinbench official source · ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions
Evaluation procedure diagram
How it worksEvaluation procedure
Evaluation procedure1. Allowed inputs: Task-specific sequences, structures or complexes.. Then: 2. Splits: The framework distinguishes natural in-distribution structures from generated-backbone out-of-distribution evaluations.. Then: 3. Metrics: Metric sets vary by task and include structure-predictor confidence, structural similarity and diversity measures.Evaluation procedure1. Allowed inputs: Task-specific sequences, structures or complexes.. Then: 2. Splits: The framework distinguishes natural in-distribution structures from generated-backbone out-of-distribution evaluations.. Then: 3. Metrics: Metric sets vary by task and include structure-predictor confidence, structural similarity and diversity measures.Evaluation procedure1. Allowed inputs: Task-specific sequences, structures or complexes.. Then: 2. Splits: The framework distinguishes natural in-distribution structures from generated-backbone out-of-distribution evaluations.. Then: 3. Metrics: Metric sets vary by task and include structure-predictor confidence, structural similarity and diversity measures.

Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.

Sourcesproteinbench official source · ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions

Source reviewed · Automated source review, 2026-09-16. All specifications and missing details

Results

Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.

ProteinBench AB-ACCURACY-AAR: Accuracy AAR

aar (percent) · Higher values are better.

ProteinBench AB-ACCURACY-AAR: Accuracy AAR · RAbD, 55 antibody-antigen complexes (ProteinBench split)

Evidence origin: Author-reported evaluation.

proteinbench primary benchmark evidence · Table 6, column(Accuracy AAR ↑)
  • Many of these metrics are better when lower, including scRMSD, perplexity and binding energy. The direction comes from the arrow in the paper's own header.
  • Rows such as Native PDBs and RAbD are the natural reference the authors include for scale, not a competing method.
Comparison details and limitations

Every method ProteinBench reports on Accuracy AAR, scored with AAR on RAbD, 55 antibody-antigen complexes.

  • Author-reported numbers, source checked but not independently reproduced.

Automated source review: 2026-09-19. Numerical source review does not establish independent reproduction.

Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.

Showing 8 of 8 matching rows.

Methods and evaluation design

Procedure, tasks and evaluated configurations

How it works

Evaluation methodology

ProteinBench evaluates protein sequence and structure methods across design, folding and conformational tasks. It distinguishes quality, diversity and novelty instead of treating all protein capabilities as one accuracy score. Each task has its own reference data, sampling procedure and baseline family.

Sourcesproteinbench primary benchmark evidence · Task definitions; ATLAS evaluation; antibody-design setup; result tables

Evaluation design

Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.

Tasks

These source-backed links do not make different protocols or scores interchangeable.

Baseline coverage

Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.

No concrete protocols are explicitly linked to this suite. Protocol identification and baseline selection are outstanding.

Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums

Coverage is derived from release 2026-09-29-06401fd5b220. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.

Run this benchmark

The primary project page describes multiple protein evaluation tasks and their results, but the inspected page does not establish a single reproducible run command with data and checkpoint setup. Choose a concrete task/protocol and its implementation rather than treating the summary website as an executable suite.

A maintained rewire runner has not been verified for this benchmark. Check data access, weights, licences, dependencies and hardware in the linked official documentation; requirements have not been fully extracted.

proteinbench official run documentation · Official project page: benchmark task descriptions, results and model/resource links
Strengths, limitations and unresolved questions

Strengths and limitations

Strengths supported by sources

  • Separate task collections expose which capability a result measures.
    Sourcesproteinbench official source · ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions

Limitations and conditions

  • Architectures receive different inputs and can generate different numbers of candidates. Task-specific exclusion rules and sampling budgets must accompany any comparison.
    Sourcesproteinbench primary benchmark evidence · Task definitions; ATLAS evaluation; antibody-design setup; result tables
Profile review details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Stable record: discovery-benchmark-proteinbench

Specifications

Inputs, training, access and other details

Explanatory profile: source reviewed · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetsTask-specific structure, sequence and complex evaluation collections.
Sourcesproteinbench official source · ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions
SplitsThe framework distinguishes natural in-distribution structures from generated-backbone out-of-distribution evaluations.
Sourcesproteinbench official source · ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions
MetricsMetric sets vary by task and include structure-predictor confidence, structural similarity and diversity measures.
Sourcesproteinbench official source · ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions
BaselinesComparators vary by task: structure predictors, inverse-folding/design models, Rosetta-based antibody methods and molecular-dynamics or ensemble references. The paper defines the eligible method and input information separately for each task.
Sourcesproteinbench primary benchmark evidence · Task definitions; ATLAS evaluation; antibody-design setup; result tables
Leakage controlsThe antibody setup clusters CDR-H3 sequences at 40% similarity and excludes clusters containing RAbD test complexes from training/validation. The ATLAS ensemble task also applies a held-out protocol for models trained on ATLAS; these are task-specific controls, not a universal suite split.
Sourcesproteinbench primary benchmark evidence · Task definitions; ATLAS evaluation; antibody-design setup; result tables
UncertaintySome task tables report repeated-experiment averages and standard deviations; others report medians.
Sourcesproteinbench official source · ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions
Entity typeProtein-model evaluation framework spanning several tasks.
Sourcesproteinbench official source · ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions
OrganismsNo single organism defines its sequence, structure and complex task collections. · Not applicable
Sourcesproteinbench official source · ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions
AssaysTask-dependent protein structure and sequence/property references.
Sourcesproteinbench official source · ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions
Allowed inputsTask-specific sequences, structures or complexes.
Sourcesproteinbench official source · ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions
AdaptationTasks define their own generation, prediction or adaptation regimes.
Sourcesproteinbench official source · ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions
Applicable tests and references

Applicability is distinct from a completed evaluation.

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Primary-paper discovery and source inspection. Source-checked results are not independently reproduced experiments.

Paper or primary resourceVersionReference
proteinbench primary benchmark evidence2409.06744v1Read source
DOI: 10.48550/arXiv.2409.06744
Historical gaps recorded on 2026-09-17

The catalogue now holds 556 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.

  • Mean/median slash pairs in Table 7 are not uncertainty.
  • EigenFold removes unknown amino acids and lacks a peptide-bond metric; preserve applicability caveats.
  • Do not transfer Table 8 bootstrap uncertainty to Table 7.
Search and extraction details

primary protocol screened

Searches

  • Genomic Benchmarks collection genomic sequence classification PMC10150520
  • mRNABench PMC12265608
  • PFMBench 2506.14796
  • ProteinBench 2409.06744

Evidence locations

  • Section 3.2.1
  • Table 7
  • Table 8 caption

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

18 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Individual claims
proteinbench official source

Original source ↗

ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions

Version: Retrieved website snapshot sha256:2e488850a6557bb57407615f2df9194351718b3dc0298a03c0c97d8e93460712
Retrieved: 2026-09-16T10:31:54.349409+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 2e488850a6557bb57407615f2df9194351718b3dc0298a03c0c97d8e93460712

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps
  • Allowed inputs: Task-specific sequences, structures or complexes.
  • Splits: The framework distinguishes natural in-distribution structures from generated-backbone out-of-distribution evaluations.
  • Metrics: Metric sets vary by task and include structure-predictor confidence, structural similarity and diversity measures.
Individual claims
proteinbench official source

Original source ↗

ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions

Version: Retrieved website snapshot sha256:2e488850a6557bb57407615f2df9194351718b3dc0298a03c0c97d8e93460712
Retrieved: 2026-09-16T10:31:54.349409+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 2e488850a6557bb57407615f2df9194351718b3dc0298a03c0c97d8e93460712

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title
Evaluation procedure
Individual claims
proteinbench official source

Original source ↗

ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions

Version: Retrieved website snapshot sha256:2e488850a6557bb57407615f2df9194351718b3dc0298a03c0c97d8e93460712
Retrieved: 2026-09-16T10:31:54.349409+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: 2e488850a6557bb57407615f2df9194351718b3dc0298a03c0c97d8e93460712

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Datasets
Task-specific structure, sequence and complex evaluation collections.
Individual claims
proteinbench official source

Original source ↗

ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions

Version: Retrieved website snapshot sha256:2e488850a6557bb57407615f2df9194351718b3dc0298a03c0c97d8e93460712
Retrieved: 2026-09-16T10:31:54.349409+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: 2e488850a6557bb57407615f2df9194351718b3dc0298a03c0c97d8e93460712

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Splits
The framework distinguishes natural in-distribution structures from generated-backbone out-of-distribution evaluations.
Individual claims
proteinbench official source

Original source ↗

ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions

Version: Retrieved website snapshot sha256:2e488850a6557bb57407615f2df9194351718b3dc0298a03c0c97d8e93460712
Retrieved: 2026-09-16T10:31:54.349409+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: 2e488850a6557bb57407615f2df9194351718b3dc0298a03c0c97d8e93460712

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Adaptation
Tasks define their own generation, prediction or adaptation regimes.
Individual claims
proteinbench official source

Original source ↗

ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions

Version: Retrieved website snapshot sha256:2e488850a6557bb57407615f2df9194351718b3dc0298a03c0c97d8e93460712
Retrieved: 2026-09-16T10:31:54.349409+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: 2e488850a6557bb57407615f2df9194351718b3dc0298a03c0c97d8e93460712

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Metrics
Metric sets vary by task and include structure-predictor confidence, structural similarity and diversity measures.
Individual claims
proteinbench official source

Original source ↗

ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions

Version: Retrieved website snapshot sha256:2e488850a6557bb57407615f2df9194351718b3dc0298a03c0c97d8e93460712
Retrieved: 2026-09-16T10:31:54.349409+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: 2e488850a6557bb57407615f2df9194351718b3dc0298a03c0c97d8e93460712

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Baselines
Comparators vary by task: structure predictors, inverse-folding/design models, Rosetta-based antibody methods and molecular-dynamics or ensemble references. The paper defines the eligible method and input information separately for each task.
Individual claims
proteinbench primary benchmark evidence

Original source ↗

Task definitions; ATLAS evaluation; antibody-design setup; result tables

Version: 2409.06744v1
Retrieved: 2026-09-16T21:07:13.231727+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.3.value

Source artifact SHA-256: 4334d636223ad42bfb9ae68aae03f5a255c29ba1cebe7b8f9588fb3b9b5453b2

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Leakage controls
The antibody setup clusters CDR-H3 sequences at 40% similarity and excludes clusters containing RAbD test complexes from training/validation. The ATLAS ensemble task also applies a held-out protocol for models trained on ATLAS; these are task-specific controls, not a universal suite split.
Individual claims
proteinbench primary benchmark evidence

Original source ↗

Task definitions; ATLAS evaluation; antibody-design setup; result tables

Version: 2409.06744v1
Retrieved: 2026-09-16T21:07:13.231727+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: 4334d636223ad42bfb9ae68aae03f5a255c29ba1cebe7b8f9588fb3b9b5453b2

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Uncertainty
Some task tables report repeated-experiment averages and standard deviations; others report medians.
Individual claims
proteinbench official source

Original source ↗

ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions

Version: Retrieved website snapshot sha256:2e488850a6557bb57407615f2df9194351718b3dc0298a03c0c97d8e93460712
Retrieved: 2026-09-16T10:31:54.349409+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.5.value

Source artifact SHA-256: 2e488850a6557bb57407615f2df9194351718b3dc0298a03c0c97d8e93460712

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: discovered

5 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: discovery-benchmark-proteinbench

areas
protein-structure
entity level
suite
scope note
Specialist molecular or omics evaluation; protocol details require review before numerical comparison.
task
Protein prediction, design and dynamics evaluation
version
Not reported
benchmark research
review date: 2026-09-17; status: primary_protocol_screened; primary sources: expansion-p3-proteinbench; inspected locators: Section 3.2.1; Table 7; Table 8 caption; searched queries: Genomic Benchmarks collection genomic sequence classification PMC10150520; mRNABench PMC12265608; PFMBench 2506.14796; ProteinBench 2409.06744; gaps: Mean/median slash pairs in Table 7 are not uncertainty.; EigenFold removes unknown amino acids and lacks a peptide-bond metric; preserve applicability caveats.; Do not transfer Table 8 bootstrap uncertainty to Table 7.; claim scope: Primary-paper discovery and source inspection. Source-checked results are not independently reproduced experiments.
historical missing metadata
dataset release: unextracted; metric implementation: unextracted; split manifest: unextracted; version: unextracted
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
entity classification
review date: 2026-09-17; rationale: The cited profile describes a collection of evaluation tasks or protocols; retain it as the top-level benchmark suite. Its datasets and individual protocols remain separate records.; source ids: evidence-benchmark-proteinbench-snapshot; source locator: ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions; ambiguities: None recorded
run documentation
record id: discovery-benchmark-proteinbench; source ids: run-doc-proteinbench-official-20260917; status: official_documentation_linked; summary: The primary project page describes multiple protein evaluation tasks and their results, but the inspected page does not establish a single reproducible run command with data and checkpoint setup. Choose a concrete task/protocol and its implementation rather than treating the summary website as an executable suite.; source locator: Official project page: benchmark task descriptions, results and model/resource links
Related records

Suggest a correction