Datasets
Task-specific structure, sequence and complex evaluation collections.
ProteinBench assesses multiple protein-model tasks using quality, novelty, diversity and robustness dimensions.
Task-specific structure, sequence and complex evaluation collections.
Metric sets vary by task and include structure-predictor confidence, structural similarity and diversity measures.
Task-specific sequences, structures or complexes.
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Source reviewed · Automated source review, 2026-09-16. All specifications and missing details
Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.
aar (percent) · Higher values are better.
ProteinBench AB-ACCURACY-AAR: Accuracy AAR · RAbD, 55 antibody-antigen complexes (ProteinBench split)
Evidence origin: Author-reported evaluation.
proteinbench primary benchmark evidence · Table 6, column(Accuracy AAR ↑)Every method ProteinBench reports on Accuracy AAR, scored with AAR on RAbD, 55 antibody-antigen complexes.
Automated source review: 2026-09-19. Numerical source review does not establish independent reproduction.
Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.
Showing 8 of 8 matching rows.
ProteinBench evaluates protein sequence and structure methods across design, folding and conformational tasks. It distinguishes quality, diversity and novelty instead of treating all protein capabilities as one accuracy score. Each task has its own reference data, sampling procedure and baseline family.
Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.
These source-backed links do not make different protocols or scores interchangeable.
Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.
No concrete protocols are explicitly linked to this suite. Protocol identification and baseline selection are outstanding.
Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums
Coverage is derived from release 2026-09-29-06401fd5b220. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.
The primary project page describes multiple protein evaluation tasks and their results, but the inspected page does not establish a single reproducible run command with data and checkpoint setup. Choose a concrete task/protocol and its implementation rather than treating the summary website as an executable suite.
A maintained rewire runner has not been verified for this benchmark. Check data access, weights, licences, dependencies and hardware in the linked official documentation; requirements have not been fully extracted.
proteinbench official run documentation · Official project page: benchmark task descriptions, results and model/resource linksPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.
Stable record: discovery-benchmark-proteinbenchExplanatory profile: source reviewed · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Datasets | Task-specific structure, sequence and complex evaluation collections.Sourcesproteinbench official source · ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions |
| Splits | The framework distinguishes natural in-distribution structures from generated-backbone out-of-distribution evaluations.Sourcesproteinbench official source · ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions |
| Metrics | Metric sets vary by task and include structure-predictor confidence, structural similarity and diversity measures.Sourcesproteinbench official source · ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions |
| Baselines | Comparators vary by task: structure predictors, inverse-folding/design models, Rosetta-based antibody methods and molecular-dynamics or ensemble references. The paper defines the eligible method and input information separately for each task.Sourcesproteinbench primary benchmark evidence · Task definitions; ATLAS evaluation; antibody-design setup; result tables |
| Leakage controls | The antibody setup clusters CDR-H3 sequences at 40% similarity and excludes clusters containing RAbD test complexes from training/validation. The ATLAS ensemble task also applies a held-out protocol for models trained on ATLAS; these are task-specific controls, not a universal suite split.Sourcesproteinbench primary benchmark evidence · Task definitions; ATLAS evaluation; antibody-design setup; result tables |
| Uncertainty | Some task tables report repeated-experiment averages and standard deviations; others report medians.Sourcesproteinbench official source · ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions |
| Entity type | Protein-model evaluation framework spanning several tasks.Sourcesproteinbench official source · ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions |
| Organisms | No single organism defines its sequence, structure and complex task collections. · Not applicableSourcesproteinbench official source · ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions |
| Assays | Task-dependent protein structure and sequence/property references.Sourcesproteinbench official source · ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions |
| Allowed inputs | Task-specific sequences, structures or complexes.Sourcesproteinbench official source · ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions |
| Adaptation | Tasks define their own generation, prediction or adaptation regimes.Sourcesproteinbench official source · ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions |
Applicability is distinct from a completed evaluation.
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Last literature check: 2026-09-17. Primary-paper discovery and source inspection. Source-checked results are not independently reproduced experiments.
| Paper or primary resource | Version | Reference |
|---|---|---|
| proteinbench primary benchmark evidence | 2409.06744v1 | Read source DOI: 10.48550/arXiv.2409.06744 |
The catalogue now holds 556 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.
primary protocol screened
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
18 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions. Individual claims | proteinbench official source ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions Version: Retrieved website snapshot sha256:2e488850a6557bb57407615f2df9194351718b3dc0298a03c0c97d8e93460712 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| proteinbench official source ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions Version: Retrieved website snapshot sha256:2e488850a6557bb57407615f2df9194351718b3dc0298a03c0c97d8e93460712 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Evaluation procedure Individual claims | proteinbench official source ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions Version: Retrieved website snapshot sha256:2e488850a6557bb57407615f2df9194351718b3dc0298a03c0c97d8e93460712 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Datasets Task-specific structure, sequence and complex evaluation collections. Individual claims | proteinbench official source ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions Version: Retrieved website snapshot sha256:2e488850a6557bb57407615f2df9194351718b3dc0298a03c0c97d8e93460712 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Splits The framework distinguishes natural in-distribution structures from generated-backbone out-of-distribution evaluations. Individual claims | proteinbench official source ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions Version: Retrieved website snapshot sha256:2e488850a6557bb57407615f2df9194351718b3dc0298a03c0c97d8e93460712 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Adaptation Tasks define their own generation, prediction or adaptation regimes. Individual claims | proteinbench official source ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions Version: Retrieved website snapshot sha256:2e488850a6557bb57407615f2df9194351718b3dc0298a03c0c97d8e93460712 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Metrics Metric sets vary by task and include structure-predictor confidence, structural similarity and diversity measures. Individual claims | proteinbench official source ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions Version: Retrieved website snapshot sha256:2e488850a6557bb57407615f2df9194351718b3dc0298a03c0c97d8e93460712 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Baselines Comparators vary by task: structure predictors, inverse-folding/design models, Rosetta-based antibody methods and molecular-dynamics or ensemble references. The paper defines the eligible method and input information separately for each task. Individual claims | proteinbench primary benchmark evidence Task definitions; ATLAS evaluation; antibody-design setup; result tables Version: 2409.06744v1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Leakage controls The antibody setup clusters CDR-H3 sequences at 40% similarity and excludes clusters containing RAbD test complexes from training/validation. The ATLAS ensemble task also applies a held-out protocol for models trained on ATLAS; these are task-specific controls, not a universal suite split. Individual claims | proteinbench primary benchmark evidence Task definitions; ATLAS evaluation; antibody-design setup; result tables Version: 2409.06744v1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Uncertainty Some task tables report repeated-experiment averages and standard deviations; others report medians. Individual claims | proteinbench official source ProteinBench official website: Abstract; sequence/structure task evaluations; metric descriptions Version: Retrieved website snapshot sha256:2e488850a6557bb57407615f2df9194351718b3dc0298a03c0c97d8e93460712 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
View linked audit checks and correction history
Release 2026-09-29-06401fd5b220 · Record review: discovered
Stable ID: discovery-benchmark-proteinbench