rewirebio.iobenchmarks
Evaluation

Boltz-2 on Best DockQ against sampling depth on cognate nanobody-antigen complexes (Smorodina et al. 2026)

Published comparison; transcribed, not reproduced.

Research readiness

These checks assess whether the evidence supports a reproducible investigation. A source-checked score alone does not meet these requirements.

Release 2026-10-10-6e93f504adfc · Evidence verified: Not verified

Evidence incomplete

Replay metrics

Exact outcomes, predictions, identifiers and evaluator are connected.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • metric replay: verification is missing

Verified: Not verified

Evidence incomplete

Investigate discrepancies

Replay evidence includes annotations and an assessment of dependence. Unknown independence permits descriptive analysis only.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • metric replay: verification is missing
  • annotations: verification is missing
  • dependence: verification is missing

Verified: Not verified

Evidence incomplete

Run locally

A pinned recipe describes the inputs, environment and resource requirements.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • recipe pinned: verification is missing
  • resource estimate: verification is missing

Verified: Not verified

Evidence incomplete

Validate independently

Separate data and exposure records support an independent test.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • independent validation: verification is missing
  • overlap checked: verification is missing

Verified: Not verified

Readiness describes the evidence in this release. Availability on your computer is checked separately when an investigation runs. Existing data exposure can prevent independent validation even when files are available.

Artifacts and reproduction

No verified artifact manifest is connected to this record yet. The gaps above identify what is needed before analysis can begin.

Read reviewed discrepancy investigations

Evaluation results

1 evaluation · 2 results. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: Boltz-2 via Boltz CLI v2.2.0, 50 diffusion samples, seed 42, MSA server (Smorodina et al. 2026)Protocol: Best DockQ against sampling depth on cognate nanobody-antigen complexes
Dataset: Smorodina et al. 2026 nanobody-antigen benchmark: 106 cognate VHH-antigen complexes
0.57 dockq
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

Boltz-2 on Best DockQ against sampling depth on cognate nanobody-antigen complexes (Smorodina et al. 2026)

structural-20261009-protocol-smorodina2026-vhh-dockq-vs-sampling

Aggregation: Not reported

Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores · Results P36, Boltz-2 (0.57 to 0.80), N = 1
Configuration: Boltz-2 via Boltz CLI v2.2.0, 50 diffusion samples, seed 42, MSA server (Smorodina et al. 2026)Protocol: Best DockQ against sampling depth on cognate nanobody-antigen complexes
Dataset: Smorodina et al. 2026 nanobody-antigen benchmark: 106 cognate VHH-antigen complexes
0.8 dockq
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

Boltz-2 on Best DockQ against sampling depth on cognate nanobody-antigen complexes (Smorodina et al. 2026)

structural-20261009-protocol-smorodina2026-vhh-dockq-vs-sampling

Aggregation: Not reported

Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores · Results P36, Boltz-2 (0.57 to 0.80), N = 100

Source checking is not independent reproduction. Release 2026-10-10-6e93f504adfc.

Evaluation procedure

structural-20261009-protocol-smorodina2026-vhh-dockq-vs-sampling

Configuration
Boltz-2 via Boltz CLI v2.2.0, 50 diffusion samples, seed 42, MSA server (Smorodina et al. 2026)
Protocol
Best DockQ against sampling depth on cognate nanobody-antigen complexes
Dataset
Smorodina et al. 2026 nanobody-antigen benchmark: 106 cognate VHH-antigen complexes
origin
Independent external evaluation
configuration
Primary source as retrieved 2026-10-09
protocol id
structural-20261009-protocol-smorodina2026-vhh-dockq-vs-sampling
dataset version
bioRxiv v1 benchmark
split
Mixed: systems inside and outside the tool's training data
inputs
VHH and antigen sequences
adaptation
None
population
106 cognate complexes
metric implementation
DockQ v2.1.3
aggregation
Median over complexes of the per-complex maximum
budget
5 seeds; N = 1 and N = 100 samples

Metadata review: source checked. Unreported conditions prevent automatic comparisons.

Reproduction

Split
Mixed: systems inside and outside the tool's training data
Adaptation
None
Scoring implementation
DockQ v2.1.3

No execution recipe has been verified for this exact configuration and evaluation. A benchmark's general instructions may use different inputs, splits or model settings.

Reproducing this published result requires matching its model configuration, data, split and scorer. Source checking or a successful smoke test does not establish score reproduction.

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

19 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-10-10-6e93f504adfc
Property and statementOriginal source and locationReview and provenance
attributes.comparison.adaptation
None
Context-only references
Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores

Original source ↗

Results P36, 'Boltz-2'

Version: bioRxiv version 1, posted 2026-03-03; Europe PMC preprint full text PPR1221387 (manuscript EMS215481); not peer reviewed
Retrieved: 2026-10-09T21:25:48Z

not individually reviewed

No individual claim review recorded

independent paper

Audit details

Field: attributes.comparison.adaptation

Source artifact SHA-256: 0ad24054d98fd1888fa8bf3120f76e22f7bf6b1965261b16f045986d72c50732

Hash scope: JATS XML parse (xml.etree), by extract/extract_structural.py

Inspected artifact

attributes.comparison.aggregation
Median over complexes of the per-complex maximum
Context-only references
Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores

Original source ↗

Results P36, 'Boltz-2'

Version: bioRxiv version 1, posted 2026-03-03; Europe PMC preprint full text PPR1221387 (manuscript EMS215481); not peer reviewed
Retrieved: 2026-10-09T21:25:48Z

not individually reviewed

No individual claim review recorded

independent paper

Audit details

Field: attributes.comparison.aggregation

Source artifact SHA-256: 0ad24054d98fd1888fa8bf3120f76e22f7bf6b1965261b16f045986d72c50732

Hash scope: JATS XML parse (xml.etree), by extract/extract_structural.py

Inspected artifact

attributes.comparison.budget
5 seeds; N = 1 and N = 100 samples
Context-only references
Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores

Original source ↗

Results P36, 'Boltz-2'

Version: bioRxiv version 1, posted 2026-03-03; Europe PMC preprint full text PPR1221387 (manuscript EMS215481); not peer reviewed
Retrieved: 2026-10-09T21:25:48Z

not individually reviewed

No individual claim review recorded

independent paper

Audit details

Field: attributes.comparison.budget

Source artifact SHA-256: 0ad24054d98fd1888fa8bf3120f76e22f7bf6b1965261b16f045986d72c50732

Hash scope: JATS XML parse (xml.etree), by extract/extract_structural.py

Inspected artifact

attributes.comparison.dataset_version
bioRxiv v1 benchmark
Context-only references
Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores

Original source ↗

Results P36, 'Boltz-2'

Version: bioRxiv version 1, posted 2026-03-03; Europe PMC preprint full text PPR1221387 (manuscript EMS215481); not peer reviewed
Retrieved: 2026-10-09T21:25:48Z

not individually reviewed

No individual claim review recorded

independent paper

Audit details

Field: attributes.comparison.dataset_version

Source artifact SHA-256: 0ad24054d98fd1888fa8bf3120f76e22f7bf6b1965261b16f045986d72c50732

Hash scope: JATS XML parse (xml.etree), by extract/extract_structural.py

Inspected artifact

attributes.comparison.inputs
VHH and antigen sequences
Context-only references
Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores

Original source ↗

Results P36, 'Boltz-2'

Version: bioRxiv version 1, posted 2026-03-03; Europe PMC preprint full text PPR1221387 (manuscript EMS215481); not peer reviewed
Retrieved: 2026-10-09T21:25:48Z

not individually reviewed

No individual claim review recorded

independent paper

Audit details

Field: attributes.comparison.inputs

Source artifact SHA-256: 0ad24054d98fd1888fa8bf3120f76e22f7bf6b1965261b16f045986d72c50732

Hash scope: JATS XML parse (xml.etree), by extract/extract_structural.py

Inspected artifact

attributes.comparison.metric_implementation
DockQ v2.1.3
Context-only references
Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores

Original source ↗

Results P36, 'Boltz-2'

Version: bioRxiv version 1, posted 2026-03-03; Europe PMC preprint full text PPR1221387 (manuscript EMS215481); not peer reviewed
Retrieved: 2026-10-09T21:25:48Z

not individually reviewed

No individual claim review recorded

independent paper

Audit details

Field: attributes.comparison.metric_implementation

Source artifact SHA-256: 0ad24054d98fd1888fa8bf3120f76e22f7bf6b1965261b16f045986d72c50732

Hash scope: JATS XML parse (xml.etree), by extract/extract_structural.py

Inspected artifact

attributes.comparison.population
106 cognate complexes
Context-only references
Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores

Original source ↗

Results P36, 'Boltz-2'

Version: bioRxiv version 1, posted 2026-03-03; Europe PMC preprint full text PPR1221387 (manuscript EMS215481); not peer reviewed
Retrieved: 2026-10-09T21:25:48Z

not individually reviewed

No individual claim review recorded

independent paper

Audit details

Field: attributes.comparison.population

Source artifact SHA-256: 0ad24054d98fd1888fa8bf3120f76e22f7bf6b1965261b16f045986d72c50732

Hash scope: JATS XML parse (xml.etree), by extract/extract_structural.py

Inspected artifact

attributes.comparison.protocol_id
structural-20261009-protocol-smorodina2026-vhh-dockq-vs-sampling
Context-only references
Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores

Original source ↗

Results P36, 'Boltz-2'

Version: bioRxiv version 1, posted 2026-03-03; Europe PMC preprint full text PPR1221387 (manuscript EMS215481); not peer reviewed
Retrieved: 2026-10-09T21:25:48Z

not individually reviewed

No individual claim review recorded

independent paper

Audit details

Field: attributes.comparison.protocol_id

Source artifact SHA-256: 0ad24054d98fd1888fa8bf3120f76e22f7bf6b1965261b16f045986d72c50732

Hash scope: JATS XML parse (xml.etree), by extract/extract_structural.py

Inspected artifact

attributes.comparison.split
Mixed: systems inside and outside the tool's training data
Context-only references
Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores

Original source ↗

Results P36, 'Boltz-2'

Version: bioRxiv version 1, posted 2026-03-03; Europe PMC preprint full text PPR1221387 (manuscript EMS215481); not peer reviewed
Retrieved: 2026-10-09T21:25:48Z

not individually reviewed

No individual claim review recorded

independent paper

Audit details

Field: attributes.comparison.split

Source artifact SHA-256: 0ad24054d98fd1888fa8bf3120f76e22f7bf6b1965261b16f045986d72c50732

Hash scope: JATS XML parse (xml.etree), by extract/extract_structural.py

Inspected artifact

attributes.limitations
1 values
  • Run by authors who did not develop the tool.
Context-only references
Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores

Original source ↗

Results P36, 'Boltz-2'

Version: bioRxiv version 1, posted 2026-03-03; Europe PMC preprint full text PPR1221387 (manuscript EMS215481); not peer reviewed
Retrieved: 2026-10-09T21:25:48Z

not individually reviewed

No individual claim review recorded

independent paper

Audit details

Field: attributes.limitations

Source artifact SHA-256: 0ad24054d98fd1888fa8bf3120f76e22f7bf6b1965261b16f045986d72c50732

Hash scope: JATS XML parse (xml.etree), by extract/extract_structural.py

Inspected artifact

Sources and history

Release 2026-10-10-6e93f504adfc · Record review: source checked

1 source records and release historyDownload this release (gzip)
Technical metadata and extraction receipts

Stable ID: structural-20261009-eval-smorodina2026-vhh-dockq-vs-sampling-boltz2

areas
proteins-complexes
contexts
research
origin
independent_paper
protocol
structural-20261009-protocol-smorodina2026-vhh-dockq-vs-sampling
version
Primary source as retrieved 2026-10-09
comparison
protocol id: structural-20261009-protocol-smorodina2026-vhh-dockq-vs-sampling; dataset version: bioRxiv v1 benchmark; split: Mixed: systems inside and outside the tool's training data; inputs: VHH and antigen sequences; adaptation: None; population: 106 cognate complexes; metric implementation: DockQ v2.1.3; aggregation: Median over complexes of the per-complex maximum; budget: 5 seeds; N = 1 and N = 100 samples
source locator
Results P36, 'Boltz-2'
limitations
Run by authors who did not develop the tool.
Related records

Suggest a correction