rewirebio.iobenchmarks
Protocol

How well ipTM tracks DockQ on cognate nanobody-antigen complexes

On the 106 cognate complexes, 50 samples per complex; correlation of ipTM with DockQ, the share of confident failures and unconfident successes, and whether ipTM follows DockQ gains from sampling.

3 evaluations · 12 results

Overview

On the 106 cognate complexes, 50 samples per complex; correlation of ipTM with DockQ, the share of confident failures and unconfident successes, and whether ipTM follows DockQ gains from sampling.

Consult the linked sources for architecture or protocol details. Missing evidence is not evidence of a missing capability.

3 recorded evaluations, 12 metric rows. A comparison chart has not yet been validated for these results. The table retains the individual findings and their sources.

View coverage and remaining gaps across all benchmarks

Results

Results are available, but no reviewed comparison panel is linked in this release.

All evaluations

3 evaluations · 12 results. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: AlphaFold3 3.0.1, 50 diffusion samples, seed 1, built-in data pipeline (Smorodina et al. 2026)Protocol: How well ipTM tracks DockQ on cognate nanobody-antigen complexes
Dataset: Smorodina et al. 2026 nanobody-antigen benchmark: 106 cognate VHH-antigen complexes
0.736 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

AF3 on How well ipTM tracks DockQ on cognate nanobody-antigen complexes (Smorodina et al. 2026)

structural-20261009-protocol-smorodina2026-vhh-iptm-dockq-calibration

Aggregation: Not reported

Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores · Results P19, AF3: ipTM against DockQ for the best-DockQ sample of each complex
Configuration: AlphaFold3 3.0.1, 50 diffusion samples, seed 1, built-in data pipeline (Smorodina et al. 2026)Protocol: How well ipTM tracks DockQ on cognate nanobody-antigen complexes
Dataset: Smorodina et al. 2026 nanobody-antigen benchmark: 106 cognate VHH-antigen complexes
-0.027 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

AF3 on How well ipTM tracks DockQ on cognate nanobody-antigen complexes (Smorodina et al. 2026)

structural-20261009-protocol-smorodina2026-vhh-iptm-dockq-calibration

Aggregation: Not reported

Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores · Results P21, AF3: change in ipTM against change in DockQ under saturation sampling
Configuration: AlphaFold3 3.0.1, 50 diffusion samples, seed 1, built-in data pipeline (Smorodina et al. 2026)Protocol: How well ipTM tracks DockQ on cognate nanobody-antigen complexes
Dataset: Smorodina et al. 2026 nanobody-antigen benchmark: 106 cognate VHH-antigen complexes
2% proportion
percent · lower

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

AF3 on How well ipTM tracks DockQ on cognate nanobody-antigen complexes (Smorodina et al. 2026)

structural-20261009-protocol-smorodina2026-vhh-iptm-dockq-calibration

Aggregation: Not reported

Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores · Results P19, AF3: share of predictions with ipTM at least 0.5 and DockQ below 0.23 (confident failures, Q2)
Configuration: AlphaFold3 3.0.1, 50 diffusion samples, seed 1, built-in data pipeline (Smorodina et al. 2026)Protocol: How well ipTM tracks DockQ on cognate nanobody-antigen complexes
Dataset: Smorodina et al. 2026 nanobody-antigen benchmark: 106 cognate VHH-antigen complexes
28% proportion
percent · lower

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

AF3 on How well ipTM tracks DockQ on cognate nanobody-antigen complexes (Smorodina et al. 2026)

structural-20261009-protocol-smorodina2026-vhh-iptm-dockq-calibration

Aggregation: Not reported

Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores · Results P19, AF3: share of predictions with ipTM below 0.5 and DockQ at least 0.23 (unconfident successes, Q4)
Configuration: AlphaFold3 3.0.1, 50 diffusion samples, seed 1, built-in data pipeline (Smorodina et al. 2026)Protocol: How well ipTM tracks DockQ on cognate nanobody-antigen complexes
Dataset: Smorodina et al. 2026 nanobody-antigen benchmark: 106 cognate VHH-antigen complexes
0.888 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

AF3 on How well ipTM tracks DockQ on cognate nanobody-antigen complexes (Smorodina et al. 2026)

structural-20261009-protocol-smorodina2026-vhh-iptm-dockq-calibration

Aggregation: Not reported

Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores · Results P19, AF3: ipTM against DockQ for the first sample (sample0) of each complex
Configuration: Boltz-2 via Boltz CLI v2.2.0, 50 diffusion samples, seed 42, MSA server (Smorodina et al. 2026)Protocol: How well ipTM tracks DockQ on cognate nanobody-antigen complexes
Dataset: Smorodina et al. 2026 nanobody-antigen benchmark: 106 cognate VHH-antigen complexes
0.665 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

Boltz-2 on How well ipTM tracks DockQ on cognate nanobody-antigen complexes (Smorodina et al. 2026)

structural-20261009-protocol-smorodina2026-vhh-iptm-dockq-calibration

Aggregation: Not reported

Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores · Results P19, Boltz-2: ipTM against DockQ for the best-DockQ sample of each complex
Configuration: Boltz-2 via Boltz CLI v2.2.0, 50 diffusion samples, seed 42, MSA server (Smorodina et al. 2026)Protocol: How well ipTM tracks DockQ on cognate nanobody-antigen complexes
Dataset: Smorodina et al. 2026 nanobody-antigen benchmark: 106 cognate VHH-antigen complexes
-0.04 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

Boltz-2 on How well ipTM tracks DockQ on cognate nanobody-antigen complexes (Smorodina et al. 2026)

structural-20261009-protocol-smorodina2026-vhh-iptm-dockq-calibration

Aggregation: Not reported

Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores · Results P21, Boltz-2: change in ipTM against change in DockQ under saturation sampling
Configuration: Boltz-2 via Boltz CLI v2.2.0, 50 diffusion samples, seed 42, MSA server (Smorodina et al. 2026)Protocol: How well ipTM tracks DockQ on cognate nanobody-antigen complexes
Dataset: Smorodina et al. 2026 nanobody-antigen benchmark: 106 cognate VHH-antigen complexes
18% proportion
percent · lower

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

Boltz-2 on How well ipTM tracks DockQ on cognate nanobody-antigen complexes (Smorodina et al. 2026)

structural-20261009-protocol-smorodina2026-vhh-iptm-dockq-calibration

Aggregation: Not reported

Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores · Results P19, Boltz-2: share of predictions with ipTM at least 0.5 and DockQ below 0.23 (confident failures, Q2)
Configuration: Boltz-2 via Boltz CLI v2.2.0, 50 diffusion samples, seed 42, MSA server (Smorodina et al. 2026)Protocol: How well ipTM tracks DockQ on cognate nanobody-antigen complexes
Dataset: Smorodina et al. 2026 nanobody-antigen benchmark: 106 cognate VHH-antigen complexes
0.493 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

Boltz-2 on How well ipTM tracks DockQ on cognate nanobody-antigen complexes (Smorodina et al. 2026)

structural-20261009-protocol-smorodina2026-vhh-iptm-dockq-calibration

Aggregation: Not reported

Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores · Results P19, Boltz-2: ipTM against DockQ for the first sample (sample0) of each complex
Configuration: Chai-1 0.6.1, 5 trunk x 10 diffusion samples, seed 42, ESM embeddings without MSAs (Smorodina et al. 2026)Protocol: How well ipTM tracks DockQ on cognate nanobody-antigen complexes
Dataset: Smorodina et al. 2026 nanobody-antigen benchmark: 106 cognate VHH-antigen complexes
0.612 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

Chai-1 on How well ipTM tracks DockQ on cognate nanobody-antigen complexes (Smorodina et al. 2026)

structural-20261009-protocol-smorodina2026-vhh-iptm-dockq-calibration

Aggregation: Not reported

Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores · Results P19, Chai-1: ipTM against DockQ for the best-DockQ sample of each complex
Configuration: Chai-1 0.6.1, 5 trunk x 10 diffusion samples, seed 42, ESM embeddings without MSAs (Smorodina et al. 2026)Protocol: How well ipTM tracks DockQ on cognate nanobody-antigen complexes
Dataset: Smorodina et al. 2026 nanobody-antigen benchmark: 106 cognate VHH-antigen complexes
-0.019 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

Chai-1 on How well ipTM tracks DockQ on cognate nanobody-antigen complexes (Smorodina et al. 2026)

structural-20261009-protocol-smorodina2026-vhh-iptm-dockq-calibration

Aggregation: Not reported

Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores · Results P21, Chai-1: change in ipTM against change in DockQ under saturation sampling
Configuration: Chai-1 0.6.1, 5 trunk x 10 diffusion samples, seed 42, ESM embeddings without MSAs (Smorodina et al. 2026)Protocol: How well ipTM tracks DockQ on cognate nanobody-antigen complexes
Dataset: Smorodina et al. 2026 nanobody-antigen benchmark: 106 cognate VHH-antigen complexes
24% proportion
percent · lower

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

Chai-1 on How well ipTM tracks DockQ on cognate nanobody-antigen complexes (Smorodina et al. 2026)

structural-20261009-protocol-smorodina2026-vhh-iptm-dockq-calibration

Aggregation: Not reported

Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores · Results P19, Chai-1: share of predictions with ipTM below 0.5 and DockQ at least 0.23 (unconfident successes, Q4)

Source checking is not independent reproduction. Release 2026-10-10-6e93f504adfc.

Methods and evaluation design

Procedure, tasks and evaluated configurations

Recorded evaluations

Each evaluation records what was tested and under which conditions.

Baseline coverage

Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.

0 of 2 active baseline roles have published Rewire measurements in this release. Measurements on a selected protocol do not establish coverage of an entire suite.

No execution recipe linked to this protocol. Recipe availability does not establish a completed evaluation.

External evaluations
3

Literature evidence is not a Rewire measurement. Executed but unpublished runs and private review status are not included.

Null control

Proposed control: requires review

Select a task-valid null control after reviewing inputs and metric

Protocol-specific applicability, permitted inputs, access, split, evaluator and execution requirements need review before implementation or execution.

This is a suggested selection rule, not a validated method or a measured score.

Conventional reference

Proposed control: requires review

Select an upstream conventional reference after reviewing the full protocol

Protocol-specific applicability, permitted inputs, access, split, evaluator and execution requirements need review before implementation or execution.

This is a suggested selection rule, not a validated method or a measured score.

Protocol coverage CSV (gzip) · Model evaluation matrix (gzip) · Source table (gzip) · Release and checksums (gzip)

Coverage is derived from release 2026-10-10-6e93f504adfc. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.

Run instructions

No runnable recipe has been reviewed for this protocol. Dataset access, model requirements, licences and compute requirements must be checked against its sources before execution.

Strengths, limitations and unresolved questions

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

0 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-10-10-6e93f504adfc
Property and statementOriginal source and locationReview and provenance

No evidence rows match these filters. Choose another scope or clear the search.

Sources and history

Release 2026-10-10-6e93f504adfc · Record review: source checked

1 source records and release historyDownload this release (gzip)
Technical metadata and extraction receipts

Stable ID: structural-20261009-protocol-smorodina2026-vhh-iptm-dockq-calibration

areas
proteins-complexes
contexts
research
protocol
For each cognate complex, 50 samples are predicted and scored by DockQ against the deposited structure. Pearson correlation of ipTM with DockQ is printed for the best-DockQ sample and for the first sample (sample0). Quadrants at DockQ 0.23 and ipTM 0.5: Q2 is ipTM at least 0.5 with DockQ below 0.23 (confident failure), Q4 is ipTM below 0.5 with DockQ at least 0.23 (unconfident success). Change correlation: Pearson correlation between the change in DockQ and the change in ipTM under saturation sampling.
version
Smorodina et al. 2026 bioRxiv v1, Results P18 to P22, Figure 3
metric
pearson-correlation
metric direction
higher
unit
unitless
metric definition
Pearson correlation between ipTM and DockQ across the cognate complexes. Quadrant shares use DockQ 0.23 (CAPRI acceptable) and ipTM 0.5 (confident) as thresholds.
limitations
Calibration is computed on cognate complexes only; it says nothing about whether ipTM separates binders from non-binders.; Quadrant thresholds (DockQ 0.23, ipTM 0.5) are the source's choice, and the text does not say which sample (best, first or worst) the quadrant shares refer to.; Not every tool has every value printed in the text.; Nanobody (VHH)-antigen complexes only; results do not transfer to conventional antibodies or other complex classes.; The set mixes systems inside and outside each tool's training data (AF3 30 of 106 in training, Chai-1 25, Boltz-2 64); the printed values are over all systems, so they are not post-cutoff results, and Boltz-2's are the most exposed.; Preprint, not peer reviewed. Supplementary tables were not read.; No uncertainty is printed for these values.; The dataset paragraph (P62) says post-October 2021 depositions were kept, 'corresponding to the earliest training cutoff among the evaluated tools (Boltz-2)'. The methods (P76) give Chai-1 about 12 January 2021, AF3 about 30 September 2021 and Boltz-2 about 1 June 2023, so Chai-1 is the earliest and Boltz-2 the latest, and the October 2021 date matches AF3. The per-tool train and test counts follow the methods, and no printed value depends on the mis-stated sentence.; Preprint, not peer reviewed.
source locator
Results P18 to P22; Figure 3; Supplementary Table 3 (not read)
Related records

Suggest a correction