rewirebio.iobenchmarks
Evaluation

Chai-1 on How well ipTM tracks DockQ on cognate nanobody-antigen complexes (Smorodina et al. 2026)

Published comparison; transcribed, not reproduced.

Research readiness

These checks assess whether the evidence supports a reproducible investigation. A source-checked score alone does not meet these requirements.

Release 2026-10-10-6e93f504adfc · Evidence verified: Not verified

Evidence incomplete

Replay metrics

Exact outcomes, predictions, identifiers and evaluator are connected.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • metric replay: verification is missing

Verified: Not verified

Evidence incomplete

Investigate discrepancies

Replay evidence includes annotations and an assessment of dependence. Unknown independence permits descriptive analysis only.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • metric replay: verification is missing
  • annotations: verification is missing
  • dependence: verification is missing

Verified: Not verified

Evidence incomplete

Run locally

A pinned recipe describes the inputs, environment and resource requirements.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • recipe pinned: verification is missing
  • resource estimate: verification is missing

Verified: Not verified

Evidence incomplete

Validate independently

Separate data and exposure records support an independent test.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • independent validation: verification is missing
  • overlap checked: verification is missing

Verified: Not verified

Readiness describes the evidence in this release. Availability on your computer is checked separately when an investigation runs. Existing data exposure can prevent independent validation even when files are available.

Artifacts and reproduction

No verified artifact manifest is connected to this record yet. The gaps above identify what is needed before analysis can begin.

Read reviewed discrepancy investigations

Evaluation results

1 evaluation · 3 results. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: Chai-1 0.6.1, 5 trunk x 10 diffusion samples, seed 42, ESM embeddings without MSAs (Smorodina et al. 2026)Protocol: How well ipTM tracks DockQ on cognate nanobody-antigen complexes
Dataset: Smorodina et al. 2026 nanobody-antigen benchmark: 106 cognate VHH-antigen complexes
0.612 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

Chai-1 on How well ipTM tracks DockQ on cognate nanobody-antigen complexes (Smorodina et al. 2026)

structural-20261009-protocol-smorodina2026-vhh-iptm-dockq-calibration

Aggregation: Not reported

Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores · Results P19, Chai-1: ipTM against DockQ for the best-DockQ sample of each complex
Configuration: Chai-1 0.6.1, 5 trunk x 10 diffusion samples, seed 42, ESM embeddings without MSAs (Smorodina et al. 2026)Protocol: How well ipTM tracks DockQ on cognate nanobody-antigen complexes
Dataset: Smorodina et al. 2026 nanobody-antigen benchmark: 106 cognate VHH-antigen complexes
-0.019 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

Chai-1 on How well ipTM tracks DockQ on cognate nanobody-antigen complexes (Smorodina et al. 2026)

structural-20261009-protocol-smorodina2026-vhh-iptm-dockq-calibration

Aggregation: Not reported

Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores · Results P21, Chai-1: change in ipTM against change in DockQ under saturation sampling
Configuration: Chai-1 0.6.1, 5 trunk x 10 diffusion samples, seed 42, ESM embeddings without MSAs (Smorodina et al. 2026)Protocol: How well ipTM tracks DockQ on cognate nanobody-antigen complexes
Dataset: Smorodina et al. 2026 nanobody-antigen benchmark: 106 cognate VHH-antigen complexes
24% proportion
percent · lower

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

Chai-1 on How well ipTM tracks DockQ on cognate nanobody-antigen complexes (Smorodina et al. 2026)

structural-20261009-protocol-smorodina2026-vhh-iptm-dockq-calibration

Aggregation: Not reported

Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores · Results P19, Chai-1: share of predictions with ipTM below 0.5 and DockQ at least 0.23 (unconfident successes, Q4)

Source checking is not independent reproduction. Release 2026-10-10-6e93f504adfc.

Evaluation procedure

structural-20261009-protocol-smorodina2026-vhh-iptm-dockq-calibration

Configuration
Chai-1 0.6.1, 5 trunk x 10 diffusion samples, seed 42, ESM embeddings without MSAs (Smorodina et al. 2026)
Protocol
How well ipTM tracks DockQ on cognate nanobody-antigen complexes
Dataset
Smorodina et al. 2026 nanobody-antigen benchmark: 106 cognate VHH-antigen complexes
origin
Independent external evaluation
configuration
Primary source as retrieved 2026-10-09
protocol id
structural-20261009-protocol-smorodina2026-vhh-iptm-dockq-calibration
dataset version
bioRxiv v1 benchmark
split
Mixed: systems inside and outside the tool's training data
inputs
VHH and antigen sequences
adaptation
None
population
106 cognate complexes
metric implementation
Pearson correlation; DockQ v2.1.3
aggregation
Across complexes
budget
50 samples per complex

Metadata review: source checked. Unreported conditions prevent automatic comparisons.

Reproduction

Split
Mixed: systems inside and outside the tool's training data
Adaptation
None
Scoring implementation
Pearson correlation; DockQ v2.1.3

No execution recipe has been verified for this exact configuration and evaluation. A benchmark's general instructions may use different inputs, splits or model settings.

Reproducing this published result requires matching its model configuration, data, split and scorer. Source checking or a successful smoke test does not establish score reproduction.

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

19 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-10-10-6e93f504adfc
Property and statementOriginal source and locationReview and provenance
attributes.comparison.adaptation
None
Context-only references
Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores

Original source ↗

Results P19 and P21, values for 'Chai-1'

Version: bioRxiv version 1, posted 2026-03-03; Europe PMC preprint full text PPR1221387 (manuscript EMS215481); not peer reviewed
Retrieved: 2026-10-09T21:25:48Z

not individually reviewed

No individual claim review recorded

independent paper

Audit details

Field: attributes.comparison.adaptation

Source artifact SHA-256: 0ad24054d98fd1888fa8bf3120f76e22f7bf6b1965261b16f045986d72c50732

Hash scope: JATS XML parse (xml.etree), by extract/extract_structural.py

Inspected artifact

attributes.comparison.aggregation
Across complexes
Context-only references
Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores

Original source ↗

Results P19 and P21, values for 'Chai-1'

Version: bioRxiv version 1, posted 2026-03-03; Europe PMC preprint full text PPR1221387 (manuscript EMS215481); not peer reviewed
Retrieved: 2026-10-09T21:25:48Z

not individually reviewed

No individual claim review recorded

independent paper

Audit details

Field: attributes.comparison.aggregation

Source artifact SHA-256: 0ad24054d98fd1888fa8bf3120f76e22f7bf6b1965261b16f045986d72c50732

Hash scope: JATS XML parse (xml.etree), by extract/extract_structural.py

Inspected artifact

attributes.comparison.budget
50 samples per complex
Context-only references
Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores

Original source ↗

Results P19 and P21, values for 'Chai-1'

Version: bioRxiv version 1, posted 2026-03-03; Europe PMC preprint full text PPR1221387 (manuscript EMS215481); not peer reviewed
Retrieved: 2026-10-09T21:25:48Z

not individually reviewed

No individual claim review recorded

independent paper

Audit details

Field: attributes.comparison.budget

Source artifact SHA-256: 0ad24054d98fd1888fa8bf3120f76e22f7bf6b1965261b16f045986d72c50732

Hash scope: JATS XML parse (xml.etree), by extract/extract_structural.py

Inspected artifact

attributes.comparison.dataset_version
bioRxiv v1 benchmark
Context-only references
Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores

Original source ↗

Results P19 and P21, values for 'Chai-1'

Version: bioRxiv version 1, posted 2026-03-03; Europe PMC preprint full text PPR1221387 (manuscript EMS215481); not peer reviewed
Retrieved: 2026-10-09T21:25:48Z

not individually reviewed

No individual claim review recorded

independent paper

Audit details

Field: attributes.comparison.dataset_version

Source artifact SHA-256: 0ad24054d98fd1888fa8bf3120f76e22f7bf6b1965261b16f045986d72c50732

Hash scope: JATS XML parse (xml.etree), by extract/extract_structural.py

Inspected artifact

attributes.comparison.inputs
VHH and antigen sequences
Context-only references
Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores

Original source ↗

Results P19 and P21, values for 'Chai-1'

Version: bioRxiv version 1, posted 2026-03-03; Europe PMC preprint full text PPR1221387 (manuscript EMS215481); not peer reviewed
Retrieved: 2026-10-09T21:25:48Z

not individually reviewed

No individual claim review recorded

independent paper

Audit details

Field: attributes.comparison.inputs

Source artifact SHA-256: 0ad24054d98fd1888fa8bf3120f76e22f7bf6b1965261b16f045986d72c50732

Hash scope: JATS XML parse (xml.etree), by extract/extract_structural.py

Inspected artifact

attributes.comparison.metric_implementation
Pearson correlation; DockQ v2.1.3
Context-only references
Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores

Original source ↗

Results P19 and P21, values for 'Chai-1'

Version: bioRxiv version 1, posted 2026-03-03; Europe PMC preprint full text PPR1221387 (manuscript EMS215481); not peer reviewed
Retrieved: 2026-10-09T21:25:48Z

not individually reviewed

No individual claim review recorded

independent paper

Audit details

Field: attributes.comparison.metric_implementation

Source artifact SHA-256: 0ad24054d98fd1888fa8bf3120f76e22f7bf6b1965261b16f045986d72c50732

Hash scope: JATS XML parse (xml.etree), by extract/extract_structural.py

Inspected artifact

attributes.comparison.population
106 cognate complexes
Context-only references
Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores

Original source ↗

Results P19 and P21, values for 'Chai-1'

Version: bioRxiv version 1, posted 2026-03-03; Europe PMC preprint full text PPR1221387 (manuscript EMS215481); not peer reviewed
Retrieved: 2026-10-09T21:25:48Z

not individually reviewed

No individual claim review recorded

independent paper

Audit details

Field: attributes.comparison.population

Source artifact SHA-256: 0ad24054d98fd1888fa8bf3120f76e22f7bf6b1965261b16f045986d72c50732

Hash scope: JATS XML parse (xml.etree), by extract/extract_structural.py

Inspected artifact

attributes.comparison.protocol_id
structural-20261009-protocol-smorodina2026-vhh-iptm-dockq-calibration
Context-only references
Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores

Original source ↗

Results P19 and P21, values for 'Chai-1'

Version: bioRxiv version 1, posted 2026-03-03; Europe PMC preprint full text PPR1221387 (manuscript EMS215481); not peer reviewed
Retrieved: 2026-10-09T21:25:48Z

not individually reviewed

No individual claim review recorded

independent paper

Audit details

Field: attributes.comparison.protocol_id

Source artifact SHA-256: 0ad24054d98fd1888fa8bf3120f76e22f7bf6b1965261b16f045986d72c50732

Hash scope: JATS XML parse (xml.etree), by extract/extract_structural.py

Inspected artifact

attributes.comparison.split
Mixed: systems inside and outside the tool's training data
Context-only references
Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores

Original source ↗

Results P19 and P21, values for 'Chai-1'

Version: bioRxiv version 1, posted 2026-03-03; Europe PMC preprint full text PPR1221387 (manuscript EMS215481); not peer reviewed
Retrieved: 2026-10-09T21:25:48Z

not individually reviewed

No individual claim review recorded

independent paper

Audit details

Field: attributes.comparison.split

Source artifact SHA-256: 0ad24054d98fd1888fa8bf3120f76e22f7bf6b1965261b16f045986d72c50732

Hash scope: JATS XML parse (xml.etree), by extract/extract_structural.py

Inspected artifact

attributes.limitations
2 values
  • Run by authors who did not develop the tool.
  • Not printed in the text for this tool: ipTM against DockQ for the first sample (sample0) of each complex; share of predictions with ipTM at least 0.5 and DockQ below 0.23 (confident failures, Q2).
Context-only references
Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores

Original source ↗

Results P19 and P21, values for 'Chai-1'

Version: bioRxiv version 1, posted 2026-03-03; Europe PMC preprint full text PPR1221387 (manuscript EMS215481); not peer reviewed
Retrieved: 2026-10-09T21:25:48Z

not individually reviewed

No individual claim review recorded

independent paper

Audit details

Field: attributes.limitations

Source artifact SHA-256: 0ad24054d98fd1888fa8bf3120f76e22f7bf6b1965261b16f045986d72c50732

Hash scope: JATS XML parse (xml.etree), by extract/extract_structural.py

Inspected artifact

Sources and history

Release 2026-10-10-6e93f504adfc · Record review: source checked

1 source records and release historyDownload this release (gzip)
Technical metadata and extraction receipts

Stable ID: structural-20261009-eval-smorodina2026-vhh-iptm-dockq-calibration-chai1

areas
proteins-complexes
contexts
research
origin
independent_paper
protocol
structural-20261009-protocol-smorodina2026-vhh-iptm-dockq-calibration
version
Primary source as retrieved 2026-10-09
comparison
protocol id: structural-20261009-protocol-smorodina2026-vhh-iptm-dockq-calibration; dataset version: bioRxiv v1 benchmark; split: Mixed: systems inside and outside the tool's training data; inputs: VHH and antigen sequences; adaptation: None; population: 106 cognate complexes; metric implementation: Pearson correlation; DockQ v2.1.3; aggregation: Across complexes; budget: 50 samples per complex
source locator
Results P19 and P21, values for 'Chai-1'
limitations
Run by authors who did not develop the tool.; Not printed in the text for this tool: ipTM against DockQ for the first sample (sample0) of each complex; share of predictions with ipTM at least 0.5 and DockQ below 0.23 (confident failures, Q2).
Related records

Suggest a correction