rewirebio.iobenchmarks
Evidence claim

control_result: tgtval-20261009-protocol-gupta2025-cumulative-hits-round5

Descriptive fact transcribed from the pinned source.

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

6 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-10-10-6e93f504adfc
Property and statementOriginal source and locationReview and provenance
attributes.field
control_result
Context-only references
LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet?

Original source ↗

Abstract; section 5; Table 1 and its caption

Version: Findings of the Association for Computational Linguistics: EMNLP 2025, pages 15482-15510
Retrieved: 2026-10-09T20:55:36Z

not individually reviewed

No individual claim review recorded

Audit details

Field: attributes.field

Source artifact SHA-256: 7dda0b590f2736b7d30e48b97167a0868a2b4bde253589d8aff7929d01506ff9

Hash scope: pdftotext -layout text layer, parsed by extract/extract_target_validation.py

Inspected artifact

attributes.source_locator
Abstract; section 5; Table 1 and its caption
Context-only references
LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet?

Original source ↗

Abstract; section 5; Table 1 and its caption

Version: Findings of the Association for Computational Linguistics: EMNLP 2025, pages 15482-15510
Retrieved: 2026-10-09T20:55:36Z

not individually reviewed

No individual claim review recorded

Audit details

Field: attributes.source_locator

Source artifact SHA-256: 7dda0b590f2736b7d30e48b97167a0868a2b4bde253589d8aff7929d01506ff9

Hash scope: pdftotext -layout text layer, parsed by extract/extract_target_validation.py

Inspected artifact

attributes.value
The source states that replacing true outcomes with randomly permuted labels has no impact on performance, and concludes that the language models tested fail to perform in-context experimental design. It performs two levels of randomisation: random measurement values, and random hit or not-hit feedback. In Table 1 the permuted variant scores at least as high as the agent on three of the five screens with each of the three backbones: Llama-3.1-8B, Qwen-2-7B, and Claude 3.5 Sonnet compared with the replicated row.
Context-only references
LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet?

Original source ↗

Abstract; section 5; Table 1 and its caption

Version: Findings of the Association for Computational Linguistics: EMNLP 2025, pages 15482-15510
Retrieved: 2026-10-09T20:55:36Z

not individually reviewed

No individual claim review recorded

Audit details

Field: attributes.value

Source artifact SHA-256: 7dda0b590f2736b7d30e48b97167a0868a2b4bde253589d8aff7929d01506ff9

Hash scope: pdftotext -layout text layer, parsed by extract/extract_target_validation.py

Inspected artifact

description
Descriptive fact transcribed from the pinned source.
Context-only references
LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet?

Original source ↗

Abstract; section 5; Table 1 and its caption

Version: Findings of the Association for Computational Linguistics: EMNLP 2025, pages 15482-15510
Retrieved: 2026-10-09T20:55:36Z

not individually reviewed

No individual claim review recorded

Audit details

Field: description

Source artifact SHA-256: 7dda0b590f2736b7d30e48b97167a0868a2b4bde253589d8aff7929d01506ff9

Hash scope: pdftotext -layout text layer, parsed by extract/extract_target_validation.py

Inspected artifact

Relationship: subject
tgtval-20261009-protocol-gupta2025-cumulative-hits-round5
Context-only references
LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet?

Original source ↗

Abstract; section 5; Table 1 and its caption

Version: Findings of the Association for Computational Linguistics: EMNLP 2025, pages 15482-15510
Retrieved: 2026-10-09T20:55:36Z

not individually reviewed

No individual claim review recorded

Audit details

Field: links:subject:tgtval-20261009-protocol-gupta2025-cumulative-hits-round5

Source artifact SHA-256: 7dda0b590f2736b7d30e48b97167a0868a2b4bde253589d8aff7929d01506ff9

Hash scope: pdftotext -layout text layer, parsed by extract/extract_target_validation.py

Inspected artifact

name
control_result: tgtval-20261009-protocol-gupta2025-cumulative-hits-round5
Context-only references
LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet?

Original source ↗

Abstract; section 5; Table 1 and its caption

Version: Findings of the Association for Computational Linguistics: EMNLP 2025, pages 15482-15510
Retrieved: 2026-10-09T20:55:36Z

not individually reviewed

No individual claim review recorded

Audit details

Field: name

Source artifact SHA-256: 7dda0b590f2736b7d30e48b97167a0868a2b4bde253589d8aff7929d01506ff9

Hash scope: pdftotext -layout text layer, parsed by extract/extract_target_validation.py

Inspected artifact

Sources and history

Release 2026-10-10-6e93f504adfc · Record review: source checked

1 source records and release historyDownload this release (gzip)
Technical metadata and extraction receipts

Stable ID: tgtval-20261009-claim-gupta2025-permuted-feedback

field
control_result
value
The source states that replacing true outcomes with randomly permuted labels has no impact on performance, and concludes that the language models tested fail to perform in-context experimental design. It performs two levels of randomisation: random measurement values, and random hit or not-hit feedback. In Table 1 the permuted variant scores at least as high as the agent on three of the five screens with each of the three backbones: Llama-3.1-8B, Qwen-2-7B, and Claude 3.5 Sonnet compared with the replicated row.
source locator
Abstract; section 5; Table 1 and its caption
review
method: source-hash-verification; ai-assisted-source-review; reviewer: claude; reviewer note: Separate Claude review agent, independent of the extractor; no human review claimed; reviewed at: 2026-10-09T21:22:54Z; artifact sha256: 7dda0b590f2736b7d30e48b97167a0868a2b4bde253589d8aff7929d01506ff9; retrieval url: https://aclanthology.org/2025.findings-emnlp.838.pdf; note: Hand transcription from the source text. Independent review 2026-10-09: checked against the re-downloaded PDF text; see docs/reviews/use-cases/therapeutic-target-validation-2026-10-09.md.
Related records

Suggest a correction