control_result: tgtval-20261009-protocol-gupta2025-cumulative-hits-round5
Descriptive fact transcribed from the pinned source.
Evidence
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Evidence table
Inspect claims, sources and review details
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
6 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| attributes.field control_result Context-only references | LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet? Abstract; section 5; Table 1 and its caption Version: Findings of the Association for Computational Linguistics: EMNLP 2025, pages 15482-15510 | not individually reviewed No individual claim review recorded Audit detailsField: Source artifact SHA-256: Hash scope: pdftotext -layout text layer, parsed by extract/extract_target_validation.py |
| attributes.source_locator Abstract; section 5; Table 1 and its caption Context-only references | LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet? Abstract; section 5; Table 1 and its caption Version: Findings of the Association for Computational Linguistics: EMNLP 2025, pages 15482-15510 | not individually reviewed No individual claim review recorded Audit detailsField: Source artifact SHA-256: Hash scope: pdftotext -layout text layer, parsed by extract/extract_target_validation.py |
| attributes.value The source states that replacing true outcomes with randomly permuted labels has no impact on performance, and concludes that the language models tested fail to perform in-context experimental design. It performs two levels of randomisation: random measurement values, and random hit or not-hit feedback. In Table 1 the permuted variant scores at least as high as the agent on three of the five screens with each of the three backbones: Llama-3.1-8B, Qwen-2-7B, and Claude 3.5 Sonnet compared with the replicated row. Context-only references | LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet? Abstract; section 5; Table 1 and its caption Version: Findings of the Association for Computational Linguistics: EMNLP 2025, pages 15482-15510 | not individually reviewed No individual claim review recorded Audit detailsField: Source artifact SHA-256: Hash scope: pdftotext -layout text layer, parsed by extract/extract_target_validation.py |
| description Descriptive fact transcribed from the pinned source. Context-only references | LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet? Abstract; section 5; Table 1 and its caption Version: Findings of the Association for Computational Linguistics: EMNLP 2025, pages 15482-15510 | not individually reviewed No individual claim review recorded Audit detailsField: Source artifact SHA-256: Hash scope: pdftotext -layout text layer, parsed by extract/extract_target_validation.py |
| Relationship: subject tgtval-20261009-protocol-gupta2025-cumulative-hits-round5 Context-only references | LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet? Abstract; section 5; Table 1 and its caption Version: Findings of the Association for Computational Linguistics: EMNLP 2025, pages 15482-15510 | not individually reviewed No individual claim review recorded Audit detailsField: Source artifact SHA-256: Hash scope: pdftotext -layout text layer, parsed by extract/extract_target_validation.py |
| name control_result: tgtval-20261009-protocol-gupta2025-cumulative-hits-round5 Context-only references | LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet? Abstract; section 5; Table 1 and its caption Version: Findings of the Association for Computational Linguistics: EMNLP 2025, pages 15482-15510 | not individually reviewed No individual claim review recorded Audit detailsField: Source artifact SHA-256: Hash scope: pdftotext -layout text layer, parsed by extract/extract_target_validation.py |
Sources and history
Release 2026-10-10-6e93f504adfc · Record review: source checked
1 source records and release history
- LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet? · Original source · Findings of the Association for Computational Linguistics: EMNLP 2025, pages 15482-15510
Technical metadata and extraction receipts
Stable ID: tgtval-20261009-claim-gupta2025-permuted-feedback
- field
- control_result
- value
- The source states that replacing true outcomes with randomly permuted labels has no impact on performance, and concludes that the language models tested fail to perform in-context experimental design. It performs two levels of randomisation: random measurement values, and random hit or not-hit feedback. In Table 1 the permuted variant scores at least as high as the agent on three of the five screens with each of the three backbones: Llama-3.1-8B, Qwen-2-7B, and Claude 3.5 Sonnet compared with the replicated row.
- source locator
- Abstract; section 5; Table 1 and its caption
- review
- method: source-hash-verification; ai-assisted-source-review; reviewer: claude; reviewer note: Separate Claude review agent, independent of the extractor; no human review claimed; reviewed at: 2026-10-09T21:22:54Z; artifact sha256: 7dda0b590f2736b7d30e48b97167a0868a2b4bde253589d8aff7929d01506ff9; retrieval url: https://aclanthology.org/2025.findings-emnlp.838.pdf; note: Hand transcription from the source text. Independent review 2026-10-09: checked against the re-downloaded PDF text; see docs/reviews/use-cases/therapeutic-target-validation-2026-10-09.md.