rewirebio.iobenchmarks
Result

148 true-positive-count

gupta2025-gp-llama-3-1-8b-schmidt2022-il2 true-positive-count

Tested configuration
GP over llama-3-1-8b embeddings (Gupta et al. 2025)
Protocol
Independent replication of 1-gene perturbation design: cumulative hits after 5 rounds of 128 genes
Dataset
Schmidt et al. 2022 screen, interleukin-2 production in primary human T cells
Procedure
tgtval-20261009-protocol-gupta2025-cumulative-hits-round5
Evaluation
GP (Llama-3.1-8B backbone) on IL2 (Gupta et al. 2025)
Coverage
Not reported scored / Not reported eligible
Uncertainty
Not reported by the source
Evidence
Independent external evaluation · source checkedLLMs for Bayesian Optimization in Scientific Domains: Are We There Yet? · Table 2, 'Llama-3.1-8B backbone' block, row 'GP', column 'IL2'

A source-checked result verifies the numerical transcription, not every model or protocol detail. Evaluation metadata: source checked. Source checked does not mean independently reproduced.

Reproduction

Split
No training split; 5 sequential selection rounds
Adaptation
Selection only
Scoring implementation
Not reported

No execution recipe has been verified for this exact configuration and evaluation. A benchmark's general instructions may use different inputs, splits or model settings.

Reproducing this published result requires matching its model configuration, data, split and scorer. Source checking or a successful smoke test does not establish score reproduction.

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

1 evidence row matching the loaded filters

Claims, original sources and review scope · Release 2026-10-10-6e93f504adfc
Property and statementOriginal source and locationReview and provenance
Reported result
147.8
Individual claims
LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet?

Original source ↗

Table 2, 'Llama-3.1-8B backbone' block, row 'GP', column 'IL2'

Version: Findings of the Association for Computational Linguistics: EMNLP 2025, pages 15482-15510
Retrieved: 2026-10-09T20:55:36Z

source checked

["source-hash-verification","pdf-text-parse","independent-cell-check"] · 2026-10-09T21:22:54Z

independent paper

Audit details

Extracted by deterministic parse of the PDF text layer (pdftotext -layout) of Tables 1 and 2, asserting the column header, the backbone section headings, the row labels and five values per row. Independent review 2026-10-09: value, metric, unit, direction, locator and configuration, protocol and dataset identity match the source.

Field: attributes.printed_value

Source artifact SHA-256: 7dda0b590f2736b7d30e48b97167a0868a2b4bde253589d8aff7929d01506ff9

Hash scope: pdftotext -layout text layer, parsed by extract/extract_target_validation.py

Inspected artifact

Extraction artifact SHA-256: 7dda0b590f2736b7d30e48b97167a0868a2b4bde253589d8aff7929d01506ff9

Extraction artifact

Sources and history

Release 2026-10-10-6e93f504adfc · Record review: source checked

1 source records and release historyDownload this release (gzip)
Technical metadata and extraction receipts

Stable ID: tgtval-20261009-result-gupta2025-gp-llama-3-1-8b-schmidt2022-il2-hits

metric
true-positive-count
metric qualifier
cumulative hits after round 5; mean of 5 runs
metric direction
higher
unit
count
printed value
147.8
numeric value
147.8
source locator
Table 2, 'Llama-3.1-8B backbone' block, row 'GP', column 'IL2'
review
method: source-hash-verification; pdf-text-parse; independent-cell-check; method note: Re-downloaded the PDF and matched its SHA-256. Converted with pdftotext -layout and parsed Tables 1 and 2 with a separate script written for this review, asserting the column header, the ground-truth row (identical in both tables), the backbone blocks and that Table 2's BDA rows equal Table 1's. Checked printed and numeric value, metric, unit, direction, locator, denominator, and the linked configuration, protocol and dataset.; reviewer: claude; reviewer note: Separate Claude review agent, independent of the extractor; no human review claimed; reviewed at: 2026-10-09T21:22:54Z; artifact sha256: 7dda0b590f2736b7d30e48b97167a0868a2b4bde253589d8aff7929d01506ff9; retrieval url: https://aclanthology.org/2025.findings-emnlp.838.pdf; note: Extracted by deterministic parse of the PDF text layer (pdftotext -layout) of Tables 1 and 2, asserting the column header, the backbone section headings, the row labels and five values per row. Independent review 2026-10-09: value, metric, unit, direction, locator and configuration, protocol and dataset identity match the source.
missing metadata
uncertainty: reason: unreported
denominator
654
denominator note
Printed ground-truth hit count for this screen: 654
source anomaly
Table 2 prints identical GP values for the Llama-3.1-8B and Qwen-2-7B blocks (147.8, 23, 22.2, 27.6, 30), although the embeddings differ between backbones.
Related records

Suggest a correction