rewirebio.iobenchmarks
Dataset

Schmidt et al. 2022 screen, interferon-gamma production in primary human T cells

Genome-wide perturbation screen measuring change in interferon-gamma (IFNG) production.

Research readiness

These checks assess whether the evidence supports a reproducible investigation. A source-checked score alone does not meet these requirements.

Release 2026-10-10-7fcc3e48a123 · Evidence verified: Not verified

Evidence incomplete

Replay metrics

Exact outcomes, predictions, identifiers and evaluator are connected.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • metric replay: verification is missing

Verified: Not verified

Evidence incomplete

Investigate discrepancies

Replay evidence includes annotations and an assessment of dependence. Unknown independence permits descriptive analysis only.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • metric replay: verification is missing
  • annotations: verification is missing
  • dependence: verification is missing

Verified: Not verified

Evidence incomplete

Run locally

A pinned recipe describes the inputs, environment and resource requirements.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • recipe pinned: verification is missing
  • resource estimate: verification is missing

Verified: Not verified

Evidence incomplete

Validate independently

Separate data and exposure records support an independent test.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • independent validation: verification is missing
  • overlap checked: verification is missing

Verified: Not verified

Readiness describes the evidence in this release. Availability on your computer is checked separately when an investigation runs. Existing data exposure can prevent independent validation even when files are available.

Artifacts and reproduction

No verified artifact manifest is connected to this record yet. The gaps above identify what is needed before analysis can begin.

Read reviewed discrepancy investigations

Evaluation results

31 evaluations · 51 results. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: BDA, llama-3-1-8b backbone (Gupta et al. 2025)Protocol: Independent replication of 1-gene perturbation design: cumulative hits after 5 rounds of 128 genes
Dataset: Schmidt et al. 2022 screen, interferon-gamma production in primary human T cells
44 true-positive-count
count · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

BDA (Llama-3.1-8B backbone) on IFNG (Gupta et al. 2025)

tgtval-20261009-protocol-gupta2025-cumulative-hits-round5

Aggregation: Not reported

LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet? · Table 1, 'Llama-3.1-8B backbone' block, row 'BDA', column 'IFNG'
Configuration: BDA, qwen-2-7b backbone (Gupta et al. 2025)Protocol: Independent replication of 1-gene perturbation design: cumulative hits after 5 rounds of 128 genes
Dataset: Schmidt et al. 2022 screen, interferon-gamma production in primary human T cells
26.2 true-positive-count
count · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

BDA (Qwen-2-7B backbone) on IFNG (Gupta et al. 2025)

tgtval-20261009-protocol-gupta2025-cumulative-hits-round5

Aggregation: Not reported

LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet? · Table 1, 'Qwen-2-7B backbone' block, row 'BDA', column 'IFNG'
Configuration: BDA-Rand, claude-3-5-sonnet backbone (Gupta et al. 2025)Protocol: Independent replication of 1-gene perturbation design: cumulative hits after 5 rounds of 128 genes
Dataset: Schmidt et al. 2022 screen, interferon-gamma production in primary human T cells
79.4 true-positive-count
count · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

BDA-Rand (Claude 3.5 Sonnet backbone) on IFNG (Gupta et al. 2025)

tgtval-20261009-protocol-gupta2025-cumulative-hits-round5

Aggregation: Not reported

LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet? · Table 1, 'Claude 3.5 Sonnet backbone' block, row 'BDA-Rand', column 'IFNG'
Configuration: BDA-Rand, llama-3-1-8b backbone (Gupta et al. 2025)Protocol: Independent replication of 1-gene perturbation design: cumulative hits after 5 rounds of 128 genes
Dataset: Schmidt et al. 2022 screen, interferon-gamma production in primary human T cells
51 true-positive-count
count · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

BDA-Rand (Llama-3.1-8B backbone) on IFNG (Gupta et al. 2025)

tgtval-20261009-protocol-gupta2025-cumulative-hits-round5

Aggregation: Not reported

LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet? · Table 1, 'Llama-3.1-8B backbone' block, row 'BDA-Rand', column 'IFNG'
Configuration: BDA-Rand, qwen-2-7b backbone (Gupta et al. 2025)Protocol: Independent replication of 1-gene perturbation design: cumulative hits after 5 rounds of 128 genes
Dataset: Schmidt et al. 2022 screen, interferon-gamma production in primary human T cells
32.4 true-positive-count
count · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

BDA-Rand (Qwen-2-7B backbone) on IFNG (Gupta et al. 2025)

tgtval-20261009-protocol-gupta2025-cumulative-hits-round5

Aggregation: Not reported

LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet? · Table 1, 'Qwen-2-7B backbone' block, row 'BDA-Rand', column 'IFNG'
Configuration: BDA (Replicated), claude-3-5-sonnet backbone (Gupta et al. 2025)Protocol: Independent replication of 1-gene perturbation design: cumulative hits after 5 rounds of 128 genes
Dataset: Schmidt et al. 2022 screen, interferon-gamma production in primary human T cells
78.8 true-positive-count
count · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

BDA (Replicated) (Claude 3.5 Sonnet backbone) on IFNG (Gupta et al. 2025)

tgtval-20261009-protocol-gupta2025-cumulative-hits-round5

Aggregation: Not reported

LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet? · Table 1, 'Claude 3.5 Sonnet backbone' block, row 'BDA (Replicated)', column 'IFNG'
Configuration: BDA (Reported Numbers), claude-3-5-sonnet backbone (Gupta et al. 2025)Protocol: Independent replication of 1-gene perturbation design: cumulative hits after 5 rounds of 128 genes
Dataset: Schmidt et al. 2022 screen, interferon-gamma production in primary human T cells
87.4 true-positive-count
count · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Result quoted from another source · Source checked
Methods, coverage and source

BDA (Reported Numbers) (Claude 3.5 Sonnet backbone) on IFNG (Gupta et al. 2025)

tgtval-20261009-protocol-gupta2025-cumulative-hits-round5

Aggregation: Not reported

LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet? · Table 1, 'Claude 3.5 Sonnet backbone' block, row 'BDA (Reported Numbers)', column 'IFNG'
Configuration: GP over llama-3-1-8b embeddings (Gupta et al. 2025)Protocol: Independent replication of 1-gene perturbation design: cumulative hits after 5 rounds of 128 genes
Dataset: Schmidt et al. 2022 screen, interferon-gamma production in primary human T cells
23 true-positive-count
count · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

GP (Llama-3.1-8B backbone) on IFNG (Gupta et al. 2025)

tgtval-20261009-protocol-gupta2025-cumulative-hits-round5

Aggregation: Not reported

LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet? · Table 2, 'Llama-3.1-8B backbone' block, row 'GP', column 'IFNG'
Configuration: GP over qwen-2-7b embeddings (Gupta et al. 2025)Protocol: Independent replication of 1-gene perturbation design: cumulative hits after 5 rounds of 128 genes
Dataset: Schmidt et al. 2022 screen, interferon-gamma production in primary human T cells
23 true-positive-count
count · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

GP (Qwen-2-7B backbone) on IFNG (Gupta et al. 2025)

tgtval-20261009-protocol-gupta2025-cumulative-hits-round5

Aggregation: Not reported

LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet? · Table 2, 'Qwen-2-7B backbone' block, row 'GP', column 'IFNG'
Configuration: Linear UCB over llama-3-1-8b embeddings (Gupta et al. 2025)Protocol: Independent replication of 1-gene perturbation design: cumulative hits after 5 rounds of 128 genes
Dataset: Schmidt et al. 2022 screen, interferon-gamma production in primary human T cells
72 true-positive-count
count · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

Linear UCB (Llama-3.1-8B backbone) on IFNG (Gupta et al. 2025)

tgtval-20261009-protocol-gupta2025-cumulative-hits-round5

Aggregation: Not reported

LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet? · Table 2, 'Llama-3.1-8B backbone' block, row 'Linear UCB', column 'IFNG'
Configuration: Linear UCB over qwen-2-7b embeddings (Gupta et al. 2025)Protocol: Independent replication of 1-gene perturbation design: cumulative hits after 5 rounds of 128 genes
Dataset: Schmidt et al. 2022 screen, interferon-gamma production in primary human T cells
74 true-positive-count
count · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

Linear UCB (Qwen-2-7B backbone) on IFNG (Gupta et al. 2025)

tgtval-20261009-protocol-gupta2025-cumulative-hits-round5

Aggregation: Not reported

LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet? · Table 2, 'Qwen-2-7B backbone' block, row 'Linear UCB', column 'IFNG'
Configuration: Badge acquisition function on an MLP surrogate (Roohani et al. 2025)Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes
Dataset: Schmidt et al. 2022 screen, interferon-gamma production in primary human T cells
0.06 recall
fraction · higher

Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass.

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

Badge on Schmidt1 (Roohani et al. 2025)

tgtval-20261009-protocol-roohani2025-hitratio-round5

Aggregation: Not reported

BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Badge', column 'All' under 'Schmidt1'
Configuration: Badge acquisition function on an MLP surrogate (Roohani et al. 2025)Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes
Dataset: Schmidt et al. 2022 screen, interferon-gamma production in primary human T cells
0.05 recall
fraction · higher

Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass.

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

Badge on Schmidt1 (Roohani et al. 2025)

tgtval-20261009-protocol-roohani2025-hitratio-round5

Aggregation: Not reported

BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Badge', column 'N/E' under 'Schmidt1'
Configuration: BioDiscoveryAgent (No-Tools), Claude 3.5 Sonnet (Roohani et al. 2025)Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes
Dataset: Schmidt et al. 2022 screen, interferon-gamma production in primary human T cells
0.095 recall
fraction · higher

Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass.

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

Claude 3.5 Sonnet on Schmidt1 (Roohani et al. 2025)

tgtval-20261009-protocol-roohani2025-hitratio-round5

Aggregation: Not reported

BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Claude 3.5 Sonnet', column 'All' under 'Schmidt1'
Configuration: BioDiscoveryAgent (No-Tools), Claude 3.5 Sonnet (Roohani et al. 2025)Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes
Dataset: Schmidt et al. 2022 screen, interferon-gamma production in primary human T cells
0.107 recall
fraction · higher

Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass.

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

Claude 3.5 Sonnet on Schmidt1 (Roohani et al. 2025)

tgtval-20261009-protocol-roohani2025-hitratio-round5

Aggregation: Not reported

BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Claude 3.5 Sonnet', column 'N/E' under 'Schmidt1'
Configuration: BioDiscoveryAgent (No-Tools), Claude 3 Haiku (Roohani et al. 2025)Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes
Dataset: Schmidt et al. 2022 screen, interferon-gamma production in primary human T cells
0.064 recall
fraction · higher

Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass.

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

Claude 3 Haiku on Schmidt1 (Roohani et al. 2025)

tgtval-20261009-protocol-roohani2025-hitratio-round5

Aggregation: Not reported

BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Claude 3 Haiku', column 'All' under 'Schmidt1'
Configuration: BioDiscoveryAgent (No-Tools), Claude 3 Haiku (Roohani et al. 2025)Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes
Dataset: Schmidt et al. 2022 screen, interferon-gamma production in primary human T cells
0.072 recall
fraction · higher

Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass.

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

Claude 3 Haiku on Schmidt1 (Roohani et al. 2025)

tgtval-20261009-protocol-roohani2025-hitratio-round5

Aggregation: Not reported

BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Claude 3 Haiku', column 'N/E' under 'Schmidt1'
Configuration: BioDiscoveryAgent (No-Tools), Claude 3 Opus (Roohani et al. 2025)Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes
Dataset: Schmidt et al. 2022 screen, interferon-gamma production in primary human T cells
0.094 recall
fraction · higher

Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass.

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

Claude 3 Opus on Schmidt1 (Roohani et al. 2025)

tgtval-20261009-protocol-roohani2025-hitratio-round5

Aggregation: Not reported

BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Claude 3 Opus', column 'All' under 'Schmidt1'
Configuration: BioDiscoveryAgent (No-Tools), Claude 3 Opus (Roohani et al. 2025)Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes
Dataset: Schmidt et al. 2022 screen, interferon-gamma production in primary human T cells
0.106 recall
fraction · higher

Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass.

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

Claude 3 Opus on Schmidt1 (Roohani et al. 2025)

tgtval-20261009-protocol-roohani2025-hitratio-round5

Aggregation: Not reported

BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Claude 3 Opus', column 'N/E' under 'Schmidt1'
Configuration: BioDiscoveryAgent (No-Tools), Claude 3 Sonnet (Roohani et al. 2025)Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes
Dataset: Schmidt et al. 2022 screen, interferon-gamma production in primary human T cells
0.076 recall
fraction · higher

Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass.

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

Claude 3 Sonnet on Schmidt1 (Roohani et al. 2025)

tgtval-20261009-protocol-roohani2025-hitratio-round5

Aggregation: Not reported

BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Claude 3 Sonnet', column 'All' under 'Schmidt1'
Configuration: BioDiscoveryAgent (No-Tools), Claude 3 Sonnet (Roohani et al. 2025)Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes
Dataset: Schmidt et al. 2022 screen, interferon-gamma production in primary human T cells
0.082 recall
fraction · higher

Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass.

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

Claude 3 Sonnet on Schmidt1 (Roohani et al. 2025)

tgtval-20261009-protocol-roohani2025-hitratio-round5

Aggregation: Not reported

BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Claude 3 Sonnet', column 'N/E' under 'Schmidt1'
Configuration: BioDiscoveryAgent (No-Tools), Claude v1 (Roohani et al. 2025)Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes
Dataset: Schmidt et al. 2022 screen, interferon-gamma production in primary human T cells
0.067 recall
fraction · higher

Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass.

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

Claude v1 on Schmidt1 (Roohani et al. 2025)

tgtval-20261009-protocol-roohani2025-hitratio-round5

Aggregation: Not reported

BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Claude v1', column 'All' under 'Schmidt1'
Configuration: BioDiscoveryAgent (No-Tools), Claude v1 (Roohani et al. 2025)Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes
Dataset: Schmidt et al. 2022 screen, interferon-gamma production in primary human T cells
0.086 recall
fraction · higher

Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass.

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

Claude v1 on Schmidt1 (Roohani et al. 2025)

tgtval-20261009-protocol-roohani2025-hitratio-round5

Aggregation: Not reported

BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Claude v1', column 'N/E' under 'Schmidt1'
Configuration: Coreset acquisition function on an MLP surrogate (Roohani et al. 2025)Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes
Dataset: Schmidt et al. 2022 screen, interferon-gamma production in primary human T cells
0.072 recall
fraction · higher

Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass.

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

Coreset on Schmidt1 (Roohani et al. 2025)

tgtval-20261009-protocol-roohani2025-hitratio-round5

Aggregation: Not reported

BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Coreset', column 'All' under 'Schmidt1'
Configuration: Coreset acquisition function on an MLP surrogate (Roohani et al. 2025)Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes
Dataset: Schmidt et al. 2022 screen, interferon-gamma production in primary human T cells
0.066 recall
fraction · higher

Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass.

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

Coreset on Schmidt1 (Roohani et al. 2025)

tgtval-20261009-protocol-roohani2025-hitratio-round5

Aggregation: Not reported

BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Coreset', column 'N/E' under 'Schmidt1'

Source checking is not independent reproduction. Release 2026-10-10-7fcc3e48a123.

Dataset and evaluation context

A dataset supplies biological observations. The evaluation protocol defines how those observations are split, used and scored.

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

14 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-10-10-7fcc3e48a123
Property and statementOriginal source and locationReview and provenance
attributes.assay
Pooled CRISPR perturbation screen with a phenotypic readout
Context-only references
LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet?

Original source ↗

Roohani et al. 2025 section 4.1; Gupta et al. 2025 appendix B.1.1

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: Findings of the Association for Computational Linguistics: EMNLP 2025, pages 15482-15510
Retrieved: 2026-10-09T20:55:36Z

not individually reviewed

No individual claim review recorded

Audit details

Field: attributes.assay

Source artifact SHA-256: 7dda0b590f2736b7d30e48b97167a0868a2b4bde253589d8aff7929d01506ff9

Hash scope: pdftotext -layout text layer, parsed by extract/extract_target_validation.py

Inspected artifact

attributes.assay
Pooled CRISPR perturbation screen with a phenotypic readout
Context-only references
BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments

Original source ↗

Roohani et al. 2025 section 4.1; Gupta et al. 2025 appendix B.1.1

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: arXiv:2405.17631 version 3, updated 2025-03-09; published as a conference paper at ICLR 2025
Retrieved: 2026-10-09T20:51:27Z

not individually reviewed

No individual claim review recorded

Audit details

Field: attributes.assay

Source artifact SHA-256: dd94d80ec75c9bb84ec3989c910a313ef1a7aec57c772d3dd35ecadd250132a7

Hash scope: pdftotext -layout text layer, parsed by extract/extract_target_validation.py

Inspected artifact

attributes.population
Over 18,000 genes, each knocked down in a distinct cell (Roohani et al. 2025 section 4.1)
Context-only references
LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet?

Original source ↗

Roohani et al. 2025 section 4.1; Gupta et al. 2025 appendix B.1.1

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: Findings of the Association for Computational Linguistics: EMNLP 2025, pages 15482-15510
Retrieved: 2026-10-09T20:55:36Z

not individually reviewed

No individual claim review recorded

Audit details

Field: attributes.population

Source artifact SHA-256: 7dda0b590f2736b7d30e48b97167a0868a2b4bde253589d8aff7929d01506ff9

Hash scope: pdftotext -layout text layer, parsed by extract/extract_target_validation.py

Inspected artifact

attributes.population
Over 18,000 genes, each knocked down in a distinct cell (Roohani et al. 2025 section 4.1)
Context-only references
BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments

Original source ↗

Roohani et al. 2025 section 4.1; Gupta et al. 2025 appendix B.1.1

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: arXiv:2405.17631 version 3, updated 2025-03-09; published as a conference paper at ICLR 2025
Retrieved: 2026-10-09T20:51:27Z

not individually reviewed

No individual claim review recorded

Audit details

Field: attributes.population

Source artifact SHA-256: dd94d80ec75c9bb84ec3989c910a313ef1a7aec57c772d3dd35ecadd250132a7

Hash scope: pdftotext -layout text layer, parsed by extract/extract_target_validation.py

Inspected artifact

attributes.scope_note
Replayed offline: the benchmark looks up the measured phenotypic response of each selected gene instead of running a new experiment. Genes the screen did not assay are absent from the candidate pool, so they are neither hits nor non-hits. Schmidt et al. 2022 report both CRISPR activation and interference screens; Roohani et al. 2025 describe the benchmark data as knockdown, and neither benchmark says which screen was used. Two Roohani et al. authors (Steinhart, Marson) are authors of Schmidt et al. 2022.
Context-only references
LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet?

Original source ↗

Roohani et al. 2025 section 4.1; Gupta et al. 2025 appendix B.1.1

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: Findings of the Association for Computational Linguistics: EMNLP 2025, pages 15482-15510
Retrieved: 2026-10-09T20:55:36Z

not individually reviewed

No individual claim review recorded

Audit details

Field: attributes.scope_note

Source artifact SHA-256: 7dda0b590f2736b7d30e48b97167a0868a2b4bde253589d8aff7929d01506ff9

Hash scope: pdftotext -layout text layer, parsed by extract/extract_target_validation.py

Inspected artifact

attributes.scope_note
Replayed offline: the benchmark looks up the measured phenotypic response of each selected gene instead of running a new experiment. Genes the screen did not assay are absent from the candidate pool, so they are neither hits nor non-hits. Schmidt et al. 2022 report both CRISPR activation and interference screens; Roohani et al. 2025 describe the benchmark data as knockdown, and neither benchmark says which screen was used. Two Roohani et al. authors (Steinhart, Marson) are authors of Schmidt et al. 2022.
Context-only references
BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments

Original source ↗

Roohani et al. 2025 section 4.1; Gupta et al. 2025 appendix B.1.1

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: arXiv:2405.17631 version 3, updated 2025-03-09; published as a conference paper at ICLR 2025
Retrieved: 2026-10-09T20:51:27Z

not individually reviewed

No individual claim review recorded

Audit details

Field: attributes.scope_note

Source artifact SHA-256: dd94d80ec75c9bb84ec3989c910a313ef1a7aec57c772d3dd35ecadd250132a7

Hash scope: pdftotext -layout text layer, parsed by extract/extract_target_validation.py

Inspected artifact

attributes.source_locator
Roohani et al. 2025 section 4.1; Gupta et al. 2025 appendix B.1.1
Context-only references
LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet?

Original source ↗

Roohani et al. 2025 section 4.1; Gupta et al. 2025 appendix B.1.1

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: Findings of the Association for Computational Linguistics: EMNLP 2025, pages 15482-15510
Retrieved: 2026-10-09T20:55:36Z

not individually reviewed

No individual claim review recorded

Audit details

Field: attributes.source_locator

Source artifact SHA-256: 7dda0b590f2736b7d30e48b97167a0868a2b4bde253589d8aff7929d01506ff9

Hash scope: pdftotext -layout text layer, parsed by extract/extract_target_validation.py

Inspected artifact

attributes.source_locator
Roohani et al. 2025 section 4.1; Gupta et al. 2025 appendix B.1.1
Context-only references
BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments

Original source ↗

Roohani et al. 2025 section 4.1; Gupta et al. 2025 appendix B.1.1

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: arXiv:2405.17631 version 3, updated 2025-03-09; published as a conference paper at ICLR 2025
Retrieved: 2026-10-09T20:51:27Z

not individually reviewed

No individual claim review recorded

Audit details

Field: attributes.source_locator

Source artifact SHA-256: dd94d80ec75c9bb84ec3989c910a313ef1a7aec57c772d3dd35ecadd250132a7

Hash scope: pdftotext -layout text layer, parsed by extract/extract_target_validation.py

Inspected artifact

attributes.version
As replayed by the cited benchmark from the original screen
Context-only references
LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet?

Original source ↗

Roohani et al. 2025 section 4.1; Gupta et al. 2025 appendix B.1.1

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: Findings of the Association for Computational Linguistics: EMNLP 2025, pages 15482-15510
Retrieved: 2026-10-09T20:55:36Z

not individually reviewed

No individual claim review recorded

Audit details

Field: attributes.version

Source artifact SHA-256: 7dda0b590f2736b7d30e48b97167a0868a2b4bde253589d8aff7929d01506ff9

Hash scope: pdftotext -layout text layer, parsed by extract/extract_target_validation.py

Inspected artifact

attributes.version
As replayed by the cited benchmark from the original screen
Context-only references
BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments

Original source ↗

Roohani et al. 2025 section 4.1; Gupta et al. 2025 appendix B.1.1

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: arXiv:2405.17631 version 3, updated 2025-03-09; published as a conference paper at ICLR 2025
Retrieved: 2026-10-09T20:51:27Z

not individually reviewed

No individual claim review recorded

Audit details

Field: attributes.version

Source artifact SHA-256: dd94d80ec75c9bb84ec3989c910a313ef1a7aec57c772d3dd35ecadd250132a7

Hash scope: pdftotext -layout text layer, parsed by extract/extract_target_validation.py

Inspected artifact

Sources and history

Release 2026-10-10-7fcc3e48a123 · Record review: source checked

2 source records and release historyDownload this release (gzip)
Technical metadata and extraction receipts

Stable ID: tgtval-20261009-data-schmidt2022-ifng

areas
cells-tissues
contexts
research
version
As replayed by the cited benchmark from the original screen
population
Over 18,000 genes, each knocked down in a distinct cell (Roohani et al. 2025 section 4.1)
assay
Pooled CRISPR perturbation screen with a phenotypic readout
scope note
Replayed offline: the benchmark looks up the measured phenotypic response of each selected gene instead of running a new experiment. Genes the screen did not assay are absent from the candidate pool, so they are neither hits nor non-hits. Schmidt et al. 2022 report both CRISPR activation and interference screens; Roohani et al. 2025 describe the benchmark data as knockdown, and neither benchmark says which screen was used. Two Roohani et al. authors (Steinhart, Marson) are authors of Schmidt et al. 2022.
source locator
Roohani et al. 2025 section 4.1; Gupta et al. 2025 appendix B.1.1
missing metadata
positives: reason: unreported; note: Roohani et al. do not print the hit count or the threshold tau per screen; Gupta et al. print their own ground-truth hit counts, recorded as a claim on their protocol; total: reason: unreported; note: Stated as 'over 18,000 genes' rather than an exact count
Related records

Suggest a correction