BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes
Offline replay of six CRISPR screens: each method selects 128 genes per round for 5 rounds and is scored on the fraction of that screen's hits it recovered.
Overview
Offline replay of six CRISPR screens: each method selects 128 genes per round for 5 rounds and is scored on the fraction of that screen's hits it recovered.
Consult the linked sources for architecture or protocol details. Missing evidence is not evidence of a missing capability.
120 recorded evaluations, 240 metric rows. A comparison chart has not yet been validated for these results. The table retains the individual findings and their sources.
Results
Results are available, but no reviewed comparison panel is linked in this release.
All evaluations
120 evaluations · 240 results. Different protocols are not a single leaderboard.
Filter evaluations
Applied filters: All linked evaluations
| Tested configuration | Protocol and dataset | Finding | Evidence and details |
|---|---|---|---|
| Configuration: Badge acquisition function on an MLP surrogate (Roohani et al. 2025) | Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes Dataset: Carnevale et al. 2022 screen, T-cell resistance to tumour-microenvironment inhibitory signals | 0.044 recall fraction · higher Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass. Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceBadge on Carnev. (Roohani et al. 2025) tgtval-20261009-protocol-roohani2025-hitratio-round5 Aggregation: Not reported BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Badge', column 'All' under 'Carnev.' |
| Configuration: Badge acquisition function on an MLP surrogate (Roohani et al. 2025) | Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes Dataset: Carnevale et al. 2022 screen, T-cell resistance to tumour-microenvironment inhibitory signals | 0.036 recall fraction · higher Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass. Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceBadge on Carnev. (Roohani et al. 2025) tgtval-20261009-protocol-roohani2025-hitratio-round5 Aggregation: Not reported BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Badge', column 'N/E' under 'Carnev.' |
| Configuration: Badge acquisition function on an MLP surrogate (Roohani et al. 2025) | Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes Dataset: CAR-T proliferation screen (unpublished dataset used by Roohani et al. 2025) | 0.042 recall fraction · higher Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass. Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceBadge on CAR-T (Roohani et al. 2025) tgtval-20261009-protocol-roohani2025-hitratio-round5 Aggregation: Not reported BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Badge', column 'All' under 'CAR-T' |
| Configuration: Badge acquisition function on an MLP surrogate (Roohani et al. 2025) | Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes Dataset: CAR-T proliferation screen (unpublished dataset used by Roohani et al. 2025) | 0.038 recall fraction · higher Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass. Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceBadge on CAR-T (Roohani et al. 2025) tgtval-20261009-protocol-roohani2025-hitratio-round5 Aggregation: Not reported BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Badge', column 'N/E' under 'CAR-T' |
| Configuration: Badge acquisition function on an MLP surrogate (Roohani et al. 2025) | Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes Dataset: Sanchez et al. 2021 screen, endogenous tau protein level in neurons | 0.039 recall fraction · higher Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass. Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceBadge on Sanchez (Roohani et al. 2025) tgtval-20261009-protocol-roohani2025-hitratio-round5 Aggregation: Not reported BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Badge', column 'All' under 'Sanchez' |
| Configuration: Badge acquisition function on an MLP surrogate (Roohani et al. 2025) | Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes Dataset: Sanchez et al. 2021 screen, endogenous tau protein level in neurons | 0.035 recall fraction · higher Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass. Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceBadge on Sanchez (Roohani et al. 2025) tgtval-20261009-protocol-roohani2025-hitratio-round5 Aggregation: Not reported BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Badge', column 'N/E' under 'Sanchez' |
| Configuration: Badge acquisition function on an MLP surrogate (Roohani et al. 2025) | Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes Dataset: Scharenberg et al. 2023 screen, lysosomal choline recycling in pancreatic cells | 0.258 recall fraction · higher Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass. Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceBadge on Scharen. (Roohani et al. 2025) tgtval-20261009-protocol-roohani2025-hitratio-round5 Aggregation: Not reported BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Badge', column 'All' under 'Scharen.' |
| Configuration: Badge acquisition function on an MLP surrogate (Roohani et al. 2025) | Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes Dataset: Scharenberg et al. 2023 screen, lysosomal choline recycling in pancreatic cells | 0.211 recall fraction · higher Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass. Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceBadge on Scharen. (Roohani et al. 2025) tgtval-20261009-protocol-roohani2025-hitratio-round5 Aggregation: Not reported BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Badge', column 'N/E' under 'Scharen.' |
| Configuration: Badge acquisition function on an MLP surrogate (Roohani et al. 2025) | Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes Dataset: Schmidt et al. 2022 screen, interferon-gamma production in primary human T cells | 0.06 recall fraction · higher Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass. Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceBadge on Schmidt1 (Roohani et al. 2025) tgtval-20261009-protocol-roohani2025-hitratio-round5 Aggregation: Not reported BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Badge', column 'All' under 'Schmidt1' |
| Configuration: Badge acquisition function on an MLP surrogate (Roohani et al. 2025) | Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes Dataset: Schmidt et al. 2022 screen, interferon-gamma production in primary human T cells | 0.05 recall fraction · higher Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass. Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceBadge on Schmidt1 (Roohani et al. 2025) tgtval-20261009-protocol-roohani2025-hitratio-round5 Aggregation: Not reported BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Badge', column 'N/E' under 'Schmidt1' |
| Configuration: Badge acquisition function on an MLP surrogate (Roohani et al. 2025) | Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes Dataset: Schmidt et al. 2022 screen, interleukin-2 production in primary human T cells | 0.077 recall fraction · higher Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass. Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceBadge on Schmidt2 (Roohani et al. 2025) tgtval-20261009-protocol-roohani2025-hitratio-round5 Aggregation: Not reported BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Badge', column 'All' under 'Schmidt2' |
| Configuration: Badge acquisition function on an MLP surrogate (Roohani et al. 2025) | Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes Dataset: Schmidt et al. 2022 screen, interleukin-2 production in primary human T cells | 0.058 recall fraction · higher Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass. Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceBadge on Schmidt2 (Roohani et al. 2025) tgtval-20261009-protocol-roohani2025-hitratio-round5 Aggregation: Not reported BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Badge', column 'N/E' under 'Schmidt2' |
| Configuration: BioDiscoveryAgent (No-Tools), Claude 3.5 Sonnet (Roohani et al. 2025) | Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes Dataset: Carnevale et al. 2022 screen, T-cell resistance to tumour-microenvironment inhibitory signals | 0.042 recall fraction · higher Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass. Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceClaude 3.5 Sonnet on Carnev. (Roohani et al. 2025) tgtval-20261009-protocol-roohani2025-hitratio-round5 Aggregation: Not reported BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Claude 3.5 Sonnet', column 'All' under 'Carnev.' |
| Configuration: BioDiscoveryAgent (No-Tools), Claude 3.5 Sonnet (Roohani et al. 2025) | Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes Dataset: Carnevale et al. 2022 screen, T-cell resistance to tumour-microenvironment inhibitory signals | 0.044 recall fraction · higher Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass. Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceClaude 3.5 Sonnet on Carnev. (Roohani et al. 2025) tgtval-20261009-protocol-roohani2025-hitratio-round5 Aggregation: Not reported BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Claude 3.5 Sonnet', column 'N/E' under 'Carnev.' |
| Configuration: BioDiscoveryAgent (No-Tools), Claude 3.5 Sonnet (Roohani et al. 2025) | Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes Dataset: CAR-T proliferation screen (unpublished dataset used by Roohani et al. 2025) | 0.13 recall fraction · higher Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass. Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceClaude 3.5 Sonnet on CAR-T (Roohani et al. 2025) tgtval-20261009-protocol-roohani2025-hitratio-round5 Aggregation: Not reported BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Claude 3.5 Sonnet', column 'All' under 'CAR-T' |
| Configuration: BioDiscoveryAgent (No-Tools), Claude 3.5 Sonnet (Roohani et al. 2025) | Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes Dataset: CAR-T proliferation screen (unpublished dataset used by Roohani et al. 2025) | 0.133 recall fraction · higher Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass. Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceClaude 3.5 Sonnet on CAR-T (Roohani et al. 2025) tgtval-20261009-protocol-roohani2025-hitratio-round5 Aggregation: Not reported BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Claude 3.5 Sonnet', column 'N/E' under 'CAR-T' |
| Configuration: BioDiscoveryAgent (No-Tools), Claude 3.5 Sonnet (Roohani et al. 2025) | Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes Dataset: Sanchez et al. 2021 screen, endogenous tau protein level in neurons | 0.066 recall fraction · higher Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass. Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceClaude 3.5 Sonnet on Sanchez (Roohani et al. 2025) tgtval-20261009-protocol-roohani2025-hitratio-round5 Aggregation: Not reported BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Claude 3.5 Sonnet', column 'All' under 'Sanchez' |
| Configuration: BioDiscoveryAgent (No-Tools), Claude 3.5 Sonnet (Roohani et al. 2025) | Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes Dataset: Sanchez et al. 2021 screen, endogenous tau protein level in neurons | 0.063 recall fraction · higher Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass. Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceClaude 3.5 Sonnet on Sanchez (Roohani et al. 2025) tgtval-20261009-protocol-roohani2025-hitratio-round5 Aggregation: Not reported BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Claude 3.5 Sonnet', column 'N/E' under 'Sanchez' |
| Configuration: BioDiscoveryAgent (No-Tools), Claude 3.5 Sonnet (Roohani et al. 2025) | Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes Dataset: Scharenberg et al. 2023 screen, lysosomal choline recycling in pancreatic cells | 0.326 recall fraction · higher Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass. Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceClaude 3.5 Sonnet on Scharen. (Roohani et al. 2025) tgtval-20261009-protocol-roohani2025-hitratio-round5 Aggregation: Not reported BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Claude 3.5 Sonnet', column 'All' under 'Scharen.' |
| Configuration: BioDiscoveryAgent (No-Tools), Claude 3.5 Sonnet (Roohani et al. 2025) | Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes Dataset: Scharenberg et al. 2023 screen, lysosomal choline recycling in pancreatic cells | 0.292 recall fraction · higher Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass. Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceClaude 3.5 Sonnet on Scharen. (Roohani et al. 2025) tgtval-20261009-protocol-roohani2025-hitratio-round5 Aggregation: Not reported BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Claude 3.5 Sonnet', column 'N/E' under 'Scharen.' |
| Configuration: BioDiscoveryAgent (No-Tools), Claude 3.5 Sonnet (Roohani et al. 2025) | Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes Dataset: Schmidt et al. 2022 screen, interferon-gamma production in primary human T cells | 0.095 recall fraction · higher Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass. Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceClaude 3.5 Sonnet on Schmidt1 (Roohani et al. 2025) tgtval-20261009-protocol-roohani2025-hitratio-round5 Aggregation: Not reported BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Claude 3.5 Sonnet', column 'All' under 'Schmidt1' |
| Configuration: BioDiscoveryAgent (No-Tools), Claude 3.5 Sonnet (Roohani et al. 2025) | Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes Dataset: Schmidt et al. 2022 screen, interferon-gamma production in primary human T cells | 0.107 recall fraction · higher Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass. Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceClaude 3.5 Sonnet on Schmidt1 (Roohani et al. 2025) tgtval-20261009-protocol-roohani2025-hitratio-round5 Aggregation: Not reported BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Claude 3.5 Sonnet', column 'N/E' under 'Schmidt1' |
| Configuration: BioDiscoveryAgent (No-Tools), Claude 3.5 Sonnet (Roohani et al. 2025) | Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes Dataset: Schmidt et al. 2022 screen, interleukin-2 production in primary human T cells | 0.104 recall fraction · higher Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass. Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceClaude 3.5 Sonnet on Schmidt2 (Roohani et al. 2025) tgtval-20261009-protocol-roohani2025-hitratio-round5 Aggregation: Not reported BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Claude 3.5 Sonnet', column 'All' under 'Schmidt2' |
| Configuration: BioDiscoveryAgent (No-Tools), Claude 3.5 Sonnet (Roohani et al. 2025) | Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes Dataset: Schmidt et al. 2022 screen, interleukin-2 production in primary human T cells | 0.122 recall fraction · higher Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass. Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceClaude 3.5 Sonnet on Schmidt2 (Roohani et al. 2025) tgtval-20261009-protocol-roohani2025-hitratio-round5 Aggregation: Not reported BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Claude 3.5 Sonnet', column 'N/E' under 'Schmidt2' |
| Configuration: BioDiscoveryAgent (No-Tools), Claude 3 Haiku (Roohani et al. 2025) | Protocol: BioDiscoveryAgent 1-gene perturbation design: hit ratio after 5 rounds of 128 genes Dataset: Carnevale et al. 2022 screen, T-cell resistance to tumour-microenvironment inhibitory signals | 0.032 recall fraction · higher Uncertainty: Not yet extracted: Appendix Table 7 prints one standard deviation over 10 runs for these values; not extracted in this pass. Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceClaude 3 Haiku on Carnev. (Roohani et al. 2025) tgtval-20261009-protocol-roohani2025-hitratio-round5 Aggregation: Not reported BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Table 1, row 'Claude 3 Haiku', column 'All' under 'Carnev.' |
Source checking is not independent reproduction. Release 2026-10-10-7fcc3e48a123.
Methods and evaluation design
Procedure, tasks and evaluated configurations
Recorded evaluations
Each evaluation records what was tested and under which conditions.
- Badge on Carnev. (Roohani et al. 2025)
- Badge on CAR-T (Roohani et al. 2025)
- Badge on Sanchez (Roohani et al. 2025)
- Badge on Scharen. (Roohani et al. 2025)
- Badge on Schmidt1 (Roohani et al. 2025)
- Badge on Schmidt2 (Roohani et al. 2025)
- Claude 3.5 Sonnet on Carnev. (Roohani et al. 2025)
- Claude 3.5 Sonnet on CAR-T (Roohani et al. 2025)
- Claude 3.5 Sonnet on Sanchez (Roohani et al. 2025)
- Claude 3.5 Sonnet on Scharen. (Roohani et al. 2025)
- Claude 3.5 Sonnet on Schmidt1 (Roohani et al. 2025)
- Claude 3.5 Sonnet on Schmidt2 (Roohani et al. 2025)
Baseline coverage
Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.
0 of 2 active baseline roles have published Rewire measurements in this release. Measurements on a selected protocol do not establish coverage of an entire suite.
No execution recipe linked to this protocol. Recipe availability does not establish a completed evaluation.
- Author-reported evaluations
- 72
- External evaluations
- 48
Literature evidence is not a Rewire measurement. Executed but unpublished runs and private review status are not included.
Null control
Proposed control: requires review
No-change prediction under matched control conditions
Protocol-specific applicability, permitted inputs, access, split, evaluator and execution requirements need review before implementation or execution.
This is a suggested selection rule, not a validated method or a measured score.
Conventional reference
Proposed control: requires review
Training-only mean-effect or linear prediction
Protocol-specific applicability, permitted inputs, access, split, evaluator and execution requirements need review before implementation or execution.
This is a suggested selection rule, not a validated method or a measured score.
Protocol coverage CSV (gzip) · Model evaluation matrix (gzip) · Source table (gzip) · Release and checksums (gzip)
Coverage is derived from release 2026-10-10-7fcc3e48a123. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.
Run instructions
No runnable recipe has been reviewed for this protocol. Dataset access, model requirements, licences and compute requirements must be checked against its sources before execution.
Strengths, limitations and unresolved questions
Evidence
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Evidence table
Inspect claims, sources and review details
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
0 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|
No evidence rows match these filters. Choose another scope or clear the search.
Sources and history
Release 2026-10-10-7fcc3e48a123 · Record review: source checked
1 source records and release history
- BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments · Original source · arXiv:2405.17631 version 3, updated 2025-03-09; published as a conference paper at ICLR 2025
Technical metadata and extraction receipts
Stable ID: tgtval-20261009-protocol-roohani2025-hitratio-round5
- areas
- cells-tissues
- contexts
- research
- protocol
- At each of 5 rounds a method selects 128 genes (32 for Scharenberg et al. 2023, whose pool is 1,061 perturbations); the measured phenotypic response of each selected gene is then revealed from the screen. Score: hit ratio after round 5, the fraction of the screen's true hits that were selected in any round, averaged over 10 runs. Reported twice per screen: over all genes, and over non-essential genes only.
- version
- Roohani et al. 2025 (ICLR 2025) sections 2 and 4; Table 1
- metric
- recall
- metric direction
- higher
- unit
- fraction
- metric definition
- Hit ratio = |genes selected in rounds 1..5 whose response exceeds the threshold tau| / |all genes in the screen whose response exceeds tau|. This is recall over the screen's hit set. Non-hits are genes the screen assayed whose response did not exceed tau; genes the screen did not assay are not in the pool.
- selection
- Candidate pool is the genes assayed by each screen (over 18,000 per screen; 1,061 for Scharenberg et al. 2023)
- limitations
- Offline replay of six existing screens: the candidate pool is the genes each screen assayed, so untested genes are absent rather than negative.; Hit ratio measures recovery of that screen's own hits, not whether a target is useful.; Hits are defined by a threshold tau on the screen's phenotypic response; the source does not print tau or the hit count per screen.; All rows were run by the agent's developers. Two authors (Steinhart, Marson) are authors of the Schmidt et al. 2022 screens, the CAR-T screen is unpublished, and the conflicts statement says patent applications have been filed on the findings.; Gupta et al. 2025 report that the same agent performs about the same when its experimental feedback is replaced by randomly permuted outcomes: in their Table 1 the permuted variant scores at least as high on three of five screens with each of three backbones. These numbers should not be read as evidence that the agent learns from each round's results.
- source locator
- Sections 2 and 4.1; Table 1 and its caption; appendix Table 7 (error intervals)
Related records
- uses data: Schmidt et al. 2022 screen, interferon-gamma production in primary human T cells
- uses data: Schmidt et al. 2022 screen, interleukin-2 production in primary human T cells
- uses data: CAR-T proliferation screen (unpublished dataset used by Roohani et al. 2025)
- uses data: Scharenberg et al. 2023 screen, lysosomal choline recycling in pancreatic cells
- uses data: Carnevale et al. 2022 screen, T-cell resistance to tumour-microenvironment inhibitory signals
- uses data: Sanchez et al. 2021 screen, endogenous tau protein level in neurons
- subject: metric_definition: tgtval-20261009-protocol-roohani2025-hitratio-round5
- subject: scoring_population: tgtval-20261009-protocol-roohani2025-hitratio-round5
- assessment: Badge on Carnev. (Roohani et al. 2025)
- assessment: Badge on CAR-T (Roohani et al. 2025)
- assessment: Badge on Sanchez (Roohani et al. 2025)
- assessment: Badge on Scharen. (Roohani et al. 2025)
- assessment: Badge on Schmidt1 (Roohani et al. 2025)
- assessment: Badge on Schmidt2 (Roohani et al. 2025)
- assessment: Claude 3.5 Sonnet on Carnev. (Roohani et al. 2025)
- assessment: Claude 3.5 Sonnet on CAR-T (Roohani et al. 2025)
- assessment: Claude 3.5 Sonnet on Sanchez (Roohani et al. 2025)
- assessment: Claude 3.5 Sonnet on Scharen. (Roohani et al. 2025)
- assessment: Claude 3.5 Sonnet on Schmidt1 (Roohani et al. 2025)
- assessment: Claude 3.5 Sonnet on Schmidt2 (Roohani et al. 2025)
- assessment: Claude 3 Haiku on Carnev. (Roohani et al. 2025)
- assessment: Claude 3 Haiku on CAR-T (Roohani et al. 2025)
- assessment: Claude 3 Haiku on Sanchez (Roohani et al. 2025)
- assessment: Claude 3 Haiku on Scharen. (Roohani et al. 2025)
- assessment: Claude 3 Haiku on Schmidt1 (Roohani et al. 2025)
- assessment: Claude 3 Haiku on Schmidt2 (Roohani et al. 2025)
- assessment: Claude 3 Opus on Carnev. (Roohani et al. 2025)
- assessment: Claude 3 Opus on CAR-T (Roohani et al. 2025)
- assessment: Claude 3 Opus on Sanchez (Roohani et al. 2025)
- assessment: Claude 3 Opus on Scharen. (Roohani et al. 2025)
- assessment: Claude 3 Opus on Schmidt1 (Roohani et al. 2025)
- assessment: Claude 3 Opus on Schmidt2 (Roohani et al. 2025)
- assessment: Claude 3 Sonnet on Carnev. (Roohani et al. 2025)
- assessment: Claude 3 Sonnet on CAR-T (Roohani et al. 2025)
- assessment: Claude 3 Sonnet on Sanchez (Roohani et al. 2025)
- assessment: Claude 3 Sonnet on Scharen. (Roohani et al. 2025)
- assessment: Claude 3 Sonnet on Schmidt1 (Roohani et al. 2025)
- assessment: Claude 3 Sonnet on Schmidt2 (Roohani et al. 2025)
- assessment: Claude v1 on Carnev. (Roohani et al. 2025)
- assessment: Claude v1 on CAR-T (Roohani et al. 2025)
- assessment: Claude v1 on Sanchez (Roohani et al. 2025)
- assessment: Claude v1 on Scharen. (Roohani et al. 2025)
- assessment: Claude v1 on Schmidt1 (Roohani et al. 2025)
- assessment: Claude v1 on Schmidt2 (Roohani et al. 2025)
- assessment: Coreset on Carnev. (Roohani et al. 2025)
- assessment: Coreset on CAR-T (Roohani et al. 2025)
- assessment: Coreset on Sanchez (Roohani et al. 2025)
- assessment: Coreset on Scharen. (Roohani et al. 2025)
- assessment: Coreset on Schmidt1 (Roohani et al. 2025)
- assessment: Coreset on Schmidt2 (Roohani et al. 2025)
- assessment: DiscoBax on Carnev. (Roohani et al. 2025)
- assessment: DiscoBax on CAR-T (Roohani et al. 2025)
- assessment: DiscoBax on Sanchez (Roohani et al. 2025)
- assessment: DiscoBax on Scharen. (Roohani et al. 2025)
- assessment: DiscoBax on Schmidt1 (Roohani et al. 2025)
- assessment: DiscoBax on Schmidt2 (Roohani et al. 2025)
- assessment: GPT-3.5-Turbo on Carnev. (Roohani et al. 2025)
- assessment: GPT-3.5-Turbo on CAR-T (Roohani et al. 2025)
- assessment: GPT-3.5-Turbo on Sanchez (Roohani et al. 2025)
- assessment: GPT-3.5-Turbo on Scharen. (Roohani et al. 2025)
- assessment: GPT-3.5-Turbo on Schmidt1 (Roohani et al. 2025)
- assessment: GPT-3.5-Turbo on Schmidt2 (Roohani et al. 2025)
- assessment: GPT-4o on Carnev. (Roohani et al. 2025)
- assessment: GPT-4o on CAR-T (Roohani et al. 2025)
- assessment: GPT-4o on Sanchez (Roohani et al. 2025)
- assessment: GPT-4o on Scharen. (Roohani et al. 2025)
- assessment: GPT-4o on Schmidt1 (Roohani et al. 2025)
- assessment: GPT-4o on Schmidt2 (Roohani et al. 2025)
- assessment: Human on Carnev. (Roohani et al. 2025)
- assessment: Human on CAR-T (Roohani et al. 2025)
- assessment: Human on Sanchez (Roohani et al. 2025)
- assessment: Human on Scharen. (Roohani et al. 2025)
- assessment: Human on Schmidt1 (Roohani et al. 2025)
- assessment: Human on Schmidt2 (Roohani et al. 2025)
- assessment: K-Means (D) on Carnev. (Roohani et al. 2025)
- assessment: K-Means (D) on CAR-T (Roohani et al. 2025)
- assessment: K-Means (D) on Sanchez (Roohani et al. 2025)
- assessment: K-Means (D) on Scharen. (Roohani et al. 2025)
- assessment: K-Means (D) on Schmidt1 (Roohani et al. 2025)
- assessment: K-Means (D) on Schmidt2 (Roohani et al. 2025)
- assessment: K-Means (E) on Carnev. (Roohani et al. 2025)
- assessment: K-Means (E) on CAR-T (Roohani et al. 2025)
- assessment: K-Means (E) on Sanchez (Roohani et al. 2025)
- assessment: K-Means (E) on Scharen. (Roohani et al. 2025)
- assessment: K-Means (E) on Schmidt1 (Roohani et al. 2025)
- assessment: K-Means (E) on Schmidt2 (Roohani et al. 2025)
- assessment: Margin Sample on Carnev. (Roohani et al. 2025)
- assessment: Margin Sample on CAR-T (Roohani et al. 2025)
- assessment: Margin Sample on Sanchez (Roohani et al. 2025)
- assessment: Margin Sample on Scharen. (Roohani et al. 2025)
- assessment: Margin Sample on Schmidt1 (Roohani et al. 2025)
- assessment: Margin Sample on Schmidt2 (Roohani et al. 2025)
- assessment: o1-mini on Carnev. (Roohani et al. 2025)
- assessment: o1-mini on CAR-T (Roohani et al. 2025)
- assessment: o1-mini on Sanchez (Roohani et al. 2025)
- assessment: o1-mini on Scharen. (Roohani et al. 2025)
- assessment: o1-mini on Schmidt1 (Roohani et al. 2025)
- assessment: o1-mini on Schmidt2 (Roohani et al. 2025)
- assessment: o1-preview on Carnev. (Roohani et al. 2025)
- assessment: o1-preview on CAR-T (Roohani et al. 2025)
- assessment: o1-preview on Sanchez (Roohani et al. 2025)
- assessment: o1-preview on Scharen. (Roohani et al. 2025)
- assessment: o1-preview on Schmidt1 (Roohani et al. 2025)
- assessment: o1-preview on Schmidt2 (Roohani et al. 2025)
- assessment: Random on Carnev. (Roohani et al. 2025)
- assessment: Random on CAR-T (Roohani et al. 2025)
- assessment: Random on Sanchez (Roohani et al. 2025)
- assessment: Random on Scharen. (Roohani et al. 2025)
- assessment: Random on Schmidt1 (Roohani et al. 2025)
- assessment: Random on Schmidt2 (Roohani et al. 2025)
- assessment: Soft Uncertain on Carnev. (Roohani et al. 2025)
- assessment: Soft Uncertain on CAR-T (Roohani et al. 2025)
- assessment: Soft Uncertain on Sanchez (Roohani et al. 2025)
- assessment: Soft Uncertain on Scharen. (Roohani et al. 2025)
- assessment: Soft Uncertain on Schmidt1 (Roohani et al. 2025)
- assessment: Soft Uncertain on Schmidt2 (Roohani et al. 2025)
- assessment: Sonnet + Coreset on Carnev. (Roohani et al. 2025)
- assessment: Sonnet + Coreset on CAR-T (Roohani et al. 2025)
- assessment: Sonnet + Coreset on Sanchez (Roohani et al. 2025)
- assessment: Sonnet + Coreset on Scharen. (Roohani et al. 2025)
- assessment: Sonnet + Coreset on Schmidt1 (Roohani et al. 2025)
- assessment: Sonnet + Coreset on Schmidt2 (Roohani et al. 2025)
- assessment: Top Uncertain on Carnev. (Roohani et al. 2025)
- assessment: Top Uncertain on CAR-T (Roohani et al. 2025)
- assessment: Top Uncertain on Sanchez (Roohani et al. 2025)
- assessment: Top Uncertain on Scharen. (Roohani et al. 2025)
- assessment: Top Uncertain on Schmidt1 (Roohani et al. 2025)
- assessment: Top Uncertain on Schmidt2 (Roohani et al. 2025)
- assessed by: Select therapeutic targets for validation