AssayBench: rank 100 candidate genes for a described CRISPR screen
Given a screen description and its significance criterion, a system returns a ranked list of 100 genes, scored against the screen's own hit labels.
Overview
Given a screen description and its significance criterion, a system returns a ranked list of 100 genes, scored against the screen's own hit labels.
Consult the linked sources for architecture or protocol details. Missing evidence is not evidence of a missing capability.
48 recorded evaluations, 144 metric rows. A comparison chart has not yet been validated for these results. The table retains the individual findings and their sources.
Results
Results are available, but no reviewed comparison panel is linked in this release.
All evaluations
48 evaluations · 144 results. Different protocols are not a single leaderboard.
Filter evaluations
Applied filters: All linked evaluations
| Tested configuration | Protocol and dataset | Finding | Evidence and details |
|---|---|---|---|
| Configuration: Biomni A1 (Claude 4) (AssayBench) | Protocol: AssayBench: rank 100 candidate genes for a described CRISPR screen Dataset: AssayBench post-cutoff cohort of human CRISPR screens | 0.0728 adjusted-ndcg-at-100 unitless · higher Uncertainty: Not reported by the source Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceBiomni A1 (Claude 4) on the AssayBench post-cutoff cohort tgtval-20261009-protocol-debrouwer2026-screen-gene-ranking Aggregation: Not reported AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents · Table 3, row 'Biomni A1 (Claude 4)' cohort 'LaTest', column 'AnDCG@100' |
| Configuration: Biomni A1 (Claude 4) (AssayBench) | Protocol: AssayBench: rank 100 candidate genes for a described CRISPR screen Dataset: AssayBench post-cutoff cohort of human CRISPR screens | NA directional-false-discovery-rate-at-100 fraction · lower Uncertainty: Not reported by the source Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceBiomni A1 (Claude 4) on the AssayBench post-cutoff cohort tgtval-20261009-protocol-debrouwer2026-screen-gene-ranking Aggregation: Not reported AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents · Table 3, row 'Biomni A1 (Claude 4)' cohort 'LaTest', column 'dFDR@100' |
| Configuration: Biomni A1 (Claude 4) (AssayBench) | Protocol: AssayBench: rank 100 candidate genes for a described CRISPR screen Dataset: AssayBench post-cutoff cohort of human CRISPR screens | 0.0214 precision-at-100 fraction · higher Uncertainty: Not reported by the source Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceBiomni A1 (Claude 4) on the AssayBench post-cutoff cohort tgtval-20261009-protocol-debrouwer2026-screen-gene-ranking Aggregation: Not reported AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents · Table 3, row 'Biomni A1 (Claude 4)' cohort 'LaTest', column 'Precision@100' |
| Configuration: Biomni A1 (Claude 4) (AssayBench) | Protocol: AssayBench: rank 100 candidate genes for a described CRISPR screen Dataset: AssayBench test cohort of human CRISPR screens | 0.108 adjusted-ndcg-at-100 unitless · higher Uncertainty: Not reported by the source Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceBiomni A1 (Claude 4) on the AssayBench test cohort tgtval-20261009-protocol-debrouwer2026-screen-gene-ranking Aggregation: Not reported AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents · Table 3, row 'Biomni A1 (Claude 4)' cohort 'test', column 'AnDCG@100' |
| Configuration: Biomni A1 (Claude 4) (AssayBench) | Protocol: AssayBench: rank 100 candidate genes for a described CRISPR screen Dataset: AssayBench test cohort of human CRISPR screens | 0.0176 directional-false-discovery-rate-at-100 fraction · lower Uncertainty: Not reported by the source Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceBiomni A1 (Claude 4) on the AssayBench test cohort tgtval-20261009-protocol-debrouwer2026-screen-gene-ranking Aggregation: Not reported AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents · Table 3, row 'Biomni A1 (Claude 4)' cohort 'test', column 'dFDR@100' |
| Configuration: Biomni A1 (Claude 4) (AssayBench) | Protocol: AssayBench: rank 100 candidate genes for a described CRISPR screen Dataset: AssayBench test cohort of human CRISPR screens | 0.17 precision-at-100 fraction · higher Uncertainty: Not reported by the source Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceBiomni A1 (Claude 4) on the AssayBench test cohort tgtval-20261009-protocol-debrouwer2026-screen-gene-ranking Aggregation: Not reported AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents · Table 3, row 'Biomni A1 (Claude 4)' cohort 'test', column 'Precision@100' |
| Configuration: Biomni A1 (Claude 4) (AssayBench) | Protocol: AssayBench: rank 100 candidate genes for a described CRISPR screen Dataset: AssayBench validation cohort of human CRISPR screens | 0.141 adjusted-ndcg-at-100 unitless · higher Uncertainty: Not reported by the source Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceBiomni A1 (Claude 4) on the AssayBench validation cohort tgtval-20261009-protocol-debrouwer2026-screen-gene-ranking Aggregation: Not reported AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents · Table 3, row 'Biomni A1 (Claude 4)' cohort 'val', column 'AnDCG@100' |
| Configuration: Biomni A1 (Claude 4) (AssayBench) | Protocol: AssayBench: rank 100 candidate genes for a described CRISPR screen Dataset: AssayBench validation cohort of human CRISPR screens | 0.0164 directional-false-discovery-rate-at-100 fraction · lower Uncertainty: Not reported by the source Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceBiomni A1 (Claude 4) on the AssayBench validation cohort tgtval-20261009-protocol-debrouwer2026-screen-gene-ranking Aggregation: Not reported AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents · Table 3, row 'Biomni A1 (Claude 4)' cohort 'val', column 'dFDR@100' |
| Configuration: Biomni A1 (Claude 4) (AssayBench) | Protocol: AssayBench: rank 100 candidate genes for a described CRISPR screen Dataset: AssayBench validation cohort of human CRISPR screens | 0.247 precision-at-100 fraction · higher Uncertainty: Not reported by the source Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceBiomni A1 (Claude 4) on the AssayBench validation cohort tgtval-20261009-protocol-debrouwer2026-screen-gene-ranking Aggregation: Not reported AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents · Table 3, row 'Biomni A1 (Claude 4)' cohort 'val', column 'Precision@100' |
| Configuration: C2S (Gemma-2B) (AssayBench) | Protocol: AssayBench: rank 100 candidate genes for a described CRISPR screen Dataset: AssayBench post-cutoff cohort of human CRISPR screens | 0.0078 adjusted-ndcg-at-100 unitless · higher Uncertainty: Not reported by the source Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceC2S (Gemma-2B) on the AssayBench post-cutoff cohort tgtval-20261009-protocol-debrouwer2026-screen-gene-ranking Aggregation: Not reported AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents · Table 3, row 'C2S (Gemma-2B)' cohort 'LaTest', column 'AnDCG@100' |
| Configuration: C2S (Gemma-2B) (AssayBench) | Protocol: AssayBench: rank 100 candidate genes for a described CRISPR screen Dataset: AssayBench post-cutoff cohort of human CRISPR screens | NA directional-false-discovery-rate-at-100 fraction · lower Uncertainty: Not reported by the source Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceC2S (Gemma-2B) on the AssayBench post-cutoff cohort tgtval-20261009-protocol-debrouwer2026-screen-gene-ranking Aggregation: Not reported AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents · Table 3, row 'C2S (Gemma-2B)' cohort 'LaTest', column 'dFDR@100' |
| Configuration: C2S (Gemma-2B) (AssayBench) | Protocol: AssayBench: rank 100 candidate genes for a described CRISPR screen Dataset: AssayBench post-cutoff cohort of human CRISPR screens | 0.0784 precision-at-100 fraction · higher Uncertainty: Not reported by the source Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceC2S (Gemma-2B) on the AssayBench post-cutoff cohort tgtval-20261009-protocol-debrouwer2026-screen-gene-ranking Aggregation: Not reported AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents · Table 3, row 'C2S (Gemma-2B)' cohort 'LaTest', column 'Precision@100' |
| Configuration: C2S (Gemma-2B) (AssayBench) | Protocol: AssayBench: rank 100 candidate genes for a described CRISPR screen Dataset: AssayBench test cohort of human CRISPR screens | 0.0355 adjusted-ndcg-at-100 unitless · higher Uncertainty: Not reported by the source Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceC2S (Gemma-2B) on the AssayBench test cohort tgtval-20261009-protocol-debrouwer2026-screen-gene-ranking Aggregation: Not reported AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents · Table 3, row 'C2S (Gemma-2B)' cohort 'test', column 'AnDCG@100' |
| Configuration: C2S (Gemma-2B) (AssayBench) | Protocol: AssayBench: rank 100 candidate genes for a described CRISPR screen Dataset: AssayBench test cohort of human CRISPR screens | 0.0218 directional-false-discovery-rate-at-100 fraction · lower Uncertainty: Not reported by the source Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceC2S (Gemma-2B) on the AssayBench test cohort tgtval-20261009-protocol-debrouwer2026-screen-gene-ranking Aggregation: Not reported AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents · Table 3, row 'C2S (Gemma-2B)' cohort 'test', column 'dFDR@100' |
| Configuration: C2S (Gemma-2B) (AssayBench) | Protocol: AssayBench: rank 100 candidate genes for a described CRISPR screen Dataset: AssayBench test cohort of human CRISPR screens | 0.0863 precision-at-100 fraction · higher Uncertainty: Not reported by the source Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceC2S (Gemma-2B) on the AssayBench test cohort tgtval-20261009-protocol-debrouwer2026-screen-gene-ranking Aggregation: Not reported AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents · Table 3, row 'C2S (Gemma-2B)' cohort 'test', column 'Precision@100' |
| Configuration: C2S (Gemma-2B) (AssayBench) | Protocol: AssayBench: rank 100 candidate genes for a described CRISPR screen Dataset: AssayBench validation cohort of human CRISPR screens | 0.0842 adjusted-ndcg-at-100 unitless · higher Uncertainty: Not reported by the source Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceC2S (Gemma-2B) on the AssayBench validation cohort tgtval-20261009-protocol-debrouwer2026-screen-gene-ranking Aggregation: Not reported AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents · Table 3, row 'C2S (Gemma-2B)' cohort 'val', column 'AnDCG@100' |
| Configuration: C2S (Gemma-2B) (AssayBench) | Protocol: AssayBench: rank 100 candidate genes for a described CRISPR screen Dataset: AssayBench validation cohort of human CRISPR screens | 0.0429 directional-false-discovery-rate-at-100 fraction · lower Uncertainty: Not reported by the source Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceC2S (Gemma-2B) on the AssayBench validation cohort tgtval-20261009-protocol-debrouwer2026-screen-gene-ranking Aggregation: Not reported AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents · Table 3, row 'C2S (Gemma-2B)' cohort 'val', column 'dFDR@100' |
| Configuration: C2S (Gemma-2B) (AssayBench) | Protocol: AssayBench: rank 100 candidate genes for a described CRISPR screen Dataset: AssayBench validation cohort of human CRISPR screens | 0.12 precision-at-100 fraction · higher Uncertainty: Not reported by the source Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceC2S (Gemma-2B) on the AssayBench validation cohort tgtval-20261009-protocol-debrouwer2026-screen-gene-ranking Aggregation: Not reported AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents · Table 3, row 'C2S (Gemma-2B)' cohort 'val', column 'Precision@100' |
| Configuration: Embedding kNN (AssayBench) | Protocol: AssayBench: rank 100 candidate genes for a described CRISPR screen Dataset: AssayBench post-cutoff cohort of human CRISPR screens | 0.0271 adjusted-ndcg-at-100 unitless · higher Uncertainty: Not reported by the source Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceEmbedding kNN on the AssayBench post-cutoff cohort tgtval-20261009-protocol-debrouwer2026-screen-gene-ranking Aggregation: Not reported AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents · Table 3, row 'Embedding kNN' cohort 'LaTest', column 'AnDCG@100' |
| Configuration: Embedding kNN (AssayBench) | Protocol: AssayBench: rank 100 candidate genes for a described CRISPR screen Dataset: AssayBench post-cutoff cohort of human CRISPR screens | NA directional-false-discovery-rate-at-100 fraction · lower Uncertainty: Not reported by the source Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceEmbedding kNN on the AssayBench post-cutoff cohort tgtval-20261009-protocol-debrouwer2026-screen-gene-ranking Aggregation: Not reported AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents · Table 3, row 'Embedding kNN' cohort 'LaTest', column 'dFDR@100' |
| Configuration: Embedding kNN (AssayBench) | Protocol: AssayBench: rank 100 candidate genes for a described CRISPR screen Dataset: AssayBench post-cutoff cohort of human CRISPR screens | 0.113 precision-at-100 fraction · higher Uncertainty: Not reported by the source Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceEmbedding kNN on the AssayBench post-cutoff cohort tgtval-20261009-protocol-debrouwer2026-screen-gene-ranking Aggregation: Not reported AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents · Table 3, row 'Embedding kNN' cohort 'LaTest', column 'Precision@100' |
| Configuration: Embedding kNN (AssayBench) | Protocol: AssayBench: rank 100 candidate genes for a described CRISPR screen Dataset: AssayBench test cohort of human CRISPR screens | 0.0646 adjusted-ndcg-at-100 unitless · higher Uncertainty: Not reported by the source Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceEmbedding kNN on the AssayBench test cohort tgtval-20261009-protocol-debrouwer2026-screen-gene-ranking Aggregation: Not reported AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents · Table 3, row 'Embedding kNN' cohort 'test', column 'AnDCG@100' |
| Configuration: Embedding kNN (AssayBench) | Protocol: AssayBench: rank 100 candidate genes for a described CRISPR screen Dataset: AssayBench test cohort of human CRISPR screens | 0.0226 directional-false-discovery-rate-at-100 fraction · lower Uncertainty: Not reported by the source Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceEmbedding kNN on the AssayBench test cohort tgtval-20261009-protocol-debrouwer2026-screen-gene-ranking Aggregation: Not reported AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents · Table 3, row 'Embedding kNN' cohort 'test', column 'dFDR@100' |
| Configuration: Embedding kNN (AssayBench) | Protocol: AssayBench: rank 100 candidate genes for a described CRISPR screen Dataset: AssayBench test cohort of human CRISPR screens | 0.106 precision-at-100 fraction · higher Uncertainty: Not reported by the source Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceEmbedding kNN on the AssayBench test cohort tgtval-20261009-protocol-debrouwer2026-screen-gene-ranking Aggregation: Not reported AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents · Table 3, row 'Embedding kNN' cohort 'test', column 'Precision@100' |
| Configuration: Embedding kNN (AssayBench) | Protocol: AssayBench: rank 100 candidate genes for a described CRISPR screen Dataset: AssayBench validation cohort of human CRISPR screens | 0.0697 adjusted-ndcg-at-100 unitless · higher Uncertainty: Not reported by the source Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceEmbedding kNN on the AssayBench validation cohort tgtval-20261009-protocol-debrouwer2026-screen-gene-ranking Aggregation: Not reported AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents · Table 3, row 'Embedding kNN' cohort 'val', column 'AnDCG@100' |
Source checking is not independent reproduction. Release 2026-10-10-7fcc3e48a123.
Methods and evaluation design
Procedure, tasks and evaluated configurations
Recorded evaluations
Each evaluation records what was tested and under which conditions.
- Biomni A1 (Claude 4) on the AssayBench post-cutoff cohort
- Biomni A1 (Claude 4) on the AssayBench test cohort
- Biomni A1 (Claude 4) on the AssayBench validation cohort
- C2S (Gemma-2B) on the AssayBench post-cutoff cohort
- C2S (Gemma-2B) on the AssayBench test cohort
- C2S (Gemma-2B) on the AssayBench validation cohort
- Embedding kNN on the AssayBench post-cutoff cohort
- Embedding kNN on the AssayBench test cohort
- Embedding kNN on the AssayBench validation cohort
- Gemini 3 Flash (GEPA) on the AssayBench post-cutoff cohort
- Gemini 3 Flash (GEPA) on the AssayBench test cohort
- Gemini 3 Flash (GEPA) on the AssayBench validation cohort
Baseline coverage
Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.
0 of 2 active baseline roles have published Rewire measurements in this release. Measurements on a selected protocol do not establish coverage of an entire suite.
No execution recipe linked to this protocol. Recipe availability does not establish a completed evaluation.
- Author-reported evaluations
- 24
- External evaluations
- 24
Literature evidence is not a Rewire measurement. Executed but unpublished runs and private review status are not included.
Null control
Proposed control: requires review
Select a task-valid null control after reviewing inputs and metric
Protocol-specific applicability, permitted inputs, access, split, evaluator and execution requirements need review before implementation or execution.
This is a suggested selection rule, not a validated method or a measured score.
Conventional reference
Proposed control: requires review
Select an upstream conventional reference after reviewing the full protocol
Protocol-specific applicability, permitted inputs, access, split, evaluator and execution requirements need review before implementation or execution.
This is a suggested selection rule, not a validated method or a measured score.
Protocol coverage CSV (gzip) · Model evaluation matrix (gzip) · Source table (gzip) · Release and checksums (gzip)
Coverage is derived from release 2026-10-10-7fcc3e48a123. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.
Run instructions
No runnable recipe has been reviewed for this protocol. Dataset access, model requirements, licences and compute requirements must be checked against its sources before execution.
Strengths, limitations and unresolved questions
Evidence
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Evidence table
Inspect claims, sources and review details
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
0 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|
No evidence rows match these filters. Choose another scope or clear the search.
Sources and history
Release 2026-10-10-7fcc3e48a123 · Record review: source checked
1 source records and release history
- AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents · Original source · arXiv:2605.10876 version 1, posted 2026-05-11; not peer reviewed
Technical metadata and extraction receipts
Stable ID: tgtval-20261009-protocol-debrouwer2026-screen-gene-ranking
- areas
- cells-tissues
- contexts
- research
- protocol
- For each benchmark entry the system is given the phenotype, cell line, cell type, library and perturbation modality, treatment and duration, plus the ranking objective, and returns 100 ranked gene symbols. Metrics: AnDCG@100 (condensed NDCG adjusted so 0 is a random ranking for that screen), Precision@100, and dFDR@100.
- version
- AssayBench, arXiv:2605.10876v1, sections 2 to 4
- metric
- adjusted-ndcg-at-100
- metric direction
- higher
- unit
- unitless
- metric definition
- AnDCG@100 removes unassayed genes from the ranked list, computes NDCG@100 against percentile-rank relevance, and rescales by the screen's analytic random baseline. Precision@100 counts hits among the top 100 predictions and divides by min(100, positive-relevance genes), so for screens with fewer than 100 hits it is the fraction of the screen's hits recovered. dFDR@100 is the share of the top 100 scored genes that are significant in the opposite direction.
- limitations
- AnDCG@100 removes genes the screen did not assay before scoring, so for the headline metric untested genes are not negatives. The printed Precision@100 and dFDR@100 formulas sum over the top 100 predictions without that step, so an unassayed gene adds nothing to Precision@100 and counts in the dFDR@100 denominator; the text calls them measures over in-screen predictions, so whether they were condensed is unclear.; Hits follow each screen's own significance criterion, so the phenotype and threshold vary between entries.; Retrospective replay of published screens; no new experiment was run.; Phenotype descriptions and effect directions come from an LLM-assisted curation step over BioGRID metadata.; The post-cutoff cohort has 19 screens against 334 in the test cohort. Most systems score lower on it, which the source reads as partly prior exposure to the literature; the trained gene-relevance predictor (0.0660 against 0.0565) and Qwen3.5-2B (0.0324 against 0.0284) score higher.; Precision@100 divides by min(100, number of positive-relevance genes), so it is not precision over a fixed 100 predictions when a screen has fewer than 100 hits.; The gene-frequency baseline is competitive mainly through fitness and viability screens, where frequently recurring essential genes give a phenotype-level prior (section 5.4).
- source locator
- Sections 2.4, 3 and 4; Tables 1 to 3
Related records
- uses data: AssayBench validation cohort of human CRISPR screens
- uses data: AssayBench test cohort of human CRISPR screens
- uses data: AssayBench post-cutoff cohort of human CRISPR screens
- assessment: Biomni A1 (Claude 4) on the AssayBench post-cutoff cohort
- assessment: Biomni A1 (Claude 4) on the AssayBench test cohort
- assessment: Biomni A1 (Claude 4) on the AssayBench validation cohort
- assessment: C2S (Gemma-2B) on the AssayBench post-cutoff cohort
- assessment: C2S (Gemma-2B) on the AssayBench test cohort
- assessment: C2S (Gemma-2B) on the AssayBench validation cohort
- assessment: Embedding kNN on the AssayBench post-cutoff cohort
- assessment: Embedding kNN on the AssayBench test cohort
- assessment: Embedding kNN on the AssayBench validation cohort
- assessment: Gemini 3 Flash (GEPA) on the AssayBench post-cutoff cohort
- assessment: Gemini 3 Flash (GEPA) on the AssayBench test cohort
- assessment: Gemini 3 Flash (GEPA) on the AssayBench validation cohort
- assessment: Gemini 3 Flash on the AssayBench post-cutoff cohort
- assessment: Gemini 3 Flash on the AssayBench test cohort
- assessment: Gemini 3 Flash on the AssayBench validation cohort
- assessment: Gemini 3 Pro (Few-shot) on the AssayBench post-cutoff cohort
- assessment: Gemini 3 Pro (Few-shot) on the AssayBench test cohort
- assessment: Gemini 3 Pro (Few-shot) on the AssayBench validation cohort
- assessment: Gemini 3 Pro on the AssayBench post-cutoff cohort
- assessment: Gemini 3 Pro on the AssayBench test cohort
- assessment: Gemini 3 Pro on the AssayBench validation cohort
- assessment: Gene-frequency (by phenotype) on the AssayBench post-cutoff cohort
- assessment: Gene-frequency (by phenotype) on the AssayBench test cohort
- assessment: Gene-frequency (by phenotype) on the AssayBench validation cohort
- assessment: Gene-relevance predictor on the AssayBench post-cutoff cohort
- assessment: Gene-relevance predictor on the AssayBench test cohort
- assessment: Gene-relevance predictor on the AssayBench validation cohort
- assessment: GPT-5.4 on the AssayBench post-cutoff cohort
- assessment: GPT-5.4 on the AssayBench test cohort
- assessment: GPT-5.4 on the AssayBench validation cohort
- assessment: GPT-OSS-120B on the AssayBench post-cutoff cohort
- assessment: GPT-OSS-120B (SFT + GRPO) on the AssayBench post-cutoff cohort
- assessment: GPT-OSS-120B (SFT + GRPO) on the AssayBench test cohort
- assessment: GPT-OSS-120B (SFT + GRPO) on the AssayBench validation cohort
- assessment: GPT-OSS-120B (SFT) on the AssayBench post-cutoff cohort
- assessment: GPT-OSS-120B (SFT) on the AssayBench test cohort
- assessment: GPT-OSS-120B (SFT) on the AssayBench validation cohort
- assessment: GPT-OSS-120B on the AssayBench test cohort
- assessment: GPT-OSS-120B on the AssayBench validation cohort
- assessment: LLM Ensemble on the AssayBench post-cutoff cohort
- assessment: LLM Ensemble on the AssayBench test cohort
- assessment: LLM Ensemble on the AssayBench validation cohort
- assessment: Oracle kNN on the AssayBench post-cutoff cohort
- assessment: Oracle kNN on the AssayBench test cohort
- assessment: Oracle kNN on the AssayBench validation cohort
- assessment: Qwen3.5-2B on the AssayBench post-cutoff cohort
- assessment: Qwen3.5-2B on the AssayBench test cohort
- assessment: Qwen3.5-2B on the AssayBench validation cohort
- assessed by: Select therapeutic targets for validation