rewirebio.iobenchmarks
Evaluation

Gemini 3 Pro (Few-shot) on the AssayBench test cohort

Published comparison; transcribed, not reproduced.

Research readiness

These checks assess whether the evidence supports a reproducible investigation. A source-checked score alone does not meet these requirements.

Release 2026-10-10-6e93f504adfc · Evidence verified: Not verified

Evidence incomplete

Replay metrics

Exact outcomes, predictions, identifiers and evaluator are connected.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • metric replay: verification is missing

Verified: Not verified

Evidence incomplete

Investigate discrepancies

Replay evidence includes annotations and an assessment of dependence. Unknown independence permits descriptive analysis only.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • metric replay: verification is missing
  • annotations: verification is missing
  • dependence: verification is missing

Verified: Not verified

Evidence incomplete

Run locally

A pinned recipe describes the inputs, environment and resource requirements.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • recipe pinned: verification is missing
  • resource estimate: verification is missing

Verified: Not verified

Evidence incomplete

Validate independently

Separate data and exposure records support an independent test.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • independent validation: verification is missing
  • overlap checked: verification is missing

Verified: Not verified

Readiness describes the evidence in this release. Availability on your computer is checked separately when an investigation runs. Existing data exposure can prevent independent validation even when files are available.

Artifacts and reproduction

No verified artifact manifest is connected to this record yet. The gaps above identify what is needed before analysis can begin.

Read reviewed discrepancy investigations

Evaluation results

1 evaluation · 3 results. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: Gemini 3 Pro (Few-shot) (AssayBench)Protocol: AssayBench: rank 100 candidate genes for a described CRISPR screen
Dataset: AssayBench test cohort of human CRISPR screens
0.154 adjusted-ndcg-at-100
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

Gemini 3 Pro (Few-shot) on the AssayBench test cohort

tgtval-20261009-protocol-debrouwer2026-screen-gene-ranking

Aggregation: Not reported

AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents · Table 3, row 'Gemini 3 Pro (Few-shot)' cohort 'test', column 'AnDCG@100'
Configuration: Gemini 3 Pro (Few-shot) (AssayBench)Protocol: AssayBench: rank 100 candidate genes for a described CRISPR screen
Dataset: AssayBench test cohort of human CRISPR screens
0.0213 directional-false-discovery-rate-at-100
fraction · lower

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

Gemini 3 Pro (Few-shot) on the AssayBench test cohort

tgtval-20261009-protocol-debrouwer2026-screen-gene-ranking

Aggregation: Not reported

AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents · Table 3, row 'Gemini 3 Pro (Few-shot)' cohort 'test', column 'dFDR@100'
Configuration: Gemini 3 Pro (Few-shot) (AssayBench)Protocol: AssayBench: rank 100 candidate genes for a described CRISPR screen
Dataset: AssayBench test cohort of human CRISPR screens
0.213 precision-at-100
fraction · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

Gemini 3 Pro (Few-shot) on the AssayBench test cohort

tgtval-20261009-protocol-debrouwer2026-screen-gene-ranking

Aggregation: Not reported

AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents · Table 3, row 'Gemini 3 Pro (Few-shot)' cohort 'test', column 'Precision@100'

Source checking is not independent reproduction. Release 2026-10-10-6e93f504adfc.

Evaluation procedure

tgtval-20261009-protocol-debrouwer2026-screen-gene-ranking

Configuration
Gemini 3 Pro (Few-shot) (AssayBench)
Protocol
AssayBench: rank 100 candidate genes for a described CRISPR screen
Dataset
AssayBench test cohort of human CRISPR screens
origin
Independent external evaluation
configuration
Primary source as retrieved 2026-10-09
protocol id
tgtval-20261009-protocol-debrouwer2026-screen-gene-ranking
dataset version
AssayBench v1
split
Temporal split, test cohort
population
334 benchmark entries from screens published after 2021
inputs
Screen description, significance criterion and ranking objective
adaptation
10 nearest-neighbour training examples in context
metric implementation
AnDCG@100, Precision@100 and dFDR@100 as defined in section 3
aggregation
Mean over benchmark entries in the cohort
budget
One ranked list of 100 genes per entry

Metadata review: source checked. Unreported conditions prevent automatic comparisons.

Reproduction

Split
Temporal split, test cohort
Adaptation
10 nearest-neighbour training examples in context
Scoring implementation
AnDCG@100, Precision@100 and dFDR@100 as defined in section 3

No execution recipe has been verified for this exact configuration and evaluation. A benchmark's general instructions may use different inputs, splits or model settings.

Reproducing this published result requires matching its model configuration, data, split and scorer. Source checking or a successful smoke test does not establish score reproduction.

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

19 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-10-10-6e93f504adfc
Property and statementOriginal source and locationReview and provenance
attributes.comparison.adaptation
10 nearest-neighbour training examples in context
Context-only references
AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents

Original source ↗

Table 3, rows for 'Gemini 3 Pro (Few-shot)', cohort 'test'

Version: arXiv:2605.10876 version 1, posted 2026-05-11; not peer reviewed
Retrieved: 2026-10-09T20:56:07Z

not individually reviewed

No individual claim review recorded

independent paper

Audit details

Field: attributes.comparison.adaptation

Source artifact SHA-256: 805b402c28e0daa186415af202e04df56bf0cd7c7b1f872a6db170ee3c6c623d

Hash scope: pdftotext -layout text layer, parsed by extract/extract_target_validation.py

Inspected artifact

attributes.comparison.aggregation
Mean over benchmark entries in the cohort
Context-only references
AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents

Original source ↗

Table 3, rows for 'Gemini 3 Pro (Few-shot)', cohort 'test'

Version: arXiv:2605.10876 version 1, posted 2026-05-11; not peer reviewed
Retrieved: 2026-10-09T20:56:07Z

not individually reviewed

No individual claim review recorded

independent paper

Audit details

Field: attributes.comparison.aggregation

Source artifact SHA-256: 805b402c28e0daa186415af202e04df56bf0cd7c7b1f872a6db170ee3c6c623d

Hash scope: pdftotext -layout text layer, parsed by extract/extract_target_validation.py

Inspected artifact

attributes.comparison.budget
One ranked list of 100 genes per entry
Context-only references
AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents

Original source ↗

Table 3, rows for 'Gemini 3 Pro (Few-shot)', cohort 'test'

Version: arXiv:2605.10876 version 1, posted 2026-05-11; not peer reviewed
Retrieved: 2026-10-09T20:56:07Z

not individually reviewed

No individual claim review recorded

independent paper

Audit details

Field: attributes.comparison.budget

Source artifact SHA-256: 805b402c28e0daa186415af202e04df56bf0cd7c7b1f872a6db170ee3c6c623d

Hash scope: pdftotext -layout text layer, parsed by extract/extract_target_validation.py

Inspected artifact

attributes.comparison.dataset_version
AssayBench v1
Context-only references
AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents

Original source ↗

Table 3, rows for 'Gemini 3 Pro (Few-shot)', cohort 'test'

Version: arXiv:2605.10876 version 1, posted 2026-05-11; not peer reviewed
Retrieved: 2026-10-09T20:56:07Z

not individually reviewed

No individual claim review recorded

independent paper

Audit details

Field: attributes.comparison.dataset_version

Source artifact SHA-256: 805b402c28e0daa186415af202e04df56bf0cd7c7b1f872a6db170ee3c6c623d

Hash scope: pdftotext -layout text layer, parsed by extract/extract_target_validation.py

Inspected artifact

attributes.comparison.inputs
Screen description, significance criterion and ranking objective
Context-only references
AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents

Original source ↗

Table 3, rows for 'Gemini 3 Pro (Few-shot)', cohort 'test'

Version: arXiv:2605.10876 version 1, posted 2026-05-11; not peer reviewed
Retrieved: 2026-10-09T20:56:07Z

not individually reviewed

No individual claim review recorded

independent paper

Audit details

Field: attributes.comparison.inputs

Source artifact SHA-256: 805b402c28e0daa186415af202e04df56bf0cd7c7b1f872a6db170ee3c6c623d

Hash scope: pdftotext -layout text layer, parsed by extract/extract_target_validation.py

Inspected artifact

attributes.comparison.metric_implementation
AnDCG@100, Precision@100 and dFDR@100 as defined in section 3
Context-only references
AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents

Original source ↗

Table 3, rows for 'Gemini 3 Pro (Few-shot)', cohort 'test'

Version: arXiv:2605.10876 version 1, posted 2026-05-11; not peer reviewed
Retrieved: 2026-10-09T20:56:07Z

not individually reviewed

No individual claim review recorded

independent paper

Audit details

Field: attributes.comparison.metric_implementation

Source artifact SHA-256: 805b402c28e0daa186415af202e04df56bf0cd7c7b1f872a6db170ee3c6c623d

Hash scope: pdftotext -layout text layer, parsed by extract/extract_target_validation.py

Inspected artifact

attributes.comparison.population
334 benchmark entries from screens published after 2021
Context-only references
AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents

Original source ↗

Table 3, rows for 'Gemini 3 Pro (Few-shot)', cohort 'test'

Version: arXiv:2605.10876 version 1, posted 2026-05-11; not peer reviewed
Retrieved: 2026-10-09T20:56:07Z

not individually reviewed

No individual claim review recorded

independent paper

Audit details

Field: attributes.comparison.population

Source artifact SHA-256: 805b402c28e0daa186415af202e04df56bf0cd7c7b1f872a6db170ee3c6c623d

Hash scope: pdftotext -layout text layer, parsed by extract/extract_target_validation.py

Inspected artifact

attributes.comparison.protocol_id
tgtval-20261009-protocol-debrouwer2026-screen-gene-ranking
Context-only references
AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents

Original source ↗

Table 3, rows for 'Gemini 3 Pro (Few-shot)', cohort 'test'

Version: arXiv:2605.10876 version 1, posted 2026-05-11; not peer reviewed
Retrieved: 2026-10-09T20:56:07Z

not individually reviewed

No individual claim review recorded

independent paper

Audit details

Field: attributes.comparison.protocol_id

Source artifact SHA-256: 805b402c28e0daa186415af202e04df56bf0cd7c7b1f872a6db170ee3c6c623d

Hash scope: pdftotext -layout text layer, parsed by extract/extract_target_validation.py

Inspected artifact

attributes.comparison.split
Temporal split, test cohort
Context-only references
AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents

Original source ↗

Table 3, rows for 'Gemini 3 Pro (Few-shot)', cohort 'test'

Version: arXiv:2605.10876 version 1, posted 2026-05-11; not peer reviewed
Retrieved: 2026-10-09T20:56:07Z

not individually reviewed

No individual claim review recorded

independent paper

Audit details

Field: attributes.comparison.split

Source artifact SHA-256: 805b402c28e0daa186415af202e04df56bf0cd7c7b1f872a6db170ee3c6c623d

Hash scope: pdftotext -layout text layer, parsed by extract/extract_target_validation.py

Inspected artifact

attributes.limitations
1 values
  • Reported by the benchmark's authors for a system they did not build.
Context-only references
AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents

Original source ↗

Table 3, rows for 'Gemini 3 Pro (Few-shot)', cohort 'test'

Version: arXiv:2605.10876 version 1, posted 2026-05-11; not peer reviewed
Retrieved: 2026-10-09T20:56:07Z

not individually reviewed

No individual claim review recorded

independent paper

Audit details

Field: attributes.limitations

Source artifact SHA-256: 805b402c28e0daa186415af202e04df56bf0cd7c7b1f872a6db170ee3c6c623d

Hash scope: pdftotext -layout text layer, parsed by extract/extract_target_validation.py

Inspected artifact

Sources and history

Release 2026-10-10-6e93f504adfc · Record review: source checked

1 source records and release historyDownload this release (gzip)
Technical metadata and extraction receipts

Stable ID: tgtval-20261009-eval-debrouwer2026-gemini-3-pro-few-shot-test

areas
cells-tissues
contexts
research
origin
independent_paper
protocol
tgtval-20261009-protocol-debrouwer2026-screen-gene-ranking
version
Primary source as retrieved 2026-10-09
comparison
protocol id: tgtval-20261009-protocol-debrouwer2026-screen-gene-ranking; dataset version: AssayBench v1; split: Temporal split, test cohort; population: 334 benchmark entries from screens published after 2021; inputs: Screen description, significance criterion and ranking objective; adaptation: 10 nearest-neighbour training examples in context; metric implementation: AnDCG@100, Precision@100 and dFDR@100 as defined in section 3; aggregation: Mean over benchmark entries in the cohort; budget: One ranked list of 100 genes per entry
source locator
Table 3, rows for 'Gemini 3 Pro (Few-shot)', cohort 'test'
limitations
Reported by the benchmark's authors for a system they did not build.
Related records

Suggest a correction