rewire.itbenchmarks
Dataset subset

GENEB representative task subset (GENEB split)

The split of GENEB representative task subset that GENEB evaluated on. The upstream dataset release is not catalogued here, so no claim is made that this matches its original splits.

Research readiness

These checks assess whether the evidence supports a reproducible investigation. A source-checked score alone does not meet these requirements.

Release 2026-09-29-06401fd5b220 · Evidence verified: Not verified

Evidence incomplete

Replay metrics

Exact outcomes, predictions, identifiers and evaluator are connected.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • metric replay: verification is missing

Verified: Not verified

Evidence incomplete

Investigate discrepancies

Replay evidence includes annotations and an assessment of dependence. Unknown independence permits descriptive analysis only.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • metric replay: verification is missing
  • annotations: verification is missing
  • dependence: verification is missing

Verified: Not verified

Evidence incomplete

Run locally

A pinned recipe describes the inputs, environment and resource requirements.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • recipe pinned: verification is missing
  • resource estimate: verification is missing

Verified: Not verified

Evidence incomplete

Validate independently

Separate data and exposure records support an independent test.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • independent validation: verification is missing
  • overlap checked: verification is missing

Verified: Not verified

Readiness describes the evidence in this release. Availability on your computer is checked separately when an investigation runs. Existing data exposure can prevent independent validation even when files are available.

Artifacts and reproduction

No verified artifact manifest is connected to this record yet. The gaps above identify what is needed before analysis can begin.

Read reviewed discrepancy investigations

Evaluation results

22 evaluations · 22 results. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: GENA-LM-Large-T2TTask: GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe
Dataset subset: GENEB representative task subset (GENEB split)
0.53 macro_mcc
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GENA-LM-Large-T2T on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe

Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7).

Aggregation: Not reported

GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GENA-LM-Large-T2T), column(Linear MCC)
Configuration: GENA-LM-Large-T2TTask: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe
Dataset subset: GENEB representative task subset (GENEB split)
0.535 macro_mcc
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GENA-LM-Large-T2T on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe

Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7).

Aggregation: Not reported

GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GENA-LM-Large-T2T), column(MLP MCC)
Configuration: GENERator-Eukaryote-1.2BTask: GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe
Dataset subset: GENEB representative task subset (GENEB split)
0.579 macro_mcc
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GENERator-Eukaryote-1.2B on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe

Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7).

Aggregation: Not reported

GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GENERator-Eukaryote-1.2B), column(Linear MCC)
Configuration: GENERator-Eukaryote-1.2BTask: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe
Dataset subset: GENEB representative task subset (GENEB split)
0.595 macro_mcc
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GENERator-Eukaryote-1.2B on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe

Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7).

Aggregation: Not reported

GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GENERator-Eukaryote-1.2B), column(MLP MCC)
Configuration: GENERator-Eukaryote-3BTask: GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe
Dataset subset: GENEB representative task subset (GENEB split)
0.605 macro_mcc
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GENERator-Eukaryote-3B on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe

Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7).

Aggregation: Not reported

GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GENERator-Eukaryote-3B), column(Linear MCC)
Configuration: GENERator-Eukaryote-3BTask: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe
Dataset subset: GENEB representative task subset (GENEB split)
0.609 macro_mcc
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GENERator-Eukaryote-3B on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe

Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7).

Aggregation: Not reported

GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GENERator-Eukaryote-3B), column(MLP MCC)
Configuration: GenomeOcean-4BTask: GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe
Dataset subset: GENEB representative task subset (GENEB split)
0.552 macro_mcc
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GenomeOcean-4B on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe

Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7).

Aggregation: Not reported

GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GenomeOcean-4B), column(Linear MCC)
Configuration: GenomeOcean-4BTask: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe
Dataset subset: GENEB representative task subset (GENEB split)
0.552 macro_mcc
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GenomeOcean-4B on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe

Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7).

Aggregation: Not reported

GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GenomeOcean-4B), column(MLP MCC)
Configuration: GenomeOcean-500MTask: GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe
Dataset subset: GENEB representative task subset (GENEB split)
0.536 macro_mcc
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GenomeOcean-500M on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe

Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7).

Aggregation: Not reported

GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GenomeOcean-500M), column(Linear MCC)
Configuration: GenomeOcean-500MTask: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe
Dataset subset: GENEB representative task subset (GENEB split)
0.535 macro_mcc
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GenomeOcean-500M on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe

Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7).

Aggregation: Not reported

GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GenomeOcean-500M), column(MLP MCC)
Configuration: GROVERTask: GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe
Dataset subset: GENEB representative task subset (GENEB split)
0.466 macro_mcc
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GROVER on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe

Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7).

Aggregation: Not reported

GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GROVER), column(Linear MCC)
Configuration: GROVERTask: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe
Dataset subset: GENEB representative task subset (GENEB split)
0.477 macro_mcc
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GROVER on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe

Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7).

Aggregation: Not reported

GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GROVER), column(MLP MCC)
Configuration: HyenaDNA-Large-1MTask: GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe
Dataset subset: GENEB representative task subset (GENEB split)
0.427 macro_mcc
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

HyenaDNA-Large-1M on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe

Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7).

Aggregation: Not reported

GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(HyenaDNA-Large-1M), column(Linear MCC)
Configuration: HyenaDNA-Large-1MTask: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe
Dataset subset: GENEB representative task subset (GENEB split)
0.479 macro_mcc
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

HyenaDNA-Large-1M on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe

Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7).

Aggregation: Not reported

GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(HyenaDNA-Large-1M), column(MLP MCC)
Configuration: LucaOneTask: GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe
Dataset subset: GENEB representative task subset (GENEB split)
0.573 macro_mcc
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

LucaOne on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe

Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7).

Aggregation: Not reported

GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(LucaOne), column(Linear MCC)
Configuration: LucaOneTask: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe
Dataset subset: GENEB representative task subset (GENEB split)
0.6 macro_mcc
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

LucaOne on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe

Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7).

Aggregation: Not reported

GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(LucaOne), column(MLP MCC)
Configuration: MutBERTTask: GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe
Dataset subset: GENEB representative task subset (GENEB split)
0.516 macro_mcc
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

MutBERT on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe

Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7).

Aggregation: Not reported

GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(MutBERT), column(Linear MCC)
Configuration: MutBERTTask: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe
Dataset subset: GENEB representative task subset (GENEB split)
0.517 macro_mcc
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

MutBERT on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe

Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7).

Aggregation: Not reported

GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(MutBERT), column(MLP MCC)
Configuration: NT-v2-50M-MSTask: GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe
Dataset subset: GENEB representative task subset (GENEB split)
0.511 macro_mcc
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-v2-50M-MS on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe

Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7).

Aggregation: Not reported

GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(NT-v2-50M-MS), column(Linear MCC)
Configuration: NT-v2-50M-MSTask: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe
Dataset subset: GENEB representative task subset (GENEB split)
0.521 macro_mcc
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-v2-50M-MS on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe

Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7).

Aggregation: Not reported

GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(NT-v2-50M-MS), column(MLP MCC)
Configuration: Omni-DNA-1BTask: GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe
Dataset subset: GENEB representative task subset (GENEB split)
0.55 macro_mcc
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

Omni-DNA-1B on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe

Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7).

Aggregation: Not reported

GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(Omni-DNA-1B), column(Linear MCC)
Configuration: Omni-DNA-1BTask: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe
Dataset subset: GENEB representative task subset (GENEB split)
0.542 macro_mcc
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

Omni-DNA-1B on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe

Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7).

Aggregation: Not reported

GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(Omni-DNA-1B), column(MLP MCC)

Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.

Subset and evaluation context

This record describes a particular subset or cohort used in an evaluation. Its results do not describe the full dataset.

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

2 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
description
The split of GENEB representative task subset that GENEB evaluated on. The upstream dataset release is not catalogued here, so no claim is made that this matches its original splits.
Context-only references
GENEB: Why Genomic Models Are Hard to Compare

Original source ↗

No field-specific location recorded

Version: 2606.04525v1
Retrieved: 2026-09-17T07:56:09.460928+00:00

not individually reviewed

No individual claim review recorded

Audit details

Field: description

Source artifact SHA-256: 47975089c0ca738d5e2d6e6ea91cd4e7b80498c3f175803d8ec39ca77aa41629

Hash scope: Exact retrieved primary paper artifact bytes.

Inspected artifact

name
GENEB representative task subset (GENEB split)
Context-only references
GENEB: Why Genomic Models Are Hard to Compare

Original source ↗

No field-specific location recorded

Version: 2606.04525v1
Retrieved: 2026-09-17T07:56:09.460928+00:00

not individually reviewed

No individual claim review recorded

Audit details

Field: name

Source artifact SHA-256: 47975089c0ca738d5e2d6e6ea91cd4e7b80498c3f175803d8ec39ca77aa41629

Hash scope: Exact retrieved primary paper artifact bytes.

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: source checked

1 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: geneb-dataset-geneb-representative-task-subset

areas
dna-genomes
missing metadata
version: unreported; url: unextracted
Related records

Suggest a correction