rewire.itbenchmarks
Benchmark

GENEB

GENEB compares frozen genomic representations across classification tasks and label-budget regimes.

Sourcesdarlednik/GENEB official source · Pinned README: Overview; data regimes; leaderboard aggregation definition

22 evaluations · 22 results

Overview

Datasets

DNA classification tasks grouped into functional categories such as promoters, enhancers, methylation and splice sites.

Metrics

MCC; macro aggregation averages category scores, while the labelled micro aggregate averages task scores.

Allowed inputs

DNA sequences and classification labels.

Sourcesdarlednik/GENEB official source · Pinned README: Overview; data regimes; leaderboard aggregation definition
Evaluation procedure diagram
How it worksEvaluation procedure
Evaluation procedure1. Allowed inputs: DNA sequences and classification labels.. Then: 2. Splits: Full-data, ten-shot and one-shot settings use a common frozen-embedding protocol.. Then: 3. Metrics: MCC; macro aggregation averages category scores, while the labelled micro aggregate averages task scores.Evaluation procedure1. Allowed inputs: DNA sequences and classification labels.. Then: 2. Splits: Full-data, ten-shot and one-shot settings use a common frozen-embedding protocol.. Then: 3. Metrics: MCC; macro aggregation averages category scores, while the labelled micro aggregate averages task scores.Evaluation procedure1. Allowed inputs: DNA sequences and classification labels.. Then: 2. Splits: Full-data, ten-shot and one-shot settings use a common frozen-embedding protocol.. Then: 3. Metrics: MCC; macro aggregation averages category scores, while the labelled micro aggregate averages task scores.

Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.

Sourcesdarlednik/GENEB official source · Pinned README: Overview; data regimes; leaderboard aggregation definition

limited source coverage · Automated source review, 2026-09-16. All specifications and missing details

Results

Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.

GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe

macro_mcc (correlation) · Higher values are better.

GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe · GENEB representative task subset (GENEB split)

Evidence origin: Author-reported evaluation.

GENEB: Why Genomic Models Are Hard to Compare · Table 8, column(Linear MCC)
  • These are averages over the thirteen representative tasks, not a score on any single task.
  • The paper's own point is that these rankings shift with the probe and the protocol, so a position here is not a general ranking.
Comparison details and limitations

Every method GENEB reports on Average macro-MCC across the 13 representative tasks, linear probe, scored with Macro-MCC on GENEB representative task subset.

  • Author-reported numbers, source checked but not independently reproduced.

Automated source review: 2026-09-18. Numerical source review does not establish independent reproduction.

Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.

Showing 11 of 11 matching rows.

Methods and evaluation design

Procedure, tasks and evaluated configurations

How it works

Evaluation methodology

GENEB compares frozen genomic model representations under standardized probe and sample-budget settings. It keeps preprocessing, partitions and random seeds consistent across models and summarizes several biological task categories. Organism and dataset coverage are uneven, so category-specific results remain necessary.

Sourcesgeneb primary benchmark evidence · Evaluation protocol; benchmark construction appendix; Limitations

Evaluation design

Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.

These source-backed links do not make different protocols or scores interchangeable.

Baseline coverage

Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.

No concrete protocols are explicitly linked to this suite. Protocol identification and baseline selection are outstanding.

Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums

Coverage is derived from release 2026-09-29-06401fd5b220. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.

Run this benchmark

The official harness documents pinned task-data retrieval and a five-task smoke-test invocation. Its example uses a user-supplied extractor/module and model identity; it cannot run unchanged without implementing that extractor. Full reference extractors are linked on a separate dev branch and need their own revision/dependency pin.

A maintained rewire runner has not been verified for this benchmark. Check data access, weights, licences, dependencies and hardware in the linked official documentation; requirements have not been fully extracted.

darlednik/GENEB / README.md · README.md lines 22–25 and 127–167 (Evaluate your model and Evaluation protocol)
Strengths, limitations and unresolved questions

Strengths and limitations

Strengths and considerations

  • Task and functional-group aggregation distinguish data-rich groups from balanced coverage.
    Sourcesdarlednik/GENEB official source · Pinned README: Overview; data regimes; leaderboard aggregation definition

Limitations and conditions

  • The paper acknowledges concentration on human and well-studied organisms. A consistent probe protocol cannot establish an unseen-genome boundary for every pretrained model.
    Sourcesgeneb primary benchmark evidence · Evaluation protocol; benchmark construction appendix; Limitations
Profile review details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Stable record: discovery-benchmark-geneb

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetsDNA classification tasks grouped into functional categories such as promoters, enhancers, methylation and splice sites.
Sourcesdarlednik/GENEB official source · Pinned README: Overview; data regimes; leaderboard aggregation definition
SplitsFull-data, ten-shot and one-shot settings use a common frozen-embedding protocol.
Sourcesdarlednik/GENEB official source · Pinned README: Overview; data regimes; leaderboard aggregation definition
MetricsMCC; macro aggregation averages category scores, while the labelled micro aggregate averages task scores.
Sourcesdarlednik/GENEB official source · Pinned README: Overview; data regimes; leaderboard aggregation definition
BaselinesA broad set of genomic foundation models evaluated with the same representation protocol.
Sourcesdarlednik/GENEB official source · Pinned README: Overview; data regimes; leaderboard aggregation definition
Leakage controlsThe protocol fixes preprocessing, partitions and seeds across models, but the inspected paper does not establish a benchmark-wide homology exclusion or pretraining-contamination audit. Shared evaluation settings are not proof that training corpora exclude test sequences. · Not reported in inspected sources
Sourcesgeneb primary benchmark evidence · Evaluation protocol; benchmark construction appendix; Limitations
UncertaintyThe stated protocol averages over five fixed random seeds for probing in the 1-shot, 10-shot and full-data regimes. Averaging seeds is not itself a confidence interval.
Sourcesgeneb primary benchmark evidence · Evaluation protocol; benchmark construction appendix; Limitations
Entity typeDNA-model classification benchmark suite.
Sourcesdarlednik/GENEB official source · Pinned README: Overview; data regimes; leaderboard aggregation definition
OrganismsThe collection combines tasks from human and other well-studied organisms, including mouse and plant categories. Its limitations explicitly note this organism bias; individual dataset provenance remains the appropriate species definition.
Sourcesgeneb primary benchmark evidence · Evaluation protocol; benchmark construction appendix; Limitations
AssaysPromoter, enhancer, methylation and splice-site classification labels.
Sourcesdarlednik/GENEB official source · Pinned README: Overview; data regimes; leaderboard aggregation definition
Allowed inputsDNA sequences and classification labels.
Sourcesdarlednik/GENEB official source · Pinned README: Overview; data regimes; leaderboard aggregation definition
AdaptationBenchmark-specific supervised evaluation of pretrained DNA models.
Sourcesdarlednik/GENEB official source · Pinned README: Overview; data regimes; leaderboard aggregation definition
Applicable tests and references

Applicability is distinct from a completed evaluation.

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.

Paper or primary resourceVersionReference
GENEB: Why Genomic Models Are Hard to Compare2606.04525v1Read source
Historical gaps recorded on 2026-09-17

The catalogue now holds 22 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.

  • Broad benchmark suite with task-specific cohorts and adaptations. Main artifact acquired; complete per-task table and supplement extraction remains pending.
Search and extraction details

source found structured extraction pending

Searches

  • GENEB primary paper benchmark results

Evidence locations

  • arXiv2606.04525v1 main tables and benchmark protocols

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

18 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Individual claims
darlednik/GENEB official source

Original source ↗

Pinned README: Overview; data regimes; leaderboard aggregation definition

Version: 9642d481e40c0af23995dcd162b779613f789f97
Retrieved: 2026-09-16T10:30:21.083930+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 0f2074b432ba2516d0ea05fafab833e953201a3fbe0129c5b6beed9f5f089e81

Hash scope: Hash scope not separately documented; inspect source record

Diagram steps
  • Allowed inputs: DNA sequences and classification labels.
  • Splits: Full-data, ten-shot and one-shot settings use a common frozen-embedding protocol.
  • Metrics: MCC; macro aggregation averages category scores, while the labelled micro aggregate averages task scores.
Individual claims
darlednik/GENEB official source

Original source ↗

Pinned README: Overview; data regimes; leaderboard aggregation definition

Version: 9642d481e40c0af23995dcd162b779613f789f97
Retrieved: 2026-09-16T10:30:21.083930+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 0f2074b432ba2516d0ea05fafab833e953201a3fbe0129c5b6beed9f5f089e81

Hash scope: Hash scope not separately documented; inspect source record

Diagram title
Evaluation procedure
Individual claims
darlednik/GENEB official source

Original source ↗

Pinned README: Overview; data regimes; leaderboard aggregation definition

Version: 9642d481e40c0af23995dcd162b779613f789f97
Retrieved: 2026-09-16T10:30:21.083930+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: 0f2074b432ba2516d0ea05fafab833e953201a3fbe0129c5b6beed9f5f089e81

Hash scope: Hash scope not separately documented; inspect source record

Datasets
DNA classification tasks grouped into functional categories such as promoters, enhancers, methylation and splice sites.
Individual claims
darlednik/GENEB official source

Original source ↗

Pinned README: Overview; data regimes; leaderboard aggregation definition

Version: 9642d481e40c0af23995dcd162b779613f789f97
Retrieved: 2026-09-16T10:30:21.083930+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: 0f2074b432ba2516d0ea05fafab833e953201a3fbe0129c5b6beed9f5f089e81

Hash scope: Hash scope not separately documented; inspect source record

Splits
Full-data, ten-shot and one-shot settings use a common frozen-embedding protocol.
Individual claims
darlednik/GENEB official source

Original source ↗

Pinned README: Overview; data regimes; leaderboard aggregation definition

Version: 9642d481e40c0af23995dcd162b779613f789f97
Retrieved: 2026-09-16T10:30:21.083930+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: 0f2074b432ba2516d0ea05fafab833e953201a3fbe0129c5b6beed9f5f089e81

Hash scope: Hash scope not separately documented; inspect source record

Adaptation
Benchmark-specific supervised evaluation of pretrained DNA models.
Individual claims
darlednik/GENEB official source

Original source ↗

Pinned README: Overview; data regimes; leaderboard aggregation definition

Version: 9642d481e40c0af23995dcd162b779613f789f97
Retrieved: 2026-09-16T10:30:21.083930+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: 0f2074b432ba2516d0ea05fafab833e953201a3fbe0129c5b6beed9f5f089e81

Hash scope: Hash scope not separately documented; inspect source record

Metrics
MCC; macro aggregation averages category scores, while the labelled micro aggregate averages task scores.
Individual claims
darlednik/GENEB official source

Original source ↗

Pinned README: Overview; data regimes; leaderboard aggregation definition

Version: 9642d481e40c0af23995dcd162b779613f789f97
Retrieved: 2026-09-16T10:30:21.083930+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: 0f2074b432ba2516d0ea05fafab833e953201a3fbe0129c5b6beed9f5f089e81

Hash scope: Hash scope not separately documented; inspect source record

Baselines
A broad set of genomic foundation models evaluated with the same representation protocol.
Individual claims
darlednik/GENEB official source

Original source ↗

Pinned README: Overview; data regimes; leaderboard aggregation definition

Version: 9642d481e40c0af23995dcd162b779613f789f97
Retrieved: 2026-09-16T10:30:21.083930+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.3.value

Source artifact SHA-256: 0f2074b432ba2516d0ea05fafab833e953201a3fbe0129c5b6beed9f5f089e81

Hash scope: Hash scope not separately documented; inspect source record

Leakage controls
The protocol fixes preprocessing, partitions and seeds across models, but the inspected paper does not establish a benchmark-wide homology exclusion or pretraining-contamination audit. Shared evaluation settings are not proof that training corpora exclude test sequences.
Individual claims
geneb primary benchmark evidence

Original source ↗

Evaluation protocol; benchmark construction appendix; Limitations

Version: 2606.04525v1
Retrieved: 2026-09-16T21:06:29.746844+00:00

unreported

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: 47975089c0ca738d5e2d6e6ea91cd4e7b80498c3f175803d8ec39ca77aa41629

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Uncertainty
The stated protocol averages over five fixed random seeds for probing in the 1-shot, 10-shot and full-data regimes. Averaging seeds is not itself a confidence interval.
Individual claims
geneb primary benchmark evidence

Original source ↗

Evaluation protocol; benchmark construction appendix; Limitations

Version: 2606.04525v1
Retrieved: 2026-09-16T21:06:29.746844+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.5.value

Source artifact SHA-256: 47975089c0ca738d5e2d6e6ea91cd4e7b80498c3f175803d8ec39ca77aa41629

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: discovered

4 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: discovery-benchmark-geneb

areas
genomics
entity level
suite
scope note
Specialist molecular or omics evaluation; protocol details require review before numerical comparison.
task
Frozen genomic representations with linear probes
version
Not reported
benchmark research
review date: 2026-09-17; status: source_found_structured_extraction_pending; primary sources: evidence-expansion-p2-evidence-discovery-final-geneb-47975089c0ca; inspected locators: arXiv2606.04525v1 main tables and benchmark protocols; searched queries: GENEB primary paper benchmark results; gaps: Broad benchmark suite with task-specific cohorts and adaptations. Main artifact acquired; complete per-task table and supplement extraction remains pending.; claim scope: Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.
historical missing metadata
dataset release: unextracted; metric implementation: unextracted; split manifest: unextracted; version: unextracted
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
entity classification
review date: 2026-09-17; rationale: The cited profile describes a collection of evaluation tasks or protocols; retain it as the top-level benchmark suite. Its datasets and individual protocols remain separate records.; source ids: src-discovery-darlednik-geneb; source locator: Pinned README: Overview; data regimes; leaderboard aggregation definition; ambiguities: None recorded
run documentation
record id: discovery-benchmark-geneb; source ids: run-doc-geneb-readme-md-9642d481; status: official_documentation_linked; summary: The official harness documents pinned task-data retrieval and a five-task smoke-test invocation. Its example uses a user-supplied extractor/module and model identity; it cannot run unchanged without implementing that extractor. Full reference extractors are linked on a separate dev branch and need their own revision/dependency pin.; source locator: README.md lines 22–25 and 127–167 (Evaluate your model and Evaluation protocol)
Related records

Suggest a correction