rewire.itbenchmarks
Benchmark

Genomic Benchmarks

Genomic Benchmarks packages genomic sequence-classification datasets with explicit versions and supplied train/test folders.

SourcesML-Bioinfo-CEITEC/genomic_benchmarks official source · Pinned README: repository purpose; info and download_dataset examples

18 evaluations · 36 results

Overview

Evaluation procedure diagram
How it worksEvaluation procedure
Evaluation procedure1. Allowed inputs: Sequences and categorical labels through the dataset loader.. Then: 2. Splits: The download API delivers prepartitioned train/test data organized by class.. Then: 3. Metrics: The original paper reports classification accuracy and F1 for its CNN baseline on each dataset.Evaluation procedure1. Allowed inputs: Sequences and categorical labels through the dataset loader.. Then: 2. Splits: The download API delivers prepartitioned train/test data organized by class.. Then: 3. Metrics: The original paper reports classification accuracy and F1 for its CNN baseline on each dataset.Evaluation procedure1. Allowed inputs: Sequences and categorical labels through the dataset loader.. Then: 2. Splits: The download API delivers prepartitioned train/test data organized by class.. Then: 3. Metrics: The original paper reports classification accuracy and F1 for its CNN baseline on each dataset.

Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.

Sources (2)ML-Bioinfo-CEITEC/genomic_benchmarks official source; genomic-benchmarks primary benchmark evidence · Pinned README: repository purpose; info and download_dataset examples; Methods: Reproducibility and Baseline model; Table 2

limited source coverage · Automated source review, 2026-09-16. All specifications and missing details

Results

Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.

Genomic Benchmarks DEMO-CODING-VS-INTERGENOMIC-SEQS-ACCURACY: demo_coding_vs_intergenomic_seqs, Accuracy

accuracy (percent) · Higher values are better.

Genomic Benchmarks DEMO-CODING-VS-INTERGENOMIC-SEQS-ACCURACY: demo_coding_vs_intergenomic_seqs, Accuracy · demo_coding_vs_intergenomic_seqs (Genomic Benchmarks split)

Evidence origin: Author-reported evaluation.

Genomic benchmarks: a collection of datasets for genomic sequence classification · Table 2, row(demo_coding_vs_intergenomic_seqs)
  • Both columns are the same baseline architecture in two frameworks, not two competing models.
  • These are the paper's reference baselines, not a leaderboard of the best available models.
Comparison details and limitations

Every method Genomic Benchmarks reports on demo_coding_vs_intergenomic_seqs, Accuracy, scored with Accuracy on demo_coding_vs_intergenomic_seqs.

  • Author-reported numbers, source checked but not independently reproduced.

Automated source review: 2026-09-18. Numerical source review does not establish independent reproduction.

Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.

Showing 2 of 2 matching rows.

Tested configuration
0255075100
Reported score
  1. Baseline CNN (TensorFlow)89.6
  2. Baseline CNN (PyTorch)87.6

Methods and evaluation design

Procedure, tasks and evaluated configurations

How it works

Evaluation methodology

Genomic Benchmarks packages sequence-classification datasets together with versioned genomic coordinates, construction notebooks and a small CNN baseline. Each dataset supplies its own train and test subsets. Dataset-specific negative sampling and split provenance matter as much as model architecture for interpreting performance.

Sourcesgenomic-benchmarks primary benchmark evidence · Methods: Reproducibility and Baseline model; Table 2

Evaluation design

Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.

Tasks

These source-backed links do not make different protocols or scores interchangeable.

Baseline coverage

Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.

No concrete protocols are explicitly linked to this suite. Protocol identification and baseline selection are outstanding.

Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums

Coverage is derived from release 2026-09-29-06401fd5b220. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.

Run this benchmark

Load a Genomic Benchmarks dataset

Install the package and load one of the nine sequence classification datasets whose baseline scores this page records.

Generate predictions and evaluate them. This recipe does not establish reproduction of a particular published score.

Dataset access
Downloaded by the package; the splits are fixed by the release.
Model and weights
None. The published baseline is a small convolutional network.
Licences
Project licence: Apache-2.0. Upstream data licences are separate and unreported here.
Software
Python with the genomic-benchmarks package from PyPI.
Hardware
CPU is enough to load the data; training the baseline benefits from a GPU.
Required inputs and expected outputs

Inputs

  • A sequence classifier you want to train and score.

Outputs

  • Train and test splits as the package ships them.

Execution steps

  1. 1. Install (Command line)

    Source reviewed; these instructions have not been executed by rewire.

    pip install genomic-benchmarks
    Genomic Benchmarks: repository README · README.md at 605d8539, Install, lines 13-13
  2. 2. List the datasets (Python)

    Source reviewed; these instructions have not been executed by rewire.

    >>> from genomic_benchmarks.data_check import list_datasets
    >>> 
    >>> list_datasets()
    ['demo_coding_vs_intergenomic_seqs', 'demo_human_or_worm', 'dummy_mouse_enhancers_ensembl', 'human_enhancers_cohn', 'human_enhancers_ensembl', 'human_ensembl_regulatory',  'human_nontata_promoters', 'human_ocr_ensembl']
    Genomic Benchmarks: repository README · README.md at 605d8539, Usage, lines 40-43

Use your own model

Run your model locally and return predictions keyed by the input IDs. The evaluator supplies biological inputs without test labels and owns scoring. This interface is not a sandbox for model code.

Pass your existing prediction function into this adapter. Its output direction must match the selected protocol.

class MyModelAdapter:
    def __init__(self, score):
        self.score = score

    def predict(self, inputs):
        return {row["id"]: float(self.score(row)) for row in inputs}

# adapter = MyModelAdapter(your_prediction_function)
# report = rewirebench.run(prepared, adapter, output="runs/my-model")

Alternatively, generate a keyed prediction file in your existing model environment and use the score-only recipe. Your model code and weights do not need to be shared.

Genomic Benchmarks: repository README · README.md at 605d8539
Scope and limitations
  • Quoted from the project's README and not executed by rewire, so the commands are evidence of what the project documents rather than a verified run.
  • The project may have changed since the pinned commit.
  • The scores here are the paper's own baseline in two frameworks, not a leaderboard.

Contribute a result for review. The library can submit an exported evaluation for private review when intake is open. Check the contribution page for access and sign-in.

Original repository instructions

Run this benchmark

Official data-loading examples and linked TensorFlow/PyTorch training notebooks are available. Dataset download alone produces train/test sequence files rather than a model score. The optional dependency examples contain unquoted >= specifiers, so do not copy those lines directly into a shell without quoting the version requirement.

A maintained rewire runner has not been verified for this benchmark. Check data access, weights, licences, dependencies and hardware in the linked official documentation; requirements have not been fully extracted.

ML-Bioinfo-CEITEC/genomic_benchmarks / README.md · README.md lines 8–34 and 36–108 (Install and Usage)
Strengths, limitations and unresolved questions

Strengths and limitations

Limitations and conditions

  • Non-overlapping positive and negative genomic intervals do not guarantee independence between related sequences across the train/test boundary. Baseline implementation differences and data versions should remain explicit.
    Sourcesgenomic-benchmarks primary benchmark evidence · Methods: Reproducibility and Baseline model; Table 2
Profile review details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Stable record: discovery-benchmark-genomic-benchmarks

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetsNamed genomic classification datasets; the README illustrates a non-TATA human-promoter task.
SourcesML-Bioinfo-CEITEC/genomic_benchmarks official source · Pinned README: repository purpose; info and download_dataset examples
SplitsThe download API delivers prepartitioned train/test data organized by class.
SourcesML-Bioinfo-CEITEC/genomic_benchmarks official source · Pinned README: repository purpose; info and download_dataset examples
MetricsThe original paper reports classification accuracy and F1 for its CNN baseline on each dataset.
Sourcesgenomic-benchmarks primary benchmark evidence · Methods: Reproducibility and Baseline model; Table 2
BaselinesThe repository provides neural-network training helpers and links experiment reports.
SourcesML-Bioinfo-CEITEC/genomic_benchmarks official source · Pinned README: repository purpose; info and download_dataset examples
Leakage controlsDataset-construction notebooks use fixed seeds. Generated negative regions match positive lengths and are rejected if they overlap positives; this does not establish a suite-wide chromosome-held-out or homology-filtered partition.
Sourcesgenomic-benchmarks primary benchmark evidence · Methods: Reproducibility and Baseline model; Table 2
UncertaintyTable 2 gives point estimates for PyTorch and TensorFlow baseline implementations. The inspected Methods do not specify repeated-training or bootstrap uncertainty for those values. · Not reported in inspected sources
Sourcesgenomic-benchmarks primary benchmark evidence · Methods: Reproducibility and Baseline model; Table 2
Entity typeRepository of genomic sequence classification datasets.
SourcesML-Bioinfo-CEITEC/genomic_benchmarks official source · Pinned README: repository purpose; info and download_dataset examples
OrganismsDataset-specific organisms; a human non-TATA promoter dataset is documented in the README.
SourcesML-Bioinfo-CEITEC/genomic_benchmarks official source · Pinned README: repository purpose; info and download_dataset examples
AssaysCurated genomic classification labels from linked source datasets.
SourcesML-Bioinfo-CEITEC/genomic_benchmarks official source · Pinned README: repository purpose; info and download_dataset examples
Allowed inputsSequences and categorical labels through the dataset loader.
SourcesML-Bioinfo-CEITEC/genomic_benchmarks official source · Pinned README: repository purpose; info and download_dataset examples
AdaptationSupervised train/test classification with a documented CNN example.
SourcesML-Bioinfo-CEITEC/genomic_benchmarks official source · Pinned README: repository purpose; info and download_dataset examples
Applicable tests and references

Applicability is distinct from a completed evaluation.

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Primary-paper discovery and source inspection. Source-checked results are not independently reproduced experiments.

Paper or primary resourceVersionReference
Genomic benchmarks: a collection of datasets for genomic sequence classificationPMC10150520Read source
DOI: 10.1186/s12863-023-01123-8
Historical gaps recorded on 2026-09-17

The catalogue now holds 36 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.

  • The two implementations are not distinct foundation-model families.
  • Dataset totals in Table 1 are not test-set denominators.
  • Table 2 does not enumerate dataset-version numbers, split file hashes or replicate uncertainty.
Search and extraction details

primary comparison table screened

Searches

  • Genomic Benchmarks collection genomic sequence classification PMC10150520

Evidence locations

  • Table 2
  • Training models section
  • Dataset construction and Table 1

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

42 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Individual claims
genomic-benchmarks primary benchmark evidence

Original source ↗

Pinned README: repository purpose; info and download_dataset examples; Methods: Reproducibility and Baseline model; Table 2

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: PMC10150520
Retrieved: 2026-09-16T21:07:13.681402+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: bda6fe51e3363a5d2fc8d265ca536897d3e83eb76458fc95066c21933e3bd0c0

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram caption
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Individual claims
ML-Bioinfo-CEITEC/genomic_benchmarks official source

Original source ↗

Pinned README: repository purpose; info and download_dataset examples; Methods: Reproducibility and Baseline model; Table 2

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 605d8539830e16c85abe7826990958303ffc5e1c
Retrieved: 2026-09-16T10:30:20.990833+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 926f0f196439564b095cbabe65f6a0acee3f22afbb0f6c8fe50bb3964082d15b

Hash scope: Hash scope not separately documented; inspect source record

Diagram steps
  • Allowed inputs: Sequences and categorical labels through the dataset loader.
  • Splits: The download API delivers prepartitioned train/test data organized by class.
  • Metrics: The original paper reports classification accuracy and F1 for its CNN baseline on each dataset.
Individual claims
genomic-benchmarks primary benchmark evidence

Original source ↗

Pinned README: repository purpose; info and download_dataset examples; Methods: Reproducibility and Baseline model; Table 2

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: PMC10150520
Retrieved: 2026-09-16T21:07:13.681402+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: bda6fe51e3363a5d2fc8d265ca536897d3e83eb76458fc95066c21933e3bd0c0

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps
  • Allowed inputs: Sequences and categorical labels through the dataset loader.
  • Splits: The download API delivers prepartitioned train/test data organized by class.
  • Metrics: The original paper reports classification accuracy and F1 for its CNN baseline on each dataset.
Individual claims
ML-Bioinfo-CEITEC/genomic_benchmarks official source

Original source ↗

Pinned README: repository purpose; info and download_dataset examples; Methods: Reproducibility and Baseline model; Table 2

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 605d8539830e16c85abe7826990958303ffc5e1c
Retrieved: 2026-09-16T10:30:20.990833+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 926f0f196439564b095cbabe65f6a0acee3f22afbb0f6c8fe50bb3964082d15b

Hash scope: Hash scope not separately documented; inspect source record

Diagram title
Evaluation procedure
Individual claims
genomic-benchmarks primary benchmark evidence

Original source ↗

Pinned README: repository purpose; info and download_dataset examples; Methods: Reproducibility and Baseline model; Table 2

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: PMC10150520
Retrieved: 2026-09-16T21:07:13.681402+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: bda6fe51e3363a5d2fc8d265ca536897d3e83eb76458fc95066c21933e3bd0c0

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title
Evaluation procedure
Individual claims
ML-Bioinfo-CEITEC/genomic_benchmarks official source

Original source ↗

Pinned README: repository purpose; info and download_dataset examples; Methods: Reproducibility and Baseline model; Table 2

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 605d8539830e16c85abe7826990958303ffc5e1c
Retrieved: 2026-09-16T10:30:20.990833+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: 926f0f196439564b095cbabe65f6a0acee3f22afbb0f6c8fe50bb3964082d15b

Hash scope: Hash scope not separately documented; inspect source record

Datasets
Named genomic classification datasets; the README illustrates a non-TATA human-promoter task.
Individual claims
ML-Bioinfo-CEITEC/genomic_benchmarks official source

Original source ↗

Pinned README: repository purpose; info and download_dataset examples

Version: 605d8539830e16c85abe7826990958303ffc5e1c
Retrieved: 2026-09-16T10:30:20.990833+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: 926f0f196439564b095cbabe65f6a0acee3f22afbb0f6c8fe50bb3964082d15b

Hash scope: Hash scope not separately documented; inspect source record

Splits
The download API delivers prepartitioned train/test data organized by class.
Individual claims
ML-Bioinfo-CEITEC/genomic_benchmarks official source

Original source ↗

Pinned README: repository purpose; info and download_dataset examples

Version: 605d8539830e16c85abe7826990958303ffc5e1c
Retrieved: 2026-09-16T10:30:20.990833+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: 926f0f196439564b095cbabe65f6a0acee3f22afbb0f6c8fe50bb3964082d15b

Hash scope: Hash scope not separately documented; inspect source record

Adaptation
Supervised train/test classification with a documented CNN example.
Individual claims
ML-Bioinfo-CEITEC/genomic_benchmarks official source

Original source ↗

Pinned README: repository purpose; info and download_dataset examples

Version: 605d8539830e16c85abe7826990958303ffc5e1c
Retrieved: 2026-09-16T10:30:20.990833+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: 926f0f196439564b095cbabe65f6a0acee3f22afbb0f6c8fe50bb3964082d15b

Hash scope: Hash scope not separately documented; inspect source record

Metrics
The original paper reports classification accuracy and F1 for its CNN baseline on each dataset.
Individual claims
genomic-benchmarks primary benchmark evidence

Original source ↗

Methods: Reproducibility and Baseline model; Table 2

Version: PMC10150520
Retrieved: 2026-09-16T21:07:13.681402+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: bda6fe51e3363a5d2fc8d265ca536897d3e83eb76458fc95066c21933e3bd0c0

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: discovered

9 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: discovery-benchmark-genomic-benchmarks

areas
genomics
entity level
suite
scope note
Specialist molecular or omics evaluation; protocol details require review before numerical comparison.
task
Genomic sequence classification
version
Not reported
benchmark research
review date: 2026-09-17; status: primary_comparison_table_screened; primary sources: expansion-p3-genomic-benchmarks; inspected locators: Table 2; Training models section; Dataset construction and Table 1; searched queries: Genomic Benchmarks collection genomic sequence classification PMC10150520; gaps: The two implementations are not distinct foundation-model families.; Dataset totals in Table 1 are not test-set denominators.; Table 2 does not enumerate dataset-version numbers, split file hashes or replicate uncertainty.; claim scope: Primary-paper discovery and source inspection. Source-checked results are not independently reproduced experiments.
historical missing metadata
dataset release: unextracted; metric implementation: unextracted; split manifest: unextracted; version: unextracted
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
entity classification
review date: 2026-09-17; rationale: The cited profile describes a collection of evaluation tasks or protocols; retain it as the top-level benchmark suite. Its datasets and individual protocols remain separate records.; source ids: src-discovery-ml-bioinfo-ceitec-genomic-benchmarks; source locator: Pinned README: repository purpose; info and download_dataset examples; ambiguities: None recorded
run documentation
record id: discovery-benchmark-genomic-benchmarks; source ids: run-doc-genomic-benchmarks-readme-md-605d8539; status: official_documentation_linked; summary: Official data-loading examples and linked TensorFlow/PyTorch training notebooks are available. Dataset download alone produces train/test sequence files rather than a model score. The optional dependency examples contain unquoted >= specifiers, so do not copy those lines directly into a shell without quoting the version requirement.; source locator: README.md lines 8–34 and 36–108 (Install and Usage)
run recipes
id: genomic-benchmarks-official; protocol id: discovery-benchmark-genomic-benchmarks; version: 605d8539830e16c85abe7826990958303ffc5e1c; title: Load a Genomic Benchmarks dataset; purpose: generate_and_evaluate; summary: Install the package and load one of the nine sequence classification datasets whose baseline scores this page records.; inputs: A sequence classifier you want to train and score.; outputs: Train and test splits as the package ships them.; requirements: data: Downloaded by the package; the splits are fixed by the release.; weights: None. The published baseline is a small convolutional network.; licence: Project licence: Apache-2.0. Upstream data licences are separate and unreported here.; software: Python with the genomic-benchmarks package from PyPI.; hardware: CPU is enough to load the data; training the baseline benefits from a GPU.; instructions: runtime: command_line; title: Install; code: pip install genomic-benchmarks; status: source_reviewed_not_executed; source ids: project-recipe-genomic-benchmarks-605d8539; source locator: README.md at 605d8539, Install, lines 13-13; runtime: python; title: List the datasets; code: >>> from genomic_benchmarks.data_check import list_datasets >>> >>> list_datasets() ['demo_coding_vs_intergenomic_seqs', 'demo_human_or_worm', 'dummy_mouse_enhancers_ensembl', 'human_enhancers_cohn', 'human_enhancers_ensembl', 'human_ensembl_regulatory', 'human_nontata_promoters', 'human_ocr_ensembl']; status: source_reviewed_not_executed; source ids: project-recipe-genomic-benchmarks-605d8539; source locator: README.md at 605d8539, Usage, lines 40-43; limitations: Quoted from the project's README and not executed by rewire, so the commands are evidence of what the project documents rather than a verified run.; The project may have changed since the pinned commit.; The scores here are the paper's own baseline in two frameworks, not a leaderboard.; source ids: project-recipe-genomic-benchmarks-605d8539; source locator: README.md at 605d8539; id: genomic-benchmarks-rewirebench; protocol id: genomic-benchmarks-v2; version: aecb9e79a2a5e83b59e482212b1a8b812dd16079; title: Score your own classifier with rewirebench; purpose: generate_and_evaluate; summary: Prepare a downloaded dataset and score an adapter on accuracy and F1 over the packaged test split.; inputs: The dataset directory the genomic-benchmarks package downloaded.; An adapter returning a class index, or a probability for a binary dataset.; outputs: A local report with accuracy, F1, the coverage and the digest of the sequences it read.; requirements: data: Whatever the package downloaded locally, read from its train and test directories.; weights: Whatever your own model needs; the runner supplies none.; licence: Runner code is MIT. The benchmark's own data terms are upstream and unreported here.; software: Python 3.11 with the pinned rewirebench environment.; hardware: CPU for scoring. Your own model decides what it needs.; instructions: runtime: command_line; title: Prepare a dataset; code: rewirebench prepare genomic-benchmarks-v2 \ --source ./genomic_benchmarks --options '{"dataset": "human_nontata_promoters"}' \ --output ./prepared-promoters; status: source_reviewed_not_executed; source ids: project-recipe-runner-genomic-benchmarks-genomic-benchmarks-md-aecb9e79; project-recipe-runner-genomic-benchmarks-genomic-benchmarks-py-aecb9e79; source locator: docs/genomic-benchmarks.md at aecb9e79, Prepare, run and score, lines 51-53; runtime: command_line; title: Run and score an adapter; code: rewirebench run --prepared ./prepared-promoters \ --adapter my_models.dna:MyAdapter \ --model-name 'my model' --training-overlap 'unreported' \ --output ./scored-promoters; status: source_reviewed_not_executed; source ids: project-recipe-runner-genomic-benchmarks-genomic-benchmarks-md-aecb9e79; project-recipe-runner-genomic-benchmarks-genomic-benchmarks-py-aecb9e79; source locator: docs/genomic-benchmarks.md at aecb9e79, Prepare, run and score, lines 61-64; limitations: Quoted from the runner's documentation and not executed by this repository.; Scoring follows the benchmark's own evaluator; running it does not by itself reproduce a published number.; Class labels are assigned from the sorted class name, because upstream takes them from filesystem order, which differs between machines. Compare class names, not indices.; For the one dataset with three classes the paper does not say which F1 averaging it used, so macro and weighted are both reported and neither is the paper's number.; Genomic Benchmarks v1 prepared artifacts exposed class information in row IDs and must be discarded. V2 uses opaque IDs and label-independent ordering; prepare the dataset again.; This scores a hashed local data copy; it does not certify the official cohort or reproduce a published experiment.; source ids: project-recipe-runner-genomic-benchmarks-genomic-benchmarks-md-aecb9e79; project-recipe-runner-genomic-benchmarks-genomic-benchmarks-py-aecb9e79; source locator: docs/genomic-benchmarks.md at aecb9e79
Related records

Suggest a correction