Datasets
Named genomic classification datasets; the README illustrates a non-TATA human-promoter task.
Genomic Benchmarks packages genomic sequence-classification datasets with explicit versions and supplied train/test folders.
Named genomic classification datasets; the README illustrates a non-TATA human-promoter task.
The original paper reports classification accuracy and F1 for its CNN baseline on each dataset.
Sequences and categorical labels through the dataset loader.
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
limited source coverage · Automated source review, 2026-09-16. All specifications and missing details
Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.
accuracy (percent) · Higher values are better.
Genomic Benchmarks DEMO-CODING-VS-INTERGENOMIC-SEQS-ACCURACY: demo_coding_vs_intergenomic_seqs, Accuracy · demo_coding_vs_intergenomic_seqs (Genomic Benchmarks split)
Evidence origin: Author-reported evaluation.
Genomic benchmarks: a collection of datasets for genomic sequence classification · Table 2, row(demo_coding_vs_intergenomic_seqs)Every method Genomic Benchmarks reports on demo_coding_vs_intergenomic_seqs, Accuracy, scored with Accuracy on demo_coding_vs_intergenomic_seqs.
Automated source review: 2026-09-18. Numerical source review does not establish independent reproduction.
Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.
Showing 2 of 2 matching rows.
Genomic Benchmarks packages sequence-classification datasets together with versioned genomic coordinates, construction notebooks and a small CNN baseline. Each dataset supplies its own train and test subsets. Dataset-specific negative sampling and split provenance matter as much as model architecture for interpreting performance.
Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.
These source-backed links do not make different protocols or scores interchangeable.
Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.
No concrete protocols are explicitly linked to this suite. Protocol identification and baseline selection are outstanding.
Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums
Coverage is derived from release 2026-09-29-06401fd5b220. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.
Install the package and load one of the nine sequence classification datasets whose baseline scores this page records.
Generate predictions and evaluate them. This recipe does not establish reproduction of a particular published score.
Source reviewed; these instructions have not been executed by rewire.
pip install genomic-benchmarksGenomic Benchmarks: repository README · README.md at 605d8539, Install, lines 13-13Source reviewed; these instructions have not been executed by rewire.
>>> from genomic_benchmarks.data_check import list_datasets
>>>
>>> list_datasets()
['demo_coding_vs_intergenomic_seqs', 'demo_human_or_worm', 'dummy_mouse_enhancers_ensembl', 'human_enhancers_cohn', 'human_enhancers_ensembl', 'human_ensembl_regulatory', 'human_nontata_promoters', 'human_ocr_ensembl']Genomic Benchmarks: repository README · README.md at 605d8539, Usage, lines 40-43Run your model locally and return predictions keyed by the input IDs. The evaluator supplies biological inputs without test labels and owns scoring. This interface is not a sandbox for model code.
Pass your existing prediction function into this adapter. Its output direction must match the selected protocol.
class MyModelAdapter:
def __init__(self, score):
self.score = score
def predict(self, inputs):
return {row["id"]: float(self.score(row)) for row in inputs}
# adapter = MyModelAdapter(your_prediction_function)
# report = rewirebench.run(prepared, adapter, output="runs/my-model")Alternatively, generate a keyed prediction file in your existing model environment and use the score-only recipe. Your model code and weights do not need to be shared.
Genomic Benchmarks: repository README · README.md at 605d8539Contribute a result for review. The library can submit an exported evaluation for private review when intake is open. Check the contribution page for access and sign-in.
Official data-loading examples and linked TensorFlow/PyTorch training notebooks are available. Dataset download alone produces train/test sequence files rather than a model score. The optional dependency examples contain unquoted >= specifiers, so do not copy those lines directly into a shell without quoting the version requirement.
A maintained rewire runner has not been verified for this benchmark. Check data access, weights, licences, dependencies and hardware in the linked official documentation; requirements have not been fully extracted.
ML-Bioinfo-CEITEC/genomic_benchmarks / README.md · README.md lines 8–34 and 36–108 (Install and Usage)Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.
Stable record: discovery-benchmark-genomic-benchmarksExplanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Datasets | Named genomic classification datasets; the README illustrates a non-TATA human-promoter task.SourcesML-Bioinfo-CEITEC/genomic_benchmarks official source · Pinned README: repository purpose; info and download_dataset examples |
| Splits | The download API delivers prepartitioned train/test data organized by class.SourcesML-Bioinfo-CEITEC/genomic_benchmarks official source · Pinned README: repository purpose; info and download_dataset examples |
| Metrics | The original paper reports classification accuracy and F1 for its CNN baseline on each dataset.Sourcesgenomic-benchmarks primary benchmark evidence · Methods: Reproducibility and Baseline model; Table 2 |
| Baselines | The repository provides neural-network training helpers and links experiment reports.SourcesML-Bioinfo-CEITEC/genomic_benchmarks official source · Pinned README: repository purpose; info and download_dataset examples |
| Leakage controls | Dataset-construction notebooks use fixed seeds. Generated negative regions match positive lengths and are rejected if they overlap positives; this does not establish a suite-wide chromosome-held-out or homology-filtered partition.Sourcesgenomic-benchmarks primary benchmark evidence · Methods: Reproducibility and Baseline model; Table 2 |
| Uncertainty | Table 2 gives point estimates for PyTorch and TensorFlow baseline implementations. The inspected Methods do not specify repeated-training or bootstrap uncertainty for those values. · Not reported in inspected sourcesSourcesgenomic-benchmarks primary benchmark evidence · Methods: Reproducibility and Baseline model; Table 2 |
| Entity type | Repository of genomic sequence classification datasets.SourcesML-Bioinfo-CEITEC/genomic_benchmarks official source · Pinned README: repository purpose; info and download_dataset examples |
| Organisms | Dataset-specific organisms; a human non-TATA promoter dataset is documented in the README.SourcesML-Bioinfo-CEITEC/genomic_benchmarks official source · Pinned README: repository purpose; info and download_dataset examples |
| Assays | Curated genomic classification labels from linked source datasets.SourcesML-Bioinfo-CEITEC/genomic_benchmarks official source · Pinned README: repository purpose; info and download_dataset examples |
| Allowed inputs | Sequences and categorical labels through the dataset loader.SourcesML-Bioinfo-CEITEC/genomic_benchmarks official source · Pinned README: repository purpose; info and download_dataset examples |
| Adaptation | Supervised train/test classification with a documented CNN example.SourcesML-Bioinfo-CEITEC/genomic_benchmarks official source · Pinned README: repository purpose; info and download_dataset examples |
Applicability is distinct from a completed evaluation.
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Last literature check: 2026-09-17. Primary-paper discovery and source inspection. Source-checked results are not independently reproduced experiments.
| Paper or primary resource | Version | Reference |
|---|---|---|
| Genomic benchmarks: a collection of datasets for genomic sequence classification | PMC10150520 | Read source DOI: 10.1186/s12863-023-01123-8 |
The catalogue now holds 36 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.
primary comparison table screened
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
42 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions. Individual claims | genomic-benchmarks primary benchmark evidence Pinned README: repository purpose; info and download_dataset examples; Methods: Reproducibility and Baseline model; Table 2 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: PMC10150520 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram caption Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions. Individual claims | ML-Bioinfo-CEITEC/genomic_benchmarks official source Pinned README: repository purpose; info and download_dataset examples; Methods: Reproducibility and Baseline model; Table 2 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 605d8539830e16c85abe7826990958303ffc5e1c | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| genomic-benchmarks primary benchmark evidence Pinned README: repository purpose; info and download_dataset examples; Methods: Reproducibility and Baseline model; Table 2 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: PMC10150520 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| ML-Bioinfo-CEITEC/genomic_benchmarks official source Pinned README: repository purpose; info and download_dataset examples; Methods: Reproducibility and Baseline model; Table 2 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 605d8539830e16c85abe7826990958303ffc5e1c | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Evaluation procedure Individual claims | genomic-benchmarks primary benchmark evidence Pinned README: repository purpose; info and download_dataset examples; Methods: Reproducibility and Baseline model; Table 2 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: PMC10150520 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Evaluation procedure Individual claims | ML-Bioinfo-CEITEC/genomic_benchmarks official source Pinned README: repository purpose; info and download_dataset examples; Methods: Reproducibility and Baseline model; Table 2 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 605d8539830e16c85abe7826990958303ffc5e1c | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Datasets Named genomic classification datasets; the README illustrates a non-TATA human-promoter task. Individual claims | ML-Bioinfo-CEITEC/genomic_benchmarks official source Pinned README: repository purpose; info and download_dataset examples Version: 605d8539830e16c85abe7826990958303ffc5e1c | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Splits The download API delivers prepartitioned train/test data organized by class. Individual claims | ML-Bioinfo-CEITEC/genomic_benchmarks official source Pinned README: repository purpose; info and download_dataset examples Version: 605d8539830e16c85abe7826990958303ffc5e1c | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Adaptation Supervised train/test classification with a documented CNN example. Individual claims | ML-Bioinfo-CEITEC/genomic_benchmarks official source Pinned README: repository purpose; info and download_dataset examples Version: 605d8539830e16c85abe7826990958303ffc5e1c | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Metrics The original paper reports classification accuracy and F1 for its CNN baseline on each dataset. Individual claims | genomic-benchmarks primary benchmark evidence Methods: Reproducibility and Baseline model; Table 2 Version: PMC10150520 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
View linked audit checks and correction history
Release 2026-09-29-06401fd5b220 · Record review: discovered
Stable ID: discovery-benchmark-genomic-benchmarks