Datasets
Named transcript datasets include translation efficiency, ribosome loading and RNA half-life.
mRNABench assesses genomic-model embeddings on transcript-specific expression, stability and regulatory tasks.
Named transcript datasets include translation efficiency, ribosome loading and RNA half-life.
Task-specific AUPRC for classification and Pearson correlation for continuous RNA properties. Cross-task summaries Fisher-transform correlations before z-scoring; these derived summaries are distinct from the original per-assay metric.
Transcript sequences; some feature baselines additionally use coding-region and splice annotations.
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Source reviewed · Automated source review, 2026-09-16. All specifications and missing details
Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.
auprc (fraction) · Higher values are better.
MIMIC mRNABench probes eCLIP: eCLIP prediction · mRNABench eCLIP as reported in MIMIC Table S11 (MIMIC mRNABench probes split)
Evidence origin: Author-reported evaluation, Result quoted from another source.
MIMIC v1: Table S11 and Appendix D.3 · Table S11 (HTML A4.T11), eCLIP column; Appendix D.3Every method MIMIC mRNABench probes reports on eCLIP prediction, scored with AUPR on mRNABench eCLIP as reported in MIMIC Table S11.
Automated source review: 2026-09-23. Numerical source review does not establish independent reproduction.
Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.
Showing 12 of 12 matching rows.
mRNABench evaluates mature-transcript representations on local sequence effects and global RNA properties. Linear probes use task-specific labels, with homology-based partitions where applicable. Chromosomal, k-mer and homology grouping are compared explicitly because random splits can overstate generalization.
Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.
These source-backed links do not make different protocols or scores interchangeable.
Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.
2 of 32 active baseline roles have published Rewire measurements in this release. Measurements on a selected protocol do not establish coverage of an entire suite.
Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums
Coverage is derived from release 2026-09-29-06401fd5b220. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.
Choose a concrete protocol before running an evaluation. Its inputs, split and scoring rules determine which results can be compared.
Predict assay mean ribosome load from the full processed sequence, including reporter context, for one chemical target and dataset. Test MSE, Pearson and Spearman correlations; fixed upstream RidgeCV procedure for frozen sequence embeddings.
Recompute metrics from supplied predictions. This recipe does not establish reproduction of a particular published score.
Choose one way to run this recipe. These instruction formats are alternatives.
Source reviewed; these instructions have not been executed by rewire.
import rewirebench
prepared = rewirebench.prepare(
'mrnabench-sample-mrl-v1', source='inputs/mrl-sample-egfp.parquet', output='prepared-mrnabench-sample', dataset='egfp', target='target_mrl_egfp_unmod', require_official=True, limit=32
)
report = rewirebench.evaluate(
prepared, 'predictions.json', output='results-mrnabench-sample-rescore',
model={'name': 'My private model', 'training_overlap': 'Unreported'},
)rewirebench 0.4 library interface; rewirebench container and HPC instructions; mRNABench Sample mean-ribosome-load targets runner guide; mRNABench Sample mean-ribosome-load targets protocol implementation; mRNABench Sample mean-ribosome-load targets source manifest; mRNABench Sample mean-ribosome-load targets automated validation receipt · docs/mrnabench.md and protocol implementation; linked validation receipt concerns its stated tests, not execution of this exact snippetRun your model locally and return predictions keyed by the input IDs. The evaluator supplies biological inputs without test labels and owns scoring. This interface is not a sandbox for model code.
Pass your existing prediction function into this adapter. Its output direction must match the selected protocol.
class MyModelAdapter:
def __init__(self, score):
self.score = score
def predict(self, inputs):
return {row["id"]: float(self.score(row)) for row in inputs}
# adapter = MyModelAdapter(your_prediction_function)
# report = rewirebench.run(prepared, adapter, output="runs/my-model")Alternatively, generate a keyed prediction file in your existing model environment and use the score-only recipe. Your model code and weights do not need to be shared.
rewirebench 0.4 library interface; rewirebench container and HPC instructions; mRNABench Sample mean-ribosome-load targets runner guide; mRNABench Sample mean-ribosome-load targets protocol implementation; mRNABench Sample mean-ribosome-load targets source manifest; mRNABench Sample mean-ribosome-load targets automated validation receipt · docs/mrnabench.md; source manifest; protocol score function; automated receipt scopeContribute a result for review. The library can submit an exported evaluation for private review when intake is open. Check the contribution page for access and sign-in.
Official package setup and an Orthrus frozen-embedding/linear-probe example are available. Base-model installation requires the documented CUDA/PyTorch stack; Evo2 has separate instructions. Data and weight paths are placeholders, and reproducing a particular paper result requires its model, split and seed configuration rather than a generic example.
A maintained rewire runner has not been verified for this benchmark. Check data access, weights, licences, dependencies and hardware in the linked official documentation; requirements have not been fully extracted.
morrislab/mRNABench / README.md · README.md lines 27–111 (Setup and Usage)Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.
Stable record: discovery-benchmark-mrnabenchExplanatory profile: source reviewed · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Datasets | Named transcript datasets include translation efficiency, ribosome loading and RNA half-life.Sourcesmorrislab/mRNABench official source · Pinned README: Overview; catalogue; dataset implementation requirements |
| Splits | The library includes training split logic and supports homology-aware splitting using gene identifiers.Sourcesmorrislab/mRNABench official source · Pinned README: Overview; catalogue; dataset implementation requirements |
| Metrics | Task-specific AUPRC for classification and Pearson correlation for continuous RNA properties. Cross-task summaries Fisher-transform correlations before z-scoring; these derived summaries are distinct from the original per-assay metric.Sourcesmrnabench primary benchmark evidence · Sections 3–4; Table 2; Appendix A and D |
| Baselines | NaiveBaseline uses k-mer/GC/sequence statistics, with a six-track variant adding CDS length and exon count; NaiveMamba is an untrained fixed-seed reference.Sourcesmorrislab/mRNABench official source · Pinned README: Overview; catalogue; dataset implementation requirements |
| Leakage controls | Homology splitting requires explicit gene metadata; its presence in the library does not establish that every dataset uses it.Sourcesmorrislab/mRNABench official source · Pinned README: Overview; catalogue; dataset implementation requirements |
| Uncertainty | Linear-probe results are means over ten random splits/seeds. Appendix A and D document standard errors and the selected configurations; Table 2 also uses a Wilcoxon signed-rank comparison.Sourcesmrnabench primary benchmark evidence · Sections 3–4; Table 2; Appendix A and D |
| Entity type | mRNA representation benchmark suite.Sourcesmorrislab/mRNABench official source · Pinned README: Overview; catalogue; dataset implementation requirements |
| Organisms | Human and other dataset-specific transcript collections; the splitter example explicitly conditions on human homology.Sourcesmorrislab/mRNABench official source · Pinned README: Overview; catalogue; dataset implementation requirements |
| Assays | Translation efficiency, ribosome load, half-life and transcript-associated annotations.Sourcesmorrislab/mRNABench official source · Pinned README: Overview; catalogue; dataset implementation requirements |
| Allowed inputs | Transcript sequences; some feature baselines additionally use coding-region and splice annotations.Sourcesmorrislab/mRNABench official source · Pinned README: Overview; catalogue; dataset implementation requirements |
| Adaptation | Frozen embeddings with linear probes; split logic is selected independently of the embedding model.Sourcesmorrislab/mRNABench official source · Pinned README: Overview; catalogue; dataset implementation requirements |
Applicability is distinct from a completed evaluation.
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Last literature check: 2026-09-17. Primary-paper discovery and source inspection. Source-checked results are not independently reproduced experiments.
| Paper or primary resource | Version | Reference |
|---|---|---|
| mRNABench: A curated benchmark for mature mRNA property and function prediction | preprint archived 2025-07-08 | Read source DOI: 10.1101/2025.07.05.662870 |
The catalogue now holds 700 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.
primary comparison table screened
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
131 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions. Individual claims | mrnabench primary benchmark evidence Pinned README: Overview; catalogue; dataset implementation requirements; Sections 3–4; Table 2; Appendix A and D Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: preprint archived 2025-07-08 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram caption Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions. Individual claims | morrislab/mRNABench official source Pinned README: Overview; catalogue; dataset implementation requirements; Sections 3–4; Table 2; Appendix A and D Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 74f96b8e6ae9f41cc3cccff089d826a62d5604b8 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| mrnabench primary benchmark evidence Pinned README: Overview; catalogue; dataset implementation requirements; Sections 3–4; Table 2; Appendix A and D Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: preprint archived 2025-07-08 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| morrislab/mRNABench official source Pinned README: Overview; catalogue; dataset implementation requirements; Sections 3–4; Table 2; Appendix A and D Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 74f96b8e6ae9f41cc3cccff089d826a62d5604b8 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Evaluation procedure Individual claims | mrnabench primary benchmark evidence Pinned README: Overview; catalogue; dataset implementation requirements; Sections 3–4; Table 2; Appendix A and D Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: preprint archived 2025-07-08 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Evaluation procedure Individual claims | morrislab/mRNABench official source Pinned README: Overview; catalogue; dataset implementation requirements; Sections 3–4; Table 2; Appendix A and D Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 74f96b8e6ae9f41cc3cccff089d826a62d5604b8 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Datasets Named transcript datasets include translation efficiency, ribosome loading and RNA half-life. Individual claims | morrislab/mRNABench official source Pinned README: Overview; catalogue; dataset implementation requirements Version: 74f96b8e6ae9f41cc3cccff089d826a62d5604b8 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Splits The library includes training split logic and supports homology-aware splitting using gene identifiers. Individual claims | morrislab/mRNABench official source Pinned README: Overview; catalogue; dataset implementation requirements Version: 74f96b8e6ae9f41cc3cccff089d826a62d5604b8 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Adaptation Frozen embeddings with linear probes; split logic is selected independently of the embedding model. Individual claims | morrislab/mRNABench official source Pinned README: Overview; catalogue; dataset implementation requirements Version: 74f96b8e6ae9f41cc3cccff089d826a62d5604b8 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Metrics Task-specific AUPRC for classification and Pearson correlation for continuous RNA properties. Cross-task summaries Fisher-transform correlations before z-scoring; these derived summaries are distinct from the original per-assay metric. Individual claims | mrnabench primary benchmark evidence Sections 3–4; Table 2; Appendix A and D Version: preprint archived 2025-07-08 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
View linked audit checks and correction history
Release 2026-09-29-06401fd5b220 · Record review: discovered
Stable ID: discovery-benchmark-mrnabench