rewire.itbenchmarks
Benchmark

mRNABench

mRNABench assesses genomic-model embeddings on transcript-specific expression, stability and regulatory tasks.

Sourcesmorrislab/mRNABench official source · Pinned README: Overview; catalogue; dataset implementation requirements

696 evaluations · 700 results

Overview

Datasets

Named transcript datasets include translation efficiency, ribosome loading and RNA half-life.

Sourcesmorrislab/mRNABench official source · Pinned README: Overview; catalogue; dataset implementation requirements

Metrics

Task-specific AUPRC for classification and Pearson correlation for continuous RNA properties. Cross-task summaries Fisher-transform correlations before z-scoring; these derived summaries are distinct from the original per-assay metric.

Sourcesmrnabench primary benchmark evidence · Sections 3–4; Table 2; Appendix A and D

Allowed inputs

Transcript sequences; some feature baselines additionally use coding-region and splice annotations.

Sourcesmorrislab/mRNABench official source · Pinned README: Overview; catalogue; dataset implementation requirements
Evaluation procedure diagram
How it worksEvaluation procedure
Evaluation procedure1. Allowed inputs: Transcript sequences; some feature baselines additionally use coding-region and splice annotations.. Then: 2. Splits: The library includes training split logic and supports homology-aware splitting using gene identifiers.. Then: 3. Metrics: Task-specific AUPRC for classification and Pearson correlation for continuous RNA properties. Cross-task summaries Fisher-transform correlations before z-scoring; these derived summaries are distinct from the original per-assay metric.Evaluation procedure1. Allowed inputs: Transcript sequences; some feature baselines additionally use coding-region and splice annotations.. Then: 2. Splits: The library includes training split logic and supports homology-aware splitting using gene identifiers.. Then: 3. Metrics: Task-specific AUPRC for classification and Pearson correlation for continuous RNA properties. Cross-task summaries Fisher-transform correlations before z-scoring; these derived summaries are distinct from the original per-assay metric.Evaluation procedure1. Allowed inputs: Transcript sequences; some feature baselines additionally use coding-region and splice annotations.. Then: 2. Splits: The library includes training split logic and supports homology-aware splitting using gene identifiers.. Then: 3. Metrics: Task-specific AUPRC for classification and Pearson correlation for continuous RNA properties. Cross-task summaries Fisher-transform correlations before z-scoring; these derived summaries are distinct from the original per-assay metric.

Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.

Sources (2)morrislab/mRNABench official source; mrnabench primary benchmark evidence · Pinned README: Overview; catalogue; dataset implementation requirements; Sections 3–4; Table 2; Appendix A and D

Source reviewed · Automated source review, 2026-09-16. All specifications and missing details

Results

Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.

MIMIC mRNABench probes eCLIP: eCLIP prediction

auprc (fraction) · Higher values are better.

MIMIC mRNABench probes eCLIP: eCLIP prediction · mRNABench eCLIP as reported in MIMIC Table S11 (MIMIC mRNABench probes split)

Evidence origin: Author-reported evaluation, Result quoted from another source.

MIMIC v1: Table S11 and Appendix D.3 · Table S11 (HTML A4.T11), eCLIP column; Appendix D.3
  • MIMIC results are reported here; every comparator is explicitly quoted from prior mRNABench work (Appendix D.3), not rerun. Do not count quoted rows as independent experiments.
  • MIMIC can use additional coding/protein modalities. This is not a matched nucleotide-only comparison.
  • The table prints AUPR for RNA localization; the prior mRNABench appendix has a conflicting metric label. Do not pool this source-specific comparison with that appendix.
  • Exact fitted checkpoints, scored counts and uncertainty are not supplied by Table S11.
Comparison details and limitations

Every method MIMIC mRNABench probes reports on eCLIP prediction, scored with AUPR on mRNABench eCLIP as reported in MIMIC Table S11.

Automated source review: 2026-09-23. Numerical source review does not establish independent reproduction.

Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.

Showing 12 of 12 matching rows.

Methods and evaluation design

Procedure, tasks and evaluated configurations

How it works

Evaluation methodology

mRNABench evaluates mature-transcript representations on local sequence effects and global RNA properties. Linear probes use task-specific labels, with homology-based partitions where applicable. Chromosomal, k-mer and homology grouping are compared explicitly because random splits can overstate generalization.

Sourcesmrnabench primary benchmark evidence · Sections 3–4; Table 2; Appendix A and D

Evaluation design

Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.

These source-backed links do not make different protocols or scores interchangeable.

Baseline coverage

Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.

2 of 32 active baseline roles have published Rewire measurements in this release. Measurements on a selected protocol do not establish coverage of an entire suite.

Baseline status by linked protocol

Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums

Coverage is derived from release 2026-09-29-06401fd5b220. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.

Run this benchmark

Choose a concrete protocol before running an evaluation. Its inputs, split and scoring rules determine which results can be compared.

mRNABench Sample mean-ribosome-load targets: score supplied predictions

Predict assay mean ribosome load from the full processed sequence, including reporter context, for one chemical target and dataset. Test MSE, Pearson and Spearman correlations; fixed upstream RidgeCV procedure for frozen sequence embeddings.

Recompute metrics from supplied predictions. This recipe does not establish reproduction of a particular published score.

Dataset access
Obtain pinned morrislab/mrl-sample parquets at ef67f7cf8a999bb1c412ad6551aa7d9f901cbb95. Four datasets, six target columns. Source hashes and each target split are checked; default splitter70/15/15 with seed2541.
Model and weights
No weights needed for rescoring or the supplied toy/composition control. Private-model weights remain local; their access requirements depend on the model.
Licences
Runner MIT. Publisher data card says licence unknown; data are not redistributed. Upstream software AGPL-3.0 is not bundled by this independent implementation.
Software
Python3.11, pinned rewirebench0.4 environment; sequence extra for parquet and HDF5. Podman or Apptainer is optional.
Hardware
CPU scoring and small controls; memory/accelerator needs for real private-model inference are model-dependent and unreported here.
Required inputs and expected outputs

Inputs

  • Prepared protocol inputs with evaluator-owned labels and opaque IDs.
  • Keyed finite scalar scores; missing predictions must be explicitly permitted and remain in coverage.

Outputs

  • Local report.json with metrics, source/version identity and coverage.
  • Local predictions.json and unscored.json; nothing submitted automatically.

Choose one way to run this recipe. These instruction formats are alternatives.

Execution steps

  1. 1. Score supplied keyed predictions (Python)

    Source reviewed; these instructions have not been executed by rewire.

    import rewirebench
    prepared = rewirebench.prepare(
        'mrnabench-sample-mrl-v1', source='inputs/mrl-sample-egfp.parquet', output='prepared-mrnabench-sample', dataset='egfp', target='target_mrl_egfp_unmod', require_official=True, limit=32
    )
    report = rewirebench.evaluate(
        prepared, 'predictions.json', output='results-mrnabench-sample-rescore',
        model={'name': 'My private model', 'training_overlap': 'Unreported'},
    )
    rewirebench 0.4 library interface; rewirebench container and HPC instructions; mRNABench Sample mean-ribosome-load targets runner guide; mRNABench Sample mean-ribosome-load targets protocol implementation; mRNABench Sample mean-ribosome-load targets source manifest; mRNABench Sample mean-ribosome-load targets automated validation receipt · docs/mrnabench.md and protocol implementation; linked validation receipt concerns its stated tests, not execution of this exact snippet

Use your own model

Run your model locally and return predictions keyed by the input IDs. The evaluator supplies biological inputs without test labels and owns scoring. This interface is not a sandbox for model code.

Pass your existing prediction function into this adapter. Its output direction must match the selected protocol.

class MyModelAdapter:
    def __init__(self, score):
        self.score = score

    def predict(self, inputs):
        return {row["id"]: float(self.score(row)) for row in inputs}

# adapter = MyModelAdapter(your_prediction_function)
# report = rewirebench.run(prepared, adapter, output="runs/my-model")

Alternatively, generate a keyed prediction file in your existing model environment and use the score-only recipe. Your model code and weights do not need to be shared.

rewirebench 0.4 library interface; rewirebench container and HPC instructions; mRNABench Sample mean-ribosome-load targets runner guide; mRNABench Sample mean-ribosome-load targets protocol implementation; mRNABench Sample mean-ribosome-load targets source manifest; mRNABench Sample mean-ribosome-load targets automated validation receipt · docs/mrnabench.md; source manifest; protocol score function; automated receipt scope
Scope and limitations
  • Only four Sample datasets are supported, with three eGFP targets plus one each for mCherry, designed and varying-length sequences. There is no aggregate across targets.
  • Sequence-only input excludes CDS, splice, uORF and Kozak tracks; multi-track published results use different information. Source DNA letters for RNA constructs are preserved.
  • Missing target values are excluded before splitting. Nine eGFP methylpseudouridine labels are missing; original source and eligible denominators are retained.
  • Frozen RidgeCV selects from alphas0.001,0.01,0.1,1,10 using training data only. Validation/test labels do not select the penalty.
  • All four source files and six target splits were checked. Small80-source-row probes matched the actual pinned upstream splitter and metrics within1e-9; full-model published performance was not reproduced.
  • These snippets are source-reviewed instructions. The linked automated receipt supports only the tests it records; it does not certify the exact snippet or published-model reproduction.
  • Smoke examples and partial predictions retain original denominators and cannot establish a complete suite. Prepared execution is offline; exports and submissions are explicit.

Contribute a result for review. The library can submit an exported evaluation for private review when intake is open. Check the contribution page for access and sign-in.

Original repository instructions

Run this benchmark

Official package setup and an Orthrus frozen-embedding/linear-probe example are available. Base-model installation requires the documented CUDA/PyTorch stack; Evo2 has separate instructions. Data and weight paths are placeholders, and reproducing a particular paper result requires its model, split and seed configuration rather than a generic example.

A maintained rewire runner has not been verified for this benchmark. Check data access, weights, licences, dependencies and hardware in the linked official documentation; requirements have not been fully extracted.

morrislab/mRNABench / README.md · README.md lines 27–111 (Setup and Usage)
Strengths, limitations and unresolved questions

Strengths and limitations

Strengths supported by sources

  • Naive sequence-feature and randomly initialized model baselines test whether pretraining adds useful information.
    Sourcesmorrislab/mRNABench official source · Pinned README: Overview; catalogue; dataset implementation requirements

Limitations and conditions

  • Homology grouping is not used for every assay: the paper retains random splits for MRL-MPRA, MRL-HL-PAIR and variant effects. Derived cross-task z-scores should not replace the original biological metrics.
    Sourcesmrnabench primary benchmark evidence · Sections 3–4; Table 2; Appendix A and D
Profile review details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Stable record: discovery-benchmark-mrnabench

Specifications

Inputs, training, access and other details

Explanatory profile: source reviewed · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetsNamed transcript datasets include translation efficiency, ribosome loading and RNA half-life.
Sourcesmorrislab/mRNABench official source · Pinned README: Overview; catalogue; dataset implementation requirements
SplitsThe library includes training split logic and supports homology-aware splitting using gene identifiers.
Sourcesmorrislab/mRNABench official source · Pinned README: Overview; catalogue; dataset implementation requirements
MetricsTask-specific AUPRC for classification and Pearson correlation for continuous RNA properties. Cross-task summaries Fisher-transform correlations before z-scoring; these derived summaries are distinct from the original per-assay metric.
Sourcesmrnabench primary benchmark evidence · Sections 3–4; Table 2; Appendix A and D
BaselinesNaiveBaseline uses k-mer/GC/sequence statistics, with a six-track variant adding CDS length and exon count; NaiveMamba is an untrained fixed-seed reference.
Sourcesmorrislab/mRNABench official source · Pinned README: Overview; catalogue; dataset implementation requirements
Leakage controlsHomology splitting requires explicit gene metadata; its presence in the library does not establish that every dataset uses it.
Sourcesmorrislab/mRNABench official source · Pinned README: Overview; catalogue; dataset implementation requirements
UncertaintyLinear-probe results are means over ten random splits/seeds. Appendix A and D document standard errors and the selected configurations; Table 2 also uses a Wilcoxon signed-rank comparison.
Sourcesmrnabench primary benchmark evidence · Sections 3–4; Table 2; Appendix A and D
Entity typemRNA representation benchmark suite.
Sourcesmorrislab/mRNABench official source · Pinned README: Overview; catalogue; dataset implementation requirements
OrganismsHuman and other dataset-specific transcript collections; the splitter example explicitly conditions on human homology.
Sourcesmorrislab/mRNABench official source · Pinned README: Overview; catalogue; dataset implementation requirements
AssaysTranslation efficiency, ribosome load, half-life and transcript-associated annotations.
Sourcesmorrislab/mRNABench official source · Pinned README: Overview; catalogue; dataset implementation requirements
Allowed inputsTranscript sequences; some feature baselines additionally use coding-region and splice annotations.
Sourcesmorrislab/mRNABench official source · Pinned README: Overview; catalogue; dataset implementation requirements
AdaptationFrozen embeddings with linear probes; split logic is selected independently of the embedding model.
Sourcesmorrislab/mRNABench official source · Pinned README: Overview; catalogue; dataset implementation requirements
Applicable tests and references

Applicability is distinct from a completed evaluation.

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Primary-paper discovery and source inspection. Source-checked results are not independently reproduced experiments.

Paper or primary resourceVersionReference
mRNABench: A curated benchmark for mature mRNA property and function predictionpreprint archived 2025-07-08Read source
DOI: 10.1101/2025.07.05.662870
Historical gaps recorded on 2026-09-17

The catalogue now holds 700 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.

  • Table 5 labels localization columns Pearson R while Table 2 labels them AUPRC; quarantine those metric identities pending reconciliation.
  • Appendix C claims ten splits but enumerates nine seeds. Record ten as reported, with discrepancy, not an inferred tenth seed.
  • Tables 5–6 uncertainties are 95% confidence intervals, not standard deviations.
  • Task/subtask pooling and transformed aggregate rankings must not be conflated with printed raw metrics.
Search and extraction details

primary comparison table screened

Searches

  • mRNABench PMC12265608

Evidence locations

  • Tables 1–2, 5–6
  • Appendix B: Mean Ribosome Load – MPRA
  • Appendix C: Linear Probing Experimental Setup
  • Appendix D

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

131 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Individual claims
mrnabench primary benchmark evidence

Original source ↗

Pinned README: Overview; catalogue; dataset implementation requirements; Sections 3–4; Table 2; Appendix A and D

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: preprint archived 2025-07-08
Retrieved: 2026-09-16T10:41:16.497221+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 79f6264ee883535203c63a313547e7c57baa85585f76b42f8d899eb17fb7e600

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram caption
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Individual claims
morrislab/mRNABench official source

Original source ↗

Pinned README: Overview; catalogue; dataset implementation requirements; Sections 3–4; Table 2; Appendix A and D

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 74f96b8e6ae9f41cc3cccff089d826a62d5604b8
Retrieved: 2026-09-16T10:30:21.823561+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: f0c67304e20ced42938829dfee39480cef51ee3a4357ee8b53eafc89c305fa60

Hash scope: Hash scope not separately documented; inspect source record

Diagram steps
  • Allowed inputs: Transcript sequences; some feature baselines additionally use coding-region and splice annotations.
  • Splits: The library includes training split logic and supports homology-aware splitting using gene identifiers.
  • Metrics: Task-specific AUPRC for classification and Pearson correlation for continuous RNA properties. Cross-task summaries Fisher-transform correlations before z-scoring; these derived summaries are distinct from the original per-assay metric.
Individual claims
mrnabench primary benchmark evidence

Original source ↗

Pinned README: Overview; catalogue; dataset implementation requirements; Sections 3–4; Table 2; Appendix A and D

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: preprint archived 2025-07-08
Retrieved: 2026-09-16T10:41:16.497221+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 79f6264ee883535203c63a313547e7c57baa85585f76b42f8d899eb17fb7e600

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps
  • Allowed inputs: Transcript sequences; some feature baselines additionally use coding-region and splice annotations.
  • Splits: The library includes training split logic and supports homology-aware splitting using gene identifiers.
  • Metrics: Task-specific AUPRC for classification and Pearson correlation for continuous RNA properties. Cross-task summaries Fisher-transform correlations before z-scoring; these derived summaries are distinct from the original per-assay metric.
Individual claims
morrislab/mRNABench official source

Original source ↗

Pinned README: Overview; catalogue; dataset implementation requirements; Sections 3–4; Table 2; Appendix A and D

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 74f96b8e6ae9f41cc3cccff089d826a62d5604b8
Retrieved: 2026-09-16T10:30:21.823561+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: f0c67304e20ced42938829dfee39480cef51ee3a4357ee8b53eafc89c305fa60

Hash scope: Hash scope not separately documented; inspect source record

Diagram title
Evaluation procedure
Individual claims
mrnabench primary benchmark evidence

Original source ↗

Pinned README: Overview; catalogue; dataset implementation requirements; Sections 3–4; Table 2; Appendix A and D

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: preprint archived 2025-07-08
Retrieved: 2026-09-16T10:41:16.497221+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: 79f6264ee883535203c63a313547e7c57baa85585f76b42f8d899eb17fb7e600

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title
Evaluation procedure
Individual claims
morrislab/mRNABench official source

Original source ↗

Pinned README: Overview; catalogue; dataset implementation requirements; Sections 3–4; Table 2; Appendix A and D

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 74f96b8e6ae9f41cc3cccff089d826a62d5604b8
Retrieved: 2026-09-16T10:30:21.823561+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: f0c67304e20ced42938829dfee39480cef51ee3a4357ee8b53eafc89c305fa60

Hash scope: Hash scope not separately documented; inspect source record

Datasets
Named transcript datasets include translation efficiency, ribosome loading and RNA half-life.
Individual claims
morrislab/mRNABench official source

Original source ↗

Pinned README: Overview; catalogue; dataset implementation requirements

Version: 74f96b8e6ae9f41cc3cccff089d826a62d5604b8
Retrieved: 2026-09-16T10:30:21.823561+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: f0c67304e20ced42938829dfee39480cef51ee3a4357ee8b53eafc89c305fa60

Hash scope: Hash scope not separately documented; inspect source record

Splits
The library includes training split logic and supports homology-aware splitting using gene identifiers.
Individual claims
morrislab/mRNABench official source

Original source ↗

Pinned README: Overview; catalogue; dataset implementation requirements

Version: 74f96b8e6ae9f41cc3cccff089d826a62d5604b8
Retrieved: 2026-09-16T10:30:21.823561+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: f0c67304e20ced42938829dfee39480cef51ee3a4357ee8b53eafc89c305fa60

Hash scope: Hash scope not separately documented; inspect source record

Adaptation
Frozen embeddings with linear probes; split logic is selected independently of the embedding model.
Individual claims
morrislab/mRNABench official source

Original source ↗

Pinned README: Overview; catalogue; dataset implementation requirements

Version: 74f96b8e6ae9f41cc3cccff089d826a62d5604b8
Retrieved: 2026-09-16T10:30:21.823561+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: f0c67304e20ced42938829dfee39480cef51ee3a4357ee8b53eafc89c305fa60

Hash scope: Hash scope not separately documented; inspect source record

Metrics
Task-specific AUPRC for classification and Pearson correlation for continuous RNA properties. Cross-task summaries Fisher-transform correlations before z-scoring; these derived summaries are distinct from the original per-assay metric.
Individual claims
mrnabench primary benchmark evidence

Original source ↗

Sections 3–4; Table 2; Appendix A and D

Version: preprint archived 2025-07-08
Retrieved: 2026-09-16T10:41:16.497221+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: 79f6264ee883535203c63a313547e7c57baa85585f76b42f8d899eb17fb7e600

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: discovered

11 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: discovery-benchmark-mrnabench

areas
rna
entity level
suite
scope note
Specialist molecular or omics evaluation; protocol details require review before numerical comparison.
task
mRNA embedding quality on downstream tasks
version
Not reported
benchmark research
review date: 2026-09-17; status: primary_comparison_table_screened; primary sources: expansion-p3-mrnabench-2025; inspected locators: Tables 1–2, 5–6; Appendix B: Mean Ribosome Load – MPRA; Appendix C: Linear Probing Experimental Setup; Appendix D; searched queries: mRNABench PMC12265608; gaps: Table 5 labels localization columns Pearson R while Table 2 labels them AUPRC; quarantine those metric identities pending reconciliation.; Appendix C claims ten splits but enumerates nine seeds. Record ten as reported, with discrepancy, not an inferred tenth seed.; Tables 5–6 uncertainties are 95% confidence intervals, not standard deviations.; Task/subtask pooling and transformed aggregate rankings must not be conflated with printed raw metrics.; claim scope: Primary-paper discovery and source inspection. Source-checked results are not independently reproduced experiments.
historical missing metadata
dataset release: unextracted; metric implementation: unextracted; split manifest: unextracted; version: unextracted
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
entity classification
review date: 2026-09-17; rationale: The cited profile describes a collection of evaluation tasks or protocols; retain it as the top-level benchmark suite. Its datasets and individual protocols remain separate records.; source ids: src-discovery-morrislab-mrnabench; source locator: Pinned README: Overview; catalogue; dataset implementation requirements; ambiguities: None recorded
run documentation
record id: discovery-benchmark-mrnabench; source ids: run-doc-mrnabench-readme-md-74f96b8e; status: official_documentation_linked; summary: Official package setup and an Orthrus frozen-embedding/linear-probe example are available. Base-model installation requires the documented CUDA/PyTorch stack; Evo2 has separate instructions. Data and weight paths are placeholders, and reproducing a particular paper result requires its model, split and seed configuration rather than a generic example.; source locator: README.md lines 27–111 (Setup and Usage)
run recipes
id: mrnabench-sample-runner-rescore-v1; protocol id: mrnabench-sample-mrl-v1; version: fcccbbcdbe3d5cd64a1a312d536615273320f7b4; title: mRNABench Sample mean-ribosome-load targets: score supplied predictions; purpose: rescore_predictions; summary: Predict assay mean ribosome load from the full processed sequence, including reporter context, for one chemical target and dataset. Test MSE, Pearson and Spearman correlations; fixed upstream RidgeCV procedure for frozen sequence embeddings.; inputs: Prepared protocol inputs with evaluator-owned labels and opaque IDs.; Keyed finite scalar scores; missing predictions must be explicitly permitted and remain in coverage.; outputs: Local report.json with metrics, source/version identity and coverage.; Local predictions.json and unscored.json; nothing submitted automatically.; requirements: data: Obtain pinned morrislab/mrl-sample parquets at ef67f7cf8a999bb1c412ad6551aa7d9f901cbb95. Four datasets, six target columns. Source hashes and each target split are checked; default splitter70/15/15 with seed2541.; weights: No weights needed for rescoring or the supplied toy/composition control. Private-model weights remain local; their access requirements depend on the model.; licence: Runner MIT. Publisher data card says licence unknown; data are not redistributed. Upstream software AGPL-3.0 is not bundled by this independent implementation.; software: Python3.11, pinned rewirebench0.4 environment; sequence extra for parquet and HDF5. Podman or Apptainer is optional.; hardware: CPU scoring and small controls; memory/accelerator needs for real private-model inference are model-dependent and unreported here.; instructions: runtime: python; title: Score supplied keyed predictions; code: import rewirebench prepared = rewirebench.prepare( 'mrnabench-sample-mrl-v1', source='inputs/mrl-sample-egfp.parquet', output='prepared-mrnabench-sample', dataset='egfp', target='target_mrl_egfp_unmod', require_official=True, limit=32 ) report = rewirebench.evaluate( prepared, 'predictions.json', output='results-mrnabench-sample-rescore', model={'name': 'My private model', 'training_overlap': 'Unreported'}, ); status: source_reviewed_not_executed; source ids: runner-04-docs-sdk-md; runner-04-docs-hpc-md; runner-04-docs-mrnabench-md; runner-04-packages-rewirebench-src-rewirebench-protocols-mrnabench-py; runner-04-packages-rewirebench-src-rewirebench-resources-mrnabench-sources-json; runner-04-packages-rewirebench-src-rewirebench-resources-mrnabench-validation-receipt-json; source locator: docs/mrnabench.md and protocol implementation; linked validation receipt concerns its stated tests, not execution of this exact snippet; runtime: command_line; title: Prepare inputs and score supplied predictions; code: git clone https://github.com/rewire-bio/rewire-benchmarks.git cd rewire-benchmarks git checkout --detach fcccbbcdbe3d5cd64a1a312d536615273320f7b4 uv sync --locked --package rewirebench --extra sequence --python 3.11 uv run --package rewirebench rewirebench prepare mrnabench-sample-mrl-v1 \ --source inputs/mrl-sample-egfp.parquet --output prepared-mrnabench-sample \ --options '{"dataset":"egfp","target":"target_mrl_egfp_unmod","require_official":true,"limit":32}' uv run --package rewirebench rewirebench evaluate \ --prepared prepared-mrnabench-sample --predictions predictions.json \ --output results-mrnabench-sample-rescore --model-name 'My private model'; status: source_reviewed_not_executed; source ids: runner-04-docs-sdk-md; runner-04-docs-hpc-md; runner-04-docs-mrnabench-md; runner-04-packages-rewirebench-src-rewirebench-protocols-mrnabench-py; runner-04-packages-rewirebench-src-rewirebench-resources-mrnabench-sources-json; runner-04-packages-rewirebench-src-rewirebench-resources-mrnabench-validation-receipt-json; source locator: docs/mrnabench.md and protocol implementation; linked validation receipt concerns its stated tests, not execution of this exact snippet; runtime: podman; title: Score prepared predictions offline; code: # Obtain/build the pinned core OCI image following the linked HPC guide. mkdir -p runs podman run --rm --userns=keep-id --network=none \ -v "$PWD/prepared-mrnabench-sample:/prepared:ro" \ -v "$PWD/predictions:/predictions:ro" -v "$PWD/runs:/outputs:rw" \ localhost/rewirebench:core evaluate --prepared /prepared \ --predictions /predictions/predictions.json --output /outputs/mrnabench-sample-rescore; status: source_reviewed_not_executed; source ids: runner-04-docs-sdk-md; runner-04-docs-hpc-md; runner-04-docs-mrnabench-md; runner-04-packages-rewirebench-src-rewirebench-protocols-mrnabench-py; runner-04-packages-rewirebench-src-rewirebench-resources-mrnabench-sources-json; runner-04-packages-rewirebench-src-rewirebench-resources-mrnabench-validation-receipt-json; source locator: docs/mrnabench.md and protocol implementation; linked validation receipt concerns its stated tests, not execution of this exact snippet; runtime: apptainer; title: Score prepared predictions offline; code: # Obtain/build the pinned SIF following the linked HPC guide. mkdir -p runs apptainer run --cleanenv --containall \ --bind "$PWD/prepared-mrnabench-sample:/prepared:ro" \ --bind "$PWD/predictions:/predictions:ro" --bind "$PWD/runs:/outputs:rw" \ rewirebench-core.sif evaluate --prepared /prepared \ --predictions /predictions/predictions.json --output /outputs/mrnabench-sample-rescore; status: source_reviewed_not_executed; source ids: runner-04-docs-sdk-md; runner-04-docs-hpc-md; runner-04-docs-mrnabench-md; runner-04-packages-rewirebench-src-rewirebench-protocols-mrnabench-py; runner-04-packages-rewirebench-src-rewirebench-resources-mrnabench-sources-json; runner-04-packages-rewirebench-src-rewirebench-resources-mrnabench-validation-receipt-json; source locator: docs/mrnabench.md and protocol implementation; linked validation receipt concerns its stated tests, not execution of this exact snippet; runtime: slurm; title: Schedule prepared scoring with site-specific resources; code: #!/bin/bash set -euo pipefail # Set these for your cluster; no performance or resource estimate is implied. : "${REWIRE_ACCOUNT:?Set your Slurm account}" : "${REWIRE_PARTITION:?Set your Slurm partition}" : "${REWIRE_CPUS:?Set the requested CPU count}" : "${REWIRE_MEMORY:?Set the requested memory}" : "${REWIRE_WALLTIME:?Set the requested time limit}" : "${REWIRE_JOB_ROOT:?Set a shared absolute directory with prepared data and SIF}" export REWIRE_JOB_ROOT sbatch --account="$REWIRE_ACCOUNT" --partition="$REWIRE_PARTITION" \ --cpus-per-task="$REWIRE_CPUS" --mem="$REWIRE_MEMORY" \ --time="$REWIRE_WALLTIME" --export=ALL <<'REWIRE_JOB' #!/bin/bash set -euo pipefail cd "$REWIRE_JOB_ROOT" mkdir -p runs apptainer run --cleanenv --containall \ --bind "$PWD/prepared-mrnabench-sample:/prepared:ro" \ --bind "$PWD/predictions:/predictions:ro" \ --bind "$PWD/runs:/outputs:rw" \ rewirebench-core.sif evaluate \ --prepared /prepared --predictions /predictions/predictions.json \ --output /outputs/mrnabench-sample-rescore REWIRE_JOB; status: source_reviewed_not_executed; source ids: runner-04-docs-sdk-md; runner-04-docs-hpc-md; runner-04-docs-mrnabench-md; runner-04-packages-rewirebench-src-rewirebench-protocols-mrnabench-py; runner-04-packages-rewirebench-src-rewirebench-resources-mrnabench-sources-json; runner-04-packages-rewirebench-src-rewirebench-resources-mrnabench-validation-receipt-json; source locator: docs/hpc.md: Slurm template, caller-supplied account/partition/resources; protocol guide: evaluate prepared predictions; limitations: Only four Sample datasets are supported, with three eGFP targets plus one each for mCherry, designed and varying-length sequences. There is no aggregate across targets.; Sequence-only input excludes CDS, splice, uORF and Kozak tracks; multi-track published results use different information. Source DNA letters for RNA constructs are preserved.; Missing target values are excluded before splitting. Nine eGFP methylpseudouridine labels are missing; original source and eligible denominators are retained.; Frozen RidgeCV selects from alphas0.001,0.01,0.1,1,10 using training data only. Validation/test labels do not select the penalty.; All four source files and six target splits were checked. Small80-source-row probes matched the actual pinned upstream splitter and metrics within1e-9; full-model published performance was not reproduced.; These snippets are source-reviewed instructions. The linked automated receipt supports only the tests it records; it does not certify the exact snippet or published-model reproduction.; Smoke examples and partial predictions retain original denominators and cannot establish a complete suite. Prepared execution is offline; exports and submissions are explicit.; source ids: runner-04-docs-sdk-md; runner-04-docs-hpc-md; runner-04-docs-mrnabench-md; runner-04-packages-rewirebench-src-rewirebench-protocols-mrnabench-py; runner-04-packages-rewirebench-src-rewirebench-resources-mrnabench-sources-json; runner-04-packages-rewirebench-src-rewirebench-resources-mrnabench-validation-receipt-json; source locator: docs/mrnabench.md; source manifest; protocol score function; automated receipt scope; id: mrnabench-sample-runner-control-v1; protocol id: mrnabench-sample-mrl-v1; version: fcccbbcdbe3d5cd64a1a312d536615273320f7b4; title: mRNABench Sample mean-ribosome-load targets: local model or software control; purpose: generate_and_evaluate; summary: Predict assay mean ribosome load from the full processed sequence, including reporter context, for one chemical target and dataset. Test MSE, Pearson and Spearman correlations; fixed upstream RidgeCV procedure for frozen sequence embeddings.; inputs: Same prepared biological inputs; no evaluation labels are exposed to adapters.; A private local adapter or the documented lightweight software control.; outputs: Local report.json with metrics, source/version identity and coverage.; Local predictions.json and unscored.json; nothing submitted automatically.; requirements: data: Obtain pinned morrislab/mrl-sample parquets at ef67f7cf8a999bb1c412ad6551aa7d9f901cbb95. Four datasets, six target columns. Source hashes and each target split are checked; default splitter70/15/15 with seed2541.; weights: No weights needed for rescoring or the supplied toy/composition control. Private-model weights remain local; their access requirements depend on the model.; licence: Runner MIT. Publisher data card says licence unknown; data are not redistributed. Upstream software AGPL-3.0 is not bundled by this independent implementation.; software: Python3.11, pinned rewirebench0.4 environment; sequence extra for parquet and HDF5. Podman or Apptainer is optional.; hardware: CPU scoring and small controls; memory/accelerator needs for real private-model inference are model-dependent and unreported here.; instructions: runtime: python; title: Run a small local software control; code: import rewirebench from rewirebench.adapters.sequence import SequenceComposition prepared = rewirebench.prepare( 'mrnabench-sample-mrl-v1', source='inputs/mrl-sample-egfp.parquet', output='prepared-mrnabench-sample', dataset='egfp', target='target_mrl_egfp_unmod', require_official=True, limit=32 ) report = rewirebench.run( prepared, SequenceComposition(), output='results-mrnabench-sample-control', model={'name': 'Local software control; not a published baseline'}, ); status: source_reviewed_not_executed; source ids: runner-04-docs-sdk-md; runner-04-docs-hpc-md; runner-04-docs-mrnabench-md; runner-04-packages-rewirebench-src-rewirebench-protocols-mrnabench-py; runner-04-packages-rewirebench-src-rewirebench-resources-mrnabench-sources-json; runner-04-packages-rewirebench-src-rewirebench-resources-mrnabench-validation-receipt-json; source locator: docs/mrnabench.md and protocol implementation; linked validation receipt concerns its stated tests, not execution of this exact snippet; runtime: command_line; title: Run a small composition control; code: git clone https://github.com/rewire-bio/rewire-benchmarks.git cd rewire-benchmarks git checkout --detach fcccbbcdbe3d5cd64a1a312d536615273320f7b4 uv sync --locked --package rewirebench --extra sequence --python 3.11 uv run --package rewirebench rewirebench prepare mrnabench-sample-mrl-v1 \ --source inputs/mrl-sample-egfp.parquet --output prepared-mrnabench-sample \ --options '{"dataset":"egfp","target":"target_mrl_egfp_unmod","require_official":true,"limit":32}' uv run --package rewirebench rewirebench run --prepared prepared-mrnabench-sample \ --adapter rewirebench.adapters.sequence:SequenceComposition --prediction-type embedding \ --output results-mrnabench-sample-control; status: source_reviewed_not_executed; source ids: runner-04-docs-sdk-md; runner-04-docs-hpc-md; runner-04-docs-mrnabench-md; runner-04-packages-rewirebench-src-rewirebench-protocols-mrnabench-py; runner-04-packages-rewirebench-src-rewirebench-resources-mrnabench-sources-json; runner-04-packages-rewirebench-src-rewirebench-resources-mrnabench-validation-receipt-json; source locator: docs/mrnabench.md and protocol implementation; linked validation receipt concerns its stated tests, not execution of this exact snippet; limitations: Only four Sample datasets are supported, with three eGFP targets plus one each for mCherry, designed and varying-length sequences. There is no aggregate across targets.; Sequence-only input excludes CDS, splice, uORF and Kozak tracks; multi-track published results use different information. Source DNA letters for RNA constructs are preserved.; Missing target values are excluded before splitting. Nine eGFP methylpseudouridine labels are missing; original source and eligible denominators are retained.; Frozen RidgeCV selects from alphas0.001,0.01,0.1,1,10 using training data only. Validation/test labels do not select the penalty.; All four source files and six target splits were checked. Small80-source-row probes matched the actual pinned upstream splitter and metrics within1e-9; full-model published performance was not reproduced.; These snippets are source-reviewed instructions. The linked automated receipt supports only the tests it records; it does not certify the exact snippet or published-model reproduction.; Smoke examples and partial predictions retain original denominators and cannot establish a complete suite. Prepared execution is offline; exports and submissions are explicit.; source ids: runner-04-docs-sdk-md; runner-04-docs-hpc-md; runner-04-docs-mrnabench-md; runner-04-packages-rewirebench-src-rewirebench-protocols-mrnabench-py; runner-04-packages-rewirebench-src-rewirebench-resources-mrnabench-sources-json; runner-04-packages-rewirebench-src-rewirebench-resources-mrnabench-validation-receipt-json; source locator: docs/mrnabench.md; source manifest; protocol score function; automated receipt scope; id: mrnabench-official; protocol id: discovery-benchmark-mrnabench; version: 74f96b8e6ae9f41cc3cccff089d826a62d5604b8; title: Run a linear probe over mRNA embeddings; purpose: generate_and_evaluate; summary: Install the package, embed a dataset with a supported model, and score the linear probe this page reports.; inputs: A supported embedding model, or your own embeddings.; outputs: Per-task probe scores over the benchmark's splits.; requirements: data: Downloaded by the package.; weights: A published mRNA or nucleotide model checkpoint.; licence: Project licence: AGPL-3.0. Upstream data licences are separate and unreported here.; software: Python with the mrna-bench package and the model's own environment.; hardware: Not stated in the cited section. Several of these steps expect a GPU.; instructions: runtime: command_line; title: Install; code: pip install mrna-bench; status: source_reviewed_not_executed; source ids: project-recipe-mrnabench-74f96b8e; source locator: README.md at 74f96b8e, Datasets Only, lines 34-34; runtime: python; title: Load a dataset; code: import mrna_bench as mb dataset = mb.load_dataset("go-mf") data_df = dataset.data_df; status: source_reviewed_not_executed; source ids: project-recipe-mrnabench-74f96b8e; source locator: README.md at 74f96b8e, Usage, lines 77-80; runtime: python; title: Embed and evaluate; code: import torch import mrna_bench as mb from mrna_bench.embedder import DatasetEmbedder from mrna_bench.linear_probe import LinearProbeBuilder device = torch.device("cuda") dataset = mb.load_dataset("go-mf") model = mb.load_model("Orthrus", "orthrus-large-6-track", device) embedder = DatasetEmbedder(model, dataset) embeddings = embedder.embed_dataset() embeddings = embeddings.detach().cpu().numpy() prober = (LinearProbeBuilder(dataset) .fetch_embedding_by_embedding_instance("orthrus-large-6", embeddings) .build_splitter("homology", species="human", eval_all_splits=False) .build_evaluator("multilabel") .set_target("target") .build() ) metrics = prober.run_linear_probe(2541) print(metrics); status: source_reviewed_not_executed; source ids: project-recipe-mrnabench-74f96b8e; source locator: README.md at 74f96b8e, Usage, lines 85-109; limitations: Quoted from the project's README and not executed by rewire, so the commands are evidence of what the project documents rather than a verified run.; The project may have changed since the pinned commit.; Metrics alternate between AUPRC on a percentage scale and Pearson R.; Each row on this page is the best checkpoint of a family, chosen by the authors.; source ids: project-recipe-mrnabench-74f96b8e; source locator: README.md at 74f96b8e
Related records

Suggest a correction