rewire.itbenchmarks
Configuration

RNA-FM

RNA-FM learns contextual representations of RNA nucleotides for downstream RNA analyses.

Sources (2)ml4bio/RNA-FM: README.md; ml4bio/RNA-FM: fm/pretrained.py · README.md: Quick Start, embedding examples and RNA foundation model comparison table

13 evaluations · 13 results

How it worksRNA-FM workflow
RNA-FM workflow1. RNA sequence. Then: 2. Nucleotide tokenizer. Then: 3. 12-layer transformer encoder. Then: 4. Contextual nucleotide representationsRNA-FM workflow1. RNA sequence. Then: 2. Nucleotide tokenizer. Then: 3. 12-layer transformer encoder. Then: 4. Contextual nucleotide representationsRNA-FM workflow1. RNA sequence. Then: 2. Nucleotide tokenizer. Then: 3. 12-layer transformer encoder. Then: 4. Contextual nucleotide representations

Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.

Sources (2)ml4bio/RNA-FM: README.md; ml4bio/RNA-FM: fm/pretrained.py · README.md: Quick Start, embedding examples and RNA foundation model comparison table

Overview

Model type

Masked-token RNA transformer encoder

Inputs

RNA sequences tokenized at nucleotide resolution.

Outputs

Contextual token embeddings for a specified downstream RNA task.

Sources (2)ml4bio/RNA-FM: README.md; ml4bio/RNA-FM: fm/pretrained.py · README.md: Quick Start, embedding examples and RNA foundation model comparison table

limited source coverage · Automated source review, 2026-09-23. All specifications and missing details

Evaluations and results

13 evaluations · 13 results. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: RNA-FMTask: BEACON APA: Alternative polyadenylation isoform prediction
Dataset subset: APARENT (BEACON split)
70.32(0.97)% r2
percent · higher

Uncertainty: type: standard_deviation; value: 0.97

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

RNA-FM on BEACON APA: Alternative polyadenylation isoform prediction

BEACON harness, fixed downstream head per task. Train/validation/test 145,463/33,170/49,755; dataset APARENT.

Aggregation: Not reported

BEACON: Benchmark for Comprehensive RNA Tasks and Language Models (arXiv:2406.10391v2) · Table3,p.8,row(RNA-FM),column(APA)
Configuration: RNA-FMTask: BEACON CMP: Contact map prediction
Dataset subset: RNAcontact (BEACON split)
47.56(6.73)% precision_at_l
percent · higher

Uncertainty: type: standard_deviation; value: 6.73

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

RNA-FM on BEACON CMP: Contact map prediction

BEACON harness, fixed downstream head per task. Train/validation/test 188/23/80; dataset RNAcontact.

Aggregation: Not reported

BEACON: Benchmark for Comprehensive RNA Tasks and Language Models (arXiv:2406.10391v2) · Table3,p.8,row(RNA-FM),column(CMP)
Configuration: RNA-FMTask: BEACON CRI-Off: CRISPR off-target effect prediction
Dataset subset: DeepCRISPR (BEACON split)
2.49(1.56)% spearman_corr
percent · higher

Uncertainty: type: standard_deviation; value: 1.56

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

RNA-FM on BEACON CRI-Off: CRISPR off-target effect prediction

BEACON harness, fixed downstream head per task. Train/validation/test 14,223/2,032/4,064; dataset DeepCRISPR.

Aggregation: Not reported

BEACON: Benchmark for Comprehensive RNA Tasks and Language Models (arXiv:2406.10391v2) · Table3,p.8,row(RNA-FM),column(CRI-Off)
Configuration: RNA-FMTask: BEACON CRI-On: CRISPR on-target efficiency prediction
Dataset subset: DeepCRISPR (BEACON split)
31.62(1.16)% spearman_corr
percent · higher

Uncertainty: type: standard_deviation; value: 1.16

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

RNA-FM on BEACON CRI-On: CRISPR on-target efficiency prediction

BEACON harness, fixed downstream head per task. Train/validation/test 1,453/207/416; dataset DeepCRISPR.

Aggregation: Not reported

BEACON: Benchmark for Comprehensive RNA Tasks and Language Models (arXiv:2406.10391v2) · Table3,p.8,row(RNA-FM),column(CRI-On)
Configuration: RNA-FMTask: BEACON DMP: Distance map prediction
Dataset subset: RNAcontact (BEACON split)
51.45(0.51)% r2
percent · higher

Uncertainty: type: standard_deviation; value: 0.51

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

RNA-FM on BEACON DMP: Distance map prediction

BEACON harness, fixed downstream head per task. Train/validation/test 188/23/80; dataset RNAcontact.

Aggregation: Not reported

BEACON: Benchmark for Comprehensive RNA Tasks and Language Models (arXiv:2406.10391v2) · Table3,p.8,row(RNA-FM),column(DMP)
Configuration: RNA-FMTask: BEACON Modif: RNA modification site prediction
Dataset subset: MultiRM (BEACON split)
94.98(0.042)% auc
percent · higher

Uncertainty: type: standard_deviation; value: 0.042

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

RNA-FM on BEACON Modif: RNA modification site prediction

BEACON harness, fixed downstream head per task. Train/validation/test 304,661/3,599/1,200; dataset MultiRM.

Aggregation: Not reported

BEACON: Benchmark for Comprehensive RNA Tasks and Language Models (arXiv:2406.10391v2) · Table3,p.8,row(RNA-FM),column(Modif)
Configuration: RNA-FMTask: BEACON MRL: Mean ribosome loading prediction
Dataset subset: Optimus (BEACON split)
79.47(0.47)% r2
percent · higher

Uncertainty: type: standard_deviation; value: 0.47

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

RNA-FM on BEACON MRL: Mean ribosome loading prediction

BEACON harness, fixed downstream head per task. Train/validation/test 76,319/7,600/7,600; dataset Optimus.

Aggregation: Not reported

BEACON: Benchmark for Comprehensive RNA Tasks and Language Models (arXiv:2406.10391v2) · Table3,p.8,row(RNA-FM),column(MRL)
Configuration: RNA-FMTask: BEACON ncRNA: Non-coding RNA family classification
Dataset subset: Noorul's ncRNA set (BEACON split)
96.81(0.061)% accuracy
percent · higher

Uncertainty: type: standard_deviation; value: 0.061

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

RNA-FM on BEACON ncRNA: Non-coding RNA family classification

BEACON harness, fixed downstream head per task. Train/validation/test 5,679/650/2,400; dataset Noorul's ncRNA set.

Aggregation: Not reported

BEACON: Benchmark for Comprehensive RNA Tasks and Language Models (arXiv:2406.10391v2) · Table3,p.8,row(RNA-FM),column(ncRNA)
Configuration: RNA-FMTask: BEACON PRS: Programmable RNA switch prediction
Dataset subset: Angenent-Mari's switch set (BEACON split)
55.98(0.09)% r2
percent · higher

Uncertainty: type: standard_deviation; value: 0.09

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

RNA-FM on BEACON PRS: Programmable RNA switch prediction

BEACON harness, fixed downstream head per task. Train/validation/test 73,227/9,153/9,154; dataset Angenent-Mari's switch set.

Aggregation: Not reported

BEACON: Benchmark for Comprehensive RNA Tasks and Language Models (arXiv:2406.10391v2) · Table3,p.8,row(RNA-FM),column(PRS)
Configuration: RNA-FMTask: BEACON SPL: Splice site prediction
Dataset subset: SpliceAI (BEACON split)
34.84(0.87)% top_k_accuracy
percent · higher

Uncertainty: type: standard_deviation; value: 0.87

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

RNA-FM on BEACON SPL: Splice site prediction

BEACON harness, fixed downstream head per task. Train/validation/test 144,628/18,078/16,505; dataset SpliceAI.

Aggregation: Not reported

BEACON: Benchmark for Comprehensive RNA Tasks and Language Models (arXiv:2406.10391v2) · Table3,p.8,row(RNA-FM),column(SPL)
Configuration: RNA-FMTask: BEACON SSI: Structure score imputation
Dataset subset: StructureImpute (BEACON split)
42.36(0.24)% r2
percent · higher

Uncertainty: type: standard_deviation; value: 0.24

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

RNA-FM on BEACON SSI: Structure score imputation

BEACON harness, fixed downstream head per task. Train/validation/test 14,049/1,756/3,095; dataset StructureImpute.

Aggregation: Not reported

BEACON: Benchmark for Comprehensive RNA Tasks and Language Models (arXiv:2406.10391v2) · Table3,p.8,row(RNA-FM),column(SSI)
Configuration: RNA-FMTask: BEACON SSP: Secondary structure prediction
Dataset subset: bpRNA-1m (BEACON split)
68.50(0.54)% f1
percent · higher

Uncertainty: type: standard_deviation; value: 0.54

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

RNA-FM on BEACON SSP: Secondary structure prediction

BEACON harness, fixed downstream head per task. Train/validation/test 10,814/1,300/1,305; dataset bpRNA-1m.

Aggregation: Not reported

BEACON: Benchmark for Comprehensive RNA Tasks and Language Models (arXiv:2406.10391v2) · Table3,p.8,row(RNA-FM),column(SSP)
Configuration: RNA-FMTask: BEACON VDP: Vaccine degradation prediction
Dataset subset: OpenVaccine (BEACON split)
0.347(0.003) mcrmse
error · lower

Uncertainty: type: standard_deviation; value: 0.003

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

RNA-FM on BEACON VDP: Vaccine degradation prediction

BEACON harness, fixed downstream head per task. Train/validation/test 2,155/245/629; dataset OpenVaccine.

Aggregation: Not reported

BEACON: Benchmark for Comprehensive RNA Tasks and Language Models (arXiv:2406.10391v2) · Table3,p.8,row(RNA-FM),column(VDP)

Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.

Use this model

How it works, versions and access

Related profile: RNA-FM. This page retains the exact record and its evaluation context.

This configuration

Published RNA language model backbone, fine-tuned per task by the BEACON authors under their protocol.

record
RNA-FM
configuration
Not reported
entity type
Configuration

How it works

How it works

RNA-FM converts an RNA sequence into one token per nucleotide. Twelve transformer encoder blocks use self-attention to produce a 640-dimensional representation at each position. During pretraining, the model learns to recover masked nucleotides from their surrounding sequence. The resulting representations can be supplied to a separately specified downstream model; they are not, by themselves, a structure or functional prediction.

SourcesRNA-FM original methods, arXiv v5 · arXiv:2204.00300v5, Methods: ncRNA data collection and preprocessing and RNA foundation model training details (p.22)
Versions and reproducibility

The nucleotide-based rna_fm_t12 and codon-based mrna_fm_t12 interfaces are distinct. The original RNA-FM paper sets a training input-length limit of 1,024 and describes a usable limit of 1,022 nucleotides. That limit should not be assigned to mRNA-FM without checking its separate configuration.

Sources (3)ml4bio/RNA-FM: README.md; ml4bio/RNA-FM: fm/pretrained.py; RNA-FM original methods, arXiv v5 · Official README: Quick Start and RNA-FM/mRNA-FM examples; arXiv:2204.00300v5, Methods: training input length and SARS-CoV-2 genome embedding extraction (pp.22–23)
Strengths, limitations and unresolved questions

Strengths and limitations

Limitations and conditions

  • Base-level RNA-FM and codon-level mRNA-FM are not interchangeable. The task head and tokenization must be specified in each evaluation.
    Sources (2)ml4bio/RNA-FM: README.md; ml4bio/RNA-FM: fm/pretrained.py · README.md: Quick Start, embedding examples and RNA foundation model comparison table
Profile review details

Follow-up review of Checkpoint identification, Training cutoff, Weights licence. Source locations and before/after decisions are recorded in the 23 September profile-evidence audit. Other explanatory content retains its earlier source scope. No human scientific review or independent reproduction is implied.

Stable record: discovery-model-rna-fm

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-23. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Inputs, outputs and configuration
PropertyDescription and evidence
Model typeMasked-token RNA transformer encoder
Sources (2)ml4bio/RNA-FM: README.md; ml4bio/RNA-FM: fm/pretrained.py · README.md: Quick Start, embedding examples and RNA foundation model comparison table
Architecture12-layer masked-token transformer encoder with hidden width 640 and 20 attention heads; nucleotide tokens produce contextual representations.
SourcesRNA-FM original methods, arXiv v5 · arXiv:2204.00300v5, Methods: ncRNA data collection and preprocessing and RNA foundation model training details (p.22)
InputsRNA sequences tokenized at nucleotide resolution.
Sources (2)ml4bio/RNA-FM: README.md; ml4bio/RNA-FM: fm/pretrained.py · README.md: Quick Start, embedding examples and RNA foundation model comparison table
OutputsContextual token embeddings for a specified downstream RNA task.
Sources (2)ml4bio/RNA-FM: README.md; ml4bio/RNA-FM: fm/pretrained.py · README.md: Quick Start, embedding examples and RNA foundation model comparison table
Parameters99M, as printed in the official Foundation Models table.
Sources (2)ml4bio/RNA-FM: README.md; ml4bio/RNA-FM: fm/pretrained.py · README.md: Quick Start, embedding examples and RNA foundation model comparison table
Known versionsrna_fm_t12 and mrna_fm_t12 are separate pretrained interfaces.
Sources (2)ml4bio/RNA-FM: README.md; ml4bio/RNA-FM: fm/pretrained.py · README.md: Quick Start, embedding examples and RNA foundation model comparison table
Training data23.7 million non-coding RNA sequences collected from RNAcentral. The authors replace T with U and remove identical sequences using CD-HIT-EST at 100% identity, naming the resulting corpus RNAcentral100.
SourcesRNA-FM original methods, arXiv v5 · arXiv:2204.00300v5, Methods: ncRNA data collection and preprocessing and RNA foundation model training details (p.22)
Training cutoffThe original Methods describe RNAcentral100 preprocessing but do not establish a dated RNAcentral release. RNAcentral100 is a processed-corpus label, not a release number. · Not reported in inspected sources
SourcesRNA-FM original methods, arXiv v5 · arXiv2204.00300v5, Methods: Large-scale pre-training dataset, p22
Context limitsThe original paper sets a training input-length limit of 1,024 and describes a usable input limit of 1,022 nucleotides. These are the original RNA-FM settings, not a validated limit for later codon-based mRNA-FM checkpoints.
SourcesRNA-FM original methods, arXiv v5 · arXiv:2204.00300v5, Methods: RNA foundation model training details (pp.22–23) and RNA-FM application input-limit statement (p.23)
Weights licenceThe official model-comparison table labels RNA-FM MIT, while the README footer explicitly refers to source code. Separate checkpoint-distribution terms remain unverified; the linked weight repository could not be retrieved during this review. · Not reported in inspected sources
Sources (3)ml4bio/RNA-FM: README.md; ml4bio/RNA-FM: LICENSE; rnafm loader: primary artifact · README: Related RNA Language Models, RNA-FM row; License footer; LICENSE. Linked https://huggingface.co/ml4bio/RNA-FM returned HTTP401 during the review.
AccessOfficial project documentation and implementation: https://github.com/ml4bio/RNA-FM
Sources (2)ml4bio/RNA-FM: README.md; ml4bio/RNA-FM: fm/pretrained.py · README.md: Quick Start, embedding examples and RNA foundation model comparison table
Code licenceMIT
Sourcesml4bio/RNA-FM: LICENSE · LICENSE: licence text
Checkpoint identificationThe official rna_fm_t12 loader downloads RNA-FM_pretrained.pth from the authors’ server. The inspected loader does not pin a checkpoint revision or publish a digest; a code commit alone cannot identify the bytes used by a published evaluation. · Not reported in inspected sources
Sourcesrnafm loader: primary artifact · load_fm_model_and_alphabet_hub: rna_fm_t12 branch, lines155–162

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

2 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Relationship: family
discovery-model-rna-fm
Individual claims
ml4bio/RNA-FM: README.md

Original source ↗

BEACON Table 2 and Appendix A.3.2 RNAcentral ncRNA model, Table 3; mRNABench Table 2 and Appendix inventory; NABench model inventory; source-labelled configuration RNA-FM | Existing reviewed locator: Table3,p.8,row(RNA-FM)

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 348951516e0963d22bbb33b3c9fc18c89081d38e
Retrieved: 2026-09-16T19:46:19.769679+00:00

source checked

automated source review · 2026-09-23

Audit details

Source review establishes this relationship only. Exact evaluated configurations and original numerical review status remain unchanged. RNA-FM ncRNA backbone explicitly identified; mRNA-FM is distinct and must not inherit these scores.

Field: links:family:discovery-model-rna-fm

Claim: model-evaluation-identity-bef319624c59d1475fe2

Source artifact SHA-256: f9f1c1d62adc471661ca98b30c0250e9f3ce0cff7433830f149f5f48ea41c3da

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Relationship: family
discovery-model-rna-fm
Individual claims
BEACON: Benchmark for Comprehensive RNA Tasks and Language Models (arXiv:2406.10391v2)

Original source ↗

BEACON Table 2 and Appendix A.3.2 RNAcentral ncRNA model, Table 3; mRNABench Table 2 and Appendix inventory; NABench model inventory; source-labelled configuration RNA-FM | Existing reviewed locator: Table3,p.8,row(RNA-FM)

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: v2, 12 December 2024
Retrieved: 2026-09-18

source checked

automated source review · 2026-09-23

Audit details

Source review establishes this relationship only. Exact evaluated configurations and original numerical review status remain unchanged. RNA-FM ncRNA backbone explicitly identified; mRNA-FM is distinct and must not inherit these scores.

Field: links:family:discovery-model-rna-fm

Claim: model-evaluation-identity-bef319624c59d1475fe2

Source artifact SHA-256: 1370d75fe591bb8f994bc67f100a1a3fd557a62b9edb962b21f2fb038e9dea16

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: source checked

1 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: beacon-model-rna-fm

areas
rna-transcriptomes
family
pretrained
source locator
Table3,p.8,row(RNA-FM)
missing metadata
checkpoint revision: unreported; parameters: unextracted
Related records

Suggest a correction