Datasets
Tasks include secondary structure, contact/distance maps, RNA-family classification, modification and expression-related outcomes.
BEACON compares RNA representations across structural and functional downstream tasks.
Tasks include secondary structure, contact/distance maps, RNA-family classification, modification and expression-related outcomes.
F1 for secondary structure; top-L precision for contacts; R² for distance maps, imputation, APA, ribosome loading and switches; top-k accuracy for splice sites; accuracy for ncRNA class; AUC for modification; MCRMSE for degradation; weighted Spearman correlation for CRISPR tasks.
RNA sequences; task-specific targets are supplied in the benchmark datasets.
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
limited source coverage · Automated source review, 2026-09-16. All specifications and missing details
Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.
r2 (percent) · Higher values are better.
BEACON APA: Alternative polyadenylation isoform prediction · APARENT (BEACON split)
Evidence origin: Author-reported evaluation.
BEACON: Benchmark for Comprehensive RNA Tasks and Language Models (arXiv:2406.10391v2) · Table3,p.8,column(APA)Every method in BEACON Table 3 on Alternative polyadenylation isoform prediction, scored with R2 on APARENT with the split 145,463/33,170/49,755.
Automated source review: 2026-09-18. Numerical source review does not establish independent reproduction.
Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.
Showing 12 of 17 matching rows.
BEACON tests RNA structure, function and engineering with 13 separately labelled datasets. Models predict nucleotide-level labels, pairwise structural maps or sequence-level properties. Its supplied task partitions and metrics must be preserved, and results are reported over three training seeds.
Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.
These source-backed links do not make different protocols or scores interchangeable.
Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.
No concrete protocols are explicitly linked to this suite. Protocol identification and baseline selection are outstanding.
Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums
Coverage is derived from release 2026-09-29-06401fd5b220. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.
Clone the benchmark, then fine-tune a model on one of the thirteen RNA tasks scored on this page.
Generate predictions and evaluate them. This recipe does not establish reproduction of a particular published score.
Source reviewed; these instructions have not been executed by rewire.
git clone https://github.com/terry-r123/RNABenchmark.git
cd RNABenchmark
conda create -n beacon python=3.8
pip install -r requirements.txtBEACON: repository README · README.md at da7f9c7a, Installation, lines 20-23Source reviewed; these instructions have not been executed by rewire.
cd RNABenchmark
bash ./scripts/BEACON-B/all_task.shBEACON: repository README · README.md at da7f9c7a, Finetuning, lines 148-149Source reviewed; these instructions have not been executed by rewire.
import os, sys
current_path = os.path.dirname(os.path.abspath(__file__))
parent_dir = os.path.dirname(current_path)
sys.path.append(parent_dir)
from model.utrlm.modeling_utrlm import UtrLmModel
from tokenizer.tokenization_opensource import OpenRnaLMTokenizer
tokenizer = OpenRnaLMTokenizer.from_pretrained('./checkpoint/opensource/utr-lm-mrl', model_max_length=1026, padding_side="right", use_fast=True,)
model = UtrLmModel.from_pretrained('./checkpoint/opensource/utr-lm-mrl')
sequences = ["AUUCCGAUUCCGAUUCCG"]
output = tokenizer.batch_encode_plus(sequences, return_tensors="pt", padding="longest", max_length = 1026, truncation=True)
input_ids = output["input_ids"]
attention_mask = output["attention_mask"]
embedding = model(input_ids=input_ids,attention_mask=attention_mask)[0] # shape [bz,length, hidden_size]
print(embedding.shape)BEACON: repository README · README.md at da7f9c7a, Computing embeddings, lines 155-170Run your model locally and return predictions keyed by the input IDs. The evaluator supplies biological inputs without test labels and owns scoring. This interface is not a sandbox for model code.
Pass your existing prediction function into this adapter. Its output direction must match the selected protocol.
class MyModelAdapter:
def __init__(self, score):
self.score = score
def predict(self, inputs):
return {row["id"]: float(self.score(row)) for row in inputs}
# adapter = MyModelAdapter(your_prediction_function)
# report = rewirebench.run(prepared, adapter, output="runs/my-model")Alternatively, generate a keyed prediction file in your existing model environment and use the score-only recipe. Your model code and weights do not need to be shared.
BEACON: repository README · README.md at da7f9c7aContribute a result for review. The library can submit an exported evaluation for private review when intake is open. Check the contribution page for access and sign-in.
The official RNA benchmark links downloadable data/checkpoints and all-task fine-tuning scripts. The documented stack includes Python 3.8, Torch 1.13.1+cu117 and Transformers 4.38.1; configure local data/model paths before executing the scripts. A minimal independent command sequence has not been validated in this pass.
A maintained rewire runner has not been verified for this benchmark. Check data access, weights, licences, dependencies and hardware in the linked official documentation; requirements have not been fully extracted.
terry-r123/RNABenchmark / README.md · README.md lines 13–32 and 144–172 (Prerequisites, data/checkpoints, Usage)Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.
Stable record: discovery-benchmark-beaconExplanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Datasets | Tasks include secondary structure, contact/distance maps, RNA-family classification, modification and expression-related outcomes.Sourcesterry-r123/RNABenchmark official source · Pinned README: Dataset; task list; Models and Model settings |
| Splits | Table 1 publishes separate training, validation and test sizes for all 13 tasks; these reuse different source datasets and are not one shared RNA partition. Structure-map tasks share their 188/23/80 partition, while other tasks use their own supplied folds.Sourcesbeacon primary benchmark evidence · Sections 3.1–3.3, 4 and 5.1; Table 1; Appendix A |
| Metrics | F1 for secondary structure; top-L precision for contacts; R² for distance maps, imputation, APA, ribosome loading and switches; top-k accuracy for splice sites; accuracy for ncRNA class; AUC for modification; MCRMSE for degradation; weighted Spearman correlation for CRISPR tasks.Sourcesbeacon primary benchmark evidence · Sections 3.1–3.3, 4 and 5.1; Table 1; Appendix A |
| Baselines | RNA-FM, RNABERT, RNA-MSM, SpliceBERT, UTR-LM, UTRBERT and BEACON variants are listed.Sourcesterry-r123/RNABenchmark official source · Pinned README: Dataset; task list; Models and Model settings |
| Leakage controls | The paper specifies source datasets and per-task partitions, but does not establish one suite-wide homology or RNA-family exclusion rule. A train/test size table alone does not demonstrate independence from model pretraining. · Not reported in inspected sourcesSourcesbeacon primary benchmark evidence · Sections 3.1–3.3, 4 and 5.1; Table 1; Appendix A |
| Uncertainty | Experiments are repeated with three random seeds; Section 5.1 reports their mean and sample standard deviation.Sourcesbeacon primary benchmark evidence · Sections 3.1–3.3, 4 and 5.1; Table 1; Appendix A |
| Entity type | RNA model benchmark suite (BEACON).Sourcesterry-r123/RNABenchmark official source · Pinned README: Dataset; task list; Models and Model settings |
| Organisms | Task dependent: human HEK293 icSHAPE data and human splice/UTR assays coexist with RNA structure collections and synthetic constructs. The programmable-switch dataset includes sequences from viral genomes and human transcription factors; the suite is not a single-organism assay.Sourcesbeacon primary benchmark evidence · Sections 3.1–3.3, 4 and 5.1; Table 1; Appendix A |
| Assays | Secondary structure, contact/distance, family, modification and expression-related labels.Sourcesterry-r123/RNABenchmark official source · Pinned README: Dataset; task list; Models and Model settings |
| Allowed inputs | RNA sequences; task-specific targets are supplied in the benchmark datasets.Sourcesterry-r123/RNABenchmark official source · Pinned README: Dataset; task list; Models and Model settings |
| Adaptation | Task-specific supervised evaluation using the supplied training scripts/configurations.Sourcesterry-r123/RNABenchmark official source · Pinned README: Dataset; task list; Models and Model settings |
Applicability is distinct from a completed evaluation.
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Last literature check: 2026-09-17. Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.
| Paper or primary resource | Version | Reference |
|---|---|---|
| BEACON: Benchmark for Comprehensive RNA Tasks and Language Models | 2406.10391v1 | Read source |
The catalogue now holds 221 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.
source found structured extraction pending
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
29 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions. Individual claims | beacon primary benchmark evidence Pinned README: Dataset; task list; Models and Model settings; Sections 3.1–3.3, 4 and 5.1; Table 1; Appendix A Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 2406.10391v1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram caption Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions. Individual claims | terry-r123/RNABenchmark official source Pinned README: Dataset; task list; Models and Model settings; Sections 3.1–3.3, 4 and 5.1; Table 1; Appendix A Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: da7f9c7ac3f39605af27e1dfcdf879adba963d79 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| beacon primary benchmark evidence Pinned README: Dataset; task list; Models and Model settings; Sections 3.1–3.3, 4 and 5.1; Table 1; Appendix A Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 2406.10391v1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| terry-r123/RNABenchmark official source Pinned README: Dataset; task list; Models and Model settings; Sections 3.1–3.3, 4 and 5.1; Table 1; Appendix A Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: da7f9c7ac3f39605af27e1dfcdf879adba963d79 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Evaluation procedure Individual claims | beacon primary benchmark evidence Pinned README: Dataset; task list; Models and Model settings; Sections 3.1–3.3, 4 and 5.1; Table 1; Appendix A Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 2406.10391v1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Evaluation procedure Individual claims | terry-r123/RNABenchmark official source Pinned README: Dataset; task list; Models and Model settings; Sections 3.1–3.3, 4 and 5.1; Table 1; Appendix A Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: da7f9c7ac3f39605af27e1dfcdf879adba963d79 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Datasets Tasks include secondary structure, contact/distance maps, RNA-family classification, modification and expression-related outcomes. Individual claims | terry-r123/RNABenchmark official source Pinned README: Dataset; task list; Models and Model settings Version: da7f9c7ac3f39605af27e1dfcdf879adba963d79 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Splits Table 1 publishes separate training, validation and test sizes for all 13 tasks; these reuse different source datasets and are not one shared RNA partition. Structure-map tasks share their 188/23/80 partition, while other tasks use their own supplied folds. Individual claims | beacon primary benchmark evidence Sections 3.1–3.3, 4 and 5.1; Table 1; Appendix A Version: 2406.10391v1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Adaptation Task-specific supervised evaluation using the supplied training scripts/configurations. Individual claims | terry-r123/RNABenchmark official source Pinned README: Dataset; task list; Models and Model settings Version: da7f9c7ac3f39605af27e1dfcdf879adba963d79 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Metrics F1 for secondary structure; top-L precision for contacts; R² for distance maps, imputation, APA, ribosome loading and switches; top-k accuracy for splice sites; accuracy for ncRNA class; AUC for modification; MCRMSE for degradation; weighted Spearman correlation for CRISPR tasks. Individual claims | beacon primary benchmark evidence Sections 3.1–3.3, 4 and 5.1; Table 1; Appendix A Version: 2406.10391v1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
View linked audit checks and correction history
Release 2026-09-29-06401fd5b220 · Record review: discovered
Stable ID: discovery-benchmark-beacon