rewire.itbenchmarks
Benchmark

BEACON

BEACON compares RNA representations across structural and functional downstream tasks.

Sourcesterry-r123/RNABenchmark official source · Pinned README: Dataset; task list; Models and Model settings

221 evaluations · 221 results

Overview

Datasets

Tasks include secondary structure, contact/distance maps, RNA-family classification, modification and expression-related outcomes.

Sourcesterry-r123/RNABenchmark official source · Pinned README: Dataset; task list; Models and Model settings

Metrics

F1 for secondary structure; top-L precision for contacts; R² for distance maps, imputation, APA, ribosome loading and switches; top-k accuracy for splice sites; accuracy for ncRNA class; AUC for modification; MCRMSE for degradation; weighted Spearman correlation for CRISPR tasks.

Sourcesbeacon primary benchmark evidence · Sections 3.1–3.3, 4 and 5.1; Table 1; Appendix A

Allowed inputs

RNA sequences; task-specific targets are supplied in the benchmark datasets.

Sourcesterry-r123/RNABenchmark official source · Pinned README: Dataset; task list; Models and Model settings
Evaluation procedure diagram
How it worksEvaluation procedure
Evaluation procedure1. Allowed inputs: RNA sequences; task-specific targets are supplied in the benchmark datasets.. Then: 2. Splits: Table 1 publishes separate training, validation and test sizes for all 13 tasks; these reuse different source datasets and are not one shared RNA partition. Structure-map tasks share their 188/23/80 partition, while other tasks use their own supplied folds.. Then: 3. Metrics: F1 for secondary structure; top-L precision for contacts; R² for distance maps, imputation, APA, ribosome loading and switches; top-k accuracy for splice sites; accuracy for ncRNA class; AUC for modification; MCRMSE for degradation; weighted Spearman correlation for CRISPR tasks.Evaluation procedure1. Allowed inputs: RNA sequences; task-specific targets are supplied in the benchmark datasets.. Then: 2. Splits: Table 1 publishes separate training, validation and test sizes for all 13 tasks; these reuse different source datasets and are not one shared RNA partition. Structure-map tasks share their 188/23/80 partition, while other tasks use their own supplied folds.. Then: 3. Metrics: F1 for secondary structure; top-L precision for contacts; R² for distance maps, imputation, APA, ribosome loading and switches; top-k accuracy for splice sites; accuracy for ncRNA class; AUC for modification; MCRMSE for degradation; weighted Spearman correlation for CRISPR tasks.Evaluation procedure1. Allowed inputs: RNA sequences; task-specific targets are supplied in the benchmark datasets.. Then: 2. Splits: Table 1 publishes separate training, validation and test sizes for all 13 tasks; these reuse different source datasets and are not one shared RNA partition. Structure-map tasks share their 188/23/80 partition, while other tasks use their own supplied folds.. Then: 3. Metrics: F1 for secondary structure; top-L precision for contacts; R² for distance maps, imputation, APA, ribosome loading and switches; top-k accuracy for splice sites; accuracy for ncRNA class; AUC for modification; MCRMSE for degradation; weighted Spearman correlation for CRISPR tasks.

Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.

Sources (2)terry-r123/RNABenchmark official source; beacon primary benchmark evidence · Pinned README: Dataset; task list; Models and Model settings; Sections 3.1–3.3, 4 and 5.1; Table 1; Appendix A

limited source coverage · Automated source review, 2026-09-16. All specifications and missing details

Results

Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.

BEACON APA: Alternative polyadenylation isoform prediction

r2 (percent) · Higher values are better.

BEACON APA: Alternative polyadenylation isoform prediction · APARENT (BEACON split)

Evidence origin: Author-reported evaluation.

BEACON: Benchmark for Comprehensive RNA Tasks and Language Models (arXiv:2406.10391v2) · Table3,p.8,column(APA)
  • The Literature SOTA row is excluded: those numbers come from other papers under their own protocols.
  • Metrics differ between tasks, so these figures cannot be averaged into one RNA score.
Comparison details and limitations

Every method in BEACON Table 3 on Alternative polyadenylation isoform prediction, scored with R2 on APARENT with the split 145,463/33,170/49,755.

  • Author-reported numbers, source checked but not independently reproduced.

Automated source review: 2026-09-18. Numerical source review does not establish independent reproduction.

Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.

Showing 12 of 17 matching rows.

Methods and evaluation design

Procedure, tasks and evaluated configurations

How it works

Evaluation methodology

BEACON tests RNA structure, function and engineering with 13 separately labelled datasets. Models predict nucleotide-level labels, pairwise structural maps or sequence-level properties. Its supplied task partitions and metrics must be preserved, and results are reported over three training seeds.

Sourcesbeacon primary benchmark evidence · Sections 3.1–3.3, 4 and 5.1; Table 1; Appendix A

Evaluation design

Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.

These source-backed links do not make different protocols or scores interchangeable.

Baseline coverage

Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.

No concrete protocols are explicitly linked to this suite. Protocol identification and baseline selection are outstanding.

Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums

Coverage is derived from release 2026-09-29-06401fd5b220. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.

Run this benchmark

Fine-tune a model on a BEACON task

Clone the benchmark, then fine-tune a model on one of the thirteen RNA tasks scored on this page.

Generate predictions and evaluate them. This recipe does not establish reproduction of a particular published score.

Dataset access
Laid out under the repository's data directory as the README describes.
Model and weights
A published RNA language model checkpoint, or none for the supervised baselines.
Licences
Project licence: Apache-2.0. Upstream data licences are separate and unreported here.
Software
Python with the repository's environment.
Hardware
Not stated in the cited section. Several of these steps expect a GPU.
Required inputs and expected outputs

Inputs

  • A model checkpoint and the task data laid out as the README shows.

Outputs

  • A fine-tuned model and its score on the task's test split.

Execution steps

  1. 1. Clone and install (Command line)

    Source reviewed; these instructions have not been executed by rewire.

    git clone https://github.com/terry-r123/RNABenchmark.git
    cd RNABenchmark
    conda create -n beacon python=3.8
    pip install -r requirements.txt
    BEACON: repository README · README.md at da7f9c7a, Installation, lines 20-23
  2. 2. Fine-tune on a task (Command line)

    Source reviewed; these instructions have not been executed by rewire.

    cd RNABenchmark
    bash ./scripts/BEACON-B/all_task.sh
    BEACON: repository README · README.md at da7f9c7a, Finetuning, lines 148-149
  3. 3. Compute embeddings (Python)

    Source reviewed; these instructions have not been executed by rewire.

    import os, sys
    current_path = os.path.dirname(os.path.abspath(__file__))
    parent_dir = os.path.dirname(current_path)
    sys.path.append(parent_dir)
    from model.utrlm.modeling_utrlm import UtrLmModel
    from tokenizer.tokenization_opensource import OpenRnaLMTokenizer
    
    tokenizer = OpenRnaLMTokenizer.from_pretrained('./checkpoint/opensource/utr-lm-mrl', model_max_length=1026, padding_side="right", use_fast=True,)
    model = UtrLmModel.from_pretrained('./checkpoint/opensource/utr-lm-mrl')     
    sequences = ["AUUCCGAUUCCGAUUCCG"]
    output = tokenizer.batch_encode_plus(sequences, return_tensors="pt", padding="longest", max_length = 1026, truncation=True)
    input_ids = output["input_ids"]
    attention_mask = output["attention_mask"]
    
    embedding = model(input_ids=input_ids,attention_mask=attention_mask)[0] # shape [bz,length, hidden_size]
    print(embedding.shape)
    BEACON: repository README · README.md at da7f9c7a, Computing embeddings, lines 155-170

Use your own model

Run your model locally and return predictions keyed by the input IDs. The evaluator supplies biological inputs without test labels and owns scoring. This interface is not a sandbox for model code.

Pass your existing prediction function into this adapter. Its output direction must match the selected protocol.

class MyModelAdapter:
    def __init__(self, score):
        self.score = score

    def predict(self, inputs):
        return {row["id"]: float(self.score(row)) for row in inputs}

# adapter = MyModelAdapter(your_prediction_function)
# report = rewirebench.run(prepared, adapter, output="runs/my-model")

Alternatively, generate a keyed prediction file in your existing model environment and use the score-only recipe. Your model code and weights do not need to be shared.

BEACON: repository README · README.md at da7f9c7a
Scope and limitations
  • Quoted from the project's README and not executed by rewire, so the commands are evidence of what the project documents rather than a verified run.
  • The project may have changed since the pinned commit.
  • Metrics differ by task, and VDP is an error where lower is better.
  • The Literature SOTA row in the paper is not part of this benchmark's own runs.

Contribute a result for review. The library can submit an exported evaluation for private review when intake is open. Check the contribution page for access and sign-in.

Original repository instructions

Run this benchmark

The official RNA benchmark links downloadable data/checkpoints and all-task fine-tuning scripts. The documented stack includes Python 3.8, Torch 1.13.1+cu117 and Transformers 4.38.1; configure local data/model paths before executing the scripts. A minimal independent command sequence has not been validated in this pass.

A maintained rewire runner has not been verified for this benchmark. Check data access, weights, licences, dependencies and hardware in the linked official documentation; requirements have not been fully extracted.

terry-r123/RNABenchmark / README.md · README.md lines 13–32 and 144–172 (Prerequisites, data/checkpoints, Usage)
Strengths, limitations and unresolved questions

Strengths and limitations

Strengths and considerations

  • Multiple RNA endpoints test distinct representation properties rather than one aggregate biological claim.
    Sourcesterry-r123/RNABenchmark official source · Pinned README: Dataset; task list; Models and Model settings

Limitations and conditions

  • BEACON combines heterogeneous source assays. Its published task partitions do not establish a common pretraining-overlap audit or a single family-held-out generalization test.
    Sourcesbeacon primary benchmark evidence · Sections 3.1–3.3, 4 and 5.1; Table 1; Appendix A
Profile review details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Stable record: discovery-benchmark-beacon

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetsTasks include secondary structure, contact/distance maps, RNA-family classification, modification and expression-related outcomes.
Sourcesterry-r123/RNABenchmark official source · Pinned README: Dataset; task list; Models and Model settings
SplitsTable 1 publishes separate training, validation and test sizes for all 13 tasks; these reuse different source datasets and are not one shared RNA partition. Structure-map tasks share their 188/23/80 partition, while other tasks use their own supplied folds.
Sourcesbeacon primary benchmark evidence · Sections 3.1–3.3, 4 and 5.1; Table 1; Appendix A
MetricsF1 for secondary structure; top-L precision for contacts; R² for distance maps, imputation, APA, ribosome loading and switches; top-k accuracy for splice sites; accuracy for ncRNA class; AUC for modification; MCRMSE for degradation; weighted Spearman correlation for CRISPR tasks.
Sourcesbeacon primary benchmark evidence · Sections 3.1–3.3, 4 and 5.1; Table 1; Appendix A
BaselinesRNA-FM, RNABERT, RNA-MSM, SpliceBERT, UTR-LM, UTRBERT and BEACON variants are listed.
Sourcesterry-r123/RNABenchmark official source · Pinned README: Dataset; task list; Models and Model settings
Leakage controlsThe paper specifies source datasets and per-task partitions, but does not establish one suite-wide homology or RNA-family exclusion rule. A train/test size table alone does not demonstrate independence from model pretraining. · Not reported in inspected sources
Sourcesbeacon primary benchmark evidence · Sections 3.1–3.3, 4 and 5.1; Table 1; Appendix A
UncertaintyExperiments are repeated with three random seeds; Section 5.1 reports their mean and sample standard deviation.
Sourcesbeacon primary benchmark evidence · Sections 3.1–3.3, 4 and 5.1; Table 1; Appendix A
Entity typeRNA model benchmark suite (BEACON).
Sourcesterry-r123/RNABenchmark official source · Pinned README: Dataset; task list; Models and Model settings
OrganismsTask dependent: human HEK293 icSHAPE data and human splice/UTR assays coexist with RNA structure collections and synthetic constructs. The programmable-switch dataset includes sequences from viral genomes and human transcription factors; the suite is not a single-organism assay.
Sourcesbeacon primary benchmark evidence · Sections 3.1–3.3, 4 and 5.1; Table 1; Appendix A
AssaysSecondary structure, contact/distance, family, modification and expression-related labels.
Sourcesterry-r123/RNABenchmark official source · Pinned README: Dataset; task list; Models and Model settings
Allowed inputsRNA sequences; task-specific targets are supplied in the benchmark datasets.
Sourcesterry-r123/RNABenchmark official source · Pinned README: Dataset; task list; Models and Model settings
AdaptationTask-specific supervised evaluation using the supplied training scripts/configurations.
Sourcesterry-r123/RNABenchmark official source · Pinned README: Dataset; task list; Models and Model settings
Applicable tests and references

Applicability is distinct from a completed evaluation.

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.

Paper or primary resourceVersionReference
BEACON: Benchmark for Comprehensive RNA Tasks and Language Models2406.10391v1Read source
Historical gaps recorded on 2026-09-17

The catalogue now holds 221 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.

  • Multi-task suite; model adaptation and metric differ by task. Full comparison table extraction remains pending; no composite RNA score inferred.
Search and extraction details

source found structured extraction pending

Searches

  • BEACON RNA benchmark 2406.10391 Table 2

Evidence locations

  • Full arXiv v1 PDF; benchmark task definitions and comparison tables

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

29 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Individual claims
beacon primary benchmark evidence

Original source ↗

Pinned README: Dataset; task list; Models and Model settings; Sections 3.1–3.3, 4 and 5.1; Table 1; Appendix A

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2406.10391v1
Retrieved: 2026-09-16T21:04:55.172530+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: c2496bff164b87e635ba253ea5a5edc55c94c4673e81c071bc8f0db1e288d006

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram caption
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Individual claims
terry-r123/RNABenchmark official source

Original source ↗

Pinned README: Dataset; task list; Models and Model settings; Sections 3.1–3.3, 4 and 5.1; Table 1; Appendix A

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: da7f9c7ac3f39605af27e1dfcdf879adba963d79
Retrieved: 2026-09-16T10:30:21.025462+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 0403f84453aace301c7d02895d94a977ccbbec77c3a94a49a23d0d529dd48d31

Hash scope: Hash scope not separately documented; inspect source record

Diagram steps
  • Allowed inputs: RNA sequences; task-specific targets are supplied in the benchmark datasets.
  • Splits: Table 1 publishes separate training, validation and test sizes for all 13 tasks; these reuse different source datasets and are not one shared RNA partition. Structure-map tasks share their 188/23/80 partition, while other tasks use their own supplied folds.
  • Metrics: F1 for secondary structure; top-L precision for contacts; R² for distance maps, imputation, APA, ribosome loading and switches; top-k accuracy for splice sites; accuracy for ncRNA class; AUC for modification; MCRMSE for degradation; weighted Spearman correlation for CRISPR tasks.
Individual claims
beacon primary benchmark evidence

Original source ↗

Pinned README: Dataset; task list; Models and Model settings; Sections 3.1–3.3, 4 and 5.1; Table 1; Appendix A

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2406.10391v1
Retrieved: 2026-09-16T21:04:55.172530+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: c2496bff164b87e635ba253ea5a5edc55c94c4673e81c071bc8f0db1e288d006

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps
  • Allowed inputs: RNA sequences; task-specific targets are supplied in the benchmark datasets.
  • Splits: Table 1 publishes separate training, validation and test sizes for all 13 tasks; these reuse different source datasets and are not one shared RNA partition. Structure-map tasks share their 188/23/80 partition, while other tasks use their own supplied folds.
  • Metrics: F1 for secondary structure; top-L precision for contacts; R² for distance maps, imputation, APA, ribosome loading and switches; top-k accuracy for splice sites; accuracy for ncRNA class; AUC for modification; MCRMSE for degradation; weighted Spearman correlation for CRISPR tasks.
Individual claims
terry-r123/RNABenchmark official source

Original source ↗

Pinned README: Dataset; task list; Models and Model settings; Sections 3.1–3.3, 4 and 5.1; Table 1; Appendix A

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: da7f9c7ac3f39605af27e1dfcdf879adba963d79
Retrieved: 2026-09-16T10:30:21.025462+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 0403f84453aace301c7d02895d94a977ccbbec77c3a94a49a23d0d529dd48d31

Hash scope: Hash scope not separately documented; inspect source record

Diagram title
Evaluation procedure
Individual claims
beacon primary benchmark evidence

Original source ↗

Pinned README: Dataset; task list; Models and Model settings; Sections 3.1–3.3, 4 and 5.1; Table 1; Appendix A

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2406.10391v1
Retrieved: 2026-09-16T21:04:55.172530+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: c2496bff164b87e635ba253ea5a5edc55c94c4673e81c071bc8f0db1e288d006

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title
Evaluation procedure
Individual claims
terry-r123/RNABenchmark official source

Original source ↗

Pinned README: Dataset; task list; Models and Model settings; Sections 3.1–3.3, 4 and 5.1; Table 1; Appendix A

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: da7f9c7ac3f39605af27e1dfcdf879adba963d79
Retrieved: 2026-09-16T10:30:21.025462+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: 0403f84453aace301c7d02895d94a977ccbbec77c3a94a49a23d0d529dd48d31

Hash scope: Hash scope not separately documented; inspect source record

Datasets
Tasks include secondary structure, contact/distance maps, RNA-family classification, modification and expression-related outcomes.
Individual claims
terry-r123/RNABenchmark official source

Original source ↗

Pinned README: Dataset; task list; Models and Model settings

Version: da7f9c7ac3f39605af27e1dfcdf879adba963d79
Retrieved: 2026-09-16T10:30:21.025462+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: 0403f84453aace301c7d02895d94a977ccbbec77c3a94a49a23d0d529dd48d31

Hash scope: Hash scope not separately documented; inspect source record

Splits
Table 1 publishes separate training, validation and test sizes for all 13 tasks; these reuse different source datasets and are not one shared RNA partition. Structure-map tasks share their 188/23/80 partition, while other tasks use their own supplied folds.
Individual claims
beacon primary benchmark evidence

Original source ↗

Sections 3.1–3.3, 4 and 5.1; Table 1; Appendix A

Version: 2406.10391v1
Retrieved: 2026-09-16T21:04:55.172530+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: c2496bff164b87e635ba253ea5a5edc55c94c4673e81c071bc8f0db1e288d006

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Adaptation
Task-specific supervised evaluation using the supplied training scripts/configurations.
Individual claims
terry-r123/RNABenchmark official source

Original source ↗

Pinned README: Dataset; task list; Models and Model settings

Version: da7f9c7ac3f39605af27e1dfcdf879adba963d79
Retrieved: 2026-09-16T10:30:21.025462+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: 0403f84453aace301c7d02895d94a977ccbbec77c3a94a49a23d0d529dd48d31

Hash scope: Hash scope not separately documented; inspect source record

Metrics
F1 for secondary structure; top-L precision for contacts; R² for distance maps, imputation, APA, ribosome loading and switches; top-k accuracy for splice sites; accuracy for ncRNA class; AUC for modification; MCRMSE for degradation; weighted Spearman correlation for CRISPR tasks.
Individual claims
beacon primary benchmark evidence

Original source ↗

Sections 3.1–3.3, 4 and 5.1; Table 1; Appendix A

Version: 2406.10391v1
Retrieved: 2026-09-16T21:04:55.172530+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: c2496bff164b87e635ba253ea5a5edc55c94c4673e81c071bc8f0db1e288d006

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: discovered

5 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: discovery-benchmark-beacon

areas
rna
entity level
suite
scope note
Specialist molecular or omics evaluation; protocol details require review before numerical comparison.
task
RNA structure, function and engineering tasks
version
Not reported
benchmark research
review date: 2026-09-17; status: source_found_structured_extraction_pending; primary sources: evidence-expansion-p2-evidence-discovery-final-beacon-c2496bff164b; inspected locators: Full arXiv v1 PDF; benchmark task definitions and comparison tables; searched queries: BEACON RNA benchmark 2406.10391 Table 2; gaps: Multi-task suite; model adaptation and metric differ by task. Full comparison table extraction remains pending; no composite RNA score inferred.; claim scope: Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.
historical missing metadata
dataset release: unextracted; metric implementation: unextracted; split manifest: unextracted; version: unextracted
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
entity classification
review date: 2026-09-17; rationale: The cited profile describes a collection of evaluation tasks or protocols; retain it as the top-level benchmark suite. Its datasets and individual protocols remain separate records.; source ids: src-discovery-terry-r123-rnabenchmark; source locator: Pinned README: Dataset; task list; Models and Model settings; ambiguities: None recorded
run documentation
record id: discovery-benchmark-beacon; source ids: run-doc-beacon-readme-md-da7f9c7a; status: official_documentation_linked; summary: The official RNA benchmark links downloadable data/checkpoints and all-task fine-tuning scripts. The documented stack includes Python 3.8, Torch 1.13.1+cu117 and Transformers 4.38.1; configure local data/model paths before executing the scripts. A minimal independent command sequence has not been validated in this pass.; source locator: README.md lines 13–32 and 144–172 (Prerequisites, data/checkpoints, Usage)
run recipes
id: beacon-official; protocol id: discovery-benchmark-beacon; version: da7f9c7ac3f39605af27e1dfcdf879adba963d79; title: Fine-tune a model on a BEACON task; purpose: generate_and_evaluate; summary: Clone the benchmark, then fine-tune a model on one of the thirteen RNA tasks scored on this page.; inputs: A model checkpoint and the task data laid out as the README shows.; outputs: A fine-tuned model and its score on the task's test split.; requirements: data: Laid out under the repository's data directory as the README describes.; weights: A published RNA language model checkpoint, or none for the supervised baselines.; licence: Project licence: Apache-2.0. Upstream data licences are separate and unreported here.; software: Python with the repository's environment.; hardware: Not stated in the cited section. Several of these steps expect a GPU.; instructions: runtime: command_line; title: Clone and install; code: git clone https://github.com/terry-r123/RNABenchmark.git cd RNABenchmark conda create -n beacon python=3.8 pip install -r requirements.txt; status: source_reviewed_not_executed; source ids: project-recipe-beacon-da7f9c7a; source locator: README.md at da7f9c7a, Installation, lines 20-23; runtime: command_line; title: Fine-tune on a task; code: cd RNABenchmark bash ./scripts/BEACON-B/all_task.sh; status: source_reviewed_not_executed; source ids: project-recipe-beacon-da7f9c7a; source locator: README.md at da7f9c7a, Finetuning, lines 148-149; runtime: python; title: Compute embeddings; code: import os, sys current_path = os.path.dirname(os.path.abspath(__file__)) parent_dir = os.path.dirname(current_path) sys.path.append(parent_dir) from model.utrlm.modeling_utrlm import UtrLmModel from tokenizer.tokenization_opensource import OpenRnaLMTokenizer tokenizer = OpenRnaLMTokenizer.from_pretrained('./checkpoint/opensource/utr-lm-mrl', model_max_length=1026, padding_side="right", use_fast=True,) model = UtrLmModel.from_pretrained('./checkpoint/opensource/utr-lm-mrl') sequences = ["AUUCCGAUUCCGAUUCCG"] output = tokenizer.batch_encode_plus(sequences, return_tensors="pt", padding="longest", max_length = 1026, truncation=True) input_ids = output["input_ids"] attention_mask = output["attention_mask"] embedding = model(input_ids=input_ids,attention_mask=attention_mask)[0] # shape [bz,length, hidden_size] print(embedding.shape); status: source_reviewed_not_executed; source ids: project-recipe-beacon-da7f9c7a; source locator: README.md at da7f9c7a, Computing embeddings, lines 155-170; limitations: Quoted from the project's README and not executed by rewire, so the commands are evidence of what the project documents rather than a verified run.; The project may have changed since the pinned commit.; Metrics differ by task, and VDP is an error where lower is better.; The Literature SOTA row in the paper is not part of this benchmark's own runs.; source ids: project-recipe-beacon-da7f9c7a; source locator: README.md at da7f9c7a
Related records

Suggest a correction