rewire.itbenchmarks
Task

extremely long RNA species classification

Long-RNA species classification evaluates representation quality on an RNAcentral-derived seven-species dataset.

SourcesBiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Results: BiRNA-BERT significantly outperforms in extremely long sequence task; Table 2; cached text lines 37–40

1 evaluation · 1 result

Overview

Datasets

Long noncoding RNA sequences from RNAcentral with species labels.

Metrics

F1 score; the reviewed task description does not establish the averaging convention.

Allowed inputs

Long RNA sequences.

SourcesBiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Results: BiRNA-BERT significantly outperforms in extremely long sequence task; Table 2; cached text lines 37–40
Evaluation procedure diagram
How it worksComputational evaluation flow
Computational evaluation flow1. Input: Long RNA sequences.. Then: 2. Task: Long-RNA species classification evaluates representation quality on an RNAcentral-derived seven-species dataset.Computational evaluation flow1. Input: Long RNA sequences.. Then: 2. Task: Long-RNA species classification evaluates representation quality on an RNAcentral-derived seven-species dataset.Computational evaluation flow1. Input: Long RNA sequences.. Then: 2. Task: Long-RNA species classification evaluates representation quality on an RNAcentral-derived seven-species dataset.

Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.

SourcesBiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Results: BiRNA-BERT significantly outperforms in extremely long sequence task; Table 2; cached text lines 37–40

limited source coverage · Automated source review, 2026-09-16. All specifications and missing details

Results

Results are available, but no reviewed comparison panel is linked in this release.

All evaluations

1 evaluation · 1 result. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: BiRNA-BERTTask: extremely long RNA species classification
Dataset: extremely long-sequence species classification
0.804 F1
fraction · unknown

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

BiRNA-BERT: extremely long RNA species classification

adaptive tokenization on full-length long RNA sequences

Aggregation: Not reported

BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Table 2, BiRNA-BERT row, F1 Score column

Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.

Methods and evaluation design

Procedure, tasks and evaluated configurations

How it works

Evaluation methodology

Long noncoding RNA sequences from RNAcentral with species labels. The long-sequence species-classification Results paragraph specifies the dataset and F1 comparisons but does not give a train/validation/test assignment rule. F1 score; the reviewed task description does not establish the averaging convention. RNA-FM and RiNALMo are compared under their sequence-length constraints. The species-classification paragraph does not specify RNAcentral overlap exclusion between its benchmark sequences and model pretraining. Structure-task deduplication elsewhere is not this task.

SourcesBiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Results: BiRNA-BERT significantly outperforms in extremely long sequence task; Table 2; cached text lines 37–40

Recorded evaluations

Each evaluation records what was tested and under which conditions.

Run instructions

No runnable recipe has been reviewed for this task. Dataset access, model requirements, licences and compute requirements must be checked against its sources before execution.

A task describes a biological question. Choose a linked protocol to obtain concrete split and scoring instructions.

Strengths, limitations and unresolved questions

Strengths and limitations

Strengths and considerations

No source-reviewed explanatory claims are recorded here yet.

Limitations and conditions

Profile review details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Stable record: reported-task-c40dac20d9af66

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetsLong noncoding RNA sequences from RNAcentral with species labels.
SourcesBiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Results: BiRNA-BERT significantly outperforms in extremely long sequence task; Table 2; cached text lines 37–40
SplitsThe long-sequence species-classification Results paragraph specifies the dataset and F1 comparisons but does not give a train/validation/test assignment rule. · Not reported in inspected sources
SourcesBiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Results: BiRNA-BERT significantly outperforms in extremely long sequence task; Table 2; cached text lines 37–40
MetricsF1 score; the reviewed task description does not establish the averaging convention.
SourcesBiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Results: BiRNA-BERT significantly outperforms in extremely long sequence task; Table 2; cached text lines 37–40
BaselinesRNA-FM and RiNALMo are compared under their sequence-length constraints.
SourcesBiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Results: BiRNA-BERT significantly outperforms in extremely long sequence task; Table 2; cached text lines 37–40
Leakage controlsThe species-classification paragraph does not specify RNAcentral overlap exclusion between its benchmark sequences and model pretraining. Structure-task deduplication elsewhere is not this task. · Not reported in inspected sources
SourcesBiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Results: BiRNA-BERT significantly outperforms in extremely long sequence task; Table 2; cached text lines 37–40
UncertaintyTable 2 reports one F1 score per model for long-sequence species classification. The corresponding main-text section and Supplementary Information do not define repeated runs, confidence intervals or a statistical comparison for this particular task. · Not reported in inspected sources
Sources (2)BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization; birna journal Supplementary Information — pinned PDF · Results: extremely long sequence task and Table 2; Supplementary Information §§1–3
Entity typePaper-specific computational evaluation protocol.
SourcesBiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Results: BiRNA-BERT significantly outperforms in extremely long sequence task; Table 2; cached text lines 37–40
OrganismsBos taurus, Gallus gallus, Gorilla gorilla, Homo sapiens, Mus musculus, Pan troglodytes and Rattus norvegicus.
SourcesBiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Results: BiRNA-BERT significantly outperforms in extremely long sequence task; Table 2; cached text lines 37–40
AssaysRNAcentral sequence/species annotations.
SourcesBiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Results: BiRNA-BERT significantly outperforms in extremely long sequence task; Table 2; cached text lines 37–40
Allowed inputsLong RNA sequences.
SourcesBiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Results: BiRNA-BERT significantly outperforms in extremely long sequence task; Table 2; cached text lines 37–40
AdaptationThe long-sequence benchmark compares BiRNA-BERT, RiNALMo and RNA-FM, with comparator truncation stated. Its main-text description and Supplementary Information do not specify the classification head or whether each encoder is frozen or fine-tuned for this species task. · Not reported in inspected sources
Sources (2)BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization; birna journal Supplementary Information — pinned PDF · Results: extremely long sequence task, Table 2; Supplementary Information §§1–3

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Primary-paper discovery and source inspection. Source-checked results are not independently reproduced experiments.

Paper or primary resourceVersionReference
BiRNA-BERT allows efficient RNA language modeling with adaptive tokenizationjournal full text in PMCRead source
DOI: 10.1038/s42003-025-08982-0
Historical gaps recorded on 2026-09-17

The catalogue now holds 1 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.

  • No uncertainty in Table 2; exact long-sequence window/aggregation details remain tied to this paper's downstream classifiers.
  • Species classification scores do not measure RNA-structure accuracy.
Search and extraction details

primary comparison table screened

Searches

  • "PMC12635123"

Evidence locations

  • Table 2
  • Long sequence classification task
  • Supplement task description

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

20 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.
Individual claims
BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization

Original source ↗

Results: BiRNA-BERT significantly outperforms in extremely long sequence task; Table 2; cached text lines 37–40

Version: journal full text in PMC
Retrieved: 2026-09-16T10:33:38.292Z

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: bf7dbc52b6515301c77010c513f13e676c38395ddc82c20310171f518690c152

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps
  • Input: Long RNA sequences.
  • Task: Long-RNA species classification evaluates representation quality on an RNAcentral-derived seven-species dataset.
Individual claims
BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization

Original source ↗

Results: BiRNA-BERT significantly outperforms in extremely long sequence task; Table 2; cached text lines 37–40

Version: journal full text in PMC
Retrieved: 2026-09-16T10:33:38.292Z

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: bf7dbc52b6515301c77010c513f13e676c38395ddc82c20310171f518690c152

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title
Computational evaluation flow
Individual claims
BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization

Original source ↗

Results: BiRNA-BERT significantly outperforms in extremely long sequence task; Table 2; cached text lines 37–40

Version: journal full text in PMC
Retrieved: 2026-09-16T10:33:38.292Z

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.diagram.title

Source artifact SHA-256: bf7dbc52b6515301c77010c513f13e676c38395ddc82c20310171f518690c152

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Datasets
Long noncoding RNA sequences from RNAcentral with species labels.
Individual claims
BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization

Original source ↗

Results: BiRNA-BERT significantly outperforms in extremely long sequence task; Table 2; cached text lines 37–40

Version: journal full text in PMC
Retrieved: 2026-09-16T10:33:38.292Z

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: bf7dbc52b6515301c77010c513f13e676c38395ddc82c20310171f518690c152

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Splits
The long-sequence species-classification Results paragraph specifies the dataset and F1 comparisons but does not give a train/validation/test assignment rule.
Individual claims
BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization

Original source ↗

Results: BiRNA-BERT significantly outperforms in extremely long sequence task; Table 2; cached text lines 37–40

Version: journal full text in PMC
Retrieved: 2026-09-16T10:33:38.292Z

unreported

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: bf7dbc52b6515301c77010c513f13e676c38395ddc82c20310171f518690c152

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Adaptation
The long-sequence benchmark compares BiRNA-BERT, RiNALMo and RNA-FM, with comparator truncation stated. Its main-text description and Supplementary Information do not specify the classification head or whether each encoder is frozen or fine-tuned for this species task.
Individual claims
BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization

Original source ↗

Results: extremely long sequence task, Table 2; Supplementary Information §§1–3

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: journal full text in PMC
Retrieved: 2026-09-16T10:33:38.292Z

unreported

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: bf7dbc52b6515301c77010c513f13e676c38395ddc82c20310171f518690c152

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Adaptation
The long-sequence benchmark compares BiRNA-BERT, RiNALMo and RNA-FM, with comparator truncation stated. Its main-text description and Supplementary Information do not specify the classification head or whether each encoder is frozen or fine-tuned for this species task.
Individual claims
birna journal Supplementary Information — pinned PDF

Original source ↗

Results: extremely long sequence task, Table 2; Supplementary Information §§1–3

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: Published supplementary PDF 42003_2025_8982_MOESM2_ESM.pdf; sha256:7897d4dcf456d1f22b0631beabf7c5fd8678d0b8765cd40325195eb5addf4493
Retrieved: 2026-09-16T21:06:10.816352+00:00

unreported

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: 7897d4dcf456d1f22b0631beabf7c5fd8678d0b8765cd40325195eb5addf4493

Hash scope: Hash scope not separately documented; inspect source record

Archive member: 42003_2025_8982_MOESM2_ESM.pdf

Inspected artifact

Metrics
F1 score; the reviewed task description does not establish the averaging convention.
Individual claims
BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization

Original source ↗

Results: BiRNA-BERT significantly outperforms in extremely long sequence task; Table 2; cached text lines 37–40

Version: journal full text in PMC
Retrieved: 2026-09-16T10:33:38.292Z

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: bf7dbc52b6515301c77010c513f13e676c38395ddc82c20310171f518690c152

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Baselines
RNA-FM and RiNALMo are compared under their sequence-length constraints.
Individual claims
BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization

Original source ↗

Results: BiRNA-BERT significantly outperforms in extremely long sequence task; Table 2; cached text lines 37–40

Version: journal full text in PMC
Retrieved: 2026-09-16T10:33:38.292Z

source checked

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.facts.3.value

Source artifact SHA-256: bf7dbc52b6515301c77010c513f13e676c38395ddc82c20310171f518690c152

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Leakage controls
The species-classification paragraph does not specify RNAcentral overlap exclusion between its benchmark sequences and model pretraining. Structure-task deduplication elsewhere is not this task.
Individual claims
BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization

Original source ↗

Results: BiRNA-BERT significantly outperforms in extremely long sequence task; Table 2; cached text lines 37–40

Version: journal full text in PMC
Retrieved: 2026-09-16T10:33:38.292Z

unreported

automated source review · 2026-09-16

Audit details

Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: bf7dbc52b6515301c77010c513f13e676c38395ddc82c20310171f518690c152

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: needs review

3 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: reported-task-c40dac20d9af66

areas
rna-transcriptomes
tasks
extremely long RNA species classification
entity level
task
version
Not reported
task
extremely long RNA species classification
scope note
Paper-specific evaluation task; protocol completeness requires further extraction.
benchmark research
review date: 2026-09-17; status: primary_comparison_table_screened; primary sources: expansion-p3-birna-bert-2025; inspected locators: Table 2; Long sequence classification task; Supplement task description; searched queries: "PMC12635123"; gaps: No uncertainty in Table 2; exact long-sequence window/aggregation details remain tied to this paper's downstream classifiers.; Species classification scores do not measure RNA-structure accuracy.; claim scope: Primary-paper discovery and source inspection. Source-checked results are not independently reproduced experiments.
historical missing metadata
protocol version: not_reported_in_legacy_extract
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
legacy kinds
benchmark
entity classification
review date: 2026-09-17; rationale: This source-scoped record identifies the biological prediction task and holds its paper context. Preserve the existing task identity; exact split, model adaptation and scoring remain in linked evaluations or separate protocol records.; source ids: birna-bert-2025; source locator: Results: BiRNA-BERT significantly outperforms in extremely long sequence task; Table 2; cached text lines 37–40; ambiguities: A paper- or suite-specific task may constrain some inputs or metrics; that alone does not make it interchangeable with a complete versioned protocol. No protocol equivalence is inferred.; Some legacy profile Entity type facts use the generic phrase computational evaluation protocol. That boilerplate is not sufficient to establish a single fixed protocol identity or to merge this task with another protocol record.
Related records

Suggest a correction