rewire.itbenchmarks
Configuration

BiRNA-BERT

BiRNA-BERT is an RNA transformer that combines nucleotide and byte-pair tokenisation for structural and longer-sequence tasks.

SourcesBiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Abstract (paragraph 1); Discussion (paragraph 2)

1 evaluation · 1 result

How it worksEvaluated procedure (conceptual)
Evaluated procedure (conceptual)1. RNA nucleotide sequences. Then: 2. BiRNA-BERT. Then: 3. RNA sequence or nucleotide representations for downstream prediction headsEvaluated procedure (conceptual)1. RNA nucleotide sequences. Then: 2. BiRNA-BERT. Then: 3. RNA sequence or nucleotide representations for downstream prediction headsEvaluated procedure (conceptual)1. RNA nucleotide sequences. Then: 2. BiRNA-BERT. Then: 3. RNA sequence or nucleotide representations for downstream prediction heads

Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings.

SourcesBiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Introduction (paragraph 4); Methods/Adaptive tokenization and dual pretraining (paragraph 1)

Overview

limited source coverage · Automated source review, 2026-09-16. All specifications and missing details

Evaluations and results

1 evaluation · 1 result. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: BiRNA-BERTTask: extremely long RNA species classification
Dataset: extremely long-sequence species classification
0.804 F1
fraction · unknown

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

BiRNA-BERT: extremely long RNA species classification

adaptive tokenization on full-length long RNA sequences

Aggregation: Not reported

BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Table 2, BiRNA-BERT row, F1 Score column

Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.

Use this model

How it works, versions and access

How it works

How the evaluated method works

An ALiBi-equipped transformer encoder is pretrained with two tokenisation schemes. Nucleotide tokens retain position resolution; BPE compresses longer sequences. Tokenisation is selected according to input length.

SourcesBiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Introduction (paragraph 4); Methods/Adaptive tokenization and dual pretraining (paragraph 1)
What was evaluated

The linked evaluation record identifies BiRNA-BERT: extremely long RNA species classification. Its dataset, split, adaptation and evidence origin remain attached to the reported results.

SourcesBiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · The named evaluation’s methods and comparison table; exact preserved evaluation IDs: evaluation-b2-birna-bert-2025
Strengths, limitations and unresolved questions

Strengths and limitations

Limitations and conditions

  • Relative positional encoding and compressed tokens do not prove unlimited useful biological context; practical memory and task-specific validation still constrain use.
    SourcesBiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Methods/Tokenization strategies for biological foundational models (paragraph 1); Methods/Positional encoding in the transformer architecture (paragraph 1)
Profile review details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Stable record: reported-model-d3fd83835a2d44

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Inputs, outputs and configuration
PropertyDescription and evidence
Model typeTransformer representation pipeline; this record is the paper-specific evaluated configuration.
SourcesBiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Introduction (paragraph 4); Methods/Adaptive tokenization and dual pretraining (paragraph 1)
Architecture / procedureAn ALiBi-equipped transformer encoder is pretrained with two tokenisation schemes. Nucleotide tokens retain position resolution; BPE compresses longer sequences. Tokenisation is selected according to input length.
SourcesBiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Introduction (paragraph 4); Methods/Adaptive tokenization and dual pretraining (paragraph 1)
Biological inputsRNA nucleotide sequences
SourcesBiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Methods/Adaptive tokenization and dual pretraining (paragraph 3); Methods/Adaptive tokenization and dual pretraining (paragraph 5)
OutputsRNA sequence or nucleotide representations for downstream prediction heads
SourcesBiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Results (paragraph 4); Methods/Statistics and reproducibility (paragraph 1)
ParametersThe paper is internally inconsistent: model descriptions use 117M, while Results / Computational resources for pretraining / BiRNA-BERT calls it 116M. An exact reconciled count is unreported. · Not reported in inspected sources
SourcesBiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Results (paragraph 2); Results/Comparison of computational efficiency/Computational resources for pretraining/BiRNA-BERT (paragraph 2)
Known versions / configurationBiRNA-BERT is the comparison-table label; that label does not specify an immutable weight revision. · Not reported in inspected sources
SourcesBiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Model identification in the comparison table and corresponding Methods; immutable checkpoint revision is not supplied by the table label.
Training data / fittingApproximately 36 million RNAcentral ncRNA sequences; the Methods describe two epochs and 26.42 billion training tokens. Hardware and elapsed-time statements differ across sections and are not treated as a reproducible compute budget.
SourcesBiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Methods/Pretraining dataset (paragraph 1); Methods/Pretraining configuration (paragraph 1)
Context limitsFor the long-sequence comparison, nucleotide-token inputs are truncated to 1,022 tokens for comparability; BPE inputs are not truncated and reach 807 tokens. ALiBi support does not establish an unlimited practical context.
SourcesBiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Results/Impact of dual tokenization is more significant in long sequences (paragraph 2); Results/Empirical perplexity analysis of different RNA language models (paragraph 6)
AccessOfficial study implementation and usage documentation: https://github.com/buetnlpbio/BiRNA-BERT/blob/14dc86b1b44c266f01025fc425103f2878646b39/README.md. This pinned documentation revision is not automatically the evaluated weight revision.
Sourcesbuetnlpbio/BiRNA-BERT README.md · README.md; installation, model download and usage instructions
Code licenceNo explicit code licence was established from the paper’s availability statement and inspected repository-root documentation. · Not reported in inspected sources
Sourcesbuetnlpbio/BiRNA-BERT README.md · README.md and repository-root licence-file search
Weights licenceThe inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately. · Not reported in inspected sources
Sourcesbuetnlpbio/BiRNA-BERT README.md · README.md; checkpoint/access documentation and licence scope

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

19 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings.
Individual claims
BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization

Original source ↗

Introduction (paragraph 4); Methods/Adaptive tokenization and dual pretraining (paragraph 1)

Version: journal full text in PMC
Retrieved: 2026-09-16T10:33:38.292Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: bf7dbc52b6515301c77010c513f13e676c38395ddc82c20310171f518690c152

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps
  • RNA nucleotide sequences
  • BiRNA-BERT
  • RNA sequence or nucleotide representations for downstream prediction heads
Individual claims
BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization

Original source ↗

Introduction (paragraph 4); Methods/Adaptive tokenization and dual pretraining (paragraph 1)

Version: journal full text in PMC
Retrieved: 2026-09-16T10:33:38.292Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: bf7dbc52b6515301c77010c513f13e676c38395ddc82c20310171f518690c152

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title
Evaluated procedure (conceptual)
Individual claims
BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization

Original source ↗

Introduction (paragraph 4); Methods/Adaptive tokenization and dual pretraining (paragraph 1)

Version: journal full text in PMC
Retrieved: 2026-09-16T10:33:38.292Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: bf7dbc52b6515301c77010c513f13e676c38395ddc82c20310171f518690c152

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Model type
Transformer representation pipeline; this record is the paper-specific evaluated configuration.
Individual claims
BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization

Original source ↗

Introduction (paragraph 4); Methods/Adaptive tokenization and dual pretraining (paragraph 1)

Version: journal full text in PMC
Retrieved: 2026-09-16T10:33:38.292Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: bf7dbc52b6515301c77010c513f13e676c38395ddc82c20310171f518690c152

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Architecture / procedure
An ALiBi-equipped transformer encoder is pretrained with two tokenisation schemes. Nucleotide tokens retain position resolution; BPE compresses longer sequences. Tokenisation is selected according to input length.
Individual claims
BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization

Original source ↗

Introduction (paragraph 4); Methods/Adaptive tokenization and dual pretraining (paragraph 1)

Version: journal full text in PMC
Retrieved: 2026-09-16T10:33:38.292Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: bf7dbc52b6515301c77010c513f13e676c38395ddc82c20310171f518690c152

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Weights licence
The inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately.
Individual claims
buetnlpbio/BiRNA-BERT README.md

Original source ↗

README.md; checkpoint/access documentation and licence scope

Version: 14dc86b1b44c266f01025fc425103f2878646b39
Retrieved: 2026-09-16T19:54:12.014900+00:00

unreported

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: 00f3835cae38a2307d8095fc5240496e01a9a981c8450230f9ae5722153ac2cc

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Biological inputs
RNA nucleotide sequences
Individual claims
BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization

Original source ↗

Methods/Adaptive tokenization and dual pretraining (paragraph 3); Methods/Adaptive tokenization and dual pretraining (paragraph 5)

Version: journal full text in PMC
Retrieved: 2026-09-16T10:33:38.292Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: bf7dbc52b6515301c77010c513f13e676c38395ddc82c20310171f518690c152

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Outputs
RNA sequence or nucleotide representations for downstream prediction heads
Individual claims
BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization

Original source ↗

Results (paragraph 4); Methods/Statistics and reproducibility (paragraph 1)

Version: journal full text in PMC
Retrieved: 2026-09-16T10:33:38.292Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.3.value

Source artifact SHA-256: bf7dbc52b6515301c77010c513f13e676c38395ddc82c20310171f518690c152

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Parameters
The paper is internally inconsistent: model descriptions use 117M, while Results / Computational resources for pretraining / BiRNA-BERT calls it 116M. An exact reconciled count is unreported.
Individual claims
BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization

Original source ↗

Results (paragraph 2); Results/Comparison of computational efficiency/Computational resources for pretraining/BiRNA-BERT (paragraph 2)

Version: journal full text in PMC
Retrieved: 2026-09-16T10:33:38.292Z

unreported

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: bf7dbc52b6515301c77010c513f13e676c38395ddc82c20310171f518690c152

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Known versions / configuration
BiRNA-BERT is the comparison-table label; that label does not specify an immutable weight revision.
Individual claims
BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization

Original source ↗

Model identification in the comparison table and corresponding Methods; immutable checkpoint revision is not supplied by the table label.

Version: journal full text in PMC
Retrieved: 2026-09-16T10:33:38.292Z

unreported

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.5.value

Source artifact SHA-256: bf7dbc52b6515301c77010c513f13e676c38395ddc82c20310171f518690c152

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: needs review

2 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: reported-model-d3fd83835a2d44

areas
rna-transcriptomes
entity level
method
version
not stated in table
reported name
BiRNA-BERT
historical missing metadata
checkpoint revision: not_reported_in_legacy_extract; training data: not_reported_in_legacy_extract; licence: not_reported_in_legacy_extract
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
legacy kinds
model
entity classification
review date: 2026-09-17; rationale: This source-scoped entry preserves the method/configuration actually named in an evaluation. It is neither a global family identity nor proof of an immutable checkpoint; the linked evaluation retains adaptation, fitting and scoring details.; source ids: birna-bert-2025; source locator: Introduction (paragraph 4); Methods/Adaptive tokenization and dual pretraining (paragraph 1) | Abstract (paragraph 1); Discussion (paragraph 2); ambiguities: Configuration means the source-labelled evaluated identity. It does not establish missing checkpoint hashes, default settings or equivalence to same-named records in other papers.
Related records

Suggest a correction