Model type
Transformer representation pipeline; this record is the paper-specific evaluated configuration.
BiRNA-BERT is an RNA transformer that combines nucleotide and byte-pair tokenisation for structural and longer-sequence tasks.
Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings.
Transformer representation pipeline; this record is the paper-specific evaluated configuration.
RNA nucleotide sequences
RNA sequence or nucleotide representations for downstream prediction heads
Official study implementation and usage documentation: https://github.com/buetnlpbio/BiRNA-BERT/blob/14dc86b1b44c266f01025fc425103f2878646b39/README.md. This pinned documentation revision is not automatically the evaluated weight revision.
limited source coverage · Automated source review, 2026-09-16. All specifications and missing details
1 evaluation · 1 result. Different protocols are not a single leaderboard.
Applied filters: All linked evaluations
| Tested configuration | Protocol and dataset | Finding | Evidence and details |
|---|---|---|---|
| Configuration: BiRNA-BERT | Task: extremely long RNA species classification Dataset: extremely long-sequence species classification | 0.804 F1 fraction · unknown Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceBiRNA-BERT: extremely long RNA species classification adaptive tokenization on full-length long RNA sequences Aggregation: Not reported BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Table 2, BiRNA-BERT row, F1 Score column |
Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.
An ALiBi-equipped transformer encoder is pretrained with two tokenisation schemes. Nucleotide tokens retain position resolution; BPE compresses longer sequences. Tokenisation is selected according to input length.
The linked evaluation record identifies BiRNA-BERT: extremely long RNA species classification. Its dataset, split, adaptation and evidence origin remain attached to the reported results.
Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.
Stable record: reported-model-d3fd83835a2d44Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Model type | Transformer representation pipeline; this record is the paper-specific evaluated configuration.SourcesBiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Introduction (paragraph 4); Methods/Adaptive tokenization and dual pretraining (paragraph 1) |
| Architecture / procedure | An ALiBi-equipped transformer encoder is pretrained with two tokenisation schemes. Nucleotide tokens retain position resolution; BPE compresses longer sequences. Tokenisation is selected according to input length.SourcesBiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Introduction (paragraph 4); Methods/Adaptive tokenization and dual pretraining (paragraph 1) |
| Biological inputs | RNA nucleotide sequencesSourcesBiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Methods/Adaptive tokenization and dual pretraining (paragraph 3); Methods/Adaptive tokenization and dual pretraining (paragraph 5) |
| Outputs | RNA sequence or nucleotide representations for downstream prediction headsSourcesBiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Results (paragraph 4); Methods/Statistics and reproducibility (paragraph 1) |
| Parameters | The paper is internally inconsistent: model descriptions use 117M, while Results / Computational resources for pretraining / BiRNA-BERT calls it 116M. An exact reconciled count is unreported. · Not reported in inspected sourcesSourcesBiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Results (paragraph 2); Results/Comparison of computational efficiency/Computational resources for pretraining/BiRNA-BERT (paragraph 2) |
| Known versions / configuration | BiRNA-BERT is the comparison-table label; that label does not specify an immutable weight revision. · Not reported in inspected sourcesSourcesBiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Model identification in the comparison table and corresponding Methods; immutable checkpoint revision is not supplied by the table label. |
| Training data / fitting | Approximately 36 million RNAcentral ncRNA sequences; the Methods describe two epochs and 26.42 billion training tokens. Hardware and elapsed-time statements differ across sections and are not treated as a reproducible compute budget.SourcesBiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Methods/Pretraining dataset (paragraph 1); Methods/Pretraining configuration (paragraph 1) |
| Context limits | For the long-sequence comparison, nucleotide-token inputs are truncated to 1,022 tokens for comparability; BPE inputs are not truncated and reach 807 tokens. ALiBi support does not establish an unlimited practical context.SourcesBiRNA-BERT allows efficient RNA language modeling with adaptive tokenization · Results/Impact of dual tokenization is more significant in long sequences (paragraph 2); Results/Empirical perplexity analysis of different RNA language models (paragraph 6) |
| Access | Official study implementation and usage documentation: https://github.com/buetnlpbio/BiRNA-BERT/blob/14dc86b1b44c266f01025fc425103f2878646b39/README.md. This pinned documentation revision is not automatically the evaluated weight revision.Sourcesbuetnlpbio/BiRNA-BERT README.md · README.md; installation, model download and usage instructions |
| Code licence | No explicit code licence was established from the paper’s availability statement and inspected repository-root documentation. · Not reported in inspected sourcesSourcesbuetnlpbio/BiRNA-BERT README.md · README.md and repository-root licence-file search |
| Weights licence | The inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately. · Not reported in inspected sourcesSourcesbuetnlpbio/BiRNA-BERT README.md · README.md; checkpoint/access documentation and licence scope |
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
19 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings. Individual claims | BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization Introduction (paragraph 4); Methods/Adaptive tokenization and dual pretraining (paragraph 1) Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization Introduction (paragraph 4); Methods/Adaptive tokenization and dual pretraining (paragraph 1) Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Evaluated procedure (conceptual) Individual claims | BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization Introduction (paragraph 4); Methods/Adaptive tokenization and dual pretraining (paragraph 1) Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Model type Transformer representation pipeline; this record is the paper-specific evaluated configuration. Individual claims | BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization Introduction (paragraph 4); Methods/Adaptive tokenization and dual pretraining (paragraph 1) Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Architecture / procedure An ALiBi-equipped transformer encoder is pretrained with two tokenisation schemes. Nucleotide tokens retain position resolution; BPE compresses longer sequences. Tokenisation is selected according to input length. Individual claims | BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization Introduction (paragraph 4); Methods/Adaptive tokenization and dual pretraining (paragraph 1) Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Weights licence The inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately. Individual claims | buetnlpbio/BiRNA-BERT README.md README.md; checkpoint/access documentation and licence scope Version: 14dc86b1b44c266f01025fc425103f2878646b39 | unreported automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Biological inputs RNA nucleotide sequences Individual claims | BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization Methods/Adaptive tokenization and dual pretraining (paragraph 3); Methods/Adaptive tokenization and dual pretraining (paragraph 5) Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Outputs RNA sequence or nucleotide representations for downstream prediction heads Individual claims | BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization Results (paragraph 4); Methods/Statistics and reproducibility (paragraph 1) Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Parameters The paper is internally inconsistent: model descriptions use 117M, while Results / Computational resources for pretraining / BiRNA-BERT calls it 116M. An exact reconciled count is unreported. Individual claims | BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization Results (paragraph 2); Results/Comparison of computational efficiency/Computational resources for pretraining/BiRNA-BERT (paragraph 2) Version: journal full text in PMC | unreported automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Known versions / configuration BiRNA-BERT is the comparison-table label; that label does not specify an immutable weight revision. Individual claims | BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization Model identification in the comparison table and corresponding Methods; immutable checkpoint revision is not supplied by the table label. Version: journal full text in PMC | unreported automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
View linked audit checks and correction history
Release 2026-09-29-06401fd5b220 · Record review: needs review
Stable ID: reported-model-d3fd83835a2d44