rewire.itbenchmarks
Configuration

ProkBERT-mini

ProkBERT-mini learns microbial DNA representations with local-context-aware tokenisation.

SourcesProkBERT family: genomic language models for microbiome applications · 2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); 4 Conclusion (paragraph 5)

1 evaluation · 4 results

How it worksEvaluated procedure (conceptual)
Evaluated procedure (conceptual)1. Microbial DNA sequences. Then: 2. ProkBERT-mini. Then: 3. DNA embeddings and task-specific genomic classificationsEvaluated procedure (conceptual)1. Microbial DNA sequences. Then: 2. ProkBERT-mini. Then: 3. DNA embeddings and task-specific genomic classificationsEvaluated procedure (conceptual)1. Microbial DNA sequences. Then: 2. ProkBERT-mini. Then: 3. DNA embeddings and task-specific genomic classifications

Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings.

SourcesProkBERT family: genomic language models for microbiome applications · 2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); 2 Materials and methods (paragraph 1)

Overview

limited source coverage · Automated source review, 2026-09-16. All specifications and missing details

Evaluations and results

1 evaluation · 4 results. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: ProkBERT-miniProtocol: E. coli sigma70 independent promoter test (E. coli sigma70 promoter prediction)
Dataset: E. coli sigma70 promoter dataset
0.87 Accuracy
unitless · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

ProkBERT-mini: E. coli sigma70 promoter prediction

Test-only evaluation on Cassiano and Silva-Rocha 2020 data; methods have different training histories. 865 high-evidence RegulonDB 10.5 promoters and 1,000 nucleotide-distribution-matched negative sequences. Promoter exact matches removed from model training.

Aggregation: Not reported

ProkBERT family: genomic language models for microbiome applications; ProkBERT family: genomic language models for microbiome applications · Table 3, ProkBERT-mini row, Accuracy column
Configuration: ProkBERT-miniProtocol: E. coli sigma70 independent promoter test (E. coli sigma70 promoter prediction)
Dataset: E. coli sigma70 promoter dataset
0.9 Sensitivity
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

ProkBERT-mini: E. coli sigma70 promoter prediction

Test-only evaluation on Cassiano and Silva-Rocha 2020 data; methods have different training histories. 865 high-evidence RegulonDB 10.5 promoters and 1,000 nucleotide-distribution-matched negative sequences. Promoter exact matches removed from model training.

Aggregation: Not reported

ProkBERT family: genomic language models for microbiome applications · Table 3, row ProkBERT-mini, column Sensitivity; XML row2 column4
Configuration: ProkBERT-miniProtocol: E. coli sigma70 independent promoter test (E. coli sigma70 promoter prediction)
Dataset: E. coli sigma70 promoter dataset
0.85 Specificity
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

ProkBERT-mini: E. coli sigma70 promoter prediction

Test-only evaluation on Cassiano and Silva-Rocha 2020 data; methods have different training histories. 865 high-evidence RegulonDB 10.5 promoters and 1,000 nucleotide-distribution-matched negative sequences. Promoter exact matches removed from model training.

Aggregation: Not reported

ProkBERT family: genomic language models for microbiome applications · Table 3, row ProkBERT-mini, column Specificity; XML row2 column5
Configuration: ProkBERT-miniProtocol: E. coli sigma70 independent promoter test (E. coli sigma70 promoter prediction)
Dataset: E. coli sigma70 promoter dataset
0.74 MCC
dimensionless · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

ProkBERT-mini: E. coli sigma70 promoter prediction

Test-only evaluation on Cassiano and Silva-Rocha 2020 data; methods have different training histories. 865 high-evidence RegulonDB 10.5 promoters and 1,000 nucleotide-distribution-matched negative sequences. Promoter exact matches removed from model training.

Aggregation: Not reported

ProkBERT family: genomic language models for microbiome applications · Table 3, row ProkBERT-mini, column MCC; XML row2 column3

Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.

Use this model

How it works, versions and access

Related profile: ProkBERT. This page retains the exact record and its evaluation context.

How it works

How the evaluated method works

A transformer encoder uses local-context-aware 6-mer tokenisation for the ProkBERT-mini variant, followed by self-supervised pretraining and task-specific fine-tuning.

SourcesProkBERT family: genomic language models for microbiome applications · 2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); 2 Materials and methods (paragraph 1)
What was evaluated

The linked evaluation record identifies ProkBERT-mini: E. coli sigma70 promoter prediction. Its dataset, split, adaptation and evidence origin remain attached to the reported results.

SourcesProkBERT family: genomic language models for microbiome applications · The named evaluation’s methods and comparison table; exact preserved evaluation IDs: evaluation-lit-033
Strengths, limitations and unresolved questions

Strengths and limitations

Strengths and considerations

  • Targets microbial sequence distributions and evaluates promoter-related downstream tasks.
    SourcesProkBERT family: genomic language models for microbiome applications · 2 Materials and methods/2.3 Application I: bacterial promoter prediction/2.3.1 Dataset overview/2.3.1.2 Dataset construction for multispecies train, test and validation sets (paragraph 6); 3 Results and discussion/3.4 ProkBERT swiftly and accurately identifies phage sequences, even in challenging settings (paragraph 2)

Limitations and conditions

  • The mini checkpoint and tokenizer settings define the evaluated system; results cannot be transferred to every ProkBERT variant.
    SourcesProkBERT family: genomic language models for microbiome applications · 2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); 3 Results and discussion/3.3 ProkBERT performs accurately and robustly in promoter sequence recognition (paragraph 8)
Profile review details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Stable record: reported-model-4438513d9cd42c

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Inputs, outputs and configuration
PropertyDescription and evidence
Model typeTransformer representation pipeline; this record is the paper-specific evaluated configuration.
SourcesProkBERT family: genomic language models for microbiome applications · 2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); 2 Materials and methods (paragraph 1)
Architecture / procedureA transformer encoder uses local-context-aware 6-mer tokenisation for the ProkBERT-mini variant, followed by self-supervised pretraining and task-specific fine-tuning.
SourcesProkBERT family: genomic language models for microbiome applications · 2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); 2 Materials and methods (paragraph 1)
Biological inputsMicrobial DNA sequences
SourcesProkBERT family: genomic language models for microbiome applications · 1 Introduction (paragraph 7); 4 Conclusion (paragraph 10)
OutputsDNA embeddings and task-specific genomic classifications
SourcesProkBERT family: genomic language models for microbiome applications · 2 Materials and methods/2.5 Applied metrics (paragraph 1); 3 Results and discussion/3.1 ProkBERT's learned representations capture genomic structure and phylogeny (paragraph 4)
ParametersApproximately 20 million parameters.
SourcesProkBERT family: genomic language models for microbiome applications · 4 Conclusion (paragraph 8); 2 Materials and methods/2.2 Pretraining and learning sequence representations/2.2.2 Training process/2.2.2.2 Training phases and configuration (paragraph 1)
Known versions / configurationProkBERT-mini is the comparison-table label; that label does not specify an immutable weight revision. · Not reported in inspected sources
SourcesProkBERT family: genomic language models for microbiome applications · Model identification in the comparison table and corresponding Methods; immutable checkpoint revision is not supplied by the table label.
Training data / fittingUnlabelled microbial genome sequences followed by supervised task data, as specified in the paper.
SourcesProkBERT family: genomic language models for microbiome applications · 2 Materials and methods/2.4 Application II: phage sequence analysis/2.4.2 Model training for phage sequence analysis (paragraph 1); 2 Materials and methods/2.2 Pretraining and learning sequence representations/2.2.4 Analysis of encoder outputs (paragraph 6)
Context limitsApproximately 1 kb for ProkBERT-mini; the distinct mini-long variant supports approximately 2 kb.
SourcesProkBERT family: genomic language models for microbiome applications · 2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); Table T3 (paragraph 1)
AccessOfficial study implementation and usage documentation: https://github.com/nbrg-ppcu/prokbert/blob/8670ae92b816cff158a0b85647a8dea122e251eb/README.md. This pinned documentation revision is not automatically the evaluated weight revision.
Sourcesnbrg-ppcu/prokbert README.md · README.md; installation, model download and usage instructions
Code licenceMIT (study repository code at the cited revision; this does not establish every dependency or historical checkpoint licence).
Sourcesnbrg-ppcu/prokbert LICENSE · LICENSE; complete licence text
Weights licenceThe inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately. · Not reported in inspected sources
Sourcesnbrg-ppcu/prokbert README.md · README.md; checkpoint/access documentation and licence scope

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

21 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings.
Individual claims
ProkBERT family: genomic language models for microbiome applications

Original source ↗

2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); 2 Materials and methods (paragraph 1)

Version: PMC10810988.1
Retrieved: 2026-09-16T10:33:36.197Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 8610e2a54aa877c8dc565a9cdb6e82099f284c5e0907a52cab18d994ea732436

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps
  • Microbial DNA sequences
  • ProkBERT-mini
  • DNA embeddings and task-specific genomic classifications
Individual claims
ProkBERT family: genomic language models for microbiome applications

Original source ↗

2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); 2 Materials and methods (paragraph 1)

Version: PMC10810988.1
Retrieved: 2026-09-16T10:33:36.197Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 8610e2a54aa877c8dc565a9cdb6e82099f284c5e0907a52cab18d994ea732436

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title
Evaluated procedure (conceptual)
Individual claims
ProkBERT family: genomic language models for microbiome applications

Original source ↗

2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); 2 Materials and methods (paragraph 1)

Version: PMC10810988.1
Retrieved: 2026-09-16T10:33:36.197Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: 8610e2a54aa877c8dc565a9cdb6e82099f284c5e0907a52cab18d994ea732436

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Model type
Transformer representation pipeline; this record is the paper-specific evaluated configuration.
Individual claims
ProkBERT family: genomic language models for microbiome applications

Original source ↗

2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); 2 Materials and methods (paragraph 1)

Version: PMC10810988.1
Retrieved: 2026-09-16T10:33:36.197Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: 8610e2a54aa877c8dc565a9cdb6e82099f284c5e0907a52cab18d994ea732436

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Architecture / procedure
A transformer encoder uses local-context-aware 6-mer tokenisation for the ProkBERT-mini variant, followed by self-supervised pretraining and task-specific fine-tuning.
Individual claims
ProkBERT family: genomic language models for microbiome applications

Original source ↗

2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); 2 Materials and methods (paragraph 1)

Version: PMC10810988.1
Retrieved: 2026-09-16T10:33:36.197Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: 8610e2a54aa877c8dc565a9cdb6e82099f284c5e0907a52cab18d994ea732436

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Weights licence
The inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately.
Individual claims
nbrg-ppcu/prokbert README.md

Original source ↗

README.md; checkpoint/access documentation and licence scope

Version: 8670ae92b816cff158a0b85647a8dea122e251eb
Retrieved: 2026-09-16T19:54:21.078483+00:00

unreported

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: 29de39c6ad006ce704ab14240cfd97af93da411fb63ec89ebe499c2646928cfc

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Biological inputs
Microbial DNA sequences
Individual claims
ProkBERT family: genomic language models for microbiome applications

Original source ↗

1 Introduction (paragraph 7); 4 Conclusion (paragraph 10)

Version: PMC10810988.1
Retrieved: 2026-09-16T10:33:36.197Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: 8610e2a54aa877c8dc565a9cdb6e82099f284c5e0907a52cab18d994ea732436

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Outputs
DNA embeddings and task-specific genomic classifications
Individual claims
ProkBERT family: genomic language models for microbiome applications

Original source ↗

2 Materials and methods/2.5 Applied metrics (paragraph 1); 3 Results and discussion/3.1 ProkBERT's learned representations capture genomic structure and phylogeny (paragraph 4)

Version: PMC10810988.1
Retrieved: 2026-09-16T10:33:36.197Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.3.value

Source artifact SHA-256: 8610e2a54aa877c8dc565a9cdb6e82099f284c5e0907a52cab18d994ea732436

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Parameters
Approximately 20 million parameters.
Individual claims
ProkBERT family: genomic language models for microbiome applications

Original source ↗

4 Conclusion (paragraph 8); 2 Materials and methods/2.2 Pretraining and learning sequence representations/2.2.2 Training process/2.2.2.2 Training phases and configuration (paragraph 1)

Version: PMC10810988.1
Retrieved: 2026-09-16T10:33:36.197Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: 8610e2a54aa877c8dc565a9cdb6e82099f284c5e0907a52cab18d994ea732436

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Known versions / configuration
ProkBERT-mini is the comparison-table label; that label does not specify an immutable weight revision.
Individual claims
ProkBERT family: genomic language models for microbiome applications

Original source ↗

Model identification in the comparison table and corresponding Methods; immutable checkpoint revision is not supplied by the table label.

Version: PMC10810988.1
Retrieved: 2026-09-16T10:33:36.197Z

unreported

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.5.value

Source artifact SHA-256: 8610e2a54aa877c8dc565a9cdb6e82099f284c5e0907a52cab18d994ea732436

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: needs review

3 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: reported-model-4438513d9cd42c

areas
microbes-communities
entity level
method
version
Not reported
reported name
ProkBERT-mini
historical missing metadata
version: not_reported_in_legacy_extract; checkpoint revision: not_reported_in_legacy_extract; training data: not_reported_in_legacy_extract; licence: not_reported_in_legacy_extract
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
legacy kinds
model
entity classification
review date: 2026-09-17; rationale: This source-scoped entry preserves the method/configuration actually named in an evaluation. It is neither a global family identity nor proof of an immutable checkpoint; the linked evaluation retains adaptation, fitting and scoring details.; source ids: prokbert-2024; source locator: 2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); 2 Materials and methods (paragraph 1) | 2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); 4 Conclusion (paragraph 5); ambiguities: Configuration means the source-labelled evaluated identity. It does not establish missing checkpoint hashes, default settings or equivalence to same-named records in other papers.
Related records

Suggest a correction