Model type
DNA transformer encoder family
Nucleotide Transformer is a family of DNA encoders pretrained on human or multispecies sequence corpora.
Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.
DNA transformer encoder family
DNA sequences tokenized into 6-mers, with single-base handling of N and remainder bases.
Contextual representations used in specified downstream prediction workflows.
Official project documentation and implementation: https://github.com/instadeepai/nucleotide-transformer
limited source coverage · Automated source review, 2026-09-16. All specifications and missing details
28 evaluations · 28 results. Different protocols are not a single leaderboard.
Applied filters: All linked evaluations
| Tested configuration | Protocol and dataset | Finding | Evidence and details |
|---|---|---|---|
| Configuration: NT-2500M-1000g | Task: GUE CORE-PROMOTER-DETECTION-ALL: Core promoter detection, dataset all Dataset subset: GUE Core promoter detection, all (GUE split) | 67.4% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNT-2500M-1000g on GUE CORE-PROMOTER-DETECTION-ALL: Core promoter detection, dataset all Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-2500M-1000g), column(Core promoter detection all) |
| Configuration: NT-2500M-1000g | Task: GUE CORE-PROMOTER-DETECTION-NOTATA: Core promoter detection, dataset notata Dataset subset: GUE Core promoter detection, notata (GUE split) | 67.5% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNT-2500M-1000g on GUE CORE-PROMOTER-DETECTION-NOTATA: Core promoter detection, dataset notata Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-2500M-1000g), column(Core promoter detection notata) |
| Configuration: NT-2500M-1000g | Task: GUE CORE-PROMOTER-DETECTION-TATA: Core promoter detection, dataset tata Dataset subset: GUE Core promoter detection, tata (GUE split) | 69.7% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNT-2500M-1000g on GUE CORE-PROMOTER-DETECTION-TATA: Core promoter detection, dataset tata Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-2500M-1000g), column(Core promoter detection tata) |
| Configuration: NT-2500M-1000g | Task: GUE COVID-VARIANT-CLASSIFICATION-COVID: Covid variant classification, dataset Covid Dataset subset: GUE Covid variant classification, Covid (GUE split) | 66.7% f1 percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-2500M-1000g), column(Covid variant classification Covid) |
| Configuration: NT-2500M-1000g | Task: GUE EPIGENETIC-MARKS-PREDICTION-H3: Epigenetic marks prediction, dataset H3 Dataset subset: GUE Epigenetic marks prediction, H3 (GUE split) | 74.6% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNT-2500M-1000g on GUE EPIGENETIC-MARKS-PREDICTION-H3: Epigenetic marks prediction, dataset H3 Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-2500M-1000g), column(Epigenetic marks prediction H3) |
| Configuration: NT-2500M-1000g | Task: GUE EPIGENETIC-MARKS-PREDICTION-H3K14AC: Epigenetic marks prediction, dataset H3K14ac Dataset subset: GUE Epigenetic marks prediction, H3K14ac (GUE split) | 44.1% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-2500M-1000g), column(Epigenetic marks prediction H3K14ac) |
| Configuration: NT-2500M-1000g | Task: GUE EPIGENETIC-MARKS-PREDICTION-H3K36ME3: Epigenetic marks prediction, dataset H3K36me3 Dataset subset: GUE Epigenetic marks prediction, H3K36me3 (GUE split) | 50.9% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-2500M-1000g), column(Epigenetic marks prediction H3K36me3) |
| Configuration: NT-2500M-1000g | Task: GUE EPIGENETIC-MARKS-PREDICTION-H3K4ME1: Epigenetic marks prediction, dataset H3K4me1 Dataset subset: GUE Epigenetic marks prediction, H3K4me1 (GUE split) | 43.1% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-2500M-1000g), column(Epigenetic marks prediction H3K4me1) |
| Configuration: NT-2500M-1000g | Task: GUE EPIGENETIC-MARKS-PREDICTION-H3K4ME2: Epigenetic marks prediction, dataset H3K4me2 Dataset subset: GUE Epigenetic marks prediction, H3K4me2 (GUE split) | 30.3% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-2500M-1000g), column(Epigenetic marks prediction H3K4me2) |
| Configuration: NT-2500M-1000g | Task: GUE EPIGENETIC-MARKS-PREDICTION-H3K4ME3: Epigenetic marks prediction, dataset H3K4me3 Dataset subset: GUE Epigenetic marks prediction, H3K4me3 (GUE split) | 30.9% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-2500M-1000g), column(Epigenetic marks prediction H3K4me3) |
| Configuration: NT-2500M-1000g | Task: GUE EPIGENETIC-MARKS-PREDICTION-H3K79ME3: Epigenetic marks prediction, dataset H3K79me3 Dataset subset: GUE Epigenetic marks prediction, H3K79me3 (GUE split) | 61.2% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-2500M-1000g), column(Epigenetic marks prediction H3K79me3) |
| Configuration: NT-2500M-1000g | Task: GUE EPIGENETIC-MARKS-PREDICTION-H3K9AC: Epigenetic marks prediction, dataset H3K9ac Dataset subset: GUE Epigenetic marks prediction, H3K9ac (GUE split) | 52.4% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-2500M-1000g), column(Epigenetic marks prediction H3K9ac) |
| Configuration: NT-2500M-1000g | Task: GUE EPIGENETIC-MARKS-PREDICTION-H4: Epigenetic marks prediction, dataset H4 Dataset subset: GUE Epigenetic marks prediction, H4 (GUE split) | 79.8% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNT-2500M-1000g on GUE EPIGENETIC-MARKS-PREDICTION-H4: Epigenetic marks prediction, dataset H4 Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-2500M-1000g), column(Epigenetic marks prediction H4) |
| Configuration: NT-2500M-1000g | Task: GUE EPIGENETIC-MARKS-PREDICTION-H4AC: Epigenetic marks prediction, dataset H4ac Dataset subset: GUE Epigenetic marks prediction, H4ac (GUE split) | 41.5% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNT-2500M-1000g on GUE EPIGENETIC-MARKS-PREDICTION-H4AC: Epigenetic marks prediction, dataset H4ac Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-2500M-1000g), column(Epigenetic marks prediction H4ac) |
| Configuration: NT-2500M-1000g | Task: GUE PROMOTER-DETECTION-ALL: Promoter detection, dataset all Dataset subset: GUE Promoter detection, all (GUE split) | 91% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNT-2500M-1000g on GUE PROMOTER-DETECTION-ALL: Promoter detection, dataset all Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-2500M-1000g), column(Promoter detection all) |
| Configuration: NT-2500M-1000g | Task: GUE PROMOTER-DETECTION-NOTATA: Promoter detection, dataset notata Dataset subset: GUE Promoter detection, notata (GUE split) | 93.1% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNT-2500M-1000g on GUE PROMOTER-DETECTION-NOTATA: Promoter detection, dataset notata Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-2500M-1000g), column(Promoter detection notata) |
| Configuration: NT-2500M-1000g | Task: GUE PROMOTER-DETECTION-TATA: Promoter detection, dataset tata Dataset subset: GUE Promoter detection, tata (GUE split) | 75.8% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNT-2500M-1000g on GUE PROMOTER-DETECTION-TATA: Promoter detection, dataset tata Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-2500M-1000g), column(Promoter detection tata) |
| Configuration: NT-2500M-1000g | Task: GUE SPLICE-SITE-PREDICTION-RECONSTRUCT: Splice site prediction, dataset Reconstruct Dataset subset: GUE Splice site prediction, Reconstruct (GUE split) | 85.8% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-2500M-1000g), column(Splice site prediction Reconstruct) |
| Configuration: NT-2500M-1000g | Task: GUE TRANSCRIPTION-FACTOR-PREDICTION-HUMAN-0: Transcription factor prediction (human), dataset 0 Dataset subset: GUE Transcription factor prediction (human), 0 (GUE split) | 66.3% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-2500M-1000g), column(Transcription factor prediction (human) 0) |
| Configuration: NT-2500M-1000g | Task: GUE TRANSCRIPTION-FACTOR-PREDICTION-HUMAN-1: Transcription factor prediction (human), dataset 1 Dataset subset: GUE Transcription factor prediction (human), 1 (GUE split) | 68.3% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-2500M-1000g), column(Transcription factor prediction (human) 1) |
| Configuration: NT-2500M-1000g | Task: GUE TRANSCRIPTION-FACTOR-PREDICTION-HUMAN-2: Transcription factor prediction (human), dataset 2 Dataset subset: GUE Transcription factor prediction (human), 2 (GUE split) | 58.7% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-2500M-1000g), column(Transcription factor prediction (human) 2) |
| Configuration: NT-2500M-1000g | Task: GUE TRANSCRIPTION-FACTOR-PREDICTION-HUMAN-3: Transcription factor prediction (human), dataset 3 Dataset subset: GUE Transcription factor prediction (human), 3 (GUE split) | 49.1% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-2500M-1000g), column(Transcription factor prediction (human) 3) |
| Configuration: NT-2500M-1000g | Task: GUE TRANSCRIPTION-FACTOR-PREDICTION-HUMAN-4: Transcription factor prediction (human), dataset 4 Dataset subset: GUE Transcription factor prediction (human), 4 (GUE split) | 67.6% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-2500M-1000g), column(Transcription factor prediction (human) 4) |
| Configuration: NT-2500M-1000g | Task: GUE TRANSCRIPTION-FACTOR-PREDICTION-MOUSE-0: Transcription factor prediction (mouse), dataset 0 Dataset subset: GUE Transcription factor prediction (mouse), 0 (GUE split) | 48.3% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-2500M-1000g), column(Transcription factor prediction (mouse) 0) |
| Configuration: NT-2500M-1000g | Task: GUE TRANSCRIPTION-FACTOR-PREDICTION-MOUSE-1: Transcription factor prediction (mouse), dataset 1 Dataset subset: GUE Transcription factor prediction (mouse), 1 (GUE split) | 80% mcc percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceFine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12. Aggregation: Not reported DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-2500M-1000g), column(Transcription factor prediction (mouse) 1) |
Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.
Related profile: Nucleotide Transformer. This page retains the exact record and its evaluation context.
Genome language model fine-tuned on each GUE dataset by the DNABERT-2 authors.
Nucleotide Transformer is a family of DNA encoders pretrained on human or multispecies sequence corpora. Encoder-only transformers with 6-mer tokens; v1 uses learned positional encodings and v2 uses rotary positions and SwiGLU. The documented inputs are DNA sequences tokenized into 6-mers, with single-base handling of N and remainder bases. The output consists of contextual representations used in specified downstream prediction workflows.
NT-v1 human-reference/1000G/multispecies variants and NT-v2 50M/100M/250M/500M. NT-v3 is a separate architecture described elsewhere in the repository. v1: approximately 6kb; v2: 2,048 tokens, approximately 12kb. Exact base count depends on special and ambiguous tokens.
Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.
Stable record: discovery-model-nucleotide-transformerExplanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Model type | DNA transformer encoder familySources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training |
| Architecture | Encoder-only transformers with 6-mer tokens; v1 uses learned positional encodings and v2 uses rotary positions and SwiGLU.Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training |
| Inputs | DNA sequences tokenized into 6-mers, with single-base handling of N and remainder bases.Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training |
| Outputs | Contextual representations used in specified downstream prediction workflows.Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training |
| Parameters | 50M to 2.5B across the documented v1/v2 family.Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training |
| Known versions | NT-v1 human-reference/1000G/multispecies variants and NT-v2 50M/100M/250M/500M. NT-v3 is a separate architecture described elsewhere in the repository.Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training |
| Training data | v1 variants use GRCh38, 3,202 human genomes or 850 multispecies genomes; v2 uses the multispecies corpus.Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training |
| Training cutoff | The paper specifies human-reference, 1000 Genomes and multispecies training collections by variant. A single latest-deposition date for all sequences is not supplied in the inspected pretraining-data section. · Not reported in inspected sourcesSources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training |
| Context limits | v1: approximately 6kb; v2: 2,048 tokens, approximately 12kb. Exact base count depends on special and ambiguous tokens.Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training |
| Weights licence | Separate checkpoint-distribution terms are not stated in the inspected release documentation and licence material. The source-code licence alone is not recorded as an explicit weight grant. · Not reported in inspected sourcesSources (6)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML; instadeepai/nucleotide-transformer: LICENSE.md · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training; LICENSE.md: licence text |
| Access | Official project documentation and implementation: https://github.com/instadeepai/nucleotide-transformerSources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training |
| Code licence | CC-BY-NC-SA-4.0Sourcesinstadeepai/nucleotide-transformer: LICENSE.md · LICENSE.md: licence text |
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
2 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Relationship: family discovery-model-nucleotide-transformer Individual claims | DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes DNABERT-2 paper GUE tables labelled NT-500M/2500M; official NT documentation Pre-trained models; source-labelled configuration NT-2500M-1000g | Existing reviewed locator: Table 6, row(NT-2500M-1000g) Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: Primary full-text snapshot retrieved 2026-09-17; exact bytes pinned by SHA-256 | source checked automated source review · 2026-09-23 Audit detailsSource review establishes this relationship only. Exact evaluated configurations and original numerical review status remain unchanged. Named NT training-corpus/parameter variants belong to broad NT family, not the catalog 50M v2 checkpoint. Field: Claim: model-evaluation-identity-1e3ecac98a4253894e00 Source artifact SHA-256: Hash scope: Exact retrieved primary paper artifact bytes. |
| Relationship: family discovery-model-nucleotide-transformer Individual claims | instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md DNABERT-2 paper GUE tables labelled NT-500M/2500M; official NT documentation Pre-trained models; source-labelled configuration NT-2500M-1000g | Existing reviewed locator: Table 6, row(NT-2500M-1000g) Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 2dc37b86e16a6970fbc731751f7719d9f676f7f9 | source checked automated source review · 2026-09-23 Audit detailsSource review establishes this relationship only. Exact evaluated configurations and original numerical review status remain unchanged. Named NT training-corpus/parameter variants belong to broad NT family, not the catalog 50M v2 checkpoint. Field: Claim: model-evaluation-identity-1e3ecac98a4253894e00 Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
View linked audit checks and correction history
Release 2026-09-29-06401fd5b220 · Record review: source checked
Stable ID: gue-method-nt-2500m-1000g