rewire.itbenchmarks
Configuration

NT-500M-human

Nucleotide Transformer is a family of DNA encoders pretrained on human or multispecies sequence corpora.

Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training

28 evaluations · 28 results

How it worksNucleotide Transformer workflow
Nucleotide Transformer workflow1. DNA sequence. Then: 2. 6-mer tokenization. Then: 3. Selected NT encoder. Then: 4. Representation or adapted predictorNucleotide Transformer workflow1. DNA sequence. Then: 2. 6-mer tokenization. Then: 3. Selected NT encoder. Then: 4. Representation or adapted predictorNucleotide Transformer workflow1. DNA sequence. Then: 2. 6-mer tokenization. Then: 3. Selected NT encoder. Then: 4. Representation or adapted predictor

Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.

Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training

Overview

Model type

DNA transformer encoder family

Inputs

DNA sequences tokenized into 6-mers, with single-base handling of N and remainder bases.

Outputs

Contextual representations used in specified downstream prediction workflows.

Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training

limited source coverage · Automated source review, 2026-09-16. All specifications and missing details

Evaluations and results

28 evaluations · 28 results. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: NT-500M-humanTask: GUE CORE-PROMOTER-DETECTION-ALL: Core promoter detection, dataset all
Dataset subset: GUE Core promoter detection, all (GUE split)
63.5% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-500M-human on GUE CORE-PROMOTER-DETECTION-ALL: Core promoter detection, dataset all

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-500M-human), column(Core promoter detection all)
Configuration: NT-500M-humanTask: GUE CORE-PROMOTER-DETECTION-NOTATA: Core promoter detection, dataset notata
Dataset subset: GUE Core promoter detection, notata (GUE split)
64.8% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-500M-human on GUE CORE-PROMOTER-DETECTION-NOTATA: Core promoter detection, dataset notata

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-500M-human), column(Core promoter detection notata)
Configuration: NT-500M-humanTask: GUE CORE-PROMOTER-DETECTION-TATA: Core promoter detection, dataset tata
Dataset subset: GUE Core promoter detection, tata (GUE split)
71.3% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-500M-human on GUE CORE-PROMOTER-DETECTION-TATA: Core promoter detection, dataset tata

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-500M-human), column(Core promoter detection tata)
Configuration: NT-500M-humanTask: GUE COVID-VARIANT-CLASSIFICATION-COVID: Covid variant classification, dataset Covid
Dataset subset: GUE Covid variant classification, Covid (GUE split)
50.8% f1
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-500M-human on GUE COVID-VARIANT-CLASSIFICATION-COVID: Covid variant classification, dataset Covid

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-500M-human), column(Covid variant classification Covid)
Configuration: NT-500M-humanTask: GUE EPIGENETIC-MARKS-PREDICTION-H3: Epigenetic marks prediction, dataset H3
Dataset subset: GUE Epigenetic marks prediction, H3 (GUE split)
69.7% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-500M-human on GUE EPIGENETIC-MARKS-PREDICTION-H3: Epigenetic marks prediction, dataset H3

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-500M-human), column(Epigenetic marks prediction H3)
Configuration: NT-500M-humanTask: GUE EPIGENETIC-MARKS-PREDICTION-H3K14AC: Epigenetic marks prediction, dataset H3K14ac
Dataset subset: GUE Epigenetic marks prediction, H3K14ac (GUE split)
33.5% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-500M-human on GUE EPIGENETIC-MARKS-PREDICTION-H3K14AC: Epigenetic marks prediction, dataset H3K14ac

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-500M-human), column(Epigenetic marks prediction H3K14ac)
Configuration: NT-500M-humanTask: GUE EPIGENETIC-MARKS-PREDICTION-H3K36ME3: Epigenetic marks prediction, dataset H3K36me3
Dataset subset: GUE Epigenetic marks prediction, H3K36me3 (GUE split)
44.1% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-500M-human on GUE EPIGENETIC-MARKS-PREDICTION-H3K36ME3: Epigenetic marks prediction, dataset H3K36me3

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-500M-human), column(Epigenetic marks prediction H3K36me3)
Configuration: NT-500M-humanTask: GUE EPIGENETIC-MARKS-PREDICTION-H3K4ME1: Epigenetic marks prediction, dataset H3K4me1
Dataset subset: GUE Epigenetic marks prediction, H3K4me1 (GUE split)
37.1% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-500M-human on GUE EPIGENETIC-MARKS-PREDICTION-H3K4ME1: Epigenetic marks prediction, dataset H3K4me1

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-500M-human), column(Epigenetic marks prediction H3K4me1)
Configuration: NT-500M-humanTask: GUE EPIGENETIC-MARKS-PREDICTION-H3K4ME2: Epigenetic marks prediction, dataset H3K4me2
Dataset subset: GUE Epigenetic marks prediction, H3K4me2 (GUE split)
30.9% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-500M-human on GUE EPIGENETIC-MARKS-PREDICTION-H3K4ME2: Epigenetic marks prediction, dataset H3K4me2

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-500M-human), column(Epigenetic marks prediction H3K4me2)
Configuration: NT-500M-humanTask: GUE EPIGENETIC-MARKS-PREDICTION-H3K4ME3: Epigenetic marks prediction, dataset H3K4me3
Dataset subset: GUE Epigenetic marks prediction, H3K4me3 (GUE split)
24.1% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-500M-human on GUE EPIGENETIC-MARKS-PREDICTION-H3K4ME3: Epigenetic marks prediction, dataset H3K4me3

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-500M-human), column(Epigenetic marks prediction H3K4me3)
Configuration: NT-500M-humanTask: GUE EPIGENETIC-MARKS-PREDICTION-H3K79ME3: Epigenetic marks prediction, dataset H3K79me3
Dataset subset: GUE Epigenetic marks prediction, H3K79me3 (GUE split)
58.4% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-500M-human on GUE EPIGENETIC-MARKS-PREDICTION-H3K79ME3: Epigenetic marks prediction, dataset H3K79me3

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-500M-human), column(Epigenetic marks prediction H3K79me3)
Configuration: NT-500M-humanTask: GUE EPIGENETIC-MARKS-PREDICTION-H3K9AC: Epigenetic marks prediction, dataset H3K9ac
Dataset subset: GUE Epigenetic marks prediction, H3K9ac (GUE split)
45.8% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-500M-human on GUE EPIGENETIC-MARKS-PREDICTION-H3K9AC: Epigenetic marks prediction, dataset H3K9ac

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-500M-human), column(Epigenetic marks prediction H3K9ac)
Configuration: NT-500M-humanTask: GUE EPIGENETIC-MARKS-PREDICTION-H4: Epigenetic marks prediction, dataset H4
Dataset subset: GUE Epigenetic marks prediction, H4 (GUE split)
76.2% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-500M-human on GUE EPIGENETIC-MARKS-PREDICTION-H4: Epigenetic marks prediction, dataset H4

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-500M-human), column(Epigenetic marks prediction H4)
Configuration: NT-500M-humanTask: GUE EPIGENETIC-MARKS-PREDICTION-H4AC: Epigenetic marks prediction, dataset H4ac
Dataset subset: GUE Epigenetic marks prediction, H4ac (GUE split)
33.7% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-500M-human on GUE EPIGENETIC-MARKS-PREDICTION-H4AC: Epigenetic marks prediction, dataset H4ac

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-500M-human), column(Epigenetic marks prediction H4ac)
Configuration: NT-500M-humanTask: GUE PROMOTER-DETECTION-ALL: Promoter detection, dataset all
Dataset subset: GUE Promoter detection, all (GUE split)
87.7% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-500M-human on GUE PROMOTER-DETECTION-ALL: Promoter detection, dataset all

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-500M-human), column(Promoter detection all)
Configuration: NT-500M-humanTask: GUE PROMOTER-DETECTION-NOTATA: Promoter detection, dataset notata
Dataset subset: GUE Promoter detection, notata (GUE split)
90.8% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-500M-human on GUE PROMOTER-DETECTION-NOTATA: Promoter detection, dataset notata

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-500M-human), column(Promoter detection notata)
Configuration: NT-500M-humanTask: GUE PROMOTER-DETECTION-TATA: Promoter detection, dataset tata
Dataset subset: GUE Promoter detection, tata (GUE split)
78.1% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-500M-human on GUE PROMOTER-DETECTION-TATA: Promoter detection, dataset tata

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-500M-human), column(Promoter detection tata)
Configuration: NT-500M-humanTask: GUE SPLICE-SITE-PREDICTION-RECONSTRUCT: Splice site prediction, dataset Reconstruct
Dataset subset: GUE Splice site prediction, Reconstruct (GUE split)
79.7% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-500M-human on GUE SPLICE-SITE-PREDICTION-RECONSTRUCT: Splice site prediction, dataset Reconstruct

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-500M-human), column(Splice site prediction Reconstruct)
Configuration: NT-500M-humanTask: GUE TRANSCRIPTION-FACTOR-PREDICTION-HUMAN-0: Transcription factor prediction (human), dataset 0
Dataset subset: GUE Transcription factor prediction (human), 0 (GUE split)
61.6% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-500M-human on GUE TRANSCRIPTION-FACTOR-PREDICTION-HUMAN-0: Transcription factor prediction (human), dataset 0

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-500M-human), column(Transcription factor prediction (human) 0)
Configuration: NT-500M-humanTask: GUE TRANSCRIPTION-FACTOR-PREDICTION-HUMAN-1: Transcription factor prediction (human), dataset 1
Dataset subset: GUE Transcription factor prediction (human), 1 (GUE split)
66.8% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-500M-human on GUE TRANSCRIPTION-FACTOR-PREDICTION-HUMAN-1: Transcription factor prediction (human), dataset 1

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-500M-human), column(Transcription factor prediction (human) 1)
Configuration: NT-500M-humanTask: GUE TRANSCRIPTION-FACTOR-PREDICTION-HUMAN-2: Transcription factor prediction (human), dataset 2
Dataset subset: GUE Transcription factor prediction (human), 2 (GUE split)
53.6% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-500M-human on GUE TRANSCRIPTION-FACTOR-PREDICTION-HUMAN-2: Transcription factor prediction (human), dataset 2

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-500M-human), column(Transcription factor prediction (human) 2)
Configuration: NT-500M-humanTask: GUE TRANSCRIPTION-FACTOR-PREDICTION-HUMAN-3: Transcription factor prediction (human), dataset 3
Dataset subset: GUE Transcription factor prediction (human), 3 (GUE split)
43% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-500M-human on GUE TRANSCRIPTION-FACTOR-PREDICTION-HUMAN-3: Transcription factor prediction (human), dataset 3

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-500M-human), column(Transcription factor prediction (human) 3)
Configuration: NT-500M-humanTask: GUE TRANSCRIPTION-FACTOR-PREDICTION-HUMAN-4: Transcription factor prediction (human), dataset 4
Dataset subset: GUE Transcription factor prediction (human), 4 (GUE split)
60.8% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-500M-human on GUE TRANSCRIPTION-FACTOR-PREDICTION-HUMAN-4: Transcription factor prediction (human), dataset 4

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-500M-human), column(Transcription factor prediction (human) 4)
Configuration: NT-500M-humanTask: GUE TRANSCRIPTION-FACTOR-PREDICTION-MOUSE-0: Transcription factor prediction (mouse), dataset 0
Dataset subset: GUE Transcription factor prediction (mouse), 0 (GUE split)
31% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-500M-human on GUE TRANSCRIPTION-FACTOR-PREDICTION-MOUSE-0: Transcription factor prediction (mouse), dataset 0

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-500M-human), column(Transcription factor prediction (mouse) 0)
Configuration: NT-500M-humanTask: GUE TRANSCRIPTION-FACTOR-PREDICTION-MOUSE-1: Transcription factor prediction (mouse), dataset 1
Dataset subset: GUE Transcription factor prediction (mouse), 1 (GUE split)
75% mcc
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-500M-human on GUE TRANSCRIPTION-FACTOR-PREDICTION-MOUSE-1: Transcription factor prediction (mouse), dataset 1

Fine-tuned on the GUE training split, scored on its test split. Split sizes are in Table 12.

Aggregation: Not reported

DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes · Table 6, row(NT-500M-human), column(Transcription factor prediction (mouse) 1)

Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.

Use this model

How it works, versions and access

Related profile: Nucleotide Transformer. This page retains the exact record and its evaluation context.

This configuration

Genome language model fine-tuned on each GUE dataset by the DNABERT-2 authors.

record
NT-500M-human
configuration
Not reported
entity type
Configuration

How it works

How it works

Nucleotide Transformer is a family of DNA encoders pretrained on human or multispecies sequence corpora. Encoder-only transformers with 6-mer tokens; v1 uses learned positional encodings and v2 uses rotary positions and SwiGLU. The documented inputs are DNA sequences tokenized into 6-mers, with single-base handling of N and remainder bases. The output consists of contextual representations used in specified downstream prediction workflows.

Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training
Versions and reproducibility

NT-v1 human-reference/1000G/multispecies variants and NT-v2 50M/100M/250M/500M. NT-v3 is a separate architecture described elsewhere in the repository. v1: approximately 6kb; v2: 2,048 tokens, approximately 12kb. Exact base count depends on special and ambiguous tokens.

Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training
Strengths, limitations and unresolved questions

Strengths and limitations

Strengths and considerations

Limitations and conditions

Profile review details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Stable record: discovery-model-nucleotide-transformer

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Inputs, outputs and configuration
PropertyDescription and evidence
Model typeDNA transformer encoder family
Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training
ArchitectureEncoder-only transformers with 6-mer tokens; v1 uses learned positional encodings and v2 uses rotary positions and SwiGLU.
Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training
InputsDNA sequences tokenized into 6-mers, with single-base handling of N and remainder bases.
Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training
OutputsContextual representations used in specified downstream prediction workflows.
Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training
Parameters50M to 2.5B across the documented v1/v2 family.
Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training
Known versionsNT-v1 human-reference/1000G/multispecies variants and NT-v2 50M/100M/250M/500M. NT-v3 is a separate architecture described elsewhere in the repository.
Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training
Training datav1 variants use GRCh38, 3,202 human genomes or 850 multispecies genomes; v2 uses the multispecies corpus.
Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training
Training cutoffThe paper specifies human-reference, 1000 Genomes and multispecies training collections by variant. A single latest-deposition date for all sequences is not supplied in the inspected pretraining-data section. · Not reported in inspected sources
Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training
Context limitsv1: approximately 6kb; v2: 2,048 tokens, approximately 12kb. Exact base count depends on special and ambiguous tokens.
Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training
Weights licenceSeparate checkpoint-distribution terms are not stated in the inspected release documentation and licence material. The source-code licence alone is not recorded as an explicit weight grant. · Not reported in inspected sources
Sources (6)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML; instadeepai/nucleotide-transformer: LICENSE.md · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training; LICENSE.md: licence text
AccessOfficial project documentation and implementation: https://github.com/instadeepai/nucleotide-transformer
Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training
Code licenceCC-BY-NC-SA-4.0
Sourcesinstadeepai/nucleotide-transformer: LICENSE.md · LICENSE.md: licence text

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

2 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Relationship: family
discovery-model-nucleotide-transformer
Individual claims
DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genomes

Original source ↗

DNABERT-2 paper GUE tables labelled NT-500M/2500M; official NT documentation Pre-trained models; source-labelled configuration NT-500M-human | Existing reviewed locator: Table 6, row(NT-500M-human)

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: Primary full-text snapshot retrieved 2026-09-17; exact bytes pinned by SHA-256
Retrieved: 2026-09-17T08:06:28.183387+00:00

source checked

automated source review · 2026-09-23

Audit details

Source review establishes this relationship only. Exact evaluated configurations and original numerical review status remain unchanged. Named NT training-corpus/parameter variants belong to broad NT family, not the catalog 50M v2 checkpoint.

Field: links:family:discovery-model-nucleotide-transformer

Claim: model-evaluation-identity-d8ee0333abe753526b9c

Source artifact SHA-256: 49300acee3e4afd44bebc3de9893c3bc310d331bd4805374e0952fdfbf366f06

Hash scope: Exact retrieved primary paper artifact bytes.

Inspected artifact

Relationship: family
discovery-model-nucleotide-transformer
Individual claims
instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md

Original source ↗

DNABERT-2 paper GUE tables labelled NT-500M/2500M; official NT documentation Pre-trained models; source-labelled configuration NT-500M-human | Existing reviewed locator: Table 6, row(NT-500M-human)

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2dc37b86e16a6970fbc731751f7719d9f676f7f9
Retrieved: 2026-09-16T19:46:19.364532+00:00

source checked

automated source review · 2026-09-23

Audit details

Source review establishes this relationship only. Exact evaluated configurations and original numerical review status remain unchanged. Named NT training-corpus/parameter variants belong to broad NT family, not the catalog 50M v2 checkpoint.

Field: links:family:discovery-model-nucleotide-transformer

Claim: model-evaluation-identity-d8ee0333abe753526b9c

Source artifact SHA-256: ab16d582de98652526b5cebb120eec969328f9db29dc741826bcd81c397e0672

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: source checked

1 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: gue-method-nt-500m-human

areas
dna-genomes
source locator
Table 6, row(NT-500M-human)
missing metadata
checkpoint revision: unreported; parameters: unextracted
Related records

Suggest a correction