rewire.itbenchmarks
Model

Nucleotide Transformer

Nucleotide Transformer is a family of DNA encoders pretrained on human or multispecies sequence corpora.

Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training

263 evaluations · 265 results · 20 evaluated configurations using this model

How it worksNucleotide Transformer workflow
Nucleotide Transformer workflow1. DNA sequence. Then: 2. 6-mer tokenization. Then: 3. Selected NT encoder. Then: 4. Representation or adapted predictorNucleotide Transformer workflow1. DNA sequence. Then: 2. 6-mer tokenization. Then: 3. Selected NT encoder. Then: 4. Representation or adapted predictorNucleotide Transformer workflow1. DNA sequence. Then: 2. 6-mer tokenization. Then: 3. Selected NT encoder. Then: 4. Representation or adapted predictor

Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.

Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training

Overview

Model type

DNA transformer encoder family

Inputs

DNA sequences tokenized into 6-mers, with single-base handling of N and remainder bases.

Outputs

Contextual representations used in specified downstream prediction workflows.

Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training

limited source coverage · Automated source review, 2026-09-16. All specifications and missing details

Evaluations and results

263 evaluations · 265 results. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: NT-1000GTask: BEND CHROMATIN: Chromatin accessibility
Dataset subset: ENCODE chromatin accessibility (BEND split)
0.77 auroc
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-1000G on BEND CHROMATIN: Chromatin accessibility

A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1.

Aggregation: Not reported

BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(NT-1000G), column(Chromatin accessibility)
Configuration: NT-1000GTask: BEND CPG: CpG methylation
Dataset subset: ENCODE CpG methylation (BEND split)
0.89 auroc
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-1000G on BEND CPG: CpG methylation

A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1.

Aggregation: Not reported

BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(NT-1000G), column(CpG methylation)
Configuration: NT-1000GTask: BEND ENHANCER: Enhancer annotation
Dataset subset: Fulco 2019, Gasperini 2019 and Enformer enhancer set (BEND split)
0.04 auprc
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-1000G on BEND ENHANCER: Enhancer annotation

A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1.

Aggregation: Not reported

BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(NT-1000G), column(Enhancer annotation)
Configuration: NT-1000GTask: BEND GENE-FINDING: Gene finding
Dataset subset: GENCODE (BEND split)
0.49 mcc
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-1000G on BEND GENE-FINDING: Gene finding

A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1.

Aggregation: Not reported

BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(NT-1000G), column(Gene finding)
Configuration: NT-1000GTask: BEND HISTONE: Histone modification
Dataset subset: ENCODE histone modification (BEND split)
0.77 auroc
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-1000G on BEND HISTONE: Histone modification

A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1.

Aggregation: Not reported

BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(NT-1000G), column(Histone modification)
Configuration: NT-1000GTask: BEND VARIANT-DISEASE: Noncoding variant effects on disease
Dataset subset: ClinVar disease variants (BEND split)
0.49 auroc
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-1000G on BEND VARIANT-DISEASE: Noncoding variant effects on disease

A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1.

Aggregation: Not reported

BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(NT-1000G), column(Noncoding variant effects on disease)
Configuration: NT-1000GTask: BEND VARIANT-EXPRESSION: Noncoding variant effects on expression
Dataset subset: DeepSEA expression variants (BEND split)
0.45 auroc
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-1000G on BEND VARIANT-EXPRESSION: Noncoding variant effects on expression

A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1.

Aggregation: Not reported

BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(NT-1000G), column(Noncoding variant effects on expression)
Configuration: NT-V2Task: BEND CHROMATIN: Chromatin accessibility
Dataset subset: ENCODE chromatin accessibility (BEND split)
0.8 auroc
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-V2 on BEND CHROMATIN: Chromatin accessibility

A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1.

Aggregation: Not reported

BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(NT-V2), column(Chromatin accessibility)
Configuration: NT-V2Task: BEND CPG: CpG methylation
Dataset subset: ENCODE CpG methylation (BEND split)
0.91 auroc
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-V2 on BEND CPG: CpG methylation

A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1.

Aggregation: Not reported

BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(NT-V2), column(CpG methylation)
Configuration: NT-V2Task: BEND ENHANCER: Enhancer annotation
Dataset subset: Fulco 2019, Gasperini 2019 and Enformer enhancer set (BEND split)
0.05 auprc
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-V2 on BEND ENHANCER: Enhancer annotation

A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1.

Aggregation: Not reported

BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(NT-V2), column(Enhancer annotation)
Configuration: NT-V2Task: BEND GENE-FINDING: Gene finding
Dataset subset: GENCODE (BEND split)
0.64 mcc
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-V2 on BEND GENE-FINDING: Gene finding

A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1.

Aggregation: Not reported

BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(NT-V2), column(Gene finding)
Configuration: NT-V2Task: BEND HISTONE: Histone modification
Dataset subset: ENCODE histone modification (BEND split)
0.76 auroc
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-V2 on BEND HISTONE: Histone modification

A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1.

Aggregation: Not reported

BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(NT-V2), column(Histone modification)
Configuration: NT-V2Task: BEND VARIANT-DISEASE: Noncoding variant effects on disease
Dataset subset: ClinVar disease variants (BEND split)
0.48 auroc
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-V2 on BEND VARIANT-DISEASE: Noncoding variant effects on disease

A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1.

Aggregation: Not reported

BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(NT-V2), column(Noncoding variant effects on disease)
Configuration: NT-V2Task: BEND VARIANT-EXPRESSION: Noncoding variant effects on expression
Dataset subset: DeepSEA expression variants (BEND split)
0.48 auroc
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NT-V2 on BEND VARIANT-EXPRESSION: Noncoding variant effects on expression

A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1.

Aggregation: Not reported

BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(NT-V2), column(Noncoding variant effects on expression)
Configuration: Nucleotide Transformer (fine-tuned)Task: DART-Eval CA-AUROC-GM12878: Chromatin activity prediction, GM12878, positives against negatives
Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split)
0.938 auroc
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

Nucleotide Transformer (fine-tuned) on DART-Eval CA-AUROC-GM12878: Chromatin activity prediction, GM12878, positives against negatives

Separating positive GM12878 peaks from matched negatives.

Aggregation: Not reported

DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 5, row(Fine-Tuned NT), column(CA-AUROC-GM12878)
Configuration: Nucleotide Transformer (fine-tuned)Task: DART-Eval CA-AUROC-H1ESC: Chromatin activity prediction, H1ESC, positives against negatives
Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split)
0.958 auroc
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

Nucleotide Transformer (fine-tuned) on DART-Eval CA-AUROC-H1ESC: Chromatin activity prediction, H1ESC, positives against negatives

Separating positive H1ESC peaks from matched negatives.

Aggregation: Not reported

DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 5, row(Fine-Tuned NT), column(CA-AUROC-H1ESC)
Configuration: Nucleotide Transformer (fine-tuned)Task: DART-Eval CA-AUROC-HEPG2: Chromatin activity prediction, HEPG2, positives against negatives
Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split)
0.922 auroc
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

Nucleotide Transformer (fine-tuned) on DART-Eval CA-AUROC-HEPG2: Chromatin activity prediction, HEPG2, positives against negatives

Separating positive HEPG2 peaks from matched negatives.

Aggregation: Not reported

DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 5, row(Fine-Tuned NT), column(CA-AUROC-HEPG2)
Configuration: Nucleotide Transformer (fine-tuned)Task: DART-Eval CA-AUROC-IMR90: Chromatin activity prediction, IMR90, positives against negatives
Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split)
0.975 auroc
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

Nucleotide Transformer (fine-tuned) on DART-Eval CA-AUROC-IMR90: Chromatin activity prediction, IMR90, positives against negatives

Separating positive IMR90 peaks from matched negatives.

Aggregation: Not reported

DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 5, row(Fine-Tuned NT), column(CA-AUROC-IMR90)
Configuration: Nucleotide Transformer (fine-tuned)Task: DART-Eval CA-AUROC-K562: Chromatin activity prediction, K562, positives against negatives
Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split)
0.941 auroc
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

Nucleotide Transformer (fine-tuned) on DART-Eval CA-AUROC-K562: Chromatin activity prediction, K562, positives against negatives

Separating positive K562 peaks from matched negatives.

Aggregation: Not reported

DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 5, row(Fine-Tuned NT), column(CA-AUROC-K562)
Configuration: Nucleotide Transformer (fine-tuned)Task: DART-Eval CA-SPEARMAN-GM12878: Chromatin activity prediction, GM12878, positives only
Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split)
0.515 spearman_r
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

Nucleotide Transformer (fine-tuned) on DART-Eval CA-SPEARMAN-GM12878: Chromatin activity prediction, GM12878, positives only

Rank correlation with measured accessibility among positive GM12878 peaks.

Aggregation: Not reported

DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 5, row(Fine-Tuned NT), column(CA-SPEARMAN-GM12878)
Configuration: Nucleotide Transformer (fine-tuned)Task: DART-Eval CA-SPEARMAN-H1ESC: Chromatin activity prediction, H1ESC, positives only
Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split)
0.737 spearman_r
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

Nucleotide Transformer (fine-tuned) on DART-Eval CA-SPEARMAN-H1ESC: Chromatin activity prediction, H1ESC, positives only

Rank correlation with measured accessibility among positive H1ESC peaks.

Aggregation: Not reported

DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 5, row(Fine-Tuned NT), column(CA-SPEARMAN-H1ESC)
Configuration: Nucleotide Transformer (fine-tuned)Task: DART-Eval CA-SPEARMAN-HEPG2: Chromatin activity prediction, HEPG2, positives only
Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split)
0.513 spearman_r
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

Nucleotide Transformer (fine-tuned) on DART-Eval CA-SPEARMAN-HEPG2: Chromatin activity prediction, HEPG2, positives only

Rank correlation with measured accessibility among positive HEPG2 peaks.

Aggregation: Not reported

DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 5, row(Fine-Tuned NT), column(CA-SPEARMAN-HEPG2)
Configuration: Nucleotide Transformer (fine-tuned)Task: DART-Eval CA-SPEARMAN-IMR90: Chromatin activity prediction, IMR90, positives only
Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split)
0.489 spearman_r
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

Nucleotide Transformer (fine-tuned) on DART-Eval CA-SPEARMAN-IMR90: Chromatin activity prediction, IMR90, positives only

Rank correlation with measured accessibility among positive IMR90 peaks.

Aggregation: Not reported

DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 5, row(Fine-Tuned NT), column(CA-SPEARMAN-IMR90)
Configuration: Nucleotide Transformer (fine-tuned)Task: DART-Eval CA-SPEARMAN-K562: Chromatin activity prediction, K562, positives only
Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split)
0.583 spearman_r
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

Nucleotide Transformer (fine-tuned) on DART-Eval CA-SPEARMAN-K562: Chromatin activity prediction, K562, positives only

Rank correlation with measured accessibility among positive K562 peaks.

Aggregation: Not reported

DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 5, row(Fine-Tuned NT), column(CA-SPEARMAN-K562)
Configuration: Nucleotide Transformer (fine-tuned)Task: DART-Eval CTS-ACC: Cell-type-specific element classification, overall accuracy
Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split)
0.632 accuracy
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

Nucleotide Transformer (fine-tuned) on DART-Eval CTS-ACC: Cell-type-specific element classification, overall accuracy

Classify which of five cell lines a accessible element belongs to.

Aggregation: Not reported

DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 4, row(Fine-Tuned Nucleotide Transformer), column(CTS-ACC)

Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.

Related configurations, pipelines and services

These configurations, services and pipelines use this model within their own configurations. Their results, where available, are not assigned to the underlying model.

Use this model

How it works, versions and access

Versions and evaluated configurations

How it works

How it works

Nucleotide Transformer is a family of DNA encoders pretrained on human or multispecies sequence corpora. Encoder-only transformers with 6-mer tokens; v1 uses learned positional encodings and v2 uses rotary positions and SwiGLU. The documented inputs are DNA sequences tokenized into 6-mers, with single-base handling of N and remainder bases. The output consists of contextual representations used in specified downstream prediction workflows.

Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training
Versions and reproducibility

NT-v1 human-reference/1000G/multispecies variants and NT-v2 50M/100M/250M/500M. NT-v3 is a separate architecture described elsewhere in the repository. v1: approximately 6kb; v2: 2,048 tokens, approximately 12kb. Exact base count depends on special and ambiguous tokens.

Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training
Strengths, limitations and unresolved questions

Strengths and limitations

Strengths and considerations

Limitations and conditions

Profile review details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Stable record: discovery-model-nucleotide-transformer

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Inputs, outputs and configuration
PropertyDescription and evidence
Model typeDNA transformer encoder family
Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training
ArchitectureEncoder-only transformers with 6-mer tokens; v1 uses learned positional encodings and v2 uses rotary positions and SwiGLU.
Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training
InputsDNA sequences tokenized into 6-mers, with single-base handling of N and remainder bases.
Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training
OutputsContextual representations used in specified downstream prediction workflows.
Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training
Parameters50M to 2.5B across the documented v1/v2 family.
Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training
Known versionsNT-v1 human-reference/1000G/multispecies variants and NT-v2 50M/100M/250M/500M. NT-v3 is a separate architecture described elsewhere in the repository.
Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training
Training datav1 variants use GRCh38, 3,202 human genomes or 850 multispecies genomes; v2 uses the multispecies corpus.
Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training
Training cutoffThe paper specifies human-reference, 1000 Genomes and multispecies training collections by variant. A single latest-deposition date for all sequences is not supplied in the inspected pretraining-data section. · Not reported in inspected sources
Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training
Context limitsv1: approximately 6kb; v2: 2,048 tokens, approximately 12kb. Exact base count depends on special and ambiguous tokens.
Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training
Weights licenceSeparate checkpoint-distribution terms are not stated in the inspected release documentation and licence material. The source-code licence alone is not recorded as an explicit weight grant. · Not reported in inspected sources
Sources (6)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML; instadeepai/nucleotide-transformer: LICENSE.md · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training; LICENSE.md: licence text
AccessOfficial project documentation and implementation: https://github.com/instadeepai/nucleotide-transformer
Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training
Code licenceCC-BY-NC-SA-4.0
Sourcesinstadeepai/nucleotide-transformer: LICENSE.md · LICENSE.md: licence text
Applicable tests and references

Applicability is distinct from a completed evaluation.

  • GENEB · Proposed association

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

97 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.
Individual claims
instadeepai/nucleotide-transformer: README.md

Original source ↗

docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2dc37b86e16a6970fbc731751f7719d9f676f7f9
Retrieved: 2026-09-16T19:46:19.364532+00:00

source checked

automated source review · 2026-09-16

Audit details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 9f51bbb20c4c5c36e77fb03ca1c5c36236e287c48a1ee31f53150545d421ec25

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram caption
Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.
Individual claims
instadeepai/nucleotide-transformer: docs/segment_nt.md

Original source ↗

docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2dc37b86e16a6970fbc731751f7719d9f676f7f9
Retrieved: 2026-09-16T19:46:19.364532+00:00

source checked

automated source review · 2026-09-16

Audit details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 8eec4580ba64ab944fb9b42674be70ffe793135f3503cce8640f5b08f8290f7a

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram caption
Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.
Individual claims
instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md

Original source ↗

docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2dc37b86e16a6970fbc731751f7719d9f676f7f9
Retrieved: 2026-09-16T19:46:19.364532+00:00

source checked

automated source review · 2026-09-16

Audit details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 23e27d40473bbabadac45c56e8e282349f0b053da85fda89ff2125b5fe381fc6

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram caption
Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.
Individual claims
instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md

Original source ↗

docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2dc37b86e16a6970fbc731751f7719d9f676f7f9
Retrieved: 2026-09-16T19:46:19.364532+00:00

source checked

automated source review · 2026-09-16

Audit details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: ab16d582de98652526b5cebb120eec969328f9db29dc741826bcd81c397e0672

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram caption
Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.
Individual claims
nt: Journal full-text XML

Original source ↗

docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: Retrieved page snapshot; no immutable publisher revision supplied
Retrieved: 2026-09-16T20:16:14.422628+00:00

source checked

automated source review · 2026-09-16

Audit details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 7c3b78a4f38ef053a08e466222e5662dd9711d91c535512d0b35d459d1fc7249

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram steps
  • DNA sequence
  • 6-mer tokenization
  • Selected NT encoder
  • Representation or adapted predictor
Individual claims
instadeepai/nucleotide-transformer: README.md

Original source ↗

docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2dc37b86e16a6970fbc731751f7719d9f676f7f9
Retrieved: 2026-09-16T19:46:19.364532+00:00

source checked

automated source review · 2026-09-16

Audit details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 9f51bbb20c4c5c36e77fb03ca1c5c36236e287c48a1ee31f53150545d421ec25

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram steps
  • DNA sequence
  • 6-mer tokenization
  • Selected NT encoder
  • Representation or adapted predictor
Individual claims
instadeepai/nucleotide-transformer: docs/segment_nt.md

Original source ↗

docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2dc37b86e16a6970fbc731751f7719d9f676f7f9
Retrieved: 2026-09-16T19:46:19.364532+00:00

source checked

automated source review · 2026-09-16

Audit details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 8eec4580ba64ab944fb9b42674be70ffe793135f3503cce8640f5b08f8290f7a

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram steps
  • DNA sequence
  • 6-mer tokenization
  • Selected NT encoder
  • Representation or adapted predictor
Individual claims
instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md

Original source ↗

docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2dc37b86e16a6970fbc731751f7719d9f676f7f9
Retrieved: 2026-09-16T19:46:19.364532+00:00

source checked

automated source review · 2026-09-16

Audit details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 23e27d40473bbabadac45c56e8e282349f0b053da85fda89ff2125b5fe381fc6

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram steps
  • DNA sequence
  • 6-mer tokenization
  • Selected NT encoder
  • Representation or adapted predictor
Individual claims
instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md

Original source ↗

docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2dc37b86e16a6970fbc731751f7719d9f676f7f9
Retrieved: 2026-09-16T19:46:19.364532+00:00

source checked

automated source review · 2026-09-16

Audit details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: ab16d582de98652526b5cebb120eec969328f9db29dc741826bcd81c397e0672

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram steps
  • DNA sequence
  • 6-mer tokenization
  • Selected NT encoder
  • Representation or adapted predictor
Individual claims
nt: Journal full-text XML

Original source ↗

docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: Retrieved page snapshot; no immutable publisher revision supplied
Retrieved: 2026-09-16T20:16:14.422628+00:00

source checked

automated source review · 2026-09-16

Audit details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 7c3b78a4f38ef053a08e466222e5662dd9711d91c535512d0b35d459d1fc7249

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: discovered

7 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: discovery-model-nucleotide-transformer

areas
genomics
access
official_source_linked
benchmark applicability
candidate; not evidence of a reported evaluation
candidate benchmark ids
discovery-benchmark-geneb
entity level
family
reported name
Nucleotide Transformer
version
Not reported
historical missing metadata
checkpoint: unextracted; code licence: unextracted; parameters: unextracted; training cutoff: unextracted; training data: unextracted; version: unextracted; weights licence: unextracted
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
entity classification
review date: 2026-09-17; rationale: The cited profile describes a named learned biological predictor or representation model/family. Preserve this identity separately from task-specific fitting, individual checkpoints, pipelines and hosted access.; source ids: evidence-official-18d8f4d5fc3f6616922d; evidence-official-657e83427ab59f3aec83; evidence-official-56f02d45976d011d80aa; evidence-official-4920952f9b3c4b8909a0; evidence-official-aadbeb0f10ec55d34f5f; source locator: docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training; ambiguities: None recorded
Related records

Suggest a correction