Model type
DNA transformer encoder family
Nucleotide Transformer is a family of DNA encoders pretrained on human or multispecies sequence corpora.
263 evaluations · 265 results · 20 evaluated configurations using this model
Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.
DNA transformer encoder family
DNA sequences tokenized into 6-mers, with single-base handling of N and remainder bases.
Contextual representations used in specified downstream prediction workflows.
Official project documentation and implementation: https://github.com/instadeepai/nucleotide-transformer
limited source coverage · Automated source review, 2026-09-16. All specifications and missing details
263 evaluations · 265 results. Different protocols are not a single leaderboard.
Applied filters: All linked evaluations
| Tested configuration | Protocol and dataset | Finding | Evidence and details |
|---|---|---|---|
| Configuration: NT-1000G | Task: BEND CHROMATIN: Chromatin accessibility Dataset subset: ENCODE chromatin accessibility (BEND split) | 0.77 auroc fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNT-1000G on BEND CHROMATIN: Chromatin accessibility A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1. Aggregation: Not reported BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(NT-1000G), column(Chromatin accessibility) |
| Configuration: NT-1000G | Task: BEND CPG: CpG methylation Dataset subset: ENCODE CpG methylation (BEND split) | 0.89 auroc fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNT-1000G on BEND CPG: CpG methylation A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1. Aggregation: Not reported BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(NT-1000G), column(CpG methylation) |
| Configuration: NT-1000G | Task: BEND ENHANCER: Enhancer annotation Dataset subset: Fulco 2019, Gasperini 2019 and Enformer enhancer set (BEND split) | 0.04 auprc fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNT-1000G on BEND ENHANCER: Enhancer annotation A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1. Aggregation: Not reported BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(NT-1000G), column(Enhancer annotation) |
| Configuration: NT-1000G | Task: BEND GENE-FINDING: Gene finding Dataset subset: GENCODE (BEND split) | 0.49 mcc correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNT-1000G on BEND GENE-FINDING: Gene finding A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1. Aggregation: Not reported BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(NT-1000G), column(Gene finding) |
| Configuration: NT-1000G | Task: BEND HISTONE: Histone modification Dataset subset: ENCODE histone modification (BEND split) | 0.77 auroc fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNT-1000G on BEND HISTONE: Histone modification A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1. Aggregation: Not reported BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(NT-1000G), column(Histone modification) |
| Configuration: NT-1000G | Task: BEND VARIANT-DISEASE: Noncoding variant effects on disease Dataset subset: ClinVar disease variants (BEND split) | 0.49 auroc fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNT-1000G on BEND VARIANT-DISEASE: Noncoding variant effects on disease A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1. Aggregation: Not reported BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(NT-1000G), column(Noncoding variant effects on disease) |
| Configuration: NT-1000G | Task: BEND VARIANT-EXPRESSION: Noncoding variant effects on expression Dataset subset: DeepSEA expression variants (BEND split) | 0.45 auroc fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNT-1000G on BEND VARIANT-EXPRESSION: Noncoding variant effects on expression A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1. Aggregation: Not reported BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(NT-1000G), column(Noncoding variant effects on expression) |
| Configuration: NT-V2 | Task: BEND CHROMATIN: Chromatin accessibility Dataset subset: ENCODE chromatin accessibility (BEND split) | 0.8 auroc fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNT-V2 on BEND CHROMATIN: Chromatin accessibility A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1. Aggregation: Not reported BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(NT-V2), column(Chromatin accessibility) |
| Configuration: NT-V2 | Task: BEND CPG: CpG methylation Dataset subset: ENCODE CpG methylation (BEND split) | 0.91 auroc fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNT-V2 on BEND CPG: CpG methylation A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1. Aggregation: Not reported BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(NT-V2), column(CpG methylation) |
| Configuration: NT-V2 | Task: BEND ENHANCER: Enhancer annotation Dataset subset: Fulco 2019, Gasperini 2019 and Enformer enhancer set (BEND split) | 0.05 auprc fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNT-V2 on BEND ENHANCER: Enhancer annotation A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1. Aggregation: Not reported BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(NT-V2), column(Enhancer annotation) |
| Configuration: NT-V2 | Task: BEND GENE-FINDING: Gene finding Dataset subset: GENCODE (BEND split) | 0.64 mcc correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNT-V2 on BEND GENE-FINDING: Gene finding A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1. Aggregation: Not reported BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(NT-V2), column(Gene finding) |
| Configuration: NT-V2 | Task: BEND HISTONE: Histone modification Dataset subset: ENCODE histone modification (BEND split) | 0.76 auroc fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNT-V2 on BEND HISTONE: Histone modification A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1. Aggregation: Not reported BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(NT-V2), column(Histone modification) |
| Configuration: NT-V2 | Task: BEND VARIANT-DISEASE: Noncoding variant effects on disease Dataset subset: ClinVar disease variants (BEND split) | 0.48 auroc fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNT-V2 on BEND VARIANT-DISEASE: Noncoding variant effects on disease A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1. Aggregation: Not reported BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(NT-V2), column(Noncoding variant effects on disease) |
| Configuration: NT-V2 | Task: BEND VARIANT-EXPRESSION: Noncoding variant effects on expression Dataset subset: DeepSEA expression variants (BEND split) | 0.48 auroc fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNT-V2 on BEND VARIANT-EXPRESSION: Noncoding variant effects on expression A downstream head trained on frozen embeddings, except for the expert methods and the fully supervised baselines, which are trained end to end. Metric and splits are from Table 1. Aggregation: Not reported BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks · Table 3, row(NT-V2), column(Noncoding variant effects on expression) |
| Configuration: Nucleotide Transformer (fine-tuned) | Task: DART-Eval CA-AUROC-GM12878: Chromatin activity prediction, GM12878, positives against negatives Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split) | 0.938 auroc fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceSeparating positive GM12878 peaks from matched negatives. Aggregation: Not reported DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 5, row(Fine-Tuned NT), column(CA-AUROC-GM12878) |
| Configuration: Nucleotide Transformer (fine-tuned) | Task: DART-Eval CA-AUROC-H1ESC: Chromatin activity prediction, H1ESC, positives against negatives Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split) | 0.958 auroc fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceSeparating positive H1ESC peaks from matched negatives. Aggregation: Not reported DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 5, row(Fine-Tuned NT), column(CA-AUROC-H1ESC) |
| Configuration: Nucleotide Transformer (fine-tuned) | Task: DART-Eval CA-AUROC-HEPG2: Chromatin activity prediction, HEPG2, positives against negatives Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split) | 0.922 auroc fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceSeparating positive HEPG2 peaks from matched negatives. Aggregation: Not reported DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 5, row(Fine-Tuned NT), column(CA-AUROC-HEPG2) |
| Configuration: Nucleotide Transformer (fine-tuned) | Task: DART-Eval CA-AUROC-IMR90: Chromatin activity prediction, IMR90, positives against negatives Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split) | 0.975 auroc fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceSeparating positive IMR90 peaks from matched negatives. Aggregation: Not reported DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 5, row(Fine-Tuned NT), column(CA-AUROC-IMR90) |
| Configuration: Nucleotide Transformer (fine-tuned) | Task: DART-Eval CA-AUROC-K562: Chromatin activity prediction, K562, positives against negatives Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split) | 0.941 auroc fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceSeparating positive K562 peaks from matched negatives. Aggregation: Not reported DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 5, row(Fine-Tuned NT), column(CA-AUROC-K562) |
| Configuration: Nucleotide Transformer (fine-tuned) | Task: DART-Eval CA-SPEARMAN-GM12878: Chromatin activity prediction, GM12878, positives only Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split) | 0.515 spearman_r correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceRank correlation with measured accessibility among positive GM12878 peaks. Aggregation: Not reported DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 5, row(Fine-Tuned NT), column(CA-SPEARMAN-GM12878) |
| Configuration: Nucleotide Transformer (fine-tuned) | Task: DART-Eval CA-SPEARMAN-H1ESC: Chromatin activity prediction, H1ESC, positives only Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split) | 0.737 spearman_r correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceRank correlation with measured accessibility among positive H1ESC peaks. Aggregation: Not reported DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 5, row(Fine-Tuned NT), column(CA-SPEARMAN-H1ESC) |
| Configuration: Nucleotide Transformer (fine-tuned) | Task: DART-Eval CA-SPEARMAN-HEPG2: Chromatin activity prediction, HEPG2, positives only Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split) | 0.513 spearman_r correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceRank correlation with measured accessibility among positive HEPG2 peaks. Aggregation: Not reported DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 5, row(Fine-Tuned NT), column(CA-SPEARMAN-HEPG2) |
| Configuration: Nucleotide Transformer (fine-tuned) | Task: DART-Eval CA-SPEARMAN-IMR90: Chromatin activity prediction, IMR90, positives only Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split) | 0.489 spearman_r correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceRank correlation with measured accessibility among positive IMR90 peaks. Aggregation: Not reported DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 5, row(Fine-Tuned NT), column(CA-SPEARMAN-IMR90) |
| Configuration: Nucleotide Transformer (fine-tuned) | Task: DART-Eval CA-SPEARMAN-K562: Chromatin activity prediction, K562, positives only Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split) | 0.583 spearman_r correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceRank correlation with measured accessibility among positive K562 peaks. Aggregation: Not reported DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 5, row(Fine-Tuned NT), column(CA-SPEARMAN-K562) |
| Configuration: Nucleotide Transformer (fine-tuned) | Task: DART-Eval CTS-ACC: Cell-type-specific element classification, overall accuracy Dataset subset: ENCODE chromatin accessibility peaks in five cell lines (DART-Eval split) | 0.632 accuracy fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceClassify which of five cell lines a accessible element belongs to. Aggregation: Not reported DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 4, row(Fine-Tuned Nucleotide Transformer), column(CTS-ACC) |
Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.
These configurations, services and pipelines use this model within their own configurations. Their results, where available, are not assigned to the underlying model.
Nucleotide Transformer is a family of DNA encoders pretrained on human or multispecies sequence corpora. Encoder-only transformers with 6-mer tokens; v1 uses learned positional encodings and v2 uses rotary positions and SwiGLU. The documented inputs are DNA sequences tokenized into 6-mers, with single-base handling of N and remainder bases. The output consists of contextual representations used in specified downstream prediction workflows.
NT-v1 human-reference/1000G/multispecies variants and NT-v2 50M/100M/250M/500M. NT-v3 is a separate architecture described elsewhere in the repository. v1: approximately 6kb; v2: 2,048 tokens, approximately 12kb. Exact base count depends on special and ambiguous tokens.
Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.
Stable record: discovery-model-nucleotide-transformerExplanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Model type | DNA transformer encoder familySources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training |
| Architecture | Encoder-only transformers with 6-mer tokens; v1 uses learned positional encodings and v2 uses rotary positions and SwiGLU.Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training |
| Inputs | DNA sequences tokenized into 6-mers, with single-base handling of N and remainder bases.Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training |
| Outputs | Contextual representations used in specified downstream prediction workflows.Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training |
| Parameters | 50M to 2.5B across the documented v1/v2 family.Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training |
| Known versions | NT-v1 human-reference/1000G/multispecies variants and NT-v2 50M/100M/250M/500M. NT-v3 is a separate architecture described elsewhere in the repository.Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training |
| Training data | v1 variants use GRCh38, 3,202 human genomes or 850 multispecies genomes; v2 uses the multispecies corpus.Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training |
| Training cutoff | The paper specifies human-reference, 1000 Genomes and multispecies training collections by variant. A single latest-deposition date for all sequences is not supplied in the inspected pretraining-data section. · Not reported in inspected sourcesSources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training |
| Context limits | v1: approximately 6kb; v2: 2,048 tokens, approximately 12kb. Exact base count depends on special and ambiguous tokens.Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training |
| Weights licence | Separate checkpoint-distribution terms are not stated in the inspected release documentation and licence material. The source-code licence alone is not recorded as an explicit weight grant. · Not reported in inspected sourcesSources (6)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML; instadeepai/nucleotide-transformer: LICENSE.md · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training; LICENSE.md: licence text |
| Access | Official project documentation and implementation: https://github.com/instadeepai/nucleotide-transformerSources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training |
| Code licence | CC-BY-NC-SA-4.0Sourcesinstadeepai/nucleotide-transformer: LICENSE.md · LICENSE.md: licence text |
Applicability is distinct from a completed evaluation.
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
97 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation. Individual claims | instadeepai/nucleotide-transformer: README.md docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 2dc37b86e16a6970fbc731751f7719d9f676f7f9 | source checked automated source review · 2026-09-16 Audit detailsInspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied. Field: Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
| Diagram caption Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation. Individual claims | instadeepai/nucleotide-transformer: docs/segment_nt.md docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 2dc37b86e16a6970fbc731751f7719d9f676f7f9 | source checked automated source review · 2026-09-16 Audit detailsInspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied. Field: Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
| Diagram caption Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation. Individual claims | instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 2dc37b86e16a6970fbc731751f7719d9f676f7f9 | source checked automated source review · 2026-09-16 Audit detailsInspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied. Field: Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
| Diagram caption Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation. Individual claims | instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 2dc37b86e16a6970fbc731751f7719d9f676f7f9 | source checked automated source review · 2026-09-16 Audit detailsInspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied. Field: Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
| Diagram caption Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation. Individual claims | nt: Journal full-text XML docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: Retrieved page snapshot; no immutable publisher revision supplied | source checked automated source review · 2026-09-16 Audit detailsInspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied. Field: Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
Diagram steps
| instadeepai/nucleotide-transformer: README.md docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 2dc37b86e16a6970fbc731751f7719d9f676f7f9 | source checked automated source review · 2026-09-16 Audit detailsInspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied. Field: Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
Diagram steps
| instadeepai/nucleotide-transformer: docs/segment_nt.md docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 2dc37b86e16a6970fbc731751f7719d9f676f7f9 | source checked automated source review · 2026-09-16 Audit detailsInspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied. Field: Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
Diagram steps
| instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 2dc37b86e16a6970fbc731751f7719d9f676f7f9 | source checked automated source review · 2026-09-16 Audit detailsInspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied. Field: Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
Diagram steps
| instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 2dc37b86e16a6970fbc731751f7719d9f676f7f9 | source checked automated source review · 2026-09-16 Audit detailsInspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied. Field: Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
Diagram steps
| nt: Journal full-text XML docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: Retrieved page snapshot; no immutable publisher revision supplied | source checked automated source review · 2026-09-16 Audit detailsInspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied. Field: Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
View linked audit checks and correction history
Release 2026-09-29-06401fd5b220 · Record review: discovered
Stable ID: discovery-model-nucleotide-transformer