Model type
DNA transformer encoder family
Nucleotide Transformer is a family of DNA encoders pretrained on human or multispecies sequence corpora.
Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.
DNA transformer encoder family
DNA sequences tokenized into 6-mers, with single-base handling of N and remainder bases.
Contextual representations used in specified downstream prediction workflows.
Official project documentation and implementation: https://github.com/instadeepai/nucleotide-transformer
limited source coverage · Automated source review, 2026-09-16. All specifications and missing details
42 evaluations · 42 results. Different protocols are not a single leaderboard.
Applied filters: All linked evaluations
| Tested configuration | Protocol and dataset | Finding | Evidence and details |
|---|---|---|---|
| Configuration: N.T.-v2-50m | Task: NABench CCV-CORR-APTAMER: Fitness prediction on aptamer assays, supervised, contiguous cross validation Dataset subset: NABench aptamer assays (NABench split) | 0.056 spearman correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceScored supervised, contiguous cross validation across the NABench aptamer assays. Aggregation: Not reported NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 11, row(N.T.-v2-50m), column(aptamer) |
| Configuration: N.T.-v2-50m | Task: NABench CCV-CORR-ENHANCER: Fitness prediction on enhancer assays, supervised, contiguous cross validation Dataset subset: NABench enhancer assays (NABench split) | 0.116 spearman correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceScored supervised, contiguous cross validation across the NABench enhancer assays. Aggregation: Not reported NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 11, row(N.T.-v2-50m), column(enhancer) |
| Configuration: N.T.-v2-50m | Task: NABench CCV-CORR-MRNA: Fitness prediction on mRNA assays, supervised, contiguous cross validation Dataset subset: NABench mRNA assays (NABench split) | 0.073 spearman correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceScored supervised, contiguous cross validation across the NABench mRNA assays. Aggregation: Not reported NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 11, row(N.T.-v2-50m), column(mRNA) |
| Configuration: N.T.-v2-50m | Task: NABench CCV-CORR-PROMOTER: Fitness prediction on promoter assays, supervised, contiguous cross validation Dataset subset: NABench promoter assays (NABench split) | 0.22 spearman correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceScored supervised, contiguous cross validation across the NABench promoter assays. Aggregation: Not reported NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 11, row(N.T.-v2-50m), column(promoter) |
| Configuration: N.T.-v2-50m | Task: NABench CCV-CORR-RIBOZYME: Fitness prediction on ribozyme assays, supervised, contiguous cross validation Dataset subset: NABench ribozyme assays (NABench split) | 0.3 spearman correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceScored supervised, contiguous cross validation across the NABench ribozyme assays. Aggregation: Not reported NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 11, row(N.T.-v2-50m), column(ribozyme) |
| Configuration: N.T.-v2-50m | Task: NABench DMS-CCV: Overall fitness prediction on NABench deep mutational scanning assays, Contiguous cross validation Spearman ρ Dataset subset: NABench deep mutational scanning assays (NABench split) | 0.22 spearman correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceAggregated by the NABench authors across NABench deep mutational scanning assays. Aggregation: Not reported NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 6, row(N.T.-v2-50m), column(Contiguous cross validation Spearman ρ) |
| Configuration: N.T.-v2-50m | Task: NABench DMS-FS: Overall fitness prediction on NABench deep mutational scanning assays, Few-shot Spearman ρ Dataset subset: NABench deep mutational scanning assays (NABench split) | 0.11 spearman correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceAggregated by the NABench authors across NABench deep mutational scanning assays. Aggregation: Not reported NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 6, row(N.T.-v2-50m), column(Few-shot Spearman ρ) |
| Configuration: N.T.-v2-50m | Task: NABench DMS-RCV: Overall fitness prediction on NABench deep mutational scanning assays, Random cross validation Spearman ρ Dataset subset: NABench deep mutational scanning assays (NABench split) | 0.465 spearman correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceAggregated by the NABench authors across NABench deep mutational scanning assays. Aggregation: Not reported NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 6, row(N.T.-v2-50m), column(Random cross validation Spearman ρ) |
| Configuration: N.T.-v2-50m | Task: NABench DMS-ZS-AUC: Overall fitness prediction on NABench deep mutational scanning assays, Zero-shot AUC Dataset subset: NABench deep mutational scanning assays (NABench split) | 0.512 auc fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceAggregated by the NABench authors across NABench deep mutational scanning assays. Aggregation: Not reported NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 6, row(N.T.-v2-50m), column(Zero-shot AUC) |
| Configuration: N.T.-v2-50m | Task: NABench DMS-ZS-CORR: Overall fitness prediction on NABench deep mutational scanning assays, Zero-shot Spearman ρ Dataset subset: NABench deep mutational scanning assays (NABench split) | 0.097 spearman correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceAggregated by the NABench authors across NABench deep mutational scanning assays. Aggregation: Not reported NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 6, row(N.T.-v2-50m), column(Zero-shot Spearman ρ) |
| Configuration: N.T.-v2-50m | Task: NABench DMS-ZS-MCC: Overall fitness prediction on NABench deep mutational scanning assays, Zero-shot MCC Dataset subset: NABench deep mutational scanning assays (NABench split) | 0.044 mcc correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceAggregated by the NABench authors across NABench deep mutational scanning assays. Aggregation: Not reported NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 6, row(N.T.-v2-50m), column(Zero-shot MCC) |
| Configuration: N.T.-v2-50m | Task: NABench DMS-ZS-NDCG: Overall fitness prediction on NABench deep mutational scanning assays, Zero-shot NDCG Dataset subset: NABench deep mutational scanning assays (NABench split) | 0.349 ndcg fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceAggregated by the NABench authors across NABench deep mutational scanning assays. Aggregation: Not reported NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 6, row(N.T.-v2-50m), column(Zero-shot NDCG) |
| Configuration: N.T.-v2-50m | Task: NABench FS-CORR-APTAMER: Fitness prediction on aptamer assays, few-shot Dataset subset: NABench aptamer assays (NABench split) | 0.23 spearman correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceN.T.-v2-50m on NABench FS-CORR-APTAMER: Fitness prediction on aptamer assays, few-shot Scored few-shot across the NABench aptamer assays. Aggregation: Not reported NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 9, row(N.T.-v2-50m), column(aptamer) |
| Configuration: N.T.-v2-50m | Task: NABench FS-CORR-ENHANCER: Fitness prediction on enhancer assays, few-shot Dataset subset: NABench enhancer assays (NABench split) | 0.026 spearman correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceN.T.-v2-50m on NABench FS-CORR-ENHANCER: Fitness prediction on enhancer assays, few-shot Scored few-shot across the NABench enhancer assays. Aggregation: Not reported NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 9, row(N.T.-v2-50m), column(enhancer) |
| Configuration: N.T.-v2-50m | Task: NABench FS-CORR-MRNA: Fitness prediction on mRNA assays, few-shot Dataset subset: NABench mRNA assays (NABench split) | 0.214 spearman correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceN.T.-v2-50m on NABench FS-CORR-MRNA: Fitness prediction on mRNA assays, few-shot Scored few-shot across the NABench mRNA assays. Aggregation: Not reported NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 9, row(N.T.-v2-50m), column(mRNA) |
| Configuration: N.T.-v2-50m | Task: NABench FS-CORR-PROMOTER: Fitness prediction on promoter assays, few-shot Dataset subset: NABench promoter assays (NABench split) | 0.061 spearman correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceN.T.-v2-50m on NABench FS-CORR-PROMOTER: Fitness prediction on promoter assays, few-shot Scored few-shot across the NABench promoter assays. Aggregation: Not reported NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 9, row(N.T.-v2-50m), column(promoter) |
| Configuration: N.T.-v2-50m | Task: NABench FS-CORR-RIBOZYME: Fitness prediction on ribozyme assays, few-shot Dataset subset: NABench ribozyme assays (NABench split) | 0.09 spearman correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceN.T.-v2-50m on NABench FS-CORR-RIBOZYME: Fitness prediction on ribozyme assays, few-shot Scored few-shot across the NABench ribozyme assays. Aggregation: Not reported NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 9, row(N.T.-v2-50m), column(ribozyme) |
| Configuration: N.T.-v2-50m | Task: NABench FS-CORR-TRNA: Fitness prediction on tRNA assays, few-shot Dataset subset: NABench tRNA assays (NABench split) | 0.327 spearman correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceN.T.-v2-50m on NABench FS-CORR-TRNA: Fitness prediction on tRNA assays, few-shot Scored few-shot across the NABench tRNA assays. Aggregation: Not reported NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 9, row(N.T.-v2-50m), column(tRNA) |
| Configuration: N.T.-v2-50m | Task: NABench RCV-CORR-APTAMER: Fitness prediction on aptamer assays, supervised, random cross validation Dataset subset: NABench aptamer assays (NABench split) | 0.462 spearman correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceScored supervised, random cross validation across the NABench aptamer assays. Aggregation: Not reported NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 10, row(N.T.-v2-50m), column(aptamer) |
| Configuration: N.T.-v2-50m | Task: NABench RCV-CORR-ENHANCER: Fitness prediction on enhancer assays, supervised, random cross validation Dataset subset: NABench enhancer assays (NABench split) | 0.148 spearman correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceScored supervised, random cross validation across the NABench enhancer assays. Aggregation: Not reported NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 10, row(N.T.-v2-50m), column(enhancer) |
| Configuration: N.T.-v2-50m | Task: NABench RCV-CORR-MRNA: Fitness prediction on mRNA assays, supervised, random cross validation Dataset subset: NABench mRNA assays (NABench split) | 0.57 spearman correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceScored supervised, random cross validation across the NABench mRNA assays. Aggregation: Not reported NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 10, row(N.T.-v2-50m), column(mRNA) |
| Configuration: N.T.-v2-50m | Task: NABench RCV-CORR-PROMOTER: Fitness prediction on promoter assays, supervised, random cross validation Dataset subset: NABench promoter assays (NABench split) | 0.632 spearman correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceScored supervised, random cross validation across the NABench promoter assays. Aggregation: Not reported NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 10, row(N.T.-v2-50m), column(promoter) |
| Configuration: N.T.-v2-50m | Task: NABench RCV-CORR-RIBOZYME: Fitness prediction on ribozyme assays, supervised, random cross validation Dataset subset: NABench ribozyme assays (NABench split) | 0.464 spearman correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceScored supervised, random cross validation across the NABench ribozyme assays. Aggregation: Not reported NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 10, row(N.T.-v2-50m), column(ribozyme) |
| Configuration: N.T.-v2-50m | Task: NABench RCV-CORR-TRNA: Fitness prediction on tRNA assays, supervised, random cross validation Dataset subset: NABench tRNA assays (NABench split) | 0.47 spearman correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceScored supervised, random cross validation across the NABench tRNA assays. Aggregation: Not reported NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 10, row(N.T.-v2-50m), column(tRNA) |
| Configuration: N.T.-v2-50m | Task: NABench SELEX-FS: Overall fitness prediction on NABench SELEX assays, Few-shot Spearman ρ Dataset subset: NABench SELEX assays (NABench split) | 0.578 spearman correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceAggregated by the NABench authors across NABench SELEX assays. Aggregation: Not reported NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 12, row(N.T.-v2-50m), column(Few-shot Spearman ρ) |
Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.
Related profile: Nucleotide Transformer. This page retains the exact record and its evaluation context.
Nucleotide foundation model evaluated by the NABench authors under their fitness prediction protocol.
Nucleotide Transformer is a family of DNA encoders pretrained on human or multispecies sequence corpora. Encoder-only transformers with 6-mer tokens; v1 uses learned positional encodings and v2 uses rotary positions and SwiGLU. The documented inputs are DNA sequences tokenized into 6-mers, with single-base handling of N and remainder bases. The output consists of contextual representations used in specified downstream prediction workflows.
NT-v1 human-reference/1000G/multispecies variants and NT-v2 50M/100M/250M/500M. NT-v3 is a separate architecture described elsewhere in the repository. v1: approximately 6kb; v2: 2,048 tokens, approximately 12kb. Exact base count depends on special and ambiguous tokens.
Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.
Stable record: discovery-model-nucleotide-transformerExplanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Model type | DNA transformer encoder familySources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training |
| Architecture | Encoder-only transformers with 6-mer tokens; v1 uses learned positional encodings and v2 uses rotary positions and SwiGLU.Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training |
| Inputs | DNA sequences tokenized into 6-mers, with single-base handling of N and remainder bases.Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training |
| Outputs | Contextual representations used in specified downstream prediction workflows.Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training |
| Parameters | 50M to 2.5B across the documented v1/v2 family.Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training |
| Known versions | NT-v1 human-reference/1000G/multispecies variants and NT-v2 50M/100M/250M/500M. NT-v3 is a separate architecture described elsewhere in the repository.Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training |
| Training data | v1 variants use GRCh38, 3,202 human genomes or 850 multispecies genomes; v2 uses the multispecies corpus.Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training |
| Training cutoff | The paper specifies human-reference, 1000 Genomes and multispecies training collections by variant. A single latest-deposition date for all sequences is not supplied in the inspected pretraining-data section. · Not reported in inspected sourcesSources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training |
| Context limits | v1: approximately 6kb; v2: 2,048 tokens, approximately 12kb. Exact base count depends on special and ambiguous tokens.Sources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training |
| Weights licence | Separate checkpoint-distribution terms are not stated in the inspected release documentation and licence material. The source-code licence alone is not recorded as an explicit weight grant. · Not reported in inspected sourcesSources (6)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML; instadeepai/nucleotide-transformer: LICENSE.md · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training; LICENSE.md: licence text |
| Access | Official project documentation and implementation: https://github.com/instadeepai/nucleotide-transformerSources (5)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · docs/nucleotide_transformer.md: Model Variants and Sizes, Tokenization and How to use; Nucleotide Transformer paper Methods: Architecture, Pre-training datasets and Training |
| Code licence | CC-BY-NC-SA-4.0Sourcesinstadeepai/nucleotide-transformer: LICENSE.md · LICENSE.md: licence text |
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
3 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Relationship: family catalog-model-nt-v2 Individual claims | NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction GENEB Table 8 NT-v2-50M-MS; NABench model inventory N.T.v2 50M; official v2-50m-multi-species model card | Existing reviewed locator: Table 7, row(N.T.-v2-50m) Version: Primary full-text snapshot retrieved 2026-09-17; exact bytes pinned by SHA-256 | source checked automated source review · 2026-09-23 Audit detailsSource review establishes this relationship only. Exact evaluated configurations and original numerical review status remain unchanged. Same named size/version checkpoint family supported, but exact revision remains unreported. Do not equate weights or adaptations. Field: Claim: model-evaluation-identity-69eae16b949d932fd882 Source artifact SHA-256: Hash scope: Exact retrieved primary paper artifact bytes. |
| Relationship: family discovery-model-nucleotide-transformer Individual claims | NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction mRNABench Table 2/Appendix model inventory; GenomeOcean Table 2; respective model methods; NABench model inventory; GENEB Table 8; source-labelled configuration N.T.-v2-50m | Existing reviewed locator: Table 7, row(N.T.-v2-50m) Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: Primary full-text snapshot retrieved 2026-09-17; exact bytes pinned by SHA-256 | source checked automated source review · 2026-09-23 Audit detailsSource review establishes this relationship only. Exact evaluated configurations and original numerical review status remain unchanged. Only family identity is shared. Preserve size, v1/v2, corpus and adaptation distinctions; generic NT summary selects a best-overall family variant, without supplying checkpoint identity. Field: Claim: model-evaluation-identity-317f83f1c64bff1d891e Source artifact SHA-256: Hash scope: Exact retrieved primary paper artifact bytes. |
| Relationship: family discovery-model-nucleotide-transformer Individual claims | instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md mRNABench Table 2/Appendix model inventory; GenomeOcean Table 2; respective model methods; NABench model inventory; GENEB Table 8; source-labelled configuration N.T.-v2-50m | Existing reviewed locator: Table 7, row(N.T.-v2-50m) Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 2dc37b86e16a6970fbc731751f7719d9f676f7f9 | source checked automated source review · 2026-09-23 Audit detailsSource review establishes this relationship only. Exact evaluated configurations and original numerical review status remain unchanged. Only family identity is shared. Preserve size, v1/v2, corpus and adaptation distinctions; generic NT summary selects a best-overall family variant, without supplying checkpoint identity. Field: Claim: model-evaluation-identity-317f83f1c64bff1d891e Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
View linked audit checks and correction history
Release 2026-09-29-06401fd5b220 · Record review: source checked
Stable ID: nabench-method-n-t-v2-50m