rewire.itbenchmarks
Model

SegmentNT

SegmentNT labels genomic elements at individual nucleotide positions using a pretrained DNA backbone.

Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata

98 evaluations · 196 results

How it worksSegmentNT workflow
SegmentNT workflow1. DNA sequence. Then: 2. NT backbone with position rescaling. Then: 3. 1D U-Net head. Then: 4. Per-base element predictionsSegmentNT workflow1. DNA sequence. Then: 2. NT backbone with position rescaling. Then: 3. 1D U-Net head. Then: 4. Per-base element predictionsSegmentNT workflow1. DNA sequence. Then: 2. NT backbone with position rescaling. Then: 3. 1D U-Net head. Then: 4. Per-base element predictions

Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.

Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata

Overview

Model type

DNA encoder with nucleotide-level segmentation head

Inputs

DNA sequences without N bases, tokenized into 6-mers under the documented length constraints.

Outputs

Per-base probabilities for 14 genomic-element classes.

Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata

Source reviewed · Automated source review, 2026-09-16. All specifications and missing details

Evaluations and results

98 evaluations · 196 results. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: SegmentNT-10kbProtocol: SegmentNT human genome annotation 3UTR auPRC: 3UTR: per-nucleotide annotation (auPRC)
Dataset subset: Human genome 3UTR test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.64 (± 0.008) auprc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.008

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

SegmentNT-10kb on SegmentNT human genome annotation 3UTR auPRC: 3UTR: per-nucleotide annotation (auPRC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 1; data row 15; model SegmentNT-10kb; column 3UTR
Configuration: SegmentNT-10kbProtocol: SegmentNT human genome annotation 3UTR MCC: 3UTR: per-nucleotide annotation (MCC)
Dataset subset: Human genome 3UTR test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.66 (± 0.007) mcc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.007

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

SegmentNT-10kb on SegmentNT human genome annotation 3UTR MCC: 3UTR: per-nucleotide annotation (MCC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 1; data row 15; model SegmentNT-10kb; column 3UTR
Configuration: SegmentNT-10kbProtocol: SegmentNT human genome annotation 5UTR auPRC: 5UTR: per-nucleotide annotation (auPRC)
Dataset subset: Human genome 5UTR test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.43 (± 0.009) auprc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.009

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

SegmentNT-10kb on SegmentNT human genome annotation 5UTR auPRC: 5UTR: per-nucleotide annotation (auPRC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 1; data row 15; model SegmentNT-10kb; column 5UTR
Configuration: SegmentNT-10kbProtocol: SegmentNT human genome annotation 5UTR MCC: 5UTR: per-nucleotide annotation (MCC)
Dataset subset: Human genome 5UTR test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.49 (± 0.007) mcc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.007

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

SegmentNT-10kb on SegmentNT human genome annotation 5UTR MCC: 5UTR: per-nucleotide annotation (MCC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 1; data row 15; model SegmentNT-10kb; column 5UTR
Configuration: SegmentNT-10kbProtocol: SegmentNT human genome annotation CTCF-bound auPRC: CTCF-bound: per-nucleotide annotation (auPRC)
Dataset subset: Human genome CTCF-bound test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.04 (± 0.001) auprc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.001

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

SegmentNT-10kb on SegmentNT human genome annotation CTCF-bound auPRC: CTCF-bound: per-nucleotide annotation (auPRC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 1; data row 15; model SegmentNT-10kb; column CTCF-bound
Configuration: SegmentNT-10kbProtocol: SegmentNT human genome annotation CTCF-bound MCC: CTCF-bound: per-nucleotide annotation (MCC)
Dataset subset: Human genome CTCF-bound test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.09 (± 0.003) mcc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.003

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

SegmentNT-10kb on SegmentNT human genome annotation CTCF-bound MCC: CTCF-bound: per-nucleotide annotation (MCC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 1; data row 15; model SegmentNT-10kb; column CTCF-bound
Configuration: SegmentNT-10kbProtocol: SegmentNT human genome annotation enhancer tissue-invariant auPRC: enhancer tissue-invariant: per-nucleotide annotation (auPRC)
Dataset subset: Human genome enhancer tissue-invariant test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.10 (± 0.004) auprc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.004

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

SegmentNT-10kb on SegmentNT human genome annotation enhancer tissue-invariant auPRC: enhancer tissue-invariant: per-nucleotide annotation (auPRC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 1; data row 15; model SegmentNT-10kb; column enhancer tissue-invariant
Configuration: SegmentNT-10kbProtocol: SegmentNT human genome annotation enhancer tissue-invariant MCC: enhancer tissue-invariant: per-nucleotide annotation (MCC)
Dataset subset: Human genome enhancer tissue-invariant test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.19 (± 0.007) mcc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.007

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

SegmentNT-10kb on SegmentNT human genome annotation enhancer tissue-invariant MCC: enhancer tissue-invariant: per-nucleotide annotation (MCC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 1; data row 15; model SegmentNT-10kb; column enhancer tissue-invariant
Configuration: SegmentNT-10kbProtocol: SegmentNT human genome annotation enhancer tissue-specific auPRC: enhancer tissue-specific: per-nucleotide annotation (auPRC)
Dataset subset: Human genome enhancer tissue-specific test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.31 (± 0.001) auprc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.001

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

SegmentNT-10kb on SegmentNT human genome annotation enhancer tissue-specific auPRC: enhancer tissue-specific: per-nucleotide annotation (auPRC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 1; data row 15; model SegmentNT-10kb; column enhancer tissue-specific
Configuration: SegmentNT-10kbProtocol: SegmentNT human genome annotation enhancer tissue-specific MCC: enhancer tissue-specific: per-nucleotide annotation (MCC)
Dataset subset: Human genome enhancer tissue-specific test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.27 (± 0.001) mcc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.001

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

SegmentNT-10kb on SegmentNT human genome annotation enhancer tissue-specific MCC: enhancer tissue-specific: per-nucleotide annotation (MCC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 1; data row 15; model SegmentNT-10kb; column enhancer tissue-specific
Configuration: SegmentNT-10kbProtocol: SegmentNT human genome annotation exon auPRC: exon: per-nucleotide annotation (auPRC)
Dataset subset: Human genome exon test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.56 (± 0.003) auprc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.003

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

SegmentNT-10kb on SegmentNT human genome annotation exon auPRC: exon: per-nucleotide annotation (auPRC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 1; data row 15; model SegmentNT-10kb; column exon
Configuration: SegmentNT-10kbProtocol: SegmentNT human genome annotation exon MCC: exon: per-nucleotide annotation (MCC)
Dataset subset: Human genome exon test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.59 (± 0.003) mcc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.003

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

SegmentNT-10kb on SegmentNT human genome annotation exon MCC: exon: per-nucleotide annotation (MCC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 1; data row 15; model SegmentNT-10kb; column exon
Configuration: SegmentNT-10kbProtocol: SegmentNT human genome annotation intron auPRC: intron: per-nucleotide annotation (auPRC)
Dataset subset: Human genome intron test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.79 (± 0.003) auprc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.003

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

SegmentNT-10kb on SegmentNT human genome annotation intron auPRC: intron: per-nucleotide annotation (auPRC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 1; data row 15; model SegmentNT-10kb; column intron
Configuration: SegmentNT-10kbProtocol: SegmentNT human genome annotation intron MCC: intron: per-nucleotide annotation (MCC)
Dataset subset: Human genome intron test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.40 (± 0.004) mcc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.004

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

SegmentNT-10kb on SegmentNT human genome annotation intron MCC: intron: per-nucleotide annotation (MCC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 1; data row 15; model SegmentNT-10kb; column intron
Configuration: SegmentNT-10kbProtocol: SegmentNT human genome annotation lncRNA auPRC: lncRNA: per-nucleotide annotation (auPRC)
Dataset subset: Human genome lncRNA test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.23 (± 0.007) auprc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.007

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

SegmentNT-10kb on SegmentNT human genome annotation lncRNA auPRC: lncRNA: per-nucleotide annotation (auPRC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 2; data row 15; model SegmentNT-10kb; column lncRNA
Configuration: SegmentNT-10kbProtocol: SegmentNT human genome annotation lncRNA MCC: lncRNA: per-nucleotide annotation (MCC)
Dataset subset: Human genome lncRNA test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.06 (± 0.013) mcc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.013

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

SegmentNT-10kb on SegmentNT human genome annotation lncRNA MCC: lncRNA: per-nucleotide annotation (MCC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 2; data row 15; model SegmentNT-10kb; column lncRNA
Configuration: SegmentNT-10kbProtocol: SegmentNT human genome annotation polyA signal auPRC: polyA signal: per-nucleotide annotation (auPRC)
Dataset subset: Human genome polyA signal test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.28 (± 0.006) auprc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.006

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

SegmentNT-10kb on SegmentNT human genome annotation polyA signal auPRC: polyA signal: per-nucleotide annotation (auPRC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 2; data row 15; model SegmentNT-10kb; column polyA signal
Configuration: SegmentNT-10kbProtocol: SegmentNT human genome annotation polyA signal MCC: polyA signal: per-nucleotide annotation (MCC)
Dataset subset: Human genome polyA signal test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.36 (± 0.007) mcc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.007

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

SegmentNT-10kb on SegmentNT human genome annotation polyA signal MCC: polyA signal: per-nucleotide annotation (MCC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 2; data row 15; model SegmentNT-10kb; column polyA signal
Configuration: SegmentNT-10kbProtocol: SegmentNT human genome annotation promoter tissue-invariant auPRC: promoter tissue-invariant: per-nucleotide annotation (auPRC)
Dataset subset: Human genome promoter tissue-invariant test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.57 (± 0.012) auprc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.012

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

SegmentNT-10kb on SegmentNT human genome annotation promoter tissue-invariant auPRC: promoter tissue-invariant: per-nucleotide annotation (auPRC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 2; data row 15; model SegmentNT-10kb; column promoter tissue-invariant
Configuration: SegmentNT-10kbProtocol: SegmentNT human genome annotation promoter tissue-invariant MCC: promoter tissue-invariant: per-nucleotide annotation (MCC)
Dataset subset: Human genome promoter tissue-invariant test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.58 (± 0.009) mcc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.009

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

SegmentNT-10kb on SegmentNT human genome annotation promoter tissue-invariant MCC: promoter tissue-invariant: per-nucleotide annotation (MCC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 2; data row 15; model SegmentNT-10kb; column promoter tissue-invariant
Configuration: SegmentNT-10kbProtocol: SegmentNT human genome annotation promoter tissue-specific auPRC: promoter tissue-specific: per-nucleotide annotation (auPRC)
Dataset subset: Human genome promoter tissue-specific test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.15 (± 0.007) auprc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.007

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

SegmentNT-10kb on SegmentNT human genome annotation promoter tissue-specific auPRC: promoter tissue-specific: per-nucleotide annotation (auPRC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 2; data row 15; model SegmentNT-10kb; column promoter tissue-specific
Configuration: SegmentNT-10kbProtocol: SegmentNT human genome annotation promoter tissue-specific MCC: promoter tissue-specific: per-nucleotide annotation (MCC)
Dataset subset: Human genome promoter tissue-specific test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.23 (± 0.009) mcc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.009

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

SegmentNT-10kb on SegmentNT human genome annotation promoter tissue-specific MCC: promoter tissue-specific: per-nucleotide annotation (MCC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 2; data row 15; model SegmentNT-10kb; column promoter tissue-specific
Configuration: SegmentNT-10kbProtocol: SegmentNT human genome annotation protein coding gene auPRC: protein coding gene: per-nucleotide annotation (auPRC)
Dataset subset: Human genome protein coding gene test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.83 (± 0.004) auprc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.004

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

SegmentNT-10kb on SegmentNT human genome annotation protein coding gene auPRC: protein coding gene: per-nucleotide annotation (auPRC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 2; data row 15; model SegmentNT-10kb; column protein coding gene
Configuration: SegmentNT-10kbProtocol: SegmentNT human genome annotation protein coding gene MCC: protein coding gene: per-nucleotide annotation (MCC)
Dataset subset: Human genome protein coding gene test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.58 (± 0.008) mcc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.008

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

SegmentNT-10kb on SegmentNT human genome annotation protein coding gene MCC: protein coding gene: per-nucleotide annotation (MCC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 2; data row 15; model SegmentNT-10kb; column protein coding gene
Configuration: SegmentNT-10kbProtocol: SegmentNT human genome annotation splice acceptor auPRC: splice acceptor: per-nucleotide annotation (auPRC)
Dataset subset: Human genome splice acceptor test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.65 (± 0.006) auprc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.006

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

SegmentNT-10kb on SegmentNT human genome annotation splice acceptor auPRC: splice acceptor: per-nucleotide annotation (auPRC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 2; data row 15; model SegmentNT-10kb; column splice acceptor

Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.

Use this model

How it works, versions and access

Versions and evaluated configurations

How it works

How it works

SegmentNT labels genomic elements at individual nucleotide positions using a pretrained DNA backbone. Nucleotide Transformer backbone with a one-dimensional U-Net segmentation head; YaRN rescales positions for longer inputs. The documented inputs are DNA sequences without N bases, tokenized into 6-mers under the documented length constraints. The output consists of per-base probabilities for 14 genomic-element classes.

Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata
Versions and reproducibility

segment_nt and segment_nt_multi_species; SegmentEnformer and SegmentBorzoi are distinct pipelines. Trained on 30kb; inference up to 50kb requires the documented rescaling. Input token count also has divisibility constraints.

Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata
Strengths, limitations and unresolved questions

Strengths and limitations

Strengths supported by sources

Limitations and conditions

Profile review details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Stable record: discovery-model-segmentnt

Specifications

Inputs, training, access and other details

Explanatory profile: source reviewed · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Inputs, outputs and configuration
PropertyDescription and evidence
Model typeDNA encoder with nucleotide-level segmentation head
Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata
ArchitectureNucleotide Transformer backbone with a one-dimensional U-Net segmentation head; YaRN rescales positions for longer inputs.
Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata
InputsDNA sequences without N bases, tokenized into 6-mers under the documented length constraints.
Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata
OutputsPer-base probabilities for 14 genomic-element classes.
Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata
Parameters563M for NT-v2 500M plus the 63M segmentation head; alternative encoder ablations are different complete pipelines.
Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata
Known versionssegment_nt and segment_nt_multi_species; SegmentEnformer and SegmentBorzoi are distinct pipelines.
Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata
Training dataHuman 14-element segmentation labels from GENCODE V44 and ENCODE SCREEN/DHS annotations. Human chromosomes 20 and 21 are held out for testing and 22 for validation. A separate multispecies model adds mouse, chicken, fly, zebrafish and worm.
Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata
Training cutoffGENCODE V44 and the named ENCODE SCREEN/DHS resources define label provenance. The inspected paper does not give one latest-experiment date shared by every resource.
Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata
Context limitsTrained on 30kb; inference up to 50kb requires the documented rescaling. Input token count also has divisibility constraints.
Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata
Weights licenceCC-BY-NC-SA-4.0 declared by the official InstaDeepAI/segment_nt model card.
Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata
AccessOfficial project documentation and implementation: https://github.com/instadeepai/nucleotide-transformer
Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata
Code licenceCC-BY-NC-SA-4.0
Sourcesinstadeepai/nucleotide-transformer: LICENSE.md · LICENSE.md: licence text

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

134 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.
Individual claims
instadeepai/nucleotide-transformer: README.md

Original source ↗

SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2dc37b86e16a6970fbc731751f7719d9f676f7f9
Retrieved: 2026-09-16T19:46:19.364532+00:00

source checked

automated source review · 2026-09-16

Audit details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 9f51bbb20c4c5c36e77fb03ca1c5c36236e287c48a1ee31f53150545d421ec25

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram caption
Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.
Individual claims
instadeepai/nucleotide-transformer: docs/segment_nt.md

Original source ↗

SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2dc37b86e16a6970fbc731751f7719d9f676f7f9
Retrieved: 2026-09-16T19:46:19.364532+00:00

source checked

automated source review · 2026-09-16

Audit details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 8eec4580ba64ab944fb9b42674be70ffe793135f3503cce8640f5b08f8290f7a

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram caption
Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.
Individual claims
instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md

Original source ↗

SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2dc37b86e16a6970fbc731751f7719d9f676f7f9
Retrieved: 2026-09-16T19:46:19.364532+00:00

source checked

automated source review · 2026-09-16

Audit details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 23e27d40473bbabadac45c56e8e282349f0b053da85fda89ff2125b5fe381fc6

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram caption
Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.
Individual claims
instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md

Original source ↗

SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2dc37b86e16a6970fbc731751f7719d9f676f7f9
Retrieved: 2026-09-16T19:46:19.364532+00:00

source checked

automated source review · 2026-09-16

Audit details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: ab16d582de98652526b5cebb120eec969328f9db29dc741826bcd81c397e0672

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram caption
Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.
Individual claims
InstaDeepAI/segment_nt: README.md

Original source ↗

SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 1048ad869036bb8e50a15d82b323ebeb52e8ab92
Retrieved: 2026-09-16T20:12:20.913226+00:00

source checked

automated source review · 2026-09-16

Audit details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 6e19e9c9293b0c392c053f92360b674b86c86c4457608b00c2fcb962bbd1d1bc

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram caption
Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.
Individual claims
InstaDeepAI/segment_nt: config.json

Original source ↗

SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 1048ad869036bb8e50a15d82b323ebeb52e8ab92
Retrieved: 2026-09-16T20:12:20.913226+00:00

source checked

automated source review · 2026-09-16

Audit details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 1412a48b99cfe79e6f34f84131ed964fd3ad5c4f3b3ea7db7ba2b0ac3466eadc

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram caption
Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.
Individual claims
segmentnt: Journal full-text XML

Original source ↗

SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: Retrieved page snapshot; no immutable publisher revision supplied
Retrieved: 2026-09-16T20:16:14.422701+00:00

source checked

automated source review · 2026-09-16

Audit details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: b9bddff89c5f8898fc677008205a47b960ea569f3887cc15845e2547b92aaa70

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram steps
  • DNA sequence
  • NT backbone with position rescaling
  • 1D U-Net head
  • Per-base element predictions
Individual claims
instadeepai/nucleotide-transformer: README.md

Original source ↗

SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2dc37b86e16a6970fbc731751f7719d9f676f7f9
Retrieved: 2026-09-16T19:46:19.364532+00:00

source checked

automated source review · 2026-09-16

Audit details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 9f51bbb20c4c5c36e77fb03ca1c5c36236e287c48a1ee31f53150545d421ec25

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram steps
  • DNA sequence
  • NT backbone with position rescaling
  • 1D U-Net head
  • Per-base element predictions
Individual claims
instadeepai/nucleotide-transformer: docs/segment_nt.md

Original source ↗

SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2dc37b86e16a6970fbc731751f7719d9f676f7f9
Retrieved: 2026-09-16T19:46:19.364532+00:00

source checked

automated source review · 2026-09-16

Audit details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 8eec4580ba64ab944fb9b42674be70ffe793135f3503cce8640f5b08f8290f7a

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram steps
  • DNA sequence
  • NT backbone with position rescaling
  • 1D U-Net head
  • Per-base element predictions
Individual claims
instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md

Original source ↗

SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2dc37b86e16a6970fbc731751f7719d9f676f7f9
Retrieved: 2026-09-16T19:46:19.364532+00:00

source checked

automated source review · 2026-09-16

Audit details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 23e27d40473bbabadac45c56e8e282349f0b053da85fda89ff2125b5fe381fc6

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: discovered

9 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: discovery-model-segmentnt

areas
genomics
access
official_source_linked
benchmark applicability
candidate; not evidence of a reported evaluation
candidate benchmark ids
None recorded
entity level
family
reported name
SegmentNT
version
Not reported
historical missing metadata
checkpoint: unextracted; code licence: unextracted; parameters: unextracted; training cutoff: unextracted; training data: unextracted; version: unextracted; weights licence: unextracted
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
entity classification
review date: 2026-09-17; rationale: The cited profile describes a named learned biological predictor or representation model/family. Preserve this identity separately from task-specific fitting, individual checkpoints, pipelines and hosted access.; source ids: evidence-official-18d8f4d5fc3f6616922d; evidence-official-657e83427ab59f3aec83; evidence-official-56f02d45976d011d80aa; evidence-official-4920952f9b3c4b8909a0; evidence-official-7e59c59bfea79722fc33; evidence-official-97071500fc3422c426d4; evidence-official-da4566a88ed9fa42fdcb; source locator: SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata; ambiguities: None recorded
Related records

Suggest a correction