Model type
Autoregressive metagenomic transformer
METAGENE-1 is an autoregressive DNA/RNA sequence model trained on wastewater metagenomic data.
Results are available for configurations using this model. Their fitted heads, extra inputs and evaluation settings are kept separate below.
Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.
Autoregressive metagenomic transformer
DNA or RNA nucleotide sequences.
Sequence generation and representations for downstream metagenomic analyses.
Official downloadable model/card and usage examples: https://huggingface.co/metagene-ai/METAGENE-1
limited source coverage · Automated source review, 2026-09-16. All specifications and missing details
These configurations, services and pipelines use this model within their own configurations. Their results, where available, are not assigned to the underlying model.
METAGENE-1 is an autoregressive DNA/RNA sequence model trained on wastewater metagenomic data. Llama-style autoregressive transformer with 32 layers, width 4,096, 32 attention heads and a 1,024-token vocabulary in the inspected configuration. The documented inputs are DNA or RNA nucleotide sequences. The output consists of sequence generation and representations for downstream metagenomic analyses.
METAGENE-1; exact model-card and configuration revision pinned in sources. Configuration fields differ: max_position_embeddings=512 and max_sequence_length=2,048. The effective supported window remains unresolved; neither value is silently promoted to a validated inference limit.
Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.
Stable record: catalog-model-metagene-1Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Model type | Autoregressive metagenomic transformerSources (4)metagene-ai/METAGENE-1: README.md; metagene-ai/METAGENE-1: config.json; metagene-ai/metagene-pretrain: README.md; metagene-ai/metagene-pretrain: train/config_hub/pretrain/genomicsllama.yml · README.md: Model Overview and Usage; config.json; config.json: model_type, hidden_size, num_hidden_layers, num_attention_heads, vocab_size and sequence-length fields; official pretraining train/config_hub/pretrain/genomicsllama.yml: train.max_seq_length |
| Architecture | Llama-style autoregressive transformer with 32 layers, width 4,096, 32 attention heads and a 1,024-token vocabulary in the inspected configuration.Sources (4)metagene-ai/METAGENE-1: README.md; metagene-ai/METAGENE-1: config.json; metagene-ai/metagene-pretrain: README.md; metagene-ai/metagene-pretrain: train/config_hub/pretrain/genomicsllama.yml · README.md: Model Overview and Usage; config.json; config.json: model_type, hidden_size, num_hidden_layers, num_attention_heads, vocab_size and sequence-length fields; official pretraining train/config_hub/pretrain/genomicsllama.yml: train.max_seq_length |
| Inputs | DNA or RNA nucleotide sequences.Sources (4)metagene-ai/METAGENE-1: README.md; metagene-ai/METAGENE-1: config.json; metagene-ai/metagene-pretrain: README.md; metagene-ai/metagene-pretrain: train/config_hub/pretrain/genomicsllama.yml · README.md: Model Overview and Usage; config.json; config.json: model_type, hidden_size, num_hidden_layers, num_attention_heads, vocab_size and sequence-length fields; official pretraining train/config_hub/pretrain/genomicsllama.yml: train.max_seq_length |
| Outputs | Sequence generation and representations for downstream metagenomic analyses.Sources (4)metagene-ai/METAGENE-1: README.md; metagene-ai/METAGENE-1: config.json; metagene-ai/metagene-pretrain: README.md; metagene-ai/metagene-pretrain: train/config_hub/pretrain/genomicsllama.yml · README.md: Model Overview and Usage; config.json; config.json: model_type, hidden_size, num_hidden_layers, num_attention_heads, vocab_size and sequence-length fields; official pretraining train/config_hub/pretrain/genomicsllama.yml: train.max_seq_length |
| Parameters | 7 billion.Sources (4)metagene-ai/METAGENE-1: README.md; metagene-ai/METAGENE-1: config.json; metagene-ai/metagene-pretrain: README.md; metagene-ai/metagene-pretrain: train/config_hub/pretrain/genomicsllama.yml · README.md: Model Overview and Usage; config.json; config.json: model_type, hidden_size, num_hidden_layers, num_attention_heads, vocab_size and sequence-length fields; official pretraining train/config_hub/pretrain/genomicsllama.yml: train.max_seq_length |
| Known versions | METAGENE-1; exact model-card and configuration revision pinned in sources.Sources (4)metagene-ai/METAGENE-1: README.md; metagene-ai/METAGENE-1: config.json; metagene-ai/metagene-pretrain: README.md; metagene-ai/metagene-pretrain: train/config_hub/pretrain/genomicsllama.yml · README.md: Model Overview and Usage; config.json; config.json: model_type, hidden_size, num_hidden_layers, num_attention_heads, vocab_size and sequence-length fields; official pretraining train/config_hub/pretrain/genomicsllama.yml: train.max_seq_length |
| Training data | More than 1.5 trillion base pairs sequenced from human wastewater samples, according to the model card.Sources (4)metagene-ai/METAGENE-1: README.md; metagene-ai/METAGENE-1: config.json; metagene-ai/metagene-pretrain: README.md; metagene-ai/metagene-pretrain: train/config_hub/pretrain/genomicsllama.yml · README.md: Model Overview and Usage; config.json; config.json: model_type, hidden_size, num_hidden_layers, num_attention_heads, vocab_size and sequence-length fields; official pretraining train/config_hub/pretrain/genomicsllama.yml: train.max_seq_length |
| Training cutoff | The official pretraining README states that the wastewater corpus is not yet publicly released; a latest sample-collection date is not provided there. · Not reported in inspected sourcesSources (4)metagene-ai/METAGENE-1: README.md; metagene-ai/METAGENE-1: config.json; metagene-ai/metagene-pretrain: README.md; metagene-ai/metagene-pretrain: train/config_hub/pretrain/genomicsllama.yml · README.md: Model Overview and Usage; config.json; config.json: model_type, hidden_size, num_hidden_layers, num_attention_heads, vocab_size and sequence-length fields; official pretraining train/config_hub/pretrain/genomicsllama.yml: train.max_seq_length |
| Context limits | The released pretraining YAML sets max_seq_length=512. The model config lists max_position_embeddings=512 and max_sequence_length=2,048; these differing configuration fields do not establish a single validated inference maximum.Sources (4)metagene-ai/METAGENE-1: README.md; metagene-ai/METAGENE-1: config.json; metagene-ai/metagene-pretrain: README.md; metagene-ai/metagene-pretrain: train/config_hub/pretrain/genomicsllama.yml · README.md: Model Overview and Usage; config.json; config.json: model_type, hidden_size, num_hidden_layers, num_attention_heads, vocab_size and sequence-length fields; official pretraining train/config_hub/pretrain/genomicsllama.yml: train.max_seq_length |
| Weights licence | Apache-2.0 declared in the model card.Sources (4)metagene-ai/METAGENE-1: README.md; metagene-ai/METAGENE-1: config.json; metagene-ai/metagene-pretrain: README.md; metagene-ai/metagene-pretrain: train/config_hub/pretrain/genomicsllama.yml · README.md: Model Overview and Usage; config.json; config.json: model_type, hidden_size, num_hidden_layers, num_attention_heads, vocab_size and sequence-length fields; official pretraining train/config_hub/pretrain/genomicsllama.yml: train.max_seq_length |
| Access | Official downloadable model/card and usage examples: https://huggingface.co/metagene-ai/METAGENE-1Sources (4)metagene-ai/METAGENE-1: README.md; metagene-ai/METAGENE-1: config.json; metagene-ai/metagene-pretrain: README.md; metagene-ai/metagene-pretrain: train/config_hub/pretrain/genomicsllama.yml · README.md: Model Overview and Usage; config.json; config.json: model_type, hidden_size, num_hidden_layers, num_attention_heads, vocab_size and sequence-length fields; official pretraining train/config_hub/pretrain/genomicsllama.yml: train.max_seq_length |
| Code licence | Apache-2.0Sourcesmetagene-ai/metagene-pretrain: train/LICENSE · train/LICENSE: licence text |
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
77 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation. Individual claims | metagene-ai/METAGENE-1: README.md README.md: Model Overview and Usage; config.json; config.json: model_type, hidden_size, num_hidden_layers, num_attention_heads, vocab_size and sequence-length fields; official pretraining train/config_hub/pretrain/genomicsllama.yml: train.max_seq_length Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: ad8a1e0ee62b85058bfc05d823d8e8d4759edc48 | source checked automated source review · 2026-09-16 Audit detailsInspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied. Field: Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
| Diagram caption Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation. Individual claims | metagene-ai/metagene-pretrain: README.md README.md: Model Overview and Usage; config.json; config.json: model_type, hidden_size, num_hidden_layers, num_attention_heads, vocab_size and sequence-length fields; official pretraining train/config_hub/pretrain/genomicsllama.yml: train.max_seq_length Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 82b9e142db2c7e0a268346d53344d6c3bf223066 | source checked automated source review · 2026-09-16 Audit detailsInspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied. Field: Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
| Diagram caption Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation. Individual claims | metagene-ai/metagene-pretrain: train/config_hub/pretrain/genomicsllama.yml README.md: Model Overview and Usage; config.json; config.json: model_type, hidden_size, num_hidden_layers, num_attention_heads, vocab_size and sequence-length fields; official pretraining train/config_hub/pretrain/genomicsllama.yml: train.max_seq_length Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 82b9e142db2c7e0a268346d53344d6c3bf223066 | source checked automated source review · 2026-09-16 Audit detailsInspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied. Field: Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
| Diagram caption Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation. Individual claims | metagene-ai/METAGENE-1: config.json README.md: Model Overview and Usage; config.json; config.json: model_type, hidden_size, num_hidden_layers, num_attention_heads, vocab_size and sequence-length fields; official pretraining train/config_hub/pretrain/genomicsllama.yml: train.max_seq_length Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: ad8a1e0ee62b85058bfc05d823d8e8d4759edc48 | source checked automated source review · 2026-09-16 Audit detailsInspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied. Field: Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
Diagram steps
| metagene-ai/METAGENE-1: README.md README.md: Model Overview and Usage; config.json; config.json: model_type, hidden_size, num_hidden_layers, num_attention_heads, vocab_size and sequence-length fields; official pretraining train/config_hub/pretrain/genomicsllama.yml: train.max_seq_length Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: ad8a1e0ee62b85058bfc05d823d8e8d4759edc48 | source checked automated source review · 2026-09-16 Audit detailsInspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied. Field: Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
Diagram steps
| metagene-ai/metagene-pretrain: README.md README.md: Model Overview and Usage; config.json; config.json: model_type, hidden_size, num_hidden_layers, num_attention_heads, vocab_size and sequence-length fields; official pretraining train/config_hub/pretrain/genomicsllama.yml: train.max_seq_length Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 82b9e142db2c7e0a268346d53344d6c3bf223066 | source checked automated source review · 2026-09-16 Audit detailsInspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied. Field: Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
Diagram steps
| metagene-ai/metagene-pretrain: train/config_hub/pretrain/genomicsllama.yml README.md: Model Overview and Usage; config.json; config.json: model_type, hidden_size, num_hidden_layers, num_attention_heads, vocab_size and sequence-length fields; official pretraining train/config_hub/pretrain/genomicsllama.yml: train.max_seq_length Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 82b9e142db2c7e0a268346d53344d6c3bf223066 | source checked automated source review · 2026-09-16 Audit detailsInspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied. Field: Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
Diagram steps
| metagene-ai/METAGENE-1: config.json README.md: Model Overview and Usage; config.json; config.json: model_type, hidden_size, num_hidden_layers, num_attention_heads, vocab_size and sequence-length fields; official pretraining train/config_hub/pretrain/genomicsllama.yml: train.max_seq_length Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: ad8a1e0ee62b85058bfc05d823d8e8d4759edc48 | source checked automated source review · 2026-09-16 Audit detailsInspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied. Field: Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
| Diagram title METAGENE-1 workflow Individual claims | metagene-ai/METAGENE-1: README.md README.md: Model Overview and Usage; config.json; config.json: model_type, hidden_size, num_hidden_layers, num_attention_heads, vocab_size and sequence-length fields; official pretraining train/config_hub/pretrain/genomicsllama.yml: train.max_seq_length Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: ad8a1e0ee62b85058bfc05d823d8e8d4759edc48 | source checked automated source review · 2026-09-16 Audit detailsInspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied. Field: Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
| Diagram title METAGENE-1 workflow Individual claims | metagene-ai/metagene-pretrain: README.md README.md: Model Overview and Usage; config.json; config.json: model_type, hidden_size, num_hidden_layers, num_attention_heads, vocab_size and sequence-length fields; official pretraining train/config_hub/pretrain/genomicsllama.yml: train.max_seq_length Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 82b9e142db2c7e0a268346d53344d6c3bf223066 | source checked automated source review · 2026-09-16 Audit detailsInspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied. Field: Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
View linked audit checks and correction history
Release 2026-09-29-06401fd5b220 · Record review: discovered
Stable ID: catalog-model-metagene-1