rewire.itbenchmarks
Dataset

GenomeOcean natural/artificial sequence test

Research readiness

These checks assess whether the evidence supports a reproducible investigation. A source-checked score alone does not meet these requirements.

Release 2026-09-29-06401fd5b220 · Evidence verified: Not verified

Evidence incomplete

Replay metrics

Exact outcomes, predictions, identifiers and evaluator are connected.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • metric replay: verification is missing

Verified: Not verified

Evidence incomplete

Investigate discrepancies

Replay evidence includes annotations and an assessment of dependence. Unknown independence permits descriptive analysis only.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • metric replay: verification is missing
  • annotations: verification is missing
  • dependence: verification is missing

Verified: Not verified

Evidence incomplete

Run locally

A pinned recipe describes the inputs, environment and resource requirements.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • recipe pinned: verification is missing
  • resource estimate: verification is missing

Verified: Not verified

Evidence incomplete

Validate independently

Separate data and exposure records support an independent test.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • independent validation: verification is missing
  • overlap checked: verification is missing

Verified: Not verified

Readiness describes the evidence in this release. Availability on your computer is checked separately when an investigation runs. Existing data exposure can prevent independent validation even when files are available.

Artifacts and reproduction

No verified artifact manifest is connected to this record yet. The gaps above identify what is needed before analysis can begin.

Read reviewed discrepancy investigations

Evaluation results

3 evaluations · 9 results. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: GenomeOceanProtocol: Natural versus GenomeOcean-generated DNA classification (Natural vs artificial microbial genome sequence)
Dataset: GenomeOcean natural/artificial sequence test
99 F1
% · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GenomeOcean: Natural vs artificial microbial genome sequence

DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.

Aggregation: Not reported

GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies; GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2, GenomeOcean row, F1 column
Configuration: DNABERT-2Protocol: Natural versus GenomeOcean-generated DNA classification (Natural vs artificial microbial genome sequence)
Dataset: GenomeOcean natural/artificial sequence test
85.1 F1
% · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

DNABERT-2: Natural vs artificial microbial genome sequence

DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.

Aggregation: Not reported

GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies; GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2, DNABERT-2 row, F1 column
Configuration: Nucleotide Transformers V2Protocol: Natural versus GenomeOcean-generated DNA classification (Natural vs artificial microbial genome sequence)
Dataset: GenomeOcean natural/artificial sequence test
83.1 F1
% · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

Nucleotide Transformers V2: Natural versus GenomeOcean-generated DNA classification

DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.

Aggregation: Not reported

GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row Nucleotide Transformers V2, column F1; XML row3 column4
Configuration: DNABERT-2Protocol: Natural versus GenomeOcean-generated DNA classification (Natural vs artificial microbial genome sequence)
Dataset: GenomeOcean natural/artificial sequence test
85% Recall
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

DNABERT-2: Natural vs artificial microbial genome sequence

DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.

Aggregation: Not reported

GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row DNABERT-2, column Recall; XML row2 column3
Configuration: GenomeOceanProtocol: Natural versus GenomeOcean-generated DNA classification (Natural vs artificial microbial genome sequence)
Dataset: GenomeOcean natural/artificial sequence test
99% Precision
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GenomeOcean: Natural vs artificial microbial genome sequence

DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.

Aggregation: Not reported

GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row GenomeOcean, column Precision; XML row4 column2
Configuration: Nucleotide Transformers V2Protocol: Natural versus GenomeOcean-generated DNA classification (Natural vs artificial microbial genome sequence)
Dataset: GenomeOcean natural/artificial sequence test
83% Recall
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

Nucleotide Transformers V2: Natural versus GenomeOcean-generated DNA classification

DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.

Aggregation: Not reported

GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row Nucleotide Transformers V2, column Recall; XML row3 column3
Configuration: DNABERT-2Protocol: Natural versus GenomeOcean-generated DNA classification (Natural vs artificial microbial genome sequence)
Dataset: GenomeOcean natural/artificial sequence test
85.2% Precision
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

DNABERT-2: Natural vs artificial microbial genome sequence

DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.

Aggregation: Not reported

GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row DNABERT-2, column Precision; XML row2 column2
Configuration: GenomeOceanProtocol: Natural versus GenomeOcean-generated DNA classification (Natural vs artificial microbial genome sequence)
Dataset: GenomeOcean natural/artificial sequence test
99% Recall
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

GenomeOcean: Natural vs artificial microbial genome sequence

DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.

Aggregation: Not reported

GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row GenomeOcean, column Recall; XML row4 column3
Configuration: Nucleotide Transformers V2Protocol: Natural versus GenomeOcean-generated DNA classification (Natural vs artificial microbial genome sequence)
Dataset: GenomeOcean natural/artificial sequence test
83.3% Precision
percent · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

Nucleotide Transformers V2: Natural versus GenomeOcean-generated DNA classification

DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.

Aggregation: Not reported

GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row Nucleotide Transformers V2, column Precision; XML row3 column2

Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.

Dataset and evaluation context

A dataset supplies biological observations. The evaluation protocol defines how those observations are split, used and scored.

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

4 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
attributes.split
Not reported
Context-only references
GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies

Original source ↗

No field-specific location recorded

Version: preprint archived 2025-02-05
Retrieved: 2026-09-16T10:33:55.224Z

missing or unspecified

No individual claim review recorded

Audit details

Field: attributes.split

Source artifact SHA-256: 3cc0df52522fccda23e3958f069c916b87ee50bb5c9a992fa37e25256546e145

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

attributes.version
Not reported
Context-only references
GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies

Original source ↗

No field-specific location recorded

Version: preprint archived 2025-02-05
Retrieved: 2026-09-16T10:33:55.224Z

missing or unspecified

No individual claim review recorded

Audit details

Field: attributes.version

Source artifact SHA-256: 3cc0df52522fccda23e3958f069c916b87ee50bb5c9a992fa37e25256546e145

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

description
No value recorded
Context-only references
GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies

Original source ↗

No field-specific location recorded

Version: preprint archived 2025-02-05
Retrieved: 2026-09-16T10:33:55.224Z

missing or unspecified

No individual claim review recorded

Audit details

Field: description

Source artifact SHA-256: 3cc0df52522fccda23e3958f069c916b87ee50bb5c9a992fa37e25256546e145

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

name
GenomeOcean natural/artificial sequence test
Context-only references
GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies

Original source ↗

No field-specific location recorded

Version: preprint archived 2025-02-05
Retrieved: 2026-09-16T10:33:55.224Z

not individually reviewed

No individual claim review recorded

Audit details

Field: name

Source artifact SHA-256: 3cc0df52522fccda23e3958f069c916b87ee50bb5c9a992fa37e25256546e145

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: needs review

1 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: reported-dataset-0c3ac7efe99c37

areas
microbes-communities
version
Not reported
split
Not reported
missing metadata
version: not_reported_in_legacy_extract; split: not_reported_in_legacy_extract; accession: not_reported_in_legacy_extract
entity classification
review date: 2026-09-17; rationale: This record identifies a biological data collection or source-labelled evaluation cohort. Keep its dataset identity; split, assay, taxonomic level, candidate restrictions and comparison context remain attributes rather than automatically becoming new entity kinds.; source ids: genomeocean-2025; source locator: Methods 4.2.5 Generated Sequence Discrimination; Table 2; ambiguities: None recorded
Related records

Suggest a correction