rewire.itbenchmarks
Protocol

Natural versus GenomeOcean-generated DNA classification (Natural vs artificial microbial genome sequence)

Natural versus GenomeOcean-generated DNA classification · Table 2:. DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.

SourcesGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:: Precision, Natural versus GenomeOcean-generated DNA classification

3 evaluations · 9 results

Overview

Key specifications have not been extracted for this record. See the linked evaluation and sources for the reported setup.

limited source coverage · Automated source review, 2026-09-17. All specifications and missing details

Results

Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.

Natural versus GenomeOcean-generated DNA classification · Table 2:

Precision (percent) · Higher values are better.

Natural versus GenomeOcean-generated DNA classification (Natural vs artificial microbial genome sequence) · GenomeOcean natural/artificial sequence test

Evidence origin: Independent external evaluation, Author-reported evaluation.

GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:: Precision, Natural versus GenomeOcean-generated DNA classification
  • Self-generator recognition is not generic biological sequence quality.
  • Counts are cohort sizes, not verified per-method successful-prediction coverage.
Comparison details and limitations

DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.

  • No interval assigned unless printed in source cell.

Automated source review: 2026-09-17. Numerical source review does not establish independent reproduction.

Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.

Showing 3 of 3 matching rows.

Tested configuration
0255075100
Reported score
  1. GenomeOcean99
  2. DNABERT-285.2
  3. Nucleotide Transformers V283.3

Methods and evaluation design

Procedure, tasks and evaluated configurations

How it works

Evaluation in this paper

DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.

SourcesGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:: Precision, Natural versus GenomeOcean-generated DNA classification

Evaluation design

Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.

These source-backed links do not make different protocols or scores interchangeable.

Recorded evaluations

Each evaluation records what was tested and under which conditions.

Baseline coverage

Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.

0 of 2 active baseline roles have published Rewire measurements in this release. Measurements on a selected protocol do not establish coverage of an entire suite.

No execution recipe linked to this protocol. Recipe availability does not establish a completed evaluation.

No reviewed evaluations with results linked in this release.

Literature evidence is not a Rewire measurement. Executed but unpublished runs and private review status are not included.

Null control

Proposed control: requires review

Protocol-valid abundance or assignment control

Protocol-specific applicability, permitted inputs, access, split, evaluator and execution requirements need review before implementation or execution.

This is a suggested selection rule, not a validated method or a measured score.

Conventional reference

Proposed control: requires review

Conventional reference-database method with pinned taxonomy/database

Protocol-specific applicability, permitted inputs, access, split, evaluator and execution requirements need review before implementation or execution.

This is a suggested selection rule, not a validated method or a measured score.

Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums

Coverage is derived from release 2026-09-29-06401fd5b220. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.

Run instructions

No runnable recipe has been reviewed for this protocol. Dataset access, model requirements, licences and compute requirements must be checked against its sources before execution.

Strengths, limitations and unresolved questions

Strengths and limitations

Strengths and considerations

No source-reviewed explanatory claims are recorded here yet.

Limitations and conditions

No source-reviewed explanatory claims are recorded here yet.

Profile review details

Primary-source transcription and separate automated review. No human sign-off or experimental reproduction.

Stable record: paper-protocol-a8c67b393443253db5

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-17. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetsNot extracted or verified for this record.
OrganismsNot extracted or verified for this record.
AssaysNot extracted or verified for this record.
SplitsNot extracted or verified for this record.
Allowed inputsNot extracted or verified for this record.
AdaptationNot extracted or verified for this record.
MetricsNot extracted or verified for this record.
BaselinesNot extracted or verified for this record.

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.

Paper or primary resourceVersionReference
GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assembliespreprint archived 2025-02-05Read source
DOI: 10.1101/2025.01.30.635558
Historical gaps recorded on 2026-09-17

The catalogue now holds 9 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.

  • Independent batch review before import; preserve existing observation identities.
Search and extraction details

complete comparison tables extracted pending publication review

Searches

  • GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies primary paper benchmark results

Evidence locations

  • Table 2:: Precision, Natural versus GenomeOcean-generated DNA classification

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

3 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Evaluation in this paper
DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.
Individual claims
GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies

Original source ↗

Table 2:: Precision, Natural versus GenomeOcean-generated DNA classification

Version: preprint archived 2025-02-05
Retrieved: 2026-09-17T07:56:18.823068+00:00

source checked

automated source review · 2026-09-17

Audit details

Primary-source transcription and separate automated review. No human sign-off or experimental reproduction.

Field: attributes.profile.sections.0.body

Source artifact SHA-256: 3cc0df52522fccda23e3958f069c916b87ee50bb5c9a992fa37e25256546e145

Hash scope: Exact retrieved primary paper artifact bytes.

Inspected artifact

Introduction
Natural versus GenomeOcean-generated DNA classification · Table 2:. DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.
Individual claims
GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies

Original source ↗

Table 2:: Precision, Natural versus GenomeOcean-generated DNA classification

Version: preprint archived 2025-02-05
Retrieved: 2026-09-17T07:56:18.823068+00:00

source checked

automated source review · 2026-09-17

Audit details

Primary-source transcription and separate automated review. No human sign-off or experimental reproduction.

Field: attributes.profile.summary

Source artifact SHA-256: 3cc0df52522fccda23e3958f069c916b87ee50bb5c9a992fa37e25256546e145

Hash scope: Exact retrieved primary paper artifact bytes.

Inspected artifact

Relationship: evaluates task
reported-task-9f9ab0090f6522
Individual claims
GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies

Original source ↗

Table 2:: Precision, Natural versus GenomeOcean-generated DNA classification

Version: preprint archived 2025-02-05
Retrieved: 2026-09-17T07:56:18.823068+00:00

source checked

automated source review · 2026-09-17

Audit details

Field: links:evaluates_task:reported-task-9f9ab0090f6522

Claim: paper-claim-734285a668bd8dea20

Source artifact SHA-256: 3cc0df52522fccda23e3958f069c916b87ee50bb5c9a992fa37e25256546e145

Hash scope: Exact retrieved primary paper artifact bytes.

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: needs review

1 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: paper-protocol-a8c67b393443253db5

areas
microbes-communities
tasks
Natural vs artificial microbial genome sequence
entity level
protocol
protocol
DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.
comparison panels
id: part2-genomeocean-2025-T2-d2237c539e; title: Natural versus GenomeOcean-generated DNA classification · Table 2:; protocol id: paper-protocol-a8c67b393443253db5; dataset id: reported-dataset-0c3ac7efe99c37; metric: Precision; unit: percent; direction: higher; result ids: paper-result-da2786542bcec5cf7f; paper-result-fae558ea3909a73b15; paper-result-af3d9eb456d1404344; source ids: part2-genomeocean-2025; source locator: Table 2:: Precision, Natural versus GenomeOcean-generated DNA classification; context: DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.; caveats: Self-generator recognition is not generic biological sequence quality.; Counts are cohort sizes, not verified per-method successful-prediction coverage.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17; id: part2-genomeocean-2025-T2-3c62335b1d; title: Natural versus GenomeOcean-generated DNA classification · Table 2:; protocol id: paper-protocol-a8c67b393443253db5; dataset id: reported-dataset-0c3ac7efe99c37; metric: Recall; unit: percent; direction: higher; result ids: paper-result-415548dbb583463a98; paper-result-b3378778c23879d6db; paper-result-f356ae1c7c4122384a; source ids: part2-genomeocean-2025; source locator: Table 2:: Recall, Natural versus GenomeOcean-generated DNA classification; context: DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.; caveats: Self-generator recognition is not generic biological sequence quality.; Counts are cohort sizes, not verified per-method successful-prediction coverage.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17; id: part2-genomeocean-2025-T2-92317bff43; title: Natural versus GenomeOcean-generated DNA classification · Table 2:; protocol id: paper-protocol-a8c67b393443253db5; dataset id: reported-dataset-0c3ac7efe99c37; metric: F1; unit: %; direction: higher; result ids: lit-b3-028; paper-result-0209741f5d0e66b495; lit-b3-027; source ids: part2-genomeocean-2025; source locator: Table 2:: F1, Natural versus GenomeOcean-generated DNA classification; context: DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.; caveats: Self-generator recognition is not generic biological sequence quality.; Counts are cohort sizes, not verified per-method successful-prediction coverage.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17
benchmark research
review date: 2026-09-17; status: complete_comparison_tables_extracted_pending_publication_review; primary sources: part2-genomeocean-2025; inspected locators: Table 2:: Precision, Natural versus GenomeOcean-generated DNA classification; searched queries: GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies primary paper benchmark results; gaps: Independent batch review before import; preserve existing observation identities.; claim scope: Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.
legacy kinds
benchmark
entity classification
review date: 2026-09-17; rationale: The source-backed record identifies a specified evaluated procedure and its dataset/split/scoring context. Classify it as a protocol while preserving version and comparison restrictions.; source ids: part2-genomeocean-2025; source locator: Table 2:: Precision, Natural versus GenomeOcean-generated DNA classification; ambiguities: None recorded
Related records

Suggest a correction