rewire.itbenchmarks
Task

CAMI II phylum read classification

CAMI II read classification evaluates a computationally limited subsample, with taxonomic-rank-specific interpretation.

SourcesNormalized compression distance for DNA classification · Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116

1 evaluation · 1 result

Overview

Datasets

A 10,000-read subset of CAMI II Sample_0 from the human-microbiome collection.

Metrics

Macro-F1 is the arithmetic mean of classwise F1; only classes present in the evaluated subset enter the macro calculation.

Allowed inputs

DNA reads and a separately assembled reference/training collection.

SourcesNormalized compression distance for DNA classification · Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116
Evaluation procedure diagram
How it worksComputational evaluation flow
Computational evaluation flow1. Input: DNA reads and a separately assembled reference/training collection.. Then: 2. Evaluation: Queries are classified against the paper’s separately assembled metagenomic reference/training data.. Then: 3. Readout: Macro-F1 is the arithmetic mean of classwise F1; only classes present in the evaluated subset enter the macro calculation.Computational evaluation flow1. Input: DNA reads and a separately assembled reference/training collection.. Then: 2. Evaluation: Queries are classified against the paper’s separately assembled metagenomic reference/training data.. Then: 3. Readout: Macro-F1 is the arithmetic mean of classwise F1; only classes present in the evaluated subset enter the macro calculation.Computational evaluation flow1. Input: DNA reads and a separately assembled reference/training collection.. Then: 2. Evaluation: Queries are classified against the paper’s separately assembled metagenomic reference/training data.. Then: 3. Readout: Macro-F1 is the arithmetic mean of classwise F1; only classes present in the evaluated subset enter the macro calculation.

Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.

SourcesNormalized compression distance for DNA classification · Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116

limited source coverage · Automated source review, 2026-09-16. All specifications and missing details

Results

Results are available, but no reviewed comparison panel is linked in this release.

All evaluations

1 evaluation · 1 result. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: NCD-gzipTask: CAMI II phylum read classification
Dataset: CAMI II Sample_0 10,000-read subsample
0.126 Macro F1
unitless · unknown

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

NCD-gzip: CAMI II phylum read classification

Phylum-level macro-averaged F1; distinct taxonomic rank from the other row.

Aggregation: Not reported

Normalized compression distance for DNA classification · Table 5, NCD Phylum row, F1 column

Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.

Methods and evaluation design

Procedure, tasks and evaluated configurations

How it works

Evaluation methodology

A 10,000-read subset of CAMI II Sample_0 from the human-microbiome collection. Queries are classified against the paper’s separately assembled metagenomic reference/training data. Macro-F1 is the arithmetic mean of classwise F1; only classes present in the evaluated subset enter the macro calculation. NCD-gzip and Kraken2. The CAMI II test uses 10,000 reads from Sample_0 and the paper’s RefSeq metagenomic training collection. A disjoint genome split is described for the separate simulated RefSeq experiment; the CAMI passage does not establish reference-genome or homology exclusion against this external sample.

SourcesNormalized compression distance for DNA classification · Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116; Methods: Datasets/Metagenomic reads; Results: CAMI dataset

Recorded evaluations

Each evaluation records what was tested and under which conditions.

Run instructions

No runnable recipe has been reviewed for this task. Dataset access, model requirements, licences and compute requirements must be checked against its sources before execution.

A task describes a biological question. Choose a linked protocol to obtain concrete split and scoring instructions.

Strengths, limitations and unresolved questions

Strengths and limitations

Strengths and considerations

No source-reviewed explanatory claims are recorded here yet.

Limitations and conditions

  • The subsample contains only Bacteria at superkingdom level, so that result does not establish broad multiclass discrimination. Phylum and superkingdom are separate catalogue outcomes.
    SourcesNormalized compression distance for DNA classification · Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116
Profile review details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Stable record: reported-task-45105e1c486251

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetsA 10,000-read subset of CAMI II Sample_0 from the human-microbiome collection.
SourcesNormalized compression distance for DNA classification · Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116
SplitsQueries are classified against the paper’s separately assembled metagenomic reference/training data.
SourcesNormalized compression distance for DNA classification · Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116
MetricsMacro-F1 is the arithmetic mean of classwise F1; only classes present in the evaluated subset enter the macro calculation.
SourcesNormalized compression distance for DNA classification · Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116
BaselinesNCD-gzip and Kraken2.
SourcesNormalized compression distance for DNA classification · Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116
Leakage controlsThe CAMI II test uses 10,000 reads from Sample_0 and the paper’s RefSeq metagenomic training collection. A disjoint genome split is described for the separate simulated RefSeq experiment; the CAMI passage does not establish reference-genome or homology exclusion against this external sample. · Not reported in inspected sources
SourcesNormalized compression distance for DNA classification · Methods: Datasets/Metagenomic reads; Results: CAMI dataset
UncertaintyTable 5 describes a single 10,000-read CAMI II subsample. It does not provide confidence intervals, repeated-subsample variation or random-seed uncertainty for the phylum metrics. · Not reported in inspected sources
SourcesNormalized compression distance for DNA classification · Results: CAMI dataset; Table 5
Entity typePaper-specific computational evaluation protocol.
SourcesNormalized compression distance for DNA classification · Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116
OrganismsCAMI II human-microbiome community taxa.
SourcesNormalized compression distance for DNA classification · Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116
AssaysRead-origin taxonomy at the selected evaluation rank.
SourcesNormalized compression distance for DNA classification · Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116
Allowed inputsDNA reads and a separately assembled reference/training collection.
SourcesNormalized compression distance for DNA classification · Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116
AdaptationReference-based classification using NCD-gzip or Kraken2.
SourcesNormalized compression distance for DNA classification · Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Primary-paper discovery and source inspection. Source-checked results are not independently reproduced experiments.

Paper or primary resourceVersionReference
Normalized compression distance for DNA classificationversion of recordRead source
DOI: 10.7717/peerj.20677
Historical gaps recorded on 2026-09-17

The catalogue now holds 1 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.

  • Table 5 prints NCD only. Narrative compares Kraken2 but supplies no paired numerical Kraken2 row; do not manufacture a two-model chart.
  • Do not attach five-fold Human DNA results from Table 3 to CAMI II.
Search and extraction details

primary comparison table screened

Searches

  • "Normalized compression distance for DNA classification"

Evidence locations

  • Table 5
  • Results: CAMI analysis
  • Methods: Metagenomic reads and Genome fragmentation

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

17 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.
Individual claims
Normalized compression distance for DNA classification

Original source ↗

Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116

Version: version of record
Retrieved: 2026-09-16T10:44:03.414806+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 10e9ba780c45e7787baff9b81ef7b45c014d7fe6c716d6c759e14d89a813dc1c

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps
  • Input: DNA reads and a separately assembled reference/training collection.
  • Evaluation: Queries are classified against the paper’s separately assembled metagenomic reference/training data.
  • Readout: Macro-F1 is the arithmetic mean of classwise F1; only classes present in the evaluated subset enter the macro calculation.
Individual claims
Normalized compression distance for DNA classification

Original source ↗

Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116

Version: version of record
Retrieved: 2026-09-16T10:44:03.414806+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 10e9ba780c45e7787baff9b81ef7b45c014d7fe6c716d6c759e14d89a813dc1c

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title
Computational evaluation flow
Individual claims
Normalized compression distance for DNA classification

Original source ↗

Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116

Version: version of record
Retrieved: 2026-09-16T10:44:03.414806+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.diagram.title

Source artifact SHA-256: 10e9ba780c45e7787baff9b81ef7b45c014d7fe6c716d6c759e14d89a813dc1c

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Datasets
A 10,000-read subset of CAMI II Sample_0 from the human-microbiome collection.
Individual claims
Normalized compression distance for DNA classification

Original source ↗

Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116

Version: version of record
Retrieved: 2026-09-16T10:44:03.414806+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: 10e9ba780c45e7787baff9b81ef7b45c014d7fe6c716d6c759e14d89a813dc1c

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Splits
Queries are classified against the paper’s separately assembled metagenomic reference/training data.
Individual claims
Normalized compression distance for DNA classification

Original source ↗

Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116

Version: version of record
Retrieved: 2026-09-16T10:44:03.414806+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: 10e9ba780c45e7787baff9b81ef7b45c014d7fe6c716d6c759e14d89a813dc1c

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Adaptation
Reference-based classification using NCD-gzip or Kraken2.
Individual claims
Normalized compression distance for DNA classification

Original source ↗

Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116

Version: version of record
Retrieved: 2026-09-16T10:44:03.414806+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: 10e9ba780c45e7787baff9b81ef7b45c014d7fe6c716d6c759e14d89a813dc1c

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Metrics
Macro-F1 is the arithmetic mean of classwise F1; only classes present in the evaluated subset enter the macro calculation.
Individual claims
Normalized compression distance for DNA classification

Original source ↗

Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116

Version: version of record
Retrieved: 2026-09-16T10:44:03.414806+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: 10e9ba780c45e7787baff9b81ef7b45c014d7fe6c716d6c759e14d89a813dc1c

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Baselines
NCD-gzip and Kraken2.
Individual claims
Normalized compression distance for DNA classification

Original source ↗

Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116

Version: version of record
Retrieved: 2026-09-16T10:44:03.414806+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.3.value

Source artifact SHA-256: 10e9ba780c45e7787baff9b81ef7b45c014d7fe6c716d6c759e14d89a813dc1c

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Leakage controls
The CAMI II test uses 10,000 reads from Sample_0 and the paper’s RefSeq metagenomic training collection. A disjoint genome split is described for the separate simulated RefSeq experiment; the CAMI passage does not establish reference-genome or homology exclusion against this external sample.
Individual claims
Normalized compression distance for DNA classification

Original source ↗

Methods: Datasets/Metagenomic reads; Results: CAMI dataset

Version: version of record
Retrieved: 2026-09-16T10:44:03.414806+00:00

unreported

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: 10e9ba780c45e7787baff9b81ef7b45c014d7fe6c716d6c759e14d89a813dc1c

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Uncertainty
Table 5 describes a single 10,000-read CAMI II subsample. It does not provide confidence intervals, repeated-subsample variation or random-seed uncertainty for the phylum metrics.
Individual claims
Normalized compression distance for DNA classification

Original source ↗

Results: CAMI dataset; Table 5

Version: version of record
Retrieved: 2026-09-16T10:44:03.414806+00:00

unreported

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.5.value

Source artifact SHA-256: 10e9ba780c45e7787baff9b81ef7b45c014d7fe6c716d6c759e14d89a813dc1c

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: needs review

2 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: reported-task-45105e1c486251

areas
microbes-communities
tasks
CAMI II phylum read classification
entity level
task
version
Not reported
task
CAMI II phylum read classification
scope note
Paper-specific evaluation task; protocol completeness requires further extraction.
benchmark research
review date: 2026-09-17; status: primary_comparison_table_screened; primary sources: expansion-p3-ncd-metagenomics-2026; inspected locators: Table 5; Results: CAMI analysis; Methods: Metagenomic reads and Genome fragmentation; searched queries: "Normalized compression distance for DNA classification"; gaps: Table 5 prints NCD only. Narrative compares Kraken2 but supplies no paired numerical Kraken2 row; do not manufacture a two-model chart.; Do not attach five-fold Human DNA results from Table 3 to CAMI II.; claim scope: Primary-paper discovery and source inspection. Source-checked results are not independently reproduced experiments.
historical missing metadata
protocol version: not_reported_in_legacy_extract; split: not_reported_in_legacy_extract
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
legacy kinds
benchmark
entity classification
review date: 2026-09-17; rationale: This source-scoped record identifies the biological prediction task and holds its paper context. Preserve the existing task identity; exact split, model adaptation and scoring remain in linked evaluations or separate protocol records.; source ids: ncd-metagenomics-2026; source locator: Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116; ambiguities: A paper- or suite-specific task may constrain some inputs or metrics; that alone does not make it interchangeable with a complete versioned protocol. No protocol equivalence is inferred.; Some legacy profile Entity type facts use the generic phrase computational evaluation protocol. That boilerplate is not sufficient to establish a single fixed protocol identity or to merge this task with another protocol record.
Related records

Suggest a correction