Datasets
A 10,000-read subset of CAMI II Sample_0 from the human-microbiome collection.
CAMI II read classification evaluates a computationally limited subsample, with taxonomic-rank-specific interpretation.
A 10,000-read subset of CAMI II Sample_0 from the human-microbiome collection.
Macro-F1 is the arithmetic mean of classwise F1; only classes present in the evaluated subset enter the macro calculation.
DNA reads and a separately assembled reference/training collection.
Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.
limited source coverage · Automated source review, 2026-09-16. All specifications and missing details
Results are available, but no reviewed comparison panel is linked in this release.
1 evaluation · 1 result. Different protocols are not a single leaderboard.
Applied filters: All linked evaluations
| Tested configuration | Protocol and dataset | Finding | Evidence and details |
|---|---|---|---|
| Configuration: NCD-gzip | Task: CAMI II phylum read classification Dataset: CAMI II Sample_0 10,000-read subsample | 0.126 Macro F1 unitless · unknown Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNCD-gzip: CAMI II phylum read classification Phylum-level macro-averaged F1; distinct taxonomic rank from the other row. Aggregation: Not reported Normalized compression distance for DNA classification · Table 5, NCD Phylum row, F1 column |
Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.
A 10,000-read subset of CAMI II Sample_0 from the human-microbiome collection. Queries are classified against the paper’s separately assembled metagenomic reference/training data. Macro-F1 is the arithmetic mean of classwise F1; only classes present in the evaluated subset enter the macro calculation. NCD-gzip and Kraken2. The CAMI II test uses 10,000 reads from Sample_0 and the paper’s RefSeq metagenomic training collection. A disjoint genome split is described for the separate simulated RefSeq experiment; the CAMI passage does not establish reference-genome or homology exclusion against this external sample.
Each evaluation records what was tested and under which conditions.
No runnable recipe has been reviewed for this task. Dataset access, model requirements, licences and compute requirements must be checked against its sources before execution.
A task describes a biological question. Choose a linked protocol to obtain concrete split and scoring instructions.
No source-reviewed explanatory claims are recorded here yet.
Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.
Stable record: reported-task-45105e1c486251Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Datasets | A 10,000-read subset of CAMI II Sample_0 from the human-microbiome collection.SourcesNormalized compression distance for DNA classification · Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116 |
| Splits | Queries are classified against the paper’s separately assembled metagenomic reference/training data.SourcesNormalized compression distance for DNA classification · Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116 |
| Metrics | Macro-F1 is the arithmetic mean of classwise F1; only classes present in the evaluated subset enter the macro calculation.SourcesNormalized compression distance for DNA classification · Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116 |
| Baselines | NCD-gzip and Kraken2.SourcesNormalized compression distance for DNA classification · Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116 |
| Leakage controls | The CAMI II test uses 10,000 reads from Sample_0 and the paper’s RefSeq metagenomic training collection. A disjoint genome split is described for the separate simulated RefSeq experiment; the CAMI passage does not establish reference-genome or homology exclusion against this external sample. · Not reported in inspected sourcesSourcesNormalized compression distance for DNA classification · Methods: Datasets/Metagenomic reads; Results: CAMI dataset |
| Uncertainty | Table 5 describes a single 10,000-read CAMI II subsample. It does not provide confidence intervals, repeated-subsample variation or random-seed uncertainty for the phylum metrics. · Not reported in inspected sourcesSourcesNormalized compression distance for DNA classification · Results: CAMI dataset; Table 5 |
| Entity type | Paper-specific computational evaluation protocol.SourcesNormalized compression distance for DNA classification · Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116 |
| Organisms | CAMI II human-microbiome community taxa.SourcesNormalized compression distance for DNA classification · Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116 |
| Assays | Read-origin taxonomy at the selected evaluation rank.SourcesNormalized compression distance for DNA classification · Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116 |
| Allowed inputs | DNA reads and a separately assembled reference/training collection.SourcesNormalized compression distance for DNA classification · Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116 |
| Adaptation | Reference-based classification using NCD-gzip or Kraken2.SourcesNormalized compression distance for DNA classification · Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116 |
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Last literature check: 2026-09-17. Primary-paper discovery and source inspection. Source-checked results are not independently reproduced experiments.
| Paper or primary resource | Version | Reference |
|---|---|---|
| Normalized compression distance for DNA classification | version of record | Read source DOI: 10.7717/peerj.20677 |
The catalogue now holds 1 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.
primary comparison table screened
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
17 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol. Individual claims | Normalized compression distance for DNA classification Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116 Version: version of record | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| Normalized compression distance for DNA classification Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116 Version: version of record | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Computational evaluation flow Individual claims | Normalized compression distance for DNA classification Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116 Version: version of record | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Datasets A 10,000-read subset of CAMI II Sample_0 from the human-microbiome collection. Individual claims | Normalized compression distance for DNA classification Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116 Version: version of record | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Splits Queries are classified against the paper’s separately assembled metagenomic reference/training data. Individual claims | Normalized compression distance for DNA classification Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116 Version: version of record | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Adaptation Reference-based classification using NCD-gzip or Kraken2. Individual claims | Normalized compression distance for DNA classification Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116 Version: version of record | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Metrics Macro-F1 is the arithmetic mean of classwise F1; only classes present in the evaluated subset enter the macro calculation. Individual claims | Normalized compression distance for DNA classification Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116 Version: version of record | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Baselines NCD-gzip and Kraken2. Individual claims | Normalized compression distance for DNA classification Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116 Version: version of record | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Leakage controls The CAMI II test uses 10,000 reads from Sample_0 and the paper’s RefSeq metagenomic training collection. A disjoint genome split is described for the separate simulated RefSeq experiment; the CAMI passage does not establish reference-genome or homology exclusion against this external sample. Individual claims | Normalized compression distance for DNA classification Methods: Datasets/Metagenomic reads; Results: CAMI dataset Version: version of record | unreported automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Uncertainty Table 5 describes a single 10,000-read CAMI II subsample. It does not provide confidence intervals, repeated-subsample variation or random-seed uncertainty for the phylum metrics. Individual claims | Normalized compression distance for DNA classification Results: CAMI dataset; Table 5 Version: version of record | unreported automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
View linked audit checks and correction history
Release 2026-09-29-06401fd5b220 · Record review: needs review
Stable ID: reported-task-45105e1c486251