rewire.itbenchmarks
Benchmark

CAFA

CAFA evaluates prospective protein-function predictions against annotations that become available after prediction submission.

SourcesCAFA official description — reviewed snapshot 2026-09-16 · Official CAFA page: The problem; The solution; challenge timeline

438 evaluations · 438 results

Overview

Metrics

CAFA3 protein-centric evaluation uses Fmax and semantic distance Smin, with prediction coverage reported. Term-centric assays also use AUROC. Ontology, knowledge category and full/partial evaluation mode must accompany the score.

Sourcescafa3 primary benchmark evidence · Methods: Protein-centric and term-centric evaluation; Figures 3–4
Evaluation procedure diagram
How it worksEvaluation procedure
Evaluation procedure1. Allowed inputs: Challenge target protein sequences and permitted pre-deadline knowledge.. Then: 2. Splits: Prospective temporal assessment separates prediction submission from subsequent annotation growth.. Then: 3. Metrics: CAFA3 protein-centric evaluation uses Fmax and semantic distance Smin, with prediction coverage reported. Term-centric assays also use AUROC. Ontology, knowledge category and full/partial evaluation mode must accompany the score.Evaluation procedure1. Allowed inputs: Challenge target protein sequences and permitted pre-deadline knowledge.. Then: 2. Splits: Prospective temporal assessment separates prediction submission from subsequent annotation growth.. Then: 3. Metrics: CAFA3 protein-centric evaluation uses Fmax and semantic distance Smin, with prediction coverage reported. Term-centric assays also use AUROC. Ontology, knowledge category and full/partial evaluation mode must accompany the score.Evaluation procedure1. Allowed inputs: Challenge target protein sequences and permitted pre-deadline knowledge.. Then: 2. Splits: Prospective temporal assessment separates prediction submission from subsequent annotation growth.. Then: 3. Metrics: CAFA3 protein-centric evaluation uses Fmax and semantic distance Smin, with prediction coverage reported. Term-centric assays also use AUROC. Ontology, knowledge category and full/partial evaluation mode must accompany the score.

Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.

Sources (2)CAFA official description — reviewed snapshot 2026-09-16; cafa3 primary benchmark evidence · Official CAFA page: The problem; The solution; challenge timeline; Methods: Protein-centric and term-centric evaluation; Figures 3–4

Source reviewed · Automated source review, 2026-09-16. All specifications and missing details

Results

Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.

CAFA3 MFO all organisms; type1 no-knowledge; mode1 · CAFA3 final benchmark; MFO; all; type1; mode1: F1-max (source entries 1–80 of 146)

F1-max (fraction) · Higher values are better.

CAFA3 MFO all organisms; type1 no-knowledge; mode1 · CAFA3 final benchmark; MFO; all; type1; mode1 · CAFA3 final benchmark; MFO; all; type1; mode1

Evidence origin: Independent external evaluation.

CAFA: Figshare 8135393 v3, file 17519846 · supplementary_data/cafa3/sheets/mfo_all_type1_mode1_all_fmax_sheet.csv; line 2; ID-model=BB7U; column F1-max through supplementary_data/cafa3/sheets/mfo_all_type1_mode1_all_fmax_sheet.csv; line 81; ID-model=M078; column F1-max
  • Missing source cells and quarantined conflicts are recorded in acquisition and audit tables. Per-result scoring denominators may be unreported.
Comparison details and limitations

Complete selected source table is retained across source-order panels. These point estimates do not establish statistical significance or a universal ranking.

  • Source-specific evaluation. No equivalence to other releases, protocols or model families is inferred.
  • Exact source-defined evaluation scope; reported scores are not rewire reproductions.

Automated source review: 2026-09-19. Numerical source review does not establish independent reproduction.

Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.

Showing 12 of 80 matching rows.

Methods and evaluation design

Procedure, tasks and evaluated configurations

How it works

Evaluation methodology

CAFA is a time-delayed protein-function challenge: predictions are submitted before new experimental annotations become available. Evaluation then compares predictions with those newly added labels, separately by ontology and evaluation mode. CAFA3 provides explicit frequency and homology-transfer baselines and protein-level bootstrap intervals.

Sourcescafa3 primary benchmark evidence · CAFA3 Methods: Protein-centric evaluation; Figures 3–4

Evaluation design

Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.

These source-backed links do not make different protocols or scores interchangeable.

Baseline coverage

Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.

0 of 6 active baseline roles have published Rewire measurements in this release. Measurements on a selected protocol do not establish coverage of an entire suite.

Baseline status by linked protocol

Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums

Coverage is derived from release 2026-09-29-06401fd5b220. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.

Run this benchmark

Choose a concrete protocol before running an evaluation. Its inputs, split and scoring rules determine which results can be compared.

Run this benchmark

The CAFA-evaluator authors document a CLI taking ontology, prediction directory and ground-truth files. Pin the challenge edition, ontology and ground truth separately; running the evaluator is not equivalent to entering a blind challenge. The challenge landing page was unavailable during retrieval, but the evaluator README was inspected and pinned.

A maintained rewire runner has not been verified for this benchmark. Check data access, weights, licences, dependencies and hardware in the linked official documentation; requirements have not been fully extracted.

BioComputingUP/CAFA-evaluator / README.md · CAFA-evaluator README.md lines 22–72 (Installation and Usage); challenge landing page retrieval failed
Strengths, limitations and unresolved questions

Strengths and limitations

Limitations and conditions

  • Annotation incompleteness, prediction coverage and the challenge’s ontology version affect interpretation. New CAFA rounds may revise the eligible proteins, metrics and uncertainty procedure.
    Sourcescafa3 primary benchmark evidence · CAFA3 Methods: Protein-centric evaluation; Figures 3–4
Profile review details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Stable record: discovery-benchmark-cafa

Specifications

Inputs, training, access and other details

Explanatory profile: source reviewed · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetsProtein sequences supplied for a challenge; proteins gaining experimental annotations after the deadline become evaluation targets.
SourcesCAFA official description — reviewed snapshot 2026-09-16 · Official CAFA page: The problem; The solution; challenge timeline
SplitsProspective temporal assessment separates prediction submission from subsequent annotation growth.
SourcesCAFA official description — reviewed snapshot 2026-09-16 · Official CAFA page: The problem; The solution; challenge timeline
MetricsCAFA3 protein-centric evaluation uses Fmax and semantic distance Smin, with prediction coverage reported. Term-centric assays also use AUROC. Ontology, knowledge category and full/partial evaluation mode must accompany the score.
Sourcescafa3 primary benchmark evidence · Methods: Protein-centric and term-centric evaluation; Figures 3–4
BaselinesIn CAFA3 protein-centric evaluation, Naïve predicts each term’s training-set frequency and BLAST transfers annotations using the highest matching sequence identity. Term-centric experiments also include expression-based comparators where available.
Sourcescafa3 primary benchmark evidence · CAFA3 Methods: Protein-centric evaluation; Figures 3–4
Leakage controlsNewly acquired experimental annotations are used for later assessment; exact knowledge-cutoff controls depend on the challenge edition.
SourcesCAFA official description — reviewed snapshot 2026-09-16 · Official CAFA page: The problem; The solution; challenge timeline
UncertaintyCAFA3 Figures 3–4 estimate 95% confidence intervals with 10,000 bootstrap samples of benchmark proteins. This is a version-specific evaluation, not a universal rule for every CAFA round.
Sourcescafa3 primary benchmark evidence · CAFA3 Methods: Protein-centric evaluation; Figures 3–4
Entity typeProspective protein-function prediction challenge.
SourcesCAFA official description — reviewed snapshot 2026-09-16 · Official CAFA page: The problem; The solution; challenge timeline
OrganismsProtein targets spanning challenge-selected taxa.
SourcesCAFA official description — reviewed snapshot 2026-09-16 · Official CAFA page: The problem; The solution; challenge timeline
AssaysExperimental functional annotations acquired after the prediction deadline.
SourcesCAFA official description — reviewed snapshot 2026-09-16 · Official CAFA page: The problem; The solution; challenge timeline
Allowed inputsChallenge target protein sequences and permitted pre-deadline knowledge.
SourcesCAFA official description — reviewed snapshot 2026-09-16 · Official CAFA page: The problem; The solution; challenge timeline
AdaptationMethods submit predictions before the target proteins gain evaluation annotations.
SourcesCAFA official description — reviewed snapshot 2026-09-16 · Official CAFA page: The problem; The solution; challenge timeline
Applicable tests and references

Applicability is distinct from a completed evaluation.

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.

Historical gaps recorded on 2026-09-17

The catalogue now holds 438 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.

  • Main tables describe participation and benchmark construction; numerical performance is mainly in figures/supplementary files. CAFA versions, ontology branch, full/partial mode and no/limited-knowledge cohorts must remain separate. No scores estimated from plot pixels.
Search and extraction details

source found structured extraction pending

Searches

  • CAFA primary paper benchmark results

Evidence locations

  • Tables1–3; protein-function prediction evaluation and challenge results

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

21 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Individual claims
CAFA official description — reviewed snapshot 2026-09-16

Original source ↗

Official CAFA page: The problem; The solution; challenge timeline; Methods: Protein-centric and term-centric evaluation; Figures 3–4

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2026-09-16 website snapshot sha256:d7a5fa76bea98551f8322aab9da965c383104c79d83992e723c0c5836b2fdcd1
Retrieved: 2026-09-16T19:47:27.547276+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: d7a5fa76bea98551f8322aab9da965c383104c79d83992e723c0c5836b2fdcd1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram caption
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Individual claims
cafa3 primary benchmark evidence

Original source ↗

Official CAFA page: The problem; The solution; challenge timeline; Methods: Protein-centric and term-centric evaluation; Figures 3–4

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: PMC6864930
Retrieved: 2026-09-16T21:05:50.876204+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 3802f37548d77addac829127bf03322cbe6c69eba944d88a775b534311804fa4

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps
  • Allowed inputs: Challenge target protein sequences and permitted pre-deadline knowledge.
  • Splits: Prospective temporal assessment separates prediction submission from subsequent annotation growth.
  • Metrics: CAFA3 protein-centric evaluation uses Fmax and semantic distance Smin, with prediction coverage reported. Term-centric assays also use AUROC. Ontology, knowledge category and full/partial evaluation mode must accompany the score.
Individual claims
CAFA official description — reviewed snapshot 2026-09-16

Original source ↗

Official CAFA page: The problem; The solution; challenge timeline; Methods: Protein-centric and term-centric evaluation; Figures 3–4

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2026-09-16 website snapshot sha256:d7a5fa76bea98551f8322aab9da965c383104c79d83992e723c0c5836b2fdcd1
Retrieved: 2026-09-16T19:47:27.547276+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: d7a5fa76bea98551f8322aab9da965c383104c79d83992e723c0c5836b2fdcd1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps
  • Allowed inputs: Challenge target protein sequences and permitted pre-deadline knowledge.
  • Splits: Prospective temporal assessment separates prediction submission from subsequent annotation growth.
  • Metrics: CAFA3 protein-centric evaluation uses Fmax and semantic distance Smin, with prediction coverage reported. Term-centric assays also use AUROC. Ontology, knowledge category and full/partial evaluation mode must accompany the score.
Individual claims
cafa3 primary benchmark evidence

Original source ↗

Official CAFA page: The problem; The solution; challenge timeline; Methods: Protein-centric and term-centric evaluation; Figures 3–4

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: PMC6864930
Retrieved: 2026-09-16T21:05:50.876204+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 3802f37548d77addac829127bf03322cbe6c69eba944d88a775b534311804fa4

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title
Evaluation procedure
Individual claims
CAFA official description — reviewed snapshot 2026-09-16

Original source ↗

Official CAFA page: The problem; The solution; challenge timeline; Methods: Protein-centric and term-centric evaluation; Figures 3–4

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2026-09-16 website snapshot sha256:d7a5fa76bea98551f8322aab9da965c383104c79d83992e723c0c5836b2fdcd1
Retrieved: 2026-09-16T19:47:27.547276+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: d7a5fa76bea98551f8322aab9da965c383104c79d83992e723c0c5836b2fdcd1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title
Evaluation procedure
Individual claims
cafa3 primary benchmark evidence

Original source ↗

Official CAFA page: The problem; The solution; challenge timeline; Methods: Protein-centric and term-centric evaluation; Figures 3–4

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: PMC6864930
Retrieved: 2026-09-16T21:05:50.876204+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: 3802f37548d77addac829127bf03322cbe6c69eba944d88a775b534311804fa4

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Datasets
Protein sequences supplied for a challenge; proteins gaining experimental annotations after the deadline become evaluation targets.
Individual claims
CAFA official description — reviewed snapshot 2026-09-16

Original source ↗

Official CAFA page: The problem; The solution; challenge timeline

Version: 2026-09-16 website snapshot sha256:d7a5fa76bea98551f8322aab9da965c383104c79d83992e723c0c5836b2fdcd1
Retrieved: 2026-09-16T19:47:27.547276+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: d7a5fa76bea98551f8322aab9da965c383104c79d83992e723c0c5836b2fdcd1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Splits
Prospective temporal assessment separates prediction submission from subsequent annotation growth.
Individual claims
CAFA official description — reviewed snapshot 2026-09-16

Original source ↗

Official CAFA page: The problem; The solution; challenge timeline

Version: 2026-09-16 website snapshot sha256:d7a5fa76bea98551f8322aab9da965c383104c79d83992e723c0c5836b2fdcd1
Retrieved: 2026-09-16T19:47:27.547276+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: d7a5fa76bea98551f8322aab9da965c383104c79d83992e723c0c5836b2fdcd1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Adaptation
Methods submit predictions before the target proteins gain evaluation annotations.
Individual claims
CAFA official description — reviewed snapshot 2026-09-16

Original source ↗

Official CAFA page: The problem; The solution; challenge timeline

Version: 2026-09-16 website snapshot sha256:d7a5fa76bea98551f8322aab9da965c383104c79d83992e723c0c5836b2fdcd1
Retrieved: 2026-09-16T19:47:27.547276+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: d7a5fa76bea98551f8322aab9da965c383104c79d83992e723c0c5836b2fdcd1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Metrics
CAFA3 protein-centric evaluation uses Fmax and semantic distance Smin, with prediction coverage reported. Term-centric assays also use AUROC. Ontology, knowledge category and full/partial evaluation mode must accompany the score.
Individual claims
cafa3 primary benchmark evidence

Original source ↗

Methods: Protein-centric and term-centric evaluation; Figures 3–4

Version: PMC6864930
Retrieved: 2026-09-16T21:05:50.876204+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: 3802f37548d77addac829127bf03322cbe6c69eba944d88a775b534311804fa4

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: discovered

5 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: discovery-benchmark-cafa

areas
protein-function
entity level
challenge
scope note
Specialist molecular or omics evaluation; protocol details require review before numerical comparison.
task
Protein function prediction challenge
version
Not reported
benchmark research
review date: 2026-09-17; status: source_found_structured_extraction_pending; primary sources: evidence-expansion-p2-evidence-discovery-final-cafa3-3802f37548d7; inspected locators: Tables1–3; protein-function prediction evaluation and challenge results; searched queries: CAFA primary paper benchmark results; gaps: Main tables describe participation and benchmark construction; numerical performance is mainly in figures/supplementary files. CAFA versions, ontology branch, full/partial mode and no/limited-knowledge cohorts must remain separate. No scores estimated from plot pixels.; claim scope: Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.
historical missing metadata
dataset release: unextracted; metric implementation: unextracted; split manifest: unextracted; version: unextracted
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
entity classification
review date: 2026-09-17; rationale: The cited profile describes an organized challenge with category- or round-specific evaluation rules; retain it as a top-level benchmark, without conflating different editions or protocols.; source ids: evidence-benchmark-cafa-20260916; source locator: Official CAFA page: The problem; The solution; challenge timeline; ambiguities: None recorded
run documentation
record id: discovery-benchmark-cafa; source ids: run-doc-cafa-evaluator-readme-md-d09ba823; status: official_documentation_linked; summary: The CAFA-evaluator authors document a CLI taking ontology, prediction directory and ground-truth files. Pin the challenge edition, ontology and ground truth separately; running the evaluator is not equivalent to entering a blind challenge. The challenge landing page was unavailable during retrieval, but the evaluator README was inspected and pinned.; source locator: CAFA-evaluator README.md lines 22–72 (Installation and Usage); challenge landing page retrieval failed
Related records

Suggest a correction