rewire.itbenchmarks
Benchmark

Open Problems

Open Problems is an extensible platform hosting benchmark tasks and their datasets.

Sourcesopenproblems-bio/openproblems official source · Pinned README: platform description and benchmark/dataset resource links

384 evaluations · 384 results

Overview

Datasets

The README links benchmark and dataset catalogues rather than fixing one data release.

Metrics

No platform-wide scientific metric applies: hosted task definitions select their own evaluator.

inapplicable

Allowed inputs

Inputs are specified by each hosted task.

Sourcesopenproblems-bio/openproblems official source · Pinned README: platform description and benchmark/dataset resource links
Evaluation procedure diagram
How it worksEvaluation procedure
Evaluation procedure1. Allowed inputs: Inputs are specified by each hosted task.. Then: 2. Splits: No platform-wide split exists: select a hosted benchmark task and its dataset protocol.. Then: 3. Metrics: No platform-wide scientific metric applies: hosted task definitions select their own evaluator.Evaluation procedure1. Allowed inputs: Inputs are specified by each hosted task.. Then: 2. Splits: No platform-wide split exists: select a hosted benchmark task and its dataset protocol.. Then: 3. Metrics: No platform-wide scientific metric applies: hosted task definitions select their own evaluator.Evaluation procedure1. Allowed inputs: Inputs are specified by each hosted task.. Then: 2. Splits: No platform-wide split exists: select a hosted benchmark task and its dataset protocol.. Then: 3. Metrics: No platform-wide scientific metric applies: hosted task definitions select their own evaluator.

Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.

Sourcesopenproblems-bio/openproblems official source · Pinned README: platform description and benchmark/dataset resource links

limited source coverage · Automated source review, 2026-09-16. All specifications and missing details

Results

Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.

Open Problems label projection CENGEN-BATCH-ACCURACY: Label projection on CeNGEN (split by batch), Accuracy

accuracy (fraction) · Higher values are better.

Open Problems label projection CENGEN-BATCH-ACCURACY: Label projection on CeNGEN (split by batch), Accuracy · CeNGEN (split by batch) (Open Problems label projection split)

Evidence origin: Author-reported evaluation.

openproblems-label primary benchmark evidence · results, dataset(cengen_batch), metric(accuracy)
  • Results published by the Open Problems project, source checked but not independently reproduced.
  • This covers the label projection task at v1.0.0 only, not the whole Open Problems suite.
  • true_labels and random_labels are controls that bound the scale, not competing methods.
  • Preprocessing is part of the run, so the same method appears once per parameter set.
Comparison details and limitations

Every method Open Problems label projection reports on Label projection on CeNGEN (split by batch), Accuracy, scored with Accuracy on CeNGEN (split by batch).

Automated source review: 2026-09-18. Numerical source review does not establish independent reproduction.

Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.

Showing 12 of 16 matching rows.

Methods and evaluation design

Procedure, tasks and evaluated configurations

How it works

Evaluation methodology

Open Problems is a collection of separately versioned single-cell evaluation tasks. Each task defines its inputs, reference data, methods, controls and metrics. For example, label projection learns from reference labels and predicts a held-out dataset, whereas integration measures how supplied batches are combined.

Sourcesopenproblems-label primary benchmark evidence · Label Projection v1.0.0: task description, dataset variants and controls

Evaluation design

Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.

Tasks

These source-backed links do not make different protocols or scores interchangeable.

Baseline coverage

Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.

No concrete protocols are explicitly linked to this suite. Protocol identification and baseline selection are outstanding.

Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums

Coverage is derived from release 2026-09-29-06401fd5b220. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.

Run this benchmark

Open Problems links task-specific repositories and contribution documentation. The inspected label-projection repository specifies AnnData inputs, prediction outputs and evaluation components, but these READMEs do not supply a complete end-to-end shell run. Select a released task workflow, executor and resource bundle before running; the platform itself has no universal single command.

A maintained rewire runner has not been verified for this benchmark. Check data access, weights, licences, dependencies and hardware in the linked official documentation; requirements have not been fully extracted.

openproblems-bio/openproblems / README.md; openproblems-bio/task_label_projection / README.md · Platform README.md lines 9–18; task_label_projection/README.md Description and API sections
Strengths, limitations and unresolved questions

Strengths and limitations

Strengths and considerations

Limitations and conditions

  • Results from different Open Problems tasks or split variants are not interchangeable. Perfect-label controls and random-label controls calibrate particular metrics; they are not measured biological performance ceilings.
    Sourcesopenproblems-label primary benchmark evidence · Label Projection v1.0.0: task description, dataset variants and controls
Profile review details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Stable record: discovery-benchmark-open-problems

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetsThe README links benchmark and dataset catalogues rather than fixing one data release.
Sourcesopenproblems-bio/openproblems official source · Pinned README: platform description and benchmark/dataset resource links
SplitsNo platform-wide split exists: select a hosted benchmark task and its dataset protocol. · Not applicable
Sourcesopenproblems-bio/openproblems official source · Pinned README: platform description and benchmark/dataset resource links
MetricsNo platform-wide scientific metric applies: hosted task definitions select their own evaluator. · Not applicable
Sourcesopenproblems-bio/openproblems official source · Pinned README: platform description and benchmark/dataset resource links
BaselinesNo single platform-wide baseline: methods and references belong to each hosted task. · Not applicable
Sourcesopenproblems-bio/openproblems official source · Pinned README: platform description and benchmark/dataset resource links
Leakage controlsControls are task dependent. Label Projection v1.0.0 distinguishes training/test batches and includes both random and batch-based CeNGEN dataset variants. Other tasks, such as batch integration, evaluate transductive processing of the supplied cells; a universal held-out-cell rule would be misleading.
Sourcesopenproblems-label primary benchmark evidence · Label Projection v1.0.0: task description, dataset variants and controls
UncertaintyThe inspected task pages and reporting configuration do not prescribe one platform-wide bootstrap or repeated-seed interval. Individual task versions define datasets, metrics and runs; model uncertainty from a method such as scANVI is not benchmark-score uncertainty. · Not reported in inspected sources
Sourcesopenproblems-label primary benchmark evidence · Label Projection v1.0.0: task description, dataset variants and controls
Entity typePlatform hosting computational biology benchmark tasks.
Sourcesopenproblems-bio/openproblems official source · Pinned README: platform description and benchmark/dataset resource links
OrganismsOrganisms belong to the selected task and dataset; the platform defines no unique organism. · Not applicable
Sourcesopenproblems-bio/openproblems official source · Pinned README: platform description and benchmark/dataset resource links
AssaysThe platform does not prescribe one assay. · Not applicable
Sourcesopenproblems-bio/openproblems official source · Pinned README: platform description and benchmark/dataset resource links
Allowed inputsInputs are specified by each hosted task.
Sourcesopenproblems-bio/openproblems official source · Pinned README: platform description and benchmark/dataset resource links
AdaptationAdaptation rules belong to the task implementation; the platform is not a fitted predictor. · Not applicable
Sourcesopenproblems-bio/openproblems official source · Pinned README: platform description and benchmark/dataset resource links

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Primary-paper discovery and source inspection. Source-checked results are not independently reproduced experiments.

Paper or primary resourceVersionReference
Defining and benchmarking open problems in single-cell analysisPMC full-text XML snapshot PMC11030530 at retrieval; byte-pinned by SHA-256Read source
DOI: 10.21203/rs.3.rs-4181617/v1
openproblems-label primary benchmark evidencev1.0.0Read source
Historical gaps recorded on 2026-09-17

The catalogue now holds 384 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.

  • Task release, dataset, preprocessing and metric must be fixed before charting.
  • The page lists multiple releases; a downloaded release artifact is needed to pin exact numerical leaderboard observations.
  • No values estimated from chart coordinates.
Search and extraction details

primary protocol screened

Searches

  • Open Problems single cell label projection benchmark v1.0.0 paper
  • CAPRI assessment rounds 46 54 protein docking results
  • GEARS predicting transcriptional outcomes multigene perturbations 2023
  • scVI scANVI reference mapping benchmark primary paper

Evidence locations

  • Primary paper: Label projection
  • Versioned label projection page: task definition, normalization, leaderboard

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

19 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Individual claims
openproblems-bio/openproblems official source

Original source ↗

Pinned README: platform description and benchmark/dataset resource links

Version: 0ca5d0cd040b741c1b6cc2e4cad7230cb2c50131
Retrieved: 2026-09-16T10:30:22.429615+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: ee26d791b60868c3e7701357831a49d8325f8d4d7cb0769b5177e5a67713114b

Hash scope: Hash scope not separately documented; inspect source record

Diagram steps
  • Allowed inputs: Inputs are specified by each hosted task.
  • Splits: No platform-wide split exists: select a hosted benchmark task and its dataset protocol.
  • Metrics: No platform-wide scientific metric applies: hosted task definitions select their own evaluator.
Individual claims
openproblems-bio/openproblems official source

Original source ↗

Pinned README: platform description and benchmark/dataset resource links

Version: 0ca5d0cd040b741c1b6cc2e4cad7230cb2c50131
Retrieved: 2026-09-16T10:30:22.429615+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: ee26d791b60868c3e7701357831a49d8325f8d4d7cb0769b5177e5a67713114b

Hash scope: Hash scope not separately documented; inspect source record

Diagram title
Evaluation procedure
Individual claims
openproblems-bio/openproblems official source

Original source ↗

Pinned README: platform description and benchmark/dataset resource links

Version: 0ca5d0cd040b741c1b6cc2e4cad7230cb2c50131
Retrieved: 2026-09-16T10:30:22.429615+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: ee26d791b60868c3e7701357831a49d8325f8d4d7cb0769b5177e5a67713114b

Hash scope: Hash scope not separately documented; inspect source record

Datasets
The README links benchmark and dataset catalogues rather than fixing one data release.
Individual claims
openproblems-bio/openproblems official source

Original source ↗

Pinned README: platform description and benchmark/dataset resource links

Version: 0ca5d0cd040b741c1b6cc2e4cad7230cb2c50131
Retrieved: 2026-09-16T10:30:22.429615+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: ee26d791b60868c3e7701357831a49d8325f8d4d7cb0769b5177e5a67713114b

Hash scope: Hash scope not separately documented; inspect source record

Splits
No platform-wide split exists: select a hosted benchmark task and its dataset protocol.
Individual claims
openproblems-bio/openproblems official source

Original source ↗

Pinned README: platform description and benchmark/dataset resource links

Version: 0ca5d0cd040b741c1b6cc2e4cad7230cb2c50131
Retrieved: 2026-09-16T10:30:22.429615+00:00

inapplicable

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: ee26d791b60868c3e7701357831a49d8325f8d4d7cb0769b5177e5a67713114b

Hash scope: Hash scope not separately documented; inspect source record

Adaptation
Adaptation rules belong to the task implementation; the platform is not a fitted predictor.
Individual claims
openproblems-bio/openproblems official source

Original source ↗

Pinned README: platform description and benchmark/dataset resource links

Version: 0ca5d0cd040b741c1b6cc2e4cad7230cb2c50131
Retrieved: 2026-09-16T10:30:22.429615+00:00

inapplicable

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: ee26d791b60868c3e7701357831a49d8325f8d4d7cb0769b5177e5a67713114b

Hash scope: Hash scope not separately documented; inspect source record

Metrics
No platform-wide scientific metric applies: hosted task definitions select their own evaluator.
Individual claims
openproblems-bio/openproblems official source

Original source ↗

Pinned README: platform description and benchmark/dataset resource links

Version: 0ca5d0cd040b741c1b6cc2e4cad7230cb2c50131
Retrieved: 2026-09-16T10:30:22.429615+00:00

inapplicable

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: ee26d791b60868c3e7701357831a49d8325f8d4d7cb0769b5177e5a67713114b

Hash scope: Hash scope not separately documented; inspect source record

Baselines
No single platform-wide baseline: methods and references belong to each hosted task.
Individual claims
openproblems-bio/openproblems official source

Original source ↗

Pinned README: platform description and benchmark/dataset resource links

Version: 0ca5d0cd040b741c1b6cc2e4cad7230cb2c50131
Retrieved: 2026-09-16T10:30:22.429615+00:00

inapplicable

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.3.value

Source artifact SHA-256: ee26d791b60868c3e7701357831a49d8325f8d4d7cb0769b5177e5a67713114b

Hash scope: Hash scope not separately documented; inspect source record

Leakage controls
Controls are task dependent. Label Projection v1.0.0 distinguishes training/test batches and includes both random and batch-based CeNGEN dataset variants. Other tasks, such as batch integration, evaluate transductive processing of the supplied cells; a universal held-out-cell rule would be misleading.
Individual claims
openproblems-label primary benchmark evidence

Original source ↗

Label Projection v1.0.0: task description, dataset variants and controls

Version: v1.0.0
Retrieved: 2026-09-16T21:16:30.026457+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: e223ab712ff55997a3abe659f280d4ea2952700e767b87e02c434701da9833c1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Uncertainty
The inspected task pages and reporting configuration do not prescribe one platform-wide bootstrap or repeated-seed interval. Individual task versions define datasets, metrics and runs; model uncertainty from a method such as scANVI is not benchmark-score uncertainty.
Individual claims
openproblems-label primary benchmark evidence

Original source ↗

Label Projection v1.0.0: task description, dataset variants and controls

Version: v1.0.0
Retrieved: 2026-09-16T21:16:30.026457+00:00

unreported

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.5.value

Source artifact SHA-256: e223ab712ff55997a3abe659f280d4ea2952700e767b87e02c434701da9833c1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: discovered

6 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: discovery-benchmark-open-problems

areas
single-cell
entity level
suite
scope note
Specialist molecular or omics evaluation; protocol details require review before numerical comparison.
task
Community single-cell analysis benchmarks
version
Not reported
benchmark research
review date: 2026-09-17; status: primary_protocol_screened; primary sources: expansion-p3-open-problems-paper; expansion-p3-open-problems; inspected locators: Primary paper: Label projection; Versioned label projection page: task definition, normalization, leaderboard; searched queries: Open Problems single cell label projection benchmark v1.0.0 paper; CAPRI assessment rounds 46 54 protein docking results; GEARS predicting transcriptional outcomes multigene perturbations 2023; scVI scANVI reference mapping benchmark primary paper; gaps: Task release, dataset, preprocessing and metric must be fixed before charting.; The page lists multiple releases; a downloaded release artifact is needed to pin exact numerical leaderboard observations.; No values estimated from chart coordinates.; claim scope: Primary-paper discovery and source inspection. Source-checked results are not independently reproduced experiments.
historical missing metadata
dataset release: unextracted; metric implementation: unextracted; split manifest: unextracted; version: unextracted
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
entity classification
review date: 2026-09-17; rationale: The cited profile describes a collection of evaluation tasks or protocols; retain it as the top-level benchmark suite. Its datasets and individual protocols remain separate records.; source ids: src-discovery-openproblems-bio-openproblems; source locator: Pinned README: platform description and benchmark/dataset resource links; ambiguities: None recorded
run documentation
record id: discovery-benchmark-open-problems; source ids: run-doc-open-problems-readme-md-0ca5d0cd; run-doc-openproblems-label-readme-md-e14ce24c; status: official_documentation_linked; summary: Open Problems links task-specific repositories and contribution documentation. The inspected label-projection repository specifies AnnData inputs, prediction outputs and evaluation components, but these READMEs do not supply a complete end-to-end shell run. Select a released task workflow, executor and resource bundle before running; the platform itself has no universal single command.; source locator: Platform README.md lines 9–18; task_label_projection/README.md Description and API sections
Related records

Suggest a correction