rewire.itbenchmarks
Task

Combinatorial cell-label classification

Combinatorial cell-label classification tests exact labels and partial correctness across single-cell and bulk-expression datasets.

SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113

2 evaluations · 2 results

Overview

Datasets

Cytokine stimulation, L1000 and GTEx datasets with multipart metadata labels.

Metrics

Accuracy and AUROC are reported separately for exact full-label matching and partial-label credit.

Allowed inputs

Gene-expression-derived cell representations.

SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113
Evaluation procedure diagram
How it worksComputational evaluation flow
Computational evaluation flow1. Input: Gene-expression-derived cell representations.. Then: 2. Evaluation: The cytokine label-classification setting includes all label combinations during training; held-out combinations belong to a separate generation task.. Then: 3. Readout: Accuracy and AUROC are reported separately for exact full-label matching and partial-label credit.Computational evaluation flow1. Input: Gene-expression-derived cell representations.. Then: 2. Evaluation: The cytokine label-classification setting includes all label combinations during training; held-out combinations belong to a separate generation task.. Then: 3. Readout: Accuracy and AUROC are reported separately for exact full-label matching and partial-label credit.Computational evaluation flow1. Input: Gene-expression-derived cell representations.. Then: 2. Evaluation: The cytokine label-classification setting includes all label combinations during training; held-out combinations belong to a separate generation task.. Then: 3. Readout: Accuracy and AUROC are reported separately for exact full-label matching and partial-label credit.

Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.

SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113

Source reviewed · Automated source review, 2026-09-16. All specifications and missing details

Results

Results are available, but no reviewed comparison panel is linked in this release.

All evaluations

2 evaluations · 2 results. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: C2S (GPT-2 Large)Task: Combinatorial cell-label classification
Dataset: L1000
0.631 Partial-label accuracy
unitless · unknown

Uncertainty: ± 0.0031

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

C2S (GPT-2 Large): Combinatorial cell-label classification

Partial-credit labels including cell type, perturbation, and dose.

Aggregation: Not reported

Cell2Sentence: Teaching Large Language Models the Language of Biology · Table 3, Partial label / C2S (GPT-2 Large) row, L1000 Acc column
Configuration: GeneformerTask: Combinatorial cell-label classification
Dataset: L1000
0.419 Partial-label accuracy
unitless · unknown

Uncertainty: ± 0.0153

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

Geneformer: Combinatorial cell-label classification

Partial-credit labels including cell type, perturbation, and dose.

Aggregation: Not reported

Cell2Sentence: Teaching Large Language Models the Language of Biology · Table 3, Partial label / Geneformer row, L1000 Acc column

Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.

Methods and evaluation design

Procedure, tasks and evaluated configurations

How it works

Evaluation methodology

Cytokine stimulation, L1000 and GTEx datasets with multipart metadata labels. The cytokine label-classification setting includes all label combinations during training; held-out combinations belong to a separate generation task. Accuracy and AUROC are reported separately for exact full-label matching and partial-label credit. k-nearest-neighbour, XGBoost, Geneformer and scGPT comparators. For combinatorial label classification, all combinations of cytokine perturbations are included in training. Holding out 10 of 140 cytokine combinations belongs to the separate perturbed-cell-generation experiment, not this classifier. L1000 and GTEx provide bulk-expression evaluation outside the single-cell fine-tuning distribution.

SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113; Fine-Tuning Datasets; Experiment 2: combinatorial label classification

Recorded evaluations

Each evaluation records what was tested and under which conditions.

Run instructions

No runnable recipe has been reviewed for this task. Dataset access, model requirements, licences and compute requirements must be checked against its sources before execution.

A task describes a biological question. Choose a linked protocol to obtain concrete split and scoring instructions.

Strengths, limitations and unresolved questions

Strengths and limitations

Strengths supported by sources

No source-reviewed explanatory claims are recorded here yet.

Limitations and conditions

  • The combinatorial classifier includes every cytokine-label combination during training. A held-out-combination claim belongs to the separate generation task; exact and partial label credit are also different metrics.
    SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Fine-Tuning Datasets; Experiment 2: combinatorial label classification; Experiments: Experiment 2/cell label prediction/Methodology; Fine-Tuning Datasets
Profile review details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Stable record: reported-task-7efe245cc94ee5

Specifications

Inputs, training, access and other details

Explanatory profile: source reviewed · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetsCytokine stimulation, L1000 and GTEx datasets with multipart metadata labels.
SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113
SplitsThe cytokine label-classification setting includes all label combinations during training; held-out combinations belong to a separate generation task.
SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113
MetricsAccuracy and AUROC are reported separately for exact full-label matching and partial-label credit.
SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113
Baselinesk-nearest-neighbour, XGBoost, Geneformer and scGPT comparators.
SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113
Leakage controlsFor combinatorial label classification, all combinations of cytokine perturbations are included in training. Holding out 10 of 140 cytokine combinations belongs to the separate perturbed-cell-generation experiment, not this classifier. L1000 and GTEx provide bulk-expression evaluation outside the single-cell fine-tuning distribution.
SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Fine-Tuning Datasets; Experiment 2: combinatorial label classification
UncertaintyThree experimental repeats are reported for each dataset.
SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113
Entity typePaper-specific computational evaluation protocol.
SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113
OrganismsThe cytokine-stimulation arm explicitly uses human peripheral blood mononuclear cells. The two additional arms use L1000 and GTEx bulk-expression collections, with their cell-line and tissue identities retained in combinatorial labels rather than treated as the same PBMC population.
SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Experiments: Experiment 2/cell label prediction/Methodology; Fine-Tuning Datasets
AssaysExpression measurements with multipart cell/tissue/stimulation labels.
SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113
Allowed inputsGene-expression-derived cell representations.
SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113
AdaptationSupervised label classification includes all label combinations in training; compositional generation is a different experiment.
SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Primary-paper discovery and source inspection. Source-checked results are not independently reproduced experiments.

Paper or primary resourceVersionReference
Cell2Sentence: Teaching Large Language Models the Language of Biologypreprint archived 2024-10-29Read source
DOI: 10.1101/2023.09.11.557287
Historical gaps recorded on 2026-09-17

The catalogue now holds 2 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.

  • Partial-credit and full-label metrics cannot share a ranking; keep dataset and label regime explicit.
  • Version is the archived preprint, not a later model family release.
Search and extraction details

primary comparison table screened

Searches

  • "PMC11565894"

Evidence locations

  • Table 3
  • Fine-Tuning Datasets
  • Classification task methodology

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

17 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.
Individual claims
Cell2Sentence: Teaching Large Language Models the Language of Biology

Original source ↗

Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113

Version: preprint archived 2024-10-29
Retrieved: 2026-09-16T10:41:16.533640+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: e088727d6e04857fccb7033a9b074e1850f775e86e7d2e99e603dde09558ab02

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps
  • Input: Gene-expression-derived cell representations.
  • Evaluation: The cytokine label-classification setting includes all label combinations during training; held-out combinations belong to a separate generation task.
  • Readout: Accuracy and AUROC are reported separately for exact full-label matching and partial-label credit.
Individual claims
Cell2Sentence: Teaching Large Language Models the Language of Biology

Original source ↗

Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113

Version: preprint archived 2024-10-29
Retrieved: 2026-09-16T10:41:16.533640+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: e088727d6e04857fccb7033a9b074e1850f775e86e7d2e99e603dde09558ab02

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title
Computational evaluation flow
Individual claims
Cell2Sentence: Teaching Large Language Models the Language of Biology

Original source ↗

Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113

Version: preprint archived 2024-10-29
Retrieved: 2026-09-16T10:41:16.533640+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.diagram.title

Source artifact SHA-256: e088727d6e04857fccb7033a9b074e1850f775e86e7d2e99e603dde09558ab02

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Datasets
Cytokine stimulation, L1000 and GTEx datasets with multipart metadata labels.
Individual claims
Cell2Sentence: Teaching Large Language Models the Language of Biology

Original source ↗

Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113

Version: preprint archived 2024-10-29
Retrieved: 2026-09-16T10:41:16.533640+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: e088727d6e04857fccb7033a9b074e1850f775e86e7d2e99e603dde09558ab02

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Splits
The cytokine label-classification setting includes all label combinations during training; held-out combinations belong to a separate generation task.
Individual claims
Cell2Sentence: Teaching Large Language Models the Language of Biology

Original source ↗

Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113

Version: preprint archived 2024-10-29
Retrieved: 2026-09-16T10:41:16.533640+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: e088727d6e04857fccb7033a9b074e1850f775e86e7d2e99e603dde09558ab02

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Adaptation
Supervised label classification includes all label combinations in training; compositional generation is a different experiment.
Individual claims
Cell2Sentence: Teaching Large Language Models the Language of Biology

Original source ↗

Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113

Version: preprint archived 2024-10-29
Retrieved: 2026-09-16T10:41:16.533640+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: e088727d6e04857fccb7033a9b074e1850f775e86e7d2e99e603dde09558ab02

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Metrics
Accuracy and AUROC are reported separately for exact full-label matching and partial-label credit.
Individual claims
Cell2Sentence: Teaching Large Language Models the Language of Biology

Original source ↗

Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113

Version: preprint archived 2024-10-29
Retrieved: 2026-09-16T10:41:16.533640+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: e088727d6e04857fccb7033a9b074e1850f775e86e7d2e99e603dde09558ab02

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Baselines
k-nearest-neighbour, XGBoost, Geneformer and scGPT comparators.
Individual claims
Cell2Sentence: Teaching Large Language Models the Language of Biology

Original source ↗

Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113

Version: preprint archived 2024-10-29
Retrieved: 2026-09-16T10:41:16.533640+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.3.value

Source artifact SHA-256: e088727d6e04857fccb7033a9b074e1850f775e86e7d2e99e603dde09558ab02

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Leakage controls
For combinatorial label classification, all combinations of cytokine perturbations are included in training. Holding out 10 of 140 cytokine combinations belongs to the separate perturbed-cell-generation experiment, not this classifier. L1000 and GTEx provide bulk-expression evaluation outside the single-cell fine-tuning distribution.
Individual claims
Cell2Sentence: Teaching Large Language Models the Language of Biology

Original source ↗

Fine-Tuning Datasets; Experiment 2: combinatorial label classification

Version: preprint archived 2024-10-29
Retrieved: 2026-09-16T10:41:16.533640+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: e088727d6e04857fccb7033a9b074e1850f775e86e7d2e99e603dde09558ab02

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Uncertainty
Three experimental repeats are reported for each dataset.
Individual claims
Cell2Sentence: Teaching Large Language Models the Language of Biology

Original source ↗

Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113

Version: preprint archived 2024-10-29
Retrieved: 2026-09-16T10:41:16.533640+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.5.value

Source artifact SHA-256: e088727d6e04857fccb7033a9b074e1850f775e86e7d2e99e603dde09558ab02

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: needs review

2 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: reported-task-7efe245cc94ee5

areas
cells-tissues
tasks
Combinatorial cell-label classification
entity level
task
version
Not reported
task
Combinatorial cell-label classification
scope note
Paper-specific evaluation task; protocol completeness requires further extraction.
benchmark research
review date: 2026-09-17; status: primary_comparison_table_screened; primary sources: expansion-p3-cell2sentence-2024; inspected locators: Table 3; Fine-Tuning Datasets; Classification task methodology; searched queries: "PMC11565894"; gaps: Partial-credit and full-label metrics cannot share a ranking; keep dataset and label regime explicit.; Version is the archived preprint, not a later model family release.; claim scope: Primary-paper discovery and source inspection. Source-checked results are not independently reproduced experiments.
historical missing metadata
protocol version: not_reported_in_legacy_extract; split: not_reported_in_legacy_extract
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
legacy kinds
benchmark
entity classification
review date: 2026-09-17; rationale: This source-scoped record identifies the biological prediction task and holds its paper context. Preserve the existing task identity; exact split, model adaptation and scoring remain in linked evaluations or separate protocol records.; source ids: cell2sentence-2024; source locator: Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113; ambiguities: A paper- or suite-specific task may constrain some inputs or metrics; that alone does not make it interchangeable with a complete versioned protocol. No protocol equivalence is inferred.; Some legacy profile Entity type facts use the generic phrase computational evaluation protocol. That boilerplate is not sufficient to establish a single fixed protocol identity or to merge this task with another protocol record.
Related records

Suggest a correction