Datasets
Cytokine stimulation, L1000 and GTEx datasets with multipart metadata labels.
Combinatorial cell-label classification tests exact labels and partial correctness across single-cell and bulk-expression datasets.
Cytokine stimulation, L1000 and GTEx datasets with multipart metadata labels.
Accuracy and AUROC are reported separately for exact full-label matching and partial-label credit.
Gene-expression-derived cell representations.
Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.
Source reviewed · Automated source review, 2026-09-16. All specifications and missing details
Results are available, but no reviewed comparison panel is linked in this release.
2 evaluations · 2 results. Different protocols are not a single leaderboard.
Applied filters: All linked evaluations
| Tested configuration | Protocol and dataset | Finding | Evidence and details |
|---|---|---|---|
| Configuration: C2S (GPT-2 Large) | Task: Combinatorial cell-label classification Dataset: L1000 | 0.631 Partial-label accuracy unitless · unknown Uncertainty: ± 0.0031 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceC2S (GPT-2 Large): Combinatorial cell-label classification Partial-credit labels including cell type, perturbation, and dose. Aggregation: Not reported Cell2Sentence: Teaching Large Language Models the Language of Biology · Table 3, Partial label / C2S (GPT-2 Large) row, L1000 Acc column |
| Configuration: Geneformer | Task: Combinatorial cell-label classification Dataset: L1000 | 0.419 Partial-label accuracy unitless · unknown Uncertainty: ± 0.0153 Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceGeneformer: Combinatorial cell-label classification Partial-credit labels including cell type, perturbation, and dose. Aggregation: Not reported Cell2Sentence: Teaching Large Language Models the Language of Biology · Table 3, Partial label / Geneformer row, L1000 Acc column |
Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.
Cytokine stimulation, L1000 and GTEx datasets with multipart metadata labels. The cytokine label-classification setting includes all label combinations during training; held-out combinations belong to a separate generation task. Accuracy and AUROC are reported separately for exact full-label matching and partial-label credit. k-nearest-neighbour, XGBoost, Geneformer and scGPT comparators. For combinatorial label classification, all combinations of cytokine perturbations are included in training. Holding out 10 of 140 cytokine combinations belongs to the separate perturbed-cell-generation experiment, not this classifier. L1000 and GTEx provide bulk-expression evaluation outside the single-cell fine-tuning distribution.
Each evaluation records what was tested and under which conditions.
No runnable recipe has been reviewed for this task. Dataset access, model requirements, licences and compute requirements must be checked against its sources before execution.
A task describes a biological question. Choose a linked protocol to obtain concrete split and scoring instructions.
No source-reviewed explanatory claims are recorded here yet.
Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.
Stable record: reported-task-7efe245cc94ee5Explanatory profile: source reviewed · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Datasets | Cytokine stimulation, L1000 and GTEx datasets with multipart metadata labels.SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113 |
| Splits | The cytokine label-classification setting includes all label combinations during training; held-out combinations belong to a separate generation task.SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113 |
| Metrics | Accuracy and AUROC are reported separately for exact full-label matching and partial-label credit.SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113 |
| Baselines | k-nearest-neighbour, XGBoost, Geneformer and scGPT comparators.SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113 |
| Leakage controls | For combinatorial label classification, all combinations of cytokine perturbations are included in training. Holding out 10 of 140 cytokine combinations belongs to the separate perturbed-cell-generation experiment, not this classifier. L1000 and GTEx provide bulk-expression evaluation outside the single-cell fine-tuning distribution.SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Fine-Tuning Datasets; Experiment 2: combinatorial label classification |
| Uncertainty | Three experimental repeats are reported for each dataset.SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113 |
| Entity type | Paper-specific computational evaluation protocol.SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113 |
| Organisms | The cytokine-stimulation arm explicitly uses human peripheral blood mononuclear cells. The two additional arms use L1000 and GTEx bulk-expression collections, with their cell-line and tissue identities retained in combinatorial labels rather than treated as the same PBMC population.SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Experiments: Experiment 2/cell label prediction/Methodology; Fine-Tuning Datasets |
| Assays | Expression measurements with multipart cell/tissue/stimulation labels.SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113 |
| Allowed inputs | Gene-expression-derived cell representations.SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113 |
| Adaptation | Supervised label classification includes all label combinations in training; compositional generation is a different experiment.SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113 |
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Last literature check: 2026-09-17. Primary-paper discovery and source inspection. Source-checked results are not independently reproduced experiments.
| Paper or primary resource | Version | Reference |
|---|---|---|
| Cell2Sentence: Teaching Large Language Models the Language of Biology | preprint archived 2024-10-29 | Read source DOI: 10.1101/2023.09.11.557287 |
The catalogue now holds 2 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.
primary comparison table screened
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
17 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol. Individual claims | Cell2Sentence: Teaching Large Language Models the Language of Biology Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113 Version: preprint archived 2024-10-29 | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| Cell2Sentence: Teaching Large Language Models the Language of Biology Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113 Version: preprint archived 2024-10-29 | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Computational evaluation flow Individual claims | Cell2Sentence: Teaching Large Language Models the Language of Biology Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113 Version: preprint archived 2024-10-29 | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Datasets Cytokine stimulation, L1000 and GTEx datasets with multipart metadata labels. Individual claims | Cell2Sentence: Teaching Large Language Models the Language of Biology Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113 Version: preprint archived 2024-10-29 | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Splits The cytokine label-classification setting includes all label combinations during training; held-out combinations belong to a separate generation task. Individual claims | Cell2Sentence: Teaching Large Language Models the Language of Biology Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113 Version: preprint archived 2024-10-29 | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Adaptation Supervised label classification includes all label combinations in training; compositional generation is a different experiment. Individual claims | Cell2Sentence: Teaching Large Language Models the Language of Biology Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113 Version: preprint archived 2024-10-29 | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Metrics Accuracy and AUROC are reported separately for exact full-label matching and partial-label credit. Individual claims | Cell2Sentence: Teaching Large Language Models the Language of Biology Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113 Version: preprint archived 2024-10-29 | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Baselines k-nearest-neighbour, XGBoost, Geneformer and scGPT comparators. Individual claims | Cell2Sentence: Teaching Large Language Models the Language of Biology Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113 Version: preprint archived 2024-10-29 | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Leakage controls For combinatorial label classification, all combinations of cytokine perturbations are included in training. Holding out 10 of 140 cytokine combinations belongs to the separate perturbed-cell-generation experiment, not this classifier. L1000 and GTEx provide bulk-expression evaluation outside the single-cell fine-tuning distribution. Individual claims | Cell2Sentence: Teaching Large Language Models the Language of Biology Fine-Tuning Datasets; Experiment 2: combinatorial label classification Version: preprint archived 2024-10-29 | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Uncertainty Three experimental repeats are reported for each dataset. Individual claims | Cell2Sentence: Teaching Large Language Models the Language of Biology Experiment 2: cell label classification; Table 3; Fine-Tuning Datasets; cached text lines 47–53, 77–78, 113 Version: preprint archived 2024-10-29 | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
View linked audit checks and correction history
Release 2026-09-29-06401fd5b220 · Record review: needs review
Stable ID: reported-task-7efe245cc94ee5