rewire.itbenchmarks
Task

Cell-type annotation

Cell-type annotation separates cross-batch prediction from a within-study random split.

SourcesscELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis · Results: Cell-type annotation; Methods: evaluation and baselines; cached text lines 23, 81, 96

2 evaluations · 2 results

Overview

Datasets

hPancreas, PBMC and Aorta annotated single-cell datasets.

Metrics

Accuracy, precision, recall and F1.

Allowed inputs

Expression-derived representations and GPT-based gene information.

SourcesscELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis · Results: Cell-type annotation; Methods: evaluation and baselines; cached text lines 23, 81, 96
Evaluation procedure diagram
How it worksComputational evaluation flow
Computational evaluation flow1. Input: Expression-derived representations and GPT-based gene information.. Then: 2. Evaluation: hPancreas/PBMC withhold one batch; Aorta uses an 80:20 within-study split.. Then: 3. Readout: Accuracy, precision, recall and F1.Computational evaluation flow1. Input: Expression-derived representations and GPT-based gene information.. Then: 2. Evaluation: hPancreas/PBMC withhold one batch; Aorta uses an 80:20 within-study split.. Then: 3. Readout: Accuracy, precision, recall and F1.Computational evaluation flow1. Input: Expression-derived representations and GPT-based gene information.. Then: 2. Evaluation: hPancreas/PBMC withhold one batch; Aorta uses an 80:20 within-study split.. Then: 3. Readout: Accuracy, precision, recall and F1.

Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.

SourcesscELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis · Results: Cell-type annotation; Methods: evaluation and baselines; cached text lines 23, 81, 96

limited source coverage · Automated source review, 2026-09-16. All specifications and missing details

Results

Results are available, but no reviewed comparison panel is linked in this release.

All evaluations

2 evaluations · 2 results. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: scGPTTask: Cell-type annotation
Dataset: hPancreas
0.55 F1
unitless · unknown

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Result quoted from another source · Source checked
Methods, coverage and source

scGPT: Cell-type annotation

Zero-shot setting; source caption says some comparator rows come from GenePT.

Aggregation: Not reported

scELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis · Table 1, hPancreas zero-shot / scGPT (z) row, F1 column
Configuration: GeneformerTask: Cell-type annotation
Dataset: hPancreas
0.27 F1
unitless · unknown

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Result quoted from another source · Source checked
Methods, coverage and source

Geneformer: Cell-type annotation

Zero-shot setting; source caption says some comparator rows come from GenePT.

Aggregation: Not reported

scELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis · Table 1, hPancreas zero-shot / Geneformer (z) row, F1 column

Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.

Methods and evaluation design

Procedure, tasks and evaluated configurations

How it works

Evaluation methodology

hPancreas, PBMC and Aorta annotated single-cell datasets. hPancreas/PBMC withhold one batch; Aorta uses an 80:20 within-study split. Accuracy, precision, recall and F1. GPT-based classifiers, scGPT, Geneformer, GPTCelltype, MLP and PCA-derived representations. Batch-held-out evaluation is explicit for two datasets; it is not the split used for Aorta.

SourcesscELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis · Results: Cell-type annotation; Methods: evaluation and baselines; cached text lines 23, 81, 96

Recorded evaluations

Each evaluation records what was tested and under which conditions.

Run instructions

No runnable recipe has been reviewed for this task. Dataset access, model requirements, licences and compute requirements must be checked against its sources before execution.

A task describes a biological question. Choose a linked protocol to obtain concrete split and scoring instructions.

Strengths, limitations and unresolved questions

Strengths and limitations

Profile review details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Stable record: reported-task-660753ec94e631

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetshPancreas, PBMC and Aorta annotated single-cell datasets.
SourcesscELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis · Results: Cell-type annotation; Methods: evaluation and baselines; cached text lines 23, 81, 96
SplitshPancreas/PBMC withhold one batch; Aorta uses an 80:20 within-study split.
SourcesscELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis · Results: Cell-type annotation; Methods: evaluation and baselines; cached text lines 23, 81, 96
MetricsAccuracy, precision, recall and F1.
SourcesscELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis · Results: Cell-type annotation; Methods: evaluation and baselines; cached text lines 23, 81, 96
BaselinesGPT-based classifiers, scGPT, Geneformer, GPTCelltype, MLP and PCA-derived representations.
SourcesscELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis · Results: Cell-type annotation; Methods: evaluation and baselines; cached text lines 23, 81, 96
Leakage controlsBatch-held-out evaluation is explicit for two datasets; it is not the split used for Aorta.
SourcesscELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis · Results: Cell-type annotation; Methods: evaluation and baselines; cached text lines 23, 81, 96
UncertaintyTable 1 reports point metrics for annotation and identifies some rows as copied from GenePT. The cell-annotation section and table do not give repeated-run uncertainty or confidence intervals for the scELMo rows; copied comparator results must not be counted as independent replications. · Not reported in inspected sources
SourcesscELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis · Results: scELMo for cell-type annotation; Table 1 caption
Entity typePaper-specific computational evaluation protocol.
SourcesscELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis · Results: Cell-type annotation; Methods: evaluation and baselines; cached text lines 23, 81, 96
OrganismsHuman pancreas, PBMC and Aorta datasets.
SourcesscELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis · Results: Cell-type annotation; Methods: evaluation and baselines; cached text lines 23, 81, 96
AssaysSingle-cell expression and cell-type labels.
SourcesscELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis · Results: Cell-type annotation; Methods: evaluation and baselines; cached text lines 23, 81, 96
Allowed inputsExpression-derived representations and GPT-based gene information.
SourcesscELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis · Results: Cell-type annotation; Methods: evaluation and baselines; cached text lines 23, 81, 96
AdaptationSupervised annotation and classifier comparisons, with batch holdout where available.
SourcesscELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis · Results: Cell-type annotation; Methods: evaluation and baselines; cached text lines 23, 81, 96

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Dated primary-source discovery and protocol/table screening. Source checking does not mean experimental reproduction. Only separately extracted and independently reviewed numeric batches are publishable.

Paper or primary resourceVersionReference
scELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysispreprint archived 2025-08-23Read source
Historical gaps recorded on 2026-09-17

The catalogue now holds 2 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.

  • complete numerical transcription and independent cell review: Full primary artifact and table inventory preserved; no new numeric row is published from this audit alone.
  • exact checkpoint hashes and per-method scored denominators: Table labels alone do not establish these fields; do not infer checkpoint or scored count from model name or dataset size.
Search and extraction details

primary comparison tables located

Searches

  • scELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis 10.1101/2023.12.07.569910

Evidence locations

  • Table 1; XML table T1

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

18 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.
Individual claims
scELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis

Original source ↗

Results: Cell-type annotation; Methods: evaluation and baselines; cached text lines 23, 81, 96

Version: preprint archived 2025-08-23
Retrieved: 2026-09-16T10:41:16.537541+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: ef75f0d63a567f5e9d7132fd847f44838a82a9741ae55323437e1d1812d86316

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps
  • Input: Expression-derived representations and GPT-based gene information.
  • Evaluation: hPancreas/PBMC withhold one batch; Aorta uses an 80:20 within-study split.
  • Readout: Accuracy, precision, recall and F1.
Individual claims
scELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis

Original source ↗

Results: Cell-type annotation; Methods: evaluation and baselines; cached text lines 23, 81, 96

Version: preprint archived 2025-08-23
Retrieved: 2026-09-16T10:41:16.537541+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: ef75f0d63a567f5e9d7132fd847f44838a82a9741ae55323437e1d1812d86316

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title
Computational evaluation flow
Individual claims
scELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis

Original source ↗

Results: Cell-type annotation; Methods: evaluation and baselines; cached text lines 23, 81, 96

Version: preprint archived 2025-08-23
Retrieved: 2026-09-16T10:41:16.537541+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.diagram.title

Source artifact SHA-256: ef75f0d63a567f5e9d7132fd847f44838a82a9741ae55323437e1d1812d86316

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Datasets
hPancreas, PBMC and Aorta annotated single-cell datasets.
Individual claims
scELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis

Original source ↗

Results: Cell-type annotation; Methods: evaluation and baselines; cached text lines 23, 81, 96

Version: preprint archived 2025-08-23
Retrieved: 2026-09-16T10:41:16.537541+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: ef75f0d63a567f5e9d7132fd847f44838a82a9741ae55323437e1d1812d86316

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Splits
hPancreas/PBMC withhold one batch; Aorta uses an 80:20 within-study split.
Individual claims
scELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis

Original source ↗

Results: Cell-type annotation; Methods: evaluation and baselines; cached text lines 23, 81, 96

Version: preprint archived 2025-08-23
Retrieved: 2026-09-16T10:41:16.537541+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: ef75f0d63a567f5e9d7132fd847f44838a82a9741ae55323437e1d1812d86316

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Adaptation
Supervised annotation and classifier comparisons, with batch holdout where available.
Individual claims
scELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis

Original source ↗

Results: Cell-type annotation; Methods: evaluation and baselines; cached text lines 23, 81, 96

Version: preprint archived 2025-08-23
Retrieved: 2026-09-16T10:41:16.537541+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: ef75f0d63a567f5e9d7132fd847f44838a82a9741ae55323437e1d1812d86316

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Metrics
Accuracy, precision, recall and F1.
Individual claims
scELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis

Original source ↗

Results: Cell-type annotation; Methods: evaluation and baselines; cached text lines 23, 81, 96

Version: preprint archived 2025-08-23
Retrieved: 2026-09-16T10:41:16.537541+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: ef75f0d63a567f5e9d7132fd847f44838a82a9741ae55323437e1d1812d86316

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Baselines
GPT-based classifiers, scGPT, Geneformer, GPTCelltype, MLP and PCA-derived representations.
Individual claims
scELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis

Original source ↗

Results: Cell-type annotation; Methods: evaluation and baselines; cached text lines 23, 81, 96

Version: preprint archived 2025-08-23
Retrieved: 2026-09-16T10:41:16.537541+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.3.value

Source artifact SHA-256: ef75f0d63a567f5e9d7132fd847f44838a82a9741ae55323437e1d1812d86316

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Leakage controls
Batch-held-out evaluation is explicit for two datasets; it is not the split used for Aorta.
Individual claims
scELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis

Original source ↗

Results: Cell-type annotation; Methods: evaluation and baselines; cached text lines 23, 81, 96

Version: preprint archived 2025-08-23
Retrieved: 2026-09-16T10:41:16.537541+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: ef75f0d63a567f5e9d7132fd847f44838a82a9741ae55323437e1d1812d86316

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Uncertainty
Table 1 reports point metrics for annotation and identifies some rows as copied from GenePT. The cell-annotation section and table do not give repeated-run uncertainty or confidence intervals for the scELMo rows; copied comparator results must not be counted as independent replications.
Individual claims
scELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis

Original source ↗

Results: scELMo for cell-type annotation; Table 1 caption

Version: preprint archived 2025-08-23
Retrieved: 2026-09-16T10:41:16.537541+00:00

unreported

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.5.value

Source artifact SHA-256: ef75f0d63a567f5e9d7132fd847f44838a82a9741ae55323437e1d1812d86316

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: needs review

2 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: reported-task-660753ec94e631

areas
cells-tissues
tasks
Cell-type annotation
entity level
task
version
Not reported
task
Cell-type annotation
scope note
Paper-specific evaluation task; protocol completeness requires further extraction.
benchmark research
review date: 2026-09-17; status: primary_comparison_tables_located; primary sources: evidence-expansion-scelmo-2025-ef75f0d6; inspected locators: Table 1; XML table T1; searched queries: scELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis 10.1101/2023.12.07.569910; gaps: complete numerical transcription and independent cell review: Full primary artifact and table inventory preserved; no new numeric row is published from this audit alone.; exact checkpoint hashes and per-method scored denominators: Table labels alone do not establish these fields; do not infer checkpoint or scored count from model name or dataset size.; claim scope: Dated primary-source discovery and protocol/table screening. Source checking does not mean experimental reproduction. Only separately extracted and independently reviewed numeric batches are publishable.
historical missing metadata
protocol version: not_reported_in_legacy_extract; split: not_reported_in_legacy_extract
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
legacy kinds
benchmark
entity classification
review date: 2026-09-17; rationale: This source-scoped record identifies the biological prediction task and holds its paper context. Preserve the existing task identity; exact split, model adaptation and scoring remain in linked evaluations or separate protocol records.; source ids: scelmo-2025; source locator: Results: Cell-type annotation; Methods: evaluation and baselines; cached text lines 23, 81, 96; ambiguities: A paper- or suite-specific task may constrain some inputs or metrics; that alone does not make it interchangeable with a complete versioned protocol. No protocol equivalence is inferred.; Some legacy profile Entity type facts use the generic phrase computational evaluation protocol. That boilerplate is not sufficient to establish a single fixed protocol identity or to merge this task with another protocol record.
Related records

Suggest a correction