Datasets
The released MassSpecGym MS/MS/molecule dataset, with task-specific inputs and candidate sets.
MassSpecGym separates spectrum-to-structure generation, candidate retrieval and structure-to-spectrum simulation.
The released MassSpecGym MS/MS/molecule dataset, with task-specific inputs and candidate sets.
Task evaluators distinguish molecular exact match/structural similarity, candidate-retrieval hit rate and spectrum similarity. These are separate readouts, not interchangeable scores.
Spectrum-to-molecule, spectrum-plus-candidates, or molecule-to-spectrum inputs depend on the selected task.
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Source reviewed · Automated source review, 2026-09-16. All specifications and missing details
Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.
Top-1 accuracy (fraction) · Higher values are better.
MassSpecGym · main (MassSpecGym De novo molecule generation) · MassSpecGym · main
Evidence origin: Author-reported evaluation.
MassSpecGym: A benchmark for the discovery and identification of molecules · Table 2: Top-1 accuracy, mainGenerate candidate molecules from an input spectrum. Compare within the same main/formula challenge and metric. Main random generation uses precursor mass; the Transformer consumes the spectrum. MCES single-linkage molecular clustering at threshold 10; fixed held-out test split
Automated source review: 2026-09-17. Numerical source review does not establish independent reproduction.
Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.
Showing 3 of 3 matching rows.
The released MassSpecGym MS/MS/molecule dataset, with task-specific inputs and candidate sets. MCES molecular clusters are grouped into fixed training, validation and test folds, stratified by acquisition metadata. Cross-fold molecular bond-edit distance is at least 10; all spectra follow the assigned molecular fold. Task evaluators distinguish molecular exact match/structural similarity, candidate-retrieval hit rate and spectrum similarity. These are separate readouts, not interchangeable scores. Maximum common edge subgraph (MCES) clustering keeps molecules connected by a bond-edit distance below 10 in the same fold. The split additionally balances instrument, collision-energy, adduct and molecule-frequency metadata; this is stronger than simply separating 2D InChIKeys.
Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.
These source-backed links do not make different protocols or scores interchangeable.
Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.
0 of 12 active baseline roles have published Rewire measurements in this release. Measurements on a selected protocol do not establish coverage of an entire suite.
Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums
Coverage is derived from release 2026-09-29-06401fd5b220. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.
Choose a concrete protocol before running an evaluation. Its inputs, split and scoring rules determine which results can be compared.
Install the package and load the benchmark dataset that the scores on this page are measured over.
Generate predictions and evaluate them. This recipe does not establish reproduction of a particular published score.
Source reviewed; these instructions have not been executed by rewire.
pip install massspecgymMassSpecGym: repository README · README.md at f259fe37, Installation, lines 43-43Source reviewed; these instructions have not been executed by rewire.
from massspecgym.utils import load_massspecgym
df = load_massspecgym()MassSpecGym: repository README · README.md at f259fe37, Getting started with MassSpecGym, lines 76-77Run your model locally and return predictions keyed by the input IDs. The evaluator supplies biological inputs without test labels and owns scoring. This interface is not a sandbox for model code.
Pass your existing prediction function into this adapter. Its output direction must match the selected protocol.
class MyModelAdapter:
def __init__(self, score):
self.score = score
def predict(self, inputs):
return {row["id"]: float(self.score(row)) for row in inputs}
# adapter = MyModelAdapter(your_prediction_function)
# report = rewirebench.run(prepared, adapter, output="runs/my-model")Alternatively, generate a keyed prediction file in your existing model environment and use the score-only recipe. Your model code and weights do not need to be shared.
MassSpecGym: repository README · README.md at f259fe37Contribute a result for review. The library can submit an exported evaluation for private review when intake is open. Check the contribution page for access and sign-in.
Run the official CPU DeepSets retrieval tutorial: train a small spectrum-to-fingerprint model, then evaluate retrieval on the provided test split.
Checked against the official instructions on 2026-09-17. These commands have not been executed by rewire. Running them does not automatically reproduce the published scores.
Repository checkout wrapper: the detached revision selects the exact official source inspected for this guide.
git clone https://github.com/pluskal-lab/MassSpecGym.git
cd MassSpecGym
git checkout --detach f259fe3780d5bd227fc6ece36ce6f397c2eef716pluskal-lab/MassSpecGym / README.md · Pinned repository revision; README.mdThe README package-install command is unversioned. Record the installed version; the pinned documentation commit is not an installed-package lock.
conda create -n massspecgym python==3.11
conda activate massspecgym
pip install massspecgympluskal-lab/MassSpecGym / README.md · README.md lines 38–51The four official Python blocks are combined in a local script. A standard main guard is added around execution for the documented four data workers; model, CPU trainer, one device, five epochs and batch size 32 are unchanged.
cat > massspecgym_retrieval_example.py <<'PY'
import torch
import torch.nn as nn
import pytorch_lightning as pl
from pytorch_lightning import Trainer
from massspecgym.data import RetrievalDataset, MassSpecDataModule
from massspecgym.data.transforms import SpecTokenizer, MolFingerprinter
from massspecgym.models.base import Stage
from massspecgym.models.retrieval.base import RetrievalMassSpecGymModel
class MyDeepSetsRetrievalModel(RetrievalMassSpecGymModel):
def __init__(
self,
hidden_channels: int = 128,
out_channels: int = 4096, # fingerprint size
*args,
**kwargs
):
"""Implement your architecture."""
super().__init__(*args, **kwargs)
self.phi = nn.Sequential(
nn.Linear(2, hidden_channels),
nn.ReLU(),
nn.Linear(hidden_channels, hidden_channels),
nn.ReLU(),
)
self.rho = nn.Sequential(
nn.Linear(hidden_channels, hidden_channels),
nn.ReLU(),
nn.Linear(hidden_channels, out_channels),
nn.Sigmoid()
)
def forward(self, x: torch.Tensor) -> torch.Tensor:
"""Implement your prediction logic."""
x = self.phi(x)
x = x.sum(dim=-2) # sum over peaks
x = self.rho(x)
return x
def step(
self, batch: dict, stage: Stage
) -> tuple[torch.Tensor, torch.Tensor]:
"""Implement your custom logic of using predictions for training and inference."""
# Unpack inputs
x = batch["spec"] # input spectra
fp_true = batch["mol"] # true fingerprints
cands = batch["candidates"] # candidate fingerprints concatenated for a batch
batch_ptr = batch["batch_ptr"] # number of candidates per sample in a batch
# Predict fingerprint
fp_pred = self.forward(x)
# Calculate loss
loss = nn.functional.mse_loss(fp_true, fp_pred)
# Calculate final similarity scores between predicted fingerprints and retrieval candidates
fp_pred_repeated = fp_pred.repeat_interleave(batch_ptr, dim=0)
scores = nn.functional.cosine_similarity(fp_pred_repeated, cands)
return dict(loss=loss, scores=scores)
if __name__ == '__main__':
# Init hyperparameters
n_peaks = 60
fp_size = 4096
batch_size = 32
# Load dataset
dataset = RetrievalDataset(
spec_transform=SpecTokenizer(n_peaks=n_peaks),
mol_transform=MolFingerprinter(fp_size=fp_size),
)
# Init data module
data_module = MassSpecDataModule(
dataset=dataset,
batch_size=batch_size,
num_workers=4
)
# Init model
model = MyDeepSetsRetrievalModel(out_channels=fp_size)
# Init trainer
trainer = Trainer(accelerator="cpu", devices=1, max_epochs=5)
# Train
trainer.fit(model, datamodule=data_module)
# Test
trainer.test(model, datamodule=data_module)
PY
python massspecgym_retrieval_example.pypluskal-lab/MassSpecGym / README.md · README.md lines 111–218Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.
Stable record: discovery-benchmark-massspecgymExplanatory profile: source reviewed · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Datasets | The released MassSpecGym MS/MS/molecule dataset, with task-specific inputs and candidate sets.Sources (4)pluskal-lab/MassSpecGym official source; pluskal-lab/MassSpecGym massspecgym/models/de_novo/base.py; pluskal-lab/MassSpecGym massspecgym/models/retrieval/base.py; pluskal-lab/MassSpecGym massspecgym/models/simulation/base.py · Pinned README: three challenges; dataset and DataModule; evaluation base classes; pinned de_novo, retrieval and simulation base evaluation classes |
| Splits | MCES molecular clusters are grouped into fixed training, validation and test folds, stratified by acquisition metadata. Cross-fold molecular bond-edit distance is at least 10; all spectra follow the assigned molecular fold.Sourcesmassspecgym primary benchmark evidence · Section 3.4; Supplementary Information 2.5 |
| Metrics | Task evaluators distinguish molecular exact match/structural similarity, candidate-retrieval hit rate and spectrum similarity. These are separate readouts, not interchangeable scores.Sources (4)pluskal-lab/MassSpecGym official source; pluskal-lab/MassSpecGym massspecgym/models/de_novo/base.py; pluskal-lab/MassSpecGym massspecgym/models/retrieval/base.py; pluskal-lab/MassSpecGym massspecgym/models/simulation/base.py · Pinned README: three challenges; dataset and DataModule; evaluation base classes; pinned de_novo, retrieval and simulation base evaluation classes |
| Baselines | The README illustrates a DeepSets-style spectrum-to-fingerprint retrieval baseline.Sources (4)pluskal-lab/MassSpecGym official source; pluskal-lab/MassSpecGym massspecgym/models/de_novo/base.py; pluskal-lab/MassSpecGym massspecgym/models/retrieval/base.py; pluskal-lab/MassSpecGym massspecgym/models/simulation/base.py · Pinned README: three challenges; dataset and DataModule; evaluation base classes; pinned de_novo, retrieval and simulation base evaluation classes |
| Leakage controls | Maximum common edge subgraph (MCES) clustering keeps molecules connected by a bond-edit distance below 10 in the same fold. The split additionally balances instrument, collision-energy, adduct and molecule-frequency metadata; this is stronger than simply separating 2D InChIKeys.Sourcesmassspecgym primary benchmark evidence · Section 3.4; Supplementary Information 2.5; Tables 2–4 |
| Uncertainty | Tables 2–4 report 99.9% bootstrap confidence intervals using 20,000 resamples. These intervals summarize test-example sampling, not variation across independently retrained models.Sourcesmassspecgym primary benchmark evidence · Section 3.4; Supplementary Information 2.5; Tables 2–4 |
| Entity type | Small-molecule MS/MS benchmark with three task directions.Sources (4)pluskal-lab/MassSpecGym official source; pluskal-lab/MassSpecGym massspecgym/models/de_novo/base.py; pluskal-lab/MassSpecGym massspecgym/models/retrieval/base.py; pluskal-lab/MassSpecGym massspecgym/models/simulation/base.py · Pinned README: three challenges; dataset and DataModule; evaluation base classes; pinned de_novo, retrieval and simulation base evaluation classes |
| Organisms | Molecule identity rather than organism classification defines these tasks. · Not applicableSources (4)pluskal-lab/MassSpecGym official source; pluskal-lab/MassSpecGym massspecgym/models/de_novo/base.py; pluskal-lab/MassSpecGym massspecgym/models/retrieval/base.py; pluskal-lab/MassSpecGym massspecgym/models/simulation/base.py · Pinned README: three challenges; dataset and DataModule; evaluation base classes; pinned de_novo, retrieval and simulation base evaluation classes |
| Assays | Tandem mass spectra paired with molecular structures.Sources (4)pluskal-lab/MassSpecGym official source; pluskal-lab/MassSpecGym massspecgym/models/de_novo/base.py; pluskal-lab/MassSpecGym massspecgym/models/retrieval/base.py; pluskal-lab/MassSpecGym massspecgym/models/simulation/base.py · Pinned README: three challenges; dataset and DataModule; evaluation base classes; pinned de_novo, retrieval and simulation base evaluation classes |
| Allowed inputs | Spectrum-to-molecule, spectrum-plus-candidates, or molecule-to-spectrum inputs depend on the selected task.Sources (4)pluskal-lab/MassSpecGym official source; pluskal-lab/MassSpecGym massspecgym/models/de_novo/base.py; pluskal-lab/MassSpecGym massspecgym/models/retrieval/base.py; pluskal-lab/MassSpecGym massspecgym/models/simulation/base.py · Pinned README: three challenges; dataset and DataModule; evaluation base classes; pinned de_novo, retrieval and simulation base evaluation classes |
| Adaptation | Supervised train/validation/test learning; pretrained or new models use the task-specific interfaces.Sources (4)pluskal-lab/MassSpecGym official source; pluskal-lab/MassSpecGym massspecgym/models/de_novo/base.py; pluskal-lab/MassSpecGym massspecgym/models/retrieval/base.py; pluskal-lab/MassSpecGym massspecgym/models/simulation/base.py · Pinned README: three challenges; dataset and DataModule; evaluation base classes; pinned de_novo, retrieval and simulation base evaluation classes |
Applicability is distinct from a completed evaluation.
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Last literature check: 2026-09-17. Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.
| Paper or primary resource | Version | Reference |
|---|---|---|
| MassSpecGym: A benchmark for the discovery and identification of molecules | 2410.23326v1 | Read source |
The catalogue now holds 116 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.
complete comparison tables extracted pending publication review
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
74 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions. Individual claims | pluskal-lab/MassSpecGym massspecgym/models/de_novo/base.py Pinned README: three challenges; dataset and DataModule; evaluation base classes; pinned de_novo, retrieval and simulation base evaluation classes; Section 3.4; Supplementary Information 2.5 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: f259fe3780d5bd227fc6ece36ce6f397c2eef716 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram caption Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions. Individual claims | pluskal-lab/MassSpecGym massspecgym/models/retrieval/base.py Pinned README: three challenges; dataset and DataModule; evaluation base classes; pinned de_novo, retrieval and simulation base evaluation classes; Section 3.4; Supplementary Information 2.5 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: f259fe3780d5bd227fc6ece36ce6f397c2eef716 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram caption Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions. Individual claims | pluskal-lab/MassSpecGym massspecgym/models/simulation/base.py Pinned README: three challenges; dataset and DataModule; evaluation base classes; pinned de_novo, retrieval and simulation base evaluation classes; Section 3.4; Supplementary Information 2.5 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: f259fe3780d5bd227fc6ece36ce6f397c2eef716 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram caption Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions. Individual claims | massspecgym primary benchmark evidence Pinned README: three challenges; dataset and DataModule; evaluation base classes; pinned de_novo, retrieval and simulation base evaluation classes; Section 3.4; Supplementary Information 2.5 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 2410.23326v1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram caption Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions. Individual claims | pluskal-lab/MassSpecGym official source Pinned README: three challenges; dataset and DataModule; evaluation base classes; pinned de_novo, retrieval and simulation base evaluation classes; Section 3.4; Supplementary Information 2.5 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: f259fe3780d5bd227fc6ece36ce6f397c2eef716 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| pluskal-lab/MassSpecGym massspecgym/models/de_novo/base.py Pinned README: three challenges; dataset and DataModule; evaluation base classes; pinned de_novo, retrieval and simulation base evaluation classes; Section 3.4; Supplementary Information 2.5 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: f259fe3780d5bd227fc6ece36ce6f397c2eef716 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| pluskal-lab/MassSpecGym massspecgym/models/retrieval/base.py Pinned README: three challenges; dataset and DataModule; evaluation base classes; pinned de_novo, retrieval and simulation base evaluation classes; Section 3.4; Supplementary Information 2.5 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: f259fe3780d5bd227fc6ece36ce6f397c2eef716 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| pluskal-lab/MassSpecGym massspecgym/models/simulation/base.py Pinned README: three challenges; dataset and DataModule; evaluation base classes; pinned de_novo, retrieval and simulation base evaluation classes; Section 3.4; Supplementary Information 2.5 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: f259fe3780d5bd227fc6ece36ce6f397c2eef716 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| massspecgym primary benchmark evidence Pinned README: three challenges; dataset and DataModule; evaluation base classes; pinned de_novo, retrieval and simulation base evaluation classes; Section 3.4; Supplementary Information 2.5 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 2410.23326v1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| pluskal-lab/MassSpecGym official source Pinned README: three challenges; dataset and DataModule; evaluation base classes; pinned de_novo, retrieval and simulation base evaluation classes; Section 3.4; Supplementary Information 2.5 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: f259fe3780d5bd227fc6ece36ce6f397c2eef716 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
View linked audit checks and correction history
Release 2026-09-29-06401fd5b220 · Record review: discovered
Stable ID: discovery-benchmark-massspecgym