Datasets
Multiple task-specific molecular datasets, including grouped benchmark themes.
TDC organizes molecular prediction tasks into datasets and benchmark groups with explicit splitting and evaluation interfaces.
Multiple task-specific molecular datasets, including grouped benchmark themes.
A named evaluator computes task metrics; ROC-AUC is one documented example, not a universal TDC metric.
Named dataset inputs and labels; the benchmark group identifies the relevant molecular representation.
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
limited source coverage · Automated source review, 2026-09-16. All specifications and missing details
Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.
auroc (fraction) · Higher values are better.
TDC ADMET benchmark group TDC-AMES: Toxicity: TDC.AMES · TDC.AMES (TDC ADMET benchmark group split)
Evidence origin: Author-reported evaluation.
Therapeutics Data Commons: Machine Learning Datasets and Tasks for Drug Discovery and Development · Table 3, row(TDC.AMES)Every method TDC ADMET benchmark group reports on Toxicity: TDC.AMES, scored with AUROC on TDC.AMES.
Automated source review: 2026-09-19. Numerical source review does not establish independent reproduction.
Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.
Showing 3 of 3 matching rows.
Therapeutics Data Commons supplies separate molecular tasks and curated benchmark groups. Each group specifies datasets, prediction units, partitions and metrics; the original ADMET example uses scaffold splits and simple descriptor or sequence baselines. Only the molecular and mechanistic tasks within rewire’s scope belong in this catalogue.
Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.
These source-backed links do not make different protocols or scores interchangeable.
Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.
No concrete protocols are explicitly linked to this suite. Protocol identification and baseline selection are outstanding.
Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums
Coverage is derived from release 2026-09-29-06401fd5b220. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.
Install PyTDC and evaluate a model against the ADMET benchmark group, which is the group whose baselines this page records.
Generate predictions and evaluate them. This recipe does not establish reproduction of a particular published score.
Source reviewed; these instructions have not been executed by rewire.
pip install PyTDCTherapeutics Data Commons: repository README · README.md at c310c35f, Using `pip`, lines 73-73Source reviewed; these instructions have not been executed by rewire.
from tdc import BenchmarkGroup
group = BenchmarkGroup(name = 'ADMET_Group', path = 'data/')
predictions_list = []
for seed in [1, 2, 3, 4, 5]:
benchmark = group.get('Caco2_Wang')
# all benchmark names in a benchmark group are stored in group.dataset_names
predictions = {}
name = benchmark['name']
train_val, test = benchmark['train_val'], benchmark['test']
train, valid = group.get_train_valid_split(benchmark = name, split_type = 'default', seed = seed)
# --------------------------------------------- #
# Train your model using train, valid, test #
# Save test prediction in y_pred_test variable #
# --------------------------------------------- #
predictions[name] = y_pred_test
predictions_list.append(predictions)
results = group.evaluate_many(predictions_list)
# {'caco2_wang': [6.328, 0.101]}Therapeutics Data Commons: repository README · README.md at c310c35f, TDC Leaderboards, lines 208-229Run your model locally and return predictions keyed by the input IDs. The evaluator supplies biological inputs without test labels and owns scoring. This interface is not a sandbox for model code.
Pass your existing prediction function into this adapter. Its output direction must match the selected protocol.
class MyModelAdapter:
def __init__(self, score):
self.score = score
def predict(self, inputs):
return {row["id"]: float(self.score(row)) for row in inputs}
# adapter = MyModelAdapter(your_prediction_function)
# report = rewirebench.run(prepared, adapter, output="runs/my-model")Alternatively, generate a keyed prediction file in your existing model environment and use the score-only recipe. Your model code and weights do not need to be shared.
Therapeutics Data Commons: repository README · README.md at c310c35fContribute a result for review. The library can submit an exported evaluation for private review when intake is open. Check the contribution page for access and sign-in.
Official package setup, benchmark-group loading and evaluation examples are available. The leaderboard template leaves model training and y_pred_test unspecified; it is not executable unchanged. Choose a concrete molecular benchmark and preserve its default split/seed protocol rather than mixing TDC tasks.
A maintained rewire runner has not been verified for this benchmark. Check data access, weights, licences, dependencies and hardware in the linked official documentation; requirements have not been fully extracted.
mims-harvard/TDC / README.md · README.md lines 66–108 and 193–232 (Installation, tutorials and Leaderboards)Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.
Stable record: discovery-benchmark-tdc-molecular-tasksExplanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Datasets | Multiple task-specific molecular datasets, including grouped benchmark themes.Sourcesmims-harvard/TDC official source · Pinned README: data functions; dataset splits; evaluators; benchmark groups |
| Splits | Random and scaffold-based splits are supported with recorded seeds/fractions; benchmark groups expose default split routines.Sourcesmims-harvard/TDC official source · Pinned README: data functions; dataset splits; evaluators; benchmark groups |
| Metrics | A named evaluator computes task metrics; ROC-AUC is one documented example, not a universal TDC metric.Sourcesmims-harvard/TDC official source · Pinned README: data functions; dataset splits; evaluators; benchmark groups |
| Baselines | The original molecular ADMET example compares RDKit2D-descriptor MLPs, Morgan-fingerprint MLPs and SMILES CNNs. Other in-scope molecular tasks require their own comparator set; these are not universal baselines for all of TDC.Sourcestdc primary benchmark evidence · Section 9 and Tables 3–4: benchmark groups |
| Leakage controls | Scaffold holdout is available for relevant molecular tasks; it is not implied for every TDC dataset.Sourcesmims-harvard/TDC official source · Pinned README: data functions; dataset splits; evaluators; benchmark groups |
| Uncertainty | The original ADMET table prints ± values but its caption and accompanying protocol do not define a suite-wide resampling or uncertainty rule. A result must retain the exact benchmark-group submission protocol before those values are interpreted. · Not reported in inspected sourcesSourcestdc primary benchmark evidence · Section 9 and Tables 3–4: benchmark groups |
| Entity type | Task and benchmark-group platform for molecular prediction.Sourcesmims-harvard/TDC official source · Pinned README: data functions; dataset splits; evaluators; benchmark groups |
| Organisms | Organism scope is dataset-specific; the umbrella platform does not define one species. · Not applicableSourcesmims-harvard/TDC official source · Pinned README: data functions; dataset splits; evaluators; benchmark groups |
| Assays | Task-specific molecular and therapeutic-property labels.Sourcesmims-harvard/TDC official source · Pinned README: data functions; dataset splits; evaluators; benchmark groups |
| Allowed inputs | Named dataset inputs and labels; the benchmark group identifies the relevant molecular representation.Sourcesmims-harvard/TDC official source · Pinned README: data functions; dataset splits; evaluators; benchmark groups |
| Adaptation | Dataset-specific supervised evaluation with default or explicitly selected split methods.Sourcesmims-harvard/TDC official source · Pinned README: data functions; dataset splits; evaluators; benchmark groups |
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Last literature check: 2026-09-17. Dated primary-source discovery and protocol/table screening. Source checking does not mean experimental reproduction. Only separately extracted and independently reviewed numeric batches are publishable.
| Paper or primary resource | Version | Reference |
|---|---|---|
| Therapeutics Data Commons: Machine Learning Datasets and Tasks for Drug Discovery and Development | arXiv 2102.09548v1; later NSF v2 discovery retrieved partially but full artifact unavailable | Read source |
The catalogue now holds 66 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.
primary protocol reviewed
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
39 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions. Individual claims | mims-harvard/TDC official source Pinned README: data functions; dataset splits; evaluators; benchmark groups Version: c310c35f27e3f506411018ac43d97b8ba23ca652 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| mims-harvard/TDC official source Pinned README: data functions; dataset splits; evaluators; benchmark groups Version: c310c35f27e3f506411018ac43d97b8ba23ca652 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Evaluation procedure Individual claims | mims-harvard/TDC official source Pinned README: data functions; dataset splits; evaluators; benchmark groups Version: c310c35f27e3f506411018ac43d97b8ba23ca652 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Datasets Multiple task-specific molecular datasets, including grouped benchmark themes. Individual claims | mims-harvard/TDC official source Pinned README: data functions; dataset splits; evaluators; benchmark groups Version: c310c35f27e3f506411018ac43d97b8ba23ca652 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Splits Random and scaffold-based splits are supported with recorded seeds/fractions; benchmark groups expose default split routines. Individual claims | mims-harvard/TDC official source Pinned README: data functions; dataset splits; evaluators; benchmark groups Version: c310c35f27e3f506411018ac43d97b8ba23ca652 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Adaptation Dataset-specific supervised evaluation with default or explicitly selected split methods. Individual claims | mims-harvard/TDC official source Pinned README: data functions; dataset splits; evaluators; benchmark groups Version: c310c35f27e3f506411018ac43d97b8ba23ca652 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Metrics A named evaluator computes task metrics; ROC-AUC is one documented example, not a universal TDC metric. Individual claims | mims-harvard/TDC official source Pinned README: data functions; dataset splits; evaluators; benchmark groups Version: c310c35f27e3f506411018ac43d97b8ba23ca652 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Baselines The original molecular ADMET example compares RDKit2D-descriptor MLPs, Morgan-fingerprint MLPs and SMILES CNNs. Other in-scope molecular tasks require their own comparator set; these are not universal baselines for all of TDC. Individual claims | tdc primary benchmark evidence Section 9 and Tables 3–4: benchmark groups Version: 2102.09548v1 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Leakage controls Scaffold holdout is available for relevant molecular tasks; it is not implied for every TDC dataset. Individual claims | mims-harvard/TDC official source Pinned README: data functions; dataset splits; evaluators; benchmark groups Version: c310c35f27e3f506411018ac43d97b8ba23ca652 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Uncertainty The original ADMET table prints ± values but its caption and accompanying protocol do not define a suite-wide resampling or uncertainty rule. A result must retain the exact benchmark-group submission protocol before those values are interpreted. Individual claims | tdc primary benchmark evidence Section 9 and Tables 3–4: benchmark groups Version: 2102.09548v1 | unreported automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
View linked audit checks and correction history
Release 2026-09-29-06401fd5b220 · Record review: discovered
Stable ID: discovery-benchmark-tdc-molecular-tasks