rewire.itbenchmarks
Benchmark

TDC molecular tasks

TDC organizes molecular prediction tasks into datasets and benchmark groups with explicit splitting and evaluation interfaces.

Sourcesmims-harvard/TDC official source · Pinned README: data functions; dataset splits; evaluators; benchmark groups

66 evaluations · 66 results

Overview

Datasets

Multiple task-specific molecular datasets, including grouped benchmark themes.

Metrics

A named evaluator computes task metrics; ROC-AUC is one documented example, not a universal TDC metric.

Allowed inputs

Named dataset inputs and labels; the benchmark group identifies the relevant molecular representation.

Sourcesmims-harvard/TDC official source · Pinned README: data functions; dataset splits; evaluators; benchmark groups
Evaluation procedure diagram
How it worksEvaluation procedure
Evaluation procedure1. Allowed inputs: Named dataset inputs and labels; the benchmark group identifies the relevant molecular representation.. Then: 2. Splits: Random and scaffold-based splits are supported with recorded seeds/fractions; benchmark groups expose default split routines.. Then: 3. Metrics: A named evaluator computes task metrics; ROC-AUC is one documented example, not a universal TDC metric.Evaluation procedure1. Allowed inputs: Named dataset inputs and labels; the benchmark group identifies the relevant molecular representation.. Then: 2. Splits: Random and scaffold-based splits are supported with recorded seeds/fractions; benchmark groups expose default split routines.. Then: 3. Metrics: A named evaluator computes task metrics; ROC-AUC is one documented example, not a universal TDC metric.Evaluation procedure1. Allowed inputs: Named dataset inputs and labels; the benchmark group identifies the relevant molecular representation.. Then: 2. Splits: Random and scaffold-based splits are supported with recorded seeds/fractions; benchmark groups expose default split routines.. Then: 3. Metrics: A named evaluator computes task metrics; ROC-AUC is one documented example, not a universal TDC metric.

Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.

Sourcesmims-harvard/TDC official source · Pinned README: data functions; dataset splits; evaluators; benchmark groups

limited source coverage · Automated source review, 2026-09-16. All specifications and missing details

Results

Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.

TDC ADMET benchmark group TDC-AMES: Toxicity: TDC.AMES

auroc (fraction) · Higher values are better.

TDC ADMET benchmark group TDC-AMES: Toxicity: TDC.AMES · TDC.AMES (TDC ADMET benchmark group split)

Evidence origin: Author-reported evaluation.

Therapeutics Data Commons: Machine Learning Datasets and Tasks for Drug Discovery and Development · Table 3, row(TDC.AMES)
  • MAE datasets are better when lower. The metric comes from Table 3 and differs by dataset.
  • These are the paper's own simple baselines, not the current leaderboard for the ADMET group.
  • The pinned v1 source does not identify whether its plus-or-minus spreads are standard deviations, standard errors or another measure.
Comparison details and limitations

Every method TDC ADMET benchmark group reports on Toxicity: TDC.AMES, scored with AUROC on TDC.AMES.

  • Author-reported numbers, source checked but not independently reproduced.

Automated source review: 2026-09-19. Numerical source review does not establish independent reproduction.

Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.

Showing 3 of 3 matching rows.

Tested configuration
00.250.50.751
Reported score
  1. RDKit2D + MLP0.823 ± 0.011
  2. Morgan + MLP0.794 ± 0.008
  3. CNN0.776 ± 0.015

Methods and evaluation design

Procedure, tasks and evaluated configurations

How it works

Evaluation methodology

Therapeutics Data Commons supplies separate molecular tasks and curated benchmark groups. Each group specifies datasets, prediction units, partitions and metrics; the original ADMET example uses scaffold splits and simple descriptor or sequence baselines. Only the molecular and mechanistic tasks within rewire’s scope belong in this catalogue.

Sourcestdc primary benchmark evidence · Section 9 and Tables 3–4: benchmark groups

Evaluation design

Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.

These source-backed links do not make different protocols or scores interchangeable.

Baseline coverage

Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.

No concrete protocols are explicitly linked to this suite. Protocol identification and baseline selection are outstanding.

Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums

Coverage is derived from release 2026-09-29-06401fd5b220. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.

Run this benchmark

Run the ADMET benchmark group with PyTDC

Install PyTDC and evaluate a model against the ADMET benchmark group, which is the group whose baselines this page records.

Generate predictions and evaluate them. This recipe does not establish reproduction of a particular published score.

Dataset access
Downloaded by PyTDC on first use; the benchmark group fixes the split.
Model and weights
None. The baselines are trained from the featurisation you choose.
Licences
Project licence: MIT. Upstream data licences are separate and unreported here.
Software
Python with PyTDC installed from PyPI.
Hardware
CPU is enough for the paper's own baselines.
Required inputs and expected outputs

Inputs

  • A model that predicts a property for a SMILES string.
  • No local data download: the loader fetches each dataset on first use.

Outputs

  • A per-dataset score from the group's own evaluator, under the scaffold split the group defines.

Execution steps

  1. 1. Install (Command line)

    Source reviewed; these instructions have not been executed by rewire.

    pip install PyTDC
    Therapeutics Data Commons: repository README · README.md at c310c35f, Using `pip`, lines 73-73
  2. 2. Evaluate against the benchmark group (Python)

    Source reviewed; these instructions have not been executed by rewire.

    from tdc import BenchmarkGroup
    group = BenchmarkGroup(name = 'ADMET_Group', path = 'data/')
    predictions_list = []
    
    for seed in [1, 2, 3, 4, 5]:
        benchmark = group.get('Caco2_Wang')
        # all benchmark names in a benchmark group are stored in group.dataset_names
        predictions = {}
        name = benchmark['name']
        train_val, test = benchmark['train_val'], benchmark['test']
        train, valid = group.get_train_valid_split(benchmark = name, split_type = 'default', seed = seed)
    
            # --------------------------------------------- #
            #  Train your model using train, valid, test    #
            #  Save test prediction in y_pred_test variable #
            # --------------------------------------------- #
    
        predictions[name] = y_pred_test
        predictions_list.append(predictions)
    
    results = group.evaluate_many(predictions_list)
    # {'caco2_wang': [6.328, 0.101]}
    Therapeutics Data Commons: repository README · README.md at c310c35f, TDC Leaderboards, lines 208-229

Use your own model

Run your model locally and return predictions keyed by the input IDs. The evaluator supplies biological inputs without test labels and owns scoring. This interface is not a sandbox for model code.

Pass your existing prediction function into this adapter. Its output direction must match the selected protocol.

class MyModelAdapter:
    def __init__(self, score):
        self.score = score

    def predict(self, inputs):
        return {row["id"]: float(self.score(row)) for row in inputs}

# adapter = MyModelAdapter(your_prediction_function)
# report = rewirebench.run(prepared, adapter, output="runs/my-model")

Alternatively, generate a keyed prediction file in your existing model environment and use the score-only recipe. Your model code and weights do not need to be shared.

Therapeutics Data Commons: repository README · README.md at c310c35f
Scope and limitations
  • Quoted from the project's README and not executed by rewire, so the commands are evidence of what the project documents rather than a verified run.
  • The project may have changed since the pinned commit.
  • The group evaluates each dataset with its own metric; there is no single ADMET score.
  • The figures on this page are the paper's baselines, not the current leaderboard.

Contribute a result for review. The library can submit an exported evaluation for private review when intake is open. Check the contribution page for access and sign-in.

Original repository instructions

Run this benchmark

Official package setup, benchmark-group loading and evaluation examples are available. The leaderboard template leaves model training and y_pred_test unspecified; it is not executable unchanged. Choose a concrete molecular benchmark and preserve its default split/seed protocol rather than mixing TDC tasks.

A maintained rewire runner has not been verified for this benchmark. Check data access, weights, licences, dependencies and hardware in the linked official documentation; requirements have not been fully extracted.

mims-harvard/TDC / README.md · README.md lines 66–108 and 193–232 (Installation, tutorials and Leaderboards)
Strengths, limitations and unresolved questions

Strengths and limitations

Strengths and considerations

  • Benchmark groups expose standard split and evaluator routines while permitting explicit alternatives.
    Sourcesmims-harvard/TDC official source · Pinned README: data functions; dataset splits; evaluators; benchmark groups

Limitations and conditions

  • TDC also contains tasks outside this catalogue’s molecular remit. Shared software access does not make datasets, labels or evaluation protocols scientifically interchangeable.
    Sourcestdc primary benchmark evidence · Section 9 and Tables 3–4: benchmark groups
Profile review details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Stable record: discovery-benchmark-tdc-molecular-tasks

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetsMultiple task-specific molecular datasets, including grouped benchmark themes.
Sourcesmims-harvard/TDC official source · Pinned README: data functions; dataset splits; evaluators; benchmark groups
SplitsRandom and scaffold-based splits are supported with recorded seeds/fractions; benchmark groups expose default split routines.
Sourcesmims-harvard/TDC official source · Pinned README: data functions; dataset splits; evaluators; benchmark groups
MetricsA named evaluator computes task metrics; ROC-AUC is one documented example, not a universal TDC metric.
Sourcesmims-harvard/TDC official source · Pinned README: data functions; dataset splits; evaluators; benchmark groups
BaselinesThe original molecular ADMET example compares RDKit2D-descriptor MLPs, Morgan-fingerprint MLPs and SMILES CNNs. Other in-scope molecular tasks require their own comparator set; these are not universal baselines for all of TDC.
Sourcestdc primary benchmark evidence · Section 9 and Tables 3–4: benchmark groups
Leakage controlsScaffold holdout is available for relevant molecular tasks; it is not implied for every TDC dataset.
Sourcesmims-harvard/TDC official source · Pinned README: data functions; dataset splits; evaluators; benchmark groups
UncertaintyThe original ADMET table prints ± values but its caption and accompanying protocol do not define a suite-wide resampling or uncertainty rule. A result must retain the exact benchmark-group submission protocol before those values are interpreted. · Not reported in inspected sources
Sourcestdc primary benchmark evidence · Section 9 and Tables 3–4: benchmark groups
Entity typeTask and benchmark-group platform for molecular prediction.
Sourcesmims-harvard/TDC official source · Pinned README: data functions; dataset splits; evaluators; benchmark groups
OrganismsOrganism scope is dataset-specific; the umbrella platform does not define one species. · Not applicable
Sourcesmims-harvard/TDC official source · Pinned README: data functions; dataset splits; evaluators; benchmark groups
AssaysTask-specific molecular and therapeutic-property labels.
Sourcesmims-harvard/TDC official source · Pinned README: data functions; dataset splits; evaluators; benchmark groups
Allowed inputsNamed dataset inputs and labels; the benchmark group identifies the relevant molecular representation.
Sourcesmims-harvard/TDC official source · Pinned README: data functions; dataset splits; evaluators; benchmark groups
AdaptationDataset-specific supervised evaluation with default or explicitly selected split methods.
Sourcesmims-harvard/TDC official source · Pinned README: data functions; dataset splits; evaluators; benchmark groups

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Dated primary-source discovery and protocol/table screening. Source checking does not mean experimental reproduction. Only separately extracted and independently reviewed numeric batches are publishable.

Paper or primary resourceVersionReference
Therapeutics Data Commons: Machine Learning Datasets and Tasks for Drug Discovery and DevelopmentarXiv 2102.09548v1; later NSF v2 discovery retrieved partially but full artifact unavailableRead source
Historical gaps recorded on 2026-09-17

The catalogue now holds 66 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.

  • complete comparable numeric result batch: Specific candidate tables and protocol boundaries are documented; no graph values or incomplete winner-only selection are converted into publishable rows.
  • Fresh arXiv v1 retrieval was incomplete and primary NSF v2 retrieval timed out; cached v1 bytes were hash-verified. The v2 full text is not claimed to have been reviewed.
Search and extraction details

primary protocol reviewed

Searches

  • Therapeutics Data Commons 2102.09548 benchmark results

Evidence locations

  • Pinned arXiv v1 Table 3 (ADMET dataset statistics), Table 4 (complete ADMET model comparisons), §6.1 (DTI dataset/protocol definitions)

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

39 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Individual claims
mims-harvard/TDC official source

Original source ↗

Pinned README: data functions; dataset splits; evaluators; benchmark groups

Version: c310c35f27e3f506411018ac43d97b8ba23ca652
Retrieved: 2026-09-16T10:30:23.885647+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: e2dacff9ad56bca50c31373a1e87041eef948721b8aa0be001a95185c375c15f

Hash scope: Hash scope not separately documented; inspect source record

Diagram steps
  • Allowed inputs: Named dataset inputs and labels; the benchmark group identifies the relevant molecular representation.
  • Splits: Random and scaffold-based splits are supported with recorded seeds/fractions; benchmark groups expose default split routines.
  • Metrics: A named evaluator computes task metrics; ROC-AUC is one documented example, not a universal TDC metric.
Individual claims
mims-harvard/TDC official source

Original source ↗

Pinned README: data functions; dataset splits; evaluators; benchmark groups

Version: c310c35f27e3f506411018ac43d97b8ba23ca652
Retrieved: 2026-09-16T10:30:23.885647+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: e2dacff9ad56bca50c31373a1e87041eef948721b8aa0be001a95185c375c15f

Hash scope: Hash scope not separately documented; inspect source record

Diagram title
Evaluation procedure
Individual claims
mims-harvard/TDC official source

Original source ↗

Pinned README: data functions; dataset splits; evaluators; benchmark groups

Version: c310c35f27e3f506411018ac43d97b8ba23ca652
Retrieved: 2026-09-16T10:30:23.885647+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: e2dacff9ad56bca50c31373a1e87041eef948721b8aa0be001a95185c375c15f

Hash scope: Hash scope not separately documented; inspect source record

Datasets
Multiple task-specific molecular datasets, including grouped benchmark themes.
Individual claims
mims-harvard/TDC official source

Original source ↗

Pinned README: data functions; dataset splits; evaluators; benchmark groups

Version: c310c35f27e3f506411018ac43d97b8ba23ca652
Retrieved: 2026-09-16T10:30:23.885647+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: e2dacff9ad56bca50c31373a1e87041eef948721b8aa0be001a95185c375c15f

Hash scope: Hash scope not separately documented; inspect source record

Splits
Random and scaffold-based splits are supported with recorded seeds/fractions; benchmark groups expose default split routines.
Individual claims
mims-harvard/TDC official source

Original source ↗

Pinned README: data functions; dataset splits; evaluators; benchmark groups

Version: c310c35f27e3f506411018ac43d97b8ba23ca652
Retrieved: 2026-09-16T10:30:23.885647+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: e2dacff9ad56bca50c31373a1e87041eef948721b8aa0be001a95185c375c15f

Hash scope: Hash scope not separately documented; inspect source record

Adaptation
Dataset-specific supervised evaluation with default or explicitly selected split methods.
Individual claims
mims-harvard/TDC official source

Original source ↗

Pinned README: data functions; dataset splits; evaluators; benchmark groups

Version: c310c35f27e3f506411018ac43d97b8ba23ca652
Retrieved: 2026-09-16T10:30:23.885647+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: e2dacff9ad56bca50c31373a1e87041eef948721b8aa0be001a95185c375c15f

Hash scope: Hash scope not separately documented; inspect source record

Metrics
A named evaluator computes task metrics; ROC-AUC is one documented example, not a universal TDC metric.
Individual claims
mims-harvard/TDC official source

Original source ↗

Pinned README: data functions; dataset splits; evaluators; benchmark groups

Version: c310c35f27e3f506411018ac43d97b8ba23ca652
Retrieved: 2026-09-16T10:30:23.885647+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: e2dacff9ad56bca50c31373a1e87041eef948721b8aa0be001a95185c375c15f

Hash scope: Hash scope not separately documented; inspect source record

Baselines
The original molecular ADMET example compares RDKit2D-descriptor MLPs, Morgan-fingerprint MLPs and SMILES CNNs. Other in-scope molecular tasks require their own comparator set; these are not universal baselines for all of TDC.
Individual claims
tdc primary benchmark evidence

Original source ↗

Section 9 and Tables 3–4: benchmark groups

Version: 2102.09548v1
Retrieved: 2026-09-16T21:04:55.709744+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.3.value

Source artifact SHA-256: aaa5526f6f100093bc06c9921a8564ad08cfd5332270d614d7681051c2dd66bc

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Leakage controls
Scaffold holdout is available for relevant molecular tasks; it is not implied for every TDC dataset.
Individual claims
mims-harvard/TDC official source

Original source ↗

Pinned README: data functions; dataset splits; evaluators; benchmark groups

Version: c310c35f27e3f506411018ac43d97b8ba23ca652
Retrieved: 2026-09-16T10:30:23.885647+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: e2dacff9ad56bca50c31373a1e87041eef948721b8aa0be001a95185c375c15f

Hash scope: Hash scope not separately documented; inspect source record

Uncertainty
The original ADMET table prints ± values but its caption and accompanying protocol do not define a suite-wide resampling or uncertainty rule. A result must retain the exact benchmark-group submission protocol before those values are interpreted.
Individual claims
tdc primary benchmark evidence

Original source ↗

Section 9 and Tables 3–4: benchmark groups

Version: 2102.09548v1
Retrieved: 2026-09-16T21:04:55.709744+00:00

unreported

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.5.value

Source artifact SHA-256: aaa5526f6f100093bc06c9921a8564ad08cfd5332270d614d7681051c2dd66bc

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: discovered

9 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: discovery-benchmark-tdc-molecular-tasks

areas
molecular-interactions
entity level
suite
scope note
Specialist molecular or omics evaluation; protocol details require review before numerical comparison.
task
Molecular binding, biochemical activity and related specialist tasks
version
Not reported
benchmark research
review date: 2026-09-17; status: primary_protocol_reviewed; primary sources: evidence-expansion-tdc-cached-v1-aaa5526f; inspected locators: Pinned arXiv v1 Table 3 (ADMET dataset statistics), Table 4 (complete ADMET model comparisons), §6.1 (DTI dataset/protocol definitions); searched queries: Therapeutics Data Commons 2102.09548 benchmark results; gaps: complete comparable numeric result batch: Specific candidate tables and protocol boundaries are documented; no graph values or incomplete winner-only selection are converted into publishable rows.; Fresh arXiv v1 retrieval was incomplete and primary NSF v2 retrieval timed out; cached v1 bytes were hash-verified. The v2 full text is not claimed to have been reviewed.; claim scope: Dated primary-source discovery and protocol/table screening. Source checking does not mean experimental reproduction. Only separately extracted and independently reviewed numeric batches are publishable.
historical missing metadata
dataset release: unextracted; metric implementation: unextracted; split manifest: unextracted; version: unextracted
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
entity classification
review date: 2026-09-17; rationale: The cited profile describes a collection of evaluation tasks or protocols; retain it as the top-level benchmark suite. Its datasets and individual protocols remain separate records.; source ids: src-discovery-mims-harvard-tdc; source locator: Pinned README: data functions; dataset splits; evaluators; benchmark groups; ambiguities: None recorded
run documentation
record id: discovery-benchmark-tdc-molecular-tasks; source ids: run-doc-tdc-molecular-tasks-readme-md-c310c35f; status: official_documentation_linked; summary: Official package setup, benchmark-group loading and evaluation examples are available. The leaderboard template leaves model training and y_pred_test unspecified; it is not executable unchanged. Choose a concrete molecular benchmark and preserve its default split/seed protocol rather than mixing TDC tasks.; source locator: README.md lines 66–108 and 193–232 (Installation, tutorials and Leaderboards)
run recipes
id: tdc-official; protocol id: discovery-benchmark-tdc-molecular-tasks; version: c310c35f27e3f506411018ac43d97b8ba23ca652; title: Run the ADMET benchmark group with PyTDC; purpose: generate_and_evaluate; summary: Install PyTDC and evaluate a model against the ADMET benchmark group, which is the group whose baselines this page records.; inputs: A model that predicts a property for a SMILES string.; No local data download: the loader fetches each dataset on first use.; outputs: A per-dataset score from the group's own evaluator, under the scaffold split the group defines.; requirements: data: Downloaded by PyTDC on first use; the benchmark group fixes the split.; weights: None. The baselines are trained from the featurisation you choose.; licence: Project licence: MIT. Upstream data licences are separate and unreported here.; software: Python with PyTDC installed from PyPI.; hardware: CPU is enough for the paper's own baselines.; instructions: runtime: command_line; title: Install; code: pip install PyTDC; status: source_reviewed_not_executed; source ids: project-recipe-tdc-c310c35f; source locator: README.md at c310c35f, Using `pip`, lines 73-73; runtime: python; title: Evaluate against the benchmark group; code: from tdc import BenchmarkGroup group = BenchmarkGroup(name = 'ADMET_Group', path = 'data/') predictions_list = [] for seed in [1, 2, 3, 4, 5]: benchmark = group.get('Caco2_Wang') # all benchmark names in a benchmark group are stored in group.dataset_names predictions = {} name = benchmark['name'] train_val, test = benchmark['train_val'], benchmark['test'] train, valid = group.get_train_valid_split(benchmark = name, split_type = 'default', seed = seed) # --------------------------------------------- # # Train your model using train, valid, test # # Save test prediction in y_pred_test variable # # --------------------------------------------- # predictions[name] = y_pred_test predictions_list.append(predictions) results = group.evaluate_many(predictions_list) # {'caco2_wang': [6.328, 0.101]}; status: source_reviewed_not_executed; source ids: project-recipe-tdc-c310c35f; source locator: README.md at c310c35f, TDC Leaderboards, lines 208-229; limitations: Quoted from the project's README and not executed by rewire, so the commands are evidence of what the project documents rather than a verified run.; The project may have changed since the pinned commit.; The group evaluates each dataset with its own metric; there is no single ADMET score.; The figures on this page are the paper's baselines, not the current leaderboard.; source ids: project-recipe-tdc-c310c35f; source locator: README.md at c310c35f; id: tdc-rewirebench; protocol id: tdc-admet-group-v1; version: aecb9e79a2a5e83b59e482212b1a8b812dd16079; title: Score your own model with rewirebench; purpose: generate_and_evaluate; summary: Prepare a dataset TDC has written to disk, then score an adapter with the metric TDC assigns that dataset.; inputs: The dataset directory TDC's BenchmarkGroup wrote locally.; An adapter returning one number per Drug_ID: a predicted value, or a positive-class score for a classification dataset.; outputs: A local report with the metric, its direction, the coverage and the digest of the files it read.; requirements: data: Whatever TDC downloaded locally. The runner records its SHA-256 rather than claiming a canonical split.; weights: Whatever your own model needs; the runner supplies none.; licence: Runner code is MIT. The benchmark's own data terms are upstream and unreported here.; software: Python 3.11 with the pinned rewirebench environment.; hardware: CPU for scoring. Your own model decides what it needs.; instructions: runtime: command_line; title: Prepare a dataset; code: rewirebench prepare tdc-admet-group-v1 \ --source ./data --options '{"dataset": "caco2_wang"}' \ --output ./prepared-caco2; status: source_reviewed_not_executed; source ids: project-recipe-runner-tdc-tdc-admet-md-aecb9e79; project-recipe-runner-tdc-tdc-admet-py-aecb9e79; source locator: docs/tdc-admet.md at aecb9e79, Prepare, run and score, lines 36-38; runtime: command_line; title: Run and score an adapter; code: rewirebench run --prepared ./prepared-caco2 \ --adapter my_models.admet:MyAdapter \ --model-name 'my model' --training-overlap 'unreported' \ --output ./scored-caco2; status: source_reviewed_not_executed; source ids: project-recipe-runner-tdc-tdc-admet-md-aecb9e79; project-recipe-runner-tdc-tdc-admet-py-aecb9e79; source locator: docs/tdc-admet.md at aecb9e79, Prepare, run and score, lines 47-50; limitations: Quoted from the runner's documentation and not executed by this repository.; Scoring follows the benchmark's own evaluator; running it does not by itself reproduce a published number.; Six of the 22 datasets are errors where lower is better; the report states the direction per dataset.; A matching metric is not a matching result: the split on disk, the featurisation and the training all have to match a published number too.; This scores a hashed local data copy; it does not certify the official cohort or reproduce a published experiment.; source ids: project-recipe-runner-tdc-tdc-admet-md-aecb9e79; project-recipe-runner-tdc-tdc-admet-py-aecb9e79; source locator: docs/tdc-admet.md at aecb9e79
Related records

Suggest a correction