Datasets
SJC combines sc-PDB, JOINED and COACH420; UniProtSMB supplies a separately curated binding-site dataset.
Small-molecule binding-site classification evaluates residue predictions on structural and curated protein annotations.
SJC combines sc-PDB, JOINED and COACH420; UniProtSMB supplies a separately curated binding-site dataset.
Precision, recall, MCC, AUROC and AUPRC.
Protein amino-acid sequences.
Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.
limited source coverage · Automated source review, 2026-09-16. All specifications and missing details
Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.
Recall (unitless) · Higher values are better.
COACH420, trained on CHEN11 (protein-small molecule binding-site prediction) · COACH420, trained on CHEN11
Evidence origin: Result quoted from another source, Author-reported evaluation.
Protein-small molecule binding site prediction based on a pre-trained protein language model with contrastive learning · Comparison of CLAPE-SMB with other models trained on CHEN11 and tested on COACH420; Table 2 (Tab2), row 2 P2Rank, column 2: COACH420, trained on CHEN11 Recall; Table 2 (Tab2), row 3 GraphBind, column 2: COACH420, trained on CHEN11 Recall; Table 2 (Tab2), row 4 CLAPE-SMB, column 2: COACH420, trained on CHEN11 RecallResidue-level small-molecule binding-site classification. CHEN11 training, COACH420 test.
Automated source review: 2026-09-17. Numerical source review does not establish independent reproduction.
Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.
Showing 3 of 3 matching rows.
SJC combines sc-PDB, JOINED and COACH420; UniProtSMB supplies a separately curated binding-site dataset. UniProtSMB representative proteins are partitioned 80:10:10 for training, validation and testing. Precision, recall, MCC, AUROC and AUPRC. ESM-2 versus ProtBert embeddings and MLP/CNN/Transformer head ablations; some head/model comparisons use the test set. Protein similarity clustering and cluster-aware split comparisons are described; ESM-2 pretraining sequences are not universally excluded. Multiple random seeds are evaluated; the paper reports fold-averaged metrics and standard deviations for its robustness experiment.
Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.
These source-backed links do not make different protocols or scores interchangeable.
Each evaluation records what was tested and under which conditions.
Choose a concrete protocol before running an evaluation. Its inputs, split and scoring rules determine which results can be compared.
No runnable recipe has been reviewed for this task. Dataset access, model requirements, licences and compute requirements must be checked against its sources before execution.
A task describes a biological question. Choose a linked protocol to obtain concrete split and scoring instructions.
No source-reviewed explanatory claims are recorded here yet.
Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.
Stable record: reported-task-b181ed450cdd41Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Datasets | SJC combines sc-PDB, JOINED and COACH420; UniProtSMB supplies a separately curated binding-site dataset.SourcesProtein-small molecule binding site prediction based on a pre-trained protein language model with contrastive learning · Methods: Evaluation metrics; SJC dataset preparation; UniProtSMB dataset preparation; Discussion; cached text lines 24–25, 35–43, 93; comparative evaluation and ablation passages; uncertainty/repeat-run/statistical-comparison passages |
| Splits | UniProtSMB representative proteins are partitioned 80:10:10 for training, validation and testing.SourcesProtein-small molecule binding site prediction based on a pre-trained protein language model with contrastive learning · Methods: Evaluation metrics; SJC dataset preparation; UniProtSMB dataset preparation; Discussion; cached text lines 24–25, 35–43, 93; comparative evaluation and ablation passages; uncertainty/repeat-run/statistical-comparison passages |
| Metrics | Precision, recall, MCC, AUROC and AUPRC.SourcesProtein-small molecule binding site prediction based on a pre-trained protein language model with contrastive learning · Methods: Evaluation metrics; SJC dataset preparation; UniProtSMB dataset preparation; Discussion; cached text lines 24–25, 35–43, 93; comparative evaluation and ablation passages; uncertainty/repeat-run/statistical-comparison passages |
| Baselines | ESM-2 versus ProtBert embeddings and MLP/CNN/Transformer head ablations; some head/model comparisons use the test set.SourcesProtein-small molecule binding site prediction based on a pre-trained protein language model with contrastive learning · Methods: Evaluation metrics; SJC dataset preparation; UniProtSMB dataset preparation; Discussion; cached text lines 24–25, 35–43, 93; comparative evaluation and ablation passages; uncertainty/repeat-run/statistical-comparison passages |
| Leakage controls | Protein similarity clustering and cluster-aware split comparisons are described; ESM-2 pretraining sequences are not universally excluded.SourcesProtein-small molecule binding site prediction based on a pre-trained protein language model with contrastive learning · Methods: Evaluation metrics; SJC dataset preparation; UniProtSMB dataset preparation; Discussion; cached text lines 24–25, 35–43, 93; comparative evaluation and ablation passages; uncertainty/repeat-run/statistical-comparison passages |
| Uncertainty | Multiple random seeds are evaluated; the paper reports fold-averaged metrics and standard deviations for its robustness experiment.SourcesProtein-small molecule binding site prediction based on a pre-trained protein language model with contrastive learning · Methods: Evaluation metrics; SJC dataset preparation; UniProtSMB dataset preparation; Discussion; cached text lines 24–25, 35–43, 93; comparative evaluation and ablation passages; uncertainty/repeat-run/statistical-comparison passages |
| Entity type | Paper-specific computational evaluation protocol.SourcesProtein-small molecule binding site prediction based on a pre-trained protein language model with contrastive learning · Methods: Evaluation metrics; SJC dataset preparation; UniProtSMB dataset preparation; Discussion; cached text lines 24–25, 35–43, 93; comparative evaluation and ablation passages; uncertainty/repeat-run/statistical-comparison passages |
| Organisms | The SJC structural collections and UniProtSMB annotations are selected by binding-site evidence and protein redundancy. Their preparation sections and dataset tables do not summarize organism composition or define a species-specific test. · Not reported in inspected sourcesSourcesProtein-small molecule binding site prediction based on a pre-trained protein language model with contrastive learning · SJC dataset preparation; UniProtSMB dataset preparation; Table 1 |
| Assays | Curated protein–small-molecule binding-site annotations.SourcesProtein-small molecule binding site prediction based on a pre-trained protein language model with contrastive learning · Methods: Evaluation metrics; SJC dataset preparation; UniProtSMB dataset preparation; Discussion; cached text lines 24–25, 35–43, 93; comparative evaluation and ablation passages; uncertainty/repeat-run/statistical-comparison passages |
| Allowed inputs | Protein amino-acid sequences.SourcesProtein-small molecule binding site prediction based on a pre-trained protein language model with contrastive learning · Methods: Evaluation metrics; SJC dataset preparation; UniProtSMB dataset preparation; Discussion; cached text lines 24–25, 35–43, 93; comparative evaluation and ablation passages; uncertainty/repeat-run/statistical-comparison passages |
| Adaptation | Supervised residue-level predictor fitted on the training proteins.SourcesProtein-small molecule binding site prediction based on a pre-trained protein language model with contrastive learning · Methods: Evaluation metrics; SJC dataset preparation; UniProtSMB dataset preparation; Discussion; cached text lines 24–25, 35–43, 93; comparative evaluation and ablation passages; uncertainty/repeat-run/statistical-comparison passages |
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Last literature check: 2026-09-17. Dated primary-source discovery and protocol/table screening. Source checking does not mean experimental reproduction. Only separately extracted and independently reviewed numeric batches are publishable.
| Paper or primary resource | Version | Reference |
|---|---|---|
| Protein-small molecule binding site prediction based on a pre-trained protein language model with contrastive learning | version of record | Read source |
The catalogue now holds 45 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.
complete tables extracted
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
17 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol. Individual claims | Protein-small molecule binding site prediction based on a pre-trained protein language model with contrastive learning Methods: Evaluation metrics; SJC dataset preparation; UniProtSMB dataset preparation; Discussion; cached text lines 24–25, 35–43, 93; comparative evaluation and ablation passages; uncertainty/repeat-run/statistical-comparison passages Version: version of record | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| Protein-small molecule binding site prediction based on a pre-trained protein language model with contrastive learning Methods: Evaluation metrics; SJC dataset preparation; UniProtSMB dataset preparation; Discussion; cached text lines 24–25, 35–43, 93; comparative evaluation and ablation passages; uncertainty/repeat-run/statistical-comparison passages Version: version of record | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Computational evaluation flow Individual claims | Protein-small molecule binding site prediction based on a pre-trained protein language model with contrastive learning Methods: Evaluation metrics; SJC dataset preparation; UniProtSMB dataset preparation; Discussion; cached text lines 24–25, 35–43, 93; comparative evaluation and ablation passages; uncertainty/repeat-run/statistical-comparison passages Version: version of record | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Datasets SJC combines sc-PDB, JOINED and COACH420; UniProtSMB supplies a separately curated binding-site dataset. Individual claims | Protein-small molecule binding site prediction based on a pre-trained protein language model with contrastive learning Methods: Evaluation metrics; SJC dataset preparation; UniProtSMB dataset preparation; Discussion; cached text lines 24–25, 35–43, 93; comparative evaluation and ablation passages; uncertainty/repeat-run/statistical-comparison passages Version: version of record | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Splits UniProtSMB representative proteins are partitioned 80:10:10 for training, validation and testing. Individual claims | Protein-small molecule binding site prediction based on a pre-trained protein language model with contrastive learning Methods: Evaluation metrics; SJC dataset preparation; UniProtSMB dataset preparation; Discussion; cached text lines 24–25, 35–43, 93; comparative evaluation and ablation passages; uncertainty/repeat-run/statistical-comparison passages Version: version of record | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Adaptation Supervised residue-level predictor fitted on the training proteins. Individual claims | Protein-small molecule binding site prediction based on a pre-trained protein language model with contrastive learning Methods: Evaluation metrics; SJC dataset preparation; UniProtSMB dataset preparation; Discussion; cached text lines 24–25, 35–43, 93; comparative evaluation and ablation passages; uncertainty/repeat-run/statistical-comparison passages Version: version of record | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Metrics Precision, recall, MCC, AUROC and AUPRC. Individual claims | Protein-small molecule binding site prediction based on a pre-trained protein language model with contrastive learning Methods: Evaluation metrics; SJC dataset preparation; UniProtSMB dataset preparation; Discussion; cached text lines 24–25, 35–43, 93; comparative evaluation and ablation passages; uncertainty/repeat-run/statistical-comparison passages Version: version of record | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Baselines ESM-2 versus ProtBert embeddings and MLP/CNN/Transformer head ablations; some head/model comparisons use the test set. Individual claims | Protein-small molecule binding site prediction based on a pre-trained protein language model with contrastive learning Methods: Evaluation metrics; SJC dataset preparation; UniProtSMB dataset preparation; Discussion; cached text lines 24–25, 35–43, 93; comparative evaluation and ablation passages; uncertainty/repeat-run/statistical-comparison passages Version: version of record | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Leakage controls Protein similarity clustering and cluster-aware split comparisons are described; ESM-2 pretraining sequences are not universally excluded. Individual claims | Protein-small molecule binding site prediction based on a pre-trained protein language model with contrastive learning Methods: Evaluation metrics; SJC dataset preparation; UniProtSMB dataset preparation; Discussion; cached text lines 24–25, 35–43, 93; comparative evaluation and ablation passages; uncertainty/repeat-run/statistical-comparison passages Version: version of record | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Uncertainty Multiple random seeds are evaluated; the paper reports fold-averaged metrics and standard deviations for its robustness experiment. Individual claims | Protein-small molecule binding site prediction based on a pre-trained protein language model with contrastive learning Methods: Evaluation metrics; SJC dataset preparation; UniProtSMB dataset preparation; Discussion; cached text lines 24–25, 35–43, 93; comparative evaluation and ablation passages; uncertainty/repeat-run/statistical-comparison passages Version: version of record | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
View linked audit checks and correction history
Release 2026-09-29-06401fd5b220 · Record review: needs review
Stable ID: reported-task-b181ed450cdd41