Datasets
Swiss-Prot reaction annotations and AlphaFold Database structures; balanced same-function and different-function pairs.
Enzyme functional-identity classification predicts whether a protein pair shares its annotated reaction function.
Swiss-Prot reaction annotations and AlphaFold Database structures; balanced same-function and different-function pairs.
Accuracy, false-positive rate, MCC, precision, recall, F1, AUROC and AUPR.
Pairs of enzyme representations.
Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.
limited source coverage · Automated source review, 2026-09-16. All specifications and missing details
Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.
ACC (%) (percent) · Higher values are better.
Enzyme-pair functional identity: original held-out test (Enzyme functional identity prediction) · FUJISAN test sub-dataset
Evidence origin: Author-reported evaluation, Independent external evaluation.
Enhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1: ACC (%), Enzyme-pair functional identity: original held-out testLightGBM FUJISAN and comparison methods; predictions thresholded to maximizeF 1. Protein-pair splitting is not evidence of disjoint proteins or families. 41,600 protein pairs; balanced functional-identity classes; random 56.25/18.75/25% training/validation/test split.
Automated source review: 2026-09-17. Numerical source review does not establish independent reproduction.
Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.
Showing 5 of 5 matching rows.
Swiss-Prot reaction annotations and AlphaFold Database structures; balanced same-function and different-function pairs. Random pair partitions provide training, validation and test subsets; model hyperparameters are tuned using cross-validation. Accuracy, false-positive rate, MCC, precision, recall, F1, AUROC and AUPR. LightGBM and multiple conventional classifiers. The low-sequence-similarity evaluation excludes pairs from the original dataset; random splitting of the main pair dataset does not establish protein-identity holdout.
Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.
These source-backed links do not make different protocols or scores interchangeable.
Choose a concrete protocol before running an evaluation. Its inputs, split and scoring rules determine which results can be compared.
No runnable recipe has been reviewed for this task. Dataset access, model requirements, licences and compute requirements must be checked against its sources before execution.
A task describes a biological question. Choose a linked protocol to obtain concrete split and scoring instructions.
No source-reviewed explanatory claims are recorded here yet.
Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.
Stable record: reported-task-1ebf9b408517f9Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Datasets | Swiss-Prot reaction annotations and AlphaFold Database structures; balanced same-function and different-function pairs.SourcesEnhanced prediction of protein functional identity through the integration of sequence and structural features · Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages |
| Splits | Random pair partitions provide training, validation and test subsets; model hyperparameters are tuned using cross-validation.SourcesEnhanced prediction of protein functional identity through the integration of sequence and structural features · Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages |
| Metrics | Accuracy, false-positive rate, MCC, precision, recall, F1, AUROC and AUPR.SourcesEnhanced prediction of protein functional identity through the integration of sequence and structural features · Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages |
| Baselines | LightGBM and multiple conventional classifiers.SourcesEnhanced prediction of protein functional identity through the integration of sequence and structural features · Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages |
| Leakage controls | The low-sequence-similarity evaluation excludes pairs from the original dataset; random splitting of the main pair dataset does not establish protein-identity holdout.SourcesEnhanced prediction of protein functional identity through the integration of sequence and structural features · Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages |
| Uncertainty | The paper assesses stability using repeated bootstrap iterations; its sampling unit must remain attached to the reported interval.SourcesEnhanced prediction of protein functional identity through the integration of sequence and structural features · Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages |
| Entity type | Paper-specific computational evaluation protocol.SourcesEnhanced prediction of protein functional identity through the integration of sequence and structural features · Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages |
| Organisms | Proteins were selected through Swiss-Prot release 2022_04 entries with Rhea reaction annotations and AlphaFold DB v4 structures. Dataset construction does not enumerate organism frequencies for the sampled 100,000 protein pairs, so a species-restricted population cannot be assigned. · Not reported in inspected sourcesSourcesEnhanced prediction of protein functional identity through the integration of sequence and structural features · Materials and methods: Dataset construction |
| Assays | Swiss-Prot reaction/function annotations with AlphaFold structures.SourcesEnhanced prediction of protein functional identity through the integration of sequence and structural features · Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages |
| Allowed inputs | Pairs of enzyme representations.SourcesEnhanced prediction of protein functional identity through the integration of sequence and structural features · Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages |
| Adaptation | Supervised same-function classification; hyperparameters use cross-validation, with algorithm choice additionally compared on test data.SourcesEnhanced prediction of protein functional identity through the integration of sequence and structural features · Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages |
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Last literature check: 2026-09-17. Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.
| Paper or primary resource | Version | Reference |
|---|---|---|
| Enhanced prediction of protein functional identity through the integration of sequence and structural features | PMC11609699.1 | Read source DOI: 10.1016/j.csbj.2024.11.028 |
The catalogue now holds 40 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.
complete comparison tables extracted pending publication review
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
17 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol. Individual claims | Enhanced prediction of protein functional identity through the integration of sequence and structural features Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages Version: PMC11609699.1 | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| Enhanced prediction of protein functional identity through the integration of sequence and structural features Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages Version: PMC11609699.1 | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Computational evaluation flow Individual claims | Enhanced prediction of protein functional identity through the integration of sequence and structural features Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages Version: PMC11609699.1 | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Datasets Swiss-Prot reaction annotations and AlphaFold Database structures; balanced same-function and different-function pairs. Individual claims | Enhanced prediction of protein functional identity through the integration of sequence and structural features Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages Version: PMC11609699.1 | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Splits Random pair partitions provide training, validation and test subsets; model hyperparameters are tuned using cross-validation. Individual claims | Enhanced prediction of protein functional identity through the integration of sequence and structural features Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages Version: PMC11609699.1 | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Adaptation Supervised same-function classification; hyperparameters use cross-validation, with algorithm choice additionally compared on test data. Individual claims | Enhanced prediction of protein functional identity through the integration of sequence and structural features Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages Version: PMC11609699.1 | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Metrics Accuracy, false-positive rate, MCC, precision, recall, F1, AUROC and AUPR. Individual claims | Enhanced prediction of protein functional identity through the integration of sequence and structural features Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages Version: PMC11609699.1 | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Baselines LightGBM and multiple conventional classifiers. Individual claims | Enhanced prediction of protein functional identity through the integration of sequence and structural features Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages Version: PMC11609699.1 | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Leakage controls The low-sequence-similarity evaluation excludes pairs from the original dataset; random splitting of the main pair dataset does not establish protein-identity holdout. Individual claims | Enhanced prediction of protein functional identity through the integration of sequence and structural features Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages Version: PMC11609699.1 | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Uncertainty The paper assesses stability using repeated bootstrap iterations; its sampling unit must remain attached to the reported interval. Individual claims | Enhanced prediction of protein functional identity through the integration of sequence and structural features Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages Version: PMC11609699.1 | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
View linked audit checks and correction history
Release 2026-09-29-06401fd5b220 · Record review: needs review
Stable ID: reported-task-1ebf9b408517f9