Datasets
Le2019 and Zhang2020-derived clathrin/non-clathrin sequence collections.
Clathrin classification uses cross-validation and multiple independent benchmark datasets with different redundancy filters.
Le2019 and Zhang2020-derived clathrin/non-clathrin sequence collections.
Accuracy, AUC, Matthews correlation coefficient, F1, sensitivity and specificity are reported. The Performance evaluation subsection describes ten-fold evaluation and early stopping but does not specify a universal probability threshold for every comparator.
Protein amino-acid sequences represented by protein language models.
Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.
limited source coverage · Automated source review, 2026-09-16. All specifications and missing details
Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.
ACC (fraction) · Higher values are better.
Clathrin independent test: selected-embedding classifiers (clathrin protein classification) · Clathrin independent test: selected-embedding classifiers
Evidence origin: Author-reported evaluation.
Advancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3: ACC, Clathrin independent test: selected-embedding classifiersConventional classifiers and PLM-CLA compared using paper-selected features. CLA-IND 0.6 independent test from Shoombuatong 2024 dataset; source Table 1 carries cohort counts.
Automated source review: 2026-09-17. Numerical source review does not establish independent reproduction.
Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.
Showing 12 of 13 matching rows.
Le2019 and Zhang2020-derived clathrin/non-clathrin sequence collections. Ten-fold cross-validation on training data plus named independent CLA test sets. Accuracy, AUC, Matthews correlation coefficient, F1, sensitivity and specificity are reported. The Performance evaluation subsection describes ten-fold evaluation and early stopping but does not specify a universal probability threshold for every comparator. deep-clathrin includes literature-reported scores; DeepCLA is reimplemented for comparison. The Zhang2020-derived collection uses BLAST redundancy filtering at 0.7, and the new dataset uses CD-HIT at 0.6. These dataset filters do not establish absence from the pretrained protein encoders’ corpora. Feature subsets are evaluated on both cross-validation and independent-test performance, which limits treating that test as untouched model selection.
Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.
These source-backed links do not make different protocols or scores interchangeable.
Each evaluation records what was tested and under which conditions.
Choose a concrete protocol before running an evaluation. Its inputs, split and scoring rules determine which results can be compared.
No runnable recipe has been reviewed for this task. Dataset access, model requirements, licences and compute requirements must be checked against its sources before execution.
A task describes a biological question. Choose a linked protocol to obtain concrete split and scoring instructions.
No source-reviewed explanatory claims are recorded here yet.
Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.
Stable record: reported-task-786c09824e9bf5Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Datasets | Le2019 and Zhang2020-derived clathrin/non-clathrin sequence collections.SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54 |
| Splits | Ten-fold cross-validation on training data plus named independent CLA test sets.SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54 |
| Metrics | Accuracy, AUC, Matthews correlation coefficient, F1, sensitivity and specificity are reported. The Performance evaluation subsection describes ten-fold evaluation and early stopping but does not specify a universal probability threshold for every comparator.SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Materials and methods: Performance evaluation |
| Baselines | deep-clathrin includes literature-reported scores; DeepCLA is reimplemented for comparison.SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54 |
| Leakage controls | The Zhang2020-derived collection uses BLAST redundancy filtering at 0.7, and the new dataset uses CD-HIT at 0.6. These dataset filters do not establish absence from the pretrained protein encoders’ corpora. Feature subsets are evaluated on both cross-validation and independent-test performance, which limits treating that test as untouched model selection.SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Materials and methods: Dataset construction and Overall framework of PLM-CLA; Results: The effect of feature selection methods on the predictive performance |
| Uncertainty | The cited text-accessible evaluation sections give no confidence-interval, resampling or repeat-run error-bar specification. Image-only tables and uninspected supplements are outside this absence claim. · Not reported in inspected sourcesSourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54 |
| Entity type | Paper-specific computational evaluation protocol.SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54 |
| Organisms | The paper does not enumerate organisms. Its pinned released CSVs contain sequence identifiers and amino-acid sequences, without organism columns; one dataset uses synthetic row identifiers. Consequently, a complete source-species inventory is not established by the supplied benchmark metadata. · Not reported in inspected sourcesSources (5)Advancing the accuracy of clathrin protein prediction through multi-source protein language models; clathrin__Dataset__Clathrin0.6.csv; clathrin__Dataset__Clathrin0.7.csv; clathrin__Dataset__Clathrin1.0.csv; clathrin__README.md · Dataset construction; pinned repository README and Dataset/Clathrin0.6.csv, Clathrin0.7.csv and Clathrin1.0.csv headers |
| Assays | Clathrin/non-clathrin sequence labels.SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54 |
| Allowed inputs | Protein amino-acid sequences represented by protein language models.SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54 |
| Adaptation | Supervised classification with cross-validation and independent test collections.SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54 |
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Last literature check: 2026-09-17. Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.
| Paper or primary resource | Version | Reference |
|---|---|---|
| Advancing the accuracy of clathrin protein prediction through multi-source protein language models | journal full text in PMC | Read source DOI: 10.1038/s41598-025-08510-4 |
The catalogue now holds 79 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.
complete comparison tables extracted pending publication review
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
21 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol. Individual claims | Advancing the accuracy of clathrin protein prediction through multi-source protein language models Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54; Materials and methods: Performance evaluation Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| Advancing the accuracy of clathrin protein prediction through multi-source protein language models Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54; Materials and methods: Performance evaluation Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Computational evaluation flow Individual claims | Advancing the accuracy of clathrin protein prediction through multi-source protein language models Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54; Materials and methods: Performance evaluation Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Datasets Le2019 and Zhang2020-derived clathrin/non-clathrin sequence collections. Individual claims | Advancing the accuracy of clathrin protein prediction through multi-source protein language models Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54 Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Splits Ten-fold cross-validation on training data plus named independent CLA test sets. Individual claims | Advancing the accuracy of clathrin protein prediction through multi-source protein language models Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54 Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Adaptation Supervised classification with cross-validation and independent test collections. Individual claims | Advancing the accuracy of clathrin protein prediction through multi-source protein language models Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54 Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Metrics Accuracy, AUC, Matthews correlation coefficient, F1, sensitivity and specificity are reported. The Performance evaluation subsection describes ten-fold evaluation and early stopping but does not specify a universal probability threshold for every comparator. Individual claims | Advancing the accuracy of clathrin protein prediction through multi-source protein language models Materials and methods: Performance evaluation Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Baselines deep-clathrin includes literature-reported scores; DeepCLA is reimplemented for comparison. Individual claims | Advancing the accuracy of clathrin protein prediction through multi-source protein language models Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54 Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Leakage controls The Zhang2020-derived collection uses BLAST redundancy filtering at 0.7, and the new dataset uses CD-HIT at 0.6. These dataset filters do not establish absence from the pretrained protein encoders’ corpora. Feature subsets are evaluated on both cross-validation and independent-test performance, which limits treating that test as untouched model selection. Individual claims | Advancing the accuracy of clathrin protein prediction through multi-source protein language models Materials and methods: Dataset construction and Overall framework of PLM-CLA; Results: The effect of feature selection methods on the predictive performance Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Uncertainty The cited text-accessible evaluation sections give no confidence-interval, resampling or repeat-run error-bar specification. Image-only tables and uninspected supplements are outside this absence claim. Individual claims | Advancing the accuracy of clathrin protein prediction through multi-source protein language models Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54 Version: journal full text in PMC | unreported automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
View linked audit checks and correction history
Release 2026-09-29-06401fd5b220 · Record review: needs review
Stable ID: reported-task-786c09824e9bf5