rewire.itbenchmarks
Task

clathrin protein classification

Clathrin classification uses cross-validation and multiple independent benchmark datasets with different redundancy filters.

SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54

14 evaluations · 79 results

Overview

Evaluation procedure diagram
How it worksComputational evaluation flow
Computational evaluation flow1. Input: Protein amino-acid sequences represented by protein language models.. Then: 2. Evaluation: Ten-fold cross-validation on training data plus named independent CLA test sets.. Then: 3. Readout: Accuracy, AUC, Matthews correlation coefficient, F1, sensitivity and specificity are reported. The Performance evaluation subsection describes ten-fold evaluation and early stopping but does not specify a universal probability threshold for every comparator.Computational evaluation flow1. Input: Protein amino-acid sequences represented by protein language models.. Then: 2. Evaluation: Ten-fold cross-validation on training data plus named independent CLA test sets.. Then: 3. Readout: Accuracy, AUC, Matthews correlation coefficient, F1, sensitivity and specificity are reported. The Performance evaluation subsection describes ten-fold evaluation and early stopping but does not specify a universal probability threshold for every comparator.Computational evaluation flow1. Input: Protein amino-acid sequences represented by protein language models.. Then: 2. Evaluation: Ten-fold cross-validation on training data plus named independent CLA test sets.. Then: 3. Readout: Accuracy, AUC, Matthews correlation coefficient, F1, sensitivity and specificity are reported. The Performance evaluation subsection describes ten-fold evaluation and early stopping but does not specify a universal probability threshold for every comparator.

Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.

SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54; Materials and methods: Performance evaluation

limited source coverage · Automated source review, 2026-09-16. All specifications and missing details

Results

Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.

Clathrin independent test: selected-embedding classifiers · Table 3

ACC (fraction) · Higher values are better.

Clathrin independent test: selected-embedding classifiers (clathrin protein classification) · Clathrin independent test: selected-embedding classifiers

Evidence origin: Author-reported evaluation.

Advancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3: ACC, Clathrin independent test: selected-embedding classifiers
  • Source prose has inconsistent dataset sizes; exact cohort counts left unresolved, not inferred.
  • Feature-selection and evaluation leakage cannot be excluded by this table alone.
Comparison details and limitations

Conventional classifiers and PLM-CLA compared using paper-selected features. CLA-IND 0.6 independent test from Shoombuatong 2024 dataset; source Table 1 carries cohort counts.

  • No interval assigned unless printed in source cell.

Automated source review: 2026-09-17. Numerical source review does not establish independent reproduction.

Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.

Showing 12 of 13 matching rows.

Methods and evaluation design

Procedure, tasks and evaluated configurations

How it works

Evaluation methodology

Le2019 and Zhang2020-derived clathrin/non-clathrin sequence collections. Ten-fold cross-validation on training data plus named independent CLA test sets. Accuracy, AUC, Matthews correlation coefficient, F1, sensitivity and specificity are reported. The Performance evaluation subsection describes ten-fold evaluation and early stopping but does not specify a universal probability threshold for every comparator. deep-clathrin includes literature-reported scores; DeepCLA is reimplemented for comparison. The Zhang2020-derived collection uses BLAST redundancy filtering at 0.7, and the new dataset uses CD-HIT at 0.6. These dataset filters do not establish absence from the pretrained protein encoders’ corpora. Feature subsets are evaluated on both cross-validation and independent-test performance, which limits treating that test as untouched model selection.

SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54; Materials and methods: Performance evaluation; Materials and methods: Dataset construction and Overall framework of PLM-CLA; Results: The effect of feature selection methods on the predictive performance

Evaluation design

Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.

These source-backed links do not make different protocols or scores interchangeable.

Recorded evaluations

Each evaluation records what was tested and under which conditions.

Run this benchmark

Choose a concrete protocol before running an evaluation. Its inputs, split and scoring rules determine which results can be compared.

Run instructions

No runnable recipe has been reviewed for this task. Dataset access, model requirements, licences and compute requirements must be checked against its sources before execution.

A task describes a biological question. Choose a linked protocol to obtain concrete split and scoring instructions.

Strengths, limitations and unresolved questions

Strengths and limitations

Strengths and considerations

No source-reviewed explanatory claims are recorded here yet.

Profile review details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Stable record: reported-task-786c09824e9bf5

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetsLe2019 and Zhang2020-derived clathrin/non-clathrin sequence collections.
SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54
SplitsTen-fold cross-validation on training data plus named independent CLA test sets.
SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54
MetricsAccuracy, AUC, Matthews correlation coefficient, F1, sensitivity and specificity are reported. The Performance evaluation subsection describes ten-fold evaluation and early stopping but does not specify a universal probability threshold for every comparator.
SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Materials and methods: Performance evaluation
Baselinesdeep-clathrin includes literature-reported scores; DeepCLA is reimplemented for comparison.
SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54
Leakage controlsThe Zhang2020-derived collection uses BLAST redundancy filtering at 0.7, and the new dataset uses CD-HIT at 0.6. These dataset filters do not establish absence from the pretrained protein encoders’ corpora. Feature subsets are evaluated on both cross-validation and independent-test performance, which limits treating that test as untouched model selection.
SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Materials and methods: Dataset construction and Overall framework of PLM-CLA; Results: The effect of feature selection methods on the predictive performance
UncertaintyThe cited text-accessible evaluation sections give no confidence-interval, resampling or repeat-run error-bar specification. Image-only tables and uninspected supplements are outside this absence claim. · Not reported in inspected sources
SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54
Entity typePaper-specific computational evaluation protocol.
SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54
OrganismsThe paper does not enumerate organisms. Its pinned released CSVs contain sequence identifiers and amino-acid sequences, without organism columns; one dataset uses synthetic row identifiers. Consequently, a complete source-species inventory is not established by the supplied benchmark metadata. · Not reported in inspected sources
Sources (5)Advancing the accuracy of clathrin protein prediction through multi-source protein language models; clathrin__Dataset__Clathrin0.6.csv; clathrin__Dataset__Clathrin0.7.csv; clathrin__Dataset__Clathrin1.0.csv; clathrin__README.md · Dataset construction; pinned repository README and Dataset/Clathrin0.6.csv, Clathrin0.7.csv and Clathrin1.0.csv headers
AssaysClathrin/non-clathrin sequence labels.
SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54
Allowed inputsProtein amino-acid sequences represented by protein language models.
SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54
AdaptationSupervised classification with cross-validation and independent test collections.
SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.

Paper or primary resourceVersionReference
Advancing the accuracy of clathrin protein prediction through multi-source protein language modelsjournal full text in PMCRead source
DOI: 10.1038/s41598-025-08510-4
Historical gaps recorded on 2026-09-17

The catalogue now holds 79 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.

  • Independent batch review before import; preserve existing observation identities.
Search and extraction details

complete comparison tables extracted pending publication review

Searches

  • Advancing the accuracy of clathrin protein prediction through multi-source protein language models primary paper benchmark results

Evidence locations

  • Tables2–5; independent-test and dataset-construction sections

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

21 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.
Individual claims
Advancing the accuracy of clathrin protein prediction through multi-source protein language models

Original source ↗

Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54; Materials and methods: Performance evaluation

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558194+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 2edc86b25707c1b737d26117093ce8d856e79cc5d0b335f27c1c341f887f1c7e

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps
  • Input: Protein amino-acid sequences represented by protein language models.
  • Evaluation: Ten-fold cross-validation on training data plus named independent CLA test sets.
  • Readout: Accuracy, AUC, Matthews correlation coefficient, F1, sensitivity and specificity are reported. The Performance evaluation subsection describes ten-fold evaluation and early stopping but does not specify a universal probability threshold for every comparator.
Individual claims
Advancing the accuracy of clathrin protein prediction through multi-source protein language models

Original source ↗

Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54; Materials and methods: Performance evaluation

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558194+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 2edc86b25707c1b737d26117093ce8d856e79cc5d0b335f27c1c341f887f1c7e

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title
Computational evaluation flow
Individual claims
Advancing the accuracy of clathrin protein prediction through multi-source protein language models

Original source ↗

Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54; Materials and methods: Performance evaluation

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558194+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.diagram.title

Source artifact SHA-256: 2edc86b25707c1b737d26117093ce8d856e79cc5d0b335f27c1c341f887f1c7e

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Datasets
Le2019 and Zhang2020-derived clathrin/non-clathrin sequence collections.
Individual claims
Advancing the accuracy of clathrin protein prediction through multi-source protein language models

Original source ↗

Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558194+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: 2edc86b25707c1b737d26117093ce8d856e79cc5d0b335f27c1c341f887f1c7e

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Splits
Ten-fold cross-validation on training data plus named independent CLA test sets.
Individual claims
Advancing the accuracy of clathrin protein prediction through multi-source protein language models

Original source ↗

Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558194+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: 2edc86b25707c1b737d26117093ce8d856e79cc5d0b335f27c1c341f887f1c7e

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Adaptation
Supervised classification with cross-validation and independent test collections.
Individual claims
Advancing the accuracy of clathrin protein prediction through multi-source protein language models

Original source ↗

Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558194+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: 2edc86b25707c1b737d26117093ce8d856e79cc5d0b335f27c1c341f887f1c7e

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Metrics
Accuracy, AUC, Matthews correlation coefficient, F1, sensitivity and specificity are reported. The Performance evaluation subsection describes ten-fold evaluation and early stopping but does not specify a universal probability threshold for every comparator.
Individual claims
Advancing the accuracy of clathrin protein prediction through multi-source protein language models

Original source ↗

Materials and methods: Performance evaluation

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558194+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: 2edc86b25707c1b737d26117093ce8d856e79cc5d0b335f27c1c341f887f1c7e

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Baselines
deep-clathrin includes literature-reported scores; DeepCLA is reimplemented for comparison.
Individual claims
Advancing the accuracy of clathrin protein prediction through multi-source protein language models

Original source ↗

Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558194+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.3.value

Source artifact SHA-256: 2edc86b25707c1b737d26117093ce8d856e79cc5d0b335f27c1c341f887f1c7e

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Leakage controls
The Zhang2020-derived collection uses BLAST redundancy filtering at 0.7, and the new dataset uses CD-HIT at 0.6. These dataset filters do not establish absence from the pretrained protein encoders’ corpora. Feature subsets are evaluated on both cross-validation and independent-test performance, which limits treating that test as untouched model selection.
Individual claims
Advancing the accuracy of clathrin protein prediction through multi-source protein language models

Original source ↗

Materials and methods: Dataset construction and Overall framework of PLM-CLA; Results: The effect of feature selection methods on the predictive performance

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558194+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: 2edc86b25707c1b737d26117093ce8d856e79cc5d0b335f27c1c341f887f1c7e

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Uncertainty
The cited text-accessible evaluation sections give no confidence-interval, resampling or repeat-run error-bar specification. Image-only tables and uninspected supplements are outside this absence claim.
Individual claims
Advancing the accuracy of clathrin protein prediction through multi-source protein language models

Original source ↗

Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558194+00:00

unreported

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.5.value

Source artifact SHA-256: 2edc86b25707c1b737d26117093ce8d856e79cc5d0b335f27c1c341f887f1c7e

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: needs review

6 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: reported-task-786c09824e9bf5

areas
proteins-complexes
tasks
clathrin protein classification
entity level
task
version
Not reported
task
clathrin protein classification
scope note
Paper-specific evaluation task; protocol completeness requires further extraction.
comparison panels
id: part2-clathrin-plm-2025-Tab3-7ef65949fa; title: Clathrin independent test: selected-embedding classifiers · Table 3; protocol id: paper-protocol-bbffa94da73852557b; dataset id: paper-dataset-a1836de18515f3ab55; metric: ACC; unit: fraction; direction: higher; result ids: paper-result-0730029c96e20b98de; paper-result-a55f9924436b4f2946; paper-result-6d26a3281ddec02ba6; paper-result-199f6ba07b9c097f4f; paper-result-644ae2f042fb2f5b23; paper-result-36ae69af478636d799; paper-result-7ed01313e3f4c28c40; paper-result-7195f65c3684098379; paper-result-a969b52eed01ab25da; paper-result-4cf9ff8fa7d95ed4ce; paper-result-e677c183db2fc0a6fe; paper-result-f444ed165adbeb9c85; paper-result-48b062f1f39640cb79; source ids: part2-clathrin-plm-2025; source locator: Table 3: ACC, Clathrin independent test: selected-embedding classifiers; context: Conventional classifiers and PLM-CLA compared using paper-selected features. CLA-IND 0.6 independent test from Shoombuatong 2024 dataset; source Table 1 carries cohort counts.; caveats: Source prose has inconsistent dataset sizes; exact cohort counts left unresolved, not inferred.; Feature-selection and evaluation leakage cannot be excluded by this table alone.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17; id: part2-clathrin-plm-2025-Tab3-9059dcb8eb; title: Clathrin independent test: selected-embedding classifiers · Table 3; protocol id: paper-protocol-bbffa94da73852557b; dataset id: paper-dataset-a1836de18515f3ab55; metric: SN; unit: fraction; direction: higher; result ids: paper-result-7e844e0a18dfd2a35b; paper-result-1d8775900e322544d5; paper-result-14af9ff7e7145d4d29; paper-result-2691b7b496a2b70805; paper-result-8a421ada6b63ee530d; paper-result-9824c5bbad9176ec8e; paper-result-5667de2e439b758d52; paper-result-c5b318cd60d0cdb4a7; paper-result-ac6a602d275b3fea3c; paper-result-5a7ce3c38ddae8961d; paper-result-a2a2ddd6bf8f13daef; paper-result-387a61f6d26366959d; paper-result-d62014e2f2c00f0b19; source ids: part2-clathrin-plm-2025; source locator: Table 3: SN, Clathrin independent test: selected-embedding classifiers; context: Conventional classifiers and PLM-CLA compared using paper-selected features. CLA-IND 0.6 independent test from Shoombuatong 2024 dataset; source Table 1 carries cohort counts.; caveats: Source prose has inconsistent dataset sizes; exact cohort counts left unresolved, not inferred.; Feature-selection and evaluation leakage cannot be excluded by this table alone.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17; id: part2-clathrin-plm-2025-Tab3-1637766aec; title: Clathrin independent test: selected-embedding classifiers · Table 3; protocol id: paper-protocol-bbffa94da73852557b; dataset id: paper-dataset-a1836de18515f3ab55; metric: SP; unit: fraction; direction: higher; result ids: paper-result-21feace8cb2ce5e3eb; paper-result-e33d8392c76daadbc3; paper-result-97cc71a9c8b17a8bc1; paper-result-667c1dd0ce262e830c; paper-result-f1e415b2cfba651012; paper-result-0ce650036f050c2762; paper-result-21c44a4182b8225eb1; paper-result-a2d81dd3a62697406b; paper-result-7c339a10051c86a7c0; paper-result-ea67a0b7b24f3bb052; paper-result-9bdbc84d434b285e04; paper-result-87b08711535c28932a; paper-result-d04ecf04da1004cff8; source ids: part2-clathrin-plm-2025; source locator: Table 3: SP, Clathrin independent test: selected-embedding classifiers; context: Conventional classifiers and PLM-CLA compared using paper-selected features. CLA-IND 0.6 independent test from Shoombuatong 2024 dataset; source Table 1 carries cohort counts.; caveats: Source prose has inconsistent dataset sizes; exact cohort counts left unresolved, not inferred.; Feature-selection and evaluation leakage cannot be excluded by this table alone.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17; id: part2-clathrin-plm-2025-Tab3-657b713876; title: Clathrin independent test: selected-embedding classifiers · Table 3; protocol id: paper-protocol-bbffa94da73852557b; dataset id: paper-dataset-a1836de18515f3ab55; metric: MCC; unit: dimensionless; direction: higher; result ids: paper-result-a4a56a5352e1f4b25f; paper-result-655c72140d737d23c3; paper-result-1482c92aa694d2f027; paper-result-250b4aaac43175427d; paper-result-ddc5d5b4d806f5d683; paper-result-92d64dfea940c8e704; paper-result-1671d2e2b268f272a5; paper-result-b30374085233a5ee59; paper-result-40e463b7ca62e3c241; paper-result-6644f2850c0d5eb339; paper-result-c9dfd2f1b60d02b3df; paper-result-930a9b4bcf61fbe236; paper-result-12c7518a23c15c58d6; source ids: part2-clathrin-plm-2025; source locator: Table 3: MCC, Clathrin independent test: selected-embedding classifiers; context: Conventional classifiers and PLM-CLA compared using paper-selected features. CLA-IND 0.6 independent test from Shoombuatong 2024 dataset; source Table 1 carries cohort counts.; caveats: Source prose has inconsistent dataset sizes; exact cohort counts left unresolved, not inferred.; Feature-selection and evaluation leakage cannot be excluded by this table alone.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17; id: part2-clathrin-plm-2025-Tab3-a89d53b495; title: Clathrin independent test: selected-embedding classifiers · Table 3; protocol id: paper-protocol-bbffa94da73852557b; dataset id: paper-dataset-a1836de18515f3ab55; metric: F1; unit: fraction; direction: higher; result ids: paper-result-f92d473cd400e9e621; paper-result-96f846d80898fae95b; paper-result-9b9e600ad2446e697f; paper-result-af17b84609852bd79e; paper-result-5a6bc1834630e501e8; paper-result-2d7b6ae84bb35ea892; paper-result-a5bbea217db4b4a850; paper-result-d35402f96201b16bf8; paper-result-b30dbb80fd831aad99; paper-result-798a21dedf23053ac0; paper-result-3e7012cff605fa9f7f; paper-result-3431e8053e5b0bad25; paper-result-2b3a797967cdef7a15; source ids: part2-clathrin-plm-2025; source locator: Table 3: F1, Clathrin independent test: selected-embedding classifiers; context: Conventional classifiers and PLM-CLA compared using paper-selected features. CLA-IND 0.6 independent test from Shoombuatong 2024 dataset; source Table 1 carries cohort counts.; caveats: Source prose has inconsistent dataset sizes; exact cohort counts left unresolved, not inferred.; Feature-selection and evaluation leakage cannot be excluded by this table alone.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17; id: part2-clathrin-plm-2025-Tab3-b6e9f189ae; title: Clathrin independent test: selected-embedding classifiers · Table 3; protocol id: paper-protocol-bbffa94da73852557b; dataset id: paper-dataset-a1836de18515f3ab55; metric: AUC; unit: fraction; direction: higher; result ids: paper-result-2eaf94777fc7f857fc; paper-result-47d350f89848982a43; paper-result-f1a86ea15cd7479d5b; paper-result-e274031e8e17dc0933; paper-result-ad1e11f4f3b25e62ed; paper-result-3741a0c6925f2947e2; paper-result-ca68d167787e314805; paper-result-7ef5e7bcd2b4a87ef9; paper-result-9565fcdb350e8db673; paper-result-5250d6a1217d2812cf; paper-result-e0a382c956e9add325; paper-result-9c2aeff72feeb82f45; paper-result-ab057b14f1c421f3ca; source ids: part2-clathrin-plm-2025; source locator: Table 3: AUC, Clathrin independent test: selected-embedding classifiers; context: Conventional classifiers and PLM-CLA compared using paper-selected features. CLA-IND 0.6 independent test from Shoombuatong 2024 dataset; source Table 1 carries cohort counts.; caveats: Source prose has inconsistent dataset sizes; exact cohort counts left unresolved, not inferred.; Feature-selection and evaluation leakage cannot be excluded by this table alone.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17
benchmark research
review date: 2026-09-17; status: complete_comparison_tables_extracted_pending_publication_review; primary sources: part2-clathrin-plm-2025; inspected locators: Tables2–5; independent-test and dataset-construction sections; searched queries: Advancing the accuracy of clathrin protein prediction through multi-source protein language models primary paper benchmark results; gaps: Independent batch review before import; preserve existing observation identities.; claim scope: Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.
historical missing metadata
protocol version: not_reported_in_legacy_extract
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
legacy kinds
benchmark
entity classification
review date: 2026-09-17; rationale: This source-scoped record identifies the biological prediction task and holds its paper context. Preserve the existing task identity; exact split, model adaptation and scoring remain in linked evaluations or separate protocol records.; source ids: clathrin-plm-2025; source locator: Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54; ambiguities: A paper- or suite-specific task may constrain some inputs or metrics; that alone does not make it interchangeable with a complete versioned protocol. No protocol equivalence is inferred.; Some legacy profile Entity type facts use the generic phrase computational evaluation protocol. That boilerplate is not sufficient to establish a single fixed protocol identity or to merge this task with another protocol record.
Related records

Suggest a correction