| Configuration: GlycanAA | Protocol: GlycanML genus Accuracy: GlycanGT study: genus Accuracy Dataset subset: SugarBase taxonomy genus; GlycanML official motif split (GlycanML split) | 0.38300000000000001 ± 0.0324371700368574 accuracy fraction · higher Uncertainty: type: standard_deviation; value: 0.0324371700368574 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceGlycanAA on GlycanML genus Accuracy: GlycanGT study: genus Accuracy Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A36:E36 (mean D36, SD E36) |
|---|
| Configuration: GlycanAA | Protocol: GlycanML genus Macro-F1: GlycanGT study: genus Macro-F1 Dataset subset: SugarBase taxonomy genus; GlycanML official motif split (GlycanML split) | 0.18640000000000001 ± 0.0119478031453485 macro_f1 dimensionless · higher Uncertainty: type: standard_deviation; value: 0.0119478031453485 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceGlycanAA on GlycanML genus Macro-F1: GlycanGT study: genus Macro-F1 Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A37:E37 (mean D37, SD E37) |
|---|
| Pipeline: GlycanGT | Protocol: GlycanML genus Accuracy: GlycanGT study: genus Accuracy Dataset subset: SugarBase taxonomy genus; GlycanML official motif split (GlycanML split) | 0.44577439245556699 ± 0.0045302850913300002 accuracy fraction · higher Uncertainty: type: standard_deviation; value: 0.0045302850913300002 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceGlycanGT on GlycanML genus Accuracy: GlycanGT study: genus Accuracy Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A77:E77 (mean D77, SD E77) |
|---|
| Pipeline: GlycanGT | Protocol: GlycanML genus Macro-F1: GlycanGT study: genus Macro-F1 Dataset subset: SugarBase taxonomy genus; GlycanML official motif split (GlycanML split) | 0.188457844899896 ± 0.0007473331496721 macro_f1 dimensionless · higher Uncertainty: type: standard_deviation; value: 0.0007473331496721 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceGlycanGT on GlycanML genus Macro-F1: GlycanGT study: genus Macro-F1 Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A78:E78 (mean D78, SD E78) |
|---|
| Configuration: Graphormer | Protocol: GlycanML genus Accuracy: GlycanGT study: genus Accuracy Dataset subset: SugarBase taxonomy genus; GlycanML official motif split (GlycanML split) | 0.38520100000000002 ± 0.017274000000000001 accuracy fraction · higher Uncertainty: type: standard_deviation; value: 0.017274000000000001 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceGraphormer on GlycanML genus Accuracy: GlycanGT study: genus Accuracy Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A92:E92 (mean D92, SD E92) |
|---|
| Configuration: Graphormer | Protocol: GlycanML genus Macro-F1: GlycanGT study: genus Macro-F1 Dataset subset: SugarBase taxonomy genus; GlycanML official motif split (GlycanML split) | 0.14089399999999999 ± 0.0048110000000000002 macro_f1 dimensionless · higher Uncertainty: type: standard_deviation; value: 0.0048110000000000002 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceGraphormer on GlycanML genus Macro-F1: GlycanGT study: genus Macro-F1 Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A93:E93 (mean D93, SD E93) |
|---|
| Configuration: RGCN | Protocol: GlycanML genus Accuracy: GlycanGT study: genus Accuracy Dataset subset: SugarBase taxonomy genus; GlycanML official motif split (GlycanML split) | 0.0116068190061661 ± 0.0093815866205132006 accuracy fraction · higher Uncertainty: type: standard_deviation; value: 0.0093815866205132006 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceRGCN on GlycanML genus Accuracy: GlycanGT study: genus Accuracy Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A15:E15 (mean D15, SD E15) |
|---|
| Configuration: RGCN | Protocol: GlycanML genus Macro-F1: GlycanGT study: genus Macro-F1 Dataset subset: SugarBase taxonomy genus; GlycanML official motif split (GlycanML split) | 0.0025822707765614 ± 0.001593712031334 macro_f1 dimensionless · higher Uncertainty: type: standard_deviation; value: 0.001593712031334 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceRGCN on GlycanML genus Macro-F1: GlycanGT study: genus Macro-F1 Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A16:E16 (mean D16, SD E16) |
|---|
| Configuration: SweetNet | Protocol: GlycanML genus Accuracy: GlycanGT study: genus Accuracy Dataset subset: SugarBase taxonomy genus; GlycanML official motif split (GlycanML split) | 0.35364526659412399 ± 0.0062192340223003999 accuracy fraction · higher Uncertainty: type: standard_deviation; value: 0.0062192340223003999 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceSweetNet on GlycanML genus Accuracy: GlycanGT study: genus Accuracy Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A50:E50 (mean D50, SD E50) |
|---|
| Configuration: SweetNet | Protocol: GlycanML genus Macro-F1: GlycanGT study: genus Macro-F1 Dataset subset: SugarBase taxonomy genus; GlycanML official motif split (GlycanML split) | 0.12932981544724301 ± 0.016260805531287802 macro_f1 dimensionless · higher Uncertainty: type: standard_deviation; value: 0.016260805531287802 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceSweetNet on GlycanML genus Macro-F1: GlycanGT study: genus Macro-F1 Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A60:E60 (mean D60, SD E60) |
|---|