| Configuration: GlycanAA | Protocol: GlycanML species Accuracy: GlycanGT study: species Accuracy Dataset subset: SugarBase taxonomy species; GlycanML official motif split (GlycanML split) | 0.40549999999999897 ± 0.0054836119483420804 accuracy fraction · higher Uncertainty: type: standard_deviation; value: 0.0054836119483420804 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceGlycanAA on GlycanML species Accuracy: GlycanGT study: species Accuracy Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A32:E32 (mean D32, SD E32) |
|---|
| Configuration: GlycanAA | Protocol: GlycanML species Macro-F1: GlycanGT study: species Macro-F1 Dataset subset: SugarBase taxonomy species; GlycanML official motif split (GlycanML split) | 0.15840000000000001 ± 0.012304064369142401 macro_f1 dimensionless · higher Uncertainty: type: standard_deviation; value: 0.012304064369142401 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceGlycanAA on GlycanML species Macro-F1: GlycanGT study: species Macro-F1 Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A33:E33 (mean D33, SD E33) |
|---|
| Pipeline: GlycanGT | Protocol: GlycanML species Accuracy: GlycanGT study: species Accuracy Dataset subset: SugarBase taxonomy species; GlycanML official motif split (GlycanML split) | 0.39173014145810597 ± 0.021681021594310401 accuracy fraction · higher Uncertainty: type: standard_deviation; value: 0.021681021594310401 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceGlycanGT on GlycanML species Accuracy: GlycanGT study: species Accuracy Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A79:E79 (mean D79, SD E79) |
|---|
| Pipeline: GlycanGT | Protocol: GlycanML species Macro-F1: GlycanGT study: species Macro-F1 Dataset subset: SugarBase taxonomy species; GlycanML official motif split (GlycanML split) | 0.184571928631623 ± 0.00080738146274129995 macro_f1 dimensionless · higher Uncertainty: type: standard_deviation; value: 0.00080738146274129995 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceGlycanGT on GlycanML species Macro-F1: GlycanGT study: species Macro-F1 Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A80:E80 (mean D80, SD E80) |
|---|
| Configuration: Graphormer | Protocol: GlycanML species Accuracy: GlycanGT study: species Accuracy Dataset subset: SugarBase taxonomy species; GlycanML official motif split (GlycanML split) | 0.35691000000000001 ± 0.019796999999999999 accuracy fraction · higher Uncertainty: type: standard_deviation; value: 0.019796999999999999 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceGraphormer on GlycanML species Accuracy: GlycanGT study: species Accuracy Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A104:E104 (mean D104, SD E104) |
|---|
| Configuration: Graphormer | Protocol: GlycanML species Macro-F1: GlycanGT study: species Macro-F1 Dataset subset: SugarBase taxonomy species; GlycanML official motif split (GlycanML split) | 0.122909 ± 0.0062329999999999998 macro_f1 dimensionless · higher Uncertainty: type: standard_deviation; value: 0.0062329999999999998 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceGraphormer on GlycanML species Macro-F1: GlycanGT study: species Macro-F1 Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A105:E105 (mean D105, SD E105) |
|---|
| Configuration: RGCN | Protocol: GlycanML species Accuracy: GlycanGT study: species Accuracy Dataset subset: SugarBase taxonomy species; GlycanML official motif split (GlycanML split) | 0.0039898440333695998 ± 0.0069106125800716001 accuracy fraction · higher Uncertainty: type: standard_deviation; value: 0.0069106125800716001 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceRGCN on GlycanML species Accuracy: GlycanGT study: species Accuracy Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A17:E17 (mean D17, SD E17) |
|---|
| Configuration: RGCN | Protocol: GlycanML species Macro-F1: GlycanGT study: species Macro-F1 Dataset subset: SugarBase taxonomy species; GlycanML official motif split (GlycanML split) | 0.00069761542364280005 ± 0.001208305357893 macro_f1 dimensionless · higher Uncertainty: type: standard_deviation; value: 0.001208305357893 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceRGCN on GlycanML species Macro-F1: GlycanGT study: species Macro-F1 Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A18:E18 (mean D18, SD E18) |
|---|
| Configuration: SweetNet | Protocol: GlycanML species Accuracy: GlycanGT study: species Accuracy Dataset subset: SugarBase taxonomy species; GlycanML official motif split (GlycanML split) | 0.35545883206383699 ± 0.025755183860829499 accuracy fraction · higher Uncertainty: type: standard_deviation; value: 0.025755183860829499 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceSweetNet on GlycanML species Accuracy: GlycanGT study: species Accuracy Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A51:E51 (mean D51, SD E51) |
|---|
| Configuration: SweetNet | Protocol: GlycanML species Macro-F1: GlycanGT study: species Macro-F1 Dataset subset: SugarBase taxonomy species; GlycanML official motif split (GlycanML split) | 0.11310443922558699 ± 0.0110807900341276 macro_f1 dimensionless · higher Uncertainty: type: standard_deviation; value: 0.0110807900341276 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceSweetNet on GlycanML species Macro-F1: GlycanGT study: species Macro-F1 Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A61:E61 (mean D61, SD E61) |
|---|