Model type
Protein sequence transformer; this record is the paper-specific evaluated configuration.
This FUJISAN-study baseline compares proteins by cosine similarity between mean ESM-2 embeddings.
Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings.
Protein sequence transformer; this record is the paper-specific evaluated configuration.
Pairs of protein amino-acid sequences
Pairwise cosine similarity used as a same-function score
Official upstream implementation and usage documentation: https://github.com/facebookresearch/esm/blob/2b369911bb5b4b0dda914521b9475cad1656b2ac/README.md. This pinned documentation revision is not automatically the evaluated weight revision.
limited source coverage · Automated source review, 2026-09-16. All specifications and missing details
1 evaluation · 8 results. Different protocols are not a single leaderboard.
Applied filters: All linked evaluations
| Tested configuration | Protocol and dataset | Finding | Evidence and details |
|---|---|---|---|
| Configuration: ESM2 | Protocol: Enzyme-pair functional identity: original held-out test (Enzyme functional identity prediction) Dataset: FUJISAN test sub-dataset | 0.799 AUROC unitless · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceESM2: Enzyme functional identity prediction LightGBM FUJISAN and comparison methods; predictions thresholded to maximizeF 1. Protein-pair splitting is not evidence of disjoint proteins or families. 41,600 protein pairs; balanced functional-identity classes; random 56.25/18.75/25% training/validation/test split. Aggregation: Not reported Enhanced prediction of protein functional identity through the integration of sequence and structural features; Enhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, ESM2 row, AUROC column |
| Configuration: ESM2 | Protocol: Enzyme-pair functional identity: original held-out test (Enzyme functional identity prediction) Dataset: FUJISAN test sub-dataset | 79.3% REC (%) percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceESM2: Enzyme functional identity prediction LightGBM FUJISAN and comparison methods; predictions thresholded to maximizeF 1. Protein-pair splitting is not evidence of disjoint proteins or families. 41,600 protein pairs; balanced functional-identity classes; random 56.25/18.75/25% training/validation/test split. Aggregation: Not reported Enhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row ESM2, column REC (%); XML row5 column4 |
| Configuration: ESM2 | Protocol: Enzyme-pair functional identity: original held-out test (Enzyme functional identity prediction) Dataset: FUJISAN test sub-dataset | 0.815 AUPR dimensionless · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceESM2: Enzyme functional identity prediction LightGBM FUJISAN and comparison methods; predictions thresholded to maximizeF 1. Protein-pair splitting is not evidence of disjoint proteins or families. 41,600 protein pairs; balanced functional-identity classes; random 56.25/18.75/25% training/validation/test split. Aggregation: Not reported Enhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row ESM2, column AUPR; XML row5 column9 |
| Configuration: ESM2 | Protocol: Enzyme-pair functional identity: original held-out test (Enzyme functional identity prediction) Dataset: FUJISAN test sub-dataset | 0.432 MCC dimensionless · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceESM2: Enzyme functional identity prediction LightGBM FUJISAN and comparison methods; predictions thresholded to maximizeF 1. Protein-pair splitting is not evidence of disjoint proteins or families. 41,600 protein pairs; balanced functional-identity classes; random 56.25/18.75/25% training/validation/test split. Aggregation: Not reported Enhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row ESM2, column MCC; XML row5 column7 |
| Configuration: ESM2 | Protocol: Enzyme-pair functional identity: original held-out test (Enzyme functional identity prediction) Dataset: FUJISAN test sub-dataset | 68.4% PRE (%) percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceESM2: Enzyme functional identity prediction LightGBM FUJISAN and comparison methods; predictions thresholded to maximizeF 1. Protein-pair splitting is not evidence of disjoint proteins or families. 41,600 protein pairs; balanced functional-identity classes; random 56.25/18.75/25% training/validation/test split. Aggregation: Not reported Enhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row ESM2, column PRE (%); XML row5 column3 |
| Configuration: ESM2 | Protocol: Enzyme-pair functional identity: original held-out test (Enzyme functional identity prediction) Dataset: FUJISAN test sub-dataset | 0.735 F1 dimensionless · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceESM2: Enzyme functional identity prediction LightGBM FUJISAN and comparison methods; predictions thresholded to maximizeF 1. Protein-pair splitting is not evidence of disjoint proteins or families. 41,600 protein pairs; balanced functional-identity classes; random 56.25/18.75/25% training/validation/test split. Aggregation: Not reported Enhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row ESM2, column F1; XML row5 column6 |
| Configuration: ESM2 | Protocol: Enzyme-pair functional identity: original held-out test (Enzyme functional identity prediction) Dataset: FUJISAN test sub-dataset | 71.3% ACC (%) percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceESM2: Enzyme functional identity prediction LightGBM FUJISAN and comparison methods; predictions thresholded to maximizeF 1. Protein-pair splitting is not evidence of disjoint proteins or families. 41,600 protein pairs; balanced functional-identity classes; random 56.25/18.75/25% training/validation/test split. Aggregation: Not reported Enhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row ESM2, column ACC (%); XML row5 column2 |
| Configuration: ESM2 | Protocol: Enzyme-pair functional identity: original held-out test (Enzyme functional identity prediction) Dataset: FUJISAN test sub-dataset | 36.7% FPR (%) percent · lower Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Independent external evaluation · Source checkedMethods, coverage and sourceESM2: Enzyme functional identity prediction LightGBM FUJISAN and comparison methods; predictions thresholded to maximizeF 1. Protein-pair splitting is not evidence of disjoint proteins or families. 41,600 protein pairs; balanced functional-identity classes; random 56.25/18.75/25% training/validation/test split. Aggregation: Not reported Enhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row ESM2, column FPR (%); XML row5 column5 |
Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.
Underlying model: ESM-2. Results on this page belong to this configuration and its evaluated settings.
esm2_t33_650M_UR50D emits 1,280-dimensional residue vectors. Mean pooling gives a protein vector; cosine similarity estimates whether two proteins share an enzymatic function.
ESM-2 is a transformer protein language-model family. The official repository exposes residue embeddings, sequence-level pooling and models at several sizes; the study configuration determines which of these is evaluated.
The linked evaluation record identifies ESM2: Enzyme functional identity prediction. Its dataset, split, adaptation and evidence origin remain attached to the reported results.
Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.
Stable record: reported-model-ccd1160ad4ec27Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Model type | Protein sequence transformer; this record is the paper-specific evaluated configuration.Sourcesfacebookresearch/esm README.md · README.md model description |
| Architecture / procedure | esm2_t33_650M_UR50D emits 1,280-dimensional residue vectors. Mean pooling gives a protein vector; cosine similarity estimates whether two proteins share an enzymatic function.SourcesEnhanced prediction of protein functional identity through the integration of sequence and structural features · Materials and methods/Prediction with ESM-2 (paragraph 2); Materials and methods/Prediction with DeepFRI (paragraph 1) |
| Biological inputs | Pairs of protein amino-acid sequencesSourcesEnhanced prediction of protein functional identity through the integration of sequence and structural features · Materials and methods/Prediction with ESM-2 (paragraph 1); Introduction (paragraph 4) |
| Outputs | Pairwise cosine similarity used as a same-function scoreSourcesEnhanced prediction of protein functional identity through the integration of sequence and structural features · Materials and methods/Prediction with ESM-2 (paragraph 2); Materials and methods/Prediction with DeepFRI (paragraph 1) |
| Parameters | 650 million backbone parameters; no FUJISAN LightGBM features are part of this comparator.SourcesEnhanced prediction of protein functional identity through the integration of sequence and structural features · Results and discussion/Analysis of feature importance (paragraph 1); Materials and methods/Model training and hyperparameter optimization (paragraph 1) |
| Known versions / configuration | esm2_t33_650M_UR50D · Not reported in inspected sourcesSourcesEnhanced prediction of protein functional identity through the integration of sequence and structural features · Model identification in the comparison table and corresponding Methods; immutable checkpoint revision is not supplied by the table label. |
| Training data / fitting | The comparator uses pretrained ESM-2 sequence embeddings and cosine similarity; it does not fit FUJISAN’s LightGBM head or a task-specific interaction classifier.SourcesEnhanced prediction of protein functional identity through the integration of sequence and structural features · ESM2 embedding comparator and cosine-similarity procedure |
| Context limits | A maximum input/context length for this exact evaluated configuration is not established by the inspected sources. · Not reported in inspected sourcesSources (2)Enhanced prediction of protein functional identity through the integration of sequence and structural features; facebookresearch/esm README.md · Materials and methods/Dataset construction; Materials and methods/Feature engineering/Full-length sequence similarity features; Materials and methods/Feature engineering/Domain structural similarity features; Materials and methods/Feature engineering/Pocket similarity features; Materials and methods/Model training and hyperparameter optimization; Materials and methods/Performance assessment; Materials and methods/Prediction with DeepFRI; Materials and methods/Prediction with ESM-2; inspected for explicit maximum input length (dataset lengths and family-wide limits are not substituted); README.md at pinned repository revision |
| Access | Official upstream implementation and usage documentation: https://github.com/facebookresearch/esm/blob/2b369911bb5b4b0dda914521b9475cad1656b2ac/README.md. This pinned documentation revision is not automatically the evaluated weight revision.Sourcesfacebookresearch/esm README.md · README.md; installation, model download and usage instructions |
| Code licence | MIT (upstream repository code at the cited revision; this does not establish every dependency or historical checkpoint licence).Sourcesfacebookresearch/esm LICENSE · LICENSE; complete licence text |
| Weights licence | The inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately. · Not reported in inspected sourcesSourcesfacebookresearch/esm README.md · README.md; checkpoint/access documentation and licence scope |
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
22 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings. Individual claims | Enhanced prediction of protein functional identity through the integration of sequence and structural features Materials and methods/Prediction with ESM-2 (paragraph 2); Materials and methods/Prediction with DeepFRI (paragraph 1) Version: PMC11609699.1 | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| Enhanced prediction of protein functional identity through the integration of sequence and structural features Materials and methods/Prediction with ESM-2 (paragraph 2); Materials and methods/Prediction with DeepFRI (paragraph 1) Version: PMC11609699.1 | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Evaluated procedure (conceptual) Individual claims | Enhanced prediction of protein functional identity through the integration of sequence and structural features Materials and methods/Prediction with ESM-2 (paragraph 2); Materials and methods/Prediction with DeepFRI (paragraph 1) Version: PMC11609699.1 | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Model type Protein sequence transformer; this record is the paper-specific evaluated configuration. Individual claims | facebookresearch/esm README.md README.md model description Version: 2b369911bb5b4b0dda914521b9475cad1656b2ac | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Architecture / procedure esm2_t33_650M_UR50D emits 1,280-dimensional residue vectors. Mean pooling gives a protein vector; cosine similarity estimates whether two proteins share an enzymatic function. Individual claims | Enhanced prediction of protein functional identity through the integration of sequence and structural features Materials and methods/Prediction with ESM-2 (paragraph 2); Materials and methods/Prediction with DeepFRI (paragraph 1) Version: PMC11609699.1 | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Weights licence The inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately. Individual claims | facebookresearch/esm README.md README.md; checkpoint/access documentation and licence scope Version: 2b369911bb5b4b0dda914521b9475cad1656b2ac | unreported automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Biological inputs Pairs of protein amino-acid sequences Individual claims | Enhanced prediction of protein functional identity through the integration of sequence and structural features Materials and methods/Prediction with ESM-2 (paragraph 1); Introduction (paragraph 4) Version: PMC11609699.1 | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Outputs Pairwise cosine similarity used as a same-function score Individual claims | Enhanced prediction of protein functional identity through the integration of sequence and structural features Materials and methods/Prediction with ESM-2 (paragraph 2); Materials and methods/Prediction with DeepFRI (paragraph 1) Version: PMC11609699.1 | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Parameters 650 million backbone parameters; no FUJISAN LightGBM features are part of this comparator. Individual claims | Enhanced prediction of protein functional identity through the integration of sequence and structural features Results and discussion/Analysis of feature importance (paragraph 1); Materials and methods/Model training and hyperparameter optimization (paragraph 1) Version: PMC11609699.1 | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Known versions / configuration esm2_t33_650M_UR50D Individual claims | Enhanced prediction of protein functional identity through the integration of sequence and structural features Model identification in the comparison table and corresponding Methods; immutable checkpoint revision is not supplied by the table label. Version: PMC11609699.1 | unreported automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
View linked audit checks and correction history
Release 2026-09-29-06401fd5b220 · Record review: needs review
Stable ID: reported-model-ccd1160ad4ec27