Model type
BERT-style protein transformer encoder
The TAPE Transformer learns contextual protein representations using masked-residue pretraining and a task-specific prediction head.
Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.
BERT-style protein transformer encoder
Amino-acid sequence in the tokenizer expected by the selected implementation.
Residue/sequence representations and predictions from the chosen task head.
Official project documentation and implementation: https://github.com/songlab-cal/tape
limited source coverage · Automated source review, 2026-09-16. All specifications and missing details
7 evaluations · 7 results. Different protocols are not a single leaderboard.
Applied filters: All linked evaluations
| Tested configuration | Protocol and dataset | Finding | Evidence and details |
|---|---|---|---|
| Configuration: TAPE Transformer (ATOM3D baseline) (cited as [Rao et al., 2019]) | Task: ATOM3D MSP: Mutation stability prediction Dataset subset: ATOM3D MSP (ATOM3D split) | 0.554 auroc fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceTAPE Transformer (ATOM3D baseline) on ATOM3D MSP: Mutation stability prediction Trained and scored under the ATOM3D split for this task. Asterisks in the paper mark a run whose training data differed. Aggregation: Not reported ATOM3D: Tasks On Molecules in Three Dimensions · Table 4, row(MSP AUROC), column([Rao et al., 2019]) |
| Configuration: TAPE Transformer (ATOM3D baseline) (cited as [Rao et al., 2019]) | Task: ATOM3D RES: Residue identity Dataset subset: ATOM3D RES (ATOM3D split) | 0.3 accuracy fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceTAPE Transformer (ATOM3D baseline) on ATOM3D RES: Residue identity Trained and scored under the ATOM3D split for this task. Asterisks in the paper mark a run whose training data differed. Aggregation: Not reported ATOM3D: Tasks On Molecules in Three Dimensions · Table 4, row(RES accuracy), column([Rao et al., 2019]) |
| Model: TAPE Transformer | Task: TAPE Fluorescence Dataset: TAPE Fluorescence source dataset | 0.68 Spearman's rho correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceTAPE Fluorescence Transformer leaderboard evaluation Train/validation at most 3 mutations; test 4–15 mutations from the GFP parent Aggregation: Not reported songlab-cal/tape official source; tape primary benchmark evidence · README.md > Leaderboard > Fluorescence; row Transformer; column Spearman's rho |
| Model: TAPE Transformer | Task: TAPE Stability Dataset: TAPE Stability source dataset | 0.73 Spearman's rho correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceTAPE Stability Transformer leaderboard evaluation Test 17 one-mutation neighbourhoods around candidate proteins; train on broad experimental design rounds Aggregation: Not reported songlab-cal/tape official source; tape primary benchmark evidence · README.md > Leaderboard > Stability; row Transformer; column Spearman's rho |
| Model: TAPE Transformer | Task: TAPE Contact Prediction Dataset: ProteinNet CASP12 | 0.36 precision at L/5 fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceTAPE Transformer: ProteinNet CASP12 CASP12 test, 30% sequence-identity partition; medium/long-range contacts Aggregation: Not reported tape primary benchmark evidence · Table2,p.7,row4(Self-supervised pretraining/Transformer),columnContact prediction |
| Model: TAPE Transformer | Task: TAPE Secondary Structure Dataset: CB513 | 0.73 accuracy fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceCB513 test set; train/test proteins filtered at 25% sequence identity Aggregation: Not reported tape primary benchmark evidence · Table2,p.7,row4(Self-supervised pretraining/Transformer),columnSecondary structure |
| Model: TAPE Transformer | Task: TAPE Remote Homology Detection Dataset: SCOP1.75 fold-level test | 0.21 accuracy fraction · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceTAPE Transformer: SCOP1.75 fold-level test Held-out evolutionary groups; fold-level classification into 1,195 folds Aggregation: Not reported tape primary benchmark evidence · Table2,p.7,row4(Self-supervised pretraining/Transformer),columnRemote homology |
Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.
The TAPE Transformer learns contextual protein representations using masked-residue pretraining and a task-specific prediction head. The June 2019 preprint uses 12 transformer layers, width 512 and eight heads (38M parameters). The later PyTorch BertConfig defaults to width 768 and 12 heads; these are different configurations. The documented inputs are amino-acid sequence in the tokenizer expected by the selected implementation. The output consists of residue/sequence representations and predictions from the chosen task head.
June 2019 TAPE preprint configurations and the later PyTorch package are distinct; the current README warns it is not an exact reproduction. The inspected default BertConfig sets max_position_embeddings to 8,096; this is an implementation default, not evidence that a historical TAPE checkpoint was trained at that length.
Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.
Stable record: discovery-model-tape-transformerExplanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Model type | BERT-style protein transformer encoderSources (7)songlab-cal/tape: README.md; songlab-cal/tape: tape/models/modeling_lstm.py; songlab-cal/tape: tape/models/modeling_resnet.py; songlab-cal/tape: tape/models/modeling_unirep.py; songlab-cal/tape: tape/models/modeling_onehot.py; songlab-cal/tape: tape/models/modeling_bert.py; tape: Primary paper PDF · tape/models/modeling_bert.py: configuration and model classes; README.md: opening maintenance warning, Examples and Data; June 2019 TAPE preprint Sections 4.1,5 and AppendixA.3 |
| Architecture | The June 2019 preprint uses 12 transformer layers, width 512 and eight heads (38M parameters). The later PyTorch BertConfig defaults to width 768 and 12 heads; these are different configurations.Sources (7)songlab-cal/tape: README.md; songlab-cal/tape: tape/models/modeling_lstm.py; songlab-cal/tape: tape/models/modeling_resnet.py; songlab-cal/tape: tape/models/modeling_unirep.py; songlab-cal/tape: tape/models/modeling_onehot.py; songlab-cal/tape: tape/models/modeling_bert.py; tape: Primary paper PDF · tape/models/modeling_bert.py: configuration and model classes; README.md: opening maintenance warning, Examples and Data; June 2019 TAPE preprint Sections 4.1,5 and AppendixA.3 |
| Inputs | Amino-acid sequence in the tokenizer expected by the selected implementation.Sources (7)songlab-cal/tape: README.md; songlab-cal/tape: tape/models/modeling_lstm.py; songlab-cal/tape: tape/models/modeling_resnet.py; songlab-cal/tape: tape/models/modeling_unirep.py; songlab-cal/tape: tape/models/modeling_onehot.py; songlab-cal/tape: tape/models/modeling_bert.py; tape: Primary paper PDF · tape/models/modeling_bert.py: configuration and model classes; README.md: opening maintenance warning, Examples and Data; June 2019 TAPE preprint Sections 4.1,5 and AppendixA.3 |
| Outputs | Residue/sequence representations and predictions from the chosen task head.Sources (7)songlab-cal/tape: README.md; songlab-cal/tape: tape/models/modeling_lstm.py; songlab-cal/tape: tape/models/modeling_resnet.py; songlab-cal/tape: tape/models/modeling_unirep.py; songlab-cal/tape: tape/models/modeling_onehot.py; songlab-cal/tape: tape/models/modeling_bert.py; tape: Primary paper PDF · tape/models/modeling_bert.py: configuration and model classes; README.md: opening maintenance warning, Examples and Data; June 2019 TAPE preprint Sections 4.1,5 and AppendixA.3 |
| Parameters | 38M for the June 2019 preprint Transformer; do not apply that total to the later PyTorch defaults or every task head.Sources (7)songlab-cal/tape: README.md; songlab-cal/tape: tape/models/modeling_lstm.py; songlab-cal/tape: tape/models/modeling_resnet.py; songlab-cal/tape: tape/models/modeling_unirep.py; songlab-cal/tape: tape/models/modeling_onehot.py; songlab-cal/tape: tape/models/modeling_bert.py; tape: Primary paper PDF · tape/models/modeling_bert.py: configuration and model classes; README.md: opening maintenance warning, Examples and Data; June 2019 TAPE preprint Sections 4.1,5 and AppendixA.3 |
| Known versions | June 2019 TAPE preprint configurations and the later PyTorch package are distinct; the current README warns it is not an exact reproduction.Sources (7)songlab-cal/tape: README.md; songlab-cal/tape: tape/models/modeling_lstm.py; songlab-cal/tape: tape/models/modeling_resnet.py; songlab-cal/tape: tape/models/modeling_unirep.py; songlab-cal/tape: tape/models/modeling_onehot.py; songlab-cal/tape: tape/models/modeling_bert.py; tape: Primary paper PDF · tape/models/modeling_bert.py: configuration and model classes; README.md: opening maintenance warning, Examples and Data; June 2019 TAPE preprint Sections 4.1,5 and AppendixA.3 |
| Training data | The June 2019 TAPE preprint uses 31M Pfam domains, with held-out families and a separate random split; downstream task heads are fitted on their own task data.Sources (7)songlab-cal/tape: README.md; songlab-cal/tape: tape/models/modeling_lstm.py; songlab-cal/tape: tape/models/modeling_resnet.py; songlab-cal/tape: tape/models/modeling_unirep.py; songlab-cal/tape: tape/models/modeling_onehot.py; songlab-cal/tape: tape/models/modeling_bert.py; tape: Primary paper PDF · tape/models/modeling_bert.py: configuration and model classes; README.md: opening maintenance warning, Examples and Data; June 2019 TAPE preprint Sections 4.1,5 and AppendixA.3 |
| Training cutoff | The June 2019 TAPE preprint identifies the Pfam-domain corpus and split procedure but does not state one latest-sequence deposition date. Later package defaults are a different implementation. · Not reported in inspected sourcesSources (7)songlab-cal/tape: README.md; songlab-cal/tape: tape/models/modeling_lstm.py; songlab-cal/tape: tape/models/modeling_resnet.py; songlab-cal/tape: tape/models/modeling_unirep.py; songlab-cal/tape: tape/models/modeling_onehot.py; songlab-cal/tape: tape/models/modeling_bert.py; tape: Primary paper PDF · tape/models/modeling_bert.py: configuration and model classes; README.md: opening maintenance warning, Examples and Data; June 2019 TAPE preprint Sections 4.1,5 and AppendixA.3 |
| Context limits | The inspected default BertConfig sets max_position_embeddings to 8,096; this is an implementation default, not evidence that a historical TAPE checkpoint was trained at that length.Sources (7)songlab-cal/tape: README.md; songlab-cal/tape: tape/models/modeling_lstm.py; songlab-cal/tape: tape/models/modeling_resnet.py; songlab-cal/tape: tape/models/modeling_unirep.py; songlab-cal/tape: tape/models/modeling_onehot.py; songlab-cal/tape: tape/models/modeling_bert.py; tape: Primary paper PDF · tape/models/modeling_bert.py: configuration and model classes; README.md: opening maintenance warning, Examples and Data; June 2019 TAPE preprint Sections 4.1,5 and AppendixA.3 |
| Weights licence | Separate checkpoint-distribution terms are not stated in the inspected release documentation and licence material. The source-code licence alone is not recorded as an explicit weight grant. · Not reported in inspected sourcesSources (8)songlab-cal/tape: README.md; songlab-cal/tape: tape/models/modeling_lstm.py; songlab-cal/tape: tape/models/modeling_resnet.py; songlab-cal/tape: tape/models/modeling_unirep.py; songlab-cal/tape: tape/models/modeling_onehot.py; songlab-cal/tape: tape/models/modeling_bert.py; tape: Primary paper PDF; songlab-cal/tape: LICENSE · tape/models/modeling_bert.py: configuration and model classes; README.md: opening maintenance warning, Examples and Data; June 2019 TAPE preprint Sections 4.1,5 and AppendixA.3; LICENSE: licence text |
| Access | Official project documentation and implementation: https://github.com/songlab-cal/tapeSources (7)songlab-cal/tape: README.md; songlab-cal/tape: tape/models/modeling_lstm.py; songlab-cal/tape: tape/models/modeling_resnet.py; songlab-cal/tape: tape/models/modeling_unirep.py; songlab-cal/tape: tape/models/modeling_onehot.py; songlab-cal/tape: tape/models/modeling_bert.py; tape: Primary paper PDF · tape/models/modeling_bert.py: configuration and model classes; README.md: opening maintenance warning, Examples and Data; June 2019 TAPE preprint Sections 4.1,5 and AppendixA.3 |
| Code licence | BSD-3-ClauseSourcessonglab-cal/tape: LICENSE · LICENSE: licence text |
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
135 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation. Individual claims | songlab-cal/tape: README.md tape/models/modeling_bert.py: configuration and model classes; README.md: opening maintenance warning, Examples and Data; June 2019 TAPE preprint Sections 4.1,5 and AppendixA.3 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 6d345c2b2bbf52cd32cf179325c222afd92aec7e | source checked automated source review · 2026-09-16 Audit detailsInspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied. Field: Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
| Diagram caption Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation. Individual claims | songlab-cal/tape: tape/models/modeling_unirep.py tape/models/modeling_bert.py: configuration and model classes; README.md: opening maintenance warning, Examples and Data; June 2019 TAPE preprint Sections 4.1,5 and AppendixA.3 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 6d345c2b2bbf52cd32cf179325c222afd92aec7e | source checked automated source review · 2026-09-16 Audit detailsInspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied. Field: Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
| Diagram caption Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation. Individual claims | songlab-cal/tape: tape/models/modeling_onehot.py tape/models/modeling_bert.py: configuration and model classes; README.md: opening maintenance warning, Examples and Data; June 2019 TAPE preprint Sections 4.1,5 and AppendixA.3 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 6d345c2b2bbf52cd32cf179325c222afd92aec7e | source checked automated source review · 2026-09-16 Audit detailsInspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied. Field: Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
| Diagram caption Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation. Individual claims | songlab-cal/tape: tape/models/modeling_lstm.py tape/models/modeling_bert.py: configuration and model classes; README.md: opening maintenance warning, Examples and Data; June 2019 TAPE preprint Sections 4.1,5 and AppendixA.3 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 6d345c2b2bbf52cd32cf179325c222afd92aec7e | source checked automated source review · 2026-09-16 Audit detailsInspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied. Field: Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
| Diagram caption Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation. Individual claims | tape: Primary paper PDF tape/models/modeling_bert.py: configuration and model classes; README.md: opening maintenance warning, Examples and Data; June 2019 TAPE preprint Sections 4.1,5 and AppendixA.3 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 1906.08230v1 | source checked automated source review · 2026-09-16 Audit detailsInspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied. Field: Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
| Diagram caption Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation. Individual claims | songlab-cal/tape: tape/models/modeling_resnet.py tape/models/modeling_bert.py: configuration and model classes; README.md: opening maintenance warning, Examples and Data; June 2019 TAPE preprint Sections 4.1,5 and AppendixA.3 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 6d345c2b2bbf52cd32cf179325c222afd92aec7e | source checked automated source review · 2026-09-16 Audit detailsInspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied. Field: Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
| Diagram caption Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation. Individual claims | songlab-cal/tape: tape/models/modeling_bert.py tape/models/modeling_bert.py: configuration and model classes; README.md: opening maintenance warning, Examples and Data; June 2019 TAPE preprint Sections 4.1,5 and AppendixA.3 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 6d345c2b2bbf52cd32cf179325c222afd92aec7e | source checked automated source review · 2026-09-16 Audit detailsInspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied. Field: Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
Diagram steps
| songlab-cal/tape: README.md tape/models/modeling_bert.py: configuration and model classes; README.md: opening maintenance warning, Examples and Data; June 2019 TAPE preprint Sections 4.1,5 and AppendixA.3 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 6d345c2b2bbf52cd32cf179325c222afd92aec7e | source checked automated source review · 2026-09-16 Audit detailsInspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied. Field: Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
Diagram steps
| songlab-cal/tape: tape/models/modeling_unirep.py tape/models/modeling_bert.py: configuration and model classes; README.md: opening maintenance warning, Examples and Data; June 2019 TAPE preprint Sections 4.1,5 and AppendixA.3 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 6d345c2b2bbf52cd32cf179325c222afd92aec7e | source checked automated source review · 2026-09-16 Audit detailsInspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied. Field: Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
Diagram steps
| songlab-cal/tape: tape/models/modeling_onehot.py tape/models/modeling_bert.py: configuration and model classes; README.md: opening maintenance warning, Examples and Data; June 2019 TAPE preprint Sections 4.1,5 and AppendixA.3 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: 6d345c2b2bbf52cd32cf179325c222afd92aec7e | source checked automated source review · 2026-09-16 Audit detailsInspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied. Field: Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
View linked audit checks and correction history
Release 2026-09-29-06401fd5b220 · Record review: discovered
Stable ID: discovery-model-tape-transformer