rewire.itbenchmarks
Benchmark

TAPE

TAPE evaluates protein representations through five supervised downstream tasks.

Sourcessonglab-cal/tape official source · Pinned README: compatibility notice; Evaluating a Downstream Model; Data; task leaderboards

50 evaluations · 50 results

Overview

Datasets

Secondary structure, contacts, remote homology, fluorescence and stability; a Pfam pretraining corpus is supplied separately.

Metrics

Task leaderboards use three-class accuracy, contact ranking, top-1 homology accuracy and Spearman correlation for fluorescence/stability.

Allowed inputs

Protein amino-acid sequences and task-specific labels.

Sourcessonglab-cal/tape official source · Pinned README: compatibility notice; Evaluating a Downstream Model; Data; task leaderboards
Evaluation procedure diagram
How it worksEvaluation procedure
Evaluation procedure1. Allowed inputs: Protein amino-acid sequences and task-specific labels.. Then: 2. Splits: Secondary structure uses 25% identity filtering; contacts use ProteinNet/CASP12 with 30% filtering; remote homology holds out superfamilies; fluorescence holds out greater mutation distances; stability holds out selected mutation neighbourhoods.. Then: 3. Metrics: Task leaderboards use three-class accuracy, contact ranking, top-1 homology accuracy and Spearman correlation for fluorescence/stability.Evaluation procedure1. Allowed inputs: Protein amino-acid sequences and task-specific labels.. Then: 2. Splits: Secondary structure uses 25% identity filtering; contacts use ProteinNet/CASP12 with 30% filtering; remote homology holds out superfamilies; fluorescence holds out greater mutation distances; stability holds out selected mutation neighbourhoods.. Then: 3. Metrics: Task leaderboards use three-class accuracy, contact ranking, top-1 homology accuracy and Spearman correlation for fluorescence/stability.Evaluation procedure1. Allowed inputs: Protein amino-acid sequences and task-specific labels.. Then: 2. Splits: Secondary structure uses 25% identity filtering; contacts use ProteinNet/CASP12 with 30% filtering; remote homology holds out superfamilies; fluorescence holds out greater mutation distances; stability holds out selected mutation neighbourhoods.. Then: 3. Metrics: Task leaderboards use three-class accuracy, contact ranking, top-1 homology accuracy and Spearman correlation for fluorescence/stability.

Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.

Sources (2)songlab-cal/tape official source; tape primary benchmark evidence · Pinned README: compatibility notice; Evaluating a Downstream Model; Data; task leaderboards; Section 3; Appendix A.1

limited source coverage · Automated source review, 2026-09-16. All specifications and missing details

Results

Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.

Secondary structure in original TAPE Table 2

accuracy (fraction) · Higher values are better.

TAPE Secondary Structure · CB513

Evidence origin: Author-reported evaluation, Independent external evaluation.

tape primary benchmark evidence · Results on downstream supervised tasks; Table2,p.7,row1(No pretraining/Transformer),columnSecondary structure; Table2,p.7,row2(No pretraining/LSTM),columnSecondary structure; Table2,p.7,row3(No pretraining/ResNet),columnSecondary structure; Table2,p.7,row4(Self-supervised pretraining/Transformer),columnSecondary structure; Table2,p.7,row5(Self-supervised pretraining/LSTM),columnSecondary structure; Table2,p.7,row6(Self-supervised pretraining/ResNet),columnSecondary structure; Table2,p.7,row7(Supervised pretraining/Bepler supervised LSTM),columnSecondary structure; Table2,p.7,row8(Self-supervised pretraining/UniRep mLSTM),columnSecondary structure; Table2,p.7,row9(One-hot baseline/One-hot),columnSecondary structure; Table2,p.7,row10(Task-specific alignment/HMM baseline/Alignment),columnSecondary structure
  • Same task split and supervised downstream architecture within this paper. Input information differs for alignment/HMM features; pretraining regimes are explicit. Do not pool scores across tasks or with later TAPE implementations. Origins are labelled per method; appearance in one table does not constitute independent replication of every model.
Comparison details and limitations

CB513 test set; train/test proteins filtered at 25% sequence identity

Automated source review: 2026-09-17. Numerical source review does not establish independent reproduction.

Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.

Showing 10 of 10 matching rows.

Methods and evaluation design

Procedure, tasks and evaluated configurations

How it works

Evaluation methodology

TAPE evaluates protein representations on five tasks: local secondary structure, contacts, remote homology, fluorescence and stability. Each has a biologically motivated supervised partition, and the original paper compares pretrained representations with untrained and alignment-based controls using task-specific prediction heads.

Sourcestape primary benchmark evidence · Section 3; Appendix A.1

Evaluation design

Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.

These source-backed links do not make different protocols or scores interchangeable.

Baseline coverage

Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.

No concrete protocols are explicitly linked to this suite. Protocol identification and baseline selection are outstanding.

Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums

Coverage is derived from release 2026-09-29-06401fd5b220. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.

Official run instructions

Evaluate an already fine-tuned secondary-structure model with the official PyTorch TAPE CLI. This is not a reproduction recipe for the original TensorFlow paper.

Checked against the official instructions on 2026-09-17. These commands have not been executed by rewire. Running them does not automatically reproduce the published scores.

Before you start

  • An isolated Python environment and an existing TAPE-compatible transformer checkpoint fine-tuned for secondary structure; a pretrained language-model checkpoint alone is insufficient.
  • Place the secondary-structure LMDB dataset linked in README.md under ./data. The README reports approximately 120 MB compressed / 2 GB unpacked for all supervised datasets.
  1. 1. Check out the reviewed repository

    Repository checkout wrapper: the detached revision selects the exact official source inspected for this guide.

    git clone https://github.com/songlab-cal/tape.git
    cd tape
    git checkout --detach 6d345c2b2bbf52cd32cf179325c222afd92aec7e
    songlab-cal/tape / README.md · Pinned repository revision; README.md
  2. 2. Install the documented package

    This is the README installation command. It does not pin the PyPI package version; record the resolved environment before comparing runs.

    pip install tape_proteins
    songlab-cal/tape / README.md · README.md lines 40–46
  3. 3. Evaluate the task-trained checkpoint

    Set TAPE_TRAINED_MODEL to the directory containing your fine-tuned model. This shell variable replaces the README placeholder results/<path_to_trained_model> without inventing a checkpoint.

    tape-eval transformer secondary_structure "${TAPE_TRAINED_MODEL:?Set TAPE_TRAINED_MODEL to your task-trained model directory}" --metrics accuracy
    songlab-cal/tape / README.md · README.md lines 172–188; Data lines 240–254

Expected outputs

  • Reported overall secondary-structure accuracy and a results.pkl file written into the trained model directory.

Scope and limitations

  • The official README warns that this PyTorch repository deliberately differs from the original paper and directs exact paper reproduction to tape-neurips2019.
  • The maintainers no longer recommend the bundled training code and do not maintain newer-PyTorch training compatibility. This guide covers evaluation only.
  • The README does not pin the package environment or checkpoint/data hashes; supply and record these before treating a run as reproducible.
  • The inspected instructions do not establish a minimum RAM/VRAM requirement, wall-clock runtime or monetary cost; none is inferred.
Strengths, limitations and unresolved questions

Strengths and limitations

Strengths and considerations

  • Distinct sequence, residue and protein-level tasks avoid relying on language-model perplexity as a proxy for transfer.
    Sourcessonglab-cal/tape official source · Pinned README: compatibility notice; Evaluating a Downstream Model; Data; task leaderboards

Limitations and conditions

  • The original TAPE paper and the later PyTorch implementation are distinct versions. Preserve the task split and implementation; pretraining exposure is separate from supervised sequence-identity filtering.
    Sourcestape primary benchmark evidence · Section 3; Appendix A.1; Tables 1–2
Profile review details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Stable record: discovery-benchmark-tape

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetsSecondary structure, contacts, remote homology, fluorescence and stability; a Pfam pretraining corpus is supplied separately.
Sourcessonglab-cal/tape official source · Pinned README: compatibility notice; Evaluating a Downstream Model; Data; task leaderboards
SplitsSecondary structure uses 25% identity filtering; contacts use ProteinNet/CASP12 with 30% filtering; remote homology holds out superfamilies; fluorescence holds out greater mutation distances; stability holds out selected mutation neighbourhoods.
Sourcestape primary benchmark evidence · Section 3; Appendix A.1
MetricsTask leaderboards use three-class accuracy, contact ranking, top-1 homology accuracy and Spearman correlation for fluorescence/stability.
Sourcessonglab-cal/tape official source · Pinned README: compatibility notice; Evaluating a Downstream Model; Data; task leaderboards
BaselinesTransformer, LSTM, ResNet, UniRep and one-hot baselines.
Sourcessonglab-cal/tape official source · Pinned README: compatibility notice; Evaluating a Downstream Model; Data; task leaderboards
Leakage controlsThe five tasks use different supervised generalization boundaries. Protein identity, evolutionary groups and mutational distance are not interchangeable, and none is a universal audit of unsupervised pretraining overlap.
Sourcestape primary benchmark evidence · Section 3; Appendix A.1
UncertaintyThe original paper reports point estimates in its task result tables. Methods and task appendices do not define a suite-wide repeated-seed or bootstrap interval; individual later evaluations must supply their own uncertainty. · Not reported in inspected sources
Sourcestape primary benchmark evidence · Section 3; Appendix A.1; Tables 1–2
Entity typeProtein-representation benchmark suite.
Sourcessonglab-cal/tape official source · Pinned README: compatibility notice; Evaluating a Downstream Model; Data; task leaderboards
OrganismsMixed protein-domain and structural sources rather than a species-held-out benchmark; the engineering tasks use green fluorescent protein variants and designed stability landscapes.
Sourcestape primary benchmark evidence · Section 3; Appendix A.1; Tables 1–2
AssaysProtein structure/homology annotations and experimental fluorescence/stability measurements.
Sourcessonglab-cal/tape official source · Pinned README: compatibility notice; Evaluating a Downstream Model; Data; task leaderboards
Allowed inputsProtein amino-acid sequences and task-specific labels.
Sourcessonglab-cal/tape official source · Pinned README: compatibility notice; Evaluating a Downstream Model; Data; task leaderboards
AdaptationUnsupervised pretraining followed by supervised downstream training; the README warns that downstream hyperparameters require task-specific tuning.
Sourcessonglab-cal/tape official source · Pinned README: compatibility notice; Evaluating a Downstream Model; Data; task leaderboards
Applicable tests and references

Applicability is distinct from a completed evaluation.

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Dated primary-source discovery and protocol/table screening. Source checking does not mean experimental reproduction. Only separately extracted and independently reviewed numeric batches are publishable.

Paper or primary resourceVersionReference
Evaluating Protein Transfer Learning with TAPE1906.08230v1Read source
Search and extraction details

complete tables extracted

Searches

  • TAPE protein representation learning benchmark Rao 2019 table 1 2 original paper

Evidence locations

  • Original arXiv v1 Table 2, p.7; Table S1, p.14; Appendix A.2

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

24 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Individual claims
tape primary benchmark evidence

Original source ↗

Pinned README: compatibility notice; Evaluating a Downstream Model; Data; task leaderboards; Section 3; Appendix A.1

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 1906.08230v1
Retrieved: 2026-09-16T20:23:48.238777+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 6ee0c3e6e870635cba8fa67e0a4abc5598c0ab2a10ba127a46b67ac450ae0168

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram caption
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Individual claims
songlab-cal/tape official source

Original source ↗

Pinned README: compatibility notice; Evaluating a Downstream Model; Data; task leaderboards; Section 3; Appendix A.1

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 6d345c2b2bbf52cd32cf179325c222afd92aec7e
Retrieved: 2026-09-16T10:31:32.443187+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: b28c74fe3cd6b69a8ba6d84891d0539e54dfef882abd5ed4d11ed0b029bb477a

Hash scope: Hash scope not separately documented; inspect source record

Diagram steps
  • Allowed inputs: Protein amino-acid sequences and task-specific labels.
  • Splits: Secondary structure uses 25% identity filtering; contacts use ProteinNet/CASP12 with 30% filtering; remote homology holds out superfamilies; fluorescence holds out greater mutation distances; stability holds out selected mutation neighbourhoods.
  • Metrics: Task leaderboards use three-class accuracy, contact ranking, top-1 homology accuracy and Spearman correlation for fluorescence/stability.
Individual claims
tape primary benchmark evidence

Original source ↗

Pinned README: compatibility notice; Evaluating a Downstream Model; Data; task leaderboards; Section 3; Appendix A.1

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 1906.08230v1
Retrieved: 2026-09-16T20:23:48.238777+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 6ee0c3e6e870635cba8fa67e0a4abc5598c0ab2a10ba127a46b67ac450ae0168

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps
  • Allowed inputs: Protein amino-acid sequences and task-specific labels.
  • Splits: Secondary structure uses 25% identity filtering; contacts use ProteinNet/CASP12 with 30% filtering; remote homology holds out superfamilies; fluorescence holds out greater mutation distances; stability holds out selected mutation neighbourhoods.
  • Metrics: Task leaderboards use three-class accuracy, contact ranking, top-1 homology accuracy and Spearman correlation for fluorescence/stability.
Individual claims
songlab-cal/tape official source

Original source ↗

Pinned README: compatibility notice; Evaluating a Downstream Model; Data; task leaderboards; Section 3; Appendix A.1

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 6d345c2b2bbf52cd32cf179325c222afd92aec7e
Retrieved: 2026-09-16T10:31:32.443187+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: b28c74fe3cd6b69a8ba6d84891d0539e54dfef882abd5ed4d11ed0b029bb477a

Hash scope: Hash scope not separately documented; inspect source record

Diagram title
Evaluation procedure
Individual claims
tape primary benchmark evidence

Original source ↗

Pinned README: compatibility notice; Evaluating a Downstream Model; Data; task leaderboards; Section 3; Appendix A.1

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 1906.08230v1
Retrieved: 2026-09-16T20:23:48.238777+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: 6ee0c3e6e870635cba8fa67e0a4abc5598c0ab2a10ba127a46b67ac450ae0168

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title
Evaluation procedure
Individual claims
songlab-cal/tape official source

Original source ↗

Pinned README: compatibility notice; Evaluating a Downstream Model; Data; task leaderboards; Section 3; Appendix A.1

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 6d345c2b2bbf52cd32cf179325c222afd92aec7e
Retrieved: 2026-09-16T10:31:32.443187+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: b28c74fe3cd6b69a8ba6d84891d0539e54dfef882abd5ed4d11ed0b029bb477a

Hash scope: Hash scope not separately documented; inspect source record

Datasets
Secondary structure, contacts, remote homology, fluorescence and stability; a Pfam pretraining corpus is supplied separately.
Individual claims
songlab-cal/tape official source

Original source ↗

Pinned README: compatibility notice; Evaluating a Downstream Model; Data; task leaderboards

Version: 6d345c2b2bbf52cd32cf179325c222afd92aec7e
Retrieved: 2026-09-16T10:31:32.443187+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: b28c74fe3cd6b69a8ba6d84891d0539e54dfef882abd5ed4d11ed0b029bb477a

Hash scope: Hash scope not separately documented; inspect source record

Splits
Secondary structure uses 25% identity filtering; contacts use ProteinNet/CASP12 with 30% filtering; remote homology holds out superfamilies; fluorescence holds out greater mutation distances; stability holds out selected mutation neighbourhoods.
Individual claims
tape primary benchmark evidence

Original source ↗

Section 3; Appendix A.1

Version: 1906.08230v1
Retrieved: 2026-09-16T20:23:48.238777+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: 6ee0c3e6e870635cba8fa67e0a4abc5598c0ab2a10ba127a46b67ac450ae0168

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Adaptation
Unsupervised pretraining followed by supervised downstream training; the README warns that downstream hyperparameters require task-specific tuning.
Individual claims
songlab-cal/tape official source

Original source ↗

Pinned README: compatibility notice; Evaluating a Downstream Model; Data; task leaderboards

Version: 6d345c2b2bbf52cd32cf179325c222afd92aec7e
Retrieved: 2026-09-16T10:31:32.443187+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: b28c74fe3cd6b69a8ba6d84891d0539e54dfef882abd5ed4d11ed0b029bb477a

Hash scope: Hash scope not separately documented; inspect source record

Metrics
Task leaderboards use three-class accuracy, contact ranking, top-1 homology accuracy and Spearman correlation for fluorescence/stability.
Individual claims
songlab-cal/tape official source

Original source ↗

Pinned README: compatibility notice; Evaluating a Downstream Model; Data; task leaderboards

Version: 6d345c2b2bbf52cd32cf179325c222afd92aec7e
Retrieved: 2026-09-16T10:31:32.443187+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: b28c74fe3cd6b69a8ba6d84891d0539e54dfef882abd5ed4d11ed0b029bb477a

Hash scope: Hash scope not separately documented; inspect source record

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: discovered

4 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: discovery-benchmark-tape

areas
protein-function
entity level
suite
scope note
Specialist molecular or omics evaluation; protocol details require review before numerical comparison.
task
Protein representation learning tasks
version
Not reported
comparison panels
id: tape-2019-table2-secondary-structure; title: Secondary structure in original TAPE Table 2; protocol id: discovery-benchmark-tape-secondary-structure; dataset id: paper-dataset-c40c5d875e595c406a; metric: accuracy; unit: fraction; direction: higher; result ids: paper-result-265fa169813f898ad7; paper-result-aa372353bb508c7305; paper-result-8843b87cc984941e2b; paper-result-b1fec931ad74219d4c; paper-result-e1f2e80ac27d7d9796; paper-result-f4bb251a370309916e; paper-result-1dd55d28805a820058; paper-result-873ec90fb6395c3fce; paper-result-d400d760b9f0abd606; paper-result-51fae590bba5e35d40; source ids: evidence-discovery-final-tape; source locator: Results on downstream supervised tasks; Table2,p.7,row1(No pretraining/Transformer),columnSecondary structure; Table2,p.7,row2(No pretraining/LSTM),columnSecondary structure; Table2,p.7,row3(No pretraining/ResNet),columnSecondary structure; Table2,p.7,row4(Self-supervised pretraining/Transformer),columnSecondary structure; Table2,p.7,row5(Self-supervised pretraining/LSTM),columnSecondary structure; Table2,p.7,row6(Self-supervised pretraining/ResNet),columnSecondary structure; Table2,p.7,row7(Supervised pretraining/Bepler supervised LSTM),columnSecondary structure; Table2,p.7,row8(Self-supervised pretraining/UniRep mLSTM),columnSecondary structure; Table2,p.7,row9(One-hot baseline/One-hot),columnSecondary structure; Table2,p.7,row10(Task-specific alignment/HMM baseline/Alignment),columnSecondary structure; context: CB513 test set; train/test proteins filtered at 25% sequence identity; caveats: Same task split and supervised downstream architecture within this paper. Input information differs for alignment/HMM features; pretraining regimes are explicit. Do not pool scores across tasks or with later TAPE implementations. Origins are labelled per method; appearance in one table does not constitute independent replication of every model.; review: method: automated_source_review; date: 2026-09-17; id: tape-2019-table2-contact-prediction; title: Contact prediction in original TAPE Table 2; protocol id: discovery-benchmark-tape-contact-prediction; dataset id: paper-dataset-b2114365a5517543b7; metric: precision at L/5; unit: fraction; direction: higher; result ids: paper-result-db4756024c9a9b11a9; paper-result-82ef40ee0784b9d645; paper-result-62b3b957e3f7390e67; paper-result-0c4a6be02b3815342f; paper-result-d161cb7797686a25fc; paper-result-0311b73601922d0209; paper-result-e4f7b63768be9c9e2c; paper-result-8ea9c35dce1fcd6959; paper-result-ea124a67c83fa7bf99; paper-result-a4faa27255330fb9f2; source ids: evidence-discovery-final-tape; source locator: Results on downstream supervised tasks; Table2,p.7,row1(No pretraining/Transformer),columnContact prediction; Table2,p.7,row2(No pretraining/LSTM),columnContact prediction; Table2,p.7,row3(No pretraining/ResNet),columnContact prediction; Table2,p.7,row4(Self-supervised pretraining/Transformer),columnContact prediction; Table2,p.7,row5(Self-supervised pretraining/LSTM),columnContact prediction; Table2,p.7,row6(Self-supervised pretraining/ResNet),columnContact prediction; Table2,p.7,row7(Supervised pretraining/Bepler supervised LSTM),columnContact prediction; Table2,p.7,row8(Self-supervised pretraining/UniRep mLSTM),columnContact prediction; Table2,p.7,row9(One-hot baseline/One-hot),columnContact prediction; Table2,p.7,row10(Task-specific alignment/HMM baseline/Alignment),columnContact prediction; context: CASP12 test, 30% sequence-identity partition; medium/long-range contacts; caveats: Same task split and supervised downstream architecture within this paper. Input information differs for alignment/HMM features; pretraining regimes are explicit. Do not pool scores across tasks or with later TAPE implementations. Origins are labelled per method; appearance in one table does not constitute independent replication of every model.; review: method: automated_source_review; date: 2026-09-17; id: tape-2019-table2-remote-homology-detection; title: Remote homology in original TAPE Table 2; protocol id: discovery-benchmark-tape-remote-homology-detection; dataset id: paper-dataset-b84068b24ba74c6f4a; metric: accuracy; unit: fraction; direction: higher; result ids: paper-result-903d5d36162835d916; paper-result-e3595d7ae7d4b30746; paper-result-f0df0b60fea8d2cbd8; paper-result-eaedde3a911fefd0f0; paper-result-5cba527e65eec40f91; paper-result-5514a0ebdd6e527021; paper-result-5a6fc3ad0ed3cd1c6e; paper-result-b8ae4a359f201ae45f; paper-result-6067a3c9fb922d3c81; paper-result-6de41027c29faf11f2; source ids: evidence-discovery-final-tape; source locator: Results on downstream supervised tasks; Table2,p.7,row1(No pretraining/Transformer),columnRemote homology; Table2,p.7,row2(No pretraining/LSTM),columnRemote homology; Table2,p.7,row3(No pretraining/ResNet),columnRemote homology; Table2,p.7,row4(Self-supervised pretraining/Transformer),columnRemote homology; Table2,p.7,row5(Self-supervised pretraining/LSTM),columnRemote homology; Table2,p.7,row6(Self-supervised pretraining/ResNet),columnRemote homology; Table2,p.7,row7(Supervised pretraining/Bepler supervised LSTM),columnRemote homology; Table2,p.7,row8(Self-supervised pretraining/UniRep mLSTM),columnRemote homology; Table2,p.7,row9(One-hot baseline/One-hot),columnRemote homology; Table2,p.7,row10(Task-specific alignment/HMM baseline/Alignment),columnRemote homology; context: Held-out evolutionary groups; fold-level classification into 1,195 folds; caveats: Same task split and supervised downstream architecture within this paper. Input information differs for alignment/HMM features; pretraining regimes are explicit. Do not pool scores across tasks or with later TAPE implementations. Origins are labelled per method; appearance in one table does not constitute independent replication of every model.; review: method: automated_source_review; date: 2026-09-17; id: tape-2019-table2-fluorescence; title: Fluorescence in original TAPE Table 2; protocol id: discovery-benchmark-tape-fluorescence; dataset id: discovery-dataset-tape-fluorescence-source-dataset; metric: Spearman's rho; unit: correlation; direction: higher; result ids: paper-result-4d32e7c25ddb70606c; paper-result-674218c8708faff017; paper-result-800c13b04edf94254c; discovery-result-tape-fluorescence-transformer-spearman-rho; discovery-result-tape-fluorescence-lstm-spearman-rho; discovery-result-tape-fluorescence-resnet-spearman-rho; discovery-result-tape-fluorescence-bepler-spearman-rho; discovery-result-tape-fluorescence-unirep-spearman-rho; discovery-result-tape-fluorescence-one-hot-spearman-rho; paper-result-6f9681dfe5d73fc00d; source ids: evidence-discovery-final-tape; source locator: Results on downstream supervised tasks; Table2,p.7,row1(No pretraining/Transformer),columnFluorescence; Table2,p.7,row2(No pretraining/LSTM),columnFluorescence; Table2,p.7,row3(No pretraining/ResNet),columnFluorescence; Table2,p.7,row4(Self-supervised pretraining/Transformer),columnFluorescence; Table2,p.7,row5(Self-supervised pretraining/LSTM),columnFluorescence; Table2,p.7,row6(Self-supervised pretraining/ResNet),columnFluorescence; Table2,p.7,row7(Supervised pretraining/Bepler supervised LSTM),columnFluorescence; Table2,p.7,row8(Self-supervised pretraining/UniRep mLSTM),columnFluorescence; Table2,p.7,row9(One-hot baseline/One-hot),columnFluorescence; Table2,p.7,row10(Task-specific alignment/HMM baseline/Alignment),columnFluorescence; context: Train/validation at most 3 mutations; test 4–15 mutations from the GFP parent; caveats: Same task split and supervised downstream architecture within this paper. Input information differs for alignment/HMM features; pretraining regimes are explicit. Do not pool scores across tasks or with later TAPE implementations. Origins are labelled per method; appearance in one table does not constitute independent replication of every model.; review: method: automated_source_review; date: 2026-09-17; id: tape-2019-table2-stability; title: Stability in original TAPE Table 2; protocol id: discovery-benchmark-tape-stability; dataset id: discovery-dataset-tape-stability-source-dataset; metric: Spearman's rho; unit: correlation; direction: higher; result ids: paper-result-61c44f4b41cb5b0b1d; paper-result-440464f32880646f52; paper-result-d6e867039a2d623b60; discovery-result-tape-stability-transformer-spearman-rho; discovery-result-tape-stability-lstm-spearman-rho; discovery-result-tape-stability-resnet-spearman-rho; discovery-result-tape-stability-bepler-spearman-rho; discovery-result-tape-stability-unirep-spearman-rho; discovery-result-tape-stability-one-hot-spearman-rho; paper-result-1a0859986a2951d9fb; source ids: evidence-discovery-final-tape; source locator: Results on downstream supervised tasks; Table2,p.7,row1(No pretraining/Transformer),columnStability; Table2,p.7,row2(No pretraining/LSTM),columnStability; Table2,p.7,row3(No pretraining/ResNet),columnStability; Table2,p.7,row4(Self-supervised pretraining/Transformer),columnStability; Table2,p.7,row5(Self-supervised pretraining/LSTM),columnStability; Table2,p.7,row6(Self-supervised pretraining/ResNet),columnStability; Table2,p.7,row7(Supervised pretraining/Bepler supervised LSTM),columnStability; Table2,p.7,row8(Self-supervised pretraining/UniRep mLSTM),columnStability; Table2,p.7,row9(One-hot baseline/One-hot),columnStability; Table2,p.7,row10(Task-specific alignment/HMM baseline/Alignment),columnStability; context: Test 17 one-mutation neighbourhoods around candidate proteins; train on broad experimental design rounds; caveats: Same task split and supervised downstream architecture within this paper. Input information differs for alignment/HMM features; pretraining regimes are explicit. Do not pool scores across tasks or with later TAPE implementations. Origins are labelled per method; appearance in one table does not constitute independent replication of every model.; review: method: automated_source_review; date: 2026-09-17
benchmark research
review date: 2026-09-17; status: complete_tables_extracted; primary sources: evidence-expansion-evidence-discovery-final-tape-6ee0c3e6; inspected locators: Original arXiv v1 Table 2, p.7; Table S1, p.14; Appendix A.2; searched queries: TAPE protein representation learning benchmark Rao 2019 table 1 2 original paper; gaps: None recorded; claim scope: Dated primary-source discovery and protocol/table screening. Source checking does not mean experimental reproduction. Only separately extracted and independently reviewed numeric batches are publishable.
historical missing metadata
dataset release: unextracted; metric implementation: unextracted; split manifest: unextracted; version: unextracted
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
entity classification
review date: 2026-09-17; rationale: The cited profile describes a collection of evaluation tasks or protocols; retain it as the top-level benchmark suite. Its datasets and individual protocols remain separate records.; source ids: src-discovery-songlab-cal-tape; source locator: Pinned README: compatibility notice; Evaluating a Downstream Model; Data; task leaderboards; ambiguities: None recorded
run guide
record id: discovery-benchmark-tape; summary: Evaluate an already fine-tuned secondary-structure model with the official PyTorch TAPE CLI. This is not a reproduction recipe for the original TensorFlow paper.; status: source_reviewed_not_executed; prerequisites: An isolated Python environment and an existing TAPE-compatible transformer checkpoint fine-tuned for secondary structure; a pretrained language-model checkpoint alone is insufficient.; Place the secondary-structure LMDB dataset linked in README.md under ./data. The README reports approximately 120 MB compressed / 2 GB unpacked for all supervised datasets.; steps: title: Check out the reviewed repository; shell: git clone https://github.com/songlab-cal/tape.git cd tape git checkout --detach 6d345c2b2bbf52cd32cf179325c222afd92aec7e; explanation: Repository checkout wrapper: the detached revision selects the exact official source inspected for this guide.; source ids: run-doc-tape-readme-md-6d345c2b; source locator: Pinned repository revision; README.md; title: Install the documented package; shell: pip install tape_proteins; explanation: This is the README installation command. It does not pin the PyPI package version; record the resolved environment before comparing runs.; source ids: run-doc-tape-readme-md-6d345c2b; source locator: README.md lines 40–46; title: Evaluate the task-trained checkpoint; shell: tape-eval transformer secondary_structure "${TAPE_TRAINED_MODEL:?Set TAPE_TRAINED_MODEL to your task-trained model directory}" --metrics accuracy; explanation: Set TAPE_TRAINED_MODEL to the directory containing your fine-tuned model. This shell variable replaces the README placeholder results/<path_to_trained_model> without inventing a checkpoint.; source ids: run-doc-tape-readme-md-6d345c2b; source locator: README.md lines 172–188; Data lines 240–254; outputs: Reported overall secondary-structure accuracy and a results.pkl file written into the trained model directory.; limitations: The official README warns that this PyTorch repository deliberately differs from the original paper and directs exact paper reproduction to tape-neurips2019.; The maintainers no longer recommend the bundled training code and do not maintain newer-PyTorch training compatibility. This guide covers evaluation only.; The README does not pin the package environment or checkpoint/data hashes; supply and record these before treating a run as reproducible.; The inspected instructions do not establish a minimum RAM/VRAM requirement, wall-clock runtime or monetary cost; none is inferred.; source ids: run-doc-tape-readme-md-6d345c2b; review: method: official_repository_review; date: 2026-09-17
run documentation
record id: discovery-benchmark-tape; source ids: run-doc-tape-readme-md-6d345c2b; status: source_reviewed_not_executed; summary: Evaluate an already fine-tuned secondary-structure model with the official PyTorch TAPE CLI. This is not a reproduction recipe for the original TensorFlow paper. Commands were source-reviewed only. The official README warns that this PyTorch repository deliberately differs from the original paper and directs exact paper reproduction to tape-neurips2019.; source locator: Pinned repository revision; README.md; README.md lines 40–46; README.md lines 172–188; Data lines 240–254
Related records

Suggest a correction