rewire.itbenchmarks
Task

human-versus-viral protein classification

This evaluation asks whether protein representations distinguish human and viral sequence labels. It is a classification task, not a direct test of immune function.

SourcesProtein Language Models Expose Viral Immune Mimicry · Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1

8 evaluations · 32 results

Overview

Datasets

Reviewed human and vertebrate-host viral protein records from Swiss-Prot/UniProtKB, with redundancy filtering described in Methods 2.1.

Metrics

AUROC, log loss, accuracy, precision and recall are described; precision and recall use macro averaging. The linked result retains its original percentage unit.

Allowed inputs

Protein sequence representations paired with the study’s human/viral labels.

SourcesProtein Language Models Expose Viral Immune Mimicry · Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1
Evaluation procedure diagram
How it worksConceptual assessment outline
Conceptual assessment outline1. Curated sequence-origin labels. Then: 2. Keep sequence clusters separate. Then: 3. Assess held-out classifications. Then: 4. Report classification metricsConceptual assessment outline1. Curated sequence-origin labels. Then: 2. Keep sequence clusters separate. Then: 3. Assess held-out classifications. Then: 4. Report classification metricsConceptual assessment outline1. Curated sequence-origin labels. Then: 2. Keep sequence clusters separate. Then: 3. Assess held-out classifications. Then: 4. Report classification metrics

Conceptual overview of the published statistical assessment.

SourcesProtein Language Models Expose Viral Immune Mimicry · Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1

limited source coverage · Automated source review, 2026-09-16. All specifications and missing details

Results

Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.

Held-out human-versus-virus protein classification · Table 1

AUROC (percent) · Higher values are better.

Held-out human-versus-virus protein classification (human-versus-viral protein classification) · human and viral proteins

Evidence origin: Author-reported evaluation.

Protein Language Models Expose Viral Immune Mimicry · Table 1: AUC (%), Held-out human-versus-virus protein classification
  • Origin classification does not establish immune mimicry.
  • Table footnote says all values are percent but log-loss values are dimensionless; log-loss units quarantined.
  • Narrative n-gram AUC91.9 differs from Table1 value91.5; preserve table value.
Comparison details and limitations

Compare protein origin classifiers; T 5 embeddings plus linear/tree models vs ESM 2 fine-tuning and simple controls. UniRef90 deduplication; proteins longer than1,600 residues excluded; UniRef50 clusters assigned80% training and20% test with no cluster shared. Separate four-fold error-analysis experiment not assigned toTable1.

  • No interval assigned unless printed in source cell.

Automated source review: 2026-09-17. Numerical source review does not establish independent reproduction.

Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.

Showing 8 of 8 matching rows.

Methods and evaluation design

Procedure, tasks and evaluated configurations

How it works

What the evaluation establishes

The source uses reviewed protein records and separates training and test data by sequence clusters. That reduces direct overlap between related examples under the stated clustering rule. It reports classification metrics for different representations; classification errors and biological explanations of those errors are separate claims.

SourcesProtein Language Models Expose Viral Immune Mimicry · Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1

Evaluation design

Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.

These source-backed links do not make different protocols or scores interchangeable.

Run this benchmark

Choose a concrete protocol before running an evaluation. Its inputs, split and scoring rules determine which results can be compared.

Run instructions

No runnable recipe has been reviewed for this task. Dataset access, model requirements, licences and compute requirements must be checked against its sources before execution.

A task describes a biological question. Choose a linked protocol to obtain concrete split and scoring instructions.

Strengths, limitations and unresolved questions

Strengths and limitations

Strengths and considerations

Limitations and conditions

  • Distinguishing sequence origin does not establish an immune mechanism. Database selection, similarity filtering and label composition constrain generalisation.
    SourcesProtein Language Models Expose Viral Immune Mimicry · Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1
Profile review details

Inspected the cited primary abstract and descriptive computational-evaluation sections. Review covers the descriptive claims shown; no executable protocol was reconstructed. Original numerical records retain their prior transcription review.

Stable record: reported-task-53506fe386e4a1

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
Entity typePaper-specific evaluation task; this profile is a descriptive evidence summary.
SourcesProtein Language Models Expose Viral Immune Mimicry · Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1
DatasetsReviewed human and vertebrate-host viral protein records from Swiss-Prot/UniProtKB, with redundancy filtering described in Methods 2.1.
SourcesProtein Language Models Expose Viral Immune Mimicry · Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1
OrganismsHuman proteins and proteins from viruses with a known vertebrate host.
SourcesProtein Language Models Expose Viral Immune Mimicry · Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1
AssaysSequence-origin labels from curated database records. This classification endpoint is not an experimental immune-response assay.
SourcesProtein Language Models Expose Viral Immune Mimicry · Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1
SplitsMethods 2.1 assigns whole UniRef50 clusters to training or test sets. This profile records the split principle, not a verified membership manifest.
SourcesProtein Language Models Expose Viral Immune Mimicry · Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1
Allowed inputsProtein sequence representations paired with the study’s human/viral labels.
SourcesProtein Language Models Expose Viral Immune Mimicry · Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1
AdaptationPretrained protein representations are evaluated through a study-specific classifier; a backbone name alone does not identify the full fitted pipeline.
SourcesProtein Language Models Expose Viral Immune Mimicry · Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1
MetricsAUROC, log loss, accuracy, precision and recall are described; precision and recall use macro averaging. The linked result retains its original percentage unit.
SourcesProtein Language Models Expose Viral Immune Mimicry · Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1
BaselinesTable 1 compares the reported protein-representation configurations. They are classification comparators, not experimental immune-function controls.
SourcesProtein Language Models Expose Viral Immune Mimicry · Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.

Paper or primary resourceVersionReference
Protein Language Models Expose Viral Immune Mimicryversion of recordRead source
DOI: 10.3390/v17091199
Historical gaps recorded on 2026-09-17

The catalogue now holds 32 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.

  • Independent batch review before import; preserve existing observation identities.
Search and extraction details

complete comparison tables extracted pending publication review

Searches

  • Protein Language Models Expose Viral Immune Mimicry primary paper benchmark results

Evidence locations

  • Table1; held-out protein-origin classification; separate error-analysis CV

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

16 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual overview of the published statistical assessment.
Individual claims
Protein Language Models Expose Viral Immune Mimicry

Original source ↗

Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1

Version: version of record
Retrieved: 2026-09-16T10:33:57.274Z

source checked

automated source review · 2026-09-16

Audit details

Inspected the cited primary abstract and descriptive computational-evaluation sections. Review covers the descriptive claims shown; no executable protocol was reconstructed. Original numerical records retain their prior transcription review.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 15250af2f75f70e2b6a3725d00bc7e276ed9d2d4a54f0ae9d6eaabf6be13e4a1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps
  • Curated sequence-origin labels
  • Keep sequence clusters separate
  • Assess held-out classifications
  • Report classification metrics
Individual claims
Protein Language Models Expose Viral Immune Mimicry

Original source ↗

Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1

Version: version of record
Retrieved: 2026-09-16T10:33:57.274Z

source checked

automated source review · 2026-09-16

Audit details

Inspected the cited primary abstract and descriptive computational-evaluation sections. Review covers the descriptive claims shown; no executable protocol was reconstructed. Original numerical records retain their prior transcription review.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 15250af2f75f70e2b6a3725d00bc7e276ed9d2d4a54f0ae9d6eaabf6be13e4a1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title
Conceptual assessment outline
Individual claims
Protein Language Models Expose Viral Immune Mimicry

Original source ↗

Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1

Version: version of record
Retrieved: 2026-09-16T10:33:57.274Z

source checked

automated source review · 2026-09-16

Audit details

Inspected the cited primary abstract and descriptive computational-evaluation sections. Review covers the descriptive claims shown; no executable protocol was reconstructed. Original numerical records retain their prior transcription review.

Field: attributes.profile.diagram.title

Source artifact SHA-256: 15250af2f75f70e2b6a3725d00bc7e276ed9d2d4a54f0ae9d6eaabf6be13e4a1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Entity type
Paper-specific evaluation task; this profile is a descriptive evidence summary.
Individual claims
Protein Language Models Expose Viral Immune Mimicry

Original source ↗

Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1

Version: version of record
Retrieved: 2026-09-16T10:33:57.274Z

source checked

automated source review · 2026-09-16

Audit details

Inspected the cited primary abstract and descriptive computational-evaluation sections. Review covers the descriptive claims shown; no executable protocol was reconstructed. Original numerical records retain their prior transcription review.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: 15250af2f75f70e2b6a3725d00bc7e276ed9d2d4a54f0ae9d6eaabf6be13e4a1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Datasets
Reviewed human and vertebrate-host viral protein records from Swiss-Prot/UniProtKB, with redundancy filtering described in Methods 2.1.
Individual claims
Protein Language Models Expose Viral Immune Mimicry

Original source ↗

Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1

Version: version of record
Retrieved: 2026-09-16T10:33:57.274Z

source checked

automated source review · 2026-09-16

Audit details

Inspected the cited primary abstract and descriptive computational-evaluation sections. Review covers the descriptive claims shown; no executable protocol was reconstructed. Original numerical records retain their prior transcription review.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: 15250af2f75f70e2b6a3725d00bc7e276ed9d2d4a54f0ae9d6eaabf6be13e4a1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Organisms
Human proteins and proteins from viruses with a known vertebrate host.
Individual claims
Protein Language Models Expose Viral Immune Mimicry

Original source ↗

Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1

Version: version of record
Retrieved: 2026-09-16T10:33:57.274Z

source checked

automated source review · 2026-09-16

Audit details

Inspected the cited primary abstract and descriptive computational-evaluation sections. Review covers the descriptive claims shown; no executable protocol was reconstructed. Original numerical records retain their prior transcription review.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: 15250af2f75f70e2b6a3725d00bc7e276ed9d2d4a54f0ae9d6eaabf6be13e4a1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Assays
Sequence-origin labels from curated database records. This classification endpoint is not an experimental immune-response assay.
Individual claims
Protein Language Models Expose Viral Immune Mimicry

Original source ↗

Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1

Version: version of record
Retrieved: 2026-09-16T10:33:57.274Z

source checked

automated source review · 2026-09-16

Audit details

Inspected the cited primary abstract and descriptive computational-evaluation sections. Review covers the descriptive claims shown; no executable protocol was reconstructed. Original numerical records retain their prior transcription review.

Field: attributes.profile.facts.3.value

Source artifact SHA-256: 15250af2f75f70e2b6a3725d00bc7e276ed9d2d4a54f0ae9d6eaabf6be13e4a1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Splits
Methods 2.1 assigns whole UniRef50 clusters to training or test sets. This profile records the split principle, not a verified membership manifest.
Individual claims
Protein Language Models Expose Viral Immune Mimicry

Original source ↗

Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1

Version: version of record
Retrieved: 2026-09-16T10:33:57.274Z

source checked

automated source review · 2026-09-16

Audit details

Inspected the cited primary abstract and descriptive computational-evaluation sections. Review covers the descriptive claims shown; no executable protocol was reconstructed. Original numerical records retain their prior transcription review.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: 15250af2f75f70e2b6a3725d00bc7e276ed9d2d4a54f0ae9d6eaabf6be13e4a1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Allowed inputs
Protein sequence representations paired with the study’s human/viral labels.
Individual claims
Protein Language Models Expose Viral Immune Mimicry

Original source ↗

Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1

Version: version of record
Retrieved: 2026-09-16T10:33:57.274Z

source checked

automated source review · 2026-09-16

Audit details

Inspected the cited primary abstract and descriptive computational-evaluation sections. Review covers the descriptive claims shown; no executable protocol was reconstructed. Original numerical records retain their prior transcription review.

Field: attributes.profile.facts.5.value

Source artifact SHA-256: 15250af2f75f70e2b6a3725d00bc7e276ed9d2d4a54f0ae9d6eaabf6be13e4a1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Adaptation
Pretrained protein representations are evaluated through a study-specific classifier; a backbone name alone does not identify the full fitted pipeline.
Individual claims
Protein Language Models Expose Viral Immune Mimicry

Original source ↗

Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1

Version: version of record
Retrieved: 2026-09-16T10:33:57.274Z

source checked

automated source review · 2026-09-16

Audit details

Inspected the cited primary abstract and descriptive computational-evaluation sections. Review covers the descriptive claims shown; no executable protocol was reconstructed. Original numerical records retain their prior transcription review.

Field: attributes.profile.facts.6.value

Source artifact SHA-256: 15250af2f75f70e2b6a3725d00bc7e276ed9d2d4a54f0ae9d6eaabf6be13e4a1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: needs review

2 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: reported-task-53506fe386e4a1

areas
proteins-complexes
tasks
human-versus-viral protein classification
entity level
task
version
Not reported
task
human-versus-viral protein classification
scope note
Paper-specific evaluation task; protocol completeness requires further extraction.
comparison panels
id: part2-viral-immune-mimicry-2025-viruses-17-01199-t001-3a83e9bebb; title: Held-out human-versus-virus protein classification · Table 1; protocol id: paper-protocol-1352834b9391c1dacb; dataset id: reported-dataset-43f24c4dfb7351; metric: AUROC; unit: percent; direction: higher; result ids: paper-result-e2941183edbcf0f806; paper-result-088aaf71a44cdc5709; paper-result-a19e0533c00420b50c; paper-result-13ca62957ffa02eef1; paper-result-a541b9f3a77eec62e2; lit-b4-009; paper-result-b740d08b806206779f; paper-result-6536626a03a3ef62f0; source ids: part2-viral-immune-mimicry-2025; source locator: Table 1: AUC (%), Held-out human-versus-virus protein classification; context: Compare protein origin classifiers; T 5 embeddings plus linear/tree models vs ESM 2 fine-tuning and simple controls. UniRef90 deduplication; proteins longer than1,600 residues excluded; UniRef50 clusters assigned80% training and20% test with no cluster shared. Separate four-fold error-analysis experiment not assigned toTable1.; caveats: Origin classification does not establish immune mimicry.; Table footnote says all values are percent but log-loss values are dimensionless; log-loss units quarantined.; Narrative n-gram AUC91.9 differs from Table1 value91.5; preserve table value.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17; id: part2-viral-immune-mimicry-2025-viruses-17-01199-t001-668a472b03; title: Held-out human-versus-virus protein classification · Table 1; protocol id: paper-protocol-1352834b9391c1dacb; dataset id: reported-dataset-43f24c4dfb7351; metric: Accur.; unit: percent; direction: higher; result ids: paper-result-91e7ac901444849ecd; paper-result-ca12e3e749c612d6d8; paper-result-ee7f4e8a864f523eaf; paper-result-544d2c1cd2bd46af9b; paper-result-32da435487ad9e4e7b; paper-result-7ffc4775000c61b1e4; paper-result-87b3b45c365edc42f3; paper-result-439ed50c8779070823; source ids: part2-viral-immune-mimicry-2025; source locator: Table 1: Accur., Held-out human-versus-virus protein classification; context: Compare protein origin classifiers; T 5 embeddings plus linear/tree models vs ESM 2 fine-tuning and simple controls. UniRef90 deduplication; proteins longer than1,600 residues excluded; UniRef50 clusters assigned80% training and20% test with no cluster shared. Separate four-fold error-analysis experiment not assigned toTable1.; caveats: Origin classification does not establish immune mimicry.; Table footnote says all values are percent but log-loss values are dimensionless; log-loss units quarantined.; Narrative n-gram AUC91.9 differs from Table1 value91.5; preserve table value.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17; id: part2-viral-immune-mimicry-2025-viruses-17-01199-t001-e488608662; title: Held-out human-versus-virus protein classification · Table 1; protocol id: paper-protocol-1352834b9391c1dacb; dataset id: reported-dataset-43f24c4dfb7351; metric: Prec.; unit: percent; direction: higher; result ids: paper-result-a950e05fb8768ed4e8; paper-result-0c1a3836d475c5a506; paper-result-00f7697f28c706f26b; paper-result-1f9cd7bc0d3be4259f; paper-result-24e80b6223e7bd1109; paper-result-0c5e3578112cc7247a; paper-result-c583dfa48d139dadc8; paper-result-f67a4402bb1e2c96ba; source ids: part2-viral-immune-mimicry-2025; source locator: Table 1: Prec., Held-out human-versus-virus protein classification; context: Compare protein origin classifiers; T 5 embeddings plus linear/tree models vs ESM 2 fine-tuning and simple controls. UniRef90 deduplication; proteins longer than1,600 residues excluded; UniRef50 clusters assigned80% training and20% test with no cluster shared. Separate four-fold error-analysis experiment not assigned toTable1.; caveats: Origin classification does not establish immune mimicry.; Table footnote says all values are percent but log-loss values are dimensionless; log-loss units quarantined.; Narrative n-gram AUC91.9 differs from Table1 value91.5; preserve table value.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17; id: part2-viral-immune-mimicry-2025-viruses-17-01199-t001-a4943fb2ad; title: Held-out human-versus-virus protein classification · Table 1; protocol id: paper-protocol-1352834b9391c1dacb; dataset id: reported-dataset-43f24c4dfb7351; metric: Recall; unit: percent; direction: higher; result ids: paper-result-b36e1a943ebc5dfb9a; paper-result-ccac699408a1229570; paper-result-4f2c34cb0e09f0167a; paper-result-4014edf9066fc226f1; paper-result-9d9d31080b47613e46; paper-result-a42799c550a1ae61e7; paper-result-67bb8fdb6692f4336f; paper-result-2aafdf0e911ae9cff5; source ids: part2-viral-immune-mimicry-2025; source locator: Table 1: Recall, Held-out human-versus-virus protein classification; context: Compare protein origin classifiers; T 5 embeddings plus linear/tree models vs ESM 2 fine-tuning and simple controls. UniRef90 deduplication; proteins longer than1,600 residues excluded; UniRef50 clusters assigned80% training and20% test with no cluster shared. Separate four-fold error-analysis experiment not assigned toTable1.; caveats: Origin classification does not establish immune mimicry.; Table footnote says all values are percent but log-loss values are dimensionless; log-loss units quarantined.; Narrative n-gram AUC91.9 differs from Table1 value91.5; preserve table value.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17
benchmark research
review date: 2026-09-17; status: complete_comparison_tables_extracted_pending_publication_review; primary sources: part2-viral-immune-mimicry-2025; inspected locators: Table1; held-out protein-origin classification; separate error-analysis CV; searched queries: Protein Language Models Expose Viral Immune Mimicry primary paper benchmark results; gaps: Independent batch review before import; preserve existing observation identities.; claim scope: Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.
historical missing metadata
protocol version: not_reported_in_legacy_extract; split: not_reported_in_legacy_extract
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
legacy kinds
benchmark
entity classification
review date: 2026-09-17; rationale: This source-scoped record identifies the biological prediction task and holds its paper context. Preserve the existing task identity; exact split, model adaptation and scoring remain in linked evaluations or separate protocol records.; source ids: viral-immune-mimicry-2025; source locator: Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1; ambiguities: A paper- or suite-specific task may constrain some inputs or metrics; that alone does not make it interchangeable with a complete versioned protocol. No protocol equivalence is inferred.; Some legacy profile Entity type facts use the generic phrase computational evaluation protocol. That boilerplate is not sufficient to establish a single fixed protocol identity or to merge this task with another protocol record.
Related records

Suggest a correction