rewire.itbenchmarks
Benchmark

FLIP2

FLIP2 evaluates protein fitness prediction under deliberately shifted training and test distributions.

Sourcesflip2 official source · FLIP2 official website: Overview; benchmark features; datasets and split categories

149 evaluations · 293 results

Overview

Datasets

Datasets span enzymatic activity, protein interactions and other measured protein properties.

Sourcesflip2 official source · FLIP2 official website: Overview; benchmark features; datasets and split categories

Metrics

Spearman rank correlation is the primary metric, with NDCG additionally measuring prioritization of high-fitness variants.

Sourcesflip2 primary benchmark evidence · Sections 2–4; Table 1; supplementary result tables

Allowed inputs

Protein sequence representations and dataset-specific labels.

Sourcesflip2 official source · FLIP2 official website: Overview; benchmark features; datasets and split categories
Evaluation procedure diagram
How it worksEvaluation procedure
Evaluation procedure1. Allowed inputs: Protein sequence representations and dataset-specific labels.. Then: 2. Splits: Split categories separate mutation counts, positions, mutation identities, fitness ranges or reference proteins.. Then: 3. Metrics: Spearman rank correlation is the primary metric, with NDCG additionally measuring prioritization of high-fitness variants.Evaluation procedure1. Allowed inputs: Protein sequence representations and dataset-specific labels.. Then: 2. Splits: Split categories separate mutation counts, positions, mutation identities, fitness ranges or reference proteins.. Then: 3. Metrics: Spearman rank correlation is the primary metric, with NDCG additionally measuring prioritization of high-fitness variants.Evaluation procedure1. Allowed inputs: Protein sequence representations and dataset-specific labels.. Then: 2. Splits: Split categories separate mutation counts, positions, mutation identities, fitness ranges or reference proteins.. Then: 3. Metrics: Spearman rank correlation is the primary metric, with NDCG additionally measuring prioritization of high-fitness variants.

Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.

Sources (2)flip2 official source; flip2 primary benchmark evidence · FLIP2 official website: Overview; benchmark features; datasets and split categories; Sections 2–4; Table 1; supplementary result tables

Source reviewed · Automated source review, 2026-09-16. All specifications and missing details

Results

Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.

FLIP2 NucB two-to-many; held-out test set · NucB: spearman

spearman (correlation) · Higher values are better.

FLIP2 NucB two-to-many; held-out test set · NucB · NucB

Evidence origin: Author-reported evaluation.

FLIP2: Expanding Protein Fitness Landscape Benchmarks for Real-World Machine Learning Applications · Table A6; row Ridge (one-hot); column spearman through Table A6; row ESMC-300M naive supervised; column spearman
  • Missing source cells and quarantined conflicts are recorded in acquisition and audit tables. Per-result scoring denominators may be unreported.
Comparison details and limitations

Complete selected source table is retained across source-order panels. These point estimates do not establish statistical significance or a universal ranking.

  • Source-specific evaluation. No equivalence to other releases, protocols or model families is inferred.
  • Exact source-defined evaluation scope; reported scores are not rewire reproductions.

Automated source review: 2026-09-19. Numerical source review does not establish independent reproduction.

Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.

Showing 9 of 9 matching rows.

Methods and evaluation design

Procedure, tasks and evaluated configurations

How it works

Evaluation methodology

FLIP2 extends protein fitness testing to enzyme activity, spectral properties, hydrophobic-core effects and protein interactions. Its partitions hold out mutation positions, mutation numbers, fitness ranges or parent sequences. Simple one-hot ridge models and pretrained versus randomly initialized networks make the baseline comparison interpretable.

Sourcesflip2 primary benchmark evidence · Sections 2–4; Table 1; supplementary result tables

Evaluation design

Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.

These source-backed links do not make different protocols or scores interchangeable.

Baseline coverage

Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.

2 of 34 active baseline roles have published Rewire measurements in this release. Measurements on a selected protocol do not establish coverage of an entire suite.

Baseline status by linked protocol

Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums

Coverage is derived from release 2026-09-29-06401fd5b220. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.

Run this benchmark

Choose a concrete protocol before running an evaluation. Its inputs, split and scoring rules determine which results can be compared.

FLIP2 archived fitness splits: score supplied predictions

Rank protein variants by measured fitness on one archived train/validation/test split. Spearman correlation and full-ranking NDCG using the upstream minimum-shifted test targets and tie handling.

Recompute metrics from supplied predictions. This recipe does not establish reproduction of a particular published score.

Dataset access
Download the selected CSV from pinned Zenodo18433203 v3. Preparation checks exact decompressed CSV SHA-256. The packaged manifest lists all16 splits; this example selects Rhomax by-wild-type.
Model and weights
No weights needed for rescoring or the supplied toy/composition control. Private-model weights remain local; their access requirements depend on the model.
Licences
Runner MIT. FLIP2 archive CC-BY-4.0; retain per-dataset attribution. Original Amylase MIT and TrpB CC0-1.0 according to source attribution. Upstream evaluator AFL-3.0; formulas independently implemented.
Software
Python3.11, pinned rewirebench0.4 environment; sequence extra for parquet and HDF5. Podman or Apptainer is optional.
Hardware
CPU scoring and small controls; memory/accelerator needs for real private-model inference are model-dependent and unreported here.
Required inputs and expected outputs

Inputs

  • Prepared protocol inputs with evaluator-owned labels and opaque IDs.
  • Keyed finite scalar scores; missing predictions must be explicitly permitted and remain in coverage.

Outputs

  • Local report.json with metrics, source/version identity and coverage.
  • Local predictions.json and unscored.json; nothing submitted automatically.

Choose one way to run this recipe. These instruction formats are alternatives.

Execution steps

  1. 1. Score supplied keyed predictions (Python)

    Source reviewed; these instructions have not been executed by rewire.

    import rewirebench
    prepared = rewirebench.prepare(
        'flip2-fitness-v1', source='flip2-data', output='prepared-flip2', dataset='flip2-rhomax-by-wild-type', limit=8
    )
    report = rewirebench.evaluate(
        prepared, 'predictions.json', output='results-flip2-rescore',
        model={'name': 'My private model', 'training_overlap': 'Unreported'},
    )
    rewirebench 0.4 library interface; rewirebench container and HPC instructions; FLIP2 archived fitness splits runner guide; FLIP2 archived fitness splits protocol implementation; FLIP2 archived fitness splits source manifest; FLIP2 archived fitness splits automated validation receipt; FLIP2 CPU composition learning control · docs/flip2.md and protocol implementation; linked validation receipt concerns its stated tests, not execution of this exact snippet

Use your own model

Run your model locally and return predictions keyed by the input IDs. The evaluator supplies biological inputs without test labels and owns scoring. This interface is not a sandbox for model code.

Pass your existing prediction function into this adapter. Its output direction must match the selected protocol.

class MyModelAdapter:
    def __init__(self, score):
        self.score = score

    def predict(self, inputs):
        return {row["id"]: float(self.score(row)) for row in inputs}

# adapter = MyModelAdapter(your_prediction_function)
# report = rewirebench.run(prepared, adapter, output="runs/my-model")

Alternatively, generate a keyed prediction file in your existing model environment and use the score-only recipe. Your model code and weights do not need to be shared.

rewirebench 0.4 library interface; rewirebench container and HPC instructions; FLIP2 archived fitness splits runner guide; FLIP2 archived fitness splits protocol implementation; FLIP2 archived fitness splits source manifest; FLIP2 archived fitness splits automated validation receipt; FLIP2 CPU composition learning control · docs/flip2.md; source manifest; protocol score function; automated receipt scope
Scope and limitations
  • A run is one of16 archived splits across7 datasets; no suite-wide score is computed.
  • Archived Amylase one_to_many and NucB split counts conflict with manuscript Table1. IRED printed total also differs. The runner preserves archive assignments and records each conflict; it does not claim paper-cohort reproduction.
  • PDZ3 inputs preserve protein:peptide notation including empty partners. An adapter must handle this representation explicitly.
  • Frozen-embedding alpha10 ridge is a Rewire extension with train-only target scaling, not the published one-hot ridge baseline.
  • Native CPU smoke checks covered all16 archived splits with8 rows per train/validation/test group; pinned upstream metric expressions agreed within1e-12. No full model runs or published-score reproduction.
  • These snippets are source-reviewed instructions. The linked automated receipt supports only the tests it records; it does not certify the exact snippet or published-model reproduction.
  • Smoke examples and partial predictions retain original denominators and cannot establish a complete suite. Prepared execution is offline; exports and submissions are explicit.

Contribute a result for review. The library can submit an exported evaluation for private review when intake is open. Check the contribution page for access and sign-in.

Original repository instructions

Run this benchmark

Official FLIP2 page links CSV/FASTA datasets and the shared FLIP repository. Choose a FLIP2 dataset and split explicitly; the repository also contains earlier FLIP material, so its root documentation alone does not identify an exact FLIP2 run configuration.

A maintained rewire runner has not been verified for this benchmark. Check data access, weights, licences, dependencies and hardware in the linked official documentation; requirements have not been fully extracted.

flip2 official run documentation · Official FLIP2 site: Data formats, dataset download and GitHub repository sections
Strengths, limitations and unresolved questions

Strengths and limitations

Strengths supported by sources

  • Explicit task diversity avoids treating fitness as a single interchangeable measurement.
    Sourcesflip2 official source · FLIP2 official website: Overview; benchmark features; datasets and split categories

Limitations and conditions

  • The same number of training examples does not make random and engineering-oriented splits equivalent. Per-landscape ranking and top-fitness retrieval should be retained alongside any aggregate.
    Sourcesflip2 primary benchmark evidence · Sections 2–4; Table 1; supplementary result tables
Profile review details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Stable record: discovery-benchmark-flip2

Specifications

Inputs, training, access and other details

Explanatory profile: source reviewed · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetsDatasets span enzymatic activity, protein interactions and other measured protein properties.
Sourcesflip2 official source · FLIP2 official website: Overview; benchmark features; datasets and split categories
SplitsSplit categories separate mutation counts, positions, mutation identities, fitness ranges or reference proteins.
Sourcesflip2 official source · FLIP2 official website: Overview; benchmark features; datasets and split categories
MetricsSpearman rank correlation is the primary metric, with NDCG additionally measuring prioritization of high-fitness variants.
Sourcesflip2 primary benchmark evidence · Sections 2–4; Table 1; supplementary result tables
BaselinesZero-shot protein models, ridge regression and fine-tuned models.
Sourcesflip2 official source · FLIP2 official website: Overview; benchmark features; datasets and split categories
Leakage controlsDistribution-shift splits are explicit; exact homology/pretraining overlap depends on the dataset.
Sourcesflip2 official source · FLIP2 official website: Overview; benchmark features; datasets and split categories
UncertaintyFine-tuned language models are run five times with different random seeds and validation-based early stopping; reported metrics are averaged. The zero-shot and ridge evaluations follow separate procedures.
Sourcesflip2 primary benchmark evidence · Sections 2–4; Table 1; supplementary result tables
Entity typeProtein fitness benchmark collection.
Sourcesflip2 official source · FLIP2 official website: Overview; benchmark features; datasets and split categories
OrganismsMixed natural and designed protein landscapes: amylase, imine reductase, NucB, TrpB, hydrophobic-core variants, microbial rhodopsins and a PDZ3–peptide interaction system. These are not a single-species benchmark.
Sourcesflip2 primary benchmark evidence · Sections 2–4; Table 1; supplementary result tables
AssaysEnzymatic activity, protein interactions and other measured properties.
Sourcesflip2 official source · FLIP2 official website: Overview; benchmark features; datasets and split categories
Allowed inputsProtein sequence representations and dataset-specific labels.
Sourcesflip2 official source · FLIP2 official website: Overview; benchmark features; datasets and split categories
AdaptationSupervised prediction under task-specific data splits.
Sourcesflip2 official source · FLIP2 official website: Overview; benchmark features; datasets and split categories

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Dated primary-source discovery and protocol/table screening. Source checking does not mean experimental reproduction. Only separately extracted and independently reviewed numeric batches are publishable.

Paper or primary resourceVersionReference
FLIP2: Expanding Protein Fitness Landscape Benchmarks for Real-World Machine Learning ApplicationsPrimary full-text snapshot retrieved 2026-09-17; exact bytes pinned by SHA-256Read source
Historical gaps recorded on 2026-09-17

The catalogue now holds 293 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.

  • complete comparable numeric result batch: Specific candidate tables and protocol boundaries are documented; no graph values or incomplete winner-only selection are converted into publishable rows.
Search and extraction details

primary protocol reviewed

Searches

  • FLIP2 protein benchmark paper flip.protein.properties

Evidence locations

  • Table 1; Appendix Tables A1–A16; A17 is best-method summary only

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

140 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Individual claims
flip2 official source

Original source ↗

FLIP2 official website: Overview; benchmark features; datasets and split categories; Sections 2–4; Table 1; supplementary result tables

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: Retrieved website snapshot sha256:cd991ae6e76a5e84ea5449f91c4ed86ba4f57942682dc3c50d872c366bbbd4b7
Retrieved: 2026-09-16T10:31:54.381289+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: cd991ae6e76a5e84ea5449f91c4ed86ba4f57942682dc3c50d872c366bbbd4b7

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram caption
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Individual claims
flip2 primary benchmark evidence

Original source ↗

FLIP2 official website: Overview; benchmark features; datasets and split categories; Sections 2–4; Table 1; supplementary result tables

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: ICML2026 manuscript
Retrieved: 2026-09-16T21:05:54.225877+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: d0e61ca27863c023adddcc225a3e4af1718f2e3fdfe3642bef89c24c93621a17

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps
  • Allowed inputs: Protein sequence representations and dataset-specific labels.
  • Splits: Split categories separate mutation counts, positions, mutation identities, fitness ranges or reference proteins.
  • Metrics: Spearman rank correlation is the primary metric, with NDCG additionally measuring prioritization of high-fitness variants.
Individual claims
flip2 official source

Original source ↗

FLIP2 official website: Overview; benchmark features; datasets and split categories; Sections 2–4; Table 1; supplementary result tables

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: Retrieved website snapshot sha256:cd991ae6e76a5e84ea5449f91c4ed86ba4f57942682dc3c50d872c366bbbd4b7
Retrieved: 2026-09-16T10:31:54.381289+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: cd991ae6e76a5e84ea5449f91c4ed86ba4f57942682dc3c50d872c366bbbd4b7

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps
  • Allowed inputs: Protein sequence representations and dataset-specific labels.
  • Splits: Split categories separate mutation counts, positions, mutation identities, fitness ranges or reference proteins.
  • Metrics: Spearman rank correlation is the primary metric, with NDCG additionally measuring prioritization of high-fitness variants.
Individual claims
flip2 primary benchmark evidence

Original source ↗

FLIP2 official website: Overview; benchmark features; datasets and split categories; Sections 2–4; Table 1; supplementary result tables

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: ICML2026 manuscript
Retrieved: 2026-09-16T21:05:54.225877+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: d0e61ca27863c023adddcc225a3e4af1718f2e3fdfe3642bef89c24c93621a17

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title
Evaluation procedure
Individual claims
flip2 official source

Original source ↗

FLIP2 official website: Overview; benchmark features; datasets and split categories; Sections 2–4; Table 1; supplementary result tables

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: Retrieved website snapshot sha256:cd991ae6e76a5e84ea5449f91c4ed86ba4f57942682dc3c50d872c366bbbd4b7
Retrieved: 2026-09-16T10:31:54.381289+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: cd991ae6e76a5e84ea5449f91c4ed86ba4f57942682dc3c50d872c366bbbd4b7

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title
Evaluation procedure
Individual claims
flip2 primary benchmark evidence

Original source ↗

FLIP2 official website: Overview; benchmark features; datasets and split categories; Sections 2–4; Table 1; supplementary result tables

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: ICML2026 manuscript
Retrieved: 2026-09-16T21:05:54.225877+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: d0e61ca27863c023adddcc225a3e4af1718f2e3fdfe3642bef89c24c93621a17

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Datasets
Datasets span enzymatic activity, protein interactions and other measured protein properties.
Individual claims
flip2 official source

Original source ↗

FLIP2 official website: Overview; benchmark features; datasets and split categories

Version: Retrieved website snapshot sha256:cd991ae6e76a5e84ea5449f91c4ed86ba4f57942682dc3c50d872c366bbbd4b7
Retrieved: 2026-09-16T10:31:54.381289+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: cd991ae6e76a5e84ea5449f91c4ed86ba4f57942682dc3c50d872c366bbbd4b7

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Splits
Split categories separate mutation counts, positions, mutation identities, fitness ranges or reference proteins.
Individual claims
flip2 official source

Original source ↗

FLIP2 official website: Overview; benchmark features; datasets and split categories

Version: Retrieved website snapshot sha256:cd991ae6e76a5e84ea5449f91c4ed86ba4f57942682dc3c50d872c366bbbd4b7
Retrieved: 2026-09-16T10:31:54.381289+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: cd991ae6e76a5e84ea5449f91c4ed86ba4f57942682dc3c50d872c366bbbd4b7

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Adaptation
Supervised prediction under task-specific data splits.
Individual claims
flip2 official source

Original source ↗

FLIP2 official website: Overview; benchmark features; datasets and split categories

Version: Retrieved website snapshot sha256:cd991ae6e76a5e84ea5449f91c4ed86ba4f57942682dc3c50d872c366bbbd4b7
Retrieved: 2026-09-16T10:31:54.381289+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: cd991ae6e76a5e84ea5449f91c4ed86ba4f57942682dc3c50d872c366bbbd4b7

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Metrics
Spearman rank correlation is the primary metric, with NDCG additionally measuring prioritization of high-fitness variants.
Individual claims
flip2 primary benchmark evidence

Original source ↗

Sections 2–4; Table 1; supplementary result tables

Version: ICML2026 manuscript
Retrieved: 2026-09-16T21:05:54.225877+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: d0e61ca27863c023adddcc225a3e4af1718f2e3fdfe3642bef89c24c93621a17

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: discovered

12 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: discovery-benchmark-flip2

areas
protein-function
entity level
suite
scope note
Specialist molecular or omics evaluation; protocol details require review before numerical comparison.
task
Expanded protein fitness landscapes
version
Not reported
benchmark research
review date: 2026-09-17; status: primary_protocol_reviewed; primary sources: evidence-expansion-flip2-d0e61ca2; inspected locators: Table 1; Appendix Tables A1–A16; A17 is best-method summary only; searched queries: FLIP2 protein benchmark paper flip.protein.properties; gaps: complete comparable numeric result batch: Specific candidate tables and protocol boundaries are documented; no graph values or incomplete winner-only selection are converted into publishable rows.; claim scope: Dated primary-source discovery and protocol/table screening. Source checking does not mean experimental reproduction. Only separately extracted and independently reviewed numeric batches are publishable.
historical missing metadata
dataset release: unextracted; metric implementation: unextracted; split manifest: unextracted; version: unextracted
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
entity classification
review date: 2026-09-17; rationale: The cited profile describes a collection of evaluation tasks or protocols; retain it as the top-level benchmark suite. Its datasets and individual protocols remain separate records.; source ids: evidence-benchmark-flip2-snapshot; source locator: FLIP2 official website: Overview; benchmark features; datasets and split categories; ambiguities: None recorded
run documentation
record id: discovery-benchmark-flip2; source ids: run-doc-flip2-official-20260917; status: official_documentation_linked; summary: Official FLIP2 page links CSV/FASTA datasets and the shared FLIP repository. Choose a FLIP2 dataset and split explicitly; the repository also contains earlier FLIP material, so its root documentation alone does not identify an exact FLIP2 run configuration.; source locator: Official FLIP2 site: Data formats, dataset download and GitHub repository sections
run recipes
id: flip2-runner-rescore-v1; protocol id: flip2-fitness-v1; version: fcccbbcdbe3d5cd64a1a312d536615273320f7b4; title: FLIP2 archived fitness splits: score supplied predictions; purpose: rescore_predictions; summary: Rank protein variants by measured fitness on one archived train/validation/test split. Spearman correlation and full-ranking NDCG using the upstream minimum-shifted test targets and tie handling.; inputs: Prepared protocol inputs with evaluator-owned labels and opaque IDs.; Keyed finite scalar scores; missing predictions must be explicitly permitted and remain in coverage.; outputs: Local report.json with metrics, source/version identity and coverage.; Local predictions.json and unscored.json; nothing submitted automatically.; requirements: data: Download the selected CSV from pinned Zenodo18433203 v3. Preparation checks exact decompressed CSV SHA-256. The packaged manifest lists all16 splits; this example selects Rhomax by-wild-type.; weights: No weights needed for rescoring or the supplied toy/composition control. Private-model weights remain local; their access requirements depend on the model.; licence: Runner MIT. FLIP2 archive CC-BY-4.0; retain per-dataset attribution. Original Amylase MIT and TrpB CC0-1.0 according to source attribution. Upstream evaluator AFL-3.0; formulas independently implemented.; software: Python3.11, pinned rewirebench0.4 environment; sequence extra for parquet and HDF5. Podman or Apptainer is optional.; hardware: CPU scoring and small controls; memory/accelerator needs for real private-model inference are model-dependent and unreported here.; instructions: runtime: python; title: Score supplied keyed predictions; code: import rewirebench prepared = rewirebench.prepare( 'flip2-fitness-v1', source='flip2-data', output='prepared-flip2', dataset='flip2-rhomax-by-wild-type', limit=8 ) report = rewirebench.evaluate( prepared, 'predictions.json', output='results-flip2-rescore', model={'name': 'My private model', 'training_overlap': 'Unreported'}, ); status: source_reviewed_not_executed; source ids: runner-04-docs-sdk-md; runner-04-docs-hpc-md; runner-04-docs-flip2-md; runner-04-packages-rewirebench-src-rewirebench-protocols-flip2-py; runner-04-packages-rewirebench-src-rewirebench-resources-flip2-sources-json; runner-04-packages-rewirebench-src-rewirebench-resources-flip2-validation-2026-09-20-json; runner-04-packages-rewirebench-src-rewirebench-resources-flip2-example-adapter-py; source locator: docs/flip2.md and protocol implementation; linked validation receipt concerns its stated tests, not execution of this exact snippet; runtime: command_line; title: Prepare inputs and score supplied predictions; code: git clone https://github.com/rewire-bio/rewire-benchmarks.git cd rewire-benchmarks git checkout --detach fcccbbcdbe3d5cd64a1a312d536615273320f7b4 uv sync --locked --package rewirebench --extra sequence --python 3.11 uv run --package rewirebench rewirebench prepare flip2-fitness-v1 \ --source flip2-data --output prepared-flip2 \ --options '{"dataset":"flip2-rhomax-by-wild-type","limit":8}' uv run --package rewirebench rewirebench evaluate \ --prepared prepared-flip2 --predictions predictions.json \ --output results-flip2-rescore --model-name 'My private model'; status: source_reviewed_not_executed; source ids: runner-04-docs-sdk-md; runner-04-docs-hpc-md; runner-04-docs-flip2-md; runner-04-packages-rewirebench-src-rewirebench-protocols-flip2-py; runner-04-packages-rewirebench-src-rewirebench-resources-flip2-sources-json; runner-04-packages-rewirebench-src-rewirebench-resources-flip2-validation-2026-09-20-json; runner-04-packages-rewirebench-src-rewirebench-resources-flip2-example-adapter-py; source locator: docs/flip2.md and protocol implementation; linked validation receipt concerns its stated tests, not execution of this exact snippet; runtime: podman; title: Score prepared predictions offline; code: # Obtain/build the pinned core OCI image following the linked HPC guide. mkdir -p runs podman run --rm --userns=keep-id --network=none \ -v "$PWD/prepared-flip2:/prepared:ro" \ -v "$PWD/predictions:/predictions:ro" -v "$PWD/runs:/outputs:rw" \ localhost/rewirebench:core evaluate --prepared /prepared \ --predictions /predictions/predictions.json --output /outputs/flip2-rescore; status: source_reviewed_not_executed; source ids: runner-04-docs-sdk-md; runner-04-docs-hpc-md; runner-04-docs-flip2-md; runner-04-packages-rewirebench-src-rewirebench-protocols-flip2-py; runner-04-packages-rewirebench-src-rewirebench-resources-flip2-sources-json; runner-04-packages-rewirebench-src-rewirebench-resources-flip2-validation-2026-09-20-json; runner-04-packages-rewirebench-src-rewirebench-resources-flip2-example-adapter-py; source locator: docs/flip2.md and protocol implementation; linked validation receipt concerns its stated tests, not execution of this exact snippet; runtime: apptainer; title: Score prepared predictions offline; code: # Obtain/build the pinned SIF following the linked HPC guide. mkdir -p runs apptainer run --cleanenv --containall \ --bind "$PWD/prepared-flip2:/prepared:ro" \ --bind "$PWD/predictions:/predictions:ro" --bind "$PWD/runs:/outputs:rw" \ rewirebench-core.sif evaluate --prepared /prepared \ --predictions /predictions/predictions.json --output /outputs/flip2-rescore; status: source_reviewed_not_executed; source ids: runner-04-docs-sdk-md; runner-04-docs-hpc-md; runner-04-docs-flip2-md; runner-04-packages-rewirebench-src-rewirebench-protocols-flip2-py; runner-04-packages-rewirebench-src-rewirebench-resources-flip2-sources-json; runner-04-packages-rewirebench-src-rewirebench-resources-flip2-validation-2026-09-20-json; runner-04-packages-rewirebench-src-rewirebench-resources-flip2-example-adapter-py; source locator: docs/flip2.md and protocol implementation; linked validation receipt concerns its stated tests, not execution of this exact snippet; runtime: slurm; title: Schedule prepared scoring with site-specific resources; code: #!/bin/bash set -euo pipefail # Set these for your cluster; no performance or resource estimate is implied. : "${REWIRE_ACCOUNT:?Set your Slurm account}" : "${REWIRE_PARTITION:?Set your Slurm partition}" : "${REWIRE_CPUS:?Set the requested CPU count}" : "${REWIRE_MEMORY:?Set the requested memory}" : "${REWIRE_WALLTIME:?Set the requested time limit}" : "${REWIRE_JOB_ROOT:?Set a shared absolute directory with prepared data and SIF}" export REWIRE_JOB_ROOT sbatch --account="$REWIRE_ACCOUNT" --partition="$REWIRE_PARTITION" \ --cpus-per-task="$REWIRE_CPUS" --mem="$REWIRE_MEMORY" \ --time="$REWIRE_WALLTIME" --export=ALL <<'REWIRE_JOB' #!/bin/bash set -euo pipefail cd "$REWIRE_JOB_ROOT" mkdir -p runs apptainer run --cleanenv --containall \ --bind "$PWD/prepared-flip2:/prepared:ro" \ --bind "$PWD/predictions:/predictions:ro" \ --bind "$PWD/runs:/outputs:rw" \ rewirebench-core.sif evaluate \ --prepared /prepared --predictions /predictions/predictions.json \ --output /outputs/flip2-rescore REWIRE_JOB; status: source_reviewed_not_executed; source ids: runner-04-docs-sdk-md; runner-04-docs-hpc-md; runner-04-docs-flip2-md; runner-04-packages-rewirebench-src-rewirebench-protocols-flip2-py; runner-04-packages-rewirebench-src-rewirebench-resources-flip2-sources-json; runner-04-packages-rewirebench-src-rewirebench-resources-flip2-validation-2026-09-20-json; runner-04-packages-rewirebench-src-rewirebench-resources-flip2-example-adapter-py; source locator: docs/hpc.md: Slurm template, caller-supplied account/partition/resources; protocol guide: evaluate prepared predictions; limitations: A run is one of16 archived splits across7 datasets; no suite-wide score is computed.; Archived Amylase one_to_many and NucB split counts conflict with manuscript Table1. IRED printed total also differs. The runner preserves archive assignments and records each conflict; it does not claim paper-cohort reproduction.; PDZ3 inputs preserve protein:peptide notation including empty partners. An adapter must handle this representation explicitly.; Frozen-embedding alpha10 ridge is a Rewire extension with train-only target scaling, not the published one-hot ridge baseline.; Native CPU smoke checks covered all16 archived splits with8 rows per train/validation/test group; pinned upstream metric expressions agreed within1e-12. No full model runs or published-score reproduction.; These snippets are source-reviewed instructions. The linked automated receipt supports only the tests it records; it does not certify the exact snippet or published-model reproduction.; Smoke examples and partial predictions retain original denominators and cannot establish a complete suite. Prepared execution is offline; exports and submissions are explicit.; source ids: runner-04-docs-sdk-md; runner-04-docs-hpc-md; runner-04-docs-flip2-md; runner-04-packages-rewirebench-src-rewirebench-protocols-flip2-py; runner-04-packages-rewirebench-src-rewirebench-resources-flip2-sources-json; runner-04-packages-rewirebench-src-rewirebench-resources-flip2-validation-2026-09-20-json; runner-04-packages-rewirebench-src-rewirebench-resources-flip2-example-adapter-py; source locator: docs/flip2.md; source manifest; protocol score function; automated receipt scope; id: flip2-runner-control-v1; protocol id: flip2-fitness-v1; version: fcccbbcdbe3d5cd64a1a312d536615273320f7b4; title: FLIP2 archived fitness splits: local model or software control; purpose: generate_and_evaluate; summary: Rank protein variants by measured fitness on one archived train/validation/test split. Spearman correlation and full-ranking NDCG using the upstream minimum-shifted test targets and tie handling.; inputs: Same prepared biological inputs; no evaluation labels are exposed to adapters.; A private local adapter or the documented lightweight software control.; outputs: Local report.json with metrics, source/version identity and coverage.; Local predictions.json and unscored.json; nothing submitted automatically.; requirements: data: Download the selected CSV from pinned Zenodo18433203 v3. Preparation checks exact decompressed CSV SHA-256. The packaged manifest lists all16 splits; this example selects Rhomax by-wild-type.; weights: No weights needed for rescoring or the supplied toy/composition control. Private-model weights remain local; their access requirements depend on the model.; licence: Runner MIT. FLIP2 archive CC-BY-4.0; retain per-dataset attribution. Original Amylase MIT and TrpB CC0-1.0 according to source attribution. Upstream evaluator AFL-3.0; formulas independently implemented.; software: Python3.11, pinned rewirebench0.4 environment; sequence extra for parquet and HDF5. Podman or Apptainer is optional.; hardware: CPU scoring and small controls; memory/accelerator needs for real private-model inference are model-dependent and unreported here.; instructions: runtime: python; title: Run a small local software control; code: import rewirebench from rewirebench.resources.flip2.example_adapter import CompositionEmbeddings prepared = rewirebench.prepare( 'flip2-fitness-v1', source='flip2-data', output='prepared-flip2', dataset='flip2-rhomax-by-wild-type', limit=8 ) report = rewirebench.run( prepared, CompositionEmbeddings(), output='results-flip2-control', model={'name': 'Local software control; not a published baseline'}, ); status: source_reviewed_not_executed; source ids: runner-04-docs-sdk-md; runner-04-docs-hpc-md; runner-04-docs-flip2-md; runner-04-packages-rewirebench-src-rewirebench-protocols-flip2-py; runner-04-packages-rewirebench-src-rewirebench-resources-flip2-sources-json; runner-04-packages-rewirebench-src-rewirebench-resources-flip2-validation-2026-09-20-json; runner-04-packages-rewirebench-src-rewirebench-resources-flip2-example-adapter-py; source locator: docs/flip2.md and protocol implementation; linked validation receipt concerns its stated tests, not execution of this exact snippet; runtime: command_line; title: Run a small composition control; code: git clone https://github.com/rewire-bio/rewire-benchmarks.git cd rewire-benchmarks git checkout --detach fcccbbcdbe3d5cd64a1a312d536615273320f7b4 uv sync --locked --package rewirebench --extra sequence --python 3.11 uv run --package rewirebench rewirebench prepare flip2-fitness-v1 \ --source flip2-data --output prepared-flip2 \ --options '{"dataset":"flip2-rhomax-by-wild-type","limit":8}' uv run --package rewirebench rewirebench run --prepared prepared-flip2 \ --adapter rewirebench.resources.flip2.example_adapter:CompositionEmbeddings --prediction-type embedding \ --output results-flip2-control; status: source_reviewed_not_executed; source ids: runner-04-docs-sdk-md; runner-04-docs-hpc-md; runner-04-docs-flip2-md; runner-04-packages-rewirebench-src-rewirebench-protocols-flip2-py; runner-04-packages-rewirebench-src-rewirebench-resources-flip2-sources-json; runner-04-packages-rewirebench-src-rewirebench-resources-flip2-validation-2026-09-20-json; runner-04-packages-rewirebench-src-rewirebench-resources-flip2-example-adapter-py; source locator: docs/flip2.md and protocol implementation; linked validation receipt concerns its stated tests, not execution of this exact snippet; limitations: A run is one of16 archived splits across7 datasets; no suite-wide score is computed.; Archived Amylase one_to_many and NucB split counts conflict with manuscript Table1. IRED printed total also differs. The runner preserves archive assignments and records each conflict; it does not claim paper-cohort reproduction.; PDZ3 inputs preserve protein:peptide notation including empty partners. An adapter must handle this representation explicitly.; Frozen-embedding alpha10 ridge is a Rewire extension with train-only target scaling, not the published one-hot ridge baseline.; Native CPU smoke checks covered all16 archived splits with8 rows per train/validation/test group; pinned upstream metric expressions agreed within1e-12. No full model runs or published-score reproduction.; These snippets are source-reviewed instructions. The linked automated receipt supports only the tests it records; it does not certify the exact snippet or published-model reproduction.; Smoke examples and partial predictions retain original denominators and cannot establish a complete suite. Prepared execution is offline; exports and submissions are explicit.; source ids: runner-04-docs-sdk-md; runner-04-docs-hpc-md; runner-04-docs-flip2-md; runner-04-packages-rewirebench-src-rewirebench-protocols-flip2-py; runner-04-packages-rewirebench-src-rewirebench-resources-flip2-sources-json; runner-04-packages-rewirebench-src-rewirebench-resources-flip2-validation-2026-09-20-json; runner-04-packages-rewirebench-src-rewirebench-resources-flip2-example-adapter-py; source locator: docs/flip2.md; source manifest; protocol score function; automated receipt scope
Related records

Suggest a correction