Datasets
Datasets span enzymatic activity, protein interactions and other measured protein properties.
FLIP2 evaluates protein fitness prediction under deliberately shifted training and test distributions.
Datasets span enzymatic activity, protein interactions and other measured protein properties.
Spearman rank correlation is the primary metric, with NDCG additionally measuring prioritization of high-fitness variants.
Protein sequence representations and dataset-specific labels.
Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.
Source reviewed · Automated source review, 2026-09-16. All specifications and missing details
Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.
spearman (correlation) · Higher values are better.
FLIP2 NucB two-to-many; held-out test set · NucB · NucB
Evidence origin: Author-reported evaluation.
FLIP2: Expanding Protein Fitness Landscape Benchmarks for Real-World Machine Learning Applications · Table A6; row Ridge (one-hot); column spearman through Table A6; row ESMC-300M naive supervised; column spearmanComplete selected source table is retained across source-order panels. These point estimates do not establish statistical significance or a universal ranking.
Automated source review: 2026-09-19. Numerical source review does not establish independent reproduction.
Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.
Showing 9 of 9 matching rows.
FLIP2 extends protein fitness testing to enzyme activity, spectral properties, hydrophobic-core effects and protein interactions. Its partitions hold out mutation positions, mutation numbers, fitness ranges or parent sequences. Simple one-hot ridge models and pretrained versus randomly initialized networks make the baseline comparison interpretable.
Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.
These source-backed links do not make different protocols or scores interchangeable.
Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.
2 of 34 active baseline roles have published Rewire measurements in this release. Measurements on a selected protocol do not establish coverage of an entire suite.
Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums
Coverage is derived from release 2026-09-29-06401fd5b220. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.
Choose a concrete protocol before running an evaluation. Its inputs, split and scoring rules determine which results can be compared.
Rank protein variants by measured fitness on one archived train/validation/test split. Spearman correlation and full-ranking NDCG using the upstream minimum-shifted test targets and tie handling.
Recompute metrics from supplied predictions. This recipe does not establish reproduction of a particular published score.
Choose one way to run this recipe. These instruction formats are alternatives.
Source reviewed; these instructions have not been executed by rewire.
import rewirebench
prepared = rewirebench.prepare(
'flip2-fitness-v1', source='flip2-data', output='prepared-flip2', dataset='flip2-rhomax-by-wild-type', limit=8
)
report = rewirebench.evaluate(
prepared, 'predictions.json', output='results-flip2-rescore',
model={'name': 'My private model', 'training_overlap': 'Unreported'},
)rewirebench 0.4 library interface; rewirebench container and HPC instructions; FLIP2 archived fitness splits runner guide; FLIP2 archived fitness splits protocol implementation; FLIP2 archived fitness splits source manifest; FLIP2 archived fitness splits automated validation receipt; FLIP2 CPU composition learning control · docs/flip2.md and protocol implementation; linked validation receipt concerns its stated tests, not execution of this exact snippetRun your model locally and return predictions keyed by the input IDs. The evaluator supplies biological inputs without test labels and owns scoring. This interface is not a sandbox for model code.
Pass your existing prediction function into this adapter. Its output direction must match the selected protocol.
class MyModelAdapter:
def __init__(self, score):
self.score = score
def predict(self, inputs):
return {row["id"]: float(self.score(row)) for row in inputs}
# adapter = MyModelAdapter(your_prediction_function)
# report = rewirebench.run(prepared, adapter, output="runs/my-model")Alternatively, generate a keyed prediction file in your existing model environment and use the score-only recipe. Your model code and weights do not need to be shared.
rewirebench 0.4 library interface; rewirebench container and HPC instructions; FLIP2 archived fitness splits runner guide; FLIP2 archived fitness splits protocol implementation; FLIP2 archived fitness splits source manifest; FLIP2 archived fitness splits automated validation receipt; FLIP2 CPU composition learning control · docs/flip2.md; source manifest; protocol score function; automated receipt scopeContribute a result for review. The library can submit an exported evaluation for private review when intake is open. Check the contribution page for access and sign-in.
Official FLIP2 page links CSV/FASTA datasets and the shared FLIP repository. Choose a FLIP2 dataset and split explicitly; the repository also contains earlier FLIP material, so its root documentation alone does not identify an exact FLIP2 run configuration.
A maintained rewire runner has not been verified for this benchmark. Check data access, weights, licences, dependencies and hardware in the linked official documentation; requirements have not been fully extracted.
flip2 official run documentation · Official FLIP2 site: Data formats, dataset download and GitHub repository sectionsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.
Stable record: discovery-benchmark-flip2Explanatory profile: source reviewed · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Datasets | Datasets span enzymatic activity, protein interactions and other measured protein properties.Sourcesflip2 official source · FLIP2 official website: Overview; benchmark features; datasets and split categories |
| Splits | Split categories separate mutation counts, positions, mutation identities, fitness ranges or reference proteins.Sourcesflip2 official source · FLIP2 official website: Overview; benchmark features; datasets and split categories |
| Metrics | Spearman rank correlation is the primary metric, with NDCG additionally measuring prioritization of high-fitness variants.Sourcesflip2 primary benchmark evidence · Sections 2–4; Table 1; supplementary result tables |
| Baselines | Zero-shot protein models, ridge regression and fine-tuned models.Sourcesflip2 official source · FLIP2 official website: Overview; benchmark features; datasets and split categories |
| Leakage controls | Distribution-shift splits are explicit; exact homology/pretraining overlap depends on the dataset.Sourcesflip2 official source · FLIP2 official website: Overview; benchmark features; datasets and split categories |
| Uncertainty | Fine-tuned language models are run five times with different random seeds and validation-based early stopping; reported metrics are averaged. The zero-shot and ridge evaluations follow separate procedures.Sourcesflip2 primary benchmark evidence · Sections 2–4; Table 1; supplementary result tables |
| Entity type | Protein fitness benchmark collection.Sourcesflip2 official source · FLIP2 official website: Overview; benchmark features; datasets and split categories |
| Organisms | Mixed natural and designed protein landscapes: amylase, imine reductase, NucB, TrpB, hydrophobic-core variants, microbial rhodopsins and a PDZ3–peptide interaction system. These are not a single-species benchmark.Sourcesflip2 primary benchmark evidence · Sections 2–4; Table 1; supplementary result tables |
| Assays | Enzymatic activity, protein interactions and other measured properties.Sourcesflip2 official source · FLIP2 official website: Overview; benchmark features; datasets and split categories |
| Allowed inputs | Protein sequence representations and dataset-specific labels.Sourcesflip2 official source · FLIP2 official website: Overview; benchmark features; datasets and split categories |
| Adaptation | Supervised prediction under task-specific data splits.Sourcesflip2 official source · FLIP2 official website: Overview; benchmark features; datasets and split categories |
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Last literature check: 2026-09-17. Dated primary-source discovery and protocol/table screening. Source checking does not mean experimental reproduction. Only separately extracted and independently reviewed numeric batches are publishable.
| Paper or primary resource | Version | Reference |
|---|---|---|
| FLIP2: Expanding Protein Fitness Landscape Benchmarks for Real-World Machine Learning Applications | Primary full-text snapshot retrieved 2026-09-17; exact bytes pinned by SHA-256 | Read source |
The catalogue now holds 293 result rows for this benchmark. A note below about pending extraction describes the state on 2026-09-17 and may since have been answered by a later batch. The result rows and their sources are the current record.
primary protocol reviewed
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
140 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions. Individual claims | flip2 official source FLIP2 official website: Overview; benchmark features; datasets and split categories; Sections 2–4; Table 1; supplementary result tables Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: Retrieved website snapshot sha256:cd991ae6e76a5e84ea5449f91c4ed86ba4f57942682dc3c50d872c366bbbd4b7 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram caption Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions. Individual claims | flip2 primary benchmark evidence FLIP2 official website: Overview; benchmark features; datasets and split categories; Sections 2–4; Table 1; supplementary result tables Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: ICML2026 manuscript | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| flip2 official source FLIP2 official website: Overview; benchmark features; datasets and split categories; Sections 2–4; Table 1; supplementary result tables Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: Retrieved website snapshot sha256:cd991ae6e76a5e84ea5449f91c4ed86ba4f57942682dc3c50d872c366bbbd4b7 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Diagram steps
| flip2 primary benchmark evidence FLIP2 official website: Overview; benchmark features; datasets and split categories; Sections 2–4; Table 1; supplementary result tables Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: ICML2026 manuscript | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Evaluation procedure Individual claims | flip2 official source FLIP2 official website: Overview; benchmark features; datasets and split categories; Sections 2–4; Table 1; supplementary result tables Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: Retrieved website snapshot sha256:cd991ae6e76a5e84ea5449f91c4ed86ba4f57942682dc3c50d872c366bbbd4b7 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Evaluation procedure Individual claims | flip2 primary benchmark evidence FLIP2 official website: Overview; benchmark features; datasets and split categories; Sections 2–4; Table 1; supplementary result tables Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: ICML2026 manuscript | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Datasets Datasets span enzymatic activity, protein interactions and other measured protein properties. Individual claims | flip2 official source FLIP2 official website: Overview; benchmark features; datasets and split categories Version: Retrieved website snapshot sha256:cd991ae6e76a5e84ea5449f91c4ed86ba4f57942682dc3c50d872c366bbbd4b7 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Splits Split categories separate mutation counts, positions, mutation identities, fitness ranges or reference proteins. Individual claims | flip2 official source FLIP2 official website: Overview; benchmark features; datasets and split categories Version: Retrieved website snapshot sha256:cd991ae6e76a5e84ea5449f91c4ed86ba4f57942682dc3c50d872c366bbbd4b7 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Adaptation Supervised prediction under task-specific data splits. Individual claims | flip2 official source FLIP2 official website: Overview; benchmark features; datasets and split categories Version: Retrieved website snapshot sha256:cd991ae6e76a5e84ea5449f91c4ed86ba4f57942682dc3c50d872c366bbbd4b7 | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Metrics Spearman rank correlation is the primary metric, with NDCG additionally measuring prioritization of high-fitness variants. Individual claims | flip2 primary benchmark evidence Sections 2–4; Table 1; supplementary result tables Version: ICML2026 manuscript | source checked automated source review · 2026-09-16 Audit detailsPrimary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
View linked audit checks and correction history
Release 2026-09-29-06401fd5b220 · Record review: discovered
Stable ID: discovery-benchmark-flip2