rewire.itbenchmarks
Protocol

MFASS v2

MFASS v2 evaluates splice-variant prioritisation using a corrected sequence baseline and a fixed grouped holdout.

SourcesMFASS computational protocol README at bee9133b83f3aedaf2bbb9013f1875515845607e · Pinned README: correction notice; Dataset; Cohort reconciliation; Split; MFASS-v2 results; Paired comparisons; Limits; Archived mfass-v1 snapshot

9 evaluations · 31 results

Overview

Metrics

Precision at a fixed review capacity, average precision and AUROC. Each method’s point estimates use its scored subset; paired comparisons use common scored variants.

SourcesMFASS computational protocol README at bee9133b83f3aedaf2bbb9013f1875515845607e · Pinned README: correction notice; Dataset; Cohort reconciliation; Split; MFASS-v2 results; Paired comparisons; Limits; Archived mfass-v1 snapshot

Allowed inputs

Sequence baselines use validated 170-base natural_seq reference and original_seq mutant pairs, differing at rel_position minus one with strand-checked alleles. SpliceAI and Pangolin instead use genomic context; their predictions are not evaluations of the same 170-base input.

SourcesMFASS computational protocol README at bee9133b83f3aedaf2bbb9013f1875515845607e · README.md lines 59–64, 141–179
Evaluation procedure diagram
How it worksComputational evaluation flow
Computational evaluation flow1. Input: Validated assay-oriented sequence pairs for baseline/DNABERT-2; specialists use genomic context.. Then: 2. Evaluation: Baseline and logistic head use MFASS training labels; specialists are zero-shot on the assay.. Then: 3. Readout: Precision at a fixed review capacity, average precision and AUROC. Each method’s point estimates use its scored subset; paired comparisons use common scored variants.Computational evaluation flow1. Input: Validated assay-oriented sequence pairs for baseline/DNABERT-2; specialists use genomic context.. Then: 2. Evaluation: Baseline and logistic head use MFASS training labels; specialists are zero-shot on the assay.. Then: 3. Readout: Precision at a fixed review capacity, average precision and AUROC. Each method’s point estimates use its scored subset; paired comparisons use common scored variants.Computational evaluation flow1. Input: Validated assay-oriented sequence pairs for baseline/DNABERT-2; specialists use genomic context.. Then: 2. Evaluation: Baseline and logistic head use MFASS training labels; specialists are zero-shot on the assay.. Then: 3. Readout: Precision at a fixed review capacity, average precision and AUROC. Each method’s point estimates use its scored subset; paired comparisons use common scored variants.

Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.

SourcesMFASS computational protocol README at bee9133b83f3aedaf2bbb9013f1875515845607e · Pinned README: correction notice; Dataset; Cohort reconciliation; Split; MFASS-v2 results; Paired comparisons; Limits; Archived mfass-v1 snapshot

limited source coverage · Automated source review, 2026-09-23. All specifications and missing details

Results

Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.

MFASS matched canonical annotation: Precision at 100

Precision at 100 (dimensionless) · Higher values are better.

MFASS: matched GENCODE 44 canonical annotation · MFASS v2 test: matched canonical annotation coverage

Evidence origin: Rewire evaluation.

MFASS matched canonical annotation v1: report.json; MFASS matched canonical annotation v1: manifest-v1.json; MFASS matched canonical annotation v1: verification.json; MFASS matched canonical annotation v1: provenance.json; MFASS matched canonical annotation v1: exclusion-verification.json · report.json: conditions.*.metrics.precision_at_capacity
  • 8,297 of 8,324 held-out variants scored in every configuration (314 positives, 460 exon groups). The same 27 rows are excluded: 23 assembly-orientation mismatches and four outside the selected canonical transcript spans. Missing scores are not zero or negative predictions.
  • Exploratory comparison: prior results were known. Paired contrast intervals are unadjusted and do not establish a universal model ranking.
  • Pangolin uses the recorded per-gene masking patch; these are exact configurations, not unqualified upstream model scores.
  • Precision at 100 is sensitive to tied-score ordering, especially P1. Numerical source checking is automated, not human review or independent reproduction.
  • Assembly-orientation issue reported at https://github.com/KosuriLab/MFASS/issues/1. Original v1 outputs remain unchanged; corrections require a new version.
Comparison details and limitations

8,297 of 8,324 held-out variants scored in every configuration (314 positives, 460 exon groups). The same 27 rows are excluded: 23 assembly-orientation mismatches and four outside the selected canonical transcript spans. Missing scores are not zero or negative predictions.

Automated source review: 2026-09-25. Numerical source review does not establish independent reproduction.

Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.

Showing 4 of 4 matching rows.

Methods and evaluation design

Procedure, tasks and evaluated configurations

How it works

Evaluation methodology

Functional exon-recognition labels from the MFASS assay; validated assay-oriented reference/mutant sequence pairs define the sequence-model inputs. The canonical split assigns whole connected exon/gene groups to training or test. Prevalence matching is selected before model runs and uses no predictions. Precision at a fixed review capacity, average precision and AUROC. Each method’s point estimates use its scored subset; paired comparisons use common scored variants. Corrected k-mer/position baseline, frozen DNABERT-2 pair embeddings with a fixed logistic head, and SpliceAI/Pangolin genomic-context specialists. Baseline and logistic-head fitting use training labels only. Exon/gene connected components prevent linked variants crossing arms; exact DNABERT-2 pretraining overlap has not been checked. Paired intervals resample whole connected exon/gene groups; precision intervals account for fixed review-list fraction during group resampling.

SourcesMFASS computational protocol README at bee9133b83f3aedaf2bbb9013f1875515845607e · Pinned README: correction notice; Dataset; Cohort reconciliation; Split; MFASS-v2 results; Paired comparisons; Limits; Archived mfass-v1 snapshot

Evaluation design

Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.

These source-backed links do not make different protocols or scores interchangeable.

Recorded evaluations

Each evaluation records what was tested and under which conditions.

Baseline coverage

Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.

2 of 2 active baseline roles have published Rewire measurements in this release. Measurements on a selected protocol do not establish coverage of an entire suite.

3 execution recipes Recipe availability does not establish a completed evaluation.

Published Rewire evaluations
5

Literature evidence is not a Rewire measurement. Executed but unpublished runs and private review status are not included.

Null control

Published Rewire reference

Training class prior / majority (MFASS v2 canonical held-out split)

Training class prior / majority on MFASS v2 canonical held-out split: methods, coverage and results

Evidence: Reviewed baseline-runs-2026-09-22: MFASS training-prior report metrics, canonical split and coverage; fixed tie-break, not ranking ability

Conventional reference

Published Rewire reference

Corrected k-mer / position baseline

Corrected k-mer / position baseline on MFASS v2: methods, coverage and results

Evidence: benchmarks/mfass/results/baseline-kmer-position-v2.json at bee9133b83f3aedaf2bbb9013f1875515845607e; existing release results and corrected split-v2

Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums

Coverage is derived from release 2026-09-29-06401fd5b220. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.

Run this benchmark

Choose a concrete protocol before running an evaluation. Its inputs, split and scoring rules determine which results can be compared.

MFASS v2: score existing predictions

Evaluate keyed splice-disruption scores on the fixed MFASS v2 test split. Higher scores mean greater splice disruption.

Recompute metrics from supplied predictions. This recipe does not establish reproduction of a particular published score.

Dataset access
Obtain the validated MFASS v2 cohort.tsv. Preparation checks its SHA-256 and the canonical 27,733-variant split from bee9133b83f3aedaf2bbb9013f1875515845607e.
Model and weights
No model or weights are required to rescore supplied predictions.
Licences
Runner code is MIT. The upstream MFASS data reuse licence is unreported; original data are not bundled.
Software
Python 3.11 with the pinned rewirebench environment. Container runtimes are optional.
Hardware
CPU scoring; minimum memory and runtime have not been established for this recipe.
Required inputs and expected outputs

Inputs

  • Validated cohort.tsv; the canonical split is packaged with the library.
  • CSV/TSV columns id, score and optional reason; every test variant needs a score or an explicit unscored reason.

Outputs

  • Local report.json with metrics, coverage, protocol and artifact hashes.
  • Local predictions.json and unscored.json; neither is submitted automatically.

Choose one way to run this recipe. These instruction formats are alternatives.

Execution steps

  1. 1. Prepare and score local predictions (Python)

    Source reviewed; these instructions have not been executed by rewire.

    import rewirebench
    prepared = rewirebench.prepare(
        "mfass-v2", source="data/mfass", output="prepared"
    )
    report = rewirebench.evaluate(
        prepared, "predictions/predictions.tsv", output="runs/mfass-rescore",
        model={"name": "My model", "training_overlap": "Unreported"},
    )
    rewirebench: Library guide; rewirebench: MFASS guide; rewirebench: MFASS protocol code; rewirebench: MFASS adapter code · docs/sdk.md: prepare/evaluate; docs/mfass.md: canonical cohort and keyed predictions

Use your own model

Run your model locally and return predictions keyed by the input IDs. The evaluator supplies biological inputs without test labels and owns scoring. This interface is not a sandbox for model code.

Pass your existing prediction function into this adapter. Its output direction must match the selected protocol.

class MyModelAdapter:
    def __init__(self, score):
        self.score = score

    def predict(self, inputs):
        return {row["id"]: float(self.score(row)) for row in inputs}

# adapter = MyModelAdapter(your_prediction_function)
# report = rewirebench.run(prepared, adapter, output="runs/my-model")

Alternatively, generate a keyed prediction file in your existing model environment and use the score-only recipe. Your model code and weights do not need to be shared.

rewirebench: Library guide; rewirebench: MFASS guide; rewirebench: MFASS protocol code; rewirebench: MFASS adapter code · Pinned runner documentation, protocol implementation and adapter implementation
Scope and limitations
  • This preserves the corrected assay-oriented variant placement and canonical train/test identities.
  • Scores use held-out variants; missing predictions retain the original denominator.
  • Rescoring supplied SpliceAI or Pangolin predictions does not recreate their genomic-context inference.
  • Smoke inputs and partial coverage do not establish a full benchmark or reproduce a published result.

Contribute a result for review. The library can submit an exported evaluation for private review when intake is open. Check the contribution page for access and sign-in.

Original repository instructions

Official run instructions

Run the corrected MFASS v2 cohort validation and baseline, with an optional DNABERT-2 resource pilot. Keep this protocol separate from the superseded v1 run.

Checked against the official instructions on 2026-09-17. These commands have not been executed by rewire. Running them does not automatically reproduce the published scores.

Before you start

  • Git, curl and uv; the package declares Python >=3.11. The README commands run from the rewire-benchmarks repository root.
  • The checked-in split-v2.tsv is canonical. The README records SHA-256 hashes for both source tables and the split; inspect them before treating a rerun as matching the published protocol.
  1. 1. Check out the reviewed repository

    Repository checkout wrapper: the detached revision selects the exact official source inspected for this guide.

    git clone https://github.com/rewire-bio/rewire-benchmarks.git
    cd rewire-benchmarks
    git checkout --detach 68d1ee53a0fcd1104d9f427d68dedffdccf2c47f
    rewire-bio/rewire-benchmarks / benchmarks/mfass/README.md · Pinned repository revision; benchmarks/mfass/README.md
  2. 2. Fetch the two published input tables

    Official download commands. These upstream master URLs are mutable; the source README records the expected input hashes.

    mkdir -p benchmarks/mfass/data
    curl -L -o benchmarks/mfass/data/snv_data_clean.txt \
      https://raw.githubusercontent.com/KosuriLab/MFASS/master/processed_data/snv/snv_data_clean.txt
    curl -L -o benchmarks/mfass/data/snv_func_annot.txt \
      https://raw.githubusercontent.com/KosuriLab/MFASS/master/processed_data/snv/snv_func_annot.txt
    rewire-bio/rewire-benchmarks / benchmarks/mfass/README.md · benchmarks/mfass/README.md lines 90–106
  3. 3. Install, validate the cohort and run the corrected baseline

    Preserves the official v2 command sequence. The baseline and encoder head are supervised on the predeclared training split.

    uv sync --package mfass --extra dnabert2-pilot
    uv run --package mfass --extra dnabert2-pilot mfass-build --check
    uv run --package mfass --extra dnabert2-pilot mfass-build
    uv run --package mfass --extra dnabert2-pilot mfass-baseline
    rewire-bio/rewire-benchmarks / benchmarks/mfass/README.md; rewire-bio/rewire-benchmarks / benchmarks/mfass/pyproject.toml · benchmarks/mfass/README.md lines 107–110; benchmarks/mfass/pyproject.toml lines 1–33
  4. 4. Optionally run the resource pilot

    The pilot measures feasibility and validates checkpoint provenance; it does not calculate accuracy. The README gates any subsequent full encoder run on projected runtime and free disk; the full run is intentionally a separate decision.

    uv run --package mfass --extra dnabert2-pilot mfass-dnabert2-pilot
    rewire-bio/rewire-benchmarks / benchmarks/mfass/README.md · benchmarks/mfass/README.md lines 111–139

Expected outputs

  • Validated assay-oriented cohort and corrected baseline results.
  • If selected, pilot runtime/memory/disk/failure measurements and model/code provenance checks; no pilot accuracy score.

Scope and limitations

  • The source reports a local run without paid APIs or cloud compute; that is not a runtime or cost guarantee for a different machine.
  • MFASS labels reflect exon recognition in an assay construct; they are not patient-RNA outcomes.
  • This command sequence does not rerun SpliceAI or Pangolin and must not relabel their previously published genomic-context predictions.
  • Do not use mfass-v1 baseline or split-cost results as current evidence.
Strengths, limitations and unresolved questions

Strengths and limitations

Limitations and conditions

Profile review details

Follow-up review of Datasets, Splits, Scoring coverage, Uncertainty, Allowed inputs, Adaptation. Source locations and before/after decisions are recorded in the 23 September profile-evidence audit. Other explanatory content retains its earlier source scope. No human scientific review or independent reproduction is implied.

Stable record: rewire-mfass-v2

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-23. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetsThe eligibility filter category=mutant with a reported strong_lof label selects 27,733 variants and 1,050 disrupting variants from 32,669 source rows. The eligible cohort spans 2,185 exons; the paper’s 2,198-exon count precedes the label filter.
SourcesMFASS computational protocol README at bee9133b83f3aedaf2bbb9013f1875515845607e · README.md lines 40–64
SplitsCanonical split-v2 assigns connected exon/gene components whole: 19,409 training variants (735 positives; 1,127 groups) and 8,324 test variants (315 positives; 463 groups). Assignment uses seed 20260914, 200 candidate draws and a requested 30% test fraction, matching prevalence before predictions. The split TSV SHA-256 is 999ebcb7e63a5c5eaa8780fa468e59ac1f934260ad50102814174c396317f052.
Sources (3)MFASS computational protocol README at bee9133b83f3aedaf2bbb9013f1875515845607e; Protocol metadata evidence: mfass-splits-split-v2.manifest.json; Protocol metadata evidence: mfass-splits-split-v2.tsv · README.md lines 66–99; split-v2.manifest.json: seed, restarts, test_frac_requested, train, test; split-v2.tsv bytes
MetricsPrecision at a fixed review capacity, average precision and AUROC. Each method’s point estimates use its scored subset; paired comparisons use common scored variants.
SourcesMFASS computational protocol README at bee9133b83f3aedaf2bbb9013f1875515845607e · Pinned README: correction notice; Dataset; Cohort reconciliation; Split; MFASS-v2 results; Paired comparisons; Limits; Archived mfass-v1 snapshot
BaselinesCorrected k-mer/position baseline, frozen DNABERT-2 pair embeddings with a fixed logistic head, and SpliceAI/Pangolin genomic-context specialists.
SourcesMFASS computational protocol README at bee9133b83f3aedaf2bbb9013f1875515845607e · Pinned README: correction notice; Dataset; Cohort reconciliation; Split; MFASS-v2 results; Paired comparisons; Limits; Archived mfass-v1 snapshot
Leakage controlsBaseline and logistic-head fitting use training labels only. Exon/gene connected components prevent linked variants crossing arms; exact DNABERT-2 pretraining overlap has not been checked.
SourcesMFASS computational protocol README at bee9133b83f3aedaf2bbb9013f1875515845607e · Pinned README: correction notice; Dataset; Cohort reconciliation; Split; MFASS-v2 results; Paired comparisons; Limits; Archived mfass-v1 snapshot
UncertaintyThe three pinned candidate-versus-corrected-baseline comparisons use 2,000 paired whole-group bootstrap draws at seed 20260914 and percentile 95% intervals. All report zero single-class draws skipped. The point estimate is the observed candidate-minus-baseline difference on the common subset, not the resample mean. Precision resamples preserve the review-list fraction corresponding to capacity 100, so their realised capacities vary.
Sources (5)Protocol metadata evidence: mfass-metrics.py; Protocol metadata evidence: mfass-src-mfass-compare.py; Protocol metadata evidence: mfass-results-compare-baseline-v2-vs-spliceai.json; Protocol metadata evidence: mfass-results-compare-baseline-v2-vs-pangolin-maskFalse.json; Protocol metadata evidence: mfass-results-compare-baseline-v2-vs-dnabert2.json · metrics.py lines 65–140; compare.py lines 33–35, 41–74; comparison JSON: capacity, seed and paired.*
Entity typePaper-specific computational evaluation protocol.
SourcesMFASS computational protocol README at bee9133b83f3aedaf2bbb9013f1875515845607e · Pinned README: correction notice; Dataset; Cohort reconciliation; Split; MFASS-v2 results; Paired comparisons; Limits; Archived mfass-v1 snapshot
OrganismsHuman (Homo sapiens): naturally occurring ExAC variants in or beside human exons, measured in the MFASS minigene assay.
SourcesMFASS original assay accession GSE120695 · GSE120695 SOFT: Series_summary and Sample_organism_ch1 fields
AssaysMFASS functional exon-recognition labels.
SourcesMFASS computational protocol README at bee9133b83f3aedaf2bbb9013f1875515845607e · Pinned README: correction notice; Dataset; Cohort reconciliation; Split; MFASS-v2 results; Paired comparisons; Limits; Archived mfass-v1 snapshot
Allowed inputsSequence baselines use validated 170-base natural_seq reference and original_seq mutant pairs, differing at rel_position minus one with strand-checked alleles. SpliceAI and Pangolin instead use genomic context; their predictions are not evaluations of the same 170-base input.
SourcesMFASS computational protocol README at bee9133b83f3aedaf2bbb9013f1875515845607e · README.md lines 59–64, 141–179
AdaptationThe corrected k-mer baseline and DNABERT-2 logistic head use the training arm. DNABERT-2 stays frozen: masked-mean last-hidden-state embeddings feed reference and mutant-minus-reference vectors to a fixed balanced L2 logistic head (C=0.1). Specialist scores are zero-shot with respect to MFASS training labels; input context and annotation versions differ.
SourcesMFASS computational protocol README at bee9133b83f3aedaf2bbb9013f1875515845607e · README.md lines 116–138 and 141–179
Scoring coverageAgainst the 8,324-variant held-out denominator, the corrected baseline and DNABERT-2 pipeline each score 8,324 variants, SpliceAI scores 8,194, and Pangolin mask=False scores 8,301. Paired comparisons use the common scored variants: 8,194 in 454 groups for SpliceAI, 8,301 in 461 groups for Pangolin, and 8,324 in 463 groups for DNABERT-2. Missing predictions remain outside the paired subset.
Sources (3)Protocol metadata evidence: mfass-results-compare-baseline-v2-vs-spliceai.json; Protocol metadata evidence: mfass-results-compare-baseline-v2-vs-pangolin-maskFalse.json; Protocol metadata evidence: mfass-results-compare-baseline-v2-vs-dnabert2.json · Each comparison JSON: denominators, independent_groups and on_common_subset; candidate settings retained in filenames

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Dated primary-source discovery and protocol/table screening. Source checking does not mean experimental reproduction. Only separately extracted and independently reviewed numeric batches are publishable.

Paper or primary resourceVersionReference
MFASS benchmark: corrected v2 protocol and reproducibility instructionsPrimary full-text snapshot retrieved 2026-09-17; exact bytes pinned by SHA-256Read source
Search and extraction details

historical results preserved

Searches

  • Pangolin SpliceAI MFASS splice variant benchmark rewire

Evidence locations

  • Pinned README baseline-v2 and comparison commands at bee9133b83f3aedaf2bbb9013f1875515845607e

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

127 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.
Individual claims
MFASS computational protocol README at bee9133b83f3aedaf2bbb9013f1875515845607e

Original source ↗

Pinned README: correction notice; Dataset; Cohort reconciliation; Split; MFASS-v2 results; Paired comparisons; Limits; Archived mfass-v1 snapshot

Version: bee9133b83f3aedaf2bbb9013f1875515845607e
Retrieved: 2026-09-16T20:55:04.446365+00:00

source checked

automated source review · 2026-09-23

Audit details

Follow-up review of Datasets, Splits, Scoring coverage, Uncertainty, Allowed inputs, Adaptation. Source locations and before/after decisions are recorded in the 23 September profile-evidence audit. Other explanatory content retains its earlier source scope. No human scientific review or independent reproduction is implied.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 62a9381484fd2360767e78b114d9aa6cd8a4fc6487a981881a4c4999f0ab1b23

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps
  • Input: Validated assay-oriented sequence pairs for baseline/DNABERT-2; specialists use genomic context.
  • Evaluation: Baseline and logistic head use MFASS training labels; specialists are zero-shot on the assay.
  • Readout: Precision at a fixed review capacity, average precision and AUROC. Each method’s point estimates use its scored subset; paired comparisons use common scored variants.
Individual claims
MFASS computational protocol README at bee9133b83f3aedaf2bbb9013f1875515845607e

Original source ↗

Pinned README: correction notice; Dataset; Cohort reconciliation; Split; MFASS-v2 results; Paired comparisons; Limits; Archived mfass-v1 snapshot

Version: bee9133b83f3aedaf2bbb9013f1875515845607e
Retrieved: 2026-09-16T20:55:04.446365+00:00

source checked

automated source review · 2026-09-23

Audit details

Follow-up review of Datasets, Splits, Scoring coverage, Uncertainty, Allowed inputs, Adaptation. Source locations and before/after decisions are recorded in the 23 September profile-evidence audit. Other explanatory content retains its earlier source scope. No human scientific review or independent reproduction is implied.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 62a9381484fd2360767e78b114d9aa6cd8a4fc6487a981881a4c4999f0ab1b23

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title
Computational evaluation flow
Individual claims
MFASS computational protocol README at bee9133b83f3aedaf2bbb9013f1875515845607e

Original source ↗

Pinned README: correction notice; Dataset; Cohort reconciliation; Split; MFASS-v2 results; Paired comparisons; Limits; Archived mfass-v1 snapshot

Version: bee9133b83f3aedaf2bbb9013f1875515845607e
Retrieved: 2026-09-16T20:55:04.446365+00:00

source checked

automated source review · 2026-09-23

Audit details

Follow-up review of Datasets, Splits, Scoring coverage, Uncertainty, Allowed inputs, Adaptation. Source locations and before/after decisions are recorded in the 23 September profile-evidence audit. Other explanatory content retains its earlier source scope. No human scientific review or independent reproduction is implied.

Field: attributes.profile.diagram.title

Source artifact SHA-256: 62a9381484fd2360767e78b114d9aa6cd8a4fc6487a981881a4c4999f0ab1b23

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Datasets
The eligibility filter category=mutant with a reported strong_lof label selects 27,733 variants and 1,050 disrupting variants from 32,669 source rows. The eligible cohort spans 2,185 exons; the paper’s 2,198-exon count precedes the label filter.
Individual claims
MFASS computational protocol README at bee9133b83f3aedaf2bbb9013f1875515845607e

Original source ↗

README.md lines 40–64

Version: bee9133b83f3aedaf2bbb9013f1875515845607e
Retrieved: 2026-09-16T20:55:04.446365+00:00

source checked

automated source review · 2026-09-23

Audit details

Follow-up review of Datasets, Splits, Scoring coverage, Uncertainty, Allowed inputs, Adaptation. Source locations and before/after decisions are recorded in the 23 September profile-evidence audit. Other explanatory content retains its earlier source scope. No human scientific review or independent reproduction is implied.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: 62a9381484fd2360767e78b114d9aa6cd8a4fc6487a981881a4c4999f0ab1b23

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Splits
Canonical split-v2 assigns connected exon/gene components whole: 19,409 training variants (735 positives; 1,127 groups) and 8,324 test variants (315 positives; 463 groups). Assignment uses seed 20260914, 200 candidate draws and a requested 30% test fraction, matching prevalence before predictions. The split TSV SHA-256 is 999ebcb7e63a5c5eaa8780fa468e59ac1f934260ad50102814174c396317f052.
Individual claims
MFASS computational protocol README at bee9133b83f3aedaf2bbb9013f1875515845607e

Original source ↗

README.md lines 66–99; split-v2.manifest.json: seed, restarts, test_frac_requested, train, test; split-v2.tsv bytes

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: bee9133b83f3aedaf2bbb9013f1875515845607e
Retrieved: 2026-09-16T20:55:04.446365+00:00

source checked

automated source review · 2026-09-23

Audit details

Follow-up review of Datasets, Splits, Scoring coverage, Uncertainty, Allowed inputs, Adaptation. Source locations and before/after decisions are recorded in the 23 September profile-evidence audit. Other explanatory content retains its earlier source scope. No human scientific review or independent reproduction is implied.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: 62a9381484fd2360767e78b114d9aa6cd8a4fc6487a981881a4c4999f0ab1b23

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Splits
Canonical split-v2 assigns connected exon/gene components whole: 19,409 training variants (735 positives; 1,127 groups) and 8,324 test variants (315 positives; 463 groups). Assignment uses seed 20260914, 200 candidate draws and a requested 30% test fraction, matching prevalence before predictions. The split TSV SHA-256 is 999ebcb7e63a5c5eaa8780fa468e59ac1f934260ad50102814174c396317f052.
Individual claims
Protocol metadata evidence: mfass-splits-split-v2.manifest.json

Original source ↗

README.md lines 66–99; split-v2.manifest.json: seed, restarts, test_frac_requested, train, test; split-v2.tsv bytes

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: bee9133b83f3aedaf2bbb9013f1875515845607e
Retrieved: 2026-09-23T18:39:03.003753+00:00

source checked

automated source review · 2026-09-23

Audit details

Follow-up review of Datasets, Splits, Scoring coverage, Uncertainty, Allowed inputs, Adaptation. Source locations and before/after decisions are recorded in the 23 September profile-evidence audit. Other explanatory content retains its earlier source scope. No human scientific review or independent reproduction is implied.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: 3a6e0754cab398f1bfc7ba6080067f065c6b70ab8af53cd27e016d11b0b8f618

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Splits
Canonical split-v2 assigns connected exon/gene components whole: 19,409 training variants (735 positives; 1,127 groups) and 8,324 test variants (315 positives; 463 groups). Assignment uses seed 20260914, 200 candidate draws and a requested 30% test fraction, matching prevalence before predictions. The split TSV SHA-256 is 999ebcb7e63a5c5eaa8780fa468e59ac1f934260ad50102814174c396317f052.
Individual claims
Protocol metadata evidence: mfass-splits-split-v2.tsv

Original source ↗

README.md lines 66–99; split-v2.manifest.json: seed, restarts, test_frac_requested, train, test; split-v2.tsv bytes

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: bee9133b83f3aedaf2bbb9013f1875515845607e
Retrieved: 2026-09-23T18:39:03.130785+00:00

source checked

automated source review · 2026-09-23

Audit details

Follow-up review of Datasets, Splits, Scoring coverage, Uncertainty, Allowed inputs, Adaptation. Source locations and before/after decisions are recorded in the 23 September profile-evidence audit. Other explanatory content retains its earlier source scope. No human scientific review or independent reproduction is implied.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: 999ebcb7e63a5c5eaa8780fa468e59ac1f934260ad50102814174c396317f052

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Adaptation
The corrected k-mer baseline and DNABERT-2 logistic head use the training arm. DNABERT-2 stays frozen: masked-mean last-hidden-state embeddings feed reference and mutant-minus-reference vectors to a fixed balanced L2 logistic head (C=0.1). Specialist scores are zero-shot with respect to MFASS training labels; input context and annotation versions differ.
Individual claims
MFASS computational protocol README at bee9133b83f3aedaf2bbb9013f1875515845607e

Original source ↗

README.md lines 116–138 and 141–179

Version: bee9133b83f3aedaf2bbb9013f1875515845607e
Retrieved: 2026-09-16T20:55:04.446365+00:00

source checked

automated source review · 2026-09-23

Audit details

Follow-up review of Datasets, Splits, Scoring coverage, Uncertainty, Allowed inputs, Adaptation. Source locations and before/after decisions are recorded in the 23 September profile-evidence audit. Other explanatory content retains its earlier source scope. No human scientific review or independent reproduction is implied.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: 62a9381484fd2360767e78b114d9aa6cd8a4fc6487a981881a4c4999f0ab1b23

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Scoring coverage
Against the 8,324-variant held-out denominator, the corrected baseline and DNABERT-2 pipeline each score 8,324 variants, SpliceAI scores 8,194, and Pangolin mask=False scores 8,301. Paired comparisons use the common scored variants: 8,194 in 454 groups for SpliceAI, 8,301 in 461 groups for Pangolin, and 8,324 in 463 groups for DNABERT-2. Missing predictions remain outside the paired subset.
Individual claims
Protocol metadata evidence: mfass-results-compare-baseline-v2-vs-dnabert2.json

Original source ↗

Each comparison JSON: denominators, independent_groups and on_common_subset; candidate settings retained in filenames

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: bee9133b83f3aedaf2bbb9013f1875515845607e
Retrieved: 2026-09-23T18:39:03.239354+00:00

source checked

automated source review · 2026-09-23

Audit details

Follow-up review of Datasets, Splits, Scoring coverage, Uncertainty, Allowed inputs, Adaptation. Source locations and before/after decisions are recorded in the 23 September profile-evidence audit. Other explanatory content retains its earlier source scope. No human scientific review or independent reproduction is implied.

Field: attributes.profile.facts.11.value

Source artifact SHA-256: 52216e003f4e5332ecb43511864c9f0b3df7b1ed0f731c631412dc7311f505c3

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Scoring coverage
Against the 8,324-variant held-out denominator, the corrected baseline and DNABERT-2 pipeline each score 8,324 variants, SpliceAI scores 8,194, and Pangolin mask=False scores 8,301. Paired comparisons use the common scored variants: 8,194 in 454 groups for SpliceAI, 8,301 in 461 groups for Pangolin, and 8,324 in 463 groups for DNABERT-2. Missing predictions remain outside the paired subset.
Individual claims
Protocol metadata evidence: mfass-results-compare-baseline-v2-vs-pangolin-maskFalse.json

Original source ↗

Each comparison JSON: denominators, independent_groups and on_common_subset; candidate settings retained in filenames

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: bee9133b83f3aedaf2bbb9013f1875515845607e
Retrieved: 2026-09-23T18:39:03.230617+00:00

source checked

automated source review · 2026-09-23

Audit details

Follow-up review of Datasets, Splits, Scoring coverage, Uncertainty, Allowed inputs, Adaptation. Source locations and before/after decisions are recorded in the 23 September profile-evidence audit. Other explanatory content retains its earlier source scope. No human scientific review or independent reproduction is implied.

Field: attributes.profile.facts.11.value

Source artifact SHA-256: fd938e1414fa1c6611c6b62597c7b7a3c2df0078301e47e23ba82183095c1097

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: source checked

19 source records and release history

Supersedes MFASS v1 (superseded)

Download this release
Technical metadata and extraction receipts

Stable ID: rewire-mfass-v2

areas
dna-genomes
entity level
protocol
version
bee9133b83f3aedaf2bbb9013f1875515845607e
task
Splice-variant prioritisation
scope note
The mfass-v1 baseline used a mis-centred k-mer window for 7,770 assay variants whose raw sequence was reverse-complemented. mfass-v2 validates assay-oriented reference and mutant pairs and rebuilds the baseline; mfass-v1 remains a historical record.
benchmark research
review date: 2026-09-17; status: historical_results_preserved; primary sources: evidence-expansion-mfass-readme-62a93814; inspected locators: Pinned README baseline-v2 and comparison commands at bee9133b83f3aedaf2bbb9013f1875515845607e; searched queries: Pangolin SpliceAI MFASS splice variant benchmark rewire; gaps: None recorded; claim scope: Dated primary-source discovery and protocol/table screening. Source checking does not mean experimental reproduction. Only separately extracted and independently reviewed numeric batches are publishable.
historical missing metadata
None recorded
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
legacy kinds
benchmark
entity classification
review date: 2026-09-17; rationale: The source-backed record identifies a specified evaluated procedure and its dataset/split/scoring context. Classify it as a protocol while preserving version and comparison restrictions.; source ids: evidence-benchmark-mfass-pinned-readme; source locator: Pinned README: correction notice; Dataset; Cohort reconciliation; Split; MFASS-v2 results; Paired comparisons; Limits; Archived mfass-v1 snapshot; ambiguities: None recorded
run guide
record id: rewire-mfass-v2; summary: Run the corrected MFASS v2 cohort validation and baseline, with an optional DNABERT-2 resource pilot. Keep this protocol separate from the superseded v1 run.; status: source_reviewed_not_executed; prerequisites: Git, curl and uv; the package declares Python >=3.11. The README commands run from the rewire-benchmarks repository root.; The checked-in split-v2.tsv is canonical. The README records SHA-256 hashes for both source tables and the split; inspect them before treating a rerun as matching the published protocol.; steps: title: Check out the reviewed repository; shell: git clone https://github.com/rewire-bio/rewire-benchmarks.git cd rewire-benchmarks git checkout --detach 68d1ee53a0fcd1104d9f427d68dedffdccf2c47f; explanation: Repository checkout wrapper: the detached revision selects the exact official source inspected for this guide.; source ids: run-doc-mfass-readme-md-68d1ee53; source locator: Pinned repository revision; benchmarks/mfass/README.md; title: Fetch the two published input tables; shell: mkdir -p benchmarks/mfass/data curl -L -o benchmarks/mfass/data/snv_data_clean.txt \ https://raw.githubusercontent.com/KosuriLab/MFASS/master/processed_data/snv/snv_data_clean.txt curl -L -o benchmarks/mfass/data/snv_func_annot.txt \ https://raw.githubusercontent.com/KosuriLab/MFASS/master/processed_data/snv/snv_func_annot.txt; explanation: Official download commands. These upstream master URLs are mutable; the source README records the expected input hashes.; source ids: run-doc-mfass-readme-md-68d1ee53; source locator: benchmarks/mfass/README.md lines 90–106; title: Install, validate the cohort and run the corrected baseline; shell: uv sync --package mfass --extra dnabert2-pilot uv run --package mfass --extra dnabert2-pilot mfass-build --check uv run --package mfass --extra dnabert2-pilot mfass-build uv run --package mfass --extra dnabert2-pilot mfass-baseline; explanation: Preserves the official v2 command sequence. The baseline and encoder head are supervised on the predeclared training split.; source ids: run-doc-mfass-readme-md-68d1ee53; run-doc-mfass-pyproject-toml-68d1ee53; source locator: benchmarks/mfass/README.md lines 107–110; benchmarks/mfass/pyproject.toml lines 1–33; title: Optionally run the resource pilot; shell: uv run --package mfass --extra dnabert2-pilot mfass-dnabert2-pilot; explanation: The pilot measures feasibility and validates checkpoint provenance; it does not calculate accuracy. The README gates any subsequent full encoder run on projected runtime and free disk; the full run is intentionally a separate decision.; source ids: run-doc-mfass-readme-md-68d1ee53; source locator: benchmarks/mfass/README.md lines 111–139; outputs: Validated assay-oriented cohort and corrected baseline results.; If selected, pilot runtime/memory/disk/failure measurements and model/code provenance checks; no pilot accuracy score.; limitations: The source reports a local run without paid APIs or cloud compute; that is not a runtime or cost guarantee for a different machine.; MFASS labels reflect exon recognition in an assay construct; they are not patient-RNA outcomes.; This command sequence does not rerun SpliceAI or Pangolin and must not relabel their previously published genomic-context predictions.; Do not use mfass-v1 baseline or split-cost results as current evidence.; source ids: run-doc-mfass-readme-md-68d1ee53; run-doc-mfass-pyproject-toml-68d1ee53; review: method: official_repository_review; date: 2026-09-17
run recipes
id: mfass-v2-rescore; protocol id: mfass-v2; version: f80cef7f818bec33e51b7f43ad499eb5078c8d87; title: MFASS v2: score existing predictions; purpose: rescore_predictions; summary: Evaluate keyed splice-disruption scores on the fixed MFASS v2 test split. Higher scores mean greater splice disruption.; inputs: Validated cohort.tsv; the canonical split is packaged with the library.; CSV/TSV columns id, score and optional reason; every test variant needs a score or an explicit unscored reason.; outputs: Local report.json with metrics, coverage, protocol and artifact hashes.; Local predictions.json and unscored.json; neither is submitted automatically.; requirements: data: Obtain the validated MFASS v2 cohort.tsv. Preparation checks its SHA-256 and the canonical 27,733-variant split from bee9133b83f3aedaf2bbb9013f1875515845607e.; weights: No model or weights are required to rescore supplied predictions.; licence: Runner code is MIT. The upstream MFASS data reuse licence is unreported; original data are not bundled.; software: Python 3.11 with the pinned rewirebench environment. Container runtimes are optional.; hardware: CPU scoring; minimum memory and runtime have not been established for this recipe.; instructions: runtime: python; title: Prepare and score local predictions; code: import rewirebench prepared = rewirebench.prepare( "mfass-v2", source="data/mfass", output="prepared" ) report = rewirebench.evaluate( prepared, "predictions/predictions.tsv", output="runs/mfass-rescore", model={"name": "My model", "training_overlap": "Unreported"}, ); status: source_reviewed_not_executed; source ids: runner-recipe-docs-sdk-md-f80cef7f; runner-recipe-docs-mfass-md-f80cef7f; runner-recipe-packages-rewirebench-src-rewirebench-protocols-mfass-py-f80cef7f; runner-recipe-packages-rewirebench-src-rewirebench-adapters-mfass-py-f80cef7f; source locator: docs/sdk.md: prepare/evaluate; docs/mfass.md: canonical cohort and keyed predictions; runtime: command_line; title: Prepare and score local predictions; code: git clone https://github.com/rewire-bio/rewire-benchmarks.git cd rewire-benchmarks git checkout --detach f80cef7f818bec33e51b7f43ad499eb5078c8d87 uv sync --locked --package rewirebench --python 3.11 uv run --package rewirebench rewirebench prepare mfass-v2 \ --source data/mfass --output prepared uv run --package rewirebench rewirebench evaluate \ --prepared prepared --predictions predictions/predictions.tsv \ --output runs/mfass-rescore --model-name "My model"; status: source_reviewed_not_executed; source ids: runner-recipe-docs-sdk-md-f80cef7f; runner-recipe-docs-mfass-md-f80cef7f; runner-recipe-packages-rewirebench-src-rewirebench-protocols-mfass-py-f80cef7f; runner-recipe-packages-rewirebench-src-rewirebench-adapters-mfass-py-f80cef7f; source locator: docs/sdk.md: command-line interface; docs/mfass.md: input preparation; runtime: podman; title: Score in a local Podman image; code: # Build the pinned core image as described in the linked HPC guide first. mkdir -p runs podman run --rm --userns=keep-id --network=none \ -v "$PWD/prepared:/prepared:ro" \ -v "$PWD/predictions:/predictions:ro" \ -v "$PWD/runs:/outputs:rw" \ localhost/rewirebench:core evaluate \ --prepared /prepared --predictions /predictions/predictions.tsv \ --output /outputs/mfass-v2-rescore; status: source_reviewed_not_executed; source ids: runner-recipe-docs-hpc-md-f80cef7f; runner-recipe-docs-sdk-md-f80cef7f; source locator: docs/hpc.md: Podman build, read-only input mounts and offline execution; docs/sdk.md: evaluate; runtime: apptainer; title: Score in an Apptainer image; code: # Build or obtain the pinned SIF as described in the linked HPC guide first. mkdir -p runs apptainer run --cleanenv --containall \ --bind "$PWD/prepared:/prepared:ro" \ --bind "$PWD/predictions:/predictions:ro" \ --bind "$PWD/runs:/outputs:rw" \ rewirebench-core.sif evaluate \ --prepared /prepared --predictions /predictions/predictions.tsv \ --output /outputs/mfass-v2-rescore; status: source_reviewed_not_executed; source ids: runner-recipe-docs-hpc-md-f80cef7f; runner-recipe-docs-sdk-md-f80cef7f; source locator: docs/hpc.md: Apptainer execution and source image identity; docs/sdk.md: evaluate; runtime: slurm; title: Submit the prepared scoring job to Slurm; code: #!/bin/bash set -euo pipefail # Set these for your cluster; no performance or resource estimate is implied. : "${REWIRE_ACCOUNT:?Set your Slurm account}" : "${REWIRE_PARTITION:?Set your Slurm partition}" : "${REWIRE_CPUS:?Set the requested CPU count}" : "${REWIRE_MEMORY:?Set the requested memory}" : "${REWIRE_WALLTIME:?Set the requested time limit}" : "${REWIRE_JOB_ROOT:?Set a shared absolute directory with prepared data and SIF}" export REWIRE_JOB_ROOT sbatch --account="$REWIRE_ACCOUNT" --partition="$REWIRE_PARTITION" \ --cpus-per-task="$REWIRE_CPUS" --mem="$REWIRE_MEMORY" \ --time="$REWIRE_WALLTIME" --export=ALL <<'REWIRE_JOB' #!/bin/bash set -euo pipefail cd "$REWIRE_JOB_ROOT" mkdir -p runs apptainer run --cleanenv --containall \ --bind "$PWD/prepared:/prepared:ro" \ --bind "$PWD/predictions:/predictions:ro" \ --bind "$PWD/runs:/outputs:rw" \ rewirebench-core.sif evaluate \ --prepared /prepared --predictions /predictions/predictions.tsv \ --output /outputs/mfass-v2-rescore REWIRE_JOB; status: source_reviewed_not_executed; source ids: runner-recipe-docs-hpc-md-f80cef7f; runner-recipe-docs-sdk-md-f80cef7f; source locator: docs/hpc.md: Slurm template, preparation outside compute nodes and site-specific resources; limitations: This preserves the corrected assay-oriented variant placement and canonical train/test identities.; Scores use held-out variants; missing predictions retain the original denominator.; Rescoring supplied SpliceAI or Pangolin predictions does not recreate their genomic-context inference.; Smoke inputs and partial coverage do not establish a full benchmark or reproduce a published result.; source ids: runner-recipe-docs-sdk-md-f80cef7f; runner-recipe-docs-mfass-md-f80cef7f; runner-recipe-packages-rewirebench-src-rewirebench-protocols-mfass-py-f80cef7f; runner-recipe-packages-rewirebench-src-rewirebench-adapters-mfass-py-f80cef7f; source locator: Pinned runner documentation, protocol implementation and adapter implementation; id: mfass-v2-baseline; protocol id: mfass-v2; version: f80cef7f818bec33e51b7f43ad499eb5078c8d87; title: MFASS v2: corrected k-mer baseline; purpose: generate_and_evaluate; summary: Fit the corrected positional and k-mer baseline using training labels only, then score held-out variants.; inputs: Validated MFASS v2 cohort.tsv.; outputs: Local report.json with metrics, coverage, protocol and artifact hashes.; Local predictions.json and unscored.json; neither is submitted automatically.; requirements: data: Obtain the validated MFASS v2 cohort.tsv. Preparation checks its SHA-256 and the canonical 27,733-variant split from bee9133b83f3aedaf2bbb9013f1875515845607e.; weights: No pretrained weights; the baseline is fitted on the fixed training split.; licence: Runner code is MIT. The upstream MFASS data reuse licence is unreported; original data are not bundled.; software: Python 3.11 with the pinned rewirebench environment. Container runtimes are optional.; hardware: CPU execution; a full baseline fit is a separate computation from score-only evaluation.; instructions: runtime: python; title: Fit the baseline and evaluate; code: import rewirebench from rewirebench.adapters.mfass import KmerBaseline prepared = rewirebench.prepare("mfass-v2", source="data/mfass", output="prepared") report = rewirebench.run( prepared, KmerBaseline(), output="runs/mfass-baseline", model={"name": "Corrected k-mer baseline", "training_overlap": "Canonical training split only"}, ); status: source_reviewed_not_executed; source ids: runner-recipe-docs-sdk-md-f80cef7f; runner-recipe-docs-mfass-md-f80cef7f; runner-recipe-packages-rewirebench-src-rewirebench-protocols-mfass-py-f80cef7f; runner-recipe-packages-rewirebench-src-rewirebench-adapters-mfass-py-f80cef7f; source locator: docs/mfass.md: baseline; adapters/mfass.py: KmerBaseline; runtime: command_line; title: Run the local baseline; code: git clone https://github.com/rewire-bio/rewire-benchmarks.git cd rewire-benchmarks git checkout --detach f80cef7f818bec33e51b7f43ad499eb5078c8d87 uv sync --locked --package rewirebench --python 3.11 uv run --package rewirebench rewirebench prepare mfass-v2 --source data/mfass --output prepared uv run --package rewirebench rewirebench run --prepared prepared \ --adapter rewirebench.adapters.mfass:KmerBaseline \ --model-name "Corrected k-mer baseline" --output runs/mfass-baseline; status: source_reviewed_not_executed; source ids: runner-recipe-docs-sdk-md-f80cef7f; runner-recipe-docs-mfass-md-f80cef7f; runner-recipe-packages-rewirebench-src-rewirebench-protocols-mfass-py-f80cef7f; runner-recipe-packages-rewirebench-src-rewirebench-adapters-mfass-py-f80cef7f; source locator: docs/sdk.md: run; docs/mfass.md: baseline; limitations: This preserves the corrected assay-oriented variant placement and canonical train/test identities.; Scores use held-out variants; missing predictions retain the original denominator.; Rescoring supplied SpliceAI or Pangolin predictions does not recreate their genomic-context inference.; Smoke inputs and partial coverage do not establish a full benchmark or reproduce a published result.; The input includes positional and conservation features, not sequence alone.; source ids: runner-recipe-docs-sdk-md-f80cef7f; runner-recipe-docs-mfass-md-f80cef7f; runner-recipe-packages-rewirebench-src-rewirebench-protocols-mfass-py-f80cef7f; runner-recipe-packages-rewirebench-src-rewirebench-adapters-mfass-py-f80cef7f; source locator: Pinned runner documentation, protocol implementation and adapter implementation; id: mfass-v2-dnabert2; protocol id: mfass-v2-frozen-encoder; version: f80cef7f818bec33e51b7f43ad499eb5078c8d87; title: MFASS v2: frozen DNABERT-2 example; purpose: generate_and_evaluate; summary: Embed validated reference/mutant assay pairs, then fit the fixed logistic-regression head on training rows only.; inputs: Validated MFASS v2 cohort.tsv.; A local DNABERT-2 checkpoint with the exact model and executable-code hashes documented by the adapter.; outputs: Local report.json with metrics, coverage, protocol and artifact hashes.; Local predictions.json and unscored.json; neither is submitted automatically.; requirements: data: Obtain the validated MFASS v2 cohort.tsv. Preparation checks its SHA-256 and the canonical 27,733-variant split from bee9133b83f3aedaf2bbb9013f1875515845607e.; weights: Obtain the pinned DNABERT-2 117M weights and reviewed code before execution. The adapter verifies hashes and loads local files only.; licence: Runner code is MIT. Check upstream DNABERT-2 code/weight terms and MFASS data conditions; weights and data are not bundled.; software: Python 3.11 with rewirebench[dnabert2]; use a separate environment from ESM.; hardware: CPU adapter. Full encoder execution can be expensive; start with an explicitly limited smoke input.; instructions: runtime: python; title: Run a limited frozen-encoder example; code: import rewirebench from rewirebench.adapters.mfass import DNABERT2 prepared = rewirebench.prepare( "mfass-v2-frozen-encoder", source="data/mfass", output="prepared-smoke", limit=200, ) report = rewirebench.run( prepared, DNABERT2(checkpoint="weights/dnabert2"), output="runs/dnabert2-smoke", model={"name": "DNABERT-2 frozen pair + logistic regression", "training_overlap": "Unreported"}, ); status: source_reviewed_not_executed; source ids: runner-recipe-docs-sdk-md-f80cef7f; runner-recipe-docs-mfass-md-f80cef7f; runner-recipe-packages-rewirebench-src-rewirebench-protocols-mfass-py-f80cef7f; runner-recipe-packages-rewirebench-src-rewirebench-adapters-mfass-py-f80cef7f; source locator: docs/mfass.md: frozen encoder example; adapters/mfass.py: DNABERT2; protocols/mfass.py: fit_embeddings; runtime: command_line; title: Run a limited frozen-encoder example; code: git clone https://github.com/rewire-bio/rewire-benchmarks.git cd rewire-benchmarks git checkout --detach f80cef7f818bec33e51b7f43ad499eb5078c8d87 uv sync --locked --package rewirebench --extra dnabert2 --python 3.11 uv run --package rewirebench --extra dnabert2 rewirebench prepare \ mfass-v2-frozen-encoder --source data/mfass --output prepared-smoke \ --options '{"limit":200}' uv run --package rewirebench --extra dnabert2 rewirebench run \ --prepared prepared-smoke --adapter rewirebench.adapters.mfass:DNABERT2 \ --adapter-options '{"checkpoint":"weights/dnabert2"}' \ --output runs/dnabert2-smoke --model-name "DNABERT-2 frozen pair + logistic regression"; status: source_reviewed_not_executed; source ids: runner-recipe-docs-sdk-md-f80cef7f; runner-recipe-docs-mfass-md-f80cef7f; runner-recipe-packages-rewirebench-src-rewirebench-protocols-mfass-py-f80cef7f; runner-recipe-packages-rewirebench-src-rewirebench-adapters-mfass-py-f80cef7f; source locator: docs/sdk.md: CLI; docs/mfass.md: local checkpoint and smoke example; limitations: This preserves the corrected assay-oriented variant placement and canonical train/test identities.; Scores use held-out variants; missing predictions retain the original denominator.; Rescoring supplied SpliceAI or Pangolin predictions does not recreate their genomic-context inference.; Smoke inputs and partial coverage do not establish a full benchmark or reproduce a published result.; The frozen-pair head is a evaluated pipeline, distinct from the bare DNABERT-2 model.; A smoke input may not contain both training classes; increase its size without selecting on test labels.; source ids: runner-recipe-docs-sdk-md-f80cef7f; runner-recipe-docs-mfass-md-f80cef7f; runner-recipe-packages-rewirebench-src-rewirebench-protocols-mfass-py-f80cef7f; runner-recipe-packages-rewirebench-src-rewirebench-adapters-mfass-py-f80cef7f; source locator: Pinned runner documentation, protocol implementation and adapter implementation
Related records

Suggest a correction