rewirebio.iobenchmarks
Dataset

CAGI5 saturation mutagenesis MPRA, four CREs (as used by Tang et al. 2025)

Saturation mutagenesis MPRA of regulatory elements from the CAGI5 regulation challenge: HepG2 elements printed as 'LDLT', SORT1 and F9 (Methods 'CAGI dataset'), and K562 element PKLR. 'LDLT' is not a gene symbol; it is presumably LDLR, but the source does not say so.

Research readiness

These checks assess whether the evidence supports a reproducible investigation. A source-checked score alone does not meet these requirements.

Release 2026-10-10-6e93f504adfc · Evidence verified: Not verified

Evidence incomplete

Replay metrics

Exact outcomes, predictions, identifiers and evaluator are connected.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • metric replay: verification is missing

Verified: Not verified

Evidence incomplete

Investigate discrepancies

Replay evidence includes annotations and an assessment of dependence. Unknown independence permits descriptive analysis only.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • metric replay: verification is missing
  • annotations: verification is missing
  • dependence: verification is missing

Verified: Not verified

Evidence incomplete

Run locally

A pinned recipe describes the inputs, environment and resource requirements.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • recipe pinned: verification is missing
  • resource estimate: verification is missing

Verified: Not verified

Evidence incomplete

Validate independently

Separate data and exposure records support an independent test.

Missing or unresolved evidence

  • No verified artifact manifest is linked to this exact record.
  • artifact hashes: verification is missing
  • join integrity: verification is missing
  • score semantics: verification is missing
  • independent validation: verification is missing
  • overlap checked: verification is missing

Verified: Not verified

Readiness describes the evidence in this release. Availability on your computer is checked separately when an investigation runs. Existing data exposure can prevent independent validation even when files are available.

Artifacts and reproduction

No verified artifact manifest is connected to this record yet. The gaps above identify what is needed before analysis can begin.

Read reviewed discrepancy investigations

Evaluation results

28 evaluations · 28 results. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: CNN-GPN, LentiMPRA-embedding (Tang et al. 2025)Protocol: CAGI5 saturation MPRA variant effect correlation, HepG2 (Tang et al. 2025 Table 1)
Dataset: CAGI5 saturation mutagenesis MPRA, four CREs (as used by Tang et al. 2025)
0.332 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CNN-GPN (LentiMPRA-embedding) on CAGI5 HepG2

regulatory-variant-20261009-protocol-tang2025-cagi5-hepg2

Aggregation: Not reported

Evaluating the representational power of pre-trained DNA language models for regulatory genomics · Table 1 row 'LentiMPRA-embedding' / 'CNN-GPN', column 'HepG2'
Configuration: CNN-NT, LentiMPRA-embedding (Tang et al. 2025)Protocol: CAGI5 saturation MPRA variant effect correlation, HepG2 (Tang et al. 2025 Table 1)
Dataset: CAGI5 saturation mutagenesis MPRA, four CREs (as used by Tang et al. 2025)
0.185 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CNN-NT (LentiMPRA-embedding) on CAGI5 HepG2

regulatory-variant-20261009-protocol-tang2025-cagi5-hepg2

Aggregation: Not reported

Evaluating the representational power of pre-trained DNA language models for regulatory genomics · Table 1 row 'LentiMPRA-embedding' / 'CNN-NT', column 'HepG2'
Configuration: CNN-SEI, LentiMPRA-embedding (Tang et al. 2025)Protocol: CAGI5 saturation MPRA variant effect correlation, HepG2 (Tang et al. 2025 Table 1)
Dataset: CAGI5 saturation mutagenesis MPRA, four CREs (as used by Tang et al. 2025)
0.579 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CNN-SEI (LentiMPRA-embedding) on CAGI5 HepG2

regulatory-variant-20261009-protocol-tang2025-cagi5-hepg2

Aggregation: Not reported

Evaluating the representational power of pre-trained DNA language models for regulatory genomics · Table 1 row 'LentiMPRA-embedding' / 'CNN-SEI', column 'HepG2'
Configuration: CNN, LentiMPRA-one-hot (Tang et al. 2025)Protocol: CAGI5 saturation MPRA variant effect correlation, HepG2 (Tang et al. 2025 Table 1)
Dataset: CAGI5 saturation mutagenesis MPRA, four CREs (as used by Tang et al. 2025)
0.324 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CNN (LentiMPRA-one-hot) on CAGI5 HepG2

regulatory-variant-20261009-protocol-tang2025-cagi5-hepg2

Aggregation: Not reported

Evaluating the representational power of pre-trained DNA language models for regulatory genomics · Table 1 row 'LentiMPRA-one-hot' / 'CNN', column 'HepG2'
Configuration: MPRAnn, LentiMPRA-one-hot (Tang et al. 2025)Protocol: CAGI5 saturation MPRA variant effect correlation, HepG2 (Tang et al. 2025 Table 1)
Dataset: CAGI5 saturation mutagenesis MPRA, four CREs (as used by Tang et al. 2025)
0.381 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

MPRAnn (LentiMPRA-one-hot) on CAGI5 HepG2

regulatory-variant-20261009-protocol-tang2025-cagi5-hepg2

Aggregation: Not reported

Evaluating the representational power of pre-trained DNA language models for regulatory genomics · Table 1 row 'LentiMPRA-one-hot' / 'MPRAnn', column 'HepG2'
Configuration: Residualbind, LentiMPRA-one-hot (Tang et al. 2025)Protocol: CAGI5 saturation MPRA variant effect correlation, HepG2 (Tang et al. 2025 Table 1)
Dataset: CAGI5 saturation mutagenesis MPRA, four CREs (as used by Tang et al. 2025)
0.485 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

Residualbind (LentiMPRA-one-hot) on CAGI5 HepG2

regulatory-variant-20261009-protocol-tang2025-cagi5-hepg2

Aggregation: Not reported

Evaluating the representational power of pre-trained DNA language models for regulatory genomics · Table 1 row 'LentiMPRA-one-hot' / 'Residualbind', column 'HepG2'
Configuration: GPN (human), Self-supervised pre-training (Tang et al. 2025)Protocol: CAGI5 saturation MPRA variant effect correlation, HepG2 (Tang et al. 2025 Table 1)
Dataset: CAGI5 saturation mutagenesis MPRA, four CREs (as used by Tang et al. 2025)
0.002 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

GPN (human) (Self-supervised pre-training) on CAGI5 HepG2

regulatory-variant-20261009-protocol-tang2025-cagi5-hepg2

Aggregation: Not reported

Evaluating the representational power of pre-trained DNA language models for regulatory genomics · Table 1 row 'Self-supervised pre-training' / 'GPN (human)', column 'HepG2'
Configuration: HyenaDNA, Self-supervised pre-training (Tang et al. 2025)Protocol: CAGI5 saturation MPRA variant effect correlation, HepG2 (Tang et al. 2025 Table 1)
Dataset: CAGI5 saturation mutagenesis MPRA, four CREs (as used by Tang et al. 2025)
0.064 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

HyenaDNA (Self-supervised pre-training) on CAGI5 HepG2

regulatory-variant-20261009-protocol-tang2025-cagi5-hepg2

Aggregation: Not reported

Evaluating the representational power of pre-trained DNA language models for regulatory genomics · Table 1 row 'Self-supervised pre-training' / 'HyenaDNA', column 'HepG2'
Configuration: NT (2B51000G), Self-supervised pre-training (Tang et al. 2025)Protocol: CAGI5 saturation MPRA variant effect correlation, HepG2 (Tang et al. 2025 Table 1)
Dataset: CAGI5 saturation mutagenesis MPRA, four CREs (as used by Tang et al. 2025)
0.125 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

NT (2B51000G) (Self-supervised pre-training) on CAGI5 HepG2

regulatory-variant-20261009-protocol-tang2025-cagi5-hepg2

Aggregation: Not reported

Evaluating the representational power of pre-trained DNA language models for regulatory genomics · Table 1 row 'Self-supervised pre-training' / 'NT (2B51000G)', column 'HepG2'
Configuration: NT (2B5Species), Self-supervised pre-training (Tang et al. 2025)Protocol: CAGI5 saturation MPRA variant effect correlation, HepG2 (Tang et al. 2025 Table 1)
Dataset: CAGI5 saturation mutagenesis MPRA, four CREs (as used by Tang et al. 2025)
0.112 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

NT (2B5Species) (Self-supervised pre-training) on CAGI5 HepG2

regulatory-variant-20261009-protocol-tang2025-cagi5-hepg2

Aggregation: Not reported

Evaluating the representational power of pre-trained DNA language models for regulatory genomics · Table 1 row 'Self-supervised pre-training' / 'NT (2B5Species)', column 'HepG2'
Configuration: NT (500M1000G), Self-supervised pre-training (Tang et al. 2025)Protocol: CAGI5 saturation MPRA variant effect correlation, HepG2 (Tang et al. 2025 Table 1)
Dataset: CAGI5 saturation mutagenesis MPRA, four CREs (as used by Tang et al. 2025)
0.041 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

NT (500M1000G) (Self-supervised pre-training) on CAGI5 HepG2

regulatory-variant-20261009-protocol-tang2025-cagi5-hepg2

Aggregation: Not reported

Evaluating the representational power of pre-trained DNA language models for regulatory genomics · Table 1 row 'Self-supervised pre-training' / 'NT (500M1000G)', column 'HepG2'
Configuration: NT (500MHuman), Self-supervised pre-training (Tang et al. 2025)Protocol: CAGI5 saturation MPRA variant effect correlation, HepG2 (Tang et al. 2025 Table 1)
Dataset: CAGI5 saturation mutagenesis MPRA, four CREs (as used by Tang et al. 2025)
0.02 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

NT (500MHuman) (Self-supervised pre-training) on CAGI5 HepG2

regulatory-variant-20261009-protocol-tang2025-cagi5-hepg2

Aggregation: Not reported

Evaluating the representational power of pre-trained DNA language models for regulatory genomics · Table 1 row 'Self-supervised pre-training' / 'NT (500MHuman)', column 'HepG2'
Configuration: Enformer (DNase), Supervised one-hot (Tang et al. 2025)Protocol: CAGI5 saturation MPRA variant effect correlation, HepG2 (Tang et al. 2025 Table 1)
Dataset: CAGI5 saturation mutagenesis MPRA, four CREs (as used by Tang et al. 2025)
0.51 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

Enformer (DNase) (Supervised one-hot) on CAGI5 HepG2

regulatory-variant-20261009-protocol-tang2025-cagi5-hepg2

Aggregation: Not reported

Evaluating the representational power of pre-trained DNA language models for regulatory genomics · Table 1 row 'Supervised one-hot' / 'Enformer (DNase)', column 'HepG2'
Configuration: SEI, Supervised one-hot (Tang et al. 2025)Protocol: CAGI5 saturation MPRA variant effect correlation, HepG2 (Tang et al. 2025 Table 1)
Dataset: CAGI5 saturation mutagenesis MPRA, four CREs (as used by Tang et al. 2025)
0.545 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

SEI (Supervised one-hot) on CAGI5 HepG2

regulatory-variant-20261009-protocol-tang2025-cagi5-hepg2

Aggregation: Not reported

Evaluating the representational power of pre-trained DNA language models for regulatory genomics · Table 1 row 'Supervised one-hot' / 'SEI', column 'HepG2'
Configuration: CNN-GPN, LentiMPRA-embedding (Tang et al. 2025)Protocol: CAGI5 saturation MPRA variant effect correlation, K562 (Tang et al. 2025 Table 1)
Dataset: CAGI5 saturation mutagenesis MPRA, four CREs (as used by Tang et al. 2025)
0.437 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CNN-GPN (LentiMPRA-embedding) on CAGI5 K562

regulatory-variant-20261009-protocol-tang2025-cagi5-k562

Aggregation: Not reported

Evaluating the representational power of pre-trained DNA language models for regulatory genomics · Table 1 row 'LentiMPRA-embedding' / 'CNN-GPN', column 'K562'
Configuration: CNN-NT, LentiMPRA-embedding (Tang et al. 2025)Protocol: CAGI5 saturation MPRA variant effect correlation, K562 (Tang et al. 2025 Table 1)
Dataset: CAGI5 saturation mutagenesis MPRA, four CREs (as used by Tang et al. 2025)
0.198 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CNN-NT (LentiMPRA-embedding) on CAGI5 K562

regulatory-variant-20261009-protocol-tang2025-cagi5-k562

Aggregation: Not reported

Evaluating the representational power of pre-trained DNA language models for regulatory genomics · Table 1 row 'LentiMPRA-embedding' / 'CNN-NT', column 'K562'
Configuration: CNN-SEI, LentiMPRA-embedding (Tang et al. 2025)Protocol: CAGI5 saturation MPRA variant effect correlation, K562 (Tang et al. 2025 Table 1)
Dataset: CAGI5 saturation mutagenesis MPRA, four CREs (as used by Tang et al. 2025)
0.701 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CNN-SEI (LentiMPRA-embedding) on CAGI5 K562

regulatory-variant-20261009-protocol-tang2025-cagi5-k562

Aggregation: Not reported

Evaluating the representational power of pre-trained DNA language models for regulatory genomics · Table 1 row 'LentiMPRA-embedding' / 'CNN-SEI', column 'K562'
Configuration: CNN, LentiMPRA-one-hot (Tang et al. 2025)Protocol: CAGI5 saturation MPRA variant effect correlation, K562 (Tang et al. 2025 Table 1)
Dataset: CAGI5 saturation mutagenesis MPRA, four CREs (as used by Tang et al. 2025)
0.365 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

CNN (LentiMPRA-one-hot) on CAGI5 K562

regulatory-variant-20261009-protocol-tang2025-cagi5-k562

Aggregation: Not reported

Evaluating the representational power of pre-trained DNA language models for regulatory genomics · Table 1 row 'LentiMPRA-one-hot' / 'CNN', column 'K562'
Configuration: MPRAnn, LentiMPRA-one-hot (Tang et al. 2025)Protocol: CAGI5 saturation MPRA variant effect correlation, K562 (Tang et al. 2025 Table 1)
Dataset: CAGI5 saturation mutagenesis MPRA, four CREs (as used by Tang et al. 2025)
0.437 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

MPRAnn (LentiMPRA-one-hot) on CAGI5 K562

regulatory-variant-20261009-protocol-tang2025-cagi5-k562

Aggregation: Not reported

Evaluating the representational power of pre-trained DNA language models for regulatory genomics · Table 1 row 'LentiMPRA-one-hot' / 'MPRAnn', column 'K562'
Configuration: Residualbind, LentiMPRA-one-hot (Tang et al. 2025)Protocol: CAGI5 saturation MPRA variant effect correlation, K562 (Tang et al. 2025 Table 1)
Dataset: CAGI5 saturation mutagenesis MPRA, four CREs (as used by Tang et al. 2025)
0.601 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

Residualbind (LentiMPRA-one-hot) on CAGI5 K562

regulatory-variant-20261009-protocol-tang2025-cagi5-k562

Aggregation: Not reported

Evaluating the representational power of pre-trained DNA language models for regulatory genomics · Table 1 row 'LentiMPRA-one-hot' / 'Residualbind', column 'K562'
Configuration: GPN (human), Self-supervised pre-training (Tang et al. 2025)Protocol: CAGI5 saturation MPRA variant effect correlation, K562 (Tang et al. 2025 Table 1)
Dataset: CAGI5 saturation mutagenesis MPRA, four CREs (as used by Tang et al. 2025)
0.037 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

GPN (human) (Self-supervised pre-training) on CAGI5 K562

regulatory-variant-20261009-protocol-tang2025-cagi5-k562

Aggregation: Not reported

Evaluating the representational power of pre-trained DNA language models for regulatory genomics · Table 1 row 'Self-supervised pre-training' / 'GPN (human)', column 'K562'
Configuration: HyenaDNA, Self-supervised pre-training (Tang et al. 2025)Protocol: CAGI5 saturation MPRA variant effect correlation, K562 (Tang et al. 2025 Table 1)
Dataset: CAGI5 saturation mutagenesis MPRA, four CREs (as used by Tang et al. 2025)
0.021 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

HyenaDNA (Self-supervised pre-training) on CAGI5 K562

regulatory-variant-20261009-protocol-tang2025-cagi5-k562

Aggregation: Not reported

Evaluating the representational power of pre-trained DNA language models for regulatory genomics · Table 1 row 'Self-supervised pre-training' / 'HyenaDNA', column 'K562'
Configuration: NT (2B51000G), Self-supervised pre-training (Tang et al. 2025)Protocol: CAGI5 saturation MPRA variant effect correlation, K562 (Tang et al. 2025 Table 1)
Dataset: CAGI5 saturation mutagenesis MPRA, four CREs (as used by Tang et al. 2025)
0.007 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

NT (2B51000G) (Self-supervised pre-training) on CAGI5 K562

regulatory-variant-20261009-protocol-tang2025-cagi5-k562

Aggregation: Not reported

Evaluating the representational power of pre-trained DNA language models for regulatory genomics · Table 1 row 'Self-supervised pre-training' / 'NT (2B51000G)', column 'K562'
Configuration: NT (2B5Species), Self-supervised pre-training (Tang et al. 2025)Protocol: CAGI5 saturation MPRA variant effect correlation, K562 (Tang et al. 2025 Table 1)
Dataset: CAGI5 saturation mutagenesis MPRA, four CREs (as used by Tang et al. 2025)
0.135 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

NT (2B5Species) (Self-supervised pre-training) on CAGI5 K562

regulatory-variant-20261009-protocol-tang2025-cagi5-k562

Aggregation: Not reported

Evaluating the representational power of pre-trained DNA language models for regulatory genomics · Table 1 row 'Self-supervised pre-training' / 'NT (2B5Species)', column 'K562'
Configuration: NT (500M1000G), Self-supervised pre-training (Tang et al. 2025)Protocol: CAGI5 saturation MPRA variant effect correlation, K562 (Tang et al. 2025 Table 1)
Dataset: CAGI5 saturation mutagenesis MPRA, four CREs (as used by Tang et al. 2025)
0.068 pearson-correlation
unitless · higher

Uncertainty: Not reported by the source

Coverage: Not reported scored / Not reported eligible

Independent external evaluation · Source checked
Methods, coverage and source

NT (500M1000G) (Self-supervised pre-training) on CAGI5 K562

regulatory-variant-20261009-protocol-tang2025-cagi5-k562

Aggregation: Not reported

Evaluating the representational power of pre-trained DNA language models for regulatory genomics · Table 1 row 'Self-supervised pre-training' / 'NT (500M1000G)', column 'K562'

Source checking is not independent reproduction. Release 2026-10-10-6e93f504adfc.

Dataset and evaluation context

A dataset supplies biological observations. The evaluation protocol defines how those observations are split, used and scored.

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

6 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-10-10-6e93f504adfc
Property and statementOriginal source and locationReview and provenance
attributes.population
Single-nucleotide variants in 230 nt windows centred on four CREs
Context-only references
Evaluating the representational power of pre-trained DNA language models for regulatory genomics

Original source ↗

Methods 'CAGI dataset'

Version: Genome Biology 26:203, published 2025-07-14; PMC12261763 full-text XML
Retrieved: 2026-10-09T21:04:28Z

not individually reviewed

No individual claim review recorded

Audit details

Field: attributes.population

Source artifact SHA-256: f6925cc2d93d0694ccc689970207c2df9b7b0d2562d276ede99f0c78d413b08b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

attributes.source_locator
Methods 'CAGI dataset'
Context-only references
Evaluating the representational power of pre-trained DNA language models for regulatory genomics

Original source ↗

Methods 'CAGI dataset'

Version: Genome Biology 26:203, published 2025-07-14; PMC12261763 full-text XML
Retrieved: 2026-10-09T21:04:28Z

not individually reviewed

No individual claim review recorded

Audit details

Field: attributes.source_locator

Source artifact SHA-256: f6925cc2d93d0694ccc689970207c2df9b7b0d2562d276ede99f0c78d413b08b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

attributes.split
Zero-shot test; not used for training
Context-only references
Evaluating the representational power of pre-trained DNA language models for regulatory genomics

Original source ↗

Methods 'CAGI dataset'

Version: Genome Biology 26:203, published 2025-07-14; PMC12261763 full-text XML
Retrieved: 2026-10-09T21:04:28Z

not individually reviewed

No individual claim review recorded

Audit details

Field: attributes.split

Source artifact SHA-256: f6925cc2d93d0694ccc689970207c2df9b7b0d2562d276ede99f0c78d413b08b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

attributes.version
CAGI5 challenge data (as used)
Context-only references
Evaluating the representational power of pre-trained DNA language models for regulatory genomics

Original source ↗

Methods 'CAGI dataset'

Version: Genome Biology 26:203, published 2025-07-14; PMC12261763 full-text XML
Retrieved: 2026-10-09T21:04:28Z

not individually reviewed

No individual claim review recorded

Audit details

Field: attributes.version

Source artifact SHA-256: f6925cc2d93d0694ccc689970207c2df9b7b0d2562d276ede99f0c78d413b08b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

description
Saturation mutagenesis MPRA of regulatory elements from the CAGI5 regulation challenge: HepG2 elements printed as 'LDLT', SORT1 and F9 (Methods 'CAGI dataset'), and K562 element PKLR. 'LDLT' is not a gene symbol; it is presumably LDLR, but the source does not say so.
Context-only references
Evaluating the representational power of pre-trained DNA language models for regulatory genomics

Original source ↗

Methods 'CAGI dataset'

Version: Genome Biology 26:203, published 2025-07-14; PMC12261763 full-text XML
Retrieved: 2026-10-09T21:04:28Z

not individually reviewed

No individual claim review recorded

Audit details

Field: description

Source artifact SHA-256: f6925cc2d93d0694ccc689970207c2df9b7b0d2562d276ede99f0c78d413b08b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

name
CAGI5 saturation mutagenesis MPRA, four CREs (as used by Tang et al. 2025)
Context-only references
Evaluating the representational power of pre-trained DNA language models for regulatory genomics

Original source ↗

Methods 'CAGI dataset'

Version: Genome Biology 26:203, published 2025-07-14; PMC12261763 full-text XML
Retrieved: 2026-10-09T21:04:28Z

not individually reviewed

No individual claim review recorded

Audit details

Field: name

Source artifact SHA-256: f6925cc2d93d0694ccc689970207c2df9b7b0d2562d276ede99f0c78d413b08b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

Release 2026-10-10-6e93f504adfc · Record review: source checked

1 source records and release historyDownload this release (gzip)
Technical metadata and extraction receipts

Stable ID: regulatory-variant-20261009-data-tang2025-cagi5-saturation-mpra

areas
dna-genomes
contexts
research
version
CAGI5 challenge data (as used)
population
Single-nucleotide variants in 230 nt windows centred on four CREs
split
Zero-shot test; not used for training
source locator
Methods 'CAGI dataset'
Related records

Suggest a correction