rewire.itbenchmarks
Configuration

ENBED

ENBED is a byte-level encoder–decoder transformer for genomic sequence representation and sequence-to-sequence tasks.

SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · 2 Methods/2.5 Application domains (paragraph 1); 5 Discussion (paragraph 1)

1 evaluation · 1 result

How it worksEvaluated procedure (conceptual)
Evaluated procedure (conceptual)1. DNA sequences at single-byte nucleotide resolution. Then: 2. ENBED. Then: 3. Task-specific classifications or generated DNA sequencesEvaluated procedure (conceptual)1. DNA sequences at single-byte nucleotide resolution. Then: 2. ENBED. Then: 3. Task-specific classifications or generated DNA sequencesEvaluated procedure (conceptual)1. DNA sequences at single-byte nucleotide resolution. Then: 2. ENBED. Then: 3. Task-specific classifications or generated DNA sequences

Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings.

SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · 1 Introduction/1.2 Our contributions (paragraph 1); 4 Ablation studies/4.1 Encoder–decoder architecture (paragraph 2)

Overview

limited source coverage · Automated source review, 2026-09-16. All specifications and missing details

Evaluations and results

1 evaluation · 1 result. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: ENBEDTask: Enhancer classification
Dataset: Genomic Benchmarks Mouse Enhancers
90.3 Accuracy
% · unknown

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · Source checked
Methods, coverage and source

ENBED: Enhancer classification

Reported Genomic Benchmarks classification accuracy.

Aggregation: Not reported

Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Table 2, Mouse Enhancers row, ENBED column

Source checking is not independent reproduction. Release 2026-09-29-06401fd5b220.

Use this model

How it works, versions and access

How it works

How the evaluated method works

Byte-level nucleotide inputs feed encoder and decoder transformer blocks with a subquadratic attention implementation. Masked-language pretraining precedes task-specific adaptation.

SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · 1 Introduction/1.2 Our contributions (paragraph 1); 4 Ablation studies/4.1 Encoder–decoder architecture (paragraph 2)
What was evaluated

The linked evaluation record identifies ENBED: Enhancer classification. Its dataset, split, adaptation and evidence origin remain attached to the reported results.

SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · The named evaluation’s methods and comparison table; exact preserved evaluation IDs: evaluation-lit-003
Strengths, limitations and unresolved questions

Strengths and limitations

Profile review details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Stable record: reported-model-8db190bee6aae5

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Inputs, outputs and configuration
PropertyDescription and evidence
Model typeTransformer representation pipeline; this record is the paper-specific evaluated configuration.
SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · 1 Introduction/1.2 Our contributions (paragraph 1); 4 Ablation studies/4.1 Encoder–decoder architecture (paragraph 2)
Architecture / procedureByte-level nucleotide inputs feed encoder and decoder transformer blocks with a subquadratic attention implementation. Masked-language pretraining precedes task-specific adaptation.
SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · 1 Introduction/1.2 Our contributions (paragraph 1); 4 Ablation studies/4.1 Encoder–decoder architecture (paragraph 2)
Biological inputsDNA sequences at single-byte nucleotide resolution
SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · 1 Introduction/1.2 Our contributions (paragraph 2); 5 Discussion (paragraph 1)
OutputsTask-specific classifications or generated DNA sequences
SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · 2 Methods/2.4 Applications of foundation models using transfer learning/2.4.2 Fine-tuning for downstream tasks (paragraph 1); 2 Methods/2.5 Application domains/2.5.1 Genomic benchmarks (paragraph 1)
Parameters1.2 billion trainable parameters in the full encoder–decoder model.
SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · 4 Ablation studies (paragraph 1); 4 Ablation studies/4.1 Encoder–decoder architecture (paragraph 1)
Known versions / configurationENBED is the comparison-table label; that label does not specify an immutable weight revision. · Not reported in inspected sources
SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Model identification in the comparison table and corresponding Methods; immutable checkpoint revision is not supplied by the table label.
Training data / fittingReference-genome sequences; the GRCh38-labelled row is a distinct human-reference configuration.
SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · 1 Introduction/1.2 Our contributions/1.2.1 Evaluation of performance on genomic benchmark datasets (paragraph 1); 3 Results/3.1 ENBED outperforms state-of-the-art models on GB datasets (paragraph 2)
Context limits16,384 input/output tokens using local sliding-window plus global attention. The 512-token value in Methods describes the dense-attention hardware baseline, not ENBED’s final context.
SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · 2 Methods/2.3 Attention (paragraph 2); 2 Methods/2.3 Attention/2.3.1 Sliding-window attention (paragraph 1)
AccessThe authors provide implementation code at https://github.itap.purdue.edu/Clan-labs/ENBED and model weights through https://huggingface.co/malusare. A table-specific checkpoint hash is not supplied by these account-level links.
SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Data availability
Code licenceNo explicit code licence was established from the paper’s availability statement and inspected repository-root documentation. · Not reported in inspected sources
SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Abstract (paragraph 1); 2 Methods/2.4 Applications of foundation models using transfer learning/2.4.1 Building the foundation model (paragraph 1)
Weights licenceThe inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately. · Not reported in inspected sources
SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Abstract (paragraph 1); 2 Methods/2.1 Encoder–decoder model architecture (paragraph 1)

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

19 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-29-06401fd5b220
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings.
Individual claims
Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision

Original source ↗

1 Introduction/1.2 Our contributions (paragraph 1); 4 Ablation studies/4.1 Encoder–decoder architecture (paragraph 2)

Version: version of record
Retrieved: 2026-09-16T10:33:35.378Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: e95d4be70d32e61af5a92eda8ea66f25a2cc629e3f83e7b5241b13cde8bdb83b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps
  • DNA sequences at single-byte nucleotide resolution
  • ENBED
  • Task-specific classifications or generated DNA sequences
Individual claims
Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision

Original source ↗

1 Introduction/1.2 Our contributions (paragraph 1); 4 Ablation studies/4.1 Encoder–decoder architecture (paragraph 2)

Version: version of record
Retrieved: 2026-09-16T10:33:35.378Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: e95d4be70d32e61af5a92eda8ea66f25a2cc629e3f83e7b5241b13cde8bdb83b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title
Evaluated procedure (conceptual)
Individual claims
Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision

Original source ↗

1 Introduction/1.2 Our contributions (paragraph 1); 4 Ablation studies/4.1 Encoder–decoder architecture (paragraph 2)

Version: version of record
Retrieved: 2026-09-16T10:33:35.378Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: e95d4be70d32e61af5a92eda8ea66f25a2cc629e3f83e7b5241b13cde8bdb83b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Model type
Transformer representation pipeline; this record is the paper-specific evaluated configuration.
Individual claims
Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision

Original source ↗

1 Introduction/1.2 Our contributions (paragraph 1); 4 Ablation studies/4.1 Encoder–decoder architecture (paragraph 2)

Version: version of record
Retrieved: 2026-09-16T10:33:35.378Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: e95d4be70d32e61af5a92eda8ea66f25a2cc629e3f83e7b5241b13cde8bdb83b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Architecture / procedure
Byte-level nucleotide inputs feed encoder and decoder transformer blocks with a subquadratic attention implementation. Masked-language pretraining precedes task-specific adaptation.
Individual claims
Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision

Original source ↗

1 Introduction/1.2 Our contributions (paragraph 1); 4 Ablation studies/4.1 Encoder–decoder architecture (paragraph 2)

Version: version of record
Retrieved: 2026-09-16T10:33:35.378Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: e95d4be70d32e61af5a92eda8ea66f25a2cc629e3f83e7b5241b13cde8bdb83b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Weights licence
The inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately.
Individual claims
Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision

Original source ↗

Abstract (paragraph 1); 2 Methods/2.1 Encoder–decoder model architecture (paragraph 1)

Version: version of record
Retrieved: 2026-09-16T10:33:35.378Z

unreported

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: e95d4be70d32e61af5a92eda8ea66f25a2cc629e3f83e7b5241b13cde8bdb83b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Biological inputs
DNA sequences at single-byte nucleotide resolution
Individual claims
Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision

Original source ↗

1 Introduction/1.2 Our contributions (paragraph 2); 5 Discussion (paragraph 1)

Version: version of record
Retrieved: 2026-09-16T10:33:35.378Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: e95d4be70d32e61af5a92eda8ea66f25a2cc629e3f83e7b5241b13cde8bdb83b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Outputs
Task-specific classifications or generated DNA sequences
Individual claims
Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision

Original source ↗

2 Methods/2.4 Applications of foundation models using transfer learning/2.4.2 Fine-tuning for downstream tasks (paragraph 1); 2 Methods/2.5 Application domains/2.5.1 Genomic benchmarks (paragraph 1)

Version: version of record
Retrieved: 2026-09-16T10:33:35.378Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.3.value

Source artifact SHA-256: e95d4be70d32e61af5a92eda8ea66f25a2cc629e3f83e7b5241b13cde8bdb83b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Parameters
1.2 billion trainable parameters in the full encoder–decoder model.
Individual claims
Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision

Original source ↗

4 Ablation studies (paragraph 1); 4 Ablation studies/4.1 Encoder–decoder architecture (paragraph 1)

Version: version of record
Retrieved: 2026-09-16T10:33:35.378Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: e95d4be70d32e61af5a92eda8ea66f25a2cc629e3f83e7b5241b13cde8bdb83b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Known versions / configuration
ENBED is the comparison-table label; that label does not specify an immutable weight revision.
Individual claims
Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision

Original source ↗

Model identification in the comparison table and corresponding Methods; immutable checkpoint revision is not supplied by the table label.

Version: version of record
Retrieved: 2026-09-16T10:33:35.378Z

unreported

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.5.value

Source artifact SHA-256: e95d4be70d32e61af5a92eda8ea66f25a2cc629e3f83e7b5241b13cde8bdb83b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-29-06401fd5b220 · Record review: needs review

1 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: reported-model-8db190bee6aae5

areas
dna-genomes
entity level
method
version
Not reported
reported name
ENBED
historical missing metadata
version: not_reported_in_legacy_extract; checkpoint revision: not_reported_in_legacy_extract; training data: not_reported_in_legacy_extract; licence: not_reported_in_legacy_extract
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
legacy kinds
model
entity classification
review date: 2026-09-17; rationale: This source-scoped entry preserves the method/configuration actually named in an evaluation. It is neither a global family identity nor proof of an immutable checkpoint; the linked evaluation retains adaptation, fitting and scoring details.; source ids: enbed-2024; source locator: 1 Introduction/1.2 Our contributions (paragraph 1); 4 Ablation studies/4.1 Encoder–decoder architecture (paragraph 2) | 2 Methods/2.5 Application domains (paragraph 1); 5 Discussion (paragraph 1); ambiguities: Configuration means the source-labelled evaluated identity. It does not establish missing checkpoint hashes, default settings or equivalence to same-named records in other papers.
Related records

Suggest a correction