| Configuration: GENA-LM-Large-T2T | Task: GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe Dataset subset: GENEB representative task subset (GENEB split) | 0.53 macro_mcc correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceGENA-LM-Large-T2T on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Aggregation: Not reported GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GENA-LM-Large-T2T), column(Linear MCC) |
|---|
| Configuration: GENA-LM-Large-T2T | Task: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Dataset subset: GENEB representative task subset (GENEB split) | 0.535 macro_mcc correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceGENA-LM-Large-T2T on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Aggregation: Not reported GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GENA-LM-Large-T2T), column(MLP MCC) |
|---|
| Configuration: GENERator-Eukaryote-1.2B | Task: GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe Dataset subset: GENEB representative task subset (GENEB split) | 0.579 macro_mcc correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceGENERator-Eukaryote-1.2B on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Aggregation: Not reported GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GENERator-Eukaryote-1.2B), column(Linear MCC) |
|---|
| Configuration: GENERator-Eukaryote-1.2B | Task: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Dataset subset: GENEB representative task subset (GENEB split) | 0.595 macro_mcc correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceGENERator-Eukaryote-1.2B on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Aggregation: Not reported GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GENERator-Eukaryote-1.2B), column(MLP MCC) |
|---|
| Configuration: GENERator-Eukaryote-3B | Task: GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe Dataset subset: GENEB representative task subset (GENEB split) | 0.605 macro_mcc correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceGENERator-Eukaryote-3B on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Aggregation: Not reported GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GENERator-Eukaryote-3B), column(Linear MCC) |
|---|
| Configuration: GENERator-Eukaryote-3B | Task: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Dataset subset: GENEB representative task subset (GENEB split) | 0.609 macro_mcc correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceGENERator-Eukaryote-3B on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Aggregation: Not reported GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GENERator-Eukaryote-3B), column(MLP MCC) |
|---|
| Configuration: GenomeOcean-4B | Task: GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe Dataset subset: GENEB representative task subset (GENEB split) | 0.552 macro_mcc correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceGenomeOcean-4B on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Aggregation: Not reported GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GenomeOcean-4B), column(Linear MCC) |
|---|
| Configuration: GenomeOcean-4B | Task: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Dataset subset: GENEB representative task subset (GENEB split) | 0.552 macro_mcc correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceGenomeOcean-4B on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Aggregation: Not reported GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GenomeOcean-4B), column(MLP MCC) |
|---|
| Configuration: GenomeOcean-500M | Task: GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe Dataset subset: GENEB representative task subset (GENEB split) | 0.536 macro_mcc correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceGenomeOcean-500M on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Aggregation: Not reported GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GenomeOcean-500M), column(Linear MCC) |
|---|
| Configuration: GenomeOcean-500M | Task: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Dataset subset: GENEB representative task subset (GENEB split) | 0.535 macro_mcc correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceGenomeOcean-500M on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Aggregation: Not reported GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GenomeOcean-500M), column(MLP MCC) |
|---|
| Configuration: GROVER | Task: GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe Dataset subset: GENEB representative task subset (GENEB split) | 0.466 macro_mcc correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceGROVER on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Aggregation: Not reported GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GROVER), column(Linear MCC) |
|---|
| Configuration: GROVER | Task: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Dataset subset: GENEB representative task subset (GENEB split) | 0.477 macro_mcc correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceGROVER on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Aggregation: Not reported GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GROVER), column(MLP MCC) |
|---|
| Configuration: HyenaDNA-Large-1M | Task: GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe Dataset subset: GENEB representative task subset (GENEB split) | 0.427 macro_mcc correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceHyenaDNA-Large-1M on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Aggregation: Not reported GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(HyenaDNA-Large-1M), column(Linear MCC) |
|---|
| Configuration: HyenaDNA-Large-1M | Task: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Dataset subset: GENEB representative task subset (GENEB split) | 0.479 macro_mcc correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceHyenaDNA-Large-1M on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Aggregation: Not reported GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(HyenaDNA-Large-1M), column(MLP MCC) |
|---|
| Configuration: LucaOne | Task: GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe Dataset subset: GENEB representative task subset (GENEB split) | 0.573 macro_mcc correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceLucaOne on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Aggregation: Not reported GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(LucaOne), column(Linear MCC) |
|---|
| Configuration: LucaOne | Task: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Dataset subset: GENEB representative task subset (GENEB split) | 0.6 macro_mcc correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceLucaOne on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Aggregation: Not reported GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(LucaOne), column(MLP MCC) |
|---|
| Configuration: MutBERT | Task: GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe Dataset subset: GENEB representative task subset (GENEB split) | 0.516 macro_mcc correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceMutBERT on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Aggregation: Not reported GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(MutBERT), column(Linear MCC) |
|---|
| Configuration: MutBERT | Task: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Dataset subset: GENEB representative task subset (GENEB split) | 0.517 macro_mcc correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceMutBERT on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Aggregation: Not reported GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(MutBERT), column(MLP MCC) |
|---|
| Configuration: NT-v2-50M-MS | Task: GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe Dataset subset: GENEB representative task subset (GENEB split) | 0.511 macro_mcc correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNT-v2-50M-MS on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Aggregation: Not reported GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(NT-v2-50M-MS), column(Linear MCC) |
|---|
| Configuration: NT-v2-50M-MS | Task: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Dataset subset: GENEB representative task subset (GENEB split) | 0.521 macro_mcc correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceNT-v2-50M-MS on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Aggregation: Not reported GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(NT-v2-50M-MS), column(MLP MCC) |
|---|
| Configuration: Omni-DNA-1B | Task: GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe Dataset subset: GENEB representative task subset (GENEB split) | 0.55 macro_mcc correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceOmni-DNA-1B on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Aggregation: Not reported GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(Omni-DNA-1B), column(Linear MCC) |
|---|
| Configuration: Omni-DNA-1B | Task: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Dataset subset: GENEB representative task subset (GENEB split) | 0.542 macro_mcc correlation · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · Source checkedMethods, coverage and sourceOmni-DNA-1B on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Aggregation: Not reported GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(Omni-DNA-1B), column(MLP MCC) |
|---|