A Benchmark of Evo2 Genomic AI Models for Efficient and Practical Deployment

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

The rapid advancement of DNA foundation language models has brought about a transformative shift in genomics, allowing for the deciphering of intricate patterns and regulatory mechanisms embedded within DNA sequences. The genomic foundation model Evo2 demonstrates remarkable capabilities in decoding DNA functional patterns through cross-species pretraining. However, despite the great potential of Evo2 in basic genomics research, there is currently no clear and systematic guidance on its specific application scenarios, performance, and optimization directions in the field of tumor genomics, and its performance dependency on specialized hardware (such as FP8 precision on H800 GPUs) has not been empirically benchmarked. Here, we present a focused validation of Evo2 using two independent cancer genomic datasets (Bladder Urothelial Carcinoma and Ovarian Cancer), we tested the downstream tasks of Evo2, including the prediction of tumor pathogenic variants and the prediction of mutational effects, and compared its performance on A100 and H800 GPUs. The results show that critical importance of FP8 precision, enabling the H800 to achieve a 4× faster inference speed than the A100 with stable accuracy (AUC 0.88-0.95). The 7B-parameter model emerged as the top performer, whereas the 40B model experienced a severe performance drop (AUC to 0.48) on non-FP8 hardware like the A100. These findings empirically validated Evo2’s hardware specifications and provided practical insights for researchers implementing the model with similar computational resources. Futhermore, our findings provide a framework for the application and optimization of downstream tasks of the DNA language model Evo2 in cancer, and can guide researchers in effectively applying it in genomic studies. Key Points Hardware Precision Impact: FP8 precision on H800 GPUs is critical for Evo 2’s performance, enabling 4× faster inference than A100 (without FP8 support) while maintaining high accuracy (AUC 0.88–0.95). Model Scale Optimization: The 7B-parameter model outperformed larger variants (e.g., 40B), which suffered severe accuracy drops (AUC as low as 0.48) on non-FP8 hardware, highlighting a balance between efficiency and performance. Practical Guidelines: We provide a framework for deploying Evo 2 in cancer genomics, including hardware recommendations, dataset curation, and downstream task optimization—valuable for researchers with varied computational resources.
Full text 2,782 characters · extracted from oa-doi-fallback · click to expand
Abstract The rapid advancement of DNA foundation language models has brought about a transformative shift in genomics, allowing for the deciphering of intricate patterns and regulatory mechanisms embedded within DNA sequences. The genomic foundation model Evo2 demonstrates remarkable capabilities in decoding DNA functional patterns through cross-species pretraining. However, despite the great potential of Evo2 in basic genomics research, there is currently no clear and systematic guidance on its specific application scenarios, performance, and optimization directions in the field of tumor genomics, and its performance dependency on specialized hardware (such as FP8 precision on H800 GPUs) has not been empirically benchmarked. Here, we present a focused validation of Evo2 using two independent cancer genomic datasets (Bladder Urothelial Carcinoma and Ovarian Cancer), we tested the downstream tasks of Evo2, including the prediction of tumor pathogenic variants and the prediction of mutational effects, and compared its performance on A100 and H800 GPUs. The results show that critical importance of FP8 precision, enabling the H800 to achieve a 4× faster inference speed than the A100 with stable accuracy (AUC 0.88-0.95). The 7B-parameter model emerged as the top performer, whereas the 40B model experienced a severe performance drop (AUC to 0.48) on non-FP8 hardware like the A100. These findings empirically validated Evo2’s hardware specifications and provided practical insights for researchers implementing the model with similar computational resources. Futhermore, our findings provide a framework for the application and optimization of downstream tasks of the DNA language model Evo2 in cancer, and can guide researchers in effectively applying it in genomic studies. Key Points Hardware Precision Impact: FP8 precision on H800 GPUs is critical for Evo 2’s performance, enabling 4× faster inference than A100 (without FP8 support) while maintaining high accuracy (AUC 0.88–0.95). Model Scale Optimization: The 7B-parameter model outperformed larger variants (e.g., 40B), which suffered severe accuracy drops (AUC as low as 0.48) on non-FP8 hardware, highlighting a balance between efficiency and performance. Practical Guidelines: We provide a framework for deploying Evo 2 in cancer genomics, including hardware recommendations, dataset curation, and downstream task optimization—valuable for researchers with varied computational resources. Competing Interest Statement The authors have declared no competing interest. Footnotes ↵* Authors share co-first authorship lhm_sl_1015{at}163.com jihongyi{at}him.cas.cn cengyuchen{at}him.cas.cn wei_lv2024{at}163.com jianminwu81{at}foxmail.com liusheng{at}him.cas.cn chunhua.lin{at}qdu.edu.cn yanghm{at}genomics.cn

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-22T02:00:06.705733+00:00
License: CC-BY-4.0