Predicting evolutionary rate as a pretraining task improves genome language model representations
preprint
OA: closed
CC-BY-4.0
Abstract
Genome language models (gLM) have the potential to further understanding of regulatory genomics without requiring labeled data. Most gLMs are pretrained using sequence reconstruction tasks inspired by natural language processing, but recent studies have shown that these gLMs often fail to capture biological signal. To overcome this, we introduce pretraining tasks that predict the rate of evolution. These tasks are designed so that they can be composed with sequence reconstruction, enabling a controlled comparison of predicting sequence only, evolutionary rate only, or both. To address gaps in existing evaluations, we developed a suite of biologically grounded benchmarks. Across these tasks, and for established variant effect prediction benchmarks, models pretrained on both sequence and evolutionary rate outperform those trained on sequence alone, and training on evolutionary rate can make the even the relatively small models in our work competitive with much larger existing gLMs for some tasks. These results establish evolution as a key training target for genome-scale models.
Full text
1,329 characters
· extracted from
oa-doi-fallback
· click to expand
Abstract
Genome language models (gLM) have the potential to further understanding of regulatory genomics without requiring labeled data. Most gLMs are pretrained using sequence reconstruction tasks inspired by natural language processing, but recent studies have shown that these gLMs often fail to capture biological signal. To overcome this, we introduce pretraining tasks that predict the rate of evolution. These tasks are designed so that they can be composed with sequence reconstruction, enabling a controlled comparison of predicting sequence only, evolutionary rate only, or both. To address gaps in existing evaluations, we developed a suite of biologically grounded benchmarks. Across these tasks, and for established variant effect prediction benchmarks, models pretrained on both sequence and evolutionary rate outperform those trained on sequence alone, and training on evolutionary rate can make the even the relatively small models in our work competitive with much larger existing gLMs for some tasks. These results establish evolution as a key training target for genome-scale models.
Competing Interest Statement
B.W. is currently employed as the Senior Vice President and Head of Biomedical AI at Xaira Therapeutics. He also serves as a scientific advisor to Deep Genomics, Shift Biosciences, and VieCure Inc.
Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.
My notes (saved in your browser only)
Ask this paper
Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works
Funding
- funders
- [{'doi': '10.13039/501100000038', 'name': 'NSERC', 'awards': ['CGS-D']}]
Citation neighborhood (sparse)
Too few in-corpus citations on either side for a chart; here are the lists.
Cites (4)
- Genome modeling and design across all domains of life with Evo 2 2025
- Predicting functional constraints across evolutionary timescales with phylogeny-informed genomic language models 2025
- Genomic Foundationless Models: Pretraining Does Not Promise Performance 2024
- GPN-MSA: an alignment-based DNA language model for genome-wide variant effect prediction 2023
References (27)
- Genome modeling and design across all domains of life with Evo 2 via crossref
- Genomic Foundationless Models: Pretraining Does Not Promise Performance via crossref
- GPN-MSA: an alignment-based DNA language model for genome-wide variant effect prediction via crossref
- Predicting functional constraints across evolutionary timescales with phylogeny-informed genomic language models via crossref
- doi:10.1093/nar/gks1092 via crossref
- doi:10.1093/nar/gkae1071 via crossref
- doi:10.1093/gbe/evac055 via crossref
- doi:10.1186/s13100-020-00208-w via crossref
- doi:10.1093/nar/gkaa1087 via crossref
- doi:10.1186/s12863-023-01123-8 via crossref
- doi:10.1101/gr.227819.117 via crossref
- doi:10.1145/3107411.3107425 via crossref
- doi:10.1186/s13059-016-0974-4 via crossref
- doi:10.1038/s41586-022-04558-8 via crossref
- doi:10.1093/nar/gkae974 via crossref
- doi:10.1093/nar/28.1.302 via crossref
- doi:10.1101/gr.097857.109 via crossref
- doi:10.1093/nar/gky1016 via crossref
- doi:10.1186/s13059-025-03674-8 via crossref
- doi:10.1093/nar/gkl822 via crossref
- doi:10.1073/pnas.2406285121 via crossref
- doi:10.1038/nmeth.3547 via crossref
- doi:10.1007/978-3-031-90252-9_7 via crossref
- doi:10.1093/nar/gkr485 via crossref
- doi:10.1038/s42256-025-01007-9 via crossref
- doi:10.1038/s41586-020-2876-6 via crossref
- doi:10.1038/s41592-024-02523-z via crossref
Source provenance
- crossref
- last seen: 2026-07-11T06:39:50.356103+00:00
- europepmc
- last seen: 2026-05-20T01:45:00.602351+00:00
- unpaywall
- last seen: 2026-05-22T02:00:06.705733+00:00
License: CC-BY-4.0