ProteomeLM: A proteome-scale language model allowing fast prediction of protein-protein interactions and gene essentiality across taxa

preprint OA: gold CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Language models trained on biological sequences are advancing inference tasks from the scale of single proteins to that of genomic neighborhoods. Here, we introduce ProteomeLM, a transformer-based language model that uniquely operates on entire proteomes from species spanning the tree of life. ProteomeLM is trained to reconstruct masked protein embeddings using the whole proteomic context, yielding contextualized protein representations that reflect proteome-scale functional constraints. Notably, ProteomeLM’s attention coefficients encode protein-protein interactions (PPI), despite being trained without interaction labels. Furthermore, it enables interactome-wide PPI screening that is substantially more accurate, and orders of magnitude faster, than amino-acid coevolution-based methods. We further develop ProteomeLM-PPI, a supervised model that combines ProteomeLM embeddings and attention coefficients to achieve state-of-the-art PPI prediction across benchmarks and species. Finally, we introduce ProteomeLM-Ess, a supervised gene essentiality predictor that generalizes across diverse taxa. Our results demonstrate the potential of proteome-scale language models for addressing function and interactions at the organism level. Significance statement Predicting protein interactions and functions is a key challenge in biology. Although deep learning-based language models are advancing the analysis of individual protein sequences and of genomic neighborhoods, they struggle to capture properties involving all the proteins expressed in a cell, such as protein–protein interactions (PPI) and gene essentiality. We present ProteomeLM, a language model that reasons on entire proteomes across diverse species. ProteomeLM captures PPI without supervision, and enables more accurate and faster screening of entire interactomes than current sequence-based approaches. ProteomeLM also delivers state-of-the-art supervised PPI prediction, and improves supervised prediction of gene essentiality compared to protein language models. These results demonstrate the potential of proteome-scale language models to reveal system-level organization and functional relationships.
Full text 2,628 characters · extracted from oa-doi-fallback · click to expand
Abstract Language models trained on biological sequences are advancing inference tasks from the scale of single proteins to that of genomic neighborhoods. Here, we introduce ProteomeLM, a transformer-based language model that uniquely operates on entire proteomes from species spanning the tree of life. ProteomeLM is trained to reconstruct masked protein embeddings using the whole proteomic context, yielding contextualized protein representations that reflect proteome-scale functional constraints. Notably, ProteomeLM’s attention coefficients encode protein-protein interactions (PPI), despite being trained without interaction labels. Furthermore, it enables interactome-wide PPI screening that is substantially more accurate, and orders of magnitude faster, than amino-acid coevolution-based methods. We further develop ProteomeLM-PPI, a supervised model that combines ProteomeLM embeddings and attention coefficients to achieve state-of-the-art PPI prediction across benchmarks and species. Finally, we introduce ProteomeLM-Ess, a supervised gene essentiality predictor that generalizes across diverse taxa. Our results demonstrate the potential of proteome-scale language models for addressing function and interactions at the organism level. Significance statement Predicting protein interactions and functions is a key challenge in biology. Although deep learning-based language models are advancing the analysis of individual protein sequences and of genomic neighborhoods, they struggle to capture properties involving all the proteins expressed in a cell, such as protein–protein interactions (PPI) and gene essentiality. We present ProteomeLM, a language model that reasons on entire proteomes across diverse species. ProteomeLM captures PPI without supervision, and enables more accurate and faster screening of entire interactomes than current sequence-based approaches. ProteomeLM also delivers state-of-the-art supervised PPI prediction, and improves supervised prediction of gene essentiality compared to protein language models. These results demonstrate the potential of proteome-scale language models to reveal system-level organization and functional relationships. Competing Interest Statement The authors have declared no competing interest. Footnotes Diverse revision and new results were added following reviewers recommendation after a first round of reviews: - We have added results on distinguishing direct PPIs from functional links - We have added examples on resolved structural complexes - We have added more benchmarking on gene essentiality - We have improved some explanations in the methods

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-21T05:10:58.409756+00:00
License: CC-BY-4.0