Pretrained protein language models choose between sequence novelty and structural completeness

preprint OA: closed
📄 Open PDF Full text JSON View at publisher

Abstract

Protein language models (PLMs) have gained increasing acceptance in tasks ranging from variant effect prediction in disease to optimization and de novo design of proteins with improved stability, target-binding affinity, and catalytic performance. Despite encouraging performance in such applications, little is understood as far as the degree to which PLM-generated sequences – putative novel protein outputs – recapitulate the broad biophysical rules and diversity of sequence, structure, and function that defines natural protein-space, vital knowledge for boosting the design capacity of PLMs in ever-more-complex systems. Towards this end, we computationally profile and characterize the sequence and structure statistics and properties of hundreds of thousands of potential small proteins proposed through free unconstrained generation from architecturally distinct PLMs. We show that although these models exhibit a prodigious latent capacity to access novel amino-acid sequences, they struggle to approach the structural variation that exists on plain display in nature. Moreover, we uncover a stark tradeoff between prioritizing sequence novelty or structural breadth, exemplified by a “helical bundle trap” that dominates model output when aiming outside the comfortable bounds and evolutionary organization of natural sequences. These findings underscore a critical need for strategies that can rapidly guide PLMs into unlocking through generation the full richness of protein sequence, structure, and function that is consistent with governing biophysics but tantalizingly untapped as of yet in design contexts. Author summary Large language models (LLMs) like GPT aren’t just for human text. Spinoff versions that treat protein and DNA sequences as special “languages” of their own, complete with preferred words and grammars, are being used to identify disease-causing genes and mutations and to design new treatments and drugs for clinical testing. But as anyone who has used an LLM chatbot has probably experienced at one time or another, these models can act narrow-minded or nonsensical, reducing their generality and utility. Protein language models are no exception to such flaws of reasoning. We show that protein language models face a fundamental choice between suggesting novel sequences that look nothing like natural ones and capturing the full range of three-dimensional shapes and structures responsible for the diverse functioning of molecular machines. Managing or bypassing this tradeoff is consequently of high import for designing proteins that impart novel functions and activities for therapeutic targeting and beyond.
Full text 2,740 characters · extracted from oa-doi-fallback · click to expand
Abstract Protein language models (PLMs) have gained increasing acceptance in tasks ranging from variant effect prediction in disease to optimization and de novo design of proteins with improved stability, target-binding affinity, and catalytic performance. Despite encouraging performance in such applications, little is understood as far as the degree to which PLM-generated sequences – putative novel protein outputs – recapitulate the broad biophysical rules and diversity of sequence, structure, and function that defines natural protein-space, vital knowledge for boosting the design capacity of PLMs in ever-more-complex systems. Towards this end, we computationally profile and characterize the sequence and structure statistics and properties of hundreds of thousands of potential small proteins proposed through free unconstrained generation from architecturally distinct PLMs. We show that although these models exhibit a prodigious latent capacity to access novel amino-acid sequences, they struggle to approach the structural variation that exists on plain display in nature. Moreover, we uncover a stark tradeoff between prioritizing sequence novelty or structural breadth, exemplified by a “helical bundle trap” that dominates model output when aiming outside the comfortable bounds and evolutionary organization of natural sequences. These findings underscore a critical need for strategies that can rapidly guide PLMs into unlocking through generation the full richness of protein sequence, structure, and function that is consistent with governing biophysics but tantalizingly untapped as of yet in design contexts. Author summary Large language models (LLMs) like GPT aren’t just for human text. Spinoff versions that treat protein and DNA sequences as special “languages” of their own, complete with preferred words and grammars, are being used to identify disease-causing genes and mutations and to design new treatments and drugs for clinical testing. But as anyone who has used an LLM chatbot has probably experienced at one time or another, these models can act narrow-minded or nonsensical, reducing their generality and utility. Protein language models are no exception to such flaws of reasoning. We show that protein language models face a fundamental choice between suggesting novel sequences that look nothing like natural ones and capturing the full range of three-dimensional shapes and structures responsible for the diverse functioning of molecular machines. Managing or bypassing this tradeoff is consequently of high import for designing proteins that impart novel functions and activities for therapeutic targeting and beyond. Competing Interest Statement The authors have declared no competing interest.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00