Structure-informed Language Models Are Protein Designers

preprint OA: closed CC-BY-NC-ND-4.0
📄 Open PDF View at publisher

Abstract

This paper demonstrates that language models are strong structure-based protein designers. We present LM-D esign , a generic approach to reprogramming sequence-based protein language models ( p LMs), that have learned massive sequential evolutionary knowledge from the universe of natural protein sequences, to acquire an immediate capability to design preferable protein sequences for given folds. We conduct a structural surgery on p LMs, where a lightweight structural adapter is implanted into p LMs and endows it with structural awareness. During inference, iterative refinement is performed to effectively optimize the generated protein sequences. Experiments show that LM-D esign improves the state-of-the-art results by a large margin, leading to 4% to 12% accuracy gains in sequence recovery ( e . g ., 55.65%/56.63% on CATH 4.2/4.3 single-chain benchmarks, and > 60% when designing protein complexes). We provide extensive and in-depth analyses, which verify that LM-D esign can (1) indeed leverage both structural and sequential knowledge to accurately handle structurally non-deterministic regions, (2) benefit from scaling data and model size, and (3) generalize to other proteins ( e . g ., antibodies and de novo proteins).

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00
unpaywall
last seen: 2026-05-22T02:00:06.705733+00:00
License: CC-BY-NC-ND-4.0