Structure-informed direct coupling analysis improves protein mutational landscape predictions

preprint OA: closed
Full text JSON View at publisher

Abstract

Direct Coupling Analysis has been instrumental over the past decade in leveraging evolutionary information and advancing our understanding of biomolecular structure and function. Here, we introduce sparse extensions of this method that explicitly incorporate structural information. StructureDCA focuses on physically relevant interactions by selectively retaining couplings between residues in spatial contact, and StructureDCA[RSA] additionally incorporates per-residue relative solvent accessibility. These models outperform state-of-the-art approaches in describing mutational landscapes, as they more effectively integrate structural context. Moreover, their sparse formulation enables orders-of-magnitude improvements in computational efficiency while preserving interpretability, providing a powerful framework for gaining mechanistic insights into mutation effects and advancing protein design. The StructureDCA models are available as a user-friendly Python package via the PyPI repository. The source code is freely accessible at https://github.com/3BioCompBio/StructureDCA , which also includes a Colab Notebook interface.
Full text 1,290 characters · extracted from oa-doi-fallback · click to expand
Abstract Direct Coupling Analysis has been instrumental over the past decade in leveraging evolutionary information and advancing our understanding of biomolecular structure and function. Here, we introduce sparse extensions of this method that explicitly incorporate structural information. StructureDCA focuses on physically relevant interactions by selectively retaining couplings between residues in spatial contact, and StructureDCA[RSA] additionally incorporates per-residue relative solvent accessibility. These models outperform state-of-the-art approaches in describing mutational landscapes, as they more effectively integrate structural context. Moreover, their sparse formulation enables orders-of-magnitude improvements in computational efficiency while preserving interpretability, providing a powerful framework for gaining mechanistic insights into mutation effects and advancing protein design. The StructureDCA models are available as a user-friendly Python package via the PyPI repository. The source code is freely accessible at https://github.com/3BioCompBio/StructureDCA, which also includes a Colab Notebook interface. Competing Interest Statement The authors have declared no competing interest. Footnotes Contact: matsvei.tsishyn{at}ulb.be, fabrizio.pucci{at}ulb.be

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00