⚙
AI-generated deep summary
by claude@2026-07, 2026-07-14
· read from full text
ⓘ
The paper presents CASPE (Critical Amino acids Streamline Protein Evolution), a lightweight protein engineering platform that uses protein large language models with gradient activation mapping and multi-layer attention to localize critical amino acid sites and predict beneficial residue changes for target properties. It describes a workflow combining CAS (site importance indicators without requiring additional structural information or prior knowledge) and APCNet (residue prediction via amino-acid point cloud classification), then evaluates performance on thermostability and pH tolerance, reporting hit rates of 31.3–60% and 40–80% respectively. A directed evolution test on phytase supports generalizability, with a reported 33.3% success rate for beneficial mutants versus FoldX (6.7%) and ESM2-t33 (13.3%). The study does not discuss any explicit limitation related to disease biology. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.
Abstract
Protein large language models (PLMs) provide a novel computational paradigm for the deep mining of sequence co-evolutionary information, significantly accelerating the generation of functional proteins for biotechnological and medical applications. However, the misalignment between zero-shot predicted evolutionary fitness and industrial application requirements leads to a limited success rate in acquiring beneficial mutations, while the high training cost presents another drawback of large models. Here, we developed CASPE (Critical Amino acids Streamline Protein Evolution), a lightweight protein engineering platform for the precise localization and adaptation of critical residues, consisting of the CAS (Critical amino acid sites) and APCNet (Amino acid Point Cloud Classification Network). CAS utilizes gradient activation mapping and multi-layer attention matrices to directly extract key information determining target properties from PLMs and transform it into explicit site importance indicators, without relying on additional structural information or prior knowledge. Working in tandem with APCNet, CASPE establishes a workflow encompassing the entire trajectory from site localization to residue prediction. CASPE achieved remarkable hit rates in identifying beneficial variants for thermostability (31.3-60%) and pH tolerance (40-80%), further uncovering the potential mechanisms of action at these key target sites. Directed evolution of phytase further validated the generalizability of CASPE. CASPET-Phytase achieved a 33.3% success rate in obtaining beneficial mutants, which was significantly better than FoldX (6.7%) and ESM2-t33 (13.3%). CASPE guides enzyme evolution towards precise, site-targeted optimization, providing an efficient computational framework for developing industrial enzymes.
Full text
2,041 characters
· extracted from
oa-doi-fallback
· click to expand
Abstract
Protein large language models (PLMs) provide a novel computational paradigm for the deep mining of sequence co-evolutionary information, significantly accelerating the generation of functional proteins for biotechnological and medical applications. However, the misalignment between zero-shot predicted evolutionary fitness and industrial application requirements leads to a limited success rate in acquiring beneficial mutations, while the high training cost presents another drawback of large models. Here, we developed CASPE (Critical Amino acids Streamline Protein Evolution), a lightweight protein engineering platform for the precise localization and adaptation of critical residues, consisting of the CAS (Critical amino acid sites) and APCNet (Amino acid Point Cloud Classification Network). CAS utilizes gradient activation mapping and multi-layer attention matrices to directly extract key information determining target properties from PLMs and transform it into explicit site importance indicators, without relying on additional structural information or prior knowledge. Working in tandem with APCNet, CASPE establishes a workflow encompassing the entire trajectory from site localization to residue prediction. CASPE achieved remarkable hit rates in identifying beneficial variants for thermostability (31.3-60%) and pH tolerance (40-80%), further uncovering the potential mechanisms of action at these key target sites. Directed evolution of phytase further validated the generalizability of CASPE. CASPET-Phytase achieved a 33.3% success rate in obtaining beneficial mutants, which was significantly better than FoldX (6.7%) and ESM2-t33 (13.3%). CASPE guides enzyme evolution towards precise, site-targeted optimization, providing an efficient computational framework for developing industrial enzymes.
Competing Interest Statement
The authors have declared no competing interest.
Footnotes
Section on title, abstract, introduction and results updated to clarify; Figure 1 and 5 revised; Supplemental files updated.
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.