Computational optimization of DEK1 calpain domain solubility through integrated structural modelling and targeted mutagenesis | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Computational optimization of DEK1 calpain domain solubility through integrated structural modelling and targeted mutagenesis Mohammad Dabiri, Zdenko Levarski, Eva Struharnanska, Viktor Demko, and 3 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7559073/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 08 Feb, 2026 Read the published version in Scientific Reports → Version 1 posted 13 You are reading this latest preprint version Abstract The DEFECTIVE KERNEL 1 (DEK1) protein plays essential functions throughout plant development. DEK1 is a multidomain 250 kDa protein with yet unsolved 3D structure. To facilitate structural and functional studies of DEK1, here we investigate its calpain protease core domain (CysPc) from Physcomitrium patens . Using integrated structural modelling we propose targeted mutagenesis of CysPc to enhance its solubility during recombinant protein production. We created a pipeline to predict the topology of the CysPc domain with improved precision, providing a robust framework for further exploration. We evaluated the native and mutant structures by MD simulations, concentrating on several solubility-related parameters. Following these features, we implemented specific single, double, and triple amino acid mutagenesis to select variants with improved solubility. Our method preserves overall structural integrity while reducing aggregation-prone traits. We advocate for the utilization of reinforcement learning method that can effectively traverse the extensive combinatorial space and prioritize mutation sets with the greatest potential for enhancing solubility. This framework provides a logical, data-driven approach to improving protein solubility, particularly beneficial in situations lacking high-resolution structural data. Biological sciences/Biochemistry Biological sciences/Biophysics Biological sciences/Computational biology and bioinformatics Biological sciences/Structural biology Calpain Protein solubility Mutagenesis Molecular Dynamics Structure prediction Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Introduction The high solubility and native structure of proteins is a desired outcome of their production in heterologous hosts, particularly Escherichia coli with its well-known drawbacks such as restricted post-translational modification capability, recurrent development of inclusion bodies, and the propensity to generate insoluble or misfolded proteins under overexpression conditions. Low solubility can generate inclusion bodies, which seriously impede structural studies, functional tests, and protein recovery 1 . Clearly showing this issue for eukaryotic and membrane-associated proteins produced in bacterial systems are variations in codon usage, lack of post-translational modifications, and predisposition for misfolding in complex protein domains 2 . In both academic and industrial biotechnology, protein insolubility thus presents a major obstacle that calls for the development of methodical and successful solutions to improve solubility. Structure determination of eukaryotic calpain proteases has been challenging due to insolubility and aggregation issues. Currently, the structure of mammalian calpains has been solved (Refs). Calpains however represent diverse family with neither the 3D structure nor the biological role determined for majority of them (Zhao et al., 2012). The solubility issue is particularly relevant to the study of DEFECTIVE KERNEL1 (DEK1), a large, multi-domain membrane protein. DEK1 is a 240 kDa protein with a 23-spanning transmembrane domain, a cytosolic linker segment, and a C-terminal calpain protease 25 CysPc. Genetic analyses in cereals and diverse model plant species showed that DEK1 plays essential functions throughout plant development via regulation of meristematic cell divisions and cell fate control 21 , 23 . In addition, point mutagenesis and genetic complementation studies indicated that the calpain moiety of DEK1 is essential for the protein function and represents the protein’s effector 26 . Despite its essential function, the 3D structure of DEK1 has not been solved yet. Consequently, performing mutagenesis studies on CycPc’s structure is particularly challenging due to the absence of crystallographic data or any experimentally determined structures in public databases, as this limits the accuracy of structural predictions and the rational design of beneficial mutations. Traditional methods for addressing this difficulty encompass empirical mutagenesis 3 , co-expression with molecular chaperones 2 , 4 , codon optimization 5 , and fusion with solubility-enhancing tags 6 – 8 . Nevertheless, these processes are frequently protracted, laborious, and contingent upon circumstances. In recent years, computational approaches have emerged as valuable tools for targeted mutagenesis aimed at enhancing protein solubility 9 . Still, these processes are often time-consuming, labour-intensive, dependent on circumstance. Recent years have seen computational methods become useful tools for targeted mutagenesis to increase protein solubility 10 – 13 . By use of structural modelling, molecular dynamics (MD) simulations, and machine learning-generated solubility predictions, these approaches identify alterations enhancing folding stability and reducing aggregation tendency 14 – 16 . For eukaryotic proteins produced in prokaryotic systems like E. coli , computational optimization provides a systematic and effective approach to enhance solubility by targeted mutagenesis which related studies revealed surface patch analysis was used to design rHuEPO variants in E. coli , where reducing positively charged patches (e.g., F48D, R150D) enhanced solubility by up to 60%, while increasing them (e.g., E13K) led to significant solubility loss compared to the wild type 17 . Also, it has been reported that these methodologies applied integrate structural modelling, MD simulations, and solubility prediction methods to systematically diminish aggregation-prone areas 18 – 20 . Based on recent research regarding the solubility engineering of aggregation-prone eukaryotic proteins and the de novo structural modelling framework and MD; we introduce a computational framework designed to enhance the solubility of the DEK1 calpain-like (CysPc) domain from P. patens by rational mutagenesis. The P. patens has been chosen as this model plant enables to use effective targeted mutagenesis to link protein structure with function 27 , 28 . Through this manner, we identify surface-exposed residues exhibiting significant aggregation propensity or detrimental solvation patterns. We suggest specific changes that will improve solubility by making the surface less hydrophobic and increasing local flexibility while keeping the structure stable. This method uses tried-and-true methods for improving the solubility of eukaryotic proteins that tend to clump together and provides a way to study the structure and function of DEK1 domains that have been hard to study experimentally because they don't dissolve well enough 29 – 33 . Here. First we developed a workflow procedure to increase the accuracy of prediction the structure and based on provided reconstructed structure we come up with a computational method using focused mutagenesis to maximize the solubility of DEK1’s CysPc domain. We find residues causing low solubility using structural modelling and MD simulations, then we offer mutation candidates expected to increase solubility while preserving domain stability. This work opens the route for functional and structural studies of DEK1 and offers a paradigm for solubility-oriented protein engineering of challenging eukaryotic domains. Results Modelling and validation of natural-type and mutated protein structures We employed various structural assessment techniques to meticulously evaluate all modelled structures, encompassing both native and mutant protein variations (Fig. 1 ). We additionally contrasted our integrated homology modelling with direct threading modelling. The projected models' TM-score, US-score, QMEANDisCo, ProQ3D, and Z-score indicate that the structure generated by AlphaFold2 is the most dependable among those examined. The continuous improvement of these quality parameters indicates that the AlphaFold2 model is the most precise and dependable method for representing the target protein's structure. The model's consistently elevated values demonstrate this (Table. 1). The SAVES v6.1 server's assessment of stereochemical quality revealed that most residues in all structures resided within the favoured and allowed sections of the Ramachandran plot, signifying appropriate backbone geometry. The QMEAN scoring and analysis conducted using the SWISS-MODEL Assessment tools verified that both the normal and mutants’ models exhibited satisfactory overall quality and reliability ratings. Result revealed that average score of QMEAN for all analysed structure with integrated homology modelling was 1.02 which considered as an acceptable score while this score was 1.53 in direct threading method of structure prediction of CysPc domain. Also, the ERRAT analysis corroborated these findings, revealing that all models demonstrated good overall quality factors, indicating a lack of substantial non-bonded interaction defects. The amalgamation of these validation techniques verifies that all produced protein structures both natural-type and mutated exhibit good structural integrity and are appropriate for further analysis. All tests which carried out with SAVES v6.1 server for integrated homology modelling indicate that the average of overall quality factor of predicted structures was 97.32 and for verified 3D structure reached to 86.53% while these metrics for direct method of structure prediction was 91.33 and 74.26% respectively which significantly increased. In addition, the Ramachandran plot analyses of natural structure of CysPc versus direct threading model demonstrated that the structured model applied method of prediction exhibited a greater proportion of residues within the acceptable (favoured and allowed) regions than those produced by direct threading structure prediction methods, signifying enhanced stereochemical quality and backbone geometry in the validated models (Fig. 2 ). MD simulations were performed on the natural CysPc domain and each individual mutant to thoroughly evaluate the structural and dynamic effects of sequence alterations. For each structure, we extracted a collection of pertinent features: SASA, H-bond, RDF, Rg, Diffusion coefficient, RMSD, RMSF, minDist. We utilized a weighted scoring approach to statistically compare and prioritize individual mutants based on these criteria. The weights for each characteristic were established using statistical study of the link between feature alterations (from natural to single mutant) and their correlation with favourable enhancements in protein solubility. The resultant set of experimentally derived weights is as follows: SASA (1.2), H-bond (1.2), RDF (1.0), Rg (1.1), Diffusion (0.8), RMSD (-1.4), RMSF (-1.4), and minDist (1.0). The weighted differences were normalized to facilitate direct comparison among features with varying scales, and the resultant Weight(x) score (Eq. 1) permitted efficient ranking of all individual mutants. Mutants with the highest composite scores were recognized as the most promising candidates for the subsequent phase of combinatorial mutant type and optimization. Evaluation of single mutants; structural and dynamic changes Single-point mutations in the CysPc sequence were predicted via AlphaFold2 and underwent 200 ns MD simulations to assess their structural and dynamic characteristics. The backbone RMSD values of all mutants stabilized post-equilibration and remained analogous to the natural type, signifying that the overall conformation of the CysPc domain was maintained. Analyses of RMSF indicated that mutations in surface-exposed or loop regions generally enhanced local flexibility, whereas substitutions within core secondary structures typically diminished RMSF, signifying localized stabilizing. Alterations SASA were noted based on the characteristics of the inserted residue: hydrophilic or charged alterations at the surface typically augmented local SASA, indicating enhanced solubility, while hydrophobic replacements diminished solvent exposure. The total count of intra-protein hydrogen bonds remained comparable to the natural type; however, certain mutants displayed novel surface-exposed hydrogen bonds that could potentially improve solubility or stability. The radius of gyration values for all mutants exhibited negligible variation, affirming that domain compactness was preserved. The results as a whole show that most of the individual mutations in the CysPc domain, as shown by AlphaFold2 and confirmed by long-timescale MD simulations, do not change the overall structure of the domain. However, they may change how flexible it is and how much solvent it is exposed to, which are important for improving stability and solubility. Iterative generation and assessment of double and triple mutants We meticulously developed and evaluated a library of double mutants in the CysPc domain of DEK1 by extensive in silico screening. Subsequently, we selected the 25 out of 45 most promising candidates from double mutants based on various MD features associated with solubility. We meticulously selected each double mutant by amalgamating two single-point mutations that had previously demonstrated enhanced solubility metrics independently. MD simulations of the double mutants demonstrated a notable decrease in RMSD, signifying enhanced structural stability relative to the natural-type and single mutants. Concurrently, RMSF analyses indicated less residue-level fluctuations in critical locations, corroborating the notion of a more stable conformation. These data aligned with signs of increased solubility, since the mutants seemed more solvent-accessible and exhibited improved hydration. To corroborate these findings, we additionally assessed other structural and solubility-related characteristics. Certain double mutants exhibited enhancements in various solubility descriptors concurrently. These mutants are excellent subjects for further mutational investigation or experimental validation. In summary, these 25 double mutants exhibited several enhanced solubility indicators and are promising candidates for subsequent triple mutants’ selection. Table 2 presents their ranking and fitness of MD-derived profiles, demonstrating the efficacy of iterative, data-driven mutagenesis techniques in systematically enhancing protein solubility in silico . Identification of top mutant candidates for improved solubility The identification and creation of triple mutants were executed via a complex computational pipeline that utilized both empirical MD data and sophisticated reinforcement learning, specifically employing a Proximal Policy Optimization (PPO) framework. Initially, a comprehensive dataset was created through in silico mutagenesis of the CysPc domain, producing single and double mutants and simulating the dynamics of each variant to derive critical solubility-related descriptors: RMSD, RMSF, SASA, H-bond, Rg, radial distribution RDF, minimum residue distance, and diffusion coefficient. The descriptors were quantitatively assessed for each mutation and normalized against the natural-type protein to create a comprehensive feature collection. The PPO methodology was subsequently employed to systematically explore the extensive combinatorial space of potential triple mutants. The reinforcement learning agent was taught using a bespoke reward function that incorporated enhancements in all pertinent solubility parameters rewarding additive or synergistic increases and penalizing unstable or redundant alterations. The agent's action space was dynamically limited: each time, it only looked at positions and substitutions that had been shown to improve solubility in single or double mutants. At the same time, it looked at the natural-type conformation to make sure that no mutations were introduced that would break important native contacts or structural motifs. For every prospective triple mutant suggested by the PPO agent, an in silico MD simulation was conducted, and the resultant structural and solubility metrics were incorporated into the algorithm to enhance its exploration approach. The algorithm's policy was enhanced by integrating knowledge of epistatic effects situations where one mutation's influence is altered by another's presence achieved through a systematic comparison of the triple mutant's characteristics with those of its individual single and double mutants. This method facilitated the identification of combinations in which the third mutation yielded a genuinely advantageous effect just when associated with double mutant backgrounds, rather than independently (Fig. 3 ). Furthermore, the natural sequence was used at each evaluation stage as a reference to confirm that the proposed triple mutants would not deviate significantly from the protein's original conformation or stability characteristics which all resulted in MUT347 as rank 1 of triple mutant selection. We ran extra MD analyses on the triple mutants that obtained the best scores to make sure they did what they were meant to do and to see how they would change the structure and flexibility of proteins. This process made sure that only the best and most lasting alternatives made it to the final phase of testing. It combined the benefits of data-driven prediction and biological plausibility. Stability and flexibility of natural-type and mutated protein structures Over a 200 ns trajectory MD models were run to see how the surface mutations affected the stability and dynamic behaviour of both the natural-type and mutant proteins. RMSD plots (Fig. 4 , B) show that the mutant protein (MUT347) reached equilibrium faster and had consistently lower RMSD values than the natural-type protein. The natural-type structure went up steadily and stopped at about 0.70 Å, but the mutant form stopped moving earlier and stayed at 0.49 Å throughout the exercise. The results show that the triple mutation made the structure more stable overall, most likely by reducing the changes in shape that come from surface-solvent interactions. We also applied RMSF analysis to check the flexibility of the area (Fig. 4 , A). As expected, both protein versions showed similar fluctuation patterns, with most of the flexible areas being found at the N and C-termini. The mutant structure had slightly lower RMSF values in several loop areas, especially between residues 0:20, 50:57, 190:209, and 322:354, which meant that the area was stiffer. This stability could mean that the surface was remodelled effectively, adding more polar or charged contacts that made solvent engagement and structural compactness better. Also, persistent fluctuation in active site of both structures still high, suggests that applied mutants have negligible impact on the structure integrity on catalytic area. Additionally, Rg analysis between natural and mutated structure supports this matter that mutated structure has better compactness and has higher integrity over time of simulation and lower fluctuation (Fig. 4 , C). The RMSD and RMSF results support the idea that surface-targeted mutations can make proteins more stable in solution, especially when there are a lot of them and depletion-induced aggregation can happen because of short interactions between molecules. Implications of secondary structure stability and solubility DSSP was used for secondary structure analysis during the exercise to find out if the changes seen in the dynamic behaviour led to changes in the structure. The time-resolved secondary structure maps (Fig. 5 ) showed that the overall secondary structure makeup of both the native and mutant forms stayed the same over the 200 ns trajectory. When looked at more closely, it was found that the mutant protein kept its α-helix and β-sheet structure more consistently, while the natural-type protein changed between structured and unstructured states more often, especially in loop-rich areas. This was confirmed by the quantitative study of the average secondary structure makeup. The mutant had a slightly higher content of helixes and turns compared to the normal variant, but a significantly lower content of coils and bends. There were also fewer standard deviations across simulation blocks. This means that the structure is generally more stable and denser, which these traits are usually linked to protein’s higher solubility. These results are important because they match the decrease in RMSF seen in the mutant. This means that the designed surface mutations made the area less flexible while also making the native secondary structure motifs more stable. From a physiological point of view, this drop in surface-exposed dynamic disorder may limit the ability for depletion to cause aggregation, especially when there is a lot going on inside cells or in formulations. The combined study of structural and secondary structures shows that changes made to the surface have a good effect on solubility while keeping the structure's integrity. In addition, minimum distance matrix (mdmat) of last 20 ns for residue-residue contact pattern has a higher proof on this matter that applied mutants has a preserved global contact map and no major rearrangement and misfolding occurred and there isn’t any disrupt domain architecture on internal topology (Fig. 5 , B). Analyses of hydrogen bonding and solvent accessibility area We examined protein-solvent hydrogen bonds (between structure and surrounding solvent) and intra-protein hydrogen (between residues) bonds and SASA during the simulation (dynamic mode) and in static mode to investigate the structural basis for the enhanced solubility of the natural-type and mutated proteins. The overall count of intra-protein hydrogen bonds was persistently elevated in the mutant structure throughout the trajectory, averaging 274.98 ± 9.6 H-bonds, in contrast to 229.17 ± 7.6 in the natural-type (Fig. 6 , A). Also, slight reduction of hydrogen bond between structures and solvent was not statistically significant but higher formation of hydrogen bond between residues (intra-molecular) detected which this matter reflects more internal stabilization and tighter fold in mutated structure. Furthermore, static mode of SASA profiles showed that the two protein variants had very different amounts of surface contact which statistically significant. The mutated structure had a higher average SASA (21108 Ų), but the natural protein had a lower surface area (17821 Ų), which supports the successful surface design and presents a greater surface area available for interaction with water molecules. This rise means that the mutations either improved the packing inside or set up connections that make the structure more stable, which is what caused the RMSD and RMSF values were reduced. There aren’t any significant differences between SASA dynamic mode of natural-type and mutant in all over trajectories, which where 185.59 ± 1.97 and 183.19 ± 3.7 nm/S 2 /N respectively. Notwithstanding, better stable fluctuation and higher accessible area were detected on mutated structure at finals frames. On the other hand, analyses of free energy surface (FES) in basis of RMSD and SASA (protein-surface) a slight better reduction of energy resulted for mutant structure in comparison with natural type (Fig. 6 , B). A bigger solvent accessible surface area on static mode and stronger intra-molecular hydrogen bonds supports the idea that the mutant has changed into a better shape that makes it easier to dissolve. Higher SASA means that more polar and hydrophilic surface residues are exposed, which lets them interact with water molecules around them in a wider area. Furthermore, analyses of H-bond among natural type and mutated structure revealed no significant difference but along with overall solvation, per-residue H-bond analysis showed that some residues close to the mutation sites were more regularly involved in water interactions in the mutant structure (Fig. 6 , A). This localized rise in hydration may work as a buffer against nonspecific inter-protein contacts, making it less likely that proteins will clump together when they are crowded. These results add to the SASA data, which showed a moderate drop in global solvent exposure but a better distribution of polar residues that could be accessed by solvents. In opposition, a significant difference of H-bond discovered in intra residues level which could improve with potential of reduce in aggregation and supports solubility via structure better compactness which Rg analyses also supports this matter. Also, studies of hydrogen bond lifetimes showed no difference between natural-type and mutant. This suggests that the applied mutant has no impact on structure destabilization and flexibility. Figure 7 , Panel A illustrates that alterations in pH result in negligible, limited variations, particularly at mutated sites, with the traces for both the natural-type and Mut347 predominantly overlapping. The principal flexibility is restricted to the N- and C-terminal peaks, and no new hotspots arise throughout the pH series. In contrast, Panels B and C (NaCl and (NH₄)₂SO₄, within a concentration range of 0.2–1.2 mol/L) demonstrate distinct ionic-strength effects: as salt concentration increases, the overall RMSF envelope diminishes, terminal peaks reduce in size, and loop areas become more uniform, with these enhancements consistently more evident in Mut347. Significantly, at the mutation site and its proximal neighbours in Mut347, the coloured traces at elevated salt concentrations are positioned beneath the Natural-type curves, suggesting enhanced fluctuations for both salts; this decrease persists into adjacent loop segments, where RMSF is markedly diminished compared to low-salt conditions. The reduction of terminal mobility (both N and C) is more pronounced in Mut347, resulting in cleaner, lower-amplitude endpoints compared to the Natural type. Although NaCl and (NH₄)₂SO₄ exhibit a similar qualitative trend, the reduction in RMSF—at the altered residue, in adjacent loops, and at the termini is somewhat more significant in (NH₄)₂SO₄ at equivalent doses. Discussion The integrated modelling and validation technique employed merging template searches (Dali/COFACTOR, RCSB PDB/UniProt) with various predictors and rigorous quality assessments—yielded a dependable structural foundation for molecular dynamics. Dali and COFACTOR are recognized for structure/function transfer from homologs, whilst RCSB PDB and UniProt function as the principal, curated repositories for 3D structures and sequence annotations, respectively 34 , 35 . Across independent quality metrics (e.g., QMEANDisCo, ProQ3D, Ramachandran/PROCHECK, ERRAT), the AlphaFold2 model consistently scored best, aligning with broad evidence that AlphaFold2 attains near-experimental accuracy for many single-domain proteins. Specifically, QMEANDisCo and ProQ3D are cutting-edge model-quality evaluators, while the SAVES/ERRAT toolbox continues to serve as a benchmark for stereochemical assessments and the identification of non-bonded interaction anomalies 36 , 37 . This consensus de novo approach is supported by recent benchmarks, which demonstrate that the combination of models from several predictors can perform better than any one approach, especially for challenging targeted structure especially when homology modelling with single template is insufficient 38 , 39 . Our findings corroborate this pattern: our combined, consensus-based approach as an integrated method achieved 86% prediction accuracy for the CysPc domain, compared to 74% for AlphaFold2 as direct-threading method 38 , 39 . The molecular dynamics investigations (RMSD, RMSF, Rg, SASA) demonstrate that the designed surface replacements enhanced global dynamics while maintaining the structural conformation. RMSD indicates the overall deviation from the reference conformation, whereas RMSF identifies flexibility along the sequence; the observed first: accelerated equilibration/lower RMSD and second: attenuated RMSF at termini and loops are characteristic indicators of enhanced conformational stability. The minor decrease and refinement of Rg indicate a more compact ensemble, aligning with conventional interpretations of protein compactness in simulations and statistical evaluations of PDB structures 40 , 41 . Additionally, alterations in SASA and hydrogen-bond counts indicate enhanced internal packing and a restructured hydration shell; increased intra-protein hydrogen bonding, together with slight modifications in protein–solvent hydrogen bonds, is a prevalent mechanism for preserving the native state without altering the fold 42 . At the sequence level, the Arg-containing substitutions near the surface likely enhance both stability and solubility by augmenting local hydrogen-bonding opportunities (including weak yet frequent C–H···O interactions) and redistributing surface charge to promote hydration and diminish aggregation-prone positive regions. The stabilizing function of canonical and weak (C–H···O) hydrogen bonds at interfaces is well established, and Arg-rich motifs can engage in both canonical hydrogen bonding and advantageous CH···O interactions 43 – 46 . Furthermore, surface electrostatics significantly affect solubility; enhanced negative/polar exposure and a reduction in extensive positive regions are associated with improved soluble expression aligning with SASA/H-bond patterns and the diminished variations in exposed loops 47 . The mutant demonstrates an elevated number of intramolecular hydrogen bonds without an increase in hydrogen bond lifetimes, indicating a denser yet not necessarily slower network precisely the type of incremental reinforcement that contributes several kJ·mol⁻¹ of stability without excessively restricting dynamics. Comprehensive mutational thermodynamics indicate that backbone and side-chain hydrogen bonding positively influence stability in a context-dependent manner, aligning with our finding that an increased number of internal hydrogen bonds correlates with decreased stability 42 . The MD data indicate a consistent stabilizing mechanism for the surface-engineered mutant. The accelerated equilibration and consistently reduced RMSD of MUT347 compared to the natural-type suggest a more restricted conformational ensemble, a result anticipated when surface electrostatics are positively altered by mutation. Previous research indicates that targeting solvent-exposed regions and improving surface charge can enhance thermodynamic stability without compromising the fold, aligning with our observations 48 – 50 . In molecular dynamics, RMSD and RMSF serve as conventional metrics for stability and stiffness; a reduction in global RMSD accompanied by a decrease in local RMSF often indicates a more stable, compact ensemble 51 , 52 . Furthermore, RMSF profiles related to implemented surface design indicate that the most significant reductions occur in loops and at the termini, whereas the structured core maintains rigidity specifically where stabilization typically initially emerges. The DSSP analysis verifies that secondary-structure content is maintained or somewhat enhanced (increased helix/turn, decreased coil/bend), suggesting that the mutant's diminished fluctuations result in improved preservation of native motifs rather than a modification of fold 53 , 54 . The subdued pH dependence and pronounced salt-dependent attenuation of RMSF most evident at the N/C-terminal peaks, solvent-exposed loops, and the mutated site align with traditional ionic-strength screening: elevated salt concentrations diminish the Debye length, mitigate long-range charge–charge repulsion, and promote short-range interactions, consequently reducing backbone fluctuations in the most unstable regions 55 .In accordance with ion specificity, the marginally enhanced damping observed with (NH₄)₂SO₄ compared to NaCl is anticipated for a more kosmotropic salt positioned higher in the Hofmeister series, which facilitates compaction and salting-out at similar ionic strengths 56 , 57 .Notably, Mut347 exhibits superior enhancement (reduced RMSF) with elevated salt concentrations compared to the Natural type particularly at and near the mutated site and demonstrates a slightly improved global profile; such sequence-dependent advancements are expected when local charge distributions are optimized by screened electrostatics and when substitutions improve interfacial interactions 58 , 59 . In this context, mutations at residue 21 and residue 342 may enhance stabilization by facilitating stronger hydrogen bonds with adjacent chains, a recognized factor in complex/interface stability, including weak yet frequent CH-O and canonical H-bonds at residue-residue interfaces 60 , 61 . Although, RDF, Diffusion coefficient and minDist had no substantial changes in all over mutants since the primary aim of mutagenesis process was to enhance solubility while conserving total structure integrity Consequently, significant alterations in these parameters were not expected. Our integrated computational pipeline, which includes iterative mutagenesis, extensive MD analyses, and data-driven through proximal policy optimization, is much better than traditional methods for structure-based protein engineering and increasing solubility. In most cases, mutagenesis and improving protein solubility use single-point mutations or random library screening. These methods are time-consuming and don't work well for finding complex epistatic interactions between many residues 62 , 63 . On the other hand, our method uses MD simulations in a systematic way to get high-dimensional structural and solubility descriptors from different single and double mutants’ library. This creates a strong data base for the future use of any data-driven techniques. Our implemented input information to find its way through the complex world of triple mutants and focuses on mutation combinations that are most likely to improve solubility in an additive or synergistic way, while also avoiding substitutions that are harmful to structure. Taken together, data supports the change in surface physiochemical properties, along with the higher internal hydrogen bonding, suggests that the mutation leads to better solvation and hydration shell formation, which are both necessary to keep the solubility and stop nonspecific aggregation when conditions are tight. These findings indicate that the surface-engineered triple mutation enhances structural compaction and internal stabilization, and more accessible surface presumably diminishing the tendency for depletion-induced aggregation in dense situations. Methods Structure prediction, preparation and validation The amino acid sequence of the DEK1 calpain-type cysteine protease domain was retrieved from the NCBI protein database ( https://www.ncbi.nlm.nih.gov/ ), specifically from P. patens (accession number: XP_024400265.1). This sequence corresponds to the predicted CysPc domain of DEK1, which plays a critical role in the protein's proteolytic activity. The sequence was used as the basis for homology modelling and subsequent structural analyses. Prior to structure prediction, the sequence was analysed to identify domain boundaries, surface residues, and potential disordered regions using standard bioinformatics tools. For finding reference structure to carry out homology modelling, threading (fold recognition) model of CysPc domain which was predicted by C-I-ITASSER server ( https://zhanggroup.org/C-I-TASSER ) 64 , and then applied in different blast engine on uniport, (against targeted data base of UniProtKB with 3D structure) ( https://www.uniprot.org/blast ), RCSB PDB structure similarity search ( https://www.rcsb.org/search/advanced ), Dali server ( http://ekhidna2.biocenter.helsinki.fi ) 35 and COFACTOR ( https://zhanggroup.org/COFACTOR ) 34 , 65 . All search engine resulted in structure of mu-like calpain (PDB: 1QXP) and related structure retrieved from protein data bank (RCSB). Because experimental structure of 1QXP has 1571 modelled residue count and deposited structure had 1800 residues, closest structure of 1QXP retrieved from RCSB protein data bank according to the result of structural search engines (1KEU, 2P0R, and 6BDT)) and Modeller 10.7 with multimode prediction strategy applied to predict complete structure with reference structures to reaching to the best model possible as reference structure 66 , 67 . The 3D structure of CysPc domain was modelled using AlphaFold2 68,69 , SWISS-MODEL ( http://swissmodel.expasy.org ) 70 , and I-TASSER ( https://zhanggroup.org/I-TASSER/ ) 71 based on homology modelling to the experimental crystal structure of mu-like calpain (PDB: 1QXP), ensuring high coverage and confidence in the predicted conformation. The resulting model was then refined by Galaxy Refine for structure refinement ( https://galaxy.seoklab.org ) 72 and ModLoop ( https://modbase.compbio.ucsf.edu/ ) for loops refinements 73 . Furthermore, the achieved 3D protein structure minimized any analysis about the structural integrity of the protein structure by SAVES v6.1 server ( https://saves.mbi.ucla.edu/ ) for 3D verification, ERRAT analyses 74 and SWISS-MODEL assess ( https://swissmodel.expasy.org/assess ) for QMEAN per residues errors 70 , QMEANDisCo for distance limitation on model quality estimation 37 , TM-score (with superimposed predicted model against each other) and US-Score (predicted model against 1QXP) to assess the similarity structure of predicted model 75 , 76 ,Z-Score for overall model quality with ProSA-web 77 , ProQ3D for improved model quality assessment ( https://proq3.bioinfo.se/pred/ ) 78 , and also, Molprobity server ( http://molprobity.biochem.duke.edu/ ) applied for Ramachandran analyses 79 – 81 . All process related to the prediction of structures and preparing structural library are presented in Fig. 2 . MD simulation and trajectory condition design The predicted 3D conformation of the protein was generated for MD simulation utilizing the CHARMM-GUI server ( https://www.charmm-gui.org ) 82 , configured under typical parameters, including solvation and ion neutralization where necessary. The system was parameterized using the CHARMM36m force field and transformed into GROMACS-compatible input files. Solvation was conducted utilizing the TIP3P water model 83 , and the system was neutralized with suitable counterions. The configured system was subsequently employed for MD simulations in GROMACS 2022, with production runs conducted for 200 nanoseconds, and 100M steps according to the duration of selected simulation and ensuring stabilization of soluble-related parameters and dependable time averaged data across root means square deviation (RMSD) and root means square fluctuation (RMSF) 84 , 85 . The simulations were conducted to examine the solubility behaviour and structural dynamics of the protein under physiological settings. In addition, MD simulation carried out in ionic conditions, sodium chloride (NaCl) as physiological ionic strength agent and ammonium sulfate ((NH 4 ) 2 SO 4 ) as protein participant agent in concentration 0.2, 0.6, 0.9 and 1.2 mol/L by manual setup of input files 86 – 89 . Also, for different pH condition behaviour, PDB files were adjusted with APBS-PDB2PQR server ( https://pdb2pqr.readthedocs.io ) 90 in range of 4, 5.5, 7 and 9. Assessment and scoring of single-point mutants for solubility improvement Following the initial 200 ns MD simulation, residue-level analyses were conducted to identify structurally unstable or flexible regions that could serve as potential mutant targets. Important metrics included RMSF to detect residues with high positional variability, solvent accessible surface area (SASA) to identify exposed areas that may affect solubility, and hydrogen bond (H-bond) occupancy and secondary structure stability analyses to determine structural integrity throughout the simulation. Residues demonstrating persistently elevated RMSF values, augmented solvent exposure, or unstable secondary structural configurations were identified as mutation hotspots. The substitution of residues aimed to generate a novel mutated sequence by replacing residues characterized by high flexibility and hydrophobicity with those of the same category that exhibit greater similarity and enhanced hydrophilicity 49 . In Table 1 , all applied mutations with replaced residues are presented. To preserve functional activity, all selected residues were visually examined using UCSF ChimeraX version 1.8 software, and only those situated on the protein's surface and distanced from the active site were selected for mutation. After 28 potential mutants were chosen based on RMSF, SASA, H-bond analyses, and the Ramachandran plot, they were all run through CamSol ( https://www-cohsoftware.ch.cam.ac.uk/ ) 91 , a server that calculates an intrinsic solubility profile and solubility gain upon mutagenesis, and Protein-Sol ( https://protein-sol.manchester.ac.uk/ ) 92 , a tool that estimates overall solubility propensity based on sequence and structural features. For re-prioritization structure and MD simulation, the top 10 highest scoring candidate models were chosen from all that were analysed and subsequent to the identification of the 10 premier mutant candidates, an exhaustive MD study was conducted on all associated protein structures. The evaluation encompassed RMSD, RMSF, SASA, numbers of H-bond, diffusion coefficient, radial distribution function (RDF), radius of gyration (Rg), and minimum distance (minDist) measurements. Finally assigning secondary structure to the residues (DSSP) carried out in order to final check and comparison of mutated structures against natural structure has not any structural modification. Also, on each cycle of mutant selection to library, structures with more than 5% changes in residue counts removed from selection and only structures with higher confidence than 95% with standard deviation of SE: ±0.05, and RMSD score less than 1.5 Å (highly similar structure against natural type) advanced to the final round of selection and sorted in library according to the DSSP analyses results. In addition, all selected structure of final library, computationally sequence alignment scored and visually matched against the reference structure (natural type). Ultimately, top scored triple mutants analysed in aspect of H-bond 93 (protein-solvent and intra-protein state), RMSD, RMSF, Rg and SASA (static mode and dynamic mode) 94 . $$\:\left(1\right)\:Weight\left(x\right)=\sum\:_{i=1}^{8}{w}_{i}.({f}_{i,\:mutant}-{f}_{i,\:\:natural})$$ Synergistic mutagenesis design; identifying optimal mutant’s combinational state We utilized a systematic combinatorial design to uncover optimal synergistic mutation combinations that enhance protein solubility, focusing on the top 10 single-point mutants previously selected by solubility-related MD characteristics. All conceivable double and triple combinations were selected and generated as described in follow, and for each, short MD simulations in 200 ns were conducted to derive solubility-relevant descriptors, including SASA, number of hydrogen bonds, radial distribution function (RDF), radius of gyration (Rg), diffusion coefficient, RMSD, RMSF, and minimum distance to solvent (minDist). in next, top 25 promising combination according to the complementary effects on solubility selected and finally a library consists of natural type, single and combined (double and triple combination) mutants generated for next phase of analyses to find the best possible combination. Each combination was assessed utilizing a weighted solubility scoring algorithm, with weights allocated to each attribute according to their established or presumed impact on solubility. The unprocessed solubility score was calculated utilizing: $$\:\left(2\right)\:Solubility\:Score=\sum\:_{i=1}^{8}{w}_{i}.{f}_{i}$$ where f i represents the normalized value of feature i and w i denotes the associated feature weight. Positive weights were allocated to qualities linked to enhanced solubility (e.g., SASA, H-bond, RDF, Rg, diffusion), whilst negative weights were designated to features related to structural instability (e.g., RMSD, RMSF, minDist). This synergistic strategy aimed to uncover mutation combinations that yield improved solubility effects surpassing those of single-point variants, therefore informing the design of optimal multi-point mutants with superior biophysical features. The ultimate solubility score was standardized to a 0–1 scale utilizing: (3) \(\:Scaled\:score=100\times\:\frac{\sum\:_{i=1}^{8}{w}_{i}.{f}_{i}-min\:score}{\text{max}score-\text{min}score}\) A reinforcement learning method called Proximal Policy Optimization (PPO) 95 used to quickly move through the combinatorial space and find the best mutation sets. To get the highest solubility score, PPO agent was taught to pick mutation combinations that work well together based on input from simulated MD-derived attributes. This integrated design approach makes it easier to choose multi-point mutants that are logical, based on data, and have solubility properties that are noticeably better than those possible with single mutations alone. Schematic diagram of applied method is presented in Fig. 3 . Declarations Funding This work is the result of implementation of Slovak Research and Development Agency grants APVV-21-0227, APVV-21-0215, APVV-22-0161 and by implementation of the project 101160008 “Fostering Excellence in Advanced Genomics and Proteomics Research at Comenius University in Bratislava – FORGENOM II” funded by the Horizon Europe program. Acknowledgments MD is thankful to Daniel Kráľ and Milan Melicherčík, Faculty of Mathematics, Physics, and Informatics (FMPI) for their kind recommendation and guidance. Authur Contributions MD and ZL: conceived and designed the study, MD: developed the structural prediction and mutagenesis workflow, and performed all molecular dynamics simulations and solubility analyses. Data processing, interpretation, and manuscript writing and editing were carried out by MD and ZL. ES: Provided recommendations and scientific guidance of research. All aspects of the research were conducted under the academic supervision of JT , VB and SS who provided critical feedback on the study design and manuscript. Competing interests The authors declare no competing interests. Data Availability The data generated and analysed during this study, including molecular dynamics simulation outputs, structural models, and solubility feature datasets, are available from the corresponding author upon reasonable request. Due to the size and computational nature of the datasets, they are not hosted in a public repository. References Villaverde, A. Mar Carrió, M. Protein aggregation in recombinant bacteria: biological role of inclusion bodies. Biotechnol. Lett. 25 , 1385–1395 (2003). Baneyx, F. & Mujacic, M. Recombinant protein folding and misfolding in Escherichia coli. Nat. Biotechnol. 22 , 1399–1408 (2004). Pantophlet, R., Wilson, I. A. & Burton, D. R. Improved design of an antigen with enhanced specificity for the broadly HIV-neutralizing antibody b12. Protein Eng. Des. Selection . 17 , 749–758 (2004). De Marco, A., Deuerling, E., Mogk, A., Tomoyasu, T. & Bukau, B. Chaperone-based procedure to increase yields of soluble recombinant proteins produced in E. coli. BMC Biotechnol 7 , (2007). Gustafsson, C., Govindarajan, S. & Minshull, J. Codon bias and heterologous protein expression. Trends Biotechnol. 22 , 346–353 (2004). Chatterjee, D. K. & Esposito, D. Enhanced soluble protein expression using two new fusion tags. Protein Exp. Purif. 46 , 122–129 (2006). Esposito, D. & Chatterjee, D. K. Enhancement of soluble protein expression through the use of fusion tags. Curr. Opin. Biotechnol. 17 , 353–358 (2006). Sachdev, D. & Chirgwin, J. M. Properties of Soluble Fusions Between Mammalian Aspartic Proteinases and Bacterial Maltose-Binding Protein. J. Protein Chem. 18 , 127–136 (1999). Agostini, F., Cirillo, D., Bolognesi, B. & Tartaglia, G. G. X-inactivation: quantitative predictions of protein interactions in the Xist network. Nucleic Acids Res. 41 , e31–e31 (2013). Kulshreshtha, S., Chaudhary, V., Goswami, G. K. & Mathur, N. Computational approaches for predicting mutant protein stability. J. Comput. Aided Mol. Des. 30 , 401–412 (2016). Damborsky, J. & Brezovsky, J. Computational tools for designing and engineering enzymes. Curr. Opin. Chem. Biol. 19 , 8–16 (2014). Ebert, M. C. & Pelletier, J. N. Computational tools for enzyme improvement: why everyone can – and should – use them. Curr. Opin. Chem. Biol. 37 , 89–96 (2017). Broom, A., Jacobi, Z., Trainor, K. & Meiering, E. M. Computational tools help improve protein stability but with a solubility tradeoff. J. Biol. Chem. 292 , 14349–14361 (2017). Childers, M. C. & Daggett, V. Insights from molecular dynamics simulations for computational protein design. Mol. Syst. Des. Eng. 2 , 9–33 (2017). Rouhani, M., Khodabakhsh, F., Norouzian, D., Cohan, R. A. & Valizadeh, V. Molecular dynamics simulation for rational protein engineering: Present and future prospectus. J. Mol. Graph. Model. 84 , 43–53 (2018). Pikkemaat, M. G., Linssen, A. B. M., Berendsen, H. J. C. & Janssen, D. B. Molecular dynamics simulations as a tool for improving protein stability. Protein Eng. Des. Selection . 15 , 185–192 (2002). Carballo-Amador, M. A., McKenzie, E. A., Dickson, A. J. & Warwicker, J. Surface patches on recombinant erythropoietin predict protein solubility: engineering proteins to minimise aggregation. BMC Biotechnol 19 , (2019). Kumar, S., Kumar Bhardwaj, V., Singh, R. & Purohit, R. Explicit-solvent molecular dynamics simulations revealed conformational regain and aggregation inhibition of I113T SOD1 by Himalayan bioactive molecules. J. Mol. Liq. 339 , 116798 (2021). Chennamsetty, N., Voynov, V., Kayser, V., Helk, B. & Trout, B. L. Prediction of Aggregation Prone Regions of Therapeutic Proteins. J. Phys. Chem. B . 114 , 6614–6624 (2010). Agrawal, N. J. et al. Aggregation in Protein-Based Biotherapeutics: Computational Studies and Tools to Identify Aggregation-Prone Regions. J. Pharm. Sci. 100 , 5081–5095 (2011). Johnson, K. L., Faulkner, C., Jeffree, C. E. & Ingram, G. C. The Phytocalpain Defective Kernel 1 Is a Novel Arabidopsis Growth Regulator Whose Activity Is Regulated by Proteolytic Processing. Plant. Cell. 20 , 2619–2630 (2008). Demko, V. et al. Genetic Analysis of DEFECTIVE KERNEL1 Loop Function in Three-Dimensional Body Patterning in Physcomitrella patens . Plant Physiol. 166 , 903–919 (2014). Lid, S. E. et al. The defective kernel 1 ( dek1 ) gene required for aleurone cell development in the endosperm of maize grains encodes a membrane protein of the calpain gene superfamily. Proc. Natl. Acad. Sci. U.S.A. 99, 5460–5465 (2002). Olsen, O. A., Perroud, P. F., Johansen, W. & Demko, V. DEK1; missing piece in puzzle of plant development. Trends Plant Sci. 20 , 70–71 (2015). Johansen, W. et al. The DEK1 calpain Linker functions in three-dimensional body patterning in Physcomitrella patens. Plant Physiol. pp.00925. (2016) (2016). 10.1104/pp.16.00925 Demko, V. et al. Regulation of developmental gatekeeping and cell fate transition by the calpain protease DEK1 in Physcomitrium patens. Commun. Biol. 7 , 261 (2024). Ako, A. E. et al. An intragenic mutagenesis strategy in Physcomitrella patens to preserve intron splicing. Sci. Rep. 7 , 5111 (2017). Perroud, P. et al. Defective Kernel 1 (DEK 1) is required for three-dimensional growth in P hyscomitrella patens . New Phytol. 203 , 794–804 (2014). Navarro, S. & Ventura, S. Computational re-design of protein structures to improve solubility. Expert Opin. Drug Discov. 14 , 1077–1088 (2019). Trainor, K., Broom, A. & Meiering, E. M. Exploring the relationships between protein sequence, structure and solubility. Curr. Opin. Struct. Biol. 42 , 136–146 (2017). Gupta, J., Nunes, C., Vyas, S. & Jonnalagadda, S. Prediction of Solubility Parameters and Miscibility of Pharmaceutical Compounds by Molecular Dynamics Simulations. J. Phys. Chem. B . 115 , 2014–2023 (2011). Ganugapati, J. & Akash, S. Multi-template homology based structure prediction and molecular docking studies of protein ‘L’ of Zaire ebolavirus (EBOV). Inf. Med. Unlocked . 9 , 68–75 (2017). Lu, H., Cheng, Z., Hu, Y. & Tang, L. V. What Can De Novo Protein Design Bring to the Treatment of Hematological Disorders? Biology 12, 166 (2023). Roy, A., Yang, J. & Zhang, Y. COFACTOR: an accurate comparative algorithm for structure-based protein function annotation. Nucleic Acids Res. 40 , W471–W477 (2012). Holm, L., Laiho, A., Törönen, P. & Salgado, M. DALI shines a light on remote homologs: One hundred discoveries. Protein Sci. 32 , e4519 (2023). Chan, P., Curtis, R. A. & Warwicker, J. Soluble expression of proteins correlates with a lack of positively-charged surface. Sci. Rep. 3 , 3333 (2013). Studer, G. et al. QMEANDisCo—distance constraints applied on model quality estimation. Bioinformatics 36 , 1765–1771 (2020). Van Den Bedem, H. & Fraser, J. S. Integrative, dynamic structural biology at atomic resolution—it’s about time. Nat. Methods . 12 , 307–318 (2015). Heo, L. & Feig, M. High-accuracy protein structures by combining machine‐learning with physics‐based refinement. Proteins 88 , 637–642 (2020). Lobanov, M. Y., Bogatyreva, N. S. & Galzitskaya, O. V. Radius of gyration as an indicator of protein structure compactness. Mol. Biol. 42 , 623–628 (2008). Abouzied, A. S. et al. Structural and free energy landscape analysis for the discovery of antiviral compounds targeting the cap-binding domain of influenza polymerase PB2. Sci. Rep. 14 , 25441 (2024). Pace, C. N. et al. Contribution of hydrogen bonds to protein stability. Protein Sci. 23 , 652–661 (2014). Jiang, L. & Lai, L. CH···O Hydrogen Bonds at Protein-Protein Interfaces. J. Biol. Chem. 277 , 37732–37740 (2002). Tsumoto, K. et al. Role of Arginine in Protein Refolding, Solubilization, and Purification. Biotechnol. Prog . 20 , 1301–1308 (2004). Strub, C. et al. Mutation of exposed hydrophobic amino acids to arginine to increase protein stability. BMC Biochem. 5 , 9 (2004). Warwicker, J., Charonis, S. & Curtis, R. A. Lysine and Arginine Content of Proteins: Computational Analysis Suggests a New Tool for Solubility Design. Mol. Pharm. 11 , 294–303 (2014). Kramer, R. M., Shende, V. R., Motl, N., Pace, C. N. & Scholtz, J. M. Toward a Molecular Understanding of Protein Solubility: Increased Negative Surface Charge Correlates with Increased Solubility. Biophys. J. 102 , 1907–1915 (2012). Mills, B. J. & Laurence Chadwick, J. S. Effects of localized interactions and surface properties on stability of protein-based therapeutics. J. Pharm. Pharmacol. 70 , 609–624 (2018). Trevino, S. R., Scholtz, J. M. & Pace, C. N. Measuring and Increasing Protein Solubility. J. Pharm. Sci. 97 , 4155–4166 (2008). Kuhn, A. B. et al. Improved Solution-State Properties of Monoclonal Antibodies by Targeted Mutations. J. Phys. Chem. B . 121 , 10818–10827 (2017). Ghahremanian, S., Rashidi, M. M., Raeisi, K. & Toghraie, D. Molecular dynamics simulation approach for discovering potential inhibitors against SARS-CoV-2: A structural review. J. Mol. Liq. 354 , 118901 (2022). Xiao, S. et al. Rational modification of protein stability by targeting surface sites leads to complicated results. Proc. Natl. Acad. Sci. U.S.A. 110, 11337–11342 (2013). Kabsch, W. & Sander, C. Dictionary of protein secondary structure: Pattern recognition of hydrogen-bonded and geometrical features. Biopolymers 22 , 2577–2637 (1983). Mokmak, W., Chunsrivirot, S., Assawamakin, A., Choowongkomon, K. & Tongsima, S. Molecular dynamics simulations reveal structural instability of human trypsin inhibitor upon D50E and Y54H mutations. J. Mol. Model. 19 , 521–528 (2013). Zhou, H. X. & Pang, X. Electrostatic Interactions in Protein Structure, Folding, Binding, and Condensation. Chem. Rev. 118 , 1691–1741 (2018). Gregory, K. P. et al. Understanding specific ion effects and the Hofmeister series. Phys. Chem. Chem. Phys. 24 , 12682–12718 (2022). Hyde, A. M. et al. General Principles and Strategies for Salting-Out Informed by the Hofmeister Series. Org. Process. Res. Dev. 21 , 1355–1370 (2017). Tadeo, X., López-Méndez, B., Castaño, D., Trigueros, T. & Millet, O. Protein Stabilization and the Hofmeister Effect: The Role of Hydrophobic Solvation. Biophys. J. 97 , 2595–2603 (2009). Tadeo, X., Pons, M. & Millet, O. Influence of the Hofmeister Anions on Protein Stability As Studied by Thermal Denaturation and Chemical Shift Perturbation. Biochemistry 46 , 917–923 (2007). Sammond, D. W. et al. Structure-based Protocol for Identifying Mutations that Enhance Protein–Protein Binding Affinities. J. Mol. Biol. 371 , 1392–1404 (2007). Xu, D., Tsai, C. J. & Nussinov, R. Hydrogen bonds and salt bridges across protein-protein interfaces. Protein Eng. Des. Selection . 10 , 999–1012 (1997). Goldenzweig, A. & Fleishman, S. J. Principles of Protein Stability and Their Application in Computational Design. Annu. Rev. Biochem. 87 , 105–129 (2018). Rosano, G. L. & Ceccarelli, E. A. Recombinant protein expression in Escherichia coli: advances and challenges. Front Microbiol 5 , (2014). Zheng, W. et al. Folding non-homologous proteins by coupling deep-learning contact maps with I-TASSER assembly simulations. Cell. Rep. Methods . 1 , 100014 (2021). Zhang, C., Freddolino, L. & Zhang, Y. COFACTOR: improved protein function prediction by combining structure, sequence and protein–protein interaction information. Nucleic Acids Res. 45 , W291–W299 (2017). Šali, A. & Blundell, T. L. Comparative Protein Modelling by Satisfaction of Spatial Restraints. J. Mol. Biol. 234 , 779–815 (1993). Fiser, A., Šali, A. & Modeller Generation and Refinement of Homology-Based Protein Structure Models. in Methods in Enzymology vol. 374 461–491 (Elsevier, (2003). Jumper, J. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596 , 583–589 (2021). Mirdita, M. et al. ColabFold: making protein folding accessible to all. Nat. Methods . 19 , 679–682 (2022). Waterhouse, A. M. et al. The structure assessment web server: for proteins, complexes and more. Nucleic Acids Res. 52 , W318–W323 (2024). Zheng, W. et al. Deep-learning-based single-domain and multidomain protein structure prediction with D-I-TASSER. Nat. Biotechnol. 10.1038/s41587-025-02654-4 (2025). Ko, J., Park, H., Heo, L. & Seok, C. GalaxyWEB server for protein structure prediction and refinement. Nucleic Acids Res. 40 , W294–W297 (2012). Fiser, A. & Sali, A. ModLoop: automated modeling of loops in protein structures. Bioinformatics 19 , 2500–2501 (2003). Lüthy, R., Bowie, J. U. & Eisenberg, D. Assessment of protein models with three-dimensional profiles. Nature 356 , 83–85 (1992). Zhang, C., Shine, M., Pyle, A. M. & Zhang, Y. US-align: universal structure alignments of proteins, nucleic acids, and macromolecular complexes. Nat. Methods . 19 , 1109–1115 (2022). Xu, J. & Zhang, Y. How significant is a protein structure similarity with TM-score = 0.5? Bioinformatics 26 , 889–895 (2010). Wiederstein, M. & Sippl, M. J. ProSA-web: interactive web service for the recognition of errors in three-dimensional structures of proteins. Nucleic Acids Res. 35 , W407–W410 (2007). Uziela, K., Menéndez Hurtado, D., Shu, N., Wallner, B. & Elofsson, A. ProQ3D: improved model quality assessments using deep learning. Bioinformatics 33 , 1578–1580 (2017). Ramachandran, G. N., Ramakrishnan, C. & Sasisekharan, V. Stereochemistry of polypeptide chain configurations. J. Mol. Biol. 7 , 95–99 (1963). Lovell, S. C. et al. Structure validation by Cα geometry: ϕ,ψ and Cβ deviation. Proteins 50 , 437–450 (2003). Williams, C. J. et al. MolProbity: More and better reference data for improved all-atom structure validation. Protein Sci. 27 , 293–315 (2018). Brooks, B. R. et al. The biomolecular simulation program. J. Comput. Chem. 30 , 1545–1614 (2009). MacKerell, A. D. et al. All-Atom Empirical Potential for Molecular Modeling and Dynamics Studies of Proteins. J. Phys. Chem. B . 102 , 3586–3616 (1998). Lee, J. et al. CHARMM-GUI Input Generator for NAMD, GROMACS, AMBER, OpenMM, and CHARMM/OpenMM Simulations Using the CHARMM36 Additive Force Field. J. Chem. Theory Comput. 12 , 405–413 (2016). Bauer, P., Hess, B. & Lindahl, E. GROMACS 2022 Source code. Zenodo https://doi.org/10.5281/ZENODO.6103835 (2022). Galm, L., Amrhein, S. & Hubbuch, J. Predictive approach for protein aggregation: Correlation of protein surface characteristics and conformational flexibility to protein aggregation propensity. Biotech Bioengineering . 114 , 1170–1183 (2017). Soares, C. M., Teixeira, V. H. & Baptista, A. M. Protein Structure and Dynamics in Nonaqueous Solvents: Insights from Molecular Dynamics Simulation Studies. Biophys. J. 84 , 1628–1641 (2003). Friedman, R., Nachliel, E. & Gutman, M. Molecular Dynamics of a Protein Surface: Ion-Residues Interactions. Biophys. J. 89 , 768–781 (2005). Zhang, Y. & Cremer, P. Interactions between macromolecules and ions: the Hofmeister series. Curr. Opin. Chem. Biol. 10 , 658–663 (2006). Jurrus, E. et al. Improvements to the APBS biomolecular solvation software suite. Protein Sci. 27 , 112–128 (2018). Sormanni, P., Aprile, F. A. & Vendruscolo, M. The CamSol Method of Rational Design of Protein Mutants with Enhanced Solubility. J. Mol. Biol. 427 , 478–490 (2015). Hebditch, M., Carballo-Amador, M. A., Charonis, S., Curtis, R. & Warwicker, J. Protein–Sol: a web tool for predicting protein solubility from sequence. Bioinformatics 33 , 3098–3100 (2017). Van Der Spoel, D., Van Maaren, P. J., Larsson, P. & Tîmneanu, N. Thermodynamics of Hydrogen Bonding in Hydrophilic and Hydrophobic Media. J. Phys. Chem. B . 110 , 4393–4398 (2006). van der Bondi, A. Waals Volumes and Radii. J. Phys. Chem. 68 , 441–451 (1964). Schulman, J., Wolski, F., Dhariwal, P., Radford, A. & Klimov, O. Proximal Policy Optimization Algorithms. Preprint at (2017). https://doi.org/10.48550/ARXIV.1707.06347 Tables Table 1: Benchmarking C-I-TASSER, AlphaFold2, and SWISS-MODEL based on structural accuracy scores. Modelling software QMEANDisCo Global TM-Score US-Score Z-Score ProQ3D C-I-ITASSER 0.59±0.05 0.70 0.70 -9.64 0.701 Alphafold2 0.69±0.05 0.72 0.73 -6.57 0.713 SWISS-MODEL 0.64±0.05 0.66 0.68 -.8.51 0.688 Table. 2: Candidate residues and position number and replaced residues for single-pointed mutants. Mutant Original residue position mutated residue Structural type MUT1 Valine 17 Arginine α-Helix MUT2 Leucine 19 Serine α-Helix MUT3 Isoleucine 21 Serine α-Helix MUT4 Valine 47 Arginine Loop MUT5 Phenylalanine 291 Tyrosine Loop MUT6 Leucine 323 Serine α-Helix MUT7 Valine 342 Asparagine Loop MUT8 Lysine 7 Arginine α-Helix MUT9 Glutamine 54 Arginine α-Helix MUT10 Lysine 339 Arginine Loop Table 3: Fitness-based ranking of protein mutants across mutation levels (description: mutant identifiers follow the format: MUTxyz, where each digit (x, y, z) represents the point of a mutation in the protein sequence. For example, MUT347 indicates mutations includes single mutants of 3, 4, and 7 (as triple mutant). Rank Single Mutant Fitness Double Mutant Fitness Triple Mutant Fitness 1 MUT3 0.1622 MUT37 0.1432 MUT347 01368 2 MUT7 0.1719 MUT34 0.1443 MUT246 0.1475 3 MUT2 0.1759 MUT67 0.1498 MUT367 0.1498 4 MUT5 0.1775 MUT35 0.1519 MUT345 0.1519 5 MUT4 0.1803 MUT24 0.1696 MUT245 0.1705 6 MUT6 0.1850 MUT25 0.1705 MUT136 0.1776 7 MUT1 0.2163 MUT26 0.1712 MUT247 0.1909 8 MUT8 0.2184 MUT36 0.1741 MUT236 0.1950 9 MUT10 0.2209 MUT45 0.1775 MUT457 0.2101 10 MUT9 0.2226 MUT13 0.1776 MUT146 0.2551 Additional Declarations No competing interests reported. Cite Share Download PDF Status: Published Journal Publication published 08 Feb, 2026 Read the published version in Scientific Reports → Version 1 posted Editorial decision: Revision requested 30 Oct, 2025 Reviews received at journal 28 Oct, 2025 Reviews received at journal 24 Oct, 2025 Reviews received at journal 19 Oct, 2025 Reviewers agreed at journal 08 Oct, 2025 Reviewers agreed at journal 06 Oct, 2025 Reviewers agreed at journal 03 Oct, 2025 Reviewers agreed at journal 03 Oct, 2025 Reviewers invited by journal 22 Sep, 2025 Editor invited by journal 12 Sep, 2025 Editor assigned by journal 10 Sep, 2025 Submission checks completed at journal 09 Sep, 2025 First submitted to journal 07 Sep, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7559073","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":523981617,"identity":"99cd4389-470b-42d8-bc3d-aeacb0645e74","order_by":0,"name":"Mohammad Dabiri","email":"","orcid":"","institution":"Comenius University in Bratislava","correspondingAuthor":false,"prefix":"","firstName":"Mohammad","middleName":"","lastName":"Dabiri","suffix":""},{"id":523981621,"identity":"b51928ff-6d4e-4deb-b289-671c0610e41b","order_by":1,"name":"Zdenko Levarski","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABCUlEQVRIiWNgGAWjYFCCBAbGBgYGGQjngIQcAwNjG4hpQEgLD0yLMclaGBKBHDa8WuTbk489nFHBwMMvkfzs448zFukbbje3PfjBUGeMSwtjz7N0ww1nGHgkZ6QZz+a5IZG74c7BdsMehsNmuLQwS+SYST5sY+AxOHPAmJnhA1DLjcQ2CR6GAza4tLBJ5H8Da7E/c/wz448PEukGQC2SfxjqcGrhkchhk9wIsoW9x5gB6LAEkBZpHgZmnA6T4HlmJjnjjASPxPGeYmaeMxKGM28kthvLGBzG6X1giD2T7KmwkeNvZt/M+ONYnTzfjfRnD99U1Bk24NIDtQxdAE9EjoJRMApGwSggDAB7slOwjt0ezgAAAABJRU5ErkJggg==","orcid":"","institution":"Comenius University in Bratislava","correspondingAuthor":true,"prefix":"","firstName":"Zdenko","middleName":"","lastName":"Levarski","suffix":""},{"id":523981622,"identity":"b557a35a-0f52-480c-a873-990c4c53da1d","order_by":2,"name":"Eva Struharnanska","email":"","orcid":"","institution":"Comenius University in Bratislava","correspondingAuthor":false,"prefix":"","firstName":"Eva","middleName":"","lastName":"Struharnanska","suffix":""},{"id":523981625,"identity":"0dff7b3f-5b27-4dcd-9981-52ee48f7f94e","order_by":3,"name":"Viktor Demko","email":"","orcid":"","institution":"Comenius University in Bratislava","correspondingAuthor":false,"prefix":"","firstName":"Viktor","middleName":"","lastName":"Demko","suffix":""},{"id":523981630,"identity":"488b7e6a-5c1f-442d-8813-dc055b1e6f26","order_by":4,"name":"Vladimir Benes","email":"","orcid":"","institution":"EMBL Heidelberg","correspondingAuthor":false,"prefix":"","firstName":"Vladimir","middleName":"","lastName":"Benes","suffix":""},{"id":523981634,"identity":"207db6b7-61e0-49c2-90b5-e42157aa9ea9","order_by":5,"name":"Jan Turna","email":"","orcid":"","institution":"Comenius University in Bratislava","correspondingAuthor":false,"prefix":"","firstName":"Jan","middleName":"","lastName":"Turna","suffix":""},{"id":523981639,"identity":"f75eb22e-1988-406f-b5ff-6e2e5b44d263","order_by":6,"name":"Stanislav Stuchlík","email":"","orcid":"","institution":"Comenius University in Bratislava","correspondingAuthor":false,"prefix":"","firstName":"Stanislav","middleName":"","lastName":"Stuchlík","suffix":""}],"badges":[],"createdAt":"2025-09-08 01:53:20","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7559073/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7559073/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1038/s41598-026-38805-z","type":"published","date":"2026-02-08T15:58:44+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":92725753,"identity":"39926f31-394f-4cb7-8d8a-8e1d9c294358","added_by":"auto","created_at":"2025-10-03 14:32:59","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":107793,"visible":true,"origin":"","legend":"","description":"","filename":"Manuscript.docx","url":"https://assets-eu.researchsquare.com/files/rs-7559073/v1/c9e237e3416f227a4b9b2aaf.docx"},{"id":92725759,"identity":"a4862139-ceaf-4ef1-9b89-1efabd1d08a5","added_by":"auto","created_at":"2025-10-03 14:33:00","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":1836001,"visible":true,"origin":"","legend":"","description":"","filename":"FiguresandTables.docx","url":"https://assets-eu.researchsquare.com/files/rs-7559073/v1/befe6faec435e0972c5a6599.docx"},{"id":92725751,"identity":"18ec6eae-d72e-46ef-9eff-6d91dba916a0","added_by":"auto","created_at":"2025-10-03 14:32:59","extension":"json","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":8388,"visible":true,"origin":"","legend":"","description":"","filename":"f6099fa58d5f4856885ec668c5697556.json","url":"https://assets-eu.researchsquare.com/files/rs-7559073/v1/077c2f6a01b502c20da596ff.json"},{"id":92726581,"identity":"c0529401-001d-4214-801a-ad970e3b33e2","added_by":"auto","created_at":"2025-10-03 14:49:00","extension":"xml","order_by":3,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":197036,"visible":true,"origin":"","legend":"","description":"","filename":"f6099fa58d5f4856885ec668c56975561enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-7559073/v1/f37052357acd9840a33b00b8.xml"},{"id":92725764,"identity":"bac98b9e-3b71-44d6-a188-241cc5e6f095","added_by":"auto","created_at":"2025-10-03 14:33:00","extension":"jpeg","order_by":4,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":271706,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7559073/v1/3b78324f3eadf8810c8376ba.jpeg"},{"id":92726290,"identity":"92827334-07ed-41c7-ad4e-8172b6b36906","added_by":"auto","created_at":"2025-10-03 14:41:00","extension":"jpeg","order_by":5,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":237178,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7559073/v1/f6a6a5fb211eb96e300b8516.jpeg"},{"id":92725766,"identity":"f7e5f878-6792-4ec6-8afc-82d65ff8e405","added_by":"auto","created_at":"2025-10-03 14:33:00","extension":"jpeg","order_by":6,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":110904,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage3.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7559073/v1/37b78fc6ca8709c1b51fdd53.jpeg"},{"id":92725761,"identity":"76d3f14a-5acf-45d3-8a76-48f3dc9fe198","added_by":"auto","created_at":"2025-10-03 14:33:00","extension":"jpeg","order_by":7,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":83248,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage4.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7559073/v1/1b6fa9913d2107bda1131b30.jpeg"},{"id":92726295,"identity":"41acf036-c21f-47b2-aee7-e7c025b60a55","added_by":"auto","created_at":"2025-10-03 14:41:00","extension":"jpeg","order_by":8,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":394788,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage5.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7559073/v1/473052ce0c72c27cda3d03b2.jpeg"},{"id":92725775,"identity":"4db1e279-16de-4927-86a2-f1338c3175a8","added_by":"auto","created_at":"2025-10-03 14:33:00","extension":"jpeg","order_by":9,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":597500,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage6.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7559073/v1/a583d64d4e2be40b3a57c0b5.jpeg"},{"id":92726293,"identity":"858d41d5-8fd5-4be5-95fb-51f8daed86f7","added_by":"auto","created_at":"2025-10-03 14:41:00","extension":"jpeg","order_by":10,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":110182,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage7.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7559073/v1/cf124ea4678bca5a1c94cc56.jpeg"},{"id":92726292,"identity":"32ed9150-9276-480d-8358-303da5d04fbf","added_by":"auto","created_at":"2025-10-03 14:41:00","extension":"png","order_by":11,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":53116,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-7559073/v1/8f26391646b16e59f54ba192.png"},{"id":92725756,"identity":"d3dd0af1-4450-44e7-aa55-00dbc4a3ccbe","added_by":"auto","created_at":"2025-10-03 14:32:59","extension":"png","order_by":12,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":38904,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-7559073/v1/3498c9b42de4f2592af2266a.png"},{"id":92726297,"identity":"415865df-738e-4082-b8e4-bdf2a2d9d686","added_by":"auto","created_at":"2025-10-03 14:41:00","extension":"png","order_by":13,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":34802,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-7559073/v1/9f1a298f7cb2f67d512c67ab.png"},{"id":92725774,"identity":"df7ec16e-1d29-4be4-a42f-8de18fc89539","added_by":"auto","created_at":"2025-10-03 14:33:00","extension":"png","order_by":14,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":13554,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-7559073/v1/d3c758e6c99939b0d546d539.png"},{"id":92725768,"identity":"4e8333bb-9874-4d1f-b818-1a04fd62151c","added_by":"auto","created_at":"2025-10-03 14:33:00","extension":"png","order_by":15,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":63752,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-7559073/v1/8dc743489215bf035da3c483.png"},{"id":92725769,"identity":"5fabd9c6-cdf6-4083-a0bb-204eb15d37b1","added_by":"auto","created_at":"2025-10-03 14:33:00","extension":"png","order_by":16,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":92880,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-7559073/v1/02bdfa67be7d290f3833fe3d.png"},{"id":92725772,"identity":"b8b105cb-e6a7-4323-a489-0510227e125b","added_by":"auto","created_at":"2025-10-03 14:33:00","extension":"png","order_by":17,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":18688,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage7.png","url":"https://assets-eu.researchsquare.com/files/rs-7559073/v1/cc0a1f4e6123ca6df31ad00d.png"},{"id":92725773,"identity":"16792717-9c89-42dd-a4d4-48537f7618df","added_by":"auto","created_at":"2025-10-03 14:33:00","extension":"xml","order_by":18,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":196967,"visible":true,"origin":"","legend":"","description":"","filename":"f6099fa58d5f4856885ec668c56975561structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-7559073/v1/f7b5421061fdd27cc6eac40c.xml"},{"id":92725771,"identity":"d1bcaca5-5d2d-4b24-a0a1-0279c3ce198d","added_by":"auto","created_at":"2025-10-03 14:33:00","extension":"html","order_by":19,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":212924,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-7559073/v1/909f587bea07a2d3be73f7b8.html"},{"id":92726288,"identity":"61d0cde9-d4e5-4aa3-91aa-48cb5090caf2","added_by":"auto","created_at":"2025-10-03 14:40:59","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":194567,"visible":true,"origin":"","legend":"\u003cp\u003eprocess related to the finding best reference sequence and integrated homology modelling of CysPc domain and library preparation.\u003c/p\u003e","description":"","filename":"image1.png","url":"https://assets-eu.researchsquare.com/files/rs-7559073/v1/ee8dc331ba06b7d5c6d5e5cf.png"},{"id":92726287,"identity":"027fcca8-b1fb-4606-abc6-8441387920d3","added_by":"auto","created_at":"2025-10-03 14:40:59","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":167169,"visible":true,"origin":"","legend":"\u003cp\u003eRamachandran, pLDDT and QMEAN diagram of Cyspc natural protein in A: integrated homology method and B: direct threading method.\u003c/p\u003e","description":"","filename":"image2.png","url":"https://assets-eu.researchsquare.com/files/rs-7559073/v1/2ae8db2b8f703eee9cb458d7.png"},{"id":92725749,"identity":"2bf3b25a-bc66-4a84-914f-c245535def17","added_by":"auto","created_at":"2025-10-03 14:32:59","extension":"jpeg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":110904,"visible":true,"origin":"","legend":"\u003cp\u003eWorkflow for multi-point mutation protein optimization using MD features and iterative training. This diagram illustrates a stepwise approach for optimizing protein solubility (or other properties) via single, double, and triple mutants, guided by MD simulation features and iterative learning.\u003c/p\u003e","description":"","filename":"image3.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7559073/v1/c1fdfd21318c9d375fa0efcb.jpeg"},{"id":92725754,"identity":"7fa90185-ef1e-4329-aaba-057bdbc394b9","added_by":"auto","created_at":"2025-10-03 14:32:59","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":26503,"visible":true,"origin":"","legend":"\u003cp\u003eA: MD simulation RMSF plot per residue for the natural-type (green) and mutant (red) proteins. Most residues show comparable fluctuation patterns in both profiles, indicating local flexibility. Both variants have slightly greater RMSF values at the C-terminal region, with the natural-type showing slightly more flexibility. Some loop areas of the mutated protein have much less variation, which suggests that changes to the surface made the protein more stable in those areas. B: RMSD of backbone atoms over time (ns) for proteins that are normally formed (green) and proteins that have been changed (red). The RMSD of the changed structure stays lower than that of the natural-type structure after it is stabilized quickly. The smaller RMSD of the changed protein supports the idea that certain surface changes make structures more compact and less likely to change. C: Rg analysis between natural-type and mutated structures reveals better compactness and integrity all over trajectory.\u003c/p\u003e","description":"","filename":"image4.png","url":"https://assets-eu.researchsquare.com/files/rs-7559073/v1/352a39e241f55fe8700dc832.png"},{"id":92726289,"identity":"495e69a9-e37a-4348-8428-4b4efadc7b9b","added_by":"auto","created_at":"2025-10-03 14:40:59","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":338524,"visible":true,"origin":"","legend":"\u003cp\u003eComparison of the global folded structures and secondary structures of the native protein and the triple-mutant variant. A: Structure alignment of natural type (green) and triple-mutated (red) type and superimposed structure (middle) of CysPc domain with RMSD of 1.174 Å. B: mdmat analyses of natural-type and MUT347 in residue-residue level.\u0026nbsp; \u0026nbsp;C: time resolved DSSP secondary structures of natural and triple-mutated type in 200 ns time-laps. D: Quantitative analysis of secondary structure composition over blocks of 50 ps; block averaging with SE±0.05. Colour coded as follows: coil (white), β-sheet (red), β-bridge (black), bend (green), turn (yellow), α-helix (blue) 3\u003csub\u003e10\u003c/sub\u003e-helix (dark gray), chain separator (light gray); Natural type (patterned bars), Mutated type (solid bars).\u003c/p\u003e","description":"","filename":"image5.png","url":"https://assets-eu.researchsquare.com/files/rs-7559073/v1/27f5d9c5c4228ff1780bbf53.png"},{"id":92727648,"identity":"690edef1-d085-4cb3-adcd-2403aff252b4","added_by":"auto","created_at":"2025-10-03 14:57:04","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":525743,"visible":true,"origin":"","legend":"\u003cp\u003eStructural comparison and surface study relates to solubility of the natural-type and mutant protein. A: Surface and water molecules distribution representations illustrate mutation-induced structural alterations (represent in yellow), with the mutant variant displaying enhanced intra-molecular hydrogen bonding, as assessed in the accompanying bar chart. B: SASA mapping by residue and area indicates that the mutant has a higher static surface exposure (21108 Ų) than the natural type (17821 Ų). Dynamic SASA analysis, however, reveals that the mutant sustains a slightly more compact conformation but better stability in finals steps.\u003c/p\u003e","description":"","filename":"image6.png","url":"https://assets-eu.researchsquare.com/files/rs-7559073/v1/c5524b27cf66b75c1181478f.png"},{"id":92726596,"identity":"a7ff8a93-e50e-42a9-8c3c-47e71f5c2cd0","added_by":"auto","created_at":"2025-10-03 14:49:04","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":69594,"visible":true,"origin":"","legend":"\u003cp\u003eEffect of pH and ionic strength on backbone flexibility of the Natural-type protein and the Mut347 variant. A: pH series. Per-residue Cα RMSF profiles for Natural-type (left) and Mut347 (right) across four pH conditions spanning acidic to basic. Line colours follow the top colour bar (blue to red = low high pH). Residue index is on the x-axis; RMSF (Å) is on the y-axis. B: NaCl series. RMSF profiles at increasing NaCl concentrations (0.2–1.2 M). Line colours map to the middle colour bar (black to green = low to high salt). Both variants show broadly similar shapes, with only modest salt-dependent dampening of fluctuations in solvent-exposed regions. C: (NH₄)\u003csub\u003e₂\u003c/sub\u003eSO₄ series. RMSF profiles at increasing (NH₄)\u003csub\u003e₂\u003c/sub\u003eSO₄ concentrations (0.2–1.2 M). Colores follow the bottom colour bar (black to green).\u003c/p\u003e","description":"","filename":"image7.png","url":"https://assets-eu.researchsquare.com/files/rs-7559073/v1/5a2213f3831ca1ca4578deed.png"},{"id":102234246,"identity":"acfb4152-91e5-4909-8cb2-9064b0d60a29","added_by":"auto","created_at":"2026-02-09 16:08:25","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2201449,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7559073/v1/140c293e-2413-46f8-9336-f45bac09296b.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Computational optimization of DEK1 calpain domain solubility through integrated structural modelling and targeted mutagenesis","fulltext":[{"header":"Introduction","content":"\u003cp\u003eThe high solubility and native structure of proteins is a desired outcome of their production in heterologous hosts, particularly \u003cem\u003eEscherichia coli\u003c/em\u003e with its well-known drawbacks such as restricted post-translational modification capability, recurrent development of inclusion bodies, and the propensity to generate insoluble or misfolded proteins under overexpression conditions. Low solubility can generate inclusion bodies, which seriously impede structural studies, functional tests, and protein recovery \u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u003c/sup\u003e. Clearly showing this issue for eukaryotic and membrane-associated proteins produced in bacterial systems are variations in codon usage, lack of post-translational modifications, and predisposition for misfolding in complex protein domains \u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e. In both academic and industrial biotechnology, protein insolubility thus presents a major obstacle that calls for the development of methodical and successful solutions to improve solubility.\u003c/p\u003e\u003cp\u003eStructure determination of eukaryotic calpain proteases has been challenging due to insolubility and aggregation issues. Currently, the structure of mammalian calpains has been solved (Refs). Calpains however represent diverse family with neither the 3D structure nor the biological role determined for majority of them (Zhao et al., 2012). The solubility issue is particularly relevant to the study of DEFECTIVE KERNEL1 (DEK1), a large, multi-domain membrane protein. DEK1 is a 240 kDa protein with a 23-spanning transmembrane domain, a cytosolic linker segment, and a C-terminal calpain protease\u003csup\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u003c/sup\u003e CysPc. Genetic analyses in cereals and diverse model plant species showed that DEK1 plays essential functions throughout plant development via regulation of meristematic cell divisions and cell fate control \u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e,\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e\u003c/sup\u003e. In addition, point mutagenesis and genetic complementation studies indicated that the calpain moiety of DEK1 is essential for the protein function and represents the protein\u0026rsquo;s effector \u003csup\u003e\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e\u003c/sup\u003e. Despite its essential function, the 3D structure of DEK1 has not been solved yet. Consequently, performing mutagenesis studies on CycPc\u0026rsquo;s structure is particularly challenging due to the absence of crystallographic data or any experimentally determined structures in public databases, as this limits the accuracy of structural predictions and the rational design of beneficial mutations.\u003c/p\u003e\u003cp\u003eTraditional methods for addressing this difficulty encompass empirical mutagenesis \u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e, co-expression with molecular chaperones \u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e,\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e, codon optimization \u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e, and fusion with solubility-enhancing tags \u003csup\u003e\u003cspan additionalcitationids=\"CR7\" citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e. Nevertheless, these processes are frequently protracted, laborious, and contingent upon circumstances. In recent years, computational approaches have emerged as valuable tools for targeted mutagenesis aimed at enhancing protein solubility \u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e. Still, these processes are often time-consuming, labour-intensive, dependent on circumstance. Recent years have seen computational methods become useful tools for targeted mutagenesis to increase protein solubility \u003csup\u003e\u003cspan additionalcitationids=\"CR11 CR12\" citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e. By use of structural modelling, molecular dynamics (MD) simulations, and machine learning-generated solubility predictions, these approaches identify alterations enhancing folding stability and reducing aggregation tendency \u003csup\u003e\u003cspan additionalcitationids=\"CR15\" citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u003c/sup\u003e. For eukaryotic proteins produced in prokaryotic systems like \u003cem\u003eE. coli\u003c/em\u003e, computational optimization provides a systematic and effective approach to enhance solubility by targeted mutagenesis which related studies revealed surface patch analysis was used to design rHuEPO variants in \u003cem\u003eE. coli\u003c/em\u003e, where reducing positively charged patches (e.g., F48D, R150D) enhanced solubility by up to 60%, while increasing them (e.g., E13K) led to significant solubility loss compared to the wild type \u003csup\u003e\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e. Also, it has been reported that these methodologies applied integrate structural modelling, MD simulations, and solubility prediction methods to systematically diminish aggregation-prone areas \u003csup\u003e\u003cspan additionalcitationids=\"CR19\" citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003eBased on recent research regarding the solubility engineering of aggregation-prone eukaryotic proteins and the \u003cem\u003ede novo\u003c/em\u003e structural modelling framework and MD; we introduce a computational framework designed to enhance the solubility of the DEK1 calpain-like (CysPc) domain from \u003cem\u003eP. patens\u003c/em\u003e by rational mutagenesis. The \u003cem\u003eP. patens\u003c/em\u003e has been chosen as this model plant enables to use effective targeted mutagenesis to link protein structure with function \u003csup\u003e\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e,\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e\u003c/sup\u003e. Through this manner, we identify surface-exposed residues exhibiting significant aggregation propensity or detrimental solvation patterns. We suggest specific changes that will improve solubility by making the surface less hydrophobic and increasing local flexibility while keeping the structure stable. This method uses tried-and-true methods for improving the solubility of eukaryotic proteins that tend to clump together and provides a way to study the structure and function of DEK1 domains that have been hard to study experimentally because they don't dissolve well enough \u003csup\u003e\u003cspan additionalcitationids=\"CR30 CR31 CR32\" citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e\u003c/sup\u003e. Here. First we developed a workflow procedure to increase the accuracy of prediction the structure and based on provided reconstructed structure we come up with a computational method using focused mutagenesis to maximize the solubility of DEK1\u0026rsquo;s CysPc domain. We find residues causing low solubility using structural modelling and MD simulations, then we offer mutation candidates expected to increase solubility while preserving domain stability. This work opens the route for functional and structural studies of DEK1 and offers a paradigm for solubility-oriented protein engineering of challenging eukaryotic domains.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003eModelling and validation of natural-type and mutated protein structures\u003c/p\u003e\n\u003cp\u003eWe employed various structural assessment techniques to meticulously evaluate all modelled structures, encompassing both native and mutant protein variations (Fig. \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e). We additionally contrasted our integrated homology modelling with direct threading modelling. The projected models\u0026apos; TM-score, US-score, QMEANDisCo, ProQ3D, and Z-score indicate that the structure generated by AlphaFold2 is the most dependable among those examined. The continuous improvement of these quality parameters indicates that the AlphaFold2 model is the most precise and dependable method for representing the target protein\u0026apos;s structure. The model\u0026apos;s consistently elevated values demonstrate this (Table. 1).\u003c/p\u003e\n\u003cp\u003eThe SAVES v6.1 server\u0026apos;s assessment of stereochemical quality revealed that most residues in all structures resided within the favoured and allowed sections of the Ramachandran plot, signifying appropriate backbone geometry. The QMEAN scoring and analysis conducted using the SWISS-MODEL Assessment tools verified that both the normal and mutants\u0026rsquo; models exhibited satisfactory overall quality and reliability ratings. Result revealed that average score of QMEAN for all analysed structure with integrated homology modelling was 1.02 which considered as an acceptable score while this score was 1.53 in direct threading method of structure prediction of CysPc domain. Also, the ERRAT analysis corroborated these findings, revealing that all models demonstrated good overall quality factors, indicating a lack of substantial non-bonded interaction defects. The amalgamation of these validation techniques verifies that all produced protein structures both natural-type and mutated exhibit good structural integrity and are appropriate for further analysis. All tests which carried out with SAVES v6.1 server for integrated homology modelling indicate that the average of overall quality factor of predicted structures was 97.32 and for verified 3D structure reached to 86.53% while these metrics for direct method of structure prediction was 91.33 and 74.26% respectively which significantly increased. In addition, the Ramachandran plot analyses of natural structure of CysPc versus direct threading model demonstrated that the structured model applied method of prediction exhibited a greater proportion of residues within the acceptable (favoured and allowed) regions than those produced by direct threading structure prediction methods, signifying enhanced stereochemical quality and backbone geometry in the validated models (Fig. \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e).\u003c/p\u003e\n\u003cp\u003eMD simulations were performed on the natural CysPc domain and each individual mutant to thoroughly evaluate the structural and dynamic effects of sequence alterations. For each structure, we extracted a collection of pertinent features: SASA, H-bond, RDF, Rg, Diffusion coefficient, RMSD, RMSF, minDist. We utilized a weighted scoring approach to statistically compare and prioritize individual mutants based on these criteria. The weights for each characteristic were established using statistical study of the link between feature alterations (from natural to single mutant) and their correlation with favourable enhancements in protein solubility. The resultant set of experimentally derived weights is as follows: SASA (1.2), H-bond (1.2), RDF (1.0), Rg (1.1), Diffusion (0.8), RMSD (-1.4), RMSF (-1.4), and minDist (1.0). The weighted differences were normalized to facilitate direct comparison among features with varying scales, and the resultant \u003cem\u003eWeight(x)\u003c/em\u003e score (Eq. 1) permitted efficient ranking of all individual mutants. Mutants with the highest composite scores were recognized as the most promising candidates for the subsequent phase of combinatorial mutant type and optimization.\u003c/p\u003e\n\u003cp\u003eEvaluation of single mutants; structural and dynamic changes\u003c/p\u003e\n\u003cp\u003eSingle-point mutations in the CysPc sequence were predicted via AlphaFold2 and underwent 200 ns MD simulations to assess their structural and dynamic characteristics. The backbone RMSD values of all mutants stabilized post-equilibration and remained analogous to the natural type, signifying that the overall conformation of the CysPc domain was maintained. Analyses of RMSF indicated that mutations in surface-exposed or loop regions generally enhanced local flexibility, whereas substitutions within core secondary structures typically diminished RMSF, signifying localized stabilizing. Alterations SASA were noted based on the characteristics of the inserted residue: hydrophilic or charged alterations at the surface typically augmented local SASA, indicating enhanced solubility, while hydrophobic replacements diminished solvent exposure. The total count of intra-protein hydrogen bonds remained comparable to the natural type; however, certain mutants displayed novel surface-exposed hydrogen bonds that could potentially improve solubility or stability. The radius of gyration values for all mutants exhibited negligible variation, affirming that domain compactness was preserved. The results as a whole show that most of the individual mutations in the CysPc domain, as shown by AlphaFold2 and confirmed by long-timescale MD simulations, do not change the overall structure of the domain. However, they may change how flexible it is and how much solvent it is exposed to, which are important for improving stability and solubility.\u003c/p\u003e\n\u003cp\u003eIterative generation and assessment of double and triple mutants\u003c/p\u003e\n\u003cp\u003eWe meticulously developed and evaluated a library of double mutants in the CysPc domain of DEK1 by extensive \u003cem\u003ein silico\u003c/em\u003e screening. Subsequently, we selected the 25 out of 45 most promising candidates from double mutants based on various MD features associated with solubility. We meticulously selected each double mutant by amalgamating two single-point mutations that had previously demonstrated enhanced solubility metrics independently. MD simulations of the double mutants demonstrated a notable decrease in RMSD, signifying enhanced structural stability relative to the natural-type and single mutants. Concurrently, RMSF analyses indicated less residue-level fluctuations in critical locations, corroborating the notion of a more stable conformation. These data aligned with signs of increased solubility, since the mutants seemed more solvent-accessible and exhibited improved hydration. To corroborate these findings, we additionally assessed other structural and solubility-related characteristics.\u003c/p\u003e\n\u003cp\u003eCertain double mutants exhibited enhancements in various solubility descriptors concurrently. These mutants are excellent subjects for further mutational investigation or experimental validation. In summary, these 25 double mutants exhibited several enhanced solubility indicators and are promising candidates for subsequent triple mutants\u0026rsquo; selection. Table \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e presents their ranking and fitness of MD-derived profiles, demonstrating the efficacy of iterative, data-driven mutagenesis techniques in systematically enhancing protein solubility \u003cem\u003ein silico\u003c/em\u003e.\u003c/p\u003e\n\u003cp\u003eIdentification of top mutant candidates for improved solubility\u003c/p\u003e\n\u003cp\u003eThe identification and creation of triple mutants were executed via a complex computational pipeline that utilized both empirical MD data and sophisticated reinforcement learning, specifically employing a Proximal Policy Optimization (PPO) framework. Initially, a comprehensive dataset was created through \u003cem\u003ein silico\u003c/em\u003e mutagenesis of the CysPc domain, producing single and double mutants and simulating the dynamics of each variant to derive critical solubility-related descriptors: RMSD, RMSF, SASA, H-bond, Rg, radial distribution RDF, minimum residue distance, and diffusion coefficient. The descriptors were quantitatively assessed for each mutation and normalized against the natural-type protein to create a comprehensive feature collection.\u003c/p\u003e\n\u003cp\u003eThe PPO methodology was subsequently employed to systematically explore the extensive combinatorial space of potential triple mutants. The reinforcement learning agent was taught using a bespoke reward function that incorporated enhancements in all pertinent solubility parameters rewarding additive or synergistic increases and penalizing unstable or redundant alterations. The agent\u0026apos;s action space was dynamically limited: each time, it only looked at positions and substitutions that had been shown to improve solubility in single or double mutants. At the same time, it looked at the natural-type conformation to make sure that no mutations were introduced that would break important native contacts or structural motifs.\u003c/p\u003e\n\u003cp\u003eFor every prospective triple mutant suggested by the PPO agent, an \u003cem\u003ein silico\u003c/em\u003e MD simulation was conducted, and the resultant structural and solubility metrics were incorporated into the algorithm to enhance its exploration approach. The algorithm\u0026apos;s policy was enhanced by integrating knowledge of epistatic effects situations where one mutation\u0026apos;s influence is altered by another\u0026apos;s presence achieved through a systematic comparison of the triple mutant\u0026apos;s characteristics with those of its individual single and double mutants. This method facilitated the identification of combinations in which the third mutation yielded a genuinely advantageous effect just when associated with double mutant backgrounds, rather than independently (Fig. \u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e). Furthermore, the natural sequence was used at each evaluation stage as a reference to confirm that the proposed triple mutants would not deviate significantly from the protein\u0026apos;s original conformation or stability characteristics which all resulted in MUT347 as rank 1 of triple mutant selection. We ran extra MD analyses on the triple mutants that obtained the best scores to make sure they did what they were meant to do and to see how they would change the structure and flexibility of proteins. This process made sure that only the best and most lasting alternatives made it to the final phase of testing. It combined the benefits of data-driven prediction and biological plausibility.\u003c/p\u003e\n\u003cp\u003eStability and flexibility of natural-type and mutated protein structures\u003c/p\u003e\n\u003cp\u003eOver a 200 ns trajectory MD models were run to see how the surface mutations affected the stability and dynamic behaviour of both the natural-type and mutant proteins. RMSD plots (Fig. \u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003e, B) show that the mutant protein (MUT347) reached equilibrium faster and had consistently lower RMSD values than the natural-type protein. The natural-type structure went up steadily and stopped at about 0.70 \u0026Aring;, but the mutant form stopped moving earlier and stayed at 0.49 \u0026Aring; throughout the exercise. The results show that the triple mutation made the structure more stable overall, most likely by reducing the changes in shape that come from surface-solvent interactions. We also applied RMSF analysis to check the flexibility of the area (Fig. \u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003e, A). As expected, both protein versions showed similar fluctuation patterns, with most of the flexible areas being found at the N and C-termini. The mutant structure had slightly lower RMSF values in several loop areas, especially between residues 0:20, 50:57, 190:209, and 322:354, which meant that the area was stiffer. This stability could mean that the surface was remodelled effectively, adding more polar or charged contacts that made solvent engagement and structural compactness better. Also, persistent fluctuation in active site of both structures still high, suggests that applied mutants have negligible impact on the structure integrity on catalytic area. Additionally, Rg analysis between natural and mutated structure supports this matter that mutated structure has better compactness and has higher integrity over time of simulation and lower fluctuation (Fig. \u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003e, C). The RMSD and RMSF results support the idea that surface-targeted mutations can make proteins more stable in solution, especially when there are a lot of them and depletion-induced aggregation can happen because of short interactions between molecules.\u003c/p\u003e\n\u003cp\u003eImplications of secondary structure stability and solubility\u003c/p\u003e\n\u003cp\u003eDSSP was used for secondary structure analysis during the exercise to find out if the changes seen in the dynamic behaviour led to changes in the structure. The time-resolved secondary structure maps (Fig. \u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003e) showed that the overall secondary structure makeup of both the native and mutant forms stayed the same over the 200 ns trajectory. When looked at more closely, it was found that the mutant protein kept its \u0026alpha;-helix and \u0026beta;-sheet structure more consistently, while the natural-type protein changed between structured and unstructured states more often, especially in loop-rich areas. This was confirmed by the quantitative study of the average secondary structure makeup. The mutant had a slightly higher content of helixes and turns compared to the normal variant, but a significantly lower content of coils and bends. There were also fewer standard deviations across simulation blocks. This means that the structure is generally more stable and denser, which these traits are usually linked to protein\u0026rsquo;s higher solubility. These results are important because they match the decrease in RMSF seen in the mutant. This means that the designed surface mutations made the area less flexible while also making the native secondary structure motifs more stable. From a physiological point of view, this drop in surface-exposed dynamic disorder may limit the ability for depletion to cause aggregation, especially when there is a lot going on inside cells or in formulations. The combined study of structural and secondary structures shows that changes made to the surface have a good effect on solubility while keeping the structure\u0026apos;s integrity. In addition, minimum distance matrix (mdmat) of last 20 ns for residue-residue contact pattern has a higher proof on this matter that applied mutants has a preserved global contact map and no major rearrangement and misfolding occurred and there isn\u0026rsquo;t any disrupt domain architecture on internal topology (Fig. \u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003e, B).\u003c/p\u003e\n\u003cp\u003eAnalyses of hydrogen bonding and solvent accessibility area\u003c/p\u003e\n\u003cp\u003eWe examined protein-solvent hydrogen bonds (between structure and surrounding solvent) and intra-protein hydrogen (between residues) bonds and SASA during the simulation (dynamic mode) and in static mode to investigate the structural basis for the enhanced solubility of the natural-type and mutated proteins. The overall count of intra-protein hydrogen bonds was persistently elevated in the mutant structure throughout the trajectory, averaging 274.98\u0026thinsp;\u0026plusmn;\u0026thinsp;9.6 H-bonds, in contrast to 229.17\u0026thinsp;\u0026plusmn;\u0026thinsp;7.6 in the natural-type (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e6\u003c/span\u003e, A). Also, slight reduction of hydrogen bond between structures and solvent was not statistically significant but higher formation of hydrogen bond between residues (intra-molecular) detected which this matter reflects more internal stabilization and tighter fold in mutated structure. Furthermore, static mode of SASA profiles showed that the two protein variants had very different amounts of surface contact which statistically significant. The mutated structure had a higher average SASA (21108 \u0026Aring;\u0026sup2;), but the natural protein had a lower surface area (17821 \u0026Aring;\u0026sup2;), which supports the successful surface design and presents a greater surface area available for interaction with water molecules.\u003c/p\u003e\n\u003cp\u003e\u003cbr\u003e\u003c/p\u003e\n\u003cp\u003eThis rise means that the mutations either improved the packing inside or set up connections that make the structure more stable, which is what caused the RMSD and RMSF values were reduced.\u003c/p\u003e\n\u003cp\u003eThere aren\u0026rsquo;t any significant differences between SASA dynamic mode of natural-type and mutant in all over trajectories, which where 185.59\u0026thinsp;\u0026plusmn;\u0026thinsp;1.97 and 183.19\u0026thinsp;\u0026plusmn;\u0026thinsp;3.7 nm/S\u003csup\u003e2\u003c/sup\u003e/N respectively. Notwithstanding, better stable fluctuation and higher accessible area were detected on mutated structure at finals frames. On the other hand, analyses of free energy surface (FES) in basis of RMSD and SASA (protein-surface) a slight better reduction of energy resulted for mutant structure in comparison with natural type (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e6\u003c/span\u003e, B).\u003c/p\u003e\n\u003cp\u003eA bigger solvent accessible surface area on static mode and stronger intra-molecular hydrogen bonds supports the idea that the mutant has changed into a better shape that makes it easier to dissolve. Higher SASA means that more polar and hydrophilic surface residues are exposed, which lets them interact with water molecules around them in a wider area. Furthermore, analyses of H-bond among natural type and mutated structure revealed no significant difference but along with overall solvation, per-residue H-bond analysis showed that some residues close to the mutation sites were more regularly involved in water interactions in the mutant structure (Fig. \u003cspan class=\"InternalRef\"\u003e6\u003c/span\u003e, A). This localized rise in hydration may work as a buffer against nonspecific inter-protein contacts, making it less likely that proteins will clump together when they are crowded. These results add to the SASA data, which showed a moderate drop in global solvent exposure but a better distribution of polar residues that could be accessed by solvents. In opposition, a significant difference of H-bond discovered in intra residues level which could improve with potential of reduce in aggregation and supports solubility via structure better compactness which Rg analyses also supports this matter. Also, studies of hydrogen bond lifetimes showed no difference between natural-type and mutant. This suggests that the applied mutant has no impact on structure destabilization and flexibility.\u003c/p\u003e\n\u003cp\u003eFigure \u003cspan class=\"InternalRef\"\u003e7\u003c/span\u003e, Panel A illustrates that alterations in pH result in negligible, limited variations, particularly at mutated sites, with the traces for both the natural-type and Mut347 predominantly overlapping. The principal flexibility is restricted to the N- and C-terminal peaks, and no new hotspots arise throughout the pH series. In contrast, Panels B and C (NaCl and (NH₄)₂SO₄, within a concentration range of 0.2\u0026ndash;1.2 mol/L) demonstrate distinct ionic-strength effects: as salt concentration increases, the overall RMSF envelope diminishes, terminal peaks reduce in size, and loop areas become more uniform, with these enhancements consistently more evident in Mut347. Significantly, at the mutation site and its proximal neighbours in Mut347, the coloured traces at elevated salt concentrations are positioned beneath the Natural-type curves, suggesting enhanced fluctuations for both salts; this decrease persists into adjacent loop segments, where RMSF is markedly diminished compared to low-salt conditions. The reduction of terminal mobility (both N and C) is more pronounced in Mut347, resulting in cleaner, lower-amplitude endpoints compared to the Natural type. Although NaCl and (NH₄)₂SO₄ exhibit a similar qualitative trend, the reduction in RMSF\u0026mdash;at the altered residue, in adjacent loops, and at the termini is somewhat more significant in (NH₄)₂SO₄ at equivalent doses.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eThe integrated modelling and validation technique employed merging template searches (Dali/COFACTOR, RCSB PDB/UniProt) with various predictors and rigorous quality assessments\u0026mdash;yielded a dependable structural foundation for molecular dynamics. Dali and COFACTOR are recognized for structure/function transfer from homologs, whilst RCSB PDB and UniProt function as the principal, curated repositories for 3D structures and sequence annotations, respectively \u003csup\u003e\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e,\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e\u003c/sup\u003e. Across independent quality metrics (e.g., QMEANDisCo, ProQ3D, Ramachandran/PROCHECK, ERRAT), the AlphaFold2 model consistently scored best, aligning with broad evidence that AlphaFold2 attains near-experimental accuracy for many single-domain proteins. Specifically, QMEANDisCo and ProQ3D are cutting-edge model-quality evaluators, while the SAVES/ERRAT toolbox continues to serve as a benchmark for stereochemical assessments and the identification of non-bonded interaction anomalies \u003csup\u003e\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e,\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e\u003c/sup\u003e. This consensus \u003cem\u003ede novo\u003c/em\u003e approach is supported by recent benchmarks, which demonstrate that the combination of models from several predictors can perform better than any one approach, especially for challenging targeted structure especially when homology modelling with single template is insufficient \u003csup\u003e\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e,\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e\u003c/sup\u003e. Our findings corroborate this pattern: our combined, consensus-based approach as an integrated method achieved 86% prediction accuracy for the CysPc domain, compared to 74% for AlphaFold2 as direct-threading method \u003csup\u003e\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e,\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003eThe molecular dynamics investigations (RMSD, RMSF, Rg, SASA) demonstrate that the designed surface replacements enhanced global dynamics while maintaining the structural conformation. RMSD indicates the overall deviation from the reference conformation, whereas RMSF identifies flexibility along the sequence; the observed first: accelerated equilibration/lower RMSD and second: attenuated RMSF at termini and loops are characteristic indicators of enhanced conformational stability. The minor decrease and refinement of Rg indicate a more compact ensemble, aligning with conventional interpretations of protein compactness in simulations and statistical evaluations of PDB structures \u003csup\u003e\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e,\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e\u003c/sup\u003e. Additionally, alterations in SASA and hydrogen-bond counts indicate enhanced internal packing and a restructured hydration shell; increased intra-protein hydrogen bonding, together with slight modifications in protein\u0026ndash;solvent hydrogen bonds, is a prevalent mechanism for preserving the native state without altering the fold \u003csup\u003e\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003eAt the sequence level, the Arg-containing substitutions near the surface likely enhance both stability and solubility by augmenting local hydrogen-bonding opportunities (including weak yet frequent C\u0026ndash;H\u0026middot;\u0026middot;\u0026middot;O interactions) and redistributing surface charge to promote hydration and diminish aggregation-prone positive regions. The stabilizing function of canonical and weak (C\u0026ndash;H\u0026middot;\u0026middot;\u0026middot;O) hydrogen bonds at interfaces is well established, and Arg-rich motifs can engage in both canonical hydrogen bonding and advantageous CH\u0026middot;\u0026middot;\u0026middot;O interactions \u003csup\u003e\u003cspan additionalcitationids=\"CR44 CR45\" citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e\u003c/sup\u003e. Furthermore, surface electrostatics significantly affect solubility; enhanced negative/polar exposure and a reduction in extensive positive regions are associated with improved soluble expression aligning with SASA/H-bond patterns and the diminished variations in exposed loops \u003csup\u003e\u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003eThe mutant demonstrates an elevated number of intramolecular hydrogen bonds without an increase in hydrogen bond lifetimes, indicating a denser yet not necessarily slower network precisely the type of incremental reinforcement that contributes several kJ\u0026middot;mol⁻\u0026sup1; of stability without excessively restricting dynamics. Comprehensive mutational thermodynamics indicate that backbone and side-chain hydrogen bonding positively influence stability in a context-dependent manner, aligning with our finding that an increased number of internal hydrogen bonds correlates with decreased stability \u003csup\u003e\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003eThe MD data indicate a consistent stabilizing mechanism for the surface-engineered mutant. The accelerated equilibration and consistently reduced RMSD of MUT347 compared to the natural-type suggest a more restricted conformational ensemble, a result anticipated when surface electrostatics are positively altered by mutation. Previous research indicates that targeting solvent-exposed regions and improving surface charge can enhance thermodynamic stability without compromising the fold, aligning with our observations \u003csup\u003e\u003cspan additionalcitationids=\"CR49\" citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e\u003c/sup\u003e. In molecular dynamics, RMSD and RMSF serve as conventional metrics for stability and stiffness; a reduction in global RMSD accompanied by a decrease in local RMSF often indicates a more stable, compact ensemble \u003csup\u003e\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e,\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003eFurthermore, RMSF profiles related to implemented surface design indicate that the most significant reductions occur in loops and at the termini, whereas the structured core maintains rigidity specifically where stabilization typically initially emerges. The DSSP analysis verifies that secondary-structure content is maintained or somewhat enhanced (increased helix/turn, decreased coil/bend), suggesting that the mutant's diminished fluctuations result in improved preservation of native motifs rather than a modification of fold \u003csup\u003e\u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e,\u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e54\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003eThe subdued pH dependence and pronounced salt-dependent attenuation of RMSF most evident at the N/C-terminal peaks, solvent-exposed loops, and the mutated site align with traditional ionic-strength screening: elevated salt concentrations diminish the Debye length, mitigate long-range charge\u0026ndash;charge repulsion, and promote short-range interactions, consequently reducing backbone fluctuations in the most unstable regions \u003csup\u003e\u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e55\u003c/span\u003e\u003c/sup\u003e.In accordance with ion specificity, the marginally enhanced damping observed with (NH₄)₂SO₄ compared to NaCl is anticipated for a more kosmotropic salt positioned higher in the Hofmeister series, which facilitates compaction and salting-out at similar ionic strengths \u003csup\u003e\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e,\u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e57\u003c/span\u003e\u003c/sup\u003e.Notably, Mut347 exhibits superior enhancement (reduced RMSF) with elevated salt concentrations compared to the Natural type particularly at and near the mutated site and demonstrates a slightly improved global profile; such sequence-dependent advancements are expected when local charge distributions are optimized by screened electrostatics and when substitutions improve interfacial interactions \u003csup\u003e\u003cspan citationid=\"CR58\" class=\"CitationRef\"\u003e58\u003c/span\u003e,\u003cspan citationid=\"CR59\" class=\"CitationRef\"\u003e59\u003c/span\u003e\u003c/sup\u003e. In this context, mutations at residue 21 and residue 342 may enhance stabilization by facilitating stronger hydrogen bonds with adjacent chains, a recognized factor in complex/interface stability, including weak yet frequent CH-O and canonical H-bonds at residue-residue interfaces \u003csup\u003e\u003cspan citationid=\"CR60\" class=\"CitationRef\"\u003e60\u003c/span\u003e,\u003cspan citationid=\"CR61\" class=\"CitationRef\"\u003e61\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003eAlthough, RDF, Diffusion coefficient and minDist had no substantial changes in all over mutants since the primary aim of mutagenesis process was to enhance solubility while conserving total structure integrity Consequently, significant alterations in these parameters were not expected. Our integrated computational pipeline, which includes iterative mutagenesis, extensive MD analyses, and data-driven through proximal policy optimization, is much better than traditional methods for structure-based protein engineering and increasing solubility. In most cases, mutagenesis and improving protein solubility use single-point mutations or random library screening. These methods are time-consuming and don't work well for finding complex epistatic interactions between many residues \u003csup\u003e\u003cspan citationid=\"CR62\" class=\"CitationRef\"\u003e62\u003c/span\u003e,\u003cspan citationid=\"CR63\" class=\"CitationRef\"\u003e63\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003eOn the other hand, our method uses MD simulations in a systematic way to get high-dimensional structural and solubility descriptors from different single and double mutants\u0026rsquo; library. This creates a strong data base for the future use of any data-driven techniques. Our implemented input information to find its way through the complex world of triple mutants and focuses on mutation combinations that are most likely to improve solubility in an additive or synergistic way, while also avoiding substitutions that are harmful to structure.\u003c/p\u003e\u003cp\u003eTaken together, data supports the change in surface physiochemical properties, along with the higher internal hydrogen bonding, suggests that the mutation leads to better solvation and hydration shell formation, which are both necessary to keep the solubility and stop nonspecific aggregation when conditions are tight. These findings indicate that the surface-engineered triple mutation enhances structural compaction and internal stabilization, and more accessible surface presumably diminishing the tendency for depletion-induced aggregation in dense situations.\u003c/p\u003e"},{"header":"Methods","content":"\u003cp\u003eStructure prediction, preparation and validation\u003c/p\u003e\u003cp\u003eThe amino acid sequence of the DEK1 calpain-type cysteine protease domain was retrieved from the NCBI protein database (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ncbi.nlm.nih.gov/\u003c/span\u003e\u003cspan address=\"https://www.ncbi.nlm.nih.gov/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e), specifically from \u003cem\u003eP. patens\u003c/em\u003e (accession number: XP_024400265.1). This sequence corresponds to the predicted CysPc domain of DEK1, which plays a critical role in the protein's proteolytic activity. The sequence was used as the basis for homology modelling and subsequent structural analyses. Prior to structure prediction, the sequence was analysed to identify domain boundaries, surface residues, and potential disordered regions using standard bioinformatics tools. For finding reference structure to carry out homology modelling, threading (fold recognition) model of CysPc domain which was predicted by C-I-ITASSER server (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://zhanggroup.org/C-I-TASSER\u003c/span\u003e\u003cspan address=\"https://zhanggroup.org/C-I-TASSER\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) \u003csup\u003e\u003cspan citationid=\"CR64\" class=\"CitationRef\"\u003e64\u003c/span\u003e\u003c/sup\u003e, and then applied in different blast engine on uniport, (against targeted data base of UniProtKB with 3D structure) (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.uniprot.org/blast\u003c/span\u003e\u003cspan address=\"https://www.uniprot.org/blast\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e), RCSB PDB structure similarity search (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.rcsb.org/search/advanced\u003c/span\u003e\u003cspan address=\"https://www.rcsb.org/search/advanced\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e), Dali server (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://ekhidna2.biocenter.helsinki.fi\u003c/span\u003e\u003cspan address=\"http://ekhidna2.biocenter.helsinki.fi\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) \u003csup\u003e\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e\u003c/sup\u003e and COFACTOR (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://zhanggroup.org/COFACTOR\u003c/span\u003e\u003cspan address=\"https://zhanggroup.org/COFACTOR\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) \u003csup\u003e\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e,\u003cspan citationid=\"CR65\" class=\"CitationRef\"\u003e65\u003c/span\u003e\u003c/sup\u003e. All search engine resulted in structure of mu-like calpain (PDB: 1QXP) and related structure retrieved from protein data bank (RCSB). Because experimental structure of 1QXP has 1571 modelled residue count and deposited structure had 1800 residues, closest structure of 1QXP retrieved from RCSB protein data bank according to the result of structural search engines (1KEU, 2P0R, and 6BDT)) and Modeller 10.7 with multimode prediction strategy applied to predict complete structure with reference structures to reaching to the best model possible as reference structure \u003csup\u003e\u003cspan citationid=\"CR66\" class=\"CitationRef\"\u003e66\u003c/span\u003e,\u003cspan citationid=\"CR67\" class=\"CitationRef\"\u003e67\u003c/span\u003e\u003c/sup\u003e. The 3D structure of CysPc domain was modelled using AlphaFold2 \u003csup\u003e68,69\u003c/sup\u003e, SWISS-MODEL (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://swissmodel.expasy.org\u003c/span\u003e\u003cspan address=\"http://swissmodel.expasy.org\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) \u003csup\u003e\u003cspan citationid=\"CR70\" class=\"CitationRef\"\u003e70\u003c/span\u003e\u003c/sup\u003e, and I-TASSER (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://zhanggroup.org/I-TASSER/\u003c/span\u003e\u003cspan address=\"https://zhanggroup.org/I-TASSER/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) \u003csup\u003e\u003cspan citationid=\"CR71\" class=\"CitationRef\"\u003e71\u003c/span\u003e\u003c/sup\u003e based on homology modelling to the experimental crystal structure of mu-like calpain (PDB: 1QXP), ensuring high coverage and confidence in the predicted conformation. The resulting model was then refined by Galaxy Refine for structure refinement (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://galaxy.seoklab.org\u003c/span\u003e\u003cspan address=\"https://galaxy.seoklab.org\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) \u003csup\u003e\u003cspan citationid=\"CR72\" class=\"CitationRef\"\u003e72\u003c/span\u003e\u003c/sup\u003e and ModLoop (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://modbase.compbio.ucsf.edu/\u003c/span\u003e\u003cspan address=\"https://modbase.compbio.ucsf.edu/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) for loops refinements \u003csup\u003e\u003cspan citationid=\"CR73\" class=\"CitationRef\"\u003e73\u003c/span\u003e\u003c/sup\u003e. Furthermore, the achieved 3D protein structure minimized any analysis about the structural integrity of the protein structure by SAVES v6.1 server (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://saves.mbi.ucla.edu/\u003c/span\u003e\u003cspan address=\"https://saves.mbi.ucla.edu/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) for 3D verification, ERRAT analyses \u003csup\u003e\u003cspan citationid=\"CR74\" class=\"CitationRef\"\u003e74\u003c/span\u003e\u003c/sup\u003e and SWISS-MODEL assess (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://swissmodel.expasy.org/assess\u003c/span\u003e\u003cspan address=\"https://swissmodel.expasy.org/assess\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) for QMEAN per residues errors \u003csup\u003e\u003cspan citationid=\"CR70\" class=\"CitationRef\"\u003e70\u003c/span\u003e\u003c/sup\u003e, QMEANDisCo for distance limitation on model quality estimation \u003csup\u003e\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e\u003c/sup\u003e, TM-score (with superimposed predicted model against each other) and US-Score (predicted model against 1QXP) to assess the similarity structure of predicted model \u003csup\u003e\u003cspan citationid=\"CR75\" class=\"CitationRef\"\u003e75\u003c/span\u003e,\u003cspan citationid=\"CR76\" class=\"CitationRef\"\u003e76\u003c/span\u003e\u003c/sup\u003e ,Z-Score for overall model quality with ProSA-web \u003csup\u003e\u003cspan citationid=\"CR77\" class=\"CitationRef\"\u003e77\u003c/span\u003e\u003c/sup\u003e, ProQ3D for improved model quality assessment (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://proq3.bioinfo.se/pred/\u003c/span\u003e\u003cspan address=\"https://proq3.bioinfo.se/pred/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) \u003csup\u003e\u003cspan citationid=\"CR78\" class=\"CitationRef\"\u003e78\u003c/span\u003e\u003c/sup\u003e, and also, Molprobity server (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://molprobity.biochem.duke.edu/\u003c/span\u003e\u003cspan address=\"http://molprobity.biochem.duke.edu/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) applied for Ramachandran analyses \u003csup\u003e\u003cspan additionalcitationids=\"CR80\" citationid=\"CR79\" class=\"CitationRef\"\u003e79\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR81\" class=\"CitationRef\"\u003e81\u003c/span\u003e\u003c/sup\u003e. All process related to the prediction of structures and preparing structural library are presented in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e.\u003c/p\u003e\u003cp\u003eMD simulation and trajectory condition design\u003c/p\u003e\u003cp\u003eThe predicted 3D conformation of the protein was generated for MD simulation utilizing the CHARMM-GUI server (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.charmm-gui.org\u003c/span\u003e\u003cspan address=\"https://www.charmm-gui.org\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) \u003csup\u003e\u003cspan citationid=\"CR82\" class=\"CitationRef\"\u003e82\u003c/span\u003e\u003c/sup\u003e, configured under typical parameters, including solvation and ion neutralization where necessary. The system was parameterized using the CHARMM36m force field and transformed into GROMACS-compatible input files. Solvation was conducted utilizing the TIP3P water model \u003csup\u003e\u003cspan citationid=\"CR83\" class=\"CitationRef\"\u003e83\u003c/span\u003e\u003c/sup\u003e, and the system was neutralized with suitable counterions. The configured system was subsequently employed for MD simulations in GROMACS 2022, with production runs conducted for 200 nanoseconds, and 100M steps according to the duration of selected simulation and ensuring stabilization of soluble-related parameters and dependable time averaged data across root means square deviation (RMSD) and root means square fluctuation (RMSF) \u003csup\u003e\u003cspan citationid=\"CR84\" class=\"CitationRef\"\u003e84\u003c/span\u003e,\u003cspan citationid=\"CR85\" class=\"CitationRef\"\u003e85\u003c/span\u003e\u003c/sup\u003e. The simulations were conducted to examine the solubility behaviour and structural dynamics of the protein under physiological settings. In addition, MD simulation carried out in ionic conditions, sodium chloride (NaCl) as physiological ionic strength agent and ammonium sulfate ((NH\u003csub\u003e4\u003c/sub\u003e)\u003csub\u003e2\u003c/sub\u003eSO\u003csub\u003e4\u003c/sub\u003e) as protein participant agent in concentration 0.2, 0.6, 0.9 and 1.2 mol/L by manual setup of input files \u003csup\u003e\u003cspan additionalcitationids=\"CR87 CR88\" citationid=\"CR86\" class=\"CitationRef\"\u003e86\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR89\" class=\"CitationRef\"\u003e89\u003c/span\u003e\u003c/sup\u003e. Also, for different pH condition behaviour, PDB files were adjusted with APBS-PDB2PQR server (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://pdb2pqr.readthedocs.io\u003c/span\u003e\u003cspan address=\"https://pdb2pqr.readthedocs.io\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) \u003csup\u003e\u003cspan citationid=\"CR90\" class=\"CitationRef\"\u003e90\u003c/span\u003e\u003c/sup\u003e in range of 4, 5.5, 7 and 9.\u003c/p\u003e\u003cp\u003eAssessment and scoring of single-point mutants for solubility improvement\u003c/p\u003e\u003cp\u003eFollowing the initial 200 ns MD simulation, residue-level analyses were conducted to identify structurally unstable or flexible regions that could serve as potential mutant targets. Important metrics included RMSF to detect residues with high positional variability, solvent accessible surface area (SASA) to identify exposed areas that may affect solubility, and hydrogen bond (H-bond) occupancy and secondary structure stability analyses to determine structural integrity throughout the simulation. Residues demonstrating persistently elevated RMSF values, augmented solvent exposure, or unstable secondary structural configurations were identified as mutation hotspots. The substitution of residues aimed to generate a novel mutated sequence by replacing residues characterized by high flexibility and hydrophobicity with those of the same category that exhibit greater similarity and enhanced hydrophilicity \u003csup\u003e\u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e\u003c/sup\u003e. In Table \u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e1\u003c/span\u003e, all applied mutations with replaced residues are presented. To preserve functional activity, all selected residues were visually examined using UCSF ChimeraX version 1.8 software, and only those situated on the protein's surface and distanced from the active site were selected for mutation. After 28 potential mutants were chosen based on RMSF, SASA, H-bond analyses, and the Ramachandran plot, they were all run through CamSol (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www-cohsoftware.ch.cam.ac.uk/\u003c/span\u003e\u003cspan address=\"https://www-cohsoftware.ch.cam.ac.uk/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) \u003csup\u003e\u003cspan citationid=\"CR91\" class=\"CitationRef\"\u003e91\u003c/span\u003e\u003c/sup\u003e, a server that calculates an intrinsic solubility profile and solubility gain upon mutagenesis, and Protein-Sol (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://protein-sol.manchester.ac.uk/\u003c/span\u003e\u003cspan address=\"https://protein-sol.manchester.ac.uk/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) \u003csup\u003e\u003cspan citationid=\"CR92\" class=\"CitationRef\"\u003e92\u003c/span\u003e\u003c/sup\u003e, a tool that estimates overall solubility propensity based on sequence and structural features. For re-prioritization structure and MD simulation, the top 10 highest scoring candidate models were chosen from all that were analysed and subsequent to the identification of the 10 premier mutant candidates, an exhaustive MD study was conducted on all associated protein structures. The evaluation encompassed RMSD, RMSF, SASA, numbers of H-bond, diffusion coefficient, radial distribution function (RDF), radius of gyration (Rg), and minimum distance (minDist) measurements. Finally assigning secondary structure to the residues (DSSP) carried out in order to final check and comparison of mutated structures against natural structure has not any structural modification. Also, on each cycle of mutant selection to library, structures with more than 5% changes in residue counts removed from selection and only structures with higher confidence than 95% with standard deviation of SE: \u0026plusmn;0.05, and RMSD score less than 1.5 \u0026Aring; (highly similar structure against natural type) advanced to the final round of selection and sorted in library according to the DSSP analyses results. In addition, all selected structure of final library, computationally sequence alignment scored and visually matched against the reference structure (natural type). Ultimately, top scored triple mutants analysed in aspect of H-bond \u003csup\u003e\u003cspan citationid=\"CR93\" class=\"CitationRef\"\u003e93\u003c/span\u003e\u003c/sup\u003e (protein-solvent and intra-protein state), RMSD, RMSF, Rg and SASA (static mode and dynamic mode) \u003csup\u003e\u003cspan citationid=\"CR94\" class=\"CitationRef\"\u003e94\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cdiv id=\"Equa\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equa\" name=\"EquationSource\"\u003e\n$$\\:\\left(1\\right)\\:Weight\\left(x\\right)=\\sum\\:_{i=1}^{8}{w}_{i}.({f}_{i,\\:mutant}-{f}_{i,\\:\\:natural})$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eSynergistic mutagenesis design; identifying optimal mutant\u0026rsquo;s combinational state\u003c/p\u003e\u003cp\u003eWe utilized a systematic combinatorial design to uncover optimal synergistic mutation combinations that enhance protein solubility, focusing on the top 10 single-point mutants previously selected by solubility-related MD characteristics. All conceivable double and triple combinations were selected and generated as described in follow, and for each, short MD simulations in 200 ns were conducted to derive solubility-relevant descriptors, including SASA, number of hydrogen bonds, radial distribution function (RDF), radius of gyration (Rg), diffusion coefficient, RMSD, RMSF, and minimum distance to solvent (minDist). in next, top 25 promising combination according to the complementary effects on solubility selected and finally a library consists of natural type, single and combined (double and triple combination) mutants generated for next phase of analyses to find the best possible combination.\u003c/p\u003e\u003cp\u003eEach combination was assessed utilizing a weighted solubility scoring algorithm, with weights allocated to each attribute according to their established or presumed impact on solubility. The unprocessed solubility score was calculated utilizing:\u003cdiv id=\"Equb\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equb\" name=\"EquationSource\"\u003e\n$$\\:\\left(2\\right)\\:Solubility\\:Score=\\sum\\:_{i=1}^{8}{w}_{i}.{f}_{i}$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003ewhere \u003cem\u003ef\u003c/em\u003e\u003csub\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sub\u003e represents the normalized value of feature \u003cem\u003ei\u003c/em\u003e and \u003cem\u003ew\u003c/em\u003e\u003csub\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sub\u003e denotes the associated feature weight. Positive weights were allocated to qualities linked to enhanced solubility (e.g., SASA, H-bond, RDF, Rg, diffusion), whilst negative weights were designated to features related to structural instability (e.g., RMSD, RMSF, minDist). This synergistic strategy aimed to uncover mutation combinations that yield improved solubility effects surpassing those of single-point variants, therefore informing the design of optimal multi-point mutants with superior biophysical features. The ultimate solubility score was standardized to a 0\u0026ndash;1 scale utilizing:\u003c/p\u003e\n\u003cdiv class=\"Heading\"\u003e(3) \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:Scaled\\:score=100\\times\\:\\frac{\\sum\\:_{i=1}^{8}{w}_{i}.{f}_{i}-min\\:score}{\\text{max}score-\\text{min}score}\\)\u003c/span\u003e\u003c/span\u003e\u003c/div\u003e\u003cp\u003eA reinforcement learning method called Proximal Policy Optimization (PPO)\u003csup\u003e\u003cspan citationid=\"CR95\" class=\"CitationRef\"\u003e95\u003c/span\u003e\u003c/sup\u003e used to quickly move through the combinatorial space and find the best mutation sets. To get the highest solubility score, PPO agent was taught to pick mutation combinations that work well together based on input from simulated MD-derived attributes. This integrated design approach makes it easier to choose multi-point mutants that are logical, based on data, and have solubility properties that are noticeably better than those possible with single mutations alone. Schematic diagram of applied method is presented in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis work is the result of implementation of Slovak Research and Development Agency grants APVV-21-0227, APVV-21-0215, APVV-22-0161 and by implementation of the project 101160008 \u0026ldquo;Fostering Excellence in Advanced Genomics and Proteomics Research at Comenius University in Bratislava \u0026ndash; \u0026nbsp;FORGENOM II\u0026rdquo; funded by the Horizon Europe program.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgments\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eMD is thankful to Daniel Kr\u0026aacute;ľ and Milan Melicherč\u0026iacute;k, Faculty of Mathematics,\u0026nbsp;Physics, and Informatics (FMPI) for their kind recommendation and guidance.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthur Contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eMD and ZL: conceived and designed the study, MD: developed the structural prediction and mutagenesis workflow, and performed all molecular dynamics simulations and solubility analyses. Data processing, interpretation, and manuscript writing and editing were carried out by MD and ZL. ES: Provided recommendations and scientific guidance of research. All aspects of the research were conducted under the academic supervision of JT , VB and SS who provided critical feedback on the study design and manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData Availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe data generated and analysed during this study, including molecular dynamics simulation outputs, structural models, and solubility feature datasets, are available from the corresponding author upon reasonable request. Due to the size and computational nature of the datasets, they are not hosted in a public repository.\u0026nbsp;\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eVillaverde, A. Mar Carri\u0026oacute;, M. Protein aggregation in recombinant bacteria: biological role of inclusion bodies. \u003cem\u003eBiotechnol. Lett.\u003c/em\u003e \u003cb\u003e25\u003c/b\u003e, 1385\u0026ndash;1395 (2003).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBaneyx, F. \u0026amp; Mujacic, M. Recombinant protein folding and misfolding in Escherichia coli. \u003cem\u003eNat. Biotechnol.\u003c/em\u003e \u003cb\u003e22\u003c/b\u003e, 1399\u0026ndash;1408 (2004).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePantophlet, R., Wilson, I. A. \u0026amp; Burton, D. R. Improved design of an antigen with enhanced specificity for the broadly HIV-neutralizing antibody b12. \u003cem\u003eProtein Eng. Des. Selection\u003c/em\u003e. \u003cb\u003e17\u003c/b\u003e, 749\u0026ndash;758 (2004).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDe Marco, A., Deuerling, E., Mogk, A., Tomoyasu, T. \u0026amp; Bukau, B. Chaperone-based procedure to increase yields of soluble recombinant proteins produced in E. coli. \u003cem\u003eBMC Biotechnol\u003c/em\u003e \u003cb\u003e7\u003c/b\u003e, (2007).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGustafsson, C., Govindarajan, S. \u0026amp; Minshull, J. Codon bias and heterologous protein expression. \u003cem\u003eTrends Biotechnol.\u003c/em\u003e \u003cb\u003e22\u003c/b\u003e, 346\u0026ndash;353 (2004).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eChatterjee, D. K. \u0026amp; Esposito, D. Enhanced soluble protein expression using two new fusion tags. \u003cem\u003eProtein Exp. Purif.\u003c/em\u003e \u003cb\u003e46\u003c/b\u003e, 122\u0026ndash;129 (2006).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eEsposito, D. \u0026amp; Chatterjee, D. K. Enhancement of soluble protein expression through the use of fusion tags. \u003cem\u003eCurr. Opin. Biotechnol.\u003c/em\u003e \u003cb\u003e17\u003c/b\u003e, 353\u0026ndash;358 (2006).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSachdev, D. \u0026amp; Chirgwin, J. M. Properties of Soluble Fusions Between Mammalian Aspartic Proteinases and Bacterial Maltose-Binding Protein. \u003cem\u003eJ. Protein Chem.\u003c/em\u003e \u003cb\u003e18\u003c/b\u003e, 127\u0026ndash;136 (1999).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAgostini, F., Cirillo, D., Bolognesi, B. \u0026amp; Tartaglia, G. G. X-inactivation: quantitative predictions of protein interactions in the Xist network. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cb\u003e41\u003c/b\u003e, e31\u0026ndash;e31 (2013).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKulshreshtha, S., Chaudhary, V., Goswami, G. K. \u0026amp; Mathur, N. Computational approaches for predicting mutant protein stability. \u003cem\u003eJ. Comput. Aided Mol. Des.\u003c/em\u003e \u003cb\u003e30\u003c/b\u003e, 401\u0026ndash;412 (2016).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDamborsky, J. \u0026amp; Brezovsky, J. Computational tools for designing and engineering enzymes. \u003cem\u003eCurr. Opin. Chem. Biol.\u003c/em\u003e \u003cb\u003e19\u003c/b\u003e, 8\u0026ndash;16 (2014).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eEbert, M. C. \u0026amp; Pelletier, J. N. Computational tools for enzyme improvement: why everyone can \u0026ndash; and should \u0026ndash; use them. \u003cem\u003eCurr. Opin. Chem. Biol.\u003c/em\u003e \u003cb\u003e37\u003c/b\u003e, 89\u0026ndash;96 (2017).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBroom, A., Jacobi, Z., Trainor, K. \u0026amp; Meiering, E. M. Computational tools help improve protein stability but with a solubility tradeoff. \u003cem\u003eJ. Biol. Chem.\u003c/em\u003e \u003cb\u003e292\u003c/b\u003e, 14349\u0026ndash;14361 (2017).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eChilders, M. C. \u0026amp; Daggett, V. Insights from molecular dynamics simulations for computational protein design. \u003cem\u003eMol. Syst. Des. Eng.\u003c/em\u003e \u003cb\u003e2\u003c/b\u003e, 9\u0026ndash;33 (2017).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRouhani, M., Khodabakhsh, F., Norouzian, D., Cohan, R. A. \u0026amp; Valizadeh, V. Molecular dynamics simulation for rational protein engineering: Present and future prospectus. \u003cem\u003eJ. Mol. Graph. Model.\u003c/em\u003e \u003cb\u003e84\u003c/b\u003e, 43\u0026ndash;53 (2018).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePikkemaat, M. G., Linssen, A. B. M., Berendsen, H. J. C. \u0026amp; Janssen, D. B. Molecular dynamics simulations as a tool for improving protein stability. \u003cem\u003eProtein Eng. Des. Selection\u003c/em\u003e. \u003cb\u003e15\u003c/b\u003e, 185\u0026ndash;192 (2002).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eCarballo-Amador, M. A., McKenzie, E. A., Dickson, A. J. \u0026amp; Warwicker, J. Surface patches on recombinant erythropoietin predict protein solubility: engineering proteins to minimise aggregation. \u003cem\u003eBMC Biotechnol\u003c/em\u003e \u003cb\u003e19\u003c/b\u003e, (2019).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKumar, S., Kumar Bhardwaj, V., Singh, R. \u0026amp; Purohit, R. Explicit-solvent molecular dynamics simulations revealed conformational regain and aggregation inhibition of I113T SOD1 by Himalayan bioactive molecules. \u003cem\u003eJ. Mol. Liq.\u003c/em\u003e \u003cb\u003e339\u003c/b\u003e, 116798 (2021).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eChennamsetty, N., Voynov, V., Kayser, V., Helk, B. \u0026amp; Trout, B. L. Prediction of Aggregation Prone Regions of Therapeutic Proteins. \u003cem\u003eJ. Phys. Chem. B\u003c/em\u003e. \u003cb\u003e114\u003c/b\u003e, 6614\u0026ndash;6624 (2010).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAgrawal, N. J. et al. Aggregation in Protein-Based Biotherapeutics: Computational Studies and Tools to Identify Aggregation-Prone Regions. \u003cem\u003eJ. Pharm. Sci.\u003c/em\u003e \u003cb\u003e100\u003c/b\u003e, 5081\u0026ndash;5095 (2011).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eJohnson, K. L., Faulkner, C., Jeffree, C. E. \u0026amp; Ingram, G. C. The Phytocalpain Defective Kernel 1 Is a Novel \u003cem\u003eArabidopsis\u003c/em\u003e Growth Regulator Whose Activity Is Regulated by Proteolytic Processing. \u003cem\u003ePlant. Cell.\u003c/em\u003e \u003cb\u003e20\u003c/b\u003e, 2619\u0026ndash;2630 (2008).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDemko, V. et al. Genetic Analysis of \u003cem\u003eDEFECTIVE KERNEL1\u003c/em\u003e Loop Function in Three-Dimensional Body Patterning in \u003cem\u003ePhyscomitrella patens\u003c/em\u003e. \u003cem\u003ePlant Physiol.\u003c/em\u003e \u003cb\u003e166\u003c/b\u003e, 903\u0026ndash;919 (2014).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLid, S. E. et al. The \u003cem\u003edefective kernel 1\u003c/em\u003e (\u003cem\u003edek1\u003c/em\u003e) gene required for aleurone cell development in the endosperm of maize grains encodes a membrane protein of the calpain gene superfamily. \u003cem\u003eProc. Natl. Acad. Sci. U.S.A.\u003c/em\u003e 99, 5460\u0026ndash;5465 (2002).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eOlsen, O. A., Perroud, P. F., Johansen, W. \u0026amp; Demko, V. DEK1; missing piece in puzzle of plant development. \u003cem\u003eTrends Plant Sci.\u003c/em\u003e \u003cb\u003e20\u003c/b\u003e, 70\u0026ndash;71 (2015).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eJohansen, W. et al. The DEK1 calpain Linker functions in three-dimensional body patterning in Physcomitrella patens. \u003cem\u003ePlant Physiol.\u003c/em\u003e pp.00925. (2016) (2016). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1104/pp.16.00925\u003c/span\u003e\u003cspan address=\"10.1104/pp.16.00925\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDemko, V. et al. Regulation of developmental gatekeeping and cell fate transition by the calpain protease DEK1 in Physcomitrium patens. \u003cem\u003eCommun. Biol.\u003c/em\u003e \u003cb\u003e7\u003c/b\u003e, 261 (2024).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAko, A. E. et al. An intragenic mutagenesis strategy in Physcomitrella patens to preserve intron splicing. \u003cem\u003eSci. Rep.\u003c/em\u003e \u003cb\u003e7\u003c/b\u003e, 5111 (2017).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePerroud, P. et al. Defective Kernel 1 (DEK 1) is required for three-dimensional growth in \u003cem\u003eP hyscomitrella patens\u003c/em\u003e. \u003cem\u003eNew Phytol.\u003c/em\u003e \u003cb\u003e203\u003c/b\u003e, 794\u0026ndash;804 (2014).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eNavarro, S. \u0026amp; Ventura, S. Computational re-design of protein structures to improve solubility. \u003cem\u003eExpert Opin. Drug Discov.\u003c/em\u003e \u003cb\u003e14\u003c/b\u003e, 1077\u0026ndash;1088 (2019).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eTrainor, K., Broom, A. \u0026amp; Meiering, E. M. Exploring the relationships between protein sequence, structure and solubility. \u003cem\u003eCurr. Opin. Struct. Biol.\u003c/em\u003e \u003cb\u003e42\u003c/b\u003e, 136\u0026ndash;146 (2017).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGupta, J., Nunes, C., Vyas, S. \u0026amp; Jonnalagadda, S. Prediction of Solubility Parameters and Miscibility of Pharmaceutical Compounds by Molecular Dynamics Simulations. \u003cem\u003eJ. Phys. Chem. B\u003c/em\u003e. \u003cb\u003e115\u003c/b\u003e, 2014\u0026ndash;2023 (2011).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGanugapati, J. \u0026amp; Akash, S. Multi-template homology based structure prediction and molecular docking studies of protein \u0026lsquo;L\u0026rsquo; of Zaire ebolavirus (EBOV). \u003cem\u003eInf. Med. Unlocked\u003c/em\u003e. \u003cb\u003e9\u003c/b\u003e, 68\u0026ndash;75 (2017).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLu, H., Cheng, Z., Hu, Y. \u0026amp; Tang, L. V. What Can De Novo Protein Design Bring to the Treatment of Hematological Disorders? \u003cem\u003eBiology\u003c/em\u003e 12, 166 (2023).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRoy, A., Yang, J. \u0026amp; Zhang, Y. COFACTOR: an accurate comparative algorithm for structure-based protein function annotation. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cb\u003e40\u003c/b\u003e, W471\u0026ndash;W477 (2012).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHolm, L., Laiho, A., T\u0026ouml;r\u0026ouml;nen, P. \u0026amp; Salgado, M. DALI shines a light on remote homologs: One hundred discoveries. \u003cem\u003eProtein Sci.\u003c/em\u003e \u003cb\u003e32\u003c/b\u003e, e4519 (2023).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eChan, P., Curtis, R. A. \u0026amp; Warwicker, J. Soluble expression of proteins correlates with a lack of positively-charged surface. \u003cem\u003eSci. Rep.\u003c/em\u003e \u003cb\u003e3\u003c/b\u003e, 3333 (2013).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eStuder, G. et al. QMEANDisCo\u0026mdash;distance constraints applied on model quality estimation. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cb\u003e36\u003c/b\u003e, 1765\u0026ndash;1771 (2020).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eVan Den Bedem, H. \u0026amp; Fraser, J. S. Integrative, dynamic structural biology at atomic resolution\u0026mdash;it\u0026rsquo;s about time. \u003cem\u003eNat. Methods\u003c/em\u003e. \u003cb\u003e12\u003c/b\u003e, 307\u0026ndash;318 (2015).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHeo, L. \u0026amp; Feig, M. High-accuracy protein structures by combining machine‐learning with physics‐based refinement. \u003cem\u003eProteins\u003c/em\u003e \u003cb\u003e88\u003c/b\u003e, 637\u0026ndash;642 (2020).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLobanov, M. Y., Bogatyreva, N. S. \u0026amp; Galzitskaya, O. V. Radius of gyration as an indicator of protein structure compactness. \u003cem\u003eMol. Biol.\u003c/em\u003e \u003cb\u003e42\u003c/b\u003e, 623\u0026ndash;628 (2008).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAbouzied, A. S. et al. Structural and free energy landscape analysis for the discovery of antiviral compounds targeting the cap-binding domain of influenza polymerase PB2. \u003cem\u003eSci. Rep.\u003c/em\u003e \u003cb\u003e14\u003c/b\u003e, 25441 (2024).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePace, C. N. et al. Contribution of hydrogen bonds to protein stability. \u003cem\u003eProtein Sci.\u003c/em\u003e \u003cb\u003e23\u003c/b\u003e, 652\u0026ndash;661 (2014).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eJiang, L. \u0026amp; Lai, L. CH\u0026middot;\u0026middot;\u0026middot;O Hydrogen Bonds at Protein-Protein Interfaces. \u003cem\u003eJ. Biol. Chem.\u003c/em\u003e \u003cb\u003e277\u003c/b\u003e, 37732\u0026ndash;37740 (2002).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eTsumoto, K. et al. Role of Arginine in Protein Refolding, Solubilization, and Purification. \u003cem\u003eBiotechnol. Prog\u003c/em\u003e. \u003cb\u003e20\u003c/b\u003e, 1301\u0026ndash;1308 (2004).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eStrub, C. et al. Mutation of exposed hydrophobic amino acids to arginine to increase protein stability. \u003cem\u003eBMC Biochem.\u003c/em\u003e \u003cb\u003e5\u003c/b\u003e, 9 (2004).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWarwicker, J., Charonis, S. \u0026amp; Curtis, R. A. Lysine and Arginine Content of Proteins: Computational Analysis Suggests a New Tool for Solubility Design. \u003cem\u003eMol. Pharm.\u003c/em\u003e \u003cb\u003e11\u003c/b\u003e, 294\u0026ndash;303 (2014).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKramer, R. M., Shende, V. R., Motl, N., Pace, C. N. \u0026amp; Scholtz, J. M. Toward a Molecular Understanding of Protein Solubility: Increased Negative Surface Charge Correlates with Increased Solubility. \u003cem\u003eBiophys. J.\u003c/em\u003e \u003cb\u003e102\u003c/b\u003e, 1907\u0026ndash;1915 (2012).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMills, B. J. \u0026amp; Laurence Chadwick, J. S. Effects of localized interactions and surface properties on stability of protein-based therapeutics. \u003cem\u003eJ. Pharm. Pharmacol.\u003c/em\u003e \u003cb\u003e70\u003c/b\u003e, 609\u0026ndash;624 (2018).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eTrevino, S. R., Scholtz, J. M. \u0026amp; Pace, C. N. Measuring and Increasing Protein Solubility. \u003cem\u003eJ. Pharm. Sci.\u003c/em\u003e \u003cb\u003e97\u003c/b\u003e, 4155\u0026ndash;4166 (2008).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKuhn, A. B. et al. Improved Solution-State Properties of Monoclonal Antibodies by Targeted Mutations. \u003cem\u003eJ. Phys. Chem. B\u003c/em\u003e. \u003cb\u003e121\u003c/b\u003e, 10818\u0026ndash;10827 (2017).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGhahremanian, S., Rashidi, M. M., Raeisi, K. \u0026amp; Toghraie, D. Molecular dynamics simulation approach for discovering potential inhibitors against SARS-CoV-2: A structural review. \u003cem\u003eJ. Mol. Liq.\u003c/em\u003e \u003cb\u003e354\u003c/b\u003e, 118901 (2022).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eXiao, S. et al. Rational modification of protein stability by targeting surface sites leads to complicated results. \u003cem\u003eProc. Natl. Acad. Sci. U.S.A.\u003c/em\u003e 110, 11337\u0026ndash;11342 (2013).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKabsch, W. \u0026amp; Sander, C. Dictionary of protein secondary structure: Pattern recognition of hydrogen-bonded and geometrical features. \u003cem\u003eBiopolymers\u003c/em\u003e \u003cb\u003e22\u003c/b\u003e, 2577\u0026ndash;2637 (1983).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMokmak, W., Chunsrivirot, S., Assawamakin, A., Choowongkomon, K. \u0026amp; Tongsima, S. Molecular dynamics simulations reveal structural instability of human trypsin inhibitor upon D50E and Y54H mutations. \u003cem\u003eJ. Mol. Model.\u003c/em\u003e \u003cb\u003e19\u003c/b\u003e, 521\u0026ndash;528 (2013).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eZhou, H. X. \u0026amp; Pang, X. Electrostatic Interactions in Protein Structure, Folding, Binding, and Condensation. \u003cem\u003eChem. Rev.\u003c/em\u003e \u003cb\u003e118\u003c/b\u003e, 1691\u0026ndash;1741 (2018).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGregory, K. P. et al. Understanding specific ion effects and the Hofmeister series. \u003cem\u003ePhys. Chem. Chem. Phys.\u003c/em\u003e \u003cb\u003e24\u003c/b\u003e, 12682\u0026ndash;12718 (2022).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHyde, A. M. et al. General Principles and Strategies for Salting-Out Informed by the Hofmeister Series. \u003cem\u003eOrg. Process. Res. Dev.\u003c/em\u003e \u003cb\u003e21\u003c/b\u003e, 1355\u0026ndash;1370 (2017).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eTadeo, X., L\u0026oacute;pez-M\u0026eacute;ndez, B., Casta\u0026ntilde;o, D., Trigueros, T. \u0026amp; Millet, O. Protein Stabilization and the Hofmeister Effect: The Role of Hydrophobic Solvation. \u003cem\u003eBiophys. J.\u003c/em\u003e \u003cb\u003e97\u003c/b\u003e, 2595\u0026ndash;2603 (2009).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eTadeo, X., Pons, M. \u0026amp; Millet, O. Influence of the Hofmeister Anions on Protein Stability As Studied by Thermal Denaturation and Chemical Shift Perturbation. \u003cem\u003eBiochemistry\u003c/em\u003e \u003cb\u003e46\u003c/b\u003e, 917\u0026ndash;923 (2007).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSammond, D. W. et al. Structure-based Protocol for Identifying Mutations that Enhance Protein\u0026ndash;Protein Binding Affinities. \u003cem\u003eJ. Mol. Biol.\u003c/em\u003e \u003cb\u003e371\u003c/b\u003e, 1392\u0026ndash;1404 (2007).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eXu, D., Tsai, C. J. \u0026amp; Nussinov, R. Hydrogen bonds and salt bridges across protein-protein interfaces. \u003cem\u003eProtein Eng. Des. Selection\u003c/em\u003e. \u003cb\u003e10\u003c/b\u003e, 999\u0026ndash;1012 (1997).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGoldenzweig, A. \u0026amp; Fleishman, S. J. Principles of Protein Stability and Their Application in Computational Design. \u003cem\u003eAnnu. Rev. Biochem.\u003c/em\u003e \u003cb\u003e87\u003c/b\u003e, 105\u0026ndash;129 (2018).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRosano, G. L. \u0026amp; Ceccarelli, E. A. Recombinant protein expression in Escherichia coli: advances and challenges. \u003cem\u003eFront Microbiol\u003c/em\u003e \u003cb\u003e5\u003c/b\u003e, (2014).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eZheng, W. et al. Folding non-homologous proteins by coupling deep-learning contact maps with I-TASSER assembly simulations. \u003cem\u003eCell. Rep. Methods\u003c/em\u003e. \u003cb\u003e1\u003c/b\u003e, 100014 (2021).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eZhang, C., Freddolino, L. \u0026amp; Zhang, Y. COFACTOR: improved protein function prediction by combining structure, sequence and protein\u0026ndash;protein interaction information. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cb\u003e45\u003c/b\u003e, W291\u0026ndash;W299 (2017).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eŠali, A. \u0026amp; Blundell, T. L. Comparative Protein Modelling by Satisfaction of Spatial Restraints. \u003cem\u003eJ. Mol. Biol.\u003c/em\u003e \u003cb\u003e234\u003c/b\u003e, 779\u0026ndash;815 (1993).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eFiser, A., Šali, A. \u0026amp; Modeller Generation and Refinement of Homology-Based Protein Structure Models. in Methods in Enzymology vol. 374 461\u0026ndash;491 (Elsevier, (2003).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eJumper, J. et al. Highly accurate protein structure prediction with AlphaFold. \u003cem\u003eNature\u003c/em\u003e \u003cb\u003e596\u003c/b\u003e, 583\u0026ndash;589 (2021).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMirdita, M. et al. ColabFold: making protein folding accessible to all. \u003cem\u003eNat. Methods\u003c/em\u003e. \u003cb\u003e19\u003c/b\u003e, 679\u0026ndash;682 (2022).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWaterhouse, A. M. et al. The structure assessment web server: for proteins, complexes and more. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cb\u003e52\u003c/b\u003e, W318\u0026ndash;W323 (2024).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eZheng, W. et al. Deep-learning-based single-domain and multidomain protein structure prediction with D-I-TASSER. \u003cem\u003eNat. Biotechnol.\u003c/em\u003e \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41587-025-02654-4\u003c/span\u003e\u003cspan address=\"10.1038/s41587-025-02654-4\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2025).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKo, J., Park, H., Heo, L. \u0026amp; Seok, C. GalaxyWEB server for protein structure prediction and refinement. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cb\u003e40\u003c/b\u003e, W294\u0026ndash;W297 (2012).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eFiser, A. \u0026amp; Sali, A. ModLoop: automated modeling of loops in protein structures. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cb\u003e19\u003c/b\u003e, 2500\u0026ndash;2501 (2003).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eL\u0026uuml;thy, R., Bowie, J. U. \u0026amp; Eisenberg, D. Assessment of protein models with three-dimensional profiles. \u003cem\u003eNature\u003c/em\u003e \u003cb\u003e356\u003c/b\u003e, 83\u0026ndash;85 (1992).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eZhang, C., Shine, M., Pyle, A. M. \u0026amp; Zhang, Y. US-align: universal structure alignments of proteins, nucleic acids, and macromolecular complexes. \u003cem\u003eNat. Methods\u003c/em\u003e. \u003cb\u003e19\u003c/b\u003e, 1109\u0026ndash;1115 (2022).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eXu, J. \u0026amp; Zhang, Y. How significant is a protein structure similarity with TM-score\u0026thinsp;=\u0026thinsp;0.5? \u003cem\u003eBioinformatics\u003c/em\u003e \u003cb\u003e26\u003c/b\u003e, 889\u0026ndash;895 (2010).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWiederstein, M. \u0026amp; Sippl, M. J. ProSA-web: interactive web service for the recognition of errors in three-dimensional structures of proteins. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cb\u003e35\u003c/b\u003e, W407\u0026ndash;W410 (2007).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eUziela, K., Men\u0026eacute;ndez Hurtado, D., Shu, N., Wallner, B. \u0026amp; Elofsson, A. ProQ3D: improved model quality assessments using deep learning. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cb\u003e33\u003c/b\u003e, 1578\u0026ndash;1580 (2017).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRamachandran, G. N., Ramakrishnan, C. \u0026amp; Sasisekharan, V. Stereochemistry of polypeptide chain configurations. \u003cem\u003eJ. Mol. Biol.\u003c/em\u003e \u003cb\u003e7\u003c/b\u003e, 95\u0026ndash;99 (1963).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLovell, S. C. et al. Structure validation by Cα geometry: ϕ,ψ and Cβ deviation. \u003cem\u003eProteins\u003c/em\u003e \u003cb\u003e50\u003c/b\u003e, 437\u0026ndash;450 (2003).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWilliams, C. J. et al. MolProbity: More and better reference data for improved all-atom structure validation. \u003cem\u003eProtein Sci.\u003c/em\u003e \u003cb\u003e27\u003c/b\u003e, 293\u0026ndash;315 (2018).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBrooks, B. R. et al. The biomolecular simulation program. \u003cem\u003eJ. Comput. Chem.\u003c/em\u003e \u003cb\u003e30\u003c/b\u003e, 1545\u0026ndash;1614 (2009).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMacKerell, A. D. et al. All-Atom Empirical Potential for Molecular Modeling and Dynamics Studies of Proteins. \u003cem\u003eJ. Phys. Chem. B\u003c/em\u003e. \u003cb\u003e102\u003c/b\u003e, 3586\u0026ndash;3616 (1998).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLee, J. et al. CHARMM-GUI Input Generator for NAMD, GROMACS, AMBER, OpenMM, and CHARMM/OpenMM Simulations Using the CHARMM36 Additive Force Field. \u003cem\u003eJ. Chem. Theory Comput.\u003c/em\u003e \u003cb\u003e12\u003c/b\u003e, 405\u0026ndash;413 (2016).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBauer, P., Hess, B. \u0026amp; Lindahl, E. GROMACS 2022 Source code. Zenodo \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.5281/ZENODO.6103835\u003c/span\u003e\u003cspan address=\"10.5281/ZENODO.6103835\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2022).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGalm, L., Amrhein, S. \u0026amp; Hubbuch, J. Predictive approach for protein aggregation: Correlation of protein surface characteristics and conformational flexibility to protein aggregation propensity. \u003cem\u003eBiotech Bioengineering\u003c/em\u003e. \u003cb\u003e114\u003c/b\u003e, 1170\u0026ndash;1183 (2017).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSoares, C. M., Teixeira, V. H. \u0026amp; Baptista, A. M. Protein Structure and Dynamics in Nonaqueous Solvents: Insights from Molecular Dynamics Simulation Studies. \u003cem\u003eBiophys. J.\u003c/em\u003e \u003cb\u003e84\u003c/b\u003e, 1628\u0026ndash;1641 (2003).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eFriedman, R., Nachliel, E. \u0026amp; Gutman, M. Molecular Dynamics of a Protein Surface: Ion-Residues Interactions. \u003cem\u003eBiophys. J.\u003c/em\u003e \u003cb\u003e89\u003c/b\u003e, 768\u0026ndash;781 (2005).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eZhang, Y. \u0026amp; Cremer, P. Interactions between macromolecules and ions: the Hofmeister series. \u003cem\u003eCurr. Opin. Chem. Biol.\u003c/em\u003e \u003cb\u003e10\u003c/b\u003e, 658\u0026ndash;663 (2006).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eJurrus, E. et al. Improvements to the APBS biomolecular solvation software suite. \u003cem\u003eProtein Sci.\u003c/em\u003e \u003cb\u003e27\u003c/b\u003e, 112\u0026ndash;128 (2018).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSormanni, P., Aprile, F. A. \u0026amp; Vendruscolo, M. The CamSol Method of Rational Design of Protein Mutants with Enhanced Solubility. \u003cem\u003eJ. Mol. Biol.\u003c/em\u003e \u003cb\u003e427\u003c/b\u003e, 478\u0026ndash;490 (2015).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHebditch, M., Carballo-Amador, M. A., Charonis, S., Curtis, R. \u0026amp; Warwicker, J. Protein\u0026ndash;Sol: a web tool for predicting protein solubility from sequence. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cb\u003e33\u003c/b\u003e, 3098\u0026ndash;3100 (2017).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eVan Der Spoel, D., Van Maaren, P. J., Larsson, P. \u0026amp; T\u0026icirc;mneanu, N. Thermodynamics of Hydrogen Bonding in Hydrophilic and Hydrophobic Media. \u003cem\u003eJ. Phys. Chem. B\u003c/em\u003e. \u003cb\u003e110\u003c/b\u003e, 4393\u0026ndash;4398 (2006).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003evan der Bondi, A. Waals Volumes and Radii. \u003cem\u003eJ. Phys. Chem.\u003c/em\u003e \u003cb\u003e68\u003c/b\u003e, 441\u0026ndash;451 (1964).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSchulman, J., Wolski, F., Dhariwal, P., Radford, A. \u0026amp; Klimov, O. Proximal Policy Optimization Algorithms. Preprint at (2017). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.48550/ARXIV.1707.06347\u003c/span\u003e\u003cspan address=\"10.48550/ARXIV.1707.06347\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"},{"header":"Tables","content":"\u003cp\u003eTable 1: Benchmarking C-I-TASSER, AlphaFold2, and SWISS-MODEL based on structural accuracy scores.\u003c/p\u003e\n\u003ctable border=\"0\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 132px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eModelling software\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 156px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eQMEANDisCo Global\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 120px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eTM-Score\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd colspan=\"2\" valign=\"top\" style=\"width: 72px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eUS-Score\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 60px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eZ-Score\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 62px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eProQ3D\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 132px;\"\u003e\n \u003cp\u003eC-I-ITASSER\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 156px;\"\u003e\n \u003cp\u003e0.59\u0026plusmn;0.05\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd colspan=\"2\" valign=\"top\" style=\"width: 126px;\"\u003e\n \u003cp\u003e0.70\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 66px;\"\u003e\n \u003cp\u003e0.70\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 60px;\"\u003e\n \u003cp\u003e-9.64\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 62px;\"\u003e\n \u003cp\u003e0.701\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 132px;\"\u003e\n \u003cp\u003eAlphafold2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 156px;\"\u003e\n \u003cp\u003e0.69\u0026plusmn;0.05\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd colspan=\"2\" valign=\"top\" style=\"width: 126px;\"\u003e\n \u003cp\u003e0.72\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 66px;\"\u003e\n \u003cp\u003e0.73\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 60px;\"\u003e\n \u003cp\u003e-6.57\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 62px;\"\u003e\n \u003cp\u003e0.713\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 132px;\"\u003e\n \u003cp\u003eSWISS-MODEL\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 156px;\"\u003e\n \u003cp\u003e0.64\u0026plusmn;0.05\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd colspan=\"2\" valign=\"top\" style=\"width: 126px;\"\u003e\n \u003cp\u003e0.66\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 66px;\"\u003e\n \u003cp\u003e0.68\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 60px;\"\u003e\n \u003cp\u003e-.8.51\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 62px;\"\u003e\n \u003cp\u003e0.688\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cbr\u003e\u003c/p\u003e\n\u003cp\u003eTable. 2: Candidate residues and position number and replaced residues for single-pointed mutants.\u003c/p\u003e\n\u003ctable border=\"0\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 124px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eMutant\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 133px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eOriginal residue\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 61px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eposition\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 156px;\"\u003e\n \u003cp\u003e\u003cstrong\u003emutated residue\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 128px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eStructural type\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 124px;\"\u003e\n \u003cp\u003eMUT1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 133px;\"\u003e\n \u003cp\u003eValine\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 61px;\"\u003e\n \u003cp\u003e17\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 156px;\"\u003e\n \u003cp\u003eArginine\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 128px;\"\u003e\n \u003cp\u003e\u0026alpha;-Helix\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 124px;\"\u003e\n \u003cp\u003eMUT2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 133px;\"\u003e\n \u003cp\u003eLeucine\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 61px;\"\u003e\n \u003cp\u003e19\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 156px;\"\u003e\n \u003cp\u003eSerine\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 128px;\"\u003e\n \u003cp\u003e\u0026alpha;-Helix\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 124px;\"\u003e\n \u003cp\u003eMUT3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 133px;\"\u003e\n \u003cp\u003eIsoleucine\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 61px;\"\u003e\n \u003cp\u003e21\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 156px;\"\u003e\n \u003cp\u003eSerine\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 128px;\"\u003e\n \u003cp\u003e\u0026alpha;-Helix\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 124px;\"\u003e\n \u003cp\u003eMUT4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 133px;\"\u003e\n \u003cp\u003eValine\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 61px;\"\u003e\n \u003cp\u003e47\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 156px;\"\u003e\n \u003cp\u003eArginine\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 128px;\"\u003e\n \u003cp\u003eLoop\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 124px;\"\u003e\n \u003cp\u003eMUT5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 133px;\"\u003e\n \u003cp\u003ePhenylalanine\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 61px;\"\u003e\n \u003cp\u003e291\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 156px;\"\u003e\n \u003cp\u003eTyrosine\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 128px;\"\u003e\n \u003cp\u003eLoop\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 124px;\"\u003e\n \u003cp\u003eMUT6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 133px;\"\u003e\n \u003cp\u003eLeucine\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 61px;\"\u003e\n \u003cp\u003e323\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 156px;\"\u003e\n \u003cp\u003eSerine\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 128px;\"\u003e\n \u003cp\u003e\u0026alpha;-Helix\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 124px;\"\u003e\n \u003cp\u003eMUT7\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 133px;\"\u003e\n \u003cp\u003eValine\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 61px;\"\u003e\n \u003cp\u003e342\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 156px;\"\u003e\n \u003cp\u003eAsparagine\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 128px;\"\u003e\n \u003cp\u003eLoop\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 124px;\"\u003e\n \u003cp\u003eMUT8\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 133px;\"\u003e\n \u003cp\u003eLysine\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 61px;\"\u003e\n \u003cp\u003e7\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 156px;\"\u003e\n \u003cp\u003eArginine\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 128px;\"\u003e\n \u003cp\u003e\u0026alpha;-Helix\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 124px;\"\u003e\n \u003cp\u003eMUT9\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 133px;\"\u003e\n \u003cp\u003eGlutamine\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 61px;\"\u003e\n \u003cp\u003e54\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 156px;\"\u003e\n \u003cp\u003eArginine\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 128px;\"\u003e\n \u003cp\u003e\u0026alpha;-Helix\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 124px;\"\u003e\n \u003cp\u003eMUT10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 133px;\"\u003e\n \u003cp\u003eLysine\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 61px;\"\u003e\n \u003cp\u003e339\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 156px;\"\u003e\n \u003cp\u003eArginine\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 128px;\"\u003e\n \u003cp\u003eLoop\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cbr\u003e\u003c/p\u003e\n\u003cp\u003eTable 3: Fitness-based ranking of protein mutants across mutation levels (description: mutant identifiers follow the format: MUTxyz, where each digit (x, y, z) represents the point of a mutation in the protein sequence. For example, MUT347 indicates mutations includes single mutants of 3, 4, and 7 (as triple mutant).\u003c/p\u003e\n\u003ctable border=\"0\" cellspacing=\"0\" cellpadding=\"0\" width=\"622\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 46px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eRank\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 109px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eSingle Mutant\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 74px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eFitness\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 114px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eDouble Mutant\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 84px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eFitness\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 120px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eTriple Mutant\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 76px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eFitness\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 46px;\"\u003e\n \u003cp\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 109px;\"\u003e\n \u003cp\u003eMUT3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 74px;\"\u003e\n \u003cp\u003e0.1622\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 114px;\"\u003e\n \u003cp\u003eMUT37\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 84px;\"\u003e\n \u003cp\u003e0.1432\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 120px;\"\u003e\n \u003cp\u003eMUT347\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 76px;\"\u003e\n \u003cp\u003e01368\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 46px;\"\u003e\n \u003cp\u003e2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 109px;\"\u003e\n \u003cp\u003eMUT7\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 74px;\"\u003e\n \u003cp\u003e0.1719\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 114px;\"\u003e\n \u003cp\u003eMUT34\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 84px;\"\u003e\n \u003cp\u003e0.1443\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 120px;\"\u003e\n \u003cp\u003eMUT246\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 76px;\"\u003e\n \u003cp\u003e0.1475\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 46px;\"\u003e\n \u003cp\u003e3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 109px;\"\u003e\n \u003cp\u003eMUT2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 74px;\"\u003e\n \u003cp\u003e0.1759\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 114px;\"\u003e\n \u003cp\u003eMUT67\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 84px;\"\u003e\n \u003cp\u003e0.1498\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 120px;\"\u003e\n \u003cp\u003eMUT367\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 76px;\"\u003e\n \u003cp\u003e0.1498\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 46px;\"\u003e\n \u003cp\u003e4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 109px;\"\u003e\n \u003cp\u003eMUT5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 74px;\"\u003e\n \u003cp\u003e0.1775\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 114px;\"\u003e\n \u003cp\u003eMUT35\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 84px;\"\u003e\n \u003cp\u003e0.1519\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 120px;\"\u003e\n \u003cp\u003eMUT345\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 76px;\"\u003e\n \u003cp\u003e0.1519\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 46px;\"\u003e\n \u003cp\u003e5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 109px;\"\u003e\n \u003cp\u003eMUT4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 74px;\"\u003e\n \u003cp\u003e0.1803\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 114px;\"\u003e\n \u003cp\u003eMUT24\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 84px;\"\u003e\n \u003cp\u003e0.1696\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 120px;\"\u003e\n \u003cp\u003eMUT245\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 76px;\"\u003e\n \u003cp\u003e0.1705\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 46px;\"\u003e\n \u003cp\u003e6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 109px;\"\u003e\n \u003cp\u003eMUT6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 74px;\"\u003e\n \u003cp\u003e0.1850\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 114px;\"\u003e\n \u003cp\u003eMUT25\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 84px;\"\u003e\n \u003cp\u003e0.1705\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 120px;\"\u003e\n \u003cp\u003eMUT136\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 76px;\"\u003e\n \u003cp\u003e0.1776\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 46px;\"\u003e\n \u003cp\u003e7\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 109px;\"\u003e\n \u003cp\u003eMUT1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 74px;\"\u003e\n \u003cp\u003e0.2163\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 114px;\"\u003e\n \u003cp\u003eMUT26\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 84px;\"\u003e\n \u003cp\u003e0.1712\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 120px;\"\u003e\n \u003cp\u003eMUT247\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 76px;\"\u003e\n \u003cp\u003e0.1909\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 46px;\"\u003e\n \u003cp\u003e8\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 109px;\"\u003e\n \u003cp\u003eMUT8\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 74px;\"\u003e\n \u003cp\u003e0.2184\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 114px;\"\u003e\n \u003cp\u003eMUT36\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 84px;\"\u003e\n \u003cp\u003e0.1741\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 120px;\"\u003e\n \u003cp\u003eMUT236\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 76px;\"\u003e\n \u003cp\u003e0.1950\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 46px;\"\u003e\n \u003cp\u003e9\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 109px;\"\u003e\n \u003cp\u003eMUT10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 74px;\"\u003e\n \u003cp\u003e0.2209\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 114px;\"\u003e\n \u003cp\u003eMUT45\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 84px;\"\u003e\n \u003cp\u003e0.1775\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 120px;\"\u003e\n \u003cp\u003eMUT457\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 76px;\"\u003e\n \u003cp\u003e0.2101\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 46px;\"\u003e\n \u003cp\u003e10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 109px;\"\u003e\n \u003cp\u003eMUT9\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 74px;\"\u003e\n \u003cp\u003e0.2226\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 114px;\"\u003e\n \u003cp\u003eMUT13\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 84px;\"\u003e\n \u003cp\u003e0.1776\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 120px;\"\u003e\n \u003cp\u003eMUT146\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 76px;\"\u003e\n \u003cp\u003e0.2551\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Calpain, Protein solubility, Mutagenesis, Molecular Dynamics, Structure prediction","lastPublishedDoi":"10.21203/rs.3.rs-7559073/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7559073/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThe DEFECTIVE KERNEL 1 (DEK1) protein plays essential functions throughout plant development. DEK1 is a multidomain 250 kDa protein with yet unsolved 3D structure. To facilitate structural and functional studies of DEK1, here we investigate its calpain protease core domain (CysPc) from \u003cem\u003ePhyscomitrium patens\u003c/em\u003e. Using integrated structural modelling we propose targeted mutagenesis of CysPc to enhance its solubility during recombinant protein production. We created a pipeline to predict the topology of the CysPc domain with improved precision, providing a robust framework for further exploration. We evaluated the native and mutant structures by MD simulations, concentrating on several solubility-related parameters. Following these features, we implemented specific single, double, and triple amino acid mutagenesis to select variants with improved solubility. Our method preserves overall structural integrity while reducing aggregation-prone traits. We advocate for the utilization of reinforcement learning method that can effectively traverse the extensive combinatorial space and prioritize mutation sets with the greatest potential for enhancing solubility. This framework provides a logical, data-driven approach to improving protein solubility, particularly beneficial in situations lacking high-resolution structural data.\u003c/p\u003e","manuscriptTitle":"Computational optimization of DEK1 calpain domain solubility through integrated structural modelling and targeted mutagenesis","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-10-03 14:32:55","doi":"10.21203/rs.3.rs-7559073/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2025-10-30T04:44:45+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-10-28T22:09:39+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-10-24T21:37:32+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-10-19T06:04:58+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"92025853155972569015045493046266422092","date":"2025-10-08T13:05:16+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"95109557289946644199991063485273349842","date":"2025-10-06T09:35:57+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"277814455956318454359722993711537146847","date":"2025-10-03T13:44:51+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"311752404777006880951731795325961999184","date":"2025-10-03T13:00:24+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-09-22T12:59:33+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2025-09-12T04:46:32+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-09-10T12:20:32+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-09-09T08:23:01+00:00","index":"","fulltext":""},{"type":"submitted","content":"Scientific Reports","date":"2025-09-08T01:50:15+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"667deae5-5318-47c1-9c33-658e76e6d2c2","owner":[],"postedDate":"October 3rd, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":55691025,"name":"Biological sciences/Biochemistry"},{"id":55691026,"name":"Biological sciences/Biophysics"},{"id":55691027,"name":"Biological sciences/Computational biology and bioinformatics"},{"id":55691028,"name":"Biological sciences/Structural biology"}],"tags":[],"updatedAt":"2026-02-09T16:07:02+00:00","versionOfRecord":{"articleIdentity":"rs-7559073","link":"https://doi.org/10.1038/s41598-026-38805-z","journal":{"identity":"scientific-reports","isVorOnly":false,"title":"Scientific Reports"},"publishedOn":"2026-02-08 15:58:44","publishedOnDateReadable":"February 8th, 2026"},"versionCreatedAt":"2025-10-03 14:32:55","video":"","vorDoi":"10.1038/s41598-026-38805-z","vorDoiUrl":"https://doi.org/10.1038/s41598-026-38805-z","workflowStages":[]},"version":"v1","identity":"rs-7559073","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7559073","identity":"rs-7559073","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.