Leveraging protein representations to explore uncharted fold spaces with generative models | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Leveraging protein representations to explore uncharted fold spaces with generative models Bruno Correia, Yangyang Miao, Martin Pacesa, Sandrine Georgeon, and 3 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7827245/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted You are reading this latest preprint version Abstract A major challenge in computational de novo protein design is the exploration of uncharted areas within the protein structural space. However, the large degrees of freedom of protein backbones complicate the sampling process during protein design. Machine learning-based models have made great strides in this problem, however due to their nature they tend to exploit rather than explore the data distribution used for training the neural networks. To address some of these challenges, we propose a new coarse grained protein structure representation generative method, DiffTopo , a diffusion model which increases the sampling efficiency and diversity. Combined with a backbone level protein generative model like RFdiffusion, novel protein folds can be generated rapidly, allowing for efficient exploration of the designable topology space. Interestingly, we have discovered that by mirroring the topological organization of native proteins using a pipeline named MirrorTopo , we can readily expand the known fold space. We generated and experimentally characterized 30 different novel topologies from DiffTopo and 6 different novel mirror topologies from MirrorTopo. The developed framework relying on low resolution sampling provides new means for fold exploration challenges, which could in principle enhance our knowledge of the first principles of protein structure and folding, as well as create new opportunities for functional design. Biological sciences/Computational biology and bioinformatics Biological sciences/Structural biology computational protein design dark-matter folds de novo design machine learning for molecular generation deep learning for protein design mirror topology design Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Introduction Proteins govern biological functions in nature through their three-dimensional structure, yet the vast space of possible folds remains largely unexplored 1 . While natural evolution has populated only a fraction of this space, computational methods now offer a route to systematically chart unexplored structural spaces 2 – 5 . One of the important challenges in protein design is to efficiently explore the designable protein fold space without being confined to the training data and in this way favor exploitation of that data distribution, rather than exploration. Here, we present a framework that explores the fold space using coarse-grained representations to traverse this landscape, enabling discovery of novel folds. Recent advances in deep generative models, particularly diffusion-based methods, have dramatically expanded the capabilities of de novo protein design. Recently, many deep generative models have been developed that can generate atomically detailed protein structures with complex topologies, such as RFdiffusion 6 , FrameDiff 6 , and Chroma 7 , among many other deep generative models 8 – 12 . However, despite their success, these methods still tend to sample structures that resemble natural folds, especially for small proteins. This reflects a fundamental limitation of coordinate-level generative modeling: the tendency to recapitulate training distributions and bias the design space toward naturally occurring solutions. As a result, large regions of the theoretically designable but evolutionarily unoccupied fold space, so-called "dark" structural matter, remain underexplored. We hypothesized that an efficient exploration of the fold space requires a different approach: Rather than sampling structures at the atomic level, the exploration should first occur in a low-resolution topological space where structural diversity and fold novelty is more accessible (Fig. 1 a). Drawing inspiration from representations of proteins as hierarchical assemblies of secondary structure elements 13 – 17 , we introduce a coarse-grained (CG) topology description that describes folds into spatial arrangements of \(\:\alpha\:\) -helices and \(\:\beta\:\) -strands (Fig. 1 b). This reduced-dimensionality manifold smooths the sampling landscape, allowing for more coarse sampling of the structural space and avoiding the combinatorial explosion of atomic representations. We developed DiffTopo, which learns fold topologies through diffusion modeling guided by secondary structure descriptors (a string of secondary structure types), allowing us to sample regions far from natural folds while keeping physical plausibility. Coarse-grained topological exploration of protein structure provides an approximate placement of secondary structure elements in 3D space; subsequently, obtaining a full atomic backbone requires a high-resolution generative model. Based on the CG topology length, we first assign standard secondary structure elements and place them at the corresponding spatial positions, resulting in an intermediate structure composed solely of idealized secondary elements, which we refer to as a protein sketch . Our framework bridges this gap by refining these sketches and using them to guide RFdiffusion⁵ generation, ultimately producing high-resolution, designable protein backbones. Importantly, this separation between fold discovery (DiffTopo), through CG sampling and structure generation (RFdiffusion), via atomic-level modeling, helps overcome some of the limitations of directly generating full atomic coordinates, which often tends to recapture existing topologies from known fold space. We further tested this by attempting a distinct strategy of structural generation, where we systematically generated mirrored image topologies of natural folds, which are absent from natural repertoires but are physically realistic and potentially designable, we dubbed this approach MirrorTopo. These two challenges of generating novel-to-nature dark folds and mirrored topologies demonstrate how coarse-grained sampling can reveal underexplored regions of structural space. Biochemical and structural validation confirmed the potential of this approach, where 40 designed dark folds and 22 mirror folds were experimentally characterized, with five crystal structures closely matching the computational models. These results suggest that our method can access folds beyond the capabilities of current generative approaches. By decoupling fold exploration from atomic-level sampling, we establish a new strategy to explore protein space that prioritizes discovery over optimization, a critical step toward understanding the full scope of possible protein architectures and their potential functions. Results Enhancing the sampling of structural diversity with DiffTopo To test the capabilities of the DiffTopo-RFdiffusion pipeline, we generated 3000 backbone structures for each of the following categories: I) random secondary structure descriptors guided backbones; II) low-frequency secondary structure descriptors guided backbones; III) random lengths ranging from 40 to 200 residues backbone generated by RFdiffusion, as reference. For categories I and II, we describe the secondary structure descriptor guided generation in detail in the methods section. To test whether the CG topology correctly captures the target protein fold and whether the RFdiffusion-generated backbones recover the intended structure, we compared protein sketches derived from CG topology with native structures to ensure the model learns the correct structural features, and with the corresponding RFdiffusion-generated backbones to verify that the generated outputs preserve the intended topology. Most protein sketches match well with the generated backbones (RMSD < 3 Å), confirming that the CG sketches can guide the sampling of RF-diffusion for the generation of all-atom protein backbones (Fig. 2 a). To assess whether the model effectively learned the geometric distribution of sketch data, we analyzed its geometric descriptors and evaluated their alignment with native structures. DiffTopo-generated sketches match the topological patterns observed in native structures, showing that the model effectively learns an accurate approximation of real data distributions (Supplementary Fig. 1a-c). We then assessed the designability of the generated backbones from different types of protocol by designing amino acid sequences using ProteinMPNN 19 and predicting whether the designed sequences could fold into the target topology with AlphaFold2(AF2) 20 . Following prior studies on protein backbone generation 20 , 21 , we quantified self-consistency using C \(\:\alpha\:\) RMSD (scRMSD, lower values indicate better consistency) and predicted local distance difference test (pLDDT, higher values indicate greater structural confidence). While a trade-off in scRMSD was observed compared to RFdiffusion alone, a substantial proportion of designs still achieved scRMSDs below 3 Å (Fig. 2 b). However, our primary objective was to assess backbone novelty. To this end, we selected backbones with high designability (scRMSD 90) and performed structural similarity searches against the CATH database 23 , 24 with FoldSeek 25 , reporting the highest template modeling score (Max PDB TM-score 26 , 27 ). By utilizing low-frequency secondary structure descriptor guidance, our approach produced significantly more structurally novel and highly designable backbones compared to baseline RFdiffusion (Fig. 2 c). This suggests that DiffTopo-RFdiffusion can generate both native-like folds and previously unobserved (dark) folds, demonstrating generalization beyond the structural space of the Protein Data Bank (PDB). As an example, for the secondary structure descriptor seed "EEEHHEHE", we generated 500 designable backbones which resulted in 109 structurally distinct clusters (Fig. 2 d), using a TM-score threshold of 0.5 28 .These findings highlight the ability of our framework to efficiently generate a diverse range of de novo protein structures with low structural redundancy. In summary, the DiffTopo-RFdiffusion framework enhances de novo protein design by generating designable and structurally novel backbones. Compared to standalone RFdiffusion, it significantly improves backbone novelty by leveraging sampling in low-frequency secondary structure order regions, enabling the discovery of both native-like and entirely dark folds beyond the known protein space. Exploring de novo “dark” folds with DiffTopo A primary goal of our work is to expand the accessible protein fold space by generating structures that are not only new to nature but also extend beyond the range of folds observed in natural proteins. To systematically assess the model’s capacity for fold exploration, we sampled 779 different dark folds with distinct secondary structure compositions (all-α, all-β, and α/β), and quantified their statistics on the designability, novelty, and diversity (Supplementary Table 1). At the sequence level none of the designed sequences displayed a hit below the significant E-value threshold of 0.001 (Supplementary Table 2) relative to the non-redundant database (NR) 29 , indicating that the designed sequences are novel and distinct from naturally occurring sequences. For additional validation, we applied the HHpred 30 , 31 to search for potential distant homologs. Intriguingly, all sequences showed E-values above 0.5, further supporting their novelty compared to known natural sequences (Supplementary Table 3). To further evaluate the structural quality and novelty of the generated designs, we analyzed sequence-structure consistency, architectural diversity, and similarity to known folds (Fig. 2 e–h). As shown in Fig. 2 e, DiffTopo produced backbones with high designability across a range of secondary structure compositions, indicating strong sequence-structure compatibility. To illustrate this, we selected representative examples of all-α, all-β, and α/β folds and showed their corresponding structures in the figure; corresponding secondary structure descriptors are provided in Supplementary Table 4. The resulting protein sketches exhibited substantial architectural diversity (Fig. 2 f), showing that the model explores a wide topological space. Structural alignment with AF2 predictions confirmed the internal consistency of these designs (Fig. 2 g), supporting their plausibility. The designs were compared to naturally occurring structural databases using Foldseek, in which the closest hits had very low structural similarity, underscoring the potential of the approach for discovering novel, previously uncharacterized folds (Fig. 2 h). To evaluate the novelty of the generated folds, we compared all dark folds to CATH domain representatives using TM-align and projected the TM-score matrix into 2D using classical multidimensional scaling (MDS) (Supplementary Fig. 2). The resulting distribution showed that dark folds form distinct structural clusters, separate from natural domain families. Even for well-populated classes like 7-helix bundles, with 394 CATH entries, DiffTopo still generated 16 unique topologies (Max TMscore < 0.5) not accounted for in the database (Supplementary Fig. 3, Supplementary Table 5). To further explore their global distribution in fold space, we applied the latest structural analysis method SHAPES 32 (structural and hierarchical assessment of proteins with embedding similarity), projecting structures into two orthogonal spaces, ProtDomainSegmentor embedding 33 and ProteinMPNN layer 1 embedding (see Methods: Mapping Protein structures to CATH database). As shown in Supplementary Fig. 4, each embedding space was rasterized into a 16×16 grid, and representative structures were selected per cell to visualize sampling coverage. While CATH domains densely populated specific regions, the dark folds broadly covered previously unoccupied areas, supporting that our method systematically samples novel and underexplored fold topologies. In summary, our pipeline demonstrates its ability to uncover novel, designable and complex folds that do not exist in nature, encompassing various secondary structure combinations, including all \(\:\alpha\:\) , all \(\:\beta\:\) , and mixed \(\:\alpha\:/\beta\:\:\) folds. DiffTopo dark folds designs exhibit high structural precision We computationally designed 778 dark folds that passed the designability filters. For experimental validation, we selected 12 all-α, 7 all-β, and 21 α/β dark fold designs. These proteins were expressed in Escherichia coli and purified using a Ni²⁺-NTA affinity column. The purified proteins were characterized by size-exclusion chromatography combined with multi-angle light scattering (SEC–MALS) and circular dichroism (CD) spectroscopy. Of 44 tested dark fold topologies, 22 designed proteins were well expressed and highly soluble. Additionally, all designs exhibited CD spectra typical of folded proteins and consistent with the designed secondary structure content (Supplementary Fig. 6–10). The designs were thermally very stable as is typical in de novo designed proteins, and were confirmed to be monomeric in solution by SEC–MALS (Supplementary Fig. 5a-w). From this set, we selected 6 representative dark folds, including all-α (N34), all-β (N40), and α/β (N5, N10, N30, N37), which are highlighted in Fig. 3 a. Crystal structures were solved for 3 of these designs: N30, N37, and N5. For N30 and N37, the crystal structures aligned especially well with the computational models, yielding backbone RMSDs of 0.58 Å and 0.70 Å, respectively (Fig. 3 b,c). FoldSeek analysis confirmed minimal similarity to known protein structures (Foldseek CATH max TM-score = 0.488 in), validating their structural novelty. For example, while N37 partially overlapped with ribonuclease H-like β-sheets (Foldseek CATH max TM-score = 0.55 for shared region), its overall topology, α-helices wrapping a β-sheet, is absent from the natural repertoire. Notably, N5, an α/β dark fold (max TM-score = 0.43), crystallized as a dimer with a β-strand-swapped interface, however, SEC–MALS confirmed its monomeric state in solution, and alignment of the non-swapped regions (35–161) revealed strong agreement with the design model (full atom-RMSD = 0.66 Å), suggesting that the strand swapping event is likely a crystallization induced artefact. These results demonstrate that dark fold designs generated by DiffTopo are not only highly designable and diverse, but also structurally accurate, validating the potential of our topology-first framework to generate entirely novel folds absent from nature. MirrorTopo generates new to nature protein folds Although DiffTopo enables the generation of novel folds, the sampled structures still follow natural-like spatial organization. To further expand the design space, we developed MirrorTopo, which generates symmetry-inverted (mirrored) versions of natural folds at the topological level, producing structures that are physically valid but have not been observed in current structural databases. MirrorTopo operates by mirroring CG topology sketches in 3D coordinate space prior to atomic-level generation using RFdiffusion. Mirror folds exhibit distinct pairwise patterns of secondary structure stacking, specifically in the dihedral angles and distances between helix–helix, helix–strand, and strand–strand pairs compared to natural proteins (Supplementary Fig. 1d,). These patterns fall outside the typical distribution learned by DiffTopo, suggesting that such topologies are unlikely to be sampled. Unlike previous research 34 , 35 that mirrored atomic-level coordinates using D-amino acids and right-handed helices, MirrorTopo maintains L-amino acids and right-handed helices, but systematically inverts how secondary structure elements are arranged in space (Fig. 4 a). This approach enables the sampling of mirrored topologies that DiffTopo cannot access allowing us to explore fold topologies that are physically valid yet absent from the known natural repertoire. First, we selected 8 native fold types from the CATH database and generated their mirrored topologies using our MirrorTopo pipeline. For each fold, multiple designs passed our designability filters (self-consistency RMSD 90). We then selected 21 of these designs for experimental testing. Among them, 13 designs were successfully expressed in E.-coli and showed expected physicochemical properties, indicating folded proteins. These 13 designs cover 6 different mirror folds from native proteins, spanning a variety of α/β architectures, representing diverse structural classes. All detailed description and structural annotations and CATH classifications are summarized in Supplementary Table 6. The first column of Fig. 4 b shows native protein structures from the CATH database, followed by their 2D secondary-structure schematics in the second column. Their mirrored counterparts, obtained by flipping the arrangement of secondary-structure elements, are presented in the third column, with the corresponding backbone models generated by MirrorTopo shown in the fourth column. The rightmost column displays FoldSeek search results, confirming that the mirror folds are structurally novel with no close matches in existing structural databases. Mirrored topologies show high structural precision For the 6 different mirror folds in Fig. 4 b, we selected one representative design from each, all of which showed CD spectra consistent with the expected secondary structure content (Fig. 5 a) and SEC–MALS profiles indicative of monomeric states in solution (Supplementary Fig. 11a-l). CD spectra and melting temperature profiles for the remaining successful designs are provided in Supplementary Figs. 12–14. We successfully solved the crystal structures of two mirror fold designs, M16 and M7. M16, the mirror version of the histidine kinase-like ATPase fold, exhibited remarkable agreement with its AF2 prediction (Cα RMSD = 0.42 Å; all-atom RMSD = 1.50 Å), confirming the high structural accuracy of the design (Fig. 5 b). Similarly, M7, a mirror of the TRAPP domain (β-lactamase-like topology), also closely matched its computational model (Cα RMSD = 0.67 Å; all-atom RMSD = 1.50 Å) (Fig. 5 c). These results validate that mirror topologies, that although absent from nature, can be realized with high fidelity using computational methods and experimentally validated. To evaluate whether the tertiary structure motifs in our mirror folds are also found in natural proteins, we used TERMANAL 36 , 37 . According to previous work by Mackenzie et al., about 600 tertiary structural motifs (TERMs) are sufficient to describe 50% of structures in the PDB. TERMANAL analysis returns two main results: Abundance and Structure Score. Abundance reflects how frequently a particular TERM (tertiary structural module) appears across the protein structure database, with higher values indicating that the TERM is common and likely favorable in natural proteins, whereas lower values indicate rarity and potential structural novelty. Structure Score provides a global assessment of a protein structure, integrating the abundance of all its constituent TERMs to measure overall structural reasonableness. Here, we performed TERM analysis on 50 random CATH folds, all experimentally generated dark folds, and all mirror folds. As shown in Fig. 6 a, CATH folds exhibit high motif abundance, with an average of ~ 0.8 per TERM. Dark folds have moderately lower abundance (~ 0.68), while mirror folds display the lowest values (~ 0.44). A similar trend was observed in Structure Scores. To illustrate these results at the residue level, we selected two representative dark folds (N30 and N37) and two mirror folds (M16 and M7), as shown in Fig. 6 b,c.The results showed that mirror folds (M16 and M7) have low motif abundance and low structural similarity scores across all regions, suggesting that their local structures are distinct from those commonly found in nature. In contrast, the dark folds N30 and N37 had high motif abundance and structural scores, especially N37, which was similar to native-like topologies. N30 also showed strong matches, except for its β-sheet region, which had fewer matching motifs. These results support our previous observation that DiffTopo samples structures that remain within a “natural-like” distribution, while mirror folds generated by MirrorTopo lie in a different and less explored region of the structural space. In other words, mirror folds access areas of the fold space that are not well covered by natural evolution. Both computational results and experimental validation demonstrate that the MirrorTopo pipeline is viable for generating structurally precise, stable, and novel proteins. The success of mirror-topology designs achieved using right-handed helices with canonical L-amino acids challenges conventional assumptions about foldability and suggests the existence of a vast, underexplored mirror-fold space. These findings highlight the potential of mirror topologies as a complementary approach to expand the universe of designable protein structures. Discussion Exploring and expanding the protein fold space remains an important challenge in protein design, particularly given that new structures could, in principle, enable novel functional roles for proteins. Existing generative models tend to stay close to natural-like topologies due to fine-grained representations and biases in training data. In contrast, our method uses a compact coarse-grained topology representation, that captures only the relative positions and connections of secondary structure elements. With this abstraction, DiffTopo efficiently learns spatial patterns and samples diverse and novel fold topologies beyond those seen in nature. Embedding analysis (Supplementary Fig. 4) shows that many designs occupy underexplored regions of fold space. Despite their novelty, these structures still retain high designability, allowing for sequences that are in good agreement with the structural templates resulting in structural predictions with high AF2 confidence. These results suggest a large, untapped fold space that can now be explored systematically. While DiffTopo expands the diversity of folds, it still follows the spatial patterns found in natural proteins. To move further beyond these constraints, we introduced MirrorTopo, which generates mirrored topologies that are structurally valid but evolutionarily absent. These mirror folds use standard L-amino acids and form right-handed helices, yet their global secondary structure arrangement is reversed. They do not resemble natural folds, as confirmed by FoldSeek analysis, and occupy regions rarely reached by generative models. Embedding analysis shows these mirror designs fall far outside the distribution of natural structures (Supplementary Fig. 1d), suggesting that mirror fold space represents a distinct, orthogonal region of protein architecture. The absence of mirror folds in nature could reflect evolutionary selection rather than physical constraints. Many natural folds have emerged through duplication, fusion, and recombination events that favor certain packing orientations. Mirrored architectures, while physically allowed, may be disfavored due to folding kinetics, sequence-to-structure mapping biases, or evolutionary path dependence. By systematically constructing these mirror topologies, MirrorTopo reveals a vast, underexplored region of the fold space that has remained invisible to both evolution and previous generative methods. This space is not only accessible, but it is designable, stable, and experimentally verifiable. To test whether these topologies are physically viable, we experimentally solved five crystal structures across both DiffTopo and MirrorTopo designs. These include dark folds (N30, N37, N5) and mirror folds (M16, M7). The results confirmed that these unconventional topologies can fold precisely into their intended atomic structures, supporting the robustness of our design approach. Among all dark folds, N30 displays a particularly entangled topology, while most other designs rely on linear stacking of secondary structures. In its crystal structure, one helix (residues 120–143) threads through a closed cavity formed by two others (61–76 and 81–95), creating a topological constraint unlikely to arise via simple stacking. Sequential AF2 prediction analysis suggests a nonlinear folding pathway: the first two helices fold and pack tightly, and only later open to accommodate the third helix (Supplementary Fig. 15 and Supplementary information: Sequential AF2 predictions analysis for N30). This behavior expands the known repertoire of foldable structures and suggests that topological complexity—such as threading or knot-like arrangements—can be encoded and accessed through designed sequences. N30 highlights the potential of our framework to generate not only novel folds, but also nontrivial folding mechanisms. While this work primarily focuses on structural diversity, our method also supports functional design through motif scaffolding (see Methods: Motif-Scaffolding via Conditional Inpainting with Diffusion Models). By conditioning on existing secondary structure motifs, DiffTopo can also generate diverse scaffolds that preserve motif geometry), highlighting the potential of our framework to support functional scaffold generation, expanding the possibilities for designing proteins with tailored interactions. Overall, our study shows the potential of utilizing different representations to enhance the exploration of ML-based methods for generative protein design showing that under sampled regions of this landscape can be accessed through this approach. We anticipate that these new methods will in future enable the design of proteins with novel biological functions. Methods Coarse-grained topology representation Inspired by the FORM representation 13 and protein sketches 14 , 15 , here protein structures are built by the assembly of their secondary structures, we adopt a simplified representation referred to as a "coarse-grained topology (CG topology)". The topology reduces the intricate protein structure into a stacked arrangement of Secondary Structure Elements (SSEs), achieved by representing secondary structures with three carbon alpha centroids, as illustrated in Fig. 1 a. All the secondary structures in each residue number are assigned by STRIDE 38 program. CG topology involves capturing the geometric position of helices and strands through centroids. For helices, centroids include the first and last four C \(\:\alpha\:\) atoms, as well as the total C \(\:\alpha\:\) atom centroid. Strands are represented by centroids of the first and last two C \(\:\alpha\:\) atoms, along with the centroid C \(\:\alpha\:\) atom within the strand. CG topology can also be easily represented as a protein sketch, which is a rough 3D approximation of a native protein structure with standard SSEs, lacking loops, and AA side chains. Compared with forms and protein sketches, this representation method of CG topology removes the concept of layers, offering a higher degree of freedom in structural representation to represent structures such as beta barrels, while remaining simple with few degrees of freedom. Dataset Using the CG topology approach we can easily convert a protein structure database, such as PDB, from standard protein models to CG topology structures. In this paper, all PDB structural data are sourced from CATH 23 , a database organized as a classification of protein structures. To mitigate the influence of structural redundancy on the data distribution, we employ the CATH-dataset-nonredundant S40 for training data, yielding 31,886 non-redundant protein structures. 90% of structures are used for training and 10% are used for independent validation. DiffTopo/MirrorTopo + RFdiffusion framework First, we introduce DiffTopo, a Equivariant diffusion model 39 , 40 for generating CG topology conditioned on SSEs strings. Based on previous works on denoising diffusion, given a data point sampled from a real data distribution \(\:{z}_{0}\sim\:q\left(z\right)\) , the forward diffusion process is defined by adding Gaussian noise gradually to the data point \(\:{z}_{0}\) from corrupted samples \(\:{z}_{t}\) . For one CG topology data \(\:{z}_{0}\) , there are three points \(\:x\) and corresponding secondary structure type \(\:u\) . At time step \(\:t\) \(\:=0,\dots\:,T\) , the conditional distribution of the intermediate data state \(\:{z}_{t}\) given the previous state is defined by the multivariate normal distribution, $$\:q\left({z}_{t}\mid\:{z}_{t-1}\right)=N\left({z}_{t};\sqrt{1-{\beta\:}_{t}}{z}_{t-1},{\beta\:}_{t}I\right)$$ The step sizes are controlled by a variance schedule \(\:{\left\{{\beta\:}_{t}\in\:\left(\text{0,1}\right)\right\}}_{t=1}^{T}\) . The process is constructed to be Markovian and if we let \(\:{\alpha\:}_{t}=1-{\beta\:}_{t}\) and \(\:\underset{\_}{{\alpha\:}_{t}}={\prod\:}_{i=1}^{t}{\alpha\:}_{i}\) , we can obtain the distribution of \(\:{z}_{t}\) given \(\:z\) $$\:q\left({z}_{t}\mid\:{z}_{0}\right)=N\left({z}_{t};\sqrt{\underset{\_}{{\alpha\:}_{t}}}{z}_{0},\left(1-\underset{\_}{{\alpha\:}_{t}}\right)I\right)$$ For the backward process, we need to learn a model \(\:p\) to approximate these conditional probabilities. $$\:p\left({z}_{t-1}\mid\:{z}_{t},u\right)=N\left({z}_{t-1};\mu\:\left({z}_{t},t,u\right),\varSigma\:\left({z}_{t},t\right)\right)$$ \(\:{\mu\:}_{\theta\:}\left({z}_{t},t,u\right)\) is predicted by the neural network. In this paper function \(\:\mu\:\) that predicts noise \(\:{ϵ}_{\theta\:}\) is implemented as geometric vector perceptron (GVP) 41 . The input to GVP is the noised version of the point coordinates \(\:{z}_{t}\) and point feature \(\:{h}_{t}\) at time \(\:t\) and context \(\:u\) . Note that the predicted noise \(\:ϵ\) includes coordinate and feature components, \(\:ϵ=[{ϵ}_{z}\) , \(\:{ϵ}_{h}]\) . The predicted noise can be calculated by $$\:{ϵ}_{\theta\:}\left({z}_{t},t\right)=GVP\left({z}_{t},{h}_{t},u,t\right)-\left[{z}_{t},{h}_{t}\right]$$ In addition to DiffTopo, we introduce MirrorTopo, a method for generating mirror-image topologies directly from native protein folds. Given a CG representation of the native fold, we obtain the mirror topology by inverting the z-coordinates while preserving the x- and y-coordinates. This transformation effectively reflects the structure across the xy-plane, yielding a topologically mirrored conformation without altering sequence or connectivity. Following the generation of the coarse-grained (CG) topology via the diffusion process or mirroring of the native CG topology, the resulting topology enables the direct extraction of the positions and lengths of standard SSEs. Since CG topology retains a certain level of fuzziness, we need a backbone level model to systematically search for the designable backbone corresponding to CG topology. Here, we leverage RFdiffusion, a highly effective backbone generation model to generate designable backbones from CG topology. Initially, we construct a protein sketch from standard SSEs based on CG topology and utilize the motif modeling function inherent in RFdiffusion to directly connect the SSEs. Subsequently, we employ the partial diffusion approach to identify a reasonable backbone structure within this interconnected structure. Then we employ the ProteinMPNN 19 algorithm for fixed backbone sequence design to obtain amino acid sequences. These sequences are then input for the structure prediction algorithm AF2 19 to predict the structure, which is use to validate the designability of generated backbones. Secondary structure descriptors Secondary structure descriptor construction - To define secondary structure conditions within a specific topological class, we constructed strings representing the desired sequence of secondary structure elements—for example, "EEEHHEHE", where each character denotes a secondary structure type (E: β-strand, H: α-helix). These SSE strings were generated either manually or through algorithmic enumeration by specifying the number, type, and order of elements. The example "EEEHHEHE" encodes a topology with eight secondary structure elements—five β-strands and three α-helices arranged in a particular sequence. These strings were used as input constraints to the generative model, directing the synthesis of backbones consistent with the specified topology. Random Secondary Structure Descriptor Sampling - We randomly selected 512 unique SSE order strings, ranging in length from 3 to 12 elements, from the CATH database. These sequences served as topological constraints to guide the generation of varied backbone conformations. Low-Frequency Secondary Structure Descriptor Sampling - To specifically target rare topologies, we compiled all possible SSE order strings of length 8 and computed their frequency across the entire CATH database. Strings with fewer than two occurrences were classified as low-frequency compositions. This procedure yielded a total of 2,471 underrepresented SSE patterns, which were used to probe the model’s ability to generalize beyond commonly observed folds. Mapping Protein structures to CATH database We computed ProteinMPNN encoder and ProtDomainSegmentor 33 embeddings for all CATH structures from Ingraham et al. 7 , selecting only those without chain breaks, with resolution better than 3.0 Å, and R free below 0.25. Both embedding types are computed at the per-residue level and averaged along the sequence dimension to obtain 128-dimensional embeddings for ProteinMPNN and 4096-dimensional embeddings for ProtDomainSegmentor. To capture structural features at different hierarchical scales, we extract local fold features using the first encoder layer of ProteinMPNN 19 , which encodes featured geometries of nearest neighbors, while global fold features are captured using the pre-final layer of ProtDomainSegmentor, which is predictive of CATH Architecture classification. We project the CATH embeddings onto their first two principal components (PCs) and apply the same transformation to embeddings of dark fold and mirror fold structures for direct comparison. TERM analysis for protein structures For TERM analysis, we randomly selected 50 representative protein structures from the CATH database, including both dark folds and mirror folds that have been experimentally expressed. All structures were queried using the TERMANAL software package against the default fragment database “bc-30-sc-20141022”, a non-redundant subset of protein chains in the Protein Data Bank (PDB) with a 30% sequence identity cutoff and determined by X-ray crystallography at high resolution. The resulting abundance and structural score files were used for subsequent analyses. Protein expression and purification DNA sequences corresponding to the amino acid sequence of the designed proteins were optimized for E. Coli codon usage and purchased from Twist Bioscience as fragments, including a solvent exposed C-terminal poly histidine tag. The DNA fragments were cloned via Gibson cloning into a pET11b backbone, and plasmids were validated by sequencing. Expression was conducted in BL21(DE3) strain: 0.5 L cultures were inoculated from an overnight culture in Luria-Bertani medium supplemented with ampicillin (100 \(\:\mu\:g/mL\) ) and incubated at 37°C and 220 rpm shaking until reaching an OD600 of 0.6–0.8; the cultures were then further incubated overnight at 18°C and inoculated with 0.5 mM IPTG to induce protein expression. Bacterial cells were harvested by centrifugation and pellets were resuspended in 30 mL lysis buffer (100 mM TRIS, pH 7.5, 500 mM NaCl, 5% glycerol, 1 mg/mL lysozyme, 1 mM PMSF, 4 \(\:\mu\:g/ml\) DNase, 0.5X Cell Lytic lysis reagent), incubated at 4°C on a rotating wheel for 1h, before the lysate was clarified by high-speed centrifugation followed by sterile filtration (0.22 \(\:\mu\:m\) ). The lysate was then applied on a 5 ml His-Trap FF column on an ÄKTA pure system (GE Healthcare), the bound proteins were first washed in a Tris based buffer (50 mM Tris, pH 7.5, 500 mM NaCl, 10 mM imidazole) and then eluted with 500 mM imidazole. Polishing purification step was performed by size exclusion chromatography on a Hiload 16/600 Superdex 75 pg column (GE Healthcare) in PBS buffer (pH 7.4). The peak corresponding to the expected molecular weight was collected and concentrated for analysis. Circular dichroism spectroscopy Far-UV circular dichroism spectra were collected between wavelengths of 200 and 250 nm on a Chirascan™ spectrometer (AppliedPhotophysics) spectrometer in a 1 mm path-length quartz cuvette. Proteins were prepared in PBS at a concentration of 0.3 mg/mL by default. Wavelength spectra were measured with a scanning speed of 20 nm/min and a response time of 0.125 s and each measurement was reference subtracted using a cuvette with PBS. The thermal denaturation curves were collected by measuring the change in ellipticity at 220 nm from 20 to 90°C with 2°C increments. Size-exclusion chromatography combined with multi-angle light scattering monodispersity and oligomeric state of the purified proteins were assessed by multi-angle light scattering, thanks to molecular weight determination in solution. 100 \(\:\mu\:g\) of protein were injected into a Superdex 75 10/300 GL column (GE Healthcare) in PBS buffer (pH 7.4) at a flow rate of 0.5 ml/min and detected on in-line multi-angle light-scattering detector (DAWN TREOS, Wyatt). Static light-scattering signals were recorded, and the scatter data were analyzed by ASTRA software (version 8.0.2.5 64-bit, Wyatt). Crystallization and X-ray structure determination The N5 design was crystallized at a concentration of 19 mg/ml using sitting drop vapor diffusion at 18°C in 0.1 M MES pH 6.5, 0.2 M KSCN, 25% w/v PEG 2000 MME buffer (Clear Strategy Screen I, Molecular Dimensions). The N30 design was crystallized at a concentration of 24.3 mg/ml using sitting drop vapor diffusion at 18°C in 0.1 M MES pH6.5; 30% v/v PEG Smear Low buffer (BCS Screen, Molecular Dimensions). The N37 design was crystallized at a concentration of 28 mg/ml using sitting drop vapor diffusion at 18°C in 0.1 M NaOAc pH5.5, 0.2 M KSCN, 25% w/v PEG 2000 MME buffer (Clear Strategy Screen I, Molecular Dimensions). The M7 design was crystallized at a concentration of 24.7 mg/ml using sitting drop vapor diffusion at 18°C in 0.1 M MES pH6.5, 0.2 M KSCN, 15% w/v PEG 4000 buffer (Clear Strategy Screen I, Molecular Dimensions). The M16 design was crystallized at a concentration of 42 mg/ml using sitting drop vapor diffusion at 18°C in 0.1 M Tris pH8.0, 0.04 M sodium format, 0.04 M CaCl2, 25% v/v PEG Smear Low buffer (BCS Screen, Molecular Dimensions). Crystals were cryoprotected in 25% glycerol and flash cooled in liquid nitrogen. Diffraction data was collected at the European Synchrotron Radiation Facility MASSIF-1 and MASSIF-3 beamlines, Grenoble, France at a temperature of 100 K. Crystallographic data was processed using the autoPROC package 43 . Phases were obtained by molecular replacement using the full or partial designed model in Phaser 44 . Atomic model rebuilding and refinement was performed using COOT 42 and Phenix refine 44 . The quality of refined models was assessed using MolProbity 46 . Declarations Data availability All data are available in the main text or as Supplementary Information. Atomic coordinates and structure factors of the reported X-ray structures have been deposited in the PDB under accession numbers 9R2K(N30), 9R2O(N5), 9R2L(N37), 9R2R(M7), 9R2V(M16). Code availability DiffTopo code and jupyter notebook is available at GitHub (https://github.com/YangyangMiao/DiffTopo ). AF2 model used for \predictions can be downloaded from (https://github.com/sokrypton/ColabFold ). ProteinMPNN is available at GitHub (https://github.com/dauparas/ProteinMPNN ). References Taylor, W. R. Exploring Protein Fold Space. Biomolecules 10 , 193 (2020). Simons, K. T., Kooperberg, C., Huang, E. & Baker, D. Assembly of protein tertiary structures from fragments with similar local sequences using simulated annealing and Bayesian scoring functions. J. Mol. Biol. 268 , 209–225 (1997). Minami, S. et al. Exploration of novel αβ-protein folds through de novo design. Nat. Struct. Mol. Biol. 1–9 (2023) doi:10.1038/s41594-023-01029-0. Bonet, J. et al. Rosetta FunFolDes – A general framework for the computational design of functional proteins. PLOS Comput. Biol. 14 , e1006623 (2018). Huang, P.-S. et al. RosettaRemodel: a generalized framework for flexible backbone protein design. PloS One 6 , e24109 (2011). Yim, J. et al. SE(3) diffusion model with application to protein backbone generation. Preprint at https://doi.org/10.48550/arXiv.2302.02277 (2023). Ingraham, J. B. et al. Illuminating protein space with a programmable generative model. Nature 623 , 1070–1078 (2023). SCUBA-D: a freshly trained diffusion model generates high-quality protein structures. Nat. Methods 21 , 1990–1991 (2024). Geffner, T. et al. Proteina: Scaling Flow-based Protein Structure Generative Models. Preprint at https://doi.org/10.48550/arXiv.2503.00710 (2025). Lin, Y. & AlQuraishi, M. Generating Novel, Designable, and Diverse Protein Structures by Equivariantly Diffusing Oriented Residue Clouds. Preprint at https://doi.org/10.48550/arXiv.2301.12485 (2023). Bose, A. J. et al. SE(3)-Stochastic Flow Matching for Protein Backbone Generation. Preprint at https://doi.org/10.48550/arXiv.2310.02391 (2024). Lu, T. et al. Conditional Protein Structure Generation with Protpardelle-1c. bioRxiv 2025.08.18.670959 (2025) doi:10.1101/2025.08.18.670959. Taylor, W. R. et al. Prediction of protein structure from ideal forms. Proteins 70 , 1610–1619 (2008). Harteveld, Z. et al. Exploring “dark-matter” protein folds using deep learning. Cell Syst. 15 , 898-910.e5 (2024). Harteveld, Z. et al. A generic framework for hierarchical de novo protein design. Proc. Natl. Acad. Sci. 119 , e2206111119 (2022). Harteveld, Z. et al. Deep sharpening of topological features for de novo protein design. in (2022). Singh, A. Bottom-up de novo protein design. Nat. Methods 18 , 233–233 (2021). Lin, Z. et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science https://doi.org/10.1126/science.ade2574 (2023) doi:10.1126/science.ade2574. Dauparas, J. et al. Robust deep learning-based protein sequence design using ProteinMPNN. Science 378 , 49–56 (2022). Jumper, J. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596 , 583–589 (2021). Yim, J. et al. Fast protein backbone generation with SE(3) flow matching. Preprint at http://arxiv.org/abs/2310.05297 (2023). Trippe, B. L. et al. Diffusion probabilistic modeling of protein backbones in 3D for the motif-scaffolding problem. Preprint at https://doi.org/10.48550/arXiv.2206.04119 (2023). Knudsen, M. & Wiuf, C. The CATH database. Hum. Genomics 4 , 207–212 (2010). Burley, S. K. et al. RCSB Protein Data Bank: powerful new tools for exploring 3D structures of biological macromolecules for basic and applied research and education in fundamental biology, biomedicine, biotechnology, bioengineering and energy sciences. Nucleic Acids Res. 49 , D437–D451 (2020). van Kempen, M. et al. Fast and accurate protein structure search with Foldseek. Nat. Biotechnol. 42 , 243–246 (2024). Zhang, Y. & Skolnick, J. Scoring function for automated assessment of protein structure template quality. Proteins 57 , 702–710 (2004). Zhang, Y. & Skolnick, J. TM-align: a protein structure alignment algorithm based on the TM-score. Nucleic Acids Res. 33 , 2302–2309 (2005). Xu, J. & Zhang, Y. How significant is a protein structure similarity with TM-score = 0.5? Bioinformatics 26 , 889–895 (2010). Altschul, S. F., Gish, W., Miller, W., Myers, E. W. & Lipman, D. J. Basic local alignment search tool. J. Mol. Biol. 215 , 403–410 (1990). Söding, J., Biegert, A. & Lupas, A. N. The HHpred interactive server for protein homology detection and structure prediction. Nucleic Acids Res. 33 , W244–W248 (2005). Gabler, F. et al. Protein Sequence Analysis Using the MPI Bioinformatics Toolkit. Curr. Protoc. Bioinforma. 72 , e108 (2020). Lu, T., Liu, M., Chen, Y., Kim, J. & Huang, P.-S. Assessing Generative Model Coverage of Protein Structures with SHAPES. 2025.01.09.632260 Preprint at https://doi.org/10.1101/2025.01.09.632260 (2025). Eguchi, R. R. & Huang, P.-S. Multi-scale structural analysis of proteins by deep semantic segmentation. Bioinformatics 36 , 1740–1749 (2020). Xu, Y. & Zhu, T. F. Mirror-image T7 transcription of chirally inverted ribosomal and functional RNAs. Science https://doi.org/10.1126/science.abm0646 (2022) doi:10.1126/science.abm0646. Zhang, G. & Zhu, T. F. Mirror-image trypsin digestion and sequencing of D-proteins. Nat. Chem. 16 , 592–598 (2024). Mackenzie, C. O., Zhou, J. & Grigoryan, G. Tertiary alphabet for the observable protein structural universe. Proc. Natl. Acad. Sci. 113 , E7438–E7447 (2016). Zheng, F., Zhang, J. & Grigoryan, G. Tertiary structural propensities reveal fundamental sequence/structure relationships. Struct. Lond. Engl. 1993 23 , 961–971 (2015). Heinig, M. & Frishman, D. STRIDE: a web server for secondary structure assignment from known atomic coordinates of proteins. Nucleic Acids Res. 32 , W500-502 (2004). Ho, J., Jain, A. & Abbeel, P. Denoising Diffusion Probabilistic Models. Preprint at https://doi.org/10.48550/arXiv.2006.11239 (2020). Satorras, V. G., Hoogeboom, E. & Welling, M. E(n) Equivariant Graph Neural Networks. Preprint at https://doi.org/10.48550/arXiv.2102.09844 (2022). Jing, B., Eismann, S., Suriana, P., Townshend, R. J. L. & Dror, R. Learning from Protein Structure with Geometric Vector Perceptrons. Preprint at https://doi.org/10.48550/arXiv.2009.01411 (2021). Emsley, P., Lohkamp, B., Scott, W. G. & Cowtan, K. Features and development of Coot. Acta Crystallogr. D Biol. Crystallogr. 66 , 486–501 (2010). Additional Declarations There is NO Competing Interest. Supplementary Files SupplementaryLeveragingproteinrepresentationstoexploreunchartedfoldspaceswithgenerativemodels.docx Supplementary information Leveraging protein representations to explore uncharted fold spaces with generative models Cite Share Download PDF Status: Under Review Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7827245","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":534987823,"identity":"1b09b4a1-1657-4fea-be96-0e2d32517f8c","order_by":0,"name":"Bruno Correia","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA2klEQVRIiWNgGAWjYBACA2TOAYYKBgb2dgYGZhK0nGFg4DlMihYGxjYitJhLnzF78IHBTs589uGDB37Os8njYWZg/FyAR4tlX4654QyGZGOZc2kJB3u3pRUDtTBLz8DnsDM8ZtI8DAcSZ/DwGBzg3XY4cT8zAxszD3Fa+D8c/Dvnf2IPCVp4GA7zNhwgrMWyh61McoZBsrEED5vBYZljyUC/MDZL49NizsO8TeJDhZ2cBA/z449vauzyeNibD37GpwXqPAQzARg7DQQ1oIAE0pSPglEwCkbBSAAA/kw/IS8qrggAAAAASUVORK5CYII=","orcid":"https://orcid.org/0000-0002-7377-8636","institution":"École Polytechnique Fédérale de Lausanne","correspondingAuthor":true,"prefix":"","firstName":"Bruno","middleName":"","lastName":"Correia","suffix":""},{"id":534987824,"identity":"07352471-2da3-48c7-98b9-6199b27de490","order_by":1,"name":"Yangyang Miao","email":"","orcid":"https://orcid.org/0009-0001-1313-189X","institution":"École Polytechnique Fédérale de Lausanne","correspondingAuthor":false,"prefix":"","firstName":"Yangyang","middleName":"","lastName":"Miao","suffix":""},{"id":534987825,"identity":"0944f1ab-0f7b-4e70-acc0-63e5f51b55be","order_by":2,"name":"Martin Pacesa","email":"","orcid":"","institution":"École Polytechnique Fédérale de Lausanne","correspondingAuthor":false,"prefix":"","firstName":"Martin","middleName":"","lastName":"Pacesa","suffix":""},{"id":534987828,"identity":"dbd34d7c-05ce-4dcc-9f56-3d193e11951b","order_by":3,"name":"Sandrine Georgeon","email":"","orcid":"https://orcid.org/0000-0002-8158-4508","institution":"École Polytechnique Fédérale de Lausanne","correspondingAuthor":false,"prefix":"","firstName":"Sandrine","middleName":"","lastName":"Georgeon","suffix":""},{"id":534987830,"identity":"18d6699e-038b-4394-8847-4ce04d022dbc","order_by":4,"name":"Joseph Schmidt","email":"","orcid":"","institution":"École Polytechnique Fédérale de Lausanne","correspondingAuthor":false,"prefix":"","firstName":"Joseph","middleName":"","lastName":"Schmidt","suffix":""},{"id":534987832,"identity":"11321340-892d-4ff8-a9aa-60aa5decc253","order_by":5,"name":"Tianyu Lu","email":"","orcid":"","institution":"Stanford University","correspondingAuthor":false,"prefix":"","firstName":"Tianyu","middleName":"","lastName":"Lu","suffix":""},{"id":534987836,"identity":"96c9e011-7142-44d8-a069-d8d72d14ca1f","order_by":6,"name":"Po-Ssu Huang","email":"","orcid":"https://orcid.org/0000-0002-7948-2895","institution":"Stanford University","correspondingAuthor":false,"prefix":"","firstName":"Po-Ssu","middleName":"","lastName":"Huang","suffix":""}],"badges":[],"createdAt":"2025-10-10 12:55:16","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7827245/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7827245/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":95079901,"identity":"f5d61531-c45a-4593-8a6a-0b3484706416","added_by":"auto","created_at":"2025-11-04 05:59:31","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":355179,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eOverview of the protein design landscape and DiffTopo-RFdiffusion pipeline.\u003c/strong\u003e \u003cstrong\u003ea,\u003c/strong\u003e Conceptual representation of the Protein Universe Sampling across structural landscapes with different resolutions. The top row illustrates our proposed topology-level sampling approach, which operates at a coarse-grained (CG) level. The rugged backbone-level landscape, where most existing methods, such as RFdiffusion and FrameDiff operate. The bottom row depicts the natural protein universe derived from ESMFold predictions atlas\u003csup\u003e18\u003c/sup\u003e. \u003cstrong\u003eb,\u003c/strong\u003eSchematic representation of the DiffTopo-RFdiffusion pipeline. The workflow begins with the conversion of secondary structure elements into CG topology representations. DiffTopo is a diffusion model that generates CG topologies that serve as structural blueprints to guide RFdiffusion on the sampling of atomic-level protein backbones.\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-7827245/v1/635ac6d4c18f925cd948c453.png"},{"id":95224380,"identity":"8e9d7477-e033-4393-bd8e-34d232eabc0c","added_by":"auto","created_at":"2025-11-05 16:23:40","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":704953,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eStructural evaluation of DiffTopo-guided sampling\u003c/strong\u003e\u003cem\u003e\u003cstrong\u003e.\u003c/strong\u003e\u003c/em\u003e\u003cem\u003e \u003c/em\u003e\u003cem\u003e\u003cstrong\u003ea,\u003c/strong\u003e\u003c/em\u003e\u003cem\u003e S\u003c/em\u003etructural comparison between native structures and their corresponding protein sketches from CG topologies (green), and between protein sketches and generated backbones (yellow). Each backbone was generated by RFdiffusion conditioned on the protein sketch\u003cem\u003e. \u003c/em\u003e\u003cstrong\u003eb,\u003c/strong\u003e Designability of generated backbones. scRMSD based on 1,000 samples per length range (40–100, 100–150, 150–200) across different methods. Purple: RFdiffusion random sampling (baseline), yellow: DiffTopo-RFDiffusion random sampling of SSE strings, green: DiffTopo-RFDiffusion sampling of low-frequency SSE strings in native proteins. \u003cstrong\u003ec,\u003c/strong\u003eNovelty of generated backbones: color scheme follows b. \u003cstrong\u003ed,\u003c/strong\u003e Diversity of generated backbones: TM-score heatmap after hierarchical clustering, showing structures from the top 5 largest clusters. \u003cstrong\u003ee,\u003c/strong\u003eTM-score between DiffTopo-generated backbones and AF2 predictions plotted against predicted pLDDT to assess the designability of the generated structures. Red points indicate highly designable sequences (scTM-score \u0026gt;0.8, pLDDT \u0026gt;80). \u003cstrong\u003ef,\u003c/strong\u003e Representative visualizations of DiffTopo-generated backbone sketches. \u003cstrong\u003eg,\u003c/strong\u003e Backbone alignment between DiffTopo-generated structures and AF2 predictions to evaluate the structural consistency. \u003cstrong\u003eh,\u003c/strong\u003e Comparison of DiffTopo-generated backbones with their nearest structural analogs identified by FoldSeek in PDB and CATH databases. Designs in row 1-3 correspond to N35, N40 and N37 in the Supplementary table 2.\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-7827245/v1/bc787b800c8e43abf2e8aa16.png"},{"id":95079902,"identity":"81aff83c-82b7-4418-bb8c-1239533c30d7","added_by":"auto","created_at":"2025-11-04 05:59:32","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":729094,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eExperimental validation of dark folds.\u003c/strong\u003e \u003cstrong\u003ea, \u003c/strong\u003estructural models of representative designs from each fold category and their corresponding circular dichroism (CD) spectra at different temperatures, confirming that the designs adopt folded structures in solution. \u003cstrong\u003eb,\u003c/strong\u003e X-ray crystal structure of N30 (colored) superimposed on the AF2 predicted model (gray). \u003cstrong\u003ec,\u003c/strong\u003e X-ray crystal structure of N37 (colored) superimposed on the design model (gray). \u003cstrong\u003ed, \u003c/strong\u003eX-ray crystal structure of N5 chain A (red) and chain B(blue) overlaid with the AF2 predicted model (gray), shown as a dimer (left) and a monomer (right). Crystal structures are available in PDB: N30(\u003ca href=\"https://www.rcsb.org/structure/9R2K\"\u003e9R2K\u003c/a\u003e), N37(\u003ca href=\"https://www.rcsb.org/structure/9R2L\"\u003e9R2L\u003c/a\u003e), N5(\u003ca href=\"https://www.rcsb.org/structure/9R2O\"\u003e9R2O\u003c/a\u003e).\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-7827245/v1/22534c57f913e63355ddfe0d.png"},{"id":95223391,"identity":"180a7f28-ccd3-47eb-a338-4dd5830d2b38","added_by":"auto","created_at":"2025-11-05 16:22:11","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":666053,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eStructural exploration in the Mirror fold space.\u003c/strong\u003e \u003cstrong\u003ea,\u003c/strong\u003e Schematic overview of the MirrorTopo pipeline. Native CG topologies are mirrored in the topological coordinate space to generate symmetry-inverted representations. These are then processed through the same downstream design pipeline, including RFdiffusion for backbone generation. The bottom row depicts the natural protein universe derived and mirror fold universe. \u003cstrong\u003eb\u003c/strong\u003e, From left to right: native protein structure from the CATH database; 2D secondary-structure connectivity of the native fold; mirrored topology reflecting a symmetry-inverted design; backbone model generated from the mirrored topology using RFdiffusion; and the top FoldSeek match from the PDB database, indicating similarity or novelty relative to known folds.\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-7827245/v1/20aec44f2e34bd6fc7675f78.png"},{"id":95079903,"identity":"0b6be320-1341-4b0e-801f-d662c1987d66","added_by":"auto","created_at":"2025-11-04 05:59:32","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":588477,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eExperimental validation of the computational designed Mirror folds.\u003c/strong\u003e \u003cstrong\u003ea,\u003c/strong\u003e Structural models of representative designs from each fold category and their corresponding circular dichroism (CD) spectra at different temperatures, confirming that the designs adopt folded structures in solution. \u003cstrong\u003eb,\u003c/strong\u003e X-ray crystal structure of M16 (colored) superimposed on the design model (gray). \u003cstrong\u003ec,\u003c/strong\u003e X-ray crystal structure of M7 (colored) superimposed on the design model (gray). Crystal structures are available in PDB :M16(\u003ca href=\"https://www.rcsb.org/structure/9R2V\"\u003e9R2V\u003c/a\u003e),M7(\u003ca href=\"https://www.rcsb.org/structure/9R2R\"\u003e9R2R\u003c/a\u003e).\u003c/p\u003e","description":"","filename":"5.png","url":"https://assets-eu.researchsquare.com/files/rs-7827245/v1/a607a344e24676dfe76023ea.png"},{"id":95079906,"identity":"90c53cac-ad25-4878-8432-75de4b3a0906","added_by":"auto","created_at":"2025-11-04 05:59:32","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":611595,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eTERM analysis of dark folds and mirror folds.\u003c/strong\u003e \u003cstrong\u003ea,\u003c/strong\u003e TERM abundance per residue(left) and representative PDB structures colored by TERM abundance, red-thin regions indicating low abundance; blue-thick regions reflecting higher abundance (right). \u003cstrong\u003eb\u003c/strong\u003e, Average abundance distribution (left) and structural scores(right) for CATH folds, dark folds, and mirror folds.\u003c/p\u003e","description":"","filename":"6.png","url":"https://assets-eu.researchsquare.com/files/rs-7827245/v1/36fee58486d87adc484e6b82.png"},{"id":95230487,"identity":"f1d07ad7-1ce7-4e11-961f-41335849f1c7","added_by":"auto","created_at":"2025-11-05 16:37:36","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":4124905,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7827245/v1/2f946386-2f53-445c-aebb-b3323aecce70.pdf"},{"id":95079907,"identity":"722a327a-1320-4f9c-bd14-6aae6f28744a","added_by":"auto","created_at":"2025-11-04 05:59:32","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":7357679,"visible":true,"origin":"","legend":"Supplementary information Leveraging protein representations to explore uncharted fold spaces with generative models","description":"","filename":"SupplementaryLeveragingproteinrepresentationstoexploreunchartedfoldspaceswithgenerativemodels.docx","url":"https://assets-eu.researchsquare.com/files/rs-7827245/v1/756e6a9139f2a2fe8297a32c.docx"}],"financialInterests":"There is \u003cb\u003eNO\u003c/b\u003e Competing Interest.","formattedTitle":"Leveraging protein representations to explore uncharted fold spaces with generative models","fulltext":[{"header":"Introduction","content":"\u003cp\u003eProteins govern biological functions in nature through their three-dimensional structure, yet the vast space of possible folds remains largely unexplored\u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u003c/sup\u003e. While natural evolution has populated only a fraction of this space, computational methods now offer a route to systematically chart unexplored structural spaces\u003csup\u003e\u003cspan additionalcitationids=\"CR3 CR4\" citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e. One of the important challenges in protein design is to efficiently explore the designable protein fold space without being confined to the training data and in this way favor exploitation of that data distribution, rather than exploration. Here, we present a framework that explores the fold space using coarse-grained representations to traverse this landscape, enabling discovery of novel folds. Recent advances in deep generative models, particularly diffusion-based methods, have dramatically expanded the capabilities of de novo protein design. Recently, many deep generative models have been developed that can generate atomically detailed protein structures with complex topologies, such as RFdiffusion\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e, FrameDiff\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e, and Chroma\u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e, among many other deep generative models\u003csup\u003e\u003cspan additionalcitationids=\"CR9 CR10 CR11\" citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e. However, despite their success, these methods still tend to sample structures that resemble natural folds, especially for small proteins. This reflects a fundamental limitation of coordinate-level generative modeling: the tendency to recapitulate training distributions and bias the design space toward naturally occurring solutions. As a result, large regions of the theoretically designable but evolutionarily unoccupied fold space, so-called \"dark\" structural matter, remain underexplored.\u003c/p\u003e\u003cp\u003eWe hypothesized that an efficient exploration of the fold space requires a different approach: Rather than sampling structures at the atomic level, the exploration should first occur in a low-resolution topological space where structural diversity and fold novelty is more accessible (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ea). Drawing inspiration from representations of proteins as hierarchical assemblies of secondary structure elements\u003csup\u003e\u003cspan additionalcitationids=\"CR14 CR15 CR16\" citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e, we introduce a coarse-grained (CG) topology description that describes folds into spatial arrangements of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\alpha\\:\\)\u003c/span\u003e\u003c/span\u003e-helices and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\beta\\:\\)\u003c/span\u003e\u003c/span\u003e-strands (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eb). This reduced-dimensionality manifold smooths the sampling landscape, allowing for more coarse sampling of the structural space and avoiding the combinatorial explosion of atomic representations. We developed DiffTopo, which learns fold topologies through diffusion modeling guided by secondary structure descriptors (a string of secondary structure types), allowing us to sample regions far from natural folds while keeping physical plausibility.\u003c/p\u003e\u003cp\u003eCoarse-grained topological exploration of protein structure provides an approximate placement of secondary structure elements in 3D space; subsequently, obtaining a full atomic backbone requires a high-resolution generative model.\u003c/p\u003e\u003cp\u003eBased on the CG topology length, we first assign standard secondary structure elements and place them at the corresponding spatial positions, resulting in an intermediate structure composed solely of idealized secondary elements, which we refer to as a \u003cem\u003eprotein sketch\u003c/em\u003e. Our framework bridges this gap by refining these sketches and using them to guide RFdiffusion⁵ generation, ultimately producing high-resolution, designable protein backbones. Importantly, this separation between fold discovery (DiffTopo), through CG sampling and structure generation (RFdiffusion), via atomic-level modeling, helps overcome some of the limitations of directly generating full atomic coordinates, which often tends to recapture existing topologies from known fold space. We further tested this by attempting a distinct strategy of structural generation, where we systematically generated mirrored image topologies of natural folds, which are absent from natural repertoires but are physically realistic and potentially designable, we dubbed this approach MirrorTopo. These two challenges of generating novel-to-nature dark folds and mirrored topologies demonstrate how coarse-grained sampling can reveal underexplored regions of structural space.\u003c/p\u003e\u003cp\u003eBiochemical and structural validation confirmed the potential of this approach, where 40 designed dark folds and 22 mirror folds were experimentally characterized, with five crystal structures closely matching the computational models. These results suggest that our method can access folds beyond the capabilities of current generative approaches. By decoupling fold exploration from atomic-level sampling, we establish a new strategy to explore protein space that prioritizes discovery over optimization, a critical step toward understanding the full scope of possible protein architectures and their potential functions.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\u003ch2\u003eEnhancing the sampling of structural diversity with DiffTopo\u003c/h2\u003e\u003cp\u003eTo test the capabilities of the DiffTopo-RFdiffusion pipeline, we generated 3000 backbone structures for each of the following categories: I) random secondary structure descriptors guided backbones; II) low-frequency secondary structure descriptors guided backbones; III) random lengths ranging from 40 to 200 residues backbone generated by RFdiffusion, as reference. For categories I and II, we describe the secondary structure descriptor guided generation in detail in the methods section. To test whether the CG topology correctly captures the target protein fold and whether the RFdiffusion-generated backbones recover the intended structure, we compared protein sketches derived from CG topology with native structures to ensure the model learns the correct structural features, and with the corresponding RFdiffusion-generated backbones to verify that the generated outputs preserve the intended topology. Most protein sketches match well with the generated backbones (RMSD\u0026thinsp;\u0026lt;\u0026thinsp;3 \u0026Aring;), confirming that the CG sketches can guide the sampling of RF-diffusion for the generation of all-atom protein backbones (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ea). To assess whether the model effectively learned the geometric distribution of sketch data, we analyzed its geometric descriptors and evaluated their alignment with native structures. DiffTopo-generated sketches match the topological patterns observed in native structures, showing that the model effectively learns an accurate approximation of real data distributions (Supplementary Fig.\u0026nbsp;1a-c).\u003c/p\u003e\u003cp\u003eWe then assessed the designability of the generated backbones from different types of protocol by designing amino acid sequences using ProteinMPNN\u003csup\u003e\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e and predicting whether the designed sequences could fold into the target topology with AlphaFold2(AF2)\u003csup\u003e\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e. Following prior studies on protein backbone generation\u003csup\u003e\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e,\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e, we quantified self-consistency using C\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\alpha\\:\\)\u003c/span\u003e\u003c/span\u003e RMSD (scRMSD, lower values indicate better consistency) and predicted local distance difference test (pLDDT, higher values indicate greater structural confidence). While a trade-off in scRMSD was observed compared to RFdiffusion alone, a substantial proportion of designs still achieved scRMSDs below 3 \u0026Aring; (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eb).\u003c/p\u003e\u003cp\u003eHowever, our primary objective was to assess backbone novelty. To this end, we selected backbones with high designability (scRMSD\u0026thinsp;\u0026lt;\u0026thinsp;3 \u0026Aring; and pLDDT\u0026thinsp;\u0026gt;\u0026thinsp;90) and performed structural similarity searches against the CATH database\u003csup\u003e\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e,\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/sup\u003e with FoldSeek\u003csup\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u003c/sup\u003e, reporting the highest template modeling score (Max PDB TM-score\u003csup\u003e\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e,\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e\u003c/sup\u003e). By utilizing low-frequency secondary structure descriptor guidance, our approach produced significantly more structurally novel and highly designable backbones compared to baseline RFdiffusion (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ec). This suggests that DiffTopo-RFdiffusion can generate both native-like folds and previously unobserved (dark) folds, demonstrating generalization beyond the structural space of the Protein Data Bank (PDB).\u003c/p\u003e\u003cp\u003eAs an example, for the secondary structure descriptor seed \"EEEHHEHE\", we generated 500 designable backbones which resulted in 109 structurally distinct clusters (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ed), using a TM-score threshold of 0.5\u003csup\u003e28\u003c/sup\u003e.These findings highlight the ability of our framework to efficiently generate a diverse range of \u003cem\u003ede novo\u003c/em\u003e protein structures with low structural redundancy.\u003c/p\u003e\u003cp\u003eIn summary, the DiffTopo-RFdiffusion framework enhances de novo protein design by generating designable and structurally novel backbones. Compared to standalone RFdiffusion, it significantly improves backbone novelty by leveraging sampling in low-frequency secondary structure order regions, enabling the discovery of both native-like and entirely dark folds beyond the known protein space.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003c/div\u003e\n\u003ch3\u003eExploring de novo “dark” folds with DiffTopo\u003c/h3\u003e\n\u003cp\u003eA primary goal of our work is to expand the accessible protein fold space by generating structures that are not only new to nature but also extend beyond the range of folds observed in natural proteins. To systematically assess the model\u0026rsquo;s capacity for fold exploration, we sampled 779 different dark folds with distinct secondary structure compositions (all-α, all-β, and α/β), and quantified their statistics on the designability, novelty, and diversity (Supplementary Table\u0026nbsp;1). At the sequence level none of the designed sequences displayed a hit below the significant E-value threshold of 0.001 (Supplementary Table\u0026nbsp;2) relative to the non-redundant database (NR)\u003csup\u003e\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e\u003c/sup\u003e, indicating that the designed sequences are novel and distinct from naturally occurring sequences. For additional validation, we applied the HHpred\u003csup\u003e\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e,\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e\u003c/sup\u003e to search for potential distant homologs. Intriguingly, all sequences showed E-values above 0.5, further supporting their novelty compared to known natural sequences (Supplementary Table\u0026nbsp;3).\u003c/p\u003e\u003cp\u003eTo further evaluate the structural quality and novelty of the generated designs, we analyzed sequence-structure consistency, architectural diversity, and similarity to known folds (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ee\u0026ndash;h). As shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ee, DiffTopo produced backbones with high designability across a range of secondary structure compositions, indicating strong sequence-structure compatibility. To illustrate this, we selected representative examples of all-α, all-β, and α/β folds and showed their corresponding structures in the figure; corresponding secondary structure descriptors are provided in Supplementary Table\u0026nbsp;4. The resulting protein sketches exhibited substantial architectural diversity (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ef), showing that the model explores a wide topological space. Structural alignment with AF2 predictions confirmed the internal consistency of these designs (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eg), supporting their plausibility. The designs were compared to naturally occurring structural databases using Foldseek, in which the closest hits had very low structural similarity, underscoring the potential of the approach for discovering novel, previously uncharacterized folds (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eh).\u003c/p\u003e\u003cp\u003eTo evaluate the novelty of the generated folds, we compared all dark folds to CATH domain representatives using TM-align and projected the TM-score matrix into 2D using classical multidimensional scaling (MDS) (Supplementary Fig.\u0026nbsp;2). The resulting distribution showed that dark folds form distinct structural clusters, separate from natural domain families. Even for well-populated classes like 7-helix bundles, with 394 CATH entries, DiffTopo still generated 16 unique topologies (Max TMscore\u0026thinsp;\u0026lt;\u0026thinsp;0.5) not accounted for in the database (Supplementary Fig.\u0026nbsp;3, Supplementary Table\u0026nbsp;5). To further explore their global distribution in fold space, we applied the latest structural analysis method SHAPES\u003csup\u003e\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e\u003c/sup\u003e (structural and hierarchical assessment of proteins with embedding similarity), projecting structures into two orthogonal spaces, ProtDomainSegmentor embedding\u003csup\u003e\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e\u003c/sup\u003e and ProteinMPNN layer 1 embedding (see Methods: Mapping Protein structures to CATH database). As shown in Supplementary Fig.\u0026nbsp;4, each embedding space was rasterized into a 16\u0026times;16 grid, and representative structures were selected per cell to visualize sampling coverage. While CATH domains densely populated specific regions, the dark folds broadly covered previously unoccupied areas, supporting that our method systematically samples novel and underexplored fold topologies.\u003c/p\u003e\u003cp\u003eIn summary, our pipeline demonstrates its ability to uncover novel, designable and complex folds that do not exist in nature, encompassing various secondary structure combinations, including all \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\alpha\\:\\)\u003c/span\u003e\u003c/span\u003e, all \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\beta\\:\\)\u003c/span\u003e\u003c/span\u003e, and mixed \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\alpha\\:/\\beta\\:\\:\\)\u003c/span\u003e\u003c/span\u003efolds.\u003c/p\u003e\n\u003ch3\u003eDiffTopo dark folds designs exhibit high structural precision\u003c/h3\u003e\n\u003cp\u003eWe computationally designed 778 dark folds that passed the designability filters. For experimental validation, we selected 12 all-α, 7 all-β, and 21 α/β dark fold designs. These proteins were expressed in \u003cem\u003eEscherichia coli\u003c/em\u003e and purified using a Ni\u0026sup2;⁺-NTA affinity column. The purified proteins were characterized by size-exclusion chromatography combined with multi-angle light scattering (SEC\u0026ndash;MALS) and circular dichroism (CD) spectroscopy. Of 44 tested dark fold topologies, 22 designed proteins were well expressed and highly soluble. Additionally, all designs exhibited CD spectra typical of folded proteins and consistent with the designed secondary structure content (Supplementary Fig.\u0026nbsp;6\u0026ndash;10). The designs were thermally very stable as is typical in de novo designed proteins, and were confirmed to be monomeric in solution by SEC\u0026ndash;MALS (Supplementary Fig.\u0026nbsp;5a-w).\u003c/p\u003e\u003cp\u003eFrom this set, we selected 6 representative dark folds, including all-α (N34), all-β (N40), and α/β (N5, N10, N30, N37), which are highlighted in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003ea. Crystal structures were solved for 3 of these designs: N30, N37, and N5. For N30 and N37, the crystal structures aligned especially well with the computational models, yielding backbone RMSDs of 0.58 \u0026Aring; and 0.70 \u0026Aring;, respectively (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eb,c). FoldSeek analysis confirmed minimal similarity to known protein structures (Foldseek CATH max TM-score\u0026thinsp;=\u0026thinsp;0.488 in), validating their structural novelty. For example, while N37 partially overlapped with ribonuclease H-like β-sheets (Foldseek CATH max TM-score\u0026thinsp;=\u0026thinsp;0.55 for shared region), its overall topology, α-helices wrapping a β-sheet, is absent from the natural repertoire. Notably, N5, an α/β dark fold (max TM-score\u0026thinsp;=\u0026thinsp;0.43), crystallized as a dimer with a β-strand-swapped interface, however, SEC\u0026ndash;MALS confirmed its monomeric state in solution, and alignment of the non-swapped regions (35\u0026ndash;161) revealed strong agreement with the design model (full atom-RMSD\u0026thinsp;=\u0026thinsp;0.66 \u0026Aring;), suggesting that the strand swapping event is likely a crystallization induced artefact.\u003c/p\u003e\u003cp\u003eThese results demonstrate that dark fold designs generated by DiffTopo are not only highly designable and diverse, but also structurally accurate, validating the potential of our topology-first framework to generate entirely novel folds absent from nature.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\n\u003ch3\u003eMirrorTopo generates new to nature protein folds\u003c/h3\u003e\n\u003cp\u003eAlthough DiffTopo enables the generation of novel folds, the sampled structures still follow natural-like spatial organization. To further expand the design space, we developed MirrorTopo, which generates symmetry-inverted (mirrored) versions of natural folds at the topological level, producing structures that are physically valid but have not been observed in current structural databases.\u003c/p\u003e\u003cp\u003eMirrorTopo operates by mirroring CG topology sketches in 3D coordinate space prior to atomic-level generation using RFdiffusion. Mirror folds exhibit distinct pairwise patterns of secondary structure stacking, specifically in the dihedral angles and distances between helix\u0026ndash;helix, helix\u0026ndash;strand, and strand\u0026ndash;strand pairs compared to natural proteins (Supplementary Fig.\u0026nbsp;1d,). These patterns fall outside the typical distribution learned by DiffTopo, suggesting that such topologies are unlikely to be sampled. Unlike previous research\u003csup\u003e\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e,\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e\u003c/sup\u003e that mirrored atomic-level coordinates using D-amino acids and right-handed helices, MirrorTopo maintains L-amino acids and right-handed helices, but systematically inverts how secondary structure elements are arranged in space (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003ea). This approach enables the sampling of mirrored topologies that DiffTopo cannot access allowing us to explore fold topologies that are physically valid yet absent from the known natural repertoire.\u003c/p\u003e\u003cp\u003eFirst, we selected 8 native fold types from the CATH database and generated their mirrored topologies using our MirrorTopo pipeline. For each fold, multiple designs passed our designability filters (self-consistency RMSD\u0026thinsp;\u0026lt;\u0026thinsp;3 \u0026Aring; and AF2 pLDDT\u0026thinsp;\u0026gt;\u0026thinsp;90). We then selected 21 of these designs for experimental testing. Among them, 13 designs were successfully expressed in \u003cem\u003eE.-coli\u003c/em\u003e and showed expected physicochemical properties, indicating folded proteins. These 13 designs cover 6 different mirror folds from native proteins, spanning a variety of α/β architectures, representing diverse structural classes. All detailed description and structural annotations and CATH classifications are summarized in Supplementary Table\u0026nbsp;6. The first column of Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eb shows native protein structures from the CATH database, followed by their 2D secondary-structure schematics in the second column. Their mirrored counterparts, obtained by flipping the arrangement of secondary-structure elements, are presented in the third column, with the corresponding backbone models generated by MirrorTopo shown in the fourth column. The rightmost column displays FoldSeek search results, confirming that the mirror folds are structurally novel with no close matches in existing structural databases.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\n\u003ch3\u003eMirrored topologies show high structural precision\u003c/h3\u003e\n\u003cp\u003eFor the 6 different mirror folds in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eb, we selected one representative design from each, all of which showed CD spectra consistent with the expected secondary structure content (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003ea) and SEC\u0026ndash;MALS profiles indicative of monomeric states in solution (Supplementary Fig.\u0026nbsp;11a-l). CD spectra and melting temperature profiles for the remaining successful designs are provided in Supplementary Figs.\u0026nbsp;12\u0026ndash;14. We successfully solved the crystal structures of two mirror fold designs, M16 and M7. M16, the mirror version of the histidine kinase-like ATPase fold, exhibited remarkable agreement with its AF2 prediction (Cα RMSD\u0026thinsp;=\u0026thinsp;0.42 \u0026Aring;; all-atom RMSD\u0026thinsp;=\u0026thinsp;1.50 \u0026Aring;), confirming the high structural accuracy of the design (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eb). Similarly, M7, a mirror of the TRAPP domain (β-lactamase-like topology), also closely matched its computational model (Cα RMSD\u0026thinsp;=\u0026thinsp;0.67 \u0026Aring;; all-atom RMSD\u0026thinsp;=\u0026thinsp;1.50 \u0026Aring;) (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003ec). These results validate that mirror topologies, that although absent from nature, can be realized with high fidelity using computational methods and experimentally validated.\u003c/p\u003e\u003cp\u003eTo evaluate whether the tertiary structure motifs in our mirror folds are also found in natural proteins, we used TERMANAL\u003csup\u003e\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e,\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e\u003c/sup\u003e. According to previous work by Mackenzie et al., about 600 tertiary structural motifs (TERMs) are sufficient to describe 50% of structures in the PDB. TERMANAL analysis returns two main results: Abundance and Structure Score. Abundance reflects how frequently a particular TERM (tertiary structural module) appears across the protein structure database, with higher values indicating that the TERM is common and likely favorable in natural proteins, whereas lower values indicate rarity and potential structural novelty. Structure Score provides a global assessment of a protein structure, integrating the abundance of all its constituent TERMs to measure overall structural reasonableness. Here, we performed TERM analysis on 50 random CATH folds, all experimentally generated dark folds, and all mirror folds. As shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003ea, CATH folds exhibit high motif abundance, with an average of ~\u0026thinsp;0.8 per TERM. Dark folds have moderately lower abundance (~\u0026thinsp;0.68), while mirror folds display the lowest values (~\u0026thinsp;0.44). A similar trend was observed in Structure Scores. To illustrate these results at the residue level, we selected two representative dark folds (N30 and N37) and two mirror folds (M16 and M7), as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eb,c.The results showed that mirror folds (M16 and M7) have low motif abundance and low structural similarity scores across all regions, suggesting that their local structures are distinct from those commonly found in nature. In contrast, the dark folds N30 and N37 had high motif abundance and structural scores, especially N37, which was similar to native-like topologies. N30 also showed strong matches, except for its β-sheet region, which had fewer matching motifs. These results support our previous observation that DiffTopo samples structures that remain within a \u0026ldquo;natural-like\u0026rdquo; distribution, while mirror folds generated by MirrorTopo lie in a different and less explored region of the structural space. In other words, mirror folds access areas of the fold space that are not well covered by natural evolution.\u003c/p\u003e\u003cp\u003eBoth computational results and experimental validation demonstrate that the MirrorTopo pipeline is viable for generating structurally precise, stable, and novel proteins. The success of mirror-topology designs achieved using right-handed helices with canonical L-amino acids challenges conventional assumptions about foldability and suggests the existence of a vast, underexplored mirror-fold space. These findings highlight the potential of mirror topologies as a complementary approach to expand the universe of designable protein structures.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003e\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eExploring and expanding the protein fold space remains an important challenge in protein design, particularly given that new structures could, in principle, enable novel functional roles for proteins. Existing generative models tend to stay close to natural-like topologies due to fine-grained representations and biases in training data. In contrast, our method uses a compact coarse-grained topology representation, that captures only the relative positions and connections of secondary structure elements. With this abstraction, DiffTopo efficiently learns spatial patterns and samples diverse and novel fold topologies beyond those seen in nature. Embedding analysis (Supplementary Fig.\u0026nbsp;4) shows that many designs occupy underexplored regions of fold space. Despite their novelty, these structures still retain high designability, allowing for sequences that are in good agreement with the structural templates resulting in structural predictions with high AF2 confidence. These results suggest a large, untapped fold space that can now be explored systematically.\u003c/p\u003e\u003cp\u003eWhile DiffTopo expands the diversity of folds, it still follows the spatial patterns found in natural proteins. To move further beyond these constraints, we introduced MirrorTopo, which generates mirrored topologies that are structurally valid but evolutionarily absent. These mirror folds use standard L-amino acids and form right-handed helices, yet their global secondary structure arrangement is reversed. They do not resemble natural folds, as confirmed by FoldSeek analysis, and occupy regions rarely reached by generative models. Embedding analysis shows these mirror designs fall far outside the distribution of natural structures (Supplementary Fig.\u0026nbsp;1d), suggesting that mirror fold space represents a distinct, orthogonal region of protein architecture.\u003c/p\u003e\u003cp\u003eThe absence of mirror folds in nature could reflect evolutionary selection rather than physical constraints. Many natural folds have emerged through duplication, fusion, and recombination events that favor certain packing orientations. Mirrored architectures, while physically allowed, may be disfavored due to folding kinetics, sequence-to-structure mapping biases, or evolutionary path dependence. By systematically constructing these mirror topologies, MirrorTopo reveals a vast, underexplored region of the fold space that has remained invisible to both evolution and previous generative methods. This space is not only accessible, but it is designable, stable, and experimentally verifiable.\u003c/p\u003e\u003cp\u003eTo test whether these topologies are physically viable, we experimentally solved five crystal structures across both DiffTopo and MirrorTopo designs. These include dark folds (N30, N37, N5) and mirror folds (M16, M7). The results confirmed that these unconventional topologies can fold precisely into their intended atomic structures, supporting the robustness of our design approach.\u003c/p\u003e\u003cp\u003eAmong all dark folds, N30 displays a particularly entangled topology, while most other designs rely on linear stacking of secondary structures. In its crystal structure, one helix (residues 120\u0026ndash;143) threads through a closed cavity formed by two others (61\u0026ndash;76 and 81\u0026ndash;95), creating a topological constraint unlikely to arise via simple stacking. Sequential AF2 prediction analysis suggests a nonlinear folding pathway: the first two helices fold and pack tightly, and only later open to accommodate the third helix (Supplementary Fig.\u0026nbsp;15 and Supplementary information: Sequential AF2 predictions analysis for N30). This behavior expands the known repertoire of foldable structures and suggests that topological complexity\u0026mdash;such as threading or knot-like arrangements\u0026mdash;can be encoded and accessed through designed sequences. N30 highlights the potential of our framework to generate not only novel folds, but also nontrivial folding mechanisms.\u003c/p\u003e\u003cp\u003eWhile this work primarily focuses on structural diversity, our method also supports functional design through motif scaffolding (see Methods: Motif-Scaffolding via Conditional Inpainting with Diffusion Models). By conditioning on existing secondary structure motifs, DiffTopo can also generate diverse scaffolds that preserve motif geometry), highlighting the potential of our framework to support functional scaffold generation, expanding the possibilities for designing proteins with tailored interactions. Overall, our study shows the potential of utilizing different representations to enhance the exploration of ML-based methods for generative protein design showing that under sampled regions of this landscape can be accessed through this approach. We anticipate that these new methods will in future enable the design of proteins with novel biological functions.\u003c/p\u003e"},{"header":"Methods","content":"\u003cp\u003eCoarse-grained topology representation\u003c/p\u003e\u003cp\u003eInspired by the FORM representation\u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e and protein sketches\u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e,\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e, here protein structures are built by the assembly of their secondary structures, we adopt a simplified representation referred to as a \"coarse-grained topology (CG topology)\". The topology reduces the intricate protein structure into a stacked arrangement of Secondary Structure Elements (SSEs), achieved by representing secondary structures with three carbon alpha centroids, as illustrated in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ea. All the secondary structures in each residue number are assigned by STRIDE\u003csup\u003e\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e\u003c/sup\u003e program. CG topology involves capturing the geometric position of helices and strands through centroids. For helices, centroids include the first and last four C\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\alpha\\:\\)\u003c/span\u003e\u003c/span\u003e atoms, as well as the total C\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\alpha\\:\\)\u003c/span\u003e\u003c/span\u003e atom centroid. Strands are represented by centroids of the first and last two C\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\alpha\\:\\)\u003c/span\u003e\u003c/span\u003e atoms, along with the centroid C\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\alpha\\:\\)\u003c/span\u003e\u003c/span\u003e atom within the strand. CG topology can also be easily represented as a protein sketch, which is a rough 3D approximation of a native protein structure with standard SSEs, lacking loops, and AA side chains. Compared with forms and protein sketches, this representation method of CG topology removes the concept of layers, offering a higher degree of freedom in structural representation to represent structures such as beta barrels, while remaining simple with few degrees of freedom.\u003c/p\u003e\u003cp\u003eDataset\u003c/p\u003e\u003cp\u003eUsing the CG topology approach we can easily convert a protein structure database, such as PDB, from standard protein models to CG topology structures. In this paper, all PDB structural data are sourced from CATH\u003csup\u003e\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e\u003c/sup\u003e, a database organized as a classification of protein structures. To mitigate the influence of structural redundancy on the data distribution, we employ the CATH-dataset-nonredundant S40 for training data, yielding 31,886 non-redundant protein structures. 90% of structures are used for training and 10% are used for independent validation.\u003c/p\u003e\u003cp\u003eDiffTopo/MirrorTopo\u0026thinsp;+\u0026thinsp;RFdiffusion framework\u003c/p\u003e\u003cp\u003eFirst, we introduce DiffTopo, a Equivariant diffusion model\u003csup\u003e\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e,\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e\u003c/sup\u003e for generating CG topology conditioned on SSEs strings. Based on previous works on denoising diffusion, given a data point sampled from a real data distribution \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{z}_{0}\\sim\\:q\\left(z\\right)\\)\u003c/span\u003e\u003c/span\u003e, the forward diffusion process is defined by adding Gaussian noise gradually to the data point \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{z}_{0}\\)\u003c/span\u003e\u003c/span\u003e from corrupted samples \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{z}_{t}\\)\u003c/span\u003e\u003c/span\u003e. For one CG topology data \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{z}_{0}\\)\u003c/span\u003e\u003c/span\u003e, there are three points \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:x\\)\u003c/span\u003e\u003c/span\u003e and corresponding secondary structure type \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:u\\)\u003c/span\u003e\u003c/span\u003e. At time step \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:t\\)\u003c/span\u003e\u003c/span\u003e \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:=0,\\dots\\:,T\\)\u003c/span\u003e\u003c/span\u003e, the conditional distribution of the intermediate data state \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{z}_{t}\\)\u003c/span\u003e\u003c/span\u003e given the previous state is defined by the multivariate normal distribution,\u003cdiv id=\"Equa\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equa\" name=\"EquationSource\"\u003e\n$$\\:q\\left({z}_{t}\\mid\\:{z}_{t-1}\\right)=N\\left({z}_{t};\\sqrt{1-{\\beta\\:}_{t}}{z}_{t-1},{\\beta\\:}_{t}I\\right)$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eThe step sizes are controlled by a variance schedule \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\left\\{{\\beta\\:}_{t}\\in\\:\\left(\\text{0,1}\\right)\\right\\}}_{t=1}^{T}\\)\u003c/span\u003e\u003c/span\u003e. The process is constructed to be Markovian and if we let \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\alpha\\:}_{t}=1-{\\beta\\:}_{t}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\underset{\\_}{{\\alpha\\:}_{t}}={\\prod\\:}_{i=1}^{t}{\\alpha\\:}_{i}\\)\u003c/span\u003e\u003c/span\u003e, we can obtain the distribution of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{z}_{t}\\)\u003c/span\u003e\u003c/span\u003e given \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:z\\)\u003c/span\u003e\u003c/span\u003e\u003cdiv id=\"Equb\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equb\" name=\"EquationSource\"\u003e\n$$\\:q\\left({z}_{t}\\mid\\:{z}_{0}\\right)=N\\left({z}_{t};\\sqrt{\\underset{\\_}{{\\alpha\\:}_{t}}}{z}_{0},\\left(1-\\underset{\\_}{{\\alpha\\:}_{t}}\\right)I\\right)$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eFor the backward process, we need to learn a model \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:p\\)\u003c/span\u003e\u003c/span\u003e to approximate these conditional probabilities.\u003cdiv id=\"Equc\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equc\" name=\"EquationSource\"\u003e\n$$\\:p\\left({z}_{t-1}\\mid\\:{z}_{t},u\\right)=N\\left({z}_{t-1};\\mu\\:\\left({z}_{t},t,u\\right),\\varSigma\\:\\left({z}_{t},t\\right)\\right)$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\mu\\:}_{\\theta\\:}\\left({z}_{t},t,u\\right)\\)\u003c/span\u003e\u003c/span\u003e is predicted by the neural network. In this paper function \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\mu\\:\\)\u003c/span\u003e\u003c/span\u003e that predicts noise \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{ϵ}_{\\theta\\:}\\)\u003c/span\u003e\u003c/span\u003e is implemented as geometric vector perceptron (GVP)\u003csup\u003e\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e\u003c/sup\u003e. The input to GVP is the noised version of the point coordinates \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{z}_{t}\\)\u003c/span\u003e\u003c/span\u003e and point feature \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{h}_{t}\\)\u003c/span\u003e\u003c/span\u003e at time \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:t\\)\u003c/span\u003e\u003c/span\u003e and context \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:u\\)\u003c/span\u003e\u003c/span\u003e. Note that the predicted noise \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:ϵ\\)\u003c/span\u003e\u003c/span\u003e includes coordinate and feature components, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:ϵ=[{ϵ}_{z}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{ϵ}_{h}]\\)\u003c/span\u003e\u003c/span\u003e. The predicted noise can be calculated by\u003cdiv id=\"Equd\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equd\" name=\"EquationSource\"\u003e\n$$\\:{ϵ}_{\\theta\\:}\\left({z}_{t},t\\right)=GVP\\left({z}_{t},{h}_{t},u,t\\right)-\\left[{z}_{t},{h}_{t}\\right]$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eIn addition to DiffTopo, we introduce MirrorTopo, a method for generating mirror-image topologies directly from native protein folds. Given a CG representation of the native fold, we obtain the mirror topology by inverting the z-coordinates while preserving the x- and y-coordinates. This transformation effectively reflects the structure across the xy-plane, yielding a topologically mirrored conformation without altering sequence or connectivity.\u003c/p\u003e\u003cp\u003eFollowing the generation of the coarse-grained (CG) topology via the diffusion process or mirroring of the native CG topology, the resulting topology enables the direct extraction of the positions and lengths of standard SSEs. Since CG topology retains a certain level of fuzziness, we need a backbone level model to systematically search for the designable backbone corresponding to CG topology. Here, we leverage RFdiffusion, a highly effective backbone generation model to generate designable backbones from CG topology. Initially, we construct a protein sketch from standard SSEs based on CG topology and utilize the motif modeling function inherent in RFdiffusion to directly connect the SSEs. Subsequently, we employ the partial diffusion approach to identify a reasonable backbone structure within this interconnected structure. Then we employ the ProteinMPNN\u003csup\u003e\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e algorithm for fixed backbone sequence design to obtain amino acid sequences. These sequences are then input for the structure prediction algorithm AF2\u003csup\u003e19\u003c/sup\u003e to predict the structure, which is use to validate the designability of generated backbones.\u003c/p\u003e\u003cp\u003eSecondary structure descriptors\u003c/p\u003e\u003cp\u003e\u003cem\u003eSecondary structure descriptor construction\u003c/em\u003e - To define secondary structure conditions within a specific topological class, we constructed strings representing the desired sequence of secondary structure elements\u0026mdash;for example, \"EEEHHEHE\", where each character denotes a secondary structure type (E: β-strand, H: α-helix). These SSE strings were generated either manually or through algorithmic enumeration by specifying the number, type, and order of elements. The example \"EEEHHEHE\" encodes a topology with eight secondary structure elements\u0026mdash;five β-strands and three α-helices arranged in a particular sequence. These strings were used as input constraints to the generative model, directing the synthesis of backbones consistent with the specified topology.\u003c/p\u003e\u003cp\u003e\u003cem\u003eRandom Secondary Structure Descriptor Sampling -\u003c/em\u003e We randomly selected 512 unique SSE order strings, ranging in length from 3 to 12 elements, from the CATH database. These sequences served as topological constraints to guide the generation of varied backbone conformations.\u003c/p\u003e\u003cp\u003e\u003cem\u003eLow-Frequency Secondary Structure Descriptor Sampling -\u003c/em\u003e To specifically target rare topologies, we compiled all possible SSE order strings of length 8 and computed their frequency across the entire CATH database. Strings with fewer than two occurrences were classified as low-frequency compositions. This procedure yielded a total of 2,471 underrepresented SSE patterns, which were used to probe the model\u0026rsquo;s ability to generalize beyond commonly observed folds.\u003c/p\u003e\u003cp\u003eMapping Protein structures to CATH database\u003c/p\u003e\u003cp\u003eWe computed ProteinMPNN encoder and ProtDomainSegmentor\u003csup\u003e\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e\u003c/sup\u003e embeddings for all CATH structures from Ingraham et al.\u003csup\u003e7\u003c/sup\u003e, selecting only those without chain breaks, with resolution better than 3.0 \u0026Aring;, and R\u003csub\u003efree\u003c/sub\u003e below 0.25. Both embedding types are computed at the per-residue level and averaged along the sequence dimension to obtain 128-dimensional embeddings for ProteinMPNN and 4096-dimensional embeddings for ProtDomainSegmentor. To capture structural features at different hierarchical scales, we extract local fold features using the first encoder layer of ProteinMPNN\u003csup\u003e\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e, which encodes featured geometries of nearest neighbors, while global fold features are captured using the pre-final layer of ProtDomainSegmentor, which is predictive of CATH Architecture classification. We project the CATH embeddings onto their first two principal components (PCs) and apply the same transformation to embeddings of dark fold and mirror fold structures for direct comparison.\u003c/p\u003e\u003cp\u003eTERM analysis for protein structures\u003c/p\u003e\u003cp\u003eFor TERM analysis, we randomly selected 50 representative protein structures from the CATH database, including both dark folds and mirror folds that have been experimentally expressed. All structures were queried using the TERMANAL software package against the default fragment database \u0026ldquo;bc-30-sc-20141022\u0026rdquo;, a non-redundant subset of protein chains in the Protein Data Bank (PDB) with a 30% sequence identity cutoff and determined by X-ray crystallography at high resolution. The resulting abundance and structural score files were used for subsequent analyses.\u003c/p\u003e\u003cp\u003eProtein expression and purification\u003c/p\u003e\u003cp\u003eDNA sequences corresponding to the amino acid sequence of the designed proteins were optimized for E. Coli codon usage and purchased from Twist Bioscience as fragments, including a solvent exposed C-terminal poly histidine tag. The DNA fragments were cloned via Gibson cloning into a pET11b backbone, and plasmids were validated by sequencing. Expression was conducted in BL21(DE3) strain: 0.5 L cultures were inoculated from an overnight culture in Luria-Bertani medium supplemented with ampicillin (100 \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\mu\\:g/mL\\)\u003c/span\u003e\u003c/span\u003e) and incubated at 37\u0026deg;C and 220 rpm shaking until reaching an OD600 of 0.6\u0026ndash;0.8; the cultures were then further incubated overnight at 18\u0026deg;C and inoculated with 0.5 mM IPTG to induce protein expression. Bacterial cells were harvested by centrifugation and pellets were resuspended in 30 mL lysis buffer (100 mM TRIS, pH 7.5, 500 mM NaCl, 5% glycerol, 1 mg/mL lysozyme, 1 mM PMSF, 4 \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\mu\\:g/ml\\)\u003c/span\u003e\u003c/span\u003e DNase, 0.5X Cell Lytic lysis reagent), incubated at 4\u0026deg;C on a rotating wheel for 1h, before the lysate was clarified by high-speed centrifugation followed by sterile filtration (0.22 \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\mu\\:m\\)\u003c/span\u003e\u003c/span\u003e). The lysate was then applied on a 5 ml His-Trap FF column on an \u0026Auml;KTA pure system (GE Healthcare), the bound proteins were first washed in a Tris based buffer (50 mM Tris, pH 7.5, 500 mM NaCl, 10 mM imidazole) and then eluted with 500 mM imidazole. Polishing purification step was performed by size exclusion chromatography on a Hiload 16/600 Superdex 75 pg column (GE Healthcare) in PBS buffer (pH 7.4). The peak corresponding to the expected molecular weight was collected and concentrated for analysis.\u003c/p\u003e\u003cp\u003eCircular dichroism spectroscopy\u003c/p\u003e\u003cp\u003eFar-UV circular dichroism spectra were collected between wavelengths of 200 and 250 nm on a Chirascan\u0026trade; spectrometer (AppliedPhotophysics) spectrometer in a 1 mm path-length quartz cuvette. Proteins were prepared in PBS at a concentration of 0.3 mg/mL by default. Wavelength spectra were measured with a scanning speed of 20 nm/min and a response time of 0.125 s and each measurement was reference subtracted using a cuvette with PBS. The thermal denaturation curves were collected by measuring the change in ellipticity at 220 nm from 20 to 90\u0026deg;C with 2\u0026deg;C increments. Size-exclusion chromatography combined with multi-angle light scattering monodispersity and oligomeric state of the purified proteins were assessed by multi-angle light scattering, thanks to molecular weight determination in solution. 100 \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\mu\\:g\\)\u003c/span\u003e\u003c/span\u003e of protein were injected into a Superdex 75 10/300 GL column (GE Healthcare) in PBS buffer (pH 7.4) at a flow rate of 0.5 ml/min and detected on in-line multi-angle light-scattering detector (DAWN TREOS, Wyatt). Static light-scattering signals were recorded, and the scatter data were analyzed by ASTRA software (version 8.0.2.5 64-bit, Wyatt).\u003c/p\u003e\u003cp\u003eCrystallization and X-ray structure determination\u003c/p\u003e\u003cp\u003eThe N5 design was crystallized at a concentration of 19 mg/ml using sitting drop vapor diffusion at 18\u0026deg;C in 0.1 M MES pH 6.5, 0.2 M KSCN, 25% w/v PEG 2000 MME buffer (Clear Strategy Screen I, Molecular Dimensions). The N30 design was crystallized at a concentration of 24.3 mg/ml using sitting drop vapor diffusion at 18\u0026deg;C in 0.1 M MES pH6.5; 30% v/v PEG Smear Low buffer (BCS Screen, Molecular Dimensions). The N37 design was crystallized at a concentration of 28 mg/ml using sitting drop vapor diffusion at 18\u0026deg;C in 0.1 M NaOAc pH5.5, 0.2 M KSCN, 25% w/v PEG 2000 MME buffer (Clear Strategy Screen I, Molecular Dimensions). The M7 design was crystallized at a concentration of 24.7 mg/ml using sitting drop vapor diffusion at 18\u0026deg;C in 0.1 M MES pH6.5, 0.2 M KSCN, 15% w/v PEG 4000 buffer (Clear Strategy Screen I, Molecular Dimensions). The M16 design was crystallized at a concentration of 42 mg/ml using sitting drop vapor diffusion at 18\u0026deg;C in 0.1 M Tris pH8.0, 0.04 M sodium format, 0.04 M CaCl2, 25% v/v PEG Smear Low buffer (BCS Screen, Molecular Dimensions). Crystals were cryoprotected in 25% glycerol and flash cooled in liquid nitrogen. Diffraction data was collected at the European Synchrotron Radiation Facility MASSIF-1 and MASSIF-3 beamlines, Grenoble, France at a temperature of 100 K. Crystallographic data was processed using the autoPROC package\u003csup\u003e43\u003c/sup\u003e. Phases were obtained by molecular replacement using the full or partial designed model in Phaser\u003csup\u003e44\u003c/sup\u003e. Atomic model rebuilding and refinement was performed using COOT \u003csup\u003e\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e\u003c/sup\u003eand Phenix refine\u003csup\u003e44\u003c/sup\u003e. The quality of refined models was assessed using MolProbity\u003csup\u003e46\u003c/sup\u003e.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eData availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll data are available in the main text or as Supplementary Information. Atomic coordinates and structure factors of the reported X-ray structures have been deposited in the PDB under accession numbers 9R2K(N30), 9R2O(N5), 9R2L(N37), 9R2R(M7), 9R2V(M16).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCode availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eDiffTopo code and jupyter notebook is available at GitHub (https://github.com/YangyangMiao/DiffTopo ). AF2 model used for \\predictions can be downloaded from (https://github.com/sokrypton/ColabFold ). ProteinMPNN is available at GitHub (https://github.com/dauparas/ProteinMPNN ).\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eTaylor, W. R. Exploring Protein Fold Space. \u003cem\u003eBiomolecules\u003c/em\u003e \u003cstrong\u003e10\u003c/strong\u003e, 193 (2020).\u003c/li\u003e\n\u003cli\u003eSimons, K. T., Kooperberg, C., Huang, E. \u0026amp; Baker, D. Assembly of protein tertiary structures from fragments with similar local sequences using simulated annealing and Bayesian scoring functions. \u003cem\u003eJ. Mol. Biol.\u003c/em\u003e \u003cstrong\u003e268\u003c/strong\u003e, 209\u0026ndash;225 (1997).\u003c/li\u003e\n\u003cli\u003eMinami, S. \u003cem\u003eet al.\u003c/em\u003e Exploration of novel \u0026alpha;\u0026beta;-protein folds through de novo design. \u003cem\u003eNat. Struct. Mol. Biol.\u003c/em\u003e 1\u0026ndash;9 (2023) doi:10.1038/s41594-023-01029-0.\u003c/li\u003e\n\u003cli\u003eBonet, J. \u003cem\u003eet al.\u003c/em\u003e Rosetta FunFolDes \u0026ndash; A general framework for the computational design of functional proteins. \u003cem\u003ePLOS Comput. Biol.\u003c/em\u003e \u003cstrong\u003e14\u003c/strong\u003e, e1006623 (2018).\u003c/li\u003e\n\u003cli\u003eHuang, P.-S. \u003cem\u003eet al.\u003c/em\u003e RosettaRemodel: a generalized framework for flexible backbone protein design. \u003cem\u003ePloS One\u003c/em\u003e \u003cstrong\u003e6\u003c/strong\u003e, e24109 (2011).\u003c/li\u003e\n\u003cli\u003eYim, J. \u003cem\u003eet al.\u003c/em\u003e SE(3) diffusion model with application to protein backbone generation. Preprint at https://doi.org/10.48550/arXiv.2302.02277 (2023).\u003c/li\u003e\n\u003cli\u003eIngraham, J. B. \u003cem\u003eet al.\u003c/em\u003e Illuminating protein space with a programmable generative model. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e623\u003c/strong\u003e, 1070\u0026ndash;1078 (2023).\u003c/li\u003e\n\u003cli\u003eSCUBA-D: a freshly trained diffusion model generates high-quality protein structures. \u003cem\u003eNat. Methods\u003c/em\u003e \u003cstrong\u003e21\u003c/strong\u003e, 1990\u0026ndash;1991 (2024).\u003c/li\u003e\n\u003cli\u003eGeffner, T. \u003cem\u003eet al.\u003c/em\u003e Proteina: Scaling Flow-based Protein Structure Generative Models. Preprint at https://doi.org/10.48550/arXiv.2503.00710 (2025).\u003c/li\u003e\n\u003cli\u003eLin, Y. \u0026amp; AlQuraishi, M. Generating Novel, Designable, and Diverse Protein Structures by Equivariantly Diffusing Oriented Residue Clouds. Preprint at https://doi.org/10.48550/arXiv.2301.12485 (2023).\u003c/li\u003e\n\u003cli\u003eBose, A. J. \u003cem\u003eet al.\u003c/em\u003e SE(3)-Stochastic Flow Matching for Protein Backbone Generation. Preprint at https://doi.org/10.48550/arXiv.2310.02391 (2024).\u003c/li\u003e\n\u003cli\u003eLu, T. \u003cem\u003eet al.\u003c/em\u003e Conditional Protein Structure Generation with Protpardelle-1c. \u003cem\u003ebioRxiv\u003c/em\u003e 2025.08.18.670959 (2025) doi:10.1101/2025.08.18.670959.\u003c/li\u003e\n\u003cli\u003eTaylor, W. R. \u003cem\u003eet al.\u003c/em\u003e Prediction of protein structure from ideal forms. \u003cem\u003eProteins\u003c/em\u003e \u003cstrong\u003e70\u003c/strong\u003e, 1610\u0026ndash;1619 (2008).\u003c/li\u003e\n\u003cli\u003eHarteveld, Z. \u003cem\u003eet al.\u003c/em\u003e Exploring \u0026ldquo;dark-matter\u0026rdquo; protein folds using deep learning. \u003cem\u003eCell Syst.\u003c/em\u003e \u003cstrong\u003e15\u003c/strong\u003e, 898-910.e5 (2024).\u003c/li\u003e\n\u003cli\u003eHarteveld, Z. \u003cem\u003eet al.\u003c/em\u003e A generic framework for hierarchical de novo protein design. \u003cem\u003eProc. Natl. Acad. Sci.\u003c/em\u003e \u003cstrong\u003e119\u003c/strong\u003e, e2206111119 (2022).\u003c/li\u003e\n\u003cli\u003eHarteveld, Z. \u003cem\u003eet al.\u003c/em\u003e Deep sharpening of topological features for de novo protein design. in (2022).\u003c/li\u003e\n\u003cli\u003eSingh, A. Bottom-up de novo protein design. \u003cem\u003eNat. Methods\u003c/em\u003e \u003cstrong\u003e18\u003c/strong\u003e, 233\u0026ndash;233 (2021).\u003c/li\u003e\n\u003cli\u003eLin, Z. \u003cem\u003eet al.\u003c/em\u003e Evolutionary-scale prediction of atomic-level protein structure with a language model. \u003cem\u003eScience\u003c/em\u003e https://doi.org/10.1126/science.ade2574 (2023) doi:10.1126/science.ade2574.\u003c/li\u003e\n\u003cli\u003eDauparas, J. \u003cem\u003eet al.\u003c/em\u003e Robust deep learning-based protein sequence design using ProteinMPNN. \u003cem\u003eScience\u003c/em\u003e \u003cstrong\u003e378\u003c/strong\u003e, 49\u0026ndash;56 (2022).\u003c/li\u003e\n\u003cli\u003eJumper, J. \u003cem\u003eet al.\u003c/em\u003e Highly accurate protein structure prediction with AlphaFold. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e596\u003c/strong\u003e, 583\u0026ndash;589 (2021).\u003c/li\u003e\n\u003cli\u003eYim, J. \u003cem\u003eet al.\u003c/em\u003e Fast protein backbone generation with SE(3) flow matching. Preprint at http://arxiv.org/abs/2310.05297 (2023).\u003c/li\u003e\n\u003cli\u003eTrippe, B. L. \u003cem\u003eet al.\u003c/em\u003e Diffusion probabilistic modeling of protein backbones in 3D for the motif-scaffolding problem. Preprint at https://doi.org/10.48550/arXiv.2206.04119 (2023).\u003c/li\u003e\n\u003cli\u003eKnudsen, M. \u0026amp; Wiuf, C. The CATH database. \u003cem\u003eHum. Genomics\u003c/em\u003e \u003cstrong\u003e4\u003c/strong\u003e, 207\u0026ndash;212 (2010).\u003c/li\u003e\n\u003cli\u003eBurley, S. K. \u003cem\u003eet al.\u003c/em\u003e RCSB Protein Data Bank: powerful new tools for exploring 3D structures of biological macromolecules for basic and applied research and education in fundamental biology, biomedicine, biotechnology, bioengineering and energy sciences. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cstrong\u003e49\u003c/strong\u003e, D437\u0026ndash;D451 (2020).\u003c/li\u003e\n\u003cli\u003evan Kempen, M. \u003cem\u003eet al.\u003c/em\u003e Fast and accurate protein structure search with Foldseek. \u003cem\u003eNat. Biotechnol.\u003c/em\u003e \u003cstrong\u003e42\u003c/strong\u003e, 243\u0026ndash;246 (2024).\u003c/li\u003e\n\u003cli\u003eZhang, Y. \u0026amp; Skolnick, J. Scoring function for automated assessment of protein structure template quality. \u003cem\u003eProteins\u003c/em\u003e \u003cstrong\u003e57\u003c/strong\u003e, 702\u0026ndash;710 (2004).\u003c/li\u003e\n\u003cli\u003eZhang, Y. \u0026amp; Skolnick, J. TM-align: a protein structure alignment algorithm based on the TM-score. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cstrong\u003e33\u003c/strong\u003e, 2302\u0026ndash;2309 (2005).\u003c/li\u003e\n\u003cli\u003eXu, J. \u0026amp; Zhang, Y. How significant is a protein structure similarity with TM-score = 0.5? \u003cem\u003eBioinformatics\u003c/em\u003e \u003cstrong\u003e26\u003c/strong\u003e, 889\u0026ndash;895 (2010).\u003c/li\u003e\n\u003cli\u003eAltschul, S. F., Gish, W., Miller, W., Myers, E. W. \u0026amp; Lipman, D. J. Basic local alignment search tool. \u003cem\u003eJ. Mol. Biol.\u003c/em\u003e \u003cstrong\u003e215\u003c/strong\u003e, 403\u0026ndash;410 (1990).\u003c/li\u003e\n\u003cli\u003eS\u0026ouml;ding, J., Biegert, A. \u0026amp; Lupas, A. N. The HHpred interactive server for protein homology detection and structure prediction. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cstrong\u003e33\u003c/strong\u003e, W244\u0026ndash;W248 (2005).\u003c/li\u003e\n\u003cli\u003eGabler, F. \u003cem\u003eet al.\u003c/em\u003e Protein Sequence Analysis Using the MPI Bioinformatics Toolkit. \u003cem\u003eCurr. Protoc. Bioinforma.\u003c/em\u003e \u003cstrong\u003e72\u003c/strong\u003e, e108 (2020).\u003c/li\u003e\n\u003cli\u003eLu, T., Liu, M., Chen, Y., Kim, J. \u0026amp; Huang, P.-S. Assessing Generative Model Coverage of Protein Structures with SHAPES. 2025.01.09.632260 Preprint at https://doi.org/10.1101/2025.01.09.632260 (2025).\u003c/li\u003e\n\u003cli\u003eEguchi, R. R. \u0026amp; Huang, P.-S. Multi-scale structural analysis of proteins by deep semantic segmentation. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cstrong\u003e36\u003c/strong\u003e, 1740\u0026ndash;1749 (2020).\u003c/li\u003e\n\u003cli\u003eXu, Y. \u0026amp; Zhu, T. F. Mirror-image T7 transcription of chirally inverted ribosomal and functional RNAs. \u003cem\u003eScience\u003c/em\u003e https://doi.org/10.1126/science.abm0646 (2022) doi:10.1126/science.abm0646.\u003c/li\u003e\n\u003cli\u003eZhang, G. \u0026amp; Zhu, T. F. Mirror-image trypsin digestion and sequencing of D-proteins. \u003cem\u003eNat. Chem.\u003c/em\u003e \u003cstrong\u003e16\u003c/strong\u003e, 592\u0026ndash;598 (2024).\u003c/li\u003e\n\u003cli\u003eMackenzie, C. O., Zhou, J. \u0026amp; Grigoryan, G. Tertiary alphabet for the observable protein structural universe. \u003cem\u003eProc. Natl. Acad. Sci.\u003c/em\u003e \u003cstrong\u003e113\u003c/strong\u003e, E7438\u0026ndash;E7447 (2016).\u003c/li\u003e\n\u003cli\u003eZheng, F., Zhang, J. \u0026amp; Grigoryan, G. Tertiary structural propensities reveal fundamental sequence/structure relationships. \u003cem\u003eStruct. Lond. Engl. 1993\u003c/em\u003e \u003cstrong\u003e23\u003c/strong\u003e, 961\u0026ndash;971 (2015).\u003c/li\u003e\n\u003cli\u003eHeinig, M. \u0026amp; Frishman, D. STRIDE: a web server for secondary structure assignment from known atomic coordinates of proteins. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cstrong\u003e32\u003c/strong\u003e, W500-502 (2004).\u003c/li\u003e\n\u003cli\u003eHo, J., Jain, A. \u0026amp; Abbeel, P. Denoising Diffusion Probabilistic Models. Preprint at https://doi.org/10.48550/arXiv.2006.11239 (2020).\u003c/li\u003e\n\u003cli\u003eSatorras, V. G., Hoogeboom, E. \u0026amp; Welling, M. E(n) Equivariant Graph Neural Networks. Preprint at https://doi.org/10.48550/arXiv.2102.09844 (2022).\u003c/li\u003e\n\u003cli\u003eJing, B., Eismann, S., Suriana, P., Townshend, R. J. L. \u0026amp; Dror, R. Learning from Protein Structure with Geometric Vector Perceptrons. Preprint at https://doi.org/10.48550/arXiv.2009.01411 (2021).\u003c/li\u003e\n\u003cli\u003eEmsley, P., Lohkamp, B., Scott, W. G. \u0026amp; Cowtan, K. Features and development of Coot. \u003cem\u003eActa Crystallogr. D Biol. Crystallogr.\u003c/em\u003e \u003cstrong\u003e66\u003c/strong\u003e, 486\u0026ndash;501 (2010).\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"nature-portfolio","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"","title":"Nature Portfolio","twitterHandle":"","acdcEnabled":false,"dfaEnabled":false,"editorialSystem":"ejp","reportingPortfolio":"","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"computational protein design, dark-matter folds, de novo design, machine learning for molecular generation, deep learning for protein design, mirror topology design","lastPublishedDoi":"10.21203/rs.3.rs-7827245/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7827245/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eA major challenge in computational \u003cem\u003ede novo\u003c/em\u003e protein design is the exploration of uncharted areas within the protein structural space. However, the large degrees of freedom of protein backbones complicate the sampling process during protein design. Machine learning-based models have made great strides in this problem, however due to their nature they tend to exploit rather than explore the data distribution used for training the neural networks. To address some of these challenges, we propose a new coarse grained protein structure representation generative method, \u003cb\u003eDiffTopo\u003c/b\u003e, a diffusion model which increases the sampling efficiency and diversity. Combined with a backbone level protein generative model like RFdiffusion, novel protein folds can be generated rapidly, allowing for efficient exploration of the designable topology space. Interestingly, we have discovered that by mirroring the topological organization of native proteins using a pipeline named \u003cb\u003eMirrorTopo\u003c/b\u003e, we can readily expand the known fold space. We generated and experimentally characterized 30 different novel topologies from DiffTopo and 6 different novel mirror topologies from MirrorTopo. The developed framework relying on low resolution sampling provides new means for fold exploration challenges, which could in principle enhance our knowledge of the first principles of protein structure and folding, as well as create new opportunities for functional design.\u003c/p\u003e","manuscriptTitle":"Leveraging protein representations to explore uncharted fold spaces with generative models","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-11-04 05:59:27","doi":"10.21203/rs.3.rs-7827245/v1","editorialEvents":[],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"nature-structural-and-molecular-biology","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"nsmb","sideBox":"Learn more about [Nature Structural \u0026 Molecular Biology](http://www.nature.com/nsmb/)","snPcode":"","submissionUrl":"","title":"Nature Structural \u0026 Molecular Biology","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"Nature Research","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"6a14bd89-c71a-4c14-9799-6c1189b11c85","owner":[],"postedDate":"November 4th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[{"id":56879267,"name":"Biological sciences/Computational biology and bioinformatics"},{"id":56879268,"name":"Biological sciences/Structural biology"}],"tags":[],"updatedAt":"2026-02-11T16:21:47+00:00","versionOfRecord":[],"versionCreatedAt":"2025-11-04 05:59:27","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7827245","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7827245","identity":"rs-7827245","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.