Sequence Complexity Dictates Polymer Mixing

preprint OA: closed CC-BY-NC-ND-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Polymer association in confined spaces governs diverse phenomena from protein aggregation to DNA condensation. We investigate the sequence-level mechanisms underlying this behavior through exact enumeration of two confined AB-type block copolymers on a two-dimensional lattice, revealing how sequence complexity controls mixing against de-mixing. We reveal that sequence complexity (heterogeneity), quantified by Shannon entropy H 2 , acts as a determinant between self-folded low-mixed state and high-mixed state. High-complexity sequences ( H 2 = 1.86 bits) with short-length repeats achieve near-complete inter-chain overlap through cooperative chain collapse, while low-complexity blocky sequences ( H 2 = 1.26 bits) maintain extended conformations with less overlapping. Free energy analysis reveals a steep increase in mixing barriers for low-complexity compared to surmountable barrier for high-complexity sequences. We find that geometric confinement modulates but does not override these sequence-dependent behaviors, with asymmetric confinement enhancing heterogeneity in overlapping for both sequence type (pronounced in high- H 2 sequences). Our theory could be applicable in assessing the phase behavior of repeat protein and nucleic acid sequences.
Full text 51,528 characters · extracted from oa-pdf · click to expand
Sequence Complexity Dictates Polymer Mixing Dibyajyoti Mohanta,1 Manish Dwivedi, 2 Hiranmay Maity, 3 and Debaprasad Giri 4 1)Department of Chemistry, University at Buffalo, Buffalo, NY, USA a) 2)Department of Physics, Institute of Science, Banaras Hindu University, Varanasi, India 3)Department of General Science, BITS, Pilani, Dubai Campus, Dubai International Academic City, Dubai, UAE 4)Department of Physics, Indian Institute of Technology (BHU), Varanasi, India Polymer association in confined spaces governs diverse phenomena from protein aggregation to DNA conden- sation. We investigate the sequence-level mechanisms underlying this behavior through exact enumeration of two confined AB-type block copolymers on a two-dimensional lattice, revealing how sequence complex- ity controls mixing against de-mixing. We reveal that sequence complexity (heterogeneity), quantified by Shannon entropy H2, acts as a determinant between self-folded low-mixed state and high-mixed state. High- complexity sequences (H 2 = 1 .86 bits) with short-length repeats achieve near-complete inter-chain overlap through cooperative chain collapse, while low-complexity blocky sequences ( H2 = 1.26 bits) maintain ex- tended conformations with less overlapping. Free energy analysis reveals a steep increase in mixing barriers for low-complexity compared to surmountable barrier for high-complexity sequences. We find that geometric confinement modulates but does not override these sequence-dependent behaviors, with asymmetric confine- ment enhancing heterogeneity in overlapping for both sequence type (pronounced in high-H2 sequences). Our theory could be applicable in assessing the phase behavior of repeat protein and nucleic acid sequences. I. INTRODUCTION Liquid-Liquid Phase Separation (LLPS) mediated by weak interactions provided the framework for compart- mentalization of biomolecules, including protein and RNA, even without any membrane 1. The formation and properties of such membraneless compartments known as biocondensates depend not only on the overall com- position of constituent proteins and nucleic acids but are also sensitive to the underlying sequence pattern- ing of interaction motifs 2,3. Such sequence-encoded ef- fects exhibit both in single-molecule and multicompo- nent systems, where sequence composition and pattern- ing shape the conformational landscape of biopolymers 4. For example, intrinsically disordered proteins (IDPs) and regions (IDRs), which lack well-defined tertiary struc- tures, adopt conformational ensembles directly encoded by their sequences 5,6. Similarly, nucleic acid sequences also control the formation of secondary structures such as hairpin loops, bulges in RNA, bubbles in DNA, through the specific arrangement of nucleotides 4,7. Polyam- pholytes with well-mixed positive and negative charges readily collapse, while polyelectrolytes like poly(A)-RNA or poly(U)-RNA resist condensation under physiological conditions7,8. This sequence dependent organization and association of biomolecules serve as a fundamental deter- minant of phase separation behavior 9. In vivo, sequence effects never act in isolation as biomolecules fold and function within crowded, geomet- rically confined environments that fundamentally alter their behavior 10,11. The cellular environment contains 20 −40% macromolecules by volume, creating steric con- a)Electronic mail: [email protected] straints that restrict conformational entropy and mod- ulate interaction energetics 12,13. Recent studies have demonstrated that confinement can reduce critical con- centrations for phase separation by 4-10 fold, making LLPS thermodynamically accessible at physiological pro- tein concentrations 14,15. Confinement creates a deli- cate balance: it restricts the conformational entropy of RNA and proteins, while at the same time enhancing local concentration effects that strengthen intermolecu- lar interactions16,17. Understanding how geometric con- finement couples with sequence patterning to determine phase separation driven multi-polymer association is es- sential. Biomolecule aggregation often features hierarchy of as- sembly that begins with monomers, dimers, oligomers, which subsequently grow into network-like topolo- gies characteristic of condensates beyond the critical density4,18–20. In this context, we leverage insights from previous theoretical and simulation studies on confined two-polymer systems (dimers) to identify the key factors that control their folding, mixing, and segregation 21–28. However, the majority of these studies have focused on homopolymers, which are largely dominated by en- tropic effects 24,26 . By design, these models neglect the sequence-specific enthalpic interactions that drive the microphase separation essential in facilitating polymer association29. Thus, a fundamental gap remains in our understanding of how sequence patterning and geomet- ric confinement jointly regulate the initial association of polymers. To remedy this, we present an exactly solvable minimal model that captures interplay between sequence complexity, geometric confinement, and polymer associ- ation, through exhaustive enumeration. A central challenge in predicting sequence-dependent polymer interactions is the quantitative characteriza- tion of sequence patterning. Shannon entropy, calcu- .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint 2 FIG. 1. Schematic diagram of the two block co-polymer (BCP) model. (Top) Examples of balanced and unbalanced BCP sequences with A-type marked as blue and B-type as yellow. The sequence complexity (H 2) is shown for each pair. (Bottom) The two BCPs are confined within following geome- tries: a square box (S s) of 10 × 10 and a rectangular box (S r) of 20 × 5 lattice units. The shaded middle layer indicates the set of possible starting positions for the polymers. In Ss , chains are enumerated from one of the two shaded layers, as 5 th and 6 th layers (Lys) are mirror images to each other. Pink diamonds represent an example ”origin-pair” for the two polymers (labeled 1 and 2). Strength of interaction between same type of beads is taken as ϵ, whereas non-specific inter- action strength is zero. lated from the probability distribution of n-letter words (n-grams), provides a robust measure of sequence pat- terning defined as sequence complexity. It correlates strongly with conformational heterogeneity and phase separation propensity 19,30–32. To this end, we em- ploy a system of two AB-type block copolymers where sequence complexity is quantified by the bi-gram en- tropy, H2 . This approach is motivated by recent stud- ies on block copolymer systems composed of two dis- tinct bead types, investigating sequence-controlled self- assembly and aggregation 29,33. This study unfolds in three key parts. First, we map out the role of sequence complexity and confinement anisotropy in shaping mixing of two polymers. Next, we in-depth analyze the free energy landscape of mixing for different H2 sequences within varying confinement. Finally, we explore how sequence complexity affects the actual size of two-polymer systems during the overlap- ping pathway. We conclude our study by investigating the mixing behavior for a diverse pair of sequences. II. MODEL AND METHOD We model a system of two block co-polymers (BCPs), each consisting of N + 1 = 18 beads, as self(mutual)- attracting-self(mutual)-avoiding walks (S(M)AS(M)A W) within a symmetric confinement Ss with side length Lxs = Lys = 10 lattice units, and an asymmetric con- finement Sr with Lxr = 20 and Lyr = 5 lattice units. Both confinements have an equal area of 100 square lat- tice units, maintaining a constant volume fraction of (2 × 18)/100 = 0.36 [Fig. 1]. We define a specific, at- tractive interaction ϵsp = ϵ between non-bonded nearest- neighbor beads of same type ( A − A or B − B). Inter- actions between different beads type ( A − B) termed as non-specific contacts and are set to zero (shown as × in Fig. 1). A. Sequence Entropy H2 of Block copolymers The n-gram entropy measures the randomness of over- lapping subsequences (‘words’) in a polymer chain, where beads are treated as letters (e.g., A/B for a diblock copolymer or hydrophobic/polar in HP models). Follow- ing Herzel34, the bi-gram (2-letter word) entropy H2 for a given 18-letter sequence of A and B is computed from the frequencies of all overlapping 2-letter words, yielding 17 such bi-grams per sequence. H2 = − X w∈{AA,AB,BA,BB} p(w) log2 p(w) (1) where the probability of each bi-gram w in the 18- bead sequence is estimated as p(w) = n(w)/17, with n(w) being the count of that bi-gram. All four bi-grams are equal-a-priori, but their empirical probabilities p(w) reflect the actual bi-gram composition of the specific sequence35. Label Sequence (length = 18) H 2 (in bits) Balanced AAABBBAAABBBAAABBB 1.86 Balanced AAAAAAAAABBBBBBBBB 1.26 Unbalanced AABBAABBAABBAABBAA 1.99 Unbalanced AAAAAABBBBBBAAAAAA 1.45 TABLE I. Bi-gram entropy H2 (in bits) for four designed 18-bead polymer sequences composed of A and B beads. “Bal- anced” sequences have equal numbers of A and B beads, whereas “Unbalanced” sequences have unequal A/B compo- sition. Larger H2 indicates greater sequence complexity. Our study focuses on four distinct two-BCP systems characterized by their sequence complexity H2 and over- all A/B composition (Table I). Each system consists of a sequence listed in Table I together with its complemen- tary partner obtained by exchanging A and B at every position. For example, the system H2 = 1.86 consists of the AAABBBAAABBBAAABBB sequence and its com- plementary partner BBBAAABBBAAABBBAAA with A/B = 1 in individual sequence and overall two-BCP system. In balanced systems ( H2 = 1.86 and H2 = 1.26), each chain has equal numbers of A and B beads, while in un- balanced systems (H2 = 1.99 and H2 = 1.45), the chains .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint 3 have unequal A/B composition. Although many distinct sequences share similar H2 values, this work focuses on four representative systems chosen to represent the high and low ends of the complexity spectrum, from short pe- riodic patterns to more blocky ones, with H2 values rang- ing approximately from 1.99 to 1.26 bits. B. Enumeration procedure The exact partition function of this binary BCP system can be calculated by exact enumeration of all conforma- tions of two polymer chains of length N = 17 (i.e., N + 1 beads)36. At this fixed chain length and volume frac- tion (ϕc = 0.36), the full conformational space is already large, so extending to longer chains would be computa- tionally expensive26,36. To obtain reliable statistics at fixedN without increas- ing chain length, we systematically scan the accessible volume by varying the positions of the polymer origins in- side the confinement. Specifically, we restrict the starting positions (origins) of both polymers to the middle layer of the box, as indicated by the shaded line in [Fig. 1]. For each possible unique pair of origin positions on this line (an “origin-pair”), we enumerate all conformations of the two BCPs; providing an exhaustive conformational ensemble for each origin-pair. This procedure yields 10 × 9 = 90 distinct origin-pairs for the Ss box and 20 × 19 = 380 origin-pairs for the Sr box. Throughout the study, observables are reported both as averages over individual origin-pairs and as en- semble averages over all origin-pairs, thereby ensuring a uniform scan of the confinement’s effect while remaining computationally feasible. The system is evaluated in the canonical ensemble. For a origin-pair, the partition function Z is computed as a sum over all possible conformations: Z = X msp,mnsp,ssp,snsp C(msp, mnsp, ssp, snsp) (exp(−βϵ))msp (exp(−βϵ))ssp (2) Here, C(msp, mnsp, ssp, snsp) denotes the number of conformations with: msp: inter-polymer specific contacts (A–A or B–B be- tween polymers) mnsp: inter-polymer non-specific contacts (A–B between polymers) ssp: intra-polymer specific contacts (A–A or B–B within each polymer, summed) snsp: intra-polymer non-specific contacts (A–B within each polymer, summed) In this model (see [Fig. 1]), only specific contacts are energetically weighted (i.e. ϵnsp = 0). So, the sys- tem energy depends solely on the total number of spe- cific contacts (m sp + ssp) (Eq. 2). On the other hand, polymer overlap or mixing involves contributions from both specific ( msp) and non-specific ( mnsp) contacts. β = 1/(kBT ) is the inverse temperature. For this study, we used reduced units where the Boltzmann constant kB is set to 1 and the temperature T is fixed at 0 .5. To cal- culate the overall ensemble average for a given sequence type, we followed Eq. 2 across all origin pairs. III. RESULTS 1. Statistical Mechanics of Contact Distribution Multiple polymer assemblies such as condensates formed by LLPS, chromatin organization, and crowded macromolecular complexes, exhibit network-like archi- tectures that typically rely on strength of multivalent inter-chain interactions4,37. In [Fig.2], we analyze how se- quence complexity and confinement jointly control BCP mixing by examining the fraction of inter-chain contacts per polymer, ϕnc = Nc/(N+1) with Nc = msp+mnsp, for high-complexity (H 2 = 1.86, short-length repeats) and low-complexity (H2 = 1.26, blocky) balanced sequences under symmetric (S s, panels a,b,c) and asymmetric ( Sr, panels d,e,f) confinement. Each radially outward bar in the polar plots (Fig.2) represents ϕnc for a distinct origin pair of the two BCPs within the confinement. At weak interaction strength ϵ = −0.5 [Fig.2 (a), (d)] , both sequences exhibit a modest and compara- ble overlap ( ϕnc < 0.5) with H2 = 1.26 being slightly higher. This indicates that conformational entropy dom- inates over low specific interaction energy which is in- sufficient to elicit sequence-dependent differences 38. As attraction strength increases to ϵ = −1.0, the distribu- tions bifurcate [Fig. 2(b,e)]. Polymers with higher H2 display a much broader spread of ϕnc values and access configurations with substantially higher overlap. Under strong attractive condition ϵ = −3.0 [Fig. 2(c,f)], the distinction is maximized. H2 = 1.86 sequences approach near-complete overlap ( ϕnc ∼ 1.1–1.2), whereas blocky H2 = 1.26 sequences saturate at significantly lower ϕnc values (≈ 0.6), in both confinements 30,33. Ensemble-averaged ϕnc over all origin-pairs, reinforce these trends [Fig. 2(g,h)]. At low |ϵ|, H2 = 1 .26 ex- hibit mild higher ϕnc than H2 = 1.86 sequences, reflect- ing their greater sensitivity to weak attractions. How- ever, as |ϵ| increases, high-H2 sequences overtake low-H2 sequences in both confinements, achieving much higher .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint 4 0 0.5 1 00.51 0 0.5 1 H2=1.86 H2=1.26 0 0.5 1 0 0.5 1 0 0.5 1 (i) Ss Sr 0 0.2 0.4<(?nc) (j) Ss Sr 0 0.2 0.4<(?nc) (k) Ss Sr 0 0.2 0.4<(?nc) -3 -2 -1 000 0.5 1 ?nc (g) H2=1.86 H2=1.26 -3 -2 -1 000 0.5 1 ?nc (h) (l) 0.450.48 0.891.07 Ss Sr 0 0.5 1jd?nc=d0j H2=1.86 H2=1.26 0 = -0.5 0 = -1.0 0 = -3.0 (a) (b) (c) (d) (e) (f) Sr Ss 0=-0.5 0=-1.0 0=-3.0 FIG. 2. Sequence-dependent inter-chain overlapping across origin pairs for balanced sequences. Polar plots of fraction of overlap contacts ( ϕnc = Nc/(N + 1)) for all origin-pairs at interaction strengths ϵ = −0.5, −1.0, −3.0 (a,b,c) for high-complexity (H2 = 1.86), and (d,e,f) for low-complexity (H 2 = 1.26) sequences in square confinement Ss and rectangular confinement Sr . Compared to Ss, Sr shows more heterogeneity in ϕnc, especially in Sr and at stronger ϵ. Ensemble-averaged ⟨ϕnc⟩ over origin pairs versus interaction strengthϵ in Ss (g) and Sr (h), showing higher mixing for high-H2 at strong interactions for both confinements. Heterogeneity in mixing, quantified as standard deviation of ϕnc across origin-pairs in Ss and Sr (i, j,k for different values of ϵ), highlighting enhanced variability in high-H2 and Sr. Initial mixing rates dϕnc/d|ϵ| in the linear regime (l), revealing faster but shallower response for low-H 2, while slow but progressive mixing for high– H2 (pronounced in Sr). saturation overlap. The contiguous arrangement of iden- tical beads in low H2 blocky chains make intra-chain fold- ing entropically favorable, allowing the BCPs to satisfy most of their valency via compact self-collapsed states (see [Fig. S1] in SI) 39–41. In contrast, the high com- plexity H2 sequences make intra-chain folding entropi- cally costly due to short-repeat lengths, leaving attrac- tive beads exposed to form the multivalent inter-chain mixing19,30,42. To quantify how confinement anisotropy modulates the heterogeneity in fraction of inter-contact distribu- tion, we computed the standard deviation of ϕnc across all origin-pairs at different strengths [Fig. 2(i,j,k)]. At all interaction strengths, both type of sequences within Sr exhibits higher heterogeneity than Ss, reflecting the broader spectrum of accessible conformations and over- lap states in anisotropic confinement 22,43. Moreover, among the sequences, high-complexity sequences consis- .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint 5 tently show larger heterogeneity than their low-H2 coun- terparts in both geometries, indicating richer landscape of partially and highly mixed configurations for low and high complexity sequences respectively. The linear fits to ϕnc in the linear regime (low |ϵ|) re- veal distinct “mixing rates” for the two sequence types [Fig. 2(l)]. Low-H 2 sequences display a steeper initial slope, presenting a faster but shallow response to in- creasing attraction that quickly saturates at low overlap. In contrast, high-H 2 sequences exhibit a more moder- ate initial increase but ultimately progress to complete overlap, with effective mixing rates |dϕnc/dϵ| ≈ 0.45 in Ss and ≈ 0.89 in Sr. Thus, asymmetric confinement nearly doubles the mixing rate for high-complexity se- quences by pre-aligning chains and reducing the entropic penalty of association 44. This is consistent with prior studies showing that confinement can shift phase bound- aries and reduce the critical concentration required for LLPS 14,15,44,45. To explore the molecular origins of sequence complexity-dependent overlapping behavior, we present the fraction of total specific contacts ϕsp = (m sp + ssp)/(N + 1) in [Fig. 3 (a)]. A key finding is that low- H2 = 1.26 sequences consistently form a higher fraction of ϕsp compared to their high-H 2 counterparts across both confinements . This observation appears counterin- tuitive at first, as one might expect that a greater num- ber of specific contacts would promote enhanced mix- ing for H2 = 1.26. However, this can be explained by decomposing total specific contacts into their inter- and intra-chain components expressed asϕmsp = msp/(N+1) and ϕssp = ssp/(N + 1) in [Fig. 3 (b) and (c)] respec- tively. Blocky, low-H2 sequences sequences have contigu- ous blocks of identical beads (more homotypic contacts per segment), making intra-chain folding energetically fa- vorable (with higher ϕssp values for H2 = 1.26 in [Fig. 3 (c)]. This reduces inter-chain overlap driven by specific contacts defined by ϕmsp [Fig. 3 (b)] for H2 = 1.26. In sharp contrast, short-length repeats in high- H2 chains maintain more open conformations because intra-chain folding incurs a large entropic penalty when attempting to bring distant non-contiguous attractive beads into con- tact 19,33. This leaves many attractive sites exposed and available to form inter-chain-specific contacts featuring higher ϕmsp for H2 = 1 .86 in [Fig. 3 (b)]. This high- lights how low-H 2 sequences are energetically favored (via more total contacts) but entropically disfavored (due to self-folding) for inter-chain mixing, whereas high-H 2 sequences trade some energetic gains for entropic advan- tages through enhanced inter-chain association [Fig. 3 (b)]41. These trends align with studies demonstrating that strong homotypic interactions in proteins promote self-association over intermolecular mixing 42. The fraction of inter-chain non-specific contacts (ϕmnsp = mnsp/(N + 1)), which do not contribute to the Boltzmann weight but still takes part in mixing, is low overall but also remains higher for short-range se- quences ( H2 = 1.86) in both confinements [Fig. 3 (d)]. This arises because, in high-H 2 chains, non-interacting blocks interspersed between favorable interaction blocks can remain overlapped between two chains due to the reduced entropic cost of constraining these short seg- ments 40. This behavior mirrors observations in sticker- spacer models of LLPS, where short spacers tend to re- main in proximity due to the strong attachments of adja- cent stickers 19,40. Finally, intra-chain non-specific con- tacts ( ϕsnsp = snsp/(N + 1)) provide minimal contri- butions to folding, thereby emphasizing the dominance of energy-driven specific interactions [Fig. 3(e)]. Collec- tively, Figs. 2 and 3 demonstrate that the spatial ar- rangement of interaction sites, rather than their total number or binding strength, governs whether polymers assemble into self-collapsed or form cooperative polymer organization30,33. 2. Free Energy Profile Analysis To determine the thermodynamic driving forces gov- erning sequence-dependent mixing, we analyzed the free energy profiles F (Nc), calculated from partition function with Nc inter-chain contacts Z(Nc) [Figure 4]: F (Nc) = −kBT ln Z(Nc) (3) where Z(Nc) is expressed as 22, Z(Nc = msp + mnsp) = X ssp,snsp C(msp, mnsp, ssp, snsp) exp(−βϵ(msp + ssp)) (4) At both weak [Fig. 4 (a)] and strong [Fig. 4 (b)] in- teraction strengths, high-H 2 exhibit higher free energy in the low to moderate contact regime (N c < 20) com- pared to low-complexity sequences. This initial penalty arises because the H2 = 1.86 chains must overcome a translational and orientational entropic penalty to align due to their short repeats, which is not the case for low- H2 blocky chains 46,47. However, when sufficient mix- ing (N c > 20) has taken place in high-H 2 sequences, the network of inter-chain contacts becomes sufficiently connected (enthalpy dominated) to overcome the en- tropic cost of chain alignment. So, the threshold which arises near Nc ≈ 20 indicates the surmountable free en- ergy barrier, beyond which the free energy landscape in- .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint 6 -3 -2 -1 000 0.2 0.4 0.6 0.8 1 1.2?sp -3 -2 -1 000 0.2 0.4 0.6 0.8 1 1.2?mspH2(Ss)=1.86 H2(Ss)=1.26 H2(Sr)=1.86 H2(Sr)=1.26 h H2(Ss)i=1.86 h H2(Ss)i=1.26 h H2(Sr)i=1.86 h H2(Sr)i=1.26 -3 -2 -1 000 0.2 0.4 0.6 0.8 1 1.2?ssp -3 -2 -1 000 0.2 0.4 0.6 0.8 1 1.2?mnsp -3 -2 -1 000 0.2 0.4 0.6 0.8 1 1.2?snsp (b) (c) (e)(d) (a) FIG. 3. Decomposition of contact fractions reveals sequence-dependent preferences for self-association versus mixing (a) For both confinements total phase-separated (i.e. ϕsp = ϕmsp + ϕssp) contact fraction are higher for H2 = 1.26 than H2 = 1.86, representing H2 = 1.26 energetically dominated. (b) Inter-chain specific contact fraction ϕmsp are higher for high-H 2, promoting sequence-specific mixing. (c) Intra-chain specific contact fraction ϕssp dominate for low-H 2, favoring self-association. (d) Inter non-specific contact fraction ϕmnsp (A-B) are higher for high- H2, contributing to overall overlap. (e) Intra-chain non-specific contact fraction ϕsnsp remain minimal for both sequences. verts (F H2=1.86 < F H2=1.26), indicating that the highly mixed network becomes the thermodynamic ground state (global minima) for high-H 248. The free energy barrier (yellow bar) and the low-mixed and high-mixed (global) minima marked as green bars are shown in [Fig. 4 (b)]. This mixing threshold at (Nc ≈ 20) basically corresponds to the valley region in density of stateD(Nc) between the double-peaks [Fig.S2 (c) and (g) in SI]. In contrast, low- H2 sequences cross their free energy minima at Nc ≈ 15 and beyond higher mixing incurs elevated free energy (unfavorable), indicating their unimodal nature in D(Nc) [Fig. S2 (b),(f) (d) and (h) in SI]. To quantify the resistance to mixing, we computed the local gradient of the free energy landscape, ∂F/∂N c at Nc [Fig. 4 (c)-(f)]. A steep positive gradient indicates a strong thermodynamic penalty against mixing, while a flat or negative gradient implies accessible transitions 49. In both confinements, the local gradient is negative at low Nc due to the fact that the confinement itself gives stability to low mixed states 44. There is a crossover of ∂F/∂N c from negative to positive values at Nc ≈ 12 −13 indicating beyond this point overlapping or mixing re- quires overcoming an increasing thermodynamic penalty. At low strength ϵ = −1.0, H2 = 1.26 sequence’s ∂F/∂N c lead over high-H2 sequences in both Ss and Sr at high Nc regime indicating their high resistance in mixing. At high strength ϵ = −3.0, [Fig. 4 (e),(f)], after crossing zero in both Ss and Sr, high H2 = 1 .86 sequences constantly feature low gradient values until a spike in ∂F/∂N c ≈ 20 (marked by vertical down arrow), indicating the free en- ergy barrier or mixing threshold separating the minimas in [Fig. 4 (b)]. Beyond this peak local gradient, we see ∂F/∂N c going near-zero or negative again (especially in Sr) showing the propensity of mixing of high- H2 se- quences. The effect of confinement is also apparent since the local peak gradient in Ss is ∂F/∂N c|Nc=20 is 2.15 kBT while for Sr, it is 1.55 kBT making Sr the preferred environment for mixing. While low- H2 sequences show maximum local gradient at higher Nc revealing the mix- ing is highly unfavorable. The negative ∂F/∂N c values on both sides of the spike for high- H2 sequences indicate a double-well free energy landscape. In such landscapes, the system can undergo spontaneous transitions between basins (mixed to demixed) via thermal fluctuations 50,51. Song et al.52 reported such results as tau-RNA complexes reversibly cross LLPS boundaries, with sequence charge patterns tuning the effective barrier between dilute and dense phases. We further quantified the thermodynamic preference for the mixed state by calculating the global free energy difference ∆F mixing between the mixed (N c > N ∗ c ) and less-mixed phases [Fig. 4 (g)]. The probability of the mixed phase is given by the normalized partition sum: PN ∗ c = P Nc>N ∗ c Ω(Nc)e−βϵ(msp+ssp) Z (5) where N ∗ c = 20 is the identified nucleation threshold [Fig. S2, SI] and used for both low and high H2 sequences . The relative stability is then: ∆Fmixing = −kBT ln  PN ∗ c 1 − PN ∗ c  (6) ∆Fmixing for H2 = 1.26 remains positive (disfavoring mixing) even at high |ϵ| before saturating just below 0. However, the ∆Fmixing value for the H2 = 1.86 sequences drops gradually and reaching ∼ −3 for Ss and ∼ −4 for Sr at ϵ = −3, confirming a cooperative transition to the mixed phase 43. .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint 7 0 5 10 15 20 25Nc -30 -28 -26 -24 -22 -20 -18 -16F(Nc) 0 5 10 15 20 25Nc -70 -65 -60 -55 -50 -45 -40F(Nc) -3 -2 -1 00-4 -3 -2 -1 0 1 2 3"F H2=1.86[Ss] H2=1.26[Ss] H2=1.86[Sr] H2=1.26[Sr] -1 0 1dF(Nc)/dNc 0 10 20Nc -1 0 1dF(Nc)/dNc Ss -2 0 2 4 Sr 0 10 20Nc -2 0 2 H2=1.26 H2=1.86 (a) 0=-1.0 0=-3.0 (d) (e)(c) (g) (f) 0=-3.00=-1.0 (b) barrierminima FIG. 4. Free energy landscapes and barriers reveal the thermodynamic basis for sequence-dependent mixing : (a, b) Free energy profiles F (Nc) as a function of inter-chain contacts Nc at weak (ϵ = −1.0) and strong ( ϵ = −3.0) interaction strengths, respectively, for high-complexity ( H2 = 1.86, blue) and low-complexity ( H2 = 1.26, red) sequences in square ( Ss) and rectangular ( Sr,triangles) confinements. High- H2 sequences exhibit an initial entropic penalty at low Nc but a favorable crossover to lower free energy at high Nc, indicative of a surmountable barrier leading to stable mixed states. The barrier is marked as yellow bar separating the two minimas marked as green bars in (b). Whereas, low- H2 sequences show a monotonic increase, disfavoring mixing. (c,d,e,f) Local gradients ∂F/∂N c (in units of kBT ) highlighting barriers and driving forces: (c, e) at ϵ = −1.0 for Ss and Sr; (d, f) at ϵ = −3.0 for Ss and Sr. High- H2 sequences display a characteristic spike around Nc ≈ 20 (downward arrows), marking the barrier between demixed and mixed minima, with negative gradients post-barrier favoring mixing more pronounced in asymmetric Sr confinement. Low-H 2 sequences show persistently positive gradients, resisting overlap. (g) Free energy difference ∆F mixing between mixed and demixed states as a function of interaction strength ϵ, confirming cooperative transitions to mixing for high- H2 (decreasing ∆ F ) versus persistent demixing for low-H 2 (positive ∆F ) in both confinements 3. Effect of sequence complexity on polymer size The joint probability distribution P (Rend, Nc), over all origin pairs with Rend as the end-to-end distance of polymer 1 within the binary polymer system, reveals sequence-dependent coupling between chain compactness and intermolecular contacts [Fig. 5]. At weak inter- actions (ϵ = −1.0), both sequences in Ss exhibit low- probability distributions (p ≈ 0.05 − 0.1) spanning mod- erate Rend (4 − 10) and Nc (10 − 22), indicating limited mixing [Fig. 5 a,b]. In Sr, high-H 2 show emerging het- erogeneity with broader Rend (10 − 20), consistent with anisotropy enhancing alignment opportunities as seen in contact heterogeneity [Fig. 2] and broader Rend distri- butions (SI [Fig. S1 e]) 43,45. At strong interactions (ϵ = −3.0), high-H 2 sequences display discrete conformational clusters, echoing double- well free energy landscapes [Fig. 4 (b)] that enable coop- erative transitions to mixed states. At higher strength, H2 = 1.86 [Fig. 5 (e),(g)] promotes mixing as well as col- lapse with higher P (Rend, Nc) than H2 = 1.26 [Fig. 5 (f),(h)]. For better identification of mixed-collapsed .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint 8 FIG. 5. Joint probability distributions P (Rend, Nc) reveal sequence and confinement dependent conformational ensembles. Top: Distributions at ϵ = −1.0 (a-d) and Middle: at ϵ = −3.0 (e-h) with Ss (square; a,b,e,f) and Sr (rectangular; c,d,g,h). Bottom: k-means clusters at ϵ = −3.0 (i-l), with centroids (circles) and colors denote clusters. High- H2 shows discrete basins favoring low Rend/high Nc, indicative of co-collapse; low-H 2 remains extended compared to H2 = 1.86 with moderate Nc. phase population, we analyze P (Rend, Nc) by k-cluster analysis, where we focus on high mixed cluster. In Ss [Fig. 5 (i)], k-means clustering (k=2) identifies a highly mixed cluster (41.22%) at low Rend (4.09 ± 0.8), indi- cating inter-chain-driven co-collapse. In Sr [Fig. 5 g,k], k=3 reveals greater heterogeneity, with the H2 = 1.86 sequences exhibit high-mixed ( Nc = 21 .43 ± 1.78) clus- ter centroid at Rend = 3.78 ± 0.4, confirming anisotropy amplifies sequence-encoded mixing without overriding it. However, low-H2 sequences (H 2 = 1.26) remain largely unimodal in Ss [Fig. 5 f,j]; centered at moderate Rend (7 − 10) and Nc (10 − 15), resisting compaction due to self-association. In Sr [Fig. 5 h,l] k=3, confinement in- duces mild heterogeneity, but higher Nc cluster occurs at extended Rend (11.59 ± 1.61, Nc = 16 .13± 1.00), sug- gesting expansion to accommodate contacts which is op- posite to high-H2 co-collapse33,40. This shows the mixing pathway of high-H 2 sequences exhibit co-collapse state while low-H2 remains in low-mix-expanded state. 4. Unbalanced Sequence Analysis The extension to unbalanced sequences H2 = 1.99 and H2 = 1.45 [Fig. 1], where the ratio of beads type in each FIG. 6. Sequence complexity governs mixing across compositions. Unbalanced sequences with higher H2 = 1.99 ( blue) exhibit increased inter-chain overlapping compared to low- H2 = 1.45 (orange) in both square ( Ss) and rectangular (Sr) confinements. However, the absolute fraction of inter- chain contacts is systematically reduced for unbalanced se- quences relative to balanced high- H2 sequences. polymer chain deviates from unity ( A/B ̸= 1), confirms the robustness of complexity-dependent behavior across distinct composition (stoichiometry)29. We compared en- semble average of fraction of inter-chain contact forma- tion /mixing between a high-complexity unbalanced se- quence ( H2 = 1.99) with a low-complexity unbalanced .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint 9 FIG. 7. Shannon bi-gram entropy (H 2) of poly-XY regions across eukaryotic proteomes (Human, Mouse, Drosophila, Zebrafish, and Plant). All proteomes exhibit clustered H2 values (mean ≈ 1.29 bits) with comparable in- terquartile ranges. Natural polyXY sequences fall between our designed low-complexity (H2 = 1.26) and high-complexity (H2 = 1.86) variants, occupying an intermediate mixing regime. sequence (H2 = 1.45) illustrated in [Fig 6 (a) and (b)]. Importantly, while overall mixing levels are reduced for unbalanced compared to balanced sequences, still high- H2 = 1 .99 continues to favor mixing ϕ1.99 N c > ϕ 1.45 N c in both confinements. This reduction arises because unbal- anced sequences have a higher translational entropy than balanced ones, thereby resulting less alignment driven mixing53 . Our observation that ’unbalanced’ stoichiome- tries (A ̸= B) reduce polymer mixing warrants careful distinction from the ’magic-ratio effect’ recently estab- lished for multicomponent condensates 53. In systems driven by heterotypic interactions (e.g., sticker-spacer models where A binds B), phase separation is often sup- pressed at 1:1 stoichiometry because the formation of small, entropy-favored monomers (self folding)53. In con- trast, our system is governed by homotypic attractions (A binds A, B binds B), where the 1:1 ’balanced’ com- position promotes inter-chain mixing. This analysis rein- forces that sequence specific interaction primarily governs polymer associations, while overall sequence composition modulates the extent of mixing but cannot override the effects of patterning. FIG. 8. Schematic free energy landscape linking se- quence complexity to mixing pathways. Free energy landscape schematic for block copolymers ( H2 = 1.26 and H2 = 1.86) as a function of inter-chain contacts Nc at ϵ = −3. The blue curve illustrates the landscape for H2 = 1.26 se- quence, which exhibit a free energy minimum at low mixing states favoring self folding over mixing due to the blockiness in H2 = 1.26. The orange curve represents H2 = 1.86 variants, where the free energy minimum shifts to higher Nc values in the mixed state. The short length repeats in H2 = 1.86 sup- presses intrachain self-folding and promotes intermixing with the neighboring chain. IV. DISCUSSION This study establishes sequence complexity, quantified by bigram Shannon entropy H2, as a molecular determi- nant dictating the balance between intermolecular mix- ing and de-mixing in confined AB-type block copolymers (BCPs). Our systematic study employing exact enumer- ation of two short BCP chains on a two-dimensional lat- tice under distinct confinement geometries reveals four critical aspects: i) high-complexity sequences ( H2 = 1.86) not only favor enhanced mixing but also introduce greater heterogeneity in overlapping configurations, par- ticularly under asymmetric confinement ( Sr). This het- erogeneity in Sr arises from distributed short-length re- peats, providing entropic flexibility that stabilizes diverse high Nc states, as evidenced by bimodal D(Nc) distribu- tions [Fig. S2 , SI] and multiple conformational clusters in P (Rend, Nc) (Fig. 5). In contrast, low-complexity blocky sequences (H 2 = 1.26, 1.45) exhibit unimodal, low-mix ensembles, where self-folding via intra-chain con- tacts yields energetic stability (higher phase-separated contacts) but resists extensive inter-chain overlap [Fig. S1] . ii) Free energy landscapes further reveal this tran- sition as high-H 2 sequences feature surmountable bar- riers ( 1-2 kBT ) separating low and high mix minima, enabling cooperative transitions to near-complete over- lap. Whereas, low- H2 landscapes show steep free energy .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint 10 barrier at elevated Nc. iii) The high-mixed high complex- ity sequences undergo co-collapsed state (coupled coil-to- globule transition) where as low H2 sequences remain ex- panded and low-mixed across confinements. iv) ’High H2 favoring high mixing’ scenario holds also for unbalanced compositions (A/B ̸= 1), emphasizing sequence pattern- ing primacy over composition. Fig. 8 illustrates this sequence-dependent behavior: low-complexity sequences (H2 = 1 .26) favor self-folded conformations with min- ima in low-mixed states, while high-complexity sequences (H2 = 1 .86) undergo cooperative mixing with minima shifted to higher Nc. To assess whether the sequence-complexity regimes ex- plored by our model are relevant to real proteins, we an- alyzed the bi-gram entropy H2 of polyXY regions (two- residue low-complexity tracts) across multiple proteomes [Fig. 7] 54. Remarkably, the distribution of H2 values is centered around 1.3 ± 0.3 bits, which overlaps precisely with the range in which our model predicts a transition between self-folded, weakly mixed states and inter-chain mixed, co-collapsed states. This correspondence suggests that many polyXY segments in vivo may be naturally tuned to function near the boundary separating mixed and de-mixed states, enabling cells to regulate conden- sate assembly and disassembly in response to changing conditions. Although our enumeration-based framework simplifies the vast complexity of single proteins and condensates on a two dimensional lattice, its strength lies in its exact calculation of the partition function, which pro- vides a precise thermodynamic characterization of the interplay between sequence patterning and confinement. More broadly, the sequence–structure relationship un- covered here reflects a fundamental organizing principle that spans from single-chain folding to higher-order con- densate assembly. By employing a minimal but exactly solvable model, we demonstrated that the patterning of interaction sites, not just their overall abundance, is a critical determinant of polymer association. 1. Acknowledgement DM thanks Erik Clarkson for reading the manuscript. 1S. F. Banani, H. O. Lee, A. A. Hyman, and M. K. Rosen, Nature Reviews Molecular Cell Biology 18, 285–298 (2017). 2T. Kaur, M. Raju, I. Alshareedah, R. B. Davis, D. A. Po- toyan, and P. R. Banerjee, Nature Communications 12 (2021), 10.1038/s41467-021-21089-4. 3T. de Vries, M. Novakovic, Y. Ni, I. Smok, C. Inghel- ram, M. Bikaki, C. P. Sarnowski, Y. Han, L. Emmanouilidis, G. Padroni, A. Leitner, and F. H.-T. Allain, Science Advances 10 (2024), 10.1126/sciadv.adm7435. 4H. T. Nguyen, N. Hori, and D. Thirumalai, Nat. Chem. 14, 775 (2022). 5R. K. Das and R. V. Pappu, Proceedings of the National Academy of Sciences 110, 13392–13397 (2013). 6A. S. Holehouse and R. V. Pappu, Annual Review of Biophysics 47, 19–39 (2018). 7A. Jain and R. D. Vale, Nature 546, 243 (2017). 8L. Sawle and K. Ghosh, The Journal of Chemical Physics 143 (2015), 10.1063/1.4929391. 9A. A. Hyman, C. A. Weber, and F. J¨ ulicher, Annual Review of Cell and Developmental Biology 30, 39–58 (2014). 10D. Lucent, V. Vishal, and V. S. Pande, Proceedings of the Na- tional Academy of Sciences 104, 10430–10434 (2007). 11J. Mateos-Langerak, M. Bohn, W. de Leeuw, O. Giromus, E. M. M. Manders, P. J. Verschure, M. H. G. Indemans, H. J. Gierman, D. W. Heermann, R. van Driel, and S. Goetze, Pro- ceedings of the National Academy of Sciences 106, 3812–3817 (2009). 12S. Mittal, R. K. Chowhan, and L. R. Singh, Biochimica et Bio- physica Acta (BBA) - General Subjects 1850, 1822–1831 (2015). 13A. A. M. Andr´ e and E. Spruijt, International Journal of Molec- ular Sciences 21, 5908 (2020). 14L. K. Davis, I. J. Ford, and B. W. Hoogenboom, eLife 11 (2022), 10.7554/elife.72627. 15B. Monterroso, W. Margolin, A. J. Boersma, G. Rivas, B. Pool- man, and S. Zorrilla, Chemical Reviews 124, 1899–1949 (2024). 16J. Ahlers, E. M. Adams, V. Bader, S. Pezzotti, K. F. Winklhofer, J. Tatzelt, and M. Havenith, Biophysical Journal120, 1266–1275 (2021). 17S. Mondal, K. Narayan, S. Botterbusch, I. Powers, J. Zheng, H. P. James, R. Jin, and T. Baumgart, Nature Communications 13 (2022), 10.1038/s41467-022-32529-0. 18S. Jeon, Y. Jeon, J.-Y. Lim, Y. Kim, B. Cha, and W. Kim, Signal Transduction and Targeted Therapy 10 (2025), 10.1038/s41392- 024-02070-1. 19J.-M. Choi, A. A. Hyman, and R. V. Pappu, Phys. Rev. E 102, 042403 (2020). 20J. D. Schmit, M. Feric, and M. Dundr, Trends in Biochemical Sciences 46, 525–534 (2021). 21Y. Liu and B. Chakraborty, Physical Biology 9, 066005 (2012). 22J. M. Polson and D. A. Rehel, Soft Matter 17, 5792–5805 (2021). 23S. Jun, A. Arnold, and B.-Y. Ha, Phys. Rev. Lett. 98, 128303 (2007). 24E. Minina and A. Arnold, Soft Matter 10, 5836 (2014). 25Z. Liu, X. Capaldi, L. Zeng, Y. Zhang, R. Reyes-Lamothe, and W. Reisner, Nature Communications 13 (2022), 10.1038/s41467- 022-31398-x. 26D. Mohanta, Soft Matter 19, 4991–5000 (2023). 27D. A. Rehel and J. M. Polson, Soft Matter 19, 1092–1108 (2023). 28L. Zeng, X. Capaldi, Z. Liu, and W. W. Reisner, Physical Review E 109 (2024), 10.1103/physreve.109.024501. 29B. G. Weiner, A. G. T. Pyo, Y. Meir, and N. S. Wingreen, PLOS Computational Biology 17, e1009748 (2021). 30R. Takaki and D. Thirumalai, Proceedings of the Na- tional Academy of Sciences 121, e2409973121 (2024), https://www.pnas.org/doi/pdf/10.1073/pnas.2409973121. 31S. Biswas and D. A. Potoyan, PLOS Computational Biology 21, e1012826 (2025). 32Y.-f. Wang, C.-l. Ren, and Y.-q. Ma, Physical Review E 112 (2025), 10.1103/783t-5xqx. 33G. L. Dignon, W. Zheng, Y. C. Kim, R. B. Best, and J. Mittal, PLOS Computational Biology 14, 1 (2018). 34H. Herzel, W. Ebeling, and A. O. Schmitt, Phys. Rev. E 50, 5061 (1994). 35L. Yu, D. K. Tanwar, E. D. S. Penha, Y. I. Wolf, E. V. Koonin, and M. K. Basu, Proceedings of the National Academy of Sciences 116, 3636 (2019), https://www.pnas.org/doi/pdf/10.1073/pnas.1814684116. 36C. Vanderzande, in Lattice Models of Polymers (Cambridge Uni- versity Press, Cambridge, 1998) pp. xi–xiv. 37D. M. Mitrea and R. W. Kriwacki, Cell Communication and Sig- naling 14, 1 (2016). 38D. Mohanta, D. Giri, and S. Kumar, Journal of Statistical Me- chanics: Theory and Experiment 2019, 043501 (2019). 39A. Sood and B. Zhang, Biophysical Journal 123, 1815–1826 (2024). .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint 11 40T. S. Harmon, A. S. Holehouse, M. K. Rosen, and R. V. Pappu, eLife 6 (2017), 10.7554/elife.30294. 41J. A. Joseph, J. R. Espinosa, I. Sanchez-Burgos, A. Garaizar, D. Frenkel, and R. Collepardo-Guevara, Biophysical Journal 120, 1219–1230 (2021). 42A. C. Murthy, G. L. Dignon, Y. Kan, G. H. Zerze, S. H. Parekh, J. Mittal, and N. L. Fawzi, Nature Structural amp; Molecular Biology 26, 637–648 (2019). 43D. Mohanta, Soft Matter 18, 2790–2799 (2022). 44M. P. Taylor, T. M. Prunty, and C. M. O’Neil, The Journal of Chemical Physics 152, 094901 (2020), https://pubs.aip.org/aip/jcp/article- pdf/doi/10.1063/1.5144818/15574842/0949011online.pd f. 45E. J. Saltzman and M. Muthukumar, The Journal of Chemical Physics 131 (2009), 10.1063/1.3267487. 46A. M. Rumyantsev, A. Johner, and J. J. de Pablo, ACS Macro Letters 10, 1048 (2021). 47F. Yu and S. Sukenik, The Journal of Physical Chemistry B 127, 4235 (2023), pMID: 37155239. 48E. W. Martin, T. S. Harmon, J. B. Hopkins, S. Chakravarthy, J. J. Incicco, P. Schuck, A. Soranno, and T. Mittag, Nat. Com- mun. 12, 4513 (2021). 49D. Wales, Energy Landscapes: Applications to Clusters, Biomolecules and Glasses (Cambridge University Press, 2001). 50K. A. Dill and H. S. Chan, Nature Structural amp; Molecular Biology 4, 10–19 (1997). 51R. P. Sear, Journal of Physics: Condensed Matter 19, 033101 (2007). 52Y. Lin, J. McCarty, J. N. Rauch, K. T. Delaney, K. S. Kosik, G. H. Fredrickson, J.-E. Shea, and S. Han, eLife8, e42571 (2019). 53Y. Zhang, B. Xu, B. G. Weiner, Y. Meir, and N. S. Wingreen, eLife 10 (2021), 10.7554/elife.62403. 54P. Mier and M. A. Andrade-Navarro, Computational and Struc- tural Biotechnology Journal 20, 5516–5523 (2022). .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint Low mixed state High mixed state Nc F(Nc) High sequence complexity Low sequence complexity .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint 0 5 10 15 20 25Nc -30 -28 -26 -24 -22 -20 -18 -16F(Nc) 0 5 10 15 20 25Nc -70 -65 -60 -55 -50 -45 -40 F(Nc) -3 -2 -1 00 -4 -3 -2 -1 0 1 2 3"F H2=1.86[Ss] H2=1.26[Ss] H2=1.86[Sr] H2=1.26[Sr] -1 0 1 dF(Nc)/dNc 0 10 20Nc -1 0 1 dF(Nc)/dNc Ss -2 0 2 4 Sr 0 10 20Nc -2 0 2 H2=1.26 H2=1.86 (a) 0=-1.0 0=-3.0 (d) (e)(c) (g) (f) 0=-3.00=-1.0 (b) barrier minima .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint -3 -2 -1 00 0 0.2 0.4 0.6 0.8 1 1.2?sp -3 -2 -1 00 0 0.2 0.4 0.6 0.8 1 1.2 ?mspH2(Ss)=1.86 H2(Ss)=1.26 H2(Sr)=1.86 H2(Sr)=1.26 h H2(Ss)i=1.86 h H2(Ss)i=1.26 h H2(Sr)i=1.86 h H2(Sr)i=1.26 -3 -2 -1 00 0 0.2 0.4 0.6 0.8 1 1.2 ?ssp -3 -2 -1 00 0 0.2 0.4 0.6 0.8 1 1.2 ?mnsp -3 -2 -1 00 0 0.2 0.4 0.6 0.8 1 1.2 ?snsp (b) (c) (e)(d) (a) .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint 0 0.5 1 0 0.5 1 0 0.5 1 H2=1.86 H2=1.26 0 0.5 1 0 0.5 1 0 0.5 1 (i) Ss Sr 0 0.2 0.4 <(?nc) (j) Ss Sr 0 0.2 0.4<(?nc) (k) Ss Sr 0 0.2 0.4 <(?nc) -3 -2 -1 00 0 0.5 1 ?nc (g) H2=1.86 H2=1.26 -3 -2 -1 00 0 0.5 1 ?nc (h) (l) 0.45 0.48 0.89 1.07 Ss Sr 0 0.5 1 jd?nc=d0j H2=1.86 H2=1.26 0 = -0.5 0 = -1.0 0 = -3.0 (a) (b) (c) (d) (e) (f) Sr Ss 0=-0.5 0=-1.0 0=-3.0 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-pdf

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-26T02:00:01.498150+00:00
License: CC-BY-NC-ND-4.0