Sequence Complexity Dictates Polymer Mixing
Dibyajyoti Mohanta,1 Manish Dwivedi, 2 Hiranmay Maity, 3 and Debaprasad Giri 4
1)Department of Chemistry, University at Buffalo, Buffalo, NY, USA a)
2)Department of Physics, Institute of Science, Banaras Hindu University, Varanasi,
India
3)Department of General Science, BITS, Pilani, Dubai Campus, Dubai International Academic City, Dubai,
UAE
4)Department of Physics, Indian Institute of Technology (BHU), Varanasi, India
Polymer association in confined spaces governs diverse phenomena from protein aggregation to DNA conden-
sation. We investigate the sequence-level mechanisms underlying this behavior through exact enumeration
of two confined AB-type block copolymers on a two-dimensional lattice, revealing how sequence complex-
ity controls mixing against de-mixing. We reveal that sequence complexity (heterogeneity), quantified by
Shannon entropy H2, acts as a determinant between self-folded low-mixed state and high-mixed state. High-
complexity sequences (H 2 = 1 .86 bits) with short-length repeats achieve near-complete inter-chain overlap
through cooperative chain collapse, while low-complexity blocky sequences ( H2 = 1.26 bits) maintain ex-
tended conformations with less overlapping. Free energy analysis reveals a steep increase in mixing barriers
for low-complexity compared to surmountable barrier for high-complexity sequences. We find that geometric
confinement modulates but does not override these sequence-dependent behaviors, with asymmetric confine-
ment enhancing heterogeneity in overlapping for both sequence type (pronounced in high-H2 sequences). Our
theory could be applicable in assessing the phase behavior of repeat protein and nucleic acid sequences.
I. INTRODUCTION
Liquid-Liquid Phase Separation (LLPS) mediated by
weak interactions provided the framework for compart-
mentalization of biomolecules, including protein and
RNA, even without any membrane 1. The formation and
properties of such membraneless compartments known
as biocondensates depend not only on the overall com-
position of constituent proteins and nucleic acids but
are also sensitive to the underlying sequence pattern-
ing of interaction motifs 2,3. Such sequence-encoded ef-
fects exhibit both in single-molecule and multicompo-
nent systems, where sequence composition and pattern-
ing shape the conformational landscape of biopolymers 4.
For example, intrinsically disordered proteins (IDPs) and
regions (IDRs), which lack well-defined tertiary struc-
tures, adopt conformational ensembles directly encoded
by their sequences 5,6. Similarly, nucleic acid sequences
also control the formation of secondary structures such as
hairpin loops, bulges in RNA, bubbles in DNA, through
the specific arrangement of nucleotides 4,7. Polyam-
pholytes with well-mixed positive and negative charges
readily collapse, while polyelectrolytes like poly(A)-RNA
or poly(U)-RNA resist condensation under physiological
conditions7,8. This sequence dependent organization and
association of biomolecules serve as a fundamental deter-
minant of phase separation behavior 9.
In vivo, sequence effects never act in isolation as
biomolecules fold and function within crowded, geomet-
rically confined environments that fundamentally alter
their behavior 10,11. The cellular environment contains
20 −40% macromolecules by volume, creating steric con-
a)Electronic mail:
[email protected]
straints that restrict conformational entropy and mod-
ulate interaction energetics 12,13. Recent studies have
demonstrated that confinement can reduce critical con-
centrations for phase separation by 4-10 fold, making
LLPS thermodynamically accessible at physiological pro-
tein concentrations 14,15. Confinement creates a deli-
cate balance: it restricts the conformational entropy of
RNA and proteins, while at the same time enhancing
local concentration effects that strengthen intermolecu-
lar interactions16,17. Understanding how geometric con-
finement couples with sequence patterning to determine
phase separation driven multi-polymer association is es-
sential.
Biomolecule aggregation often features hierarchy of as-
sembly that begins with monomers, dimers, oligomers,
which subsequently grow into network-like topolo-
gies characteristic of condensates beyond the critical
density4,18–20. In this context, we leverage insights from
previous theoretical and simulation studies on confined
two-polymer systems (dimers) to identify the key factors
that control their folding, mixing, and segregation 21–28.
However, the majority of these studies have focused
on homopolymers, which are largely dominated by en-
tropic effects 24,26 . By design, these models neglect the
sequence-specific enthalpic interactions that drive the
microphase separation essential in facilitating polymer
association29. Thus, a fundamental gap remains in our
understanding of how sequence patterning and geomet-
ric confinement jointly regulate the initial association of
polymers. To remedy this, we present an exactly solvable
minimal model that captures interplay between sequence
complexity, geometric confinement, and polymer associ-
ation, through exhaustive enumeration.
A central challenge in predicting sequence-dependent
polymer interactions is the quantitative characteriza-
tion of sequence patterning. Shannon entropy, calcu-
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint
2
FIG. 1. Schematic diagram of the two block co-polymer
(BCP) model. (Top) Examples of balanced and unbalanced
BCP sequences with A-type marked as blue and B-type as
yellow. The sequence complexity (H 2) is shown for each pair.
(Bottom) The two BCPs are confined within following geome-
tries: a square box (S s) of 10 × 10 and a rectangular box (S r)
of 20 × 5 lattice units. The shaded middle layer indicates
the set of possible starting positions for the polymers. In Ss
, chains are enumerated from one of the two shaded layers,
as 5 th and 6 th layers (Lys) are mirror images to each other.
Pink diamonds represent an example ”origin-pair” for the two
polymers (labeled 1 and 2). Strength of interaction between
same type of beads is taken as ϵ, whereas non-specific inter-
action strength is zero.
lated from the probability distribution of n-letter words
(n-grams), provides a robust measure of sequence pat-
terning defined as sequence complexity. It correlates
strongly with conformational heterogeneity and phase
separation propensity 19,30–32. To this end, we em-
ploy a system of two AB-type block copolymers where
sequence complexity is quantified by the bi-gram en-
tropy, H2 . This approach is motivated by recent stud-
ies on block copolymer systems composed of two dis-
tinct bead types, investigating sequence-controlled self-
assembly and aggregation 29,33.
This study unfolds in three key parts. First, we map
out the role of sequence complexity and confinement
anisotropy in shaping mixing of two polymers. Next,
we in-depth analyze the free energy landscape of mixing
for different H2 sequences within varying confinement.
Finally, we explore how sequence complexity affects the
actual size of two-polymer systems during the overlap-
ping pathway. We conclude our study by investigating
the mixing behavior for a diverse pair of sequences.
II. MODEL AND METHOD
We model a system of two block co-polymers (BCPs),
each consisting of N + 1 = 18 beads, as self(mutual)-
attracting-self(mutual)-avoiding walks (S(M)AS(M)A W)
within a symmetric confinement Ss with side length
Lxs = Lys = 10 lattice units, and an asymmetric con-
finement Sr with Lxr = 20 and Lyr = 5 lattice units.
Both confinements have an equal area of 100 square lat-
tice units, maintaining a constant volume fraction of
(2 × 18)/100 = 0.36 [Fig. 1]. We define a specific, at-
tractive interaction ϵsp = ϵ between non-bonded nearest-
neighbor beads of same type ( A − A or B − B). Inter-
actions between different beads type ( A − B) termed as
non-specific contacts and are set to zero (shown as × in
Fig. 1).
A. Sequence Entropy H2 of Block copolymers
The n-gram entropy measures the randomness of over-
lapping subsequences (‘words’) in a polymer chain, where
beads are treated as letters (e.g., A/B for a diblock
copolymer or hydrophobic/polar in HP models). Follow-
ing Herzel34, the bi-gram (2-letter word) entropy H2 for
a given 18-letter sequence of A and B is computed from
the frequencies of all overlapping 2-letter words, yielding
17 such bi-grams per sequence.
H2 = −
X
w∈{AA,AB,BA,BB}
p(w) log2 p(w) (1)
where the probability of each bi-gram w in the 18-
bead sequence is estimated as p(w) = n(w)/17, with
n(w) being the count of that bi-gram. All four bi-grams
are equal-a-priori, but their empirical probabilities p(w)
reflect the actual bi-gram composition of the specific
sequence35.
Label Sequence (length = 18) H 2 (in bits)
Balanced AAABBBAAABBBAAABBB 1.86
Balanced AAAAAAAAABBBBBBBBB 1.26
Unbalanced AABBAABBAABBAABBAA 1.99
Unbalanced AAAAAABBBBBBAAAAAA 1.45
TABLE I. Bi-gram entropy H2 (in bits) for four designed
18-bead polymer sequences composed of A and B beads. “Bal-
anced” sequences have equal numbers of A and B beads,
whereas “Unbalanced” sequences have unequal A/B compo-
sition. Larger H2 indicates greater sequence complexity.
Our study focuses on four distinct two-BCP systems
characterized by their sequence complexity H2 and over-
all A/B composition (Table I). Each system consists of a
sequence listed in Table I together with its complemen-
tary partner obtained by exchanging A and B at every
position. For example, the system H2 = 1.86 consists of
the AAABBBAAABBBAAABBB sequence and its com-
plementary partner BBBAAABBBAAABBBAAA with
A/B = 1 in individual sequence and overall two-BCP
system.
In balanced systems ( H2 = 1.86 and H2 = 1.26), each
chain has equal numbers of A and B beads, while in un-
balanced systems (H2 = 1.99 and H2 = 1.45), the chains
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint
3
have unequal A/B composition. Although many distinct
sequences share similar H2 values, this work focuses on
four representative systems chosen to represent the high
and low ends of the complexity spectrum, from short pe-
riodic patterns to more blocky ones, with H2 values rang-
ing approximately from 1.99 to 1.26 bits.
B. Enumeration procedure
The exact partition function of this binary BCP system
can be calculated by exact enumeration of all conforma-
tions of two polymer chains of length N = 17 (i.e., N + 1
beads)36. At this fixed chain length and volume frac-
tion (ϕc = 0.36), the full conformational space is already
large, so extending to longer chains would be computa-
tionally expensive26,36.
To obtain reliable statistics at fixedN without increas-
ing chain length, we systematically scan the accessible
volume by varying the positions of the polymer origins in-
side the confinement. Specifically, we restrict the starting
positions (origins) of both polymers to the middle layer
of the box, as indicated by the shaded line in [Fig. 1].
For each possible unique pair of origin positions on this
line (an “origin-pair”), we enumerate all conformations
of the two BCPs; providing an exhaustive conformational
ensemble for each origin-pair.
This procedure yields 10 × 9 = 90 distinct origin-pairs
for the Ss box and 20 × 19 = 380 origin-pairs for the
Sr box. Throughout the study, observables are reported
both as averages over individual origin-pairs and as en-
semble averages over all origin-pairs, thereby ensuring a
uniform scan of the confinement’s effect while remaining
computationally feasible.
The system is evaluated in the canonical ensemble. For
a origin-pair, the partition function Z is computed as a
sum over all possible conformations:
Z =
X
msp,mnsp,ssp,snsp
C(msp, mnsp, ssp, snsp) (exp(−βϵ))msp (exp(−βϵ))ssp (2)
Here, C(msp, mnsp, ssp, snsp) denotes the number of
conformations with:
msp: inter-polymer specific contacts (A–A or B–B be-
tween polymers)
mnsp: inter-polymer non-specific contacts (A–B between
polymers)
ssp: intra-polymer specific contacts (A–A or B–B within
each polymer, summed)
snsp: intra-polymer non-specific contacts (A–B within
each polymer, summed)
In this model (see [Fig. 1]), only specific contacts
are energetically weighted (i.e. ϵnsp = 0). So, the sys-
tem energy depends solely on the total number of spe-
cific contacts (m sp + ssp) (Eq. 2). On the other hand,
polymer overlap or mixing involves contributions from
both specific ( msp) and non-specific ( mnsp) contacts.
β = 1/(kBT ) is the inverse temperature. For this study,
we used reduced units where the Boltzmann constant kB
is set to 1 and the temperature T is fixed at 0 .5. To cal-
culate the overall ensemble average for a given sequence
type, we followed Eq. 2 across all origin pairs.
III. RESULTS
1. Statistical Mechanics of Contact Distribution
Multiple polymer assemblies such as condensates
formed by LLPS, chromatin organization, and crowded
macromolecular complexes, exhibit network-like archi-
tectures that typically rely on strength of multivalent
inter-chain interactions4,37. In [Fig.2], we analyze how se-
quence complexity and confinement jointly control BCP
mixing by examining the fraction of inter-chain contacts
per polymer, ϕnc = Nc/(N+1) with Nc = msp+mnsp, for
high-complexity (H 2 = 1.86, short-length repeats) and
low-complexity (H2 = 1.26, blocky) balanced sequences
under symmetric (S s, panels a,b,c) and asymmetric ( Sr,
panels d,e,f) confinement. Each radially outward bar in
the polar plots (Fig.2) represents ϕnc for a distinct origin
pair of the two BCPs within the confinement.
At weak interaction strength ϵ = −0.5 [Fig.2 (a),
(d)] , both sequences exhibit a modest and compara-
ble overlap ( ϕnc < 0.5) with H2 = 1.26 being slightly
higher. This indicates that conformational entropy dom-
inates over low specific interaction energy which is in-
sufficient to elicit sequence-dependent differences 38. As
attraction strength increases to ϵ = −1.0, the distribu-
tions bifurcate [Fig. 2(b,e)]. Polymers with higher H2
display a much broader spread of ϕnc values and access
configurations with substantially higher overlap. Under
strong attractive condition ϵ = −3.0 [Fig. 2(c,f)], the
distinction is maximized. H2 = 1.86 sequences approach
near-complete overlap ( ϕnc ∼ 1.1–1.2), whereas blocky
H2 = 1.26 sequences saturate at significantly lower ϕnc
values (≈ 0.6), in both confinements 30,33.
Ensemble-averaged ϕnc over all origin-pairs, reinforce
these trends [Fig. 2(g,h)]. At low |ϵ|, H2 = 1 .26 ex-
hibit mild higher ϕnc than H2 = 1.86 sequences, reflect-
ing their greater sensitivity to weak attractions. How-
ever, as |ϵ| increases, high-H2 sequences overtake low-H2
sequences in both confinements, achieving much higher
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint
4
0
0.5
1
00.51
0
0.5
1
H2=1.86
H2=1.26
0
0.5
1
0
0.5
1
0
0.5
1
(i)
Ss Sr
0
0.2
0.4<(?nc)
(j)
Ss Sr
0
0.2
0.4<(?nc)
(k)
Ss Sr
0
0.2
0.4<(?nc)
-3 -2 -1 000
0.5
1
?nc
(g)
H2=1.86
H2=1.26
-3 -2 -1 000
0.5
1
?nc
(h)
(l)
0.450.48
0.891.07
Ss Sr
0
0.5
1jd?nc=d0j
H2=1.86
H2=1.26
0 = -0.5 0 = -1.0 0 = -3.0
(a) (b) (c)
(d) (e) (f)
Sr
Ss
0=-0.5 0=-1.0 0=-3.0
FIG. 2. Sequence-dependent inter-chain overlapping across origin pairs for balanced sequences. Polar plots of
fraction of overlap contacts ( ϕnc = Nc/(N + 1)) for all origin-pairs at interaction strengths ϵ = −0.5, −1.0, −3.0 (a,b,c) for
high-complexity (H2 = 1.86), and (d,e,f) for low-complexity (H 2 = 1.26) sequences in square confinement Ss and rectangular
confinement Sr . Compared to Ss, Sr shows more heterogeneity in ϕnc, especially in Sr and at stronger ϵ. Ensemble-averaged
⟨ϕnc⟩ over origin pairs versus interaction strengthϵ in Ss (g) and Sr (h), showing higher mixing for high-H2 at strong interactions
for both confinements. Heterogeneity in mixing, quantified as standard deviation of ϕnc across origin-pairs in Ss and Sr (i, j,k
for different values of ϵ), highlighting enhanced variability in high-H2 and Sr. Initial mixing rates dϕnc/d|ϵ| in the linear regime
(l), revealing faster but shallower response for low-H 2, while slow but progressive mixing for high– H2 (pronounced in Sr).
saturation overlap. The contiguous arrangement of iden-
tical beads in low H2 blocky chains make intra-chain fold-
ing entropically favorable, allowing the BCPs to satisfy
most of their valency via compact self-collapsed states
(see [Fig. S1] in SI) 39–41. In contrast, the high com-
plexity H2 sequences make intra-chain folding entropi-
cally costly due to short-repeat lengths, leaving attrac-
tive beads exposed to form the multivalent inter-chain
mixing19,30,42.
To quantify how confinement anisotropy modulates
the heterogeneity in fraction of inter-contact distribu-
tion, we computed the standard deviation of ϕnc across
all origin-pairs at different strengths [Fig. 2(i,j,k)]. At
all interaction strengths, both type of sequences within
Sr exhibits higher heterogeneity than Ss, reflecting the
broader spectrum of accessible conformations and over-
lap states in anisotropic confinement 22,43. Moreover,
among the sequences, high-complexity sequences consis-
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint
5
tently show larger heterogeneity than their low-H2 coun-
terparts in both geometries, indicating richer landscape
of partially and highly mixed configurations for low and
high complexity sequences respectively.
The linear fits to ϕnc in the linear regime (low |ϵ|) re-
veal distinct “mixing rates” for the two sequence types
[Fig. 2(l)]. Low-H 2 sequences display a steeper initial
slope, presenting a faster but shallow response to in-
creasing attraction that quickly saturates at low overlap.
In contrast, high-H 2 sequences exhibit a more moder-
ate initial increase but ultimately progress to complete
overlap, with effective mixing rates |dϕnc/dϵ| ≈ 0.45 in
Ss and ≈ 0.89 in Sr. Thus, asymmetric confinement
nearly doubles the mixing rate for high-complexity se-
quences by pre-aligning chains and reducing the entropic
penalty of association 44. This is consistent with prior
studies showing that confinement can shift phase bound-
aries and reduce the critical concentration required for
LLPS 14,15,44,45.
To explore the molecular origins of sequence
complexity-dependent overlapping behavior, we present
the fraction of total specific contacts ϕsp = (m sp +
ssp)/(N + 1) in [Fig. 3 (a)]. A key finding is that low-
H2 = 1.26 sequences consistently form a higher fraction
of ϕsp compared to their high-H 2 counterparts across
both confinements . This observation appears counterin-
tuitive at first, as one might expect that a greater num-
ber of specific contacts would promote enhanced mix-
ing for H2 = 1.26. However, this can be explained by
decomposing total specific contacts into their inter- and
intra-chain components expressed asϕmsp = msp/(N+1)
and ϕssp = ssp/(N + 1) in [Fig. 3 (b) and (c)] respec-
tively. Blocky, low-H2 sequences sequences have contigu-
ous blocks of identical beads (more homotypic contacts
per segment), making intra-chain folding energetically fa-
vorable (with higher ϕssp values for H2 = 1.26 in [Fig. 3
(c)]. This reduces inter-chain overlap driven by specific
contacts defined by ϕmsp [Fig. 3 (b)] for H2 = 1.26. In
sharp contrast, short-length repeats in high- H2 chains
maintain more open conformations because intra-chain
folding incurs a large entropic penalty when attempting
to bring distant non-contiguous attractive beads into con-
tact 19,33. This leaves many attractive sites exposed and
available to form inter-chain-specific contacts featuring
higher ϕmsp for H2 = 1 .86 in [Fig. 3 (b)]. This high-
lights how low-H 2 sequences are energetically favored
(via more total contacts) but entropically disfavored (due
to self-folding) for inter-chain mixing, whereas high-H 2
sequences trade some energetic gains for entropic advan-
tages through enhanced inter-chain association [Fig. 3
(b)]41. These trends align with studies demonstrating
that strong homotypic interactions in proteins promote
self-association over intermolecular mixing 42.
The fraction of inter-chain non-specific contacts
(ϕmnsp = mnsp/(N + 1)), which do not contribute to
the Boltzmann weight but still takes part in mixing, is
low overall but also remains higher for short-range se-
quences ( H2 = 1.86) in both confinements [Fig. 3 (d)].
This arises because, in high-H 2 chains, non-interacting
blocks interspersed between favorable interaction blocks
can remain overlapped between two chains due to the
reduced entropic cost of constraining these short seg-
ments 40. This behavior mirrors observations in sticker-
spacer models of LLPS, where short spacers tend to re-
main in proximity due to the strong attachments of adja-
cent stickers 19,40. Finally, intra-chain non-specific con-
tacts ( ϕsnsp = snsp/(N + 1)) provide minimal contri-
butions to folding, thereby emphasizing the dominance
of energy-driven specific interactions [Fig. 3(e)]. Collec-
tively, Figs. 2 and 3 demonstrate that the spatial ar-
rangement of interaction sites, rather than their total
number or binding strength, governs whether polymers
assemble into self-collapsed or form cooperative polymer
organization30,33.
2. Free Energy Profile Analysis
To determine the thermodynamic driving forces gov-
erning sequence-dependent mixing, we analyzed the free
energy profiles F (Nc), calculated from partition function
with Nc inter-chain contacts Z(Nc) [Figure 4]:
F (Nc) = −kBT ln Z(Nc) (3)
where Z(Nc) is expressed as 22,
Z(Nc = msp + mnsp) =
X
ssp,snsp
C(msp, mnsp, ssp, snsp) exp(−βϵ(msp + ssp)) (4)
At both weak [Fig. 4 (a)] and strong [Fig. 4 (b)] in-
teraction strengths, high-H 2 exhibit higher free energy
in the low to moderate contact regime (N c < 20) com-
pared to low-complexity sequences. This initial penalty
arises because the H2 = 1.86 chains must overcome a
translational and orientational entropic penalty to align
due to their short repeats, which is not the case for low-
H2 blocky chains 46,47. However, when sufficient mix-
ing (N c > 20) has taken place in high-H 2 sequences,
the network of inter-chain contacts becomes sufficiently
connected (enthalpy dominated) to overcome the en-
tropic cost of chain alignment. So, the threshold which
arises near Nc ≈ 20 indicates the surmountable free en-
ergy barrier, beyond which the free energy landscape in-
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint
6
-3 -2 -1 000
0.2
0.4
0.6
0.8
1
1.2?sp
-3 -2 -1 000
0.2
0.4
0.6
0.8
1
1.2?mspH2(Ss)=1.86
H2(Ss)=1.26
H2(Sr)=1.86
H2(Sr)=1.26
h H2(Ss)i=1.86
h H2(Ss)i=1.26
h H2(Sr)i=1.86
h H2(Sr)i=1.26
-3 -2 -1 000
0.2
0.4
0.6
0.8
1
1.2?ssp
-3 -2 -1 000
0.2
0.4
0.6
0.8
1
1.2?mnsp
-3 -2 -1 000
0.2
0.4
0.6
0.8
1
1.2?snsp
(b) (c) (e)(d)
(a)
FIG. 3. Decomposition of contact fractions reveals sequence-dependent preferences for self-association versus
mixing (a) For both confinements total phase-separated (i.e. ϕsp = ϕmsp + ϕssp) contact fraction are higher for H2 = 1.26
than H2 = 1.86, representing H2 = 1.26 energetically dominated. (b) Inter-chain specific contact fraction ϕmsp are higher
for high-H 2, promoting sequence-specific mixing. (c) Intra-chain specific contact fraction ϕssp dominate for low-H 2, favoring
self-association. (d) Inter non-specific contact fraction ϕmnsp (A-B) are higher for high- H2, contributing to overall overlap. (e)
Intra-chain non-specific contact fraction ϕsnsp remain minimal for both sequences.
verts (F H2=1.86 < F H2=1.26), indicating that the highly
mixed network becomes the thermodynamic ground state
(global minima) for high-H 248. The free energy barrier
(yellow bar) and the low-mixed and high-mixed (global)
minima marked as green bars are shown in [Fig. 4 (b)].
This mixing threshold at (Nc ≈ 20) basically corresponds
to the valley region in density of stateD(Nc) between the
double-peaks [Fig.S2 (c) and (g) in SI]. In contrast, low-
H2 sequences cross their free energy minima at Nc ≈ 15
and beyond higher mixing incurs elevated free energy
(unfavorable), indicating their unimodal nature in D(Nc)
[Fig. S2 (b),(f) (d) and (h) in SI].
To quantify the resistance to mixing, we computed the
local gradient of the free energy landscape, ∂F/∂N c at
Nc [Fig. 4 (c)-(f)]. A steep positive gradient indicates
a strong thermodynamic penalty against mixing, while a
flat or negative gradient implies accessible transitions 49.
In both confinements, the local gradient is negative at
low Nc due to the fact that the confinement itself gives
stability to low mixed states 44. There is a crossover of
∂F/∂N c from negative to positive values at Nc ≈ 12 −13
indicating beyond this point overlapping or mixing re-
quires overcoming an increasing thermodynamic penalty.
At low strength ϵ = −1.0, H2 = 1.26 sequence’s ∂F/∂N c
lead over high-H2 sequences in both Ss and Sr at high Nc
regime indicating their high resistance in mixing. At high
strength ϵ = −3.0, [Fig. 4 (e),(f)], after crossing zero in
both Ss and Sr, high H2 = 1 .86 sequences constantly
feature low gradient values until a spike in ∂F/∂N c ≈ 20
(marked by vertical down arrow), indicating the free en-
ergy barrier or mixing threshold separating the minimas
in [Fig. 4 (b)]. Beyond this peak local gradient, we
see ∂F/∂N c going near-zero or negative again (especially
in Sr) showing the propensity of mixing of high- H2 se-
quences. The effect of confinement is also apparent since
the local peak gradient in Ss is ∂F/∂N c|Nc=20 is 2.15
kBT while for Sr, it is 1.55 kBT making Sr the preferred
environment for mixing. While low- H2 sequences show
maximum local gradient at higher Nc revealing the mix-
ing is highly unfavorable. The negative ∂F/∂N c values
on both sides of the spike for high- H2 sequences indicate
a double-well free energy landscape. In such landscapes,
the system can undergo spontaneous transitions between
basins (mixed to demixed) via thermal fluctuations 50,51.
Song et al.52 reported such results as tau-RNA complexes
reversibly cross LLPS boundaries, with sequence charge
patterns tuning the effective barrier between dilute and
dense phases.
We further quantified the thermodynamic preference
for the mixed state by calculating the global free energy
difference ∆F mixing between the mixed (N c > N ∗
c ) and
less-mixed phases [Fig. 4 (g)]. The probability of the
mixed phase is given by the normalized partition sum:
PN ∗
c =
P
Nc>N ∗
c
Ω(Nc)e−βϵ(msp+ssp)
Z (5)
where N ∗
c = 20 is the identified nucleation threshold [Fig.
S2, SI] and used for both low and high H2 sequences .
The relative stability is then:
∆Fmixing = −kBT ln
PN ∗
c
1 − PN ∗
c
(6)
∆Fmixing for H2 = 1.26 remains positive (disfavoring
mixing) even at high |ϵ| before saturating just below 0.
However, the ∆Fmixing value for the H2 = 1.86 sequences
drops gradually and reaching ∼ −3 for Ss and ∼ −4 for
Sr at ϵ = −3, confirming a cooperative transition to the
mixed phase 43.
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint
7
0 5 10 15 20 25Nc
-30
-28
-26
-24
-22
-20
-18
-16F(Nc)
0 5 10 15 20 25Nc
-70
-65
-60
-55
-50
-45
-40F(Nc)
-3 -2 -1 00-4
-3
-2
-1
0
1
2
3"F
H2=1.86[Ss]
H2=1.26[Ss]
H2=1.86[Sr]
H2=1.26[Sr]
-1
0
1dF(Nc)/dNc
0 10 20Nc
-1
0
1dF(Nc)/dNc
Ss
-2
0
2
4
Sr
0 10 20Nc
-2
0
2
H2=1.26
H2=1.86
(a)
0=-1.0 0=-3.0
(d)
(e)(c) (g)
(f)
0=-3.00=-1.0
(b)
barrierminima
FIG. 4. Free energy landscapes and barriers reveal the thermodynamic basis for sequence-dependent mixing :
(a, b) Free energy profiles F (Nc) as a function of inter-chain contacts Nc at weak (ϵ = −1.0) and strong ( ϵ = −3.0) interaction
strengths, respectively, for high-complexity ( H2 = 1.86, blue) and low-complexity ( H2 = 1.26, red) sequences in square ( Ss)
and rectangular ( Sr,triangles) confinements. High- H2 sequences exhibit an initial entropic penalty at low Nc but a favorable
crossover to lower free energy at high Nc, indicative of a surmountable barrier leading to stable mixed states. The barrier is
marked as yellow bar separating the two minimas marked as green bars in (b). Whereas, low- H2 sequences show a monotonic
increase, disfavoring mixing. (c,d,e,f) Local gradients ∂F/∂N c (in units of kBT ) highlighting barriers and driving forces: (c,
e) at ϵ = −1.0 for Ss and Sr; (d, f) at ϵ = −3.0 for Ss and Sr. High- H2 sequences display a characteristic spike around
Nc ≈ 20 (downward arrows), marking the barrier between demixed and mixed minima, with negative gradients post-barrier
favoring mixing more pronounced in asymmetric Sr confinement. Low-H 2 sequences show persistently positive gradients,
resisting overlap. (g) Free energy difference ∆F mixing between mixed and demixed states as a function of interaction strength
ϵ, confirming cooperative transitions to mixing for high- H2 (decreasing ∆ F ) versus persistent demixing for low-H 2 (positive
∆F ) in both confinements
3. Effect of sequence complexity on polymer size
The joint probability distribution P (Rend, Nc), over
all origin pairs with Rend as the end-to-end distance of
polymer 1 within the binary polymer system, reveals
sequence-dependent coupling between chain compactness
and intermolecular contacts [Fig. 5]. At weak inter-
actions (ϵ = −1.0), both sequences in Ss exhibit low-
probability distributions (p ≈ 0.05 − 0.1) spanning mod-
erate Rend (4 − 10) and Nc (10 − 22), indicating limited
mixing [Fig. 5 a,b]. In Sr, high-H 2 show emerging het-
erogeneity with broader Rend (10 − 20), consistent with
anisotropy enhancing alignment opportunities as seen in
contact heterogeneity [Fig. 2] and broader Rend distri-
butions (SI [Fig. S1 e]) 43,45.
At strong interactions (ϵ = −3.0), high-H 2 sequences
display discrete conformational clusters, echoing double-
well free energy landscapes [Fig. 4 (b)] that enable coop-
erative transitions to mixed states. At higher strength,
H2 = 1.86 [Fig. 5 (e),(g)] promotes mixing as well as col-
lapse with higher P (Rend, Nc) than H2 = 1.26 [Fig. 5
(f),(h)]. For better identification of mixed-collapsed
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint
8
FIG. 5. Joint probability distributions P (Rend, Nc) reveal sequence and confinement dependent conformational
ensembles. Top: Distributions at ϵ = −1.0 (a-d) and Middle: at ϵ = −3.0 (e-h) with Ss (square; a,b,e,f) and Sr (rectangular;
c,d,g,h). Bottom: k-means clusters at ϵ = −3.0 (i-l), with centroids (circles) and colors denote clusters. High- H2 shows discrete
basins favoring low Rend/high Nc, indicative of co-collapse; low-H 2 remains extended compared to H2 = 1.86 with moderate
Nc.
phase population, we analyze P (Rend, Nc) by k-cluster
analysis, where we focus on high mixed cluster. In Ss
[Fig. 5 (i)], k-means clustering (k=2) identifies a highly
mixed cluster (41.22%) at low Rend (4.09 ± 0.8), indi-
cating inter-chain-driven co-collapse. In Sr [Fig. 5 g,k],
k=3 reveals greater heterogeneity, with the H2 = 1.86
sequences exhibit high-mixed ( Nc = 21 .43 ± 1.78) clus-
ter centroid at Rend = 3.78 ± 0.4, confirming anisotropy
amplifies sequence-encoded mixing without overriding it.
However, low-H2 sequences (H 2 = 1.26) remain largely
unimodal in Ss [Fig. 5 f,j]; centered at moderate Rend
(7 − 10) and Nc (10 − 15), resisting compaction due to
self-association. In Sr [Fig. 5 h,l] k=3, confinement in-
duces mild heterogeneity, but higher Nc cluster occurs at
extended Rend (11.59 ± 1.61, Nc = 16 .13± 1.00), sug-
gesting expansion to accommodate contacts which is op-
posite to high-H2 co-collapse33,40. This shows the mixing
pathway of high-H 2 sequences exhibit co-collapse state
while low-H2 remains in low-mix-expanded state.
4. Unbalanced Sequence Analysis
The extension to unbalanced sequences H2 = 1.99 and
H2 = 1.45 [Fig. 1], where the ratio of beads type in each
FIG. 6. Sequence complexity governs mixing across
compositions. Unbalanced sequences with higher H2 = 1.99
( blue) exhibit increased inter-chain overlapping compared to
low- H2 = 1.45 (orange) in both square ( Ss) and rectangular
(Sr) confinements. However, the absolute fraction of inter-
chain contacts is systematically reduced for unbalanced se-
quences relative to balanced high- H2 sequences.
polymer chain deviates from unity ( A/B ̸= 1), confirms
the robustness of complexity-dependent behavior across
distinct composition (stoichiometry)29. We compared en-
semble average of fraction of inter-chain contact forma-
tion /mixing between a high-complexity unbalanced se-
quence ( H2 = 1.99) with a low-complexity unbalanced
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint
9
FIG. 7. Shannon bi-gram entropy (H 2) of poly-XY
regions across eukaryotic proteomes (Human, Mouse,
Drosophila, Zebrafish, and Plant). All proteomes exhibit
clustered H2 values (mean ≈ 1.29 bits) with comparable in-
terquartile ranges. Natural polyXY sequences fall between
our designed low-complexity (H2 = 1.26) and high-complexity
(H2 = 1.86) variants, occupying an intermediate mixing
regime.
sequence (H2 = 1.45) illustrated in [Fig 6 (a) and (b)].
Importantly, while overall mixing levels are reduced for
unbalanced compared to balanced sequences, still high-
H2 = 1 .99 continues to favor mixing ϕ1.99
N c > ϕ 1.45
N c in
both confinements. This reduction arises because unbal-
anced sequences have a higher translational entropy than
balanced ones, thereby resulting less alignment driven
mixing53 . Our observation that ’unbalanced’ stoichiome-
tries (A ̸= B) reduce polymer mixing warrants careful
distinction from the ’magic-ratio effect’ recently estab-
lished for multicomponent condensates 53. In systems
driven by heterotypic interactions (e.g., sticker-spacer
models where A binds B), phase separation is often sup-
pressed at 1:1 stoichiometry because the formation of
small, entropy-favored monomers (self folding)53. In con-
trast, our system is governed by homotypic attractions
(A binds A, B binds B), where the 1:1 ’balanced’ com-
position promotes inter-chain mixing. This analysis rein-
forces that sequence specific interaction primarily governs
polymer associations, while overall sequence composition
modulates the extent of mixing but cannot override the
effects of patterning.
FIG. 8. Schematic free energy landscape linking se-
quence complexity to mixing pathways. Free energy
landscape schematic for block copolymers ( H2 = 1.26 and
H2 = 1.86) as a function of inter-chain contacts Nc at ϵ = −3.
The blue curve illustrates the landscape for H2 = 1.26 se-
quence, which exhibit a free energy minimum at low mixing
states favoring self folding over mixing due to the blockiness in
H2 = 1.26. The orange curve represents H2 = 1.86 variants,
where the free energy minimum shifts to higher Nc values in
the mixed state. The short length repeats in H2 = 1.86 sup-
presses intrachain self-folding and promotes intermixing with
the neighboring chain.
IV. DISCUSSION
This study establishes sequence complexity, quantified
by bigram Shannon entropy H2, as a molecular determi-
nant dictating the balance between intermolecular mix-
ing and de-mixing in confined AB-type block copolymers
(BCPs). Our systematic study employing exact enumer-
ation of two short BCP chains on a two-dimensional lat-
tice under distinct confinement geometries reveals four
critical aspects: i) high-complexity sequences ( H2 =
1.86) not only favor enhanced mixing but also introduce
greater heterogeneity in overlapping configurations, par-
ticularly under asymmetric confinement ( Sr). This het-
erogeneity in Sr arises from distributed short-length re-
peats, providing entropic flexibility that stabilizes diverse
high Nc states, as evidenced by bimodal D(Nc) distribu-
tions [Fig. S2 , SI] and multiple conformational clusters
in P (Rend, Nc) (Fig. 5). In contrast, low-complexity
blocky sequences (H 2 = 1.26, 1.45) exhibit unimodal,
low-mix ensembles, where self-folding via intra-chain con-
tacts yields energetic stability (higher phase-separated
contacts) but resists extensive inter-chain overlap [Fig.
S1] . ii) Free energy landscapes further reveal this tran-
sition as high-H 2 sequences feature surmountable bar-
riers ( 1-2 kBT ) separating low and high mix minima,
enabling cooperative transitions to near-complete over-
lap. Whereas, low- H2 landscapes show steep free energy
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint
10
barrier at elevated Nc. iii) The high-mixed high complex-
ity sequences undergo co-collapsed state (coupled coil-to-
globule transition) where as low H2 sequences remain ex-
panded and low-mixed across confinements. iv) ’High H2
favoring high mixing’ scenario holds also for unbalanced
compositions (A/B ̸= 1), emphasizing sequence pattern-
ing primacy over composition. Fig. 8 illustrates this
sequence-dependent behavior: low-complexity sequences
(H2 = 1 .26) favor self-folded conformations with min-
ima in low-mixed states, while high-complexity sequences
(H2 = 1 .86) undergo cooperative mixing with minima
shifted to higher Nc.
To assess whether the sequence-complexity regimes ex-
plored by our model are relevant to real proteins, we an-
alyzed the bi-gram entropy H2 of polyXY regions (two-
residue low-complexity tracts) across multiple proteomes
[Fig. 7] 54. Remarkably, the distribution of H2 values is
centered around 1.3 ± 0.3 bits, which overlaps precisely
with the range in which our model predicts a transition
between self-folded, weakly mixed states and inter-chain
mixed, co-collapsed states. This correspondence suggests
that many polyXY segments in vivo may be naturally
tuned to function near the boundary separating mixed
and de-mixed states, enabling cells to regulate conden-
sate assembly and disassembly in response to changing
conditions.
Although our enumeration-based framework simplifies
the vast complexity of single proteins and condensates
on a two dimensional lattice, its strength lies in its
exact calculation of the partition function, which pro-
vides a precise thermodynamic characterization of the
interplay between sequence patterning and confinement.
More broadly, the sequence–structure relationship un-
covered here reflects a fundamental organizing principle
that spans from single-chain folding to higher-order con-
densate assembly. By employing a minimal but exactly
solvable model, we demonstrated that the patterning of
interaction sites, not just their overall abundance, is a
critical determinant of polymer association.
1. Acknowledgement
DM thanks Erik Clarkson for reading the manuscript.
1S. F. Banani, H. O. Lee, A. A. Hyman, and M. K. Rosen, Nature
Reviews Molecular Cell Biology 18, 285–298 (2017).
2T. Kaur, M. Raju, I. Alshareedah, R. B. Davis, D. A. Po-
toyan, and P. R. Banerjee, Nature Communications 12 (2021),
10.1038/s41467-021-21089-4.
3T. de Vries, M. Novakovic, Y. Ni, I. Smok, C. Inghel-
ram, M. Bikaki, C. P. Sarnowski, Y. Han, L. Emmanouilidis,
G. Padroni, A. Leitner, and F. H.-T. Allain, Science Advances
10 (2024), 10.1126/sciadv.adm7435.
4H. T. Nguyen, N. Hori, and D. Thirumalai, Nat. Chem. 14, 775
(2022).
5R. K. Das and R. V. Pappu, Proceedings of the National
Academy of Sciences 110, 13392–13397 (2013).
6A. S. Holehouse and R. V. Pappu, Annual Review of Biophysics
47, 19–39 (2018).
7A. Jain and R. D. Vale, Nature 546, 243 (2017).
8L. Sawle and K. Ghosh, The Journal of Chemical Physics 143
(2015), 10.1063/1.4929391.
9A. A. Hyman, C. A. Weber, and F. J¨ ulicher, Annual Review of
Cell and Developmental Biology 30, 39–58 (2014).
10D. Lucent, V. Vishal, and V. S. Pande, Proceedings of the Na-
tional Academy of Sciences 104, 10430–10434 (2007).
11J. Mateos-Langerak, M. Bohn, W. de Leeuw, O. Giromus,
E. M. M. Manders, P. J. Verschure, M. H. G. Indemans, H. J.
Gierman, D. W. Heermann, R. van Driel, and S. Goetze, Pro-
ceedings of the National Academy of Sciences 106, 3812–3817
(2009).
12S. Mittal, R. K. Chowhan, and L. R. Singh, Biochimica et Bio-
physica Acta (BBA) - General Subjects 1850, 1822–1831 (2015).
13A. A. M. Andr´ e and E. Spruijt, International Journal of Molec-
ular Sciences 21, 5908 (2020).
14L. K. Davis, I. J. Ford, and B. W. Hoogenboom, eLife 11 (2022),
10.7554/elife.72627.
15B. Monterroso, W. Margolin, A. J. Boersma, G. Rivas, B. Pool-
man, and S. Zorrilla, Chemical Reviews 124, 1899–1949 (2024).
16J. Ahlers, E. M. Adams, V. Bader, S. Pezzotti, K. F. Winklhofer,
J. Tatzelt, and M. Havenith, Biophysical Journal120, 1266–1275
(2021).
17S. Mondal, K. Narayan, S. Botterbusch, I. Powers, J. Zheng,
H. P. James, R. Jin, and T. Baumgart, Nature Communications
13 (2022), 10.1038/s41467-022-32529-0.
18S. Jeon, Y. Jeon, J.-Y. Lim, Y. Kim, B. Cha, and W. Kim, Signal
Transduction and Targeted Therapy 10 (2025), 10.1038/s41392-
024-02070-1.
19J.-M. Choi, A. A. Hyman, and R. V. Pappu, Phys. Rev. E 102,
042403 (2020).
20J. D. Schmit, M. Feric, and M. Dundr, Trends in Biochemical
Sciences 46, 525–534 (2021).
21Y. Liu and B. Chakraborty, Physical Biology 9, 066005 (2012).
22J. M. Polson and D. A. Rehel, Soft Matter 17, 5792–5805 (2021).
23S. Jun, A. Arnold, and B.-Y. Ha, Phys. Rev. Lett. 98, 128303
(2007).
24E. Minina and A. Arnold, Soft Matter 10, 5836 (2014).
25Z. Liu, X. Capaldi, L. Zeng, Y. Zhang, R. Reyes-Lamothe, and
W. Reisner, Nature Communications 13 (2022), 10.1038/s41467-
022-31398-x.
26D. Mohanta, Soft Matter 19, 4991–5000 (2023).
27D. A. Rehel and J. M. Polson, Soft Matter 19, 1092–1108 (2023).
28L. Zeng, X. Capaldi, Z. Liu, and W. W. Reisner, Physical Review
E 109 (2024), 10.1103/physreve.109.024501.
29B. G. Weiner, A. G. T. Pyo, Y. Meir, and N. S. Wingreen, PLOS
Computational Biology 17, e1009748 (2021).
30R. Takaki and D. Thirumalai, Proceedings of the Na-
tional Academy of Sciences 121, e2409973121 (2024),
https://www.pnas.org/doi/pdf/10.1073/pnas.2409973121.
31S. Biswas and D. A. Potoyan, PLOS Computational Biology 21,
e1012826 (2025).
32Y.-f. Wang, C.-l. Ren, and Y.-q. Ma, Physical Review E 112
(2025), 10.1103/783t-5xqx.
33G. L. Dignon, W. Zheng, Y. C. Kim, R. B. Best, and J. Mittal,
PLOS Computational Biology 14, 1 (2018).
34H. Herzel, W. Ebeling, and A. O. Schmitt, Phys. Rev. E 50,
5061 (1994).
35L. Yu, D. K. Tanwar, E. D. S. Penha, Y. I. Wolf,
E. V. Koonin, and M. K. Basu, Proceedings of
the National Academy of Sciences 116, 3636 (2019),
https://www.pnas.org/doi/pdf/10.1073/pnas.1814684116.
36C. Vanderzande, in Lattice Models of Polymers (Cambridge Uni-
versity Press, Cambridge, 1998) pp. xi–xiv.
37D. M. Mitrea and R. W. Kriwacki, Cell Communication and Sig-
naling 14, 1 (2016).
38D. Mohanta, D. Giri, and S. Kumar, Journal of Statistical Me-
chanics: Theory and Experiment 2019, 043501 (2019).
39A. Sood and B. Zhang, Biophysical Journal 123, 1815–1826
(2024).
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint
11
40T. S. Harmon, A. S. Holehouse, M. K. Rosen, and R. V. Pappu,
eLife 6 (2017), 10.7554/elife.30294.
41J. A. Joseph, J. R. Espinosa, I. Sanchez-Burgos, A. Garaizar,
D. Frenkel, and R. Collepardo-Guevara, Biophysical Journal
120, 1219–1230 (2021).
42A. C. Murthy, G. L. Dignon, Y. Kan, G. H. Zerze, S. H. Parekh,
J. Mittal, and N. L. Fawzi, Nature Structural amp; Molecular
Biology 26, 637–648 (2019).
43D. Mohanta, Soft Matter 18, 2790–2799 (2022).
44M. P. Taylor, T. M. Prunty, and C. M.
O’Neil, The Journal of Chemical Physics 152,
094901 (2020), https://pubs.aip.org/aip/jcp/article-
pdf/doi/10.1063/1.5144818/15574842/0949011online.pd f.
45E. J. Saltzman and M. Muthukumar, The Journal of Chemical
Physics 131 (2009), 10.1063/1.3267487.
46A. M. Rumyantsev, A. Johner, and J. J. de Pablo, ACS Macro
Letters 10, 1048 (2021).
47F. Yu and S. Sukenik, The Journal of Physical Chemistry B 127,
4235 (2023), pMID: 37155239.
48E. W. Martin, T. S. Harmon, J. B. Hopkins, S. Chakravarthy,
J. J. Incicco, P. Schuck, A. Soranno, and T. Mittag, Nat. Com-
mun. 12, 4513 (2021).
49D. Wales, Energy Landscapes: Applications to Clusters,
Biomolecules and Glasses (Cambridge University Press, 2001).
50K. A. Dill and H. S. Chan, Nature Structural amp; Molecular
Biology 4, 10–19 (1997).
51R. P. Sear, Journal of Physics: Condensed Matter 19, 033101
(2007).
52Y. Lin, J. McCarty, J. N. Rauch, K. T. Delaney, K. S. Kosik,
G. H. Fredrickson, J.-E. Shea, and S. Han, eLife8, e42571 (2019).
53Y. Zhang, B. Xu, B. G. Weiner, Y. Meir, and N. S. Wingreen,
eLife 10 (2021), 10.7554/elife.62403.
54P. Mier and M. A. Andrade-Navarro, Computational and Struc-
tural Biotechnology Journal 20, 5516–5523 (2022).
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint
Low mixed state High mixed state
Nc
F(Nc)
High sequence complexity
Low sequence complexity
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint
0 5 10 15 20 25Nc
-30
-28
-26
-24
-22
-20
-18
-16F(Nc)
0 5 10 15 20 25Nc
-70
-65
-60
-55
-50
-45
-40
F(Nc)
-3 -2 -1 00
-4
-3
-2
-1
0
1
2
3"F
H2=1.86[Ss]
H2=1.26[Ss]
H2=1.86[Sr]
H2=1.26[Sr]
-1
0
1
dF(Nc)/dNc
0 10 20Nc
-1
0
1
dF(Nc)/dNc
Ss
-2
0
2
4
Sr
0 10 20Nc
-2
0
2
H2=1.26
H2=1.86
(a)
0=-1.0 0=-3.0
(d)
(e)(c)
(g)
(f)
0=-3.00=-1.0
(b)
barrier
minima
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint
-3 -2 -1 00
0
0.2
0.4
0.6
0.8
1
1.2?sp
-3 -2 -1 00
0
0.2
0.4
0.6
0.8
1
1.2
?mspH2(Ss)=1.86
H2(Ss)=1.26
H2(Sr)=1.86
H2(Sr)=1.26
h H2(Ss)i=1.86
h H2(Ss)i=1.26
h H2(Sr)i=1.86
h H2(Sr)i=1.26
-3 -2 -1 00
0
0.2
0.4
0.6
0.8
1
1.2
?ssp
-3 -2 -1 00
0
0.2
0.4
0.6
0.8
1
1.2
?mnsp
-3 -2 -1 00
0
0.2
0.4
0.6
0.8
1
1.2
?snsp
(b) (c) (e)(d)
(a)
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint
0
0.5
1
0
0.5
1
0
0.5
1
H2=1.86
H2=1.26
0
0.5
1
0
0.5
1
0
0.5
1
(i)
Ss Sr
0
0.2
0.4
<(?nc)
(j)
Ss Sr
0
0.2
0.4<(?nc)
(k)
Ss Sr
0
0.2
0.4
<(?nc)
-3 -2 -1 00
0
0.5
1
?nc
(g)
H2=1.86
H2=1.26
-3 -2 -1 00
0
0.5
1
?nc
(h)
(l)
0.45 0.48
0.89
1.07
Ss Sr
0
0.5
1
jd?nc=d0j
H2=1.86
H2=1.26
0 = -0.5 0 = -1.0 0 = -3.0
(a) (b) (c)
(d) (e) (f)
Sr
Ss
0=-0.5 0=-1.0 0=-3.0
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 25, 2025. ; https://doi.org/10.64898/2025.12.23.696237doi: bioRxiv preprint
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.