Geometrically encoded positioning of introns, intergenic segments, and exons in the human genome

preprint OA: closed CC-BY-NC-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Human tissues require a mechanism to generate durable, yet modifiable, transcriptional memories to sustain cell function across a lifetime. Previously, we demonstrated that nanoscale packing domains couple heterochromatin (cores) and euchromatin (outer zone) into unified reaction volumes that can generate transcriptional memory. In prior work, this framework demonstrated that RNA synthesis occurred within the ideal zone (intermediate density) portions of the domain. Naturally, this creates a question of where genes are positioned in relation to the packing domain architecture and which genetic material fills the domain core to sustain transcription. Here we propose that this could be solved by the encoded positioning of introns, intergenic segments, and exons as a projection of the functional packing layers of domains. This suggests that introns and intergenic segments are coupled to adjacent exons to generate coherent packing domain volumes. We illustrate how this organization would reconcile contradictions in epigenetic patterns, non-randomness in oncogenic mutations, and produce durable transcriptional memory. We conclude by showing that this genome geometry might have coincided with the rapid evolution of body-plan complexity, suggesting that chromatin geometry could be fundamental to metazoan evolution.
Full text 120,064 characters · extracted from oa-pdf · 9 sections · click to expand

Abstract

40 Human tissues require a mechanism to generate durable, yet modifiable, transcriptional 41 memories to sustain cell function across a lifetime. Previously, we demonstrated that nanoscale 42 packing domains couple heterochromatin (cores) and euchromatin (outer zone) into unified 43 reaction volumes that can generate transcriptional memory. In prior work, this framework 44 demonstrated that RNA synthesis occurred within the ideal zone (intermediate density) portions 45 of the domain. Naturally, this creates a question of where genes are positioned in relation to the 46 packing domain architecture and which genetic material fills the domain core to sustain 47 transcription. Here we propose that this could be solved by the encoded positioning of introns, 48 intergenic segments, and exons as a projection of the functional packing layers of domains. This 49 suggests that introns and intergenic segments are coupled to adjacent exons to generate 50 coherent packing domain volumes. We illustrate how this organization would reconcile 51 contradictions in epigenetic patterns, non-randomness in oncogenic mutations, and produce 52 durable transcriptional memory. We conclude by showing that this genome geometry might have 53 coincided with the rapid evolution of body-plan complexity, suggesting that chromatin geometry 54 could be fundamental to metazoan evolution. 55 56 Main Text 57 58

Introduction

59 60 Although most cells in a multicellular organism share the same genome, they differentiate into 61 hundreds of cell types. Even cells of the same lineage can perform different functions, shifting 62 over timescales from minutes to decades, while many disease states arise from altered gene 63 expression. A key underlying question is how genetic information gives rise to diverse and 64 dynamic cell phenotypes across development, aging, and disease. Many sequence-dependent 65 transcriptional regulatory elements are well studied, including enhancers, promoters, insulators, 66 silencers, etc [1–4]. These provide significant insight into transcriptional regulation. Paired with 67 imaging data, an emerging view is that chromatin organizes into packed structures throughout the 68 nucleus. We recently described the structure-function lifecycle of nanoscale packing domains to 69

Result

in a functional coupling between heterochromatin and euchromatin into a unified functional 70 volume[5,6]. The formation and function of chromatin packing domains intersects physical and 71 biological processes. This process is self-assembling, with transcriptional reactions having the 72 capacity to initiate domain formation. As a result, it introduces a mechanism for transcriptionally 73 mediated cellular memory encoded as a physical structure. An interesting physical feature of 74 these packing domains is that density exists as a continuous gradient. This gradient, when paired 75 with the size of enzymes (heterochromatin enzymes are small, ~2nm compared to 6nm for 76 euchromatin enzymes), generates a mechanism to geometrically guide enzymes to different 77 locations within the volume[6]. Interestingly, an intermediate zone is generated that appears to be 78 optimal for the positioning of RNA polymerase II (Pol-II) and transcription factors. Paired with 79 modeling and super resolution imaging, we showed that this efficient packing – heterochromatin 80 deposition into cores - acts to increase transcriptional output at distal regions [6]. This suggests 81 that the genome is folded in a manner that can couple heterochromatin and euchromatin together 82 into unified functional structures. 83 Applying this information to muscle differentiation, we observed that the activation of transcription 84 on exons on the genes Myh1 and Myh2 was associated with the deposition of heterochromatin 85 within non-exonic (NE) segments (introns, intergenic segments) of gene bodies [6]. Naturally, we 86 wondered if this observation provides a system for genes to pack into nanoscale structures where 87 NE elements act to produce the packing domain volumes that could position exons into the ideal 88 zones. Using myogenesis again as a model system, to our surprise, we observed that 58 of the 89 top 100 differentially activated genes in muscle differentiation are associated with decreased 90 accessibility in NE segments ( Figure 1a-c ). This behavior contrasts with accessibility at the 91 .CC-BY-NC 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 29, 2025. ; https://doi.org/10.1101/2025.05.29.656862doi: bioRxiv preprint 3 transcription start site, which is generally stable ( Figure 1b-c ). Given this finding, we then 92 examined whether heterochromatin resides in NE segments with adjacent RNA synthesis across 93 human tissues broadly by analyzing data available through ENCODE [7–9]. In multiple tissues, 94 critical functional genes contain heterochromatin within NE segments adjacent to actively 95 transcribed exons ( Figure 1d ). For example, H3K9me3 (a marker of constitutive 96 heterochromatin) is deposited within TNNI3 (cardiac troponin) in cardiac samples. This 97 biallelically expressed gene must be transcribed throughout the human lifespan to maintain heart 98 contractility (SI Figure 1)[10–13]. 99 We wish to propose a radical hypothesis – described within this work – that nanoscale packing 100 geometry is encoded in the positioning of exons, introns, and intergenic segments as projections 101 of the functional layers of domains observed on ChromSTEM tomography. This would pair NE 102 segments with adjacent exons as a linear projection of the 3-D volumes observed 103 experimentally[6,14–16]. For this hypothesis to be valid, it requires (1) a length-dependent pairing of 104 NE segments with exons to generate volumetric ratios, and (2) that they be oriented in the 105 direction of gene transcription. We demonstrate that both properties are observed. We then 106 provide experimental evidence from nascent-RNA sequencing, ChIP-Seq, ATAC-Seq, paired-end 107 tag sequencing (ChiaPET), and ChromSTEM tomography to support this hypothesis. While the 108 proposed geometric system produces a mechanism to generate transcriptional efficiency, 109 memory, and complexity, we show that it comes at increased mutation risk of the exposed branch 110 points between domains. By analyzing oncogenes, we show mutation frequencies across all 111 cancers correlating with genes containing branch point segments. Of particular interest is that this 112 geometric system may have emerged in parallel with increased body plan complexity of 113 metazoans, suggesting that geometry introduces a possible novel, non-mutagenic system to 114 increase information sampling based on physical properties. In sum, this theory generates new 115 avenues of research across diseases of aging, species evolution, and complex organ 116 development rooted in physical genomics. 117

Results

118 119 Chromatin domain geometry optimizes transcription in the human nucleus 120 All methods encounter limitations in measuring chromatin at the smallest length-scales due to the 121 problem of missing information. Many tools provide sequence-dependent chromatin interactions 122 including proximity capture methods (Hi-C, Sprite, etc) [17–19] and in situ hybridization (FISH) 123 imaging[20–22]. However, connectivity in a population may not equate to geometry in a single cell 124 (SI Figure 1 ). To complement connectivity measurements with nanoscale density, most studies 125 employ super-resolution imaging, accessibility assays, or variants of chromatin 126 immunoprecipitation[14,15,23–30]. For example, due to probe size, it appears that antibodies used for 127 super-resolution imaging of chromatin cannot penetrate high density domains in situ. This results 128 in high-density and low-density regions devoid of signal ( SI Figure 2a-d) [31,32]. A separate 129 problem is present for super-resolution imaging using FISH probes that rely on formamide 130 dehybridization. The denaturation process swells nanoscale domains, causing a loss of volume 131 and density information [33–35]. Due to these limitations of nanoscale methods, there is limited 132 knowledge of human gene geometry at the smallest length scales. While FISH has shown that a 133 large chain or melt phenotypes are present for a few very genes, the observed distances are 134 consistent with these genes being compressed in space (SI Materials and Methods)[36]. With the 135 advent of ChromEM/ChromSTEM imaging, it is now possible to study how chromatin transitions 136 from disordered beads-on-a-string into 3-D volumes as a general framework, but this too lacks 137 sequence-specific information [15,16]. To overcome this limitation, it is necessary to pair the 138 findings from this modality with other imaging modalities, molecular technologies, and modeling. 139 From this integrated approach, it was shown that chromatin domains observed on ChromSTEM 140 imaging appear to be self-assembling structures guided by transcription and cohesin to create 141 volumes. The geometry of these packing domains is a mass fractal: for large genes and loops, 142 .CC-BY-NC 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 29, 2025. ; https://doi.org/10.1101/2025.05.29.656862doi: bioRxiv preprint 4 the mass of a gene/loop (in nucleosomes) fills space as a function of radius by the dimension, D. 143 This is analytically described by the equation, /g18Y9/g185Y/g1871/g1871 /g4666 /g1866/g187Y/g1855/g1864/g1857 /g1867/g1871/g1867 /g1865/g1857 /g1871 /g4667 /gY404 /g1870 /g3005 ; a property confirmed on 144 ChromSTEM tomography[15,16,37,38]. 145 In many ways, it is conceptually easier to consider two limiting cases of mass fractals to 146 understand the structure of nanoscopic packing domains. A random walk without attractive 147 potentials or confinement has scaling D=2, which has a uniform distribution of density within a 148 volume (SI Figure 3a). An extension of this is a confined random walk, the fractal globule, which 149 will have a contact scaling of D=3 but also has a uniform distribution of density (SI Figure 3b)[38]. 150 These two extremes are similar in that density is unform – they resemble statistically 151 homogeneous structures that are similar to spheres with different levels of packing. It may be 152 tempting to assign chromatin as an open state to D=2 and closed state to D=3[38]. However, this 153 is not observed experimentally on ChromSTEM. Instead, chromatin throughout the nucleus 154 organizes into domains with D values ranging between 2.2 and 2.8 with a radius ranging between 155 25-175nm. This indicates non-uniform density: a gradient is present that radially decays from a 156 high-density interior to a low-density periphery [15,16]. The result is that packing at the nanoscale 157 couples high density (heterochromatin cores) with a lower density outer zone (euchromatin). An 158 intermediate density between these two regions appears to be an ideal zone for transcriptional 159 reactions to occur. Based on these zones, it was shown that disruption of heterochromatin 160 enzymes can experimentally decrease RNA synthesis due to the disruption of the ideal zone 161 supported by the core element. Furthermore, it appears that this system can self-assemble 162 through transcription-mediated loops first generating nascent (small, poorly packed) domains. 163 While these results indicate that transcription appears to happen at the intermediate region, it 164 remains to be understood where the geometric information for such a system is stored. 165 A possibility we consider is that large genes and loops are partially constrained (packed) into 166 domain volumes. We can consider muscle and myosin heavy chain 1 (Myh1) as a representative 167 example for comparison between models. Myh1 is ~26,000 basepairs, which translates into ~130 168 nucleosomes. A nucleosome is approximately a cylinder composed of ~200bp with diameter of 169 ~11nm and height of ~5.5nm. Stacking end-to-end to produce a fully stretched state produces a 170 ~715nm loop. For scale, chromosome 2 is composed of 243Mbp, contains ~1,200 genes, and in 171 mature muscle cells has a radius of ~1.5 μm (SI Figure 4) [39]. Since muscle function requires the 172 synthesis of hundreds of genes, many of which are quite large, accounting for space is crucial. 173 Outside of muscle, a similar problem arises with loops. Many loops are over 100kbp (which 174 translates into 5.5um), a length which would span the radius of the nucleus if not folded in space 175 (Figure 3a, SI Figure 5 )[40,41]. We propose a hypothesis that packing information is stored by 176 exons non-randomly pairing with NE in order to generate the functional layers observed on 177 ChromSTEM imaging (Figure 2b). 178 Human genes are a predictable, power-law geometric assembly of exons and introns 179 As a consistent convention across the different types of RNA products, we refer to DNA 180 transcribed and processed into functional RNA as ‘exons’ independent of the RNA product class 181 (mRNA, lncRNA, pseudogenes, etc). Similarly, we refer to infrequently transcribed (NE) DNA as 182 introns (within a gene body) and intergenic segments (between bodies) [42–44]. We recognize that 183 non-exonic DNA has many sequence-specific regulatory functions [45–47]. We intentionally do not 184 make sequence -specific measurements, as the focus is on packing geometry encoding novel 185 information. Instead, this investigation focuses on whether 3-D geometry, guided by transcription, 186 could project these elements as reaction volumes within the human genome. This geometric 187 theory is as follows. 188 The proposed volumetric domain functional organization results in an inverse relationship of the 189 ratio of exons-to-introns as a function of gene length. This is because the length of a segment 190 (gene/loop) folded into a domain translates into the radius as a function of D ( Figure 2d-f ) 191 [15,16,30,31,48–50]. The inverse relationship of a reaction zone to the domain volume mirrors that of the 192 .CC-BY-NC 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 29, 2025. ; https://doi.org/10.1101/2025.05.29.656862doi: bioRxiv preprint 5 surface-area-to-volume ratio observed in spheres (SI Figure 6), which decreases as a function of 193 the radius ( Figure 2b, SI Materials and Methods ). However, instead of the case of each layer 194 being a thin surface, it instead refers to content organized into zones within a total volume. In this 195 case, it indicates that genes contain almost enough NE DNA to fold into domains to position 196 exons for efficient transcription. In sharp contrast to this inverse relationship, chromatin 197 organizing into beads-on-a-string would result in a linear relationship between the exonic fraction 198 to the NE length ( SI Figure 6 ). This is because as the number of nucleosomes increases, the 199 amount of accessible DNA increases proportionally ( SI Figure 6). The null hypothesis is that no 200 volume-producing geometric relationship exists and elements within the genome are randomly 201 spaced. This would result in genes having a random amount and positioning of intronic DNA as a 202 function of length. Based on this, one would reasonably assume that exon length would be ~10% 203 of any gene body independent of the length. This is because only 10% of the gene length on 204 average is exonic. Instead, there is a broad distribution in the exon fraction within genes ( Figure 205 3a, interquartile range of 4.6% to 21.1%) that is length dependent, as we describe below. 206 Using the publicly available human reference genome GRCh38, with annotations from RefSeq, 207 we tested if the inverse relationship occurs across the human genome[51]. We calculated the ratio 208 of exons to introns vs total length for all protein-coding genes, finding that most human genes are 209 described by this geometric principle ( Figure 3b, Materials and Methods ). This relationship is 210 defined as follows by Equation (1) where p is a scalar in basepairs: 211 1) /g3006/g3051/g3042/g3041 /g3010/g3041/g3047/g3045/g3042/g3041 /g1542 /g3043 /g3013/g3032/g3041/g3034/g3047/g3035 /g3289 212 213 We observe most genes are well described by a p range from 500bp to 25,000bp, which is 214 consistent with ~2 to 125 nucleosomes when n is close to 1. This is observed both for individual 215 isoforms or all the variant isoforms within the RefSeq database ( SI Figure 7 ). An alternative 216 explanation for such a trend is that as gene length increases, the exonic fraction decreases. If this 217 were the case, this pattern would be observed even if exons are randomly redistributed. 218 However, randomization produces a length-independent constant ratio of E/I of ~0.1, indicating 219 that gene length is non-randomly associated with exon content (Figure 3c). However, to account 220 for the possibility that the observed ratio decreases due to length alone, we tested if intron length 221 is geometrically a power-law of exon length. 222 If intron content is a power law with a value between 2 and 3, it would support the hypothesis that 223 compositions of genes are related to packing geometry. Supporting our volumetric hypothesis, we 224 observed this relationship ( Figure 3d). This is captured by Equation (2) where p ranges from 225 500bp to 25,000bp: 226 2) /g18Y5/g1866/g1872/g1870/g1867/g1866 /g150Y /g18Y1/g1876/g1867 /g1866 /g3082 /g1868/gY415 227 Understanding the function of the scalar, p, and the exponent, γ, requires considering how these 228 translate linear information into 3-D volumes. Along the linear genome, p and γ are proposed to 229 control the density of exonic information. Here, a lower γ (~2) produces less spacing than larger γ 230 (~3). Likewise, the value of p indicates where along the depth of a volume a segment is being 231 positioned (a smaller value suggesting it is deeper in a volume, and a larger value the inverse). 232 The central hypothesis would suggest that both degrees of freedom could interact to 233 accommodate an effective packing configuration across multiple conditions (concentration of 234 nucleosome remodeling enzymes, ion concentrations, nuclear size, existing domains, etc) that 235 are difficult to measure in every human cell. Generally, a high p and low γ state produces high 236 information density. As a result, structural genes such as Myh1 ( p ~7,800, γ = 2.2) are proposed 237 to adopt more complex geometries with some degree of packing present. It is worth noting that in 238 .CC-BY-NC 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 29, 2025. ; https://doi.org/10.1101/2025.05.29.656862doi: bioRxiv preprint 6 the proposed hypothesis, γ is inversely proportional to D as the transformation of the chain into 239 the occupied volume, quantified by: 240 3) /g2011/gY404 1 /gYY97 /g3005 /g3005/g2879/g3081 241 where β defines the fractional dimension of exonic contents on the ideal zone ( Figure 3e-f 242

Materials

and Methods). 243 The values of these scaling exponents are indicative of whether exonic and NE regions likely 244 follow 3D volume relationships of domains. In a case when the 3D packing geometry is not 245 observed (e.g. a chain assembly), β=0 resulting in /g2011/gY404 1 , /g1855/gY404 1 , and /g1866/gY404 0 with /g18Y5/g150Y /g18Y1 and /g18Y1/g150Y /g18Y8 . 246 On the other hand, a strong packing geometry corresponds to 2/gY409/g2011 /gY4093, /g1855/gY407 1 , and /g1866 ~ 1 . Our 247 experimental data is consistent with the latter case of packing geometry as we describe in detail 248 within this manuscript. Since we propose that these properties are related to the activity of 249 transcription, we verified that this behavior is independent of transcript type by analyzing the 250 behavior of non-protein coding RNA classes (SI Figure 7). 251 Gene positions on human chromosomes reflect the projection of 3-D domain volumes 252 We next investigated if these principles generalize to the positioning of genes on chromosomes. 253 Since transcription guides domain formation experimentally, it suggests that genes oriented in cis 254 may share domain volumes. In this hypothesis, intergenic segments could therefore share similar 255 volumetric properties as introns to support domain formation. On both ChromSTEM imaging and 256 in polymer simulations, packing domains are separated by an interdomain space formed by DNA 257 [15,16,30,48,52,53]. From experimental observations, this linker segment must contain an element that 258 can recognize Pol II (for domain generation) and contain enough DNA to generate separation 259 between the two domains. We refer to this as a ‘hinge’: when engaged by actively transcribing 260 Pol II it will form separate segments into two domain volumes. A domain will be formed by the 261 span between the two hinges. When not engaged by Pol II, a hinge may revert to fold into a 262 volume element in an alternate configuration ( Figure 4a). If hinges are only composed of 263 transcribed segments, the supercoiling generated by polymerase could potentially alter the 264 composition of the domains (domains could merge or separate randomly) [54,55]. Alternatively, if 265 hinges were completely non-exonic, they would lack a mechanism for Pol II to guide domain 266 segmentation. Since hinges are not domains, they are likely no more than a few nucleosomes 267 long (<1000bp, ~50nm long). Above this range, they would potentially interact with nucleosome 268 remodelers via the mechanisms that produce domain volumes, limiting their ability to achieve 269 spacing. 270 By considering the presence of hinges, we find that every human chromosome assembles into 271 power-law packing ratios (Figure 4b). This organization resembles individual genes described by 272 Eq. 3, with p and γ adopting the same meaning: 273 4) /g1840/g18Y1 /g150Y /g18Y1/g1876/g1867/g1866 /g3082 /g1868/gY415 274 275 This suggests that exons are non-randomly linked with their adjacent NE neighbor to produce 276 reaction volumes. Since this is difficult to test experimentally, we analytically tested this 277 hypothesis by randomizing the positions of exons alone and compared to randomizing exon+NE 278 pairs ( Figure 4c). If exons are uncorrelated NE elements producing packing domain volumes, 279 then their randomization would still produce ratios that are consistent with volumes. When exons 280 are distributed alone, the power-law compositions degrade into linear (unpacked) ratios ( Figure 281 4c-d). Conversely, when exon+NE are randomly distributed together, power-law patterns are 282 maintained. This supports the hypothesis that exonic segments are non-randomly paired to an 283 adjacent NE segment in 3-D volumes. If this is intrinsically driven by transcription itself and not 284 the final product, this pattern would be present in either protein coding or non-protein-coding 285 .CC-BY-NC 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 29, 2025. ; https://doi.org/10.1101/2025.05.29.656862doi: bioRxiv preprint 7 segmentation. Independent of the type of RNA generated (protein coding RNA, non-protein 286 coding RNA), chromosomes assemble into power-law segments of exonic-/NE- elements (Figure 287 4c-d). Collectively, these findings are consistent with the hypothesis that exons, introns, and 288 intergenic segments are coupled in transcriptionally mediated reaction volumes. 289 Predictions of genome geometry are observed across methods 290 The central hypothesis proposed in this manuscript – exons, introns, and intergenic segments are 291 non-randomly coupled by the 3-D volume geometry – introduces four dependent and testable 292 hypotheses, even as testing the central hypothesis directly is challenging: 293 H1) The DNA content of predicted segments resembles experimental observations on 294 ChromSTEM, 295 H2) Exons have a higher likelihood of interacting with Pol II based on their position, 296 H3) RNA synthesis decreases from the introns of long genes from suboptimal positioning, 297 and 298 H4) Hinge segments are enriched for features of active transcription. 299 Although these hypotheses cannot definitively prove the novel positioning hypothesis, as it would 300 require manipulation and synthesis of genetic elements, the proposed central hypothesis requires 301 all of them to be sustained. Hence we tested whether these dependent hypotheses are supported 302 by experimental data. 303 To test H1, we utilized existing ChromSTEM imaging data from HCT-116 cells. The DNA content 304 observed on ChromSTEM domains can be calculated by converting intensity into content using 305 the mass-fractal relationship [5,16]. Although we cannot identify the sequence composition on 306 ChromSTEM, we can compare the distribution of sizes between theory and experimental 307 measurements of PDs. Consistent with theory predictions, the DNA totals generated by the 308 segmentation described is very similar to those experimentally observed ( Figure 5a, Mann 309 Whitney U-Test 0.89 for negative reading frame and 0.93 for positive reading frame) but not for 310 segments generated from randomly positioned exons (Figure SI 8, Mann Whitney U-Test <10-45). 311 H2 requires three observations. First, there must be a length-dependent coupling between 312 functional zones (cores, active Pol II, and euchromatin) reflecting domain volumes. Specifically, 313 the higher volumetric ratio in small domains produces a higher concentration of Pol II with a 314 shorter linear distance to a core. In contrast to packing, a chain assembly would result in 315 increasing distance between Pol II to H3K9me3. Using ChIP-Seq data available through 316 ENCODE[7–9], we tested if the coupling between transcription and heterochromatin is consistent 317 with domain volumes by measuring the distance between active Pol II (Ps2/Ps5) and the nearest 318 H3K9me3 as a function of the polymerase concentration [6]. We performed this analysis on 319 induced pluripotent stem cells (GM23338), HepG2 hepatocellular carcinoma cells, and SK-N-SH 320 neuroblastoma cells using data from ENCODE. If Pol II is guided by geometric position, an 321 inverse distance relationship is likely to be conserved across models (higher Pol II concentration 322 resulting in shorter distances). Further, because iPSCs are highly enriched in euchromatin [56,57], 323 the theory predicts them to have the smallest domains with highest volumetric ratios and shortest 324 distances. These predictions are observed: all three cell lines have the inverse distance 325 relationship and iPSCs have the shortest distances on average (Figure 5b, SI Figure 8). Second, 326 H2 indicates that exons are positioned to geometrically interact with polymerase, euchromatin, 327 and heterochromatin. Specifically, one would expect H3K4me3 and active RNA polymerase II to 328 localize in the projection of the ideal zone mainly composed of exons whereas H3K9me3 would 329 localize to deeper layers. This would be manifested in the segments by a shift in the p required to 330 position these elements such that H3K4me3 and Pol II would range between ~250-25,000 bp and 331 .CC-BY-NC 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 29, 2025. ; https://doi.org/10.1101/2025.05.29.656862doi: bioRxiv preprint 8 H3K9me3 to values below this range (1-50bp). This is indeed observed experimentally in HCT-332 116 cells (Figure 5c, SI Figure 9). 333 Third, H2 requires a Pol II preference to binding exons when accounting for the differences in 334 segment length. This is similarly involved in H3 because if Pol II continuously moves throughout a 335 gene, including long introns, it would not be uniformly distributed. In support of H2 and H3, active 336 Pol-II is nearly ten times more likely to bind to an exon than an intron in both Hep2G and HCT-337 116 cells when accounting for length (Figure 5d, p-value <10-3). H3 further requires that this bias 338 in positioning result in a gradual decay in the amount of RNA observed from introns in long genes 339 due to decreased efficiency. In bulk RNA-seq, this results in a ‘saw-tooth pattern’ which was 340 previously theorized to occur from the delayed processing of RNA by the splicing enzymes or the 341 decay in introns [58]. However, the rate of splicing can greatly exceed the rate of transcription for 342 long introns, indicating that this is not necessarily a rate-limiting event in RNA synthesis [59–61]. If 343 genes are synthesized in total, the decay in intronic RNA likely would be symmetrical with the 344 highest rates of decay in the shortest introns. This is because processing will not impact the 345 synthesis rate within intron bodies and short segments will likely have the fastest processing. 346 Utilizing nascent RNA-seq, we tested which configuration is most likely. Consistent with H3, there 347 is a length-dependent decrease in RNA synthesis in 5’ to 3’ orientation for introns of long genes 348 (Figure 5e). 349 We used a similar approach to investigate H4. We analyzed data from ENCODE and the Atlas of 350 Enhancers[2] in HCT-116 cells for hinges positioned in the human genome (HG) vs hinges 351 generated by randomly repositioning exons (R). Hinge positions behaved as hypothesized, with 352 enrichment in euchromatin marks and enhancers and depletion of heterochromatin markers 353 compared to the random genome. In the human genome, ~7.8% (interquartile range 7.1% to 354 9.5%) were bound by active RNA polymerase II (Pol II-Ps5), 9.7% (interquartile range 8.3% to 355 10.3%) were marked by H3K27ac, 8.7% (interquartile range 7.8% to 9.5%) were marked by 356 H3K4me3, and 19.8% of annotated enhancers contained at least one hinge element (interquartile 357 range 15.6% to 24.8%) (Figure 5e&f, SI Figure 9). In contrast, hinge positions were depleted of 358 heterochromatin modifications with H3K9me3 occurring 0.13% of the time (interquartile range 359 0.1% to 0.17%) and H3K27me3 occurring 1.2% of the time (interquartile range of 0.9% to 1.8%, 360 Figure 5g&h, SI Figure 9 ). Collectively this indicates, as proposed, that the act of transcription 361 could facilitate the generation of domain geometry in a predictable manner. 362 Genome geometry suggests transcriptional loops are efficiently packed 363 In sum, H1-H4 are supported by experimental evidence. While individually they can be explained 364 by alternative mechanisms, the novel hypothesis that exons are non-randomly coupled to NE to 365 generate domain volumes, in our view, best explains the sum of the evidence. Having 366 demonstrated that the volumetric pattern of domains could be projected onto the positioning of 367 exons, introns, and intergenic segments, we then studied whether they intersect with gene 368 transcription. To do so, we tested if the proposed theory would indicate that transcriptionally 369 active loops (mediated by Pol II), are packed in space. Two features we have proposed could be 370 manifested in chromatin loops generated by Pol II[62]. The first is that some loops must represent 371 durable domain volumes with compositions that mirror the ratios we described. 372 Durable structures could produce high-frequency loops within a population due to spatial 373 confinement. Therefore, we used publicly available data through ENCODE to analyze the 374 behavior of Pol II generated loops on ChiaPET [62]. We partitioned loops into very strong loops 375 (>20 events) and compared them to more transient loops (<5 loop events). We then analyzed the 376 composition of these two groups. Consistent with the novel hypothesis, very strong loops contain 377 exon/NE ratios that are consistent with volumetric ratios of domains ( SI Figure 10). Likewise, 378 more transient loops spanned a spectrum of states including a mix of exon/NE segments, 379 suggesting stretched as well as packing configurations (SI Figure 10). 380 .CC-BY-NC 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 29, 2025. ; https://doi.org/10.1101/2025.05.29.656862doi: bioRxiv preprint 9 To test if stable loops are a packed distribution of sizes, we utilized publicly available ATAC-Seq 381 and ChIP-Seq. Similar to the geometry of domains guiding chromatin enzymes, the Tn5-382 transposase would interact with the outer zone based on the packing of the loop. If loops are 383 instead a chain, then accessibility will be a linear function of loop length. Consistent with the 384 packing properties proposed, accessibility is a power-law of length observed in strong loops ( SI 385 Figure 10). This occurs even upon depletion of RAD21, indicating that the generated volumes 386 are not dependent on cohesin extrusion, which is consistent with findings on ChromSTEM 387 imaging that mature domains remain even upon RAD21 depletion [5,6] ( SI Figure 10 ). Similarly, 388 analysis of these volumes mirrored the proposed packing observed by segmenting the genome 389 from the proposed hinges ( Figure 5c, SI Figure 10 ). This is also observed experimentally 390 (Figure 5k, SI Figure 10 ): the chromatin modifications are a power-law of non-exonic length 391 within the loop segment. Features that represent the core of domains (H3K9me3, H3K27me3, 392 Figure 5k, SI Figure 10) were at the proposed shorter depths (p ranging from 1 to 50 basepairs, 393 Figure 5f) whereas Pol II and euchromatin aligned with predicted exon positions. 394 Geometry suggests packing is associated with differentiation and oncogenic risk 395 We then explored the risks and benefits of the proposed geometric system. The process of 396 domain formation and maturation echoes that of reinforcement learning systems in artificial 397 intelligence and neural networks [6,63,64]. Inputs (signaling cascades, mechanical force) intersect 398 with the current state (existing domains, ionic conditions, nuclear volume, and nucleosome 399 remodeling complex concentration) that generate an output (RNA synthesis, domain 400 activation/modification/degradation). Since the output modifies the state (domains), a future input 401 experiences a different state that was modified by the prior inputs. By organizing the human 402 genome into several thousand domains guided by transcriptional inputs determined by the 403 existing state of the cell, cells can produce coordinated behaviors across a tissue. The domains 404 encode a crucial feature for multicellular systems: a mechanism to produce coherent, 405 reinforceable, and predictable memories of prior events. This represents an efficient, non-406 mutational system to increase complexity since geometry (packing) stores states. If this is true, 407 genes active in terminally differentiated human tissues will be primarily power-law compositions, 408 based on the need to maintain dynamic function for decades. 409 To test if this is indeed observed, we analyzed differentiated tissues from each germ layer: 410 esophageal mucosa (endoderm), cardiac muscle (mesoderm), and cortical neurons (ectoderm). 411 We utilized GTEx expression data from these three sites and selected genes that were 412 preferentially associated with each tissue ( SI Table 1 )[65]. We observed conservation of the 413 domain geometry system for genes involved in maintaining tissue function across the human 414 lifespan (Figure 6a&b). Given this finding, we explored whether transcription factor families were 415 similarly organized. Embryogenic development is a complex process with the timing of 416 transcription factor activation impacting tissue formation. Therefore, we analyzed transcription 417 factors in relation to developmental timing. We found that genes of early development 418 (pluripotency factors, HOX genes) [66,67] favored exon enrichment (linear structure or very small 419 domains) whereas end-organ factors (e.g. MITF [68] or RUNX2 [69]) would generate large, power-420 law domains (Figure 6c&d, SI Table 2 ). However, by the proposed nature of a hinge guided by 421 transcriptional activity, a portion of genes are placed in relatively risky conditions as they require 422 being spanned between multiple reaction volumes. The risk to gene segments arises because 423 they are exposed to environmental conditions and are at risk for entanglement events. As a 424 result, our hypothesis predicts that these segments would be at a higher risk of mutations across 425 all cancers. To test if this is the case, we investigated the mutation frequency of Tier-1 oncogenes 426 compared to the likelihood that these genes overlap with a hinge position. We utilized the 427 reported mutation frequencies across all tumor types generated by ROSETTA, which analyzed 428 frequencies from publicly available cancer genomics sources (i.e. TCGA/TARGET Program)[70,71]. 429 As these frequencies span across all tumor tissue types, they would provide an understanding of 430 a conserved mechanism across cancers independent of tissue-specific factors. The null 431 .CC-BY-NC 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 29, 2025. ; https://doi.org/10.1101/2025.05.29.656862doi: bioRxiv preprint 10 hypothesis is that no association would be observed and that mutations would be independent of 432 the generated geometry. Remarkably, we observed that the frequency of oncogenic mutations is 433 strongly correlated with the likelihood of a gene containing a hinge position ( Figure 6 e&f ). 434 Indeed, many crucial Tier-1 genes (both oncogenes and tumor suppressors) overlap with the 435 presence of a hinge including TP53, BRCA1/2, ATM, RB1, IDH1/2, and PIK3CA ( SI Table 3 ). 436 Collectively, these findings suggest that the proposed geometry is paired with a risk to genes that 437 are required to create bifurcations between domains. While further work is necessary to 438 understand if positioning is causal, it suggests that oncogenic mutation frequency may be linked 439 to geometry. 440 Power-law geometry parallels the emergence of organo-axial development 441 Since we propose that the benefit of the proposed geometry hypothesis is to produce durable cell 442 states, we conclude by performing a comparative analysis across eukaryotic genomes to 443 understand if it is a consequence of the act of splicing or an emergent regulatory mechanism. We 444 specifically chose to analyze the following species due to their well-characterized genomes and 445 existence of nucleosomes, introns, and splicing machinery: S. cerevisiae (a model of a 446 monocellular organism with splicing and introns), C. elegans (a multi-cellular organism with 447 simple organoaxial positioning), D. melanogaster and D. rerio (multi-cellular organisms with 448 complex organoaxial positioning), and M. musculus (a non-primate mammal with complex 449 organoaxial positioning)[72]. If this system is a byproduct of the act of splicing, all the investigated 450 genomes would have similar geometries independent of complexity. Instead, if geometry built 451 upon splicing to increase durable cell states, it would parallel the complexity of metazoan cell 452 types. Consistent with the hypothesis that geometric encoding could facilitate organ complexity, 453 power-law coupling of exons with introns appears to parallel organ specification Figure 7a-h). We 454 observed the transformation from linear geometries with exon enrichment ( S. cerevisiae) first 455 toward a mix of power-law and linear structures ( C. elegans ) (Figure 7a-h ). As complexity 456 increased, there is a further transition from an equal mix of linear and small geometric assemblies 457 to primarily geometric assemblies in D. rerio onwards. From the perspective of domain volume 458 generation from the linear genomes, power-law geometry would require activation of one gene 459 coming at the expense of a neighbor. Alternatively, transcription in linear genomes could occur in 460 the span connecting domains, potentially indicating a mechanism by which the process 461 progressed from a linear to a power-law assembly. While a linear assembly has benefits for rapid 462 simultaneous synthesis, it is unlikely to efficiently produce memory and specialization. As a result 463 of these considerations, it is worth noting that compaction within S. cerevisiae effectively 464 translates into a 3-D barrier, whereas this is not the case in the assembly of domain volumes [73]. 465 Instead, genomes built on a volumetric geometry can produce complex, highly dynamic, and 466 durable states. 467 468

Discussion

469 470 We set out to understand how packing domain information is stored within the genome in order to 471 explain how a paradoxical decrease in accessibility of introns and intergenic segments is 472 observed in transcriptional activation during muscle differentiation. We arrived at the novel 473 hypothesis that the positioning of segments (exons, introns, intergenic) may encode 3-D packing 474 information to generate nanoscale volumes. While additional investigation is needed to fully prove 475 the implications of the proposed theory, it does introduce a mechanism for storing volumetric 476 information for efficient transcriptional reactions (Figure 1)[15,16,30,48,52,53]. We first highlight that our 477 findings demonstrate a crucial role for non-exonic DNA to act as ‘volumetric DNA’ in complex, 478 multicellular eukaryotes to optimize chemical reactions involved in transcription [42–44]. This theory 479 intriguingly suggests that the positional composition of human chromosomes may exist to 480 produce efficient RNA synthesis (Figure 2-5). This is because while the vast array of the genome 481 does not make RNA directly, length-based positioning would create efficient reaction volumes. 482 .CC-BY-NC 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 29, 2025. ; https://doi.org/10.1101/2025.05.29.656862doi: bioRxiv preprint 11 There are immense benefits from the proposed system: space and information optimization, 483 durability, efficiency, simultaneous enzymatic processing, and geometric modulation to increase 484 degrees of freedom (Figure 1&2). 485 Finally, it is notable that while most of the human, mouse, and zebrafish genomes appear to 486 organize by a geometric principle ( Figures 7 ), the transformation from linear, exon-rich 487 configurations to power-law volumes correlates with the degree of body plan complexity. In future 488 work investigating individual gene configurations, the framework demonstrated here could provide 489 an understanding of how a beneficial epigenetic projection (3-D domain volumes) becomes 490 encoded as a sequence-specific segment. This would require the ability to model and test specific 491 gene configurations. Likewise, it generates questions as to whether DNA repair and replication 492 enzymes behave like Pol II as a function of physiochemical conditions and the position of 493 elements. We anticipate that future work can address these questions as the central hypothesis is 494 mechanistically explored. These investigations could dissect how packing information is passed 495 through the cell cycle, and how TADs and cohesin loops interact with this system. With those 496 considerations in mind, we describe several transformative hypotheses to test across different 497 disciplines based on the presented theory in addition to mechanistically testing the hypothesis 498 proposed here: 499 Hypothesis 1) Domain geometry facilitated the rapid emergence of body plan diversity. The 500 transformation of chromosomes from linear to power-law assemblies parallels increasing body 501 plan complexity ( Figure 7). We naturally wonder if this converges at the time of the Cambrian 502 Expansion[74–76] as domain geometry can be a non-mutational system to increase degrees of 503 freedom. This occurs because changing configurations depends on modifying spacing and 504 position without requiring sequence mutations. This hypothesis can be tested by using the 505 framework described here and then measuring the change in exon sequence composition 506 compared to the change in arrangement. If evolutionary trajectories are found to be encoded in 507 the arrangement of elements, it would demonstrate that selective pressure converges on gene 508 position, orientation, and segment lengths to generate cohesive assemblies. What we perceive as 509 ‘neutral drift’ from the sequence perspective is potentially balanced by selection for maintaining 510 volumes. This allows increased sequence sampling of these regions mutationally for potentially 511 beneficial states while maintaining a ‘default’ state as volumetric elements. Such findings would 512 suggest that genomic selection is driven by the benefit of the system, and not necessarily the 513 gene alone[42,43]. 514 Hypothesis 2) Volume stabilization of mature domains defines cell response across decades. 515 Nuclear swelling and heterochromatin loss are hallmarks of aging across human tissues [53,77–79]. 516 Tissue development is defined by the non-random deposition of heterochromatin [80]. The volume 517 ratios in packing domains reflect these states [48,81,82]. With this in mind, several diseases may be 518 influenced by domain degradation over time. For example, among the risk factors for Alzheimers 519 is the read-through fusion of ApoE and Tomm40 [83] ( Figure 1 ). Could nuclear expansion and 520 heterochromatin loss shift reaction volumes to produce transcriptional ‘fusions’ by positioning 521 segments along the ideal zone as a single element [84]? Could a similar process be involved in 522 inflammatory diseases since the misfolding of domains could create transcriptional memories that 523 sustain inflammation [85]? Finally, in cancer, chromosome fusions, fragmentation, and copy 524 number variations are significant prognostic factors. We observed that oncogene mutation 525 frequencies were correlated with localization to a hinge segment ( Figure 6); could these be non-526 random events that are related to how structure responds to local conditions? The nanoscale and 527 microscale transformation of the nucleus is similarly a hallmark of malignancy and 528 chemoresistance. Increasing the total genomic content, shifting the positioning of genes, and 529 altering nuclear volumes generates unexpected geometries and the loss of coherent responses to 530 the same stimuli. Based on the principles of domain geometry, one could prevent these events by 531 targeting the physiochemical conditions that define the structure of domains[86–88]. 532 .CC-BY-NC 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 29, 2025. ; https://doi.org/10.1101/2025.05.29.656862doi: bioRxiv preprint 12 Hypothesis 3) Domain geometry accommodates latency . Among the mysteries of transcriptional 533 patterns is an unused surplus of binding sites for key transcription factors (e.g. Myod)[89–91]. While 534 it has been suggested that these are errors, we can reconsider them in terms of information 535 latency based on domains acting as a computational system. The default state generates 536 volumes when unused, which remain as a reservoir for new domains if conditions change. In 537 effect, the default positions needed for a human body plan are encoded in the measured 538 positions of exonic/volumetric pairs with some degree of error to preserve a non-default 539 adaptation strategy. If these adaptations improve fitness, they could be hard coded for the next 540 progeny. Repositioning of these elements and the deposition of transcription factor binding sites 541 based on geometry could provide insight into how sequence specificity intersects with geometric 542 specificity. 543 Hypothesis 4) Chromatin is a geometric computational system 544 By considering physical properties of the genome at the intersection with transcription reactions, 545 these elements mirror aspects of learning, computation, and neural networks [63,64]. Some of the 546 positional elements observed mirror elements of logic operators or computations, such as “AND” 547 where genes in a paired segment become co-transcribed and “OR” states where one gene 548 competes with another for volume. In this context, the predictable spacing between exons, 549 introns, and intergenic segments requires inputs from the state of the cell. Alterations in 550 mechanics[30,92], ions [88], redox state [93,94], and nucleosome remodeling enzymes [53] can provide 551 this information. Domains would then act as geometric processors, with the structures formed 552 representing the intersection of inputs and the current state. While this is more challenging to test, 553 manipulating gene positions or controlling the order of signals could produce insights into this 554 process. Finally, from the perspective of synthetic biology, controlling volume positioning during 555 the generation of artificial chromosomes may guide the ability to generate complex traits 556 Collectively, this novel hypothesis requires additional mechanistic investigation for causality. The 557 proposed system currently invites new avenues of exploration of genomics at the intersection of 558 molecular, physical, chemical, and computational properties. Grounded in the intersection of 559 these fields, exploration into the influence of domain volume assemblies on cellular fitness and 560 organism evolution could serve as a framework for future studies in critical processes such as 561 DNA replication, repair, and splicing. 562

Materials and methods

563 564 Geometric positioning of exons and volumetric DNA in relation to packing domains 565 We derive scaling relationships to assess whether non-exonic (NE) DNA is coupled with exons to 566 generate volume assemblies. Comparison between this analytical model and sequencing and 567 ChromSTEM experimental data suggests that non-exonic segments likely correspond to 568 volumetric organization surrounded by exons behaving as a “wavy” line on the reaction zone of 569 the volume provided by the non-exonic elements, with domain fractal dimension D inversely 570 related to γ. 571 572 Chromatin packing domains are nanoscopic, heterogeneous mass-fractal structures. The 573 transformation of a chromatin chain into volume is defined by how the length of a segment of the 574 chromatin polymer within a 3D domain, M, scales as a function of the radial distance of the 575 volume containing the polymer, r by /g18Y9~ /g1870 /g3005 , where D is the fractal dimension. The chromatin 576 volume fraction, φ, at the radial distance r is defined by: 577 1) /g20Y8 /g4666 /g1870 /g4667 /gY404 /g20Y8 /g2868 /g4672 /g3045 /g3278 /g3045 /g467Y /g2871/g2879/g3005 578 .CC-BY-NC 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 29, 2025. ; https://doi.org/10.1101/2025.05.29.656862doi: bioRxiv preprint 13 where /g20Y8 /g2868 is the chromatin volume fraction at r=0 (domain center) and r c is the chromatin chain 579 radius[15,16,37]. The total amount of chromatin in basepairs within the domain volume will therefore 580 be: 581 2) /g1840 /g4666 /g1870 /g4667 /gY404 /g1840 /g3030 /g1827 /g3005 /g4672 /g3045 /g3045 /g3278 /g467Y /g3005 582 with N being the basepairs contained within the radius, N c the number of basepairs within the 583 chain and A D the packing efficiency of the chain within a domain. Noting that A D =1 indicates 584 efficient packing throughout the domain volume. 585 586 For hard 3-D objects (e.g., a hard sphere), the number of basepairs, N s, within a zone shell of 587 radius ΔR that is much smaller than the radius of the volume (ΔR << r) is: 588 3) /g1840 /g3046 /g4666 /g1870 /g4667 /gY404 /g3031/g3015 /g3031/g3045 ∆/g1844 589 Substituting in Eq.(6), we therefore observe: 590 4) /g1840 /g3046 /g4666 /g1870 /g4667 /gY404/g18Y0 /g1499 /g1840 /g3030 /g1827 /g3005 /g4672 /g3045 /g3045 /g3278 /g467Y /g3005/g2879/g2869 ∆/g3019 /g3045 /g3278 591 592 Accounting for the fact that the chain elements are a polymer that can go in and out of the hard 593 shell (e.g., the “wiggly” line in Figure 1), it is instead necessary to take the fractional derivative to 594 capture the behavior at the reaction zone: 595 5) /g1840 /g3046 /g4666 /g1870 /g4667 /gY404 /g3031/g3015 /g3031/g3045 /g3329 ∆/g1844 /g3081 596 where /g2010 is the order (dimension) of the derivative and ranges from 0 < /g2010 <1. 597 598 Utilizing the chain rule for a function f(x): 599 6) /g3031/g3033 /g3031/g3051 /g3328 /gY404 /g3031/g3033 /g3031/g3051 /g3031/g3051 /g3031/g3051 /g3328 /gY404 /g2869 /g3080 /g1876 /g2869/g2879/g3080 /g3031/g3033 /g3031/g3051 600 601 and Eqs.(9, 10), we therefore observe that the composition of the ideal zone is: 602 7) /g1840 /g3046 /g4666 /g1870 /g4667 /gY404 /g2869 /g3081 /g1870 /g2869/g2879/g3081 /g3031/g3015 /g3031/g3045 ∆/g1844 /g3081 /gY404 /g3005 /g3081 /g1840 /g3030 /g1827 /g3005 /g4672 /g3045 /g3045 /g3278 /g467Y /g3005/g2879/g3081 /g4666 ∆/g3019 /g3045 /g3278 /g4667/g3081 603 604 Assuming that the content of exons, E, in basepairs is approximately that of the contents of a 605 domain ideal zone /g18Y1 /g150Y /g1840 /g3046 then the proportionality constant in this relationship is /g150Y/g2010 . From this, 606 we observe: 607 8) E /g4666 r /g4667 /gY404/g18Y0 /g186Y /g1840 /g3030 /g1827 /g3005 /g4672 /g3045 /g3045 /g3278 /g467Y /g3005/g2879/g3081 /g4666 ∆/g3019 /g3045 /g3278 /g4667/g3081 608 With k representing the fraction of the zone basepairs that are exons. Likewise, the length of a 609 segment within a domain volume is the number of basepairs within that volume, L(r) = N (r). 610 Transcriptional reactions occur at the ideal zone (the “Goldilocks zone”), the region where the 611 balance between density stabilizes the intermediate complexes without overly limiting diffusivity of 612 the reactant species[87,88,95,96]. For a domain limited by its ideal zone, the total length of the gene, 613 L(r=Rgl), and the exons, E(r= Rgl) within domains to the goldilocks radius are: 614 .CC-BY-NC 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 29, 2025. ; https://doi.org/10.1101/2025.05.29.656862doi: bioRxiv preprint 14 9) /g18Y8/gY404 /g1840 /g3030 /g1827 /g3005 /g4666 /g3019 /g3282/g3287 /g3045 /g3278 /g4667/g3005 615 which indicates that 616 10) /g4672 /g3019 /g3282/g3287 /g3045 /g3278 /g467Y/gY404/g4666 /g18Y8 /g1840 /g3030 /g1827 /g3005/gY415 /g4667 /g2869 /g3005/g3415 617 618 Solving for E using Eq.(12) utilizing the relation from Eq.(14) we can calculate the exon contents 619 by 620 11) E/gY404/g18Y0 /g186Y /g4666 /g1840 /g3030 /g1827 /g3005 /g4667/g4672 /g3081 /g3005/g3415 /g4673 /g2012/g3034/g3039 /g18Y8 /g4672 /g3253/g3127/g3329 /g3253 /g4673 with /g2012/g3034/g3039 /gY404/g4666 ∆/g3019 /g3282/g3287 /g3045 /g3278 /g4667/g3081 621 For simplicity, we can now let /g1851/gY404 /g18Y0 /g186Y /g4666 /g1840 /g3030 /g1827 /g3005 /g4667/g4672 /g3081 /g3005/g3415 /g4673 /g2012/g3034/g3039 and /g1829/gY404 /g3005/g2879/g3081 /g3005 which simplifies Eq.(14) to 622 12) E/gY404 Y /g18Y8 /g3004 with C generally bounded between 2/3 and 1 due to the limits of D in cells of 2 to 623 3. 624 625 The existence of power-law scaling between exon length, E, and total gene length, L suggests a 626 scaling relationship between the intron/intergenic segment and the exon. We observe an 627 exponent, γ in Eq.(2) and Eq.(4) such that /g18Y5/gY404 /g3006 /g3330 /g3040 . 628 If /g2011/gY404 1 /gYY97 /g2869 /g3004 then /g18Y5/gY404 /g3006 /g3117/g3126/g3117//g3278 /g3040 , this would result in /g3010 /g3006 /gY404 /g3006 /g3117//g3278 /g3040 which by Eq.(15) becomes /g3010 /g3006 /gY404 /g3026 /g3117//g3278 /g3013 /g3040 . Given 629 that Y is nearly 1, this results in /g3010 /g3006 /g1542 /g3013 /g3289 /g3040 consistent with the experimental observations in Eq.(1). 630 Translating the power-law scaling described by γ into the mass-fractal dimension of domains, we 631 observe that 632 13) /g2011/gY404 1 /gYY97 /g2869 /g3004 /gY4041 /gYY97 /g3005 /g3005/g2879/g3081 as described in Eq.(4). 633 634 For a domain that is extends outside of its ideal zone, some NE regions may be found at r > R gl. 635 Treatment similar to the one described above can be applied. From Eqs. 6 and 11, it follows that 636 14) /g18Y8/gY404 /g1840 /g4666 /g1844 /g3032 /g4667/gY404 /g2871 /g3005 /g1840 /g3030 /g20Y0 /g2868 /g4666 /g3101 /g3116 /g3101 /g3280 /g4667 /g3253 /g3119/g3127/g3253 , /g18Y1/gY404 /g1840 /g4666 /g1844 /g3032 /g4667/gY404 3/g186Y /g1840 /g3030 /g20Y0 /g2868 /g4666 /g3101 /g3116 /g3101 /g3282/g3287 /g4667 /g3253/g3127/g3329 /g3119/g3127/g3253 /g4666 ∆/g3019 /g3045 /g3278 /g4667/g3081 , 637 where /g20Y0 /g3032 is the chromatin volume fraction at the radial distance corresponding to the outer bound 638 of the domain, r=/g1844 /g3032 . In a special case of /g20Y0 /g2868 /gY404/g20Y0 /g3034/g3039 , manipulation of equations 18 leads to 639 15) /g18Y1/gY404 /g186Y /g18Y0 /g4666 /g2871 /g3005 /g20Y0 /g3034/g3039 /g4667 /g3329 /g3119 /g4666 ∆/g3019 /g3045 /g3278 /g4667/g3081 /g18Y8 /g2869/g2879 /g3329 /g3119 . 640 Again, we see power-law scaling relationships among /g18Y5, /g18Y1 , /g18Y8 : 641 16) /g3010 /g3006 /g150Y/g18Y8 /g3041 , /g18Y5/g150Y /g18Y1 /g3082 , and /g18Y1/g150Y /g18Y8 /g3004 . 642 643 We now consider positioning of elements depending on the property of the exon segments in 644 relation to the ideal zone elements defined by β. For β=1, the entire contents of the exon are 645 within the hard shell. In contrast, for β=0, the exon portion is volumetrically distributed indicating 646 that there is no geometry. 647 .CC-BY-NC 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 29, 2025. ; https://doi.org/10.1101/2025.05.29.656862doi: bioRxiv preprint 15 Indeed, in the most likely case, where exons constitute a portion of the ideal zone (a wiggly line) 648 then 0 < /g2010 <1 with D of domains ranging experimentally on ChromSTEM and in polymer modeling 649 between 2 and 3 (2 < D < 3)[15,16,48], this results in 2 < γ < 3 observed experimentally. 650 When γ=2 and β=0, exons are distributed for any observed D (no geometric relationship). 651 Conversely, when γ=3 and β=1, D =2 results in the limiting case of a chromatin polymer in a good 652 solvent. The most frequently experimentally observed γ ~ 2.3 is achieved when β~0.7, indicating 653 a substantial congruence between the ideal zone and exons, for experimentally (ChromSTEM) 654 observed most probable D ~ 2.8. It is worth noting that in this derivation the observed γ and m 655 values plotted are those of the ensemble, and not the realized values for each individual gene. In 656 essence, the entire genome is organized by this behavior, however, each individual gene may 657 require segments from a neighboring segment in order to achieve the realized domain. 658 Power-law genomic analysis 659 RefSeq genomes were obtained from Igv.org/app mapping to the UCSC genome browser 660 assemblies. The selected genome assemblies were as follows: Human (GrCH38), Mouse 661 (GRCm39), Zebrafish (GRCZ11), Drosophila melanogaster (dm6), and S. cerevisiae (sacCer3). 662 These text files were converted to xlsx extensions and then imported into Mathematica v12 for 663 subsequent analysis with custom-built code. Genes were then separated to analyze protein 664 coding with assigned prefix “NM” and non-protein coding “NR”. The first isoform was selected for 665 analysis of individual genes and compared to analysis with all isoforms. At the level of 666 chromosome analysis, all isoforms were considered. Exon start and stop positions were used for 667 segmentation and an reciprocal start/stop position for introns was generated. For simplicity, all 668 exons were assumed to be part of the gene within the isoform variant. For chromosome wide 669 analysis to account for the direction of transcriptional reading frames, genes were separated by 670 their location to the positive or negative strand orientation for analysis. Multi-start or multi-stop 671 exon overlapping events accounted for less than 7% of exon positions but were omitted in whole 672 chromosome analysis for simplicity. Hinge elements were subsequently identified either for the 673 whole chromosome in the read orientation (positive, negative) by mapping in the orientation of the 674 read-frame (exon-NT) such that their size was below the predicted threshold: ~200-300bp (1 675 nucleosome hinge) or ~500-600bp (2.5 nucleosome hinge). The length of transcribed (exonic) 676 and non-exonic (NE) sequences were summed in the intervening segments. The equations 677 above were compared as described for the observed exon/intron ratio verses gene length, exon 678 vs intron, and exonic vs NE in the respective figures. Randomization of only exons occurred by 679 utilizing the generated lengths of exons and randomly repositioning these segments throughout 680 the length of a chromosome. The remaining space was then defined as non-exonic. 681 Randomization of an exon with its associated intron occurred by random permutation of the 682 position of the shared elements along the generated lists for each chromosome. 683 Utilizing the Tissue Dashboard through the GTEx Portal, we subselected genes from the top 50 684 expressed within each respective tissues to represent different embryonic origins such that they 685 generally did not overlap with other tissue beds. The selected tissues were the cortex (ectoderm), 686 cardiac muscle (endoderm), and esophagus (endoderm). 687 EU-RNA-Seq for Nascent RNA Analysis 688 Sample Generation 689 Nascent RNA was labelled, captured, and sequenced following the protocol reported by Palozola 690 et al[97]. In brief, to selectively nascent RNA, HCT116 cells were treated with media containing 0.5 691 mM 5-ethynyluridine (EU) for 1 hour (Click-iT nascent RNA Capture Kit cat. no. c10365, Thermo 692 Fisher Scientific). These RNAs were retrieved using Click-iT chemistry to bind biotin azide the 693 ethylene group of EU-labeled RNA. The EU-labeled nascent RNA was purified using MyOne 694 Streptavidin T1 magnetic beads. Captured EU-RNA attached on streptavidin beads was 695 .CC-BY-NC 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 29, 2025. ; https://doi.org/10.1101/2025.05.29.656862doi: bioRxiv preprint 16 immediately subjected to on-bead sequencing library generation using the Universal Plus Total 696 RNA-Seq with NuQuant® (Tecan) according to the manufacturer’s protocols with modifications. 697 On-bead complementary DNA (cDNA) was synthesized by reverse transcriptase using random 698 hexamer primers. The cDNA fragments were then blunt-ended through an end-repair reaction, 699 followed by dA-tailing. Subsequently, specific double-stranded barcoded adapters were ligated 700 and library amplification for 15 cycles was performed. PCR libraries were cleaned up, measured 701 on an Agilent Bioanalyzer using the DNA1000 assay, pooled at equal concentrations and 702 sequenced on a Novogene Nova Seq X with 50 BP paired end. 703 Data Analysis 704 Paired end nascent RNA transcripts were trimmed with TrimGalore v0.6.10 and aligned using 705 Hisat2 v2.1.0. Alignment files were sorted, filtered, and converted to BAM format using Samtools 706 v1.6. Coverage files were generated using Deeptools v3.1.1 and normalized using RPGC 707 normalization. Finally, aligned filtered reads were counted using separate genome annotation 708 files for intron, exon, gene bodies, and intergenic regions with Htseq v2.0.2. Gene body coverage 709 plots were generated using the Superintronic package described in Lee et al[98] by binning each 710 gene body into 20 bins in the 5’->3’ orientation and calculating the average coverage in each bin 711 for short, average, and long genes within introns and on exons. 712 Estimation of compaction from nucleosome ‘beads on a string arrays’ compared to 713 observed sizes. 714 We consider the hypothesis that genes arrange as ‘open’ 10nm beads-on-a-string configurations 715 to facilitate the RNA synthesis from a gene body in its entirety. One can calculate the length of 716 the nucleosome array for thyroglobulin gene (TG) and Titan (TTN) compared to experimental 717 observations as follows[36]. TG is ~268,000bp and TTN is ~304,000 which converts to an array of 718 1340 and 1520 nucleosomes for 200bp increments, respectively. Using the diameter of 719 nucleosomes as 11nm, this produces chains that are 14.74 and 16.72 microns long for each 720 gene. The reported median inter-flank distance in transcriptionally active state for each gene was 721 observed as 703nm and 1104nm, respectively[36]. This 10-fold difference in length was accounted 722 for by the introduction of stiffness, indicating the need for selective compaction of these genes 723 even in their active state. A similar observation is observed within mouse neuronal cells in Rbfox1 724 (~1.527Mbp in mice, ~2.4Mbp in humans). In mice, Rbfox1 has a calculated chain length of ~84 725 microns (200bp/bead) as ‘beads on a string’ but an experimentally observed transcription start 726 site (TSS) to transcription end distance (TES) of ~1 micron[99]. 727 This indicates an 84-fold compression between the TSS-TES during active RNA synthesis. A 728 series of reaction volumes does not mean a gene does not undergo decompaction for 729 transcriptional activation. Instead, it indicates decompaction likely transitions from large domains 730 into a series of smaller domains. The smaller reaction volumes could be predictably encoded by 731 the spatial positioning between exons, introns, and intergenic segments to create efficient 732 reaction volumes. 733 Chromatin Connectivity, Enhancer, Epigenetic Modifications and Tier-1 Oncogenes: 734 The respective data of Chromatin interaction analysis with paired end tag (ChiaPET), Chromatin 735 Immunoprecipitation with Sequencing (ChIP-Seq), and RNA-Sequencing Analysis (RNA-Seq) 736 were all obtained from ENCODE[7–9]. The respective bed and bedpe files used were uploaded and 737 listed in SI Table 3 . The analysis for composition of ChiaPET loops were described as above 738 where the total content in both orientations were considered for each chromosome. For analysis 739 of element overlap with hinge positions, the identified 300bp hinges were used and the likelihood 740 of any portion overlapping with a histone modification peak or enhancer peak was calculated. As 741 a control variation in the read lengths covered by each element, the hinges observed when exons 742 were randomly scrambled across the chromosome length were used. A two-tailed t-test 743 .CC-BY-NC 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 29, 2025. ; https://doi.org/10.1101/2025.05.29.656862doi: bioRxiv preprint 17 comparing the observed frequency of overlap of the feature with a hinge per chromosome was 744 compared to the observed frequency in the hinges of randomly generated segments. 745 RefSeq gene positions used to define the start/stop position of gene annotations as described in 746 Almassalha et al[53]. Using the .bed files from ENCODE as above, we identified the mean location 747 of each peak for the different marks with a p-value cut-off of at least <0.1. For each cell line, we 748 then organized the Pol II-PS5 peaks by its density within each gene ranging from at least 1 to 10. 749 We then calculated the distance of the Pol II-PS5 peak to the nearest peak of the respective 750 histone mark. Using the cumulative density of distance, we calculated the median distance as a 751 function of the number of Pol II-Ps5 peaks on a gene body. With respect to the per chromosome 752 analysis, the total segment length of each mark was calculated for every somatic chromosome. 753 We then normalized for the difference in chromosome length and plotted the coverage of Pol II-754 PS5 against each respective chromatin mark. In the case of human tissues, we instead plotted 755 the association between heterochromatin and euchromatin since active RNA polymerase data 756 was not available. 757 Analysis of oncogene mutational frequency compared to hinge positioning was performed by 758 restricting the hinge length to 300bp or smaller (~1-2 nucleosomes). The position of the hinges 759 were mapped to protein coding genes. We then calculated the fraction of genes from a Tier-1 760 oncogene from the data within Mendiratta et al [70]. We then grouped the oncogenes into 761 ascending groups and calculated the mean mutational frequency observed for each group. The 762 values were then plotted compared to the observed frequency of hinges within genes within that 763 group. 764 Multi-color Single Molecule Localization Microscopy 765 Cell culture 766 HCT116 cells (ATCC, #CCL-247) were cultured in McCoy’s 5A Modified Medium (Thermo Fisher 767 Scientific, #16600-082, Waltham, MA). The cell media were supplemented with 10% fetal bovine 768 serum (FBS; Thermo Fisher Scientific, #16000-044, Waltham, MA) and 100 μ g/ml penicillin-769 streptomycin (Thermo Fisher Scientific, #15140-122, Waltham, MA). Cells were cultured under 770 standard conditions at 37°C in a humidified atmosphere with 5% CO2. Following detachment by 771 trypsinization, cells were allowed to re-adhere and recover for at least 24 hours before further 772 handling. Imaging was conducted when cell surface confluence ranged between 40–70%. For 773 this study, cells were used between passages 5 and 20. 774 Dual-color SMLM for EdU and histone modification 775 Cell Preparation and Fixation 776 After 48 hours from being seeded, the cells were incubated with EdU for 2 hours, followed by 777 fixation for 10 minutes at room temperature with 4% paraformaldehyde in PBS. Fixed samples 778 were washed three times in PBS for 5 minutes each. Secondary staining with AF647 via click-779 reaction chemistry was then performed according to the manufacturer’s protocol (ThermoFisher). 780 781 Permeabilization and Primary Antibody Staining 782 Samples were permeabilized and blocked using a buffer containing 3% bovine serum albumin 783 (BSA) and 0.5% Triton X-100 in PBS for 1 hour. They were then incubated with rabbit anti-784 H3K9me3 or mouse anti-H3K27me3/H3k4me3 (Abcam) diluted in blocking buffer for 1–2 hours at 785 room temperature on a shaker. Samples were washed three times in a washing buffer composed 786 of 0.2% BSA and 0.1% Triton X-100 in PBS. 787 .CC-BY-NC 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 29, 2025. ; https://doi.org/10.1101/2025.05.29.656862doi: bioRxiv preprint 18 788 Secondary Antibody Staining and Imaging 789 The samples were incubated with goat anti-rabbit or goat anti-mouse AF488 (ThermoFisher) for 790 40–60 minutes at room temperature on a shaker. After incubation, samples were washed twice in 791 PBS for 5 minutes each. Imaging was performed immediately, acquiring 10,000 frames per 792 channel for each cell. 793 Dual-color SMLM for EdU and BrdU 794 After 48 hours of seeding, the cells were incubated with EdU and BrdU simultaneously for 2 795 hours. This was followed by fixation for 10 minutes at room temperature using 4% 796 paraformaldehyde in PBS. The fixed samples were then washed three times with PBS, each 797 wash lasting 5 minutes. BrdU and EdU staining was performed according the BrdU and EdU 798 double staining protocol provided by Thermofisher. 799 SMLM reconstruction 800 The raw SMLM images were reconstructed using the built-in Thunder-STORM plugin in ImageJ. 801 The camera setup parameters for the plugin were tailored to the imaging configuration. In our 802 setup, we used a pixel size of 110 nm (calculated as the camera pixel pitch divided by the 803

Objective

magnification), photoelectrons per A/D count of 1.09, and a base level of 0. The peak 804 intensity threshold coefficient was adjusted based on the acquisition quality, typically ranging from 805 1 to 2. 806 Chromosome Paint 807 Human myoblasts were differentiated into myoblasts and chromosome painting was performed 808 on myotubes as described previously [29,39]. Images were acquired on a Nikon Confocal 809 Microscope or a Zeiss LSM 800 Confocal microscope. Imaging was done with a Plan Apo VC 100 810 × 1.4 NA oil objective as a multidimensional z-stack. The acquired 3D image stacks were then fed 811 through imaging processing pipelines utilizing standard tools on Cell Profiler and Fiji pipelines 812 performed the functions of translating images to maximal projections, calculating distance, object 813 size, nuclear size, and radius[100,101]. 814 Declaration of generative AI and AI-assisted technologies in the writing process 815 During the preparation of this work the author(s) used ChatGPT in order to ensure readability. 816 After using this tool/service, the author(s) reviewed and edited the content as needed and take(s) 817 full responsibility for the content of the publication. 818 819 Data Sharing Plan: 820 The generated analysis code and data will be made available upon reasonable request. Public 821 sources of data are included in the summary tables for reference. 822 Author Contributions: 823 Conceptualization: LMA, KLM, MC, CD, IS, VB 824 Writing – original draft: LMA, KLM 825 Writing – review & editing: LMA, KLM, MC, CD, RG, JI, LMC, WSL, RN, PSD, IS, VB 826 Methodology: LMA, RG, IS, VB 827 Resources: LMA, KLM, PSD, IS, VB 828 .CC-BY-NC 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 29, 2025. ; https://doi.org/10.1101/2025.05.29.656862doi: bioRxiv preprint 19 Funding Acquisition: LMA, KLM, PSD, IS, VB 829 Supervision: IS, VB 830 Project Administration: IS, VB 831 Investigation: LMA, KLM, RG, JI, LMC, WSL 832 Validation: LMA 833 Formal Analysis: LMA, VB 834 Visualization: LMA, KLM 835 Software: LMA 836 Data Curation: LMA 837 838 Competing Interest Statement: The authors declare no financial interests or consulting 839 interests related to this work. 840 841 Acknowledgments 842 We appreciate the thoughtful discussion and feedback on this manuscript from Dr. Sui Huang. 843 We appreciate the generous contributions from the funding sources listed below including the 844 National Institutes of Health, National Science Foundation, Hyundai Hope on Wheels, Alex’s 845 Lemonade Stand Foundation, Northwestern University, and Ann and Robert Lurie Children’s 846 Hospital. Computational analysis of ATAC-Seq and EU-Seq data was supported in part through 847 the computational resources and staff contributions provided by the Genomics Compute Cluster, 848 which is jointly supported by the Feinberg School of Medicine, the Center for Genetic Medicine, 849 and Feinberg’s Department of Biochemistry and Molecular Genetics, the Office of the Provost, 850 the Office for Research, and Northwestern Information Technology. The Genomics Compute 851 Cluster is part of Quest, Northwestern University’s high-performance computing facility, with the 852 purpose to advance research in genomics. We appreciate the generous support from the 853 ENCODE Consortium in the generation and dissemination of publicly available datasets. We 854 specifically want to thank the labs of J Michael Cherry, Charles Lee, Richard Myers, Bradley 855 Bernstein, Thomas Gingeras, Barbara Wold, Peggy Farnham, Michael Snyder, and John 856 Stamatoyannopoulo for the generation and publication of the data utilized within this manuscript. 857 We similarly appreciate the Atlas of Enhancers for their generation and dissemination of 858 enhancer datasets. 859 860 Funding: 861 National Science Foundation grant EFMA-1830961 (MC, IS, VB) 862 National Science Foundation grant EFMA-1830969 (VB) 863 National Science Foundation grant CBET-2430743 (VB) 864 National Institutes of Health grant R01CA228272 (WSL, IS, VB) 865 National Institutes of Health grant U54 CA268084 (WSL, LMC, LMA, KLM, MC, IS, VB) 866 National Institutes of Health grant U54 CA261694 (LMC, VB) 867 National Institutes of Health grant U01DK134321 (PSD) 868 National Institutes of Health grant R01DK135620 (PSD) 869 NIH Training Grant T32AI083216 (LMA) 870 .CC-BY-NC 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 29, 2025. ; https://doi.org/10.1101/2025.05.29.656862doi: bioRxiv preprint 20 NIH Training Grant T32GM132605 (LMC, VB) 871 Hyundai Hope on Wheels Hope Scholar Grant (KLM) 872 CURE Childhood Cancer Early Investigator Grant (KLM) 873 Alex’s Lemonade Stand Foundation ‘A’ award grant (KLM) #23-28271 874 Ann and Robert H. Lurie Children’s Hospital of Chicago under the Molecular and 875 Translational Cancer Biology Neighborhood (KLM) 876 National Institutes of Health National Center for Advancing Translational Sciences 877 KL2TR001424 (KLM) 878 Northwestern University Starzl Scholar Award (LMA) 879 880 881

References

882 [1] J. M. H ernán dez-He rnánd ez, E. G . Ga rcía- González, C. E. Bru n, M. A . Rudnicki, Sem in Cell 883 Dev Biol 2017 , 72, 10. 884 [2] T. Gao, J. Qian , Nu cleic A cids Res 2019 . 885 [3] J. M. B rown, N . A. Ro ber ts, B. Grah am, D. Waithe , C. Lagerholm, J . M. Tel enius, S. De 886 Ornell as, A. M . Oud elaa r, C. Scott , I. Szcz erbal, C. Ba bbs, M. T. Kassouf, J . R. Hugh es, D. R. Higgs, 887 V. J. Buckle, N at Co mmu n 2018 , 9 , 3849. 888 [4] T.-H. S. Hsieh, C. Cat toglio, E. Slob odyanyuk, A. S. Hans en, X. Darzacq, R . Tjian, Na t Gene t 889 2022 , 54, 1919. 890 [5] W. S. Li, L. M. Car ter , L. M. Almassal ha, R . Gong, E. M . Pujadas-Liwag, T. Kuo, K. L. 891 MacQuarri e, M . Carignano, C. Dunton , V. Dravid, M. T. Kanemaki, I . Szleifer , V. Ba ckman, Sci Adv 892 2025 , 11. 893 [6] L. M. Almassalha , M. Carignano , E. P. Liwag, W. S. Li, R. Gong, N . Acost a, C. L. Dunton, P. 894 C. Gonzalez, L. M. Car ter , R. Kakkaramad am, M. Kröger , K. L. MacQuar rie, J. Fr ede rick, I. C. Ye, P. 895 Su, T. Kuo, K. I. M edina, J. A . Pritchard , A. Skol, R. Nap, M . Kanemaki, V. Dravid, I . S zleifer, V. 896 Backman, Sci Adv 2025 , 11. 897 [7] B. C. Hitz, J.-W. Le e, O . Jola nki, M. S. Kag da, K. Gr aham, P. Sud, I . Gab dank, J . S. St rat tan, 898 C. A. Sloan, T. Dresz er, L. D. Rowe , N. R . Podduturi , V. S. Mall adi, E. T. Chan, J . M. D avidson, M. 899 Ho, S. Miyasat o, M. Simison , F. Tanaka, Y . Luo, I. Wh aling, E. L. Hong, B . T. Lee, R. Sandstrom, E . 900 Rynes, J. N elson, A. Nishid a, A. Ingersoll , M. Buckley, M. Fr erke r, D. S. Kim, N. B ole y, D. Trout, A. 901 Dobin, S. Rahmania n, D. Wyman, G . Bald errama- Guti err ez, F. R eese , N. C. Durand , O. 902 Dudchenko, D. Weisz, S . S. P. Rao, A. Blac kburn, D. Gkount aroulis , M. Sad r, M. Ols hansky, Y. 903 Eliaz, D. Nguyen, I . Bochkov, M. S. Sh amim, R. Mahajan, E. Aiden , T. Ginger as, S. H eath, M. Hirs t, 904 W. J . Kent, A . Kundaje, A . Mor taz avi, B. Wold, J . M. Cher ry, . 905 [8] An integr ate d encyclopedia of DNA el em ents in th e human genome , Na ture 2012 , 489 , 906 57. 907 .CC-BY-NC 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 29, 2025. ; https://doi.org/10.1101/2025.05.29.656862doi: bioRxiv preprint 21 [9] Y. Luo, B. C. Hitz, I. G abdank, J . A. Hilton, M. S. Kagda, B . Lam, Z. Myers, P. Sud, J . J ou, K. 908 Lin, U. K. Baymuradov, K. G raham, C. Lit t on, S. R. Miyasa to, J. S. S tra tt an, O . Jola n ki, J.-W. Le e, F. 909 Y. Tanaka, P. Aden ekan, E. O ’N eill, J . M. Cherry, N uclei c Aci ds Res 2020 , 48, D882. 910 [10] C. Cruchaga, Arch N eurol 2011 , 68, 1013. 911 [11] A. Lipov, S. J . Jurgens , F. Maz zaro tto , M. Allouba, J. P. Pirruccello , Y. Aguib, M. 912 Genna relli, M . H. Yacoub , P. T. Ellinor, C. R. Bezzin a, R. Walsh, N at ure Cardi ovasc u lar Researc h 913 2023 , 2 , 1078. 914 [12] I. A. E. Bollen , M. Schuld t, M. Ha rakalova, A. Vink, F. W. Asse lbergs, J. R. Pint o, M. 915 Krüger, D. W. D. Kust er, J . van der Veld e n, J Physiol 2017 , 595 , 4677. 916 [13] T. Maruth appu, C. Sco tt, D. Kelsell , Ge ne s (Basel) 2014 , 5 , 615. 917 [14] E. Miron, R . Olde nkamp, J. M . Brown , D. M. S. Pinto, C. S . Xu, A. R. F aria, H . A. Sh a ban, J. 918 D. P. Rhodes, C. Innoce nt, S . de Or nellas, H. F. Hess, V. Buckle, L. Sch ermell eh, Sci Adv 2020 , 6 . 919 [15] Y. Li, A. Eshein, R. K. A. Virk, A. Eid, W. W u, J. Fr ederick, D. VanDerway, S. Glads tei n, K. 920 Huang, A. R . Shim, N. M . An thony, G . M. Bauer , X. Zhou, V. Agrawal, E. M . Pujadas, S. Jain, G. 921 Esteve, J . E. Chandler , T.-Q. Nguyen, R . Bl eher, J. J . de Pablo, I. Szl eifer, V. P. Dravid, L. M. 922 Almassalha, V. Backman, Sci A dv 2021 , 7 . 923 [16] Y. Li, V. Agrawal, R. K. A . Virk, E. Roth , W. S. Li, A. Eshein, J . Fre derick, K. Huang, L. 924 Almassalha, R . Bleh er, M . A. Carignan o, I. Szleifer, V. P. Dravid, V. Backman, Sci Re p 2022 , 12, 925 12198. 926 [17] E. Lieberman-Aid en, N . L. van Berkum, L. Williams, M. Imaka ev, T. Ragoczy, A. Telli ng, I. 927 Amit, B. R . Lajoie, P. J . Sabo, M . O. Dorsc hner, R . Sandst rom, B. B erns tein, M . A. B ender , M. 928 Groudin e, A . Gnirk e, J . Stama toyannop o ulos, L. A. Mirny, E. S . Lander , J. Dekker , Science (1979) 929 2009 , 326 , 289. 930 [18] S. S. P. Rao, M . H. Hun tley, N . C. Durand, E. K. Stamenova , I. D. Bochkov, J . T. Robi nson, 931 A. L. Sanbo rn, I . Machol, A . D. Ome r, E. S. Lander, E. L. Aid en, Cell 2014 , 159 , 1665. 932 [19] S. S. P. Rao, S .-C. Huang, B. Glenn S t Hilai re, J . M. Engrei tz, E. M . Perez , K.-R. Kieffe r-933 Kwon, A. L. Sanborn , S. E. J ohnsto ne, G. D. Bascom, I. D. Bochkov, X. Huang, M . S. Shamim, J. 934 Shin, D. Turner, Z. Ye , A. D. Ome r, J . T. Ro binson, T. Schlick, B. E. Be rnst ein, R. Case llas, E. S. 935 Lander, E. L. Aid en, Cell 2017 , 171 , 305. 936 [20] B. Bintu , L. J. M ate o, J .-H. Su, N . A. Sinn ot t-Armstr ong, M. Parke r, S. Kinro t, K. Yam aya, 937 A. N. Boe ttige r, X. Zhuang, Scien ce (1979) 2018 , 362 . 938 [21] A. N. Boe ttige r, B. Bin tu, J. R. M offitt, S. Wang, B. J. Belive au, G . Fuden berg, M . I makaev, 939 L. A. Mirny, C. Wu, X. Zhuang, Nat ure 2016 , 529 , 418. 940 [22] A. Hafner , M. Park, S. E. Berge r, S. E. Mu r phy, E. P. Nora, A . N . Boe ttige r, Mol Cell 2023 , 941 83, 1377. 942 .CC-BY-NC 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 29, 2025. ; https://doi.org/10.1101/2025.05.29.656862doi: bioRxiv preprint 22 [23] Q. Szabo, A. Donjon, I . Je rković, G. L. Pap adopoulos, T. Cheu tin, B . Bon ev, E. P. No ra, B. 943 G. Bru neau , F. Ban tignies, G. Cavalli, Na t Gene t 2020 , 52, 1151. 944 [24] Q. Szabo, F . Ban tignies, G . Cavalli, Sci Adv 2019 , 5 . 945 [25] J. O tt erstr om, A. Cast ells-Ga rcia, C. Vicari o, P. A. Gom ez-Ga rcia, M. P. Cosma, M . 946 Lakadamyali, Nucl eic A cids Res 2019 , 47, 8470. 947 [26] A. Castells-G arcia, I. Ed-daoui, E. Gonz ále z-Almela, C. Vicario, J . O tt estrom, M. 948 Lakadamyali, M. V. Negu embor, M . P. Cosma, Nuclei c Ac ids Res 2022 , 50, 175. 949 [27] M. V. Neguemb or, J . P. Arcon , D. Buitr ago, R. Lema, J . Walthe r, X. Ga rat e, L. Ma rti n, P. 950 Romero, J. AlH aj Abed, M . Gu t, J . Blanc, M. Lakadamyali, C. Wu, I . Brun H eat h, M. Orozco, P. D. 951 Dans, M. P. Cosma, Nat Str uct Mol Biol 2022 , 29, 1011. 952 [28] M. A. Ricci, C. Man zo, M . F. Ga rcía-Parajo , M. Lakadamyali, M. P. Cosma, Cell 2015 , 160 , 953 1145. 954 [29] E. M. Pujadas Liwag, X. Wei, N. Ac osta, L. M. Carte r, J . Yang, L. M. Almassalh a, S. J ain, A. 955 Daneshkhah, S. S. P. R ao, F. S eker-Polat , K. L. MacQuarri e, J . Iba rra, V. Ag rawal, E. L. Aiden, M . T. 956 Kanemaki, V. Backman, M. Adli, Ge no me Biol 2024 , 25, 77. 957 [30] X. Wang, V. Agrawal, C. L. Dunton , Y. Liu, R. K. A. Virk, P. A. Pa tel, L. Car ter, E . M. 958 Pujadas, Y. Li, S. Jain, H. Wang, N . Ni, H.- M. Tsai, N. Rive ra-Bola nos, J . Fred erick, E . Roth, R . 959 Bleher , C. Duan, P. Ntzi achristos , T. C. He , R. R. Rei d, B. Jiang, H. Sub ramania n, V. Backman, G. A . 960 Ameer , Na t Biom ed Eng 2023 , 7 , 1514. 961 [31] J. Hong, A . D. Cavga, D. Shah, E. Laue, J . Taipale, Sc aling l aws of huma n tra nscrip ti onal 962 acti vity , 2023 . 963 [32] K. Maeshima, K. Kaizu , S. Tamura, T. Noz aki, T. Kokubo, K. Takahashi, J our nal of Physics: 964 Conde nsed M att er 2015 , 27, 064116. 965 [33] I. Solovei, A. Cavallo, L. Sche rmelleh , F. Ja unin, C. Scasselati , D. Cmarko, C. Cremer, S. 966 Fakan, T. Cremer , Exp Cell Res 2002 , 276 , 10. 967 [34] J. M. B rown, S. De O rnellas , E. Parisi, L. S chermelleh , V. J. Buckle , Na t Proto c 2022 , 17, 968 1306. 969 [35] A. R. Shim, J. Fr ederick, E. M . Pujadas, T. Kuo, I. C. Ye, J . A. Pritch ard, C. L. Dunton , P. C. 970 Gonzal ez, N . Acost a, S. J ain, N . M. A ntho ny, L. M. Almassalha, I. Szleif er, V. Backm an, PLoS One 971 2024 , 19, e0301000. 972 [36] S. Leidescher , J. Ri bisel, S. Ul lrich, Y. Fe od orova, E. Hildeb rand, A. G alitsyna, S . Bult mann, 973 S. Link, K. Thanisch, C. Mulholland, J. Dek ker, H. Leonh ardt , L. Mirny, I . Solovei, Na t Cell Biol 974 2022 , 24, 327. 975 [37] L. M. Almassalha , A. Tiwari, P. T. Ruh off, Y. Stypula-Cyrus, L. Cherkezyan, H. M ats uda, M. 976 A. Dela Cruz, J . E. Chandle r, C. Whit e, C. Maneval, H. Su bramani an, I . Szleifer , H. K. Roy, V. 977 Backman, Sci Rep 2017 , 7 , 41061. 978 .CC-BY-NC 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 29, 2025. ; https://doi.org/10.1101/2025.05.29.656862doi: bioRxiv preprint 23 [38] L. A. Mirny, Chr omoso me Resear ch 2011 , 19, 37. 979 [39] J. Iba rra , T. Hersh enhouse , L. Almassalha , D. Walte rhouse , V. Backman, K. L. 980 MacQuarri e, Fro nt Cell Dev Biol 2023 , 11 . 981 [40] T. Sabat é, B. Lelandais, M .-C. Rober t, M . Szal ay, J.-Y. Tineve z, E. Be rtr and, C. Zimmer, 982 bioRxi v 2024 , 2024.08.09.605990 . 983 [41] M. Gab riele, H. B . Brand ão, S. Gross e-Holz, A. Jh a, G . M. Dail ey, C. Cat togl io, T.-H. S. 984 Hsieh, L. Mirny, C. Zechner , A. S . Hansen, Science (1979) 2022 , 376 , 496. 985 [42] W. F. Doolit tle , C. Sapienz a, Na ture 1980 , 284 , 601. 986 [43] L. E. Orgel, F. H . C. Crick, Nat ure 1980 , 284 , 604. 987 [44] S. Ohno, Br ookh ave n Symp Bi ol. 1972 , 23 , 366. 988 [45] L. Stat ello, C.-J . Guo, L.-L. Chen, M. Huar t e, Na t Rev M ol Cell Biol 2021 , 22, 96. 989 [46] R. Shang, S. Le e, G . Senavira thn e, E. C. Lai, Nat R ev Ge net 2023 , 24, 816. 990 [47] K. V. Morris, J . S. Ma ttick, N at R ev Ge net 2014 , 15, 423. 991 [48] M. A. Carignan o, M. Kr oeger , L. M. Almas salha, V. Agrawal, W. S. Li, E. M . Pujadas-Liwag, 992 R. J. Nap, V. Backman, I. Szleifer, Elife 2024 , 13. 993 [49] A. Bancau d, S. Hue t, N . Daigle, J . Mozzico nacci, J. Be audouin , J. Elle nberg, EM BO J 2009 , 994 28, 3785. 995 [50] K. Metz e, Exp ert Rev M ol Diag n 2013 , 13, 719. 996 [51] N. A. O’Lea ry, M. W . Wright , J. R . Bris ter , S. Ciufo, D. Haddad, R. McVeigh, B . Rajpu t, B. 997 Robber tse, B . Smith-Whi te, D. Ako-A djei, A. Astashyn, A. Bad re tdin, Y. Ba o, O . Blin kova, V. 998 Brover, V. Che tvernin, J. Choi, E. Cox, O. Ermolaeva, C. M. Fa rrell , T. Goldfa rb, T. Gupta , D. Haft, 999 E. Hatche r, W. Hl avina, V. S. J oarda r, V. K. Kodali, W. Li, D. Maglo tt , P. Maste rson, K. M. 1000 McGarvey, M . R. Mur phy, K. O’N eill, S . Pujar, S. H. Ra ngwala, D. Rausch, L. D. Rid dick, C. Schoch, 1001 A. Shkeda , S. S. S torz, H . Sun, F. Thiba ud-Nissen, I . Tolstoy, R. E. Tully, A . R. Vatsa n , C. Wallin, D. 1002 Webb, W . Wu, M . J . Landrum, A . Kimchi, T. Tatusova, M. DiCuccio, P. Kitts, T. D. M urphy, K. D. 1003 Pruitt, N uclei c Aci ds Res 2016 , 44, D733. 1004 [52] W. S. Li, L. M. Car ter , L. M. Almassal ha, R . Gong, E. Pujadas Liwag, T. Kuo, K. L. 1005 MacQuarri e, M . Carignano, C. L. Dunton , V. Dravid, M. Kanemaki, I . Szleifer , V. Bac kman, Sci Adv 1006 2025 . 1007 [53] L. M. Almassalha , M. Carignano , E. Pujadas Liwag, W. S. Li, R. Gong, N. Acos ta, C. L . 1008 Dunton, P. C. Gonz alez, L. M . Cart er, R . Kakkaramadam, M . Kroger, K. L. M acQuar rie, J . 1009 Frede rick, I. C. Ye, P. Su, T. Kuo, K. M edin a, J. A . Pritcha rd, A. Skol, R. N ap, M. Kan e maki, V. 1010 Dravid, I. Szleife r, V. Backman, Sci A dv 2025 . 1011 .CC-BY-NC 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 29, 2025. ; https://doi.org/10.1101/2025.05.29.656862doi: bioRxiv preprint 24 [54] M. V. Neguemb or, L. Ma rtin , Á. Cast ells-García, P. A . Góm ez-Ga rcía, C. Vicario, D. 1012 Carnevali, J . AlHaj Abed , A. Gran ados, R. Sebasti an-Perez, F . Sot tile, J. Solo n, C. tin g Wu, M. 1013 Lakadamyali, M. P. Cosma, Mol Cell 2021 , 81, 3065. 1014 [55] R. Janisse n, R. B arth , M. Polinde r, J . van der Torr e, C. Dekker, N uclei c Aci ds Res 2024 , 52, 1015 1677. 1016 [56] A. Ma ttou t, A . Biran , E. Mesho rer , J Mol Cell Biol 2011 , 3 , 341. 1017 [57] E. Fussner, U . Djuric, M. Stra uss, A. Ho tt a , C. Perez-Ira tx eta , F. Lanne r, F. J. Dilwort h, J. 1018 Ellis, D. P. Bazett- Jones , EMBO J 2011 , 30 , 1778. 1019 [58] Y. Kawamura, S. Koyama, R. Yoshid a, Bioi nforma tics 2019 , 35, 1877. 1020 [59] F. C. Oest err eich, L. Her zel, K. S traub e, K. Hujer, J. Howa rd, K. M. Neugeb aue r, Cell 2016 , 1021 165 , 372. 1022 [60] H. Shenasa , D. L. Bentley, Trends G ene t 2023 , 39, 672. 1023 [61] Y. Zeng, B. J. Fa ir, H. Zeng, A . Krishnamoh an, Y. Hou, J . M. Hall , A. J. Ru thenbu rg, Y. I. Li, J. 1024 P. Staley, Mol Cell 2022 , 82, 4681. 1025 [62] G. Li, L. Cai, H. Chang, P. Hong, Q. Zhou, E . V Kulakova, N. A. Kolchan ov, Y. Ruan, BM C 1026 Geno mics 2014 , 15, S11. 1027 [63] L. P. Kaelbling, M. L. Littman , A. W . Moo r e, Jo urna l of Artificial In tellige nce R esear ch 1028 1996 , 4 , 237. 1029 [64] M. van Ot terlo, M. W iering, 2012 , pp . 3– 42. 1030 [65] J. Lonsdale , J. Thomas , M. Salva tore , R. Phillips, E. Lo, S. Shad, R . Hasz, G. Wal te rs, F. 1031 Garcia, N. Young, B . Fost er, M . Moser , E. Karasik, B. Gill ard, K. R amsey, S. Sullivan, J. Bridg e, H. 1032 Magazine, J. Syron , J. Fl eming, L. Siminoff, H. Traino, M. M osavel, L. Ba rker, S . Jew ell, D. Rohr er, 1033 D. Maxim, D. Filkins, P. Harbach, E . Corta dillo, B. B erghuis, L. Turne r, E. Hudson , K. Feenstr a, L. 1034 Sobin, J. R obb, P. Br anton , G. Ko rzeni ewski, C. Shive, D. Tabor, L. Qi, K. Gr och, S. N ampally, S. 1035 Buia, A. Zimmerman , A. Smit h, R. Bu rges, K. Robinson, K. Valen tino, D. Br adbury, M. Cosentino , 1036 N. Diaz-Mayoral , M. Kenn edy, T. Engel, P. Williams, K. Erickson, K. Ard lie, W . Winc kler, G . Ge tz, 1037 D. DeLuca, D. MacArthu r, M. Kellis, A. Th omson, T. Young, E. Gelfan d, M. Donova n, Y. Meng, G . 1038 Gran t, D. Mash, Y. M arcus, M . Basile, J. Li u, J. Zhu, Z. Tu, N. J. Cox, D. L. Nic olae, E . R. Gamaz on, 1039 H. K. Im, A. Konkashb aev, J . Pritchard , M. Stevens, T. Flut re, X. Wen, E. T. De rmitza kis, T. 1040 Lappalainen , R. Guigo , J. M onlong, M . Sa mmeth, D. Koller, A. Ba ttl e, S. M ostafavi, M. McCarthy, 1041 M. Rivas, J. M aller , I. Rusyn, A. No bel, F . Wright, A . Shaba lin, M. F eolo, N. Sha rop ova, A. Stu rcke, 1042 J. Paschal, J . M. And erson , E. L. Wilde r, L. K. Derr, E. D. Gr een, J. P. St ruewing, G . Temple, S. 1043 Volpi, J. T. Boye r, E. J . Thomson, M. S . Gu yer, C. Ng, A. A bdallah , D. Colantuoni , T. R. Insel, S . E. 1044 Koester , A. R . Little, P. K. Bend er, T. Lehn er, Y. Yao, C. C. Compton, J. B . Vaught, S. Sawyer, N. C. 1045 Lockhart, J . Demchok, H. F. Mo ore, Na t G enet 2013 , 45, 580. 1046 [66] K. Takahashi, S. Yamanaka , Na t Rev M ol Cell Biol 2016 , 17, 183. 1047 [67] J. C. Pearson , D. Lemons, W. Mc Ginnis, N at Rev Gene t 2005 , 6 , 893. 1048 .CC-BY-NC 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 29, 2025. ; https://doi.org/10.1101/2025.05.29.656862doi: bioRxiv preprint 25 [68] C. R. Goding, H . Arnh eit er, Ge n e s D e v 2019 , 33, 983. 1049 [69] W.-J. Kim, H.-L. Shin , B.-S. Kim, H.-J . Kim, H.-M. Ryoo, Ex p Mol Me d 2020 , 52, 1178. 1050 [70] G. Me ndira tta , E. Ke, M. Aziz, D. Liarakos, M. Tong, E. C. Stit es, Na t Com mun 2021 , 12, 1051 5961. 1052 [71] Z. Sondka, S. Bamford , C. G. Cole, S . A. W ard, I . Dunham, S. A. F orbes , Na t Rev Ca n cer 1053 2018 , 18, 696. 1054 [72] G. S. Gerh ard, Ex p Ger on tol 2003 , 38, 13 33. 1055 [73] A. Delamar re, B . Bailey, J. Yavid, R. Koche , N. Mohib ullah, I . Whi tehous e, C h r om at in 1056 archi tec ture m ap pin g by m ulti plex pr oxi mity t ag gin g , 2024 . 1057 [74] X. Zhang, D. Shu, PalZ 2021 , 95, 641. 1058 [75] M. Roy, N. Kim, Y. Xing, C. Lee, RN A 2008 , 14, 2261. 1059 [76] I. Ruiz-Trillo, K. Kin, E. Casacub ert a, A nnu Rev Micro biol 2023 , 77, 499. 1060 [77] J.-H. Lee , E. W. Kim, D. L. Crot eau, V. A . B ohr, Exp M ol Med 2020 , 52, 1466 . 1061 [78] V. R. Kedlian, Y. W ang, T. Liu, X. Chen, L. Bolt, C. Tudor , Z. Shen, E. S . Fasouli, E. 1062 Prigmore, V. Kleshchevnikov, J. P. Pet t, T. Li, J. E. G. Lawre nce, S. Per era , M. Pre te, N. Huang, Q . 1063 Guo, X. Zeng, L. Yang, K. Polański, N.- J. Chipampe, M . Dabrowska, X. Li, O. A . Bayr aktar, M . Patel , 1064 N. Kumasaka, K. T. Mah bubani, A. P. Xian g, K. B. Meyer , K. Saeb-Parsy, S. A . Teich mann, H. 1065 Zhang, Nat A gin g 2024 . 1066 [79] R. U. Path ak, M. Soujanya, R . K. Mishra , Agein g Res Rev 2021 , 67, 101264. 1067 [80] D. Nicett o, K. S. Zar et, Curr O pin G ene t Dev 2019 , 55, 1. 1068 [81] Y. Liu, J. Dekker, N at Cell Bi ol 2022 , 24, 1516. 1069 [82] J. T. Sand ers, R . Golloshi , P. Das, Y. Xu, P. H. Terry, D. G . Nash, J. Dekker, R . P. McCord, 1070 Sci Rep 2022 , 12, 4721. 1071 [83] J. Xu, J . Duan, Z. Cai, C. Arai, C. Di, C. C. Vente rs, J . Xu, M. Jo nes, B .-R. So, G. Dreyfuss, 1072 TOMM40-APOE chimera linki ng Alz heim er’s highes t risk genes: a new pa thway fo r mitoc ho ndria 1073 regula tio n an d APOE4 pat ho genesis , 2024 . 1074 [84] K. Pabis, D. Barardo , O. Si rbu, K. Selvar ajoo, J. Grub er, B . K. Kennedy, Elife 2024 , 12. 1075 [85] C. T. Stankey, C. Bourges , L. M. Haag, T. T urner-St okes, A . P. Piedade, C. Palmer- Jo nes, I. 1076 Papa, M. Silva dos Sant os, Q. Zhang, A . J. Cameron, A . Legrini, T. Zhang, C. S. Wo o d, F. N. New, L. 1077 O. Randz avola, L. Spei del, A . C. Brown, A . Hall, F. Saffioti, E. C. Parkes, W. Edwards , H. 1078 Direskeneli, P. C. Gr ayson, L. Jiang, P. A . Merkel, G. Sa ruhan-Diresken eli, A . H. Sa walha, E. 1079 Tombetti , A. Quaglia , D. Thorburn , J. C. K night, A. P. Rochfor d, C. D. Murray, P. Divakar, M. 1080 Gre en, E. Nye , J. I. MacR ae, N . B. Jamieso n, P. Skoglund, M. Z. Cader, C. W allace, D . C. Thomas, J. 1081 C. Lee, Na ture 2024 , 630 , 447. 1082 .CC-BY-NC 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 29, 2025. ; https://doi.org/10.1101/2025.05.29.656862doi: bioRxiv preprint 26 [86] C. Bailey, O. Pich, K. Thol, T. B . K. Watkin s, J. Luebeck, A . Rowan, G. St avrou, N . E. 1083 Weiser , B. Damer acharla , R. Be ntham, W .-T. Lu, J. Kitt el, S. Y. C. Yang, B . E. Howitt , N. Sharma , 1084 M. Litovchenko, R . Salgado, K. L. Hung, A. J. Cornish, D. A. M oor e, R. S . Houlsto n, V. Bafna, H. Y. 1085 Chang, S. Nik-Zainal, N. Kanu , N. Mc Gran ahan, J . C. Ambrose , P. Arumugam, R. Be vers, M. Bled a, 1086 F. Boardm an-Pret ty, C. R. Boust red, H . Br ittain , M. A . Brown, M . J. Caulfield , G. C. Chan, A. Gi ess, 1087 J. N . Griffin, A . Hamblin, S. H enders on, T. J. P. Hubbard , R. J ackson, L. J. Jones, D. K asperaviciut e, 1088 M. Kayikci, A. Kousath anas, L. Lahnst ein, A. Lakey, S. E. A . Leigh, I. U . S. Leong, F . J. Lopez, F. 1089 Maleady-Crowe, M . McEntagar t, F. Minn eci, J. Mi tchell, L. M outsian as, M. M uelle r, N. 1090 Murugaesu, A. C. Ne ed, P. O ’Donovan, C. A. Odhams, C. Patch, D. Per ez-Gil, M . B. Pereira , J. 1091 Pullinger, T. Rahim, A . Rendo n, T. Rogers , K. Savage, K. Sawant, R . H. Scot t, A . Siddi q, A. Siegha rt, 1092 S. C. Smith, A. S osinsky, A. Stuckey, M . Tanguy, A. L. Taylor Tavares, E. R. A. Thoma s, S. R. 1093 Thompson, A. Tucci, M. J . Well and, E. W ill iams, K. Witkowska, S. M . Wood , M. Zar owiecki, A. M. 1094 Flanagan, P. S. Misch el, M. Jamal-Hanjani , C. Swanton, Na tu r e 2024 , 635 , 193. 1095 [87] R. K. A. Virk, W . Wu, L. M . Almassalha, G. M. Baue r, Y. Li, D. VanDerway, J. Fr ede ri ck, D. 1096 Zhang, A. Eshein, H. K. R oy, I. Szleifer, V. Backman, Sci Adv 2020 , 6 . 1097 [88] L. M. Almassalha , G. M . Bau er, W . Wu, L. Cherkezyan, D. Zhang, A. Kend ra, S. Glad stein, 1098 J. E. Chandle r, D. VanDerway, B.-L. L. Se a gle, A. Ugolkov, D. D. Billade au, T. V. O’ H alloran, A. P. 1099 Mazar, H . K. Roy, I. S zleifer , S. Shah abi, V. Backman, N at B i om e d E ng 2017 , 1 , 902. 1100 [89] Z. Yang, K. L. MacQuarri e, E. Anal au, A . E. Tyler, F. J. Dilworth , Y. Cao, S. J . Diede, S . J. 1101 Tapscott, Genes Dev 2009 , 23, 694. 1102 [90] Y. Cao, Z. Yao, D. Sarkar, M . Lawrence , G. J. Sanchez , M. H. Park er, K. L. MacQu arri e, J. 1103 Davison, M. T. Morgan, W. L. Ruzz o, R. C. Gentleman, S. J. Tapsco tt, Dev Cell 2010 , 18, 662. 1104 [91] K. L. MacQuarri e, A . P. Fong, R. H. M orse, S. J. Tapscot t, Tren ds in Ge neti cs 2011 , 27, 1105 141. 1106 [92] S. J. H eo, S. Thakur , X. Chen, C. Loebel, B. Xia, R. McBea th, J . A. B urdick, V. B. Sh en oy, R. 1107 L. Mauck, M. Lakadamyali, N at Bi ome d Eng 2023 , 7 , 177. 1108 [93] L. M. Almassalha , G. M . Bau er, J . E. Chan dler, S. Glads tein, L. Cherk ezyan, Y. Stypu la-1109 Cyrus, S. Weinbe rg, D. Zhang, P. Thusgaard Ruhoff, H. K. Roy, H. Subr amanian , N. S. Chandel, I . 1110 Szleifer, V. Backman , Proceedi ngs of t he Nati onal Ac adem y of Scienc es 2016 , 113 . 1111 [94] F. R. Palma, D. R. Coelh o, K. Pulakanti, M . J. Sakiyama, Y. Huang, F . T. Ogat a, J . M. Danes, 1112 A. Meye r, C. M. Fur dui, D. R. Spi tz, A . P. Gomes, B . N. Gant ner , S. Rao , V. Backman, M. G . Bonini, 1113 Cell Rep 2024 , 43, 113897. 1114 [95] H. Matsud a, G . G. Put zel, V. Backman, I. Szleifer, Bi op hys J 2014 , 106 , 1801. 1115 [96] A. R. Shim, R . J. Nap, K. Hu ang, L. M. Alm assalha, H. M atusda , V. Backman, I . Szleif er, 1116 Bioph ys J 2020 , 118 , 2117. 1117 [97] K. C. Palozola, G . Donahue, K. S . Zare t, STAR Protoc 2021 , 2 , 100651. 1118 .CC-BY-NC 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 29, 2025. ; https://doi.org/10.1101/2025.05.29.656862doi: bioRxiv preprint 27 [98] S. Lee, A . Y. Zhang, S. Su, A . P. Ng, A. Z. H olik, M.-L. Asselin-Laba t, M. E. Ri tchie, C. W. 1119 Law, NAR Ge nom Bi oinform 2020 , 2 . 1120 [99] W. Winick-Ng, A. Kukalev, I. Ha rabul a , L. Zea-Redondo , D. Szabó, M . Meije r, L. Ser ebreni , 1121 Y. Zhang, S. Bianco, A . M. Chiari ello, I . Ir a storza-Azc ara te, C. J . Thieme, T. M . Spark s, S. Carvalho, 1122 L. Fiorillo, F. M usella, E. Irani , E. Torlai Tri glia, A. A . Kolodziejczyk, A. Ab entung , G. Apostol ova, E. 1123 J. Paul, V. Frank e, R. Kempfer , A. Aka lin, S. A. Teichmann , G. Dechan t, M . A. Ungl e ss, M. 1124 Nicodemi, L. W elch, G . Castelo-B ranco, A . Pombo, Nat ure 2021 , 599 , 684. 1125 [100] D. R. Stirling, M . J. Swain-B owden, A . M. Lucas, A. E. Carpen ter , B. A . Cimini, A. 1126 Goodman , BMC Bioi nforma tics 2021 , 22, 433. 1127 [101] J. Schindeli n, I. Arganda-Car reras , E. Frise , V. Kaynig, M. Longair, T. Pietzsch, S. Pr e ibisch, 1128 C. Rueden, S . Saalfeld , B. Schmid, J .-Y. Tinevez, D. J. Whit e, V. Har tens tein, K. Elic eiri, P. 1129 Tomancak, A. Cardon a, Na t Me tho ds 2012 , 9 , 676. 1130 1131 Figures and Tables 1132 1133 Figure 1. A paradox of the loss of accessibility with transcriptional amplification. A) 1134 DNAse-seq analysis demonstrates loss of accessibility with concurrent transcriptional activation 1135 during muscle differentiation. B-C) Loss of accessibility in non-exonic segments occurs across 1136 myogenic transcriptional activation. D) Transcriptional activity is associated with intronic 1137 heterochromatin deposition across tissue types. These are biallelic, constitutively expressed 1138 .CC-BY-NC 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 29, 2025. ; https://doi.org/10.1101/2025.05.29.656862doi: bioRxiv preprint 28 genes in their respective tissues. Critically, this includes TNNI3 (troponin), which is a component 1139 of the sarcomere necessary for cardiac contractility. E) Mass-fractal packing domains create a 1140 unified reaction-volume due to the continuous density gradient across domain layers. F) 1141 Transcription is non-monotonically dependent on local density due to the trade-off between 1142 entropic gain (remaining bound as an intermediate complex decreases the excluded volume to 1143 other macromolecules) and the diffusibility of reactant species. As density continues to increase, 1144 Pol II subunits can no longer penetrate deeper into a domain volume resulting in inhibition. 1145 1146 1147 1148 1149 1150 Figure 2. Packing domains as a geometric solution to optimize nuclear volume and 1151 transcriptional efficiency. A) Loops have a broad range of sizes with many larger loops 1152 (>100Kbp) generating lengths that would span the human nucleus without a system for efficient 1153 packing. B) Proposed framework that the position of exons, introns, and intergenic elements 1154 produces a system to reliably generate reaction volumes. Exons with short intronic sequences 1155 fold into an ideal zone within a volume generated by NE DNA (the ideal zone as a surface-area to 1156 volume - SA/V – of the total volume). The resulting volumes represent the structures observed on 1157 ChromSTEM imaging. The continued selection for elements across broad-timescales results in 1158 an encoding within the genome. C-E) Transformation from beads on a string into mass-fractal 1159 volumes compresses genes from micron-length chains into nanoscopic volumes. 1160 1161 1162 1163 1164 1165 1166 1167 1168 .CC-BY-NC 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 29, 2025. ; https://doi.org/10.1101/2025.05.29.656862doi: bioRxiv preprint 29 1169 Figure 3. Introns and genes are geometrically linked to exon length by physical principles. 1170 A) Histogram of protein coding genes in the human genome showing that the median fraction is 1171 ~9.8% exonic with a subset of genes that are almost completely exon. B) Plot of the ratio of exon 1172 (ideal zone)/intron (total volume) compared to the gene length of human genes. The constant, p, 1173 reflects the position along a domain volume. Here, n set to 1. As length increases, as expected, 1174 the volume ratio decreases as expected for a distribution of reaction volumes. C) Randomization 1175 of exon/intron segments results in the statistically grounded null hypothesis of no relationship 1176 between length and exon/intron ratios with a fraction approaching the median of 0.1. D) Intron 1177 length is a power-law of exon length for protein coding genes with values of p and γ as reported. 1178 E) Schematic representation of gene composition in relation to p and γ , indicating that high γ 1179 indicates more non-exonic volumetric elements are present within a segment. F) Relationship 1180 between γ and D depends on the proportion of the exons making up the ideal zone where β =1 1181 indicates the entire exon contents are confined to a hard surface. 1182 1183 1184 1185 1186 1187 1188 1189 1190 1191 1192 1193 .CC-BY-NC 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 29, 2025. ; https://doi.org/10.1101/2025.05.29.656862doi: bioRxiv preprint 30 1194 Figure 4. Exons are non-randomly coupled to adjacent volumetric DNA to generate power-1195 law segments independent of the final RNA product. A) Schematic representation of the 1196 organization of a gene on a chromosome segmented into two separate domains generated by a 1197 hinge element. The reaction volumes are produced by the volumetric DNA to guide the position of 1198 exons to ideal reaction zones. B) Analysis of the structure of human chromosomes in the positive 1199 strand orientation showing power-law assemblies of nontranscribed (volumetric) elements scaling 1200 as a power-law of exonic (ideal zone) elements. C) Randomly redistributing an exon with adjacent 1201 volumetric DNA conserves power-law distribution. In contrast, randomly distributing exons results 1202 in a linear distribution of segments. D-E) Comparison of organization generated by considering 1203 only ( D) protein coding genes compared to ( E) only non-protein coding genes in the positive 1204 strand orientation. In either case, chromosomes assemble into power-law units. 1205 1206 1207 1208 1209 1210 1211 .CC-BY-NC 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 29, 2025. ; https://doi.org/10.1101/2025.05.29.656862doi: bioRxiv preprint 31 1212 Figure 5. Exons, introns, and intergenic segments are non-randomly positioned into 1213 reaction volumes. A) Comparison of the DNA content observed in ChromSTEM packing 1214 domains compared to analytical predictions of the model when hinge length is less than 300bp 1215 demonstrating similar distributions in content. B) Analysis of distance between active Pol II (PS-5) 1216 to the nearest core element (H3K9me3) as a function of Pol II Ps-5 density. Shallower depths and 1217 smaller domains would result in a higher concentration of Pol II Ps-5 per segment and would be 1218 expected to have a shorter distance to a core as is experimentally observed. C) Analysis of the 1219 ChIP-Seq content of heterochromatin (H3K9me3) and promoter associated euchromatin 1220 (H3K4me3) within theorized volumes from the model in HCT-116. Consistent with underlying 1221 theory, chromatin segments are composed of both elements with positioning H3K4me3 coinciding 1222 with position of exon elements while heterochromatin is positioned to deeper layers. This 1223 partitioning suggests that heterochromatin is coupled with euchromatin in genetic segments. D) 1224 ChIP-Seq analysis of active isoforms Pol II Ps-5 in HCT-116 cells and Pol II Ps-2 in Hep2G cells 1225 within gene bodies (exons and introns) demonstrating preferential localization of polymerase onto 1226 exons per basepair length. If polymerases were uniformly throughout a gene body equivalent 1227 coverage per basepair would be observed. As a control for read-coverage bias, this was 1228 compared to H3K9me3 (K9) which preferentially localizes to introns. E) Nascent RNA-seq 1229 analysis in HCT-116 demonstrates nearly uniform synthesis of short introns and exonic 1230 sequences independent of gene lengths. In contrast, longer genes demonstrate decreasing 1231 synthesis of RNA in the direction of the reading frame consistent with Pol II having ideal reaction 1232 positions. F) Analysis of hinge segments in HCT-116 demonstrates an enrichment toward 1233 transcriptionally active features and enhancer positions compared to randomly generated 1234 segments (R). G) In contrast, a very small percentage (<1%) of hinge positions overlaps with 1235 constitutive heterochromatin. H) Experimentally observed RNA polymerase loops plotted as a 1236 function of their NE content (approximately the volume generated) compared to the observed 1237 ChIP-Seq content within each loop in HCT-116 cells. Within large polymerase loop domains, we 1238 observe an accumulation of heterochromatin. The total heterochromatin content increases a 1239 .CC-BY-NC 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 29, 2025. ; https://doi.org/10.1101/2025.05.29.656862doi: bioRxiv preprint 32 function of the size suggesting packing occurring as a function of the volume generated. Small 1240 loops are primarily composed of euchromatin, indicating their volume would appear to be 1241 relatively decompacted. Collectively, this suggests that transcriptional loops have a degree of 1242 packing that correlates to the packing behavior for domain volumes observed on ChromSTEM 1243 imaging. 1244 Figure 6. Power-law geometry produces a trade-off between modular durability and 1245 damage-risk. A&B) Analysis of geometric properties of genes that are primarily unique to the 1246 esophagus, cerebellum, and muscle tissue demonstrating power-law organization. C&D) Stem-1247 cell related transcription factors (Yamanaka factors, YF) generally organize as linear geometries. 1248 Similarly, HOX genes demonstrate two phenotypes: a cluster with linear organization (values of 1249 E/I >1) and a group that organizes into power-law distribution. Transcription factors such as 1250 RUNX2 that define tissue function are prim arily organized as power-law geometries. E) Analysis 1251 of the frequence of protein-coding genes containing a hinge position demonstrating that the 200 1252 most frequent Tier-1 oncogenes contain at least one hinge position. WG – whole genome 1253 compared Tier-1 oncogenes by their frequency: T50 – top 50 genes, T100 – top 100 genes, T200 1254 – top 200 genes. F) Analysis of Tier 1 oncogene frequencies demonstrates an acceleration then 1255 plateau in frequencies as the likelihood of containing a hinge increases. Genes that are less likely 1256 to contain a hinge element had a lower correlation with oncogenic mutation frequency. This effect 1257 appears to plateau at mutation frequencies occurring over 1% of the time. 1258 1259 1260 1261 1262 .CC-BY-NC 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 29, 2025. ; https://doi.org/10.1101/2025.05.29.656862doi: bioRxiv preprint 33 1263 Figure 7. Packing geometry parallels body plan complexity in metazoans. A&B) Analysis of 1264 chromosomal architecture (A) and genes (B) demonstrates that the S. cerevisiae genome is most 1265 likely organized as linear beads on a string assembly. As gene length increases, the non-exonic 1266 content increases linearly to create a chain. C&D) Analysis of chromosomal ( C) and genes ( D) 1267 demonstrating a transition toward power-law assemblies. In C. elegans, genes appear to be 1268 equally split between linear assemblies (E/I >1) and power-law assemblies. E-H) We observe a 1269 transformation with increasing body-plan complexity in both genes and chromosomes of D. rerio 1270 (E-F) and M. musculus (G-H) that their organizational structure resembles the structure of genes 1271 and chromosomes observed in humans. 1272 1273 .CC-BY-NC 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 29, 2025. ; https://doi.org/10.1101/2025.05.29.656862doi: bioRxiv preprint

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-pdf

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-26T02:00:01.498150+00:00
License: CC-BY-NC-4.0