Incomplete lineage sorting of segmental duplications defines the human chromosome 2 fusion site early during African great ape speciation

preprint OA: closed CC-BY-NC-ND-4.0
📄 Open PDF Full text JSON View at publisher
AI-generated deep summary by claude@2026-07, 2026-07-03 · read from full text

This study analyzed finished telomere-to-telomere genome sequences from great apes and macaques to reconstruct the ~109 kb human chromosome 2 fusion site using comparative genomics, single–base pair characterization, and epigenetic/structural analyses across lineages. The authors found that the fusion region is shaped by multiple pericentric inversions, segmental duplications (SDs), rapid turnover of subterminal repetitive DNA, and that three SDs originated >5 million years ago and show differential presence among African great apes due to incomplete lineage sorting and lineage-specific duplication. A human-specific fusion-associated degradation of divergent α-satellite arrays generated five structural haplotypes in humans, and CRISPR/Cas9-mediated depletion of the fusion site altered expression of 108 genes; a key caveat is that functional inference relies on cell-line perturbation rather than a direct in vivo developmental mechanism. This paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

ABSTRACT All great apes differ karyotypically from humans due to the fusion of chromosomes 2a and 2b, resulting in human chromosome 2. Yet, the structure, function, and evolutionary history of the genomic regions associated with this fusion remain poorly understood. Here, we analyze finished telomere-to-telomere chromosomes in great apes and macaques to show that the fusion was associated with multiple pericentric inversions, segmental duplications (SDs), and the rapid turnover of subterminal repetitive DNA. We characterized the fusion site at single-base-pair resolution and identified three distinct SDs that originated more than 5 million years ago. These three distinct SDs were differentially distributed among African great apes as a result of incomplete lineage sorting (ILS) and lineage-specific duplication. Most conspicuously, one of these SDs shares homology to a hypomethylated SD spacer sequence present in hundreds of copies in the subterminal heterochromatin of chimpanzees and bonobos. The fusion in human was accompanied by a systematic degradation of the three divergent α-satellite arrays representing the ancestral centromere creating five distinct structural haplotypes in humans. CRISPR/Cas9-mediated depletion of the fusion site in human cell lines significantly alters the expression of 108 genes, indicating a potential regulatory consequence to this human-specific karyotypic change.
Full text 80,771 characters · extracted from oa-pdf · 7 sections · click to expand

Abstract

37 All great apes differ karyotypically from humans due to the fusion of chromosomes 2a and 2b, 38 resulting in human chromosome 2. Yet, the structure, function, and evolutionary history of the 39 genomic regions associated with this fusion remain poorly understood. Here, we analyze finished 40 telomere-to-telomere chromosomes in great apes and macaques to show that the fusion was 41 associated with multiple pericentric inversions, segmental duplications (SDs), and the rapid 42 turnover of subterminal repetitive DNA. We characterized the fusion site at single-base-pair 43 resolution and identified three distinct SDs that originated more than 5 million years ago. These 44 three distinct SDs were differentially distributed among African great apes as a result of 45 incomplete lineage sorting (ILS) and lineage-specific duplication. Most conspicuously, one of 46 these SDs shares homology to a hypomethylated SD spacer sequence present in hundreds of 47 copies in the subterminal heterochromatin of chimpanzees and bonobos. The fusion in human 48 was accompanied by a systematic degradation of the three divergent α-satellite arrays 49 representing the ancestral centromere creating five distinct structural haplotypes in humans. 50 CRISPR/Cas9-mediated depletion of the fusion site in human cell lines significantly alters the 51 expression of 108 genes, indicating a potential regulatory consequence to this human-specific 52 karyotypic change. 53 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 4

Introduction

54 Karyotype evolution is a critical aspect of evolutionary biology because it has been associated 55 with speciation, adaptation, and disease1–7. In 1935, Wilson and Painter first observed that 56 chromosome fusion could lead to speciation in flies7. As the field advanced, numerous instances 57 of chromosome fusion have been documented in plant and animal speciation1–4,8–10. Several 58 molecular mechanisms have been proposed as key drivers of chromosome fusion in speciation 59 and oncogenesis, including telomere-telomere, telomere-centromere, tandem repeat, segmental 60 duplication (SD), and retrotransposon expansions and fusions11–15. 61 62 Recent advances in genome editing and synthetic biology have provided deeper insights into the 63 role of chromosome fusion in evolution16–18. Boeke and colleagues, for example, demonstrated 64 that chromosome fusion could induce reproductive isolation in yeast, highlighting the potential 65 of these genomic changes to influence speciation19,20. Additionally, studies on chromosome 66 engineering in mice reveal significant chromatin conformation alterations in the chromosome 67 fusion regions, further emphasizing the impact of karyotype evolution on genetic and phenotypic 68 diversity21,22. 69 70 Human chromosome 2 (chr2) was formed from the fusion of nonhuman primate (NHP) chr2a 71 and chr2b23–25, representing arguably the most significant karyotypic differences between 72 humans and NHPs. Ijdo and Wienberg were the first to identify the fusion site at human 73 chromosome 2q13-2q14 using cytogenetic banding approaches25. Subsequent studies uncovered 74 dispersed SDs associated with this fusion site26,27. These analyses relied on partial genome 75 assemblies generated using short-read technologies or Sanger sequencing of bacterial artificial 76 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 5 chromosome (BAC) clones, often resulting in incomplete genetic information at the fusion site26–77 29. Consequently, the structure, evolutionary history, and function of the fusion site have 78 remained only partially understood. 79 80 Available genomic assemblies and FISH experiments have shown an abundance of satellite 81 sequences in the subtelomeric repetitive regions of NHP chr2a and chr2b, whereas such 82 sequences are absent in humans26,30,31. During the fusion of human chr2, one of the centromeres 83 in the fused chromosome becomes inactive, and degraded with independent transposable element 84 (TE) retrotransposition events occurring at the degenerate site32. In addition, there has been 85 considerable debate regarding the timing of the chr2 fusion event. Based on SD and SVA (SINE-86 VNTR-Alu) divergence, the fusion event was estimated to have occurred early in human 87 evolution, 5-7 million years ago (mya) and 2.5-4.5 mya, respectively26,33. However, a recent 88 study, utilizing clustered substitution statistics, proposed a much more recent origin of 89 approximately 0.9 mya34. 90 91 The complex repetitive structure at the fusion site, subtelomeric repetitive regions, as well as the 92 inactive centromere have made it challenging to reconstruct the evolutionary history and 93 potential functional consequences of the fusion event. Here, we leverage the complete sequences 94 of great apes35,36 and a macaque37 to revisit the structure and evolutionary history of human chr2. 95 We aim to: (1) characterize the fusion site at a single-base-pair resolution, (2) study this in the 96 context of epigenetic and structural changes occurring at the subtelomeric repetitive regions and 97 degenerate centromere site, (3) use these data to create a model for human chr2 evolution, and 98 (4) examine the potential functional consequences by creating fusion-site depletion cell lines. 99 100 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 6

Results

101 Comparative sequence analysis of the human chromosome 2 fusion site. We performed a 102 comparative analysis of 13 finished nonhuman great ape chromosomes35 and the finished 103 macaque genome37 to human chr2 (Figure 1a, Methods). We identified numerous non-syntenic 104 segments (76-154 regions ≥10 kbp in length) between humans and NHPs, accounting for 9.86 105 Mbp (macaque) and up to 57.5 Mbp (gorilla) of unalignable chromosomal sequence per species. 106 Most of this sequence corresponded to various classes of repetitive DNA (Figure 1a, 107 Supplementary Table 1), including satellite repeats, tandem and interspersed SDs. In particular, 108 among the nonhuman African great apes, both chr2a and chr2b are acrocentric with the presence 109 of subterminal heterochromatic caps (9.6-19.6 Mbp) demarcating the ends of the chromosomes 110 and are composed of a 32 bp AT-rich satellite DNA (pCht) and SD spacer regions26,31,35,38. In 111 addition, there are multiple pericentric and paracentric evolutionary inversions distinguishing the 112 ape lineages from each other39, several of which share homology to the SD flanking regions of 113 the fusion site (Figure 1b and Figure 2a). One pericentric inversion occurred in the common 114 ancestor of African great apes after divergence from orangutans (chr2b) while another occurred 115 before the Pan–gorilla split (chr2a)39,40 (Figure 1a). Furthermore, we identified a truncated SD 116 containing the partial CBWD2 at the chr2a pericentric inversion breakpoints, and another 117 truncated SD with FOXD4L1 at the chr2b breakpoint (Supplementary Figure 1)28,29. 118 119 Evolutionary reconstruction of the ancestral fusion site. We further compared human chr2 120 with Pan chr2a and chr2b to precisely characterize the ~109 kbp fusion site in the human T2T-121 CHM13v2.0 genome assembly (2q14.1, chr2:113,940,058-114,049,496) (Figure 1b and 122 Supplementary Figure 2). Leveraging data from 47 human genomes from the Human Pangenome 123 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 7

Reference

Consortium (HPRC)41, we confirmed that there is a complete absence of large 124 structural variants within this fusion site in 94 sequence-resolved human haplotypes, suggesting 125 a fixed structural haplotype for the entire fusion site (Supplementary Figure 3). Neither 126 nucleotide diversity (π) nor Tajima’s D values show significant reductions at the fusion site in 127 African or non-African populations, relative to the chr2 average (Supplementary Figure 4) 128 consistent with the locus evolving neutrally. 129 130 Within the larger ~455 kbp duplication block (seven independent SD blocks) demarcating the 131 fusion site region, we narrowed the fusion site to a ~109 kbp segment consisting of three distinct, 132 high-identity SDs (≥98% identity and ≥20 kbp in length) (Figure 1b, Figure 2a, Supplementary 133 Figures 5-12, and Supplementary Tables 2-3). In humans, these SDs share homology with 134 sequences corresponding to chr9 (SD_fusion_A, 50 kbp), chr12/20 (SD_fusion_B, 36 kbp), and 135 chr22 (SD_fusion_C, 22 kbp) (Figure 2a, Supplementary Table 2). Among African apes, each 136 SD is highly variable in copy number showing evidence of shared ancestral locations as well as 137 lineage-specific duplications (Supplementary Table 2). Using macaque and orangutan as 138 outgroups, we reconstructed the evolutionary history of each SD estimating the divergence 139 timepoint from each SD to its nearest nonhuman great ape neighbor (Figure 2b-g, Methods). 140 141 SD_fusion_A (50 kbp) is present as a single copy in macaques and orangutans and corresponds 142 to a partial duplication of PGM529 (14 exons of an ancestral gene), a gene involved in 143 carbohydrate metabolism, originating from African great ape ancestral chr9 (phylogenetic group 144 IX) (Figure 2b-c). This SD began to duplicate in an interspersed configuration in the common 145 ancestral lineage of African great apes (~7.4 mya) with the locus expanding to highest copy 146 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 8 number in chimpanzees and bonobos where it defines the SD spacer region demarcating large 147 blocks of pCht on chr2a and chr2b and other chromosomes (chr3, chr5-chr13, chr15-chr22, and 148 chrX in chimpanzee; each chromosome in the bonobo genome) (Figure 2b-c and Supplementary 149 Figure 6). Of note, the SD_fusion_A (chr2) and other human copies share a monophyletic origin 150 (~3.3 mya) most closely related to gorilla copies ~6 mya (95% CI: 4.8-7.4 mya) instead of 151 chimpanzee or bonobo—a pattern consistent with incomplete lineage sorting (ILS). 152 153 Similarly, SD_fusion_B, is a 36 kbp segment corresponding to the partially truncated WASH242 154 (11 exons of an ancestral gene that potentially regulates actin cytoskeleton dynamics), FAM138B 155 (3 exons of an ancestral gene of unknown function), and DDX11 (3 exons of an ancestral gene 156 implicated in DNA metabolism). The segment exists as a single copy in both macaques and 157 orangutans and originated from a locus mapping to ancestral African great ape chr12 158 (phylogenetic group XII) (Figure 2d-e and Supplementary Figure 7). All human copies mapping 159 to human chr2, chr20, chr12, and chr9 show a monophyletic origin (~4.4 mya, 95% CI: 3.1-5.8 160 mya) suggesting human-specific duplications or interlocus gene conversion events43. With 161 respect to nonhuman African great apes, however, this clade diverges from a distinct 162 monophyletic clade that includes chimpanzee, gorilla, and bonobo (~7.4 mya, 95% CI: 5.4-9.3 163 mya) (Figure 2d). This topology is once again consistent with ILS. 164 165 In contrast to SD_fusion_A and SD_fusion_B, SD_fusion_C (the most distal, 22 kbp) shows 166 evidence of independent duplication in all primates, including macaques and orangutans, with 167 the exception of chimpanzees where it exists as a single copy on chimpanzee chr6 168 (Supplementary Figure 8). All human SDs show a monophyletic origin (~3.5 mya, 95% CI: 2.7-169 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 9 4.3 mya) but show a deep coalescence with another genomic segment present only on bonobo 170 chr22 (phylogenetic group XXII), dating back to ~6.9 mya (95% CI: 5.3-8.6 mya) (Figure 2f). 171 Notably, SD_fusion_C and its flanking region correspond to a larger SD that aligns exclusively 172 between the human chr2 SD_fusion_C site and bonobo chr22, but not with any other NHPs 173 (Supplementary Figure 13). Thus, SD_fusion_C and its flanking region likely represent an 174 ancestral genomic structure in the latest common ancestor (LCA) of African great apes, which 175 subsequently sorted in the human and bonobo lineages. 176 177 We extended the ILS analysis to the 5 Mbp mapping proximally (chr2a) and distally (chr2b) to 178 the fusion site in humans (Methods) using a 500 bp windowed approach44 (Supplementary Table 179 4). We observe a sharp transition in the proportion of ILS windows at the site of the chr2 fusion. 180 Specifically, we find the proportion of ILS rises to 45.6% (a 1.39-fold excess compared to the 181 chr2 average) as we approach the fusion site distally (chr2b side). In contrast, the proximal 182 portion appears depleted for ILS segments (Figure 2h-i, Supplementary Figure 14 and 183 Supplementary Table 5). There is, thus, a polarized pattern of ILS with maxima and minima 184 occurring on either side of the fusion site. 185 186 The refined analysis of telomeric sequences at the fusion site. Previous investigations 187 identified inverted telomeric sequences (TTAGGG/CCCTAA) at the fusion site, suggesting a 188 telomere-to-telomere (T2T) fusion25. We confirmed the presence of interstitial telomeric 189 sequences (CCCTAA, chr2: 114,027,659-114,028,207) at SD_fusion_C and its corresponding 190 SD on human chr22 (TTAGGG, chr22:51,254,086-51,323,279) and similar sequences in bonobo 191 chr22 (TGAGGG, chr22: 60,722,979-60,724,048). The sequence (GGGTTA) was identified at 192 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 10 SD_fusion_B (chr2: 114,027,333-114,027,658) and its corresponding SD on human chr12 193 (TAACCC, chr12:2,843-3,030) and bonobo chr20 (TAACCC, chr20:1,139,021-1,139,266), but 194 not at its corresponding SD on human chr20 (Supplementary Figure 15). These observations 195 argue that these telomeric sequences were present on SD_fusion_B and SD_fusion_C prior to the 196 fusion event. In addition, two telomeric-associated repeats (TAR1) flank the telomeric sequences 197 of SD_fusion_B and SD_fusion_C27. Given the genomic structure of TAR1 and telomeric 198 sequences on these SDs (Supplementary Figures 15-16), we propose that, in the ancestral 199 configuration, each SD contained a TAR1 element and telomeric sequences, arranged in an 200 inverted orientation on ancestral chr2a and chr2b. The end-to-end fusion of these two distinct 201 SDs may have facilitated the fusion of ancestral chr2a and chr2b, leading to the vestigial 202 presence of two TAR1 elements and the (CTAACC) and (GGGTTA) repeat motifs at the human 203 fusion site. 204 205 African great ape subterminal satellite repeat expansion and human chromosome 2 fusion. 206 Reconstructing the evolutionary history of the fusion event has been challenging due to extensive 207 SD and lineage-specific turnover of the subterminal heterochromatic caps in Pan and gorilla35 208 (Supplementary Table 6). With the exception of the short arm of gorilla chr2a, the corresponding 209 regions in both Pan and gorilla are composed of nearly continuous megabase-pair tracts of 210 satellite DNA interspersed with SD spacers35. The satellite sequence is made of a tandem 32 bp 211 repeat motif (pCht)31 punctuated on average every 287 kbp in Pan and every 389 kbp in gorilla 212 by a SD spacer (Figures 3a, Supplementary Figures 17-19). The interrupting SD spacers are 213 variable with a modal length of 32 kbp in the Pan lineage and 33.7 kbp in gorilla (Supplementary 214 Figure 17) and correspond to hypomethylated pockets flanked by the hypermethylated satellite 215 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 11 DNA35 (Figure 3a). Although both the gorilla and chimpanzee SD spacers differ in sequence 216 composition and the subterminal heterochromatic cap is largely thought to have evolved 217 independently in both lineages26,35, the net effect is that the subterminal portions of both 218 chimpanzee chr2a and chr2b share 23.6 Mbp of high-identity sequence homology involving both 219 the SD spacers and pCht satellite DNA. In the case of the gorilla, the homology is restricted to 220 the pCht satellite DNA for both arms of gorilla chr2b and the q-arm of gorilla chr2a. The 221 organization of the p-arm of gorilla chr2a differs considerably and is much more similar to the 222 organization found in orangutan, which is enriched in HSatIII-like repeat sequences (Figure 3b). 223 It is classified as an acrocentric short arm lacking a nuclear organizing region35. 224 225 Importantly, the SD spacer that expanded in the subterminal heterochromatic caps in chimpanzee 226 and bonobo shares 98.5% identity with the SD_fusion_A segment mapping at the fusion site on 227 human chr2 (Figure 3c). All 429~452 SD spacers within the Pan pCht regions, including those at 228 the ends of these chromosomal regions, are monophyletic in origin estimated to have expanded 229 approximately ~5.5 mya (95% CI: 4.2-6.8 mya). This subterminal SD spacer (32 kbp) in Pan is 230 18 kbp smaller than SD_fusion_A (50 kbp) where the shared LCA predates the human–Pan–231 gorilla divergence (Figure 2b and 2c). Comparing the phylogenies of the sequence unique to 232 SD_fusion_A and sequence shared with the subterminal heterochromatic cap SD spacers shows 233 nearly coincident ILS topology (generalized Robinson–Foulds distance=0.28, p=1.6×10-4). This 234 indicates that the SD spacers within Pan heterochromatic caps are derived from the duplicated 235 sequence that gave rise to ancestral SD_fusion_A. Thus, the hyperexpansion of subterminal 236 satellite DNA associated with subterminal heterochromatic caps in Pan and the chr2 fusion are 237 linked genetically with two different evolutionary trajectories and karyotypic consequences in 238 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 12 human and Pan (Figure 2c-d, Supplementary Figure 20 and Supplementary Table 7). 239 240 Centromere retention and degeneration in human chromosome 2. An important consequence 241 of the human chr2 fusion was that the ancestral chr2b centromere in NHPs became inactive in 242 our human lineage (Figure 4a). Chr2a differs from chr2b by the presence of large tracts of HSatII 243 arrays in humans and HSatIII arrays in Pan (Supplementary Figures 21-22). Overall, the active 244 human centromere in humans is more similar to that of gorilla with respect to suprachromosomal 245 family (SF) organization. Both humans and gorillas possess SF2, whereas chimpanzees and 246 bonobos have SF3. 247 248 We compared the active chromosome centromere of chimpanzee chr2b with the structure of the 249 vestigial centromere in humans. We identified three distinct α-satellite arrays (chr2:132,644,386–250 132,685,996) in humans45 with homology to chimpanzee that had been interrupted by various 251 simple repeats (CCTCTC) and retrotransposon elements (SVA and L1PA3) in the human lineage 252 (Figure 4b). All three satellite arrays were derived from divergent monomeric α-satellite regions 253 ancestral to human and chimpanzee, rather than from higher-order repeats (HORs). Our analysis 254 of 94 human genome assemblies39 identifies five distinct structural haplotypes due primarily to 255 length variation of the first two α-satellite arrays (Figure 4c-d and Supplementary Table 8). A 256 sequence comparison using unique k-mers from this region shows that ~97.5% (8,246/8,460) and 257 ~96.7% (8,181/8,460) are also identified in Neanderthal and Denisovan genomes, respectively 258 (Methods), confirming that the centromeric degeneration occurred long before the divergence of 259 modern and archaic humans32. 260 261 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 13 We also compared methylation patterns of the NHP chr2a and chr2b centromeres (Figure 4e). 262 The human α-satellite arrays mapping to the degenerate site show significantly lower 263 methylation levels, when compared to typical NHP HORs (excluding centromere dip regions 264 [CDRs], p=0.02), with the exception of bonobo chr2b and gorilla chr2b. HORs in bonobo chr2b 265 and gorilla chr2b are significantly hypomethylated when compared to other NHP HORs 266 (p=1.92×10-6), possibly due to the smaller size of these HORs (lengths: 71 kbp for bonobo chr2b, 267 and 106 kbp for gorilla chr2b). 268 269 Functional assessment of the fusion site by depletion. Four putative noncoding 270 genes/pseudogenes (PGM5P4, FAM138B, WASH2P, and DDX11L2) have been annotated at the 271 site of the human chr2 fusion. According to GTEx short-read RNA sequencing (RNA-seq), these 272 four genes/pseudogenes are expressed in testis, esophagus, fallopian tube, and cerebellum 273 tissues46 (Supplementary Figure 23). Further, both PGM5P4 and WASH2P are supported by 274 long-read Iso-Seq transcript data from CHM13hTERT47 and kidney tissue from ENCODE48. In 275 addition, methylation analysis demarcates a prominent CpG island showing the 276 promoters/enhancers of PGM5P4, as identified using ONT reads from the T2T-CHM13 cell 277 line47 (Supplementary Figure 24). 278 279 To explore the potential function of the fusion site, we used CRISPR/Cas9 to delete this region 280 in a kidney-derived epithelial-like cell line (HEK293T) (Figure 5a). Two independent pairs of 281 single guide RNA (sgRNA; L-sg1 and R-sg; L-sg2 and R-sg) were designed for the depletion 282 (Figure 5b), and the heterozygous depletion of the fusion site in both sgRNA groups was 283 confirmed by PCR and Sanger sequencing (Figure 5c and Supplementary Table 9). We then 284 conducted RNA-seq, generating 53.28 Gbp, 67.42 Gbp, and 122.41 Gbp of data for the L-sg1 285 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 14 depletion, L-sg2 depletion, and wild-type (WT) cell lines (three biological repeats for each 286 depletion cell line and six biological repeats for WT cell lines) (Supplementary Table 10). 287 288 Differential gene expression analysis identified 37 upregulated and 71 downregulated genes in 289 both depletion cell lines (Supplementary Figure 25 and Supplementary Table 11, Methods). 290 These genes were significantly enriched in pathways related to the regulation of transcription by 291 RNA polymerase II (p=2.2×10-3). Among the top ten differentially expressed genes (DEGs) in 292 both depletion cell lines, three were upregulated and seven were downregulated (Figure 5d-e). 293 We further validated five of ten gene expression changes using RT-PCR in the depletion cell 294 lines (Figure 5g and Supplementary Table 12). 295 296

Discussion

297 Complete T2T sequence for chromosome 2, 2a, and 2b among the great apes35–37 allowed us to 298 systematically examine the complex genomic architecture and further refine the evolutionary 299 history of the human-specific chr2 fusion event. There are three important conclusions. First, the 300 fusion event was intimately associated with SDs that have been restructuring genomes and 301 chromosomes throughout great ape evolution26,28. The 109 kbp fusion site in humans consists of 302 three independent SDs, each with distinct trajectories in different ape lineages, and these are 303 further embedded in a larger duplication block of ~455 kbp. All SDs are highly variable in copy 304 number among the great apes, and many have been reused as breakpoint sequences during great 305 ape evolution. Although cytogenetic studies previously reported two pericentric inversions on 306 ancestral chr2a and chr2b26, we observe that the flanking region at the human fusion site (chr2a, 307 proximal side) shares 98.4% identity with an SD flanking the pericentric inversion that 308 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 15 distinguishes Sumatran orangutan chr2b from gorilla chr2b. Moreover, SD_fusion_A shares 309 97.8% identity with the SD spacer sequence of the Pan subterminal heterochromatic caps where 310 it expanded to hundreds of copies in the chimpanzee and bonobo lineages but not in gorilla. 311 312 Second, we show that the fusion site has been strongly subjected to ILS. All three SDs, for 313 example, show phylogenetic signatures consistent with ILS with a maximum occurring at the 314 fusion site suggesting that the fusion event occurred >5 mya. Thus, the fusion event is not a 315 recent evolutionary event but potentially occurred during African great ape speciation when the 316 effective population size was predicted to be much larger than contemporary ape populations35,44. 317 Given fossil evidence that Australopithecus existed around 2-4 mya, Paranthropus around 1-3 318 mya, and the earliest Homo fossils around 2-3 mya49–51, we speculate that this fusion event did 319 arise in the genus Homo but rather occurred in ancestral great ape populations that would give 320 rise to humans potentially by creating a stasipatric speciation barrier52. 321 322 Third, the fusion site was associated with extensive subtelomeric satellite turnover and 323 epigenetic differences among great apes. The two satellite motifs present in the subtelomeric 324 regions of gorilla chr2a and chr2b are distinct from each other: one resembles the subtelomeric 325 repetitive regions of orangutans, while the other resembles the genomic architecture of Pan 326 species. SD_fusion_A, in part, defines the chr2 fusion site but also associates with large tracts of 327 hypermethylated pCht satellite repeats that define subterminal heterochromatin in chimpanzees 328 and bonobos35 (Supplementary Figure 26). The juxtaposition of novel SDs at the fusion site and 329 the shift from heterochromatic or acrocentric DNA at the termini of chr2a and chr2b led to 330 methylation differences among the great apes at this locus. Indeed, our depletion experiments in 331 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 16 humans show consistent gene expression changes suggesting that the euchromatization of the 332 fusion site may have had more global genome-wide regulatory effects. 333 334 We propose two evolutionary scenarios for the formation of human chr2. In the first, the large 335 effective population size of the human–Pan–gorilla ancestral population facilitated the 336 coexistence of several NHP chr2a and chr2b subtelomeric structural haplotypes 5-7 mya. In one 337 of these structural configurations, SD_fusion_A and SD_fusion_B became juxtaposed and 338 duplicated to the subtelomeric region of human–Pan–gorilla ancestral chr2a prior to the fusion 339 event (Figure 6a). Similarly, SD_fusion_C was duplicated to the subtelomeric region of ancestral 340 chr2b (where the full-length structure is still retained subtelomerically in bonobo); yet, it is no 341 longer located on other nonhuman great ape chr2b due to the rapid exchange of subtelomeric 342 SDs and evolutionary turnover of satellite DNA26,53. Subsequently, the telomeric and TAR1 343 sequences in SD_fusion_B and SD_fusion_C mediated the fusion (Figure 6a). In the Pan and 344 gorilla lineages, chr2a and chr2b (only in chimpanzee and bonobo) experienced a different 345 evolutionary trajectory associated with the formation of subterminal heterochromatic caps. 346 Previous studies26,35 have shown that the heterochromatic caps likely evolved independently in 347 both gorilla and chimpanzee, albeit convergently with a similar architectures where hundreds of 348 kilobase pairs of satellite pCht DNA are punctuated by a ~30 kbp SD spacer region that defines a 349 pocket of hypomethylation. The SD spacers in gorilla and chimpanzee subterminal 350 heterochromatin caps are distinct but in both chimpanzee and bonobo the SD spacers represent a 351 derivative of the SD_fusion_A confirming that this sequence was present subtelomerically for 352 chr2a in the human–Pan ancestor. SD_fusion_A subsequently expanded to hundreds of copies in 353 chimpanzee and bonobo in conjunction with the subterminal heterochromatic satellite. Yet, in the 354 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 17 human lineage, SD_fusion_A remains a low-copy SD with no trace of pCht satellite sequence in 355 the human genome35. 356 357 Alternatively, the SDs evolved similarly, but in the common ancestor of human and Pan, there 358 was an incipient association with rudimentary pCht satellite sequences with the SD_fusion_A 359 sequence that mediated exchange between ancestral chr2a and chr2b via nonallelic homologous 360 recombination (NAHR) or ectopic exchange. Subsequently, during NAHR between full-length 361 SD_fusion_A elements, the pCht regions were lost in the human lineage and two additional SDs 362 were inserted at the fusion breakpoint region (Figure 6b). In the Pan lineage, a portion of the 363 SD_fusion_A became hyperexpanded defining the subterminal SD spacer region of all 364 heterochromatic caps in bonobo and chimpanzee. In support of this model, it is known that in 365 both chimpanzee and gorilla the subterminal satellite DNA forms unique post-bouquet structures 366 in germ cells and are hotspots of ectopic exchange between nonhomologous chromosomes54. If 367 such exchanges occurred at the edge of the incipient subterminal heterochromatic caps before 368 there were many copies of the subterminal satellite DNA, it would help explain the absence of 369 pCht sequence in the human genome—i.e., the fusion of chr2a and chr2b helped eliminate the 370 potential for the formation of subterminal heterochromatic caps. 371 372 Irrespective of the model, there is evidence that SDs in the LCA and ILS potentially facilitated 373 by a large effective population size played a key role35,44. These aspects helped maintain and 374 diversify karyotypic structures of the ancestral chr2a and chr2b for possibly millions of years. 375 The divergent fates of the subtelomeric repetitive regions of these ancestral chromosomes, e.g., 376 pCht expansion and SD spacer insertions in the Pan lineage, or the SD insertions in the human 377 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 18 fusion site, moreover, may have driven the speciation between humans and NHPs through ILS 378 (Supplementary Figure 27). In this light, it is interesting that we document an asymmetric ILS 379 pattern, with an increase in ILS segments on the distal side (chr2b) and a decrease on the 380 proximal side (chr2a) of the fusion site. This suggests diverse evolutionary trajectories for the 381 ancestral subtelomeric regions of NHP chr2a and chr2b highlighting the potential important role 382 of both SDs and ILS in understanding chromosomal evolution, speciation, and sequence 383 turnover. 384 385

Methods

386 Data resources and comparative analysis. The great ape and macaque T2T genome assemblies 387 are publicly available (Data Availability). Syntenic relationships among primate 388 chr2/chr2a/chr2b were assessed using minimap255 (v2.24) with the following parameters: ‘-c -x 389 asm20 --secondary=no --eqx -Y -K 8G -s 1000’ and visualized using 390 SafFire (https://mrvollger.github.io/SafFire). We identified centromeres, transposons, and 391 subtelomeric satellites based on RepeatMasker (v4.1.4) annotation 392 (https://www.repeatmasker.org/). 393 394 Fusion site characterization and population analysis. To precisely define the fusion site, we 395 aligned NHP chr2a and chr2b to human chr2 using minimap2 (v2.24) with the following 396 parameters ‘-x asm20 -r500,20000 -s 2000 -p 0.01 -N 1000 --cs’. The alignments were visualized 397 using the minimiro (https://github.com/mrvollger/minimiro). We used VCFtools56 (v0.1.16) to 398 calculate nucleotide diversity (pi) and Tajima’s D for population genetic analyses with the 399 following parameters ‘--window-pi 20000 --window-pi-step 10000’ and ‘--TajimaD 20000’. 400 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 19 401 SD annotations and phylogenetic analysis. SD annotations of the fusion site in the T2T-402 CHM13v2.0 were based on the UCSC SEDEF-SD track (https://genome.ucsc.edu/). In this 403 study, we focused on the SDs with an identity of ≥98% and length ≥20 kbp. These SD sequences 404 were aligned to NHP T2T genomes to identify homologous regions using minimap2 (v2.24) with 405 the parameters: ‘-cx asm20 -p 0.5 --eqx’. Orthologous segments were identified based on 406 flanking region synteny. We aligned all primate homologous SDs with MAFFT57 (v7.515) and 407 used trimAl58 (v1.4) to remove the noise sequences with the parameter: ‘-automated1’. The 408 phylogenetic trees were constructed with IQ-TREE59 (v2.1.4), and then we used BEAST60 409 (v2.6.6) with the HKY model incorporating gamma site, calibrated yule, and relaxed log-normal 410 clock models to infer the split time. Node ages were estimated using log-normal priors. We 411 conducted three independent runs for each tree and the results were consistent. Effective sample 412 sizes exceeded 200 for all parameters in all runs. 413 414 For ILS analysis, we first truncated the fusion site flanking regions into 500 bp windows. Then, 415 we utilized transanno (v0.4.5) (https://github.com/informationsea/transanno) to align these 500 416 bp segments to NHP genomes and generated multiple alignments across primate species. We 417 then utilized IQ-TREE59 (v1.6.12) with HKY model and ete3 python package to analyze 418 phylogeny trees, as described previously44. 419 420 Subtelomeric repetitive region characterization, methylation analysis, and TE annotation. 421 The pairwise identity heatmaps of each subtelomeric region were generated by ModDotPlot61 422 (https://github.com/marbl/ModDotPlot). We used SEDEF62 (v1.1r35) to annotate SDs. To 423 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 20 analyze 5mC methylation levels in pCht regions and SD spacers, we utilized BEDTools63 424 (v2.30.0) to calculate mean methylation levels within 1 kbp windows, using a 500 bp step size. 425 Synteny between subtelomeric repetitive regions was defined using minimap2 (v2.24) with the 426 following parameters ‘-cx asm20 --secondary=no -A1 -B2 -O2,12 -s 1000 -Y -K 8G --eqx’. The 427 syntenic relationship was visualized using R package circlize64 (v0.4.16). R package dendextend 428 (v1.18.1)65 and TreeDist (v2.9.1)66 were utilized to plot the co-phylogeny and to calculate the 429 generalized Robinson–Foulds distance. 430 431 Centromere analysis. To analyze the structure of each chromosome, we first ran RepeatMasker 432 (v4.1.4) on all chromosomes and identified the α-satellite-enriched regions as centromeres67. 433 Then, we utilized HumAS-HMMER (https://github.com/fedorrik/HumAS-434 HMMER_for_AnVIL) to classify the suprachromosomal families (SFs) of α-satellites in human, 435 Pan, gorilla, and orangutan. For macaque, the OWM-SF annotation tool was used to annotate the 436 SFs37. The HOR arrays were characterized using the tool StV (https://github.com/fedorrik/stv). 437 However, for orangutan chr2a and chr2b, the tool encountered difficulties in correctly identifying 438 the HORs, likely due to the complex structure of these acrocentric chromosomes. Thus, we 439 selected the largest continuous regions containing the same SFs as their HORs. We then 440 estimated the frequency of 5mC and CpG methylation within the HOR regions. BEDTools 441 (v2.30.0) was used to count the frequency of methylation within 5 kbp windows and we 442 determined the regions with the minimum frequency and below the lower quartile among the 443 whole HOR as CDRs. Additionally, we used the previous published tool68 444 (https://github.com/altemose/chm13_hsat) to identity the HSatII/HSatIII arrays in human and 445 Pan peri/centromeric regions (α-satellite-enriched regions and 5 Mbp on the p-arm and q-arm). 446 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 21 447 We further characterized the human centromere degenerate site through synteny comparisons 448 between Pan and human, incorporating RepeatMasker annotations of ‘ALR_Alpha.’ Breakpoints 449 between chimpanzees and humans were identified through alignment and analysis of specific SF 450 organizations. ModDotPlot (https://github.com/marbl/ModDotPlot) was used to generate 451 heatmaps for chimpanzee chr2b with a window size of 5,000 bp and for the human centromere 452 degenerate site with a window size of 200 bp. 453 454 To assess the diversity of centromere degeneration across human populations, we first confirmed 455 the presence of this site in all human genomes using minimap2 (v2.24) with the parameters ‘-cx 456 asm20 --secondary=no -s 2500’. We then extracted the targeted regions from each assembly 457 from HPRC data and ran RepeatMasker (v4.1.4) on these regions. We defined three satellite 458 arrays—cenD_1, cenD_2, and cenD_3—with subtypes determined by length variations. 459 460 To confirm the fusion occurred in archaic humans, we randomly selected two individuals from 461 the five structural haplotypes (total 10 individuals) and five NHP genomes to run mrsFAST69 462 (v3.4.2) for identifying singly unique nucleotide k-mers (SUNKs) in modern humans. 463 Subsequently, we checked the counts of the SUNKs in archaic human genomes. 464 465 To compare the methylation frequencies among human chr2, α-satellite arrays at the degenerate 466 centromeric site, and NHP centromeres, we chunk the HORs (excluding CDRs; for macaques, 467 we used the whole α-satellite region, excluding CDRs) into 17.1 kbp windows by BEDTools 468 (v2.30.0) with parameters: ‘-w 171000 -s 8550’ and calculate the frequencies within windows. 469 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 22 Visualizations were generated using ggplot2. 470 471 Fusion site KO experiments and RNA-seq analysis. The design of optimal sgRNA pairs to 472 target sites was performed using the online CRISPR design tool (http://crispor.gi.ucsc.edu/). The 473 complementary oligonucleotide pairs of L-sgRNA #1/#2 and R-sgRNA were annealed at 95℃ 474 for 5 min, with ramp-down to 25℃ to generate the double-stranded DNA (dsDNA) fragment, 475 before ligation into BbsI-digested PX459 (plasmid #48139, Addgene). 476 477 HEK-293T cell lines were maintained in complete culture medium: DMEM (Dulbecco’s 478 modified Eagle’s medium) supplemented with 10% FBS (fetal bovine serum), 1% Penicillin-479 Streptomycin Solution (100×stock, Gibco). Transfection of CRISPR plasmid DNA into HEK-480 293T cells was performed using LipofectamineTM 3000 Reagent (Invitrogen). Briefly, 0.5–481 5×105/cm2 HEK293T cells in 6 well plates were transfected with 2.5 μg of CRISPR plasmid 482 DNA. Then, 24 hours after transfection, selection was conducted by adding puromycin (1 μg/ml) 483 to the media for the next 48 hr. Surviving cells were cultured for 4-7 days without selection 484 before harvesting. 485 486 Genomic DNA (gDNA) was extracted from transfected HEK-293T from each subline (1×105) 487 with the TIANamp genomic DNA Kit (DP304-02; Tiangen) according to the manufacturer’s 488 instructions. Genome deletions caused by the sgRNA pairs were detected by PCR-amplification 489 of gDNA using a primer pair flanking the deletion. The primers used for PCR screening are 490 listed in Supplementary Table 8. Genomic PCR products were detected by agarose gel 491 electrophoresis and Sanger sequencing (Shanghai, TsingKe). 492 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 23 493 Total RNA of the fusion site depletion and wild-type cell lines were reverse-transcribed using 494 HiScript II Q RT SuperMix for qPCR kit (Vazyme), respectively. Gene expression was 495 normalized to GADPH, and fold changes were calculated as described in the Figure 5. Error bars 496 in the RNA analysis represent the standard deviation of the average fold changes based on at 497 least two cell lines and/or two experimental replicates as indicated in the Figure 5. The primer 498 sequences are listed in Supplementary Table 12. 499 500 We utilized fastp70 (v0.23.2) to perform adaptor trimming of raw sequencing reads. The trimmed 501 reads were aligned to the reference genome (T2T-CHM13v2.0) using hisat271 (v2.2.1). 502 Differential expression analysis was conducted using the R package DESeq272 and edgeR73 and 503 we identified genes as significantly differentially expressed with following criteria: false 504 discovery rate (FDR) 2. The genes considered significantly 505 differentially expressed met two criteria: they were identified by both tools and demonstrated 506 significance in two independent sgRNA design experiments, forming the final set. Visualizations 507 were generated using ggplot2. We performed Gene Ontology (GO) enrichment analysis using 508 DAVID GO74 (https://david.ncifcrf.gov/). 509 510 Acknowledgments 511 We thank Tonia Brown for editing this manuscript. We thank the HPRC and Primate T2T 512 Consortium for providing the long-read human and great ape genome assemblies. The 513 computations in this study were run on the Siyuan-1 supported by the Center for High 514 Performance Computing at Shanghai Jiao Tong University. E.E.E. is an investigator of the 515 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 24 Howard Hughes Medical Institute. 516 517 This article is subject to HHMI’s Open Access to Publications policy. HHMI lab heads have 518 previously granted a nonexclusive CC BY 4.0 license to the public and a sublicensable license to 519 HHMI in their research articles. Pursuant to those licenses, the author-accepted manuscript of 520 this article can be made freely available under a CC BY 4.0 license immediately upon 521 publication. 522 523 Author contributions 524 Y.M., E.E.E., and Q.S. conceived the project; X.J., Z.Y., X.Y., K.M., S.Z., J.C., J.H., L.F., J.Z., 525 and Y.M. contributed to the syntenic comparison, fusion site characterization, ILS, and 526 subtelomeric repetitive region analyses; L.Z., X.J., Y.L., Y.N., X.B., and Q.S. contributed to the 527 fusion site depletion analysis; D.Y. and E.E.E. generated the genome assemblies of great apes; 528 X.J., G.Z., and Y.M. analyzed the centromeres. Y.M. and E.E.E. wrote the draft manuscript with 529 contributions from other authors. All authors read and approved the manuscript. 530 531 Conflict of Interest 532 E.E.E. is a scientific advisory board (SAB) member of Variant Bio, Inc. The other authors 533 declare no competing interests. 534 535 Funding 536 This work was supported, in part, by grants from the National Natural Science Foundation of 537 China (32370658), Natural Science Foundation of Chongqing, China (CSTB2024NSCQ-538 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 25 JQX0004), and Shanghai Jiao Tong University 2030 Initiative (WH510363001-7) to Y.M. This 539 work was supported, in part, by grants from the National Key Research and Development 540 Program of China (2022YFF0710901), the National Natural Science Foundation of China Grant 541 (82021001), Biological Resources Program of Chinese Academy of Sciences (KFJ-BRP-005), 542 National Science and Technology Innovation 2030 Major Program (2021ZD0200900) to Q.S. 543 This work is partially sponsored by Shanghai Rising-Star Program (24YF2721800) to K.M. This 544 work was supported, in part, by US National Institutes of Health (NIH) grant HG002385 to 545 E.E.E. 546 547 Data Availability 548 The T2T primate genomes used in this study are available from GenBank via accessions: 549 GCA_009914755.4, GCA_028858775.2, GCA_028885625.2, GCA_028885655.2, 550 GCA_029281585.2, GCA_029289425.2 and GCA_037993035.1. The T2T primate genome 551 assemblies are also available on GitHub (https://github.com/marbl/Primates and 552 https://github.com/zhang-shilong/T2T-MFA8). The Neanderthal and Denisovan genomes used 553 are available from https://www.eva.mpg.de/genetics/genome-projects. The Iso-Seq data are 554 available on ENCODE (tissues: ENCFF492BYP, ENCFF306ZPP, ENCFF318SKH, CHM13-555 T2T2 hTERT Iso-Seq: SRR12519035, SRR12519036). 556 557

References

558 1. Ferguson-Smith, M. A. & Trifonov, V. Mammalian karyotype evolution. Nat Rev Genet 8, 559 950–962 (2007). 560 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 26 2. Bush, G. L., Case, S. M., Wilson, A. C. & Patton, J. L. Rapid speciation and chromosomal 561 evolution in mammals. Proc. Natl. Acad. Sci. U.S.A. 74, 3942–3946 (1977). 562 3. Leibowitz, M. L., Zhang, C.-Z. & Pellman, D. Chromothripsis: A New Mechanism for Rapid 563 Karyotype Evolution. Annu. Rev. Genet. 49, 183–211 (2015). 564 4. Rieseberg, L. H. Chromosomal rearrangements and speciation. Trends in Ecology & Evolution 565 16, 351–358 (2001). 566 5. Watkins, T. B. K. et al. Pervasive chromosomal instability and karyotype order in tumour 567 evolution. Nature 587, 126–132 (2020). 568 6. Damas, J. et al. Evolution of the ancestral mammalian karyotype and syntenic regions. Proc. 569 Natl. Acad. Sci. U.S.A. 119, e2209139119 (2022). 570 7. Painter, T. S. & Stone, W. CHROMOSOME FUSION AND SPECIATION IN 571 DROSOPHILAE. Genetics 20, 327–341 (1935). 572 8. Schubert, I. Chromosome evolution. Current Opinion in Plant Biology 10, 109–115 (2007). 573 9. Kirkpatrick, M. & Barton, N. Chromosome Inversions, Local Adaptation and Speciation. 574 Genetics 173, 419–434 (2006). 575 10. Baker, R. J. & Bickham, J. W. Speciation by monobrachial centric fusions. Proc. Natl. Acad. 576 Sci. U.S.A. 83, 8245–8248 (1986). 577 11. Baird, D. M. Telomeres and genomic evolution. Phil. Trans. R. Soc. B 373, 20160437 578 (2018). 579 12. Zhao, N. et al. Critically short telomeres derepress retrotransposons to promote genome 580 instability in embryonic stem cells. Cell Discov 9, 45 (2023). 581 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 27 13. Lee, C., Sasi, R. & Lin, C. C. Interstitial localization of telomeric DNA sequences in the 582 Indian muntjac chromosomes: further evidence for tandem chromosome fusions in the 583 karyotypic evolution of the Asian muntjacs. Cytogenet Genome Res 63, 156–159 (1993). 584 14. Chikashige, Y. Meiotic nuclear reorganization: switching the position of centromeres and 585 telomeres in the fission yeast Schizosaccharomyces pombe. The EMBO Journal 16, 193–202 586 (1997). 587 15. Bailey, J. A. & Eichler, E. E. Primate segmental duplications: crucibles of evolution, 588 diversity and disease. Nat Rev Genet 7, 552–564 (2006). 589 16. Mudd, A. B., Bredeson, J. V., Baum, R., Hockemeyer, D. & Rokhsar, D. S. Analysis of 590 muntjac deer genome and chromatin architecture reveals rapid karyotype evolution. Commun 591 Biol 3, 480 (2020). 592 17. Augustijnen, H. et al. A macroevolutionary role for chromosomal fusion and fission in 593 Erebia butterflies. Sci. Adv. 10, eadl0989 (2024). 594 18. Wright, C. J., Stevens, L., Mackintosh, A., Lawniczak, M. & Blaxter, M. Comparative 595 genomics reveals the dynamics of chromosome evolution in Lepidoptera. Nat Ecol Evol 8, 777–596 790 (2024). 597 19. Luo, J., Sun, X., Cormack, B. P. & Boeke, J. D. Karyotype engineering by chromosome 598 fusion leads to reproductive isolation in yeast. Nature 560, 392–396 (2018). 599 20. Dymond, J. S. et al. Synthetic chromosome arms function in yeast and generate phenotypic 600 diversity by design. Nature 477, 471–476 (2011). 601 21. Wang, L.-B. et al. A sustainable mouse karyotype created by programmed chromosome 602 fusion. Science 377, 967–975 (2022). 603 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 28 22. Zhang, X. M. et al. Creation of artificial karyotypes in mice reveals robustness of genome 604 organization. Cell Res 32, 1026–1029 (2022). 605 23. Wienberg, J., Jauch, A., Stanyon, R. & Cremer, T. Molecular cytotaxonomy of primates by 606 chromosomal in situ suppression hybridization. Genomics 8, 347–350 (1990). 607 24. Dutrillaux, B., Rethoré, M. O. & Lejeune, J. [Comparison of the karyotype of the orangutan 608 (Pongo pygmaeus) to those of man, chimpanzee, and gorilla]. Ann Genet 18, 153–161 (1975). 609 25. IJdo, J. W., Baldini, A., Ward, D. C., Reeders, S. T. & Wells, R. A. Origin of human 610 chromosome 2: an ancestral telomere-telomere fusion. Proc. Natl. Acad. Sci. U.S.A. 88, 9051–611 9055 (1991). 612 26. Ventura, M. et al. The evolution of African great ape subtelomeric heterochromatin and the 613 fusion of human chromosome 2. Genome Res. 22, 1036–1049 (2012). 614 27. Poszewiecka, B., Gogolewski, K., Karolak, J. A., Stankiewicz, P. & Gambin, A. 615 PhaseDancer: a novel targeted assembler of segmental duplications unravels the complexity of 616 the human chromosome 2 fusion going from 48 to 46 chromosomes in hominin evolution. 617 Genome Biol 24, 205 (2023). 618 28. Fan, Y., Linardopoulou, E., Friedman, C., Williams, E. & Trask, B. J. Genomic Structure and 619 Evolution of the Ancestral Chromosome Fusion Site in 2q13–2q14.1 and Paralogous Regions on 620 Other Human Chromosomes. Genome Res. 12, 1651–1662 (2002). 621 29. Fan, Y., Newman, T., Linardopoulou, E. & Trask, B. J. Gene Content and Function of the 622 Ancestral Chromosome Fusion Site in Human Chromosome 2q13–2q14.1 and Paralogous 623 Regions. Genome Res. 12, 1663–1672 (2002). 624 30. Martin, C. L. et al. The Evolutionary Origin of Human Subtelomeric Homologies—or Where 625 the Ends Begin. The American Journal of Human Genetics 70, 972–984 (2002). 626 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 29 31. Royle, N. J., Baird, D. M. & Jeffreys, A. J. A subterminal satellite located adjacent to 627 telomeres in chimpanzees is absent from the human genome. Nat Genet 6, 52–56 (1994). 628 32. Miga, K. H. Chromosome-Specific Centromere Sequences Provide an Estimate of the 629 Ancestral Chromosome 2 Fusion Event in Hominin Genomes. JHERED 108, 45–52 (2017). 630 33. Wang, H. et al. SVA Elements: A Hominid-specific Retroposon Family. Journal of 631 Molecular Biology 354, 994–1007 (2005). 632 34. Poszewiecka, B., Gogolewski, K., Stankiewicz, P. & Gambin, A. Revised time estimation of 633 the ancestral human chromosome 2 fusion. BMC Genomics 23, 616 (2022). 634 35. Yoo, D. et al. Complete sequencing of ape genomes. Preprint at 635 https://doi.org/10.1101/2024.07.31.605654 (2024). 636 36. Makova, K. D. et al. The complete sequence and comparative analysis of ape sex 637 chromosomes. Nature 630, 401–411 (2024). 638 37. Zhang, S. et al. Comparative genomics of macaques and integrated insights into genetic 639 variation and population history. Preprint at https://doi.org/10.1101/2024.04.07.588379 (2024). 640 38. Koga, A., Hirai, Y., Hara, T. & Hirai, H. Repetitive sequences originating from the 641 centromere constitute large-scale heterochromatin in the telomere region in the siamang, a small 642 ape. Heredity 109, 180–187 (2012). 643 39. Dutrillaux, B. Chromosomal evolution in Primates: Tentative phylogeny from Microcebus 644 murinus (Prosimian) to man. Hum Genet 48, 251–314 (1979). 645 40. Yunis, J. J. & Prakash, O. The Origin of Man: A Chromosomal Pictorial Legacy. Science 646 215, 1525–1530 (1982). 647 41. Liao, W.-W. et al. A draft human pangenome reference. Nature 617, 312–324 (2023). 648 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 30 42. Linardopoulou, E. V. et al. Human Subtelomeric WASH Genes Encode a New Subclass of 649 the WASP Family. PLoS Genet 3, e237 (2007). 650 43. Vollger, M. R. et al. Increased mutation and gene conversion within human segmental 651 duplications. Nature 617, 325–334 (2023). 652 44. Mao, Y. et al. A high-quality bonobo genome refines the analysis of hominid evolution. 653 Nature 594, 77–81 (2021). 654 45. Chiatante, G., Giannuzzi, G., Calabrese, F. M., Eichler, E. E. & Ventura, M. Centromere 655 Destiny in Dicentric Chromosomes: New Insights from the Evolution of Human Chromosome 2 656 Ancestral Centromeric Region. Molecular Biology and Evolution 34, 1669–1681 (2017). 657 46. Lonsdale, J. et al. The Genotype-Tissue Expression (GTEx) project. Nat Genet 45, 580–585 658 (2013). 659 47. Logsdon, G. A. et al. The structure, function and evolution of a complete human 660 chromosome 8. Nature 593, 101–107 (2021). 661 48. Luo, Y. et al. New developments on the Encyclopedia of DNA Elements (ENCODE) data 662 portal. Nucleic Acids Research 48, D882–D889 (2020). 663 49. Lacruz, R. S. et al. The evolutionary history of the human face. Nat Ecol Evol 3, 726–736 664 (2019). 665 50. Alemseged, Z. Reappraising the palaeobiology of Australopithecus. Nature 617, 45–54 666 (2023). 667 51. Broeils, L. A., Ruiz-Orera, J., Snel, B., Hubner, N. & Van Heesch, S. Evolution and 668 implications of de novo genes in humans. Nat Ecol Evol 7, 804–815 (2023). 669 52. Mayr, E. Michael J. D. White., Modes of Speciation. Systematic Biology 27, 478–482 (1978). 670 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 31 53. Trask B. Fluorescence. In: Birren B, Green ED, Hieter P, Klapholz S, Myers RM, Riethman 671 H, Roskams J, editors. Situ Hybridization. In Genome analysis: A laboratory manual. Cold 672 Spring Harbor, NY: Cold Spring Harbor Laboratory Press; 1999. pp. 303–413. 673 54. Hirai, H. et al. Structural variations of subterminal satellite blocks and their source 674 mechanisms as inferred from the meiotic configurations of chimpanzee chromosome termini. 675 Chromosome Res 27, 321–332 (2019). 676 55. Li, H. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics 34, 3094–677 3100 (2018). 678 56. Danecek, P. et al. The variant call format and VCFtools. Bioinformatics 27, 2156–2158 679 (2011). 680 57. Katoh, K. & Standley, D. M. MAFFT Multiple Sequence Alignment Software Version 7: 681 Improvements in Performance and Usability. Molecular Biology and Evolution 30, 772–780 682 (2013). 683 58. Capella-Gutiérrez, S., Silla-Martínez, J. M. & Gabaldón, T. trimAl: a tool for automated 684 alignment trimming in large-scale phylogenetic analyses. Bioinformatics 25, 1972–1973 (2009). 685 59. Minh, B. Q. et al. IQ-TREE 2: New Models and Efficient Methods for Phylogenetic 686 Inference in the Genomic Era. Molecular Biology and Evolution 37, 1530–1534 (2020). 687 60. Bouckaert, R. et al. BEAST 2.5: An advanced software platform for Bayesian evolutionary 688 analysis. PLoS Comput Biol 15, e1006650 (2019). 689 61. Sweeten, A. P., Schatz, M. C. & Phillippy, A. M. ModDotPlot—rapid and interactive 690 visualization of tandem repeats. Bioinformatics 40, btae493 (2024). 691 62. Numanagić, I. et al. Fast characterization of segmental duplications in genome assemblies. 692 Bioinformatics 34, i706–i714 (2018). 693 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 32 63. Quinlan, A. R. & Hall, I. M. BEDTools: a flexible suite of utilities for comparing genomic 694 features. Bioinformatics 26, 841–842 (2010). 695 64. Gu, Z., Gu, L., Eils, R., Schlesner, M. & Brors, B. circlize implements and enhances circular 696 visualization in R. Bioinformatics 30, 2811–2812 (2014). 697 65. Galili, T. dendextend: an R package for visualizing, adjusting and comparing trees of 698 hierarchical clustering. Bioinformatics 31, 3718–3720 (2015). 699 66. Smith, M. R. Information theoretic generalized Robinson–Foulds metrics for comparing 700 phylogenetic trees. Bioinformatics 36, 5007–5013 (2020). 701 67. Altemose, N. et al. Complete genomic and epigenetic maps of human centromeres. Science 702 376, eabl4178 (2022). 703 68. Altemose, N., Miga, K. H., Maggioni, M. & Willard, H. F. Genomic Characterization of 704 Large Heterochromatic Gaps in the Human Genome Assembly. PLoS Comput Biol 10, e1003628 705 (2014). 706 69. Hach, F. et al. mrsFAST: a cache-oblivious algorithm for short-read mapping. Nat Methods 707 7, 576–577 (2010). 708 70. Chen, S. Ultrafast one‐pass FASTQ data preprocessing, quality control, and deduplication 709 using fastp. iMeta 2, e107 (2023). 710 71. Kim, D., Paggi, J. M., Park, C., Bennett, C. & Salzberg, S. L. Graph-based genome 711 alignment and genotyping with HISAT2 and HISAT-genotype. Nat Biotechnol 37, 907–915 712 (2019). 713 72. Love, M. I., Huber, W. & Anders, S. Moderated estimation of fold change and dispersion for 714 RNA-seq data with DESeq2. Genome Biol 15, 550 (2014). 715 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 33 73. Chen, Y., Lun, A. T. L. & Smyth, G. K. From reads to genes to pathways: differential 716 expression analysis of RNA-Seq experiments using Rsubread and the edgeR quasi-likelihood 717 pipeline. F1000Res 5, 1438 (2016). 718 74. Sherman, B. T. et al. DAVID: a web server for functional enrichment analysis and functional 719 annotation of gene lists (2021 update). Nucleic Acids Research 50, W216–W221 (2022). 720 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 34 721 Figure 1. The comparative sequencing analysis of primate chromosome 2. (a) The syntenic 722 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 35 comparison of chromosome 2 highlights the extent of primate chromosome 2 evolutionary 723 rearrangements. Syntenic regions conserved in order (blue) are contrasted with evolutionary 724 inversions (orange) and non-syntenic regions (breaks) corresponding to alpha satellite (dark 725 brown), pCht subterminal satellite (green), and other satellite DNA (pink). B. orangutan and S. 726 orangutan represent Bornean orangutan and Sumatran orangutan, respectively. (b) High-727 resolution analysis of the human fusion site (chr2:113,940,058-114,049,496, colored in amber) 728 shows the non-syntenic breakpoint region in the context of annotated human protein-coding 729 genes and segmental duplications (SDs). Pan SD spacers (purple blocks) in chimpanzee pCht 730 show the homology of the partial region of human fusion site. 731 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 36 732 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 37 Figure 2. The genomic structure and evolutionary history of SDs at the human fusion site. 733 (a) Human chromosome 2 fusion site and SD organization. A human genomic segment 734 (chr2:113,650,520-114,212,116) with gene annotations shows the region where chromosome 2a 735 and 2b fused (chr2:113,940,058-114,049,496, amber). The region consists of a large (~455 kbp) 736 duplication block made up of seven SDs with variable copy numbers in each ape genome. The 737 three SDs at the fusion site are designated as SD_fusion_A, SD_fusion_B, and SD_fusion_C. 738 The table indicates the copy number of the full-length homologous segments for each primate 739 genome (with brackets indicating the SD_fusion_A copy number of the derived sequence in the 740 subterminal heterochromatic caps of chimpanzees and bonobo; see Supplementary Table 2 for 741 the detailed alignments at the fusion site). (b-g) Phylogenetic trees based on a (b) 53 kbp 742 (SD_fusion_A), (d) 68 kbp (SD_fusion_B), and (f) 23 kbp (SD_fusion_C) multiple sequence 743 alignment show that these regions have been subject to incomplete lineage sorting (ILS). The 744 chromosome schematic depicts the genomic locations of SD_fusion_A (c), SD_fusion_B (e), and 745 SD_fusion_C (g) in each primate. (h) Proportion of human–gorilla and Pan–gorilla tree topology 746 in a 500 bp window of 5 Mbp flanking region near the fusion site. The red dotted line represents 747 the mean value of the whole chr2. (i) Fold change of the mean proportion of discordant 748 topologies in the flanking region compared with the whole genome average. (Note, chromosome 749 names in the great apes refer to the human homologous chromosome also known as the 750 phylogenetic group designation.) 751 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 38 752 Figure 3. Rapid turnover of subtelomeric repetitive regions in primates and evolutionary 753 connection between SD spacers in Pan lineage and SD_fusion_A at the fusion site. (a) The 754 identity heatmap of the subterminal heterochromatic cap for the p-arm of chimpanzee chr2b. SD 755 spacer elements (purple) are annotated below and are flanked by large tracts of pCht satellite 756 (green) (zoomed-in panel shows the structure at higher resolution). (b) Circos plot shows 757 sequence identity of chr2a and chr2b among the apes highlighting the rapid turnover of satellite 758 DNA. The three layers (outer to inner) represent the subtelomeric region, tandem repeat 759 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 39 satellites, and transposon element annotations. (c) Dot plot shows the synteny between a bonobo 760 SD spacer within the heterochromatic cap vs. human SD_fusion_A segment. The sequence 761 identity was calculated at a 100 bp resolution, with the dot color representing sequence identity 762 ranging from 80% (purple) to 100% (red). (d) Phylogenetic topology comparison shows nearly 763 consistent ILS evolutionary pattern of SD_fusion_A (excluding spacer sequence) and SD spacer. 764 Each node is supported by a bootstrap value of 100/100. The red lines indicate discordant tree 765 topologies within each major clade between the two trees.766 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 40 767 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 41 Figure 4. Comparative analysis of active and inactive centromeric regions of human 768 chromosome 2. (a) Genomic structure and suprachromosomal family (SF) annotation of 769 centromeres in human chr2, degenerate site, and NHP chr2a/chr2b. Different SFs are shown by 770 various colors, with centromere dip regions (CDRs) marked by arrows. (b) Comparison of the 771 centromere in chimpanzee chr2b with the human degenerate centromeric region. The middle 772 panel shows the SF annotations for the chimpanzee centromere and human centromere 773 degenerate region. Potential α-satellite deletion breakpoints (b1 and b2) are indicated with dotted 774 lines. The heatmaps of each region are shown in upper and lower, respectively. (c) The panel 775 illustrates five distinct structural haplotypes of the degenerate centromere site in humans. (d) The 776 panel shows the lengths of three α-satellite arrays in humans. (e) The methylation degree of HOR 777 across NHPs and of three α-satellite arrays at the degenerate centromere site. The red triangle 778 represents the methylation frequency of CDR in each chromosome (Met. Freq. strands for 779 methylation frequency). 780 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 42 781 Figure 5. Fusion site knockout and gene expression alteration. (a) Schematic representation 782 of the depletion experiments performed in HEK293T cell lines. (b) Two distinct sgRNA pairs 783 (L-sg1 and R-sg; L-sg2 and R-sg) were designed for the depletion experiments (left panel). Red 784 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 43 sequences indicate sgRNA, while blue sequences indicate PAM sites. (c) The PCR analysis 785 confirms heterozygous depletion in cell lines. WT represents the wild-type cell lines. (d) and (e) 786 Differentially expressed genes (DEGs) are shown in volcano plots (left panel for the L-sg1 and 787 R-sg pair, right panel for the L-sg2 and R-sg pair). Genes with log2 fold change (log2FC) 788 values >2 and FDR <0.05 are highlighted in red (upregulated) and blue (downregulated). The 789 overlap of top 10 DEGs in each group are emphasized in deeper red/blue, with gene names 790 indicated. (f) Venn diagrams illustrate the overlap of upregulated and downregulated genes 791 between both depletion groups. (g) The RT-PCR analyses confirmed 5 of 10 top-rank DEGs. WT 792 represents the relative gene expression levels normalized to GAPDH in wild-type cell lines. 793 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 44 794 Figure 6. Models for chromosome 2 evolution. Two different models are depicted for the 795 origin of the human chromosome 2 fusion. (a) In the last common ancestor of humans, bonobos, 796 chimpanzees, and gorillas, diverse ancestral genomic structures emerged at the ends of chr2a and 797 chr2b. In the human ancestral lineage, chr2a and chr2b consisted of complex SD blocks that 798 emerged as a result of ILS, and then became juxtaposed by a telomere-to-telomere fusion. (b) 799 Rudimentary pCht sequences were present in the ancestral lineage of humans and chimpanzees 800 in association with SD_fusion_A. Nonallelic homologous recombination (NAHR) occurred 801 between these in the human lineage eliminating pCht in humans but retaining SD_fusion_A with 802 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint 45 other SDs subsequently duplicating to the location. In chimpanzees, an independent expansion of 803 the SD spacers and pCht sequences occurred leading to the formation of the subterminal 804 heterochromatic caps. 805 .CC-BY-NC-ND 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-pdf

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-22T02:00:06.705733+00:00
License: CC-BY-NC-ND-4.0