Abstract
37
All great apes differ karyotypically from humans due to the fusion of chromosomes 2a and 2b, 38
resulting in human chromosome 2. Yet, the structure, function, and evolutionary history of the 39
genomic regions associated with this fusion remain poorly understood. Here, we analyze finished 40
telomere-to-telomere chromosomes in great apes and macaques to show that the fusion was 41
associated with multiple pericentric inversions, segmental duplications (SDs), and the rapid 42
turnover of subterminal repetitive DNA. We characterized the fusion site at single-base-pair 43
resolution and identified three distinct SDs that originated more than 5 million years ago. These 44
three distinct SDs were differentially distributed among African great apes as a result of 45
incomplete lineage sorting (ILS) and lineage-specific duplication. Most conspicuously, one of 46
these SDs shares homology to a hypomethylated SD spacer sequence present in hundreds of 47
copies in the subterminal heterochromatin of chimpanzees and bonobos. The fusion in human 48
was accompanied by a systematic degradation of the three divergent α-satellite arrays 49
representing the ancestral centromere creating five distinct structural haplotypes in humans. 50
CRISPR/Cas9-mediated depletion of the fusion site in human cell lines significantly alters the 51
expression of 108 genes, indicating a potential regulatory consequence to this human-specific 52
karyotypic change. 53
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
4
Introduction
54
Karyotype evolution is a critical aspect of evolutionary biology because it has been associated 55
with speciation, adaptation, and disease1–7. In 1935, Wilson and Painter first observed that 56
chromosome fusion could lead to speciation in flies7. As the field advanced, numerous instances 57
of chromosome fusion have been documented in plant and animal speciation1–4,8–10. Several 58
molecular mechanisms have been proposed as key drivers of chromosome fusion in speciation 59
and oncogenesis, including telomere-telomere, telomere-centromere, tandem repeat, segmental 60
duplication (SD), and retrotransposon expansions and fusions11–15. 61
62
Recent advances in genome editing and synthetic biology have provided deeper insights into the 63
role of chromosome fusion in evolution16–18. Boeke and colleagues, for example, demonstrated 64
that chromosome fusion could induce reproductive isolation in yeast, highlighting the potential 65
of these genomic changes to influence speciation19,20. Additionally, studies on chromosome 66
engineering in mice reveal significant chromatin conformation alterations in the chromosome 67
fusion regions, further emphasizing the impact of karyotype evolution on genetic and phenotypic 68
diversity21,22. 69
70
Human chromosome 2 (chr2) was formed from the fusion of nonhuman primate (NHP) chr2a 71
and chr2b23–25, representing arguably the most significant karyotypic differences between 72
humans and NHPs. Ijdo and Wienberg were the first to identify the fusion site at human 73
chromosome 2q13-2q14 using cytogenetic banding approaches25. Subsequent studies uncovered 74
dispersed SDs associated with this fusion site26,27. These analyses relied on partial genome 75
assemblies generated using short-read technologies or Sanger sequencing of bacterial artificial 76
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
5
chromosome (BAC) clones, often resulting in incomplete genetic information at the fusion site26–77
29. Consequently, the structure, evolutionary history, and function of the fusion site have 78
remained only partially understood. 79
80
Available genomic assemblies and FISH experiments have shown an abundance of satellite 81
sequences in the subtelomeric repetitive regions of NHP chr2a and chr2b, whereas such 82
sequences are absent in humans26,30,31. During the fusion of human chr2, one of the centromeres 83
in the fused chromosome becomes inactive, and degraded with independent transposable element 84
(TE) retrotransposition events occurring at the degenerate site32. In addition, there has been 85
considerable debate regarding the timing of the chr2 fusion event. Based on SD and SVA (SINE-86
VNTR-Alu) divergence, the fusion event was estimated to have occurred early in human 87
evolution, 5-7 million years ago (mya) and 2.5-4.5 mya, respectively26,33. However, a recent 88
study, utilizing clustered substitution statistics, proposed a much more recent origin of 89
approximately 0.9 mya34. 90
91
The complex repetitive structure at the fusion site, subtelomeric repetitive regions, as well as the 92
inactive centromere have made it challenging to reconstruct the evolutionary history and 93
potential functional consequences of the fusion event. Here, we leverage the complete sequences 94
of great apes35,36 and a macaque37 to revisit the structure and evolutionary history of human chr2. 95
We aim to: (1) characterize the fusion site at a single-base-pair resolution, (2) study this in the 96
context of epigenetic and structural changes occurring at the subtelomeric repetitive regions and 97
degenerate centromere site, (3) use these data to create a model for human chr2 evolution, and 98
(4) examine the potential functional consequences by creating fusion-site depletion cell lines. 99
100
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
6
Results
101
Comparative sequence analysis of the human chromosome 2 fusion site. We performed a 102
comparative analysis of 13 finished nonhuman great ape chromosomes35 and the finished 103
macaque genome37 to human chr2 (Figure 1a, Methods). We identified numerous non-syntenic 104
segments (76-154 regions ≥10 kbp in length) between humans and NHPs, accounting for 9.86 105
Mbp (macaque) and up to 57.5 Mbp (gorilla) of unalignable chromosomal sequence per species. 106
Most of this sequence corresponded to various classes of repetitive DNA (Figure 1a, 107
Supplementary Table 1), including satellite repeats, tandem and interspersed SDs. In particular, 108
among the nonhuman African great apes, both chr2a and chr2b are acrocentric with the presence 109
of subterminal heterochromatic caps (9.6-19.6 Mbp) demarcating the ends of the chromosomes 110
and are composed of a 32 bp AT-rich satellite DNA (pCht) and SD spacer regions26,31,35,38. In 111
addition, there are multiple pericentric and paracentric evolutionary inversions distinguishing the 112
ape lineages from each other39, several of which share homology to the SD flanking regions of 113
the fusion site (Figure 1b and Figure 2a). One pericentric inversion occurred in the common 114
ancestor of African great apes after divergence from orangutans (chr2b) while another occurred 115
before the Pan–gorilla split (chr2a)39,40 (Figure 1a). Furthermore, we identified a truncated SD 116
containing the partial CBWD2 at the chr2a pericentric inversion breakpoints, and another 117
truncated SD with FOXD4L1 at the chr2b breakpoint (Supplementary Figure 1)28,29. 118
119
Evolutionary reconstruction of the ancestral fusion site. We further compared human chr2 120
with Pan chr2a and chr2b to precisely characterize the ~109 kbp fusion site in the human T2T-121
CHM13v2.0 genome assembly (2q14.1, chr2:113,940,058-114,049,496) (Figure 1b and 122
Supplementary Figure 2). Leveraging data from 47 human genomes from the Human Pangenome 123
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
7
Reference
Consortium (HPRC)41, we confirmed that there is a complete absence of large 124
structural variants within this fusion site in 94 sequence-resolved human haplotypes, suggesting 125
a fixed structural haplotype for the entire fusion site (Supplementary Figure 3). Neither 126
nucleotide diversity (π) nor Tajima’s D values show significant reductions at the fusion site in 127
African or non-African populations, relative to the chr2 average (Supplementary Figure 4) 128
consistent with the locus evolving neutrally. 129
130
Within the larger ~455 kbp duplication block (seven independent SD blocks) demarcating the 131
fusion site region, we narrowed the fusion site to a ~109 kbp segment consisting of three distinct, 132
high-identity SDs (≥98% identity and ≥20 kbp in length) (Figure 1b, Figure 2a, Supplementary 133
Figures 5-12, and Supplementary Tables 2-3). In humans, these SDs share homology with 134
sequences corresponding to chr9 (SD_fusion_A, 50 kbp), chr12/20 (SD_fusion_B, 36 kbp), and 135
chr22 (SD_fusion_C, 22 kbp) (Figure 2a, Supplementary Table 2). Among African apes, each 136
SD is highly variable in copy number showing evidence of shared ancestral locations as well as 137
lineage-specific duplications (Supplementary Table 2). Using macaque and orangutan as 138
outgroups, we reconstructed the evolutionary history of each SD estimating the divergence 139
timepoint from each SD to its nearest nonhuman great ape neighbor (Figure 2b-g, Methods). 140
141
SD_fusion_A (50 kbp) is present as a single copy in macaques and orangutans and corresponds 142
to a partial duplication of PGM529 (14 exons of an ancestral gene), a gene involved in 143
carbohydrate metabolism, originating from African great ape ancestral chr9 (phylogenetic group 144
IX) (Figure 2b-c). This SD began to duplicate in an interspersed configuration in the common 145
ancestral lineage of African great apes (~7.4 mya) with the locus expanding to highest copy 146
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
8
number in chimpanzees and bonobos where it defines the SD spacer region demarcating large 147
blocks of pCht on chr2a and chr2b and other chromosomes (chr3, chr5-chr13, chr15-chr22, and 148
chrX in chimpanzee; each chromosome in the bonobo genome) (Figure 2b-c and Supplementary 149
Figure 6). Of note, the SD_fusion_A (chr2) and other human copies share a monophyletic origin 150
(~3.3 mya) most closely related to gorilla copies ~6 mya (95% CI: 4.8-7.4 mya) instead of 151
chimpanzee or bonobo—a pattern consistent with incomplete lineage sorting (ILS). 152
153
Similarly, SD_fusion_B, is a 36 kbp segment corresponding to the partially truncated WASH242 154
(11 exons of an ancestral gene that potentially regulates actin cytoskeleton dynamics), FAM138B 155
(3 exons of an ancestral gene of unknown function), and DDX11 (3 exons of an ancestral gene 156
implicated in DNA metabolism). The segment exists as a single copy in both macaques and 157
orangutans and originated from a locus mapping to ancestral African great ape chr12 158
(phylogenetic group XII) (Figure 2d-e and Supplementary Figure 7). All human copies mapping 159
to human chr2, chr20, chr12, and chr9 show a monophyletic origin (~4.4 mya, 95% CI: 3.1-5.8 160
mya) suggesting human-specific duplications or interlocus gene conversion events43. With 161
respect to nonhuman African great apes, however, this clade diverges from a distinct 162
monophyletic clade that includes chimpanzee, gorilla, and bonobo (~7.4 mya, 95% CI: 5.4-9.3 163
mya) (Figure 2d). This topology is once again consistent with ILS. 164
165
In contrast to SD_fusion_A and SD_fusion_B, SD_fusion_C (the most distal, 22 kbp) shows 166
evidence of independent duplication in all primates, including macaques and orangutans, with 167
the exception of chimpanzees where it exists as a single copy on chimpanzee chr6 168
(Supplementary Figure 8). All human SDs show a monophyletic origin (~3.5 mya, 95% CI: 2.7-169
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
9
4.3 mya) but show a deep coalescence with another genomic segment present only on bonobo 170
chr22 (phylogenetic group XXII), dating back to ~6.9 mya (95% CI: 5.3-8.6 mya) (Figure 2f). 171
Notably, SD_fusion_C and its flanking region correspond to a larger SD that aligns exclusively 172
between the human chr2 SD_fusion_C site and bonobo chr22, but not with any other NHPs 173
(Supplementary Figure 13). Thus, SD_fusion_C and its flanking region likely represent an 174
ancestral genomic structure in the latest common ancestor (LCA) of African great apes, which 175
subsequently sorted in the human and bonobo lineages. 176
177
We extended the ILS analysis to the 5 Mbp mapping proximally (chr2a) and distally (chr2b) to 178
the fusion site in humans (Methods) using a 500 bp windowed approach44 (Supplementary Table 179
4). We observe a sharp transition in the proportion of ILS windows at the site of the chr2 fusion. 180
Specifically, we find the proportion of ILS rises to 45.6% (a 1.39-fold excess compared to the 181
chr2 average) as we approach the fusion site distally (chr2b side). In contrast, the proximal 182
portion appears depleted for ILS segments (Figure 2h-i, Supplementary Figure 14 and 183
Supplementary Table 5). There is, thus, a polarized pattern of ILS with maxima and minima 184
occurring on either side of the fusion site. 185
186
The refined analysis of telomeric sequences at the fusion site. Previous investigations 187
identified inverted telomeric sequences (TTAGGG/CCCTAA) at the fusion site, suggesting a 188
telomere-to-telomere (T2T) fusion25. We confirmed the presence of interstitial telomeric 189
sequences (CCCTAA, chr2: 114,027,659-114,028,207) at SD_fusion_C and its corresponding 190
SD on human chr22 (TTAGGG, chr22:51,254,086-51,323,279) and similar sequences in bonobo 191
chr22 (TGAGGG, chr22: 60,722,979-60,724,048). The sequence (GGGTTA) was identified at 192
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
10
SD_fusion_B (chr2: 114,027,333-114,027,658) and its corresponding SD on human chr12 193
(TAACCC, chr12:2,843-3,030) and bonobo chr20 (TAACCC, chr20:1,139,021-1,139,266), but 194
not at its corresponding SD on human chr20 (Supplementary Figure 15). These observations 195
argue that these telomeric sequences were present on SD_fusion_B and SD_fusion_C prior to the 196
fusion event. In addition, two telomeric-associated repeats (TAR1) flank the telomeric sequences 197
of SD_fusion_B and SD_fusion_C27. Given the genomic structure of TAR1 and telomeric 198
sequences on these SDs (Supplementary Figures 15-16), we propose that, in the ancestral 199
configuration, each SD contained a TAR1 element and telomeric sequences, arranged in an 200
inverted orientation on ancestral chr2a and chr2b. The end-to-end fusion of these two distinct 201
SDs may have facilitated the fusion of ancestral chr2a and chr2b, leading to the vestigial 202
presence of two TAR1 elements and the (CTAACC) and (GGGTTA) repeat motifs at the human 203
fusion site. 204
205
African great ape subterminal satellite repeat expansion and human chromosome 2 fusion. 206
Reconstructing the evolutionary history of the fusion event has been challenging due to extensive 207
SD and lineage-specific turnover of the subterminal heterochromatic caps in Pan and gorilla35 208
(Supplementary Table 6). With the exception of the short arm of gorilla chr2a, the corresponding 209
regions in both Pan and gorilla are composed of nearly continuous megabase-pair tracts of 210
satellite DNA interspersed with SD spacers35. The satellite sequence is made of a tandem 32 bp 211
repeat motif (pCht)31 punctuated on average every 287 kbp in Pan and every 389 kbp in gorilla 212
by a SD spacer (Figures 3a, Supplementary Figures 17-19). The interrupting SD spacers are 213
variable with a modal length of 32 kbp in the Pan lineage and 33.7 kbp in gorilla (Supplementary 214
Figure 17) and correspond to hypomethylated pockets flanked by the hypermethylated satellite 215
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
11
DNA35 (Figure 3a). Although both the gorilla and chimpanzee SD spacers differ in sequence 216
composition and the subterminal heterochromatic cap is largely thought to have evolved 217
independently in both lineages26,35, the net effect is that the subterminal portions of both 218
chimpanzee chr2a and chr2b share 23.6 Mbp of high-identity sequence homology involving both 219
the SD spacers and pCht satellite DNA. In the case of the gorilla, the homology is restricted to 220
the pCht satellite DNA for both arms of gorilla chr2b and the q-arm of gorilla chr2a. The 221
organization of the p-arm of gorilla chr2a differs considerably and is much more similar to the 222
organization found in orangutan, which is enriched in HSatIII-like repeat sequences (Figure 3b). 223
It is classified as an acrocentric short arm lacking a nuclear organizing region35. 224
225
Importantly, the SD spacer that expanded in the subterminal heterochromatic caps in chimpanzee 226
and bonobo shares 98.5% identity with the SD_fusion_A segment mapping at the fusion site on 227
human chr2 (Figure 3c). All 429~452 SD spacers within the Pan pCht regions, including those at 228
the ends of these chromosomal regions, are monophyletic in origin estimated to have expanded 229
approximately ~5.5 mya (95% CI: 4.2-6.8 mya). This subterminal SD spacer (32 kbp) in Pan is 230
18 kbp smaller than SD_fusion_A (50 kbp) where the shared LCA predates the human–Pan–231
gorilla divergence (Figure 2b and 2c). Comparing the phylogenies of the sequence unique to 232
SD_fusion_A and sequence shared with the subterminal heterochromatic cap SD spacers shows 233
nearly coincident ILS topology (generalized Robinson–Foulds distance=0.28, p=1.6×10-4). This 234
indicates that the SD spacers within Pan heterochromatic caps are derived from the duplicated 235
sequence that gave rise to ancestral SD_fusion_A. Thus, the hyperexpansion of subterminal 236
satellite DNA associated with subterminal heterochromatic caps in Pan and the chr2 fusion are 237
linked genetically with two different evolutionary trajectories and karyotypic consequences in 238
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
12
human and Pan (Figure 2c-d, Supplementary Figure 20 and Supplementary Table 7). 239
240
Centromere retention and degeneration in human chromosome 2. An important consequence 241
of the human chr2 fusion was that the ancestral chr2b centromere in NHPs became inactive in 242
our human lineage (Figure 4a). Chr2a differs from chr2b by the presence of large tracts of HSatII 243
arrays in humans and HSatIII arrays in Pan (Supplementary Figures 21-22). Overall, the active 244
human centromere in humans is more similar to that of gorilla with respect to suprachromosomal 245
family (SF) organization. Both humans and gorillas possess SF2, whereas chimpanzees and 246
bonobos have SF3. 247
248
We compared the active chromosome centromere of chimpanzee chr2b with the structure of the 249
vestigial centromere in humans. We identified three distinct α-satellite arrays (chr2:132,644,386–250
132,685,996) in humans45 with homology to chimpanzee that had been interrupted by various 251
simple repeats (CCTCTC) and retrotransposon elements (SVA and L1PA3) in the human lineage 252
(Figure 4b). All three satellite arrays were derived from divergent monomeric α-satellite regions 253
ancestral to human and chimpanzee, rather than from higher-order repeats (HORs). Our analysis 254
of 94 human genome assemblies39 identifies five distinct structural haplotypes due primarily to 255
length variation of the first two α-satellite arrays (Figure 4c-d and Supplementary Table 8). A 256
sequence comparison using unique k-mers from this region shows that ~97.5% (8,246/8,460) and 257
~96.7% (8,181/8,460) are also identified in Neanderthal and Denisovan genomes, respectively 258
(Methods), confirming that the centromeric degeneration occurred long before the divergence of 259
modern and archaic humans32. 260
261
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
13
We also compared methylation patterns of the NHP chr2a and chr2b centromeres (Figure 4e). 262
The human α-satellite arrays mapping to the degenerate site show significantly lower 263
methylation levels, when compared to typical NHP HORs (excluding centromere dip regions 264
[CDRs], p=0.02), with the exception of bonobo chr2b and gorilla chr2b. HORs in bonobo chr2b 265
and gorilla chr2b are significantly hypomethylated when compared to other NHP HORs 266
(p=1.92×10-6), possibly due to the smaller size of these HORs (lengths: 71 kbp for bonobo chr2b, 267
and 106 kbp for gorilla chr2b). 268
269
Functional assessment of the fusion site by depletion. Four putative noncoding 270
genes/pseudogenes (PGM5P4, FAM138B, WASH2P, and DDX11L2) have been annotated at the 271
site of the human chr2 fusion. According to GTEx short-read RNA sequencing (RNA-seq), these 272
four genes/pseudogenes are expressed in testis, esophagus, fallopian tube, and cerebellum 273
tissues46 (Supplementary Figure 23). Further, both PGM5P4 and WASH2P are supported by 274
long-read Iso-Seq transcript data from CHM13hTERT47 and kidney tissue from ENCODE48. In 275
addition, methylation analysis demarcates a prominent CpG island showing the 276
promoters/enhancers of PGM5P4, as identified using ONT reads from the T2T-CHM13 cell 277
line47 (Supplementary Figure 24). 278
279
To explore the potential function of the fusion site, we used CRISPR/Cas9 to delete this region 280
in a kidney-derived epithelial-like cell line (HEK293T) (Figure 5a). Two independent pairs of 281
single guide RNA (sgRNA; L-sg1 and R-sg; L-sg2 and R-sg) were designed for the depletion 282
(Figure 5b), and the heterozygous depletion of the fusion site in both sgRNA groups was 283
confirmed by PCR and Sanger sequencing (Figure 5c and Supplementary Table 9). We then 284
conducted RNA-seq, generating 53.28 Gbp, 67.42 Gbp, and 122.41 Gbp of data for the L-sg1 285
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
14
depletion, L-sg2 depletion, and wild-type (WT) cell lines (three biological repeats for each 286
depletion cell line and six biological repeats for WT cell lines) (Supplementary Table 10). 287
288
Differential gene expression analysis identified 37 upregulated and 71 downregulated genes in 289
both depletion cell lines (Supplementary Figure 25 and Supplementary Table 11, Methods). 290
These genes were significantly enriched in pathways related to the regulation of transcription by 291
RNA polymerase II (p=2.2×10-3). Among the top ten differentially expressed genes (DEGs) in 292
both depletion cell lines, three were upregulated and seven were downregulated (Figure 5d-e). 293
We further validated five of ten gene expression changes using RT-PCR in the depletion cell 294
lines (Figure 5g and Supplementary Table 12). 295
296
Discussion
297
Complete T2T sequence for chromosome 2, 2a, and 2b among the great apes35–37 allowed us to 298
systematically examine the complex genomic architecture and further refine the evolutionary 299
history of the human-specific chr2 fusion event. There are three important conclusions. First, the 300
fusion event was intimately associated with SDs that have been restructuring genomes and 301
chromosomes throughout great ape evolution26,28. The 109 kbp fusion site in humans consists of 302
three independent SDs, each with distinct trajectories in different ape lineages, and these are 303
further embedded in a larger duplication block of ~455 kbp. All SDs are highly variable in copy 304
number among the great apes, and many have been reused as breakpoint sequences during great 305
ape evolution. Although cytogenetic studies previously reported two pericentric inversions on 306
ancestral chr2a and chr2b26, we observe that the flanking region at the human fusion site (chr2a, 307
proximal side) shares 98.4% identity with an SD flanking the pericentric inversion that 308
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
15
distinguishes Sumatran orangutan chr2b from gorilla chr2b. Moreover, SD_fusion_A shares 309
97.8% identity with the SD spacer sequence of the Pan subterminal heterochromatic caps where 310
it expanded to hundreds of copies in the chimpanzee and bonobo lineages but not in gorilla. 311
312
Second, we show that the fusion site has been strongly subjected to ILS. All three SDs, for 313
example, show phylogenetic signatures consistent with ILS with a maximum occurring at the 314
fusion site suggesting that the fusion event occurred >5 mya. Thus, the fusion event is not a 315
recent evolutionary event but potentially occurred during African great ape speciation when the 316
effective population size was predicted to be much larger than contemporary ape populations35,44. 317
Given fossil evidence that Australopithecus existed around 2-4 mya, Paranthropus around 1-3 318
mya, and the earliest Homo fossils around 2-3 mya49–51, we speculate that this fusion event did 319
arise in the genus Homo but rather occurred in ancestral great ape populations that would give 320
rise to humans potentially by creating a stasipatric speciation barrier52. 321
322
Third, the fusion site was associated with extensive subtelomeric satellite turnover and 323
epigenetic differences among great apes. The two satellite motifs present in the subtelomeric 324
regions of gorilla chr2a and chr2b are distinct from each other: one resembles the subtelomeric 325
repetitive regions of orangutans, while the other resembles the genomic architecture of Pan 326
species. SD_fusion_A, in part, defines the chr2 fusion site but also associates with large tracts of 327
hypermethylated pCht satellite repeats that define subterminal heterochromatin in chimpanzees 328
and bonobos35 (Supplementary Figure 26). The juxtaposition of novel SDs at the fusion site and 329
the shift from heterochromatic or acrocentric DNA at the termini of chr2a and chr2b led to 330
methylation differences among the great apes at this locus. Indeed, our depletion experiments in 331
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
16
humans show consistent gene expression changes suggesting that the euchromatization of the 332
fusion site may have had more global genome-wide regulatory effects. 333
334
We propose two evolutionary scenarios for the formation of human chr2. In the first, the large 335
effective population size of the human–Pan–gorilla ancestral population facilitated the 336
coexistence of several NHP chr2a and chr2b subtelomeric structural haplotypes 5-7 mya. In one 337
of these structural configurations, SD_fusion_A and SD_fusion_B became juxtaposed and 338
duplicated to the subtelomeric region of human–Pan–gorilla ancestral chr2a prior to the fusion 339
event (Figure 6a). Similarly, SD_fusion_C was duplicated to the subtelomeric region of ancestral 340
chr2b (where the full-length structure is still retained subtelomerically in bonobo); yet, it is no 341
longer located on other nonhuman great ape chr2b due to the rapid exchange of subtelomeric 342
SDs and evolutionary turnover of satellite DNA26,53. Subsequently, the telomeric and TAR1 343
sequences in SD_fusion_B and SD_fusion_C mediated the fusion (Figure 6a). In the Pan and 344
gorilla lineages, chr2a and chr2b (only in chimpanzee and bonobo) experienced a different 345
evolutionary trajectory associated with the formation of subterminal heterochromatic caps. 346
Previous studies26,35 have shown that the heterochromatic caps likely evolved independently in 347
both gorilla and chimpanzee, albeit convergently with a similar architectures where hundreds of 348
kilobase pairs of satellite pCht DNA are punctuated by a ~30 kbp SD spacer region that defines a 349
pocket of hypomethylation. The SD spacers in gorilla and chimpanzee subterminal 350
heterochromatin caps are distinct but in both chimpanzee and bonobo the SD spacers represent a 351
derivative of the SD_fusion_A confirming that this sequence was present subtelomerically for 352
chr2a in the human–Pan ancestor. SD_fusion_A subsequently expanded to hundreds of copies in 353
chimpanzee and bonobo in conjunction with the subterminal heterochromatic satellite. Yet, in the 354
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
17
human lineage, SD_fusion_A remains a low-copy SD with no trace of pCht satellite sequence in 355
the human genome35. 356
357
Alternatively, the SDs evolved similarly, but in the common ancestor of human and Pan, there 358
was an incipient association with rudimentary pCht satellite sequences with the SD_fusion_A 359
sequence that mediated exchange between ancestral chr2a and chr2b via nonallelic homologous 360
recombination (NAHR) or ectopic exchange. Subsequently, during NAHR between full-length 361
SD_fusion_A elements, the pCht regions were lost in the human lineage and two additional SDs 362
were inserted at the fusion breakpoint region (Figure 6b). In the Pan lineage, a portion of the 363
SD_fusion_A became hyperexpanded defining the subterminal SD spacer region of all 364
heterochromatic caps in bonobo and chimpanzee. In support of this model, it is known that in 365
both chimpanzee and gorilla the subterminal satellite DNA forms unique post-bouquet structures 366
in germ cells and are hotspots of ectopic exchange between nonhomologous chromosomes54. If 367
such exchanges occurred at the edge of the incipient subterminal heterochromatic caps before 368
there were many copies of the subterminal satellite DNA, it would help explain the absence of 369
pCht sequence in the human genome—i.e., the fusion of chr2a and chr2b helped eliminate the 370
potential for the formation of subterminal heterochromatic caps. 371
372
Irrespective of the model, there is evidence that SDs in the LCA and ILS potentially facilitated 373
by a large effective population size played a key role35,44. These aspects helped maintain and 374
diversify karyotypic structures of the ancestral chr2a and chr2b for possibly millions of years. 375
The divergent fates of the subtelomeric repetitive regions of these ancestral chromosomes, e.g., 376
pCht expansion and SD spacer insertions in the Pan lineage, or the SD insertions in the human 377
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
18
fusion site, moreover, may have driven the speciation between humans and NHPs through ILS 378
(Supplementary Figure 27). In this light, it is interesting that we document an asymmetric ILS 379
pattern, with an increase in ILS segments on the distal side (chr2b) and a decrease on the 380
proximal side (chr2a) of the fusion site. This suggests diverse evolutionary trajectories for the 381
ancestral subtelomeric regions of NHP chr2a and chr2b highlighting the potential important role 382
of both SDs and ILS in understanding chromosomal evolution, speciation, and sequence 383
turnover. 384
385
Methods
386
Data resources and comparative analysis. The great ape and macaque T2T genome assemblies 387
are publicly available (Data Availability). Syntenic relationships among primate 388
chr2/chr2a/chr2b were assessed using minimap255 (v2.24) with the following parameters: ‘-c -x 389
asm20 --secondary=no --eqx -Y -K 8G -s 1000’ and visualized using 390
SafFire (https://mrvollger.github.io/SafFire). We identified centromeres, transposons, and 391
subtelomeric satellites based on RepeatMasker (v4.1.4) annotation 392
(https://www.repeatmasker.org/). 393
394
Fusion site characterization and population analysis. To precisely define the fusion site, we 395
aligned NHP chr2a and chr2b to human chr2 using minimap2 (v2.24) with the following 396
parameters ‘-x asm20 -r500,20000 -s 2000 -p 0.01 -N 1000 --cs’. The alignments were visualized 397
using the minimiro (https://github.com/mrvollger/minimiro). We used VCFtools56 (v0.1.16) to 398
calculate nucleotide diversity (pi) and Tajima’s D for population genetic analyses with the 399
following parameters ‘--window-pi 20000 --window-pi-step 10000’ and ‘--TajimaD 20000’. 400
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
19
401
SD annotations and phylogenetic analysis. SD annotations of the fusion site in the T2T-402
CHM13v2.0 were based on the UCSC SEDEF-SD track (https://genome.ucsc.edu/). In this 403
study, we focused on the SDs with an identity of ≥98% and length ≥20 kbp. These SD sequences 404
were aligned to NHP T2T genomes to identify homologous regions using minimap2 (v2.24) with 405
the parameters: ‘-cx asm20 -p 0.5 --eqx’. Orthologous segments were identified based on 406
flanking region synteny. We aligned all primate homologous SDs with MAFFT57 (v7.515) and 407
used trimAl58 (v1.4) to remove the noise sequences with the parameter: ‘-automated1’. The 408
phylogenetic trees were constructed with IQ-TREE59 (v2.1.4), and then we used BEAST60 409
(v2.6.6) with the HKY model incorporating gamma site, calibrated yule, and relaxed log-normal 410
clock models to infer the split time. Node ages were estimated using log-normal priors. We 411
conducted three independent runs for each tree and the results were consistent. Effective sample 412
sizes exceeded 200 for all parameters in all runs. 413
414
For ILS analysis, we first truncated the fusion site flanking regions into 500 bp windows. Then, 415
we utilized transanno (v0.4.5) (https://github.com/informationsea/transanno) to align these 500 416
bp segments to NHP genomes and generated multiple alignments across primate species. We 417
then utilized IQ-TREE59 (v1.6.12) with HKY model and ete3 python package to analyze 418
phylogeny trees, as described previously44. 419
420
Subtelomeric repetitive region characterization, methylation analysis, and TE annotation. 421
The pairwise identity heatmaps of each subtelomeric region were generated by ModDotPlot61 422
(https://github.com/marbl/ModDotPlot). We used SEDEF62 (v1.1r35) to annotate SDs. To 423
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
20
analyze 5mC methylation levels in pCht regions and SD spacers, we utilized BEDTools63 424
(v2.30.0) to calculate mean methylation levels within 1 kbp windows, using a 500 bp step size. 425
Synteny between subtelomeric repetitive regions was defined using minimap2 (v2.24) with the 426
following parameters ‘-cx asm20 --secondary=no -A1 -B2 -O2,12 -s 1000 -Y -K 8G --eqx’. The 427
syntenic relationship was visualized using R package circlize64 (v0.4.16). R package dendextend 428
(v1.18.1)65 and TreeDist (v2.9.1)66 were utilized to plot the co-phylogeny and to calculate the 429
generalized Robinson–Foulds distance. 430
431
Centromere analysis. To analyze the structure of each chromosome, we first ran RepeatMasker 432
(v4.1.4) on all chromosomes and identified the α-satellite-enriched regions as centromeres67. 433
Then, we utilized HumAS-HMMER (https://github.com/fedorrik/HumAS-434
HMMER_for_AnVIL) to classify the suprachromosomal families (SFs) of α-satellites in human, 435
Pan, gorilla, and orangutan. For macaque, the OWM-SF annotation tool was used to annotate the 436
SFs37. The HOR arrays were characterized using the tool StV (https://github.com/fedorrik/stv). 437
However, for orangutan chr2a and chr2b, the tool encountered difficulties in correctly identifying 438
the HORs, likely due to the complex structure of these acrocentric chromosomes. Thus, we 439
selected the largest continuous regions containing the same SFs as their HORs. We then 440
estimated the frequency of 5mC and CpG methylation within the HOR regions. BEDTools 441
(v2.30.0) was used to count the frequency of methylation within 5 kbp windows and we 442
determined the regions with the minimum frequency and below the lower quartile among the 443
whole HOR as CDRs. Additionally, we used the previous published tool68 444
(https://github.com/altemose/chm13_hsat) to identity the HSatII/HSatIII arrays in human and 445
Pan peri/centromeric regions (α-satellite-enriched regions and 5 Mbp on the p-arm and q-arm). 446
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
21
447
We further characterized the human centromere degenerate site through synteny comparisons 448
between Pan and human, incorporating RepeatMasker annotations of ‘ALR_Alpha.’ Breakpoints 449
between chimpanzees and humans were identified through alignment and analysis of specific SF 450
organizations. ModDotPlot (https://github.com/marbl/ModDotPlot) was used to generate 451
heatmaps for chimpanzee chr2b with a window size of 5,000 bp and for the human centromere 452
degenerate site with a window size of 200 bp. 453
454
To assess the diversity of centromere degeneration across human populations, we first confirmed 455
the presence of this site in all human genomes using minimap2 (v2.24) with the parameters ‘-cx 456
asm20 --secondary=no -s 2500’. We then extracted the targeted regions from each assembly 457
from HPRC data and ran RepeatMasker (v4.1.4) on these regions. We defined three satellite 458
arrays—cenD_1, cenD_2, and cenD_3—with subtypes determined by length variations. 459
460
To confirm the fusion occurred in archaic humans, we randomly selected two individuals from 461
the five structural haplotypes (total 10 individuals) and five NHP genomes to run mrsFAST69 462
(v3.4.2) for identifying singly unique nucleotide k-mers (SUNKs) in modern humans. 463
Subsequently, we checked the counts of the SUNKs in archaic human genomes. 464
465
To compare the methylation frequencies among human chr2, α-satellite arrays at the degenerate 466
centromeric site, and NHP centromeres, we chunk the HORs (excluding CDRs; for macaques, 467
we used the whole α-satellite region, excluding CDRs) into 17.1 kbp windows by BEDTools 468
(v2.30.0) with parameters: ‘-w 171000 -s 8550’ and calculate the frequencies within windows. 469
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
22
Visualizations were generated using ggplot2. 470
471
Fusion site KO experiments and RNA-seq analysis. The design of optimal sgRNA pairs to 472
target sites was performed using the online CRISPR design tool (http://crispor.gi.ucsc.edu/). The 473
complementary oligonucleotide pairs of L-sgRNA #1/#2 and R-sgRNA were annealed at 95℃ 474
for 5 min, with ramp-down to 25℃ to generate the double-stranded DNA (dsDNA) fragment, 475
before ligation into BbsI-digested PX459 (plasmid #48139, Addgene). 476
477
HEK-293T cell lines were maintained in complete culture medium: DMEM (Dulbecco’s 478
modified Eagle’s medium) supplemented with 10% FBS (fetal bovine serum), 1% Penicillin-479
Streptomycin Solution (100×stock, Gibco). Transfection of CRISPR plasmid DNA into HEK-480
293T cells was performed using LipofectamineTM 3000 Reagent (Invitrogen). Briefly, 0.5–481
5×105/cm2 HEK293T cells in 6 well plates were transfected with 2.5 μg of CRISPR plasmid 482
DNA. Then, 24 hours after transfection, selection was conducted by adding puromycin (1 μg/ml) 483
to the media for the next 48 hr. Surviving cells were cultured for 4-7 days without selection 484
before harvesting. 485
486
Genomic DNA (gDNA) was extracted from transfected HEK-293T from each subline (1×105) 487
with the TIANamp genomic DNA Kit (DP304-02; Tiangen) according to the manufacturer’s 488
instructions. Genome deletions caused by the sgRNA pairs were detected by PCR-amplification 489
of gDNA using a primer pair flanking the deletion. The primers used for PCR screening are 490
listed in Supplementary Table 8. Genomic PCR products were detected by agarose gel 491
electrophoresis and Sanger sequencing (Shanghai, TsingKe). 492
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
23
493
Total RNA of the fusion site depletion and wild-type cell lines were reverse-transcribed using 494
HiScript II Q RT SuperMix for qPCR kit (Vazyme), respectively. Gene expression was 495
normalized to GADPH, and fold changes were calculated as described in the Figure 5. Error bars 496
in the RNA analysis represent the standard deviation of the average fold changes based on at 497
least two cell lines and/or two experimental replicates as indicated in the Figure 5. The primer 498
sequences are listed in Supplementary Table 12. 499
500
We utilized fastp70 (v0.23.2) to perform adaptor trimming of raw sequencing reads. The trimmed 501
reads were aligned to the reference genome (T2T-CHM13v2.0) using hisat271 (v2.2.1). 502
Differential expression analysis was conducted using the R package DESeq272 and edgeR73 and 503
we identified genes as significantly differentially expressed with following criteria: false 504
discovery rate (FDR) 2. The genes considered significantly 505
differentially expressed met two criteria: they were identified by both tools and demonstrated 506
significance in two independent sgRNA design experiments, forming the final set. Visualizations 507
were generated using ggplot2. We performed Gene Ontology (GO) enrichment analysis using 508
DAVID GO74 (https://david.ncifcrf.gov/). 509
510
Acknowledgments 511
We thank Tonia Brown for editing this manuscript. We thank the HPRC and Primate T2T 512
Consortium for providing the long-read human and great ape genome assemblies. The 513
computations in this study were run on the Siyuan-1 supported by the Center for High 514
Performance Computing at Shanghai Jiao Tong University. E.E.E. is an investigator of the 515
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
24
Howard Hughes Medical Institute. 516
517
This article is subject to HHMI’s Open Access to Publications policy. HHMI lab heads have 518
previously granted a nonexclusive CC BY 4.0 license to the public and a sublicensable license to 519
HHMI in their research articles. Pursuant to those licenses, the author-accepted manuscript of 520
this article can be made freely available under a CC BY 4.0 license immediately upon 521
publication. 522
523
Author contributions 524
Y.M., E.E.E., and Q.S. conceived the project; X.J., Z.Y., X.Y., K.M., S.Z., J.C., J.H., L.F., J.Z., 525
and Y.M. contributed to the syntenic comparison, fusion site characterization, ILS, and 526
subtelomeric repetitive region analyses; L.Z., X.J., Y.L., Y.N., X.B., and Q.S. contributed to the 527
fusion site depletion analysis; D.Y. and E.E.E. generated the genome assemblies of great apes; 528
X.J., G.Z., and Y.M. analyzed the centromeres. Y.M. and E.E.E. wrote the draft manuscript with 529
contributions from other authors. All authors read and approved the manuscript. 530
531
Conflict of Interest 532
E.E.E. is a scientific advisory board (SAB) member of Variant Bio, Inc. The other authors 533
declare no competing interests. 534
535
Funding 536
This work was supported, in part, by grants from the National Natural Science Foundation of 537
China (32370658), Natural Science Foundation of Chongqing, China (CSTB2024NSCQ-538
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
25
JQX0004), and Shanghai Jiao Tong University 2030 Initiative (WH510363001-7) to Y.M. This 539
work was supported, in part, by grants from the National Key Research and Development 540
Program of China (2022YFF0710901), the National Natural Science Foundation of China Grant 541
(82021001), Biological Resources Program of Chinese Academy of Sciences (KFJ-BRP-005), 542
National Science and Technology Innovation 2030 Major Program (2021ZD0200900) to Q.S. 543
This work is partially sponsored by Shanghai Rising-Star Program (24YF2721800) to K.M. This 544
work was supported, in part, by US National Institutes of Health (NIH) grant HG002385 to 545
E.E.E. 546
547
Data Availability 548
The T2T primate genomes used in this study are available from GenBank via accessions: 549
GCA_009914755.4, GCA_028858775.2, GCA_028885625.2, GCA_028885655.2, 550
GCA_029281585.2, GCA_029289425.2 and GCA_037993035.1. The T2T primate genome 551
assemblies are also available on GitHub (https://github.com/marbl/Primates and 552
https://github.com/zhang-shilong/T2T-MFA8). The Neanderthal and Denisovan genomes used 553
are available from https://www.eva.mpg.de/genetics/genome-projects. The Iso-Seq data are 554
available on ENCODE (tissues: ENCFF492BYP, ENCFF306ZPP, ENCFF318SKH, CHM13-555
T2T2 hTERT Iso-Seq: SRR12519035, SRR12519036). 556
557
References
558
1. Ferguson-Smith, M. A. & Trifonov, V. Mammalian karyotype evolution. Nat Rev Genet 8, 559
950–962 (2007). 560
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
26
2. Bush, G. L., Case, S. M., Wilson, A. C. & Patton, J. L. Rapid speciation and chromosomal 561
evolution in mammals. Proc. Natl. Acad. Sci. U.S.A. 74, 3942–3946 (1977). 562
3. Leibowitz, M. L., Zhang, C.-Z. & Pellman, D. Chromothripsis: A New Mechanism for Rapid 563
Karyotype Evolution. Annu. Rev. Genet. 49, 183–211 (2015). 564
4. Rieseberg, L. H. Chromosomal rearrangements and speciation. Trends in Ecology & Evolution 565
16, 351–358 (2001). 566
5. Watkins, T. B. K. et al. Pervasive chromosomal instability and karyotype order in tumour 567
evolution. Nature 587, 126–132 (2020). 568
6. Damas, J. et al. Evolution of the ancestral mammalian karyotype and syntenic regions. Proc. 569
Natl. Acad. Sci. U.S.A. 119, e2209139119 (2022). 570
7. Painter, T. S. & Stone, W. CHROMOSOME FUSION AND SPECIATION IN 571
DROSOPHILAE. Genetics 20, 327–341 (1935). 572
8. Schubert, I. Chromosome evolution. Current Opinion in Plant Biology 10, 109–115 (2007). 573
9. Kirkpatrick, M. & Barton, N. Chromosome Inversions, Local Adaptation and Speciation. 574
Genetics 173, 419–434 (2006). 575
10. Baker, R. J. & Bickham, J. W. Speciation by monobrachial centric fusions. Proc. Natl. Acad. 576
Sci. U.S.A. 83, 8245–8248 (1986). 577
11. Baird, D. M. Telomeres and genomic evolution. Phil. Trans. R. Soc. B 373, 20160437 578
(2018). 579
12. Zhao, N. et al. Critically short telomeres derepress retrotransposons to promote genome 580
instability in embryonic stem cells. Cell Discov 9, 45 (2023). 581
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
27
13. Lee, C., Sasi, R. & Lin, C. C. Interstitial localization of telomeric DNA sequences in the 582
Indian muntjac chromosomes: further evidence for tandem chromosome fusions in the 583
karyotypic evolution of the Asian muntjacs. Cytogenet Genome Res 63, 156–159 (1993). 584
14. Chikashige, Y. Meiotic nuclear reorganization: switching the position of centromeres and 585
telomeres in the fission yeast Schizosaccharomyces pombe. The EMBO Journal 16, 193–202 586
(1997). 587
15. Bailey, J. A. & Eichler, E. E. Primate segmental duplications: crucibles of evolution, 588
diversity and disease. Nat Rev Genet 7, 552–564 (2006). 589
16. Mudd, A. B., Bredeson, J. V., Baum, R., Hockemeyer, D. & Rokhsar, D. S. Analysis of 590
muntjac deer genome and chromatin architecture reveals rapid karyotype evolution. Commun 591
Biol 3, 480 (2020). 592
17. Augustijnen, H. et al. A macroevolutionary role for chromosomal fusion and fission in 593
Erebia butterflies. Sci. Adv. 10, eadl0989 (2024). 594
18. Wright, C. J., Stevens, L., Mackintosh, A., Lawniczak, M. & Blaxter, M. Comparative 595
genomics reveals the dynamics of chromosome evolution in Lepidoptera. Nat Ecol Evol 8, 777–596
790 (2024). 597
19. Luo, J., Sun, X., Cormack, B. P. & Boeke, J. D. Karyotype engineering by chromosome 598
fusion leads to reproductive isolation in yeast. Nature 560, 392–396 (2018). 599
20. Dymond, J. S. et al. Synthetic chromosome arms function in yeast and generate phenotypic 600
diversity by design. Nature 477, 471–476 (2011). 601
21. Wang, L.-B. et al. A sustainable mouse karyotype created by programmed chromosome 602
fusion. Science 377, 967–975 (2022). 603
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
28
22. Zhang, X. M. et al. Creation of artificial karyotypes in mice reveals robustness of genome 604
organization. Cell Res 32, 1026–1029 (2022). 605
23. Wienberg, J., Jauch, A., Stanyon, R. & Cremer, T. Molecular cytotaxonomy of primates by 606
chromosomal in situ suppression hybridization. Genomics 8, 347–350 (1990). 607
24. Dutrillaux, B., Rethoré, M. O. & Lejeune, J. [Comparison of the karyotype of the orangutan 608
(Pongo pygmaeus) to those of man, chimpanzee, and gorilla]. Ann Genet 18, 153–161 (1975). 609
25. IJdo, J. W., Baldini, A., Ward, D. C., Reeders, S. T. & Wells, R. A. Origin of human 610
chromosome 2: an ancestral telomere-telomere fusion. Proc. Natl. Acad. Sci. U.S.A. 88, 9051–611
9055 (1991). 612
26. Ventura, M. et al. The evolution of African great ape subtelomeric heterochromatin and the 613
fusion of human chromosome 2. Genome Res. 22, 1036–1049 (2012). 614
27. Poszewiecka, B., Gogolewski, K., Karolak, J. A., Stankiewicz, P. & Gambin, A. 615
PhaseDancer: a novel targeted assembler of segmental duplications unravels the complexity of 616
the human chromosome 2 fusion going from 48 to 46 chromosomes in hominin evolution. 617
Genome Biol 24, 205 (2023). 618
28. Fan, Y., Linardopoulou, E., Friedman, C., Williams, E. & Trask, B. J. Genomic Structure and 619
Evolution of the Ancestral Chromosome Fusion Site in 2q13–2q14.1 and Paralogous Regions on 620
Other Human Chromosomes. Genome Res. 12, 1651–1662 (2002). 621
29. Fan, Y., Newman, T., Linardopoulou, E. & Trask, B. J. Gene Content and Function of the 622
Ancestral Chromosome Fusion Site in Human Chromosome 2q13–2q14.1 and Paralogous 623
Regions. Genome Res. 12, 1663–1672 (2002). 624
30. Martin, C. L. et al. The Evolutionary Origin of Human Subtelomeric Homologies—or Where 625
the Ends Begin. The American Journal of Human Genetics 70, 972–984 (2002). 626
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
29
31. Royle, N. J., Baird, D. M. & Jeffreys, A. J. A subterminal satellite located adjacent to 627
telomeres in chimpanzees is absent from the human genome. Nat Genet 6, 52–56 (1994). 628
32. Miga, K. H. Chromosome-Specific Centromere Sequences Provide an Estimate of the 629
Ancestral Chromosome 2 Fusion Event in Hominin Genomes. JHERED 108, 45–52 (2017). 630
33. Wang, H. et al. SVA Elements: A Hominid-specific Retroposon Family. Journal of 631
Molecular Biology 354, 994–1007 (2005). 632
34. Poszewiecka, B., Gogolewski, K., Stankiewicz, P. & Gambin, A. Revised time estimation of 633
the ancestral human chromosome 2 fusion. BMC Genomics 23, 616 (2022). 634
35. Yoo, D. et al. Complete sequencing of ape genomes. Preprint at 635
https://doi.org/10.1101/2024.07.31.605654 (2024). 636
36. Makova, K. D. et al. The complete sequence and comparative analysis of ape sex 637
chromosomes. Nature 630, 401–411 (2024). 638
37. Zhang, S. et al. Comparative genomics of macaques and integrated insights into genetic 639
variation and population history. Preprint at https://doi.org/10.1101/2024.04.07.588379 (2024). 640
38. Koga, A., Hirai, Y., Hara, T. & Hirai, H. Repetitive sequences originating from the 641
centromere constitute large-scale heterochromatin in the telomere region in the siamang, a small 642
ape. Heredity 109, 180–187 (2012). 643
39. Dutrillaux, B. Chromosomal evolution in Primates: Tentative phylogeny from Microcebus 644
murinus (Prosimian) to man. Hum Genet 48, 251–314 (1979). 645
40. Yunis, J. J. & Prakash, O. The Origin of Man: A Chromosomal Pictorial Legacy. Science 646
215, 1525–1530 (1982). 647
41. Liao, W.-W. et al. A draft human pangenome reference. Nature 617, 312–324 (2023). 648
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
30
42. Linardopoulou, E. V. et al. Human Subtelomeric WASH Genes Encode a New Subclass of 649
the WASP Family. PLoS Genet 3, e237 (2007). 650
43. Vollger, M. R. et al. Increased mutation and gene conversion within human segmental 651
duplications. Nature 617, 325–334 (2023). 652
44. Mao, Y. et al. A high-quality bonobo genome refines the analysis of hominid evolution. 653
Nature 594, 77–81 (2021). 654
45. Chiatante, G., Giannuzzi, G., Calabrese, F. M., Eichler, E. E. & Ventura, M. Centromere 655
Destiny in Dicentric Chromosomes: New Insights from the Evolution of Human Chromosome 2 656
Ancestral Centromeric Region. Molecular Biology and Evolution 34, 1669–1681 (2017). 657
46. Lonsdale, J. et al. The Genotype-Tissue Expression (GTEx) project. Nat Genet 45, 580–585 658
(2013). 659
47. Logsdon, G. A. et al. The structure, function and evolution of a complete human 660
chromosome 8. Nature 593, 101–107 (2021). 661
48. Luo, Y. et al. New developments on the Encyclopedia of DNA Elements (ENCODE) data 662
portal. Nucleic Acids Research 48, D882–D889 (2020). 663
49. Lacruz, R. S. et al. The evolutionary history of the human face. Nat Ecol Evol 3, 726–736 664
(2019). 665
50. Alemseged, Z. Reappraising the palaeobiology of Australopithecus. Nature 617, 45–54 666
(2023). 667
51. Broeils, L. A., Ruiz-Orera, J., Snel, B., Hubner, N. & Van Heesch, S. Evolution and 668
implications of de novo genes in humans. Nat Ecol Evol 7, 804–815 (2023). 669
52. Mayr, E. Michael J. D. White., Modes of Speciation. Systematic Biology 27, 478–482 (1978). 670
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
31
53. Trask B. Fluorescence. In: Birren B, Green ED, Hieter P, Klapholz S, Myers RM, Riethman 671
H, Roskams J, editors. Situ Hybridization. In Genome analysis: A laboratory manual. Cold 672
Spring Harbor, NY: Cold Spring Harbor Laboratory Press; 1999. pp. 303–413. 673
54. Hirai, H. et al. Structural variations of subterminal satellite blocks and their source 674
mechanisms as inferred from the meiotic configurations of chimpanzee chromosome termini. 675
Chromosome Res 27, 321–332 (2019). 676
55. Li, H. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics 34, 3094–677
3100 (2018). 678
56. Danecek, P. et al. The variant call format and VCFtools. Bioinformatics 27, 2156–2158 679
(2011). 680
57. Katoh, K. & Standley, D. M. MAFFT Multiple Sequence Alignment Software Version 7: 681
Improvements in Performance and Usability. Molecular Biology and Evolution 30, 772–780 682
(2013). 683
58. Capella-Gutiérrez, S., Silla-Martínez, J. M. & Gabaldón, T. trimAl: a tool for automated 684
alignment trimming in large-scale phylogenetic analyses. Bioinformatics 25, 1972–1973 (2009). 685
59. Minh, B. Q. et al. IQ-TREE 2: New Models and Efficient Methods for Phylogenetic 686
Inference in the Genomic Era. Molecular Biology and Evolution 37, 1530–1534 (2020). 687
60. Bouckaert, R. et al. BEAST 2.5: An advanced software platform for Bayesian evolutionary 688
analysis. PLoS Comput Biol 15, e1006650 (2019). 689
61. Sweeten, A. P., Schatz, M. C. & Phillippy, A. M. ModDotPlot—rapid and interactive 690
visualization of tandem repeats. Bioinformatics 40, btae493 (2024). 691
62. Numanagić, I. et al. Fast characterization of segmental duplications in genome assemblies. 692
Bioinformatics 34, i706–i714 (2018). 693
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
32
63. Quinlan, A. R. & Hall, I. M. BEDTools: a flexible suite of utilities for comparing genomic 694
features. Bioinformatics 26, 841–842 (2010). 695
64. Gu, Z., Gu, L., Eils, R., Schlesner, M. & Brors, B. circlize implements and enhances circular 696
visualization in R. Bioinformatics 30, 2811–2812 (2014). 697
65. Galili, T. dendextend: an R package for visualizing, adjusting and comparing trees of 698
hierarchical clustering. Bioinformatics 31, 3718–3720 (2015). 699
66. Smith, M. R. Information theoretic generalized Robinson–Foulds metrics for comparing 700
phylogenetic trees. Bioinformatics 36, 5007–5013 (2020). 701
67. Altemose, N. et al. Complete genomic and epigenetic maps of human centromeres. Science 702
376, eabl4178 (2022). 703
68. Altemose, N., Miga, K. H., Maggioni, M. & Willard, H. F. Genomic Characterization of 704
Large Heterochromatic Gaps in the Human Genome Assembly. PLoS Comput Biol 10, e1003628 705
(2014). 706
69. Hach, F. et al. mrsFAST: a cache-oblivious algorithm for short-read mapping. Nat Methods 707
7, 576–577 (2010). 708
70. Chen, S. Ultrafast one‐pass FASTQ data preprocessing, quality control, and deduplication 709
using fastp. iMeta 2, e107 (2023). 710
71. Kim, D., Paggi, J. M., Park, C., Bennett, C. & Salzberg, S. L. Graph-based genome 711
alignment and genotyping with HISAT2 and HISAT-genotype. Nat Biotechnol 37, 907–915 712
(2019). 713
72. Love, M. I., Huber, W. & Anders, S. Moderated estimation of fold change and dispersion for 714
RNA-seq data with DESeq2. Genome Biol 15, 550 (2014). 715
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
33
73. Chen, Y., Lun, A. T. L. & Smyth, G. K. From reads to genes to pathways: differential 716
expression analysis of RNA-Seq experiments using Rsubread and the edgeR quasi-likelihood 717
pipeline. F1000Res 5, 1438 (2016). 718
74. Sherman, B. T. et al. DAVID: a web server for functional enrichment analysis and functional 719
annotation of gene lists (2021 update). Nucleic Acids Research 50, W216–W221 (2022). 720
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
34
721
Figure 1. The comparative sequencing analysis of primate chromosome 2. (a) The syntenic 722
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
35
comparison of chromosome 2 highlights the extent of primate chromosome 2 evolutionary 723
rearrangements. Syntenic regions conserved in order (blue) are contrasted with evolutionary 724
inversions (orange) and non-syntenic regions (breaks) corresponding to alpha satellite (dark 725
brown), pCht subterminal satellite (green), and other satellite DNA (pink). B. orangutan and S. 726
orangutan represent Bornean orangutan and Sumatran orangutan, respectively. (b) High-727
resolution analysis of the human fusion site (chr2:113,940,058-114,049,496, colored in amber) 728
shows the non-syntenic breakpoint region in the context of annotated human protein-coding 729
genes and segmental duplications (SDs). Pan SD spacers (purple blocks) in chimpanzee pCht 730
show the homology of the partial region of human fusion site. 731
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
36
732
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
37
Figure 2. The genomic structure and evolutionary history of SDs at the human fusion site. 733
(a) Human chromosome 2 fusion site and SD organization. A human genomic segment 734
(chr2:113,650,520-114,212,116) with gene annotations shows the region where chromosome 2a 735
and 2b fused (chr2:113,940,058-114,049,496, amber). The region consists of a large (~455 kbp) 736
duplication block made up of seven SDs with variable copy numbers in each ape genome. The 737
three SDs at the fusion site are designated as SD_fusion_A, SD_fusion_B, and SD_fusion_C. 738
The table indicates the copy number of the full-length homologous segments for each primate 739
genome (with brackets indicating the SD_fusion_A copy number of the derived sequence in the 740
subterminal heterochromatic caps of chimpanzees and bonobo; see Supplementary Table 2 for 741
the detailed alignments at the fusion site). (b-g) Phylogenetic trees based on a (b) 53 kbp 742
(SD_fusion_A), (d) 68 kbp (SD_fusion_B), and (f) 23 kbp (SD_fusion_C) multiple sequence 743
alignment show that these regions have been subject to incomplete lineage sorting (ILS). The 744
chromosome schematic depicts the genomic locations of SD_fusion_A (c), SD_fusion_B (e), and 745
SD_fusion_C (g) in each primate. (h) Proportion of human–gorilla and Pan–gorilla tree topology 746
in a 500 bp window of 5 Mbp flanking region near the fusion site. The red dotted line represents 747
the mean value of the whole chr2. (i) Fold change of the mean proportion of discordant 748
topologies in the flanking region compared with the whole genome average. (Note, chromosome 749
names in the great apes refer to the human homologous chromosome also known as the 750
phylogenetic group designation.) 751
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
38
752
Figure 3. Rapid turnover of subtelomeric repetitive regions in primates and evolutionary 753
connection between SD spacers in Pan lineage and SD_fusion_A at the fusion site. (a) The 754
identity heatmap of the subterminal heterochromatic cap for the p-arm of chimpanzee chr2b. SD 755
spacer elements (purple) are annotated below and are flanked by large tracts of pCht satellite 756
(green) (zoomed-in panel shows the structure at higher resolution). (b) Circos plot shows 757
sequence identity of chr2a and chr2b among the apes highlighting the rapid turnover of satellite 758
DNA. The three layers (outer to inner) represent the subtelomeric region, tandem repeat 759
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
39
satellites, and transposon element annotations. (c) Dot plot shows the synteny between a bonobo 760
SD spacer within the heterochromatic cap vs. human SD_fusion_A segment. The sequence 761
identity was calculated at a 100 bp resolution, with the dot color representing sequence identity 762
ranging from 80% (purple) to 100% (red). (d) Phylogenetic topology comparison shows nearly 763
consistent ILS evolutionary pattern of SD_fusion_A (excluding spacer sequence) and SD spacer. 764
Each node is supported by a bootstrap value of 100/100. The red lines indicate discordant tree 765
topologies within each major clade between the two trees.766
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
40
767
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
41
Figure 4. Comparative analysis of active and inactive centromeric regions of human 768
chromosome 2. (a) Genomic structure and suprachromosomal family (SF) annotation of 769
centromeres in human chr2, degenerate site, and NHP chr2a/chr2b. Different SFs are shown by 770
various colors, with centromere dip regions (CDRs) marked by arrows. (b) Comparison of the 771
centromere in chimpanzee chr2b with the human degenerate centromeric region. The middle 772
panel shows the SF annotations for the chimpanzee centromere and human centromere 773
degenerate region. Potential α-satellite deletion breakpoints (b1 and b2) are indicated with dotted 774
lines. The heatmaps of each region are shown in upper and lower, respectively. (c) The panel 775
illustrates five distinct structural haplotypes of the degenerate centromere site in humans. (d) The 776
panel shows the lengths of three α-satellite arrays in humans. (e) The methylation degree of HOR 777
across NHPs and of three α-satellite arrays at the degenerate centromere site. The red triangle 778
represents the methylation frequency of CDR in each chromosome (Met. Freq. strands for 779
methylation frequency). 780
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
42
781
Figure 5. Fusion site knockout and gene expression alteration. (a) Schematic representation 782
of the depletion experiments performed in HEK293T cell lines. (b) Two distinct sgRNA pairs 783
(L-sg1 and R-sg; L-sg2 and R-sg) were designed for the depletion experiments (left panel). Red 784
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
43
sequences indicate sgRNA, while blue sequences indicate PAM sites. (c) The PCR analysis 785
confirms heterozygous depletion in cell lines. WT represents the wild-type cell lines. (d) and (e) 786
Differentially expressed genes (DEGs) are shown in volcano plots (left panel for the L-sg1 and 787
R-sg pair, right panel for the L-sg2 and R-sg pair). Genes with log2 fold change (log2FC) 788
values >2 and FDR <0.05 are highlighted in red (upregulated) and blue (downregulated). The 789
overlap of top 10 DEGs in each group are emphasized in deeper red/blue, with gene names 790
indicated. (f) Venn diagrams illustrate the overlap of upregulated and downregulated genes 791
between both depletion groups. (g) The RT-PCR analyses confirmed 5 of 10 top-rank DEGs. WT 792
represents the relative gene expression levels normalized to GAPDH in wild-type cell lines. 793
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
44
794
Figure 6. Models for chromosome 2 evolution. Two different models are depicted for the 795
origin of the human chromosome 2 fusion. (a) In the last common ancestor of humans, bonobos, 796
chimpanzees, and gorillas, diverse ancestral genomic structures emerged at the ends of chr2a and 797
chr2b. In the human ancestral lineage, chr2a and chr2b consisted of complex SD blocks that 798
emerged as a result of ILS, and then became juxtaposed by a telomere-to-telomere fusion. (b) 799
Rudimentary pCht sequences were present in the ancestral lineage of humans and chimpanzees 800
in association with SD_fusion_A. Nonallelic homologous recombination (NAHR) occurred 801
between these in the human lineage eliminating pCht in humans but retaining SD_fusion_A with 802
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
45
other SDs subsequently duplicating to the location. In chimpanzees, an independent expansion of 803
the SD spacers and pCht sequences occurred leading to the formation of the subterminal 804
heterochromatic caps. 805
.CC-BY-NC-ND 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 17, 2024. ; https://doi.org/10.1101/2024.12.12.628057doi: bioRxiv preprint
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.