Integrated Genotyping Strategies Uncovering Detailed Haplotype Structures and Characterization of DMD duplications

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract Duchenne and Becker muscular dystrophies (DMD/BMDs) are X-linked genetic disorders caused by mutations in the dystrophin gene (DMD), characterized by progressive muscle weakness and degeneration. While DMD duplications account for approximately 10% of cases, their clinical impact varies significantly, ranging from severe phenotypes to asymptomatic presentations, posing significant challenges in determining their pathogenicity. This study investigates the molecular complexity of DMD duplications and their implications for disease progression. Through analyzing 3,842 patients using multiple sequencing platforms, we identified 39 cases with DMD duplications and characterized four distinct duplication patterns. These structure variations not only influence pathogenicity interpretation but also reflect specific mechanisms of genomic instability. Our findings reveal that conventional genetic testing methods frequently fail to accurately resolve duplication structures, limiting their predictive value for clinical outcomes. By integrating whole genome sequencing and optical genome mapping, we achieved precise haplotype resolution, substantially enhancing genotype–phenotype correlations. These results underscore the critical importance of adopting multi-platform genomic strategies to improve diagnostic accuracy, refine pathogenicity assessment, and optimize personalized genetic counseling for patients with DMD duplications.
Full text 125,467 characters · extracted from preprint-html · click to expand
Integrated Genotyping Strategies Uncovering Detailed Haplotype Structures and Characterization of DMD duplications | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Integrated Genotyping Strategies Uncovering Detailed Haplotype Structures and Characterization of DMD duplications Tingting Yu, Ruen Yao, Qihua Fu, Jin Sun, Jie Tang, lu Wei, Juan Geng, and 4 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6126136/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Duchenne and Becker muscular dystrophies (DMD/BMDs) are X-linked genetic disorders caused by mutations in the dystrophin gene ( DMD ), characterized by progressive muscle weakness and degeneration. While DMD duplications account for approximately 10% of cases, their clinical impact varies significantly, ranging from severe phenotypes to asymptomatic presentations, posing significant challenges in determining their pathogenicity. This study investigates the molecular complexity of DMD duplications and their implications for disease progression. Through analyzing 3,842 patients using multiple sequencing platforms, we identified 39 cases with DMD duplications and characterized four distinct duplication patterns. These structure variations not only influence pathogenicity interpretation but also reflect specific mechanisms of genomic instability. Our findings reveal that conventional genetic testing methods frequently fail to accurately resolve duplication structures, limiting their predictive value for clinical outcomes. By integrating whole genome sequencing and optical genome mapping, we achieved precise haplotype resolution, substantially enhancing genotype–phenotype correlations. These results underscore the critical importance of adopting multi-platform genomic strategies to improve diagnostic accuracy, refine pathogenicity assessment, and optimize personalized genetic counseling for patients with DMD duplications. Biological sciences/Genetics/Clinical genetics/Genetic testing Health sciences/Diseases/Neurological disorders/Dystonia Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Introduction Duchenne and Becker muscular dystrophies (DMD/BMDs) are the most prevalent X-linked inherited muscle disorders, with an estimated incidence ranging from 1 in 5,000 to 1 in 6,000 in newborn males 1 . DMD/BMDs are primarily caused by mutations in dystrophin gene ( DMD ), resulting in either a loss of function or insufficient dosage of dystrophin protein. DMD (OMIM: 300377) is the largest known gene in humans, spanning approximately 2.4 Mb on chromosome Xp21.1 and consisting of 79 exons. Structural variants (SVs) are the most common type of mutation in DMD/BMD, accounting for approximately 80% of cases, with duplications contributing to about 10% of all mutations 2 . While numerous studies have characterized the duplication spectrum at exon-level resolution 3 – 5 , with the Leiden Open Variation Database (LOVD) documenting nearly 1,000 distinct duplication genotypes as of June 2024 5 , the precise genomic architecture, duplication patterns, and underlying molecular mechanisms remain to be fully elucidated. The clinical manifestations of DMD/BMD demonstrate remarkable phenotypic heterogeneity. DMD patients typically present with calf muscle pseudohypertrophy and difficulties in walking or stair climbing between the ages of two and five years. The disease progressively worsens, with most patients losing ambulation by 12 years of age and experiencing life-threatening cardiomyopathy leading to mortality by the age of 20 6 . In contrast, BMD patients exhibit milder symptoms, characterized by later onset and slower progression, with ambulation typically preserved until the fifth or sixth decade 7 . Genotype-phenotype analyses have shown that mutation effects in the open reading frame (ORF) of DMD can be predictive of the disease severity, with frameshift mutations generally leading to DMD and in-frame mutations typically resulting in BMD 7 – 9 . While genotype-phenotype predictions based on deletion mutations demonstrate high reliability, duplications frequently present with inconsistent clinical outcomes, complicating prognostic assessments 10 – 12 . Notably, emerging evidence has identified benign DMD duplication variants that do not manifest clinical phenotypes, further highlighting the complexity of these genomic alterations 13 – 15 . The reasons behind the phenotypic variability associated with DMD duplications remain poorly understood. Current diagnostic approaches for DMD/BMD primarily employ multiplex ligation-dependent probe amplification (MLPA), Sanger sequencing and high throughput sequencing such as whole exome sequencing (WES), collectively enabling the identification of pathogenic variants in approximately 98% of cases 16 . However, these conventional methods exhibit significant limitations in detecting complex SVs, especially in resolving precise haplotype configurations for duplications. Whole genome sequencing (WGS) has gradually been integrated into clinical practice, offering enhanced capacity for detecting SVs 17 . Moreover, long-read sequencing (LRS) technologies have recently been applied to DMD mutation detection, enabling the analysis of large fragments extending to tens of kilobases 15 , 18 . Optical genome mapping (OGM), another long-read technology, can detect SVs on a megabase scale and has been used effectively for previously undiagnosed cases 13 , 19 . The advent of these advanced genomic technologies provides unprecedented opportunities for high-resolution molecular characterization of DMD duplications, facilitating more precise genotype-phenotype predictions in clinical diagnostics. In this study, we reanalyzed the DMD -MLPA and WES data from a cohort of 3,842 patients and identified 39 individuals with DMD duplications. To elucidate the molecular architecture of these duplications, we performed WGS on all 39 patients and subsequently applied OGM to six WGS-unresolved cases. Through comparative analysis of detection methodologies, we established an optimized framework for accurate identification and characterization of DMD duplications. Furthermore, we conducted a genotype-phenotype analysis based on the duplication haplotypes, enabling refined clinical interpretation. This investigation provided valuable insights into the diversity and complexity of DMD duplications, significantly advancing our capabilities in pathogenicity assessment, genetic counseling, and clinical screening for DMD duplications. Results Cohort Design To explore the genomic complexity of DMD duplications and their genotype-phenotype relationships, we conducted a retrospective analysis of patients who presented at Shanghai Children’s Medical Center between August 2016 and January 2024. The study cohort comprised two distinct patient populations: (1) 578 patients clinically suspected of DMD/BMD who had undergone MLPA, Sanger sequencing, or WES, and (2) 3,264 patients with suspected alternative genetic disorders who had also been evaluated through WES analysis. Among them, 39 patients with confirmed DMD duplications (26 detected by MLPA and 13 by WES) were recalled for WGS and OGM analyses to resolve their haplotype structures (Fig. 1 a). The final study cohort consisted of 31 male DMD/BMD patients, 1 female BMD patient, 1 asymptomatic male carrier, and 6 asymptomatic female carriers (Fig. 1 b, Supplementary Table 1, Supplementary Table 2). Pedigree analysis revealed that these patients originated from 38 independent families (Supplementary Fig. S1 ). Notably, WGS34 and WGS35 were derived from the same family lineage; therefore, only WGS34 was included in subsequent analyses to maintain statistical independence. High Heterogeneity of DMD Duplications Revealed by WGS A total of 38 distinct duplications were identified in our cohort (Supplementary Table 3). The duplications exhibited remarkable heterogeneity among individuals and shared little structural similarity with known pathological duplications documented in DECIPHER database (Supplementary Fig. 2a, b). The size distribution of these duplications ranged from 10,725 bp to 884,245 bp, with a median of 260,488 bp, where 50% of duplications were between 100 kb and 500 kb. Notable, 40% of duplications (n = 16) were split into two or more segments with more than two breakpoints, and 24% (n = 9) involved extragenic regions. Additionally, some duplications overlapped with other OMIM genes, such as IL1RAPL1 (OMIM: 300206), CFAP47 (OMIM: 301057), and XK (OMIM: 314850) (Supplementary Fig. 2b). Specifically, one female sample (WGS33) displayed a duplication that included a segment of chromosome Y, suggesting an interchromosomal translocation event. The exon-level resolution analysis revealed 34 unique patterns among the 38 identified DMD duplications. Only four recurrent patterns were observed: exon 2 (E2), exons 3–4 (E3-4), exons 3–5 (E3-5), and exons 46–49 (E46-49), each appearing in two samples, while the remaining patterns appeared only once in our cohort (Fig. 2 a). Six duplication genotypes were newly identified in our cohort, including exons 5–45 (E5-45), exons 8–9/28–43 (E8-9, 28–43), exons 24–44 (E24-44), exons 31–33 (E31-33), exons 45–52/55 (E45-52, 55), and exons 48–51 (E48-51) (Fig. 2 b). Cumulative duplication frequencies differed between our cohort and public database. While LOVD reported higher duplication frequencies in exons 2–9, our cohort demonstrated a predominance of duplications in exons 45–51 (Fig. 2 c). This discrepancy may be attributed to differences in cohort size and population characteristics. Using WGS, we classified duplications into four patterns: Tandem duplication (Tandem-Dup, 58%, n = 22), Duplication-Normal-Duplication (Dup-Nor-Dup, 16%, n = 6), Duplication-Inversion-Duplication (Dup-Inv-Dup, 16%, n = 6), and Intricate duplication (Intricate-Dup, 10%, n = 4) (Fig. 2 d, Supplementary Fig. 3, Supplementary Data). The Tandem-Dup exhibited a single haplotype that could be fully resolved by WGS. In contrast, the remaining three patterns, collectively referred to complex duplications, had multiple potential haplotypes, thereby complicating haplotype determination and variant interpretation. Breakpoint Features of DMD Duplication We identified 118 breakpoints across 59 breakpoint pairs, with 79% located within DMD introns. The remaining breakpoints were distributed upstream (12%), downstream (6%), interchromosomal (2%) or within DMD exons (2%) (Fig. 3 a). Notably, breakpoints in intron 2 were mainly linked to Tandem-Dup, while those in intron 44 were associated with complex duplication patterns (Supplementary Fig. 4a). Breakpoints in intron 49 frequently led to duplications extending beyond DMD (Supplementary Fig. 4b). Previous studies have shown that repeat elements 20 , GC-unbalanced regions 21 and meiotic DNA strand breaks (DSBs) 22 contribute to SV formation. To investigate their roles in DMD duplications, we defined genomic regions containing these features as genome unstable regions. Our analysis revealed that 73% of breakpoint pairs had both ends located within genome unstable regions, while the remaining 27% had one end in such regions (Fig. 3 b, Supplementary Fig. 4c). Moreover, the number of breakpoints within introns showed a positive correlation with intron length and the density of genome unstable regions (Supplementary Fig. 4d, e). Specifically, 50% of breakpoints were located within repeat elements, with SINE (Alu) and LINE (L1) being the most prevalent (Fig. 3 c). Additionally, the 5' ends of breakpoint pairs were significant enriched in GC-poor regions compared to genomic background (Fig. 3 d). Furthermore, breakpoints were highly associated with DSB hotspots, particularly those linked to complex duplication patterns (Fig. 3 e). Among the 59 breakpoint pairs, 30 showed microhomology sequences ranging from 1 to 7 bp, 18 contained no-templated insertions at the junctions, and 11 exhibited direct joins (Supplementary Table 4). Most breakpoints had a homology ratio of 60–70% in the flanking sequences, with seven pairs showing exceptionally high homology (> 80%) (Fig. 3 f). These high-homology breakpoint pairs were predominantly located in repeat elements, specifically SINE: Alu (n = 6) and LINE: L1 (n = 1) (Supplementary Fig. 4f). These pairs were more likely to induce complex duplication patterns, such as Dup-Nor-Dup (n = 2), Dup-Inv-Dup (n = 3), and Intricate-Dup (n = 2) (Fig. 3 g). Genomic annotation of these high-homology breakpoint pairs showed one end located in DMD introns (introns 17, 49, 51, 57, or 61), while the other end was located outside of DMD . This resulted in duplications extending beyond DMD and led to distinct clinical outcomes. OGM Is Essential for Precise Localization and Haplotype Determination of DMD Duplications Using WGS, we identified 13 samples with multiple possible haplotypes, each potentially leading to different effects on the DMD products. We applied OGM to six WGS-unresolved samples, demonstrating its utility in defining SV structures and assessing pathogenicity. In case WGS40, a duplication was incidentally detected during prenatal testing of a 14-week pregnant woman (Gravida 2 Para 0). Her first pregnancy was terminated due to prenatal diagnoses of Tetralogy of Fallot, diaphragmatic hernia, and spinal abnormalities. Chromosomal microarray analysis (CMA) revealed no pathogenic or likely pathogenic copy number variations (CNVs) in the fetus, but it incidentally identified a maternally inherited duplication of exons 50–55 in DMD , classified as a variant of uncertain significance (VUS). This duplication had been reported in one BMD and 11 DMD cases in LOVD, suggesting a potential risk for DMD in offspring. Genotype-phenotype co-segregation analysis found no male family members carrying the duplication (Fig. 4 a). WGS further characterized the duplication as a Dup-Nor-Dup pattern, consisting of two segments: chrX:31567450–31831842 (exons 50–55 of DMD ) and chrX:33890225–34038113 (intergenic region) (Fig. 4 b). This pattern suggested two possible haplotypes: duplication insertion into intron 55 of DMD or insertion into chrX:34038113 (extragenic). OGM clarified that the duplication was inserted into chrX:34038113, outside DMD (Fig. 4 c). As this insertion did not disrupt DMD transcription, the duplication was classified as benign, posing no risk for DMD/BMD in the offspring. In case WGS25, a 4.7-year-old boy presented with delayed motor milestones, muscle weakness, and elevated creatine kinase (CK), CK-MB, Aspartate transferase (AST), Alanine transaminase (ALT), and Lactate Dehydrogenase (LDH). Family history indicated that his maternal uncle lost ambulation at an early age (Fig. 5 a). WES identified a duplication of exons 45–51 in DMD , predicted to maintain the reading frame and typically associated with BMD. However, the patient's clinical presentation and family history were consistent with DMD. Notably, LOVD records showed phenotypic variability for this duplication, with six cases classified as DMD, two as BMD, and one reported as an asymptomatic male. WGS further characterized the duplication as a Dup-Nor-Dup pattern comprising two segments: chrX:31749465–32066307 (exons 45–51 of DMD ) and chrX:36315168–36597300 (exons 57–65 of CFAP47 ) (Fig. 5 b). Two potential haplotypes were identified: one with the duplication inserted into intron 51 of DMD and the other with insertion into intron 56 of CFAP47 . OGM revealed that the duplication inserted into intron 51 of DMD , aligning with the patient's DMD phenotype and family history (Fig. 5 c). In case WGS31, a 6-year-old male exhibited muscle weakness, hypertrophic calves, Gowers' sign, and attention deficits. MLPA identified a maternally inherited duplication of exons 46–49 in DMD , though the family history was unavailable (Fig. 6 a). WGS revealed a complex duplication with six breakpoints and three segments: chrX:28281895–28382937 (intergenic), chrX:28486473–28668996 (exon 1 of IL1RAPL1 ), and chrX:31832901–31954458 (exons 46–49 of DMD ) (Fig. 6 b). OGM resolved the structure and confirmed that these three duplicated segments were combined and inserted into intron 49 of DMD , disrupting DMD products and aligning with the observed phenotype (Fig. 6 c). OGM May Miss Detection of Certain Exons Due to Low Resolution Although OGM significantly aids in identifying duplication structures, its ability to replace conventional detection technologies remains uncertain. When mapping OGM labeling sites to DMD , it was observed that some labels skipped two exons (Fig. 7 a). Specifically, among the 559 labeling sites identified within DMD introns, 17 introns lacked labeling sites, suggesting that SVs with breakpoints located in these introns may be missed by OGM (Fig. 7 b). Additionally, the resolution of OGM on DMD is approximately 550 bp, which means smaller SVs may be undetected (Fig. 7 c). In our cohort, discrepancies between OGM and WGS results were observed in 33% of the detections (n = 2). In case WGS10, WGS identified two duplication segments: chrX:32275575–32447885 (exons 28–43 of DMD ) and chrX:32676655–32765682 (exons 8–9 of DMD ) (Fig. 7 d). However, the absence of OGM labeling sites in intron 26 resulted in uncertain boundary determination for segment A (Fig. 7 e). Similarly, in case WGS23, WGS indicated a Dup-Inv-Dup pattern (Fig. 7 f), while OGM detected a tandem duplication. This discrepancy is likely due to the small size of segment C (411 bp), which falls below OGM’s resolution limit (Fig. 7 g). Comparative Analysis of Multiple Detection Technologies for DMD duplications MLPA and WES are commonly used for detecting variants in DMD . We compared their exon-level performance with WGS to assess diagnostic consistency (Supplementary Table 5). Among the 26 samples analyzed using MLPA, 88% (n = 23) were concordant with WGS results, while 12% (n = 3) exhibited discrepancies: one sample missed a non-contiguous exon (WGS26), another involved a female sample with both a duplication and a deletion on separate alleles (MLPA generated a combined result) (WGS22), and the third missed an exon containing a breakpoint (WGS19). For the 16 samples analyzed by both WES and WGS, 69% (n = 11) were concordant. Discrepancies were observed in 31% of cases: 19% (n = 3) had one-exon discrepancies at duplication boundaries (WGS09, WGS18, and WGS27), 6% (n = 1) missed a non-contiguous exon (WGS26), and 6% (n = 1) failed to detect an exon with a breakpoint (WGS39). Among the six samples analyzed using both WGS and OGM, 67% (n = 4) showed consistent results, and discrepancies were noted in 33% (n = 2) due to OGM’s limited resolution (WGS10, WGS23). Genotype-Phenotype Analysis for DMD Duplications Accurate understanding genotype-phenotype relationships is crucial for precise clinical diagnosis. By leveraging the detailed mapping of duplication haplotypes provided by WGS and OGM, we were able to more accurately assess their effects on DMD ORF and conduct robust phenotype-correlation analyses (Table 1 , Supplementary Table 6). For duplications with multiple potential haplotypes, the most likely structure was inferred based on the patient’s phenotype or family history. The predicted effects were classified into four categories: none (no impact on DMD ), in-frame (duplication preserves the ORF), frameshift (duplication disrupts the ORF), and extragenic-insert (duplication contains large extragenic segments). Table 1 Summary of phenotype and predicted ORF’s effects in diagnosed male patients Diagnosed Male Patients none in-frame frameshift extragenic-insert DMD (n=25) 0 3 19 5 BMD (n=4) 0 4 0 0 Asymptomatic (n=1) 1 0 0 0 Concordance Rate 1 100% 57% 100% 100% ¹The concordance rate indicates the rate that the observed phenotype aligned with the predicted ORF effect. Specifically, “None” is expected to be asymptomatic, “In-frame” is expected to correspond to BMD, and both “Frameshift” and “Extragenic-insert” are expected to result in DMD. As expected, patients with duplications predicted as none were asymptomatic, and all male patients with frameshift duplications (n = 18) were diagnosed to DMD. However, among patients with in-frame duplications (n = 7), 57% (n = 4) presented with BMD, while 43% (n = 3) were diagnosed with DMD. The in-frame duplications associated with DMD involved two specific genotypes and exhibited a Tandem-Dup pattern: exons 3–4 duplication (n = 2, WGS04, WGS05) and exon 34 duplication (n = 1, WGS20). Furthermore, all duplications classified as extragenic-inserts (n = 5) were linked to DMD. Notably, extragenic-insert duplications included genotypes such as exons 45–51 duplication (WGS25) and exons 45–57 duplication (WGS29), which were initially predicted as in-frame with canonical ORF prediction tools 5 , 9 , 23 , resulted in DMD due to the involvement of large intergenic regions. This refined classification improved the accuracy of genotype-phenotype predictions. Discussion In this study, we uncovered the remarkable diversity and complexity of DMD duplications, extending beyond previous exon-level analyses 2 . Within our cohort, four exon-level duplication genotypes were recurrent, but their breakpoints varied significantly. Additionally, six novel genotypes were identified in our cohort, which had not been reported previously. We classified these duplications into four distinct patterns: Tandem-Dup (58%), Dup-Nor-Dup (16%), Dup-Inv-Dup (16%), and Intricate-Dup (11%). Notably, only approximately 60% followed the canonical tandem arrangement, while the remaining 40% displayed more complex configurations. Other potential patterns, such as inverted-insert tandem duplications 18 , may have been missed due to the limited cohort size. These findings highlight the limitations of widely used databases, such as LOVD, which primarily catalog exon-level genotypes of DMD duplications, but lack critical details information about duplication patterns, breakpoints, and orientations. Moreover, the canonical reading-frame checker assumes a tandem arrangement of duplicated exons, which may lead to inaccuracies in ORF effect predictions and clinical interpretations 5 , 9 , 23 . To improve genetic counseling and phenotype predictions, it is essential to expand these databases with comprehensive details on duplication characteristics and their associated phenotypes. Our breakpoint analysis sheds light on the potential mechanisms underlying duplication formation. Consistent with prior studies, breakpoints were predominantly located in genomic regions prone to instability 24 . We hypothesize that replication-based fork stalling and template switching (FoSTeS) is key drivers of duplication formation (Supplementary Fig. 5). Previous studies have implicated highly homologous repeats in both tandem duplications via non-allelic homologous recombination (NAHR) and more complex structures, such as duplication-triplication/inverted-duplication via multiple rounds of FoSTeS 25 . In our cohort, highly homologous repeats were associated with Dup-Nor-Dup (n = 2), Dup-Inv-Dup (n = 3), and Intricate-Dup (n = 2), suggesting the duplication formation may involve both NAHR and FoSTeS (Supplementary Fig. 6, 7). These repeat elements, particularly in DMD introns 17, 49, 51, 57, and 61, contributed to duplications that extended beyond DMD , further complicating pathogenicity interpretation. Additional data are further required to validate and refine potential hotspot regions for complex duplications beyond DMD . Case studies from our cohort underscore the importance of resolving duplication haplotypes to accurately determine clinical outcomes. WGS25 and WGS40 both involved Dup-Nor-Dup duplications, but their clinical outcomes differed. In WGS40, the duplication was located in an intergenic region upstream of DMD and was classified as benign, whereas the duplication of WGS25 disrupted DMD , leading to DMD phenotypes. Another instructive case, WGS31, involved a duplication spanning DMD and IL1RAPL1 , a gene associated with neurodevelopmental disorders. A previously reported complex duplication, combined with IL1RAPL1 (exons 3–11) and DMD (exon 1), was linked to autism spectrum disorder, rather than DMD/BMD 26 , emphasizing the significance of the disrupted gene in determining disease manifestation. Collectively, these cases highlight the critical need for comprehensive structural characterization and the adoption of a detailed nomenclature, including duplicated segments, duplication orientation, inserted locations, and all affected pathological genes, to improve clinical interpretations. For instance, the detailed haplotype of WGS31 could be described as “ DMD :intron49ins[(chrX:28281004–28383006)INV, IL1RAPL1 :exon1, DMD :exon46-49]dup”, referring to a joint duplication of three segments: inversion of chrX:28281004–28383006, exon 1 of IL1RAPL1 , and exons 46–49 of DMD , inserted into intron 49 of DMD . All duplications in our cohort were nomenclature accordingly in Supplementary Table 2.Our findings reinforced the phenotype prediction based on reading-frame rule 9 , with all frameshift duplications resulted in DMD and 57% of in-frame duplications were linked to BMD. However, 43% of in-frame duplications (n = 3) caused DMD, specifically involving exons 3–4 duplications (n = 2) and exon 34 duplications (n = 1). These genotypes are also more frequently associated with DMD in LOVD (11 DMD vs. 7 BMD for exons 3–4 duplication; 2 DMD vs. 0 BMD for exon 34 duplication),likely due to the disruption of critical functional domains, such as the N-terminal F-actin-binding domain (exons 3–4) and the R11 region of the rod domain (exon 34), impairing sarcolemma interaction and leading to severe phenotypes 27 . We also introduced a new ORF effect category, extragenic-insert, for duplications containing large extragenic segments. These duplications extend beyond conventional ORF-based predictions and are often associated with severe phenotypes. This refinement improved the accuracy of genotype-phenotype prediction. For example, the exons 45–51 duplication (WGS25) and exons 45–57 duplication (WGS29), previously predicted as BMDs under conventional ORF-based rules, were accurately reclassified as DMD due to the involvement of large extragenic regions. Further research is needed to elucidate the pathological mechanisms underlying such duplications and their effects on DMD transcription and splicing. Our study underscored the limitations of detection platforms. WGS failed to resolve complex haplotype structures in 34% of cases (13/38) due to its short-read constraints. OGM effectively resolved haplotypes in 67% of ambiguous cases (4/6), but it missed smaller segments in 33% of cases (2/6) because of its resolution thresholds. Besides, OGM demonstrated higher sample quality requirements and cost burdens compared to WGS. Combining WGS and OGM results could provide more comprehensive SV detection. Taken together, we propose a sequential testing strategy: employ WGS to analyze duplications identified as VUS by MLPA, WES, CMA or noninvasive prenatal testing (NIPT) 28 , followed by OGM for unresolved cases to identify haplotypes and predict clinical outcomes (Fig. 8 ). In the long term, the development of novel integrative approaches or more advanced technologies will be essential to fully capture the complete spectrum of DMD variants and improve diagnostic accuracy. Our study does have several limitations. The relatively modest sample size could introduce potential sampling bias, and single-exon duplications might be underestimated in the WES reanalysis. Additionally, we lacked DMD transcript data due to the constraints of our sample types. Nevertheless, our cohort is still one of the largest datasets on DMD duplications and integrates multi-technology analyses, thereby providing robust insights into the DMD duplication spectrum and advancing pathogenicity interpretations. Methods Patient information A total of 3,842 patients underwent genetic testing (MLPA & sanger sequencing, WES) at the Department of Medical Genetics and Molecular Diagnostic Laboratory, Shanghai Children’s Medical Center (SCMC) from August 2016 to January 2024. The clinical information and genetic testing data were collected and reanalyzed. Among them, 39 patients confirmed with DMD duplications were recruited for WGS and OGM. Ethical approval was obtained (SCMCIRB-K2020060-1), and written informed consent was provided by all participants. MLPA CNVs of DMD were detected via a modified MLPA as previously described 29 . Briefly, CNVplex kit (Genesky Diagnostics, Suzhou, China) were used following the manufacturer’s instructions. Data were visualized and analyzed with the corresponding CNVplex analysis software. OGM Ultra-high molecular weight DNA was extracted, labeled, and processed from peripheral blood using the Prep SP Blood and Cell DNA Isolation Kit (Bionano Genomics, San Diego, CA), Qubit Fluorometers (Bionano Genomics, San Diego, CA), and Prep DLS Labeling kits (Bionano Genomics Inc., San Diego, CA, USA). The labeled DNA was loaded for linearization and imaging using Saphyr platform (Bionano Genomics, San Diego, CA). The genome of proband was assembled using the Bionano Solve V3.7 software. SVs were called against the human reference genome (GRCh38/hg38) and visualized using Bionano Access software (V1.7). The site of OGM labeling is generated with digest_genome.py of HiC-Pro 30 (v3.1.0) with the parament of “-r ^CTTAAG”. WES and WES-based CNV detection WES analysis followed previously described methods 31 . CNVs were detected using CNVkit 32 (v0.9.9), with samples from the same batch used as controls. The results were filtered based on the following criteria: SVLEN > 50000 or 5, and FOLD_CHANGE_LOG > 0.5 or < -0.8. WGS and WGS-based CNV detection Genomic DNA was extracted and sequencing libraries constructed using the MGIEasy PCR-Free DNA Library Prep Kit (MGI Tech, Shenzhen, China). Adapters were trimmed with fastp 33 (v0.23.2), and reads were aligned to hg38 using BWA 34 (v0.2.10). SVs and CNVs were called using DELLY 35 (version 0.7.6), CNVkit 32 (v0.9.9), and Control-FREEC 36 (v11.6), then merged with SURVIVOR 37 (v1.0.7). Genotypes were called using svtyper 38 (v0.7.1). Breakpoint feature analysis Breakpoints were extracted from soft-clipped reads and validated by UCSC BLAT ( https://genome.ucsc.edu/cgi-bin/hgBlat ) 39 . GC content was calculated from 100 bp regions flanking the breakpoints, and 10,000 random controls were generated using bedtools shuffle (v2.31.0) as genomic background. The regions with GC content less than 35% were defined as GC-poor, while regions with GC content greater than 60% as GC-rich. Meiotic DSB hotspots were obtained from GSE59836 22 and random regions were generated using bedtools shuffle (v2.31.0) with 10,000 iterations to serve as randomly selected background controls. Breakpoints located within 500 bp of DSB hotspots were classified as being within DSB hotspots. Repeat elements were obtained from UCSC Table Browser. Homology ratios of flanking sequences at breakpoint pairs were calculated using Biopython (v1.85), employing a 500 bp sliding window across the 1000 bp regions surrounding each breakpoint pair, and the maximum value was selected. Effects on the DMD ORF The effects of duplications on the DMD ORF (NM_004006.3) were predicted based the reconstructed haplotype structures. For duplications with multiple potential haplotypes, predictions were made for each haplotype independently. The effect was determined by considering the cumulative length of inserted exons in the transcriptional direction. In cases where exons were invertedly inserted into DMD introns, the insertion length was assigned as "0" due to the loss of splicing donor and acceptor motifs 40 . If the cumulative insertion length was a multiple of three, the duplication was predicted to maintain the reading frame (in-frame); otherwise, it was predicted to disrupt the reading frame (frameshift). Duplications involving large segments of intergenic regions or other genes were classified under the category extragenic-insert. Public data download in LOVD DMD duplication variants and associated phenotypes were downloaded from LOVD 5 (updated 2024/06/24) ( https://databases.lovd.nl/shared/genes/DMD ). Abbreviations DMD: dystrophin DMD: Duchenne muscular dystrophy BMD: Becker muscular dystrophy MD: muscular dystrophy WES: whole exome sequencing MLPA: multiplex ligation-dependent probe amplification WGS: whole genome sequencing OGM: Optical genome mapping Tandem-Dup: tandem duplication Dup-Nor-Dup: duplication-normal-duplication Dup-Inv-Dup: duplication-inversion-duplication Intricate-Dup: intricate duplication SV: Structural variants LOVD: Leiden Open Variation Database ORF: open reading frame LRS: long-read sequencing DSB: meiotic DNA strand break NAHR: non-allelic homologous recombination FoSTeS: fork stalling and template switching CK: creatine kinase AST: Aspartate transferase ALT: Alanine transaminase LDH: Lactate Dehydrogenase CNV: copy number variation VUS: variants of uncertain significance ABD: N-terminal F-actin-binding domain R: Rod domain CR: cysteine-rich domain CT: C-terminal domain NIPT: noninvasive prenatal testing CMA: chromosomal microarray analysis Declarations Availability of data and materials The raw data have been deposited into the Genome Sequence Archive (https://ngdc.cncb.ac.cn/gsa-human/) with the accession number HRA009618. Consent for publication Not applicable Acknowledgements This study was supported by the Key Discipline Group of Pudong New Area Health Commission (PWZxq2022-07), the Shanghai Key Laboratory of Clinical Molecular Diagnostics for Pediatrics (20dz2260900), Shanghai “Rising Stars of Medical Talents” Youth Development Program, Postdoctoral Fellowship Program of CPSF (GZC20241067), and National Clinical Key Specialty Construction Project (10000015Z155080000004). We sincerely appreciate the support of BeCreative Lab in Beijing for optical genome mapping analysis and Biosan Diagnostics in Hangzhou for whole genome sequencing analysis. We are also deeply grateful to the patients and their families for their participation in this study. Author contributions Tingting Yu: Conceptualization, Resources, Writing - Review & Editing, Supervision, Funding acquisition; Ruen Yao, Qihua Fu: Conceptualization, Resources, Writing - Review & Editing, Funding acquisition; Jin Sun, Jie Tang: Methodology, Formal analysis, Visualization, Writing - Original Draft; Lu Wei: Data Curation, Writing - Review & Editing; Geng Juan, Xiao Rui: Methodology, Writing - Review & Editing; Jian Wang, Niu Li, Shuyuan Li: Resources, Writing - Review & Editing. All authors read and approved the final manuscript. Competing interests The authors declare no competing interests. References Mah JK et al (2014) A systematic review and meta-analysis on the epidemiology of Duchenne and Becker muscular dystrophy. Neuromuscul Disord 24:482–491 Zhao L et al (2024) Comprehensive analysis of 2097 patients with dystrophinopathy based on a database from 2011 to 2021. Orphanet J Rare Dis 19:311 Kong X et al (2019) Genetic analysis of 1051 Chinese families with Duchenne/Becker Muscular Dystrophy. BMC Med Genet 20:139 Tuffery-Giraud S et al (2009) Genotype-phenotype analysis in 2,405 patients with a dystrophinopathy using the UMD-DMD database: a model of nationwide knowledgebase. Hum Mutat 30:934–945 Aartsma-Rus A, Van Deutekom JC, Fokkema IF, Van Ommen GJ, Den Dunnen JT (2006) Entries in the Leiden Duchenne muscular dystrophy mutation database: an overview of mutation types and paradoxical cases that confirm the reading-frame rule. Muscle Nerve 34:135–144 Mercuri E, Muntoni F (2013) Muscular dystrophies. Lancet 381:845–860 Gorgoglione D et al (2024) Natural history of Becker muscular dystrophy: DMD gene mutations predict clinical severity. Brain Monaco AP, Bertelson CJ, Liechti-Gallati S, Moser H, Kunkel LM (1988) An explanation for the phenotypic differences between patients bearing partial deletions of the DMD locus. Genomics 2, 90 – 5 Magri F et al (2011) Genotype and phenotype characterization in a large dystrophinopathic cohort with extended follow-up. J Neurol 258:1610–1623 Lim KRQ, Nguyen Q, Yokota T (2020) Genotype-Phenotype Correlations in Duchenne and Becker Muscular Dystrophy Patients from the Canadian Neuromuscular Disease Registry. J Pers Med 10 Yang J et al (2013) MLPA-based genotype-phenotype analysis in 1053 Chinese patients with DMD/BMD. BMC Med Genet 14:29 Vengalil S et al (2017) Duchenne Muscular Dystrophy and Becker Muscular Dystrophy Confirmed by Multiplex Ligation-Dependent Probe Amplification: Genotype-Phenotype Correlation in a Large Cohort. J Clin Neurol 13:91–97 Zhang Y et al (2024) Prenatal risk assessment of Xp21.1 duplication involving the DMD gene by optical genome mapping. Life Sci Alliance 7 Shen J et al (2024) Comprehensive analysis of genomic complexity in the 5' end coding region of the DMD gene in patients of exons 1–2 duplications based on long-read sequencing. BMC Genomics 25:292 Bai Y et al (2022) Long-Read Sequencing Revealed Extragenic and Intragenic Duplications of Exons 56–61 in DMD in an Asymptomatic Male and a DMD Patient. Front Genet 13:878806 Beggs AH, Koenig M, Boyce FM, Kunkel LM (1990) Detection of 98% of DMD/BMD gene deletions by polymerase chain reaction. Hum Genet 86:45–48 Pluta N et al (2023) Whole-Genome Sequencing Identified New Structural Variations in the DMD Gene That Cause Duchenne Muscular Dystrophy in Two Girls. Int J Mol Sci 24 Ling C et al (2023) Uncovering the true features of dystrophin gene rearrangement and improving the molecular diagnosis of Duchenne and Becker muscular dystrophies. iScience 26:108365 Xiao B et al (2024) Combining optical genome mapping and RNA-seq for structural variants detection and interpretation in unsolved neurodevelopmental disorders. Genome Med 16:113 Liao X et al (2023) Repetitive DNA sequence detection and its role in the human genome. Commun Biology 6:954 Kiktev DA, Sheng Z, Lobachev KS, Petes TD (2018) GC content elevates mutation and recombination rates in the yeast Saccharomyces cerevisiae. Proceedings of the National Academy of Sciences 115, E7109-E7118 Pratto F et al (2014) DNA recombination. Recombination initiation maps of individual human genomes. Science 346:1256442 Zhou J, Xin J, Niu Y, Wu S (2017) DMDtoolkit: a tool for visualizing the mutated dystrophin protein and predicting the clinical severity in DMD. BMC Bioinformatics 18:87 Carvalho CM, Lupski JR (2016) Mechanisms underlying structural variant formation in genomic disorders. Nat Rev Genet 17:224–238 Grochowski CM et al (2024) Inverted triplications formed by iterative template switches generate structural variant diversity at genomic disorder loci. Cell Genom 4:100590 Eisfeldt J et al (2024) Resolving complex duplication variants in autism spectrum disorder using long-read genome sequencing. Genome Res 34:1763–1773 Bhosle RC, Michele DE, Campbell KP, Li Z, Robson RM (2006) Interactions of intermediate filament protein synemin with dystrophin and utrophin. Biochem Biophys Res Commun 346:768–777 Brison N et al (2019) Maternal copy-number variations in the DMD gene as secondary findings in noninvasive prenatal screening. Genet Med 21:2774–2780 Zhang X et al (2015) A modified multiplex ligation-dependent probe amplification method for the detection of 22q11.2 copy number variations in patients with congenital heart disease. BMC Genomics 16:364 Servant N et al (2015) HiC-Pro: an optimized and flexible pipeline for Hi-C data processing. Genome Biol 16:259 Yao R, Yu T, Qing Y, Wang J, Shen Y (2019) Evaluation of copy number variant detection from panel-based next-generation sequencing data. Mol Genet Genomic Med 7:e00513 Talevich E, Shain AH, Botton T, Bastian BC, CNVkit (2016) Genome-Wide Copy Number Detection and Visualization from Targeted DNA Sequencing. PLoS Comput Biol 12:e1004873 Chen S, Zhou Y, Chen Y, Gu J (2018) fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics 34:i884–i890 Li H, Durbin R (2009) Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics 25:1754–1760 Rausch T et al (2012) DELLY: structural variant discovery by integrated paired-end and split-read analysis. Bioinformatics 28:i333–i339 Boeva V et al (2012) Control-FREEC: a tool for assessing copy number and allelic content using next-generation sequencing data. Bioinformatics 28:423–425 Jeffares DC et al (2017) Transient structural variations have strong effects on quantitative traits and reproductive isolation in fission yeast. Nat Commun 8:14061 Chiang C et al (2015) SpeedSeq: ultra-fast personal genome analysis and interpretation. Nat Methods 12:966–968 Kent WJ (2002) BLAT–the BLAST-like alignment tool. Genome Res 12, 656 – 64 Xie Z et al (2020) Long-read whole-genome sequencing for the genetic diagnosis of dystrophinopathies. Ann Clin Transl Neurol 7:2041–2046 Additional Declarations There is NO Competing Interest. Supplementary Files SupplementaryInformation.docx Supplymentary Table 1, 3 and 5; Supplementary Fig. 1 to 7. SupplymentaryTable2.xlsx Supplymentary Table 2 SupplymentaryTable4.xlsx Supplymentary Table 4 SupplymentaryTable6.xlsx Supplymentary Table 6 Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6126136","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":425212721,"identity":"3e4938ae-0fe9-4e6b-9e5c-c3e6ad164158","order_by":0,"name":"Tingting Yu","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA4ElEQVRIie3PMQrCMBSA4VcKnVK7tijqER4IiotepeKq6BGcdNA66x0cdAmOKR1cKl07Cp2FQhcjHUzcrRkF80PeEPKRBECn+9GYWATAZKrAlAQFsXx1IkJ5EaoBTGdhNH+WjZ5zKHI+GYCzWvvAz1VkDNF+i6S/u1MvoGNw4+vRCOIvxN4gwfRKTYMyQHd6NI2lEomzQp2QhyDJBupKxIszjOxFR9xideVfiPxLGFSQ2mWUFaRsDjGJspzTQdNZBacbryBtJsb7Ga7/3iBysM8AoLWQsxTLqTyn0+l0/9wLnJFVsp2leHMAAAAASUVORK5CYII=","orcid":"","institution":"Shanghai Children's Medical Center, Shanghai Jiao Tong University School of Medicine","correspondingAuthor":true,"prefix":"","firstName":"Tingting","middleName":"","lastName":"Yu","suffix":""},{"id":425212722,"identity":"63ccb299-6a00-4994-bcec-ddada1291ce5","order_by":1,"name":"Ruen Yao","email":"","orcid":"","institution":"Shanghai Children's Medical Center, Shanghai Jiao Tong University School of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Ruen","middleName":"","lastName":"Yao","suffix":""},{"id":425212723,"identity":"fac270d8-4127-4e9d-ab84-4da2a846e202","order_by":2,"name":"Qihua Fu","email":"","orcid":"","institution":"Medical Genetics and Rare Diseases Center, Sichuan Academy of Medical Sciences \u0026 Sichuan Provincial People's Hospital, University of Electronic Science and Technology of China","correspondingAuthor":false,"prefix":"","firstName":"Qihua","middleName":"","lastName":"Fu","suffix":""},{"id":425212724,"identity":"82d8e39c-9ff1-4305-bd6b-46701fde999b","order_by":3,"name":"Jin Sun","email":"","orcid":"","institution":"Shanghai Children's Medical Center, School of Medicine, Shanghai Jiao Tong University","correspondingAuthor":false,"prefix":"","firstName":"Jin","middleName":"","lastName":"Sun","suffix":""},{"id":425212725,"identity":"9a18bf01-b06c-42f2-86c0-604355e0c3fb","order_by":4,"name":"Jie Tang","email":"","orcid":"","institution":"Shanghai Children's Medical Center, School of Medicine, Shanghai Jiao Tong University","correspondingAuthor":false,"prefix":"","firstName":"Jie","middleName":"","lastName":"Tang","suffix":""},{"id":425212726,"identity":"6a6aed45-65e9-4eff-8aa2-7bebfdccce3a","order_by":5,"name":"lu Wei","email":"","orcid":"","institution":"Shanghai Children's Medical Center, School of Medicine, Shanghai Jiao Tong University","correspondingAuthor":false,"prefix":"","firstName":"lu","middleName":"","lastName":"Wei","suffix":""},{"id":425212727,"identity":"a96c22ef-5580-489a-aa43-a5aefd8a5bb9","order_by":6,"name":"Juan Geng","email":"","orcid":"","institution":"the X Laboratory, and Center for Mendelian Genomics, Advanced Institute of Information Technology, Peking University","correspondingAuthor":false,"prefix":"","firstName":"Juan","middleName":"","lastName":"Geng","suffix":""},{"id":425212728,"identity":"b097d47d-3bca-425d-b9ef-aeb19c4eab2f","order_by":7,"name":"Rui Xiao","email":"","orcid":"","institution":"Zhejiang Biosan Biochemical Technologies Co. Ltd, Hangzhou, China","correspondingAuthor":false,"prefix":"","firstName":"Rui","middleName":"","lastName":"Xiao","suffix":""},{"id":425212729,"identity":"145b02e0-e579-4c2c-be12-c05d96b58868","order_by":8,"name":"Niu Li","email":"","orcid":"","institution":"Shanghai Children’s Medical Center","correspondingAuthor":false,"prefix":"","firstName":"Niu","middleName":"","lastName":"Li","suffix":""},{"id":425212730,"identity":"92a9e159-76dd-4a2b-b4ae-8fc0cc0608dc","order_by":9,"name":"Shuyuan Li","email":"","orcid":"","institution":"International Peace Maternity and Child Health Hospital, School of Medicine, Shanghai Jiao Tong University","correspondingAuthor":false,"prefix":"","firstName":"Shuyuan","middleName":"","lastName":"Li","suffix":""},{"id":425212731,"identity":"2ab92628-17c0-49ed-a1c5-d8614a5322e6","order_by":10,"name":"Jian Wang","email":"","orcid":"","institution":"Shanghai Jiaotong University","correspondingAuthor":false,"prefix":"","firstName":"Jian","middleName":"","lastName":"Wang","suffix":""}],"badges":[],"createdAt":"2025-02-28 07:15:15","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6126136/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6126136/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":78226815,"identity":"d8f443c0-8c7b-4b12-845c-16c9c4639a68","added_by":"auto","created_at":"2025-03-11 07:04:29","extension":"jpeg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":928737,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eStudy Overview\u003c/strong\u003e\u003cbr\u003e\na. Study Design.Patients with validated \u003cem\u003eDMD\u003c/em\u003e duplications detected by MLPA/WES were recruited for WGS and OGM analysis. Duplication patterns, comparisons between different technologies, and genotype-phenotype relationships were investigated in this study.\u003c/p\u003e\n\u003cp\u003eb. Cohort overview. Family history specifically refers to muscular dystrophy (MD) related phenotypes, such as calf hypertrophy, muscle weakness and atrophy, compromised or lost ambulation, and myocarditis.\u003c/p\u003e","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-6126136/v1/0586898ad8c9d6889869da1f.jpeg"},{"id":78226031,"identity":"10e2b87f-619a-44d8-8a4e-c27d5ae38412","added_by":"auto","created_at":"2025-03-11 06:56:29","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":287867,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eDuplication Patterns in Our Cohort\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003ea. Duplication spectrum on \u003cem\u003eDMD\u003c/em\u003e. The x-axis labels “0” and “80” denote genomic regions upstream and downstream of \u003cem\u003eDMD\u003c/em\u003e, respectively, whereas 1–79 indicate \u003cem\u003eDMD\u003c/em\u003eexons 1-79. The y-axis labels represent specific genotypes, and “*” indicates novel genotypes in this cohort. Points represent duplications, with a deeper color denoting recurring genotypes. ABD, R, CR, and CT refer to the four major \u003cem\u003eDMD\u003c/em\u003eprotein domains: the N-terminal F-actin-binding domain (ABD), Rod (R), cysteine-rich (CR), and C-terminal (CT).\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-6126136/v1/7a3a0df549f09b5bd2ccc6ee.png"},{"id":78226032,"identity":"fd620f0d-ab00-4d51-8889-d92c4431bf3c","added_by":"auto","created_at":"2025-03-11 06:56:29","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":275678,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eBreakpoint Features of Duplications\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003ea. Distribution of breakpoints across \u003cem\u003eDMD\u003c/em\u003e.\u003c/p\u003e\n\u003cp\u003eb. Ends of breakpoint pairs located in genome-unstable regions. 5′ and 3′ refer to the positional order of breakpoints on the forward strand.\u003c/p\u003e\n\u003cp\u003ec. Breakpoint distributions across repeat elements. Bar colors reflect the positional order of breakpoints on the forward strand.\u003c/p\u003e\n\u003cp\u003ed. GC content composition of 5′ and 3′ breakpoints vs. Background. GC-poor represents GC content less than 35%; GC-rich represents GC content greater than 60%; GC-moderate represents GC content at the range of 35%-60%. GC content is calculated for ±100 bp flanking breakpoints. Background were calculated using 10,000 randomly generated 200-bp intervals in the genome. Statistical comparisons were performed using the Fisher test.\u003c/p\u003e\n\u003cp\u003ee. Breakpoint counts in meiotic DSB hotspots vs. Background. Bar colors reflect the positional order of breakpoints on the forward strand. Genomic background regions are randomly generated by 10,000 permutations of DSB hotspots using bedtools shuffle. Statistical comparisons were performed using the Permutation test.\u003c/p\u003e\n\u003cp\u003ef. Homology ratios in the flanking regions of breakpoint pairs. Ratios are calculated using 500-bp window flanking breakpoint pairs.\u003c/p\u003e\n\u003cp\u003eg. Schematic diagram of Dup-Nor-Dup and Dup-Inv-Dup duplications involving high-homology repeats. In the Dup-Nor-Dup, the high homology repeats were in the same direction while in the Dup-Inv-Dup, the high homology repeats were in the reverse direction.\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-6126136/v1/d2219650cb7f8c122590c03a.png"},{"id":78228419,"identity":"3917cc76-8ae8-49fb-8bc8-877110167bc2","added_by":"auto","created_at":"2025-03-11 07:12:29","extension":"jpeg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":589727,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eBenign \u003c/strong\u003e\u003cem\u003e\u003cstrong\u003eDMD \u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003eDuplication of Dup-Nor-Dup Identified in Prenatal Testing\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003ea. Pedigree diagrams of WGS40. WGS40 was a female carrier (Gravida 2, Para 0) whose first pregnancy was terminated due to prenatal findings of Tetralogy of Fallot, diaphragmatic hernia, and spinal abnormalities.\u003c/p\u003e\n\u003cp\u003eb. IGV screenshots showing WGS results for WGS40.\u003c/p\u003e\n\u003cp\u003ec. Schematic representation of OGM results for WGS40. The haplotype sketch illustrated the resolved haplotype structure by integrating WGS and OGM results.\u003c/p\u003e","description":"","filename":"floatimage4.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-6126136/v1/0735401c88b8154441aee437.jpeg"},{"id":78228421,"identity":"2d6dca7b-6c37-4837-b60f-94dda964b43c","added_by":"auto","created_at":"2025-03-11 07:12:29","extension":"jpeg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":529140,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003ePathogenic \u003c/strong\u003e\u003cem\u003e\u003cstrong\u003eDMD\u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003eDuplication of Dup-Nor-Dup\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003ea. Pedigree diagram of WGS25. WGS25 was a 4.7-year-old male presenting with delayed motor milestones, muscle weakness, and elevated CK, CK-MB, AST, ALT, and LDH. His maternal uncle lost ambulation.\u003c/p\u003e\n\u003cp\u003eb. IGV screenshots showing WGS results for WGS25\u003c/p\u003e\n\u003cp\u003ec. Schematic representation of OGM results for WGS25. The haplotype sketch illustrated the resolved haplotype structure by integrating WGS and OGM results.\u003c/p\u003e","description":"","filename":"floatimage5.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-6126136/v1/030f444458a6b3a33c1d6c4c.jpeg"},{"id":78226822,"identity":"c5679ffa-24de-41d1-a784-f0334c3298e4","added_by":"auto","created_at":"2025-03-11 07:04:29","extension":"jpeg","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":753827,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eComplex \u003c/strong\u003e\u003cem\u003e\u003cstrong\u003eDMD \u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003eDuplication with Multiple Breakpoints\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003ea. Pedigree diagrams of WGS31, a 6.7-year-old male with muscle weakness, hypertrophic calves, Gowers' sign, and attention deficit.\u003c/p\u003e\n\u003cp\u003eb. IGV screenshots showing WGS results for WGS31.\u003c/p\u003e\n\u003cp\u003ec. Schematic representation of OGM results for WGS31. The haplotype sketch illustrated the resolved haplotype structure by integrating WGS and OGM results.\u003c/p\u003e","description":"","filename":"floatimage6.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-6126136/v1/1c56818b3bd253535d4f4fe4.jpeg"},{"id":78226045,"identity":"4559167d-c4eb-4dd9-a645-ee659bdf7a77","added_by":"auto","created_at":"2025-03-11 06:56:29","extension":"jpeg","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":909671,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eLimitations in OGM Detection Precision\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003ea. Schematic diagram of OGM labeling sites.\u003c/p\u003e\n\u003cp\u003eb. Cumulative count of OGM labeling sites across \u003cem\u003eDMD\u003c/em\u003e introns. Red text indicates the 17 introns without OGM labeling sites.\u003c/p\u003e\n\u003cp\u003ec. Distribution of distances between adjacent OGM labeling sites on\u003cem\u003e DMD\u003c/em\u003e, showing a resolution size of approximately 550 bp.\u003c/p\u003e\n\u003cp\u003ed, f. IGV screenshots showing WGS results for WGS10 (d) and WGS23 (f), respectively.\u003c/p\u003e\n\u003cp\u003ee, g. Schematic representation of OGM results for WGS10 (e) and WGS23 (g), respectively. The haplotype sketches illustrated the resolved haplotype structures by integrating WGS and OGM results.\u003c/p\u003e","description":"","filename":"floatimage7.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-6126136/v1/25999897a92e19e48262e4f0.jpeg"},{"id":78229283,"identity":"a04629b5-e3a1-49b3-a656-ff2aac08ee5a","added_by":"auto","created_at":"2025-03-11 07:20:29","extension":"jpeg","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":418436,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eClinical diagnosis flowchart.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eClinical diagnosis flowchart for duplication of uncertain clinical significance. 1 It is recommended that prenatal samples undergo simultaneous WGS and OGM analysis to reduce diagnostic turnaround time. For OGM, cultured amniotic fluid samples are preferred; alternatively, blood samples from a male family member with the same duplication can be utilized. Additionally, The haplotype structure of DMD in female carrier samples may be challenging to resolve due to allelic complexity. 2 In-frame duplications generally result in BMD; however, certain genotypes that disrupt core functional domains may lead to DMD. 3 Frameshift or extragenic-insert duplications predominantly cause DMD, with rare exceptions leading to BMD.\u003c/p\u003e","description":"","filename":"floatimage8.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-6126136/v1/1b9290a955c5a6775533c1ac.jpeg"},{"id":84050846,"identity":"245c19a2-e62a-40a4-a958-da9f38f37648","added_by":"auto","created_at":"2025-06-06 08:23:32","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":5506564,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6126136/v1/40ea2eec-f50c-4a9f-8b3c-6780463b3173.pdf"},{"id":78226814,"identity":"a7433325-c594-49ff-b8f6-5a4c604da199","added_by":"auto","created_at":"2025-03-11 07:04:29","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":2075099,"visible":true,"origin":"","legend":"\u003cp\u003eSupplymentary Table 1, 3 and 5; Supplementary Fig.\u0026nbsp;1 to 7.\u003c/p\u003e","description":"","filename":"SupplementaryInformation.docx","url":"https://assets-eu.researchsquare.com/files/rs-6126136/v1/04a0771a0be25a1ecc8cd461.docx"},{"id":78226028,"identity":"ad8cd650-3642-4b83-b4a2-95e85e753ac7","added_by":"auto","created_at":"2025-03-11 06:56:29","extension":"xlsx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":21759,"visible":true,"origin":"","legend":"\u003cp\u003eSupplymentary Table 2\u003c/p\u003e","description":"","filename":"SupplymentaryTable2.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-6126136/v1/21df8951155b7131370455e6.xlsx"},{"id":78226049,"identity":"25071461-ef1e-4ba8-ab79-b0ead7feefef","added_by":"auto","created_at":"2025-03-11 06:56:31","extension":"xlsx","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":18558,"visible":true,"origin":"","legend":"\u003cp\u003eSupplymentary Table 4\u003c/p\u003e","description":"","filename":"SupplymentaryTable4.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-6126136/v1/efad93c5486cb91eddd32650.xlsx"},{"id":78226027,"identity":"c7c82ec0-4889-4c71-a595-b2efaae0dd3f","added_by":"auto","created_at":"2025-03-11 06:56:29","extension":"xlsx","order_by":4,"title":"","display":"","copyAsset":false,"role":"supplement","size":14317,"visible":true,"origin":"","legend":"\u003cp\u003eSupplymentary Table 6\u003c/p\u003e","description":"","filename":"SupplymentaryTable6.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-6126136/v1/91d001b2d605a6f1ba455d06.xlsx"}],"financialInterests":"There is \u003cb\u003eNO\u003c/b\u003e Competing Interest.","formattedTitle":"Integrated Genotyping Strategies Uncovering Detailed Haplotype Structures and Characterization of DMD duplications","fulltext":[{"header":"Introduction","content":"\u003cp\u003eDuchenne and Becker muscular dystrophies (DMD/BMDs) are the most prevalent X-linked inherited muscle disorders, with an estimated incidence ranging from 1 in 5,000 to 1 in 6,000 in newborn males \u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u003c/sup\u003e. DMD/BMDs are primarily caused by mutations in dystrophin gene (\u003cem\u003eDMD\u003c/em\u003e), resulting in either a loss of function or insufficient dosage of dystrophin protein. \u003cem\u003eDMD\u003c/em\u003e (OMIM: 300377) is the largest known gene in humans, spanning approximately 2.4 Mb on chromosome Xp21.1 and consisting of 79 exons. Structural variants (SVs) are the most common type of mutation in DMD/BMD, accounting for approximately 80% of cases, with duplications contributing to about 10% of all mutations \u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e. While numerous studies have characterized the duplication spectrum at exon-level resolution \u003csup\u003e\u003cspan additionalcitationids=\"CR4\" citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e, with the Leiden Open Variation Database (LOVD) documenting nearly 1,000 distinct duplication genotypes as of June 2024 \u003csup\u003e5\u003c/sup\u003e, the precise genomic architecture, duplication patterns, and underlying molecular mechanisms remain to be fully elucidated.\u003c/p\u003e \u003cp\u003eThe clinical manifestations of DMD/BMD demonstrate remarkable phenotypic heterogeneity. DMD patients typically present with calf muscle pseudohypertrophy and difficulties in walking or stair climbing between the ages of two and five years. The disease progressively worsens, with most patients losing ambulation by 12 years of age and experiencing life-threatening cardiomyopathy leading to mortality by the age of 20 \u003csup\u003e6\u003c/sup\u003e. In contrast, BMD patients exhibit milder symptoms, characterized by later onset and slower progression, with ambulation typically preserved until the fifth or sixth decade \u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e. Genotype-phenotype analyses have shown that mutation effects in the open reading frame (ORF) of \u003cem\u003eDMD\u003c/em\u003e can be predictive of the disease severity, with frameshift mutations generally leading to DMD and in-frame mutations typically resulting in BMD \u003csup\u003e\u003cspan additionalcitationids=\"CR8\" citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e. While genotype-phenotype predictions based on deletion mutations demonstrate high reliability, duplications frequently present with inconsistent clinical outcomes, complicating prognostic assessments \u003csup\u003e\u003cspan additionalcitationids=\"CR11\" citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e. Notably, emerging evidence has identified benign \u003cem\u003eDMD\u003c/em\u003e duplication variants that do not manifest clinical phenotypes, further highlighting the complexity of these genomic alterations \u003csup\u003e\u003cspan additionalcitationids=\"CR14\" citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e. The reasons behind the phenotypic variability associated with \u003cem\u003eDMD\u003c/em\u003e duplications remain poorly understood.\u003c/p\u003e \u003cp\u003eCurrent diagnostic approaches for DMD/BMD primarily employ multiplex ligation-dependent probe amplification (MLPA), Sanger sequencing and high throughput sequencing such as whole exome sequencing (WES), collectively enabling the identification of pathogenic variants in approximately 98% of cases \u003csup\u003e\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u003c/sup\u003e. However, these conventional methods exhibit significant limitations in detecting complex SVs, especially in resolving precise haplotype configurations for duplications. Whole genome sequencing (WGS) has gradually been integrated into clinical practice, offering enhanced capacity for detecting SVs \u003csup\u003e\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e. Moreover, long-read sequencing (LRS) technologies have recently been applied to \u003cem\u003eDMD\u003c/em\u003e mutation detection, enabling the analysis of large fragments extending to tens of kilobases \u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e,\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e. Optical genome mapping (OGM), another long-read technology, can detect SVs on a megabase scale and has been used effectively for previously undiagnosed cases \u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e,\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e. The advent of these advanced genomic technologies provides unprecedented opportunities for high-resolution molecular characterization of \u003cem\u003eDMD\u003c/em\u003e duplications, facilitating more precise genotype-phenotype predictions in clinical diagnostics.\u003c/p\u003e \u003cp\u003eIn this study, we reanalyzed the \u003cem\u003eDMD\u003c/em\u003e-MLPA and WES data from a cohort of 3,842 patients and identified 39 individuals with \u003cem\u003eDMD\u003c/em\u003e duplications. To elucidate the molecular architecture of these duplications, we performed WGS on all 39 patients and subsequently applied OGM to six WGS-unresolved cases. Through comparative analysis of detection methodologies, we established an optimized framework for accurate identification and characterization of DMD duplications. Furthermore, we conducted a genotype-phenotype analysis based on the duplication haplotypes, enabling refined clinical interpretation. This investigation provided valuable insights into the diversity and complexity of \u003cem\u003eDMD\u003c/em\u003e duplications, significantly advancing our capabilities in pathogenicity assessment, genetic counseling, and clinical screening for \u003cem\u003eDMD\u003c/em\u003e duplications.\u003c/p\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eCohort Design\u003c/h2\u003e \u003cp\u003eTo explore the genomic complexity of \u003cem\u003eDMD\u003c/em\u003e duplications and their genotype-phenotype relationships, we conducted a retrospective analysis of patients who presented at Shanghai Children\u0026rsquo;s Medical Center between August 2016 and January 2024. The study cohort comprised two distinct patient populations: (1) 578 patients clinically suspected of DMD/BMD who had undergone MLPA, Sanger sequencing, or WES, and (2) 3,264 patients with suspected alternative genetic disorders who had also been evaluated through WES analysis. Among them, 39 patients with confirmed \u003cem\u003eDMD\u003c/em\u003e duplications (26 detected by MLPA and 13 by WES) were recalled for WGS and OGM analyses to resolve their haplotype structures (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ea). The final study cohort consisted of 31 male DMD/BMD patients, 1 female BMD patient, 1 asymptomatic male carrier, and 6 asymptomatic female carriers (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eb, Supplementary Table\u0026nbsp;1, Supplementary Table\u0026nbsp;2). Pedigree analysis revealed that these patients originated from 38 independent families (Supplementary Fig. \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e). Notably, WGS34 and WGS35 were derived from the same family lineage; therefore, only WGS34 was included in subsequent analyses to maintain statistical independence.\u003c/p\u003e \u003cp\u003e \u003cb\u003eHigh Heterogeneity of\u003c/b\u003e \u003cb\u003eDMD\u003c/b\u003e \u003cb\u003eDuplications Revealed by WGS\u003c/b\u003e\u003c/p\u003e \u003cp\u003eA total of 38 distinct duplications were identified in our cohort (Supplementary Table\u0026nbsp;3). The duplications exhibited remarkable heterogeneity among individuals and shared little structural similarity with known pathological duplications documented in DECIPHER database (Supplementary Fig.\u0026nbsp;2a, b). The size distribution of these duplications ranged from 10,725 bp to 884,245 bp, with a median of 260,488 bp, where 50% of duplications were between 100 kb and 500 kb. Notable, 40% of duplications (n\u0026thinsp;=\u0026thinsp;16) were split into two or more segments with more than two breakpoints, and 24% (n\u0026thinsp;=\u0026thinsp;9) involved extragenic regions. Additionally, some duplications overlapped with other OMIM genes, such as \u003cem\u003eIL1RAPL1\u003c/em\u003e (OMIM: 300206), \u003cem\u003eCFAP47\u003c/em\u003e (OMIM: 301057), and \u003cem\u003eXK\u003c/em\u003e (OMIM: 314850) (Supplementary Fig.\u0026nbsp;2b). Specifically, one female sample (WGS33) displayed a duplication that included a segment of chromosome Y, suggesting an interchromosomal translocation event.\u003c/p\u003e \u003cp\u003eThe exon-level resolution analysis revealed 34 unique patterns among the 38 identified \u003cem\u003eDMD\u003c/em\u003e duplications. Only four recurrent patterns were observed: exon 2 (E2), exons 3\u0026ndash;4 (E3-4), exons 3\u0026ndash;5 (E3-5), and exons 46\u0026ndash;49 (E46-49), each appearing in two samples, while the remaining patterns appeared only once in our cohort (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ea). Six duplication genotypes were newly identified in our cohort, including exons 5\u0026ndash;45 (E5-45), exons 8\u0026ndash;9/28\u0026ndash;43 (E8-9, 28\u0026ndash;43), exons 24\u0026ndash;44 (E24-44), exons 31\u0026ndash;33 (E31-33), exons 45\u0026ndash;52/55 (E45-52, 55), and exons 48\u0026ndash;51 (E48-51) (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eb). Cumulative duplication frequencies differed between our cohort and public database. While LOVD reported higher duplication frequencies in exons 2\u0026ndash;9, our cohort demonstrated a predominance of duplications in exons 45\u0026ndash;51 (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ec). This discrepancy may be attributed to differences in cohort size and population characteristics.\u003c/p\u003e \u003cp\u003eUsing WGS, we classified duplications into four patterns: Tandem duplication (Tandem-Dup, 58%, n\u0026thinsp;=\u0026thinsp;22), Duplication-Normal-Duplication (Dup-Nor-Dup, 16%, n\u0026thinsp;=\u0026thinsp;6), Duplication-Inversion-Duplication (Dup-Inv-Dup, 16%, n\u0026thinsp;=\u0026thinsp;6), and Intricate duplication (Intricate-Dup, 10%, n\u0026thinsp;=\u0026thinsp;4) (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ed, Supplementary Fig.\u0026nbsp;3, Supplementary Data). The Tandem-Dup exhibited a single haplotype that could be fully resolved by WGS. In contrast, the remaining three patterns, collectively referred to complex duplications, had multiple potential haplotypes, thereby complicating haplotype determination and variant interpretation.\u003c/p\u003e \u003cp\u003e \u003cb\u003eBreakpoint Features of\u003c/b\u003e \u003cb\u003eDMD\u003c/b\u003e \u003cb\u003eDuplication\u003c/b\u003e\u003c/p\u003e \u003cp\u003eWe identified 118 breakpoints across 59 breakpoint pairs, with 79% located within \u003cem\u003eDMD\u003c/em\u003e introns. The remaining breakpoints were distributed upstream (12%), downstream (6%), interchromosomal (2%) or within \u003cem\u003eDMD\u003c/em\u003e exons (2%) (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003ea). Notably, breakpoints in intron 2 were mainly linked to Tandem-Dup, while those in intron 44 were associated with complex duplication patterns (Supplementary Fig.\u0026nbsp;4a). Breakpoints in intron 49 frequently led to duplications extending beyond \u003cem\u003eDMD\u003c/em\u003e (Supplementary Fig.\u0026nbsp;4b).\u003c/p\u003e \u003cp\u003ePrevious studies have shown that repeat elements \u003csup\u003e\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e, GC-unbalanced regions \u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e and meiotic DNA strand breaks (DSBs) \u003csup\u003e\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e contribute to SV formation. To investigate their roles in \u003cem\u003eDMD\u003c/em\u003e duplications, we defined genomic regions containing these features as genome unstable regions. Our analysis revealed that 73% of breakpoint pairs had both ends located within genome unstable regions, while the remaining 27% had one end in such regions (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eb, Supplementary Fig.\u0026nbsp;4c). Moreover, the number of breakpoints within introns showed a positive correlation with intron length and the density of genome unstable regions (Supplementary Fig.\u0026nbsp;4d, e). Specifically, 50% of breakpoints were located within repeat elements, with SINE (Alu) and LINE (L1) being the most prevalent (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003ec). Additionally, the 5' ends of breakpoint pairs were significant enriched in GC-poor regions compared to genomic background (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003ed). Furthermore, breakpoints were highly associated with DSB hotspots, particularly those linked to complex duplication patterns (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003ee).\u003c/p\u003e \u003cp\u003eAmong the 59 breakpoint pairs, 30 showed microhomology sequences ranging from 1 to 7 bp, 18 contained no-templated insertions at the junctions, and 11 exhibited direct joins (Supplementary Table\u0026nbsp;4). Most breakpoints had a homology ratio of 60\u0026ndash;70% in the flanking sequences, with seven pairs showing exceptionally high homology (\u0026gt;\u0026thinsp;80%) (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003ef). These high-homology breakpoint pairs were predominantly located in repeat elements, specifically SINE: Alu (n\u0026thinsp;=\u0026thinsp;6) and LINE: L1 (n\u0026thinsp;=\u0026thinsp;1) (Supplementary Fig.\u0026nbsp;4f). These pairs were more likely to induce complex duplication patterns, such as Dup-Nor-Dup (n\u0026thinsp;=\u0026thinsp;2), Dup-Inv-Dup (n\u0026thinsp;=\u0026thinsp;3), and Intricate-Dup (n\u0026thinsp;=\u0026thinsp;2) (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eg). Genomic annotation of these high-homology breakpoint pairs showed one end located in \u003cem\u003eDMD\u003c/em\u003e introns (introns 17, 49, 51, 57, or 61), while the other end was located outside of \u003cem\u003eDMD\u003c/em\u003e. This resulted in duplications extending beyond \u003cem\u003eDMD\u003c/em\u003e and led to distinct clinical outcomes.\u003c/p\u003e \u003cp\u003e \u003cb\u003eOGM Is Essential for Precise Localization and Haplotype Determination of\u003c/b\u003e \u003cb\u003eDMD\u003c/b\u003e \u003cb\u003eDuplications\u003c/b\u003e\u003c/p\u003e \u003cp\u003eUsing WGS, we identified 13 samples with multiple possible haplotypes, each potentially leading to different effects on the \u003cem\u003eDMD\u003c/em\u003e products. We applied OGM to six WGS-unresolved samples, demonstrating its utility in defining SV structures and assessing pathogenicity.\u003c/p\u003e \u003cp\u003eIn case WGS40, a duplication was incidentally detected during prenatal testing of a 14-week pregnant woman (Gravida 2 Para 0). Her first pregnancy was terminated due to prenatal diagnoses of Tetralogy of Fallot, diaphragmatic hernia, and spinal abnormalities. Chromosomal microarray analysis (CMA) revealed no pathogenic or likely pathogenic copy number variations (CNVs) in the fetus, but it incidentally identified a maternally inherited duplication of exons 50\u0026ndash;55 in \u003cem\u003eDMD\u003c/em\u003e, classified as a variant of uncertain significance (VUS). This duplication had been reported in one BMD and 11 DMD cases in LOVD, suggesting a potential risk for DMD in offspring. Genotype-phenotype co-segregation analysis found no male family members carrying the duplication (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003ea). WGS further characterized the duplication as a Dup-Nor-Dup pattern, consisting of two segments: chrX:31567450\u0026ndash;31831842 (exons 50\u0026ndash;55 of \u003cem\u003eDMD\u003c/em\u003e) and chrX:33890225\u0026ndash;34038113 (intergenic region) (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eb). This pattern suggested two possible haplotypes: duplication insertion into intron 55 of \u003cem\u003eDMD\u003c/em\u003e or insertion into chrX:34038113 (extragenic). OGM clarified that the duplication was inserted into chrX:34038113, outside \u003cem\u003eDMD\u003c/em\u003e (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003ec). As this insertion did not disrupt \u003cem\u003eDMD\u003c/em\u003e transcription, the duplication was classified as benign, posing no risk for DMD/BMD in the offspring.\u003c/p\u003e \u003cp\u003eIn case WGS25, a 4.7-year-old boy presented with delayed motor milestones, muscle weakness, and elevated creatine kinase (CK), CK-MB, Aspartate transferase (AST), Alanine transaminase (ALT), and Lactate Dehydrogenase (LDH). Family history indicated that his maternal uncle lost ambulation at an early age (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003ea). WES identified a duplication of exons 45\u0026ndash;51 in \u003cem\u003eDMD\u003c/em\u003e, predicted to maintain the reading frame and typically associated with BMD. However, the patient's clinical presentation and family history were consistent with DMD. Notably, LOVD records showed phenotypic variability for this duplication, with six cases classified as DMD, two as BMD, and one reported as an asymptomatic male. WGS further characterized the duplication as a Dup-Nor-Dup pattern comprising two segments: chrX:31749465\u0026ndash;32066307 (exons 45\u0026ndash;51 of \u003cem\u003eDMD\u003c/em\u003e) and chrX:36315168\u0026ndash;36597300 (exons 57\u0026ndash;65 of \u003cem\u003eCFAP47\u003c/em\u003e) (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eb). Two potential haplotypes were identified: one with the duplication inserted into intron 51 of \u003cem\u003eDMD\u003c/em\u003e and the other with insertion into intron 56 of \u003cem\u003eCFAP47\u003c/em\u003e. OGM revealed that the duplication inserted into intron 51 of \u003cem\u003eDMD\u003c/em\u003e, aligning with the patient's DMD phenotype and family history (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003ec).\u003c/p\u003e \u003cp\u003eIn case WGS31, a 6-year-old male exhibited muscle weakness, hypertrophic calves, Gowers' sign, and attention deficits. MLPA identified a maternally inherited duplication of exons 46\u0026ndash;49 in \u003cem\u003eDMD\u003c/em\u003e, though the family history was unavailable (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003ea). WGS revealed a complex duplication with six breakpoints and three segments: chrX:28281895\u0026ndash;28382937 (intergenic), chrX:28486473\u0026ndash;28668996 (exon 1 of \u003cem\u003eIL1RAPL1\u003c/em\u003e), and chrX:31832901\u0026ndash;31954458 (exons 46\u0026ndash;49 of \u003cem\u003eDMD\u003c/em\u003e) (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eb). OGM resolved the structure and confirmed that these three duplicated segments were combined and inserted into intron 49 of \u003cem\u003eDMD\u003c/em\u003e, disrupting \u003cem\u003eDMD\u003c/em\u003e products and aligning with the observed phenotype (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003ec).\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eOGM May Miss Detection of Certain Exons Due to Low Resolution\u003c/h3\u003e\n\u003cp\u003eAlthough OGM significantly aids in identifying duplication structures, its ability to replace conventional detection technologies remains uncertain. When mapping OGM labeling sites to \u003cem\u003eDMD\u003c/em\u003e, it was observed that some labels skipped two exons (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003ea). Specifically, among the 559 labeling sites identified within \u003cem\u003eDMD\u003c/em\u003e introns, 17 introns lacked labeling sites, suggesting that SVs with breakpoints located in these introns may be missed by OGM (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eb). Additionally, the resolution of OGM on \u003cem\u003eDMD\u003c/em\u003e is approximately 550 bp, which means smaller SVs may be undetected (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003ec).\u003c/p\u003e \u003cp\u003eIn our cohort, discrepancies between OGM and WGS results were observed in 33% of the detections (n\u0026thinsp;=\u0026thinsp;2). In case WGS10, WGS identified two duplication segments: chrX:32275575\u0026ndash;32447885 (exons 28\u0026ndash;43 of \u003cem\u003eDMD\u003c/em\u003e) and chrX:32676655\u0026ndash;32765682 (exons 8\u0026ndash;9 of \u003cem\u003eDMD\u003c/em\u003e) (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003ed). However, the absence of OGM labeling sites in intron 26 resulted in uncertain boundary determination for segment A (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003ee). Similarly, in case WGS23, WGS indicated a Dup-Inv-Dup pattern (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003ef), while OGM detected a tandem duplication. This discrepancy is likely due to the small size of segment C (411 bp), which falls below OGM\u0026rsquo;s resolution limit (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eg).\u003c/p\u003e \u003cp\u003e \u003cb\u003eComparative Analysis of Multiple Detection Technologies for\u003c/b\u003e \u003cb\u003eDMD\u003c/b\u003e \u003cb\u003eduplications\u003c/b\u003e\u003c/p\u003e \u003cp\u003eMLPA and WES are commonly used for detecting variants in \u003cem\u003eDMD\u003c/em\u003e. We compared their exon-level performance with WGS to assess diagnostic consistency (Supplementary Table\u0026nbsp;5). Among the 26 samples analyzed using MLPA, 88% (n\u0026thinsp;=\u0026thinsp;23) were concordant with WGS results, while 12% (n\u0026thinsp;=\u0026thinsp;3) exhibited discrepancies: one sample missed a non-contiguous exon (WGS26), another involved a female sample with both a duplication and a deletion on separate alleles (MLPA generated a combined result) (WGS22), and the third missed an exon containing a breakpoint (WGS19). For the 16 samples analyzed by both WES and WGS, 69% (n\u0026thinsp;=\u0026thinsp;11) were concordant. Discrepancies were observed in 31% of cases: 19% (n\u0026thinsp;=\u0026thinsp;3) had one-exon discrepancies at duplication boundaries (WGS09, WGS18, and WGS27), 6% (n\u0026thinsp;=\u0026thinsp;1) missed a non-contiguous exon (WGS26), and 6% (n\u0026thinsp;=\u0026thinsp;1) failed to detect an exon with a breakpoint (WGS39). Among the six samples analyzed using both WGS and OGM, 67% (n\u0026thinsp;=\u0026thinsp;4) showed consistent results, and discrepancies were noted in 33% (n\u0026thinsp;=\u0026thinsp;2) due to OGM\u0026rsquo;s limited resolution (WGS10, WGS23).\u003c/p\u003e \u003cp\u003e \u003cb\u003eGenotype-Phenotype Analysis for\u003c/b\u003e \u003cb\u003eDMD\u003c/b\u003e \u003cb\u003eDuplications\u003c/b\u003e\u003c/p\u003e \u003cp\u003eAccurate understanding genotype-phenotype relationships is crucial for precise clinical diagnosis. By leveraging the detailed mapping of duplication haplotypes provided by WGS and OGM, we were able to more accurately assess their effects on \u003cem\u003eDMD\u003c/em\u003e ORF and conduct robust phenotype-correlation analyses (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e, Supplementary Table\u0026nbsp;6). For duplications with multiple potential haplotypes, the most likely structure was inferred based on the patient\u0026rsquo;s phenotype or family history. The predicted effects were classified into four categories: none (no impact on \u003cem\u003eDMD\u003c/em\u003e), in-frame (duplication preserves the ORF), frameshift (duplication disrupts the ORF), and extragenic-insert (duplication contains large extragenic segments).\u003c/p\u003e \u003cp\u003eTable 1 Summary of phenotype and predicted ORF\u0026rsquo;s effects in diagnosed male patients\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\" width=\"100%\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 26.8041%;\"\u003e\n \u003cp\u003eDiagnosed Male Patients\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 17.5258%;\"\u003e\n \u003cp\u003enone\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 17.5258%;\"\u003e\n \u003cp\u003ein-frame\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 17.5258%;\"\u003e\n \u003cp\u003eframeshift\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 20.6186%;\"\u003e\n \u003cp\u003eextragenic-insert\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 26.8041%;\"\u003e\n \u003cp\u003eDMD (n=25)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 17.5258%;\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 17.5258%;\"\u003e\n \u003cp\u003e3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 17.5258%;\"\u003e\n \u003cp\u003e19\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 20.6186%;\"\u003e\n \u003cp\u003e5\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 26.8041%;\"\u003e\n \u003cp\u003eBMD (n=4)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 17.5258%;\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 17.5258%;\"\u003e\n \u003cp\u003e4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 17.5258%;\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 20.6186%;\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 26.8041%;\"\u003e\n \u003cp\u003eAsymptomatic (n=1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 17.5258%;\"\u003e\n \u003cp\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 17.5258%;\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 17.5258%;\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 20.6186%;\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 26.8041%;\"\u003e\n \u003cp\u003eConcordance Rate \u003csup\u003e1\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 17.5258%;\"\u003e\n \u003cp\u003e100%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 17.5258%;\"\u003e\n \u003cp\u003e57%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 17.5258%;\"\u003e\n \u003cp\u003e100%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 20.6186%;\"\u003e\n \u003cp\u003e100%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u0026sup1;The concordance rate indicates the rate that the observed phenotype aligned with the predicted ORF effect. Specifically, \u0026ldquo;None\u0026rdquo; is expected to be asymptomatic, \u0026ldquo;In-frame\u0026rdquo; is expected to correspond to BMD, and both \u0026ldquo;Frameshift\u0026rdquo; and \u0026ldquo;Extragenic-insert\u0026rdquo; are expected to result in DMD.\u003c/p\u003e\u003cp\u003eAs expected, patients with duplications predicted as none were asymptomatic, and all male patients with frameshift duplications (n\u0026thinsp;=\u0026thinsp;18) were diagnosed to DMD. However, among patients with in-frame duplications (n\u0026thinsp;=\u0026thinsp;7), 57% (n\u0026thinsp;=\u0026thinsp;4) presented with BMD, while 43% (n\u0026thinsp;=\u0026thinsp;3) were diagnosed with DMD. The in-frame duplications associated with DMD involved two specific genotypes and exhibited a Tandem-Dup pattern: exons 3\u0026ndash;4 duplication (n\u0026thinsp;=\u0026thinsp;2, WGS04, WGS05) and exon 34 duplication (n\u0026thinsp;=\u0026thinsp;1, WGS20). Furthermore, all duplications classified as extragenic-inserts (n\u0026thinsp;=\u0026thinsp;5) were linked to DMD. Notably, extragenic-insert duplications included genotypes such as exons 45\u0026ndash;51 duplication (WGS25) and exons 45\u0026ndash;57 duplication (WGS29), which were initially predicted as in-frame with canonical ORF prediction tools \u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e,\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e,\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e\u003c/sup\u003e, resulted in DMD due to the involvement of large intergenic regions. This refined classification improved the accuracy of genotype-phenotype predictions.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eIn this study, we uncovered the remarkable diversity and complexity of \u003cem\u003eDMD\u003c/em\u003e duplications, extending beyond previous exon-level analyses \u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e. Within our cohort, four exon-level duplication genotypes were recurrent, but their breakpoints varied significantly. Additionally, six novel genotypes were identified in our cohort, which had not been reported previously. We classified these duplications into four distinct patterns: Tandem-Dup (58%), Dup-Nor-Dup (16%), Dup-Inv-Dup (16%), and Intricate-Dup (11%). Notably, only approximately 60% followed the canonical tandem arrangement, while the remaining 40% displayed more complex configurations. Other potential patterns, such as inverted-insert tandem duplications \u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e, may have been missed due to the limited cohort size. These findings highlight the limitations of widely used databases, such as LOVD, which primarily catalog exon-level genotypes of \u003cem\u003eDMD\u003c/em\u003e duplications, but lack critical details information about duplication patterns, breakpoints, and orientations. Moreover, the canonical reading-frame checker assumes a tandem arrangement of duplicated exons, which may lead to inaccuracies in ORF effect predictions and clinical interpretations \u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e,\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e,\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e\u003c/sup\u003e. To improve genetic counseling and phenotype predictions, it is essential to expand these databases with comprehensive details on duplication characteristics and their associated phenotypes.\u003c/p\u003e \u003cp\u003eOur breakpoint analysis sheds light on the potential mechanisms underlying duplication formation. Consistent with prior studies, breakpoints were predominantly located in genomic regions prone to instability \u003csup\u003e\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/sup\u003e. We hypothesize that replication-based fork stalling and template switching (FoSTeS) is key drivers of duplication formation (Supplementary Fig.\u0026nbsp;5). Previous studies have implicated highly homologous repeats in both tandem duplications via non-allelic homologous recombination (NAHR) and more complex structures, such as duplication-triplication/inverted-duplication via multiple rounds of FoSTeS \u003csup\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u003c/sup\u003e. In our cohort, highly homologous repeats were associated with Dup-Nor-Dup (n\u0026thinsp;=\u0026thinsp;2), Dup-Inv-Dup (n\u0026thinsp;=\u0026thinsp;3), and Intricate-Dup (n\u0026thinsp;=\u0026thinsp;2), suggesting the duplication formation may involve both NAHR and FoSTeS (Supplementary Fig.\u0026nbsp;6, 7). These repeat elements, particularly in \u003cem\u003eDMD\u003c/em\u003e introns 17, 49, 51, 57, and 61, contributed to duplications that extended beyond \u003cem\u003eDMD\u003c/em\u003e, further complicating pathogenicity interpretation. Additional data are further required to validate and refine potential hotspot regions for complex duplications beyond \u003cem\u003eDMD\u003c/em\u003e.\u003c/p\u003e \u003cp\u003eCase studies from our cohort underscore the importance of resolving duplication haplotypes to accurately determine clinical outcomes. WGS25 and WGS40 both involved Dup-Nor-Dup duplications, but their clinical outcomes differed. In WGS40, the duplication was located in an intergenic region upstream of \u003cem\u003eDMD\u003c/em\u003e and was classified as benign, whereas the duplication of WGS25 disrupted \u003cem\u003eDMD\u003c/em\u003e, leading to DMD phenotypes. Another instructive case, WGS31, involved a duplication spanning \u003cem\u003eDMD\u003c/em\u003e and \u003cem\u003eIL1RAPL1\u003c/em\u003e, a gene associated with neurodevelopmental disorders. A previously reported complex duplication, combined with \u003cem\u003eIL1RAPL1\u003c/em\u003e (exons 3\u0026ndash;11) and \u003cem\u003eDMD\u003c/em\u003e (exon 1), was linked to autism spectrum disorder, rather than DMD/BMD \u003csup\u003e\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e\u003c/sup\u003e, emphasizing the significance of the disrupted gene in determining disease manifestation. Collectively, these cases highlight the critical need for comprehensive structural characterization and the adoption of a detailed nomenclature, including duplicated segments, duplication orientation, inserted locations, and all affected pathological genes, to improve clinical interpretations. For instance, the detailed haplotype of WGS31 could be described as \u0026ldquo;\u003cem\u003eDMD\u003c/em\u003e:intron49ins[(chrX:28281004\u0026ndash;28383006)INV, \u003cem\u003eIL1RAPL1\u003c/em\u003e:exon1, \u003cem\u003eDMD\u003c/em\u003e:exon46-49]dup\u0026rdquo;, referring to a joint duplication of three segments: inversion of chrX:28281004\u0026ndash;28383006, exon 1 of \u003cem\u003eIL1RAPL1\u003c/em\u003e, and exons 46\u0026ndash;49 of \u003cem\u003eDMD\u003c/em\u003e, inserted into intron 49 of \u003cem\u003eDMD\u003c/em\u003e. All duplications in our cohort were nomenclature accordingly in Supplementary Table\u0026nbsp;2.Our findings reinforced the phenotype prediction based on reading-frame rule \u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e, with all frameshift duplications resulted in DMD and 57% of in-frame duplications were linked to BMD. However, 43% of in-frame duplications (n\u0026thinsp;=\u0026thinsp;3) caused DMD, specifically involving exons 3\u0026ndash;4 duplications (n\u0026thinsp;=\u0026thinsp;2) and exon 34 duplications (n\u0026thinsp;=\u0026thinsp;1). These genotypes are also more frequently associated with DMD in LOVD (11 DMD vs. 7 BMD for exons 3\u0026ndash;4 duplication; 2 DMD vs. 0 BMD for exon 34 duplication),likely due to the disruption of critical functional domains, such as the N-terminal F-actin-binding domain (exons 3\u0026ndash;4) and the R11 region of the rod domain (exon 34), impairing sarcolemma interaction and leading to severe phenotypes \u003csup\u003e\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e\u003c/sup\u003e. We also introduced a new ORF effect category, extragenic-insert, for duplications containing large extragenic segments. These duplications extend beyond conventional ORF-based predictions and are often associated with severe phenotypes. This refinement improved the accuracy of genotype-phenotype prediction. For example, the exons 45\u0026ndash;51 duplication (WGS25) and exons 45\u0026ndash;57 duplication (WGS29), previously predicted as BMDs under conventional ORF-based rules, were accurately reclassified as DMD due to the involvement of large extragenic regions. Further research is needed to elucidate the pathological mechanisms underlying such duplications and their effects on \u003cem\u003eDMD\u003c/em\u003e transcription and splicing.\u003c/p\u003e \u003cp\u003eOur study underscored the limitations of detection platforms. WGS failed to resolve complex haplotype structures in 34% of cases (13/38) due to its short-read constraints. OGM effectively resolved haplotypes in 67% of ambiguous cases (4/6), but it missed smaller segments in 33% of cases (2/6) because of its resolution thresholds. Besides, OGM demonstrated higher sample quality requirements and cost burdens compared to WGS. Combining WGS and OGM results could provide more comprehensive SV detection. Taken together, we propose a sequential testing strategy: employ WGS to analyze duplications identified as VUS by MLPA, WES, CMA or noninvasive prenatal testing (NIPT) \u003csup\u003e\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e\u003c/sup\u003e, followed by OGM for unresolved cases to identify haplotypes and predict clinical outcomes (Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003e). In the long term, the development of novel integrative approaches or more advanced technologies will be essential to fully capture the complete spectrum of \u003cem\u003eDMD\u003c/em\u003e variants and improve diagnostic accuracy.\u003c/p\u003e \u003cp\u003eOur study does have several limitations. The relatively modest sample size could introduce potential sampling bias, and single-exon duplications might be underestimated in the WES reanalysis. Additionally, we lacked \u003cem\u003eDMD\u003c/em\u003e transcript data due to the constraints of our sample types. Nevertheless, our cohort is still one of the largest datasets on \u003cem\u003eDMD\u003c/em\u003e duplications and integrates multi-technology analyses, thereby providing robust insights into the \u003cem\u003eDMD\u003c/em\u003e duplication spectrum and advancing pathogenicity interpretations.\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003ePatient information\u003c/h2\u003e \u003cp\u003eA total of 3,842 patients underwent genetic testing (MLPA \u0026amp; sanger sequencing, WES) at the Department of Medical Genetics and Molecular Diagnostic Laboratory, Shanghai Children\u0026rsquo;s Medical Center (SCMC) from August 2016 to January 2024. The clinical information and genetic testing data were collected and reanalyzed. Among them, 39 patients confirmed with \u003cem\u003eDMD\u003c/em\u003e duplications were recruited for WGS and OGM. Ethical approval was obtained (SCMCIRB-K2020060-1), and written informed consent was provided by all participants.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eMLPA\u003c/h2\u003e \u003cp\u003eCNVs of \u003cem\u003eDMD\u003c/em\u003e were detected via a modified MLPA as previously described\u003csup\u003e\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e\u003c/sup\u003e. Briefly, CNVplex kit (Genesky Diagnostics, Suzhou, China) were used following the manufacturer\u0026rsquo;s instructions. Data were visualized and analyzed with the corresponding CNVplex analysis software.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eOGM\u003c/h3\u003e\n\u003cp\u003eUltra-high molecular weight DNA was extracted, labeled, and processed from peripheral blood using the Prep SP Blood and Cell DNA Isolation Kit (Bionano Genomics, San Diego, CA), Qubit Fluorometers (Bionano Genomics, San Diego, CA), and Prep DLS Labeling kits (Bionano Genomics Inc., San Diego, CA, USA). The labeled DNA was loaded for linearization and imaging using Saphyr platform (Bionano Genomics, San Diego, CA). The genome of proband was assembled using the Bionano Solve V3.7 software. SVs were called against the human reference genome (GRCh38/hg38) and visualized using Bionano Access software (V1.7). The site of OGM labeling is generated with digest_genome.py of HiC-Pro \u003csup\u003e\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e (v3.1.0) with the parament of \u0026ldquo;-r ^CTTAAG\u0026rdquo;.\u003c/p\u003e\n\u003ch3\u003eWES and WES-based CNV detection\u003c/h3\u003e\n\u003cp\u003eWES analysis followed previously described methods \u003csup\u003e\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e\u003c/sup\u003e. CNVs were detected using CNVkit \u003csup\u003e\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e\u003c/sup\u003e (v0.9.9), with samples from the same batch used as controls. The results were filtered based on the following criteria: SVLEN\u0026thinsp;\u0026gt;\u0026thinsp;50000 or \u0026lt; -50000, PROBES\u0026thinsp;\u0026gt;\u0026thinsp;5, and FOLD_CHANGE_LOG\u0026thinsp;\u0026gt;\u0026thinsp;0.5 or \u0026lt; -0.8.\u003c/p\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eWGS and WGS-based CNV detection\u003c/h2\u003e \u003cp\u003eGenomic DNA was extracted and sequencing libraries constructed using the MGIEasy PCR-Free DNA Library Prep Kit (MGI Tech, Shenzhen, China). Adapters were trimmed with fastp \u003csup\u003e\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e\u003c/sup\u003e (v0.23.2), and reads were aligned to hg38 using BWA \u003csup\u003e\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e\u003c/sup\u003e (v0.2.10). SVs and CNVs were called using DELLY \u003csup\u003e\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e\u003c/sup\u003e (version 0.7.6), CNVkit \u003csup\u003e\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e\u003c/sup\u003e (v0.9.9), and Control-FREEC \u003csup\u003e\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e\u003c/sup\u003e (v11.6), then merged with SURVIVOR \u003csup\u003e\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e\u003c/sup\u003e (v1.0.7). Genotypes were called using svtyper \u003csup\u003e\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e\u003c/sup\u003e (v0.7.1).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eBreakpoint feature analysis\u003c/h2\u003e \u003cp\u003eBreakpoints were extracted from soft-clipped reads and validated by UCSC BLAT (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://genome.ucsc.edu/cgi-bin/hgBlat\u003c/span\u003e\u003cspan address=\"https://genome.ucsc.edu/cgi-bin/hgBlat\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) \u003csup\u003e\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e\u003c/sup\u003e. GC content was calculated from 100 bp regions flanking the breakpoints, and 10,000 random controls were generated using bedtools shuffle (v2.31.0) as genomic background. The regions with GC content less than 35% were defined as GC-poor, while regions with GC content greater than 60% as GC-rich. Meiotic DSB hotspots were obtained from GSE59836 \u003csup\u003e22\u003c/sup\u003e and random regions were generated using bedtools shuffle (v2.31.0) with 10,000 iterations to serve as randomly selected background controls. Breakpoints located within 500 bp of DSB hotspots were classified as being within DSB hotspots. Repeat elements were obtained from UCSC Table Browser. Homology ratios of flanking sequences at breakpoint pairs were calculated using Biopython (v1.85), employing a 500 bp sliding window across the 1000 bp regions surrounding each breakpoint pair, and the maximum value was selected.\u003c/p\u003e \u003cp\u003e \u003cb\u003eEffects on the\u003c/b\u003e \u003cb\u003eDMD\u003c/b\u003e \u003cb\u003eORF\u003c/b\u003e\u003c/p\u003e \u003cp\u003eThe effects of duplications on the \u003cem\u003eDMD\u003c/em\u003e ORF (NM_004006.3) were predicted based the reconstructed haplotype structures. For duplications with multiple potential haplotypes, predictions were made for each haplotype independently. The effect was determined by considering the cumulative length of inserted exons in the transcriptional direction. In cases where exons were invertedly inserted into \u003cem\u003eDMD\u003c/em\u003e introns, the insertion length was assigned as \"0\" due to the loss of splicing donor and acceptor motifs \u003csup\u003e\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e\u003c/sup\u003e. If the cumulative insertion length was a multiple of three, the duplication was predicted to maintain the reading frame (in-frame); otherwise, it was predicted to disrupt the reading frame (frameshift). Duplications involving large segments of intergenic regions or other genes were classified under the category extragenic-insert.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003ePublic data download in LOVD\u003c/h2\u003e \u003cp\u003e \u003cem\u003eDMD\u003c/em\u003e duplication variants and associated phenotypes were downloaded from LOVD \u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e (updated 2024/06/24) (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://databases.lovd.nl/shared/genes/DMD\u003c/span\u003e\u003cspan address=\"https://databases.lovd.nl/shared/genes/DMD\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e"},{"header":"Abbreviations","content":"\u003cp\u003e\u003cem\u003eDMD: dystrophin\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eDMD:\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003eDuchenne muscular dystrophy\u003c/p\u003e\n\u003cp\u003eBMD:\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003eBecker muscular dystrophy\u003c/p\u003e\n\u003cp\u003eMD: muscular dystrophy\u003c/p\u003e\n\u003cp\u003eWES: whole exome sequencing\u003c/p\u003e\n\u003cp\u003eMLPA: multiplex ligation-dependent probe amplification\u003c/p\u003e\n\u003cp\u003eWGS: whole genome sequencing\u003c/p\u003e\n\u003cp\u003eOGM: Optical genome mapping\u003c/p\u003e\n\u003cp\u003eTandem-Dup: tandem duplication\u003c/p\u003e\n\u003cp\u003eDup-Nor-Dup: duplication-normal-duplication\u003c/p\u003e\n\u003cp\u003eDup-Inv-Dup: duplication-inversion-duplication\u003c/p\u003e\n\u003cp\u003eIntricate-Dup: intricate duplication\u003c/p\u003e\n\u003cp\u003eSV: Structural variants\u003c/p\u003e\n\u003cp\u003eLOVD: Leiden Open Variation Database\u003c/p\u003e\n\u003cp\u003eORF: open reading frame\u003c/p\u003e\n\u003cp\u003eLRS: long-read sequencing\u003c/p\u003e\n\u003cp\u003eDSB: meiotic DNA strand break\u003c/p\u003e\n\u003cp\u003eNAHR: non-allelic homologous recombination\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eFoSTeS: fork stalling and template switching\u003c/p\u003e\n\u003cp\u003eCK: creatine kinase\u003c/p\u003e\n\u003cp\u003eAST: Aspartate transferase\u003c/p\u003e\n\u003cp\u003eALT: Alanine transaminase\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eLDH: Lactate Dehydrogenase\u003c/p\u003e\n\u003cp\u003eCNV: copy number variation\u003c/p\u003e\n\u003cp\u003eVUS: variants of uncertain significance\u003c/p\u003e\n\u003cp\u003eABD: N-terminal F-actin-binding domain\u003c/p\u003e\n\u003cp\u003eR: Rod domain\u003c/p\u003e\n\u003cp\u003eCR: cysteine-rich domain\u003c/p\u003e\n\u003cp\u003eCT: C-terminal domain\u003c/p\u003e\n\u003cp\u003eNIPT: noninvasive prenatal testing\u003c/p\u003e\n\u003cp\u003eCMA: chromosomal microarray analysis\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eAvailability of data and materials\u003c/strong\u003e\u003cbr\u003e\u0026nbsp;The raw data have been deposited into the Genome Sequence Archive (https://ngdc.cncb.ac.cn/gsa-human/) with the accession number HRA009618.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003cbr\u003e\u0026nbsp;This study was supported by the Key Discipline Group of Pudong New Area Health Commission (PWZxq2022-07), the Shanghai Key Laboratory of Clinical Molecular Diagnostics for Pediatrics (20dz2260900), Shanghai \u0026ldquo;Rising Stars of Medical Talents\u0026rdquo; Youth Development Program, Postdoctoral Fellowship Program of CPSF (GZC20241067), and National Clinical Key Specialty Construction Project (10000015Z155080000004). We sincerely appreciate the support of BeCreative Lab in Beijing for optical genome mapping analysis and Biosan Diagnostics in Hangzhou for whole genome sequencing analysis. We are also deeply grateful to the patients and their families for their participation in this study.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTingting Yu: Conceptualization, Resources, Writing - Review \u0026amp; Editing, Supervision, Funding acquisition; Ruen Yao, Qihua Fu: Conceptualization, Resources, Writing - Review \u0026amp; Editing, Funding acquisition; Jin Sun, Jie Tang: Methodology, Formal analysis, Visualization, Writing - Original Draft; Lu Wei: Data Curation, Writing - Review \u0026amp; Editing; Geng Juan, Xiao Rui: Methodology, Writing - Review \u0026amp; Editing; Jian Wang, Niu Li, Shuyuan Li: Resources, Writing - Review \u0026amp; Editing. All authors read and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003cbr\u003eThe authors declare no competing interests.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eMah JK et al (2014) A systematic review and meta-analysis on the epidemiology of Duchenne and Becker muscular dystrophy. Neuromuscul Disord 24:482\u0026ndash;491\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhao L et al (2024) Comprehensive analysis of 2097 patients with dystrophinopathy based on a database from 2011 to 2021. Orphanet J Rare Dis 19:311\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKong X et al (2019) Genetic analysis of 1051 Chinese families with Duchenne/Becker Muscular Dystrophy. BMC Med Genet 20:139\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTuffery-Giraud S et al (2009) Genotype-phenotype analysis in 2,405 patients with a dystrophinopathy using the UMD-DMD database: a model of nationwide knowledgebase. Hum Mutat 30:934\u0026ndash;945\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAartsma-Rus A, Van Deutekom JC, Fokkema IF, Van Ommen GJ, Den Dunnen JT (2006) Entries in the Leiden Duchenne muscular dystrophy mutation database: an overview of mutation types and paradoxical cases that confirm the reading-frame rule. Muscle Nerve 34:135\u0026ndash;144\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMercuri E, Muntoni F (2013) Muscular dystrophies. Lancet 381:845\u0026ndash;860\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGorgoglione D et al (2024) Natural history of Becker muscular dystrophy: DMD gene mutations predict clinical severity. Brain\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMonaco AP, Bertelson CJ, Liechti-Gallati S, Moser H, Kunkel LM (1988) An explanation for the phenotypic differences between patients bearing partial deletions of the DMD locus. \u003cem\u003eGenomics\u003c/em\u003e 2, 90\u0026thinsp;\u0026ndash;\u0026thinsp;5\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMagri F et al (2011) Genotype and phenotype characterization in a large dystrophinopathic cohort with extended follow-up. J Neurol 258:1610\u0026ndash;1623\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLim KRQ, Nguyen Q, Yokota T (2020) Genotype-Phenotype Correlations in Duchenne and Becker Muscular Dystrophy Patients from the Canadian Neuromuscular Disease Registry. J Pers Med 10\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang J et al (2013) MLPA-based genotype-phenotype analysis in 1053 Chinese patients with DMD/BMD. BMC Med Genet 14:29\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVengalil S et al (2017) Duchenne Muscular Dystrophy and Becker Muscular Dystrophy Confirmed by Multiplex Ligation-Dependent Probe Amplification: Genotype-Phenotype Correlation in a Large Cohort. J Clin Neurol 13:91\u0026ndash;97\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang Y et al (2024) Prenatal risk assessment of Xp21.1 duplication involving the DMD gene by optical genome mapping. Life Sci Alliance 7\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShen J et al (2024) Comprehensive analysis of genomic complexity in the 5' end coding region of the DMD gene in patients of exons 1\u0026ndash;2 duplications based on long-read sequencing. BMC Genomics 25:292\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBai Y et al (2022) Long-Read Sequencing Revealed Extragenic and Intragenic Duplications of Exons 56\u0026ndash;61 in DMD in an Asymptomatic Male and a DMD Patient. Front Genet 13:878806\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBeggs AH, Koenig M, Boyce FM, Kunkel LM (1990) Detection of 98% of DMD/BMD gene deletions by polymerase chain reaction. Hum Genet 86:45\u0026ndash;48\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePluta N et al (2023) Whole-Genome Sequencing Identified New Structural Variations in the DMD Gene That Cause Duchenne Muscular Dystrophy in Two Girls. Int J Mol Sci 24\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLing C et al (2023) Uncovering the true features of dystrophin gene rearrangement and improving the molecular diagnosis of Duchenne and Becker muscular dystrophies. iScience 26:108365\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXiao B et al (2024) Combining optical genome mapping and RNA-seq for structural variants detection and interpretation in unsolved neurodevelopmental disorders. Genome Med 16:113\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiao X et al (2023) Repetitive DNA sequence detection and its role in the human genome. Commun Biology 6:954\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKiktev DA, Sheng Z, Lobachev KS, Petes TD (2018) GC content elevates mutation and recombination rates in the yeast\u0026thinsp;\u0026lt;\u0026thinsp;i\u0026thinsp;\u0026gt;\u0026thinsp;Saccharomyces cerevisiae. \u003cem\u003eProceedings of the National Academy of Sciences\u003c/em\u003e 115, E7109-E7118\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePratto F et al (2014) DNA recombination. Recombination initiation maps of individual human genomes. Science 346:1256442\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhou J, Xin J, Niu Y, Wu S (2017) DMDtoolkit: a tool for visualizing the mutated dystrophin protein and predicting the clinical severity in DMD. BMC Bioinformatics 18:87\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCarvalho CM, Lupski JR (2016) Mechanisms underlying structural variant formation in genomic disorders. Nat Rev Genet 17:224\u0026ndash;238\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGrochowski CM et al (2024) Inverted triplications formed by iterative template switches generate structural variant diversity at genomic disorder loci. Cell Genom 4:100590\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEisfeldt J et al (2024) Resolving complex duplication variants in autism spectrum disorder using long-read genome sequencing. Genome Res 34:1763\u0026ndash;1773\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBhosle RC, Michele DE, Campbell KP, Li Z, Robson RM (2006) Interactions of intermediate filament protein synemin with dystrophin and utrophin. Biochem Biophys Res Commun 346:768\u0026ndash;777\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBrison N et al (2019) Maternal copy-number variations in the DMD gene as secondary findings in noninvasive prenatal screening. Genet Med 21:2774\u0026ndash;2780\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang X et al (2015) A modified multiplex ligation-dependent probe amplification method for the detection of 22q11.2 copy number variations in patients with congenital heart disease. BMC Genomics 16:364\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eServant N et al (2015) HiC-Pro: an optimized and flexible pipeline for Hi-C data processing. Genome Biol 16:259\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYao R, Yu T, Qing Y, Wang J, Shen Y (2019) Evaluation of copy number variant detection from panel-based next-generation sequencing data. Mol Genet Genomic Med 7:e00513\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTalevich E, Shain AH, Botton T, Bastian BC, CNVkit (2016) Genome-Wide Copy Number Detection and Visualization from Targeted DNA Sequencing. PLoS Comput Biol 12:e1004873\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen S, Zhou Y, Chen Y, Gu J (2018) fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics 34:i884\u0026ndash;i890\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi H, Durbin R (2009) Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics 25:1754\u0026ndash;1760\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRausch T et al (2012) DELLY: structural variant discovery by integrated paired-end and split-read analysis. Bioinformatics 28:i333\u0026ndash;i339\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBoeva V et al (2012) Control-FREEC: a tool for assessing copy number and allelic content using next-generation sequencing data. Bioinformatics 28:423\u0026ndash;425\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJeffares DC et al (2017) Transient structural variations have strong effects on quantitative traits and reproductive isolation in fission yeast. Nat Commun 8:14061\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChiang C et al (2015) SpeedSeq: ultra-fast personal genome analysis and interpretation. Nat Methods 12:966\u0026ndash;968\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKent WJ (2002) BLAT\u0026ndash;the BLAST-like alignment tool. \u003cem\u003eGenome Res\u003c/em\u003e 12, 656\u0026thinsp;\u0026ndash;\u0026thinsp;64\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXie Z et al (2020) Long-read whole-genome sequencing for the genetic diagnosis of dystrophinopathies. Ann Clin Transl Neurol 7:2041\u0026ndash;2046\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-6126136/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6126136/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eDuchenne and Becker muscular dystrophies (DMD/BMDs) are X-linked genetic disorders caused by mutations in the dystrophin gene (\u003cem\u003eDMD\u003c/em\u003e), characterized by progressive muscle weakness and degeneration. While \u003cem\u003eDMD\u003c/em\u003e duplications account for approximately 10% of cases, their clinical impact varies significantly, ranging from severe phenotypes to asymptomatic presentations, posing significant challenges in determining their pathogenicity. This study investigates the molecular complexity of DMD duplications and their implications for disease progression. Through analyzing 3,842 patients using multiple sequencing platforms, we identified 39 cases with \u003cem\u003eDMD\u003c/em\u003e duplications and characterized four distinct duplication patterns. These structure variations not only influence pathogenicity interpretation but also reflect specific mechanisms of genomic instability. Our findings reveal that conventional genetic testing methods frequently fail to accurately resolve duplication structures, limiting their predictive value for clinical outcomes. By integrating whole genome sequencing and optical genome mapping, we achieved precise haplotype resolution, substantially enhancing genotype\u0026ndash;phenotype correlations. These results underscore the critical importance of adopting multi-platform genomic strategies to improve diagnostic accuracy, refine pathogenicity assessment, and optimize personalized genetic counseling for patients with \u003cem\u003eDMD\u003c/em\u003e duplications.\u003c/p\u003e","manuscriptTitle":"Integrated Genotyping Strategies Uncovering Detailed Haplotype Structures and Characterization of DMD duplications","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-03-11 06:56:24","doi":"10.21203/rs.3.rs-6126136/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"35220908-4912-4d62-afc4-139a1dc4020f","owner":[],"postedDate":"March 11th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":45320389,"name":"Biological sciences/Genetics/Clinical genetics/Genetic testing"},{"id":45320390,"name":"Health sciences/Diseases/Neurological disorders/Dystonia"}],"tags":[],"updatedAt":"2025-06-06T08:15:24+00:00","versionOfRecord":[],"versionCreatedAt":"2025-03-11 06:56:24","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6126136","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6126136","identity":"rs-6126136","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-24T02:00:01.246996+00:00
License: CC-BY-4.0