Multi-omic re-analysis increases diagnostic yield in individuals with Cornelia de Lange syndrome

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract Background Approximately half of the individuals with a clinically diagnosed Mendelian condition do not receive a molecular diagnosis. Current standard-of-care diagnostic pipelines, which are largely focused on exonic sequence variants, may not be comprehensive enough to identify all pathogenic variants. A comprehensive analytical approach capable of identifying noncoding and structural variants is needed to bridge the diagnostic gap. Cornelia de Lange Syndrome (CdLS) is a multisystem developmental diagnosis caused primarily by pathogenic variants in one of the six genes known to cause CdLS (NIPBL, SMC3, SMC1A, HDAC8, RAD21, BRD4) , although pathogenic variants in additional phenocopy genes have also been implicated. We hypothesized that individuals with a clinical diagnosis of CdLS and no molecular diagnosis harbor pathogenic, causative variants that are not identified or prioritized by the current standard of care, exome-focused workflows. Methods We performed a re-analysis of the genome sequencing data from a previously published cohort of 173 individuals with a clinical diagnosis of CdLS (Gabriella Miller Kids First cohort) and expanded the scope of analysis to include noncoding and structural variants. We used RNA-sequencing data in a subset of individuals (n = 62) to complement the DNA workflow. Results Re-analysis, including copy-number and structural variation and using transcriptome sequencing as a complementary assay, revealed molecular etiologies in an additional 37 probands. Thus, increasing the total diagnostic yield in this previously undiagnosed cohort to 60%. The new diagnoses were enriched for variants beyond the standard exonic SNVs/ indels, including cryptic non-coding variants (promoter, deep intronic, large insertions), copy number variants, balanced rearrangements such as inversions, and variants in additional genes that phenocopy CdLS. Transcriptome aided re-analysis helped uncover cryptic noncoding variants in the DNA that lacked sufficient computational evidence for a splicing abnormality and yet produced aberrantly spliced mRNA. Conclusions Our results underscore the need for whole genome (and transcriptome) sequencing and a comprehensive, unbiased analytical protocol integrating structural and noncoding variants to exhaustively mine a phenotypically and genetically heterogeneous cohort to maximize its diagnostic yield. The additional diagnostic yield solely from noncoding and structural variants highlights the limitations of an exome-focused analysis workflow and highlights the utility of transcriptome analysis beyond the use of splicing prediction tools.
Full text 97,288 characters · extracted from preprint-html · click to expand
Multi-omic re-analysis increases diagnostic yield in individuals with Cornelia de Lange syndrome | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Multi-omic re-analysis increases diagnostic yield in individuals with Cornelia de Lange syndrome Ramakrishnan Rajagopalan, Tejas Jammihal, Tanaya Jadhav, Maninder Kaur, and 5 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9078185/v1 This work is licensed under a CC BY 4.0 License Status: Under Revision Version 1 posted 9 You are reading this latest preprint version Abstract Background Approximately half of the individuals with a clinically diagnosed Mendelian condition do not receive a molecular diagnosis. Current standard-of-care diagnostic pipelines, which are largely focused on exonic sequence variants, may not be comprehensive enough to identify all pathogenic variants. A comprehensive analytical approach capable of identifying noncoding and structural variants is needed to bridge the diagnostic gap. Cornelia de Lange Syndrome (CdLS) is a multisystem developmental diagnosis caused primarily by pathogenic variants in one of the six genes known to cause CdLS (NIPBL, SMC3, SMC1A, HDAC8, RAD21, BRD4) , although pathogenic variants in additional phenocopy genes have also been implicated. We hypothesized that individuals with a clinical diagnosis of CdLS and no molecular diagnosis harbor pathogenic, causative variants that are not identified or prioritized by the current standard of care, exome-focused workflows. Methods We performed a re-analysis of the genome sequencing data from a previously published cohort of 173 individuals with a clinical diagnosis of CdLS (Gabriella Miller Kids First cohort) and expanded the scope of analysis to include noncoding and structural variants. We used RNA-sequencing data in a subset of individuals (n = 62) to complement the DNA workflow. Results Re-analysis, including copy-number and structural variation and using transcriptome sequencing as a complementary assay, revealed molecular etiologies in an additional 37 probands. Thus, increasing the total diagnostic yield in this previously undiagnosed cohort to 60%. The new diagnoses were enriched for variants beyond the standard exonic SNVs/ indels, including cryptic non-coding variants (promoter, deep intronic, large insertions), copy number variants, balanced rearrangements such as inversions, and variants in additional genes that phenocopy CdLS. Transcriptome aided re-analysis helped uncover cryptic noncoding variants in the DNA that lacked sufficient computational evidence for a splicing abnormality and yet produced aberrantly spliced mRNA. Conclusions Our results underscore the need for whole genome (and transcriptome) sequencing and a comprehensive, unbiased analytical protocol integrating structural and noncoding variants to exhaustively mine a phenotypically and genetically heterogeneous cohort to maximize its diagnostic yield. The additional diagnostic yield solely from noncoding and structural variants highlights the limitations of an exome-focused analysis workflow and highlights the utility of transcriptome analysis beyond the use of splicing prediction tools. Biological sciences/Computational biology and bioinformatics Health sciences/Diseases Biological sciences/Genetics Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Introduction Standard exonic, sequence variant-focused analytical approaches, which currently form the standard of care in molecular diagnostics, present significant limitations[ 1 , 2 ]. These approaches predominantly identify protein-coding variants while often overlooking noncoding, regulatory, and structural variants that may contribute to Cornelia de Lange Syndrome (CdLS) pathogenesis. While the standard of care sequencing has moved from capture-based exome sequencing to whole genome sequencing, the analytical workflows are still focused on exonic regions of the genome and interpretable/ reportable variants (exome slice on a genome backbone). Variants in enhancer or promoter regions, deep intronic variants that affect splicing, or structural variants such as copy number alterations and chromosomal rearrangements can elude traditional exome sequencing or exome-focused analytical approaches[ 1 , 3 – 5 ]. As a result, individuals with clinically suspected CdLS but without identifiable exonic pathogenic variants may remain undiagnosed, emphasizing the need for broader genome-wide analyses, including whole-genome sequencing (WGS) and transcriptomic profiling, to capture the full spectrum of pathogenic variants. Routine re-analysis of genomic data has also proven valuable in increasing diagnostic yield, as newly discovered pathogenic variants and updated classification criteria can help identify previously unrecognized disease-causing mutations[ 4 , 6 , 7 ]. As bioinformatic tools advance and variant interpretation frameworks evolve, periodic re-evaluation of existing genomic data can significantly improve diagnostic accuracy and patient outcomes. Cornelia de Lange Syndrome (CdLS; MIM# 122470, 300590, 610759, 614701, 300882, 620568) is a genetically heterogeneous and clinically variable, multi-system diagnosis characterized by characteristic facial features, microcephaly, limb malformations, prenatal and postnatal growth retardation, global developmental delays and various other systemic involvement (e.g. diaphragmatic hernia, congenital heart differences, cleft palate, hypertichosis, intestinal malrotation and others)[ 8 ]. CdLS presents with a highly variable phenotype often complicating early clinical diagnosis, and the severity and spectrum of manifestations can differ significantly among affected individuals, necessitating a comprehensive diagnostic approach. Molecular confirmation of CdLS is achieved through genetic testing, which has identified six primary causative genes: NIPBL, SMC1A, SMC3, RAD21, HDAC8, and BRD4 . Among these, variants in NIPBL account for most cases (~ 80%), while variants in SMC1A, SMC3, RAD21, HDAC8, and BRD4 collectively contribute to 10–15% of cases and manifest varying phenotypic expression of the diagnosis[ 9 ]. These genes encode key components (structural and regulatory) of the cohesin complex, which plays a crucial role in chromosomal segregation, DNA repair, and gene expression regulation. Despite advances in genetic testing, the diagnostic landscape of CdLS remains challenging due to phenotypic overlap with several other syndromes, including KBG syndrome ( ANKRD11) , Coffin-Siris Syndrome ( ARID1B) , Rubinstein-Taybi Syndrome ( CREBBP, EP300) , and Wiedemann-Steiner Syndrome ( KMT2A) , among others[ 7 , 10 – 13 ]. These phenocopy syndromes exhibit clinical similarities to CdLS but arise from distinct genetic etiologies, underscoring the need for comprehensive, genome-wide, and, at times, multi-omic diagnostic approaches. We recently reported the mutational landscape of CdLS and related diagnoses in a large cohort of 716 individuals. A likely molecular etiology was identified in 59% of the probands tested (n = 422 out of 716) using standard molecular diagnostic approaches, including targeted panels, exome sequencing, and exome slice on a genome platform (GS cohort n = 177). In the total cohort of 716 probands clinically diagnosed with CdLS, 358 (50%) had variants in known CdLS-associated genes, including NIPBL (271 individuals, 64%), SMC1A (40, 9%), HDAC8 (25, 6%), SMC3 (16, 4%), and RAD21 (6, 1%). Additionally, 64 probands (9%) had variants in 32 different genes that phenocopied CdLS. Notably, in the GS cohort described by Kaur et al., a molecular diagnosis was identified in 37.5% of individuals, while the rest of the cohort remained without a molecular diagnosis[ 8 ]. We hypothesized that the individuals with a clinical diagnosis of CdLS but no molecular diagnosis from prior analysis harbor pathogenic, disease-causing variants that could not be identified or prioritized by the standard of care exome-focused workflows. In this work, we performed a comprehensive re-analysis of the genome sequencing data using an expanded workflow to include noncoding, copy-number, and structural variants, along with the transcriptome data for a subset of probands. By addressing the limitations of the exome-focused workflows and utilizing emerging methodologies, this paper aims to explore the evolving landscape of genomic diagnostics of rare Mendelian diagnoses, emphasizing the need for comprehensive analytical workflows and integration of advanced genomic technologies including the role of multi-omics analyses. We demonstrate that a comprehensive analytical workflow including noncoding and structural variants significantly increases the diagnostic yield in a Mendelian diagnosis such as CdLS. Methods Patient cohort We started with a cohort of 177 individuals with genome sequencing data described in Kaur et. al. Four samples were removed from this cohort including three with failed familial relationships and one being a duplicate resulting in 173 individuals with genome sequencing data. This cohort included 88 trios, 29 duos, and 56 singletons. These individuals went through an exome slice/panel analysis focused on genes known to cause CdLS and related neurodevelopmental diagnoses. Sixty-five individuals had a previously identified molecular diagnosis reported in the Kaur et. al., leaving 108 individuals without a molecular diagnosis. A subset of 54 individuals without a molecular diagnosis (in whom cell lines were available to isolate RNA) from this cohort had whole transcriptome data (Fig. 1). Ethics statement All probands and family members were enrolled and consented to participate in a research study approved by the Institutional Review Board at CHOP. Data production and analysis Genome sequencing (GS) was performed at the BROAD institute on the Illumina Novaseq X platform at a minimum average read depth of 30x and details of the prior analytical workflow is described elsewhere[ 8 ]. Raw sequencing data (fastq) were aligned to the human genome reference hg38 using DRAGEN v3.9. The quality control protocol included sex check as identified by the DRAGEN pipeline (Illumina Inc. San Diego, CA) and the kinship coefficients estimated by somalier [ 14 ]. Copy number variants identified using the read depth-based algorithm, cnvpytor[ 15 ], were clustered along with breakends identified by the split-read algorithm, manta [ 16 ], to identify dispersed duplications. Variant prioritization was performed by multiple computational workflows following a variant filtration heuristic including a phenotype-agnostic, weighted sum model using several variant features (an internal tool) and phenotype-based variant prioritization algorithm Exomiser[ 17 ]. Whole genome VCF was filtered to retain high quality (site quality QUAL > 20), rare variants (gnomAD v.3.1 AF < 0.5%) before getting passed to the variant prioritization algorithms[ 18 ]. The homegrown, phenotype agnostic algorithm employed a linear weighted sum model using inheritance information when family members were available, variant allele frequency from gnomAD v3.1, relevant computational scores for missense (REVEL, missense Z)[ 19 ], loss of function intolerance for the gene (LOEUF), splice prediction for intronic variants (spliceAI)[ 20 ], conservation score (GERP 3.0), and the presence in known mutation databases (HGMD, ClinVar)[ 21 ]. RNA sequencing was performed at the BROAD institute at a minimum depth of 100M reads per sample and DRAGEN v3.9 RNA-seq workflow was used to align the raw reads to the human transcriptome and estimate the abundance of transcripts defined in ENCODE V39. The RNA outlier analysis was performed using the R package OUTRIDER [ 22 ] and the aberrant splicing outliers were identified using leafcutter[ 23 ]. DNA sequencing data was investigated for variants when there was an aberrant expression or splicing event detected in the RNA analysis. All the other downstream analyses and visualizations were produced using in-house scripts written in Python/ R. Variants were classified by a clinical laboratory geneticist according to the ACMG guidelines. Variant confirmation Newly identified variants that required orthogonal confirmation were determined on a case-by-case basis. All variants were manually inspected in the CRAM alignments file for their presence and the strength of evidence. A subset of newly identified variants was confirmed using Sanger sequencing (SNV/ INDEL), droplet-digital PCR (copy-number variants), and long-range PCR followed by Gel electrophoresis (long insertions). Results Summary Kaur et. al. reported a 37.5% diagnostic yield in a genome sequencing cohort of 173 patients with CdLS and related diagnoses. The re-analysis workflow described in this paper was able to identify all the previously reported diagnoses. Expanding the genome analysis workflow to include intronic noncoding variants and structural variants and integrating transcriptome sequencing as a complementary assay in a subset of probands revealed molecular etiologies in an additional 37 new diangoses. Notably, this reanalysis identified additional diagnoses in two probands previously reported as positives in Kaur et. al. which increased the total diagnostic yield in the GS cohort to 57.8% (100 of 173). The new diagnoses from re-analysis were enriched for variants beyond standard exonic SNV/ indels. Cryptic non-coding variants (promoter, deep intronic), copy number variation including mosaic events, and chromosomal rearrangements including inversions and insertional translocations made up 26 of the 37 new diagnoses (70.3%) (Fig. 2). Four probands in the cohort had two likely causative variants identified (4/173; 2.3%). Seventeen of the 37 new diagnoses (46%) were in the previously known CdLS genes which included 12 noncoding sequence variants and five structural variants. Thirteen of the 37 new diagnoses were identified de novo where parental samples available, 17 of unknown inheritance, and the rest inherited from a parent. A parallel analysis of whole transcriptome data in 54 individuals helped guide the analysis of genome sequencing data and help uncover at least two of the noncoding, cryptic splice variants unnoticed by the primary DNA analysis. Structural variants The structural variants identified included two copy-neutral inversions, 9 deletions including two large (> 5Mb) deletions, 2 duplications including one large (> 5Mb) duplication and an intergenic duplication, and 2 long (> 1kb) insertions (Table 1). The smallest structural variant identified was an 84bp deletion in the last of exon of the gene EP300 known to cause Menke-Hennekam Syndrome 2 (OMIM #618333) and the largest was a 24Mb deletion in 3q which included multiple haploinsufficient genes STAG1, MED12L, ZIC1, and FOXL2 . Five of the structural variants involved a known CdLS gene. This included a single exon, 3.2kb deletion in RAD21 , a 350kb deletion involving HDAC8 , a 2.8Mb inversion disrupting NIPBL with a breakpoint in intron 6, a long insertion in the intron 33 of NIPBL , and a 252kb dispersed duplication from chr4 inserted into HDAC8. There were nine other probands with a structural variant identified in genes known to cause other developmental diagnoses. The three large copy-number variants included two deletions (9.3Mb in 1q and 24Mb in 3q) involving multiple syndromic genes, and a 7.6Mb duplication including a triplosensitive gene SETD1B. Other events included a 571kb inversion involving ZEB2 (Mowat-Wilson syndrome) and deletions involving PHF6 (Borjeson-Forssman-Lehmann syndrome) and HPRT1 (Lesch-Nyhan syndrome), ANKRD11 (KBG syndrome), and BCL11A (Dias-Logan syndrome). One proband had a deletion in HNRNPD (HNRNP-related neurodevelopmental disorders) and an intragenic duplication involving exons 5–6 of SETBP1 (Intellectual developmental disorder, autosomal dominant 29). Finally, a de novo , 20.5kb deletion involving exons 7 to 15 in EPC1 , a potential candidate gene, was identified in a single proband. EPC1 (Enhancer of Polycomb Homolog 1), one of the polycomb group of proteins, is not associated with a Mendelian phenotype but predicted to be constrained for loss of function variants in common population (pLI = 1 and LOEUF 0.23). Cohesin functionally interacts with the polycomb group of proteins responsible for epigenetic silencing of proteins during development making it a strong candidate gene for further exploration. Noncoding variants There were 12 noncoding, intronic splice or cryptic splice, and promoter variants. Eleven of these variants were in established CdLS genes and the remaining one was a pair of obligatory splice variants in METTL5. There were 3 patients with promoter variants in NIPBL (two c.-467C > T and one c.-324delC). These promoter variants in NIPBL (GRCh38 5:36,876,445CG > G; c.-324delC, chr5:36876791C > T; c.467C > T) were predicted to disrupt the Kozak sequence in the 5’ UTR of NIPBL . The c.-467C > T has been previously identified as de novo in an individual with CdLS and shown to produce reduced mRNA expression[ 24 , 25 ]. While c.-324delC has not been seen before, similar variants creating a frameshift in the uORF have been reported in patients with CdLS[ 24 ]. All but one cryptic intronic splice variant had a SpliceAI score > 0.2. The c.4560 + 1975G > A in NIPBL had a SpliceAI score of 0.08. Other exonic sequence variants There were 11 exonic sequence variants (SNVs/ indels) rescued by the reanalysis that were not previously described in Kaur et. al. None of the 11 variants were in the known CdLS genes. These included apparently loss of function, stop gain variants in WWOX (homozygous p.Arg264*), a frameshift in ZMYM2 (p.thr188Asnfs*12), a start loss in TAF1 (c.1A > C), and several previously published pathogenic variants in ClinVar ( CDC42 p.Ala159Val ( de novo) and p.Asp170Gly, NAA10 p.Arg83Cys ( de novo) , SETD5 p.Val305Gly). The inheritance information for all these variants is shown in Table 1. Transcriptome analysis The transcriptome analysis resolved two novel diagnoses overlooked by the DNA workflow. First one (CDL-111 and CDL-378) involved a deep intronic variant in NIPBL ( c.4560 + 1975 G > A) with a very low spliceAI score (< 0.2) yet resulting in an inclusion of a novel exon in the intron 21 of NIPBL (Fig. 3). Outlier splicing analysis using leafcutter revealed a novel junction involving exon 21, and further manual inspection suggested a novel exon inclusion between exons 21 and 22. Another proband (CDL-053) had a strong variant of uncertain significance, p.(Cys194Arg), in the gene EBF3 known to cause Hypotonia, ataxia, and delayed development syndrome (OMIM #617330). This variant is well conserved in a protein domain and adjacent to the previously published p.Lys193Asn variant. However, expression outlier analysis using OUTRIDER revealed SMC3 as an outlier with reduced expression. Further inspection of DNA sequencing data revealed a maternally inherited intronic variant in SMC3 (GRCh38 chr10:110589714T > C; NM_005445.3:c.1409 + 6T > C) and the variant was not prioritized for analysis (Fig. 4). The genome re-analysis identified a de novo insertion in intron 32 of NIPBL in CDL-200 that could not be fully resolved because it involved repetitive sequences and read pairs whose mates mapped to multiple locations across the genome (Fig. 5). RNA sequencing indicated an exon 33 skipping event supported by five reads; however, this signal could not be reliably distinguished from background noise. Further ddPCR expression assay (probe in exon 33) showed reduced expression compared to parental samples (Fig. 6). Discussion In this study, we performed a comprehensive re-analysis of genome sequencing data from 173 probands including structural and noncoding variants. The re-analysis yielded 37 new diagnoses with an enrichment for noncoding and structural variants and increased the overall diagnostic yield to 57.8%. The proportion of the noncoding and structural variants identified in the re-analysis (26 of the 37; 70.3%) highlights the limitations of an exome focused analysis approach and the opportunities for narrowing the diagnostic gap in Mendelian diagnoses. Structural variant callers for short-read sequencing data typically identify several thousands of events including a large proportion of breakends enriched for false positives originating from repetitive or low complexity regions of the genome making it hard to prioritize true and clinically relevant structural rearrangements. Insertional translocations are not typically recognized from short-read sequencing data and need specialized workflows. One of the insertional translocations identified in this study disrupted HDAC8 , a known gene for CdLS, but had a read-pair signature consistent with a balanced translocation. However, further manual review of the alignments revealed that there were two subgroups of reads, each mapping to proximal and distal ends of a 252kb duplication in chr4. The two inversions identified in this study were much smaller than 10Mb in length (2.8Mb and 571Kb) and cannot be detected by conventional Karyotyping methods. While whole genome sequencing has enabled the detection of these events, they are not routinely detected due to the challenges in differentiating true events from likely false positives. The two inversions reported in our study did not involve repetitive regions and had breakpoint spanning reads which helped recognize them as true rearrangements. The insertion in the NIPBL in CDL-200 involved a repeat element that cannot be characterized by short-read sequencing and the deletion involving the BCL11A in CDL-637 involved additional sequences at both ends suggesting a repeat element insertion at the deletion site. These complex rearrangements are often mediated by repeat or segmental duplication mediated events and the ambiguity in the read-pair signatures arising from mapping issues hamper interpretation. Long read sequencing can alleviate this challenge when there are breakpoint spanning reads that are anchored in a unique region. However, long read sequencing may encounter similar challenges if the region involved is longer than the length of the reads. The other structural variant findings including simple deletions, duplications, and cytogenetic events may speak to the disparities in the prior analytical workflows and testing protocols. The use of transcriptome data to guide the re-analysis was helpful in both identifying previously unrecognized cryptic splice variants and confirming the aberrant splicing consequences of cryptic splice variants with computational evidence. The RNA analysis for CDL-053 using OUTRIDER revealed SMC3 as an outlier for that patient and the RNA alignments showed novel junctions which prompted us to identify a maternally inherited c.1409 + 6T > C variant in SMC3. This variant was not prioritized by the variant prioritization algorithm Exomiser likely because it was inherited from an unaffected mother. Similarly, a de novo intronic variant in NIPBL (CDL-378) was previously unrecognized due to the lack of computational evidence for aberrant splicing (SpliceAI score = 0.08). Splicing outlier analysis suggested a novel junction in intron 21 of NIPBL which prompted us to recognize this de novo intronic variant as clinically relevant. While the transcriptome data was helpful in guiding the DNA analysis, it could obscure the real diagnosis when there are multiple events resulting in aberrant expression. For example, in a family with proband (CDL-315) and an affected mother, the RNA expression outlier analysis revealed a different set of genes as outliers (Fig. 5). The proband showed several genes as upregulated outliers ( BAZ1B, TBL2, RFC2) pointing to the 7q11.23 region associated with Williams-Buren syndrome typically resulting from the loss of material in 7q11.23. We did not observe the same RNA expression outlier profile in the mother who was assumed to have the same phenotype. Analysis of the copy-number data revealed a tandem duplication of the 7q11.23 region inherited from the father while identifying a 28.3kb deletion involving exons 10–14 of ANKRD11 associated with KBG syndrome. ANKRD11 is well expressed in the lymphoblastoid cells but it failed to reach statistical significance to be identified as an expression outlier (Fig. 7). We used transcriptome to guide and inform the DNA analysis, and our results did not produce a convincing argument for an RNA-first approach to Mendelian diagnostics. The clinical and diagnostic utility of transcriptome data is limited to the genes expressed in the tissue available for study and often a challenge in clinical lab settings. The substantial, tissue-dependent variability in mRNA expression profiles and isoform diversity presents a challenge in using them broad manner. In this context, our prior research has demonstrated that lymphoblastoid cell lines (LCLs) recapitulate key transcriptomic features of brain tissue for many neurodevelopmental genes. Notably, LCLs express these genes at a 1.8-fold higher level than whole blood and mirror the brain's isoform diversity for a large proportion of them. There has been reports that utilize fibroblast-to-neuron cell differentiation methods to investigate genes that are not expressed in a commonly utilized tissue such as whole blood. However, the practical utility of such protocols in the routine, standard of care clinical workflows and wider adoption is yet to be seen. Finally, RNA sequencing provides a functional readout, and a definitive diagnosis still requires identification of the causal DNA variant. The diagnostic yield of this cohort stands at 57.8% with the remaining 73 individuals undiagnosed after the re-analysis. We hypothesize that there may be several reasons including pathogenic or likely pathogenic variants 1) that are mosaic in specific tissue and not present in tested tissue, 2) in unrecognized disease genes, 3) not prioritized or filtered out by the current workflow, and 4) not tractable by the short-read sequencing data. Tissue-limited mosaicism constitutes a significant and often underrecognized etiological factor for many dominant, primarily de novo, genetic conditions. Prior studies have estimated that up to 20% of probands with CdLS may have mosaicism that is not detectable in blood. Although investigations are constrained by the practical accessibility of various tissues, the high potential for mosaicism warrants a revised diagnostic approach. Consequently, probands who present with a classic phenotype but yield negative results from genomic testing on blood or buccal samples should be prioritized for subsequent analysis of DNA from alternative sources, such as cultured fibroblasts from a skin biopsy or saliva. A gene-centric and exome approach to variant interpretation is also a major limitation in the current workflow. There may be variants affecting regulatory regions outside protein-coding genes and not recognized as related to the phenotype. Further analysis focusing on noncoding regions of the genome and third generation long-read sequencing may provide some additional answers. Our results underscore the need for the whole genome (and transcriptome) sequencing and a comprehensive, unbiased analytical protocol including structural and noncoding variants to maximally mine a phenotypically homogeneous and genetically heterogeneous cohort, to increase the diagnostic yield. Declarations Funding This work was supported by the National Institutes of Health Grants (RO1, PPG, XO1/Gabriella Miller Kids First, and grants from the National CdLS Foundation). Rajagopalan R, Jadhav T, Blair J, Szot K, and Conlin L received salary support from the NHGRI Early-Stage Investigator grant R01-HG013355. Data availability The datasets used in this article are available in the Gabriella Miller Kids First: Cornelia de Lange Syndrome cohort accessible at https://www.ncbi.nlm.nih.gov/projects/gap/cgi-bin/study.cgi?study_id=phs002174.v1.p1. Author Contribution RR designed the study, performed data analysis and the interpretation, and wrote the manuscript. LKC and IDK interpreted results and edited the manuscript. SR and IDK performed clinical correlation and interpretation of results. TJ, JB, and TJ ran bioinformatics workflows and helped with the interpretation of results. MK and KS performed all the wet lab experiments and helped interpreting the results. All authors read and approved the final manuscript. Acknowledgements We are exceptionally grateful to the individuals and families with Cornelia de Lange Syndrome who participated in this study, as well as to the referring physicians and colleagues including Salim Aftimos, Eric Haan, Maria Giovannucci Uzielli, Fred Gilbert, Elizabeth Loy, and others who have contributed samples and clinical information. We are indebted to the continued support of the National (USA) and the International Cornelia de Lange Syndrome Foundations. We are also deeply indebted to the contributions over many years of the late Dr. Laird Jackson as well as the continued support of the endowed CdLS and Related Diagnoses Multispecialty Center and the Rare Diagnoses Program in the Roberts Individualized Medical Genetics Center (RIMGC) at CHOP. Competing Interest Declaration All authors declare no competing interests or conflicts of interest. Author information Perelman School of Medicine, University of Pennsylvania, Philadelphia, PA. Division of Genomic Diagnostics, Children’s Hospital of Philadelphia, Philadelphia, PA. Ramakrishnan Rajagopalan, Laura Conlin Division of Genomic Diagnostics, Children’s Hospital of Philadelphia, Philadelphia, PA. Tanaya Jadhav, Tejas Jammihal, Justin Blair, Maninder Kaur, Kaitlyn Szot Department of Pediatrics, Children’s Hospital of Philadelphia, Philadelphia, PA. Sarah Raible Division of Medical Genetics, Cohen Children’s Medical Center/ Northwell Health and the Department of Pediatrics, Zucker School of Medicine, Hofstra University, Great Neck, NY. Ian Krantz. References Burdick KJ, Cogan JD, Rives LC, Robertson AK, Koziura ME, Brokamp E, Duncan L, Hannig V, Pfotenhauer J, Vanzo R et al : Limitations of exome sequencing in detecting rare and undiagnosed diseases. Am J Med Genet A 2020, 182(6):1400–1406. Riess O, Sturm M, Menden B, Liebmann A, Demidov G, Witt D, Casadei N, Admard J, Schutz L, Ossowski S et al : Genomes in clinical care. NPJ Genom Med 2024, 9(1):20. Marwaha S, Knowles JW, Ashley EA: A guide for the diagnosis of rare and undiagnosed disease: beyond the exome. Genome Med 2022, 14(1):23. Wojcik MH, Reuter CM, Marwaha S, Mahmoud M, Duyzend MH, Barseghyan H, Yuan B, Boone PM, Groopman EE, Delot EC et al : Beyond the exome: What's next in diagnostic testing for Mendelian conditions. Am J Hum Genet 2023, 110(8):1229–1248. Rajagopalan R, Gilbert MA, McEldrew DA, Nassur JA, Loomes KM, Piccoli DA, Krantz ID, Conlin LK, Spinner NB: Genome sequencing increases diagnostic yield in clinically diagnosed Alagille syndrome patients with previously negative test results. Genet Med 2021, 23(2):323–330. Welland MJ, Ahlquist KD, De Fazio P, Austin-Tse C, Pais L, Wedd L, Bryen S, Rius R, Franklin M, Morrison C et al : Scalable automated reanalysis of genomic data in research and clinical rare disease cohorts. medRxiv 2025. Ansari M, Halachev M, Parry D, Campos JL, D'Souza EN, Barnett C, Wilkie AOM, Barnicoat A, Patel CV, Sukarova-Angelovska E et al : Whole Genome Sequencing of "Mutation-Negative" Individuals With Cornelia de Lange Syndrome. Hum Mutat 2025, 2025:4711663. Kaur M, Blair J, Devkota B, Fortunato S, Clark D, Lawrence A, Kim J, Do W, Semeo B, Katz O et al : Genomic analyses in Cornelia de Lange Syndrome and related diagnoses: Novel candidate genes, genotype-phenotype correlations and common mechanisms. Am J Med Genet A 2023, 191(8):2113–2131. Cornelia de Lange Syndrome [ https://www.ncbi.nlm.nih.gov/books/NBK1104/ ] Olley G, Ansari M, Bengani H, Grimes GR, Rhodes J, von Kriegsheim A, Blatnik A, Stewart FJ, Wakeling E, Carroll N et al : BRD4 interacts with NIPBL and BRD4 is mutated in a Cornelia de Lange-like syndrome. Nat Genet 2018, 50(3):329–332. Woods SA, Robinson HB, Kohler LJ, Agamanolis D, Sterbenz G, Khalifa M: Exome sequencing identifies a novel EP300 frame shift mutation in a patient with features that overlap Cornelia de Lange syndrome. Am J Med Genet A 2014, 164A(1):251–258. Parenti I, Teresa-Rodrigo ME, Pozojevic J, Ruiz Gil S, Bader I, Braunholz D, Bramswig NC, Gervasini C, Larizza L, Pfeiffer L et al : Mutations in chromatin regulators functionally link Cornelia de Lange syndrome and clinically overlapping phenotypes. Hum Genet 2017, 136(3):307–320. Izumi K, Nakato R, Zhang Z, Edmondson AC, Noon S, Dulik MC, Rajagopalan R, Venditti CP, Gripp K, Samanich J et al : Germline gain-of-function mutations in AFF4 cause a developmental syndrome functionally linking the super elongation complex and cohesin. Nat Genet 2015, 47(4):338–344. Pedersen BS, Bhetariya PJ, Brown J, Kravitz SN, Marth G, Jensen RL, Bronner MP, Underhill HR, Quinlan AR: Somalier: rapid relatedness estimation for cancer and germline studies using efficient genome sketches. Genome Med 2020, 12(1):62. Suvakov M, Panda A, Diesh C, Holmes I, Abyzov A: CNVpytor: a tool for copy number variation detection and analysis from read depth and allele imbalance in whole-genome sequencing. Gigascience 2021, 10(11). Chen X, Schulz-Trieglaff O, Shaw R, Barnes B, Schlesinger F, Kallberg M, Cox AJ, Kruglyak S, Saunders CT: Manta: rapid detection of structural variants and indels for germline and cancer sequencing applications. Bioinformatics 2016, 32(8):1220–1222. Smedley D, Jacobsen JO, Jager M, Kohler S, Holtgrewe M, Schubach M, Siragusa E, Zemojtel T, Buske OJ, Washington NL et al : Next-generation diagnostics and disease-gene discovery with the Exomiser. Nat Protoc 2015, 10(12):2004–2015. Gudmundsson S, Singer-Berk M, Watts NA, Phu W, Goodrich JK, Solomonson M, Genome Aggregation Database C, Rehm HL, MacArthur DG, O'Donnell-Luria A: Variant interpretation using population databases: Lessons from gnomAD. Hum Mutat 2022, 43(8):1012–1030. Ioannidis NM, Rothstein JH, Pejaver V, Middha S, McDonnell SK, Baheti S, Musolf A, Li Q, Holzinger E, Karyadi D et al : REVEL: An Ensemble Method for Predicting the Pathogenicity of Rare Missense Variants. Am J Hum Genet 2016, 99(4):877–885. Jaganathan K, Kyriazopoulou Panagiotopoulou S, McRae JF, Darbandi SF, Knowles D, Li YI, Kosmicki JA, Arbelaez J, Cui W, Schwartz GB et al : Predicting Splicing from Primary Sequence with Deep Learning. Cell 2019, 176(3):535–548 e524. Landrum MJ, Lee JM, Benson M, Brown GR, Chao C, Chitipiralla S, Gu B, Hart J, Hoffman D, Jang W et al : ClinVar: improving access to variant interpretations and supporting evidence. Nucleic Acids Res 2018, 46(D1):D1062-D1067. Brechtmann F, Mertes C, Matuseviciute A, Yepez VA, Avsec Z, Herzog M, Bader DM, Prokisch H, Gagneur J: OUTRIDER: A Statistical Method for Detecting Aberrantly Expressed Genes in RNA Sequencing Data. Am J Hum Genet 2018, 103(6):907–917. Li YI, Knowles DA, Humphrey J, Barbeira AN, Dickinson SP, Im HK, Pritchard JK: Annotation-free quantification of RNA splicing using LeafCutter. Nat Genet 2018, 50(1):151–158. Coursimault J, Rovelet-Lecrux A, Cassinari K, Brischoux-Boucher E, Saugier-Veber P, Goldenberg A, Lecoquierre F, Drouot N, Richard AC, Vera G et al : uORF-introducing variants in the 5'UTR of the NIPBL gene as a cause of Cornelia de Lange syndrome. Hum Mutat 2022, 43(9):1239–1248. Chen Y, Chen Q, Yuan K, Zhu J, Fang Y, Yan Q, Wang C: A Novel de Novo Variant in 5' UTR of the NIPBL Associated with Cornelia de Lange Syndrome. Genes (Basel) 2022, 13(5). Additional Declarations No competing interests reported. Supplementary Files RajagopalanetalCdLSmultiomicsforNPJGenomicMedicineTable1.xlsx Cite Share Download PDF Status: Under Revision Version 1 posted Editorial decision: Revision requested 07 Apr, 2026 Reviews received at journal 04 Apr, 2026 Reviewers agreed at journal 31 Mar, 2026 Reviews received at journal 30 Mar, 2026 Reviewers agreed at journal 30 Mar, 2026 Reviewers invited by journal 30 Mar, 2026 Editor assigned by journal 29 Mar, 2026 Submission checks completed at journal 10 Mar, 2026 First submitted to journal 09 Mar, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9078185","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":604123869,"identity":"374965c0-80da-4e90-a1eb-7d68941b3ff1","order_by":0,"name":"Ramakrishnan Rajagopalan","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA80lEQVRIiWNgGAWjYHCCBDDJByI+MBwAcySI0sIGxIwziNTCANfCzEOMFvP2Aw8/fqm4I8/GfvbYZ5uaO3LmDMwHb/Pg0SJzJiFZWubMM8M2nrzk2TnHnhlbNrAlW+PTIsGQkCAt2XaYsY0hx5g5t+Fw4oYDPGbSeLXwP0j+DdRi38b/xpjZEqyF/xt+LRIJaZIf2w4ntkkAbWGE2MJGQMuDNGuGM8+S2yTeGDP2HDtsbNnMZmw5B6/DcpJv/qi4Y9vPn2PM8KPmsJw5e/PDG2/waGFg4EmARwcYGDDjVQ4C7AcYf6BoIahjFIyCUTAKRhoAADgRTSqPJWxXAAAAAElFTkSuQmCC","orcid":"","institution":"Children's Hospital of Philadelphia","correspondingAuthor":true,"prefix":"","firstName":"Ramakrishnan","middleName":"","lastName":"Rajagopalan","suffix":""},{"id":604123879,"identity":"d5f5f473-fe57-4e88-abf1-af41c74b0fb8","order_by":1,"name":"Tejas Jammihal","email":"","orcid":"","institution":"Children's Hospital of Philadelphia","correspondingAuthor":false,"prefix":"","firstName":"Tejas","middleName":"","lastName":"Jammihal","suffix":""},{"id":604123881,"identity":"5c3ebe9d-4e90-42aa-bcac-32ad6fba307d","order_by":2,"name":"Tanaya Jadhav","email":"","orcid":"","institution":"Children's Hospital of Philadelphia","correspondingAuthor":false,"prefix":"","firstName":"Tanaya","middleName":"","lastName":"Jadhav","suffix":""},{"id":604123890,"identity":"2e1c1017-ecc6-43db-ad8a-c2d0a4d1caee","order_by":3,"name":"Maninder Kaur","email":"","orcid":"","institution":"Children's Hospital of Philadelphia","correspondingAuthor":false,"prefix":"","firstName":"Maninder","middleName":"","lastName":"Kaur","suffix":""},{"id":604123893,"identity":"211dc1cf-043e-4d7c-866f-992e2094f860","order_by":4,"name":"Kaitlyn Szot","email":"","orcid":"","institution":"Children's Hospital of Philadelphia","correspondingAuthor":false,"prefix":"","firstName":"Kaitlyn","middleName":"","lastName":"Szot","suffix":""},{"id":604123894,"identity":"00f5851e-d325-448f-8396-997d5a757bed","order_by":5,"name":"Justin Blair","email":"","orcid":"","institution":"Children's Hospital of Philadelphia","correspondingAuthor":false,"prefix":"","firstName":"Justin","middleName":"","lastName":"Blair","suffix":""},{"id":604123898,"identity":"05088196-6b23-456d-8781-20baf9e02a80","order_by":6,"name":"Sarah Raible","email":"","orcid":"","institution":"University of Pennsylvania","correspondingAuthor":false,"prefix":"","firstName":"Sarah","middleName":"","lastName":"Raible","suffix":""},{"id":604123900,"identity":"0cc330d4-abb7-49d4-a730-1ab117841991","order_by":7,"name":"Laura Conlin","email":"","orcid":"","institution":"Children's Hospital of Philadelphia","correspondingAuthor":false,"prefix":"","firstName":"Laura","middleName":"","lastName":"Conlin","suffix":""},{"id":604123905,"identity":"908e0aa5-d195-44a2-b1c8-b097eda09ecc","order_by":8,"name":"Ian Krantz","email":"","orcid":"","institution":"Cohen Children's Medical Center","correspondingAuthor":false,"prefix":"","firstName":"Ian","middleName":"","lastName":"Krantz","suffix":""}],"badges":[],"createdAt":"2026-03-10 02:38:29","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9078185/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9078185/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":105186893,"identity":"9bacff0b-52df-47fa-82e6-7e5e6a2e8083","added_by":"auto","created_at":"2026-03-23 08:42:20","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":81853,"visible":true,"origin":"","legend":"\u003cp\u003eConsort diagram showing the patient population in this paper. Exclusion criteria and results of the study are indicated.\u003c/p\u003e","description":"","filename":"1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-9078185/v1/2bdbaf7defd7036880fe8ff0.jpg"},{"id":105187011,"identity":"4d691c9e-77d7-4cb0-b747-98501de5d65f","added_by":"auto","created_at":"2026-03-23 08:42:40","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":59129,"visible":true,"origin":"","legend":"\u003cp\u003eFlow diagram visualizing the classification of genetic diagnoses, stratified by variant type, for both CdLS and non-CdLS genes.\u003c/p\u003e","description":"","filename":"2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-9078185/v1/7af407c4047b7382acc70392.jpg"},{"id":105186952,"identity":"886ce7a9-6f72-4823-b896-eb7f7cbe9309","added_by":"auto","created_at":"2026-03-23 08:42:24","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":95322,"visible":true,"origin":"","legend":"\u003cp\u003eA de novo, deep intronic variant c.4560+1975G\u0026gt;A with a spliceAI score of 0.08 introducing a novel exon in the NIPBL. The table shows spliceAI and Pangolin scores and the schematic in the right shows the position of the novel exon (top right) and the sashimi plot from RNAseq data (bottom right).\u003c/p\u003e","description":"","filename":"3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-9078185/v1/0b11699443fadb20abe36238.jpg"},{"id":105186968,"identity":"129eadbb-a7d4-42bd-a7aa-76d104ca5c88","added_by":"auto","created_at":"2026-03-23 08:42:29","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":77890,"visible":true,"origin":"","legend":"\u003cp\u003eExon skipping event in the proband CDL-053. Outlier expression analysis using OUTRIDER identified aberrant, decreased expression in SMC3. Inspection of the RNAseq data revealed an exon skipping event and the short-read genome showed a maternally inherited c.1409+6T\u0026gt;C in SMC3.\u003c/p\u003e","description":"","filename":"4.jpg","url":"https://assets-eu.researchsquare.com/files/rs-9078185/v1/4973b214010461fa707fe5dd.jpg"},{"id":105186953,"identity":"e5eed4a8-418f-4b3c-af72-952087a9e4f3","added_by":"auto","created_at":"2026-03-23 08:42:24","extension":"jpg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":185846,"visible":true,"origin":"","legend":"\u003cp\u003eIGV screenshot showing a de novo insertion in intron 33 of NIPBL in patient CDL-200. The coverage pileup indicates a copy-number difference, and the discordant read pattern is consistent with insertion of nonspecific sequence, such as repetitive elements, with mates mapping to multiple chromosomes (colored reads).\u003c/p\u003e","description":"","filename":"5.jpg","url":"https://assets-eu.researchsquare.com/files/rs-9078185/v1/6fdbea51dbe13bddccf3fd34.jpg"},{"id":105186880,"identity":"cea9e19f-8666-4043-b61a-6b357daa07c3","added_by":"auto","created_at":"2026-03-23 08:42:19","extension":"jpg","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":45589,"visible":true,"origin":"","legend":"\u003cp\u003eddPCR analysis of NIPBL exon 41 expression in CDL-200, indicating markedly reduced transcript levels consistent with loss of mRNA resulting from the insertion.\u003c/p\u003e","description":"","filename":"6.jpg","url":"https://assets-eu.researchsquare.com/files/rs-9078185/v1/0e9b36e4955636b2f86c7ec0.jpg"},{"id":105187076,"identity":"e990454f-de50-4e39-ad18-8d961ed889d6","added_by":"auto","created_at":"2026-03-23 08:43:05","extension":"jpg","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":237958,"visible":true,"origin":"","legend":"\u003cp\u003eOutlier expression analysis in an affected mother-daughter duo showed different expression profiles. The proband showed overexpression of three genes in the 7q Williams syndrome region including BAZ1B and two genes with reduced expression. The mother’s data showed completely different set of genes dysregulated. The copy-number analysis in this family revealed the 28Kb deletion in ANKRD11. Panel C shows the IGV screenshot showing the abnormal read pairs (red) spanning the length of the deletion.\u003c/p\u003e","description":"","filename":"7.jpg","url":"https://assets-eu.researchsquare.com/files/rs-9078185/v1/c7dbda7fdfe23027785a928a.jpg"},{"id":105187233,"identity":"16153ea3-d193-4769-919e-7ec4e5e1685c","added_by":"auto","created_at":"2026-03-23 08:43:35","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1359728,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9078185/v1/545f9027-8a1e-40e4-9fa4-af7b84bc0676.pdf"},{"id":105186969,"identity":"c5aedafe-48a3-42f7-b08c-028de4d87128","added_by":"auto","created_at":"2026-03-23 08:42:29","extension":"xlsx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":15670,"visible":true,"origin":"","legend":"","description":"","filename":"RajagopalanetalCdLSmultiomicsforNPJGenomicMedicineTable1.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-9078185/v1/74de7dfac8bedc0d5dd679e3.xlsx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Multi-omic re-analysis increases diagnostic yield in individuals with Cornelia de Lange syndrome","fulltext":[{"header":"Introduction","content":"\u003cp\u003eStandard exonic, sequence variant-focused analytical approaches, which currently form the standard of care in molecular diagnostics, present significant limitations[\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. These approaches predominantly identify protein-coding variants while often overlooking noncoding, regulatory, and structural variants that may contribute to Cornelia de Lange Syndrome (CdLS) pathogenesis. While the standard of care sequencing has moved from capture-based exome sequencing to whole genome sequencing, the analytical workflows are still focused on exonic regions of the genome and interpretable/ reportable variants (exome slice on a genome backbone). Variants in enhancer or promoter regions, deep intronic variants that affect splicing, or structural variants such as copy number alterations and chromosomal rearrangements can elude traditional exome sequencing or exome-focused analytical approaches[\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan additionalcitationids=\"CR4\" citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. As a result, individuals with clinically suspected CdLS but without identifiable exonic pathogenic variants may remain undiagnosed, emphasizing the need for broader genome-wide analyses, including whole-genome sequencing (WGS) and transcriptomic profiling, to capture the full spectrum of pathogenic variants. Routine re-analysis of genomic data has also proven valuable in increasing diagnostic yield, as newly discovered pathogenic variants and updated classification criteria can help identify previously unrecognized disease-causing mutations[\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. As bioinformatic tools advance and variant interpretation frameworks evolve, periodic re-evaluation of existing genomic data can significantly improve diagnostic accuracy and patient outcomes.\u003c/p\u003e \u003cp\u003eCornelia de Lange Syndrome (CdLS; MIM# 122470, 300590, 610759, 614701, 300882, 620568) is a genetically heterogeneous and clinically variable, multi-system diagnosis characterized by characteristic facial features, microcephaly, limb malformations, prenatal and postnatal growth retardation, global developmental delays and various other systemic involvement (e.g. diaphragmatic hernia, congenital heart differences, cleft palate, hypertichosis, intestinal malrotation and others)[\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. CdLS presents with a highly variable phenotype often complicating early clinical diagnosis, and the severity and spectrum of manifestations can differ significantly among affected individuals, necessitating a comprehensive diagnostic approach. Molecular confirmation of CdLS is achieved through genetic testing, which has identified six primary causative genes: \u003cem\u003eNIPBL, SMC1A, SMC3, RAD21, HDAC8, and BRD4\u003c/em\u003e. Among these, variants in \u003cem\u003eNIPBL\u003c/em\u003e account for most cases (~\u0026thinsp;80%), while variants in \u003cem\u003eSMC1A, SMC3, RAD21, HDAC8, and BRD4\u003c/em\u003e collectively contribute to 10\u0026ndash;15% of cases and manifest varying phenotypic expression of the diagnosis[\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. These genes encode key components (structural and regulatory) of the cohesin complex, which plays a crucial role in chromosomal segregation, DNA repair, and gene expression regulation. Despite advances in genetic testing, the diagnostic landscape of CdLS remains challenging due to phenotypic overlap with several other syndromes, including KBG syndrome (\u003cem\u003eANKRD11)\u003c/em\u003e, Coffin-Siris Syndrome (\u003cem\u003eARID1B)\u003c/em\u003e, Rubinstein-Taybi Syndrome (\u003cem\u003eCREBBP, EP300)\u003c/em\u003e, and Wiedemann-Steiner Syndrome (\u003cem\u003eKMT2A)\u003c/em\u003e, among others[\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e, \u003cspan additionalcitationids=\"CR11 CR12\" citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. These phenocopy syndromes exhibit clinical similarities to CdLS but arise from distinct genetic etiologies, underscoring the need for comprehensive, genome-wide, and, at times, multi-omic diagnostic approaches.\u003c/p\u003e \u003cp\u003eWe recently reported the mutational landscape of CdLS and related diagnoses in a large cohort of 716 individuals. A likely molecular etiology was identified in 59% of the probands tested (n\u0026thinsp;=\u0026thinsp;422 out of 716) using standard molecular diagnostic approaches, including targeted panels, exome sequencing, and exome slice on a genome platform (GS cohort n\u0026thinsp;=\u0026thinsp;177). In the total cohort of 716 probands clinically diagnosed with CdLS, 358 (50%) had variants in known CdLS-associated genes, including \u003cem\u003eNIPBL\u003c/em\u003e (271 individuals, 64%), \u003cem\u003eSMC1A\u003c/em\u003e (40, 9%), \u003cem\u003eHDAC8\u003c/em\u003e (25, 6%), \u003cem\u003eSMC3\u003c/em\u003e (16, 4%), and \u003cem\u003eRAD21\u003c/em\u003e (6, 1%). Additionally, 64 probands (9%) had variants in 32 different genes that phenocopied CdLS. Notably, in the GS cohort described by Kaur et al., a molecular diagnosis was identified in 37.5% of individuals, while the rest of the cohort remained without a molecular diagnosis[\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. We hypothesized that the individuals with a clinical diagnosis of CdLS but no molecular diagnosis from prior analysis harbor pathogenic, disease-causing variants that could not be identified or prioritized by the standard of care exome-focused workflows. In this work, we performed a comprehensive re-analysis of the genome sequencing data using an expanded workflow to include noncoding, copy-number, and structural variants, along with the transcriptome data for a subset of probands. By addressing the limitations of the exome-focused workflows and utilizing emerging methodologies, this paper aims to explore the evolving landscape of genomic diagnostics of rare Mendelian diagnoses, emphasizing the need for comprehensive analytical workflows and integration of advanced genomic technologies including the role of multi-omics analyses. We demonstrate that a comprehensive analytical workflow including noncoding and structural variants significantly increases the diagnostic yield in a Mendelian diagnosis such as CdLS.\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003ePatient cohort\u003c/h2\u003e \u003cp\u003eWe started with a cohort of 177 individuals with genome sequencing data described in Kaur et. al. Four samples were removed from this cohort including three with failed familial relationships and one being a duplicate resulting in 173 individuals with genome sequencing data. This cohort included 88 trios, 29 duos, and 56 singletons. These individuals went through an exome slice/panel analysis focused on genes known to cause CdLS and related neurodevelopmental diagnoses. Sixty-five individuals had a previously identified molecular diagnosis reported in the Kaur et. al., leaving 108 individuals without a molecular diagnosis. A subset of 54 individuals without a molecular diagnosis (in whom cell lines were available to isolate RNA) from this cohort had whole transcriptome data (Fig.\u0026nbsp;1).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eEthics statement\u003c/h3\u003e\n\u003cp\u003eAll probands and family members were enrolled and consented to participate in a research study approved by the Institutional Review Board at CHOP.\u003c/p\u003e\n\u003ch3\u003eData production and analysis\u003c/h3\u003e\n\u003cp\u003eGenome sequencing (GS) was performed at the BROAD institute on the Illumina Novaseq X platform at a minimum average read depth of 30x and details of the prior analytical workflow is described elsewhere[\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. Raw sequencing data (fastq) were aligned to the human genome reference hg38 using DRAGEN v3.9. The quality control protocol included sex check as identified by the DRAGEN pipeline (Illumina Inc. San Diego, CA) and the kinship coefficients estimated by \u003cem\u003esomalier\u003c/em\u003e[\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. Copy number variants identified using the read depth-based algorithm, cnvpytor[\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e], were clustered along with breakends identified by the split-read algorithm, \u003cem\u003emanta\u003c/em\u003e[\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e], to identify dispersed duplications. Variant prioritization was performed by multiple computational workflows following a variant filtration heuristic including a phenotype-agnostic, weighted sum model using several variant features (an internal tool) and phenotype-based variant prioritization algorithm Exomiser[\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]. Whole genome VCF was filtered to retain high quality (site quality QUAL\u0026thinsp;\u0026gt;\u0026thinsp;20), rare variants (gnomAD v.3.1 AF\u0026thinsp;\u0026lt;\u0026thinsp;0.5%) before getting passed to the variant prioritization algorithms[\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. The homegrown, phenotype agnostic algorithm employed a linear weighted sum model using inheritance information when family members were available, variant allele frequency from gnomAD v3.1, relevant computational scores for missense (REVEL, missense Z)[\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e], loss of function intolerance for the gene (LOEUF), splice prediction for intronic variants (spliceAI)[\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e], conservation score (GERP 3.0), and the presence in known mutation databases (HGMD, ClinVar)[\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eRNA sequencing was performed at the BROAD institute at a minimum depth of 100M reads per sample and DRAGEN v3.9 RNA-seq workflow was used to align the raw reads to the human transcriptome and estimate the abundance of transcripts defined in ENCODE V39. The RNA outlier analysis was performed using the R package OUTRIDER [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e] and the aberrant splicing outliers were identified using leafcutter[\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e]. DNA sequencing data was investigated for variants when there was an aberrant expression or splicing event detected in the RNA analysis. All the other downstream analyses and visualizations were produced using in-house scripts written in Python/ R. Variants were classified by a clinical laboratory geneticist according to the ACMG guidelines.\u003c/p\u003e\n\u003ch3\u003eVariant confirmation\u003c/h3\u003e\n\u003cp\u003eNewly identified variants that required orthogonal confirmation were determined on a case-by-case basis. All variants were manually inspected in the CRAM alignments file for their presence and the strength of evidence. A subset of newly identified variants was confirmed using Sanger sequencing (SNV/ INDEL), droplet-digital PCR (copy-number variants), and long-range PCR followed by Gel electrophoresis (long insertions).\u003c/p\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eSummary\u003c/h2\u003e \u003cp\u003eKaur et. al. reported a 37.5% diagnostic yield in a genome sequencing cohort of 173 patients with CdLS and related diagnoses. The re-analysis workflow described in this paper was able to identify all the previously reported diagnoses. Expanding the genome analysis workflow to include intronic noncoding variants and structural variants and integrating transcriptome sequencing as a complementary assay in a subset of probands revealed molecular etiologies in an additional 37 new diangoses. Notably, this reanalysis identified additional diagnoses in two probands previously reported as positives in Kaur et. al. which increased the total diagnostic yield in the GS cohort to 57.8% (100 of 173).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe new diagnoses from re-analysis were enriched for variants beyond standard exonic SNV/ indels. Cryptic non-coding variants (promoter, deep intronic), copy number variation including mosaic events, and chromosomal rearrangements including inversions and insertional translocations made up 26 of the 37 new diagnoses (70.3%) (Fig.\u0026nbsp;2). Four probands in the cohort had two likely causative variants identified (4/173; 2.3%). Seventeen of the 37 new diagnoses (46%) were in the previously known CdLS genes which included 12 noncoding sequence variants and five structural variants. Thirteen of the 37 new diagnoses were identified \u003cem\u003ede novo\u003c/em\u003e where parental samples available, 17 of unknown inheritance, and the rest inherited from a parent. A parallel analysis of whole transcriptome data in 54 individuals helped guide the analysis of genome sequencing data and help uncover at least two of the noncoding, cryptic splice variants unnoticed by the primary DNA analysis.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eStructural variants\u003c/h3\u003e\n\u003cp\u003eThe structural variants identified included two copy-neutral inversions, 9 deletions including two large (\u0026gt;\u0026thinsp;5Mb) deletions, 2 duplications including one large (\u0026gt;\u0026thinsp;5Mb) duplication and an intergenic duplication, and 2 long (\u0026gt;\u0026thinsp;1kb) insertions (Table\u0026nbsp;1). The smallest structural variant identified was an 84bp deletion in the last of exon of the gene \u003cem\u003eEP300\u003c/em\u003e known to cause Menke-Hennekam Syndrome 2 (OMIM #618333) and the largest was a 24Mb deletion in 3q which included multiple haploinsufficient genes \u003cem\u003eSTAG1, MED12L, ZIC1, and FOXL2\u003c/em\u003e.\u003c/p\u003e \u003cp\u003eFive of the structural variants involved a known CdLS gene. This included a single exon, 3.2kb deletion in \u003cem\u003eRAD21\u003c/em\u003e, a 350kb deletion involving \u003cem\u003eHDAC8\u003c/em\u003e, a 2.8Mb inversion disrupting \u003cem\u003eNIPBL\u003c/em\u003e with a breakpoint in intron 6, a long insertion in the intron 33 of \u003cem\u003eNIPBL\u003c/em\u003e, and a 252kb dispersed duplication from chr4 inserted into \u003cem\u003eHDAC8.\u003c/em\u003e\u003c/p\u003e \u003cp\u003eThere were nine other probands with a structural variant identified in genes known to cause other developmental diagnoses. The three large copy-number variants included two deletions (9.3Mb in 1q and 24Mb in 3q) involving multiple syndromic genes, and a 7.6Mb duplication including a triplosensitive gene \u003cem\u003eSETD1B.\u003c/em\u003e Other events included a 571kb inversion involving \u003cem\u003eZEB2\u003c/em\u003e (Mowat-Wilson syndrome) and deletions involving \u003cem\u003ePHF6\u003c/em\u003e (Borjeson-Forssman-Lehmann syndrome) and \u003cem\u003eHPRT1\u003c/em\u003e (Lesch-Nyhan syndrome), \u003cem\u003eANKRD11\u003c/em\u003e (KBG syndrome), and \u003cem\u003eBCL11A\u003c/em\u003e (Dias-Logan syndrome). One proband had a deletion in \u003cem\u003eHNRNPD\u003c/em\u003e (HNRNP-related neurodevelopmental disorders) and an intragenic duplication involving exons 5\u0026ndash;6 of \u003cem\u003eSETBP1\u003c/em\u003e (Intellectual developmental disorder, autosomal dominant 29).\u003c/p\u003e \u003cp\u003eFinally, a \u003cem\u003ede novo\u003c/em\u003e, 20.5kb deletion involving exons 7 to 15 in \u003cem\u003eEPC1\u003c/em\u003e, a potential candidate gene, was identified in a single proband. EPC1 (Enhancer of Polycomb Homolog 1), one of the polycomb group of proteins, is not associated with a Mendelian phenotype but predicted to be constrained for loss of function variants in common population (pLI\u0026thinsp;=\u0026thinsp;1 and LOEUF 0.23). Cohesin functionally interacts with the polycomb group of proteins responsible for epigenetic silencing of proteins during development making it a strong candidate gene for further exploration.\u003c/p\u003e\n\u003ch3\u003eNoncoding variants\u003c/h3\u003e\n\u003cp\u003eThere were 12 noncoding, intronic splice or cryptic splice, and promoter variants. Eleven of these variants were in established CdLS genes and the remaining one was a pair of obligatory splice variants in \u003cem\u003eMETTL5.\u003c/em\u003e There were 3 patients with promoter variants in \u003cem\u003eNIPBL\u003c/em\u003e (two c.-467C\u0026thinsp;\u0026gt;\u0026thinsp;T and one c.-324delC). These promoter variants in \u003cem\u003eNIPBL\u003c/em\u003e (GRCh38 5:36,876,445CG\u0026thinsp;\u0026gt;\u0026thinsp;G; c.-324delC, chr5:36876791C\u0026thinsp;\u0026gt;\u0026thinsp;T; c.467C\u0026thinsp;\u0026gt;\u0026thinsp;T) were predicted to disrupt the Kozak sequence in the 5\u0026rsquo; UTR of \u003cem\u003eNIPBL\u003c/em\u003e. The c.-467C\u0026thinsp;\u0026gt;\u0026thinsp;T has been previously identified as \u003cem\u003ede novo\u003c/em\u003e in an individual with CdLS and shown to produce reduced mRNA expression[\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e, \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]. While c.-324delC has not been seen before, similar variants creating a frameshift in the uORF have been reported in patients with CdLS[\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e]. All but one cryptic intronic splice variant had a SpliceAI score\u0026thinsp;\u0026gt;\u0026thinsp;0.2. The c.4560\u0026thinsp;+\u0026thinsp;1975G\u0026thinsp;\u0026gt;\u0026thinsp;A in \u003cem\u003eNIPBL\u003c/em\u003e had a SpliceAI score of 0.08.\u003c/p\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eOther exonic sequence variants\u003c/h2\u003e \u003cp\u003eThere were 11 exonic sequence variants (SNVs/ indels) rescued by the reanalysis that were not previously described in Kaur et. al. None of the 11 variants were in the known CdLS genes. These included apparently loss of function, stop gain variants in \u003cem\u003eWWOX\u003c/em\u003e (homozygous p.Arg264*), a frameshift in \u003cem\u003eZMYM2\u003c/em\u003e (p.thr188Asnfs*12), a start loss in \u003cem\u003eTAF1\u003c/em\u003e (c.1A\u0026thinsp;\u0026gt;\u0026thinsp;C), and several previously published pathogenic variants in ClinVar (\u003cem\u003eCDC42\u003c/em\u003e p.Ala159Val (\u003cem\u003ede novo)\u003c/em\u003e and p.Asp170Gly, \u003cem\u003eNAA10\u003c/em\u003e p.Arg83Cys (\u003cem\u003ede novo)\u003c/em\u003e, \u003cem\u003eSETD5\u003c/em\u003e p.Val305Gly). The inheritance information for all these variants is shown in Table\u0026nbsp;1.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eTranscriptome analysis\u003c/h2\u003e \u003cp\u003eThe transcriptome analysis resolved two novel diagnoses overlooked by the DNA workflow. First one (CDL-111 and CDL-378) involved a deep intronic variant in \u003cem\u003eNIPBL\u003c/em\u003e (\u003cem\u003ec.4560\u0026thinsp;+\u0026thinsp;1975 G\u0026thinsp;\u0026gt;\u0026thinsp;A)\u003c/em\u003e with a very low spliceAI score (\u0026lt;\u0026thinsp;0.2) yet resulting in an inclusion of a novel exon in the intron 21 of \u003cem\u003eNIPBL\u003c/em\u003e (Fig.\u0026nbsp;3). Outlier splicing analysis using leafcutter revealed a novel junction involving exon 21, and further manual inspection suggested a novel exon inclusion between exons 21 and 22.\u003c/p\u003e\u003cp\u003eAnother proband (CDL-053) had a strong variant of uncertain significance, p.(Cys194Arg), in the gene \u003cem\u003eEBF3\u003c/em\u003e known to cause Hypotonia, ataxia, and delayed development syndrome (OMIM #617330). This variant is well conserved in a protein domain and adjacent to the previously published p.Lys193Asn variant. However, expression outlier analysis using OUTRIDER revealed \u003cem\u003eSMC3\u003c/em\u003e as an outlier with reduced expression. Further inspection of DNA sequencing data revealed a maternally inherited intronic variant in \u003cem\u003eSMC3\u003c/em\u003e (GRCh38 chr10:110589714T\u0026thinsp;\u0026gt;\u0026thinsp;C; NM_005445.3:c.1409\u0026thinsp;+\u0026thinsp;6T\u0026thinsp;\u0026gt;\u0026thinsp;C) and the variant was not prioritized for analysis (Fig.\u0026nbsp;4).\u003c/p\u003e \u003cp\u003eThe genome re-analysis identified a \u003cem\u003ede novo\u003c/em\u003e insertion in intron 32 of NIPBL in CDL-200 that could not be fully resolved because it involved repetitive sequences and read pairs whose mates mapped to multiple locations across the genome (Fig.\u0026nbsp;5). RNA sequencing indicated an exon 33 skipping event supported by five reads; however, this signal could not be reliably distinguished from background noise. Further ddPCR expression assay (probe in exon 33) showed reduced expression compared to parental samples (Fig.\u0026nbsp;6).\u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eIn this study, we performed a comprehensive re-analysis of genome sequencing data from 173 probands including structural and noncoding variants. The re-analysis yielded 37 new diagnoses with an enrichment for noncoding and structural variants and increased the overall diagnostic yield to 57.8%.\u003c/p\u003e \u003cp\u003eThe proportion of the noncoding and structural variants identified in the re-analysis (26 of the 37; 70.3%) highlights the limitations of an exome focused analysis approach and the opportunities for narrowing the diagnostic gap in Mendelian diagnoses. Structural variant callers for short-read sequencing data typically identify several thousands of events including a large proportion of breakends enriched for false positives originating from repetitive or low complexity regions of the genome making it hard to prioritize true and clinically relevant structural rearrangements. Insertional translocations are not typically recognized from short-read sequencing data and need specialized workflows. One of the insertional translocations identified in this study disrupted \u003cem\u003eHDAC8\u003c/em\u003e, a known gene for CdLS, but had a read-pair signature consistent with a balanced translocation. However, further manual review of the alignments revealed that there were two subgroups of reads, each mapping to proximal and distal ends of a 252kb duplication in chr4. The two inversions identified in this study were much smaller than 10Mb in length (2.8Mb and 571Kb) and cannot be detected by conventional Karyotyping methods. While whole genome sequencing has enabled the detection of these events, they are not routinely detected due to the challenges in differentiating true events from likely false positives. The two inversions reported in our study did not involve repetitive regions and had breakpoint spanning reads which helped recognize them as true rearrangements. The insertion in the \u003cem\u003eNIPBL in\u003c/em\u003e CDL-200 involved a repeat element that cannot be characterized by short-read sequencing and the deletion involving the \u003cem\u003eBCL11A\u003c/em\u003e in CDL-637 involved additional sequences at both ends suggesting a repeat element insertion at the deletion site. These complex rearrangements are often mediated by repeat or segmental duplication mediated events and the ambiguity in the read-pair signatures arising from mapping issues hamper interpretation. Long read sequencing can alleviate this challenge when there are breakpoint spanning reads that are anchored in a unique region. However, long read sequencing may encounter similar challenges if the region involved is longer than the length of the reads. The other structural variant findings including simple deletions, duplications, and cytogenetic events may speak to the disparities in the prior analytical workflows and testing protocols.\u003c/p\u003e \u003cp\u003eThe use of transcriptome data to guide the re-analysis was helpful in both identifying previously unrecognized cryptic splice variants and confirming the aberrant splicing consequences of cryptic splice variants with computational evidence. The RNA analysis for CDL-053 using OUTRIDER revealed \u003cem\u003eSMC3\u003c/em\u003e as an outlier for that patient and the RNA alignments showed novel junctions which prompted us to identify a maternally inherited c.1409\u0026thinsp;+\u0026thinsp;6T\u0026thinsp;\u0026gt;\u0026thinsp;C variant in \u003cem\u003eSMC3.\u003c/em\u003e This variant was not prioritized by the variant prioritization algorithm Exomiser likely because it was inherited from an unaffected mother. Similarly, a \u003cem\u003ede novo\u003c/em\u003e intronic variant in \u003cem\u003eNIPBL\u003c/em\u003e (CDL-378) was previously unrecognized due to the lack of computational evidence for aberrant splicing (SpliceAI score\u0026thinsp;=\u0026thinsp;0.08). Splicing outlier analysis suggested a novel junction in intron 21 of \u003cem\u003eNIPBL\u003c/em\u003e which prompted us to recognize this \u003cem\u003ede novo\u003c/em\u003e intronic variant as clinically relevant.\u003c/p\u003e \u003cp\u003eWhile the transcriptome data was helpful in guiding the DNA analysis, it could obscure the real diagnosis when there are multiple events resulting in aberrant expression. For example, in a family with proband (CDL-315) and an affected mother, the RNA expression outlier analysis revealed a different set of genes as outliers (Fig.\u0026nbsp;5). The proband showed several genes as upregulated outliers (\u003cem\u003eBAZ1B, TBL2, RFC2)\u003c/em\u003e pointing to the 7q11.23 region associated with Williams-Buren syndrome typically resulting from the loss of material in 7q11.23. We did not observe the same RNA expression outlier profile in the mother who was assumed to have the same phenotype. Analysis of the copy-number data revealed a tandem duplication of the 7q11.23 region inherited from the father while identifying a 28.3kb deletion involving exons 10\u0026ndash;14 of \u003cem\u003eANKRD11\u003c/em\u003e associated with KBG syndrome. \u003cem\u003eANKRD11\u003c/em\u003e is well expressed in the lymphoblastoid cells but it failed to reach statistical significance to be identified as an expression outlier (Fig.\u0026nbsp;7).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eWe used transcriptome to guide and inform the DNA analysis, and our results did not produce a convincing argument for an RNA-first approach to Mendelian diagnostics. The clinical and diagnostic utility of transcriptome data is limited to the genes expressed in the tissue available for study and often a challenge in clinical lab settings. The substantial, tissue-dependent variability in mRNA expression profiles and isoform diversity presents a challenge in using them broad manner. In this context, our prior research has demonstrated that lymphoblastoid cell lines (LCLs) recapitulate key transcriptomic features of brain tissue for many neurodevelopmental genes. Notably, LCLs express these genes at a 1.8-fold higher level than whole blood and mirror the brain's isoform diversity for a large proportion of them. There has been reports that utilize fibroblast-to-neuron cell differentiation methods to investigate genes that are not expressed in a commonly utilized tissue such as whole blood. However, the practical utility of such protocols in the routine, standard of care clinical workflows and wider adoption is yet to be seen. Finally, RNA sequencing provides a functional readout, and a definitive diagnosis still requires identification of the causal DNA variant.\u003c/p\u003e \u003cp\u003eThe diagnostic yield of this cohort stands at 57.8% with the remaining 73 individuals undiagnosed after the re-analysis. We hypothesize that there may be several reasons including pathogenic or likely pathogenic variants 1) that are mosaic in specific tissue and not present in tested tissue, 2) in unrecognized disease genes, 3) not prioritized or filtered out by the current workflow, and 4) not tractable by the short-read sequencing data. Tissue-limited mosaicism constitutes a significant and often underrecognized etiological factor for many dominant, primarily de novo, genetic conditions. Prior studies have estimated that up to 20% of probands with CdLS may have mosaicism that is not detectable in blood. Although investigations are constrained by the practical accessibility of various tissues, the high potential for mosaicism warrants a revised diagnostic approach. Consequently, probands who present with a classic phenotype but yield negative results from genomic testing on blood or buccal samples should be prioritized for subsequent analysis of DNA from alternative sources, such as cultured fibroblasts from a skin biopsy or saliva. A gene-centric and exome approach to variant interpretation is also a major limitation in the current workflow. There may be variants affecting regulatory regions outside protein-coding genes and not recognized as related to the phenotype. Further analysis focusing on noncoding regions of the genome and third generation long-read sequencing may provide some additional answers.\u003c/p\u003e \u003cp\u003eOur results underscore the need for the whole genome (and transcriptome) sequencing and a comprehensive, unbiased analytical protocol including structural and noncoding variants to maximally mine a phenotypically homogeneous and genetically heterogeneous cohort, to increase the diagnostic yield.\u003c/p\u003e "},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis work was supported by the National Institutes of Health Grants (RO1, PPG, XO1/Gabriella Miller Kids First, and grants from the National CdLS Foundation).\u0026nbsp;Rajagopalan R, Jadhav T, Blair J, Szot K, and Conlin L received salary support from the NHGRI Early-Stage Investigator grant R01-HG013355.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe datasets used in this article are available in the Gabriella Miller Kids First: Cornelia de Lange Syndrome cohort accessible at https://www.ncbi.nlm.nih.gov/projects/gap/cgi-bin/study.cgi?study_id=phs002174.v1.p1.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor Contribution\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eRR designed the study, performed data analysis and the interpretation, and wrote the manuscript. LKC and IDK interpreted results and edited the manuscript. SR and IDK performed clinical correlation and interpretation of results. TJ, JB, and TJ ran bioinformatics workflows and helped with the interpretation of results. MK and KS performed all the wet lab experiments and helped interpreting the results. All authors read and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe are exceptionally grateful to the individuals and families with Cornelia de Lange Syndrome who participated in this study, as well as to the referring physicians and colleagues including Salim Aftimos, Eric Haan, Maria Giovannucci Uzielli, Fred Gilbert, Elizabeth Loy, and others who have contributed samples and clinical information. We are indebted to the continued support of the National (USA) and the International Cornelia de Lange Syndrome Foundations. We are also deeply indebted to the contributions over many years of the late Dr. Laird Jackson as well as the continued support of the endowed CdLS and Related Diagnoses Multispecialty Center and the Rare Diagnoses Program in the Roberts Individualized Medical Genetics Center (RIMGC) at CHOP.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting Interest Declaration\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll authors declare no competing interests or conflicts of interest.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor information\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003ePerelman School of Medicine, University of Pennsylvania, Philadelphia, PA.\u003c/p\u003e\n\u003cp\u003eDivision of Genomic Diagnostics, Children\u0026rsquo;s Hospital of Philadelphia, Philadelphia, PA.\u003c/p\u003e\n\u003cp\u003eRamakrishnan Rajagopalan, Laura Conlin\u003c/p\u003e\n\u003cp\u003eDivision of Genomic Diagnostics, Children\u0026rsquo;s Hospital of Philadelphia, Philadelphia, PA.\u003c/p\u003e\n\u003cp\u003eTanaya Jadhav, Tejas Jammihal, Justin Blair, Maninder Kaur, Kaitlyn Szot\u003c/p\u003e\n\u003cp\u003eDepartment of Pediatrics, Children\u0026rsquo;s Hospital of Philadelphia, Philadelphia, PA.\u003c/p\u003e\n\u003cp\u003eSarah Raible\u003c/p\u003e\n\u003cp\u003eDivision of Medical Genetics, Cohen Children\u0026rsquo;s Medical Center/ Northwell Health and the Department of Pediatrics, Zucker School of Medicine, Hofstra University, Great Neck, NY.\u003c/p\u003e\n\u003cp\u003eIan Krantz.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eBurdick KJ, Cogan JD, Rives LC, Robertson AK, Koziura ME, Brokamp E, Duncan L, Hannig V, Pfotenhauer J, Vanzo R \u003cem\u003eet al\u003c/em\u003e: Limitations of exome sequencing in detecting rare and undiagnosed diseases. \u003cem\u003eAm J Med Genet A\u003c/em\u003e 2020, 182(6):1400\u0026ndash;1406.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRiess O, Sturm M, Menden B, Liebmann A, Demidov G, Witt D, Casadei N, Admard J, Schutz L, Ossowski S \u003cem\u003eet al\u003c/em\u003e: Genomes in clinical care. \u003cem\u003eNPJ Genom Med\u003c/em\u003e 2024, 9(1):20.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMarwaha S, Knowles JW, Ashley EA: A guide for the diagnosis of rare and undiagnosed disease: beyond the exome. \u003cem\u003eGenome Med\u003c/em\u003e 2022, 14(1):23.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWojcik MH, Reuter CM, Marwaha S, Mahmoud M, Duyzend MH, Barseghyan H, Yuan B, Boone PM, Groopman EE, Delot EC \u003cem\u003eet al\u003c/em\u003e: Beyond the exome: What's next in diagnostic testing for Mendelian conditions. \u003cem\u003eAm J Hum Genet\u003c/em\u003e 2023, 110(8):1229\u0026ndash;1248.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRajagopalan R, Gilbert MA, McEldrew DA, Nassur JA, Loomes KM, Piccoli DA, Krantz ID, Conlin LK, Spinner NB: Genome sequencing increases diagnostic yield in clinically diagnosed Alagille syndrome patients with previously negative test results. \u003cem\u003eGenet Med\u003c/em\u003e 2021, 23(2):323\u0026ndash;330.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWelland MJ, Ahlquist KD, De Fazio P, Austin-Tse C, Pais L, Wedd L, Bryen S, Rius R, Franklin M, Morrison C \u003cem\u003eet al\u003c/em\u003e: Scalable automated reanalysis of genomic data in research and clinical rare disease cohorts. \u003cem\u003emedRxiv\u003c/em\u003e 2025.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAnsari M, Halachev M, Parry D, Campos JL, D'Souza EN, Barnett C, Wilkie AOM, Barnicoat A, Patel CV, Sukarova-Angelovska E \u003cem\u003eet al\u003c/em\u003e: Whole Genome Sequencing of \"Mutation-Negative\" Individuals With Cornelia de Lange Syndrome. \u003cem\u003eHum Mutat\u003c/em\u003e 2025, 2025:4711663.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKaur M, Blair J, Devkota B, Fortunato S, Clark D, Lawrence A, Kim J, Do W, Semeo B, Katz O \u003cem\u003eet al\u003c/em\u003e: Genomic analyses in Cornelia de Lange Syndrome and related diagnoses: Novel candidate genes, genotype-phenotype correlations and common mechanisms. \u003cem\u003eAm J Med Genet A\u003c/em\u003e 2023, 191(8):2113\u0026ndash;2131.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCornelia de Lange Syndrome [\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ncbi.nlm.nih.gov/books/NBK1104/\u003c/span\u003e\u003cspan address=\"https://www.ncbi.nlm.nih.gov/books/NBK1104/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e]\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOlley G, Ansari M, Bengani H, Grimes GR, Rhodes J, von Kriegsheim A, Blatnik A, Stewart FJ, Wakeling E, Carroll N \u003cem\u003eet al\u003c/em\u003e: BRD4 interacts with NIPBL and BRD4 is mutated in a Cornelia de Lange-like syndrome. \u003cem\u003eNat Genet\u003c/em\u003e 2018, 50(3):329\u0026ndash;332.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWoods SA, Robinson HB, Kohler LJ, Agamanolis D, Sterbenz G, Khalifa M: Exome sequencing identifies a novel EP300 frame shift mutation in a patient with features that overlap Cornelia de Lange syndrome. \u003cem\u003eAm J Med Genet A\u003c/em\u003e 2014, 164A(1):251\u0026ndash;258.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eParenti I, Teresa-Rodrigo ME, Pozojevic J, Ruiz Gil S, Bader I, Braunholz D, Bramswig NC, Gervasini C, Larizza L, Pfeiffer L \u003cem\u003eet al\u003c/em\u003e: Mutations in chromatin regulators functionally link Cornelia de Lange syndrome and clinically overlapping phenotypes. \u003cem\u003eHum Genet\u003c/em\u003e 2017, 136(3):307\u0026ndash;320.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eIzumi K, Nakato R, Zhang Z, Edmondson AC, Noon S, Dulik MC, Rajagopalan R, Venditti CP, Gripp K, Samanich J \u003cem\u003eet al\u003c/em\u003e: Germline gain-of-function mutations in AFF4 cause a developmental syndrome functionally linking the super elongation complex and cohesin. \u003cem\u003eNat Genet\u003c/em\u003e 2015, 47(4):338\u0026ndash;344.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePedersen BS, Bhetariya PJ, Brown J, Kravitz SN, Marth G, Jensen RL, Bronner MP, Underhill HR, Quinlan AR: Somalier: rapid relatedness estimation for cancer and germline studies using efficient genome sketches. \u003cem\u003eGenome Med\u003c/em\u003e 2020, 12(1):62.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSuvakov M, Panda A, Diesh C, Holmes I, Abyzov A: CNVpytor: a tool for copy number variation detection and analysis from read depth and allele imbalance in whole-genome sequencing. \u003cem\u003eGigascience\u003c/em\u003e 2021, 10(11).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen X, Schulz-Trieglaff O, Shaw R, Barnes B, Schlesinger F, Kallberg M, Cox AJ, Kruglyak S, Saunders CT: Manta: rapid detection of structural variants and indels for germline and cancer sequencing applications. \u003cem\u003eBioinformatics\u003c/em\u003e 2016, 32(8):1220\u0026ndash;1222.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSmedley D, Jacobsen JO, Jager M, Kohler S, Holtgrewe M, Schubach M, Siragusa E, Zemojtel T, Buske OJ, Washington NL \u003cem\u003eet al\u003c/em\u003e: Next-generation diagnostics and disease-gene discovery with the Exomiser. \u003cem\u003eNat Protoc\u003c/em\u003e 2015, 10(12):2004\u0026ndash;2015.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGudmundsson S, Singer-Berk M, Watts NA, Phu W, Goodrich JK, Solomonson M, Genome Aggregation Database C, Rehm HL, MacArthur DG, O'Donnell-Luria A: Variant interpretation using population databases: Lessons from gnomAD. \u003cem\u003eHum Mutat\u003c/em\u003e 2022, 43(8):1012\u0026ndash;1030.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eIoannidis NM, Rothstein JH, Pejaver V, Middha S, McDonnell SK, Baheti S, Musolf A, Li Q, Holzinger E, Karyadi D \u003cem\u003eet al\u003c/em\u003e: REVEL: An Ensemble Method for Predicting the Pathogenicity of Rare Missense Variants. \u003cem\u003eAm J Hum Genet\u003c/em\u003e 2016, 99(4):877\u0026ndash;885.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJaganathan K, Kyriazopoulou Panagiotopoulou S, McRae JF, Darbandi SF, Knowles D, Li YI, Kosmicki JA, Arbelaez J, Cui W, Schwartz GB \u003cem\u003eet al\u003c/em\u003e: Predicting Splicing from Primary Sequence with Deep Learning. \u003cem\u003eCell\u003c/em\u003e 2019, 176(3):535\u0026ndash;548 e524.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLandrum MJ, Lee JM, Benson M, Brown GR, Chao C, Chitipiralla S, Gu B, Hart J, Hoffman D, Jang W \u003cem\u003eet al\u003c/em\u003e: ClinVar: improving access to variant interpretations and supporting evidence. \u003cem\u003eNucleic Acids Res\u003c/em\u003e 2018, 46(D1):D1062-D1067.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBrechtmann F, Mertes C, Matuseviciute A, Yepez VA, Avsec Z, Herzog M, Bader DM, Prokisch H, Gagneur J: OUTRIDER: A Statistical Method for Detecting Aberrantly Expressed Genes in RNA Sequencing Data. \u003cem\u003eAm J Hum Genet\u003c/em\u003e 2018, 103(6):907\u0026ndash;917.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi YI, Knowles DA, Humphrey J, Barbeira AN, Dickinson SP, Im HK, Pritchard JK: Annotation-free quantification of RNA splicing using LeafCutter. \u003cem\u003eNat Genet\u003c/em\u003e 2018, 50(1):151\u0026ndash;158.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCoursimault J, Rovelet-Lecrux A, Cassinari K, Brischoux-Boucher E, Saugier-Veber P, Goldenberg A, Lecoquierre F, Drouot N, Richard AC, Vera G \u003cem\u003eet al\u003c/em\u003e: uORF-introducing variants in the 5'UTR of the NIPBL gene as a cause of Cornelia de Lange syndrome. \u003cem\u003eHum Mutat\u003c/em\u003e 2022, 43(9):1239\u0026ndash;1248.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen Y, Chen Q, Yuan K, Zhu J, Fang Y, Yan Q, Wang C: A Novel de Novo Variant in 5' UTR of the NIPBL Associated with Cornelia de Lange Syndrome. \u003cem\u003eGenes (Basel)\u003c/em\u003e 2022, 13(5).\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"npj-genomic-medicine","isNatureJournal":false,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"npjgenmed","sideBox":"Learn more about [npj Genomic Medicine](http://www.nature.com/npjgenmed/)","snPcode":"41525","submissionUrl":"https://mts-npjgenmed.nature.com/","title":"npj Genomic Medicine","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"NPJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-9078185/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9078185/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eBackground\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eApproximately half of the individuals with a clinically diagnosed Mendelian condition do not receive a molecular diagnosis. Current standard-of-care diagnostic pipelines, which are largely focused on exonic sequence variants, may not be comprehensive enough to identify all pathogenic variants. A comprehensive analytical approach capable of identifying noncoding and structural variants is needed to bridge the diagnostic gap. Cornelia de Lange Syndrome (CdLS) is a multisystem developmental diagnosis caused primarily by pathogenic variants in one of the six genes known to cause CdLS \u003cem\u003e(NIPBL, SMC3, SMC1A, HDAC8, RAD21, BRD4)\u003c/em\u003e, although pathogenic variants in additional phenocopy genes have also been implicated. We hypothesized that individuals with a clinical diagnosis of CdLS and no molecular diagnosis harbor pathogenic, causative variants that are not identified or prioritized by the current standard of care, exome-focused workflows.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMethods\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe performed a re-analysis of the genome sequencing data from a previously published cohort of 173 individuals with a clinical diagnosis of CdLS (Gabriella Miller Kids First cohort) and expanded the scope of analysis to include noncoding and structural variants. We used RNA-sequencing data in a subset of individuals (n = 62) to complement the DNA workflow.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eResults\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eRe-analysis, including copy-number and structural variation and using transcriptome sequencing as a complementary assay, revealed molecular etiologies in an additional 37 probands. Thus, increasing the total diagnostic yield \u003cu\u003ein this previously undiagnosed cohort\u003c/u\u003e to 60%. The new diagnoses were enriched for variants beyond the standard exonic SNVs/ indels, including cryptic non-coding variants (promoter, deep intronic, large insertions), copy number variants, balanced rearrangements such as inversions, and variants in additional genes that phenocopy CdLS. Transcriptome aided re-analysis helped uncover cryptic noncoding variants in the DNA that lacked sufficient computational evidence for a splicing abnormality and yet produced aberrantly spliced mRNA.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConclusions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eOur results underscore the need for whole genome (and transcriptome) sequencing and a comprehensive, unbiased analytical protocol integrating structural and noncoding variants to exhaustively mine a phenotypically and genetically heterogeneous cohort to maximize its diagnostic yield. The additional diagnostic yield solely from noncoding and structural variants highlights the limitations of an exome-focused analysis workflow and highlights the utility of transcriptome analysis beyond the use of splicing prediction tools.\u003c/p\u003e","manuscriptTitle":"Multi-omic re-analysis increases diagnostic yield in individuals with Cornelia de Lange syndrome","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-03-23 08:38:37","doi":"10.21203/rs.3.rs-9078185/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2026-04-08T03:49:24+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-04-04T04:25:57+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"255493249664767439863078905622998202951","date":"2026-03-31T08:32:06+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-03-30T14:34:05+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"18261733638084574227355634652571242305","date":"2026-03-30T08:10:41+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-03-30T07:49:12+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-03-29T04:05:54+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-03-11T02:29:24+00:00","index":"","fulltext":""},{"type":"submitted","content":"npj Genomic Medicine","date":"2026-03-10T02:26:05+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"npj-genomic-medicine","isNatureJournal":false,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"npjgenmed","sideBox":"Learn more about [npj Genomic Medicine](http://www.nature.com/npjgenmed/)","snPcode":"41525","submissionUrl":"https://mts-npjgenmed.nature.com/","title":"npj Genomic Medicine","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"NPJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"8eb5966a-422c-4b12-acb0-432006fb27b2","owner":[],"postedDate":"March 23rd, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"in-revision","subjectAreas":[{"id":64286777,"name":"Biological sciences/Computational biology and bioinformatics"},{"id":64286778,"name":"Health sciences/Diseases"},{"id":64286779,"name":"Biological sciences/Genetics"}],"tags":[],"updatedAt":"2026-04-08T03:54:13+00:00","versionOfRecord":[],"versionCreatedAt":"2026-03-23 08:38:37","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9078185","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9078185","identity":"rs-9078185","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-24T02:00:01.246996+00:00
License: CC-BY-4.0