Allelic Prevalence and Geographic Distribution of Cerebrotendinous Xanthomatosis | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Allelic Prevalence and Geographic Distribution of Cerebrotendinous Xanthomatosis Tiziano Pramparo, Robert D. Steiner, Steve Rodems, Celia Jenkinson This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-1942700/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 17 Jan, 2023 Read the published version in Orphanet Journal of Rare Diseases → Version 1 posted 4 You are reading this latest preprint version Abstract Background: Cerebrotendinous xanthomatosis (CTX) is a rare recessive genetic disease characterized by disruption of bile acid synthesis due to inactivation of the CYP27A1 gene. Treatment is available in the form of bile acid replacement. CTX is likely underdiagnosed, and prevalence estimates based on case diagnosis are probably inaccurate. Large population-based genomic databases are a valuable resource to estimate prevalence of rare recessive diseases as an orthogonal unbiased approach building upon traditional epidemiological studies. Methods: We leveraged the Hardy-Weinberg principle and allele frequencies from gnomAD to calculate CTX prevalence. ClinVar and HGMD were used to identify high-confidence pathogenic missense variants and to calculate a disease-specific cutoff. Variant pathogenicity was also assessed by the VarSome implementation of the ACMG/AMP algorithm and the REVEL in silico predictor. Results: CTX prevalence estimates were highest in Asians (1:44,407-93,084) and lowest in the Finnish population (1:3,388,767). Intermediate estimates were found in Europeans, Americans, and Africans/African Americans (1:70,795-233,597). The REVEL-predicted pathogenic variants accounted for a greater increase in prevalence estimates for Europeans, Americans, and Africans/African Americans compared with Asians. We identified the most frequent alleles designated pathogenic in ClinVar (p.Gly472Ala, p.Arg395Cys), labeled pathogenic based on sequence consequence (p.Met1?), and predicted to be pathogenic by REVEL (p.Met383Lys, p.Arg448His) across populations. Also, we provide a prospective geographic map of estimated disease distribution based on CYP27A1 variation queries performed by healthcare providers from selected specialties. Conclusions: Prevalence estimates calculated herein support and expand upon existing evidence indicating underdiagnosis of CTX, suggesting that improved detection strategies are needed. Increased awareness of CTX is important for early diagnosis, which is essential for patients as early treatment significantly slows or prevents disease progression. allele frequency rare metabolic disease cerebrotendinous xanthomatosis (CTX) genomic variation pathogenic variants Figures Figure 1 Figure 2 Introduction Cerebrotendinous xanthomatosis (CTX, OMIM 213700) is a rare autosomal recessive condition characterized by disruption of bile acid synthesis due to inactivation of the CYP27A1 gene. Biallelic pathogenic variants are responsible for loss of enzymatic sterol-27-hydroxylase activity leading to reduced production of chenodeoxycholic acid (CDCA) and cholic acid, and accumulation of cholestanol and bile alcohols [ 1 , 2 ]. CTX is highly heterogeneous in clinical presentation and signs and symptoms can overlap with other conditions (eg, sitosterolemia, familial hypercholesterolemia). The first signs and symptoms of the classical presentation are typically non-neurological and may include diarrhea during infancy (sometimes with failure to thrive), juvenile-onset-cataracts, and tendon xanthomas [ 3 , 4 ]. Neurological dysfunction may be present in childhood with intellectual disability and/or autism spectrum disorder and frequently develops in adulthood with epilepsy, pyramidal and extrapyramidal signs, cerebellar syndrome, peripheral neuropathy and intellectual disability/cognitive decline and progressive neurodegeneration [ 4 – 10 ]. Because of the pleiotropic phenotypes, accurate diagnosis may be delayed 10–20 years [ 11 , 12 ] and is often achieved only after irreversible neurological involvement has ensued. Treatment with bile acid replacement in the form of CDCA has been successfully used for decades [ 2 , 13 ] and, alternatively, cholic acid treatment has been reported in a modest number of cases. Clinical studies have provided powerful evidence showing that early recognition and treatment are critical to prevent or ameliorate disease progression and to avoid devastating neurological complications. Improved outcome was reported in two independent studies for patients with treatment initiated before age 24–25 years as compared to those patients who started later [ 14 , 15 ]. These results further emphasize the need to improve early recognition strategies [ 16 ], stimulate newborn screening programs [ 17 – 19 ], and leverage large population genetic data to update and more accurately estimate disease prevalence [ 20 ]. A relatively recent approach of estimating disease prevalence from the causative or “predicted” causative alleles (herein labelled allelic prevalence) may contribute to increased disease awareness and improved surveillance and, ultimately, may stimulate earlier diagnosis [ 21 – 25 ]. This approach has become possible due to the availability of large population databases, such as the Genome Aggregation Database (gnomAD) [ 26 ]. Essentially, this genetic approach is unbiased and orthogonal to clinical studies that instead focus on the recognition, recruitment, and count of the number of confirmed affected individuals. This latter approach may suffer from several sources of bias, likely leading to inaccurate estimates and severe underdiagnosis of the disease. A clear limitation to disease recognition underlying failure to diagnose is the presence of “milder” phenotypes, which includes those CTX patients without neurological involvement [ 27 ]. Despite identification of the same underlying genetic variants observed in severe cases, such individuals may be overlooked and/or misdiagnosed with other, often more common, diseases. The identification and reporting of such patient subpopulations further support the notion that genotype-phenotype correlations have not been established in CTX [ 3 ]. A previously published first estimate of the genetic prevalence of CTX utilized data from the Exome Aggregation Consortium (ExAC) database, which included ~ 60,000 unrelated individuals [ 20 ]. Here we provide an update of genetic prevalence utilizing the much larger gnomAD [ 26 ] and taking advantage of additional resources, including the Human Gene Mutation Database (HGMD®) and ClinVar database as well as the VarSome variant classifier. Our approach to CTX prevalence estimate utilizes the Hardy-Weinberg principle and observed variant allele frequencies implementing a variety of variant inclusion/exclusion criteria and automated and manual genetic variant curation steps in an attempt to provide the most conservative and accurate result. In concordance with previous estimates, we report heterogeneity in prevalence across ancestries and identify the most frequent alleles driving these estimates. Finally, we present a map of worldwide geographic estimate disease distribution using a novel approach based on CYP27A1 variant queries performed by clinicians across six continents. The map and CTX prevalence estimate results indicate that the number of individuals with this treatable disorder worldwide should greatly exceed the few hundred currently reported [ 12 , 28 ]. Methods Identification of high-confidence CYP27A1 pathogenic missense variants for cutoff calculation We defined “high-confidence” (HC) pathogenic, missense variants classified as such by both the ClinVar database ( https://www.ncbi.nlm.nih.gov/clinvar/ ) and the HGMD ( http://www.hgmd.cf.ac.uk/ac/index.php ) as previously adopted [ 29 ]. Only HGMD missense variants with a disease-causing mutation (DM) label that were also present and designated (likely) pathogenic in ClinVar were included in the cutoff calculation (see supplementary material). We reclassified these HC pathogenic variants with the VarSome American College of Medical Genetics and Genomics (ACMG)/Association for Molecular Pathology (AMP)–based algorithm ( https://varsome.com/about/resources/acmg-implementation/ ) and considered only those classified (likely) pathogenic (see supplementary material). Among the HC pathogenic variants, we included the ClinVar “likely pathogenic” (LP) based on the following considerations: (1) according to the ACMG/AMP guidelines, LP correlates to a probability of 90% that a variant will be disease-causing [ 30 ]; (2) DM variants reported in HGMD also included LP according to the VarSome classification and were consistent with ClinVar labels; (3) the difference in interpretation confidence between pathogenic and LP was deemed medically insignificant [ 31 ]; and (4) retaining LP variants in our cutoff calculation did not support the hypothesis of lower confidence in pathogenic variant effects for LP. LP variants for the purposes of this study are included as pathogenic. Identification of CYP27A1 pathogenic variants in gnomAD for prevalence calculation The gnomAD v2.1.1 ( https://gnomad.broadinstitute.org/ ) was accessed on October 12, 2021, to gather CYP27A1 variant information based on the canonical transcript ENST00000258415.4 and reference sequences NM_000784.4 and NP_000775.1. ClinVar clinical significance annotation in gnomAD was based on the October 2, 2021, release. We selected CYP27A1 variants with a pathogenic ClinVar clinical significance label for all available sequence consequences. We manually verified that these variants were assigned pathogenic in the HGMD (DM) and the ClinVar database. In ClinVar, we considered reliable pathogenic submissions, including single submitters, limited to those from clinical genetic/diagnostic laboratories, since variant classification/interpretation was likely performed following ACMG/AMP guidelines [ 32 ]. All remaining frameshift, nonsense, start_lost, stop_gained, and canonical splice variants (2 bases ± splice acceptor/donor site) in gnomAD were considered pathogenic based on their high-impact consequence on the coding sequence ( see Table 1 ), as calculated by VEP ( https://m.ensembl.org/info/genome/variation/prediction/predicted_data.html ) and classified in previous studies [ 25 , 33 ]. To increase stringency, we manually reclassified and verified each pathogenic variant with the VarSome algorithm (see supplementary material). Lastly, we filtered out variants with a gnomAD label of “lc_lof” and variants lacking multiple ClinVar entries from clinical genetic/diagnostic laboratories supporting pathogenicity. Table 1 CYP27A1 gnomAD variants Variant type All N (%) ClinV-P N (%) SeqCon-P N (%) Revel-P N (%) Model N (%) missense 344 (39.8) 13 (3.8) - 28 (8.1) 41 (11.9) intron 232 (26.9) - - - - synonymous 139 (16.1) - - - - splice_region 35 (4.1) - - - - frameshift 31 (3.6) 12 (38.7) 17 (54.8) - 29 (93.5) stop_gained 23 (2.7) 11 (47.8) 10 (43.5) - 21 (91.3) 5_prime_UTR 20 (2.3) - - - - 3_prime_UTR 18 (2.1) - - - - splice_donor 8 (0.9) 5 (62.5) 3 (37.5) - 8 (100) inframe_deletion 4 (0.5) - - - 0 (0. 0) splice_acceptor 4 (0.5) 3 (75.0) 1 (25.0) - 4 (100) start_lost 3 (0.3) - 3 (100) - 3 (100) inframe_insertion 2 (0.2) - - - - stop_retained 1 (0.1) - - - - Total 864 (100) 44 (5.1) 34 (3.9) 28 (3.2) 106 (12.3) Variant count breakdown by type: all gnomAD variants (All), designated pathogenic based on ClinVar clinical significance label (ClinV-P); considered pathogenic based on sequence consequence (SeqCon-P); predicted pathogenic based on REVEL analysis (Revel-P), and final list of variants selected for the prevalence calculation (Model). Identification of CYP27A1 predicted pathogenic variants in gnomAD for prevalence calculation All remaining missense variants in gnomAD (N = 321) underwent bioinformatic analysis to predict the likelihood of pathogenicity. We used the REVEL ensemble in silico predictor, which predicts the pathogenic variant effects of missense based on a combination of scores from 13 individual tools and is considered more reliable in assessing pathogenicity [ 34 ] than the historically popular SIFT and Polyphen-2 tools [ 30 ]. REVEL is the reference algorithm for the ACMG/AMP variant interpretation as part of the standard operating procedure in the SVI-approved expert panel specifications ( https://clinicalgenome.org/working-groups/sequence-variant-interpretation/ ). We established a disease-specific cutoff using the mean REVEL score value calculated from the list of HC pathogenic missense variants defined above (see supplementary material). The REVEL mean value was calculated after removing outliers using the interquartile range formula, as previously adopted [ 20 ]. The calculated score (0.834) was then used as cutoff to filter the remaining 321 missense variants. Only gnomAD missense variants with scores equal to or greater than the calculated REVEL cutoff were predicted pathogenic and retained in the complete model for the prevalence calculation. Lastly, to avoid underestimation of the prevalence calculation, we manually investigated the excluded missense variants and recovered those with ClinVar submissions supporting evidence for pathogenicity and/or those present in HGMD or literature. Literature searches were done in LitVar ( https://www.ncbi.nlm.nih.gov/CBBresearch/Lu/Demo/LitVar/ ), Pubmed ( https://pubmed.ncbi.nlm.nih.gov/ ), and Google Scholar ( https://scholar.google.com/ ). Mutalyzer ( https://mutalyzer.nl/ ) was used to check the syntax and names or abbreviations of the variants according to updated Human Genome Variation Society (HGSV) nomenclature. CTX allelic prevalence estimation AFs of the final list of variants were used to calculate CTX prevalence across six gnomAD populations (AFR = African/African American, AMR = Latino/Admixed American, EUR = Non-Finnish European, FIN = Finnish, SAS = South Asian, EAS = East Asian) using the Hardy-Weinberg principle (p 2 + 2pq + q 2 ) [ 20 , 35 ] under the assumption of mutual independence of the rare variants and full penetrance. Prevalence was calculated as the squared sum of the carrier AF of the mutant alleles (q) with p approximated to 1 [ 20 ]. The 95% confidence interval (CI) was calculated with the binomial formula using the Wilson method for each population’s list of risk alleles, as previously adopted [ 24 , 36 , 37 ]. CTX variants search activity and geographic distribution via VarSome Insights We used the VarSome Insights (VSI) service to gather genomic queries relevant to CYP27A1 gene variation and generate an estimated CTX disease map based on the raw number of unique queries from professionals designated as healthcare providers (HCPs). Original profession labels were reclassified and grouped into fewer labels with the goal of distinguishing clinical/medical queries from research/other type of queries and to resolve word mismatches due to language or spelling differences. We further removed from the analysis HCPs queries that, based on the self-reported profession/position/title, were deemed more likely to be of research or nonclinical nature (eg, Research Scientist). Queries were gathered from January 1, 2021, to November 11, 2021. Results We identified a total of 864 CYP27A1 gene variants, of which the majority (40%) were missense. Forty-four variants were designated ClinVar pathogenic in gnomAD (12 frameshift, 13 missense, 8 canonical splice site, and 11 stop_gained), and 34 variants were considered pathogenic based on sequence consequence (17 frameshift, 3 start_lost, 10 stop_gained, and 4 canonical splice) ( see Table 1 ). We manually verified that the 44 ClinVar variants were assigned pathogenic in HGMD and ClinVar and that all 78 (44 + 34) were classified pathogenic by the VarSome algorithm (see supplementary material). Altogether, pathogenic variants represented 9% of all CYP27A1 variants in gnomAD. We then applied our REVEL bioinformatic workflow and identified 23 missense variants passing the disease-specific cutoff, thus predicted to be pathogenic. None of these 23 variants were predicted to be (likely) benign by the VarSome algorithm (see supplementary material), and the vast majority were absent from the HGMD/literature. Of the 23 REVEL-predicted pathogenic variants, p.Met383Lys (VarID: 2-219678874-T-A, HGSV: c.1148T > A) was found with a higher AF (0.12%) in the AMR population compared with the remaining missense variants. With such AF and without reliable estimates for a parameter required to apply the Max Population AF filter [ 38 ], we opted to calculate two estimates for AMR prevalence: with and without the p.Met383Lys variant. Lastly, we manually searched for missense variants with a REVEL score below the cutoff but with supporting information for pathogenicity. We identified five additional missense variants, bringing the number of predicted pathogenic variants to 28 ( see Table 1 ). The average REVEL score of these recovered variants was 0.78, which was within 0.5 standard deviations below the mean score of the HC pathogenic variants. CTX prevalence estimates The final list of variants for the CTX prevalence calculation included 44 designated ClinVar pathogenic variants, 34 considered pathogenic based on sequence consequence, and 28 (including p.Met383Lys) predicted pathogenic variants by the in silico REVEL analysis (see Table 1 and supplementary material ). Using the gnomAD carrier AF of these 106 variants, we calculated CTX prevalence estimates and pooled AF by population (Table 2 ). We found the highest estimates in the EAS and SAS populations, ranging from 2.25 to 1.07 per 100,000 or from 1 per 44,407 to 1 per 93,084, with a pooled AF ranging from 0.00475 (95% CI: 0.00187–0.01457) to 0.00328 (95% CI: 0.00172–0.00764), respectively. The lowest prevalence estimate was found in the FIN population with 0.03 per 100,000 or 1 per 3,388,767 and a pooled AF 0.00054 (95% CI: 0.00019–0.00169). These estimates were not substantially different from those calculated with only the 78 pathogenic variants excluding the variants only predicted pathogenic by the REVEL analysis ( see Table 2 ). Intermediate prevalence estimates were found in the AMR, AFR, and EUR populations. In AMR without the variant p.Met383Lys (M383K), these estimates were 0.63 per 100,000 or 1 per 157,878 with a pooled AF 0.00252 (95% CI: 0.00113–0.00653), while those with the M383K variant were 1.41 per 100,000 or 1 per 70,795, with a pooled AF 0.00376 (95% CI: 0.00206–0.0082). In AFR, these estimates were 0.6 per 100,000 or 1 per 166,440, with a pooled AF 0.00245 (95% CI: 0.00071–0.00898). In EUR, these estimates were 0.43 per 100,000 or 1 per 233,597, with a pooled AF 0.00207 (95% CI: 0.0008–0.00671). Of note, estimates calculated without the predicted pathogenic missense variants showed very similar prevalence across AFR, AMR, and EUR, ranging from 0.25 to 0.21 per 100,000 or 1 per 393,497 to 482,603 ( see Table 2 ). Figure 1 shows the most frequent alleles in each population. The missense variant G472A (VarID: 2-219679419-G-C, HGVS: c.1415G > C, p.Gly472Ala) and the nonsense variant (VarID: 2-219646907-T-C, HGVS: c.2T > C, p.Met1?) were the most frequent in the EAS and SAS populations. G472A was also recently identified in a dried blood spot sample together with the variant R405Q (VarID: 2-219679132-G-A, HGVS: c.1214G > A, p.Arg405Gln), which we also found in the top EAS alleles [ 19 ]. The missense variant R395C (VarID: 2-219678909-C-T, HGVS: c.1183C > T, p.Arg395Cys) remained the most frequent in the FIN and EUR populations, while in the AMR population was second to the missense M383K (VarID: 2-219678874-T-A, HGSV: c.1148T > A, p.Met383Lys). The missense variant R448H (VarID: 2-219679347-G-A, HGVS: c.1343G > A, p.Arg448His) was the most frequent in the AFR population. Table 2 CTX prevalence estimates across six gnomAD populations Model Estimate AFR EAS FIN EUR AMR* SAS ClinV-P + SeqCon-P per 100,000 0.21 1.41 0.03 0.25 0.21 0.95 1 per 472,468 71,089 3,946,059 393,497 482,603 105,299 ClinV-P + SeqCon-P + Revel-P per 100,000 0.6 2.25 0.03 0.43 1.41/0.63 1.07 1 per 166,440 44,407 3,388,767 233,597 70,795/157,878 93,084 Pooled AF (95% CI) 0.00245 (0.00071–0.00898) 0.00475 (0.00187–0.01457) 0.00054 (0.00019–0.00169) 0.00207 (0.0008–0.00671) 0.00376 (0.00206–0.0082)/0.00252 (0.00113–0.00653) 0.00328 (0.00172–0.00764) *In the model ClinV-P + SeqCon-P + Revel-P, the two prevalence estimates correspond to with/without the M383K variant. AF, allele frequency; AFR, African/African American; AMR, Latino Admixed American; EAS, East Asian; EUR, European non-Finnish; FIN, European Finnish; SAS, South Asian. CTX search activity geographic distribution We leveraged VSI to infer a possible CTX disease geographic distribution estimate based on genomic variant queries through the VarSome search engine. This analysis was performed under the assumption that queries were potentially linked to clinical activity ultimately aiding the diagnosis or care of prospective CTX patients. In other words, it was assumed that HCPs from the selected clinical/medical specialties using VarSome in this manner were gathering CYP27A1 variant pathogenicity information because a patient in their care was likely affected with CTX. Overall, queries were performed largely by clinical and research professionals using the platform to gather information on the CYP27A1 gene and related variants. In 2021, we identified a total of 1,243 queries, of which, 827 were from designated HCPs, representing 67% of all queries. After removing those under research or nonclinical professions, we considered 576 queries as a proxy for CTX clinical activity, of which half (n = 288) were from clinical geneticists and genetic counselors. Analysis of the final queries from HCPs showed that pathogenic, uncertain significance, and benign variants were equally represented (~ 33% each). The most queried variant was by far the missense P384L (NM_000784.4:c.1151C > T), which is consistent with its high population AF and several benign submissions in ClinVar. Among the top pathogenic variants, we found a similar number of queries for c.1184 + 1G > A, R127W (NM_000784.4:c.379C > T) and R395C/S (NM_000784.4:c.1183C > T/A), which are also found among the top alleles across gnomAD populations in our prevalence model (Fig. 1 ). Lastly, we mapped all the VarSome pathogenic HCPs queries (N = 187) to infer a possible geographic distribution of CTX disease and identified Spain as the top country with the highest number of CYP27A1 variant queries, followed by Iran, Italy, USA, and Turkey (Fig. 2 ). Discussion There is scientific consensus that CTX is clinically not well recognized, frequently has delayed diagnosis [ 17 ], and has a prevalence likely underestimated in the general population [ 20 , 39 ], with less than 600 cases reported worldwide [ 12 , 28 ]. In order to facilitate early diagnosis of this disease, we sought to calculate the most updated conservative and accurate estimates for CTX prevalence. We followed the approach of previous studies for estimating how common a monogenic autosomal recessive disease may be [ 20 , 25 , 33 , 40 , 41 ], and we leveraged the availability of a large genetic dataset, gnomAD. We developed a highly curated list of alleles in CYP27A1 that have been reported to be pathogenic and used their bioinformatic characteristics to identify additional missense alleles of CYP27A1 that are predicted to be pathogenic for CTX. We then applied the Hardy-Weinberg principle to estimate CTX risk in global populations. These analyses show that CTX is present in global populations at rates more common than previously appreciated. For the first time, we also attempted to geographically map CTX clinical activity worldwide, using a new emerging tool for genomic variation data sharing. The belief that CTX is an exceedingly rare occurrence is not an uncommon scenario in the field of rare diseases. As for CTX, this is typically due to the intrinsic nature of the disease, which can show pleomorphism in clinical and laboratory features overlapping with other diseases, having a variable clinical course and symptoms onset, and lacking any clear genotype-phenotype correlations [ 3 , 42 ]. The task of establishing an accurate estimate for how common CTX may be in the population becomes even more challenging considering the presence of a subpopulation of patients with a milder phenotype [ 27 ]. Interestingly, such individuals may carry the same variants (eg, p.Arg395Cys) that are found in patients with neurological involvement, suggesting that perhaps additional damaging mechanisms or genetic modifiers other than CYP27A1 loss of function may be at play. Reinforcing this enigmatic picture is the report of a pair of siblings carrying the same pathogenic variants but being at opposite ends of the clinical spectrum, where one sibling developed rare spinal xanthomatosis and the other developed a mild form with minor tendon xanthomas [ 43 ]. In contrast, examples of common clinical manifestation are well documented. Bilateral cataracts are found in about 80% of CTX patients [ 16 , 44 ] and interim analysis of patients recruited on the sole basis for having juvenile-onset idiopathic bilateral cataract shows a molecularly confirmed diagnosis in 1.5%-1.8% of the patients, representing a 500-fold increase in CTX prevalence in this subset of patients [ 45 , 46 ]. Our findings largely support, expand upon, and refine previous genetic estimates based on the ExAC data and more limited variant inclusion [ 20 ]. We identified three levels for CTX prevalence. The highest estimates were found in the Asian populations at 1 per 44,407 − 93,084, which are in line with previous investigations (1:36,072–75,601) [ 20 ]. In both studies, the top variant remained the missense c.1415G > C (p.Gly472Ala), but with a higher carrier AF in gnomAD compared with the ExAC (0.0014 vs 0.0010) [ 20 ]. To our knowledge, the p.Gly472Ala variant has been reported in two patients: one of Asian origin carrying a homozygous mutation [ 5 ] and one in a newborn screening [ 19 ]. In the latter study, one sample was found positive for biochemical CTX biomarkers and compound heterozygous for two pathogenic variants in CYP27A1 (c.1214G > A - p.Arg405Gln; c.1415G > C - p.Gly472Ala). While we do not have confirmation that this sample was from an individual of EAS ancestry, we speculate that this might be the case as p.Gly472Ala is unique to EAS and both alleles show high AF in EAS compared with the other ancestries in gnomAD. In such a scenario, the reported incidence of 1 per 32,000 from Hong et al. [ 19 ], might provide an independent and orthogonal validation of our findings in the EAS population. An intermediate level of CTX prevalence was found in the AMR, AFR, and EUR populations. In AMR, although we obtained very close estimates (1 per 70,795 − 157,878) compared with the ExAC study (1:71,677 − 148,914) [ 20 ], differences in the number, type, and AF of variants may be found. For instance, we identified twice as many alleles, with the top variant p.Met383Lys showing a much higher AF (0.00124 vs 0.00078). We provided two prevalence estimates for AMR since it would be difficult to confidently include/exclude this variant from our analysis. Following ACMG/AMP guidelines, we could attempt to apply the BS1 criteria that is used to classify a variant as “likely benign” when its AF is greater than expected for the disorder [ 30 , 47 , 48 ]. However, reliable estimates for population prevalence and allelic heterogeneity would be needed to use the maximum credible population AF as filter [ 38 ]. A plausible strategy to leverage this filter would be to create a distribution of prevalence values and use two estimates for allelic heterogeneity, for instance, 10% (conservative) and 30% based on the most frequent allele found in CTX patients (c.1183C > T, p.Arg395Cys) [ 27 ]. By doing this exercise, we found that at 10% allelic heterogeneity, p.Met383Lys was retained only with a disease prevalence of 1:10,000–20,000 (penetrance 100%-50%), which is currently not supported by any clinical, epidemiological, or genetic study in the AMR population. At 30% allelic heterogeneity, we found that p.Met383Lys was retained with a disease prevalence of 1:100,000-180,000 (penetrance 100%-50%), which may be currently supported by genetic estimates from the previous CTX study (~ 1:70,000-150,000) [ 20 ]. However, it could be argued that if p.Met383Lys was as frequent as p.Arg395Cys, we would have expected at least a few CTX patients described in the literature carrying this allele. Consistent with the strong literature evidence, we found p.Arg395Cys to be most prevalent in the EUR, FIN, and AMR populations. Also, this variant showed an increase in AF compared with the ExAC frequency (0.00051 vs 0.00017) [ 20 ]. Our estimates instead show a significant increase in prevalence for the AFR and EUR populations, that moved from 1:468,624 to 1:166,440 and from 1:461,358 to 1:233,597, respectively [ 20 ]. Again, the same top drivers were found between the two studies but with an increase in AF (p.Arg448His: 0.00041 vs 0.00031 in AFR; p.Arg395Cys: 0.00038 vs 0.00021 in EUR). As discussed above, such differences are not surprising, especially when considering the sizable difference in the number of individuals investigated between the two databases. In addition, we need to take into account improvements in the performance of in silico predictors and the evolving nature of clinical and functional information, which are both critical to variant classification. Similar to the AF of p.Met383Lys, we have retained in our model the start_lost variant c.2T > C (NM_000784.4:p.?), which was found to be unique to the SAS population. While this type of variant may be considered always pathogenic, as it is expected to produce no protein, the clinical impact in practice can be heterogeneous, and alternative mechanisms for protein translation should be considered. In fact, it is notable to find this variant at the heterozygous state in four siblings of a South African family with a mild (no neurological involvement) CTX phenotype (see family #8) [ 27 ]. It is possible that future functional studies will be able to clarify the role of the c.2T > C allele in the pathogenesis of CTX or perhaps to rule out its involvement with direct implications for our SAS estimate. The lowest level of prevalence was found in the FIN population, which is in line with findings from the previous CTX study [ 20 ]. Lastly, with the purpose of increasing awareness, having a proactive approach to patient identification, and aiding early treatment intervention, we leveraged the VarSome platform to create a clinically relevant geographic map based on recent query activity. We found that most of the variant queries were from clinical geneticists/genetic counselors and that the most searched variants were consistent with our findings from the analysis of the gnomAD AF. We recognize that the proposed geographic map represents only a proxy to a possible clinical distribution of prospective CTX patients and that the difference in counts per country may also depend on or be limited to factors such as access and popularity of the VarSome search engine, country size, accessibility, and costs to genetic testing, as well as socioeconomic or political efforts to boost scientific progress in the fight against rare diseases. Nonetheless, our findings line up with patient reports from around the world (USA, Israel, Italy, Japan, the Netherlands, Belgium, Brazil, Canada, France, Iran, Norway, Tunisia, Spain, China, and Sweden; see CTX at https://www.ncbi.nlm.nih.gov/books/NBK1409/ and https://rarediseases.org/ ). There are limitations and methodological assumptions underlying the approach in this study. We and others have applied the Hardy-Weinberg principle in our genetic risk calculations that assume that heterozygous individuals are not subject to selection, populations are at equilibrium with respect to allele and genotype frequencies, and random mating is observed. Also, the concept of AF and how rare an allele may be is strictly related to the size of the population under investigation; thus, with the availability of larger population-based datasets, estimates will change and become more accurate. Lastly, we recognize that although the adoption of filtering criteria and in silico tools to predict pathogenic missense with unknown or uncertain clinical significance is helpful and commonly used, there is currently no gold standard and these strategies could lead to the over- or underestimation of disease frequency. We attempted to mitigate overestimation by cross-referencing the pathogenic labels across different databases and sources and by applying strict filtering criteria with the goal of providing rather conservative estimates. We mitigated underestimation by manually curating missense variants that were filtered out by our bioinformatic workflow. Additional genetic variation that has not been accounted for in our calculation may be conferred by inframe deletions or insertions, intronic, noncanonical splice sites, and structural variants. Lasty, our approach assumed that all variants in the final model contribute to the risk of disease with 100% penetrance. While we are not aware of reports on CYP27A1 pathogenic variants of reduced penetrance, we cannot exclude their presence. In conclusion, our study, that includes additional variants, new informatics tools, and newer, expanded databases, supports and refines previous estimates for CTX disease risk at the population level and provides a novel prospective geographical map of CTX clinical activity worldwide. We confirm with this larger, more comprehensive study that CTX is more common than current worldwide patient estimates, and we highlight the most common pathogenic variants. We underscore the value of leveraging large and diversified population-based genetic databases to assess risk for inherited diseases. Such efforts, cross-referenced with other large-scale programs, such as newborn screening [ 19 ] and retrospective administrative claims studies [ 49 ], may provide the most accurate strategy to assess disease presence in a population. In turn, this will translate into greater awareness, better recognition, and early treatment intervention, which will directly benefit patients and caretakers. Declarations Ethics approval and consent to participate Not applicable. Consent for publication Not applicable. Availability of data and materials All data generated or analyzed during this study are included in this published article or uploaded as supplementary information and/or available in public open access databases described in the methods section. Competing interests TP, SR, CJ are full-time employees of Travere Therapeutics, Inc., and may have an equity or other financial interest in Travere Therapeutics, Inc. RDS is an employee of PreventionGenetics, an Exact Sciences company. He has equity interest in and has received consulting fees from Acer Therapeutics and PTC Therapeutics. He has received consulting fees from Aeglea BioTherapeutics; Alexion Pharmaceuticals, Inc; Best Doctors; Health Advances LLC; Leadiant, Precision for Value; and Travere Therapeutics, Inc; and honoraria from Medscape/WebMD and The France Foundation. He received no funding to support writing and revising this manuscript. He has received research funding from Alexion Pharmaceuticals, Inc and the Smith Lemli Opitz Foundation. Funding This study was funded by Travere Therapeutics, Inc. Authors’ contributions TP, SR, CJ conceived the study. All authors contributed to the review of the results, the writing and revision of the manuscript. TP collected the data, performed the analyses, and generated the results. Acknowledgments We thank Shawn Hayes, Patricia Bedard, and Karsten Baumgaertel for critical reading of the manuscript and/or the helpful feedback on the approach. We thank Dr. Colleen Burns for helping with the confidence interval calculation. We thank Andy Cosgrove for providing access to the queries raw data for the geographic mapping. We thank the MedVal Scientific Information Services, LLC team for editorial support. References Bjorkhem I, Hansson M. Cerebrotendinous xanthomatosis: an inborn error in bile acid synthesis with defined mutations but still a challenge. Biochem Biophys Res Commun. 2010;396(1):46–9. Berginer VM, Salen G, Shefer S. Long-term treatment of cerebrotendinous xanthomatosis with chenodeoxycholic acid. N Engl J Med. 1984;311(26):1649–52. Nie S, Chen G, Cao X, Zhang Y. Cerebrotendinous xanthomatosis: a comprehensive review of pathogenesis, clinical manifestations, diagnosis, and management. Orphanet J Rare Dis. 2014;9:179. Verrips A, van Engelen BG, Wevers RA, et al. Presence of diarrhea and absence of tendon xanthomas in patients with cerebrotendinous xanthomatosis. Arch Neurol. 2000;57(4):520–4. Verrips A, Hoefsloot LH, Steenbergen GC, et al. Clinical and molecular genetic characteristics of patients with cerebrotendinous xanthomatosis. Brain. 2000;123(Pt 5):908–19. Salen G, Steiner RD. Epidemiology, diagnosis, and treatment of cerebrotendinous xanthomatosis (CTX). J Inherit Metab Dis. 2017;40(6):771–81. Stelten BML, Bonnot O, Huidekoper HH, et al. Autism spectrum disorder: an early and frequent feature in cerebrotendinous xanthomatosis. J Inherit Metab Dis. 2018;41(4):641–6. Yunisova G, Tufekcioglu Z, Dogu O, et al. Patients with Lately Diagnosed Cerebrotendinous Xanthomatosis. Neurodegener Dis. 2019;19(5–6):218–24. Amador MDM, Masingue M, Debs R, et al. Treatment with chenodeoxycholic acid in cerebrotendinous xanthomatosis: clinical, neurophysiological, and quantitative brain structural outcomes. J Inherit Metab Dis. 2018;41(5):799–807. Fraidakis MJ. Psychiatric manifestations in cerebrotendinous xanthomatosis. Transl Psychiatry. 2013;3:e302. DeBarber AE, Duell PB. Update on cerebrotendinous xanthomatosis. Curr Opin Lipidol. 2021;32(2):123–31. Badura-Stronka M, Hirschfeld AS, Winczewska-Wiktor A, et al. First case series of Polish patients with cerebrotendinous xanthomatosis and systematic review of cases from the 21st century. Clin Genet. 2022;101(2):190–207. Salen G, Berginer V, Shore V, et al. Increased concentrations of cholestanol and apolipoprotein B in the cerebrospinal fluid of patients with cerebrotendinous xanthomatosis. Effect of chenodeoxycholic acid. N Engl J Med. 1987;316(20):1233–8. Yahalom G, Tsabari R, Molshatzki N, Ephraty L, Cohen H, Hassin-Baer S. Neurological outcome in cerebrotendinous xanthomatosis treated with chenodeoxycholic acid: early versus late diagnosis. Clin Neuropharmacol. 2013;36(3):78–83. Stelten BML, Huidekoper HH, van de Warrenburg BPC, et al. Long-term treatment effect in cerebrotendinous xanthomatosis depends on age at treatment start. Neurology. 2019;92(2):e83–95. Mignarri A, Gallus GN, Dotti MT, Federico A. A suspicion index for early diagnosis and treatment of cerebrotendinous xanthomatosis. J Inherit Metab Dis. 2014;37(3):421–9. DeBarber AE, Luo J, Star-Weinstock M, et al. A blood test for cerebrotendinous xanthomatosis with potential for disease detection in newborns. J Lipid Res. 2014;55(1):146–54. Vaz FM, Bootsma AH, Kulik W, et al. A newborn screening method for cerebrotendinous xanthomatosis using bile alcohol glucuronides and metabolite ratios. J Lipid Res. 2017;58(5):1002–7. Hong X, Daiker J, Sadilek M, et al. Toward newborn screening of cerebrotendinous xanthomatosis: results of a biomarker research study using 32,000 newborn dried blood spots. Genet Med. 2020;22(10):1606–12. Appadurai V, DeBarber A, Chiang PW, et al. Apparent underdiagnosis of Cerebrotendinous Xanthomatosis revealed by analysis of ~ 60,000 human exomes. Mol Genet Metab. 2015;116(4):298–304. Carter A, Brackley SM, Gao J, Mann JP. The global prevalence and genetic spectrum of lysosomal acid lipase deficiency: A rare condition that mimics NAFLD. J Hepatol. 2019;70(1):142–50. Nappo S, Mannucci L, Novelli G, Sangiuolo F, D'Apice MR, Botta A. Carrier frequency of CFTR variants in the non-Caucasian populations by genome aggregation database (gnomAD)-based analysis. Ann Hum Genet. 2020;84(6):463–8. Borges P, Pasqualim G, Giugliani R, Vairo F, Matte U. Estimated prevalence of mucopolysaccharidoses from population-based exomes and genomes. Orphanet J Rare Dis. 2020;15(1):324. Gao J, Brackley S, Mann JP. The global prevalence of Wilson disease from next-generation sequencing data. Genet Med. 2019;21(5):1155–63. Tan J, Wagner M, Stenton SL, et al. Lifetime risk of autosomal recessive mitochondrial disorders calculated from genetic databases. EBioMedicine. 2020;54:102730. Karczewski KJ, Francioli LC, Tiao G, et al. The mutational constraint spectrum quantified from variation in 141,456 humans. Nature. 2020;581(7809):434–43. Stelten BML, Raal FJ, Marais AD, et al. Cerebrotendinous xanthomatosis without neurological involvement. J Intern Med. 2021;290:1039–47. Stelten BML, Dotti MT, Verrips A, et al. Expert opinion on diagnosing, treating and managing patients with cerebrotendinous xanthomatosis (CTX): a modified Delphi study. Orphanet J Rare Dis. 2021;16(1):353. van Rooij J, Arp P, Broer L, et al. Reduced penetrance of pathogenic ACMG variants in a deeply phenotyped cohort study and evaluation of ClinVar classification over time. Genet Med. 2020;22(11):1812–20. Richards S, Aziz N, Bale S, et al. Standards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology. Genet Med. 2015;17(5):405–24. Harrison SM, Dolinsky JS, Knight Johnson AE, et al. Clinical laboratories collaborate to resolve differences in variant interpretations submitted to ClinVar. Genet Med. 2017;19(10):1096–104. Rehm HL, Berg JS, Brooks LD, et al. ClinGen–the Clinical Genome Resource. N Engl J Med. 2015;372(23):2235–42. Brezavar D, Bonnen PE. Incidence of PKAN determined by bioinformatic and population-based analysis of ~ 140,000 humans. Mol Genet Metab. 2019;128(4):463–9. Ioannidis NM, Rothstein JH, Pejaver V, et al. REVEL: An Ensemble Method for Predicting the Pathogenicity of Rare Missense Variants. Am J Hum Genet. 2016;99(4):877–85. Hardy GH. Mendelian Proportions in a Mixed Population. Science. 1908;28(706):49–50. Minikel EV, Vallabh SM, Lek M, et al. Quantifying prion disease penetrance using large population control cohorts. Sci Transl Med. 2016;8(322):322ra9. Schrodi SJ, DeBarber A, He M, et al. Prevalence estimation for monogenic autosomal recessive diseases using population-based genetic data. Hum Genet. 2015;134(6):659–69. Whiffin N, Minikel E, Walsh R, et al. Using high-resolution variant frequencies to empower clinical genome interpretation. Genet Med. 2017;19(10):1151–8. Raymond GV, Schiffmann R. Cerebrotendinous xanthomatosis: The rare "treatable" disease you never consider. Neurology. 2019;92(2):61–2. Coffey AJ, Durkie M, Hague S, et al. A genetic study of Wilson's disease in the United Kingdom. Brain. 2013;136(Pt 5):1476–87. Kaler SG, Ferreira CR, Yam LS. Estimated birth prevalence of Menkes disease and ATP7A-related disorders based on the Genome Aggregation Database (gnomAD). Mol Genet Metab Rep. 2020;24:100602. Zadori D, Szpisjak L, Madar L, et al. Different phenotypes in identical twins with cerebrotendinous xanthomatosis: case series. Neurol Sci. 2017;38(3):481–3. Guenzel AJ, DeBarber A, Raymond K, Dhamija R. Familial variability of cerebrotendinous xanthomatosis lacking typical biochemical findings. JIMD Rep. 2021;59(1):3–9. Wong JC, Walsh K, Hayden D, Eichler FS. Natural history of neurological abnormalities in cerebrotendinous xanthomatosis. J Inherit Metab Dis. 2018;41(4):647–56. Freedman SF, Brennand C, Chiang J, et al. Prevalence of Cerebrotendinous Xanthomatosis Among Patients Diagnosed With Acquired Juvenile-Onset Idiopathic Bilateral Cataracts. JAMA Ophthalmol. 2019;137(11):1312–6. Atilla H, Coskun T, Elibol B, Kadayifcilar S, Altinel S. Group G-E-IW. Prevalence of cerebrotendinous xanthomatosis in cases with idiopathic bilateral juvenile cataract in ophthalmology clinics in Turkey. J AAPOS. 2021;25:269.e1-269.e6. Tavtigian SV, Greenblatt MS, Harrison SM, et al. Modeling the ACMG/AMP variant classification guidelines as a Bayesian classification framework. Genet Med. 2018;20(9):1054–60. Savige J, Storey H, Watson E, et al. Consensus statement on standards and guidelines for the molecular diagnostics of Alport syndrome: refining the ACMG criteria. Eur J Hum Genet. 2021;29(8):1186–97. Sellos-Moura M, Glavin F, Lapidus D, Evans K, Lew CR, Irwin DE. Prevalence, characteristics, and costs of diagnosed homocystinuria, elevated homocysteine, and phenylketonuria in the United States: a retrospective claims-based comparison. BMC Health Serv Res. 2020;20(1):183. Supplementary Files SupplementaryTables07APR2022.xlsx Cite Share Download PDF Status: Published Journal Publication published 17 Jan, 2023 Read the published version in Orphanet Journal of Rare Diseases → Version 1 posted Reviewers agreed at journal 30 Sep, 2022 Reviewers invited by journal 18 Aug, 2022 Editor assigned by journal 10 Aug, 2022 First submitted to journal 08 Aug, 2022 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-1942700","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":129899453,"identity":"ae506dff-2985-418e-b5bb-c8c56f7567f0","order_by":0,"name":"Tiziano Pramparo","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABEklEQVRIie3Sv0vDQBTA8SeBu+WVrCnF+C9EAqmFgv/Ky3JzoSAOBQPCTbau5r8oFIpjyuG55A8oHFSKOAoWodDBokQRJYlZHe47Pu5zP+AAbLb/mAOQ/Rq4TQJLpC2bSGkS6AZyyvk62932odtRT7PBaOWHD+P1K56vfODqblp9MVqMcwG9iYjMjR6GkeZhB/NhCCjEsppkWUsqCHKIDDKK55pBO5UUJx5G1eQgWbwVhG8N7imeSebs0j1d1BMH1OcpGJnWx+ZTxpi3SYigjigG6lAK7F3hmUknFHpasJMXTcey5i38+v5x8yz7fhf53Ay25LuX2lnSiI5crnQV+T4tgOIn/Ij9sbyoTGw2m8321TuZDF23LotyIAAAAABJRU5ErkJggg==","orcid":"https://orcid.org/0000-0001-6929-1027","institution":"Travere Therapeutics Inc","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Tiziano","middleName":"","lastName":"Pramparo","suffix":""},{"id":129899454,"identity":"ecfb56a3-47b0-4b29-9514-65fe98ac68c5","order_by":1,"name":"Robert D. Steiner","email":"","orcid":"","institution":"UNIVERSITY OF WISCONSIN SCHOOL OF MEDICINE AND PUBLIC HEALTH","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Robert","middleName":"D.","lastName":"Steiner","suffix":""},{"id":129899455,"identity":"2afb27a1-3dae-42de-99dc-a1f17d7f055f","order_by":2,"name":"Steve Rodems","email":"","orcid":"","institution":"Travere Therapeutics Inc","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Steve","middleName":"","lastName":"Rodems","suffix":""},{"id":129899456,"identity":"ccf90a7e-ed51-4381-802c-cada9f0b30f4","order_by":3,"name":"Celia Jenkinson","email":"","orcid":"","institution":"Travere Therapeutics Inc","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Celia","middleName":"","lastName":"Jenkinson","suffix":""}],"badges":[],"createdAt":"2022-08-08 20:54:39","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-1942700/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-1942700/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1186/s13023-022-02578-1","type":"published","date":"2023-01-17T18:25:08+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":25561930,"identity":"b9cd8d21-d249-4cd1-be75-1f5174860e41","added_by":"auto","created_at":"2022-08-23 17:22:34","extension":"jpeg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":860916,"visible":true,"origin":"","legend":"\u003cp\u003eCTX alleles distribution from gnomAD.\u003cstrong\u003e \u003c/strong\u003eDistribution of the CTX alleles with labels for the top alleles (Allele_Freq \u0026gt; 1e-04) and color-coded by sequence consequence. Variant names follow the HGVS transcript or protein consequence. Amino acids are abbreviated by their one letter code.\u0026nbsp;\u003c/p\u003e","description":"","filename":"Figure1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-1942700/v1/ed4718fa08c3e3b1f8e5936e.jpeg"},{"id":25561245,"identity":"a0dfee82-8fc2-46c9-b55f-8c5e98f29480","added_by":"auto","created_at":"2022-08-23 17:17:34","extension":"jpeg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":1021847,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eCYP27A1\u003c/em\u003e pathogenic variants world map.\u003cstrong\u003e \u003c/strong\u003eGeographic distribution of \u003cem\u003eCYP27A1\u003c/em\u003e gene variants queries performed by HCPs of selected clinical/medical professions using the VarSome search engine. The map represents a proxy for clinical activity by country potentially linked to the diagnosis and care of prospective CTX patients. Grey color indicates no queries reported. \u003c/p\u003e\u003cp\u003e\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"Figure2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-1942700/v1/601fcd49b8743dd56077712d.jpeg"},{"id":44717218,"identity":"7e35fc67-118e-4dc5-b1d6-6713c8d604e8","added_by":"auto","created_at":"2023-10-16 18:33:18","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":574679,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-1942700/v1/aede0af8-8393-4b81-a7a3-178663068c9f.pdf"},{"id":25561243,"identity":"d3427b07-c9ec-45fe-997e-36835ed4315b","added_by":"auto","created_at":"2022-08-23 17:17:34","extension":"xlsx","order_by":7,"title":"","display":"","copyAsset":false,"role":"supplement","size":74866,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryTables07APR2022.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-1942700/v1/a717acce628e9753bea75a76.xlsx"}],"financialInterests":"","formattedTitle":"Allelic Prevalence and Geographic Distribution of Cerebrotendinous Xanthomatosis","fulltext":[{"header":"Introduction","content":"\u003cp\u003eCerebrotendinous xanthomatosis (CTX, OMIM 213700) is a rare autosomal recessive condition characterized by disruption of bile acid synthesis due to inactivation of the \u003cem\u003eCYP27A1\u003c/em\u003e gene. Biallelic pathogenic variants are responsible for loss of enzymatic sterol-27-hydroxylase activity leading to reduced production of chenodeoxycholic acid (CDCA) and cholic acid, and accumulation of cholestanol and bile alcohols [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. CTX is highly heterogeneous in clinical presentation and signs and symptoms can overlap with other conditions (eg, sitosterolemia, familial hypercholesterolemia). The first signs and symptoms of the classical presentation are typically non-neurological and may include diarrhea during infancy (sometimes with failure to thrive), juvenile-onset-cataracts, and tendon xanthomas [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. Neurological dysfunction may be present in childhood with intellectual disability and/or autism spectrum disorder and frequently develops in adulthood with epilepsy, pyramidal and extrapyramidal signs, cerebellar syndrome, peripheral neuropathy and intellectual disability/cognitive decline and progressive neurodegeneration [\u003cspan additionalcitationids=\"CR5 CR6 CR7 CR8 CR9\" citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. Because of the pleiotropic phenotypes, accurate diagnosis may be delayed 10\u0026ndash;20 years [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e] and is often achieved only after irreversible neurological involvement has ensued. Treatment with bile acid replacement in the form of CDCA has been successfully used for decades [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e] and, alternatively, cholic acid treatment has been reported in a modest number of cases.\u003c/p\u003e \u003cp\u003eClinical studies have provided powerful evidence showing that early recognition and treatment are critical to prevent or ameliorate disease progression and to avoid devastating neurological complications. Improved outcome was reported in two independent studies for patients with treatment initiated before age 24\u0026ndash;25 years as compared to those patients who started later [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]. These results further emphasize the need to improve early recognition strategies [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e], stimulate newborn screening programs [\u003cspan additionalcitationids=\"CR18\" citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e], and leverage large population genetic data to update and more accurately estimate disease prevalence [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eA relatively recent approach of estimating disease prevalence from the causative or \u0026ldquo;predicted\u0026rdquo; causative alleles (herein labelled allelic prevalence) may contribute to increased disease awareness and improved surveillance and, ultimately, may stimulate earlier diagnosis [\u003cspan additionalcitationids=\"CR22 CR23 CR24\" citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]. This approach has become possible due to the availability of large population databases, such as the Genome Aggregation Database (gnomAD) [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e]. Essentially, this genetic approach is unbiased and orthogonal to clinical studies that instead focus on the recognition, recruitment, and count of the number of confirmed affected individuals. This latter approach may suffer from several sources of bias, likely leading to inaccurate estimates and severe underdiagnosis of the disease. A clear limitation to disease recognition underlying failure to diagnose is the presence of \u0026ldquo;milder\u0026rdquo; phenotypes, which includes those CTX patients without neurological involvement [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]. Despite identification of the same underlying genetic variants observed in severe cases, such individuals may be overlooked and/or misdiagnosed with other, often more common, diseases. The identification and reporting of such patient subpopulations further support the notion that genotype-phenotype correlations have not been established in CTX [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eA previously published first estimate of the genetic prevalence of CTX utilized data from the Exome Aggregation Consortium (ExAC) database, which included\u0026thinsp;~\u0026thinsp;60,000 unrelated individuals [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. Here we provide an update of genetic prevalence utilizing the much larger gnomAD [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e] and taking advantage of additional resources, including the Human Gene Mutation Database (HGMD\u0026reg;) and ClinVar database as well as the VarSome variant classifier. Our approach to CTX prevalence estimate utilizes the Hardy-Weinberg principle and observed variant allele frequencies implementing a variety of variant inclusion/exclusion criteria and automated and manual genetic variant curation steps in an attempt to provide the most conservative and accurate result. In concordance with previous estimates, we report heterogeneity in prevalence across ancestries and identify the most frequent alleles driving these estimates. Finally, we present a map of worldwide geographic estimate disease distribution using a novel approach based on \u003cem\u003eCYP27A1\u003c/em\u003e variant queries performed by clinicians across six continents. The map and CTX prevalence estimate results indicate that the number of individuals with this treatable disorder worldwide should greatly exceed the few hundred currently reported [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e, \u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e].\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eIdentification of high-confidence CYP27A1 pathogenic missense variants for cutoff calculation\u003c/h2\u003e \u003cp\u003eWe defined \u0026ldquo;high-confidence\u0026rdquo; (HC) pathogenic, missense variants classified as such by both the ClinVar database (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ncbi.nlm.nih.gov/clinvar/\u003c/span\u003e\u003cspan address=\"https://www.ncbi.nlm.nih.gov/clinvar/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) and the HGMD (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://www.hgmd.cf.ac.uk/ac/index.php\u003c/span\u003e\u003cspan address=\"http://www.hgmd.cf.ac.uk/ac/index.php\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) as previously adopted [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e]. Only HGMD missense variants with a disease-causing mutation (DM) label that were also present and designated (likely) pathogenic in ClinVar were included in the cutoff calculation (see supplementary material). We reclassified these HC pathogenic variants with the VarSome American College of Medical Genetics and Genomics (ACMG)/Association for Molecular Pathology (AMP)\u0026ndash;based algorithm (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://varsome.com/about/resources/acmg-implementation/\u003c/span\u003e\u003cspan address=\"https://varsome.com/about/resources/acmg-implementation/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) and considered only those classified (likely) pathogenic (see supplementary material). Among the HC pathogenic variants, we included the ClinVar \u0026ldquo;likely pathogenic\u0026rdquo; (LP) based on the following considerations: (1) according to the ACMG/AMP guidelines, LP correlates to a probability of 90% that a variant will be disease-causing [\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e]; (2) DM variants reported in HGMD also included LP according to the VarSome classification and were consistent with ClinVar labels; (3) the difference in interpretation confidence between pathogenic and LP was deemed medically insignificant [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]; and (4) retaining LP variants in our cutoff calculation did not support the hypothesis of lower confidence in pathogenic variant effects for LP. LP variants for the purposes of this study are included as pathogenic.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003eIdentification of CYP27A1 pathogenic variants in gnomAD for prevalence calculation\u003c/h2\u003e \u003cp\u003eThe gnomAD v2.1.1 (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://gnomad.broadinstitute.org/\u003c/span\u003e\u003cspan address=\"https://gnomad.broadinstitute.org/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) was accessed on October 12, 2021, to gather \u003cem\u003eCYP27A1\u003c/em\u003e variant information based on the canonical transcript ENST00000258415.4 and reference sequences NM_000784.4 and NP_000775.1. ClinVar clinical significance annotation in gnomAD was based on the October 2, 2021, release. We selected \u003cem\u003eCYP27A1\u003c/em\u003e variants with a pathogenic ClinVar clinical significance label for all available sequence consequences. We manually verified that these variants were assigned pathogenic in the HGMD (DM) and the ClinVar database. In ClinVar, we considered reliable pathogenic submissions, including single submitters, limited to those from clinical genetic/diagnostic laboratories, since variant classification/interpretation was likely performed following ACMG/AMP guidelines [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e]. All remaining frameshift, nonsense, start_lost, stop_gained, and canonical splice variants (2 bases\u0026thinsp;\u0026plusmn;\u0026thinsp;splice acceptor/donor site) in gnomAD were considered pathogenic based on their high-impact consequence on the coding sequence (\u003cb\u003esee\u003c/b\u003e Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e), as calculated by VEP (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://m.ensembl.org/info/genome/variation/prediction/predicted_data.html\u003c/span\u003e\u003cspan address=\"https://m.ensembl.org/info/genome/variation/prediction/predicted_data.html\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) and classified in previous studies [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e, \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e]. To increase stringency, we manually reclassified and verified each pathogenic variant with the VarSome algorithm (see supplementary material). Lastly, we filtered out variants with a gnomAD label of \u0026ldquo;lc_lof\u0026rdquo; and variants lacking multiple ClinVar entries from clinical genetic/diagnostic laboratories supporting pathogenicity.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003e\u003cem\u003eCYP27A1\u003c/em\u003e gnomAD variants\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eVariant type\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAll\u003c/p\u003e \u003cp\u003eN (%)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eClinV-P\u003c/p\u003e \u003cp\u003eN (%)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSeqCon-P\u003c/p\u003e \u003cp\u003eN (%)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eRevel-P\u003c/p\u003e \u003cp\u003eN (%)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eModel \u003c/p\u003e \u003cp\u003eN (%)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003emissense\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e344 (39.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e13 (3.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e28 (8.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e41 (11.9)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eintron\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e232 (26.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003esynonymous\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e139 (16.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003esplice_region\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e35 (4.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eframeshift\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e31 (3.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e12 (38.7)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e17 (54.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e29 (93.5)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003estop_gained\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e23 (2.7)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e11 (47.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e10 (43.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e21 (91.3)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e5_prime_UTR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e20 (2.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e3_prime_UTR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e18 (2.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003esplice_donor\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8 (0.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e5 (62.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e3 (37.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e8 (100)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003einframe_deletion\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e4 (0.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0 (0. 0)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003esplice_acceptor\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e4 (0.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e3 (75.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1 (25.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e4 (100)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003estart_lost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e3 (0.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e3 (100)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e3 (100)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003einframe_insertion\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e2 (0.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003estop_retained\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1 (0.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTotal\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e864 (100)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e44 (5.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e34 (3.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e28 (3.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e106 (12.3)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"6\"\u003eVariant count breakdown by type: all gnomAD variants (All), designated pathogenic based on ClinVar clinical significance label (ClinV-P); considered pathogenic based on sequence consequence (SeqCon-P); predicted pathogenic based on REVEL analysis (Revel-P), and final list of variants selected for the prevalence calculation (Model).\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003eIdentification of CYP27A1 predicted pathogenic variants in gnomAD for prevalence calculation\u003c/h2\u003e \u003cp\u003eAll remaining missense variants in gnomAD (N\u0026thinsp;=\u0026thinsp;321) underwent bioinformatic analysis to predict the likelihood of pathogenicity. We used the REVEL ensemble in silico predictor, which predicts the pathogenic variant effects of missense based on a combination of scores from 13 individual tools and is considered more reliable in assessing pathogenicity [\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e] than the historically popular SIFT and Polyphen-2 tools [\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e]. REVEL is the reference algorithm for the ACMG/AMP variant interpretation as part of the standard operating procedure in the SVI-approved expert panel specifications (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://clinicalgenome.org/working-groups/sequence-variant-interpretation/\u003c/span\u003e\u003cspan address=\"https://clinicalgenome.org/working-groups/sequence-variant-interpretation/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eWe established a disease-specific cutoff using the mean REVEL score value calculated from the list of HC pathogenic missense variants defined above (see supplementary material). The REVEL mean value was calculated after removing outliers using the interquartile range formula, as previously adopted [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. The calculated score (0.834) was then used as cutoff to filter the remaining 321 missense variants. Only gnomAD missense variants with scores equal to or greater than the calculated REVEL cutoff were predicted pathogenic and retained in the complete model for the prevalence calculation. Lastly, to avoid underestimation of the prevalence calculation, we manually investigated the excluded missense variants and recovered those with ClinVar submissions supporting evidence for pathogenicity and/or those present in HGMD or literature. Literature searches were done in LitVar (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ncbi.nlm.nih.gov/CBBresearch/Lu/Demo/LitVar/\u003c/span\u003e\u003cspan address=\"https://www.ncbi.nlm.nih.gov/CBBresearch/Lu/Demo/LitVar/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e), Pubmed (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://pubmed.ncbi.nlm.nih.gov/\u003c/span\u003e\u003cspan address=\"https://pubmed.ncbi.nlm.nih.gov/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e), and Google Scholar (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://scholar.google.com/\u003c/span\u003e\u003cspan address=\"https://scholar.google.com/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). Mutalyzer (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://mutalyzer.nl/\u003c/span\u003e\u003cspan address=\"https://mutalyzer.nl/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) was used to check the syntax and names or abbreviations of the variants according to updated Human Genome Variation Society (HGSV) nomenclature.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003eCTX allelic prevalence estimation\u003c/h2\u003e \u003cp\u003eAFs of the final list of variants were used to calculate CTX prevalence across six gnomAD populations (AFR\u0026thinsp;=\u0026thinsp;African/African American, AMR\u0026thinsp;=\u0026thinsp;Latino/Admixed American, EUR\u0026thinsp;=\u0026thinsp;Non-Finnish European, FIN\u0026thinsp;=\u0026thinsp;Finnish, SAS\u0026thinsp;=\u0026thinsp;South Asian, EAS\u0026thinsp;=\u0026thinsp;East Asian) using the Hardy-Weinberg principle (p\u003csup\u003e2\u003c/sup\u003e\u0026thinsp;+\u0026thinsp;2pq\u0026thinsp;+\u0026thinsp;q\u003csup\u003e2\u003c/sup\u003e) [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e, \u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e] under the assumption of mutual independence of the rare variants and full penetrance. Prevalence was calculated as the squared sum of the carrier AF of the mutant alleles (q) with p approximated to 1 [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. The 95% confidence interval (CI) was calculated with the binomial formula using the Wilson method for each population\u0026rsquo;s list of risk alleles, as previously adopted [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e, \u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e, \u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003eCTX variants search activity and geographic distribution via VarSome Insights\u003c/h2\u003e \u003cp\u003eWe used the VarSome Insights (VSI) service to gather genomic queries relevant to \u003cem\u003eCYP27A1\u003c/em\u003e gene variation and generate an estimated CTX disease map based on the raw number of unique queries from professionals designated as healthcare providers (HCPs). Original profession labels were reclassified and grouped into fewer labels with the goal of distinguishing clinical/medical queries from research/other type of queries and to resolve word mismatches due to language or spelling differences. We further removed from the analysis HCPs queries that, based on the self-reported profession/position/title, were deemed more likely to be of research or nonclinical nature (eg, Research Scientist). Queries were gathered from January 1, 2021, to November 11, 2021.\u003c/p\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cp\u003eWe identified a total of 864 \u003cem\u003eCYP27A1\u003c/em\u003e gene variants, of which the majority (40%) were missense. Forty-four variants were designated ClinVar pathogenic in gnomAD (12 frameshift, 13 missense, 8 canonical splice site, and 11 stop_gained), and 34 variants were considered pathogenic based on sequence consequence (17 frameshift, 3 start_lost, 10 stop_gained, and 4 canonical splice) (\u003cb\u003esee\u003c/b\u003e Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eWe manually verified that the 44 ClinVar variants were assigned pathogenic in HGMD and ClinVar and that all 78 (44\u0026thinsp;+\u0026thinsp;34) were classified pathogenic by the VarSome algorithm (see supplementary material). Altogether, pathogenic variants represented 9% of all \u003cem\u003eCYP27A1\u003c/em\u003e variants in gnomAD. We then applied our REVEL bioinformatic workflow and identified 23 missense variants passing the disease-specific cutoff, thus predicted to be pathogenic. None of these 23 variants were predicted to be (likely) benign by the VarSome algorithm (see supplementary material), and the vast majority were absent from the HGMD/literature. Of the 23 REVEL-predicted pathogenic variants, p.Met383Lys (VarID: 2-219678874-T-A, HGSV: c.1148T\u0026thinsp;\u0026gt;\u0026thinsp;A) was found with a higher AF (0.12%) in the AMR population compared with the remaining missense variants. With such AF and without reliable estimates for a parameter required to apply the Max Population AF filter [\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e], we opted to calculate two estimates for AMR prevalence: with and without the p.Met383Lys variant. Lastly, we manually searched for missense variants with a REVEL score below the cutoff but with supporting information for pathogenicity. We identified five additional missense variants, bringing the number of predicted pathogenic variants to 28 (\u003cb\u003esee\u003c/b\u003e Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). The average REVEL score of these recovered variants was 0.78, which was within 0.5 standard deviations below the mean score of the HC pathogenic variants.\u003c/p\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003eCTX prevalence estimates\u003c/h2\u003e \u003cp\u003eThe final list of variants for the CTX prevalence calculation included 44 designated ClinVar pathogenic variants, 34 considered pathogenic based on sequence consequence, and 28 (including p.Met383Lys) predicted pathogenic variants by the in silico REVEL analysis (see Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e \u003cb\u003eand supplementary material\u003c/b\u003e). Using the gnomAD carrier AF of these 106 variants, we calculated CTX prevalence estimates and pooled AF by population (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). We found the highest estimates in the EAS and SAS populations, ranging from 2.25 to 1.07 per 100,000 or from 1 per 44,407 to 1 per 93,084, with a pooled AF ranging from 0.00475 (95% CI: 0.00187\u0026ndash;0.01457) to 0.00328 (95% CI: 0.00172\u0026ndash;0.00764), respectively. The lowest prevalence estimate was found in the FIN population with 0.03 per 100,000 or 1 per 3,388,767 and a pooled AF 0.00054 (95% CI: 0.00019\u0026ndash;0.00169). These estimates were not substantially different from those calculated with only the 78 pathogenic variants excluding the variants only predicted pathogenic by the REVEL analysis (\u003cb\u003esee\u003c/b\u003e Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). Intermediate prevalence estimates were found in the AMR, AFR, and EUR populations. In AMR without the variant p.Met383Lys (M383K), these estimates were 0.63 per 100,000 or 1 per 157,878 with a pooled AF 0.00252 (95% CI: 0.00113\u0026ndash;0.00653), while those with the M383K variant were 1.41 per 100,000 or 1 per 70,795, with a pooled AF 0.00376 (95% CI: 0.00206\u0026ndash;0.0082). In AFR, these estimates were 0.6 per 100,000 or 1 per 166,440, with a pooled AF 0.00245 (95% CI: 0.00071\u0026ndash;0.00898). In EUR, these estimates were 0.43 per 100,000 or 1 per 233,597, with a pooled AF 0.00207 (95% CI: 0.0008\u0026ndash;0.00671). Of note, estimates calculated without the predicted pathogenic missense variants showed very similar prevalence across AFR, AMR, and EUR, ranging from 0.25 to 0.21 per 100,000 or 1 per 393,497 to 482,603 (\u003cb\u003esee\u003c/b\u003e Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). Figure\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e shows the most frequent alleles in each population. The missense variant G472A (VarID: 2-219679419-G-C, HGVS: c.1415G\u0026thinsp;\u0026gt;\u0026thinsp;C, p.Gly472Ala) and the nonsense variant (VarID: 2-219646907-T-C, HGVS: c.2T\u0026thinsp;\u0026gt;\u0026thinsp;C, p.Met1?) were the most frequent in the EAS and SAS populations. G472A was also recently identified in a dried blood spot sample together with the variant R405Q (VarID: 2-219679132-G-A, HGVS: c.1214G\u0026thinsp;\u0026gt;\u0026thinsp;A, p.Arg405Gln), which we also found in the top EAS alleles [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. The missense variant R395C (VarID: 2-219678909-C-T, HGVS: c.1183C\u0026thinsp;\u0026gt;\u0026thinsp;T, p.Arg395Cys) remained the most frequent in the FIN and EUR populations, while in the AMR population was second to the missense M383K (VarID: 2-219678874-T-A, HGSV: c.1148T\u0026thinsp;\u0026gt;\u0026thinsp;A, p.Met383Lys). The missense variant R448H (VarID: 2-219679347-G-A, HGVS: c.1343G\u0026thinsp;\u0026gt;\u0026thinsp;A, p.Arg448His) was the most frequent in the AFR population.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eCTX prevalence estimates across six gnomAD populations\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"8\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eEstimate\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAFR\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eEAS\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eFIN\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eEUR\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eAMR*\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c8\"\u003e \u003cp\u003eSAS\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eClinV-P +\u003c/p\u003e \u003cp\u003eSeqCon-P\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eper 100,000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.21\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1.41\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.03\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.25\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.21\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e0.95\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1 per\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e472,468\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e71,089\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e3,946,059\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e393,497\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e482,603\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e105,299\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eClinV-P +\u003c/p\u003e \u003cp\u003eSeqCon-P +\u003c/p\u003e \u003cp\u003eRevel-P\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eper 100,000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e2.25\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.03\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.43\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e1.41/0.63\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e1.07\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1 per\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e166,440\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e44,407\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e3,388,767\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e233,597\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e70,795/157,878\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e93,084\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePooled AF\u003c/p\u003e \u003cp\u003e(95% CI)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.00245 (0.00071\u0026ndash;0.00898)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.00475 (0.00187\u0026ndash;0.01457)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.00054 (0.00019\u0026ndash;0.00169)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.00207 (0.0008\u0026ndash;0.00671)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.00376 (0.00206\u0026ndash;0.0082)/0.00252 (0.00113\u0026ndash;0.00653)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e0.00328 (0.00172\u0026ndash;0.00764)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"8\"\u003e*In the model ClinV-P\u0026thinsp;+\u0026thinsp;SeqCon-P\u0026thinsp;+\u0026thinsp;Revel-P, the two prevalence estimates correspond to with/without the M383K variant.\u003c/td\u003e\u003c/tr\u003e \u003ctr\u003e\u003ctd colspan=\"8\"\u003eAF, allele frequency; AFR, African/African American; AMR, Latino Admixed American; EAS, East Asian; EUR, European non-Finnish; FIN, European Finnish; SAS, South Asian.\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003eCTX search activity geographic distribution\u003c/h2\u003e \u003cp\u003eWe leveraged VSI to infer a possible CTX disease geographic distribution estimate based on genomic variant queries through the VarSome search engine. This analysis was performed under the assumption that queries were potentially linked to clinical activity ultimately aiding the diagnosis or care of prospective CTX patients. In other words, it was assumed that HCPs from the selected clinical/medical specialties using VarSome in this manner were gathering \u003cem\u003eCYP27A1\u003c/em\u003e variant pathogenicity information because a patient in their care was likely affected with CTX. Overall, queries were performed largely by clinical and research professionals using the platform to gather information on the \u003cem\u003eCYP27A1\u003c/em\u003e gene and related variants. In 2021, we identified a total of 1,243 queries, of which, 827 were from designated HCPs, representing 67% of all queries. After removing those under research or nonclinical professions, we considered 576 queries as a proxy for CTX clinical activity, of which half (n\u0026thinsp;=\u0026thinsp;288) were from clinical geneticists and genetic counselors.\u003c/p\u003e \u003cp\u003eAnalysis of the final queries from HCPs showed that pathogenic, uncertain significance, and benign variants were equally represented (~\u0026thinsp;33% each). The most queried variant was by far the missense P384L (NM_000784.4:c.1151C\u0026thinsp;\u0026gt;\u0026thinsp;T), which is consistent with its high population AF and several benign submissions in ClinVar. Among the top pathogenic variants, we found a similar number of queries for c.1184\u0026thinsp;+\u0026thinsp;1G\u0026thinsp;\u0026gt;\u0026thinsp;A, R127W (NM_000784.4:c.379C\u0026thinsp;\u0026gt;\u0026thinsp;T) and R395C/S (NM_000784.4:c.1183C\u0026thinsp;\u0026gt;\u0026thinsp;T/A), which are also found among the top alleles across gnomAD populations in our prevalence model (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). Lastly, we mapped all the VarSome pathogenic HCPs queries (N\u0026thinsp;=\u0026thinsp;187) to infer a possible geographic distribution of CTX disease and identified Spain as the top country with the highest number of \u003cem\u003eCYP27A1\u003c/em\u003e variant queries, followed by Iran, Italy, USA, and Turkey (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eThere is scientific consensus that CTX is clinically not well recognized, frequently has delayed diagnosis [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e], and has a prevalence likely underestimated in the general population [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e, \u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e], with less than 600 cases reported worldwide [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e, \u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e]. In order to facilitate early diagnosis of this disease, we sought to calculate the most updated conservative and accurate estimates for CTX prevalence. We followed the approach of previous studies for estimating how common a monogenic autosomal recessive disease may be [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e, \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e, \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e, \u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e, \u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e], and we leveraged the availability of a large genetic dataset, gnomAD. We developed a highly curated list of alleles in \u003cem\u003eCYP27A1\u003c/em\u003e that have been reported to be pathogenic and used their bioinformatic characteristics to identify additional missense alleles of \u003cem\u003eCYP27A1\u003c/em\u003e that are predicted to be pathogenic for CTX. We then applied the Hardy-Weinberg principle to estimate CTX risk in global populations. These analyses show that CTX is present in global populations at rates more common than previously appreciated. For the first time, we also attempted to geographically map CTX clinical activity worldwide, using a new emerging tool for genomic variation data sharing.\u003c/p\u003e \u003cp\u003eThe belief that CTX is an exceedingly rare occurrence is not an uncommon scenario in the field of rare diseases. As for CTX, this is typically due to the intrinsic nature of the disease, which can show pleomorphism in clinical and laboratory features overlapping with other diseases, having a variable clinical course and symptoms onset, and lacking any clear genotype-phenotype correlations [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e]. The task of establishing an accurate estimate for how common CTX may be in the population becomes even more challenging considering the presence of a subpopulation of patients with a milder phenotype [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]. Interestingly, such individuals may carry the same variants (eg, p.Arg395Cys) that are found in patients with neurological involvement, suggesting that perhaps additional damaging mechanisms or genetic modifiers other than \u003cem\u003eCYP27A1\u003c/em\u003e loss of function may be at play. Reinforcing this enigmatic picture is the report of a pair of siblings carrying the same pathogenic variants but being at opposite ends of the clinical spectrum, where one sibling developed rare spinal xanthomatosis and the other developed a mild form with minor tendon xanthomas [\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e]. In contrast, examples of common clinical manifestation are well documented. Bilateral cataracts are found in about 80% of CTX patients [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e, \u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e] and interim analysis of patients recruited on the sole basis for having juvenile-onset idiopathic bilateral cataract shows a molecularly confirmed diagnosis in 1.5%-1.8% of the patients, representing a 500-fold increase in CTX prevalence in this subset of patients [\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e, \u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eOur findings largely support, expand upon, and refine previous genetic estimates based on the ExAC data and more limited variant inclusion [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. We identified three levels for CTX prevalence. The highest estimates were found in the Asian populations at 1 per 44,407\u0026thinsp;\u0026minus;\u0026thinsp;93,084, which are in line with previous investigations (1:36,072\u0026ndash;75,601) [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. In both studies, the top variant remained the missense c.1415G\u0026thinsp;\u0026gt;\u0026thinsp;C (p.Gly472Ala), but with a higher carrier AF in gnomAD compared with the ExAC (0.0014 vs 0.0010) [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. To our knowledge, the p.Gly472Ala variant has been reported in two patients: one of Asian origin carrying a homozygous mutation [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e] and one in a newborn screening [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. In the latter study, one sample was found positive for biochemical CTX biomarkers and compound heterozygous for two pathogenic variants in CYP27A1 (c.1214G\u0026thinsp;\u0026gt;\u0026thinsp;A - p.Arg405Gln; c.1415G\u0026thinsp;\u0026gt;\u0026thinsp;C - p.Gly472Ala). While we do not have confirmation that this sample was from an individual of EAS ancestry, we speculate that this might be the case as p.Gly472Ala is unique to EAS and both alleles show high AF in EAS compared with the other ancestries in gnomAD. In such a scenario, the reported incidence of 1 per 32,000 from Hong et al. [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e], might provide an independent and orthogonal validation of our findings in the EAS population.\u003c/p\u003e \u003cp\u003eAn intermediate level of CTX prevalence was found in the AMR, AFR, and EUR populations. In AMR, although we obtained very close estimates (1 per 70,795\u0026thinsp;\u0026minus;\u0026thinsp;157,878) compared with the ExAC study (1:71,677\u0026thinsp;\u0026minus;\u0026thinsp;148,914) [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e], differences in the number, type, and AF of variants may be found. For instance, we identified twice as many alleles, with the top variant p.Met383Lys showing a much higher AF (0.00124 vs 0.00078). We provided two prevalence estimates for AMR since it would be difficult to confidently include/exclude this variant from our analysis. Following ACMG/AMP guidelines, we could attempt to apply the BS1 criteria that is used to classify a variant as \u0026ldquo;likely benign\u0026rdquo; when its AF is greater than expected for the disorder [\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e, \u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e, \u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e]. However, reliable estimates for population prevalence and allelic heterogeneity would be needed to use the maximum credible population AF as filter [\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e]. A plausible strategy to leverage this filter would be to create a distribution of prevalence values and use two estimates for allelic heterogeneity, for instance, 10% (conservative) and 30% based on the most frequent allele found in CTX patients (c.1183C\u0026thinsp;\u0026gt;\u0026thinsp;T, p.Arg395Cys) [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]. By doing this exercise, we found that at 10% allelic heterogeneity, p.Met383Lys was retained only with a disease prevalence of 1:10,000\u0026ndash;20,000 (penetrance 100%-50%), which is currently not supported by any clinical, epidemiological, or genetic study in the AMR population. At 30% allelic heterogeneity, we found that p.Met383Lys was retained with a disease prevalence of 1:100,000-180,000 (penetrance 100%-50%), which may be currently supported by genetic estimates from the previous CTX study (~\u0026thinsp;1:70,000-150,000) [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. However, it could be argued that if p.Met383Lys was as frequent as p.Arg395Cys, we would have expected at least a few CTX patients described in the literature carrying this allele.\u003c/p\u003e \u003cp\u003eConsistent with the strong literature evidence, we found p.Arg395Cys to be most prevalent in the EUR, FIN, and AMR populations. Also, this variant showed an increase in AF compared with the ExAC frequency (0.00051 vs 0.00017) [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. Our estimates instead show a significant increase in prevalence for the AFR and EUR populations, that moved from 1:468,624 to 1:166,440 and from 1:461,358 to 1:233,597, respectively [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. Again, the same top drivers were found between the two studies but with an increase in AF (p.Arg448His: 0.00041 vs 0.00031 in AFR; p.Arg395Cys: 0.00038 vs 0.00021 in EUR). As discussed above, such differences are not surprising, especially when considering the sizable difference in the number of individuals investigated between the two databases. In addition, we need to take into account improvements in the performance of in silico predictors and the evolving nature of clinical and functional information, which are both critical to variant classification.\u003c/p\u003e \u003cp\u003eSimilar to the AF of p.Met383Lys, we have retained in our model the start_lost variant c.2T\u0026thinsp;\u0026gt;\u0026thinsp;C (NM_000784.4:p.?), which was found to be unique to the SAS population. While this type of variant may be considered always pathogenic, as it is expected to produce no protein, the clinical impact in practice can be heterogeneous, and alternative mechanisms for protein translation should be considered. In fact, it is notable to find this variant at the heterozygous state in four siblings of a South African family with a mild (no neurological involvement) CTX phenotype (see family #8) [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]. It is possible that future functional studies will be able to clarify the role of the c.2T\u0026thinsp;\u0026gt;\u0026thinsp;C allele in the pathogenesis of CTX or perhaps to rule out its involvement with direct implications for our SAS estimate. The lowest level of prevalence was found in the FIN population, which is in line with findings from the previous CTX study [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eLastly, with the purpose of increasing awareness, having a proactive approach to patient identification, and aiding early treatment intervention, we leveraged the VarSome platform to create a clinically relevant geographic map based on recent query activity. We found that most of the variant queries were from clinical geneticists/genetic counselors and that the most searched variants were consistent with our findings from the analysis of the gnomAD AF. We recognize that the proposed geographic map represents only a proxy to a possible clinical distribution of prospective CTX patients and that the difference in counts per country may also depend on or be limited to factors such as access and popularity of the VarSome search engine, country size, accessibility, and costs to genetic testing, as well as socioeconomic or political efforts to boost scientific progress in the fight against rare diseases. Nonetheless, our findings line up with patient reports from around the world (USA, Israel, Italy, Japan, the Netherlands, Belgium, Brazil, Canada, France, Iran, Norway, Tunisia, Spain, China, and Sweden; see CTX at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ncbi.nlm.nih.gov/books/NBK1409/\u003c/span\u003e\u003cspan address=\"https://www.ncbi.nlm.nih.gov/books/NBK1409/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://rarediseases.org/\u003c/span\u003e\u003cspan address=\"https://rarediseases.org/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eThere are limitations and methodological assumptions underlying the approach in this study. We and others have applied the Hardy-Weinberg principle in our genetic risk calculations that assume that heterozygous individuals are not subject to selection, populations are at equilibrium with respect to allele and genotype frequencies, and random mating is observed. Also, the concept of AF and how rare an allele may be is strictly related to the size of the population under investigation; thus, with the availability of larger population-based datasets, estimates will change and become more accurate. Lastly, we recognize that although the adoption of filtering criteria and in silico tools to predict pathogenic missense with unknown or uncertain clinical significance is helpful and commonly used, there is currently no gold standard and these strategies could lead to the over- or underestimation of disease frequency. We attempted to mitigate overestimation by cross-referencing the pathogenic labels across different databases and sources and by applying strict filtering criteria with the goal of providing rather conservative estimates. We mitigated underestimation by manually curating missense variants that were filtered out by our bioinformatic workflow. Additional genetic variation that has not been accounted for in our calculation may be conferred by inframe deletions or insertions, intronic, noncanonical splice sites, and structural variants. Lasty, our approach assumed that all variants in the final model contribute to the risk of disease with 100% penetrance. While we are not aware of reports on \u003cem\u003eCYP27A1\u003c/em\u003e pathogenic variants of reduced penetrance, we cannot exclude their presence.\u003c/p\u003e \u003cp\u003eIn conclusion, our study, that includes additional variants, new informatics tools, and newer, expanded databases, supports and refines previous estimates for CTX disease risk at the population level and provides a novel prospective geographical map of CTX clinical activity worldwide. We confirm with this larger, more comprehensive study that CTX is more common than current worldwide patient estimates, and we highlight the most common pathogenic variants. We underscore the value of leveraging large and diversified population-based genetic databases to assess risk for inherited diseases. Such efforts, cross-referenced with other large-scale programs, such as newborn screening [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e] and retrospective administrative claims studies [\u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e], may provide the most accurate strategy to assess disease presence in a population. In turn, this will translate into greater awareness, better recognition, and early treatment intervention, which will directly benefit patients and caretakers.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll data generated or analyzed during this study are included in this published article or uploaded as supplementary information and/or available in public open access databases described in the methods section.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTP, SR, CJ are full-time employees of Travere Therapeutics, Inc., and may have an equity or other financial interest in Travere Therapeutics, Inc. RDS is an employee of PreventionGenetics, an Exact Sciences company. He has equity interest in and has received consulting fees from Acer Therapeutics and PTC Therapeutics. He has received consulting fees from Aeglea BioTherapeutics; Alexion Pharmaceuticals, Inc; Best Doctors; Health Advances LLC; Leadiant, Precision for Value; and Travere Therapeutics, Inc; and honoraria from Medscape/WebMD and The France Foundation. He received no funding to support writing and revising this manuscript. He has received research funding from Alexion Pharmaceuticals, Inc and the Smith Lemli Opitz Foundation.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study was funded by Travere Therapeutics, Inc.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026rsquo; contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTP, SR, CJ conceived the study. All authors contributed to the review of the results, the writing and revision of the manuscript. TP collected the data, performed the analyses, and generated the results.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgments\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe thank Shawn Hayes, Patricia Bedard, and Karsten Baumgaertel for critical reading of the manuscript and/or the helpful feedback on the approach. We thank Dr. Colleen Burns for helping with the confidence interval calculation. We thank Andy Cosgrove for providing access to the queries raw data for the geographic mapping. We thank the MedVal Scientific Information Services, LLC team for editorial support.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n \u003cli\u003e\u003cspan\u003eBjorkhem I, Hansson M. Cerebrotendinous xanthomatosis: an inborn error in bile acid synthesis with defined mutations but still a challenge. Biochem Biophys Res Commun. 2010;396(1):46\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eBerginer VM, Salen G, Shefer S. Long-term treatment of cerebrotendinous xanthomatosis with chenodeoxycholic acid. N Engl J Med. 1984;311(26):1649\u0026ndash;52.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eNie S, Chen G, Cao X, Zhang Y. Cerebrotendinous xanthomatosis: a comprehensive review of pathogenesis, clinical manifestations, diagnosis, and management. Orphanet J Rare Dis. 2014;9:179.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eVerrips A, van Engelen BG, Wevers RA, et al. Presence of diarrhea and absence of tendon xanthomas in patients with cerebrotendinous xanthomatosis. Arch Neurol. 2000;57(4):520\u0026ndash;4.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eVerrips A, Hoefsloot LH, Steenbergen GC, et al. Clinical and molecular genetic characteristics of patients with cerebrotendinous xanthomatosis. Brain. 2000;123(Pt 5):908\u0026ndash;19.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eSalen G, Steiner RD. Epidemiology, diagnosis, and treatment of cerebrotendinous xanthomatosis (CTX). J Inherit Metab Dis. 2017;40(6):771\u0026ndash;81.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eStelten BML, Bonnot O, Huidekoper HH, et al. Autism spectrum disorder: an early and frequent feature in cerebrotendinous xanthomatosis. J Inherit Metab Dis. 2018;41(4):641\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eYunisova G, Tufekcioglu Z, Dogu O, et al. Patients with Lately Diagnosed Cerebrotendinous Xanthomatosis. Neurodegener Dis. 2019;19(5\u0026ndash;6):218\u0026ndash;24.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eAmador MDM, Masingue M, Debs R, et al. Treatment with chenodeoxycholic acid in cerebrotendinous xanthomatosis: clinical, neurophysiological, and quantitative brain structural outcomes. J Inherit Metab Dis. 2018;41(5):799\u0026ndash;807.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eFraidakis MJ. Psychiatric manifestations in cerebrotendinous xanthomatosis. Transl Psychiatry. 2013;3:e302.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eDeBarber AE, Duell PB. Update on cerebrotendinous xanthomatosis. Curr Opin Lipidol. 2021;32(2):123\u0026ndash;31.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eBadura-Stronka M, Hirschfeld AS, Winczewska-Wiktor A, et al. First case series of Polish patients with cerebrotendinous xanthomatosis and systematic review of cases from the 21st century. Clin Genet. 2022;101(2):190\u0026ndash;207.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eSalen G, Berginer V, Shore V, et al. Increased concentrations of cholestanol and apolipoprotein B in the cerebrospinal fluid of patients with cerebrotendinous xanthomatosis. Effect of chenodeoxycholic acid. N Engl J Med. 1987;316(20):1233\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eYahalom G, Tsabari R, Molshatzki N, Ephraty L, Cohen H, Hassin-Baer S. Neurological outcome in cerebrotendinous xanthomatosis treated with chenodeoxycholic acid: early versus late diagnosis. Clin Neuropharmacol. 2013;36(3):78\u0026ndash;83.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eStelten BML, Huidekoper HH, van de Warrenburg BPC, et al. Long-term treatment effect in cerebrotendinous xanthomatosis depends on age at treatment start. Neurology. 2019;92(2):e83\u0026ndash;95.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eMignarri A, Gallus GN, Dotti MT, Federico A. A suspicion index for early diagnosis and treatment of cerebrotendinous xanthomatosis. J Inherit Metab Dis. 2014;37(3):421\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eDeBarber AE, Luo J, Star-Weinstock M, et al. A blood test for cerebrotendinous xanthomatosis with potential for disease detection in newborns. J Lipid Res. 2014;55(1):146\u0026ndash;54.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eVaz FM, Bootsma AH, Kulik W, et al. A newborn screening method for cerebrotendinous xanthomatosis using bile alcohol glucuronides and metabolite ratios. J Lipid Res. 2017;58(5):1002\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eHong X, Daiker J, Sadilek M, et al. Toward newborn screening of cerebrotendinous xanthomatosis: results of a biomarker research study using 32,000 newborn dried blood spots. Genet Med. 2020;22(10):1606\u0026ndash;12.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eAppadurai V, DeBarber A, Chiang PW, et al. Apparent underdiagnosis of Cerebrotendinous Xanthomatosis revealed by analysis of ~ 60,000 human exomes. Mol Genet Metab. 2015;116(4):298\u0026ndash;304.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eCarter A, Brackley SM, Gao J, Mann JP. The global prevalence and genetic spectrum of lysosomal acid lipase deficiency: A rare condition that mimics NAFLD. J Hepatol. 2019;70(1):142\u0026ndash;50.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eNappo S, Mannucci L, Novelli G, Sangiuolo F, D\u0026apos;Apice MR, Botta A. Carrier frequency of CFTR variants in the non-Caucasian populations by genome aggregation database (gnomAD)-based analysis. Ann Hum Genet. 2020;84(6):463\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eBorges P, Pasqualim G, Giugliani R, Vairo F, Matte U. Estimated prevalence of mucopolysaccharidoses from population-based exomes and genomes. Orphanet J Rare Dis. 2020;15(1):324.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eGao J, Brackley S, Mann JP. The global prevalence of Wilson disease from next-generation sequencing data. Genet Med. 2019;21(5):1155\u0026ndash;63.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eTan J, Wagner M, Stenton SL, et al. Lifetime risk of autosomal recessive mitochondrial disorders calculated from genetic databases. EBioMedicine. 2020;54:102730.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eKarczewski KJ, Francioli LC, Tiao G, et al. The mutational constraint spectrum quantified from variation in 141,456 humans. Nature. 2020;581(7809):434\u0026ndash;43.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eStelten BML, Raal FJ, Marais AD, et al. Cerebrotendinous xanthomatosis without neurological involvement. J Intern Med. 2021;290:1039\u0026ndash;47.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eStelten BML, Dotti MT, Verrips A, et al. Expert opinion on diagnosing, treating and managing patients with cerebrotendinous xanthomatosis (CTX): a modified Delphi study. Orphanet J Rare Dis. 2021;16(1):353.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003evan Rooij J, Arp P, Broer L, et al. Reduced penetrance of pathogenic ACMG variants in a deeply phenotyped cohort study and evaluation of ClinVar classification over time. Genet Med. 2020;22(11):1812\u0026ndash;20.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eRichards S, Aziz N, Bale S, et al. Standards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology. Genet Med. 2015;17(5):405\u0026ndash;24.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eHarrison SM, Dolinsky JS, Knight Johnson AE, et al. Clinical laboratories collaborate to resolve differences in variant interpretations submitted to ClinVar. Genet Med. 2017;19(10):1096\u0026ndash;104.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eRehm HL, Berg JS, Brooks LD, et al. ClinGen\u0026ndash;the Clinical Genome Resource. N Engl J Med. 2015;372(23):2235\u0026ndash;42.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eBrezavar D, Bonnen PE. Incidence of PKAN determined by bioinformatic and population-based analysis of ~ 140,000 humans. Mol Genet Metab. 2019;128(4):463\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eIoannidis NM, Rothstein JH, Pejaver V, et al. REVEL: An Ensemble Method for Predicting the Pathogenicity of Rare Missense Variants. Am J Hum Genet. 2016;99(4):877\u0026ndash;85.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eHardy GH. Mendelian Proportions in a Mixed Population. Science. 1908;28(706):49\u0026ndash;50.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eMinikel EV, Vallabh SM, Lek M, et al. Quantifying prion disease penetrance using large population control cohorts. Sci Transl Med. 2016;8(322):322ra9.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eSchrodi SJ, DeBarber A, He M, et al. Prevalence estimation for monogenic autosomal recessive diseases using population-based genetic data. Hum Genet. 2015;134(6):659\u0026ndash;69.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eWhiffin N, Minikel E, Walsh R, et al. Using high-resolution variant frequencies to empower clinical genome interpretation. Genet Med. 2017;19(10):1151\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eRaymond GV, Schiffmann R. Cerebrotendinous xanthomatosis: The rare \u0026quot;treatable\u0026quot; disease you never consider. Neurology. 2019;92(2):61\u0026ndash;2.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eCoffey AJ, Durkie M, Hague S, et al. A genetic study of Wilson\u0026apos;s disease in the United Kingdom. Brain. 2013;136(Pt 5):1476\u0026ndash;87.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eKaler SG, Ferreira CR, Yam LS. Estimated birth prevalence of Menkes disease and ATP7A-related disorders based on the Genome Aggregation Database (gnomAD). Mol Genet Metab Rep. 2020;24:100602.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eZadori D, Szpisjak L, Madar L, et al. Different phenotypes in identical twins with cerebrotendinous xanthomatosis: case series. Neurol Sci. 2017;38(3):481\u0026ndash;3.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eGuenzel AJ, DeBarber A, Raymond K, Dhamija R. Familial variability of cerebrotendinous xanthomatosis lacking typical biochemical findings. JIMD Rep. 2021;59(1):3\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eWong JC, Walsh K, Hayden D, Eichler FS. Natural history of neurological abnormalities in cerebrotendinous xanthomatosis. J Inherit Metab Dis. 2018;41(4):647\u0026ndash;56.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eFreedman SF, Brennand C, Chiang J, et al. Prevalence of Cerebrotendinous Xanthomatosis Among Patients Diagnosed With Acquired Juvenile-Onset Idiopathic Bilateral Cataracts. JAMA Ophthalmol. 2019;137(11):1312\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eAtilla H, Coskun T, Elibol B, Kadayifcilar S, Altinel S. Group G-E-IW. Prevalence of cerebrotendinous xanthomatosis in cases with idiopathic bilateral juvenile cataract in ophthalmology clinics in Turkey. J AAPOS. 2021;25:269.e1-269.e6.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eTavtigian SV, Greenblatt MS, Harrison SM, et al. Modeling the ACMG/AMP variant classification guidelines as a Bayesian classification framework. Genet Med. 2018;20(9):1054\u0026ndash;60.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eSavige J, Storey H, Watson E, et al. Consensus statement on standards and guidelines for the molecular diagnostics of Alport syndrome: refining the ACMG criteria. Eur J Hum Genet. 2021;29(8):1186\u0026ndash;97.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eSellos-Moura M, Glavin F, Lapidus D, Evans K, Lew CR, Irwin DE. Prevalence, characteristics, and costs of diagnosed homocystinuria, elevated homocysteine, and phenylketonuria in the United States: a retrospective claims-based comparison. BMC Health Serv Res. 2020;20(1):183.\u003c/span\u003e\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"orphanet-journal-of-rare-diseases","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"ojrd","sideBox":"Learn more about [Orphanet Journal of Rare Diseases](http://ojrd.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/ojrd/default.aspx","title":"Orphanet Journal of Rare Diseases","twitterHandle":"@bmc","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"BMC/SO AJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"allele frequency, rare metabolic disease, cerebrotendinous xanthomatosis (CTX), genomic variation, pathogenic variants","lastPublishedDoi":"10.21203/rs.3.rs-1942700/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-1942700/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eBackground:\u003c/strong\u003e Cerebrotendinous xanthomatosis (CTX) is a rare recessive genetic disease characterized by disruption of bile acid synthesis due to inactivation of the \u003cem\u003eCYP27A1\u003c/em\u003e gene. Treatment is available in the form of bile acid replacement. CTX is likely underdiagnosed, and prevalence estimates based on case diagnosis are probably inaccurate. Large population-based genomic databases are a valuable resource to estimate prevalence of rare recessive diseases as an orthogonal unbiased approach building upon traditional epidemiological studies. \u003c/p\u003e\u003cp\u003e\u003cstrong\u003eMethods:\u003c/strong\u003e We leveraged the Hardy-Weinberg principle and allele frequencies from gnomAD to calculate CTX prevalence. ClinVar and HGMD were used to identify high-confidence pathogenic missense variants and to calculate a disease-specific cutoff. Variant pathogenicity was also assessed by the VarSome implementation of the ACMG/AMP algorithm and the REVEL in silico predictor. \u003c/p\u003e\u003cp\u003e\u003cstrong\u003eResults:\u003c/strong\u003e CTX prevalence estimates were highest in Asians (1:44,407-93,084) and lowest in the Finnish population (1:3,388,767). Intermediate estimates were found in Europeans, Americans, and Africans/African Americans (1:70,795-233,597). The REVEL-predicted pathogenic variants accounted for a greater increase in prevalence estimates for Europeans, Americans, and Africans/African Americans compared with Asians. We identified the most frequent alleles designated pathogenic in ClinVar (p.Gly472Ala, p.Arg395Cys), labeled pathogenic based on sequence consequence (p.Met1?), and predicted to be pathogenic by REVEL (p.Met383Lys, p.Arg448His) across populations. Also, we provide a prospective geographic map of estimated disease distribution based on \u003cem\u003eCYP27A1\u003c/em\u003e variation queries performed by healthcare providers from selected specialties. \u003c/p\u003e\u003cp\u003e\u003cstrong\u003eConclusions:\u003c/strong\u003e Prevalence estimates calculated herein support and expand upon existing evidence indicating underdiagnosis of CTX, suggesting that improved detection strategies are needed. Increased awareness of CTX is important for early diagnosis, which is essential for patients as early treatment significantly slows or prevents disease progression.\u003c/p\u003e","manuscriptTitle":"Allelic Prevalence and Geographic Distribution of Cerebrotendinous Xanthomatosis","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2022-08-23 17:17:32","doi":"10.21203/rs.3.rs-1942700/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"reviewerAgreed","content":"","date":"2022-09-30T04:34:56+00:00","index":0,"fulltext":""},{"type":"reviewersInvited","content":"","date":"2022-08-18T13:44:58+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2022-08-10T07:00:05+00:00","index":"","fulltext":""},{"type":"submitted","content":"Orphanet Journal of Rare Diseases","date":"2022-08-08T16:53:22+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"orphanet-journal-of-rare-diseases","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"ojrd","sideBox":"Learn more about [Orphanet Journal of Rare Diseases](http://ojrd.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/ojrd/default.aspx","title":"Orphanet Journal of Rare Diseases","twitterHandle":"@bmc","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"BMC/SO AJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"cfe4d77a-ed9a-43e6-932a-86ccaac4f199","owner":[],"postedDate":"August 23rd, 2022","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2023-10-16T18:29:10+00:00","versionOfRecord":{"articleIdentity":"rs-1942700","link":"https://doi.org/10.1186/s13023-022-02578-1","journal":{"identity":"orphanet-journal-of-rare-diseases","isVorOnly":false,"title":"Orphanet Journal of Rare Diseases"},"publishedOn":"2023-01-17 18:25:08","publishedOnDateReadable":"January 17th, 2023"},"versionCreatedAt":"2022-08-23 17:17:32","video":"","vorDoi":"10.1186/s13023-022-02578-1","vorDoiUrl":"https://doi.org/10.1186/s13023-022-02578-1","workflowStages":[]},"version":"v1","identity":"rs-1942700","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-1942700","identity":"rs-1942700","version":["v1"]},"buildId":"cBFmMYwuxLRRLfASyISRj","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.