{"paper_id":"f41958e2-2246-4be3-bf52-72c2aed26099","body_text":"The integration of electronic health records (EHRs) with genetic data is\ntransforming biomedical research, offering unprecedented opportunities for\npreventing and managing common medical conditions 1 . Longitudinal sampling, linked with genetic and\nenvironmental data, provides advantages over standard cohort-driven\nresearch 2 . This has fueled\nthe creation of nationwide biobanks, including the UK Biobank (UKBB)  3 , 4 , All of Us (AoU) 5 , FinnGen 6  and\nTaiwan Biobank 7 , as well as several\nlarge-scale academic biobanks, such as Mt. Sinai’s BioMe 8 , Vanderbilt’s BioVU 9 , Geissinger’s MyCode 10 , and the Michigan Genomics\nInitiative 11 . Integration\nof these biobanks has permitted innovative collaborative efforts, such as the eMERGE\nconsortium 12 , COVID-19\nhost genomics initiative 13 , and\nthe Global Biobank Initiative 1 .\nAlthough these efforts have substantially advanced genetic and biomedical\ndiscovery, their concentration on participants of European (EUR) ancestry limits\ngeneralizability 14 – 18 . As PGS are validated for clinical\nuse, the importance of measuring their accuracy in diverse populations grows; recent\nanalyses show a continuous relationship between ancestral distance from the\nreference population and the utility of PGS 17 . Similarly, rare genetic variation has substantial\nancestry-specific effects and distribution; for example, in African Americans,\n APOE4  alleles have reduced impact on Alzheimer’s disease\nrisk 19  and rare protective\n PCSK9  variants are more prevalent 20 . The interpretation of clinically-relevant\nrare variation is further hampered by a bias towards EUR variants in genetic\ndatabases 21 , 22 . Including non-EUR populations reveals\nsubstantial disparities in clinically-relevant rare variant frequencies 23  and increases statistical power\nfor discovery 24 , 25 . Thus, greater ancestral diversity in\nbiobanks with detailed medical records strengthens efforts to advance precision\nhealth 14 , 26 , 27 .\nHere, we analyzed data from the UCLA ATLAS Community Health Initiative,\nlinking EHR with genomic information for 92,164 participants with array genotyping\nand 61,797 with whole-exome sequencing (WES). ATLAS reflects Los Angeles’\nancestral diversity within a single health system 28 – 30 , which reduces confounding from differing clinical practices\nand enables robust cross-population comparisons. We performed phenome-wide\nassociations using common and rare variants within continental (broad-scale) and\nsub-continental (fine-scale) ancestral groups, and quantified disease diagnoses and\ngenetic risk across these groups. Using longitudinal EHR data, we employed\nsemaglutide as a case study to identify genetic factors that alter drug efficacy. We\nidentified multiple unreported ancestry-specific risk associations, further\ndemonstrating the utility of ancestral diversity for more equitable, personalized\nmedicine research ( Figure 1a , schematic\noverview).\n\nUCLA Health serves Los Angeles County, one of the world’s most\nancestrally diverse metropolitan areas, with a population of 9.6 million. To\ndate, the UCLA ATLAS initiative has consented ~250,000 UCLA Health\npatients and has collected biomaterials from ~130,000. ATLAS enrollment\nlargely reflects the composition of UCLA Health patients, who are concentrated\nacross Los Angeles ( Figure 1b ). The EHR,\ninitiated in 2013, enables continuous longitudinal stratification of\nparticipants by disease states, with a mean and median of 8.6 and 7.9 years of\nparticipation per individual. As of November 2024, ATLAS genomic data included\nWES data from 61,797 participants and custom array genotyping using the Illumina\nGlobal Screening Array from 92,164 participants. Extensive quality control (QC)\nindicated high data quality ( Figure S1 ). We leveraged these data to interrogate social and\ngenetic factors that affect disease risk and health outcomes ( Figure 1a ).\nGenotyped cohort demographics are summarized in  Table 1  and the  Methods . Biobank\nparticipants were older in age and had a higher comorbidity index than\nnon-biobank patients (2.9  versus  1.7 mean Elixhauser index;\n1-year post-collection), consistent with a higher number of clinical visits,\nwhich favors enrollment (controlled for data-completeness,  Methods ;\n Figure 1c ). EHR-based phenotypes,\nincluding vital signs, disease diagnoses, and lab tests, were defined through\nharmonization procedures and subsequently validated ( Methods ;  Figure S2a – l ). Common disease\ndomains were led by endocrine/metabolic, cardiovascular, gastrointestinal\ndisorders and neoplasms, reflecting known global health challenges ( Figure 1d ;  Figure S2m – p ;  Methods ) 31 – 33 . The ATLAS EHR data includes a total of\n71,739,582 lab tests (considering complete blood count [CBC], lipid, metabolic,\nhemoglobin A1c [HbA1c], and 25-hydroxyvitamin D panels), with a mean of 540 lab\ntest results per participant ( Table S1 ). Detailed prescription data show a total of 5,952,958\nprescriptions; the 50 most frequently prescribed medications are listed in  Table S1 .\nAlthough we primarily consider genetically determined ancestry in our\nanalyses, rather than the social constructs of race and ethnicity 34 , 35 , we note substantial diversity based on participant\nself-reports. For instance, 62.7% of participants self-identify as White, 12% as\nAsian, 4.6% as Black or African American, 2.6% as Middle Eastern or North\nAfrican, 0.9% as American Indian or Alaska Native and 0.3% as Pacific Islander.\nA substantial proportion (14.8%) report Hispanic or Latino ethnicity ( Table 1 ).\nSelf-identified race/ancestry, a cultural-societal construct, and\ngenetic ancestry are conceptually distinct 29 , 34 , 36  and genetic ancestry must be considered\nto prevent confounding in genetic association studies 35 , 37 . We classified biobank participants into six broad-scale\nancestry populations: European (EUR), African (AFR), South Asian (SAS), East\nAsian (EAS), Admixed American (AMR) and an Unclassifiable (UNC) population,\naligning with the 1000 Genomes Project super-populations 38  ( Figure\n1e – g ;\n Methods ). ATLAS is diverse relative to most other large\nbiobanks ( Methods ). Consistent with self-reported race/ethnicity,\nalmost a third of the biobank participants (32%) were assigned to non-EUR\ngenetic ancestries. There was significant agreement between self-reported race\nand broad-scale ancestry – 99% of self-identified White participants were\nassigned to EUR or AMR populations ( Figure S3a ). Of those with\n“unknown” self-reported race, 46% were assigned to AMR ancestry\nand 45% to EUR.\nWe next asked how health system usage or disease burden varied by\nbroad-scale ancestry, observing that the mean number of encounters varied\nsignificantly across ancestries (P-value = 7.5×10 −138 ,\nANCOVA, adjusted for genetic sex, age and Barriers to Accessing Services\n[BAS] 39 ), with the\nhighest numbers of total encounters in participants from AMR ancestry (adjusted\nmean = 16.2), followed by AFR (15.8) and SAS (14.2), with the lowest values in\nEUR (12.2) and EAS (12.9) ( Figure 1h ). We\nalso assessed health burden using the comorbidity index score 40 , 41 , which was higher in EAS (adjusted mean = 3.7) and AMR\n(3.6) compared to others (all other populations 2.6-2.9;  Figure 1i ; P-value =\n2.6×10 −5 , ANCOVA, adjusted for genetic sex, age, BAS\nand Area Deprivation Index [ADI] 42 ). Similar results for disease distribution across\nancestries were obtained after applying inverse probability weighting (IPW) to\nadjust for participation bias relative to the broader UCLA Health population\n( Figure S3b ;\n Methods ). Increased encounter numbers were only partially\nexplained by elevated comorbidity index scores, indicated by modest correlations\nbetween the two (mean total encounters  vs . Elixhauser\nComorbidity Index: r = 0.32, P-value <2.2×10 −16 ;\nmean hospital encounters  vs . Elixhauser Comorbidity Index: r =\n0.26, P-value <2.2×10 −16 ;  Figure S3c – d ).\nThe population diversity of ATLAS enables the interrogation of the\ncombined genetic and social/environmental effects on the risk of disease\ndiagnoses. To illustrate this, we assessed variation in medical conditions\nacross populations, replicating multiple known associations\n( Methods ;  Table S2 ;  Figure\nS3e ). Further, we identified previously unreported associations\n( Figure 1j ), including a lower risk of\nepilepsy in EAS compared to EUR participants (odds ratio [OR] = 0.46 with 95%\nconfidence interval [0.32, 0.65], P Bonferroni  =\n2.3×10 −4 ), not previously detected 43 , 44 . Similarly, we found significantly reduced risk for\nbipolar disorder in those with AMR ancestry, clarifying conflicting findings in\nprevious smaller studies 45 , 46  (OR = 0.47 [0.40, 0.56],\nP Bonferroni  = 1.5×10 −17 ). We also show a\nsignificantly lower risk of sleep apnea in EAS participants relative to EUR (OR\n= 0.67 [0.59, 0.75], PBonferroni = 4.1×10 −10 ), which\nhas been controversial 47 – 50 , with\na weaker, but significant effect after correcting for body mass index (BMI) (OR\n= 0.84 [0.74, 0.95], P-value = 4.6×10 −3 ;  Table S2 ). These\nassociations remained significant after adjusting for socioeconomic status\n(SES), using both the BAS and ADI measures, and with IPW ( Table S2 ).\nFine-scale ancestries 37 , 51 – 53 , many of which are\nunderstudied 14 – 16 , reflect recent geographic or\ndemographic stratification and can reveal important contributions to health\ndisparities and disease risk 14 . We quantified fine-scale ancestry using identity-by-descent\n(IBD) 53 , 54  in a population nearly three times\nlarger than any previously published study 52 , 53 . This\napproach identifies genomic regions shared between individuals due to a common\nancestor and defines fine-scale ancestries, which we refer to as\n“clusters” ( Methods ) 52 . We identified 36 fine-scale ancestry\nclusters with at least 30 participants ( Figure\n2a – b ;  Figure S4a – c ;  Methods ), labeled\nwith both a numeric identifier ( e.g. , IBD-01) sorted based on\nthe cluster sample size, and a corresponding cluster name to facilitate\ninterpretation ( Methods ). Because PCA captures ancient population\nstructure and IBD segments reflect recent shared ancestry, full agreement\nbetween these approaches is not expected, as is observed for some of our IBD\nclusters, such as IBD-02 ( Figure 2a ). We\nreplicated previously identified clusters and added unreported ones, such as a\nNative Hawaiian cluster (n IBD-29  = 64) and a Bantu cluster\n(n IBD-35  = 32). The largest clusters consisted of Northern\nEuropeans (n IBD-01  = 33,675) and Southern Europeans\n(n IBD-02  = 14,841). The remaining clusters represent the\nheterogeneity of ancestral origins in Los Angeles, including Ashkenazi Jewish\n(n IBD-03  = 14,262), Filipino (n IBD-09  = 1,438),\nIranian Jewish (n IBD-11  = 707), and Armenian\n(n IBD-16,IBD-23,IBD-31  = 560) populations. The diversity of ATLAS\nis reflected not only by a relatively large portion of non-EUR individuals, but\nalso by substantial diversity within the EUR broad-scale ancestry population.\nSeveral EUR clusters are large relative to published datasets, including\nSouthern European, Ashkenazi Jewish, Armenian and Iranian Jewish 3 , 53 , 55 – 57  ( Table S3 ). Similarly, for EAS, we\nidentified multiple clusters, including Filipino, representing the largest\npublished Filipino genotyped cohort of which we are aware 58 . Despite this strength, for some\npopulations, sample sizes are limited, which leads to relatively small\nclusters.\nFor a nuanced understanding of diagnostic variation, we tested the\nprevalence of 1,253 phecodes 59 – 61  across\n24 fine-scale clusters with at least 100 participants.  Table S3  shows all associations and\n Figure 2c  highlights selected\nassociations (also available at  https://atlas-phewas.mednet.ucla.edu/ancestry ). We replicated known\nfindings ( Methods ) and uncovered numerous associations between\nfine-scale clusters and diseases. For instance, Filipinos, an understudied\npopulation in genetics research, exhibited the lowest risk for vitamin B-complex\ndeficiency (OR IBD-09  = 0.50 [0.36, 0.67], FDR =\n2.1×10 −4 ), and the highest risk for cholesterolosis\nof the gallbladder (OR IBD-09  = 4.3 [2.9, 6.2], FDR =\n5.6×10 −12 ), both previously unreported. In Iranian\nJewish subjects, another group with low research inclusion, we detected a\npreviously unreported elevated risk for glaucoma (OR IBD-11  = 1.9\n[1.5-2.4], FDR = 2.9×10 −6 ), a condition known to be\nmore common in individuals of African and East Asian ancestry. We also observed\nseveral shared health risk patterns among Jewish ancestry clusters. Both Iranian\nand Ashkenazi Jewish clusters showed the highest risk for bladder cancer\n(Iranian Jewish: CI IBD-11  = 1.4-5.6, FDR = 0.008; Ashkenazi Jewish:\nOR IBD-03  = 1.5 [1.1, 1.9], FDR = 0.02) and a high hyperplasia of\nprostate risk (Iranian Jewish: OR IBD-11  = 2.3 [1.8, 2.8], FDR =\n1.1×10 −13 ; Ashkenazi Jewish: OR  IBD-03  =\n1.7 [1.6, 1.8], FDR = 4.6×10 −66 ), and the lowest risk\nfor cirrhosis of liver (Iranian Jewish: CI IBD-11  = 0.03-0.4, FDR =\n3.1×10 −2 ; Ashkenazi Jewish: OR IBD-03  =\n0.4 [0.3, 0.5] , FDR = 6.8×10 −24 ). Mexicans and South\nAmericans suffered consistently more from hormones’ adverse effects in\ntherapeutic use, a previously unreported association (Mexican American clusters:\nOR IBD-04,-07,and-12  = 2.0-2.8 [1.3, 3.4], FDR =\n1.6×10 −2  - 2.1×10 −24 ;\ncombined South Americans: OR IBD-08  = 1.9 [1.4, 2.5], FDR =\n1.3×10 −4 ). These associations persisted after SES\nadjustment, using the BAS and ADI ranks. In most cases, associations held after\napplying IPW and also adjusting for SES, except for hormone adverse effects,\nwhere only the largest Mexican American group remained highly significant,\nlikely due to smaller sample sizes resulting from the IPW analyses ( Table S3 ). These findings\ndemonstrate the utility of fine-scale ancestry in identifying disease risk\nprofiles in underrepresented populations.\nWe next focused on cardio-metabolic diseases due to their global impact\non public health. We selected a few targeted cardio-metabolic phenotypes and\ncompared disease risk across fine-scale clusters within the same broad-scale\nancestry ( Methods ;  Figure S4d ). Strikingly, among\nAsian clusters, the Filipino cluster had an elevated risk for all tested\ncardio-metabolic conditions, which remained after adjusting for BMI. Across\nfine-scale EUR ancestries, the larger Armenian cluster showed high\ncardiometabolic disease risk. Iranian clusters, both Jewish and non-Jewish, had\na relatively higher risk of coronary atherosclerosis, hyperlipidemia and type 2\ndiabetes, but not hypertension. Adjusting for SES factors and adding IPW\nmaintained these significant patterns ( Table S3 ), except abdominal aortic\naneurysm in Filipinos, which lost significance, and hyperlipidemia in Iranian\nJewish participants, which did not remain significant only when both SES\nadjustment and IPW were applied- likely due to reduced sample size, although the\nOR remained elevated.\nTo examine the robustness of ATLAS in predicting genetic risk and to\nassess the quality of our dataset, we first tested the utility of PGS in\nstratifying risk for common disorders. We observed high PGS performance in EUR\nindividuals–on average, 18.6% of patients diagnosed with major disorders\nwere in the top PGS decile, rising to 41% for type 1 diabetes ( Figure 3 ;  STAR\nMethods ). Consistent with prior reports 62 – 65 , this predictive power was diminished in non-European\nancestries ( Figure 3 ;\n Methods ).\nNext, we performed phenome-wide association studies (PheWAS) using\nRegenie 66  to map the\ncommon variant architecture of clinical diagnoses and laboratory measurements,\nconducting separate analyses within each of our broad- and fine-scale ancestral\ncohorts ( Methods ; n = 84,110 total participants). After linkage\ndisequilibrium (LD) pruning, we identified 19,431 unique variant-phenotype\nassociations passing a genome-wide significance threshold of\n5×10 −8  ( Table S4 ), with 5,772 passing the\nmost conservative Bonferroni threshold of 6.4×10 −11 \n(5×10 −8  conditioned on 776 tested phenotypes). The\nmajority (83.8%) of associations were detected in only one ancestry at\ngenome-wide significance, though utilizing a more permissive threshold of\n1×10 −4  showed cross-ancestry support for a\nsubstantial portion of associations (54.7%;  Figure S6a – b ). Supporting the validity of the\ndata used for PheWAS, we confirmed ancestry-specific differences in APOE allele\nfrequencies and attenuated Alzheimer’s risk associated with the ε4\nallele in African Americans and replicated known associations such as\n PNPLA3  rs738409-G with non-alcoholic fatty liver in AMR\ngroups 67 \n( Methods ;  Figure S6c – e ). We identified numerous ‘previously unreported\nfindings’, which we defined as associations that have not been reported\nat genome-wide significance in the existing literature, the European Molecular\nBiology Laboratory-European Bioinformatics Institute (EMBL-EBI) GWAS\nCatalog 68 , or summary\nstatistics from UKBB 69 , Taiwan\nBioBank 7 , and\nAoU 5 . Notably, 39.5% of\nassociations were unreported in the European Molecular Biology\nLaboratory-European Bioinformatics Institute EMBL-EBI GWAS Catalog 68  ( Methods ). We\nperformed replication analyses for the top previously unreported findings in\nindependent biobanks, including AoU 5 , UKBB 69 ,\nTaiwan BioBank 7 , and\nBioMe 8  ( Table S4 ;\n Methods ). Summary statistics are available at  atlas-phewas.mednet.ucla.edu .\nWe identified a previously unreported association between\n FN3K  rs7208565-T and increased risk for intestinal\ndisaccharidase deficiency across EUR populations ( Figure 4A ; P EUR  = 3.8 × 10 −41 ;\nOR EUR  = 1.3; MAF EUR  = 33.3%; MAF IBD-01  =\n32.3%; MAF IBD-02  = 34.3%; MAF IBD-03  = 35.6%) and the AMR\npopulation (P AMR  = 4.9 × 10 −11 ;\nOR AMR  = 1.3; MAF AMR  = 42.3%;  Table S4 ). This variant was\nadditionally associated with increased risk of the “other abnormal\nglucose” phecode (P EUR  = 1.8 × 10 −42 ;\nOR EUR  = 1.2), whereas a variant in close LD (rs113373052-T) was\nassociated with increased HbA1c (P EUR  = 5.4 ×\n10 −97 ; β EUR  = 0.13), reflecting\nadditional consequences on glucose homeostasis. Intestinal disaccharidase\ndeficiency is the inability to completely digest sugars such as lactose and\nsucrose, leading to irritable bowel syndrome (IBS)-like symptoms 70 .  FN3K \nphosphorylates glycated proteins, preventing the formation of advanced glycation\nend-products 71 .\nSupportive of these findings, rs7208565 has been previously associated with type\n2 diabetes 72 .\nFurthermore, we discovered several previously unreported associations\ndriven by low-frequency variants (MAF ~1-2%) in non-EUR cohorts. For\ninstance, we linked rs115750084-G in  DPP6 , a gene implicated in\nsynaptic signaling 73  and\nneurological disorders 74 , 75 , to an increased risk of major\ndepressive disorder in the AFR cohort (P AFR  = 4.6 ×\n10 −8 ; OR AFR  = 4.2; MAF AFR  = 1.1%;\n Figure S6F ). In\nthe EAS cohort, we associated rs77742325-G in  MAS1 , a key\ncomponent of the bone-protective ACE2/Angiotensin-(1-7)/Mas axis, 76  with osteoporosis\n(P EAS  = 1.5 × 10 −8 ; OR EAS  =\n3.5; MAF EAS  = 1.8%;  Figure S6G ). We also linked an\nintronic deletion (rs202215133-ATATCATAG>A) in UBR5, an E3 ubiquitin\nligase implicated in neuroinflammatory and neurodevelopmental\nsyndromes, 77 , 78  with migraine headache in the Mexican\nAmerican cluster (P IBD-04  = 1.7 × 10 −8 ;\nOR IBD-04  = 5.6; MAF IBD-04  = 1.1%;  Figure S6H ). These associations\nwere independently replicated, albeit with more modest effect sizes in the\nreplication cohort ( Table\nS4 ). Given that the variants are low MAF, further validation through\ndirect sequencing in large cohorts will be essential to confirm their\ngeneralizability.\nSimilarly, our large AMR discovery cohort revealed several unreported\nassociations. For instance, we linked  STARD7  rs17419569-C with\nincreased risk for asthma in the Mexican American cluster ( Figure 4B ; P IBD-04  = 1.9 ×\n10 −9 ; OR IBD-04  = 2.6; MAF IBD-04  =\n2.5%). Prior work suggests that decreased  STARD7  expression is\nassociated with enhanced allergic responses in the human lung and significant\nincreases in airway hyperresponsiveness in haploinsufficient\n Stard7  mice, 79  whereas rare variants within  STARD7  were\nlinked with asthma (P UKBB  = 5.1 × 10 −3 ;\n Table S4 ). In the\nsame cluster,  GPX7/SHISAL2A  rs74744741-C was associated with\ngastrointestinal reflux disease (GERD) ( Figure\n4C ; P IBD-04  = 5.8 × 10 −10 ;\nOR IBD-04  = 1.6; MAF IBD-04  = 8.9%).\n GPX7  has been associated with carcinogenesis in the context\nof GERD-associated Barrett’s esophagus, 80  whereas  SHISAL2A  has an\nunknown function, but is highly expressed in the small intestine and in lymphoid\ntissues. 81 \nAdditionally, we identified two low-frequency AMR variants associated with\nchronic renal failure (CRF;  Figure 4D ),\nrs112680741-C (P AMR  = 2.9 × 10 −11 ;\nOR AMR  = 3.8; MAF AMR  = 1.3%) and rs2744548-C\n(P AMR  = 9.6 × 10 −11 ; OR AMR  =\n3.2; MAF AMR  = 1.7%), nominating  GPLD1 ,\n ALDH5A1 , and  KIAA0319  as potential CRF\nrisk genes. These associations did not replicate in available biobank\npopulations ( Table\nS4 ), possibly because the ATLAS AMR populations from the Los Angeles area\nrepresent distinct Mexican and South American ancestries 82 , 83 , whereas those from other biobanks (BioMe, for example) are\ncomprised of individuals with distinct ancestral cluster assignments to Puerto\nRico and the Dominican Republic. Other cohort-specific characteristics may also\ncontribute to these differences, so independent replication in additional\ncohorts is needed.\nThis first wave of the ATLAS WES catalog comprises more than 11.5\nmillion autosomal variants, with a median of 9,821 missense and 159\nloss-of-function (LOF) variants per individual, closely matching expectations\nfrom other studies 84 . Most\n(98%) WES variants were rare (MAF < 1% in any broad-scale ancestry),\nincluding 2,873,731 rare missense and 172,462 rare LOF variants, with a median\nof 15 rare LOF and 403 rare missense variants per participant ( Table S5 ). AFR participants showed\nthe greatest number of rare LOF and missense variants, matching prior\nresults 85 . To evaluate\nour WES, we analyzed a predefined set of clinically relevant rare variants known\nto have elevated frequencies in specific populations and demonstrated that ATLAS\ncaptures these established patterns ( Methods ;  Figure S7a – e ). Then, we examined differences\nin 69 ClinPGx pharmacogenomic variants 86 , 87  across\nfine-scale ancestries ( Figure 4e ;  Table S4 ). We identified\nseveral unreported enrichments in specific clusters, including a\n SLCO1B1  variant, rs4149056, which affects the metabolism\nand toxicity of statins 88  in\nAshkenazi and Iranian Jews; a  CYP4F2  variant, rs2108622, that\ncan increase the required dosage of warfarin 89 , enriched in Ashkenazi Jewish, Lebanese, Armenian 1,\nEgyptian Christian, and South Asian 2 individuals; and  PYD \nvariants, which affect the toxicity of chemotherapy drugs 90 , abundant in Mexican American, African\nAmerican and South American clusters. This suggests the importance of\nconsidering fine-scale ancestry in precision drug prescribing.\nNext, we used ATLAS’ ancestral heterogeneity to identify\nenrichments of known rare clinically relevant variants, aggregated by gene\nacross fine-scale clusters to illustrate the utility of fine-scale ancestry in\nidentifying populations at higher risk of rare, monogenic disease. We queried\nthe full set of the curated rare pathogenic/likely-pathogenic (P/LP)\nClinGen 91  variants\nthat have strong clinical and genetic evidence (n = 643), identifying 5,223\nunrelated carriers ( Figure 5a ). Alongside\nknown associations ( STAR Methods ), we\ndetected a previously unreported elevated carrier frequency for variants in the\nCenters for Disease Control and Prevention (CDC) Tier 1 gene,\n LDLR , which causes familial hypercholesterolemia (FH) in\nthe Filipino cluster (OR IBD-09  = 3.9 [1.4, 8.8], FDR = 0.03). This\nfinding is notable given the high prevalence of dyslipidemia in\nFilipinos 92  ( Figure S4D ). Indeed, LDLR\nFH variants accounted for 3.1% (population attributable risk [PAR]) of high LDL\ncases (LDL > 190 mg/dl) in the Filipino cluster. However, due to the low\nnumber of carriers in this cluster (<10), the PAR CI was wide [0.3,\n26.1], and independent replication is required to refine this finding.\nAdditionally, we identified previously unreported elevated carrier frequency of\nvariants that cause non-syndromic genetic deafness within two different genes in\ntwo separate clusters: (1)  MYO15A  in the Mexican American 1\ncluster (OR IBD-04  = 6.9 [2.6, 17.4], FDR =\n8.44×10 −4 ) and (2)  CDH23  in the\nChinese + Korean (OR IBD-05  = 6.0 [2.0, 15.6], FDR =\n6.04×10 −3 ) and Japanese (OR IBD-10  = 28.1\n[8.9, 77.1], FDR = 7.4 × 10 −6 ) clusters. We also\nidentified previously unreported elevated carrier frequencies in the Ashkenazi\nJewish cluster for Pendred syndrome variants in  SLC264A \n(OR IBD-03  = 3.8 [2.9, 5.0], FDR = 8.6 ×\n10 −19 ) and recombinase activating gene 2 deficiency in\n RAG2  (OR IBD-03  = 2.9 [1.4, 5.8], FDR = 0.02).\nFinally, we focused on the clinically actionable American College of Medical\nGenetics and Genomics (ACMG) 93 \ngenes, where ClinGen-curated variants within these genes manifested a strong EUR\nbias ( Figure 5b ;  Figure S7f ;  Methods ).\nThis bias was recently shown in AoU 23  and we provide additional supporting evidence by\ndemonstrating this bias in a single health system.\nAs clinical variant databases are primarily ascertained from EUR\npopulations, we took a conservative approach to call rare predicted deleterious\nLOF (dLOF) and missense coding variants (dMIS). High-confidence dLOF were\ndefined using the loss-of-function transcript effect estimator\n(LOFTEE) 94 , which\naccounts for biological context beyond protein truncation. dMIS variants were\ndefined using a consensus of a majority (5 of 9) of state-of-the-art\ncomputational methods based on the Critical Assessment of Genome Interpretation\n(CAGI) project 95  for\npathogenicity prediction ( Methods ;  Figure S7g – h ). This approach is distinct from\nthe one taken with ClinGen variants – variants identified in this manner\nmay not have literature evidence for their impact on gene function but are\npotentially less biased by EUR oversampling.\nFirst, we kept our focus on ACMG genes 93 , repeating the ClinGen-based analysis\ndescribed above using predicted rare dLOF and dMIS variants. We examined the\nfrequency of 1,602 rare dLOF and 6,487 rare dMIS variants within ACMG variants\nacross populations ( Figure 5c ;  Figure S7i ). The results\nwere consistent with previous work showing that AFR participants harbor\nsubstantially more dLOF compared to other ancestries 85 , and highlighted other groups with\npreviously unreported low or high ACMG putative damaging variants for further\ninvestigation. Overall, these results contrasted in comparison with rare\nvariants identified in ClinGen ( Figure 5b ),\nconsistent with the interpretation that ClinGen variant frequencies are biased\ntoward EUR participants, as is the case for most curated clinical datasets.\nSecond, to examine the influence of dLOF and dMIS on ATLAS phenotypes,\nwe performed exome-wide association studies (ExWAS) across 17,537 protein-coding\ngenes ( Methods ). Using Regenie’s unified gene burden\nassociation strategy 66 , we\nidentified 1,099 unique Bonferroni-significant gene-trait associations. Within\nthe EUR population, we replicated numerous known associations; 45 of the top 50\nassociations had been previously reported 84  ( Table S6 ). These include  PKD1  with cystic kidney\ndisease and  TTN  with primary intrinsic cardiomyopathy ( Figure 5e ) 84 . We also identified numerous previously\nunreported associations ( Table S6 ). In the Northern European cluster, we detected gene-level\nassociations between  HNRNPA1L2  and acquired absence of the\nbreast (GENE_PIBD-01 = 1.0×10 −25 ), breast cancer\n(GENE_PIBD-01 = 2.2× 10-8 ), and malignant neoplasm of the\nfemale breast (GENE_P IBD-01  = 1.5 × 10 −7 ).\nNotably,  HNRNPA1L2  is on chromosome 13, ~20 megabases\nfrom  BRCA2 , representing a different locus. Though\n HNRNPA1L2  has not been previously associated with breast\ncancer, it is highly expressed in breast invasive carcinoma 81  and a paralog  HNRNPA1 \nwas previously implicated in breast cancer progression 96 . In the Ashkenazi Jewish (IBD-03)\ncluster, we observed associations between  EPG5  and HDL\ncholesterol level (P LOF_MIS  = 3.6×10 −10 ;\nβ LOF_MIS  = 1.8 [1.2, 2.3]) and triglyceride level\n(P LOF_MIS  = 1.33×10 −8 ;\nβ LOF_MIS  = −1.81 [−2.4, −1.2]).\n EPG5  is an autophagy tethering factor classically\nassociated with Vici syndrome 97 , a severe developmental disorder, though our findings\nsuggest an additional role in lipophagy and lipid metabolism 98 .\nIn the AMR population, we identified an unreported association between\ndLOF and dMIS variation in  CLN3  and cystic kidney disease\n(GENE_P AMR  = 8.7×10 −9 ), supported by\nexperimental evidence that  CLN3  is highly expressed in\nmedullary collecting duct principal cells and plays a role in\nosmoregulation 99 . We\nadditionally identified unreported associations between  PPARG \nand viral pneumonia (P AMR,MIS  = 1.1×10 −6 ,\nOR AMR,MIS  = 53.0 [13.5, 208.5]), as well as\n NADSYN1  and abnormal lung examination findings\n(P AMR,MIS  = 1.2×10 −5 ,\nOR AMR,MIS  = 3.7 [2.1, 6.5]).  PPARG  is known to\nregulate macrophage response to pulmonary inflammation 100 , while  NADSYN1 \nknockout in mouse models was shown to cause abnormal lung development 101 , though neither has been\npreviously associated with respiratory traits in humans.\nWithin the AFR population and African American cluster, we uncovered\nunreported associations between  DDHD2  and dysphagia\n(P IBD-06,LOF  = 1.2×10 −6 ,\nOR IBD-06,LOF  = 8.3 [3.7, 18.7]) and  EFCAB13  and\nessential hypertension (P AFR,MIS  = 9.3×10 −7 ,\nOR AFR,MIS  = 13.1 [4.5, 37.7]). These findings align with prior\nstudies; variants in  DDHD2  have been shown to cause hereditary\nspastic paraplegia and symptoms of dysphagia 102 , though not in African Americans, and\nmethylation studies have prioritized  EFCAB13  as a risk factor\nfor heart failure 103 . We\ndetected an additional association between  ANKZF1  and\nperipheral vascular disease (P AFR,LOF_MIS  =\n1.6×10 −6 , OR AFR,LOF_MIS  = 58.8 [12.6,\n275.1]), which is supported by recent experimental findings 104 .\nNext, in the EAS population, we observed a previously unreported\nassociation between  AK7 a nd non-rheumatic mitral valve\ndisorders (P EAS,LOF  = 7.6×10 −7 ,\nOR EAS,LOF  = 85.7 [39.2, 187.2]). While  AK7  has\nnever been associated with cardiovascular phenotypes, it has a well-established\nrole in primary ciliary dyskinesia 106  and mutations in other ciliary genes have been\nassociated with mitral valve prolapse 107 . We further identified an unreported association\nbetween  KIF2B  and nephritis/nephropathy in the Chinese/Korean\ncluster (P IBD-05,MIS  = 7.2×10 −7 ,\nOR IBD-05,MIS  = 26.2 [9.0, 76.4]). Though  KIF2B \nhas not been linked to renal traits, other genes in the kinesin family (KIF)\nhave been implicated in the pathogenesis of renal cancer 108 , 109 . All previously unreported findings described above were\nindependently replicated in AoU 5 , UKBB 69  or\nBioMe 8  at a nominal\nP-value < 0.05 ( Table\nS6 ).\nFinally, we examined known risk associations to identify\nancestry-specific patterns of deleterious rare variation. We tested two known\n GBA1  rare missense variants, p.Glu365Lys and p.Thr408Met,\nin the AMR and EUR populations; both variants are reported to increase risk for\nParkinson’s disease 110 . Despite similar allele frequencies within each population,\np.Glu365Lys was only associated with increased risk in EUR\n(P EUR,E365K  = 0.01; P AMR,E365K  = 0.8), whereas\np.Thr408Met was only associated with increased risk in AMR participants\n(P EUR,T408M  = 0.2; P AMR,T408M  = 0.003), displaying\nevidence of ancestry-specific risk stratification ( Figure S7j ). While we hypothesize\nthat penetrance is modulated by population-specific genetic backgrounds, further\nstudies are required to identify the causes of these differential associations.\nIn total, We performed replication analyses for 21 of our top previously\nunreported PheWAS and ExWAS findings in independent biobanks and replicated 16\n(76%) at a nominal p value of 0.05.\nOne of the advantages of EHR data is the ability to consider dynamic\nchanges over time. To illustrate this, we integrated all main study components\n(genetic ancestry, common and rare genetic variants) to study the efficacy of\nglucagon-like peptide-1 (GLP-1) receptor agonists (GLP1-RAs) for weight loss. We\nfocused on semaglutide, which had the greatest number of prescriptions in ATLAS\n( Figure\nS8a – b ; 7,340 participants;  Figure 6a ).\nWe observed a steady decrease in average weight up to ~60 weeks of\nsemaglutide treatment, consistent with previous findings 111 – 113  ( Figure S8c ). We tested if the medication dose, route, sex, age, and\ninitial BMI affected semaglutide efficacy, corroborating that dose and\nsubcutaneous delivery versus oral delivery were positively correlated with\nweight loss ( Figure 6b ; linear\nmixed-effects model [LMM]: dose: P Bonferroni  =\n3.3×10 −97 , semipartial R² (explained\nvariance) = 1.8%, effect size = −1.0; route oral  vs. \nsubcutaneous: PBonferroni = 7.2×10 −37 , semi-partial\nR 2  = 1.1%, effect size = 2.2). We did not detect a significant\neffect of age and sex on efficacy.\nNext, we asked if semaglutide’s effects varied by broad-scale\nancestry, finding that ancestry significantly influenced weight loss over 60\nweeks of treatment ( Figure 6c ; ANOVA on an\nLMM; ancestry: P Bonferroni  = 0.002; ancestry×time:\nP Bonferroni  = 0.0006). Comparisons between EUR and the other\npopulations revealed less weight loss in AMR, and a slower rate of weight loss\nin AMR and EAS ( Figure 6d ; LMM;\nancestry AMR : P Bonferroni  =\n2.9×10 −3 , β = 0.8;\nancestry AMR ×time: P Bonferroni  =\n5.0×10 −3 , β = 45.2; ancestry EAS \n×time: P Bonferroni  = 3.2×10 −3 ,\nβ = 76.4), consistent with recent analysis based on a single time point,\nwhich showed reduced efficacy in AMR and AFR cohorts (here, in AFR, the result\nwas not significant but the direction of effect was consistent). 113\nGiven limited evidence on the role of inherited factors in GLP1-RAs\neffectiveness 114 , we\nused ATLAS to explore the contribution of common genetic variation to\nsemaglutide-related weight loss. We found no correlation between BMI PGS and\nsemaglutide efficacy. However, weight loss was negatively correlated with DM2\nPGS ( Figure 6e ; LMM, PGS High\n vs.  Low : P Bonferroni  =\n2.8×10 −4 , β = 0.96; PGS Med\n vs.  Low : P Bonferroni  =\n8.8×10 −3 , β = 0.73). The same relationship\nwas observed in a simplified model using only maximum weight loss\nrecorded 114  (linear\nregression, P Bonferroni  = 1.3×10 −2 , β\n= 0.31;  Figure S8d ). We\ncould not find an equally sized cohort for replication, but we used EUR AoU,\nwhich had a sample size 50% smaller than the ATLAS cohort used in this analysis\nwith array genotyping (1,578  vs . 3,165 in AoU and ATLAS). In\nAoU EUR participants, we observed a concordant direction of effect, supporting\nreplication of the finding, although statistical significance was not reached.\nThis may reflect the smaller sample size or other population differences, such\nas in SES and compliance ( Figure S8e ). Larger samples are needed for more formal replication.\nAs a second step, we conducted a genome-wide association study (GWAS)\nmeta-analysis across ancestries, but did not identify any significant loci\n( Methods ;  Figure S8f ).\nFinally, we used WES to test gene-level associations with semaglutide\nresponse (Regenie 66 ;\n Methods ). Given the modest sample size, we focused on EUR\nindividuals and limited our test to the subset of proteins whose plasma\nabundance in humans was recently shown to be altered by semaglutide 115 . We identified a significant\nassociation of weight loss on semaglutide with  PTPRU \n(P Bonferroni  = 7.6×10 −3 , β =\n−0.87) ( Figure 6f ;  Figure S8g ). For replication, we\nused AoU, with a cohort that was 21% smaller than the ATLAS cohort used (1,581\n vs . 2,012 in AoU and ATLAS with WES). This analysis\nsupported our original finding by showing a consistent direction of effect,\nalthough significance was not reached, possibly due to the smaller sample size\nor differences between the tested populations (AoU: β = −0.27,\nP-value = 2×10 −1 ). Therefore, larger sample sizes in\nfuture studies will be needed to formally replicate this finding. However, the\nBonferroni adjusted combined P from ATLAS and AoU remained significant,\nsupporting this finding (P Bonferroni  =\n2.3×10 −2 ). This association involves 37 variants,\nof which one is common (rs2235937; nominally associated with weight loss; P =\n9.7×10 −3 , β = −0.06 in EUR), and the\nrest are rare, the frequency of which varied across ancestries ( Figure S8h ;  Methods ).\nThe direction of effect indicates that  PTPRU  activity is\nnegatively associated with weight loss in patients taking semaglutide (variants\nin  PTPRU  with a predicted high or moderate functional\nconsequence contribute to weight loss on semaglutide). Maretty  et\nal.  highlighted PTPRU as a protein whose abundance is elevated in\nindividuals with higher genetic risk for increased BMI or type 2 diabetes and is\ndownregulated by semaglutide treatment 115 . Although direct causality has not been established,\nthese observations, together with our findings, support a model in which\nsemaglutide-induced downregulation of  PTPRU  and genetically\nreduced PTPRU function promote greater weight loss during semaglutide treatment.\nThis gene has no previous known functional relationship to weight loss or\nrelated metabolic functions. However, combined with the data from serum\nproteomics 115 , our\ngenetic analysis nominates this protein kinase and its pathways as candidates\nfor further investigation.\n\nLarge-scale, longitudinal population studies have accelerated our\nunderstanding of the causes and consequences of a wide variety of biomedical\nconditions. The Framingham study 116 , 117  laid the\nfoundation for population-based cohorts, followed by efforts like the UKBB and\nhealth system biobanks 3 , 8 , 9 , 11 , 53 , 72 , 84 , 118 , 119 . Despite the known sources of\nerrors and incomplete phenotyping in health system EHRs, multiple studies have shown\nthat many of these factors can be mitigated, allowing robust analyses 3 , 9 , 10 , 120 . Indeed, we were able to validate dozens of\npreviously identified genetic associations in ATLAS based on EHR-derived phenotypes\nalone. We also leverage the heterogeneity of genetic ancestries in our population to\nvalidate and extend our knowledge of disease burden and genetic risk factors.\nPrior work has mostly relied on broad-scale genetic ancestry to estimate\nhealth risk. 5 , 30 , 117 \nFine-scale ancestry discerns more recent genetic variation  via \nshared IBD, 120  but smaller\nclusters can limit power. Through PheWAS and ExWAS across broad- and fine-scale\nancestries, we identified numerous unreported genetic associations, many of which\nwere subsequently replicated. We leveraged fine-scale ancestry definitions to\nbenefit populations that are underrepresented in genetics and medical\nresearch 122 . For example,\nto our knowledge, ATLAS consists of the largest genetic dataset of Filipino\nindividuals (n = 1,438  vs.  1,028 in the next largest\ndataset 58 ), which can be\nused to improve precision care for this population and fuel discovery. Similarly,\nour Ashkenazi Jewish cluster of 14,261 participants, which is 2.8 times larger than\nany prior single-study genetic cohort 57 , enabled previously unreported associations and enhanced risk\nquantification. Future work should assess the transferability of our results across\npopulations 123 .\nDiscerning genetic variant risk frequency and effect across fine-scale\nancestries is powerful for optimized risk stratification and tailoring therapeutic\nintervention. It has been well demonstrated that diagnostic misclassification due to\ndifferences in variant frequencies across ancestries is a serious risk 22 . We emphasize that leveraging\ndiversity within a single regional biobank can reduce confounding by ancestral\ngenetic differences and non-genetic variables that differ across countries or\ngeographically distinct biobanks. We demonstrated the application of longitudinal\nEHR to gain insights into health outcomes over time, by linking genetic factors to\nweight loss on semaglutide. This involved integrating dynamic EHR changes with\ndetailed prescription information, including the medication dose and route, which we\nidentified as major confounders. We considered only a single GLP-1RA drug,\nsemaglutide, to avoid the complexity introduced by different GLP-1RAs, which have\ndiffering efficacy 124 . Lastly,\nour study highlights fine-scale populations with unreported low or high ACMG\nputative damaging variants for further investigation. We expect ATLAS to grow and\nbecome an engine for clinical intervention, allowing researchers to rapidly identify\nindividuals for clinical studies and to implement precision medicine as the field\nevolves.\nCertain limitations should be considered. First, due to the nature of\nEHRs, our study captures a partial view of patient phenotypes and care. Second,\nour analysis focuses on the risk of receiving a diagnosis for a disease, which\nis related, but not equivalent to the underlying risk of disease development.\nThis may influence how our findings should be interpreted in the context of true\ndisease incidence. Third, we relied on computational predictors for comparing\nthe abundances of predicted damaging variants in ACMG genes across ancestries.\nSome predictors may introduce ancestry-related biases 125 , 126 , potentially affecting the accuracy of the results. To\nmitigate these potential biases, we applied rigorous filtering and only retained\nvariants upon concordant agreement in a majority of nine tools. Additionally,\ncomparing the numbers of predicted damaging variants across ancestries may be\nless stable when small ancestry groups are involved. We therefore excluded\ngroups with fewer than 400 individuals and reflected uncertainty using error\nbars. Lastly, while our study introduces multiple previously unreported\nassociations, findings that could not be replicated should be further evaluated\nin independent cohorts as they become available.\n\nQuestions should be directed to the lead contact, Daniel H. Geschwind\n( dhg@mednet.ucla.edu ).\nNo materials were generated in this study.\n\nThe Institutional Review Board of the University of California, Los\nAngeles gave ethical approval for this work. The IRB Number for the ATLAS\nInitiative protocol is IRB#17-001013.\nHuman participants included in this study were from the UCLA ATLAS\nCommunity Health Initiative. Enrollment procedures, consent, and data collection\nare described in Johnson et al 30 . Participants provided informed consent for research under\nIRB#17-001013. The cohort description is provided below and summarized in  Table 1  and  Figure 1 .\nATLAS enrollment reflects a modest over-representation of females by\nself-reported sex, especially among patients aged 20-60 ( Figure 1c ;  Table\n1 ). Among ATLAS participants, the median age at the most recent time\npoint is 61 years for males and 56 years for females. Medical morbidity was\nnotably more widespread in the biobank population compared to all adult UCLA\npatients. This point was evidenced by a significantly larger number of diagnoses\nin the biobank population compared to all UCLA patients (15.6\n vs.  10.5 mean ICD codes per patient, one year after\ncollection). The table “Baseline demographic information in ATLAS and in\nnon-biobank UCLA patients” shows baseline demographic information of the\nUCLA ATLAS population, and a subset of all other adult UCLA patients (Rest of\nDDR) seen within one year of the ATLAS launch date. All table variables differed\nsignificantly between the ATLAS and non-ATLAS populations (Mann-Whitney test for\ncontinuous variables and chi-square test for categorical variables) after\nmultiple-testing correction. This included a significant difference overall, as\nwell as in the proportions of alive and deceased individuals.\nThe biobank population experienced a substantially higher number of\ndiagnoses across the most prevalent medical condition categories, including\nendocrine, cardiovascular, musculoskeletal, gastrointestinal, neoplasms,\nneurological, respiratory, mental, genitourinary and infections ( Figure S2o – p ). For instance,\nendocrine/metabolic and cardiovascular diseases were diagnosed in 52.0% and 44%\nof the biobank participants versus 38% and 36% of non-biobank adults,\nrespectively. Biobank participants also experienced substantially more neoplasms\n(30%  vs.  19%). This likely led to significantly more medical\nencounters in the biobank population (119  vs.  64 mean total\nmedical encounters per patient, and 16  vs.  12 mean inpatient\ndays), which is consistent with expectations, as greater numbers of clinical\nvisits facilitate the passive blood collection that powers our biobank. The most\ncommon diagnoses included hypertension (30% of participants), hyperlipidemia\n(25%), GERD (18%), anxiety disorder (17%), and depression (15%) ( Figure S2m ).\nAs time progresses, more diagnoses have an opportunity to be added, and\nthe rate of diagnoses is consistent across a wide spectrum of organ systems (1-2\nyears after collection;  Figure\nS2n ). The table “Characteristics over time” shows\ntime-dependent characteristics of the UCLA ATLAS population (n=59,949) at the\ntime of collection, and a subset of all other adult UCLA patients seen within\none year of the ATLAS launch date (“Rest of DDR”, within 1 year of\nATLAS launch date; n=186,895). Both cohorts were limited to a subset of\nindividuals who had an encounter between 1-2 years after the initial encounter.\nEncounter statistics (total encounters, total inpatient encounters, total\ninpatient days, types of encounters) summarize the encounters within the first\nyear after the initial encounter. The changes within each cohort were tested for\na non-zero difference using a Mann-Whitney test for variables available at two\ntimes, and a one-sample t-test for the encounter statistics. * denotes a\nsignificant difference in the change. The cohorts were also compared at each\ntime point using the Mann-Whitney test (continuous) and the Chi-square test\n(categorical). After multiple-testing correction, ** denotes a significant\ndifference between the two populations at the respective time (including\nchange).  Baseline demographic information in ATLAS and in non-biobank\nUCLA patients ATLAS-Overall ATLAS-Alive ATLAS-Deceased Rest of DDR-Overall Rest of DDR-Alive Rest of DDR-Deceased n (%) 88,436 84,593 (95.7) 3,843 (4.3) 410,041 392,915 (95.8) 17,126 (4.2) Age, mean (SD) 54.6 (17.0) 54.1 (16.8) 67.6 (14.8) 51.8 (18.5) 50.9 (18.1) 72.8 (14.8) Self-reported sex, n (%):\nFemale 49,346 (55.8) 47,688 (56.4) 1,658 (43.1) 234,189 (57.1) 225,735 (57.5) 8,454 (49.4) Self-reported sex, n (%):\nMale 39,058 (44.2) 36,873 (43.6) 2,185 (56.9) 175,841 (42.9) 167,169 (42.5) 8,672 (50.6) Self-reported sex, n (%):\nUnspecified, X 32 (0.0) 32 (0.0) 11 (0.0) 11 (0.0) Self-reported race, n (%):\nAmerican Indian, Alaska Native 785 (0.9) 762 (0.9) 23 (0.6) 2,233 (0.5) 2,175 (0.6) 58 (0.3) Self-reported race, n (%):\nAsian 10,537 (11.9) 10,180 (12.0) 357 (9.3) 42,307 (10.3) 40,619 (10.3) 1,688 (9.9) Self-reported race, n (%):\nBlack, African American 4,079 (4.6) 3,857 (4.6) 222 (5.8) 23,420 (5.7) 22,256 (5.7) 1,164 (6.8) Self-reported race, n (%):\nCaribbean/West Indian 161 (0.2) 161 (0.2) 335 (0.1) 332 (0.1) 3 (0.0) Self-reported race, n (%):\nMiddle Eastern or North African 2,276 (2.6) 2,227 (2.6) 49 (1.3) 5,086 (1.2) 5,017 (1.3) 69 (0.4) Self-reported race, n (%):\nNative Hawaiian, Guamanian or Chamorro, Samoan, Other\nPacific Islander 255 (0.3) 247 (0.3) 8 (0.2) 1,050 (0.3) 1,011 (0.3) 39 (0.2) Self-reported race, n (%):\nOther Race 3,184 (3.6) 2,737 (3.2) 447 (11.6) 42,188 (10.3) 40,137 (10.2) 2,051 (12.0)  Self-reported race, n\n(%): Unknown, declined to specify 11,329 (12.8) 11,050 (13.1) 279 (7.3) 62,735 (15.3) 61,692 (15.7) 1,043 (6.1)  Self-reported race, n\n(%): White 55,830 (63.1) 53,372 (63.1) 2,458 (64.0) 230,687 (56.3) 219,676 (55.9) 11,011 (64.3) Self-reported ethnicity, n (%):\nHispanic/Latino, Cuban, Hispanic/Spanish origin, Mexican,\nMexican American, Chicano/a, Puerto Rican 12,966 (14.7) 12,235 (14.5) 731 (19.0) 50,229 (12.2) 48,100 (12.2) 2,129 (12.4)  Self-reported\nethnicity, n (%): Non-Hispanic/Latino 68,784 (77.8) 65,829 (77.8) 2,955 (76.9) 301,251 (73.5) 287,269 (73.1) 13,982 (81.6)  Self-reported\nethnicity, n (%): Unknown, declined to specify 6,686 (7.6) 6,529 (7.7) 157 (4.1) 58,561 (14.3) 57,546 (14.6) 1,015 (5.9) Self-reported smoking, n (%):\nFormer 24,054 (27.2) 22,504 (26.6) 1,550 (40.3) 92,364 (22.5) 85,576 (21.8) 6,788 (39.6) Self-reported smoking, n (%):\nNever 60,477 (68.4) 58,366 (69.0) 2,111 (54.9) 275,883 (67.3) 266,871 (67.9) 9,012 (52.6) Self-reported smoking, n (%):\nPassive Smoke Exposure - Never Smoker 112 (0.1) 105 (0.1) 7 (0.2) 826 (0.2) 769 (0.2) 57 (0.3) Self-reported smoking, n (%):\nSmoker 3,412 (3.9) 3,286 (3.9) 126 (3.3) 27,509 (6.7) 26,796 (6.8) 713 (4.2)  Self-reported smoking,\nn (%): Unknown, declined to specify 381 (0.4) 332 (0.4) 49 (1.3) 13,459 (3.3) 12,903 (3.3) 556 (3.2) Types of Encounters, n (%):\nInpatient and Outpatient Encounters 28,879 (32.7) 25,725 (30.4) 3,154 (82.1) 72,278 (17.6) 62,279 (15.9) 9,999 (58.4)  Types of Encounters, n\n(%): Only Outpatient Encounters 59,557 (67.3) 58,868 (69.6) 689 (17.9) 337,763 (82.4) 330,636 (84.1) 7,127 (41.6) Total Encounters, mean\n(SD) 118.8 (146.3) 112.1 (137.5) 267.4 (230.9) 64.2 (95.8) 59.7 (88.1) 166.7 (174.3) Total Inpatient Encounters,\nmean (SD) 0.8 (2.2) 0.7 (1.9) 3.8 (5.0) 0.4 (1.3) 0.3 (1.1) 2.2 (3.7) Total Inpatient Days, mean\n(SD) 16.2 (33.2) 13.4 (28.6) 38.3 (53.6) 12.3 (25.8) 9.4 (18.9) 29.7 (47.2) \n Characteristics over time Variable ATLAS- At Collection ATLAS- 1 Year after\nCollection ATLAS-Change Rest of DDR-At Collection Rest of DDR-1 Year after\nCollection Rest of DDR-Change Age, mean (SD) 56.4 (16.4) 57.6 (16.3) 1.2 (0.4)* 53.5 (17.9)** 54.8 (17.8)** 1.3 (0.5)*,** BMI, mean (SD) 27.3 (6.2) 27.3 (28.2) 0.1 (28.7) 26.8 (24.0)** 26.7 (14.8)** −0.1 (27.6) Diastolic BP, mean (SD) 75.5 (9.7) 75.4 (9.4) −0.1 (10.4) 75.4 (10.2)** 75.6 (10.1) 0.2 (10.5)*,** Systolic BP, mean (SD) 125.0 (16.7) 124.8 (16.8) −0.2 (17.1) 125.3 (17.6) 125.3 (17.6) 0.0 (16.8)** Number of ICD Codes, mean\n(SD) 14.2 (11.2) 15.6 (11.9) 1.5 (6.5)* 9.4 (8.4)** 10.5 (9.1)** 1.1 (5.1)*,** Number of Phecodes, mean\n(SD) 7.1 (5.7) 7.8 (6.0) 0.7 (3.2)* 4.7 (4.3)** 5.2 (4.6)** 0.5 (2.6)*,** Charlson Comorbidity Index,\nmean (SD) 1.0 (1.8) 1.1 (1.9) 0.1 (1.0)* 0.7 (1.5)** 0.7 (1.5)** 0.1 (0.8)*,** Elixhauser Comorbidity Index,\nmean (SD) 2.6 (8.3) 2.9 (8.7) 0.3 (4.4)* 1.6 (6.8)** 1.7 (7.1)** 0.2 (3.6)** Total Encounters, mean\n(SD) 23.2 (27.2)* 12.0 (15.8)*,** Total Inpatient Encounters,\nmean (SD) 0.2 (0.7)* 0.1 (0.4)*,** Total Inpatient Days, mean\n(SD) 1.4 (7.8)* 0.4 (4.2)*,** Inpatient and Outpatient\nEncounters (%) 8,413 (14.0) 9,869 (5.3)** Only Outpatient Encounters\n(%) 51,536 (86.0) 177,026 (94.7)**\nBaseline demographic information in ATLAS and in non-biobank\nUCLA patients\nCharacteristics over time\nWe retrieved the latest results of the 39 most common laboratory tests:\nblood count, lipid panel, metabolic panel, HbA1c, and vitamin D,25-Hydroxy to\nexplore test frequencies and variation in results. The largest variability in\nadults was observed for bilirubin (median = 0.4, IQR = 0.3), followed by\ntriglycerides (median = 90, IQR = 66), alanine aminotransferase (median = 22,\nIQR = 14), neutrophils (median = 3.73, IQR = 2.07), and LDL cholesterol (median\n= 97, IQR = 50). Similarly, we retrieved information on the 50 most abundant\nfilled prescriptions to learn about treatment and prescribing tendencies. ( Table S1 ). As hospital\nvisits and surgeries were the most common types of encounters in the biobank\npopulation, it is not surprising that among the most prescribed generic\nmedications were Acetaminophen (n = 738,061 prescriptions), Ondansetron HCl (n =\n636,165), Propofol (n = 633,830), and Lidocaine HCl (n = 492,838) ( Table S1 ).\nATLAS consists of 32% participants of non-European ancestry and a\nrepresentation of five broadscale and 36 fine-scale genetic ancestry groups\nwithin a single health system (see  METHOD\nDETAILS  to understand how these were defined). By comparison, the\nvast majority of the UK biobank 4  participants are EUR (95.8% EUR, 2.1% South Asian, and 2.1%\nAFR). FinnGen 6  is ~98% Finnish\nEUR, Taiwan Biobank 7  is over\n99% EAS Han Chinese, Michigan Genomics Initiative 11  is 88% EUR, and Geisinger\nMyCode 169  is\napproximately 97% EUR. While others like All of Us, Mt. Sinai’s BioMe,\nand Vanderbilt’s BioVU reflect diverse urban populations, ATLAS captures\na wider and more detailed range of genetic ancestries in the Los Angeles\npopulation. All of Us 5  includes\na larger portion of Black participants, but a smaller portion of Asian and\nMiddle Eastern/North African self-identified individuals (under a combined race\nand ethnicity category: White 53.3%, Black or African American 21.2%, Hispanic\nor Latino 17.8%, Asian 3.1%, Middle Eastern or North African 0.6%). Mt.\nSinai’s BioMe 53 , the\nonly biobank with reported fine-scale ancestries, included 17 fine-scale\nclusters  vs.  36 in ATLAS, and substantially fewer Asian\nindividuals. Vanderbilt’s BioVU 170  and Penn Medicine Biobank 119  include small Admixed American\npopulations, and these biobanks, along with the VA Million Veterans\nProgram 72  and the\nColorado biobank include small Asian populations. The tables below show\ncomparisons of ancestry numbers and ratios across biobanks. Unclassified\nparticipants are not included. Some biobank studies have larger sample sizes but\nlack inferred genetic ancestry and are thus not included.  Numbers of biobank participants by genetic ancestry Biobank AMR AFR EUR EAS SAS Total UK Biobank (National) - 9,633 431,805 - 9,252 450,690 All of Us (National) 28,901 34,037 101,613 3,255 - 167,806 Taiwan Biobank (National) - - - 108,955 - 108,955 Mount Sinai BioMe\n(Academic) 10,638 6,983 8,477 780 617 27,495 Vanderbilt BioVU\n(Academic) 2,466 15,597 69,810 896 414 89,183 Geisinger MyCode\n(Academic) - 1,377 44,522 - - 45,899 Michigan Genomics\n(Academic) 843 5,962 75,943 2,172 1,435 86,355 VA Million Veterans Program\n(Academic) 59,048 121,117 449,042 6,702 - 635,909 Penn Medicine Biobank\n(Academic) 711 11,300 30,360 680 573 43,624 Colorado Biobank\n(Academic) 18,137 7,466 145,070 4,080 - 174,753 UCLA\nATLAS 12,822 4,061 62,902 8,080 1,761 89,626\nNumbers of biobank participants by genetic ancestry\nBelow, percentages exclude unclassified participants and therefore\ndiffer from  Figure 1 .  Percentages of biobank participants by genetic ancestry Biobank AMR AFR EUR EAS SAS UK Biobank (National) 0.0% 2.1% 95.8% 0.0% 2.1% All of Us (National) 17.2% 20.3% 60.6% 1.9% 0.0% Taiwan Biobank (National) 0.0% 0.0% 0.0% 100.0% 0.0% Mount Sinai BioMe\n(Academic) 38.7% 25.4% 30.8% 2.8% 2.2% Vanderbilt BioVU\n(Academic) 2.8% 17.5% 78.3% 1.0% 0.5% Geisinger MyCode\n(Academic) 0.0% 3.0% 97.0% 0.0% 0.0% Michigan Genomics\n(Academic) 1.0% 6.9% 87.9% 2.5% 1.7% VA Million Veterans Program\n(Academic) 9.3% 19.0% 70.6% 1.1% 0.0% Penn Medicine BioBank\n(Academic) 1.6% 25.9% 69.6% 1.6% 1.3% Colorado BioBank\n(Academic) 10.4% 4.3% 83.0% 2.3% 0.0% UCLA\nATLAS 14.3% 4.5% 70.2% 9.0% 2.0%\nPercentages of biobank participants by genetic ancestry\n\nOur study moves from characterizing the ATLAS cohort to ancestry-stratified\nanalyses of disease burden – first at the broad-scale ancestry level, then at\nthe fine-scale cluster level- followed by evaluations of common and rare genetic\nrisk, and finally, an integrative case-study using longitudinal EHR data.\nArray genotypes were obtained using the Global Screening Array. All data\nwas mapped to GRCh38 and dbSNP, build 147 171 . Common haplotypes and variants were imputed using the\nTOPMed Freeze 5 panel 29 , 30  using 668,127 observed\nsingle-nucleotide polymorphisms (SNP), resulting in a total of 50,757,223\nhigh-quality called genotypes variants following imputation, an average of\n2,048,050 per individual. Details regarding QC and the imputation procedure were\ndescribed before 29 . Minimal QC\nwas applied to the imputed genotypes. We retained only non-duplicated,\nbi-allelic variants with high imputation quality (R2 > 0.7), a minor\nallele frequency (MAF) > 0.1%, and a missing rate < 5%. Genotypes\nwith a missing rate > 5% were excluded from further analysis. Concordance\nbetween genotypes determined by observed array variants and targeted Illumina\nsequencing was approximately 99.6%, determined using the vcf-compare command in\nVCFtools (v0.1.16) 172 .\nDemographic information, vitals, lab tests, International Classification\nof Diseases (ICD) codes, and medication prescriptions were retrieved from the\nUCLA Data Discovery Repository (DDR), established on March 2, 2013 29 , containing deidentified\nparticipant EHR data from our health system.\nDemographic, vital, and lab data were up to date as of Aug 24, 2024. For\nvitals and encounter counts, only records from in-person encounters (hospital\nvisits, appointments, surgeries, office visits, and walk-ins) were considered.\nFor vital signs, median values across all encounters were used to reduce the\nimpact of potential recording errors or error-prone self-reported values. For\nlab data, the most recent values were considered. For  Figure 1c , self-reported sex and age at the most\nrecent time point at the time of analysis were used.\nYearly encounter numbers were retrieved in 2024. We included only\ncomplete yearly data up to 2023 to ensure consistency and avoid partial data\nfrom 2024. Hospital encounters during the COVID pandemic years (2019-2022) were\nexcluded.\nFor associations involving clinical phenotypes, both ICD-9 and 10 were\nextracted. To ensure harmonized and comparable phenotypes across data sources,\nwe adopted a structured, standard approach using phecodes 59 – 61 . We utilized the existing PheWAS catalog maps and\napplied standardized rules for case–control definitions. The ICD-9 codes\nwere mapped to phecodes using phecode Map v1.2 59 , while ICD-10 codes were mapped using\nphecode Map v1.2b1 60 , both of\nwhich were obtained through the PheWAS catalog 61 . Cases were defined using the widely\nused “Rule of Two” approach 61 , 173 – 176 . For each phecode,\nparticipants were classified as cases if they had at least two occurrences of\nthe same phecode, with occurrences spaced at least 30 days apart to ensure\npersistence of diagnosis. In total, there were 1,308 phecodes with at least 100\ncases in ATLAS. Controls were defined as participants who had no record of the\nphecode in their EHR data and at least two recorded encounters in the system,\nfollowing similar established rules for minimum data content 61 , 177 , 178 .\nUndecidable participants were excluded.\nWe used phecodes 59 – 61  to\ndefine phenotypes and validated the generalizability of these phenotypes using\nseveral approaches. First, we tested cross-checked information between a variety\nof vital signs and lab tests and specific case-control groups defined by\nphecodes, using Wilcoxon tests and density plots, which confirmed the expected\ndifferences for these phenotypes. This included the following: 1) phecode-lab\ntest pairs: Type 2 diabetes-Hemoglobin A1c, Hyperlipidemia-Trygllicoride,\nVitamin D deficiency-Total 25-Hydroxy vitamin D, Deficiency anemias-Mean\nCorpuscular Volume (MCV), Deficiency anemias-Hemoglobin, and 2) the\nphecode-vital sign pairs: Essential hypertension-Blood pressure diastolic,\nEssential hypertensionBlood pressure systolic, Obesity-BMI, Anorexia\nnervosa-BMI, Short stature-Height ( Figure S2a – j ). Second, we replicated many\nknown associations, including ancestry-related differences in disease prevalence\nfor 28 major phenotypes. For example, across cancer types, we observed the\nhighest risk for prostate cancer in AFR ancestry, stomach cancer in EAS, bladder\ncancer in AMR, and breast cancer in EUR 179  (see more below and in  Table S2 ). Across fine-scale\nancestries, for example, we showed increased breast cancer and Crohn’s\ndisease diagnoses in the Ashkenazi Jewish (IBD-03) cluster 53 , 180  (see more below). We also replicated 11,756\nvariant–phenotype associations, such as a lower risk of\nAlzheimer’s disease for African American participants with the\nε4/ε4 haplotype (see more below). PGS prediction of 28 traits\nrevealed significant enrichment in the top PGS decile for 89% of these traits in\nEUR (see more below). Collectively, the replication of known associations using\nATLAS-defined phecodes indicates high-quality, well-defined phenotypes that\nmatch external definitions. Third, we used EHR data on cancer diagnoses to\nvalidate prostate and breast cancer cases (the two most diagnosed cancers)\nagainst phecode-defined case-control groups. We found a high agreement of 95%,\nfor both cancers, considering decidable patients based on phecodes ( Figure S2k – l ). The following\nsections describe replication results conducted to validate the quality of the\ndata and the phenotype definitions.\nWe assessed variation in medical conditions and disease prevalence\nacross populations across preselected prevalent major conditions ( Table S2 ). This\nconfirmed known ancestry-disease associations. Across cancer types, we\nobserved the highest risk for prostate cancer in AFR ancestry, stomach\ncancer in EAS, bladder cancer in AMR, and breast cancer in EUR 179 . We also replicated\nseveral known cardiovascular disease associations. For instance, AFR\nparticipants were more affected by hypertension and myocardial infarction,\nwhile those with SAS ancestry had the lowest risk of atrial fibrillation,\ndespite the highest incidence of coronary atherosclerosis, an established\nbut seemingly contradictory risk profile 181 . Major metabolic disorders were\nmore prevalent in AFR than in other ancestries. This finding included\ndiagnostic codes (9- and 10-ICD codes) related to Type 2 diabetes,\nhypercholesterolemia, and hyperlipidemia, supported by increases in\nmetabolic-related laboratory measurements and vital signs ( Figure S3e ). Type 2 diabetes\nwas more common in all non-EUR groups relative to EUR. Previous work has\nsuggested that some of these differences reflect social and environmental,\nrather than genetic, factors 182 . With regard to neurological conditions, we find a\nhigher risk of dementia in AFR participants, but a lower risk of migraines\nand Parkinson’s Disease in AFR compared to EUR patients, consistent\nwith prior analyses 183 – 185 . Among neuropsychiatric disorders, anxiety and major\ndepressive disorders were most frequent in EUR participants, and least\nfrequent in the continental Asian cluster, as previously reported 186 , 187 . These findings demonstrate the\nutility of ATLAS in robustly replicating known associations within a single\nhealth system, reducing potential confounding due to geographic or health\nsystem-related effects.\nWe tested the prevalence of 1,253 phecodes 59 – 61  across fine-scale clusters with at least 100\nparticipants ( Table\nS3 ; selected tests,  Figure\n2c ). As a measure of integrity, we tested for known increases in\nprevalence in these fine-scale ancestries, for example, replicating an\nincrease in breast cancer and Crohn’s disease in the Ashkenazi Jewish\n(IBD-03) cluster. Similarly, we observed the known increase in gout in the\nFilipino cluster (IBD-09) 188  and Alzheimer’s disease and dementias in the\nPuerto Rican cluster (IBD-15) 189 . Consistent with the global ancestry findings,\nAfrican American (IBD-06) had the highest prevalence of hypertension.\nWe leveraged genotyping in our cohort to calculate individual PGS\nfor a range of cancer, cardiovascular, metabolic, neuro-psychiatric and\nautoimmune diseases and tested their relationship to disease risk. We\nfocused on the top end of the PGS distribution (10%), compared with the\n5 th  decile, in EUR participants ( Figure 3 ; see the table below). For Type 1\ndiabetes, the top PGS decile was the most enriched (OR = 11.6 [7.6, 19.0],\nFDR = 4.1×10 −25 ), with 41% of the diagnosed\nparticipants in ATLAS assigned to the top PGS decile. The second most\nenriched trait was Crohn’s disease, with 33% of diagnosed\nparticipants within the top PGS decile (OR = 5.5 [4.0, 7.7], FDR =\n7.9×10 −24 ), followed by gout with 27% (OR = 5.1\n[4.1, 6.4], FDR = 2.0×10 −45 ), testicular cancer\nwith 25% (OR = 3.7 [1.2, 7.3], FDR = 1.2×10 −4 ), and\nprostate cancer with 21% (OR = 3.4 [2.8, 4.0], FDR =\n2.0×10 −45 ) ( Figure S5a – e ). On average, the\ntop decile of risk accounted for 18.6% of diagnosed participants across 28\npreselected prevalent disorders. Twenty-five of 28 (89%) showed significant\nenrichment of participants within the top PGS decile (mean OR for\nsignificance tests = 2.9). Performance declined when we applied PGS to\nnon-EUR populations, as expected 55 – 58 .\nThis was due to both reduced sample sizes and the model fit ( Figure S5f – i ), identifying only\n17, 27, 36, and 38% significant associations of diseases with the top PGS\ndecile, for SAS, AFR, EAS, and AMR, respectively. Concordantly, the top PGS\ndecile accounted for fewer cases (on average, 13, 15, 16 and 16%, for SAS,\nAFR, EAS and AMR). This further supports the need for larger, more\nancestrally heterogeneous cohorts for clinical development of PGS 55 – 58 .  The OR and prevalence of cases for the top and bottom\nPGS deciles across traits in EUR Trait OR top OR bottom Cases top (%) Cases bottom (%) Cases sample size FDR top FDR bottom Atrial fibrillation 3.2 0.5 19.5 4.8 4565 6.60×10 −66 1.90×10 −13 Bipolar 1.6 0.7 15.5 7.4 1147 9.60×10 −5 0.011 Bladder cancer 1.8 0.8 15.2 6.9 698 0.00028 0.3 Breast cancer 2.4 0.5 18.0 4.9 3016 2.20×10 −26 2.40×10 −08 Cerebrovascular disease 1.3 1.1 9.9 10.5 334 0.34 0.71 Colorectal cancer 1.6 0.6 15.0 5.9 842 0.0026 0.0064 Coronary atherosclerosis 2.5 0.5 16.0 6.1 7951 2.20×10 −59 2.70×10 −24 Gout 5.1 0.4 27.2 2.7 1638 2.00×10 −45 1.10×10 −05 Hypercholesterolemia 1.4 0.6 12.7 6.8 10836 1.50×10 −14 5.70×10 −24 Hyperlipidemia 1.6 0.6 12.2 7.8 21829 6.70×10 −29 1.30×10 −30 Hypertension 2.1 0.6 12.6 7.8 21517 2.50×10 −63 1.10×10 −28 Hypertrophic obstructive\ncardiomyopathy 1.7 0.4 17.8 3.9 129 0.13 0.11 Lung cancer 1.1 0.7 10.3 7.2 976 0.65 0.062 Major depressive disorder 1.6 0.7 13.0 7.1 10793 3.00×10 −22 1.50×10 −10 Malignant neoplasm of\ntestis 3.7 0.4 24.9 2.9 173 0.00012 0.13 Melanoma 2.3 0.5 19.6 3.4 1196 8.80×10 −12 1.00×10 −4 Multiple sclerosis 2.8 0.3 26.1 2.9 379 6.00×10 −7 0.0024 Myocardial infarction 1.6 0.8 12.9 7.2 1744 2.30×10 −5 0.11 Ovarian cancer 2.1 0.8 17.9 6.7 403 0.00066 0.35 Pancreatic cancer 1.6 0.7 12.6 5.2 382 0.048 0.29 Prostate cancer 3.4 0.3 21.5 2.9 2703 2.00×10 −45 1.10×10 −17 Psoriatic arthropathy 1.5 0.9 14.1 8.0 503 0.029 0.53 Crohn’s disease 5.5 0.6 33.0 3.6 758 7.90×10 −24 0.11 Schizophrenia 2.9 0.3 21.9 2.3 128 0.01 0.11 Systemic lupus\nerythematosus 2.9 0.9 22.0 6.3 569 4.40×10 −9 0.56 Thyroid cancer 3.0 0.6 21.4 4.1 786 2.60×10 −12 0.022 Type 1 diabetes 11.6 0.4 40.8 1.5 537 4.10×10 −25 0.05 Type 2 diabetes 2.3 0.4 16.9 4.3 6302 2.20×10 −48 8.60×10 −26\nThe OR and prevalence of cases for the top and bottom\nPGS deciles across traits in EUR\nReplication of genotype-phenotype associations is important to\nassess the quality of both genetic and phenotype data, as well as to\nevaluate the power and effectiveness of the biobank in detecting true\ngenetic associations. To address this, we first examined the distribution of\n APOE  alleles across fine-scale clusters, highlighting\nthe increased frequency of  ε4  risk alleles in the\nAfrican American (IBD-06) and Bantu (IBD-35) clusters ( Figure S6c ) and replicating the\nfinding that African American participants with the\n ε4/ε4  haplotype have a lower risk of\nAlzheimer’s disease compared to other populations (see in the table\nbelow FDR-significant results). Next, we replicated two well-known genetic\nassociations 190 \nin the African American cluster, between  HBB  rs334-A ( Figure S6d ; MAFIBD-06\n= 5.2%) and a diagnosis of sickle cell anemia (PIBD-06 =\n2.0×10 −78 ; ORIBD-06 = 15.1) and between the\nDuffy null  ACKR1  rs2814778-C polymorphism ( Figure S6d ; MAFIBD-06 = 77.1%)\nand a decrease in neutrophil count (PIBD-06 =\n4.0×10 −31 ; ORIBD-06 0.7). Using ATLAS blood\nwork results, we further quantified the impact of rs334-A on mean\ncorpuscular hemoglobin concentration (PIBD-06 =\n5.2×10 −15 ; βIBD-06 = 0.4), nucleated red\nblood cell count (PIBD- 06 = 1.6×10 −10 ;\nβIBD-06 = 0.2), and mean corpuscular volume (MCV) (PIBD-06 =\n3.3×10 −8 ; βIBD-06 = − 0.3).\nAcross multiple Asian clusters, we confirmed the association between a\nvariant in high LD with the -- SEA  deletion, which causes\ninherited alpha-thalassemia 191 , and microcytic anemia. We observed a significant\ndecrease in MCV in response to  LUC7L  rs372755452-A in the\nbroad-scale EAS ancestry (PEAS = 6.1×10 −88 ;\nβEAS = −1.6; MAFEAS = 0.9%), and the fine scale Chinese +\nKorean (PIBD-05 = 2.7×10 −52 ; βIBD-05 =\n−1.6; MAFIBD-05 = 0.8%) and Filipino (PIBD-09 =\n2.2×10 −24 ; βIBD-09 = −1.8;\nMAFIBD-09 = 1.2%) clusters.\nNext, we replicated the finding that AMR participants are twice as\nlikely to carry a  PNPLA3  rs738409-G missense variant that\ngreatly increases the risk for non-alcoholic fatty liver disease ( Figure S6d ), a major\ncause of cirrhosis that often necessitates liver transplant 67 . Using fine-scale\nancestral mapping, we studied the association between rs738409-G and\nnon-alcoholic cirrhosis of liver in two Mexican American clusters ( Figure S6e ), IBD-04\n(PIBD-04 = 2.7×10 −14 ; ORIBD-04 = 1.90) and IBD-07\n(PIBD-07 = 1.5×10 −8 ; ORIBD-07 = 1.9), identifying\nsimilar risk but differing minor allele frequencies across these clusters\n(MAFIBD-04 = 45.0%; MAFIBD-07 = 52.9%). The same variant in Northern\nEuropeans had a smaller effect (PIBD-01 = 3.6×10 −9 ;\nORIBD-01 = 1.5) and was half as frequent (MAFIBD-01 = 22.9%).  Ancestry-specific effects of APOE ε4 on\nAlzheimer’s disease risk Haplotype Ancestry OR P-value CI low CI high e4e4 All 2.87 1.25×10 −41 2.45 3.28 e4e4 EUR 2.86 1.40×10 −31 2.38 3.34 e3e4 All 1.06 3.88×10 −19 0.82 1.29 e3e4 EUR 1.01 1.43×10 −13 0.74 1.28 e4e4 AMR 4.18 5.99×10 −6 2.37 5.99 e3e4 EAS 1.93 1.23×10 −4 0.95 2.92 e4e4 AFR 1.76 3.10×10 −3 0.59 2.93 e4e4 EAS 4.89 4.59×10 −3 1.51 8.26 e3e4 Unclassified 1.80 8.83×10 −3 0.45 3.15 e4e4 Unclassified 5.66 9.62×10 −3 1.37 9.94\nAncestry-specific effects of APOE ε4 on\nAlzheimer’s disease risk\nAnalyzing clinically relevant rare variants, known to have elevated\nfrequencies in specific populations, replicated known ancestry-associated\npatterns, assessing the quality of the data and the robustness of ATLAS in\ncapturing ancestry-associated allele frequency patterns. Focusing on\nFamilial Mediterranean Fever (FMF) variants, we identified 753 carriers with\nthe highest risk in the Armenian clusters 192 , 193  ( Figure S7a ). Carriers of these\nvariants also had a higher risk of the amyloidosis phecode, which is known\nto be associated with FMF (OR = 3.7 [2.2, 5.8]). The HBB:p.E7V variant,\nwhich is responsible for most sickle cell anemia cases, was carried by 273\nparticipants, with elevated frequency in the African American cluster\n(OR IBD-06  = 51.4 [39.7, 67.0];  Figure S7b ), in line with\nprevious findings 194 .\nFinally, we explored protective variants in the  PCSK9  gene\ncausing lowered LDL levels 195  and identified 49 carriers of loss-of-function\nvariants within the African American cluster (OR IBD-06  = 9.5\n[4.8, 17.8]) ( Figure\nS7c ), as is known 20 .\nWhile the above variants were selected  via  a\nliterature search, we also queried the entire ClinGen pathogenic or likely\npathogenic (P/LP) variant catalog, aggregated by gene, to identify\nenrichment in fine-scale populations. Among replicated associations,\nvariants in  BRCA1  and  BRCA2  were our first\ncandidates of interest. Ashkenazi Jewish had the primary risk for carrying\neither ClinGen P/LP variants in  BRCA1  or\n BRCA2  ( BRCA1 , OR IBD-03  =\n47.1 [20.6, 133.0];  BRCA2 , OR IBD-03  = 48.2\n[23.6, 114.2]) ( Figure 5a ). Of note,\nthese observations are based on six P/LP BRCA ClinGen variants found in\nATLAS, which include only two out of three of the Ashkenazi Jewish founder\nalleles (at the time of the analysis, the Ashkenazi Jewish founder allele\n NM_007294.3 :c.5266dup was not included in the ClinGen Evidence Repository).\nThe population attributable risk (PAR) was 3.0 [1.8, 5.0]% and 3.2 [1.9,\n5.3]% for  BRCA1  and  BRCA2  variants,\nrespectively, in the Ashkenazi Jewish cluster for breast cancer. Independent\nof ClinGen, we considered the three  BRCA  Ashkenazi Jewish\nfounder alleles and found that they were most prevalent in EUR among\nbroad-scale ancestries, and in Ashkenazi Jewish among fine-scale clusters\n( Figure\nS7d – e ). In Ashkenazi Jewish, the PAR for breast cancer, associated\nwith the three Ashkenazi Jewish founder alleles combined, was 8.2 [5.9,\n11.2]% when all ages were included and increased to 12.5 [8.5, 18.1]% in\nparticipants under 70. This increase is expected, since BRCA founder alleles\nare associated with early-onset breast cancer. In Northern Europeans, we\nobserved enrichments of P/LP variants in  MYOC \n(OR IBD-01  = 2.7 [1.7, 4.2]) that are associated with juvenile\nopen angle glaucoma, with a PAR of 0.7 [0.2, 3.4]%.\nA bias toward EUR in ClinGen variants was recently shown in\nAoU 23 . We tested\nwhether we could replicate this bias using a single health system. In this\npart, we asked if the total allele frequencies of the clinically actionable\nACMG ClinGen P/LP variants vary across broad- and fine-scale ancestries.\nOverall, 17 ACMG genes had at least one P/LP variant in ClinGen present in\nATLAS participants. For a more nuanced biological understanding, we divided\nthe ACMG variants into two groups of rare LOF (n = 53) and rare P/LP\nmissense (n = 131) variants. We defined ‘LOF’ for variants\nranked as high-confidence LOF by LOFTEE 94  and ‘missense’ based on the\nVEP 137 \n“missense_variant” annotation (see below,  Differences in total ClinGen allele frequency within ACMG\ngenes across ancestries ). We observed the highest frequency of\nrare P/LP LOF variants in EUR participants among broad-scale populations,\nand in the Ashkenazi Jewish cluster (IBD-03) among fine-scale clusters\n( Figure 5b ). This number was\nprimarily driven by BRCA P/LP alleles, which accounted for 63% of all\nClinGen LOF P/LP variants in EUR and 94% in the Ashkenazi Jewish cluster\n(OR EUR  = 3.7 CI EUR  = 2.6-5.4,\nP Bonferroni  = 6.3×10 −17 ;\nOR Ashkenazi Jewish  = 6.5, CI Ashkenazi Jewish  = 5.3\n- 8.1, P Bonferroni  = 3.5×10 −62 ). Still,\nwhen Ashkenazi Jewish individuals are removed from the broad-scale EUR\ncategory, a substantial enrichment of ClinGen missense variants is observed\nfor this group (OR EUR \n (non-Ashkenazi Jewish)  = 1.8 [1.2, 2.7], P-value = 0.002,\nP Bonferroni  = 0.01;  Figure S7f ), demonstrating that\nthe signal is not derived from Ashkenazi Jewish individuals alone. No\ndifferences in rare P/LP missense total counts were identified across\nbroad-scale ancestries, but across fine-scale ancestry clusters, Northern\nEUR participants (IBD-01) had significantly more (OR Northern EUR \n= 1.4 [1.2, 1.8], P Bonferroni  = 0.02), with a similar trend\nobserved in Southern Europeans that did not reach statistical significance.\nThis suggests that clinical datasets are particularly biased toward Northern\nEUR driven by the composition of large contributing cohorts such as the UK\nBiobank. Missense variants were depleted in Ashkenazi Jewish, possibly since\nthey are underrepresented in most EUR large scale cohorts, or due to\ntechnical differences in how studies or countries classify variants as P/LP.\nExcluding Ashkenazi Jewish as a sensitivity analysis resulted in a similar,\nalbeit weaker trend, of enrichment of missense variants in Northern EUR (OR\n= 1.3 [1.01, 1.6], P-value = 0.04, P Bonferroni  = 0.5;  Figure S6f ).\nAs all samples were collected incidentally, there was interest in\ncharacterizing the population, especially in comparison to non-biobank UCLA\npatients. To characterize the non-biobank UCLA patients while mitigating\ntime-dependent confounding, we included a subpopulation with an encounter within\none year of the ATLAS launch date. Encounters could be either inpatient or\noutpatient, and we summarized the relative proportion of patients with only\noutpatient encounters or at least one inpatient encounter (there were no\npatients with only inpatient encounters).\nTo identify clinical phenotype patterns in ATLAS and to compare these to\npatterns of non-biobank UCLA patients, we used all ICD codes from each patient\nto encapsulate past and present conditions, subject to practical\nchallenges 196 . For\nthe ATLAS population, we captured a snapshot using their encounter closest to\ntheir biobank sample collection date. For the non-biobank patients, we used\ntheir closest encounter to the ATLAS launch date. These encounters represented\nthe baseline encounters for both populations. The ICD codes at these encounters\nwere converted to phecodes and phecode groups 59 – 61 , 197  to\nrepresent meaningful categories of disease. Phecodes were extracted using pandas\nv2.2.2 with Python v3.9.19. We reported the unique phecodes with a prevalence\n≥5% in the UCLA ATLAS population. ICD codes at the baseline encounters\nwere also used to calculate the Charlson and Elixhauser Comorbidity Indices\n 40 , 41 , 198 , both measures that predict mortality. The scores were\nderived using the comorbidity R function v1.0.7 154  with R v4.1.2. To compare the\nprevalence of disease categories between UCLA Biobank participants and\nnon-biobank UCLA Health patients, we used logistic regression to test the\nassociation of each category with participant group, adjusting for age and\nsex.\nIn addition to characterizing patient populations at a baseline time, we\nalso described differences in how they changed over time, an important\nconsideration when assessing relative disease burden across populations.\nParticipants were enrolled in UCLA ATLAS and were encountered within the health\nsystem at different times; therefore, we controlled the interval over which we\nmeasured change. We identified patients with at least one encounter between one\nand two years after their baseline encounter. From this subpopulation, we\nobtained patients’ first encounter within this window and summarized the\ncumulative encounters that occurred between the two times. We also used the ICD\ncodes at the second encounter to derive new comorbidity scores, as well as\nextract updated phecodes. The change in comorbidity scores and increases in\nphecode prevalence provided insight into the evolving disposition of the UCLA\nATLAS population over time.\nThe genetic ancestry of participants in the ATLAS dataset was estimated\nby assessing their proximity to the centroids of 1000 Genomes superpopulations\nin principal component (PC) space. PCA across common genetic variants\ndemonstrates granular relationships and provides a quantitative basis for\nassessing relationships between ancestry and disease 29 , 36 . The top 20 PCs were calculated using the bed_projectPCA\nfunction (the bigsnpr 155  R\npackage v1.12.2) with default parameters. For every individual, the Euclidean\ndistance to the centroids of the five broad-scale populations (AMR, AFR, EUR,\nEAS, SAS) was computed. Participants were assigned AMR or AFR ancestry if the\nnearest centroid corresponded to one of these populations, as these groups are\nwell-separated in PC space. For EUR, EAS, and SAS ancestries, which exhibit more\ngenetic overlap, a stricter distance threshold was enforced to minimize\nmisclassification. Specifically, an individual was assigned to one of these\nancestries if their distance to the nearest centroid was less than a scaled\nthreshold, calculated as:\nThreshold=max(dist)×min(FST)/max(FST)×0.5, where max(dist) is the\nlargest squared distance among centroids and min(FST)/max(FST) accounts for\ngenetic differentiation. Participants who could not be assigned to any ancestry\ncluster under these criteria were labeled as “admixed/unknown.”\nVisualization was performed using the Boutros Plotting General (BPG) R package\nv.7.1.0 156 .\nVariation in encounter numbers across broad-scale ancestries was tested\nusing ANCOVA, adjusted for genetic sex, age, and BAS rank to account for\nhealthcare access differences. Comorbidity index values across broad-scale\nancestries were compared using ANCOVA, adjusted for genetic sex, age, and both\nBAS and ADI ranks, which provide complementary measures of socioeconomic status.\nSince not all patients had BAS and ADI values, the sample sizes for these\nanalyses were reduced (numbers are specified in the body of each figure).\nAdjusted means and CI were calculated using the R emmeans package, version\n1.10.5 132 .\nTo test differences in disease diagnosis across ancestries, phecodes\n(retrieved as described in  Retrieving phenotype\ndata ) were associated with genetic ancestry populations using\nlogistic regression, adjusting for age (age at diagnosis for cases and the\nlatest age for controls) and genetic sex if applicable. The following\ndisease-phecode pairs were used to define disease diagnosis: Type 2\ndiabetes-Type 2 diabetes, Sleep apnea-Sleep apnea, Essential\nhypertension-Essential hypertension, Hyperlipidemia-Hyperlipidemia, Anxiety\ndisorders-Anxiety disorders&Anxiety disorder&Generalized anxiety\ndisorder, Asthma-Asthma, Parkinson’s disease-Parkinson’s disease,\nType 1 diabetes-Type 1 diabetes, Schizophrenia-Schizophrenia, Crohn’s\ndisease-Regional enteritis, Chronic kidney disease-Chronic kidney disease, Stage\nI or II, Multiple sclerosis-Multiple sclerosis, Major depressive disorder-Major\ndepressive disorder, Cerebrovascular disease-Cerebrovascular disease, Atrial\nfibrillation-Atrial fibrillation, Hypercholesterolemia-Hypercholesterolemia,\natherosclerosis-Coronary atherosclerosis, Hyperlipidemia-Hyperlipidemia,\nCoronary Hypertrophic obstructive cardiomyopathy-Hypertrophic obstructive\ncardiomyopathy, Myocardial infarction-Myocardial infarction, Systemic lupus\nerythematosus- Systemic lupus erythematosus , Gout-Gout, Bipolar-Bipolar,\nPsoriatic arthropathy-Psoriatic arthropathy, Epilepsy-Epilepsy,\nNeurofibromatosis-Neurofibromatosis, Dementias-Dementias, Obesity-Obesity,\nObsessive-compulsive disorders-Obsessive-compulsive disorders, Autism- Autism,\nMigraine- Migraine, Alzheimer’s disease- Alzheimer’s disease,\nCoronary atherosclerosis-Coronary atherosclerosis, Posttraumatic stress\ndisorder-Posttraumatic stress disorder. Patients under 18 years old, with\nambiguous sex or with “unclassified” genetic ancestry were\nexcluded. Visualization was performed using the BPG R package v.7.1.0 156 .\nPGS were calculated using array data after imputation in EUR\nparticipants. Related participants based on their genetic similarity were\nexcluded (defined using PLINK v2.0a 109  with the relatedness coefficient --king-cutoff 0.05).\nPGS were calculated using pgsc_calc 158  with the default settings and --min_overlap of 0.65.\nLogistic regression was used to associate every PGS with the corresponding trait\nbased on phecodes (see  Retrieving phenotype\ndata ). The following PGS model IDs from the PGS catalog 199  and their phecode pairs were\ntested: PGS002250-Malignant neoplasm of ovary, PGS003766-Cancer of prostate,\nPGS000004-Malignant neoplasm of female breast&Breast cancer,\nPGS001794-Thyroid cancer, PGS000079-Melanomas of skin, PGS004884-Cancer of\nbronchus; lung, PGS002264-Pancreatic cancer, PGS003395-Colorectal\ncancer&Colon cancer, PGS000729-Type 2 diabetes, PGS002025-Type 1 diabetes,\nPGS004254-Regional enteritis, PGS004699-Multiple sclerosis,\nPGS000134-Schizophrenia, PGS004760-Major depressive disorder,\nPGS001806-Malignant neoplasm of testis, PGS004687-Malignant neoplasm of bladder\n& Cancer of bladder, PGS000039-Cerebrovascular disease, PGS004526-Essential\nhypertension, PGS004706-Atrial fibrillation, PGS004784-Hypercholesterolemia,\nPGS002029-Hyperlipidemia, PGS003726-Coronary atherosclerosis,\nPGS000739-Hypertrophic obstructive cardiomyopathy, PGS004528-Myocardial\ninfarction, PGS000803-Systemic lupus erythematosus, PGS001789-Gout,\nPGS002786-Bipolar, PGS000198-Psoriatic arthropathy. For prostate and testicular\ncancer, only males were included, and for breast and ovarian cancer, only\nfemales. An adjustment was made for age at diagnosis for cases and the latest\nage for controls, genetic sex when both sexes were included, and the first ten\ngenetic PCs. For  Figure 3 , the top and\nbottom PGS deciles compared with the 5 th  decile were considered,\ntesting only EUR participants. FDR was used for multiple testing correction. The\nsame process was repeated for other ancestries as presented in the  supplementary material . Visualization was\ngenerated using the BPG R package v.7.1.0 156 .\nATLAS array data were merged with genotyping data from the 1000\nGenomes Project 38 , the\nSimons Genome Diversity Project 130 , and the Human Genome Diversity Project 131 . BCFtools 128  annotate was used to\nharmonize variant reference SNP ID (RSIDs), and BCFtools 128  norm with a GRCh38\ngenome reference was used to standardize the genotyping data. Sites or\nindividuals with more than 1% missing were removed using PLINK 133  --mind and --geno. Only\nSNPs with MAF > 1% across all participants were kept.\nSHAPEIT5 134  with\ndefault parameters and the distributed GRCh38 map files were used to phase\ngenotyping data, one chromosome at a time.\nA custom Python script that converts PLINK bed files to PLINK\nped/map 133  files\nwhile conserving phasing information was used to convert genotyping data.\nCentimorgan data for the map files were generated using the same genetic map\ndata in SHAPEIT5.\nIdentity-by-descent segments were called using iLASH 135  with the following\nparameters: slice_size 350, step_size 350, perm_count 20, shingle_size 15,\nshingle_overlap 0, bucket_count 5, max_thread 20, match_threshold 0.99,\ninterest_threshold 0.70, min_length 2.9, auto_slice 1, slice_length 2.9,\ncm_overlap 1 and minhash_threshold 55. Identity-by-descent was called one\nchromosome at a time.\nIdentity-by-descent segment outliers were removed as described in\nBelbin et al. 53  and\nCaggiano et al. 52 .\nSegments overlapping centromeres or the human leukocyte antigen (HLA) region\nwere removed. Regions that may have false positive identity by descent were\nidentified using the following process and removed: total identity by\ndescent per each SNP was identified by summing across all\nidentity-by-descent segments that overlapped each SNP; SNPs with a total\nidentity by descent greater than or less than three standard deviations from\nthe genome-wide mean were removed.\nTo identify clusters, we followed the approach of Caggiano et\nal. 52  and Dai et\nal. 54  and applied\nLouvain clustering 200 . An\nundirected network is generated based on pairs of individuals who share\nidentity-by-descent segments: nodes are the individuals, edges are the\ntotal, genome-wide identity-by-descent shared as the edges. We used the\nPython package, NetworkIt 201 , to iteratively run Louvain clustering four times to\ndetect fine-scale clusters.\nTo avoid redundancy and maximize sample size, clusters were merged\nin two stages to produce a final set of consensus fine-scale clusters.\nFirst, following Caggiano et al. 52 and Dai et al. 54 , we computed pairwise\nHudson’s F ST  using PLINK v2.0a 136  across 378 clusters identified from\nthe fourth layer of Louvain clustering. After removing clusters with fewer\nthan 10 participants, to avoid unreliable F ST  estimates, we\nmerged the remaining 356 clusters into 67 clusters if the pairwise\nF ST  was less than 0.001.\nIn the second stage, we refined clusters using IBD sharing. For each\ncluster pair, we examined all inter-cluster individual pairs to calculate:\n(1) IBD mean , the average cM shared, and (2) IBD prop ,\nthe proportion of pairs sharing at least 3cM of IBD as detected by iLash. We\ndefined a composite metric for cluster merging, IBD weighted  =\nIBD mean ×IBD prop , that captures both the\nextent and prevalence of genomic IBD sharing. Next, clusters were sorted\nfrom the smallest to the largest number of ATLAS participants. For each\ncluster, we computed the IBD weighted  score with all larger\nclusters and calculated the mean of these pairwise values. The cluster was\nthen merged into the most similar larger cluster –\n i.e. , the one with the highest IBD weighted \nscore-only if that score exceeded the mean. Otherwise, the smaller cluster\nwas retained independently. This process merged 14 small clusters into\nlarger parent clusters, resulting in a total of 36 fine-scale clusters with\n≥ 30 participants each for downstream analyses 52 . These fine-scale clusters were\nassigned unique identifiers (IBD-01 through IBD-36) and manually annotated\nwith labels to ease interpretation. Due to filtering out clusters of small\nsample sizes before and after merging, not all individuals were assigned to\na fine-scale ancestry cluster.\nWe primarily relied on reference individuals, described in\n“ Fine-scale Ancestry\nPre-processing and Quality Control ”, to add labels to\nclusters. When clusters did not contain reference individuals or were\nheterogeneous, we utilized patient-reported race, ethnicity, preferred\nlanguage, and religious affiliation (in order of priority) to inform our\ncluster labeling. These aspects are not caused by identity-by-descent\nsegment sharing but can be indicative of a shared culture for individuals\nwithin a cluster; these shared practices can influence a group’s\ndemography, environment, and disease risk. Our labels are not definitive and\nare our best attempt to generate informative assignments for each\ncluster.\nWe used logistic regression to model the association between fine-scale\nancestry assignment and phecode prevalence, estimating OR with 95% confidence\nintervals using the logistf R package v1.26.0 for Firth’s bias-reduced\npenalized-likelihood logistic regression 157 . Differences were tested between every fine-scale\nancestry cluster with at least 100 participants (with no ambiguous genetic sex\nand over the age of 18) and all other ATLAS participants. For phecodes that are\nsex-specific, only participants of the corresponding sex were included in the\nanalysis. Adjustment was made for age at diagnosis for cases, and the latest age\nfor controls, and for sex when applicable. Phecodes with defined categories\naccording to the PheWAS catalog 61 , and that were diagnosed in at least 100 participants with\nno ambiguous genetic sex and over the age of 18 in total in ATLAS were tested,\nresulting in 1,253 phecodes. In situations where the number of cases in a\ncluster was small, exact ORs were not reported in the text to protect patient\nprivacy. For visualization, fine-scale clusters were grouped into four panels\n(EUR, Asian, AMR, and AFR) based on the predominant broad-scale ancestry of\nparticipants. If the predominant ancestry was “Unclassified” (as\nin the Japanese and Egyptian Christian clusters), the second most prevalent\nbroad-scale ancestry was used. The full phecode names presented in  Figure 2c  were: Hormones and synthetic\nsubstitutes causing adverse effects in therapeutic use, Cardiomegaly, Allergic\nreaction to food, Cirrhosis of liver without mention of alcohol, Vitamin\nB-complex deficiencies, Anemia of chronic disease, Cholesterolosis of\ngallbladder, Chronic renal failure [CKD], Dementias, Cancer of bladder,\nCataract, Leukemia, Hereditary hemochromatosis, Attention deficit hyperactivity\ndisorder, Glaucoma, Hyperplasia of prostate, Amyloidosis, Chronic pulmonary\nheart disease. These names were shortened in the figure and text for easier\nreading.\nThe same model was used to test differences in cardio-metabolic diseases\nusing appropriate phecodes, for each fine-scale cluster within the same\nbroad-scale continental ancestry. For this goal, clusters were assigned to\nbroad-scale ancestries based on the predominant ancestry match among\nparticipants within each cluster. The largest population cluster within each\nbroad-scale ancestry was used as the reference level for associations, adjusting\nfor BMI, sex and age at diagnosis for cases, and the latest age for controls. As\na sensitivity analysis, we tested whether the observed patterns persisted when\ncorrecting for SES factors (adjusting for ADI and BAS ranks) in addition to BMI,\nage, and sex. As a second step, we also added IPW (see  IPW analysis ).\nFDR was used for multiple testing correction. The data were visualized\nusing the BPG R package v.7.1.0 156 .\nThe following studies were used to compare ATLAS’s broad-scale\ngenetic diversity with other large scale biobanks: Halldorsson  et\nal . (UKBB) 4 , Kurki\n et al . (FinnGen) 6 , Feng  et al . (Taiwan Biobank) 7 , Zawistowski  et\nal.  (Michigan Genomics Initiative) 11 , Verma  et al .\n(Geisinger MyCode) 169 , The\nAll of Us Research Program Genomics Investigators  et\nal. 5 , Shaw\n et al . (Vanderbilt’s BioVU) 170 , Verma  et al . (Penn\nMedicine Biobank) 119 , Verma\n et al . (VA Million Veterans Program) 72  and Wiley  et al . (the\nColorado biobank).\nStudies that were used to compare sample sizes of ATLAS’s\nfine-scale ancestry clusters with other published cohorts with available genetic\ndata ( Table S3 )\nincluded: Belbin  et al . (BioMe fine-scale IBD\nclusters) 53 , Wu\n et al . and Li  et al . (Ashkenazi\nJewish) 57 , 202 , Haber  et al . and\nHovhannisyan  et al . (Armenian) 55 , 56 , Mehrjoo  et al . (Iranian) 203 , Larena  et\nal . (Filipino) 58 ,\nSohail  et al . and Ziyatdinov  et al .\n(Mexican) 204 , 205 .\nIPW were calculated using logistic regression in R with ATLAS\nparticipants as cases and other UCLA Health patients as controls. As described\nabove, we included UCLA Health patients who had at least one encounter within\none year of the ATLAS launch date. The following covariates were included:\nself-reported race, sex, age group, BAS rank, and ADI rank. Individuals of\nunknown race were excluded. For each ATLAS participant, the probabilities were\nextracted from the model, and the weights were defined as 1 divided by the\npredicted probabilities. To assess whether previously unreported ancestry and\ndisease status associations hold after applying IPW, we used the svyglm function\n(R survey package v4.4.8 164 )\nwith a quasibinomial model, adjusting for age, sex, BAS rank, and ADI rank,\nincluding the calculated weights for ancestry groups with more than 5 cases. To\ncompare comorbidity index values across broad-scale ancestries, considering IPW,\nlinear regression was used, adjusting for age, sex, BAS rank, and ADI rank,\nincluding the calculated weights. Adjusted means and CI were obtained using the\nR emmeans package, v1.10.5 132 . Patients under 18 years old, with ambiguous sex or with\n“unclassified” genetic ancestry were excluded.\nPheWAS were conducted using the Regenie v4.0 framework 66  for all five broad-scale\npopulations (AMR, AFR, EAS, EUR, SAS) and fifteen fine-scale clusters with at\nleast 400 participants (IBD-01 through IBD-15). Related individuals were removed\nusing a 0.05 kinship cutoff in PLINK v2.0a 136 . Sample sizes for each tested group can be found in\n Table S4 .\nAssociation testing was performed separately within each group\n( i.e. , population or cluster) for 1,437 binary traits and\n41 quantitative traits, including ICD-derived diagnoses and clinical laboratory\nmeasurements (see  Retrieving phenotype\ndata ). Samples were restricted to those in predefined inclusion lists\n(--keep), and trait-specific covariates were provided  via \n--covarFile, including age, sex (modeled categorically with –catCovarList\nsex), BMI, and the top 10 genetic PCs. Quantitative traits were divided into two\nanalysis groups based on missingness patterns, following Regenie’s\nrecommendation that traits with similar levels of missing data are modeled\ntogether in Step 1.\nFor each group, unimputed array genotypes in bed format were\nfiltered using PLINK v2.00a 136  to produce a minimal set of variants suitable for\nestimating genomewide polygenic effects through ridge regression in step 1\nof Regenie. We included autosomal variants with call rate ≥99%\n(--geno 0.01), minor allele frequency ≥1% (--maf 0.01), and\nHardy-Weinberg equilibrium p > 1×10 − 15\n(--hwe 1e-15). Additional linkage disequilibrium (LD) pruning was applied\nusing a sliding window of 1000 SNPs, advanced by 100 SNPs, with an r2\nthreshold of 0.9 (--indep-pairwise 1000 100 0.9), producing an average of\n389,209 SNPs per group. For binary traits, step 1 was run with a minimum\ncase count of 50 (--minCaseCount 50). For quantitative traits, step 1 was\nrun with Rank Inverse Normal Transformation (--apply-rint) enabled to\nstabilize variance across lab values with differing distributions. For all\ntraits, a block size of 1000 (--bsize 1000) was used, leave-one out cross\nvalidation (--loocv) was enabled, and sex was marked as a categorical\ncovariate (--catCovarList sex).\nPrediction files from step 1 for each trait and imputed genotypes in\nBGEN format were used to perform association testing in step 2 of Regenie.\nFor binary traits, a minimum allele count of 20 (--minMAC 20) and a minimum\ncase count of 50 (--minCaseCount 50) were required, and the Firth\napproximation was performed (--firth --approx) using a p-value threshold of\n0.01 (--pThresh 0.01). For quantitative traits, a minimum allele count of 20\n(--minMAC 20) was required and Rank Inverse Normal Transformation\n(--apply-rint) was applied. For all traits, a block size of 500 (--bsize\n500) was used and sex was marked as a categorical covariate (--catCovarList\nsex). Locus pruning was performed separately for each broad- and fine-scale\nancestry using PLINK 2.0a to identify unique variant-phenotype\nassociations.\nTo standardize phenotypes for comparison, phecodes for binary traits\nwere mapped to the Experimental Factor Ontology (EFO) using the text2term\nv4.5.0 Python package. Each phecode was assigned up to three top-matching\nEFO terms. To determine whether each unique variant–phenotype\nassociation was previously unreported, we queried the EMBL-EBI GWAS Catalog\n(June 27, 2025 freeze) for matching associations. An association from our\nstudy was classified as “previously unreported in the EMBL-EBI GWAS\nCatalog”, only if both of the following searches yielded no results.\n(1) Associations between the lead variant and any phenotype in the same\nphecode group as the associated phecode ( e.g. , for\ncolorectal cancer, we searched for associations with all EFO terms in the\nbroader “neoplasm” category). (2) Associations between the\nnearest protein-coding gene (within 10kb) and any phenotype in the same\ncategory as the associated phecode.\nGenomic DNA libraries were created by enzymatically shearing high\nmolecular weight genomic DNA to a mean fragment size of 200 base pairs.\nMultiplexity of exome capture and sequencing was achieved by adding unique\nasymmetric 10-bp barcodes to the DNA fragments of single samples during library\namplification. Equal molar amounts of DNA samples were pooled for exome capture\nusing a slightly modified version probe library of xGen exome research panel\nfrom Integrated DNA Technology (IDT). After PCR amplification and quantification\nof the captured DNA, samples were multiplexed and loaded to Illumina sequencing\nmachines for sequencing to generate 75-base-pair paired-end reads. The samples\nin this study were sequenced using the Illumina sequencing machines, including\nNovaSeq 6000 with S2 or S4 flow cells and the NovaSeqX with 25B flow cells.\nSequencing reads in FASTQ format were generated from Illumina image data\nusing the bcl2fastq program (v2.20, Illumina). Whole-exome reads alignment and\ngermline small variant detection were conducted using the Original Quality\nFunctionally Equivalent (OQFE) protocol described in Krasheninina et al.,\n2020 206 . Briefly, raw\nread files (FASTQ) were mapped to the GRCh38 reference obtained from  https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/technical/reference/GRCh38_reference_genome/ \nusing BWA-MEM v0.7.17-r1188 in an alt-aware manner 139 . Duplicate reads were then marked with\nPicard v2.21.2 127 . The final\nCRAM files were compressed with SAMtools v1.2 128 . Germline variant detection was\nperformed on each CRAM using a Parabricks accelerated version of DeepVariant\nv0.10.0 with a custom WES model 140 , resulting in a sample-level gVCF (genomic VCF).\nPer-sample gVCFs were merged with GLnexus v1.4.3 141  into a joint-genotyped multi-sample\nproject-level VCF (pVCF). Variant prediction was restricted to the exome capture\nregion and the 100 base-pairs buffer on each side of the target regions. The WES\nwas of high quality and reached an average coverage of 41.6-fold with a minimum\nof 24.2-fold, 90% having ≥ 34.7-fold in targeted regions, consistent with\na previous publication 207 .\nThe pVCF was converted to a PLINK file format using PLINK 1.9 133  for downstream analyses.\nThese data underwent extensive quality control to ensure the absence of\ncontamination, duplication, and other technical errors, as well as sufficient\nread depth to guarantee reliability and accuracy. Genetic duplicates were\ndefined based on the aggregated genotype data of all sequenced samples. Sex was\npredicted based on the ratio of read coverage on chromosome Y over the whole\nexome read coverage. VerifyBamID v1.1.3 142  was used to estimate sample contamination. Samples\nwere excluded if they showed sex discordance, duplication, cross-individual\ncontamination > 5%, or coverage < 20X in more than 20% off target\nregions. Overall, 386 samples were excluded for failing QC. Specifically, 249\nsamples were flagged for gender discordance, 12 had less than 80% of the exome\ncovered at 20X, 80 showed contamination levels above 5%, and 74 were unresolved\nduplicates. After accounting for 29 samples that appeared in more than one\nexclusion category, the final number of unique samples excluded due to QC\nfailure was 386. Alignment quality metrics were generated using Picard\nCollectHsMetrics v2.27.4 127 \nand SAMtools stats v1.15.1 with default settings. MultiQC v1.27.1 143  was used to systematically\naggregate sample-level quality metrics. RTGtools v 3.12.1 129  was used to assess variant QC metrics,\nincluding the number of variants for SNPs, small insertions and deletions;\ngenotype counts; heterozygous-to-homozygous ratios for each variant type and\ntransition/transversion (Ti/Tv) ratio. Genotypes were further confirmed using an\nalternative approach 208  with\n500 random samples, yielding a high concordance for on-target sites (median\nconcordance of 99.5% for SNPs and indels).\nFor WES-based analyses, we used samples that also had array genotyping\ndata (60,025 out of 61,797 patients with WES), which was necessary to\nconsistently define broad- and fine-scale genetic ancestry. We annotated\nvariants using Ensembl Variant Effect Predictor (VEP) v.112 137  for 58,387 samples assigned to the EUR,\nAFR, SAS, EAS, or AMR ancestry class ( i.e. , excluding\nunclassified genetic ancestry samples). Annotations were performed using the\nGRCh38 cache and the corresponding reference FASTA file. We limited our analysis\nto autosomal chromosomes to avoid technical artifacts in variant calling caused\nby the differences in ploidy between males and females, as well as the\nhigh-sequence similarity between the X and Y chromosomes in certain\nregions 209 . Counts\nfor single-nucleotide variants, indels, multi-allelic, synonymous, missense, and\nLOF variants were restricted to whole-exome sequencing (WES)-targeted regions.\nConsistent with a previous UKBB WES study 210 , we classified LOF variants as those with the\nfollowing consequences: stop_gained, start_lost, splice_donor, splice_acceptor,\nstop_lost, and frameshift. To increase reliability, we further restricted LOF\nvariants to those flagged as high confidence by LOFTEE 94 . For multi-allelic variants, the\npredicted function for each alternate allele was determined using the\n--pick-allele option based on the default ordered set of criteria defined by\nVEP 137 . To determine\nthe number of variants with MAF < 1% while accounting for\nancestry-specific allele frequency differences, we retained variants with MAF\n< 1% in at least one ancestry group.\nFor multi-allelic variants, we defined the minor allele as the second\nmost common allele (including the reference) and classified the variant as rare\nif the cumulative allele frequency of all alternate alleles was < 1%. For\nrare multi-allelic variants, we incremented a sample’s count only if it\ncarried an alternate allele with MAF < 1%. In multi-allelic cases where\nthe minor allele was the reference or the alternate alleles had AF ≥ 1%,\nwe incremented a sample’s count if it carried any alternate allele. For\nrare functional variants ( e.g. , missense), we incremented a\nsample’s count only if it carried an alternate allele with MAF <\n1% corresponding to that functional category.\nWe identified variants  via  a literature search, where\nvariants are known to be enriched within certain populations. For this goal, we\nselected: 1) Familial Mediterranean Fever (FMF) related variants in the MEFV\ngene (V726A, M694V, M694I, M680I and E148Q) 211  2) the HBB:p.E7V variant, and 3)\n PCSK9  P/LP variants 195 .\nFor FMF and  PCSK9 , carriers were identified as\nparticipants who carry at least one corresponding pathogenic variant within\n PCSK9  or FMF variant group. For HBB:p.E7V, the frequency\nwas estimated based on this variant alone. We only considered unrelated\nparticipants, using a kinship coefficient of 0.05 based on PLINK v2.0a 136  king-cutoff. Fisher’s\nExact test was used to test the carrier frequency of each cluster against the\ncarrier frequency of the other clusters grouped together. The risk of\namyloidosis (a common symptom of FMF) by carrier status was calculated\nconsidering carriers of known FMF and patients assigned the amyloidosis phecode\n(270.33).\nTo test differences in frequencies of pharmacogenomic variants, 121\nCPIC Level 1A variants were pulled from ClinPGx 86 , 87 . An overview of the analyzed variants can be found in  Table S4 . As variants\nspanned both common to rare frequencies, we used exome data for consistency.\nOverall, the variants were highly abundant in ATLAS, with 95.5% of all unrelated\nindividuals carrying at least one. Of these variants, 69 were identified in the\nUCLA ATLAS exomes. To test enrichment of single alleles across clusters,\n“carriers” were identified as participants who carry at least one\nallele of a variant within a gene associated with a pharmacogenomic drug\ninteraction or monogenic condition (unless only a single variant was involved).\nWe then calculated carrier frequency as the number of identified carriers over\nthe total number of participants with WES available and were unrelated using a\nkinship coefficient of 0.05 based on PLINK v2.0a 136 \n king-cutoff . To identify clusters with a significantly\ndifferent carrier frequency of variants, we applied a Fisher’s Exact test\nto the carrier frequency of each cluster against the carrier frequency of the\nother clusters grouped together. We then used FDR corrections for p-values\ngenerated from all Fisher’s exact tests and selected significant clusters\n(FDR ≤ 0.05) where at least 5 carriers were identified.\nTo identify clusters where participants were enriched for carriers of\nmonogenic ClinGen variants associated with disease, we used pathogenic variants\nthat were curated by experts within the field and underwent stringent review to\nbe considered pathogenic from ClinGen 91 . ClinGen variants filtered to identify P/LP variants\nwith autosomal dominant inheritance, autosomal recessive inheritance, or\nsemidominant inheritance, resulting in 3,521 variants associated with 204\nconditions overall according to ClinGen. We removed one  HNF4A \nClinGen variant due to a high allele frequency (MAF>0.01). Enrichment for\nvariants, aggregated per gene, was tested as described above for ClinPGx\nvariants with WES data. PAR was calculated for each cluster, showing significant\nenrichment of ClinGen P/LP variant frequencies aggregated by gene, focusing on\ngenes associated with monogenic diseases with autosomal dominant or semidominant\ninheritance. To ensure sufficient power, we included only clusters with at least\n30 individuals diagnosed with the disease with WES data available. We used the\nfollowing formula: PAR = (p(RR-1))/(p(RR-1)+1), where p is the carrier frequency\nand RR is the ratio of the proportion of carriers with disease to the proportion\nof noncarriers with disease. For familial hypercholesterolemia, cases were\ndefined as individuals with LDL levels exceeding 190 mg/dL, adjusted for statin\nuse. In case of statin use, the LDL levels were divided by 0.7, as previously\ndone 195 , 212 .\nTo calculate the PAR for breast cancer associated with the three\nAshkenazi Jewish founder alleles, we used the formula above. Because individuals\nin this group are often aware of their genetic status and may pursue preventive\nbreast cancer measures, we excluded individuals who underwent surgery (using the\n‘acquired absence of breast’ phecode) but were not diagnosed with\nbreast cancer, to improve the reliability of the estimate.\nThe list of all ClinGen 91  P/LP variants identified in any ACMG secondary findings\n(SF) v3.2 genes 93  was\nextracted as described above. Variants were labeled as missense or LOF, based on\nVEP v.112 137  annotations.\nMissense variants were defined for ‘missense_variant’ variants\naccording to the ‘consequence’ VEP output column. LOF was defined\nfor high-confidence “HC” LOF variants based on LOFTEE 94 . All variants were rare across\nbroad-scale ancestries. A WES plink file with all participants was filtered to\ninclude only the listed variants. Then, this file was broken into ancestry\ngroups ( i.e. , broad-scale population, or fine-scale cluster)\nusing PLINK v2.0a 136 . Related\nparticipants were removed using a 0.05 kinship cutoff in PLINK v2.0a 136 . The frequency of each\nallele in every population group was calculated with the PLINK v2.0a --freq\ncommand, and total frequencies were summed up for rare missense and LOF variants\nseparately. To test differences in the allele counts across populations, allele\ndosages were calculated with PLINK v2.0a 136  --recode A option. Dosages were summed to calculate\nthe total alternative (ALT) allele count within each population. The number of\nreference (REF) alleles was defined as twice the number of individuals in the\ngroup minus the number of ALT alleles. Fisher’s Exact tests were applied\nto test differences between the number of REF and ALT alleles in each population\ncompared to all other participants not assigned to that specific group,\nconsidering only groups with over 400 participants. Bonferroni correction was\napplied for multiple testing correction. Plotting was done using the BPG R\npackage v.7.1.0 156 .\nVariants from exome sequencing were annotated using VEP v.112 137  with the dbNSFP\nv4.9a 138  and\nLOFTEE 94  plugins\ninstalled. The LOFTEE high-confidence “HC” flag was used to select\nfor LOF variants predicted to have deleterious effects.\nMissense variants were assigned a 9-point deleteriousness score based\non a consensus of nine missense deleteriousness prediction toolkits, similar to\nmethods described in prior biobank-scale rare variant studies. We used the\nCritical Assessment of Genome Interpretation (CAGI) project 95  to prioritize well-performing tools not\ntrained on the same features. We selected five meta-predictors –\nClinPred 144 ,\nMetaRNN 145 ,\nBayesDel_addAF 146 ,\nVARITY_R 147 ,\nREVEL 148  – and\nfour stand-alone predictors – AlphaMissense 149 , MutPred2 150 , VEST4 151 , ESM-1b 152  – for use in our analysis. We\nassigned each variant a binary score per tool based on dbNSFP rank scores: 1 if\nthe variant’s score exceeded the threshold score for being more likely a\ndeleterious ClinGen or ClinVar variant than a background variant ( Figure S7g – h ), and 0 otherwise.\nSumming these binary scores produced a deleteriousness score ranging from 0 to\n9, with predicted damaging missense variants scoring ≥ 5 retained for\ndownstream analysis.\nA WES PLINK file was filtered to include computationally predicted LOF\nand predicted damaging missense variants (see above) in ACMG SF v3.2\ngenes 93 , providing a\nmore comprehensive evaluation than the one based solely on ClinGen P/LP\nvariants, which included only 17 genes. A WES plink file with all participants\nwas filtered to include only the ACMG putative damaging variants, and the file\nwas split into ancestry groups (broad- and fine-scale ancestries) using PLINK\nv2.0a 136 . Related\nparticipants were removed using a 0.05 kinship cutoff in PLINK v2.0a 136 . Allele frequencies were\ncalculated with PLINK v2.0a, and only rare variants (MAF < 1% in all\nbroad-scale ancestries) were kept. Allele dosages per individual were extracted\nusing PLINK v2.0a, and the total rare LOF and predicted damaging missense\nalleles were counted separately per participant. First, the differences between\nthe total numbers of REF and ALT alleles across ancestries were evaluated with a\nFisher’s Exact test as described above (ClinGen ACMG analysis). Second, a\nMann-Whitney U test with a Bonferroni correction was applied to test the\ndifference in the distribution of rare LOF and predicted damaging missense\ncounts per individual between any broad- or fine-scale group compared to all\nothers, with groups that included at least 100 participants.  Figure S7I  presents the\nMann-Whitney U test results, with statistically significant differences\nindicated by an asterisk (*). Visualization was done with BPG 156 .\nExWAS were conducted using the Regenie v4.0 framework 66  for all five broad-scale\npopulations (AMR, AFR, EAS, EUR, SAS) and fifteen fine-scale clusters with at\nleast 400 participants (IBD-01 through IBD-15). Related individuals were removed\nusing a 0.05 kinship cutoff in PLINK v2.0a 136 . Sample sizes for each tested group can be found in\n Table S6 . Within\neach group, imputed genotype dosages in BGEN format (--bgen) and prediction\nscores from PheWAS step 1 (--pred) were used to conduct gene-based association\ntesting for binary and quantitative traits across 17,537 genes. Samples were\nrestricted to those in predefined inclusion lists (--keep), and trait-specific\ncovariates were provided  via  --covarFile, including age, sex\n(modeled categorically with --catCovarList sex), BMI, and the top 10 genetic\nprincipal components.\nWhole-exome genotypes in bed format were filtered using PLINK v2.0a,\nretaining variants with a call rate ≥ 90% (--geno 0.1), a minor allele\ncount ≥ 1 (--mac 1), and a Hardy-Weinberg equilibrium p-value >\n1×10 −15  (--hwe 1e-15). In addition, 784,548\nvariants overlapping low-complexity regions (LCRs) were excluded prior to\nanalysis. Briefly, SNP positions were extracted from the .bim file, intersected\nusing bedtools v2.29.1 153 \nwith annotated LCRs from the Genome in a Bottle Consortium 213 , and filtered from the genotype files,\nresulting in 12,326,160 exome variants retained for burden testing.\nFiltered whole-exome genotypes in BGEN format and prediction files from\nPheWAS step 1 were used to perform gene-based association testing. Regenie-style\nannotation, set, and mask files for predicted deleterious LOF and missense\nvariants (see  Variant annotation using a\nconsensus of computational tools ) were created programmatically using\nPython v3.11.9 with the polars v1.2.1 package. For binary traits (--bt), a\nminimum case count of 50 (--minCaseCount 50), a minimum minor allele count of 5\n(--minMAC 5), and a block size of 1000 (--bsize 1000) were enforced. Firth\nlogistic regression with saddlepoint approximation (--firth --approx) was\napplied for variants with p < 0.01 (--pThresh 0.01). For quantitative\ntraits (--qt), rank inverse normal transformation (--apply-rint), a minimum\nminor allele count of 5 (--minMAC 5), and a block size of 500 (--bsize 500) were\nenforced. Regenie’s implementation of the RGC gene-based p-value test\n(–rgc-gene-p) was enabled for all analyses. Additionally, variants were\nbinned by minor allele frequency using 1% bins (--aaf-bins 0.01), and SNP-level\nmembership for each burden mask was recorded (--write-mask-snplist).\nIn addition to gene-based analyses, single variant association testing\nwas performed for selected predicted deleterious LOF and missense variants using\nfiltered whole-exome genotype dosages in BGEN format and prediction scores from\nPheWAS step 1. For both binary and quantitative traits, tests incorporated the\nsame set of covariates (age, sex modeled categorically, BMI, and the top 10\ngenetic principal components) and enforced a minimum minor allele count of 5\n(--minMAC 5). For binary traits, Firth logistic regression with saddlepoint\napproximation was applied for variants with p < 0.01, along with a block\nsize of 1000, while quantitative traits were evaluated using rank inverse normal\ntransformed phenotypes with a block size of 500. A Bonferroni-adjusted p\n< 0.05 was defined as the cutoff for significance.\nReplication analyses were conducted utilizing the Python Hail\npackage v0.2.134 165  on\nAoU Workbench using the Controlled Tier Dataset v8. Phenotypes were obtained\nby querying OMOP databases for per-patient ICD codes, which were\nsubsequently converted to phecodes v1.2 utilizing the R PheWAS package\nv0.99.6 176 .\nParticipants with both phecodes and single-read WGS data were filtered to\nremove participants flagged from genotype QC or from relatedness QC,\nresulting in 291,082 total participants. Logistic regression was performed\nwithin the given replication genetic ancestry cohort by testing case/control\nstatus for a given phenotype and SNP, utilizing age, sex, age^2, sex*age,\nsex*age^2, and genotype PCs 1-10 as covariates.\nReplications for variant-trait associations were obtained by\nquerying the Controlled Tier Dataset v8 “All by All” allele\ncount/allele frequency (ACAF) variant result tables using the Python Hail\npackage v0.2.134 on AoU Workbench.\nReplications for variant-trait associations were obtained by\nquerying the summary statistics portal at  https://taiwanview.twbiobank.org.tw/pheweb.php 7 .\nReplications for gene-trait associations were obtained by querying\nthe Controlled Tier Dataset v8 “All by All” rare variant\nresult tables using the Python Hail package v0.2.134 165  on AoU Workbench.\nReplications for gene-trait associations were obtained by querying\nthe AstraZeneca summary statistics portal at  http://azphewas.com/ 69 .\nThe BioMe Biobank consists of electronic health records and genetic\ndata from approximately 60,000 participants from the Mount Sinai Health\nSystem in New York. Participant recruitment was between 2007 and 2023. This\nstudy was approved by the Icahn School of Medicine at Mount Sinai’s\nInstitutional Review Board (Institutional Review Board 07–0529). All\nstudy participants provided written informed consent.\nBioMe participants were genotyped using the Illumina Infinium\nGlobal Diversity Array (GDA; number of participants, N=23,430; number of\nvariants, n=1,833,111) or Infinium Global Screening Array (GSA; N=32,595;\nn=635,623). Quality control consisted of removing participants with a call\nrate <95%, a mismatch between self-reported and genetic sex, and high\nheterozygosity. Duplicated sites and sites with a genotyping rate of\n<95% were removed. QC was done with PLINK2. Data was imputed with the\nTOPMED imputation server 214  with genome build hg38. Genetic PCs were calculated\nacross participants using PLINK2 for the genotyping sites after LD pruning.\nExome sequencing data were generated by the Regeneron Genetic\nCenter 215 .\nQuality control consisted of using the Goldilocks Filter 210 , and variants with\nquality scores <3 or depth of coverage scores <7 for SNPs, or\n<5 or depth of coverage <10 for indels were removed.\nMonomorphic sites were removed. Genetic ancestry was assigned using a random\nforest classifier trained on principal components derived from the 1000\nGenomes data to align with the genetic ancestry assignment in AoU 5 . Participants were assigned\na genetic ancestry using 10 PCs.\nBio Me  laboratory data were processed according to\nthe QualityLab pipeline 216 . Briefly, quantitative lab values obtained in\ndifferent clinical contexts (ambulatory, emergency, inpatient, and urgent\ncare), were cleaned. Only labs with at least 100 patients and at least 1000\nnumeric observations were considered. Only adult (> 18 years) lab\nvalues were considered. Per lab, outliers were removed, defined as values as\ngreater or less than 4 standard deviations from the mean of that lab.\nNon-numeric labs were removed. Labs must have had >70% of the\nreported units matching. After cleaning, a median value and median age was\ncalculated per person for each lab test in each clinical context and across\nall contexts. Laboratory phenotypes were manually assigned a match to UKBB\nphenotypes using available metadata.\nElectronic health phenotype data in the form of ICD-10 codes were\ncollapsed into PhecodeX phecodes 217 . Phecodes were transformed into a binary matrix,\nwhere each row was an individual and each column was a phecode. A\nparticipant had a 1 if they ever were diagnosed with that phecode, otherwise\ntheir value was set to 0. Age was calculated as current age, defined from\n01/01/2025.\nAssociation testing was performed using Regenie v4.0 66 . Burden testing was\nperformed using SKATO and the following masks: missense, missense_pLoF,\npLoF, and synonymous. PLoF variants were defined using LOFTEE 94 .\nATLAS GLP1-RAs prescriptions, including medication names, start and\nend dates, discrete dose, usage instruction, strength and route (oral or\nsubcutaneous), were retrieved and grouped based on the following simple\ngeneric names: dulaglutide, semaglutide, liraglutide, exenatide, albiglutide\nand lixisenatide. For all statistical analyses, only semaglutide users were\nconsidered. Usage start dates were defined based on the earliest\nprescription start date for each patient. In the case of a missing start\ndate, the prescription ordering date was used instead (in most cases, these\ntwo fields were identical). Overlapping medication periods were handled such\nthat when a new prescription started, the previous one was considered to\nhave ended. Similarly, missing end dates were determined using the start\ndate of the next prescription when available. In the case of completely\noverlapping prescriptions with different instructions, the combination of\nboth routes and the weighted average dose was considered. If the discrete\ndose information was missing, the medication dose was extracted from the\ninstructions’ free text field. If the instructions also omitted the\ndose information, it was imputed for a given medication type, based on the\nweighted dose average from the entire cohort on the corresponding type.\nInitial weight and BMI were defined as the median of all available weight or\nBMI measurements recorded from in-person visits, within 6 months prior to\nthe first prescription start date. Utilizing the median rather than a single\ndata point helped minimize the likelihood of recording errors in the EHR.\nThe percentage of weight change was obtained for every weight measurement\nrecorded on in-person visits within the period of active prescriptions\nbetween 4-60 weeks in total on semaglutide. The medication dose at each time\npoint was defined as the weighted sum of all medication doses (doses\nmultiplied by the number of prescription weeks) by the measurement date. In\ncase both oral and subcutaneous medications were used within a period, a\ncombined route category was defined.\nTo identify the overall weight loss patterns across time, the FPCA\nR function from the fdapace package v.0.6.0 159 , 160 , which is suited to plot smoothed longitudinal\ndata with repeated measurement, was used. This identified a consistent\nweight loss pattern up to 60 weeks, and sparse data points beyond\n~150 weeks ( Figure\nS8C ). Thus, we restricted all analyses to this period.\nFor all association tests, only participants aged 18 years or older\nwere included. For all analysis parts that involved longitudinal data with\nrepeated measurements, a linear mixed model with the bobyqa optimizer and an\nincreased function evaluation limit (maxfun = 10000) was used (lmer R\nfunction; the lmerTest R package v.3.1.3 161 ). In each case, ANOVA was used to\nidentify differences between models to define the best way to model a\npotential nonlinear relationship between weeks and weight loss, when\ntreating weeks as a fixed effect, based on the Akaike Information Criterion\n(AIC). In some cases, the best model was achieved using a restricted cubic\nspline (RCS) with the rcs R function from the rms package (v.7.0.0), applied\nto weeks. In other cases, a polynomial function of weeks provided a better\nfit. To plot smoothed longitudinal data with 95% CI, the fitted model values\n(excluding covariates) and 95% lower and upper confidence bounds, were\nextracted using the visreg package 162  v.2.7.0 and visualized using the BPG\npackage 156 .\nFor testing the effect of baseline factors, fixed effects were\ndefined for the medication dose, route, sex, age, initial BMI, first ten\ngenetic PCs (excluding PC6 due to a strong collinearity with PC5 that\ndisrupted the model convergence) and weeks on semaglutide. The model\nincluded both random intercepts and slopes for weeks on semaglutide.\nBonferroni correction was applied to control for multiple testing.\nFor testing differences across ancestries, fixed effects were\ndefined for the medication dose, route, sex, age, initial weight, a\npolynomial function of weeks on semaglutide, and an interaction between a\npolynomial function of weeks and genetic ancestry categorical class (EUR,\nAFR, EAS, SAS and AMR). The overall number of weight measurements was\n24,145, with a mean of 5 repeated measurements per patient. Ancestry sample\nsizes were: European (EUR), 3189; African (AFR), 373; admixed American\n(AMR), 913; East Asian (EAS), 291; South Asian (SAS), 107. The model\nincluded both random intercepts and slopes for the polynomial function of\nweeks on semaglutide. An ANOVA was performed on the model with a Bonferroni\ncorrection to assess the global effects of ancestry and\nancestry×weeks interaction. When a significant result was found, post\nhoc comparisons between groups were conducted using the summary lmerTest\nfunction (v.3.1.3 161 ),\napplying a Bonferroni correction to all 12 class or class × time\ninteractions.\nTo test the relationship between PGS and weight loss, related\nparticipants based on their genetic similarity were excluded (defined using\nPLINK v2.0a 136  with the\n--king-cutoff 0.05). Scaled BMI (PGS000027) and DM2 PGS (PGS000729) in EUR\nparticipants were divided into three equal bins each: low, intermediate and\nhigh scores. Then, using longitudinal data, for each trait, fixed effects\nwere defined for the categorical PGS bins, medication dose, route, sex, age,\ninitial weight, first ten genetic PCs (excluding PC6 due to a strong\ncollinearity with PC5 that disrupted the model convergence) and a polynomial\nfunction of weeks. Random intercepts were defined to account for repeated\nmeasures within individuals. Bonferroni correction was applied to control\nfor multiple testing. To test the relationship between PGS and weight loss,\nrelying on a simplified model where the maximum weight loss was considered\nfor each EUR participant, linear regression was applied. The model was\nadjusted for the medication dose and route, 10 genetic PCs, age, sex, and\ninitial weight. Visualization was made with BPG 156 , with smoothed data and 95% CI\nusing loess.as R function (fANCOVA package v.0.6.1 163 ).\nThe analysis was performed to test a relationship between the\nmaximum weight loss on semaglutide and common genetic variants, for each\nbroad-scale ancestry using SAIGE 167 . For step one, array observed variants were used\nfollowing PLINK v2.0a 127 \nfiltering, with the flags: --maf 0.01 --mind 0.1 --geno 0.1 --hwe 1e-6 . For\nthe second step, array observed and imputed variants were used following\nplink filtering with --maf 0.01 --geno 0.05 --hwe 1e-6. A quantitative\nanalysis was conducted with the traitType flag. The medication dose and\nroute, the first five genetic PCs, age, sex, and initial weight were used as\ncovariates. METAL 168  was\nused for meta-analysis.\nTo identify genes genetically associated with weight loss, we used\nRegenie 66  with an\nadditive model for gene-level tests. Only EUR semaglutide users that are not\nrelated (relatedness coefficient --king-cutoff 0.05) with WES data were used\n(n = 2,012), and for each, the maximum weight loss record with the\ncorresponding number of weeks was considered, using information on aggregate\ndose and route. The list of candidate genes was limited to\nBonferroni-significant proteins whose plasma abundance was altered by\nsemaglutide treatment 115 .\nWe considered variants within these genes with a predicted moderate or high\nimpact on the protein function, according to VEP v.1.2 137 . For step one, the variant list was\nlimited to observed array SNPs following a PLINK v2.0a 136  filtering with: --maf 0.01 --mac\n100 --indep-pairwise 1000 100 0.9 --chr 1-22 --snps-only --geno 0.1 --hwe\n1e-15. In step 2, we applied PLINK v2.0a filtering to variants across all\nEUR participants using the parameters --geno 0.05 and --hwe 1e-6.\nSubsequently, the file was filtered to include only semaglutide users and\nwas used in Step 2. The medication dose and route, the first 10 genetic PCs,\nage, sex, and initial weight were used as covariates. Bonferroni correction\nwas applied to control for multiple testing. Frequency of variants involved\nin  PTPRU  association with weight loss on semaglutide by\nancestry can be found in the table below.\nFor replication analyses in AoU the same processes and tools were\nconducted in EUR individuals. Controlled Tier Dataset v8 was accessed\nthrough the Researcher Workbench and imported in PLINK format, including\nexome variants in chromosome 1 (to extract  PTPRU  variants)\nand common ACAF variant datasets (for PGS calculation). Genetic ancestry and\nPCs derived from pre-defined assignments provided by AoU 5 .  PTPRU  VEP 137  variant annotations were\nobtained using Python Hail v0.2.134 165  to identify functional consequences. For the\nRegenie analysis, related individuals were not removed to maximise power, as\nRegenie is designed to account for relatedness. The combined P-value of\nATLAS and AoU was calculated using Fisher’s combined probability\ntest, implemented with the fisher function in the poolr R package\nv1.2.0 166 .\nMultiple testing was addressed using a Bonferroni correction that accounted\nfor the number of gene-level tests in the discovery cohort, along with one\nreplication P-value and one combined P-value.  Frequency of variants involved in\n PTPRU  association with weight loss on\nsemaglutide by ancestry. Variant AFR AMR EAS EUR SAS 1:29236667:T:C 0 0 0 1.2×10 −5 0 1:29258717:G:A 0.00018 0 0 0.00013 0 1:29259276:C:T 0 0 0 1.2×10 −5 0 1:29259313:A:G 0 0 0 4.9×10 −5 0 1:29259883:C:G 0 0 0 3.7×10 −5 0 1:29259912:A:C 0 0 0 1.2×10 −5 0 1:29260003:G:C 0 0 0 1.2×10 −5 0 1:29260895:A:C 0.00055 0.0025 0 0.0027 0 1:29275496:G:A 0 0 0 7.3×10 −5 0 1:29275715:G:T 0.0018 0.011 0.0018 0.0092 0.023 1:29279056:G:A 0 0 0 1.2×10 −5 0 1:29279527:G:T 0 6.0×10 −5 0 1.2×10 −5 0 1:29279541:C:A 0 0 0 1.2×10 −5 0 1:29282705:G:A 0.0013 0.0028 0 0.0050 0.0018 1:29282710:C:T 0 0 0 8.5×10 −5 0 1:29284754:C:T 0 0 0 9.7×10 −5 0 1:29291897:G:A 0 0 0 0.00023 0.00045 1:29291929:G:A 0 0.00030 0 2.4×10 −5 0 1:29291942:C:T 0 0.00042 0.00020 0.00052 0 1:29291958:A:G 0.0058 0.00030 0 2.4×10 −5 0 1:29291978:G:A 0 0 0 0.00015 0 1:29303853:A:G 0 0.00036 0 0.00024 0 1:29303863:C:T 0 0.00012 0 0.00032 0 1:29303875:G:T 0.094 0.0065 9.9 × 10–5 0.00078 0.00045 1:29304795:G:A 0 0 0 1.2×10 −5 0 1:29305397:A:G 0.23 0.38 0.48 0.25 0.38 1:29311488:C:T 0.00018 0 0 4.9×10 −5 0.00045 1:29311513:C:A 0 0 0 3.7×10 −5 0 1:29311706:G:C 0 0 0 1.2×10 −5 0 1:29312579:C:T 0 6.0×10 −5 0 9.7×10 −5 0 1:29312615:G:A 0 0.0011 0 0.00015 0 1:29315374:C:T 0 5.99×10 −5 0 2.4×10 −5 0.00045 1:29315448:G:A 0 0 0 1.2×10 −5 0 1:29317767:C:T 0 0.00036 0 0.0014 0 1:29317838:G:A 0 0 0 0.00016 0 1:29320709:A:G 0 0 0 6.1×10 −5 0 1:29323643:C:T 0 0 0.00030 1.2×10 −5 0\nFrequency of variants involved in\n PTPRU  association with weight loss on\nsemaglutide by ancestry.\nStatistical analyses were performed using R, PLINK, Regenie, and\nrelated tools. Logistic regression, linear regression, Firth-penalized\nregression, linear mixed models, and Fisher’s exact tests were applied in\nthe relevant analyses, adjusting for appropriate covariates. Multiple-testing\ncorrection was applied using the FDR or Bonferroni adjustment, as specified in\nthe relevant analyses. Detailed statistical models and software parameters are\nprovided in the  METHOD DETAILS \nsection.\nAn interactive web portal enables users to explore the PheWAS results,\nexamine fine-scale ancestry associations with clinical phenotypes, and download\nsupplemental data at  atlas-phewas.mednet.ucla.edu .\n\nFigure S1. Whole-exome sequencing quality control, related to\nthe \n STAR Methods .  a.  Coverage\ndistribution. The mean and median values were calculated across all samples.\nFold enrichment is the degree to which the baited region is enriched\ncompared to the background genomic region.  b.  Percentages of\nbases above different coverage thresholds. Every y-axis point shows the\npercentage of bases above the corresponding x-axis value. The blue area\nshows 95% of the data, while the surrounding gray lines represent the\nremaining 5% of the distribution.  c.  Average base counts by\nexome capture regions.  d.  Excluded bases from coverage\ncalculation based on Picard 127 \n e.  Insert size distribution.  f.  Samtools 128  summary statistics\noutput.  g . Genotype number distribution according to\nRTGtools 129 .\n h . Summary output from RTGtools. Ti/Tv is the ratio of\ntransition (Ti) to transversion (Tv). This included SNPs from off-target\nregions.  i.  The distribution of variant counts across all\nsamples after cohort re-genotyping.\nFigure S2. Phenotype validation and prevalence, related\nto \n Figure 1 \n and the \n STAR Methods .\n( A–J ) Comparison of vital signs and laboratory test\ndistributions between case and control groups defined by phecodes, with p\nvalues indicating significance from Wilcoxon tests. ( K and L )\nOverlap between cancer cases defined from EHR cancer diagnosis data\n(“Cancer Stage Fact”) and phecode-based case-control groups\nfor breast ( K ) and prostate cancer ( L ).\n( M ) Prevalence of frequent phecodes. The red bars indicate\nthe prevalence of the most frequent phecodes in subset of the UCLA ATLAS\npopulation who had an encounter between 1–2 years after their initial\nencounter, at the time of sample collection. The blue bars represent the\nprevalence of phecodes in this same subset at their encounters 1–2\nyears after their initial encounter. ( N ) The change in phecode\nprevalence over one year post-collection in the subset of the UCLA ATLAS\npopulation who had an encounter between 1–2 years after their initial\nencounter. The triangles represent the number of new patients, and the blue\nbars represent the percentage of patients. ( O and P ) The\ndifference in ( O ) phecode group and ( P ) phecode\nprevalence between participants in the UCLA biobank (purple; “UCLA\nATLAS”) at their time of collection compared with all other UCLA\nHealth patients with EHR data (green; “Rest of Data Discovery\nRepository [DDR]”) within one year of the ATLAS launch date. Logistic\nregression tests were used to compare the two groups adjusted for sex and\nage (ORs with 95% confidence intervals, shown in the right panels). After\nBonferroni multiple testing corrections, ** denotes a significant difference\nbetween the two populations.\nFigure S3. Broad-scale genetic ancestry and health\ncharacteristics, related to \n Figure 1 . a.  Agreement\nbetween self-reported race and genetic ancestry predictions.  b .\nThe Elixhauser comorbidity index varies across genetic ancestries after\napplying inverse probability weighting. ANOVA on a linear regression model\ntested the overall effect of the categorical ancestry predictor, yielding\nits P-value. Adjusted means and 95% CI are presented per ancestry group.\n c-d.  The relationship between total and hospital encounters\nand the comorbidity index. r, Pearson correlation, P, the correlation\nP-value.  e . Variation in laboratory or vital sign measurements\nacross broad-scale ancestries. ANCOVA adjusted for genetic sex and age was\nperformed separately for each tested phenotype.\nFigure S4. Fine-scale genetic ancestry supplementary results,\nrelated to \n Figure 2 . a-b.  Quality\ncontrol of identity-by-descent (IBD) segments. Distribution of IBD segment\nlength per chromosome, before ( A ) and after ( B )\nremoving human leukocyte antigens (HLA), centromere, and IBD depth outliers.\nAfter quality control, IBD segments display exponential decay of segment\nlength as expected (most noticeably for chromosomes 6, 15 and 22). The\nnumbers at the top of each plot represent the chromosome number.\n c . The distribution of genetic ancestry between fine- and\nbroad-scale populations. IBD clusters were sorted by the predominant\nbroad-scale ancestry of participants in each cluster, rather than cluster\nsize as shown in  Figure 2a , providing\nan alternative visualization.  D.  Cardio-metabolic disease risk\nfor each fine-scale group within the same broad-scale ancestry.\nRepresentative cardio-metabolic phecodes were selected, and only populations\nwith at least 100 participants were tested. In cases of small sample sizes,\n‘–’ was used instead of numeric values to protect\npatient privacy. Filled points represent significant results (FDR ≤\n0.05). Firth’s bias-reduced logistic regression adjusted for BMI,\nsex, and age was used to obtain OR. Among Asian clusters, Filipino\nindividuals were at high risk for all tested medical conditions (essential\nhypertension: OR IBD-09  =1.6 [1.4, 1.9], FDR = 8.6 ×\n10 −11 ; type 2 diabetes: OR IBD-09  = 1.5\n[1.3, 1.7], FDR = 7.4 × 10 −6 ; coronary\natherosclerosis: OR IBD-09  = 1.4 [1.1, 1.7], FDR = 5.3 ×\n10 −3 ; abdominal aortic aneurysm: OR IBD-09  =\n2.8. [1.4, 5.3], FDR = 7.1 × 10 −3 ; hyperlipidemia:\nOR IBD-09  = 1.2 [1.0, 1.4], FDR = 2.4 ×\n10 −2 ). The largest Armenian cluster and Jewish and\nnon-Jewish Iranian clusters showed a high risk for type 2 diabetes (Armenian\n1: OR IBD-16  = 2.0 [1.5, 2.7], FDR = 2.5 ×\n10 −6 ; Iranian Jewish: OR IBD-11  = 2.4 [1.9,\n2.9], FDR = 6.0 × 10 −16 ; Iranian:\nOR IBD-17  = 2.2 [1.6, 2.9], FDR = 1.6 ×\n10 −6 ), hyperlipidemia (Armenian 1: OR IBD-16 \n= 1.6 [1.2-2.0], FDR = 8.6 × 10 −4 ; Iranian Jewish:\nOR IBD-11  = 1.5 [1.2, 1.8], FDR = .1.0 ×\n10 −4 ; Iranian: OR IBD-17  = 1.8 [1.4, 2.3],\nFDR = 2.9 × 10 −5 ) and coronary atherosclerosis\n(Armenian 1: OR IBD-16  = 1.9 [1.4, 2.5], FDR = 8.8 ×\n10 −5 ; Iranian Jewish: OR IBD-11  = 1.8 [1.5,\n2.2], FDR = 4.2 × 10 −8 ; Iranian:\nOR IBD-17  = 1.9 [1.4, 2.5], FDR = 2.9 ×\n10 −5 ).\nFigure S5. polygenic score association results, related\nto \n Figure 3 \n and the \n STAR Methods . a-e \nSelected associations between polygenic scores (PGS) and diseases in\nEuropean (EUR) individuals.  f-i  The odds ratio (OR) and\nprevalence of cases for the top and bottom PGS deciles across non-EUR\nancestries. Logistic regression was used to obtain the OR and p values,\nwhich were adjusted using FDR.\nFigure S6. PhWAS supplementary results, related to \n Figure 4 .\na-b. Distribution of pruned\nvariant-trait associations. The number of genome-wide significant pruned gene-trait associations\nshared across fine-scale ancestries, using a prioritized gene within 10kb of the pruned variant\nand matching effect direction. b presents the same associations described in a, but with a more\npermissive threshold for gene-trait support from less powered ancestral clusters. c. APOE\nhaplotype frequency across fine-scale cohorts. d. Allele frequency for known risk variants across\nfine-scale cohorts; shading indicates the level of over- (red) or under- (blue) enrichment of a\nhaplotype/allele in a given cohort via Fisher’s exact test. In c-d, Fisher’s exact test was utilized to\ncompare allele frequency. Bold border indicates a significant result (FDR ≤ 0.1). e. Impact of the\nnon-alcoholic cirrhosis risk variant rs738409-G on cirrhosis and clinical sequalae across finescale\ncohorts. Odds ratios and 95% CI were calculated using logistic regression. f-h. Replicated\nlow-MAF PheWAS associations between rs115750084-G and major depressive disorder (f),\nrs77742325-G and osteoporosis with no other symptoms (g) and rs202215133-A and migraine\n(h).\nFigure S7. Known or putative rare pathogenic variants, related to Figure 5  \n Figure 5  and the STAR\nMethods.\na. Frequency differences in Familial Mediterranean Fever (FMF) known risk alleles. b.\nThe risk of carrying the HBB:p.E7V variant. c. The risk of carrying loss-of-function variants in\nPCSK9. In a-c, the error bars show the 95% Wilson score confidence intervals. d-e. The\nfrequency of BRCA Ashkenazi Jewish founder alleles across broad-scale (d) and fine-scale\nancestry groups (e). f. Differences across populations in the total numbers of rare ClinGene P/LP\nvariants in American College of Medical Genetics (ACMG) genes, excluding Ashkenazi Jewish,\nas a sensitivity analysis to one presented in Figure 5b. ‘Ref’ is the total number of reference\nalleles, and ‘Alt’ of P/LP ClinGen. Fisher’s exact tests were used to produce odds ratios. Filled\npoints represent significant results at the level of nominal P-value (≤ 0.05). Only fine-scale\nancestry clusters with more than 400 participants were included. g-h. Distributions of missense\nvariant pathogenicity rank scores for nine computational tools. Each subplot compares the\ndistribution of all missense variants (purple to green) to known pathogenic variants (red), defined\nas either: g. ClinGen curated missense variants. h. ClinVar pathogenic/likely pathogenic (P/LP)\nmissense variants. The black dashed line indicates the median rankscore across all missense\nvariants for the given tool. The red dashed line denotes the median rankscore among pathogenic\nvariants in the corresponding dataset. The purple dashed line marks the likelihood-based\nintersection cutoff derived from the point at which the pathogenic and background distributions\ncross. i. The number of rare computationally predicted damaging missense and LOF alleles per\nindividual across ancestries. Mann-Whitney U test with a Bonferroni correction was applied to\ntest the difference in the distribution of rare LOF and predicted damaging missense counts per\nindividual between any broad- or fine-scale group compared to all others. Statistically significant\ndifferences are indicated by an asterisk (*). Only fine-scale ancestry clusters with more than 100 participants were included. j. Ancestry-specific carrier counts for two known GBA1 rare LOF\nvariants (left) and their impact on Parkinson’s disease risk (right) via logistic regression.\nFigure S8 .  GLP1-RAs ATLAS users and semaglutide\ninvestigation complementary data, related to \n Figure 6 .  a . Age by sex of\nGLP-1 receptor agonist (GLP1-RAs) users.  b . Prescription\nnumbers for GLP1-RAs by simple generic names.  c . Weight loss\npatterns across time. Presented are smoothed longitudinal data using a\nfunctional boxplot approach, with median values and pointwise intervals\nbetween the 20th and 80th quantiles. Vertical tick marks along the x-axis\nshow the deciles of the data distribution, indicating where most data points\nare concentrated across weeks.  d . The relationship between bins\nof polygenic scores (PGS) for body mass index (BMI) (left) and type 2\ndiabetes mellitus (right), and weight loss in response to semaglutide. Shown\nare Loess smoothed plots with 95% CI based on the maximum weight loss for\npatients across weeks. p values and effect sizes were obtained from a linear\nregression model with covariates, and a Bonferroni correction was applied.\n e . The relationship between type 2 diabetes mellitus PGS\nand weight loss in response to semaglutide in EUR AoU participants. Scaled\nPGS were divided into groups and linear mixed model fitted values were\nplotted with 95% CI, based on longitudinal data with repeated weight\nmeasurements. P-values were obtained from a linear mixed-effects model that\nincluded covariates.  f . GWAS Q-Q plot for weight loss on\nsemaglutide displays no significant findings.  g . Gene-level\ntest Q-Q plot for weight-loss on semaglutide (genomic inflation factor:\n0.95).  h.  Differences in carrier numbers of semaglutide\nefficacy involved alleles in  PTPRU  across ancestries using\nFirth’s bias-reduced logistic regression. Only variants that\nexhibited at least one significant difference between EUR and another\nancestry are shown. Filled dots represent significant odds ratios (FDR\n≤0.05).\nTable S1. Summary of lab test results and prescription numbers for\nthe top 50 prescribed medications in ATLAS, related to the  STAR Methods .\nTable S2. Phecodes and broad-scale ancestry associations, related\nto  Figure 1  and the  STAR Methods .\nTable S3. Fine-scale ancestry cluster sample sizes compared with\npublished cohorts and phecode associations, related to  Figure 2  and the  STAR Methods .\nTable S4. PheWAS results and pharmacogenomic variants, related to\n Figure 4 .\nTable S5. Summary of WES variant counts by category and differences\nacross broad-scale ancestries, related to  Figure 5  and the  STAR\nMethods .\nTable S6. ExWAS results, related to  Figure 5 .","source_license":"CC-BY-4.0","license_restricted":false}