Advancing precision health discovery in a genetically diverse health system.

OA: closed CC-BY-4.0

Abstract

Linking genetic data with electronic health records in hospital biobanks promises to advance precision medicine, but limited ancestral diversity constrains discovery and generalizability. We analyzed 93,936 participants from the UCLA ATLAS Community Health Initiative to inform disease prevalence and genetic risk across five continental and 36 fine-scale ancestry groups. We discovered numerous unreported gene-phenotype associations, including FN3K with intestinal disaccharidase deficiency in Europeans and admixed Americans. Polygenic scores (PGS) robustly predicted common diseases, with effects markedly diminished in non-Europeans. Furthermore, we reduced the pronounced European bias in curated clinical variants using computational predictors, uncovering unreported disease-gene associations, including ANKZF1 and peripheral vascular disease in African Americans. Longitudinal data revealed that semaglutide efficacy varies across ancestries, is associated with PGS for type 2 diabetes, and is modulated by genetic variation in PTPRU. These findings illustrate how ancestrally diverse biobanks from a single health system yield robust disease associations and pharmacogenomic insights.
Full text 145,968 characters · extracted from pmc-nxml · 7 sections · click to expand

Star

The Institutional Review Board of the University of California, Los Angeles gave ethical approval for this work. The IRB Number for the ATLAS Initiative protocol is IRB#17-001013. Human participants included in this study were from the UCLA ATLAS Community Health Initiative. Enrollment procedures, consent, and data collection are described in Johnson et al 30 . Participants provided informed consent for research under IRB#17-001013. The cohort description is provided below and summarized in Table 1 and Figure 1 . ATLAS enrollment reflects a modest over-representation of females by self-reported sex, especially among patients aged 20-60 ( Figure 1c ; Table 1 ). Among ATLAS participants, the median age at the most recent time point is 61 years for males and 56 years for females. Medical morbidity was notably more widespread in the biobank population compared to all adult UCLA patients. This point was evidenced by a significantly larger number of diagnoses in the biobank population compared to all UCLA patients (15.6 vs. 10.5 mean ICD codes per patient, one year after collection). The table “Baseline demographic information in ATLAS and in non-biobank UCLA patients” shows baseline demographic information of the UCLA ATLAS population, and a subset of all other adult UCLA patients (Rest of DDR) seen within one year of the ATLAS launch date. All table variables differed significantly between the ATLAS and non-ATLAS populations (Mann-Whitney test for continuous variables and chi-square test for categorical variables) after multiple-testing correction. This included a significant difference overall, as well as in the proportions of alive and deceased individuals. The biobank population experienced a substantially higher number of diagnoses across the most prevalent medical condition categories, including endocrine, cardiovascular, musculoskeletal, gastrointestinal, neoplasms, neurological, respiratory, mental, genitourinary and infections ( Figure S2o – p ). For instance, endocrine/metabolic and cardiovascular diseases were diagnosed in 52.0% and 44% of the biobank participants versus 38% and 36% of non-biobank adults, respectively. Biobank participants also experienced substantially more neoplasms (30% vs. 19%). This likely led to significantly more medical encounters in the biobank population (119 vs. 64 mean total medical encounters per patient, and 16 vs. 12 mean inpatient days), which is consistent with expectations, as greater numbers of clinical visits facilitate the passive blood collection that powers our biobank. The most common diagnoses included hypertension (30% of participants), hyperlipidemia (25%), GERD (18%), anxiety disorder (17%), and depression (15%) ( Figure S2m ). As time progresses, more diagnoses have an opportunity to be added, and the rate of diagnoses is consistent across a wide spectrum of organ systems (1-2 years after collection; Figure S2n ). The table “Characteristics over time” shows time-dependent characteristics of the UCLA ATLAS population (n=59,949) at the time of collection, and a subset of all other adult UCLA patients seen within one year of the ATLAS launch date (“Rest of DDR”, within 1 year of ATLAS launch date; n=186,895). Both cohorts were limited to a subset of individuals who had an encounter between 1-2 years after the initial encounter. Encounter statistics (total encounters, total inpatient encounters, total inpatient days, types of encounters) summarize the encounters within the first year after the initial encounter. The changes within each cohort were tested for a non-zero difference using a Mann-Whitney test for variables available at two times, and a one-sample t-test for the encounter statistics. * denotes a significant difference in the change. The cohorts were also compared at each time point using the Mann-Whitney test (continuous) and the Chi-square test (categorical). After multiple-testing correction, ** denotes a significant difference between the two populations at the respective time (including change). Baseline demographic information in ATLAS and in non-biobank UCLA patients ATLAS-Overall ATLAS-Alive ATLAS-Deceased Rest of DDR-Overall Rest of DDR-Alive Rest of DDR-Deceased n (%) 88,436 84,593 (95.7) 3,843 (4.3) 410,041 392,915 (95.8) 17,126 (4.2) Age, mean (SD) 54.6 (17.0) 54.1 (16.8) 67.6 (14.8) 51.8 (18.5) 50.9 (18.1) 72.8 (14.8) Self-reported sex, n (%): Female 49,346 (55.8) 47,688 (56.4) 1,658 (43.1) 234,189 (57.1) 225,735 (57.5) 8,454 (49.4) Self-reported sex, n (%): Male 39,058 (44.2) 36,873 (43.6) 2,185 (56.9) 175,841 (42.9) 167,169 (42.5) 8,672 (50.6) Self-reported sex, n (%): Unspecified, X 32 (0.0) 32 (0.0) 11 (0.0) 11 (0.0) Self-reported race, n (%): American Indian, Alaska Native 785 (0.9) 762 (0.9) 23 (0.6) 2,233 (0.5) 2,175 (0.6) 58 (0.3) Self-reported race, n (%): Asian 10,537 (11.9) 10,180 (12.0) 357 (9.3) 42,307 (10.3) 40,619 (10.3) 1,688 (9.9) Self-reported race, n (%): Black, African American 4,079 (4.6) 3,857 (4.6) 222 (5.8) 23,420 (5.7) 22,256 (5.7) 1,164 (6.8) Self-reported race, n (%): Caribbean/West Indian 161 (0.2) 161 (0.2) 335 (0.1) 332 (0.1) 3 (0.0) Self-reported race, n (%): Middle Eastern or North African 2,276 (2.6) 2,227 (2.6) 49 (1.3) 5,086 (1.2) 5,017 (1.3) 69 (0.4) Self-reported race, n (%): Native Hawaiian, Guamanian or Chamorro, Samoan, Other Pacific Islander 255 (0.3) 247 (0.3) 8 (0.2) 1,050 (0.3) 1,011 (0.3) 39 (0.2) Self-reported race, n (%): Other Race 3,184 (3.6) 2,737 (3.2) 447 (11.6) 42,188 (10.3) 40,137 (10.2) 2,051 (12.0)  Self-reported race, n (%): Unknown, declined to specify 11,329 (12.8) 11,050 (13.1) 279 (7.3) 62,735 (15.3) 61,692 (15.7) 1,043 (6.1)  Self-reported race, n (%): White 55,830 (63.1) 53,372 (63.1) 2,458 (64.0) 230,687 (56.3) 219,676 (55.9) 11,011 (64.3) Self-reported ethnicity, n (%): Hispanic/Latino, Cuban, Hispanic/Spanish origin, Mexican, Mexican American, Chicano/a, Puerto Rican 12,966 (14.7) 12,235 (14.5) 731 (19.0) 50,229 (12.2) 48,100 (12.2) 2,129 (12.4)  Self-reported ethnicity, n (%): Non-Hispanic/Latino 68,784 (77.8) 65,829 (77.8) 2,955 (76.9) 301,251 (73.5) 287,269 (73.1) 13,982 (81.6)  Self-reported ethnicity, n (%): Unknown, declined to specify 6,686 (7.6) 6,529 (7.7) 157 (4.1) 58,561 (14.3) 57,546 (14.6) 1,015 (5.9) Self-reported smoking, n (%): Former 24,054 (27.2) 22,504 (26.6) 1,550 (40.3) 92,364 (22.5) 85,576 (21.8) 6,788 (39.6) Self-reported smoking, n (%): Never 60,477 (68.4) 58,366 (69.0) 2,111 (54.9) 275,883 (67.3) 266,871 (67.9) 9,012 (52.6) Self-reported smoking, n (%): Passive Smoke Exposure - Never Smoker 112 (0.1) 105 (0.1) 7 (0.2) 826 (0.2) 769 (0.2) 57 (0.3) Self-reported smoking, n (%): Smoker 3,412 (3.9) 3,286 (3.9) 126 (3.3) 27,509 (6.7) 26,796 (6.8) 713 (4.2)  Self-reported smoking, n (%): Unknown, declined to specify 381 (0.4) 332 (0.4) 49 (1.3) 13,459 (3.3) 12,903 (3.3) 556 (3.2) Types of Encounters, n (%): Inpatient and Outpatient Encounters 28,879 (32.7) 25,725 (30.4) 3,154 (82.1) 72,278 (17.6) 62,279 (15.9) 9,999 (58.4)  Types of Encounters, n (%): Only Outpatient Encounters 59,557 (67.3) 58,868 (69.6) 689 (17.9) 337,763 (82.4) 330,636 (84.1) 7,127 (41.6) Total Encounters, mean (SD) 118.8 (146.3) 112.1 (137.5) 267.4 (230.9) 64.2 (95.8) 59.7 (88.1) 166.7 (174.3) Total Inpatient Encounters, mean (SD) 0.8 (2.2) 0.7 (1.9) 3.8 (5.0) 0.4 (1.3) 0.3 (1.1) 2.2 (3.7) Total Inpatient Days, mean (SD) 16.2 (33.2) 13.4 (28.6) 38.3 (53.6) 12.3 (25.8) 9.4 (18.9) 29.7 (47.2) Characteristics over time Variable ATLAS- At Collection ATLAS- 1 Year after Collection ATLAS-Change Rest of DDR-At Collection Rest of DDR-1 Year after Collection Rest of DDR-Change Age, mean (SD) 56.4 (16.4) 57.6 (16.3) 1.2 (0.4)* 53.5 (17.9)** 54.8 (17.8)** 1.3 (0.5)*,** BMI, mean (SD) 27.3 (6.2) 27.3 (28.2) 0.1 (28.7) 26.8 (24.0)** 26.7 (14.8)** −0.1 (27.6) Diastolic BP, mean (SD) 75.5 (9.7) 75.4 (9.4) −0.1 (10.4) 75.4 (10.2)** 75.6 (10.1) 0.2 (10.5)*,** Systolic BP, mean (SD) 125.0 (16.7) 124.8 (16.8) −0.2 (17.1) 125.3 (17.6) 125.3 (17.6) 0.0 (16.8)** Number of ICD Codes, mean (SD) 14.2 (11.2) 15.6 (11.9) 1.5 (6.5)* 9.4 (8.4)** 10.5 (9.1)** 1.1 (5.1)*,** Number of Phecodes, mean (SD) 7.1 (5.7) 7.8 (6.0) 0.7 (3.2)* 4.7 (4.3)** 5.2 (4.6)** 0.5 (2.6)*,** Charlson Comorbidity Index, mean (SD) 1.0 (1.8) 1.1 (1.9) 0.1 (1.0)* 0.7 (1.5)** 0.7 (1.5)** 0.1 (0.8)*,** Elixhauser Comorbidity Index, mean (SD) 2.6 (8.3) 2.9 (8.7) 0.3 (4.4)* 1.6 (6.8)** 1.7 (7.1)** 0.2 (3.6)** Total Encounters, mean (SD) 23.2 (27.2)* 12.0 (15.8)*,** Total Inpatient Encounters, mean (SD) 0.2 (0.7)* 0.1 (0.4)*,** Total Inpatient Days, mean (SD) 1.4 (7.8)* 0.4 (4.2)*,** Inpatient and Outpatient Encounters (%) 8,413 (14.0) 9,869 (5.3)** Only Outpatient Encounters (%) 51,536 (86.0) 177,026 (94.7)** Baseline demographic information in ATLAS and in non-biobank UCLA patients Characteristics over time We retrieved the latest results of the 39 most common laboratory tests: blood count, lipid panel, metabolic panel, HbA1c, and vitamin D,25-Hydroxy to explore test frequencies and variation in results. The largest variability in adults was observed for bilirubin (median = 0.4, IQR = 0.3), followed by triglycerides (median = 90, IQR = 66), alanine aminotransferase (median = 22, IQR = 14), neutrophils (median = 3.73, IQR = 2.07), and LDL cholesterol (median = 97, IQR = 50). Similarly, we retrieved information on the 50 most abundant filled prescriptions to learn about treatment and prescribing tendencies. ( Table S1 ). As hospital visits and surgeries were the most common types of encounters in the biobank population, it is not surprising that among the most prescribed generic medications were Acetaminophen (n = 738,061 prescriptions), Ondansetron HCl (n = 636,165), Propofol (n = 633,830), and Lidocaine HCl (n = 492,838) ( Table S1 ). ATLAS consists of 32% participants of non-European ancestry and a representation of five broadscale and 36 fine-scale genetic ancestry groups within a single health system (see METHOD DETAILS to understand how these were defined). By comparison, the vast majority of the UK biobank 4 participants are EUR (95.8% EUR, 2.1% South Asian, and 2.1% AFR). FinnGen 6 is ~98% Finnish EUR, Taiwan Biobank 7 is over 99% EAS Han Chinese, Michigan Genomics Initiative 11 is 88% EUR, and Geisinger MyCode 169 is approximately 97% EUR. While others like All of Us, Mt. Sinai’s BioMe, and Vanderbilt’s BioVU reflect diverse urban populations, ATLAS captures a wider and more detailed range of genetic ancestries in the Los Angeles population. All of Us 5 includes a larger portion of Black participants, but a smaller portion of Asian and Middle Eastern/North African self-identified individuals (under a combined race and ethnicity category: White 53.3%, Black or African American 21.2%, Hispanic or Latino 17.8%, Asian 3.1%, Middle Eastern or North African 0.6%). Mt. Sinai’s BioMe 53 , the only biobank with reported fine-scale ancestries, included 17 fine-scale clusters vs. 36 in ATLAS, and substantially fewer Asian individuals. Vanderbilt’s BioVU 170 and Penn Medicine Biobank 119 include small Admixed American populations, and these biobanks, along with the VA Million Veterans Program 72 and the Colorado biobank include small Asian populations. The tables below show comparisons of ancestry numbers and ratios across biobanks. Unclassified participants are not included. Some biobank studies have larger sample sizes but lack inferred genetic ancestry and are thus not included. Numbers of biobank participants by genetic ancestry Biobank AMR AFR EUR EAS SAS Total UK Biobank (National) - 9,633 431,805 - 9,252 450,690 All of Us (National) 28,901 34,037 101,613 3,255 - 167,806 Taiwan Biobank (National) - - - 108,955 - 108,955 Mount Sinai BioMe (Academic) 10,638 6,983 8,477 780 617 27,495 Vanderbilt BioVU (Academic) 2,466 15,597 69,810 896 414 89,183 Geisinger MyCode (Academic) - 1,377 44,522 - - 45,899 Michigan Genomics (Academic) 843 5,962 75,943 2,172 1,435 86,355 VA Million Veterans Program (Academic) 59,048 121,117 449,042 6,702 - 635,909 Penn Medicine Biobank (Academic) 711 11,300 30,360 680 573 43,624 Colorado Biobank (Academic) 18,137 7,466 145,070 4,080 - 174,753 UCLA ATLAS 12,822 4,061 62,902 8,080 1,761 89,626 Numbers of biobank participants by genetic ancestry Below, percentages exclude unclassified participants and therefore differ from Figure 1 . Percentages of biobank participants by genetic ancestry Biobank AMR AFR EUR EAS SAS UK Biobank (National) 0.0% 2.1% 95.8% 0.0% 2.1% All of Us (National) 17.2% 20.3% 60.6% 1.9% 0.0% Taiwan Biobank (National) 0.0% 0.0% 0.0% 100.0% 0.0% Mount Sinai BioMe (Academic) 38.7% 25.4% 30.8% 2.8% 2.2% Vanderbilt BioVU (Academic) 2.8% 17.5% 78.3% 1.0% 0.5% Geisinger MyCode (Academic) 0.0% 3.0% 97.0% 0.0% 0.0% Michigan Genomics (Academic) 1.0% 6.9% 87.9% 2.5% 1.7% VA Million Veterans Program (Academic) 9.3% 19.0% 70.6% 1.1% 0.0% Penn Medicine BioBank (Academic) 1.6% 25.9% 69.6% 1.6% 1.3% Colorado BioBank (Academic) 10.4% 4.3% 83.0% 2.3% 0.0% UCLA ATLAS 14.3% 4.5% 70.2% 9.0% 2.0% Percentages of biobank participants by genetic ancestry

Method

Our study moves from characterizing the ATLAS cohort to ancestry-stratified analyses of disease burden – first at the broad-scale ancestry level, then at the fine-scale cluster level- followed by evaluations of common and rare genetic risk, and finally, an integrative case-study using longitudinal EHR data. Array genotypes were obtained using the Global Screening Array. All data was mapped to GRCh38 and dbSNP, build 147 171 . Common haplotypes and variants were imputed using the TOPMed Freeze 5 panel 29 , 30 using 668,127 observed single-nucleotide polymorphisms (SNP), resulting in a total of 50,757,223 high-quality called genotypes variants following imputation, an average of 2,048,050 per individual. Details regarding QC and the imputation procedure were described before 29 . Minimal QC was applied to the imputed genotypes. We retained only non-duplicated, bi-allelic variants with high imputation quality (R2 > 0.7), a minor allele frequency (MAF) > 0.1%, and a missing rate 5% were excluded from further analysis. Concordance between genotypes determined by observed array variants and targeted Illumina sequencing was approximately 99.6%, determined using the vcf-compare command in VCFtools (v0.1.16) 172 . Demographic information, vitals, lab tests, International Classification of Diseases (ICD) codes, and medication prescriptions were retrieved from the UCLA Data Discovery Repository (DDR), established on March 2, 2013 29 , containing deidentified participant EHR data from our health system. Demographic, vital, and lab data were up to date as of Aug 24, 2024. For vitals and encounter counts, only records from in-person encounters (hospital visits, appointments, surgeries, office visits, and walk-ins) were considered. For vital signs, median values across all encounters were used to reduce the impact of potential recording errors or error-prone self-reported values. For lab data, the most recent values were considered. For Figure 1c , self-reported sex and age at the most recent time point at the time of analysis were used. Yearly encounter numbers were retrieved in 2024. We included only complete yearly data up to 2023 to ensure consistency and avoid partial data from 2024. Hospital encounters during the COVID pandemic years (2019-2022) were excluded. For associations involving clinical phenotypes, both ICD-9 and 10 were extracted. To ensure harmonized and comparable phenotypes across data sources, we adopted a structured, standard approach using phecodes 59 – 61 . We utilized the existing PheWAS catalog maps and applied standardized rules for case–control definitions. The ICD-9 codes were mapped to phecodes using phecode Map v1.2 59 , while ICD-10 codes were mapped using phecode Map v1.2b1 60 , both of which were obtained through the PheWAS catalog 61 . Cases were defined using the widely used “Rule of Two” approach 61 , 173 – 176 . For each phecode, participants were classified as cases if they had at least two occurrences of the same phecode, with occurrences spaced at least 30 days apart to ensure persistence of diagnosis. In total, there were 1,308 phecodes with at least 100 cases in ATLAS. Controls were defined as participants who had no record of the phecode in their EHR data and at least two recorded encounters in the system, following similar established rules for minimum data content 61 , 177 , 178 . Undecidable participants were excluded. We used phecodes 59 – 61 to define phenotypes and validated the generalizability of these phenotypes using several approaches. First, we tested cross-checked information between a variety of vital signs and lab tests and specific case-control groups defined by phecodes, using Wilcoxon tests and density plots, which confirmed the expected differences for these phenotypes. This included the following: 1) phecode-lab test pairs: Type 2 diabetes-Hemoglobin A1c, Hyperlipidemia-Trygllicoride, Vitamin D deficiency-Total 25-Hydroxy vitamin D, Deficiency anemias-Mean Corpuscular Volume (MCV), Deficiency anemias-Hemoglobin, and 2) the phecode-vital sign pairs: Essential hypertension-Blood pressure diastolic, Essential hypertensionBlood pressure systolic, Obesity-BMI, Anorexia nervosa-BMI, Short stature-Height ( Figure S2a – j ). Second, we replicated many known associations, including ancestry-related differences in disease prevalence for 28 major phenotypes. For example, across cancer types, we observed the highest risk for prostate cancer in AFR ancestry, stomach cancer in EAS, bladder cancer in AMR, and breast cancer in EUR 179 (see more below and in Table S2 ). Across fine-scale ancestries, for example, we showed increased breast cancer and Crohn’s disease diagnoses in the Ashkenazi Jewish (IBD-03) cluster 53 , 180 (see more below). We also replicated 11,756 variant–phenotype associations, such as a lower risk of Alzheimer’s disease for African American participants with the ε4/ε4 haplotype (see more below). PGS prediction of 28 traits revealed significant enrichment in the top PGS decile for 89% of these traits in EUR (see more below). Collectively, the replication of known associations using ATLAS-defined phecodes indicates high-quality, well-defined phenotypes that match external definitions. Third, we used EHR data on cancer diagnoses to validate prostate and breast cancer cases (the two most diagnosed cancers) against phecode-defined case-control groups. We found a high agreement of 95%, for both cancers, considering decidable patients based on phecodes ( Figure S2k – l ). The following sections describe replication results conducted to validate the quality of the data and the phenotype definitions. We assessed variation in medical conditions and disease prevalence across populations across preselected prevalent major conditions ( Table S2 ). This confirmed known ancestry-disease associations. Across cancer types, we observed the highest risk for prostate cancer in AFR ancestry, stomach cancer in EAS, bladder cancer in AMR, and breast cancer in EUR 179 . We also replicated several known cardiovascular disease associations. For instance, AFR participants were more affected by hypertension and myocardial infarction, while those with SAS ancestry had the lowest risk of atrial fibrillation, despite the highest incidence of coronary atherosclerosis, an established but seemingly contradictory risk profile 181 . Major metabolic disorders were more prevalent in AFR than in other ancestries. This finding included diagnostic codes (9- and 10-ICD codes) related to Type 2 diabetes, hypercholesterolemia, and hyperlipidemia, supported by increases in metabolic-related laboratory measurements and vital signs ( Figure S3e ). Type 2 diabetes was more common in all non-EUR groups relative to EUR. Previous work has suggested that some of these differences reflect social and environmental, rather than genetic, factors 182 . With regard to neurological conditions, we find a higher risk of dementia in AFR participants, but a lower risk of migraines and Parkinson’s Disease in AFR compared to EUR patients, consistent with prior analyses 183 – 185 . Among neuropsychiatric disorders, anxiety and major depressive disorders were most frequent in EUR participants, and least frequent in the continental Asian cluster, as previously reported 186 , 187 . These findings demonstrate the utility of ATLAS in robustly replicating known associations within a single health system, reducing potential confounding due to geographic or health system-related effects. We tested the prevalence of 1,253 phecodes 59 – 61 across fine-scale clusters with at least 100 participants ( Table S3 ; selected tests, Figure 2c ). As a measure of integrity, we tested for known increases in prevalence in these fine-scale ancestries, for example, replicating an increase in breast cancer and Crohn’s disease in the Ashkenazi Jewish (IBD-03) cluster. Similarly, we observed the known increase in gout in the Filipino cluster (IBD-09) 188 and Alzheimer’s disease and dementias in the Puerto Rican cluster (IBD-15) 189 . Consistent with the global ancestry findings, African American (IBD-06) had the highest prevalence of hypertension. We leveraged genotyping in our cohort to calculate individual PGS for a range of cancer, cardiovascular, metabolic, neuro-psychiatric and autoimmune diseases and tested their relationship to disease risk. We focused on the top end of the PGS distribution (10%), compared with the 5 th decile, in EUR participants ( Figure 3 ; see the table below). For Type 1 diabetes, the top PGS decile was the most enriched (OR = 11.6 [7.6, 19.0], FDR = 4.1×10 −25 ), with 41% of the diagnosed participants in ATLAS assigned to the top PGS decile. The second most enriched trait was Crohn’s disease, with 33% of diagnosed participants within the top PGS decile (OR = 5.5 [4.0, 7.7], FDR = 7.9×10 −24 ), followed by gout with 27% (OR = 5.1 [4.1, 6.4], FDR = 2.0×10 −45 ), testicular cancer with 25% (OR = 3.7 [1.2, 7.3], FDR = 1.2×10 −4 ), and prostate cancer with 21% (OR = 3.4 [2.8, 4.0], FDR = 2.0×10 −45 ) ( Figure S5a – e ). On average, the top decile of risk accounted for 18.6% of diagnosed participants across 28 preselected prevalent disorders. Twenty-five of 28 (89%) showed significant enrichment of participants within the top PGS decile (mean OR for significance tests = 2.9). Performance declined when we applied PGS to non-EUR populations, as expected 55 – 58 . This was due to both reduced sample sizes and the model fit ( Figure S5f – i ), identifying only 17, 27, 36, and 38% significant associations of diseases with the top PGS decile, for SAS, AFR, EAS, and AMR, respectively. Concordantly, the top PGS decile accounted for fewer cases (on average, 13, 15, 16 and 16%, for SAS, AFR, EAS and AMR). This further supports the need for larger, more ancestrally heterogeneous cohorts for clinical development of PGS 55 – 58 . The OR and prevalence of cases for the top and bottom PGS deciles across traits in EUR Trait OR top OR bottom Cases top (%) Cases bottom (%) Cases sample size FDR top FDR bottom Atrial fibrillation 3.2 0.5 19.5 4.8 4565 6.60×10 −66 1.90×10 −13 Bipolar 1.6 0.7 15.5 7.4 1147 9.60×10 −5 0.011 Bladder cancer 1.8 0.8 15.2 6.9 698 0.00028 0.3 Breast cancer 2.4 0.5 18.0 4.9 3016 2.20×10 −26 2.40×10 −08 Cerebrovascular disease 1.3 1.1 9.9 10.5 334 0.34 0.71 Colorectal cancer 1.6 0.6 15.0 5.9 842 0.0026 0.0064 Coronary atherosclerosis 2.5 0.5 16.0 6.1 7951 2.20×10 −59 2.70×10 −24 Gout 5.1 0.4 27.2 2.7 1638 2.00×10 −45 1.10×10 −05 Hypercholesterolemia 1.4 0.6 12.7 6.8 10836 1.50×10 −14 5.70×10 −24 Hyperlipidemia 1.6 0.6 12.2 7.8 21829 6.70×10 −29 1.30×10 −30 Hypertension 2.1 0.6 12.6 7.8 21517 2.50×10 −63 1.10×10 −28 Hypertrophic obstructive cardiomyopathy 1.7 0.4 17.8 3.9 129 0.13 0.11 Lung cancer 1.1 0.7 10.3 7.2 976 0.65 0.062 Major depressive disorder 1.6 0.7 13.0 7.1 10793 3.00×10 −22 1.50×10 −10 Malignant neoplasm of testis 3.7 0.4 24.9 2.9 173 0.00012 0.13 Melanoma 2.3 0.5 19.6 3.4 1196 8.80×10 −12 1.00×10 −4 Multiple sclerosis 2.8 0.3 26.1 2.9 379 6.00×10 −7 0.0024 Myocardial infarction 1.6 0.8 12.9 7.2 1744 2.30×10 −5 0.11 Ovarian cancer 2.1 0.8 17.9 6.7 403 0.00066 0.35 Pancreatic cancer 1.6 0.7 12.6 5.2 382 0.048 0.29 Prostate cancer 3.4 0.3 21.5 2.9 2703 2.00×10 −45 1.10×10 −17 Psoriatic arthropathy 1.5 0.9 14.1 8.0 503 0.029 0.53 Crohn’s disease 5.5 0.6 33.0 3.6 758 7.90×10 −24 0.11 Schizophrenia 2.9 0.3 21.9 2.3 128 0.01 0.11 Systemic lupus erythematosus 2.9 0.9 22.0 6.3 569 4.40×10 −9 0.56 Thyroid cancer 3.0 0.6 21.4 4.1 786 2.60×10 −12 0.022 Type 1 diabetes 11.6 0.4 40.8 1.5 537 4.10×10 −25 0.05 Type 2 diabetes 2.3 0.4 16.9 4.3 6302 2.20×10 −48 8.60×10 −26 The OR and prevalence of cases for the top and bottom PGS deciles across traits in EUR Replication of genotype-phenotype associations is important to assess the quality of both genetic and phenotype data, as well as to evaluate the power and effectiveness of the biobank in detecting true genetic associations. To address this, we first examined the distribution of APOE alleles across fine-scale clusters, highlighting the increased frequency of ε4 risk alleles in the African American (IBD-06) and Bantu (IBD-35) clusters ( Figure S6c ) and replicating the finding that African American participants with the ε4/ε4 haplotype have a lower risk of Alzheimer’s disease compared to other populations (see in the table below FDR-significant results). Next, we replicated two well-known genetic associations 190 in the African American cluster, between HBB rs334-A ( Figure S6d ; MAFIBD-06 = 5.2%) and a diagnosis of sickle cell anemia (PIBD-06 = 2.0×10 −78 ; ORIBD-06 = 15.1) and between the Duffy null ACKR1 rs2814778-C polymorphism ( Figure S6d ; MAFIBD-06 = 77.1%) and a decrease in neutrophil count (PIBD-06 = 4.0×10 −31 ; ORIBD-06 0.7). Using ATLAS blood work results, we further quantified the impact of rs334-A on mean corpuscular hemoglobin concentration (PIBD-06 = 5.2×10 −15 ; βIBD-06 = 0.4), nucleated red blood cell count (PIBD- 06 = 1.6×10 −10 ; βIBD-06 = 0.2), and mean corpuscular volume (MCV) (PIBD-06 = 3.3×10 −8 ; βIBD-06 = − 0.3). Across multiple Asian clusters, we confirmed the association between a variant in high LD with the -- SEA deletion, which causes inherited alpha-thalassemia 191 , and microcytic anemia. We observed a significant decrease in MCV in response to LUC7L rs372755452-A in the broad-scale EAS ancestry (PEAS = 6.1×10 −88 ; βEAS = −1.6; MAFEAS = 0.9%), and the fine scale Chinese + Korean (PIBD-05 = 2.7×10 −52 ; βIBD-05 = −1.6; MAFIBD-05 = 0.8%) and Filipino (PIBD-09 = 2.2×10 −24 ; βIBD-09 = −1.8; MAFIBD-09 = 1.2%) clusters. Next, we replicated the finding that AMR participants are twice as likely to carry a PNPLA3 rs738409-G missense variant that greatly increases the risk for non-alcoholic fatty liver disease ( Figure S6d ), a major cause of cirrhosis that often necessitates liver transplant 67 . Using fine-scale ancestral mapping, we studied the association between rs738409-G and non-alcoholic cirrhosis of liver in two Mexican American clusters ( Figure S6e ), IBD-04 (PIBD-04 = 2.7×10 −14 ; ORIBD-04 = 1.90) and IBD-07 (PIBD-07 = 1.5×10 −8 ; ORIBD-07 = 1.9), identifying similar risk but differing minor allele frequencies across these clusters (MAFIBD-04 = 45.0%; MAFIBD-07 = 52.9%). The same variant in Northern Europeans had a smaller effect (PIBD-01 = 3.6×10 −9 ; ORIBD-01 = 1.5) and was half as frequent (MAFIBD-01 = 22.9%). Ancestry-specific effects of APOE ε4 on Alzheimer’s disease risk Haplotype Ancestry OR P-value CI low CI high e4e4 All 2.87 1.25×10 −41 2.45 3.28 e4e4 EUR 2.86 1.40×10 −31 2.38 3.34 e3e4 All 1.06 3.88×10 −19 0.82 1.29 e3e4 EUR 1.01 1.43×10 −13 0.74 1.28 e4e4 AMR 4.18 5.99×10 −6 2.37 5.99 e3e4 EAS 1.93 1.23×10 −4 0.95 2.92 e4e4 AFR 1.76 3.10×10 −3 0.59 2.93 e4e4 EAS 4.89 4.59×10 −3 1.51 8.26 e3e4 Unclassified 1.80 8.83×10 −3 0.45 3.15 e4e4 Unclassified 5.66 9.62×10 −3 1.37 9.94 Ancestry-specific effects of APOE ε4 on Alzheimer’s disease risk Analyzing clinically relevant rare variants, known to have elevated frequencies in specific populations, replicated known ancestry-associated patterns, assessing the quality of the data and the robustness of ATLAS in capturing ancestry-associated allele frequency patterns. Focusing on Familial Mediterranean Fever (FMF) variants, we identified 753 carriers with the highest risk in the Armenian clusters 192 , 193 ( Figure S7a ). Carriers of these variants also had a higher risk of the amyloidosis phecode, which is known to be associated with FMF (OR = 3.7 [2.2, 5.8]). The HBB:p.E7V variant, which is responsible for most sickle cell anemia cases, was carried by 273 participants, with elevated frequency in the African American cluster (OR IBD-06 = 51.4 [39.7, 67.0]; Figure S7b ), in line with previous findings 194 . Finally, we explored protective variants in the PCSK9 gene causing lowered LDL levels 195 and identified 49 carriers of loss-of-function variants within the African American cluster (OR IBD-06 = 9.5 [4.8, 17.8]) ( Figure S7c ), as is known 20 . While the above variants were selected via a literature search, we also queried the entire ClinGen pathogenic or likely pathogenic (P/LP) variant catalog, aggregated by gene, to identify enrichment in fine-scale populations. Among replicated associations, variants in BRCA1 and BRCA2 were our first candidates of interest. Ashkenazi Jewish had the primary risk for carrying either ClinGen P/LP variants in BRCA1 or BRCA2 ( BRCA1 , OR IBD-03 = 47.1 [20.6, 133.0]; BRCA2 , OR IBD-03 = 48.2 [23.6, 114.2]) ( Figure 5a ). Of note, these observations are based on six P/LP BRCA ClinGen variants found in ATLAS, which include only two out of three of the Ashkenazi Jewish founder alleles (at the time of the analysis, the Ashkenazi Jewish founder allele NM_007294.3 :c.5266dup was not included in the ClinGen Evidence Repository). The population attributable risk (PAR) was 3.0 [1.8, 5.0]% and 3.2 [1.9, 5.3]% for BRCA1 and BRCA2 variants, respectively, in the Ashkenazi Jewish cluster for breast cancer. Independent of ClinGen, we considered the three BRCA Ashkenazi Jewish founder alleles and found that they were most prevalent in EUR among broad-scale ancestries, and in Ashkenazi Jewish among fine-scale clusters ( Figure S7d – e ). In Ashkenazi Jewish, the PAR for breast cancer, associated with the three Ashkenazi Jewish founder alleles combined, was 8.2 [5.9, 11.2]% when all ages were included and increased to 12.5 [8.5, 18.1]% in participants under 70. This increase is expected, since BRCA founder alleles are associated with early-onset breast cancer. In Northern Europeans, we observed enrichments of P/LP variants in MYOC (OR IBD-01 = 2.7 [1.7, 4.2]) that are associated with juvenile open angle glaucoma, with a PAR of 0.7 [0.2, 3.4]%. A bias toward EUR in ClinGen variants was recently shown in AoU 23 . We tested whether we could replicate this bias using a single health system. In this part, we asked if the total allele frequencies of the clinically actionable ACMG ClinGen P/LP variants vary across broad- and fine-scale ancestries. Overall, 17 ACMG genes had at least one P/LP variant in ClinGen present in ATLAS participants. For a more nuanced biological understanding, we divided the ACMG variants into two groups of rare LOF (n = 53) and rare P/LP missense (n = 131) variants. We defined ‘LOF’ for variants ranked as high-confidence LOF by LOFTEE 94 and ‘missense’ based on the VEP 137 “missense_variant” annotation (see below, Differences in total ClinGen allele frequency within ACMG genes across ancestries ). We observed the highest frequency of rare P/LP LOF variants in EUR participants among broad-scale populations, and in the Ashkenazi Jewish cluster (IBD-03) among fine-scale clusters ( Figure 5b ). This number was primarily driven by BRCA P/LP alleles, which accounted for 63% of all ClinGen LOF P/LP variants in EUR and 94% in the Ashkenazi Jewish cluster (OR EUR = 3.7 CI EUR = 2.6-5.4, P Bonferroni = 6.3×10 −17 ; OR Ashkenazi Jewish = 6.5, CI Ashkenazi Jewish = 5.3 - 8.1, P Bonferroni = 3.5×10 −62 ). Still, when Ashkenazi Jewish individuals are removed from the broad-scale EUR category, a substantial enrichment of ClinGen missense variants is observed for this group (OR EUR (non-Ashkenazi Jewish) = 1.8 [1.2, 2.7], P-value = 0.002, P Bonferroni = 0.01; Figure S7f ), demonstrating that the signal is not derived from Ashkenazi Jewish individuals alone. No differences in rare P/LP missense total counts were identified across broad-scale ancestries, but across fine-scale ancestry clusters, Northern EUR participants (IBD-01) had significantly more (OR Northern EUR = 1.4 [1.2, 1.8], P Bonferroni = 0.02), with a similar trend observed in Southern Europeans that did not reach statistical significance. This suggests that clinical datasets are particularly biased toward Northern EUR driven by the composition of large contributing cohorts such as the UK Biobank. Missense variants were depleted in Ashkenazi Jewish, possibly since they are underrepresented in most EUR large scale cohorts, or due to technical differences in how studies or countries classify variants as P/LP. Excluding Ashkenazi Jewish as a sensitivity analysis resulted in a similar, albeit weaker trend, of enrichment of missense variants in Northern EUR (OR = 1.3 [1.01, 1.6], P-value = 0.04, P Bonferroni = 0.5; Figure S6f ). As all samples were collected incidentally, there was interest in characterizing the population, especially in comparison to non-biobank UCLA patients. To characterize the non-biobank UCLA patients while mitigating time-dependent confounding, we included a subpopulation with an encounter within one year of the ATLAS launch date. Encounters could be either inpatient or outpatient, and we summarized the relative proportion of patients with only outpatient encounters or at least one inpatient encounter (there were no patients with only inpatient encounters). To identify clinical phenotype patterns in ATLAS and to compare these to patterns of non-biobank UCLA patients, we used all ICD codes from each patient to encapsulate past and present conditions, subject to practical challenges 196 . For the ATLAS population, we captured a snapshot using their encounter closest to their biobank sample collection date. For the non-biobank patients, we used their closest encounter to the ATLAS launch date. These encounters represented the baseline encounters for both populations. The ICD codes at these encounters were converted to phecodes and phecode groups 59 – 61 , 197 to represent meaningful categories of disease. Phecodes were extracted using pandas v2.2.2 with Python v3.9.19. We reported the unique phecodes with a prevalence ≥5% in the UCLA ATLAS population. ICD codes at the baseline encounters were also used to calculate the Charlson and Elixhauser Comorbidity Indices 40 , 41 , 198 , both measures that predict mortality. The scores were derived using the comorbidity R function v1.0.7 154 with R v4.1.2. To compare the prevalence of disease categories between UCLA Biobank participants and non-biobank UCLA Health patients, we used logistic regression to test the association of each category with participant group, adjusting for age and sex. In addition to characterizing patient populations at a baseline time, we also described differences in how they changed over time, an important consideration when assessing relative disease burden across populations. Participants were enrolled in UCLA ATLAS and were encountered within the health system at different times; therefore, we controlled the interval over which we measured change. We identified patients with at least one encounter between one and two years after their baseline encounter. From this subpopulation, we obtained patients’ first encounter within this window and summarized the cumulative encounters that occurred between the two times. We also used the ICD codes at the second encounter to derive new comorbidity scores, as well as extract updated phecodes. The change in comorbidity scores and increases in phecode prevalence provided insight into the evolving disposition of the UCLA ATLAS population over time. The genetic ancestry of participants in the ATLAS dataset was estimated by assessing their proximity to the centroids of 1000 Genomes superpopulations in principal component (PC) space. PCA across common genetic variants demonstrates granular relationships and provides a quantitative basis for assessing relationships between ancestry and disease 29 , 36 . The top 20 PCs were calculated using the bed_projectPCA function (the bigsnpr 155 R package v1.12.2) with default parameters. For every individual, the Euclidean distance to the centroids of the five broad-scale populations (AMR, AFR, EUR, EAS, SAS) was computed. Participants were assigned AMR or AFR ancestry if the nearest centroid corresponded to one of these populations, as these groups are well-separated in PC space. For EUR, EAS, and SAS ancestries, which exhibit more genetic overlap, a stricter distance threshold was enforced to minimize misclassification. Specifically, an individual was assigned to one of these ancestries if their distance to the nearest centroid was less than a scaled threshold, calculated as: Threshold=max(dist)×min(FST)/max(FST)×0.5, where max(dist) is the largest squared distance among centroids and min(FST)/max(FST) accounts for genetic differentiation. Participants who could not be assigned to any ancestry cluster under these criteria were labeled as “admixed/unknown.” Visualization was performed using the Boutros Plotting General (BPG) R package v.7.1.0 156 . Variation in encounter numbers across broad-scale ancestries was tested using ANCOVA, adjusted for genetic sex, age, and BAS rank to account for healthcare access differences. Comorbidity index values across broad-scale ancestries were compared using ANCOVA, adjusted for genetic sex, age, and both BAS and ADI ranks, which provide complementary measures of socioeconomic status. Since not all patients had BAS and ADI values, the sample sizes for these analyses were reduced (numbers are specified in the body of each figure). Adjusted means and CI were calculated using the R emmeans package, version 1.10.5 132 . To test differences in disease diagnosis across ancestries, phecodes (retrieved as described in Retrieving phenotype data ) were associated with genetic ancestry populations using logistic regression, adjusting for age (age at diagnosis for cases and the latest age for controls) and genetic sex if applicable. The following disease-phecode pairs were used to define disease diagnosis: Type 2 diabetes-Type 2 diabetes, Sleep apnea-Sleep apnea, Essential hypertension-Essential hypertension, Hyperlipidemia-Hyperlipidemia, Anxiety disorders-Anxiety disorders&Anxiety disorder&Generalized anxiety disorder, Asthma-Asthma, Parkinson’s disease-Parkinson’s disease, Type 1 diabetes-Type 1 diabetes, Schizophrenia-Schizophrenia, Crohn’s disease-Regional enteritis, Chronic kidney disease-Chronic kidney disease, Stage I or II, Multiple sclerosis-Multiple sclerosis, Major depressive disorder-Major depressive disorder, Cerebrovascular disease-Cerebrovascular disease, Atrial fibrillation-Atrial fibrillation, Hypercholesterolemia-Hypercholesterolemia, atherosclerosis-Coronary atherosclerosis, Hyperlipidemia-Hyperlipidemia, Coronary Hypertrophic obstructive cardiomyopathy-Hypertrophic obstructive cardiomyopathy, Myocardial infarction-Myocardial infarction, Systemic lupus erythematosus- Systemic lupus erythematosus , Gout-Gout, Bipolar-Bipolar, Psoriatic arthropathy-Psoriatic arthropathy, Epilepsy-Epilepsy, Neurofibromatosis-Neurofibromatosis, Dementias-Dementias, Obesity-Obesity, Obsessive-compulsive disorders-Obsessive-compulsive disorders, Autism- Autism, Migraine- Migraine, Alzheimer’s disease- Alzheimer’s disease, Coronary atherosclerosis-Coronary atherosclerosis, Posttraumatic stress disorder-Posttraumatic stress disorder. Patients under 18 years old, with ambiguous sex or with “unclassified” genetic ancestry were excluded. Visualization was performed using the BPG R package v.7.1.0 156 . PGS were calculated using array data after imputation in EUR participants. Related participants based on their genetic similarity were excluded (defined using PLINK v2.0a 109 with the relatedness coefficient --king-cutoff 0.05). PGS were calculated using pgsc_calc 158 with the default settings and --min_overlap of 0.65. Logistic regression was used to associate every PGS with the corresponding trait based on phecodes (see Retrieving phenotype data ). The following PGS model IDs from the PGS catalog 199 and their phecode pairs were tested: PGS002250-Malignant neoplasm of ovary, PGS003766-Cancer of prostate, PGS000004-Malignant neoplasm of female breast&Breast cancer, PGS001794-Thyroid cancer, PGS000079-Melanomas of skin, PGS004884-Cancer of bronchus; lung, PGS002264-Pancreatic cancer, PGS003395-Colorectal cancer&Colon cancer, PGS000729-Type 2 diabetes, PGS002025-Type 1 diabetes, PGS004254-Regional enteritis, PGS004699-Multiple sclerosis, PGS000134-Schizophrenia, PGS004760-Major depressive disorder, PGS001806-Malignant neoplasm of testis, PGS004687-Malignant neoplasm of bladder & Cancer of bladder, PGS000039-Cerebrovascular disease, PGS004526-Essential hypertension, PGS004706-Atrial fibrillation, PGS004784-Hypercholesterolemia, PGS002029-Hyperlipidemia, PGS003726-Coronary atherosclerosis, PGS000739-Hypertrophic obstructive cardiomyopathy, PGS004528-Myocardial infarction, PGS000803-Systemic lupus erythematosus, PGS001789-Gout, PGS002786-Bipolar, PGS000198-Psoriatic arthropathy. For prostate and testicular cancer, only males were included, and for breast and ovarian cancer, only females. An adjustment was made for age at diagnosis for cases and the latest age for controls, genetic sex when both sexes were included, and the first ten genetic PCs. For Figure 3 , the top and bottom PGS deciles compared with the 5 th decile were considered, testing only EUR participants. FDR was used for multiple testing correction. The same process was repeated for other ancestries as presented in the supplementary material . Visualization was generated using the BPG R package v.7.1.0 156 . ATLAS array data were merged with genotyping data from the 1000 Genomes Project 38 , the Simons Genome Diversity Project 130 , and the Human Genome Diversity Project 131 . BCFtools 128 annotate was used to harmonize variant reference SNP ID (RSIDs), and BCFtools 128 norm with a GRCh38 genome reference was used to standardize the genotyping data. Sites or individuals with more than 1% missing were removed using PLINK 133 --mind and --geno. Only SNPs with MAF > 1% across all participants were kept. SHAPEIT5 134 with default parameters and the distributed GRCh38 map files were used to phase genotyping data, one chromosome at a time. A custom Python script that converts PLINK bed files to PLINK ped/map 133 files while conserving phasing information was used to convert genotyping data. Centimorgan data for the map files were generated using the same genetic map data in SHAPEIT5. Identity-by-descent segments were called using iLASH 135 with the following parameters: slice_size 350, step_size 350, perm_count 20, shingle_size 15, shingle_overlap 0, bucket_count 5, max_thread 20, match_threshold 0.99, interest_threshold 0.70, min_length 2.9, auto_slice 1, slice_length 2.9, cm_overlap 1 and minhash_threshold 55. Identity-by-descent was called one chromosome at a time. Identity-by-descent segment outliers were removed as described in Belbin et al. 53 and Caggiano et al. 52 . Segments overlapping centromeres or the human leukocyte antigen (HLA) region were removed. Regions that may have false positive identity by descent were identified using the following process and removed: total identity by descent per each SNP was identified by summing across all identity-by-descent segments that overlapped each SNP; SNPs with a total identity by descent greater than or less than three standard deviations from the genome-wide mean were removed. To identify clusters, we followed the approach of Caggiano et al. 52 and Dai et al. 54 and applied Louvain clustering 200 . An undirected network is generated based on pairs of individuals who share identity-by-descent segments: nodes are the individuals, edges are the total, genome-wide identity-by-descent shared as the edges. We used the Python package, NetworkIt 201 , to iteratively run Louvain clustering four times to detect fine-scale clusters. To avoid redundancy and maximize sample size, clusters were merged in two stages to produce a final set of consensus fine-scale clusters. First, following Caggiano et al. 52 and Dai et al. 54 , we computed pairwise Hudson’s F ST using PLINK v2.0a 136 across 378 clusters identified from the fourth layer of Louvain clustering. After removing clusters with fewer than 10 participants, to avoid unreliable F ST estimates, we merged the remaining 356 clusters into 67 clusters if the pairwise F ST was less than 0.001. In the second stage, we refined clusters using IBD sharing. For each cluster pair, we examined all inter-cluster individual pairs to calculate: (1) IBD mean , the average cM shared, and (2) IBD prop , the proportion of pairs sharing at least 3cM of IBD as detected by iLash. We defined a composite metric for cluster merging, IBD weighted = IBD mean ×IBD prop , that captures both the extent and prevalence of genomic IBD sharing. Next, clusters were sorted from the smallest to the largest number of ATLAS participants. For each cluster, we computed the IBD weighted score with all larger clusters and calculated the mean of these pairwise values. The cluster was then merged into the most similar larger cluster – i.e. , the one with the highest IBD weighted score-only if that score exceeded the mean. Otherwise, the smaller cluster was retained independently. This process merged 14 small clusters into larger parent clusters, resulting in a total of 36 fine-scale clusters with ≥ 30 participants each for downstream analyses 52 . These fine-scale clusters were assigned unique identifiers (IBD-01 through IBD-36) and manually annotated with labels to ease interpretation. Due to filtering out clusters of small sample sizes before and after merging, not all individuals were assigned to a fine-scale ancestry cluster. We primarily relied on reference individuals, described in “ Fine-scale Ancestry Pre-processing and Quality Control ”, to add labels to clusters. When clusters did not contain reference individuals or were heterogeneous, we utilized patient-reported race, ethnicity, preferred language, and religious affiliation (in order of priority) to inform our cluster labeling. These aspects are not caused by identity-by-descent segment sharing but can be indicative of a shared culture for individuals within a cluster; these shared practices can influence a group’s demography, environment, and disease risk. Our labels are not definitive and are our best attempt to generate informative assignments for each cluster. We used logistic regression to model the association between fine-scale ancestry assignment and phecode prevalence, estimating OR with 95% confidence intervals using the logistf R package v1.26.0 for Firth’s bias-reduced penalized-likelihood logistic regression 157 . Differences were tested between every fine-scale ancestry cluster with at least 100 participants (with no ambiguous genetic sex and over the age of 18) and all other ATLAS participants. For phecodes that are sex-specific, only participants of the corresponding sex were included in the analysis. Adjustment was made for age at diagnosis for cases, and the latest age for controls, and for sex when applicable. Phecodes with defined categories according to the PheWAS catalog 61 , and that were diagnosed in at least 100 participants with no ambiguous genetic sex and over the age of 18 in total in ATLAS were tested, resulting in 1,253 phecodes. In situations where the number of cases in a cluster was small, exact ORs were not reported in the text to protect patient privacy. For visualization, fine-scale clusters were grouped into four panels (EUR, Asian, AMR, and AFR) based on the predominant broad-scale ancestry of participants. If the predominant ancestry was “Unclassified” (as in the Japanese and Egyptian Christian clusters), the second most prevalent broad-scale ancestry was used. The full phecode names presented in Figure 2c were: Hormones and synthetic substitutes causing adverse effects in therapeutic use, Cardiomegaly, Allergic reaction to food, Cirrhosis of liver without mention of alcohol, Vitamin B-complex deficiencies, Anemia of chronic disease, Cholesterolosis of gallbladder, Chronic renal failure [CKD], Dementias, Cancer of bladder, Cataract, Leukemia, Hereditary hemochromatosis, Attention deficit hyperactivity disorder, Glaucoma, Hyperplasia of prostate, Amyloidosis, Chronic pulmonary heart disease. These names were shortened in the figure and text for easier reading. The same model was used to test differences in cardio-metabolic diseases using appropriate phecodes, for each fine-scale cluster within the same broad-scale continental ancestry. For this goal, clusters were assigned to broad-scale ancestries based on the predominant ancestry match among participants within each cluster. The largest population cluster within each broad-scale ancestry was used as the reference level for associations, adjusting for BMI, sex and age at diagnosis for cases, and the latest age for controls. As a sensitivity analysis, we tested whether the observed patterns persisted when correcting for SES factors (adjusting for ADI and BAS ranks) in addition to BMI, age, and sex. As a second step, we also added IPW (see IPW analysis ). FDR was used for multiple testing correction. The data were visualized using the BPG R package v.7.1.0 156 . The following studies were used to compare ATLAS’s broad-scale genetic diversity with other large scale biobanks: Halldorsson et al . (UKBB) 4 , Kurki et al . (FinnGen) 6 , Feng et al . (Taiwan Biobank) 7 , Zawistowski et al. (Michigan Genomics Initiative) 11 , Verma et al . (Geisinger MyCode) 169 , The All of Us Research Program Genomics Investigators et al. 5 , Shaw et al . (Vanderbilt’s BioVU) 170 , Verma et al . (Penn Medicine Biobank) 119 , Verma et al . (VA Million Veterans Program) 72 and Wiley et al . (the Colorado biobank). Studies that were used to compare sample sizes of ATLAS’s fine-scale ancestry clusters with other published cohorts with available genetic data ( Table S3 ) included: Belbin et al . (BioMe fine-scale IBD clusters) 53 , Wu et al . and Li et al . (Ashkenazi Jewish) 57 , 202 , Haber et al . and Hovhannisyan et al . (Armenian) 55 , 56 , Mehrjoo et al . (Iranian) 203 , Larena et al . (Filipino) 58 , Sohail et al . and Ziyatdinov et al . (Mexican) 204 , 205 . IPW were calculated using logistic regression in R with ATLAS participants as cases and other UCLA Health patients as controls. As described above, we included UCLA Health patients who had at least one encounter within one year of the ATLAS launch date. The following covariates were included: self-reported race, sex, age group, BAS rank, and ADI rank. Individuals of unknown race were excluded. For each ATLAS participant, the probabilities were extracted from the model, and the weights were defined as 1 divided by the predicted probabilities. To assess whether previously unreported ancestry and disease status associations hold after applying IPW, we used the svyglm function (R survey package v4.4.8 164 ) with a quasibinomial model, adjusting for age, sex, BAS rank, and ADI rank, including the calculated weights for ancestry groups with more than 5 cases. To compare comorbidity index values across broad-scale ancestries, considering IPW, linear regression was used, adjusting for age, sex, BAS rank, and ADI rank, including the calculated weights. Adjusted means and CI were obtained using the R emmeans package, v1.10.5 132 . Patients under 18 years old, with ambiguous sex or with “unclassified” genetic ancestry were excluded. PheWAS were conducted using the Regenie v4.0 framework 66 for all five broad-scale populations (AMR, AFR, EAS, EUR, SAS) and fifteen fine-scale clusters with at least 400 participants (IBD-01 through IBD-15). Related individuals were removed using a 0.05 kinship cutoff in PLINK v2.0a 136 . Sample sizes for each tested group can be found in Table S4 . Association testing was performed separately within each group ( i.e. , population or cluster) for 1,437 binary traits and 41 quantitative traits, including ICD-derived diagnoses and clinical laboratory measurements (see Retrieving phenotype data ). Samples were restricted to those in predefined inclusion lists (--keep), and trait-specific covariates were provided via --covarFile, including age, sex (modeled categorically with –catCovarList sex), BMI, and the top 10 genetic PCs. Quantitative traits were divided into two analysis groups based on missingness patterns, following Regenie’s recommendation that traits with similar levels of missing data are modeled together in Step 1. For each group, unimputed array genotypes in bed format were filtered using PLINK v2.00a 136 to produce a minimal set of variants suitable for estimating genomewide polygenic effects through ridge regression in step 1 of Regenie. We included autosomal variants with call rate ≥99% (--geno 0.01), minor allele frequency ≥1% (--maf 0.01), and Hardy-Weinberg equilibrium p > 1×10 − 15 (--hwe 1e-15). Additional linkage disequilibrium (LD) pruning was applied using a sliding window of 1000 SNPs, advanced by 100 SNPs, with an r2 threshold of 0.9 (--indep-pairwise 1000 100 0.9), producing an average of 389,209 SNPs per group. For binary traits, step 1 was run with a minimum case count of 50 (--minCaseCount 50). For quantitative traits, step 1 was run with Rank Inverse Normal Transformation (--apply-rint) enabled to stabilize variance across lab values with differing distributions. For all traits, a block size of 1000 (--bsize 1000) was used, leave-one out cross validation (--loocv) was enabled, and sex was marked as a categorical covariate (--catCovarList sex). Prediction files from step 1 for each trait and imputed genotypes in BGEN format were used to perform association testing in step 2 of Regenie. For binary traits, a minimum allele count of 20 (--minMAC 20) and a minimum case count of 50 (--minCaseCount 50) were required, and the Firth approximation was performed (--firth --approx) using a p-value threshold of 0.01 (--pThresh 0.01). For quantitative traits, a minimum allele count of 20 (--minMAC 20) was required and Rank Inverse Normal Transformation (--apply-rint) was applied. For all traits, a block size of 500 (--bsize 500) was used and sex was marked as a categorical covariate (--catCovarList sex). Locus pruning was performed separately for each broad- and fine-scale ancestry using PLINK 2.0a to identify unique variant-phenotype associations. To standardize phenotypes for comparison, phecodes for binary traits were mapped to the Experimental Factor Ontology (EFO) using the text2term v4.5.0 Python package. Each phecode was assigned up to three top-matching EFO terms. To determine whether each unique variant–phenotype association was previously unreported, we queried the EMBL-EBI GWAS Catalog (June 27, 2025 freeze) for matching associations. An association from our study was classified as “previously unreported in the EMBL-EBI GWAS Catalog”, only if both of the following searches yielded no results. (1) Associations between the lead variant and any phenotype in the same phecode group as the associated phecode ( e.g. , for colorectal cancer, we searched for associations with all EFO terms in the broader “neoplasm” category). (2) Associations between the nearest protein-coding gene (within 10kb) and any phenotype in the same category as the associated phecode. Genomic DNA libraries were created by enzymatically shearing high molecular weight genomic DNA to a mean fragment size of 200 base pairs. Multiplexity of exome capture and sequencing was achieved by adding unique asymmetric 10-bp barcodes to the DNA fragments of single samples during library amplification. Equal molar amounts of DNA samples were pooled for exome capture using a slightly modified version probe library of xGen exome research panel from Integrated DNA Technology (IDT). After PCR amplification and quantification of the captured DNA, samples were multiplexed and loaded to Illumina sequencing machines for sequencing to generate 75-base-pair paired-end reads. The samples in this study were sequenced using the Illumina sequencing machines, including NovaSeq 6000 with S2 or S4 flow cells and the NovaSeqX with 25B flow cells. Sequencing reads in FASTQ format were generated from Illumina image data using the bcl2fastq program (v2.20, Illumina). Whole-exome reads alignment and germline small variant detection were conducted using the Original Quality Functionally Equivalent (OQFE) protocol described in Krasheninina et al., 2020 206 . Briefly, raw read files (FASTQ) were mapped to the GRCh38 reference obtained from https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/technical/reference/GRCh38_reference_genome/ using BWA-MEM v0.7.17-r1188 in an alt-aware manner 139 . Duplicate reads were then marked with Picard v2.21.2 127 . The final CRAM files were compressed with SAMtools v1.2 128 . Germline variant detection was performed on each CRAM using a Parabricks accelerated version of DeepVariant v0.10.0 with a custom WES model 140 , resulting in a sample-level gVCF (genomic VCF). Per-sample gVCFs were merged with GLnexus v1.4.3 141 into a joint-genotyped multi-sample project-level VCF (pVCF). Variant prediction was restricted to the exome capture region and the 100 base-pairs buffer on each side of the target regions. The WES was of high quality and reached an average coverage of 41.6-fold with a minimum of 24.2-fold, 90% having ≥ 34.7-fold in targeted regions, consistent with a previous publication 207 . The pVCF was converted to a PLINK file format using PLINK 1.9 133 for downstream analyses. These data underwent extensive quality control to ensure the absence of contamination, duplication, and other technical errors, as well as sufficient read depth to guarantee reliability and accuracy. Genetic duplicates were defined based on the aggregated genotype data of all sequenced samples. Sex was predicted based on the ratio of read coverage on chromosome Y over the whole exome read coverage. VerifyBamID v1.1.3 142 was used to estimate sample contamination. Samples were excluded if they showed sex discordance, duplication, cross-individual contamination > 5%, or coverage < 20X in more than 20% off target regions. Overall, 386 samples were excluded for failing QC. Specifically, 249 samples were flagged for gender discordance, 12 had less than 80% of the exome covered at 20X, 80 showed contamination levels above 5%, and 74 were unresolved duplicates. After accounting for 29 samples that appeared in more than one exclusion category, the final number of unique samples excluded due to QC failure was 386. Alignment quality metrics were generated using Picard CollectHsMetrics v2.27.4 127 and SAMtools stats v1.15.1 with default settings. MultiQC v1.27.1 143 was used to systematically aggregate sample-level quality metrics. RTGtools v 3.12.1 129 was used to assess variant QC metrics, including the number of variants for SNPs, small insertions and deletions; genotype counts; heterozygous-to-homozygous ratios for each variant type and transition/transversion (Ti/Tv) ratio. Genotypes were further confirmed using an alternative approach 208 with 500 random samples, yielding a high concordance for on-target sites (median concordance of 99.5% for SNPs and indels). For WES-based analyses, we used samples that also had array genotyping data (60,025 out of 61,797 patients with WES), which was necessary to consistently define broad- and fine-scale genetic ancestry. We annotated variants using Ensembl Variant Effect Predictor (VEP) v.112 137 for 58,387 samples assigned to the EUR, AFR, SAS, EAS, or AMR ancestry class ( i.e. , excluding unclassified genetic ancestry samples). Annotations were performed using the GRCh38 cache and the corresponding reference FASTA file. We limited our analysis to autosomal chromosomes to avoid technical artifacts in variant calling caused by the differences in ploidy between males and females, as well as the high-sequence similarity between the X and Y chromosomes in certain regions 209 . Counts for single-nucleotide variants, indels, multi-allelic, synonymous, missense, and LOF variants were restricted to whole-exome sequencing (WES)-targeted regions. Consistent with a previous UKBB WES study 210 , we classified LOF variants as those with the following consequences: stop_gained, start_lost, splice_donor, splice_acceptor, stop_lost, and frameshift. To increase reliability, we further restricted LOF variants to those flagged as high confidence by LOFTEE 94 . For multi-allelic variants, the predicted function for each alternate allele was determined using the --pick-allele option based on the default ordered set of criteria defined by VEP 137 . To determine the number of variants with MAF < 1% while accounting for ancestry-specific allele frequency differences, we retained variants with MAF < 1% in at least one ancestry group. For multi-allelic variants, we defined the minor allele as the second most common allele (including the reference) and classified the variant as rare if the cumulative allele frequency of all alternate alleles was < 1%. For rare multi-allelic variants, we incremented a sample’s count only if it carried an alternate allele with MAF < 1%. In multi-allelic cases where the minor allele was the reference or the alternate alleles had AF ≥ 1%, we incremented a sample’s count if it carried any alternate allele. For rare functional variants ( e.g. , missense), we incremented a sample’s count only if it carried an alternate allele with MAF < 1% corresponding to that functional category. We identified variants via a literature search, where variants are known to be enriched within certain populations. For this goal, we selected: 1) Familial Mediterranean Fever (FMF) related variants in the MEFV gene (V726A, M694V, M694I, M680I and E148Q) 211 2) the HBB:p.E7V variant, and 3) PCSK9 P/LP variants 195 . For FMF and PCSK9 , carriers were identified as participants who carry at least one corresponding pathogenic variant within PCSK9 or FMF variant group. For HBB:p.E7V, the frequency was estimated based on this variant alone. We only considered unrelated participants, using a kinship coefficient of 0.05 based on PLINK v2.0a 136 king-cutoff. Fisher’s Exact test was used to test the carrier frequency of each cluster against the carrier frequency of the other clusters grouped together. The risk of amyloidosis (a common symptom of FMF) by carrier status was calculated considering carriers of known FMF and patients assigned the amyloidosis phecode (270.33). To test differences in frequencies of pharmacogenomic variants, 121 CPIC Level 1A variants were pulled from ClinPGx 86 , 87 . An overview of the analyzed variants can be found in Table S4 . As variants spanned both common to rare frequencies, we used exome data for consistency. Overall, the variants were highly abundant in ATLAS, with 95.5% of all unrelated individuals carrying at least one. Of these variants, 69 were identified in the UCLA ATLAS exomes. To test enrichment of single alleles across clusters, “carriers” were identified as participants who carry at least one allele of a variant within a gene associated with a pharmacogenomic drug interaction or monogenic condition (unless only a single variant was involved). We then calculated carrier frequency as the number of identified carriers over the total number of participants with WES available and were unrelated using a kinship coefficient of 0.05 based on PLINK v2.0a 136 king-cutoff . To identify clusters with a significantly different carrier frequency of variants, we applied a Fisher’s Exact test to the carrier frequency of each cluster against the carrier frequency of the other clusters grouped together. We then used FDR corrections for p-values generated from all Fisher’s exact tests and selected significant clusters (FDR ≤ 0.05) where at least 5 carriers were identified. To identify clusters where participants were enriched for carriers of monogenic ClinGen variants associated with disease, we used pathogenic variants that were curated by experts within the field and underwent stringent review to be considered pathogenic from ClinGen 91 . ClinGen variants filtered to identify P/LP variants with autosomal dominant inheritance, autosomal recessive inheritance, or semidominant inheritance, resulting in 3,521 variants associated with 204 conditions overall according to ClinGen. We removed one HNF4A ClinGen variant due to a high allele frequency (MAF>0.01). Enrichment for variants, aggregated per gene, was tested as described above for ClinPGx variants with WES data. PAR was calculated for each cluster, showing significant enrichment of ClinGen P/LP variant frequencies aggregated by gene, focusing on genes associated with monogenic diseases with autosomal dominant or semidominant inheritance. To ensure sufficient power, we included only clusters with at least 30 individuals diagnosed with the disease with WES data available. We used the following formula: PAR = (p(RR-1))/(p(RR-1)+1), where p is the carrier frequency and RR is the ratio of the proportion of carriers with disease to the proportion of noncarriers with disease. For familial hypercholesterolemia, cases were defined as individuals with LDL levels exceeding 190 mg/dL, adjusted for statin use. In case of statin use, the LDL levels were divided by 0.7, as previously done 195 , 212 . To calculate the PAR for breast cancer associated with the three Ashkenazi Jewish founder alleles, we used the formula above. Because individuals in this group are often aware of their genetic status and may pursue preventive breast cancer measures, we excluded individuals who underwent surgery (using the ‘acquired absence of breast’ phecode) but were not diagnosed with breast cancer, to improve the reliability of the estimate. The list of all ClinGen 91 P/LP variants identified in any ACMG secondary findings (SF) v3.2 genes 93 was extracted as described above. Variants were labeled as missense or LOF, based on VEP v.112 137 annotations. Missense variants were defined for ‘missense_variant’ variants according to the ‘consequence’ VEP output column. LOF was defined for high-confidence “HC” LOF variants based on LOFTEE 94 . All variants were rare across broad-scale ancestries. A WES plink file with all participants was filtered to include only the listed variants. Then, this file was broken into ancestry groups ( i.e. , broad-scale population, or fine-scale cluster) using PLINK v2.0a 136 . Related participants were removed using a 0.05 kinship cutoff in PLINK v2.0a 136 . The frequency of each allele in every population group was calculated with the PLINK v2.0a --freq command, and total frequencies were summed up for rare missense and LOF variants separately. To test differences in the allele counts across populations, allele dosages were calculated with PLINK v2.0a 136 --recode A option. Dosages were summed to calculate the total alternative (ALT) allele count within each population. The number of reference (REF) alleles was defined as twice the number of individuals in the group minus the number of ALT alleles. Fisher’s Exact tests were applied to test differences between the number of REF and ALT alleles in each population compared to all other participants not assigned to that specific group, considering only groups with over 400 participants. Bonferroni correction was applied for multiple testing correction. Plotting was done using the BPG R package v.7.1.0 156 . Variants from exome sequencing were annotated using VEP v.112 137 with the dbNSFP v4.9a 138 and LOFTEE 94 plugins installed. The LOFTEE high-confidence “HC” flag was used to select for LOF variants predicted to have deleterious effects. Missense variants were assigned a 9-point deleteriousness score based on a consensus of nine missense deleteriousness prediction toolkits, similar to methods described in prior biobank-scale rare variant studies. We used the Critical Assessment of Genome Interpretation (CAGI) project 95 to prioritize well-performing tools not trained on the same features. We selected five meta-predictors – ClinPred 144 , MetaRNN 145 , BayesDel_addAF 146 , VARITY_R 147 , REVEL 148 – and four stand-alone predictors – AlphaMissense 149 , MutPred2 150 , VEST4 151 , ESM-1b 152 – for use in our analysis. We assigned each variant a binary score per tool based on dbNSFP rank scores: 1 if the variant’s score exceeded the threshold score for being more likely a deleterious ClinGen or ClinVar variant than a background variant ( Figure S7g – h ), and 0 otherwise. Summing these binary scores produced a deleteriousness score ranging from 0 to 9, with predicted damaging missense variants scoring ≥ 5 retained for downstream analysis. A WES PLINK file was filtered to include computationally predicted LOF and predicted damaging missense variants (see above) in ACMG SF v3.2 genes 93 , providing a more comprehensive evaluation than the one based solely on ClinGen P/LP variants, which included only 17 genes. A WES plink file with all participants was filtered to include only the ACMG putative damaging variants, and the file was split into ancestry groups (broad- and fine-scale ancestries) using PLINK v2.0a 136 . Related participants were removed using a 0.05 kinship cutoff in PLINK v2.0a 136 . Allele frequencies were calculated with PLINK v2.0a, and only rare variants (MAF < 1% in all broad-scale ancestries) were kept. Allele dosages per individual were extracted using PLINK v2.0a, and the total rare LOF and predicted damaging missense alleles were counted separately per participant. First, the differences between the total numbers of REF and ALT alleles across ancestries were evaluated with a Fisher’s Exact test as described above (ClinGen ACMG analysis). Second, a Mann-Whitney U test with a Bonferroni correction was applied to test the difference in the distribution of rare LOF and predicted damaging missense counts per individual between any broad- or fine-scale group compared to all others, with groups that included at least 100 participants. Figure S7I presents the Mann-Whitney U test results, with statistically significant differences indicated by an asterisk (*). Visualization was done with BPG 156 . ExWAS were conducted using the Regenie v4.0 framework 66 for all five broad-scale populations (AMR, AFR, EAS, EUR, SAS) and fifteen fine-scale clusters with at least 400 participants (IBD-01 through IBD-15). Related individuals were removed using a 0.05 kinship cutoff in PLINK v2.0a 136 . Sample sizes for each tested group can be found in Table S6 . Within each group, imputed genotype dosages in BGEN format (--bgen) and prediction scores from PheWAS step 1 (--pred) were used to conduct gene-based association testing for binary and quantitative traits across 17,537 genes. Samples were restricted to those in predefined inclusion lists (--keep), and trait-specific covariates were provided via --covarFile, including age, sex (modeled categorically with --catCovarList sex), BMI, and the top 10 genetic principal components. Whole-exome genotypes in bed format were filtered using PLINK v2.0a, retaining variants with a call rate ≥ 90% (--geno 0.1), a minor allele count ≥ 1 (--mac 1), and a Hardy-Weinberg equilibrium p-value > 1×10 −15 (--hwe 1e-15). In addition, 784,548 variants overlapping low-complexity regions (LCRs) were excluded prior to analysis. Briefly, SNP positions were extracted from the .bim file, intersected using bedtools v2.29.1 153 with annotated LCRs from the Genome in a Bottle Consortium 213 , and filtered from the genotype files, resulting in 12,326,160 exome variants retained for burden testing. Filtered whole-exome genotypes in BGEN format and prediction files from PheWAS step 1 were used to perform gene-based association testing. Regenie-style annotation, set, and mask files for predicted deleterious LOF and missense variants (see Variant annotation using a consensus of computational tools ) were created programmatically using Python v3.11.9 with the polars v1.2.1 package. For binary traits (--bt), a minimum case count of 50 (--minCaseCount 50), a minimum minor allele count of 5 (--minMAC 5), and a block size of 1000 (--bsize 1000) were enforced. Firth logistic regression with saddlepoint approximation (--firth --approx) was applied for variants with p < 0.01 (--pThresh 0.01). For quantitative traits (--qt), rank inverse normal transformation (--apply-rint), a minimum minor allele count of 5 (--minMAC 5), and a block size of 500 (--bsize 500) were enforced. Regenie’s implementation of the RGC gene-based p-value test (–rgc-gene-p) was enabled for all analyses. Additionally, variants were binned by minor allele frequency using 1% bins (--aaf-bins 0.01), and SNP-level membership for each burden mask was recorded (--write-mask-snplist). In addition to gene-based analyses, single variant association testing was performed for selected predicted deleterious LOF and missense variants using filtered whole-exome genotype dosages in BGEN format and prediction scores from PheWAS step 1. For both binary and quantitative traits, tests incorporated the same set of covariates (age, sex modeled categorically, BMI, and the top 10 genetic principal components) and enforced a minimum minor allele count of 5 (--minMAC 5). For binary traits, Firth logistic regression with saddlepoint approximation was applied for variants with p < 0.01, along with a block size of 1000, while quantitative traits were evaluated using rank inverse normal transformed phenotypes with a block size of 500. A Bonferroni-adjusted p < 0.05 was defined as the cutoff for significance. Replication analyses were conducted utilizing the Python Hail package v0.2.134 165 on AoU Workbench using the Controlled Tier Dataset v8. Phenotypes were obtained by querying OMOP databases for per-patient ICD codes, which were subsequently converted to phecodes v1.2 utilizing the R PheWAS package v0.99.6 176 . Participants with both phecodes and single-read WGS data were filtered to remove participants flagged from genotype QC or from relatedness QC, resulting in 291,082 total participants. Logistic regression was performed within the given replication genetic ancestry cohort by testing case/control status for a given phenotype and SNP, utilizing age, sex, age^2, sex*age, sex*age^2, and genotype PCs 1-10 as covariates. Replications for variant-trait associations were obtained by querying the Controlled Tier Dataset v8 “All by All” allele count/allele frequency (ACAF) variant result tables using the Python Hail package v0.2.134 on AoU Workbench. Replications for variant-trait associations were obtained by querying the summary statistics portal at https://taiwanview.twbiobank.org.tw/pheweb.php 7 . Replications for gene-trait associations were obtained by querying the Controlled Tier Dataset v8 “All by All” rare variant result tables using the Python Hail package v0.2.134 165 on AoU Workbench. Replications for gene-trait associations were obtained by querying the AstraZeneca summary statistics portal at http://azphewas.com/ 69 . The BioMe Biobank consists of electronic health records and genetic data from approximately 60,000 participants from the Mount Sinai Health System in New York. Participant recruitment was between 2007 and 2023. This study was approved by the Icahn School of Medicine at Mount Sinai’s Institutional Review Board (Institutional Review Board 07–0529). All study participants provided written informed consent. BioMe participants were genotyped using the Illumina Infinium Global Diversity Array (GDA; number of participants, N=23,430; number of variants, n=1,833,111) or Infinium Global Screening Array (GSA; N=32,595; n=635,623). Quality control consisted of removing participants with a call rate <95%, a mismatch between self-reported and genetic sex, and high heterozygosity. Duplicated sites and sites with a genotyping rate of <95% were removed. QC was done with PLINK2. Data was imputed with the TOPMED imputation server 214 with genome build hg38. Genetic PCs were calculated across participants using PLINK2 for the genotyping sites after LD pruning. Exome sequencing data were generated by the Regeneron Genetic Center 215 . Quality control consisted of using the Goldilocks Filter 210 , and variants with quality scores <3 or depth of coverage scores <7 for SNPs, or <5 or depth of coverage <10 for indels were removed. Monomorphic sites were removed. Genetic ancestry was assigned using a random forest classifier trained on principal components derived from the 1000 Genomes data to align with the genetic ancestry assignment in AoU 5 . Participants were assigned a genetic ancestry using 10 PCs. Bio Me laboratory data were processed according to the QualityLab pipeline 216 . Briefly, quantitative lab values obtained in different clinical contexts (ambulatory, emergency, inpatient, and urgent care), were cleaned. Only labs with at least 100 patients and at least 1000 numeric observations were considered. Only adult (> 18 years) lab values were considered. Per lab, outliers were removed, defined as values as greater or less than 4 standard deviations from the mean of that lab. Non-numeric labs were removed. Labs must have had >70% of the reported units matching. After cleaning, a median value and median age was calculated per person for each lab test in each clinical context and across all contexts. Laboratory phenotypes were manually assigned a match to UKBB phenotypes using available metadata. Electronic health phenotype data in the form of ICD-10 codes were collapsed into PhecodeX phecodes 217 . Phecodes were transformed into a binary matrix, where each row was an individual and each column was a phecode. A participant had a 1 if they ever were diagnosed with that phecode, otherwise their value was set to 0. Age was calculated as current age, defined from 01/01/2025. Association testing was performed using Regenie v4.0 66 . Burden testing was performed using SKATO and the following masks: missense, missense_pLoF, pLoF, and synonymous. PLoF variants were defined using LOFTEE 94 . ATLAS GLP1-RAs prescriptions, including medication names, start and end dates, discrete dose, usage instruction, strength and route (oral or subcutaneous), were retrieved and grouped based on the following simple generic names: dulaglutide, semaglutide, liraglutide, exenatide, albiglutide and lixisenatide. For all statistical analyses, only semaglutide users were considered. Usage start dates were defined based on the earliest prescription start date for each patient. In the case of a missing start date, the prescription ordering date was used instead (in most cases, these two fields were identical). Overlapping medication periods were handled such that when a new prescription started, the previous one was considered to have ended. Similarly, missing end dates were determined using the start date of the next prescription when available. In the case of completely overlapping prescriptions with different instructions, the combination of both routes and the weighted average dose was considered. If the discrete dose information was missing, the medication dose was extracted from the instructions’ free text field. If the instructions also omitted the dose information, it was imputed for a given medication type, based on the weighted dose average from the entire cohort on the corresponding type. Initial weight and BMI were defined as the median of all available weight or BMI measurements recorded from in-person visits, within 6 months prior to the first prescription start date. Utilizing the median rather than a single data point helped minimize the likelihood of recording errors in the EHR. The percentage of weight change was obtained for every weight measurement recorded on in-person visits within the period of active prescriptions between 4-60 weeks in total on semaglutide. The medication dose at each time point was defined as the weighted sum of all medication doses (doses multiplied by the number of prescription weeks) by the measurement date. In case both oral and subcutaneous medications were used within a period, a combined route category was defined. To identify the overall weight loss patterns across time, the FPCA R function from the fdapace package v.0.6.0 159 , 160 , which is suited to plot smoothed longitudinal data with repeated measurement, was used. This identified a consistent weight loss pattern up to 60 weeks, and sparse data points beyond ~150 weeks ( Figure S8C ). Thus, we restricted all analyses to this period. For all association tests, only participants aged 18 years or older were included. For all analysis parts that involved longitudinal data with repeated measurements, a linear mixed model with the bobyqa optimizer and an increased function evaluation limit (maxfun = 10000) was used (lmer R function; the lmerTest R package v.3.1.3 161 ). In each case, ANOVA was used to identify differences between models to define the best way to model a potential nonlinear relationship between weeks and weight loss, when treating weeks as a fixed effect, based on the Akaike Information Criterion (AIC). In some cases, the best model was achieved using a restricted cubic spline (RCS) with the rcs R function from the rms package (v.7.0.0), applied to weeks. In other cases, a polynomial function of weeks provided a better fit. To plot smoothed longitudinal data with 95% CI, the fitted model values (excluding covariates) and 95% lower and upper confidence bounds, were extracted using the visreg package 162 v.2.7.0 and visualized using the BPG package 156 . For testing the effect of baseline factors, fixed effects were defined for the medication dose, route, sex, age, initial BMI, first ten genetic PCs (excluding PC6 due to a strong collinearity with PC5 that disrupted the model convergence) and weeks on semaglutide. The model included both random intercepts and slopes for weeks on semaglutide. Bonferroni correction was applied to control for multiple testing. For testing differences across ancestries, fixed effects were defined for the medication dose, route, sex, age, initial weight, a polynomial function of weeks on semaglutide, and an interaction between a polynomial function of weeks and genetic ancestry categorical class (EUR, AFR, EAS, SAS and AMR). The overall number of weight measurements was 24,145, with a mean of 5 repeated measurements per patient. Ancestry sample sizes were: European (EUR), 3189; African (AFR), 373; admixed American (AMR), 913; East Asian (EAS), 291; South Asian (SAS), 107. The model included both random intercepts and slopes for the polynomial function of weeks on semaglutide. An ANOVA was performed on the model with a Bonferroni correction to assess the global effects of ancestry and ancestry×weeks interaction. When a significant result was found, post hoc comparisons between groups were conducted using the summary lmerTest function (v.3.1.3 161 ), applying a Bonferroni correction to all 12 class or class × time interactions. To test the relationship between PGS and weight loss, related participants based on their genetic similarity were excluded (defined using PLINK v2.0a 136 with the --king-cutoff 0.05). Scaled BMI (PGS000027) and DM2 PGS (PGS000729) in EUR participants were divided into three equal bins each: low, intermediate and high scores. Then, using longitudinal data, for each trait, fixed effects were defined for the categorical PGS bins, medication dose, route, sex, age, initial weight, first ten genetic PCs (excluding PC6 due to a strong collinearity with PC5 that disrupted the model convergence) and a polynomial function of weeks. Random intercepts were defined to account for repeated measures within individuals. Bonferroni correction was applied to control for multiple testing. To test the relationship between PGS and weight loss, relying on a simplified model where the maximum weight loss was considered for each EUR participant, linear regression was applied. The model was adjusted for the medication dose and route, 10 genetic PCs, age, sex, and initial weight. Visualization was made with BPG 156 , with smoothed data and 95% CI using loess.as R function (fANCOVA package v.0.6.1 163 ). The analysis was performed to test a relationship between the maximum weight loss on semaglutide and common genetic variants, for each broad-scale ancestry using SAIGE 167 . For step one, array observed variants were used following PLINK v2.0a 127 filtering, with the flags: --maf 0.01 --mind 0.1 --geno 0.1 --hwe 1e-6 . For the second step, array observed and imputed variants were used following plink filtering with --maf 0.01 --geno 0.05 --hwe 1e-6. A quantitative analysis was conducted with the traitType flag. The medication dose and route, the first five genetic PCs, age, sex, and initial weight were used as covariates. METAL 168 was used for meta-analysis. To identify genes genetically associated with weight loss, we used Regenie 66 with an additive model for gene-level tests. Only EUR semaglutide users that are not related (relatedness coefficient --king-cutoff 0.05) with WES data were used (n = 2,012), and for each, the maximum weight loss record with the corresponding number of weeks was considered, using information on aggregate dose and route. The list of candidate genes was limited to Bonferroni-significant proteins whose plasma abundance was altered by semaglutide treatment 115 . We considered variants within these genes with a predicted moderate or high impact on the protein function, according to VEP v.1.2 137 . For step one, the variant list was limited to observed array SNPs following a PLINK v2.0a 136 filtering with: --maf 0.01 --mac 100 --indep-pairwise 1000 100 0.9 --chr 1-22 --snps-only --geno 0.1 --hwe 1e-15. In step 2, we applied PLINK v2.0a filtering to variants across all EUR participants using the parameters --geno 0.05 and --hwe 1e-6. Subsequently, the file was filtered to include only semaglutide users and was used in Step 2. The medication dose and route, the first 10 genetic PCs, age, sex, and initial weight were used as covariates. Bonferroni correction was applied to control for multiple testing. Frequency of variants involved in PTPRU association with weight loss on semaglutide by ancestry can be found in the table below. For replication analyses in AoU the same processes and tools were conducted in EUR individuals. Controlled Tier Dataset v8 was accessed through the Researcher Workbench and imported in PLINK format, including exome variants in chromosome 1 (to extract PTPRU variants) and common ACAF variant datasets (for PGS calculation). Genetic ancestry and PCs derived from pre-defined assignments provided by AoU 5 . PTPRU VEP 137 variant annotations were obtained using Python Hail v0.2.134 165 to identify functional consequences. For the Regenie analysis, related individuals were not removed to maximise power, as Regenie is designed to account for relatedness. The combined P-value of ATLAS and AoU was calculated using Fisher’s combined probability test, implemented with the fisher function in the poolr R package v1.2.0 166 . Multiple testing was addressed using a Bonferroni correction that accounted for the number of gene-level tests in the discovery cohort, along with one replication P-value and one combined P-value. Frequency of variants involved in PTPRU association with weight loss on semaglutide by ancestry. Variant AFR AMR EAS EUR SAS 1:29236667:T:C 0 0 0 1.2×10 −5 0 1:29258717:G:A 0.00018 0 0 0.00013 0 1:29259276:C:T 0 0 0 1.2×10 −5 0 1:29259313:A:G 0 0 0 4.9×10 −5 0 1:29259883:C:G 0 0 0 3.7×10 −5 0 1:29259912:A:C 0 0 0 1.2×10 −5 0 1:29260003:G:C 0 0 0 1.2×10 −5 0 1:29260895:A:C 0.00055 0.0025 0 0.0027 0 1:29275496:G:A 0 0 0 7.3×10 −5 0 1:29275715:G:T 0.0018 0.011 0.0018 0.0092 0.023 1:29279056:G:A 0 0 0 1.2×10 −5 0 1:29279527:G:T 0 6.0×10 −5 0 1.2×10 −5 0 1:29279541:C:A 0 0 0 1.2×10 −5 0 1:29282705:G:A 0.0013 0.0028 0 0.0050 0.0018 1:29282710:C:T 0 0 0 8.5×10 −5 0 1:29284754:C:T 0 0 0 9.7×10 −5 0 1:29291897:G:A 0 0 0 0.00023 0.00045 1:29291929:G:A 0 0.00030 0 2.4×10 −5 0 1:29291942:C:T 0 0.00042 0.00020 0.00052 0 1:29291958:A:G 0.0058 0.00030 0 2.4×10 −5 0 1:29291978:G:A 0 0 0 0.00015 0 1:29303853:A:G 0 0.00036 0 0.00024 0 1:29303863:C:T 0 0.00012 0 0.00032 0 1:29303875:G:T 0.094 0.0065 9.9 × 10–5 0.00078 0.00045 1:29304795:G:A 0 0 0 1.2×10 −5 0 1:29305397:A:G 0.23 0.38 0.48 0.25 0.38 1:29311488:C:T 0.00018 0 0 4.9×10 −5 0.00045 1:29311513:C:A 0 0 0 3.7×10 −5 0 1:29311706:G:C 0 0 0 1.2×10 −5 0 1:29312579:C:T 0 6.0×10 −5 0 9.7×10 −5 0 1:29312615:G:A 0 0.0011 0 0.00015 0 1:29315374:C:T 0 5.99×10 −5 0 2.4×10 −5 0.00045 1:29315448:G:A 0 0 0 1.2×10 −5 0 1:29317767:C:T 0 0.00036 0 0.0014 0 1:29317838:G:A 0 0 0 0.00016 0 1:29320709:A:G 0 0 0 6.1×10 −5 0 1:29323643:C:T 0 0 0.00030 1.2×10 −5 0 Frequency of variants involved in PTPRU association with weight loss on semaglutide by ancestry. Statistical analyses were performed using R, PLINK, Regenie, and related tools. Logistic regression, linear regression, Firth-penalized regression, linear mixed models, and Fisher’s exact tests were applied in the relevant analyses, adjusting for appropriate covariates. Multiple-testing correction was applied using the FDR or Bonferroni adjustment, as specified in the relevant analyses. Detailed statistical models and software parameters are provided in the METHOD DETAILS section. An interactive web portal enables users to explore the PheWAS results, examine fine-scale ancestry associations with clinical phenotypes, and download supplemental data at atlas-phewas.mednet.ucla.edu .

Results

UCLA Health serves Los Angeles County, one of the world’s most ancestrally diverse metropolitan areas, with a population of 9.6 million. To date, the UCLA ATLAS initiative has consented ~250,000 UCLA Health patients and has collected biomaterials from ~130,000. ATLAS enrollment largely reflects the composition of UCLA Health patients, who are concentrated across Los Angeles ( Figure 1b ). The EHR, initiated in 2013, enables continuous longitudinal stratification of participants by disease states, with a mean and median of 8.6 and 7.9 years of participation per individual. As of November 2024, ATLAS genomic data included WES data from 61,797 participants and custom array genotyping using the Illumina Global Screening Array from 92,164 participants. Extensive quality control (QC) indicated high data quality ( Figure S1 ). We leveraged these data to interrogate social and genetic factors that affect disease risk and health outcomes ( Figure 1a ). Genotyped cohort demographics are summarized in Table 1 and the Methods . Biobank participants were older in age and had a higher comorbidity index than non-biobank patients (2.9 versus 1.7 mean Elixhauser index; 1-year post-collection), consistent with a higher number of clinical visits, which favors enrollment (controlled for data-completeness, Methods ; Figure 1c ). EHR-based phenotypes, including vital signs, disease diagnoses, and lab tests, were defined through harmonization procedures and subsequently validated ( Methods ; Figure S2a – l ). Common disease domains were led by endocrine/metabolic, cardiovascular, gastrointestinal disorders and neoplasms, reflecting known global health challenges ( Figure 1d ; Figure S2m – p ; Methods ) 31 – 33 . The ATLAS EHR data includes a total of 71,739,582 lab tests (considering complete blood count [CBC], lipid, metabolic, hemoglobin A1c [HbA1c], and 25-hydroxyvitamin D panels), with a mean of 540 lab test results per participant ( Table S1 ). Detailed prescription data show a total of 5,952,958 prescriptions; the 50 most frequently prescribed medications are listed in Table S1 . Although we primarily consider genetically determined ancestry in our analyses, rather than the social constructs of race and ethnicity 34 , 35 , we note substantial diversity based on participant self-reports. For instance, 62.7% of participants self-identify as White, 12% as Asian, 4.6% as Black or African American, 2.6% as Middle Eastern or North African, 0.9% as American Indian or Alaska Native and 0.3% as Pacific Islander. A substantial proportion (14.8%) report Hispanic or Latino ethnicity ( Table 1 ). Self-identified race/ancestry, a cultural-societal construct, and genetic ancestry are conceptually distinct 29 , 34 , 36 and genetic ancestry must be considered to prevent confounding in genetic association studies 35 , 37 . We classified biobank participants into six broad-scale ancestry populations: European (EUR), African (AFR), South Asian (SAS), East Asian (EAS), Admixed American (AMR) and an Unclassifiable (UNC) population, aligning with the 1000 Genomes Project super-populations 38 ( Figure 1e – g ; Methods ). ATLAS is diverse relative to most other large biobanks ( Methods ). Consistent with self-reported race/ethnicity, almost a third of the biobank participants (32%) were assigned to non-EUR genetic ancestries. There was significant agreement between self-reported race and broad-scale ancestry – 99% of self-identified White participants were assigned to EUR or AMR populations ( Figure S3a ). Of those with “unknown” self-reported race, 46% were assigned to AMR ancestry and 45% to EUR. We next asked how health system usage or disease burden varied by broad-scale ancestry, observing that the mean number of encounters varied significantly across ancestries (P-value = 7.5×10 −138 , ANCOVA, adjusted for genetic sex, age and Barriers to Accessing Services [BAS] 39 ), with the highest numbers of total encounters in participants from AMR ancestry (adjusted mean = 16.2), followed by AFR (15.8) and SAS (14.2), with the lowest values in EUR (12.2) and EAS (12.9) ( Figure 1h ). We also assessed health burden using the comorbidity index score 40 , 41 , which was higher in EAS (adjusted mean = 3.7) and AMR (3.6) compared to others (all other populations 2.6-2.9; Figure 1i ; P-value = 2.6×10 −5 , ANCOVA, adjusted for genetic sex, age, BAS and Area Deprivation Index [ADI] 42 ). Similar results for disease distribution across ancestries were obtained after applying inverse probability weighting (IPW) to adjust for participation bias relative to the broader UCLA Health population ( Figure S3b ; Methods ). Increased encounter numbers were only partially explained by elevated comorbidity index scores, indicated by modest correlations between the two (mean total encounters vs . Elixhauser Comorbidity Index: r = 0.32, P-value <2.2×10 −16 ; mean hospital encounters vs . Elixhauser Comorbidity Index: r = 0.26, P-value <2.2×10 −16 ; Figure S3c – d ). The population diversity of ATLAS enables the interrogation of the combined genetic and social/environmental effects on the risk of disease diagnoses. To illustrate this, we assessed variation in medical conditions across populations, replicating multiple known associations ( Methods ; Table S2 ; Figure S3e ). Further, we identified previously unreported associations ( Figure 1j ), including a lower risk of epilepsy in EAS compared to EUR participants (odds ratio [OR] = 0.46 with 95% confidence interval [0.32, 0.65], P Bonferroni = 2.3×10 −4 ), not previously detected 43 , 44 . Similarly, we found significantly reduced risk for bipolar disorder in those with AMR ancestry, clarifying conflicting findings in previous smaller studies 45 , 46 (OR = 0.47 [0.40, 0.56], P Bonferroni = 1.5×10 −17 ). We also show a significantly lower risk of sleep apnea in EAS participants relative to EUR (OR = 0.67 [0.59, 0.75], PBonferroni = 4.1×10 −10 ), which has been controversial 47 – 50 , with a weaker, but significant effect after correcting for body mass index (BMI) (OR = 0.84 [0.74, 0.95], P-value = 4.6×10 −3 ; Table S2 ). These associations remained significant after adjusting for socioeconomic status (SES), using both the BAS and ADI measures, and with IPW ( Table S2 ). Fine-scale ancestries 37 , 51 – 53 , many of which are understudied 14 – 16 , reflect recent geographic or demographic stratification and can reveal important contributions to health disparities and disease risk 14 . We quantified fine-scale ancestry using identity-by-descent (IBD) 53 , 54 in a population nearly three times larger than any previously published study 52 , 53 . This approach identifies genomic regions shared between individuals due to a common ancestor and defines fine-scale ancestries, which we refer to as “clusters” ( Methods ) 52 . We identified 36 fine-scale ancestry clusters with at least 30 participants ( Figure 2a – b ; Figure S4a – c ; Methods ), labeled with both a numeric identifier ( e.g. , IBD-01) sorted based on the cluster sample size, and a corresponding cluster name to facilitate interpretation ( Methods ). Because PCA captures ancient population structure and IBD segments reflect recent shared ancestry, full agreement between these approaches is not expected, as is observed for some of our IBD clusters, such as IBD-02 ( Figure 2a ). We replicated previously identified clusters and added unreported ones, such as a Native Hawaiian cluster (n IBD-29 = 64) and a Bantu cluster (n IBD-35 = 32). The largest clusters consisted of Northern Europeans (n IBD-01 = 33,675) and Southern Europeans (n IBD-02 = 14,841). The remaining clusters represent the heterogeneity of ancestral origins in Los Angeles, including Ashkenazi Jewish (n IBD-03 = 14,262), Filipino (n IBD-09 = 1,438), Iranian Jewish (n IBD-11 = 707), and Armenian (n IBD-16,IBD-23,IBD-31 = 560) populations. The diversity of ATLAS is reflected not only by a relatively large portion of non-EUR individuals, but also by substantial diversity within the EUR broad-scale ancestry population. Several EUR clusters are large relative to published datasets, including Southern European, Ashkenazi Jewish, Armenian and Iranian Jewish 3 , 53 , 55 – 57 ( Table S3 ). Similarly, for EAS, we identified multiple clusters, including Filipino, representing the largest published Filipino genotyped cohort of which we are aware 58 . Despite this strength, for some populations, sample sizes are limited, which leads to relatively small clusters. For a nuanced understanding of diagnostic variation, we tested the prevalence of 1,253 phecodes 59 – 61 across 24 fine-scale clusters with at least 100 participants. Table S3 shows all associations and Figure 2c highlights selected associations (also available at https://atlas-phewas.mednet.ucla.edu/ancestry ). We replicated known findings ( Methods ) and uncovered numerous associations between fine-scale clusters and diseases. For instance, Filipinos, an understudied population in genetics research, exhibited the lowest risk for vitamin B-complex deficiency (OR IBD-09 = 0.50 [0.36, 0.67], FDR = 2.1×10 −4 ), and the highest risk for cholesterolosis of the gallbladder (OR IBD-09 = 4.3 [2.9, 6.2], FDR = 5.6×10 −12 ), both previously unreported. In Iranian Jewish subjects, another group with low research inclusion, we detected a previously unreported elevated risk for glaucoma (OR IBD-11 = 1.9 [1.5-2.4], FDR = 2.9×10 −6 ), a condition known to be more common in individuals of African and East Asian ancestry. We also observed several shared health risk patterns among Jewish ancestry clusters. Both Iranian and Ashkenazi Jewish clusters showed the highest risk for bladder cancer (Iranian Jewish: CI IBD-11 = 1.4-5.6, FDR = 0.008; Ashkenazi Jewish: OR IBD-03 = 1.5 [1.1, 1.9], FDR = 0.02) and a high hyperplasia of prostate risk (Iranian Jewish: OR IBD-11 = 2.3 [1.8, 2.8], FDR = 1.1×10 −13 ; Ashkenazi Jewish: OR IBD-03 = 1.7 [1.6, 1.8], FDR = 4.6×10 −66 ), and the lowest risk for cirrhosis of liver (Iranian Jewish: CI IBD-11 = 0.03-0.4, FDR = 3.1×10 −2 ; Ashkenazi Jewish: OR IBD-03 = 0.4 [0.3, 0.5] , FDR = 6.8×10 −24 ). Mexicans and South Americans suffered consistently more from hormones’ adverse effects in therapeutic use, a previously unreported association (Mexican American clusters: OR IBD-04,-07,and-12 = 2.0-2.8 [1.3, 3.4], FDR = 1.6×10 −2 - 2.1×10 −24 ; combined South Americans: OR IBD-08 = 1.9 [1.4, 2.5], FDR = 1.3×10 −4 ). These associations persisted after SES adjustment, using the BAS and ADI ranks. In most cases, associations held after applying IPW and also adjusting for SES, except for hormone adverse effects, where only the largest Mexican American group remained highly significant, likely due to smaller sample sizes resulting from the IPW analyses ( Table S3 ). These findings demonstrate the utility of fine-scale ancestry in identifying disease risk profiles in underrepresented populations. We next focused on cardio-metabolic diseases due to their global impact on public health. We selected a few targeted cardio-metabolic phenotypes and compared disease risk across fine-scale clusters within the same broad-scale ancestry ( Methods ; Figure S4d ). Strikingly, among Asian clusters, the Filipino cluster had an elevated risk for all tested cardio-metabolic conditions, which remained after adjusting for BMI. Across fine-scale EUR ancestries, the larger Armenian cluster showed high cardiometabolic disease risk. Iranian clusters, both Jewish and non-Jewish, had a relatively higher risk of coronary atherosclerosis, hyperlipidemia and type 2 diabetes, but not hypertension. Adjusting for SES factors and adding IPW maintained these significant patterns ( Table S3 ), except abdominal aortic aneurysm in Filipinos, which lost significance, and hyperlipidemia in Iranian Jewish participants, which did not remain significant only when both SES adjustment and IPW were applied- likely due to reduced sample size, although the OR remained elevated. To examine the robustness of ATLAS in predicting genetic risk and to assess the quality of our dataset, we first tested the utility of PGS in stratifying risk for common disorders. We observed high PGS performance in EUR individuals–on average, 18.6% of patients diagnosed with major disorders were in the top PGS decile, rising to 41% for type 1 diabetes ( Figure 3 ; STAR Methods ). Consistent with prior reports 62 – 65 , this predictive power was diminished in non-European ancestries ( Figure 3 ; Methods ). Next, we performed phenome-wide association studies (PheWAS) using Regenie 66 to map the common variant architecture of clinical diagnoses and laboratory measurements, conducting separate analyses within each of our broad- and fine-scale ancestral cohorts ( Methods ; n = 84,110 total participants). After linkage disequilibrium (LD) pruning, we identified 19,431 unique variant-phenotype associations passing a genome-wide significance threshold of 5×10 −8 ( Table S4 ), with 5,772 passing the most conservative Bonferroni threshold of 6.4×10 −11 (5×10 −8 conditioned on 776 tested phenotypes). The majority (83.8%) of associations were detected in only one ancestry at genome-wide significance, though utilizing a more permissive threshold of 1×10 −4 showed cross-ancestry support for a substantial portion of associations (54.7%; Figure S6a – b ). Supporting the validity of the data used for PheWAS, we confirmed ancestry-specific differences in APOE allele frequencies and attenuated Alzheimer’s risk associated with the ε4 allele in African Americans and replicated known associations such as PNPLA3 rs738409-G with non-alcoholic fatty liver in AMR groups 67 ( Methods ; Figure S6c – e ). We identified numerous ‘previously unreported findings’, which we defined as associations that have not been reported at genome-wide significance in the existing literature, the European Molecular Biology Laboratory-European Bioinformatics Institute (EMBL-EBI) GWAS Catalog 68 , or summary statistics from UKBB 69 , Taiwan BioBank 7 , and AoU 5 . Notably, 39.5% of associations were unreported in the European Molecular Biology Laboratory-European Bioinformatics Institute EMBL-EBI GWAS Catalog 68 ( Methods ). We performed replication analyses for the top previously unreported findings in independent biobanks, including AoU 5 , UKBB 69 , Taiwan BioBank 7 , and BioMe 8 ( Table S4 ; Methods ). Summary statistics are available at atlas-phewas.mednet.ucla.edu . We identified a previously unreported association between FN3K rs7208565-T and increased risk for intestinal disaccharidase deficiency across EUR populations ( Figure 4A ; P EUR = 3.8 × 10 −41 ; OR EUR = 1.3; MAF EUR = 33.3%; MAF IBD-01 = 32.3%; MAF IBD-02 = 34.3%; MAF IBD-03 = 35.6%) and the AMR population (P AMR = 4.9 × 10 −11 ; OR AMR = 1.3; MAF AMR = 42.3%; Table S4 ). This variant was additionally associated with increased risk of the “other abnormal glucose” phecode (P EUR = 1.8 × 10 −42 ; OR EUR = 1.2), whereas a variant in close LD (rs113373052-T) was associated with increased HbA1c (P EUR = 5.4 × 10 −97 ; β EUR = 0.13), reflecting additional consequences on glucose homeostasis. Intestinal disaccharidase deficiency is the inability to completely digest sugars such as lactose and sucrose, leading to irritable bowel syndrome (IBS)-like symptoms 70 . FN3K phosphorylates glycated proteins, preventing the formation of advanced glycation end-products 71 . Supportive of these findings, rs7208565 has been previously associated with type 2 diabetes 72 . Furthermore, we discovered several previously unreported associations driven by low-frequency variants (MAF ~1-2%) in non-EUR cohorts. For instance, we linked rs115750084-G in DPP6 , a gene implicated in synaptic signaling 73 and neurological disorders 74 , 75 , to an increased risk of major depressive disorder in the AFR cohort (P AFR = 4.6 × 10 −8 ; OR AFR = 4.2; MAF AFR = 1.1%; Figure S6F ). In the EAS cohort, we associated rs77742325-G in MAS1 , a key component of the bone-protective ACE2/Angiotensin-(1-7)/Mas axis, 76 with osteoporosis (P EAS = 1.5 × 10 −8 ; OR EAS = 3.5; MAF EAS = 1.8%; Figure S6G ). We also linked an intronic deletion (rs202215133-ATATCATAG>A) in UBR5, an E3 ubiquitin ligase implicated in neuroinflammatory and neurodevelopmental syndromes, 77 , 78 with migraine headache in the Mexican American cluster (P IBD-04 = 1.7 × 10 −8 ; OR IBD-04 = 5.6; MAF IBD-04 = 1.1%; Figure S6H ). These associations were independently replicated, albeit with more modest effect sizes in the replication cohort ( Table S4 ). Given that the variants are low MAF, further validation through direct sequencing in large cohorts will be essential to confirm their generalizability. Similarly, our large AMR discovery cohort revealed several unreported associations. For instance, we linked STARD7 rs17419569-C with increased risk for asthma in the Mexican American cluster ( Figure 4B ; P IBD-04 = 1.9 × 10 −9 ; OR IBD-04 = 2.6; MAF IBD-04 = 2.5%). Prior work suggests that decreased STARD7 expression is associated with enhanced allergic responses in the human lung and significant increases in airway hyperresponsiveness in haploinsufficient Stard7 mice, 79 whereas rare variants within STARD7 were linked with asthma (P UKBB = 5.1 × 10 −3 ; Table S4 ). In the same cluster, GPX7/SHISAL2A rs74744741-C was associated with gastrointestinal reflux disease (GERD) ( Figure 4C ; P IBD-04 = 5.8 × 10 −10 ; OR IBD-04 = 1.6; MAF IBD-04 = 8.9%). GPX7 has been associated with carcinogenesis in the context of GERD-associated Barrett’s esophagus, 80 whereas SHISAL2A has an unknown function, but is highly expressed in the small intestine and in lymphoid tissues. 81 Additionally, we identified two low-frequency AMR variants associated with chronic renal failure (CRF; Figure 4D ), rs112680741-C (P AMR = 2.9 × 10 −11 ; OR AMR = 3.8; MAF AMR = 1.3%) and rs2744548-C (P AMR = 9.6 × 10 −11 ; OR AMR = 3.2; MAF AMR = 1.7%), nominating GPLD1 , ALDH5A1 , and KIAA0319 as potential CRF risk genes. These associations did not replicate in available biobank populations ( Table S4 ), possibly because the ATLAS AMR populations from the Los Angeles area represent distinct Mexican and South American ancestries 82 , 83 , whereas those from other biobanks (BioMe, for example) are comprised of individuals with distinct ancestral cluster assignments to Puerto Rico and the Dominican Republic. Other cohort-specific characteristics may also contribute to these differences, so independent replication in additional cohorts is needed. This first wave of the ATLAS WES catalog comprises more than 11.5 million autosomal variants, with a median of 9,821 missense and 159 loss-of-function (LOF) variants per individual, closely matching expectations from other studies 84 . Most (98%) WES variants were rare (MAF < 1% in any broad-scale ancestry), including 2,873,731 rare missense and 172,462 rare LOF variants, with a median of 15 rare LOF and 403 rare missense variants per participant ( Table S5 ). AFR participants showed the greatest number of rare LOF and missense variants, matching prior results 85 . To evaluate our WES, we analyzed a predefined set of clinically relevant rare variants known to have elevated frequencies in specific populations and demonstrated that ATLAS captures these established patterns ( Methods ; Figure S7a – e ). Then, we examined differences in 69 ClinPGx pharmacogenomic variants 86 , 87 across fine-scale ancestries ( Figure 4e ; Table S4 ). We identified several unreported enrichments in specific clusters, including a SLCO1B1 variant, rs4149056, which affects the metabolism and toxicity of statins 88 in Ashkenazi and Iranian Jews; a CYP4F2 variant, rs2108622, that can increase the required dosage of warfarin 89 , enriched in Ashkenazi Jewish, Lebanese, Armenian 1, Egyptian Christian, and South Asian 2 individuals; and PYD variants, which affect the toxicity of chemotherapy drugs 90 , abundant in Mexican American, African American and South American clusters. This suggests the importance of considering fine-scale ancestry in precision drug prescribing. Next, we used ATLAS’ ancestral heterogeneity to identify enrichments of known rare clinically relevant variants, aggregated by gene across fine-scale clusters to illustrate the utility of fine-scale ancestry in identifying populations at higher risk of rare, monogenic disease. We queried the full set of the curated rare pathogenic/likely-pathogenic (P/LP) ClinGen 91 variants that have strong clinical and genetic evidence (n = 643), identifying 5,223 unrelated carriers ( Figure 5a ). Alongside known associations ( STAR Methods ), we detected a previously unreported elevated carrier frequency for variants in the Centers for Disease Control and Prevention (CDC) Tier 1 gene, LDLR , which causes familial hypercholesterolemia (FH) in the Filipino cluster (OR IBD-09 = 3.9 [1.4, 8.8], FDR = 0.03). This finding is notable given the high prevalence of dyslipidemia in Filipinos 92 ( Figure S4D ). Indeed, LDLR FH variants accounted for 3.1% (population attributable risk [PAR]) of high LDL cases (LDL > 190 mg/dl) in the Filipino cluster. However, due to the low number of carriers in this cluster (<10), the PAR CI was wide [0.3, 26.1], and independent replication is required to refine this finding. Additionally, we identified previously unreported elevated carrier frequency of variants that cause non-syndromic genetic deafness within two different genes in two separate clusters: (1) MYO15A in the Mexican American 1 cluster (OR IBD-04 = 6.9 [2.6, 17.4], FDR = 8.44×10 −4 ) and (2) CDH23 in the Chinese + Korean (OR IBD-05 = 6.0 [2.0, 15.6], FDR = 6.04×10 −3 ) and Japanese (OR IBD-10 = 28.1 [8.9, 77.1], FDR = 7.4 × 10 −6 ) clusters. We also identified previously unreported elevated carrier frequencies in the Ashkenazi Jewish cluster for Pendred syndrome variants in SLC264A (OR IBD-03 = 3.8 [2.9, 5.0], FDR = 8.6 × 10 −19 ) and recombinase activating gene 2 deficiency in RAG2 (OR IBD-03 = 2.9 [1.4, 5.8], FDR = 0.02). Finally, we focused on the clinically actionable American College of Medical Genetics and Genomics (ACMG) 93 genes, where ClinGen-curated variants within these genes manifested a strong EUR bias ( Figure 5b ; Figure S7f ; Methods ). This bias was recently shown in AoU 23 and we provide additional supporting evidence by demonstrating this bias in a single health system. As clinical variant databases are primarily ascertained from EUR populations, we took a conservative approach to call rare predicted deleterious LOF (dLOF) and missense coding variants (dMIS). High-confidence dLOF were defined using the loss-of-function transcript effect estimator (LOFTEE) 94 , which accounts for biological context beyond protein truncation. dMIS variants were defined using a consensus of a majority (5 of 9) of state-of-the-art computational methods based on the Critical Assessment of Genome Interpretation (CAGI) project 95 for pathogenicity prediction ( Methods ; Figure S7g – h ). This approach is distinct from the one taken with ClinGen variants – variants identified in this manner may not have literature evidence for their impact on gene function but are potentially less biased by EUR oversampling. First, we kept our focus on ACMG genes 93 , repeating the ClinGen-based analysis described above using predicted rare dLOF and dMIS variants. We examined the frequency of 1,602 rare dLOF and 6,487 rare dMIS variants within ACMG variants across populations ( Figure 5c ; Figure S7i ). The results were consistent with previous work showing that AFR participants harbor substantially more dLOF compared to other ancestries 85 , and highlighted other groups with previously unreported low or high ACMG putative damaging variants for further investigation. Overall, these results contrasted in comparison with rare variants identified in ClinGen ( Figure 5b ), consistent with the interpretation that ClinGen variant frequencies are biased toward EUR participants, as is the case for most curated clinical datasets. Second, to examine the influence of dLOF and dMIS on ATLAS phenotypes, we performed exome-wide association studies (ExWAS) across 17,537 protein-coding genes ( Methods ). Using Regenie’s unified gene burden association strategy 66 , we identified 1,099 unique Bonferroni-significant gene-trait associations. Within the EUR population, we replicated numerous known associations; 45 of the top 50 associations had been previously reported 84 ( Table S6 ). These include PKD1 with cystic kidney disease and TTN with primary intrinsic cardiomyopathy ( Figure 5e ) 84 . We also identified numerous previously unreported associations ( Table S6 ). In the Northern European cluster, we detected gene-level associations between HNRNPA1L2 and acquired absence of the breast (GENE_PIBD-01 = 1.0×10 −25 ), breast cancer (GENE_PIBD-01 = 2.2× 10-8 ), and malignant neoplasm of the female breast (GENE_P IBD-01 = 1.5 × 10 −7 ). Notably, HNRNPA1L2 is on chromosome 13, ~20 megabases from BRCA2 , representing a different locus. Though HNRNPA1L2 has not been previously associated with breast cancer, it is highly expressed in breast invasive carcinoma 81 and a paralog HNRNPA1 was previously implicated in breast cancer progression 96 . In the Ashkenazi Jewish (IBD-03) cluster, we observed associations between EPG5 and HDL cholesterol level (P LOF_MIS = 3.6×10 −10 ; β LOF_MIS = 1.8 [1.2, 2.3]) and triglyceride level (P LOF_MIS = 1.33×10 −8 ; β LOF_MIS = −1.81 [−2.4, −1.2]). EPG5 is an autophagy tethering factor classically associated with Vici syndrome 97 , a severe developmental disorder, though our findings suggest an additional role in lipophagy and lipid metabolism 98 . In the AMR population, we identified an unreported association between dLOF and dMIS variation in CLN3 and cystic kidney disease (GENE_P AMR = 8.7×10 −9 ), supported by experimental evidence that CLN3 is highly expressed in medullary collecting duct principal cells and plays a role in osmoregulation 99 . We additionally identified unreported associations between PPARG and viral pneumonia (P AMR,MIS = 1.1×10 −6 , OR AMR,MIS = 53.0 [13.5, 208.5]), as well as NADSYN1 and abnormal lung examination findings (P AMR,MIS = 1.2×10 −5 , OR AMR,MIS = 3.7 [2.1, 6.5]). PPARG is known to regulate macrophage response to pulmonary inflammation 100 , while NADSYN1 knockout in mouse models was shown to cause abnormal lung development 101 , though neither has been previously associated with respiratory traits in humans. Within the AFR population and African American cluster, we uncovered unreported associations between DDHD2 and dysphagia (P IBD-06,LOF = 1.2×10 −6 , OR IBD-06,LOF = 8.3 [3.7, 18.7]) and EFCAB13 and essential hypertension (P AFR,MIS = 9.3×10 −7 , OR AFR,MIS = 13.1 [4.5, 37.7]). These findings align with prior studies; variants in DDHD2 have been shown to cause hereditary spastic paraplegia and symptoms of dysphagia 102 , though not in African Americans, and methylation studies have prioritized EFCAB13 as a risk factor for heart failure 103 . We detected an additional association between ANKZF1 and peripheral vascular disease (P AFR,LOF_MIS = 1.6×10 −6 , OR AFR,LOF_MIS = 58.8 [12.6, 275.1]), which is supported by recent experimental findings 104 . Next, in the EAS population, we observed a previously unreported association between AK7 a nd non-rheumatic mitral valve disorders (P EAS,LOF = 7.6×10 −7 , OR EAS,LOF = 85.7 [39.2, 187.2]). While AK7 has never been associated with cardiovascular phenotypes, it has a well-established role in primary ciliary dyskinesia 106 and mutations in other ciliary genes have been associated with mitral valve prolapse 107 . We further identified an unreported association between KIF2B and nephritis/nephropathy in the Chinese/Korean cluster (P IBD-05,MIS = 7.2×10 −7 , OR IBD-05,MIS = 26.2 [9.0, 76.4]). Though KIF2B has not been linked to renal traits, other genes in the kinesin family (KIF) have been implicated in the pathogenesis of renal cancer 108 , 109 . All previously unreported findings described above were independently replicated in AoU 5 , UKBB 69 or BioMe 8 at a nominal P-value < 0.05 ( Table S6 ). Finally, we examined known risk associations to identify ancestry-specific patterns of deleterious rare variation. We tested two known GBA1 rare missense variants, p.Glu365Lys and p.Thr408Met, in the AMR and EUR populations; both variants are reported to increase risk for Parkinson’s disease 110 . Despite similar allele frequencies within each population, p.Glu365Lys was only associated with increased risk in EUR (P EUR,E365K = 0.01; P AMR,E365K = 0.8), whereas p.Thr408Met was only associated with increased risk in AMR participants (P EUR,T408M = 0.2; P AMR,T408M = 0.003), displaying evidence of ancestry-specific risk stratification ( Figure S7j ). While we hypothesize that penetrance is modulated by population-specific genetic backgrounds, further studies are required to identify the causes of these differential associations. In total, We performed replication analyses for 21 of our top previously unreported PheWAS and ExWAS findings in independent biobanks and replicated 16 (76%) at a nominal p value of 0.05. One of the advantages of EHR data is the ability to consider dynamic changes over time. To illustrate this, we integrated all main study components (genetic ancestry, common and rare genetic variants) to study the efficacy of glucagon-like peptide-1 (GLP-1) receptor agonists (GLP1-RAs) for weight loss. We focused on semaglutide, which had the greatest number of prescriptions in ATLAS ( Figure S8a – b ; 7,340 participants; Figure 6a ). We observed a steady decrease in average weight up to ~60 weeks of semaglutide treatment, consistent with previous findings 111 – 113 ( Figure S8c ). We tested if the medication dose, route, sex, age, and initial BMI affected semaglutide efficacy, corroborating that dose and subcutaneous delivery versus oral delivery were positively correlated with weight loss ( Figure 6b ; linear mixed-effects model [LMM]: dose: P Bonferroni = 3.3×10 −97 , semipartial R² (explained variance) = 1.8%, effect size = −1.0; route oral vs. subcutaneous: PBonferroni = 7.2×10 −37 , semi-partial R 2 = 1.1%, effect size = 2.2). We did not detect a significant effect of age and sex on efficacy. Next, we asked if semaglutide’s effects varied by broad-scale ancestry, finding that ancestry significantly influenced weight loss over 60 weeks of treatment ( Figure 6c ; ANOVA on an LMM; ancestry: P Bonferroni = 0.002; ancestry×time: P Bonferroni = 0.0006). Comparisons between EUR and the other populations revealed less weight loss in AMR, and a slower rate of weight loss in AMR and EAS ( Figure 6d ; LMM; ancestry AMR : P Bonferroni = 2.9×10 −3 , β = 0.8; ancestry AMR ×time: P Bonferroni = 5.0×10 −3 , β = 45.2; ancestry EAS ×time: P Bonferroni = 3.2×10 −3 , β = 76.4), consistent with recent analysis based on a single time point, which showed reduced efficacy in AMR and AFR cohorts (here, in AFR, the result was not significant but the direction of effect was consistent). 113 Given limited evidence on the role of inherited factors in GLP1-RAs effectiveness 114 , we used ATLAS to explore the contribution of common genetic variation to semaglutide-related weight loss. We found no correlation between BMI PGS and semaglutide efficacy. However, weight loss was negatively correlated with DM2 PGS ( Figure 6e ; LMM, PGS High vs. Low : P Bonferroni = 2.8×10 −4 , β = 0.96; PGS Med vs. Low : P Bonferroni = 8.8×10 −3 , β = 0.73). The same relationship was observed in a simplified model using only maximum weight loss recorded 114 (linear regression, P Bonferroni = 1.3×10 −2 , β = 0.31; Figure S8d ). We could not find an equally sized cohort for replication, but we used EUR AoU, which had a sample size 50% smaller than the ATLAS cohort used in this analysis with array genotyping (1,578 vs . 3,165 in AoU and ATLAS). In AoU EUR participants, we observed a concordant direction of effect, supporting replication of the finding, although statistical significance was not reached. This may reflect the smaller sample size or other population differences, such as in SES and compliance ( Figure S8e ). Larger samples are needed for more formal replication. As a second step, we conducted a genome-wide association study (GWAS) meta-analysis across ancestries, but did not identify any significant loci ( Methods ; Figure S8f ). Finally, we used WES to test gene-level associations with semaglutide response (Regenie 66 ; Methods ). Given the modest sample size, we focused on EUR individuals and limited our test to the subset of proteins whose plasma abundance in humans was recently shown to be altered by semaglutide 115 . We identified a significant association of weight loss on semaglutide with PTPRU (P Bonferroni = 7.6×10 −3 , β = −0.87) ( Figure 6f ; Figure S8g ). For replication, we used AoU, with a cohort that was 21% smaller than the ATLAS cohort used (1,581 vs . 2,012 in AoU and ATLAS with WES). This analysis supported our original finding by showing a consistent direction of effect, although significance was not reached, possibly due to the smaller sample size or differences between the tested populations (AoU: β = −0.27, P-value = 2×10 −1 ). Therefore, larger sample sizes in future studies will be needed to formally replicate this finding. However, the Bonferroni adjusted combined P from ATLAS and AoU remained significant, supporting this finding (P Bonferroni = 2.3×10 −2 ). This association involves 37 variants, of which one is common (rs2235937; nominally associated with weight loss; P = 9.7×10 −3 , β = −0.06 in EUR), and the rest are rare, the frequency of which varied across ancestries ( Figure S8h ; Methods ). The direction of effect indicates that PTPRU activity is negatively associated with weight loss in patients taking semaglutide (variants in PTPRU with a predicted high or moderate functional consequence contribute to weight loss on semaglutide). Maretty et al. highlighted PTPRU as a protein whose abundance is elevated in individuals with higher genetic risk for increased BMI or type 2 diabetes and is downregulated by semaglutide treatment 115 . Although direct causality has not been established, these observations, together with our findings, support a model in which semaglutide-induced downregulation of PTPRU and genetically reduced PTPRU function promote greater weight loss during semaglutide treatment. This gene has no previous known functional relationship to weight loss or related metabolic functions. However, combined with the data from serum proteomics 115 , our genetic analysis nominates this protein kinase and its pathways as candidates for further investigation.

Resource

Questions should be directed to the lead contact, Daniel H. Geschwind ( [email protected] ). No materials were generated in this study.

Discussion

Large-scale, longitudinal population studies have accelerated our understanding of the causes and consequences of a wide variety of biomedical conditions. The Framingham study 116 , 117 laid the foundation for population-based cohorts, followed by efforts like the UKBB and health system biobanks 3 , 8 , 9 , 11 , 53 , 72 , 84 , 118 , 119 . Despite the known sources of errors and incomplete phenotyping in health system EHRs, multiple studies have shown that many of these factors can be mitigated, allowing robust analyses 3 , 9 , 10 , 120 . Indeed, we were able to validate dozens of previously identified genetic associations in ATLAS based on EHR-derived phenotypes alone. We also leverage the heterogeneity of genetic ancestries in our population to validate and extend our knowledge of disease burden and genetic risk factors. Prior work has mostly relied on broad-scale genetic ancestry to estimate health risk. 5 , 30 , 117 Fine-scale ancestry discerns more recent genetic variation via shared IBD, 120 but smaller clusters can limit power. Through PheWAS and ExWAS across broad- and fine-scale ancestries, we identified numerous unreported genetic associations, many of which were subsequently replicated. We leveraged fine-scale ancestry definitions to benefit populations that are underrepresented in genetics and medical research 122 . For example, to our knowledge, ATLAS consists of the largest genetic dataset of Filipino individuals (n = 1,438 vs. 1,028 in the next largest dataset 58 ), which can be used to improve precision care for this population and fuel discovery. Similarly, our Ashkenazi Jewish cluster of 14,261 participants, which is 2.8 times larger than any prior single-study genetic cohort 57 , enabled previously unreported associations and enhanced risk quantification. Future work should assess the transferability of our results across populations 123 . Discerning genetic variant risk frequency and effect across fine-scale ancestries is powerful for optimized risk stratification and tailoring therapeutic intervention. It has been well demonstrated that diagnostic misclassification due to differences in variant frequencies across ancestries is a serious risk 22 . We emphasize that leveraging diversity within a single regional biobank can reduce confounding by ancestral genetic differences and non-genetic variables that differ across countries or geographically distinct biobanks. We demonstrated the application of longitudinal EHR to gain insights into health outcomes over time, by linking genetic factors to weight loss on semaglutide. This involved integrating dynamic EHR changes with detailed prescription information, including the medication dose and route, which we identified as major confounders. We considered only a single GLP-1RA drug, semaglutide, to avoid the complexity introduced by different GLP-1RAs, which have differing efficacy 124 . Lastly, our study highlights fine-scale populations with unreported low or high ACMG putative damaging variants for further investigation. We expect ATLAS to grow and become an engine for clinical intervention, allowing researchers to rapidly identify individuals for clinical studies and to implement precision medicine as the field evolves. Certain limitations should be considered. First, due to the nature of EHRs, our study captures a partial view of patient phenotypes and care. Second, our analysis focuses on the risk of receiving a diagnosis for a disease, which is related, but not equivalent to the underlying risk of disease development. This may influence how our findings should be interpreted in the context of true disease incidence. Third, we relied on computational predictors for comparing the abundances of predicted damaging variants in ACMG genes across ancestries. Some predictors may introduce ancestry-related biases 125 , 126 , potentially affecting the accuracy of the results. To mitigate these potential biases, we applied rigorous filtering and only retained variants upon concordant agreement in a majority of nine tools. Additionally, comparing the numbers of predicted damaging variants across ancestries may be less stable when small ancestry groups are involved. We therefore excluded groups with fewer than 400 individuals and reflected uncertainty using error bars. Lastly, while our study introduces multiple previously unreported associations, findings that could not be replicated should be further evaluated in independent cohorts as they become available.

Introduction

The integration of electronic health records (EHRs) with genetic data is transforming biomedical research, offering unprecedented opportunities for preventing and managing common medical conditions 1 . Longitudinal sampling, linked with genetic and environmental data, provides advantages over standard cohort-driven research 2 . This has fueled the creation of nationwide biobanks, including the UK Biobank (UKBB) 3 , 4 , All of Us (AoU) 5 , FinnGen 6 and Taiwan Biobank 7 , as well as several large-scale academic biobanks, such as Mt. Sinai’s BioMe 8 , Vanderbilt’s BioVU 9 , Geissinger’s MyCode 10 , and the Michigan Genomics Initiative 11 . Integration of these biobanks has permitted innovative collaborative efforts, such as the eMERGE consortium 12 , COVID-19 host genomics initiative 13 , and the Global Biobank Initiative 1 . Although these efforts have substantially advanced genetic and biomedical discovery, their concentration on participants of European (EUR) ancestry limits generalizability 14 – 18 . As PGS are validated for clinical use, the importance of measuring their accuracy in diverse populations grows; recent analyses show a continuous relationship between ancestral distance from the reference population and the utility of PGS 17 . Similarly, rare genetic variation has substantial ancestry-specific effects and distribution; for example, in African Americans, APOE4 alleles have reduced impact on Alzheimer’s disease risk 19 and rare protective PCSK9 variants are more prevalent 20 . The interpretation of clinically-relevant rare variation is further hampered by a bias towards EUR variants in genetic databases 21 , 22 . Including non-EUR populations reveals substantial disparities in clinically-relevant rare variant frequencies 23 and increases statistical power for discovery 24 , 25 . Thus, greater ancestral diversity in biobanks with detailed medical records strengthens efforts to advance precision health 14 , 26 , 27 . Here, we analyzed data from the UCLA ATLAS Community Health Initiative, linking EHR with genomic information for 92,164 participants with array genotyping and 61,797 with whole-exome sequencing (WES). ATLAS reflects Los Angeles’ ancestral diversity within a single health system 28 – 30 , which reduces confounding from differing clinical practices and enables robust cross-population comparisons. We performed phenome-wide associations using common and rare variants within continental (broad-scale) and sub-continental (fine-scale) ancestral groups, and quantified disease diagnoses and genetic risk across these groups. Using longitudinal EHR data, we employed semaglutide as a case study to identify genetic factors that alter drug efficacy. We identified multiple unreported ancestry-specific risk associations, further demonstrating the utility of ancestral diversity for more equitable, personalized medicine research ( Figure 1a , schematic overview).

Supplementary Material

Figure S1. Whole-exome sequencing quality control, related to the STAR Methods . a. Coverage distribution. The mean and median values were calculated across all samples. Fold enrichment is the degree to which the baited region is enriched compared to the background genomic region. b. Percentages of bases above different coverage thresholds. Every y-axis point shows the percentage of bases above the corresponding x-axis value. The blue area shows 95% of the data, while the surrounding gray lines represent the remaining 5% of the distribution. c. Average base counts by exome capture regions. d. Excluded bases from coverage calculation based on Picard 127 e. Insert size distribution. f. Samtools 128 summary statistics output. g . Genotype number distribution according to RTGtools 129 . h . Summary output from RTGtools. Ti/Tv is the ratio of transition (Ti) to transversion (Tv). This included SNPs from off-target regions. i. The distribution of variant counts across all samples after cohort re-genotyping. Figure S2. Phenotype validation and prevalence, related to Figure 1 and the STAR Methods . ( A–J ) Comparison of vital signs and laboratory test distributions between case and control groups defined by phecodes, with p values indicating significance from Wilcoxon tests. ( K and L ) Overlap between cancer cases defined from EHR cancer diagnosis data (“Cancer Stage Fact”) and phecode-based case-control groups for breast ( K ) and prostate cancer ( L ). ( M ) Prevalence of frequent phecodes. The red bars indicate the prevalence of the most frequent phecodes in subset of the UCLA ATLAS population who had an encounter between 1–2 years after their initial encounter, at the time of sample collection. The blue bars represent the prevalence of phecodes in this same subset at their encounters 1–2 years after their initial encounter. ( N ) The change in phecode prevalence over one year post-collection in the subset of the UCLA ATLAS population who had an encounter between 1–2 years after their initial encounter. The triangles represent the number of new patients, and the blue bars represent the percentage of patients. ( O and P ) The difference in ( O ) phecode group and ( P ) phecode prevalence between participants in the UCLA biobank (purple; “UCLA ATLAS”) at their time of collection compared with all other UCLA Health patients with EHR data (green; “Rest of Data Discovery Repository [DDR]”) within one year of the ATLAS launch date. Logistic regression tests were used to compare the two groups adjusted for sex and age (ORs with 95% confidence intervals, shown in the right panels). After Bonferroni multiple testing corrections, ** denotes a significant difference between the two populations. Figure S3. Broad-scale genetic ancestry and health characteristics, related to Figure 1 . a. Agreement between self-reported race and genetic ancestry predictions. b . The Elixhauser comorbidity index varies across genetic ancestries after applying inverse probability weighting. ANOVA on a linear regression model tested the overall effect of the categorical ancestry predictor, yielding its P-value. Adjusted means and 95% CI are presented per ancestry group. c-d. The relationship between total and hospital encounters and the comorbidity index. r, Pearson correlation, P, the correlation P-value. e . Variation in laboratory or vital sign measurements across broad-scale ancestries. ANCOVA adjusted for genetic sex and age was performed separately for each tested phenotype. Figure S4. Fine-scale genetic ancestry supplementary results, related to Figure 2 . a-b. Quality control of identity-by-descent (IBD) segments. Distribution of IBD segment length per chromosome, before ( A ) and after ( B ) removing human leukocyte antigens (HLA), centromere, and IBD depth outliers. After quality control, IBD segments display exponential decay of segment length as expected (most noticeably for chromosomes 6, 15 and 22). The numbers at the top of each plot represent the chromosome number. c . The distribution of genetic ancestry between fine- and broad-scale populations. IBD clusters were sorted by the predominant broad-scale ancestry of participants in each cluster, rather than cluster size as shown in Figure 2a , providing an alternative visualization. D. Cardio-metabolic disease risk for each fine-scale group within the same broad-scale ancestry. Representative cardio-metabolic phecodes were selected, and only populations with at least 100 participants were tested. In cases of small sample sizes, ‘–’ was used instead of numeric values to protect patient privacy. Filled points represent significant results (FDR ≤ 0.05). Firth’s bias-reduced logistic regression adjusted for BMI, sex, and age was used to obtain OR. Among Asian clusters, Filipino individuals were at high risk for all tested medical conditions (essential hypertension: OR IBD-09 =1.6 [1.4, 1.9], FDR = 8.6 × 10 −11 ; type 2 diabetes: OR IBD-09 = 1.5 [1.3, 1.7], FDR = 7.4 × 10 −6 ; coronary atherosclerosis: OR IBD-09 = 1.4 [1.1, 1.7], FDR = 5.3 × 10 −3 ; abdominal aortic aneurysm: OR IBD-09 = 2.8. [1.4, 5.3], FDR = 7.1 × 10 −3 ; hyperlipidemia: OR IBD-09 = 1.2 [1.0, 1.4], FDR = 2.4 × 10 −2 ). The largest Armenian cluster and Jewish and non-Jewish Iranian clusters showed a high risk for type 2 diabetes (Armenian 1: OR IBD-16 = 2.0 [1.5, 2.7], FDR = 2.5 × 10 −6 ; Iranian Jewish: OR IBD-11 = 2.4 [1.9, 2.9], FDR = 6.0 × 10 −16 ; Iranian: OR IBD-17 = 2.2 [1.6, 2.9], FDR = 1.6 × 10 −6 ), hyperlipidemia (Armenian 1: OR IBD-16 = 1.6 [1.2-2.0], FDR = 8.6 × 10 −4 ; Iranian Jewish: OR IBD-11 = 1.5 [1.2, 1.8], FDR = .1.0 × 10 −4 ; Iranian: OR IBD-17 = 1.8 [1.4, 2.3], FDR = 2.9 × 10 −5 ) and coronary atherosclerosis (Armenian 1: OR IBD-16 = 1.9 [1.4, 2.5], FDR = 8.8 × 10 −5 ; Iranian Jewish: OR IBD-11 = 1.8 [1.5, 2.2], FDR = 4.2 × 10 −8 ; Iranian: OR IBD-17 = 1.9 [1.4, 2.5], FDR = 2.9 × 10 −5 ). Figure S5. polygenic score association results, related to Figure 3 and the STAR Methods . a-e Selected associations between polygenic scores (PGS) and diseases in European (EUR) individuals. f-i The odds ratio (OR) and prevalence of cases for the top and bottom PGS deciles across non-EUR ancestries. Logistic regression was used to obtain the OR and p values, which were adjusted using FDR. Figure S6. PhWAS supplementary results, related to Figure 4 . a-b. Distribution of pruned variant-trait associations. The number of genome-wide significant pruned gene-trait associations shared across fine-scale ancestries, using a prioritized gene within 10kb of the pruned variant and matching effect direction. b presents the same associations described in a, but with a more permissive threshold for gene-trait support from less powered ancestral clusters. c. APOE haplotype frequency across fine-scale cohorts. d. Allele frequency for known risk variants across fine-scale cohorts; shading indicates the level of over- (red) or under- (blue) enrichment of a haplotype/allele in a given cohort via Fisher’s exact test. In c-d, Fisher’s exact test was utilized to compare allele frequency. Bold border indicates a significant result (FDR ≤ 0.1). e. Impact of the non-alcoholic cirrhosis risk variant rs738409-G on cirrhosis and clinical sequalae across finescale cohorts. Odds ratios and 95% CI were calculated using logistic regression. f-h. Replicated low-MAF PheWAS associations between rs115750084-G and major depressive disorder (f), rs77742325-G and osteoporosis with no other symptoms (g) and rs202215133-A and migraine (h). Figure S7. Known or putative rare pathogenic variants, related to Figure 5 Figure 5 and the STAR Methods. a. Frequency differences in Familial Mediterranean Fever (FMF) known risk alleles. b. The risk of carrying the HBB:p.E7V variant. c. The risk of carrying loss-of-function variants in PCSK9. In a-c, the error bars show the 95% Wilson score confidence intervals. d-e. The frequency of BRCA Ashkenazi Jewish founder alleles across broad-scale (d) and fine-scale ancestry groups (e). f. Differences across populations in the total numbers of rare ClinGene P/LP variants in American College of Medical Genetics (ACMG) genes, excluding Ashkenazi Jewish, as a sensitivity analysis to one presented in Figure 5b. ‘Ref’ is the total number of reference alleles, and ‘Alt’ of P/LP ClinGen. Fisher’s exact tests were used to produce odds ratios. Filled points represent significant results at the level of nominal P-value (≤ 0.05). Only fine-scale ancestry clusters with more than 400 participants were included. g-h. Distributions of missense variant pathogenicity rank scores for nine computational tools. Each subplot compares the distribution of all missense variants (purple to green) to known pathogenic variants (red), defined as either: g. ClinGen curated missense variants. h. ClinVar pathogenic/likely pathogenic (P/LP) missense variants. The black dashed line indicates the median rankscore across all missense variants for the given tool. The red dashed line denotes the median rankscore among pathogenic variants in the corresponding dataset. The purple dashed line marks the likelihood-based intersection cutoff derived from the point at which the pathogenic and background distributions cross. i. The number of rare computationally predicted damaging missense and LOF alleles per individual across ancestries. Mann-Whitney U test with a Bonferroni correction was applied to test the difference in the distribution of rare LOF and predicted damaging missense counts per individual between any broad- or fine-scale group compared to all others. Statistically significant differences are indicated by an asterisk (*). Only fine-scale ancestry clusters with more than 100 participants were included. j. Ancestry-specific carrier counts for two known GBA1 rare LOF variants (left) and their impact on Parkinson’s disease risk (right) via logistic regression. Figure S8 . GLP1-RAs ATLAS users and semaglutide investigation complementary data, related to Figure 6 . a . Age by sex of GLP-1 receptor agonist (GLP1-RAs) users. b . Prescription numbers for GLP1-RAs by simple generic names. c . Weight loss patterns across time. Presented are smoothed longitudinal data using a functional boxplot approach, with median values and pointwise intervals between the 20th and 80th quantiles. Vertical tick marks along the x-axis show the deciles of the data distribution, indicating where most data points are concentrated across weeks. d . The relationship between bins of polygenic scores (PGS) for body mass index (BMI) (left) and type 2 diabetes mellitus (right), and weight loss in response to semaglutide. Shown are Loess smoothed plots with 95% CI based on the maximum weight loss for patients across weeks. p values and effect sizes were obtained from a linear regression model with covariates, and a Bonferroni correction was applied. e . The relationship between type 2 diabetes mellitus PGS and weight loss in response to semaglutide in EUR AoU participants. Scaled PGS were divided into groups and linear mixed model fitted values were plotted with 95% CI, based on longitudinal data with repeated weight measurements. P-values were obtained from a linear mixed-effects model that included covariates. f . GWAS Q-Q plot for weight loss on semaglutide displays no significant findings. g . Gene-level test Q-Q plot for weight-loss on semaglutide (genomic inflation factor: 0.95). h. Differences in carrier numbers of semaglutide efficacy involved alleles in PTPRU across ancestries using Firth’s bias-reduced logistic regression. Only variants that exhibited at least one significant difference between EUR and another ancestry are shown. Filled dots represent significant odds ratios (FDR ≤0.05). Table S1. Summary of lab test results and prescription numbers for the top 50 prescribed medications in ATLAS, related to the STAR Methods . Table S2. Phecodes and broad-scale ancestry associations, related to Figure 1 and the STAR Methods . Table S3. Fine-scale ancestry cluster sample sizes compared with published cohorts and phecode associations, related to Figure 2 and the STAR Methods . Table S4. PheWAS results and pharmacogenomic variants, related to Figure 4 . Table S5. Summary of WES variant counts by category and differences across broad-scale ancestries, related to Figure 5 and the STAR Methods . Table S6. ExWAS results, related to Figure 5 .

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

⚙ Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml ⓘ

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-09-27T09:11:36.575535+00:00
License: CC-BY-4.0 · commercial use OK · attribution required
Per Europe PMC