Star
The Institutional Review Board of the University of California, Los
Angeles gave ethical approval for this work. The IRB Number for the ATLAS
Initiative protocol is IRB#17-001013.
Human participants included in this study were from the UCLA ATLAS
Community Health Initiative. Enrollment procedures, consent, and data collection
are described in Johnson et al 30 . Participants provided informed consent for research under
IRB#17-001013. The cohort description is provided below and summarized in Table 1 and Figure 1 .
ATLAS enrollment reflects a modest over-representation of females by
self-reported sex, especially among patients aged 20-60 ( Figure 1c ; Table
1 ). Among ATLAS participants, the median age at the most recent time
point is 61 years for males and 56 years for females. Medical morbidity was
notably more widespread in the biobank population compared to all adult UCLA
patients. This point was evidenced by a significantly larger number of diagnoses
in the biobank population compared to all UCLA patients (15.6
vs. 10.5 mean ICD codes per patient, one year after
collection). The table “Baseline demographic information in ATLAS and in
non-biobank UCLA patients” shows baseline demographic information of the
UCLA ATLAS population, and a subset of all other adult UCLA patients (Rest of
DDR) seen within one year of the ATLAS launch date. All table variables differed
significantly between the ATLAS and non-ATLAS populations (Mann-Whitney test for
continuous variables and chi-square test for categorical variables) after
multiple-testing correction. This included a significant difference overall, as
well as in the proportions of alive and deceased individuals.
The biobank population experienced a substantially higher number of
diagnoses across the most prevalent medical condition categories, including
endocrine, cardiovascular, musculoskeletal, gastrointestinal, neoplasms,
neurological, respiratory, mental, genitourinary and infections ( Figure S2o – p ). For instance,
endocrine/metabolic and cardiovascular diseases were diagnosed in 52.0% and 44%
of the biobank participants versus 38% and 36% of non-biobank adults,
respectively. Biobank participants also experienced substantially more neoplasms
(30% vs. 19%). This likely led to significantly more medical
encounters in the biobank population (119 vs. 64 mean total
medical encounters per patient, and 16 vs. 12 mean inpatient
days), which is consistent with expectations, as greater numbers of clinical
visits facilitate the passive blood collection that powers our biobank. The most
common diagnoses included hypertension (30% of participants), hyperlipidemia
(25%), GERD (18%), anxiety disorder (17%), and depression (15%) ( Figure S2m ).
As time progresses, more diagnoses have an opportunity to be added, and
the rate of diagnoses is consistent across a wide spectrum of organ systems (1-2
years after collection; Figure
S2n ). The table “Characteristics over time” shows
time-dependent characteristics of the UCLA ATLAS population (n=59,949) at the
time of collection, and a subset of all other adult UCLA patients seen within
one year of the ATLAS launch date (“Rest of DDR”, within 1 year of
ATLAS launch date; n=186,895). Both cohorts were limited to a subset of
individuals who had an encounter between 1-2 years after the initial encounter.
Encounter statistics (total encounters, total inpatient encounters, total
inpatient days, types of encounters) summarize the encounters within the first
year after the initial encounter. The changes within each cohort were tested for
a non-zero difference using a Mann-Whitney test for variables available at two
times, and a one-sample t-test for the encounter statistics. * denotes a
significant difference in the change. The cohorts were also compared at each
time point using the Mann-Whitney test (continuous) and the Chi-square test
(categorical). After multiple-testing correction, ** denotes a significant
difference between the two populations at the respective time (including
change). Baseline demographic information in ATLAS and in non-biobank
UCLA patients ATLAS-Overall ATLAS-Alive ATLAS-Deceased Rest of DDR-Overall Rest of DDR-Alive Rest of DDR-Deceased n (%) 88,436 84,593 (95.7) 3,843 (4.3) 410,041 392,915 (95.8) 17,126 (4.2) Age, mean (SD) 54.6 (17.0) 54.1 (16.8) 67.6 (14.8) 51.8 (18.5) 50.9 (18.1) 72.8 (14.8) Self-reported sex, n (%):
Female 49,346 (55.8) 47,688 (56.4) 1,658 (43.1) 234,189 (57.1) 225,735 (57.5) 8,454 (49.4) Self-reported sex, n (%):
Male 39,058 (44.2) 36,873 (43.6) 2,185 (56.9) 175,841 (42.9) 167,169 (42.5) 8,672 (50.6) Self-reported sex, n (%):
Unspecified, X 32 (0.0) 32 (0.0) 11 (0.0) 11 (0.0) Self-reported race, n (%):
American Indian, Alaska Native 785 (0.9) 762 (0.9) 23 (0.6) 2,233 (0.5) 2,175 (0.6) 58 (0.3) Self-reported race, n (%):
Asian 10,537 (11.9) 10,180 (12.0) 357 (9.3) 42,307 (10.3) 40,619 (10.3) 1,688 (9.9) Self-reported race, n (%):
Black, African American 4,079 (4.6) 3,857 (4.6) 222 (5.8) 23,420 (5.7) 22,256 (5.7) 1,164 (6.8) Self-reported race, n (%):
Caribbean/West Indian 161 (0.2) 161 (0.2) 335 (0.1) 332 (0.1) 3 (0.0) Self-reported race, n (%):
Middle Eastern or North African 2,276 (2.6) 2,227 (2.6) 49 (1.3) 5,086 (1.2) 5,017 (1.3) 69 (0.4) Self-reported race, n (%):
Native Hawaiian, Guamanian or Chamorro, Samoan, Other
Pacific Islander 255 (0.3) 247 (0.3) 8 (0.2) 1,050 (0.3) 1,011 (0.3) 39 (0.2) Self-reported race, n (%):
Other Race 3,184 (3.6) 2,737 (3.2) 447 (11.6) 42,188 (10.3) 40,137 (10.2) 2,051 (12.0) Self-reported race, n
(%): Unknown, declined to specify 11,329 (12.8) 11,050 (13.1) 279 (7.3) 62,735 (15.3) 61,692 (15.7) 1,043 (6.1) Self-reported race, n
(%): White 55,830 (63.1) 53,372 (63.1) 2,458 (64.0) 230,687 (56.3) 219,676 (55.9) 11,011 (64.3) Self-reported ethnicity, n (%):
Hispanic/Latino, Cuban, Hispanic/Spanish origin, Mexican,
Mexican American, Chicano/a, Puerto Rican 12,966 (14.7) 12,235 (14.5) 731 (19.0) 50,229 (12.2) 48,100 (12.2) 2,129 (12.4) Self-reported
ethnicity, n (%): Non-Hispanic/Latino 68,784 (77.8) 65,829 (77.8) 2,955 (76.9) 301,251 (73.5) 287,269 (73.1) 13,982 (81.6) Self-reported
ethnicity, n (%): Unknown, declined to specify 6,686 (7.6) 6,529 (7.7) 157 (4.1) 58,561 (14.3) 57,546 (14.6) 1,015 (5.9) Self-reported smoking, n (%):
Former 24,054 (27.2) 22,504 (26.6) 1,550 (40.3) 92,364 (22.5) 85,576 (21.8) 6,788 (39.6) Self-reported smoking, n (%):
Never 60,477 (68.4) 58,366 (69.0) 2,111 (54.9) 275,883 (67.3) 266,871 (67.9) 9,012 (52.6) Self-reported smoking, n (%):
Passive Smoke Exposure - Never Smoker 112 (0.1) 105 (0.1) 7 (0.2) 826 (0.2) 769 (0.2) 57 (0.3) Self-reported smoking, n (%):
Smoker 3,412 (3.9) 3,286 (3.9) 126 (3.3) 27,509 (6.7) 26,796 (6.8) 713 (4.2) Self-reported smoking,
n (%): Unknown, declined to specify 381 (0.4) 332 (0.4) 49 (1.3) 13,459 (3.3) 12,903 (3.3) 556 (3.2) Types of Encounters, n (%):
Inpatient and Outpatient Encounters 28,879 (32.7) 25,725 (30.4) 3,154 (82.1) 72,278 (17.6) 62,279 (15.9) 9,999 (58.4) Types of Encounters, n
(%): Only Outpatient Encounters 59,557 (67.3) 58,868 (69.6) 689 (17.9) 337,763 (82.4) 330,636 (84.1) 7,127 (41.6) Total Encounters, mean
(SD) 118.8 (146.3) 112.1 (137.5) 267.4 (230.9) 64.2 (95.8) 59.7 (88.1) 166.7 (174.3) Total Inpatient Encounters,
mean (SD) 0.8 (2.2) 0.7 (1.9) 3.8 (5.0) 0.4 (1.3) 0.3 (1.1) 2.2 (3.7) Total Inpatient Days, mean
(SD) 16.2 (33.2) 13.4 (28.6) 38.3 (53.6) 12.3 (25.8) 9.4 (18.9) 29.7 (47.2)
Characteristics over time Variable ATLAS- At Collection ATLAS- 1 Year after
Collection ATLAS-Change Rest of DDR-At Collection Rest of DDR-1 Year after
Collection Rest of DDR-Change Age, mean (SD) 56.4 (16.4) 57.6 (16.3) 1.2 (0.4)* 53.5 (17.9)** 54.8 (17.8)** 1.3 (0.5)*,** BMI, mean (SD) 27.3 (6.2) 27.3 (28.2) 0.1 (28.7) 26.8 (24.0)** 26.7 (14.8)** −0.1 (27.6) Diastolic BP, mean (SD) 75.5 (9.7) 75.4 (9.4) −0.1 (10.4) 75.4 (10.2)** 75.6 (10.1) 0.2 (10.5)*,** Systolic BP, mean (SD) 125.0 (16.7) 124.8 (16.8) −0.2 (17.1) 125.3 (17.6) 125.3 (17.6) 0.0 (16.8)** Number of ICD Codes, mean
(SD) 14.2 (11.2) 15.6 (11.9) 1.5 (6.5)* 9.4 (8.4)** 10.5 (9.1)** 1.1 (5.1)*,** Number of Phecodes, mean
(SD) 7.1 (5.7) 7.8 (6.0) 0.7 (3.2)* 4.7 (4.3)** 5.2 (4.6)** 0.5 (2.6)*,** Charlson Comorbidity Index,
mean (SD) 1.0 (1.8) 1.1 (1.9) 0.1 (1.0)* 0.7 (1.5)** 0.7 (1.5)** 0.1 (0.8)*,** Elixhauser Comorbidity Index,
mean (SD) 2.6 (8.3) 2.9 (8.7) 0.3 (4.4)* 1.6 (6.8)** 1.7 (7.1)** 0.2 (3.6)** Total Encounters, mean
(SD) 23.2 (27.2)* 12.0 (15.8)*,** Total Inpatient Encounters,
mean (SD) 0.2 (0.7)* 0.1 (0.4)*,** Total Inpatient Days, mean
(SD) 1.4 (7.8)* 0.4 (4.2)*,** Inpatient and Outpatient
Encounters (%) 8,413 (14.0) 9,869 (5.3)** Only Outpatient Encounters
(%) 51,536 (86.0) 177,026 (94.7)**
Baseline demographic information in ATLAS and in non-biobank
UCLA patients
Characteristics over time
We retrieved the latest results of the 39 most common laboratory tests:
blood count, lipid panel, metabolic panel, HbA1c, and vitamin D,25-Hydroxy to
explore test frequencies and variation in results. The largest variability in
adults was observed for bilirubin (median = 0.4, IQR = 0.3), followed by
triglycerides (median = 90, IQR = 66), alanine aminotransferase (median = 22,
IQR = 14), neutrophils (median = 3.73, IQR = 2.07), and LDL cholesterol (median
= 97, IQR = 50). Similarly, we retrieved information on the 50 most abundant
filled prescriptions to learn about treatment and prescribing tendencies. ( Table S1 ). As hospital
visits and surgeries were the most common types of encounters in the biobank
population, it is not surprising that among the most prescribed generic
medications were Acetaminophen (n = 738,061 prescriptions), Ondansetron HCl (n =
636,165), Propofol (n = 633,830), and Lidocaine HCl (n = 492,838) ( Table S1 ).
ATLAS consists of 32% participants of non-European ancestry and a
representation of five broadscale and 36 fine-scale genetic ancestry groups
within a single health system (see METHOD
DETAILS to understand how these were defined). By comparison, the
vast majority of the UK biobank 4 participants are EUR (95.8% EUR, 2.1% South Asian, and 2.1%
AFR). FinnGen 6 is ~98% Finnish
EUR, Taiwan Biobank 7 is over
99% EAS Han Chinese, Michigan Genomics Initiative 11 is 88% EUR, and Geisinger
MyCode 169 is
approximately 97% EUR. While others like All of Us, Mt. Sinai’s BioMe,
and Vanderbilt’s BioVU reflect diverse urban populations, ATLAS captures
a wider and more detailed range of genetic ancestries in the Los Angeles
population. All of Us 5 includes
a larger portion of Black participants, but a smaller portion of Asian and
Middle Eastern/North African self-identified individuals (under a combined race
and ethnicity category: White 53.3%, Black or African American 21.2%, Hispanic
or Latino 17.8%, Asian 3.1%, Middle Eastern or North African 0.6%). Mt.
Sinai’s BioMe 53 , the
only biobank with reported fine-scale ancestries, included 17 fine-scale
clusters vs. 36 in ATLAS, and substantially fewer Asian
individuals. Vanderbilt’s BioVU 170 and Penn Medicine Biobank 119 include small Admixed American
populations, and these biobanks, along with the VA Million Veterans
Program 72 and the
Colorado biobank include small Asian populations. The tables below show
comparisons of ancestry numbers and ratios across biobanks. Unclassified
participants are not included. Some biobank studies have larger sample sizes but
lack inferred genetic ancestry and are thus not included. Numbers of biobank participants by genetic ancestry Biobank AMR AFR EUR EAS SAS Total UK Biobank (National) - 9,633 431,805 - 9,252 450,690 All of Us (National) 28,901 34,037 101,613 3,255 - 167,806 Taiwan Biobank (National) - - - 108,955 - 108,955 Mount Sinai BioMe
(Academic) 10,638 6,983 8,477 780 617 27,495 Vanderbilt BioVU
(Academic) 2,466 15,597 69,810 896 414 89,183 Geisinger MyCode
(Academic) - 1,377 44,522 - - 45,899 Michigan Genomics
(Academic) 843 5,962 75,943 2,172 1,435 86,355 VA Million Veterans Program
(Academic) 59,048 121,117 449,042 6,702 - 635,909 Penn Medicine Biobank
(Academic) 711 11,300 30,360 680 573 43,624 Colorado Biobank
(Academic) 18,137 7,466 145,070 4,080 - 174,753 UCLA
ATLAS 12,822 4,061 62,902 8,080 1,761 89,626
Numbers of biobank participants by genetic ancestry
Below, percentages exclude unclassified participants and therefore
differ from Figure 1 . Percentages of biobank participants by genetic ancestry Biobank AMR AFR EUR EAS SAS UK Biobank (National) 0.0% 2.1% 95.8% 0.0% 2.1% All of Us (National) 17.2% 20.3% 60.6% 1.9% 0.0% Taiwan Biobank (National) 0.0% 0.0% 0.0% 100.0% 0.0% Mount Sinai BioMe
(Academic) 38.7% 25.4% 30.8% 2.8% 2.2% Vanderbilt BioVU
(Academic) 2.8% 17.5% 78.3% 1.0% 0.5% Geisinger MyCode
(Academic) 0.0% 3.0% 97.0% 0.0% 0.0% Michigan Genomics
(Academic) 1.0% 6.9% 87.9% 2.5% 1.7% VA Million Veterans Program
(Academic) 9.3% 19.0% 70.6% 1.1% 0.0% Penn Medicine BioBank
(Academic) 1.6% 25.9% 69.6% 1.6% 1.3% Colorado BioBank
(Academic) 10.4% 4.3% 83.0% 2.3% 0.0% UCLA
ATLAS 14.3% 4.5% 70.2% 9.0% 2.0%
Percentages of biobank participants by genetic ancestry
Method
Our study moves from characterizing the ATLAS cohort to ancestry-stratified
analyses of disease burden – first at the broad-scale ancestry level, then at
the fine-scale cluster level- followed by evaluations of common and rare genetic
risk, and finally, an integrative case-study using longitudinal EHR data.
Array genotypes were obtained using the Global Screening Array. All data
was mapped to GRCh38 and dbSNP, build 147 171 . Common haplotypes and variants were imputed using the
TOPMed Freeze 5 panel 29 , 30 using 668,127 observed
single-nucleotide polymorphisms (SNP), resulting in a total of 50,757,223
high-quality called genotypes variants following imputation, an average of
2,048,050 per individual. Details regarding QC and the imputation procedure were
described before 29 . Minimal QC
was applied to the imputed genotypes. We retained only non-duplicated,
bi-allelic variants with high imputation quality (R2 > 0.7), a minor
allele frequency (MAF) > 0.1%, and a missing rate 5% were excluded from further analysis. Concordance
between genotypes determined by observed array variants and targeted Illumina
sequencing was approximately 99.6%, determined using the vcf-compare command in
VCFtools (v0.1.16) 172 .
Demographic information, vitals, lab tests, International Classification
of Diseases (ICD) codes, and medication prescriptions were retrieved from the
UCLA Data Discovery Repository (DDR), established on March 2, 2013 29 , containing deidentified
participant EHR data from our health system.
Demographic, vital, and lab data were up to date as of Aug 24, 2024. For
vitals and encounter counts, only records from in-person encounters (hospital
visits, appointments, surgeries, office visits, and walk-ins) were considered.
For vital signs, median values across all encounters were used to reduce the
impact of potential recording errors or error-prone self-reported values. For
lab data, the most recent values were considered. For Figure 1c , self-reported sex and age at the most
recent time point at the time of analysis were used.
Yearly encounter numbers were retrieved in 2024. We included only
complete yearly data up to 2023 to ensure consistency and avoid partial data
from 2024. Hospital encounters during the COVID pandemic years (2019-2022) were
excluded.
For associations involving clinical phenotypes, both ICD-9 and 10 were
extracted. To ensure harmonized and comparable phenotypes across data sources,
we adopted a structured, standard approach using phecodes 59 – 61 . We utilized the existing PheWAS catalog maps and
applied standardized rules for case–control definitions. The ICD-9 codes
were mapped to phecodes using phecode Map v1.2 59 , while ICD-10 codes were mapped using
phecode Map v1.2b1 60 , both of
which were obtained through the PheWAS catalog 61 . Cases were defined using the widely
used “Rule of Two” approach 61 , 173 – 176 . For each phecode,
participants were classified as cases if they had at least two occurrences of
the same phecode, with occurrences spaced at least 30 days apart to ensure
persistence of diagnosis. In total, there were 1,308 phecodes with at least 100
cases in ATLAS. Controls were defined as participants who had no record of the
phecode in their EHR data and at least two recorded encounters in the system,
following similar established rules for minimum data content 61 , 177 , 178 .
Undecidable participants were excluded.
We used phecodes 59 – 61 to
define phenotypes and validated the generalizability of these phenotypes using
several approaches. First, we tested cross-checked information between a variety
of vital signs and lab tests and specific case-control groups defined by
phecodes, using Wilcoxon tests and density plots, which confirmed the expected
differences for these phenotypes. This included the following: 1) phecode-lab
test pairs: Type 2 diabetes-Hemoglobin A1c, Hyperlipidemia-Trygllicoride,
Vitamin D deficiency-Total 25-Hydroxy vitamin D, Deficiency anemias-Mean
Corpuscular Volume (MCV), Deficiency anemias-Hemoglobin, and 2) the
phecode-vital sign pairs: Essential hypertension-Blood pressure diastolic,
Essential hypertensionBlood pressure systolic, Obesity-BMI, Anorexia
nervosa-BMI, Short stature-Height ( Figure S2a – j ). Second, we replicated many
known associations, including ancestry-related differences in disease prevalence
for 28 major phenotypes. For example, across cancer types, we observed the
highest risk for prostate cancer in AFR ancestry, stomach cancer in EAS, bladder
cancer in AMR, and breast cancer in EUR 179 (see more below and in Table S2 ). Across fine-scale
ancestries, for example, we showed increased breast cancer and Crohn’s
disease diagnoses in the Ashkenazi Jewish (IBD-03) cluster 53 , 180 (see more below). We also replicated 11,756
variant–phenotype associations, such as a lower risk of
Alzheimer’s disease for African American participants with the
ε4/ε4 haplotype (see more below). PGS prediction of 28 traits
revealed significant enrichment in the top PGS decile for 89% of these traits in
EUR (see more below). Collectively, the replication of known associations using
ATLAS-defined phecodes indicates high-quality, well-defined phenotypes that
match external definitions. Third, we used EHR data on cancer diagnoses to
validate prostate and breast cancer cases (the two most diagnosed cancers)
against phecode-defined case-control groups. We found a high agreement of 95%,
for both cancers, considering decidable patients based on phecodes ( Figure S2k – l ). The following
sections describe replication results conducted to validate the quality of the
data and the phenotype definitions.
We assessed variation in medical conditions and disease prevalence
across populations across preselected prevalent major conditions ( Table S2 ). This
confirmed known ancestry-disease associations. Across cancer types, we
observed the highest risk for prostate cancer in AFR ancestry, stomach
cancer in EAS, bladder cancer in AMR, and breast cancer in EUR 179 . We also replicated
several known cardiovascular disease associations. For instance, AFR
participants were more affected by hypertension and myocardial infarction,
while those with SAS ancestry had the lowest risk of atrial fibrillation,
despite the highest incidence of coronary atherosclerosis, an established
but seemingly contradictory risk profile 181 . Major metabolic disorders were
more prevalent in AFR than in other ancestries. This finding included
diagnostic codes (9- and 10-ICD codes) related to Type 2 diabetes,
hypercholesterolemia, and hyperlipidemia, supported by increases in
metabolic-related laboratory measurements and vital signs ( Figure S3e ). Type 2 diabetes
was more common in all non-EUR groups relative to EUR. Previous work has
suggested that some of these differences reflect social and environmental,
rather than genetic, factors 182 . With regard to neurological conditions, we find a
higher risk of dementia in AFR participants, but a lower risk of migraines
and Parkinson’s Disease in AFR compared to EUR patients, consistent
with prior analyses 183 – 185 . Among neuropsychiatric disorders, anxiety and major
depressive disorders were most frequent in EUR participants, and least
frequent in the continental Asian cluster, as previously reported 186 , 187 . These findings demonstrate the
utility of ATLAS in robustly replicating known associations within a single
health system, reducing potential confounding due to geographic or health
system-related effects.
We tested the prevalence of 1,253 phecodes 59 – 61 across fine-scale clusters with at least 100
participants ( Table
S3 ; selected tests, Figure
2c ). As a measure of integrity, we tested for known increases in
prevalence in these fine-scale ancestries, for example, replicating an
increase in breast cancer and Crohn’s disease in the Ashkenazi Jewish
(IBD-03) cluster. Similarly, we observed the known increase in gout in the
Filipino cluster (IBD-09) 188 and Alzheimer’s disease and dementias in the
Puerto Rican cluster (IBD-15) 189 . Consistent with the global ancestry findings,
African American (IBD-06) had the highest prevalence of hypertension.
We leveraged genotyping in our cohort to calculate individual PGS
for a range of cancer, cardiovascular, metabolic, neuro-psychiatric and
autoimmune diseases and tested their relationship to disease risk. We
focused on the top end of the PGS distribution (10%), compared with the
5 th decile, in EUR participants ( Figure 3 ; see the table below). For Type 1
diabetes, the top PGS decile was the most enriched (OR = 11.6 [7.6, 19.0],
FDR = 4.1×10 −25 ), with 41% of the diagnosed
participants in ATLAS assigned to the top PGS decile. The second most
enriched trait was Crohn’s disease, with 33% of diagnosed
participants within the top PGS decile (OR = 5.5 [4.0, 7.7], FDR =
7.9×10 −24 ), followed by gout with 27% (OR = 5.1
[4.1, 6.4], FDR = 2.0×10 −45 ), testicular cancer
with 25% (OR = 3.7 [1.2, 7.3], FDR = 1.2×10 −4 ), and
prostate cancer with 21% (OR = 3.4 [2.8, 4.0], FDR =
2.0×10 −45 ) ( Figure S5a – e ). On average, the
top decile of risk accounted for 18.6% of diagnosed participants across 28
preselected prevalent disorders. Twenty-five of 28 (89%) showed significant
enrichment of participants within the top PGS decile (mean OR for
significance tests = 2.9). Performance declined when we applied PGS to
non-EUR populations, as expected 55 – 58 .
This was due to both reduced sample sizes and the model fit ( Figure S5f – i ), identifying only
17, 27, 36, and 38% significant associations of diseases with the top PGS
decile, for SAS, AFR, EAS, and AMR, respectively. Concordantly, the top PGS
decile accounted for fewer cases (on average, 13, 15, 16 and 16%, for SAS,
AFR, EAS and AMR). This further supports the need for larger, more
ancestrally heterogeneous cohorts for clinical development of PGS 55 – 58 . The OR and prevalence of cases for the top and bottom
PGS deciles across traits in EUR Trait OR top OR bottom Cases top (%) Cases bottom (%) Cases sample size FDR top FDR bottom Atrial fibrillation 3.2 0.5 19.5 4.8 4565 6.60×10 −66 1.90×10 −13 Bipolar 1.6 0.7 15.5 7.4 1147 9.60×10 −5 0.011 Bladder cancer 1.8 0.8 15.2 6.9 698 0.00028 0.3 Breast cancer 2.4 0.5 18.0 4.9 3016 2.20×10 −26 2.40×10 −08 Cerebrovascular disease 1.3 1.1 9.9 10.5 334 0.34 0.71 Colorectal cancer 1.6 0.6 15.0 5.9 842 0.0026 0.0064 Coronary atherosclerosis 2.5 0.5 16.0 6.1 7951 2.20×10 −59 2.70×10 −24 Gout 5.1 0.4 27.2 2.7 1638 2.00×10 −45 1.10×10 −05 Hypercholesterolemia 1.4 0.6 12.7 6.8 10836 1.50×10 −14 5.70×10 −24 Hyperlipidemia 1.6 0.6 12.2 7.8 21829 6.70×10 −29 1.30×10 −30 Hypertension 2.1 0.6 12.6 7.8 21517 2.50×10 −63 1.10×10 −28 Hypertrophic obstructive
cardiomyopathy 1.7 0.4 17.8 3.9 129 0.13 0.11 Lung cancer 1.1 0.7 10.3 7.2 976 0.65 0.062 Major depressive disorder 1.6 0.7 13.0 7.1 10793 3.00×10 −22 1.50×10 −10 Malignant neoplasm of
testis 3.7 0.4 24.9 2.9 173 0.00012 0.13 Melanoma 2.3 0.5 19.6 3.4 1196 8.80×10 −12 1.00×10 −4 Multiple sclerosis 2.8 0.3 26.1 2.9 379 6.00×10 −7 0.0024 Myocardial infarction 1.6 0.8 12.9 7.2 1744 2.30×10 −5 0.11 Ovarian cancer 2.1 0.8 17.9 6.7 403 0.00066 0.35 Pancreatic cancer 1.6 0.7 12.6 5.2 382 0.048 0.29 Prostate cancer 3.4 0.3 21.5 2.9 2703 2.00×10 −45 1.10×10 −17 Psoriatic arthropathy 1.5 0.9 14.1 8.0 503 0.029 0.53 Crohn’s disease 5.5 0.6 33.0 3.6 758 7.90×10 −24 0.11 Schizophrenia 2.9 0.3 21.9 2.3 128 0.01 0.11 Systemic lupus
erythematosus 2.9 0.9 22.0 6.3 569 4.40×10 −9 0.56 Thyroid cancer 3.0 0.6 21.4 4.1 786 2.60×10 −12 0.022 Type 1 diabetes 11.6 0.4 40.8 1.5 537 4.10×10 −25 0.05 Type 2 diabetes 2.3 0.4 16.9 4.3 6302 2.20×10 −48 8.60×10 −26
The OR and prevalence of cases for the top and bottom
PGS deciles across traits in EUR
Replication of genotype-phenotype associations is important to
assess the quality of both genetic and phenotype data, as well as to
evaluate the power and effectiveness of the biobank in detecting true
genetic associations. To address this, we first examined the distribution of
APOE alleles across fine-scale clusters, highlighting
the increased frequency of ε4 risk alleles in the
African American (IBD-06) and Bantu (IBD-35) clusters ( Figure S6c ) and replicating the
finding that African American participants with the
ε4/ε4 haplotype have a lower risk of
Alzheimer’s disease compared to other populations (see in the table
below FDR-significant results). Next, we replicated two well-known genetic
associations 190
in the African American cluster, between HBB rs334-A ( Figure S6d ; MAFIBD-06
= 5.2%) and a diagnosis of sickle cell anemia (PIBD-06 =
2.0×10 −78 ; ORIBD-06 = 15.1) and between the
Duffy null ACKR1 rs2814778-C polymorphism ( Figure S6d ; MAFIBD-06 = 77.1%)
and a decrease in neutrophil count (PIBD-06 =
4.0×10 −31 ; ORIBD-06 0.7). Using ATLAS blood
work results, we further quantified the impact of rs334-A on mean
corpuscular hemoglobin concentration (PIBD-06 =
5.2×10 −15 ; βIBD-06 = 0.4), nucleated red
blood cell count (PIBD- 06 = 1.6×10 −10 ;
βIBD-06 = 0.2), and mean corpuscular volume (MCV) (PIBD-06 =
3.3×10 −8 ; βIBD-06 = − 0.3).
Across multiple Asian clusters, we confirmed the association between a
variant in high LD with the -- SEA deletion, which causes
inherited alpha-thalassemia 191 , and microcytic anemia. We observed a significant
decrease in MCV in response to LUC7L rs372755452-A in the
broad-scale EAS ancestry (PEAS = 6.1×10 −88 ;
βEAS = −1.6; MAFEAS = 0.9%), and the fine scale Chinese +
Korean (PIBD-05 = 2.7×10 −52 ; βIBD-05 =
−1.6; MAFIBD-05 = 0.8%) and Filipino (PIBD-09 =
2.2×10 −24 ; βIBD-09 = −1.8;
MAFIBD-09 = 1.2%) clusters.
Next, we replicated the finding that AMR participants are twice as
likely to carry a PNPLA3 rs738409-G missense variant that
greatly increases the risk for non-alcoholic fatty liver disease ( Figure S6d ), a major
cause of cirrhosis that often necessitates liver transplant 67 . Using fine-scale
ancestral mapping, we studied the association between rs738409-G and
non-alcoholic cirrhosis of liver in two Mexican American clusters ( Figure S6e ), IBD-04
(PIBD-04 = 2.7×10 −14 ; ORIBD-04 = 1.90) and IBD-07
(PIBD-07 = 1.5×10 −8 ; ORIBD-07 = 1.9), identifying
similar risk but differing minor allele frequencies across these clusters
(MAFIBD-04 = 45.0%; MAFIBD-07 = 52.9%). The same variant in Northern
Europeans had a smaller effect (PIBD-01 = 3.6×10 −9 ;
ORIBD-01 = 1.5) and was half as frequent (MAFIBD-01 = 22.9%). Ancestry-specific effects of APOE ε4 on
Alzheimer’s disease risk Haplotype Ancestry OR P-value CI low CI high e4e4 All 2.87 1.25×10 −41 2.45 3.28 e4e4 EUR 2.86 1.40×10 −31 2.38 3.34 e3e4 All 1.06 3.88×10 −19 0.82 1.29 e3e4 EUR 1.01 1.43×10 −13 0.74 1.28 e4e4 AMR 4.18 5.99×10 −6 2.37 5.99 e3e4 EAS 1.93 1.23×10 −4 0.95 2.92 e4e4 AFR 1.76 3.10×10 −3 0.59 2.93 e4e4 EAS 4.89 4.59×10 −3 1.51 8.26 e3e4 Unclassified 1.80 8.83×10 −3 0.45 3.15 e4e4 Unclassified 5.66 9.62×10 −3 1.37 9.94
Ancestry-specific effects of APOE ε4 on
Alzheimer’s disease risk
Analyzing clinically relevant rare variants, known to have elevated
frequencies in specific populations, replicated known ancestry-associated
patterns, assessing the quality of the data and the robustness of ATLAS in
capturing ancestry-associated allele frequency patterns. Focusing on
Familial Mediterranean Fever (FMF) variants, we identified 753 carriers with
the highest risk in the Armenian clusters 192 , 193 ( Figure S7a ). Carriers of these
variants also had a higher risk of the amyloidosis phecode, which is known
to be associated with FMF (OR = 3.7 [2.2, 5.8]). The HBB:p.E7V variant,
which is responsible for most sickle cell anemia cases, was carried by 273
participants, with elevated frequency in the African American cluster
(OR IBD-06 = 51.4 [39.7, 67.0]; Figure S7b ), in line with
previous findings 194 .
Finally, we explored protective variants in the PCSK9 gene
causing lowered LDL levels 195 and identified 49 carriers of loss-of-function
variants within the African American cluster (OR IBD-06 = 9.5
[4.8, 17.8]) ( Figure
S7c ), as is known 20 .
While the above variants were selected via a
literature search, we also queried the entire ClinGen pathogenic or likely
pathogenic (P/LP) variant catalog, aggregated by gene, to identify
enrichment in fine-scale populations. Among replicated associations,
variants in BRCA1 and BRCA2 were our first
candidates of interest. Ashkenazi Jewish had the primary risk for carrying
either ClinGen P/LP variants in BRCA1 or
BRCA2 ( BRCA1 , OR IBD-03 =
47.1 [20.6, 133.0]; BRCA2 , OR IBD-03 = 48.2
[23.6, 114.2]) ( Figure 5a ). Of note,
these observations are based on six P/LP BRCA ClinGen variants found in
ATLAS, which include only two out of three of the Ashkenazi Jewish founder
alleles (at the time of the analysis, the Ashkenazi Jewish founder allele
NM_007294.3 :c.5266dup was not included in the ClinGen Evidence Repository).
The population attributable risk (PAR) was 3.0 [1.8, 5.0]% and 3.2 [1.9,
5.3]% for BRCA1 and BRCA2 variants,
respectively, in the Ashkenazi Jewish cluster for breast cancer. Independent
of ClinGen, we considered the three BRCA Ashkenazi Jewish
founder alleles and found that they were most prevalent in EUR among
broad-scale ancestries, and in Ashkenazi Jewish among fine-scale clusters
( Figure
S7d – e ). In Ashkenazi Jewish, the PAR for breast cancer, associated
with the three Ashkenazi Jewish founder alleles combined, was 8.2 [5.9,
11.2]% when all ages were included and increased to 12.5 [8.5, 18.1]% in
participants under 70. This increase is expected, since BRCA founder alleles
are associated with early-onset breast cancer. In Northern Europeans, we
observed enrichments of P/LP variants in MYOC
(OR IBD-01 = 2.7 [1.7, 4.2]) that are associated with juvenile
open angle glaucoma, with a PAR of 0.7 [0.2, 3.4]%.
A bias toward EUR in ClinGen variants was recently shown in
AoU 23 . We tested
whether we could replicate this bias using a single health system. In this
part, we asked if the total allele frequencies of the clinically actionable
ACMG ClinGen P/LP variants vary across broad- and fine-scale ancestries.
Overall, 17 ACMG genes had at least one P/LP variant in ClinGen present in
ATLAS participants. For a more nuanced biological understanding, we divided
the ACMG variants into two groups of rare LOF (n = 53) and rare P/LP
missense (n = 131) variants. We defined ‘LOF’ for variants
ranked as high-confidence LOF by LOFTEE 94 and ‘missense’ based on the
VEP 137
“missense_variant” annotation (see below, Differences in total ClinGen allele frequency within ACMG
genes across ancestries ). We observed the highest frequency of
rare P/LP LOF variants in EUR participants among broad-scale populations,
and in the Ashkenazi Jewish cluster (IBD-03) among fine-scale clusters
( Figure 5b ). This number was
primarily driven by BRCA P/LP alleles, which accounted for 63% of all
ClinGen LOF P/LP variants in EUR and 94% in the Ashkenazi Jewish cluster
(OR EUR = 3.7 CI EUR = 2.6-5.4,
P Bonferroni = 6.3×10 −17 ;
OR Ashkenazi Jewish = 6.5, CI Ashkenazi Jewish = 5.3
- 8.1, P Bonferroni = 3.5×10 −62 ). Still,
when Ashkenazi Jewish individuals are removed from the broad-scale EUR
category, a substantial enrichment of ClinGen missense variants is observed
for this group (OR EUR
(non-Ashkenazi Jewish) = 1.8 [1.2, 2.7], P-value = 0.002,
P Bonferroni = 0.01; Figure S7f ), demonstrating that
the signal is not derived from Ashkenazi Jewish individuals alone. No
differences in rare P/LP missense total counts were identified across
broad-scale ancestries, but across fine-scale ancestry clusters, Northern
EUR participants (IBD-01) had significantly more (OR Northern EUR
= 1.4 [1.2, 1.8], P Bonferroni = 0.02), with a similar trend
observed in Southern Europeans that did not reach statistical significance.
This suggests that clinical datasets are particularly biased toward Northern
EUR driven by the composition of large contributing cohorts such as the UK
Biobank. Missense variants were depleted in Ashkenazi Jewish, possibly since
they are underrepresented in most EUR large scale cohorts, or due to
technical differences in how studies or countries classify variants as P/LP.
Excluding Ashkenazi Jewish as a sensitivity analysis resulted in a similar,
albeit weaker trend, of enrichment of missense variants in Northern EUR (OR
= 1.3 [1.01, 1.6], P-value = 0.04, P Bonferroni = 0.5; Figure S6f ).
As all samples were collected incidentally, there was interest in
characterizing the population, especially in comparison to non-biobank UCLA
patients. To characterize the non-biobank UCLA patients while mitigating
time-dependent confounding, we included a subpopulation with an encounter within
one year of the ATLAS launch date. Encounters could be either inpatient or
outpatient, and we summarized the relative proportion of patients with only
outpatient encounters or at least one inpatient encounter (there were no
patients with only inpatient encounters).
To identify clinical phenotype patterns in ATLAS and to compare these to
patterns of non-biobank UCLA patients, we used all ICD codes from each patient
to encapsulate past and present conditions, subject to practical
challenges 196 . For
the ATLAS population, we captured a snapshot using their encounter closest to
their biobank sample collection date. For the non-biobank patients, we used
their closest encounter to the ATLAS launch date. These encounters represented
the baseline encounters for both populations. The ICD codes at these encounters
were converted to phecodes and phecode groups 59 – 61 , 197 to
represent meaningful categories of disease. Phecodes were extracted using pandas
v2.2.2 with Python v3.9.19. We reported the unique phecodes with a prevalence
≥5% in the UCLA ATLAS population. ICD codes at the baseline encounters
were also used to calculate the Charlson and Elixhauser Comorbidity Indices
40 , 41 , 198 , both measures that predict mortality. The scores were
derived using the comorbidity R function v1.0.7 154 with R v4.1.2. To compare the
prevalence of disease categories between UCLA Biobank participants and
non-biobank UCLA Health patients, we used logistic regression to test the
association of each category with participant group, adjusting for age and
sex.
In addition to characterizing patient populations at a baseline time, we
also described differences in how they changed over time, an important
consideration when assessing relative disease burden across populations.
Participants were enrolled in UCLA ATLAS and were encountered within the health
system at different times; therefore, we controlled the interval over which we
measured change. We identified patients with at least one encounter between one
and two years after their baseline encounter. From this subpopulation, we
obtained patients’ first encounter within this window and summarized the
cumulative encounters that occurred between the two times. We also used the ICD
codes at the second encounter to derive new comorbidity scores, as well as
extract updated phecodes. The change in comorbidity scores and increases in
phecode prevalence provided insight into the evolving disposition of the UCLA
ATLAS population over time.
The genetic ancestry of participants in the ATLAS dataset was estimated
by assessing their proximity to the centroids of 1000 Genomes superpopulations
in principal component (PC) space. PCA across common genetic variants
demonstrates granular relationships and provides a quantitative basis for
assessing relationships between ancestry and disease 29 , 36 . The top 20 PCs were calculated using the bed_projectPCA
function (the bigsnpr 155 R
package v1.12.2) with default parameters. For every individual, the Euclidean
distance to the centroids of the five broad-scale populations (AMR, AFR, EUR,
EAS, SAS) was computed. Participants were assigned AMR or AFR ancestry if the
nearest centroid corresponded to one of these populations, as these groups are
well-separated in PC space. For EUR, EAS, and SAS ancestries, which exhibit more
genetic overlap, a stricter distance threshold was enforced to minimize
misclassification. Specifically, an individual was assigned to one of these
ancestries if their distance to the nearest centroid was less than a scaled
threshold, calculated as:
Threshold=max(dist)×min(FST)/max(FST)×0.5, where max(dist) is the
largest squared distance among centroids and min(FST)/max(FST) accounts for
genetic differentiation. Participants who could not be assigned to any ancestry
cluster under these criteria were labeled as “admixed/unknown.”
Visualization was performed using the Boutros Plotting General (BPG) R package
v.7.1.0 156 .
Variation in encounter numbers across broad-scale ancestries was tested
using ANCOVA, adjusted for genetic sex, age, and BAS rank to account for
healthcare access differences. Comorbidity index values across broad-scale
ancestries were compared using ANCOVA, adjusted for genetic sex, age, and both
BAS and ADI ranks, which provide complementary measures of socioeconomic status.
Since not all patients had BAS and ADI values, the sample sizes for these
analyses were reduced (numbers are specified in the body of each figure).
Adjusted means and CI were calculated using the R emmeans package, version
1.10.5 132 .
To test differences in disease diagnosis across ancestries, phecodes
(retrieved as described in Retrieving phenotype
data ) were associated with genetic ancestry populations using
logistic regression, adjusting for age (age at diagnosis for cases and the
latest age for controls) and genetic sex if applicable. The following
disease-phecode pairs were used to define disease diagnosis: Type 2
diabetes-Type 2 diabetes, Sleep apnea-Sleep apnea, Essential
hypertension-Essential hypertension, Hyperlipidemia-Hyperlipidemia, Anxiety
disorders-Anxiety disorders&Anxiety disorder&Generalized anxiety
disorder, Asthma-Asthma, Parkinson’s disease-Parkinson’s disease,
Type 1 diabetes-Type 1 diabetes, Schizophrenia-Schizophrenia, Crohn’s
disease-Regional enteritis, Chronic kidney disease-Chronic kidney disease, Stage
I or II, Multiple sclerosis-Multiple sclerosis, Major depressive disorder-Major
depressive disorder, Cerebrovascular disease-Cerebrovascular disease, Atrial
fibrillation-Atrial fibrillation, Hypercholesterolemia-Hypercholesterolemia,
atherosclerosis-Coronary atherosclerosis, Hyperlipidemia-Hyperlipidemia,
Coronary Hypertrophic obstructive cardiomyopathy-Hypertrophic obstructive
cardiomyopathy, Myocardial infarction-Myocardial infarction, Systemic lupus
erythematosus- Systemic lupus erythematosus , Gout-Gout, Bipolar-Bipolar,
Psoriatic arthropathy-Psoriatic arthropathy, Epilepsy-Epilepsy,
Neurofibromatosis-Neurofibromatosis, Dementias-Dementias, Obesity-Obesity,
Obsessive-compulsive disorders-Obsessive-compulsive disorders, Autism- Autism,
Migraine- Migraine, Alzheimer’s disease- Alzheimer’s disease,
Coronary atherosclerosis-Coronary atherosclerosis, Posttraumatic stress
disorder-Posttraumatic stress disorder. Patients under 18 years old, with
ambiguous sex or with “unclassified” genetic ancestry were
excluded. Visualization was performed using the BPG R package v.7.1.0 156 .
PGS were calculated using array data after imputation in EUR
participants. Related participants based on their genetic similarity were
excluded (defined using PLINK v2.0a 109 with the relatedness coefficient --king-cutoff 0.05).
PGS were calculated using pgsc_calc 158 with the default settings and --min_overlap of 0.65.
Logistic regression was used to associate every PGS with the corresponding trait
based on phecodes (see Retrieving phenotype
data ). The following PGS model IDs from the PGS catalog 199 and their phecode pairs were
tested: PGS002250-Malignant neoplasm of ovary, PGS003766-Cancer of prostate,
PGS000004-Malignant neoplasm of female breast&Breast cancer,
PGS001794-Thyroid cancer, PGS000079-Melanomas of skin, PGS004884-Cancer of
bronchus; lung, PGS002264-Pancreatic cancer, PGS003395-Colorectal
cancer&Colon cancer, PGS000729-Type 2 diabetes, PGS002025-Type 1 diabetes,
PGS004254-Regional enteritis, PGS004699-Multiple sclerosis,
PGS000134-Schizophrenia, PGS004760-Major depressive disorder,
PGS001806-Malignant neoplasm of testis, PGS004687-Malignant neoplasm of bladder
& Cancer of bladder, PGS000039-Cerebrovascular disease, PGS004526-Essential
hypertension, PGS004706-Atrial fibrillation, PGS004784-Hypercholesterolemia,
PGS002029-Hyperlipidemia, PGS003726-Coronary atherosclerosis,
PGS000739-Hypertrophic obstructive cardiomyopathy, PGS004528-Myocardial
infarction, PGS000803-Systemic lupus erythematosus, PGS001789-Gout,
PGS002786-Bipolar, PGS000198-Psoriatic arthropathy. For prostate and testicular
cancer, only males were included, and for breast and ovarian cancer, only
females. An adjustment was made for age at diagnosis for cases and the latest
age for controls, genetic sex when both sexes were included, and the first ten
genetic PCs. For Figure 3 , the top and
bottom PGS deciles compared with the 5 th decile were considered,
testing only EUR participants. FDR was used for multiple testing correction. The
same process was repeated for other ancestries as presented in the supplementary material . Visualization was
generated using the BPG R package v.7.1.0 156 .
ATLAS array data were merged with genotyping data from the 1000
Genomes Project 38 , the
Simons Genome Diversity Project 130 , and the Human Genome Diversity Project 131 . BCFtools 128 annotate was used to
harmonize variant reference SNP ID (RSIDs), and BCFtools 128 norm with a GRCh38
genome reference was used to standardize the genotyping data. Sites or
individuals with more than 1% missing were removed using PLINK 133 --mind and --geno. Only
SNPs with MAF > 1% across all participants were kept.
SHAPEIT5 134 with
default parameters and the distributed GRCh38 map files were used to phase
genotyping data, one chromosome at a time.
A custom Python script that converts PLINK bed files to PLINK
ped/map 133 files
while conserving phasing information was used to convert genotyping data.
Centimorgan data for the map files were generated using the same genetic map
data in SHAPEIT5.
Identity-by-descent segments were called using iLASH 135 with the following
parameters: slice_size 350, step_size 350, perm_count 20, shingle_size 15,
shingle_overlap 0, bucket_count 5, max_thread 20, match_threshold 0.99,
interest_threshold 0.70, min_length 2.9, auto_slice 1, slice_length 2.9,
cm_overlap 1 and minhash_threshold 55. Identity-by-descent was called one
chromosome at a time.
Identity-by-descent segment outliers were removed as described in
Belbin et al. 53 and
Caggiano et al. 52 .
Segments overlapping centromeres or the human leukocyte antigen (HLA) region
were removed. Regions that may have false positive identity by descent were
identified using the following process and removed: total identity by
descent per each SNP was identified by summing across all
identity-by-descent segments that overlapped each SNP; SNPs with a total
identity by descent greater than or less than three standard deviations from
the genome-wide mean were removed.
To identify clusters, we followed the approach of Caggiano et
al. 52 and Dai et
al. 54 and applied
Louvain clustering 200 . An
undirected network is generated based on pairs of individuals who share
identity-by-descent segments: nodes are the individuals, edges are the
total, genome-wide identity-by-descent shared as the edges. We used the
Python package, NetworkIt 201 , to iteratively run Louvain clustering four times to
detect fine-scale clusters.
To avoid redundancy and maximize sample size, clusters were merged
in two stages to produce a final set of consensus fine-scale clusters.
First, following Caggiano et al. 52 and Dai et al. 54 , we computed pairwise
Hudson’s F ST using PLINK v2.0a 136 across 378 clusters identified from
the fourth layer of Louvain clustering. After removing clusters with fewer
than 10 participants, to avoid unreliable F ST estimates, we
merged the remaining 356 clusters into 67 clusters if the pairwise
F ST was less than 0.001.
In the second stage, we refined clusters using IBD sharing. For each
cluster pair, we examined all inter-cluster individual pairs to calculate:
(1) IBD mean , the average cM shared, and (2) IBD prop ,
the proportion of pairs sharing at least 3cM of IBD as detected by iLash. We
defined a composite metric for cluster merging, IBD weighted =
IBD mean ×IBD prop , that captures both the
extent and prevalence of genomic IBD sharing. Next, clusters were sorted
from the smallest to the largest number of ATLAS participants. For each
cluster, we computed the IBD weighted score with all larger
clusters and calculated the mean of these pairwise values. The cluster was
then merged into the most similar larger cluster –
i.e. , the one with the highest IBD weighted
score-only if that score exceeded the mean. Otherwise, the smaller cluster
was retained independently. This process merged 14 small clusters into
larger parent clusters, resulting in a total of 36 fine-scale clusters with
≥ 30 participants each for downstream analyses 52 . These fine-scale clusters were
assigned unique identifiers (IBD-01 through IBD-36) and manually annotated
with labels to ease interpretation. Due to filtering out clusters of small
sample sizes before and after merging, not all individuals were assigned to
a fine-scale ancestry cluster.
We primarily relied on reference individuals, described in
“ Fine-scale Ancestry
Pre-processing and Quality Control ”, to add labels to
clusters. When clusters did not contain reference individuals or were
heterogeneous, we utilized patient-reported race, ethnicity, preferred
language, and religious affiliation (in order of priority) to inform our
cluster labeling. These aspects are not caused by identity-by-descent
segment sharing but can be indicative of a shared culture for individuals
within a cluster; these shared practices can influence a group’s
demography, environment, and disease risk. Our labels are not definitive and
are our best attempt to generate informative assignments for each
cluster.
We used logistic regression to model the association between fine-scale
ancestry assignment and phecode prevalence, estimating OR with 95% confidence
intervals using the logistf R package v1.26.0 for Firth’s bias-reduced
penalized-likelihood logistic regression 157 . Differences were tested between every fine-scale
ancestry cluster with at least 100 participants (with no ambiguous genetic sex
and over the age of 18) and all other ATLAS participants. For phecodes that are
sex-specific, only participants of the corresponding sex were included in the
analysis. Adjustment was made for age at diagnosis for cases, and the latest age
for controls, and for sex when applicable. Phecodes with defined categories
according to the PheWAS catalog 61 , and that were diagnosed in at least 100 participants with
no ambiguous genetic sex and over the age of 18 in total in ATLAS were tested,
resulting in 1,253 phecodes. In situations where the number of cases in a
cluster was small, exact ORs were not reported in the text to protect patient
privacy. For visualization, fine-scale clusters were grouped into four panels
(EUR, Asian, AMR, and AFR) based on the predominant broad-scale ancestry of
participants. If the predominant ancestry was “Unclassified” (as
in the Japanese and Egyptian Christian clusters), the second most prevalent
broad-scale ancestry was used. The full phecode names presented in Figure 2c were: Hormones and synthetic
substitutes causing adverse effects in therapeutic use, Cardiomegaly, Allergic
reaction to food, Cirrhosis of liver without mention of alcohol, Vitamin
B-complex deficiencies, Anemia of chronic disease, Cholesterolosis of
gallbladder, Chronic renal failure [CKD], Dementias, Cancer of bladder,
Cataract, Leukemia, Hereditary hemochromatosis, Attention deficit hyperactivity
disorder, Glaucoma, Hyperplasia of prostate, Amyloidosis, Chronic pulmonary
heart disease. These names were shortened in the figure and text for easier
reading.
The same model was used to test differences in cardio-metabolic diseases
using appropriate phecodes, for each fine-scale cluster within the same
broad-scale continental ancestry. For this goal, clusters were assigned to
broad-scale ancestries based on the predominant ancestry match among
participants within each cluster. The largest population cluster within each
broad-scale ancestry was used as the reference level for associations, adjusting
for BMI, sex and age at diagnosis for cases, and the latest age for controls. As
a sensitivity analysis, we tested whether the observed patterns persisted when
correcting for SES factors (adjusting for ADI and BAS ranks) in addition to BMI,
age, and sex. As a second step, we also added IPW (see IPW analysis ).
FDR was used for multiple testing correction. The data were visualized
using the BPG R package v.7.1.0 156 .
The following studies were used to compare ATLAS’s broad-scale
genetic diversity with other large scale biobanks: Halldorsson et
al . (UKBB) 4 , Kurki
et al . (FinnGen) 6 , Feng et al . (Taiwan Biobank) 7 , Zawistowski et
al. (Michigan Genomics Initiative) 11 , Verma et al .
(Geisinger MyCode) 169 , The
All of Us Research Program Genomics Investigators et
al. 5 , Shaw
et al . (Vanderbilt’s BioVU) 170 , Verma et al . (Penn
Medicine Biobank) 119 , Verma
et al . (VA Million Veterans Program) 72 and Wiley et al . (the
Colorado biobank).
Studies that were used to compare sample sizes of ATLAS’s
fine-scale ancestry clusters with other published cohorts with available genetic
data ( Table S3 )
included: Belbin et al . (BioMe fine-scale IBD
clusters) 53 , Wu
et al . and Li et al . (Ashkenazi
Jewish) 57 , 202 , Haber et al . and
Hovhannisyan et al . (Armenian) 55 , 56 , Mehrjoo et al . (Iranian) 203 , Larena et
al . (Filipino) 58 ,
Sohail et al . and Ziyatdinov et al .
(Mexican) 204 , 205 .
IPW were calculated using logistic regression in R with ATLAS
participants as cases and other UCLA Health patients as controls. As described
above, we included UCLA Health patients who had at least one encounter within
one year of the ATLAS launch date. The following covariates were included:
self-reported race, sex, age group, BAS rank, and ADI rank. Individuals of
unknown race were excluded. For each ATLAS participant, the probabilities were
extracted from the model, and the weights were defined as 1 divided by the
predicted probabilities. To assess whether previously unreported ancestry and
disease status associations hold after applying IPW, we used the svyglm function
(R survey package v4.4.8 164 )
with a quasibinomial model, adjusting for age, sex, BAS rank, and ADI rank,
including the calculated weights for ancestry groups with more than 5 cases. To
compare comorbidity index values across broad-scale ancestries, considering IPW,
linear regression was used, adjusting for age, sex, BAS rank, and ADI rank,
including the calculated weights. Adjusted means and CI were obtained using the
R emmeans package, v1.10.5 132 . Patients under 18 years old, with ambiguous sex or with
“unclassified” genetic ancestry were excluded.
PheWAS were conducted using the Regenie v4.0 framework 66 for all five broad-scale
populations (AMR, AFR, EAS, EUR, SAS) and fifteen fine-scale clusters with at
least 400 participants (IBD-01 through IBD-15). Related individuals were removed
using a 0.05 kinship cutoff in PLINK v2.0a 136 . Sample sizes for each tested group can be found in
Table S4 .
Association testing was performed separately within each group
( i.e. , population or cluster) for 1,437 binary traits and
41 quantitative traits, including ICD-derived diagnoses and clinical laboratory
measurements (see Retrieving phenotype
data ). Samples were restricted to those in predefined inclusion lists
(--keep), and trait-specific covariates were provided via
--covarFile, including age, sex (modeled categorically with –catCovarList
sex), BMI, and the top 10 genetic PCs. Quantitative traits were divided into two
analysis groups based on missingness patterns, following Regenie’s
recommendation that traits with similar levels of missing data are modeled
together in Step 1.
For each group, unimputed array genotypes in bed format were
filtered using PLINK v2.00a 136 to produce a minimal set of variants suitable for
estimating genomewide polygenic effects through ridge regression in step 1
of Regenie. We included autosomal variants with call rate ≥99%
(--geno 0.01), minor allele frequency ≥1% (--maf 0.01), and
Hardy-Weinberg equilibrium p > 1×10 − 15
(--hwe 1e-15). Additional linkage disequilibrium (LD) pruning was applied
using a sliding window of 1000 SNPs, advanced by 100 SNPs, with an r2
threshold of 0.9 (--indep-pairwise 1000 100 0.9), producing an average of
389,209 SNPs per group. For binary traits, step 1 was run with a minimum
case count of 50 (--minCaseCount 50). For quantitative traits, step 1 was
run with Rank Inverse Normal Transformation (--apply-rint) enabled to
stabilize variance across lab values with differing distributions. For all
traits, a block size of 1000 (--bsize 1000) was used, leave-one out cross
validation (--loocv) was enabled, and sex was marked as a categorical
covariate (--catCovarList sex).
Prediction files from step 1 for each trait and imputed genotypes in
BGEN format were used to perform association testing in step 2 of Regenie.
For binary traits, a minimum allele count of 20 (--minMAC 20) and a minimum
case count of 50 (--minCaseCount 50) were required, and the Firth
approximation was performed (--firth --approx) using a p-value threshold of
0.01 (--pThresh 0.01). For quantitative traits, a minimum allele count of 20
(--minMAC 20) was required and Rank Inverse Normal Transformation
(--apply-rint) was applied. For all traits, a block size of 500 (--bsize
500) was used and sex was marked as a categorical covariate (--catCovarList
sex). Locus pruning was performed separately for each broad- and fine-scale
ancestry using PLINK 2.0a to identify unique variant-phenotype
associations.
To standardize phenotypes for comparison, phecodes for binary traits
were mapped to the Experimental Factor Ontology (EFO) using the text2term
v4.5.0 Python package. Each phecode was assigned up to three top-matching
EFO terms. To determine whether each unique variant–phenotype
association was previously unreported, we queried the EMBL-EBI GWAS Catalog
(June 27, 2025 freeze) for matching associations. An association from our
study was classified as “previously unreported in the EMBL-EBI GWAS
Catalog”, only if both of the following searches yielded no results.
(1) Associations between the lead variant and any phenotype in the same
phecode group as the associated phecode ( e.g. , for
colorectal cancer, we searched for associations with all EFO terms in the
broader “neoplasm” category). (2) Associations between the
nearest protein-coding gene (within 10kb) and any phenotype in the same
category as the associated phecode.
Genomic DNA libraries were created by enzymatically shearing high
molecular weight genomic DNA to a mean fragment size of 200 base pairs.
Multiplexity of exome capture and sequencing was achieved by adding unique
asymmetric 10-bp barcodes to the DNA fragments of single samples during library
amplification. Equal molar amounts of DNA samples were pooled for exome capture
using a slightly modified version probe library of xGen exome research panel
from Integrated DNA Technology (IDT). After PCR amplification and quantification
of the captured DNA, samples were multiplexed and loaded to Illumina sequencing
machines for sequencing to generate 75-base-pair paired-end reads. The samples
in this study were sequenced using the Illumina sequencing machines, including
NovaSeq 6000 with S2 or S4 flow cells and the NovaSeqX with 25B flow cells.
Sequencing reads in FASTQ format were generated from Illumina image data
using the bcl2fastq program (v2.20, Illumina). Whole-exome reads alignment and
germline small variant detection were conducted using the Original Quality
Functionally Equivalent (OQFE) protocol described in Krasheninina et al.,
2020 206 . Briefly, raw
read files (FASTQ) were mapped to the GRCh38 reference obtained from https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/technical/reference/GRCh38_reference_genome/
using BWA-MEM v0.7.17-r1188 in an alt-aware manner 139 . Duplicate reads were then marked with
Picard v2.21.2 127 . The final
CRAM files were compressed with SAMtools v1.2 128 . Germline variant detection was
performed on each CRAM using a Parabricks accelerated version of DeepVariant
v0.10.0 with a custom WES model 140 , resulting in a sample-level gVCF (genomic VCF).
Per-sample gVCFs were merged with GLnexus v1.4.3 141 into a joint-genotyped multi-sample
project-level VCF (pVCF). Variant prediction was restricted to the exome capture
region and the 100 base-pairs buffer on each side of the target regions. The WES
was of high quality and reached an average coverage of 41.6-fold with a minimum
of 24.2-fold, 90% having ≥ 34.7-fold in targeted regions, consistent with
a previous publication 207 .
The pVCF was converted to a PLINK file format using PLINK 1.9 133 for downstream analyses.
These data underwent extensive quality control to ensure the absence of
contamination, duplication, and other technical errors, as well as sufficient
read depth to guarantee reliability and accuracy. Genetic duplicates were
defined based on the aggregated genotype data of all sequenced samples. Sex was
predicted based on the ratio of read coverage on chromosome Y over the whole
exome read coverage. VerifyBamID v1.1.3 142 was used to estimate sample contamination. Samples
were excluded if they showed sex discordance, duplication, cross-individual
contamination > 5%, or coverage < 20X in more than 20% off target
regions. Overall, 386 samples were excluded for failing QC. Specifically, 249
samples were flagged for gender discordance, 12 had less than 80% of the exome
covered at 20X, 80 showed contamination levels above 5%, and 74 were unresolved
duplicates. After accounting for 29 samples that appeared in more than one
exclusion category, the final number of unique samples excluded due to QC
failure was 386. Alignment quality metrics were generated using Picard
CollectHsMetrics v2.27.4 127
and SAMtools stats v1.15.1 with default settings. MultiQC v1.27.1 143 was used to systematically
aggregate sample-level quality metrics. RTGtools v 3.12.1 129 was used to assess variant QC metrics,
including the number of variants for SNPs, small insertions and deletions;
genotype counts; heterozygous-to-homozygous ratios for each variant type and
transition/transversion (Ti/Tv) ratio. Genotypes were further confirmed using an
alternative approach 208 with
500 random samples, yielding a high concordance for on-target sites (median
concordance of 99.5% for SNPs and indels).
For WES-based analyses, we used samples that also had array genotyping
data (60,025 out of 61,797 patients with WES), which was necessary to
consistently define broad- and fine-scale genetic ancestry. We annotated
variants using Ensembl Variant Effect Predictor (VEP) v.112 137 for 58,387 samples assigned to the EUR,
AFR, SAS, EAS, or AMR ancestry class ( i.e. , excluding
unclassified genetic ancestry samples). Annotations were performed using the
GRCh38 cache and the corresponding reference FASTA file. We limited our analysis
to autosomal chromosomes to avoid technical artifacts in variant calling caused
by the differences in ploidy between males and females, as well as the
high-sequence similarity between the X and Y chromosomes in certain
regions 209 . Counts
for single-nucleotide variants, indels, multi-allelic, synonymous, missense, and
LOF variants were restricted to whole-exome sequencing (WES)-targeted regions.
Consistent with a previous UKBB WES study 210 , we classified LOF variants as those with the
following consequences: stop_gained, start_lost, splice_donor, splice_acceptor,
stop_lost, and frameshift. To increase reliability, we further restricted LOF
variants to those flagged as high confidence by LOFTEE 94 . For multi-allelic variants, the
predicted function for each alternate allele was determined using the
--pick-allele option based on the default ordered set of criteria defined by
VEP 137 . To determine
the number of variants with MAF < 1% while accounting for
ancestry-specific allele frequency differences, we retained variants with MAF
< 1% in at least one ancestry group.
For multi-allelic variants, we defined the minor allele as the second
most common allele (including the reference) and classified the variant as rare
if the cumulative allele frequency of all alternate alleles was < 1%. For
rare multi-allelic variants, we incremented a sample’s count only if it
carried an alternate allele with MAF < 1%. In multi-allelic cases where
the minor allele was the reference or the alternate alleles had AF ≥ 1%,
we incremented a sample’s count if it carried any alternate allele. For
rare functional variants ( e.g. , missense), we incremented a
sample’s count only if it carried an alternate allele with MAF <
1% corresponding to that functional category.
We identified variants via a literature search, where
variants are known to be enriched within certain populations. For this goal, we
selected: 1) Familial Mediterranean Fever (FMF) related variants in the MEFV
gene (V726A, M694V, M694I, M680I and E148Q) 211 2) the HBB:p.E7V variant, and 3)
PCSK9 P/LP variants 195 .
For FMF and PCSK9 , carriers were identified as
participants who carry at least one corresponding pathogenic variant within
PCSK9 or FMF variant group. For HBB:p.E7V, the frequency
was estimated based on this variant alone. We only considered unrelated
participants, using a kinship coefficient of 0.05 based on PLINK v2.0a 136 king-cutoff. Fisher’s
Exact test was used to test the carrier frequency of each cluster against the
carrier frequency of the other clusters grouped together. The risk of
amyloidosis (a common symptom of FMF) by carrier status was calculated
considering carriers of known FMF and patients assigned the amyloidosis phecode
(270.33).
To test differences in frequencies of pharmacogenomic variants, 121
CPIC Level 1A variants were pulled from ClinPGx 86 , 87 . An overview of the analyzed variants can be found in Table S4 . As variants
spanned both common to rare frequencies, we used exome data for consistency.
Overall, the variants were highly abundant in ATLAS, with 95.5% of all unrelated
individuals carrying at least one. Of these variants, 69 were identified in the
UCLA ATLAS exomes. To test enrichment of single alleles across clusters,
“carriers” were identified as participants who carry at least one
allele of a variant within a gene associated with a pharmacogenomic drug
interaction or monogenic condition (unless only a single variant was involved).
We then calculated carrier frequency as the number of identified carriers over
the total number of participants with WES available and were unrelated using a
kinship coefficient of 0.05 based on PLINK v2.0a 136
king-cutoff . To identify clusters with a significantly
different carrier frequency of variants, we applied a Fisher’s Exact test
to the carrier frequency of each cluster against the carrier frequency of the
other clusters grouped together. We then used FDR corrections for p-values
generated from all Fisher’s exact tests and selected significant clusters
(FDR ≤ 0.05) where at least 5 carriers were identified.
To identify clusters where participants were enriched for carriers of
monogenic ClinGen variants associated with disease, we used pathogenic variants
that were curated by experts within the field and underwent stringent review to
be considered pathogenic from ClinGen 91 . ClinGen variants filtered to identify P/LP variants
with autosomal dominant inheritance, autosomal recessive inheritance, or
semidominant inheritance, resulting in 3,521 variants associated with 204
conditions overall according to ClinGen. We removed one HNF4A
ClinGen variant due to a high allele frequency (MAF>0.01). Enrichment for
variants, aggregated per gene, was tested as described above for ClinPGx
variants with WES data. PAR was calculated for each cluster, showing significant
enrichment of ClinGen P/LP variant frequencies aggregated by gene, focusing on
genes associated with monogenic diseases with autosomal dominant or semidominant
inheritance. To ensure sufficient power, we included only clusters with at least
30 individuals diagnosed with the disease with WES data available. We used the
following formula: PAR = (p(RR-1))/(p(RR-1)+1), where p is the carrier frequency
and RR is the ratio of the proportion of carriers with disease to the proportion
of noncarriers with disease. For familial hypercholesterolemia, cases were
defined as individuals with LDL levels exceeding 190 mg/dL, adjusted for statin
use. In case of statin use, the LDL levels were divided by 0.7, as previously
done 195 , 212 .
To calculate the PAR for breast cancer associated with the three
Ashkenazi Jewish founder alleles, we used the formula above. Because individuals
in this group are often aware of their genetic status and may pursue preventive
breast cancer measures, we excluded individuals who underwent surgery (using the
‘acquired absence of breast’ phecode) but were not diagnosed with
breast cancer, to improve the reliability of the estimate.
The list of all ClinGen 91 P/LP variants identified in any ACMG secondary findings
(SF) v3.2 genes 93 was
extracted as described above. Variants were labeled as missense or LOF, based on
VEP v.112 137 annotations.
Missense variants were defined for ‘missense_variant’ variants
according to the ‘consequence’ VEP output column. LOF was defined
for high-confidence “HC” LOF variants based on LOFTEE 94 . All variants were rare across
broad-scale ancestries. A WES plink file with all participants was filtered to
include only the listed variants. Then, this file was broken into ancestry
groups ( i.e. , broad-scale population, or fine-scale cluster)
using PLINK v2.0a 136 . Related
participants were removed using a 0.05 kinship cutoff in PLINK v2.0a 136 . The frequency of each
allele in every population group was calculated with the PLINK v2.0a --freq
command, and total frequencies were summed up for rare missense and LOF variants
separately. To test differences in the allele counts across populations, allele
dosages were calculated with PLINK v2.0a 136 --recode A option. Dosages were summed to calculate
the total alternative (ALT) allele count within each population. The number of
reference (REF) alleles was defined as twice the number of individuals in the
group minus the number of ALT alleles. Fisher’s Exact tests were applied
to test differences between the number of REF and ALT alleles in each population
compared to all other participants not assigned to that specific group,
considering only groups with over 400 participants. Bonferroni correction was
applied for multiple testing correction. Plotting was done using the BPG R
package v.7.1.0 156 .
Variants from exome sequencing were annotated using VEP v.112 137 with the dbNSFP
v4.9a 138 and
LOFTEE 94 plugins
installed. The LOFTEE high-confidence “HC” flag was used to select
for LOF variants predicted to have deleterious effects.
Missense variants were assigned a 9-point deleteriousness score based
on a consensus of nine missense deleteriousness prediction toolkits, similar to
methods described in prior biobank-scale rare variant studies. We used the
Critical Assessment of Genome Interpretation (CAGI) project 95 to prioritize well-performing tools not
trained on the same features. We selected five meta-predictors –
ClinPred 144 ,
MetaRNN 145 ,
BayesDel_addAF 146 ,
VARITY_R 147 ,
REVEL 148 – and
four stand-alone predictors – AlphaMissense 149 , MutPred2 150 , VEST4 151 , ESM-1b 152 – for use in our analysis. We
assigned each variant a binary score per tool based on dbNSFP rank scores: 1 if
the variant’s score exceeded the threshold score for being more likely a
deleterious ClinGen or ClinVar variant than a background variant ( Figure S7g – h ), and 0 otherwise.
Summing these binary scores produced a deleteriousness score ranging from 0 to
9, with predicted damaging missense variants scoring ≥ 5 retained for
downstream analysis.
A WES PLINK file was filtered to include computationally predicted LOF
and predicted damaging missense variants (see above) in ACMG SF v3.2
genes 93 , providing a
more comprehensive evaluation than the one based solely on ClinGen P/LP
variants, which included only 17 genes. A WES plink file with all participants
was filtered to include only the ACMG putative damaging variants, and the file
was split into ancestry groups (broad- and fine-scale ancestries) using PLINK
v2.0a 136 . Related
participants were removed using a 0.05 kinship cutoff in PLINK v2.0a 136 . Allele frequencies were
calculated with PLINK v2.0a, and only rare variants (MAF < 1% in all
broad-scale ancestries) were kept. Allele dosages per individual were extracted
using PLINK v2.0a, and the total rare LOF and predicted damaging missense
alleles were counted separately per participant. First, the differences between
the total numbers of REF and ALT alleles across ancestries were evaluated with a
Fisher’s Exact test as described above (ClinGen ACMG analysis). Second, a
Mann-Whitney U test with a Bonferroni correction was applied to test the
difference in the distribution of rare LOF and predicted damaging missense
counts per individual between any broad- or fine-scale group compared to all
others, with groups that included at least 100 participants. Figure S7I presents the
Mann-Whitney U test results, with statistically significant differences
indicated by an asterisk (*). Visualization was done with BPG 156 .
ExWAS were conducted using the Regenie v4.0 framework 66 for all five broad-scale
populations (AMR, AFR, EAS, EUR, SAS) and fifteen fine-scale clusters with at
least 400 participants (IBD-01 through IBD-15). Related individuals were removed
using a 0.05 kinship cutoff in PLINK v2.0a 136 . Sample sizes for each tested group can be found in
Table S6 . Within
each group, imputed genotype dosages in BGEN format (--bgen) and prediction
scores from PheWAS step 1 (--pred) were used to conduct gene-based association
testing for binary and quantitative traits across 17,537 genes. Samples were
restricted to those in predefined inclusion lists (--keep), and trait-specific
covariates were provided via --covarFile, including age, sex
(modeled categorically with --catCovarList sex), BMI, and the top 10 genetic
principal components.
Whole-exome genotypes in bed format were filtered using PLINK v2.0a,
retaining variants with a call rate ≥ 90% (--geno 0.1), a minor allele
count ≥ 1 (--mac 1), and a Hardy-Weinberg equilibrium p-value >
1×10 −15 (--hwe 1e-15). In addition, 784,548
variants overlapping low-complexity regions (LCRs) were excluded prior to
analysis. Briefly, SNP positions were extracted from the .bim file, intersected
using bedtools v2.29.1 153
with annotated LCRs from the Genome in a Bottle Consortium 213 , and filtered from the genotype files,
resulting in 12,326,160 exome variants retained for burden testing.
Filtered whole-exome genotypes in BGEN format and prediction files from
PheWAS step 1 were used to perform gene-based association testing. Regenie-style
annotation, set, and mask files for predicted deleterious LOF and missense
variants (see Variant annotation using a
consensus of computational tools ) were created programmatically using
Python v3.11.9 with the polars v1.2.1 package. For binary traits (--bt), a
minimum case count of 50 (--minCaseCount 50), a minimum minor allele count of 5
(--minMAC 5), and a block size of 1000 (--bsize 1000) were enforced. Firth
logistic regression with saddlepoint approximation (--firth --approx) was
applied for variants with p < 0.01 (--pThresh 0.01). For quantitative
traits (--qt), rank inverse normal transformation (--apply-rint), a minimum
minor allele count of 5 (--minMAC 5), and a block size of 500 (--bsize 500) were
enforced. Regenie’s implementation of the RGC gene-based p-value test
(–rgc-gene-p) was enabled for all analyses. Additionally, variants were
binned by minor allele frequency using 1% bins (--aaf-bins 0.01), and SNP-level
membership for each burden mask was recorded (--write-mask-snplist).
In addition to gene-based analyses, single variant association testing
was performed for selected predicted deleterious LOF and missense variants using
filtered whole-exome genotype dosages in BGEN format and prediction scores from
PheWAS step 1. For both binary and quantitative traits, tests incorporated the
same set of covariates (age, sex modeled categorically, BMI, and the top 10
genetic principal components) and enforced a minimum minor allele count of 5
(--minMAC 5). For binary traits, Firth logistic regression with saddlepoint
approximation was applied for variants with p < 0.01, along with a block
size of 1000, while quantitative traits were evaluated using rank inverse normal
transformed phenotypes with a block size of 500. A Bonferroni-adjusted p
< 0.05 was defined as the cutoff for significance.
Replication analyses were conducted utilizing the Python Hail
package v0.2.134 165 on
AoU Workbench using the Controlled Tier Dataset v8. Phenotypes were obtained
by querying OMOP databases for per-patient ICD codes, which were
subsequently converted to phecodes v1.2 utilizing the R PheWAS package
v0.99.6 176 .
Participants with both phecodes and single-read WGS data were filtered to
remove participants flagged from genotype QC or from relatedness QC,
resulting in 291,082 total participants. Logistic regression was performed
within the given replication genetic ancestry cohort by testing case/control
status for a given phenotype and SNP, utilizing age, sex, age^2, sex*age,
sex*age^2, and genotype PCs 1-10 as covariates.
Replications for variant-trait associations were obtained by
querying the Controlled Tier Dataset v8 “All by All” allele
count/allele frequency (ACAF) variant result tables using the Python Hail
package v0.2.134 on AoU Workbench.
Replications for variant-trait associations were obtained by
querying the summary statistics portal at https://taiwanview.twbiobank.org.tw/pheweb.php 7 .
Replications for gene-trait associations were obtained by querying
the Controlled Tier Dataset v8 “All by All” rare variant
result tables using the Python Hail package v0.2.134 165 on AoU Workbench.
Replications for gene-trait associations were obtained by querying
the AstraZeneca summary statistics portal at http://azphewas.com/ 69 .
The BioMe Biobank consists of electronic health records and genetic
data from approximately 60,000 participants from the Mount Sinai Health
System in New York. Participant recruitment was between 2007 and 2023. This
study was approved by the Icahn School of Medicine at Mount Sinai’s
Institutional Review Board (Institutional Review Board 07–0529). All
study participants provided written informed consent.
BioMe participants were genotyped using the Illumina Infinium
Global Diversity Array (GDA; number of participants, N=23,430; number of
variants, n=1,833,111) or Infinium Global Screening Array (GSA; N=32,595;
n=635,623). Quality control consisted of removing participants with a call
rate <95%, a mismatch between self-reported and genetic sex, and high
heterozygosity. Duplicated sites and sites with a genotyping rate of
<95% were removed. QC was done with PLINK2. Data was imputed with the
TOPMED imputation server 214 with genome build hg38. Genetic PCs were calculated
across participants using PLINK2 for the genotyping sites after LD pruning.
Exome sequencing data were generated by the Regeneron Genetic
Center 215 .
Quality control consisted of using the Goldilocks Filter 210 , and variants with
quality scores <3 or depth of coverage scores <7 for SNPs, or
<5 or depth of coverage <10 for indels were removed.
Monomorphic sites were removed. Genetic ancestry was assigned using a random
forest classifier trained on principal components derived from the 1000
Genomes data to align with the genetic ancestry assignment in AoU 5 . Participants were assigned
a genetic ancestry using 10 PCs.
Bio Me laboratory data were processed according to
the QualityLab pipeline 216 . Briefly, quantitative lab values obtained in
different clinical contexts (ambulatory, emergency, inpatient, and urgent
care), were cleaned. Only labs with at least 100 patients and at least 1000
numeric observations were considered. Only adult (> 18 years) lab
values were considered. Per lab, outliers were removed, defined as values as
greater or less than 4 standard deviations from the mean of that lab.
Non-numeric labs were removed. Labs must have had >70% of the
reported units matching. After cleaning, a median value and median age was
calculated per person for each lab test in each clinical context and across
all contexts. Laboratory phenotypes were manually assigned a match to UKBB
phenotypes using available metadata.
Electronic health phenotype data in the form of ICD-10 codes were
collapsed into PhecodeX phecodes 217 . Phecodes were transformed into a binary matrix,
where each row was an individual and each column was a phecode. A
participant had a 1 if they ever were diagnosed with that phecode, otherwise
their value was set to 0. Age was calculated as current age, defined from
01/01/2025.
Association testing was performed using Regenie v4.0 66 . Burden testing was
performed using SKATO and the following masks: missense, missense_pLoF,
pLoF, and synonymous. PLoF variants were defined using LOFTEE 94 .
ATLAS GLP1-RAs prescriptions, including medication names, start and
end dates, discrete dose, usage instruction, strength and route (oral or
subcutaneous), were retrieved and grouped based on the following simple
generic names: dulaglutide, semaglutide, liraglutide, exenatide, albiglutide
and lixisenatide. For all statistical analyses, only semaglutide users were
considered. Usage start dates were defined based on the earliest
prescription start date for each patient. In the case of a missing start
date, the prescription ordering date was used instead (in most cases, these
two fields were identical). Overlapping medication periods were handled such
that when a new prescription started, the previous one was considered to
have ended. Similarly, missing end dates were determined using the start
date of the next prescription when available. In the case of completely
overlapping prescriptions with different instructions, the combination of
both routes and the weighted average dose was considered. If the discrete
dose information was missing, the medication dose was extracted from the
instructions’ free text field. If the instructions also omitted the
dose information, it was imputed for a given medication type, based on the
weighted dose average from the entire cohort on the corresponding type.
Initial weight and BMI were defined as the median of all available weight or
BMI measurements recorded from in-person visits, within 6 months prior to
the first prescription start date. Utilizing the median rather than a single
data point helped minimize the likelihood of recording errors in the EHR.
The percentage of weight change was obtained for every weight measurement
recorded on in-person visits within the period of active prescriptions
between 4-60 weeks in total on semaglutide. The medication dose at each time
point was defined as the weighted sum of all medication doses (doses
multiplied by the number of prescription weeks) by the measurement date. In
case both oral and subcutaneous medications were used within a period, a
combined route category was defined.
To identify the overall weight loss patterns across time, the FPCA
R function from the fdapace package v.0.6.0 159 , 160 , which is suited to plot smoothed longitudinal
data with repeated measurement, was used. This identified a consistent
weight loss pattern up to 60 weeks, and sparse data points beyond
~150 weeks ( Figure
S8C ). Thus, we restricted all analyses to this period.
For all association tests, only participants aged 18 years or older
were included. For all analysis parts that involved longitudinal data with
repeated measurements, a linear mixed model with the bobyqa optimizer and an
increased function evaluation limit (maxfun = 10000) was used (lmer R
function; the lmerTest R package v.3.1.3 161 ). In each case, ANOVA was used to
identify differences between models to define the best way to model a
potential nonlinear relationship between weeks and weight loss, when
treating weeks as a fixed effect, based on the Akaike Information Criterion
(AIC). In some cases, the best model was achieved using a restricted cubic
spline (RCS) with the rcs R function from the rms package (v.7.0.0), applied
to weeks. In other cases, a polynomial function of weeks provided a better
fit. To plot smoothed longitudinal data with 95% CI, the fitted model values
(excluding covariates) and 95% lower and upper confidence bounds, were
extracted using the visreg package 162 v.2.7.0 and visualized using the BPG
package 156 .
For testing the effect of baseline factors, fixed effects were
defined for the medication dose, route, sex, age, initial BMI, first ten
genetic PCs (excluding PC6 due to a strong collinearity with PC5 that
disrupted the model convergence) and weeks on semaglutide. The model
included both random intercepts and slopes for weeks on semaglutide.
Bonferroni correction was applied to control for multiple testing.
For testing differences across ancestries, fixed effects were
defined for the medication dose, route, sex, age, initial weight, a
polynomial function of weeks on semaglutide, and an interaction between a
polynomial function of weeks and genetic ancestry categorical class (EUR,
AFR, EAS, SAS and AMR). The overall number of weight measurements was
24,145, with a mean of 5 repeated measurements per patient. Ancestry sample
sizes were: European (EUR), 3189; African (AFR), 373; admixed American
(AMR), 913; East Asian (EAS), 291; South Asian (SAS), 107. The model
included both random intercepts and slopes for the polynomial function of
weeks on semaglutide. An ANOVA was performed on the model with a Bonferroni
correction to assess the global effects of ancestry and
ancestry×weeks interaction. When a significant result was found, post
hoc comparisons between groups were conducted using the summary lmerTest
function (v.3.1.3 161 ),
applying a Bonferroni correction to all 12 class or class × time
interactions.
To test the relationship between PGS and weight loss, related
participants based on their genetic similarity were excluded (defined using
PLINK v2.0a 136 with the
--king-cutoff 0.05). Scaled BMI (PGS000027) and DM2 PGS (PGS000729) in EUR
participants were divided into three equal bins each: low, intermediate and
high scores. Then, using longitudinal data, for each trait, fixed effects
were defined for the categorical PGS bins, medication dose, route, sex, age,
initial weight, first ten genetic PCs (excluding PC6 due to a strong
collinearity with PC5 that disrupted the model convergence) and a polynomial
function of weeks. Random intercepts were defined to account for repeated
measures within individuals. Bonferroni correction was applied to control
for multiple testing. To test the relationship between PGS and weight loss,
relying on a simplified model where the maximum weight loss was considered
for each EUR participant, linear regression was applied. The model was
adjusted for the medication dose and route, 10 genetic PCs, age, sex, and
initial weight. Visualization was made with BPG 156 , with smoothed data and 95% CI
using loess.as R function (fANCOVA package v.0.6.1 163 ).
The analysis was performed to test a relationship between the
maximum weight loss on semaglutide and common genetic variants, for each
broad-scale ancestry using SAIGE 167 . For step one, array observed variants were used
following PLINK v2.0a 127
filtering, with the flags: --maf 0.01 --mind 0.1 --geno 0.1 --hwe 1e-6 . For
the second step, array observed and imputed variants were used following
plink filtering with --maf 0.01 --geno 0.05 --hwe 1e-6. A quantitative
analysis was conducted with the traitType flag. The medication dose and
route, the first five genetic PCs, age, sex, and initial weight were used as
covariates. METAL 168 was
used for meta-analysis.
To identify genes genetically associated with weight loss, we used
Regenie 66 with an
additive model for gene-level tests. Only EUR semaglutide users that are not
related (relatedness coefficient --king-cutoff 0.05) with WES data were used
(n = 2,012), and for each, the maximum weight loss record with the
corresponding number of weeks was considered, using information on aggregate
dose and route. The list of candidate genes was limited to
Bonferroni-significant proteins whose plasma abundance was altered by
semaglutide treatment 115 .
We considered variants within these genes with a predicted moderate or high
impact on the protein function, according to VEP v.1.2 137 . For step one, the variant list was
limited to observed array SNPs following a PLINK v2.0a 136 filtering with: --maf 0.01 --mac
100 --indep-pairwise 1000 100 0.9 --chr 1-22 --snps-only --geno 0.1 --hwe
1e-15. In step 2, we applied PLINK v2.0a filtering to variants across all
EUR participants using the parameters --geno 0.05 and --hwe 1e-6.
Subsequently, the file was filtered to include only semaglutide users and
was used in Step 2. The medication dose and route, the first 10 genetic PCs,
age, sex, and initial weight were used as covariates. Bonferroni correction
was applied to control for multiple testing. Frequency of variants involved
in PTPRU association with weight loss on semaglutide by
ancestry can be found in the table below.
For replication analyses in AoU the same processes and tools were
conducted in EUR individuals. Controlled Tier Dataset v8 was accessed
through the Researcher Workbench and imported in PLINK format, including
exome variants in chromosome 1 (to extract PTPRU variants)
and common ACAF variant datasets (for PGS calculation). Genetic ancestry and
PCs derived from pre-defined assignments provided by AoU 5 . PTPRU VEP 137 variant annotations were
obtained using Python Hail v0.2.134 165 to identify functional consequences. For the
Regenie analysis, related individuals were not removed to maximise power, as
Regenie is designed to account for relatedness. The combined P-value of
ATLAS and AoU was calculated using Fisher’s combined probability
test, implemented with the fisher function in the poolr R package
v1.2.0 166 .
Multiple testing was addressed using a Bonferroni correction that accounted
for the number of gene-level tests in the discovery cohort, along with one
replication P-value and one combined P-value. Frequency of variants involved in
PTPRU association with weight loss on
semaglutide by ancestry. Variant AFR AMR EAS EUR SAS 1:29236667:T:C 0 0 0 1.2×10 −5 0 1:29258717:G:A 0.00018 0 0 0.00013 0 1:29259276:C:T 0 0 0 1.2×10 −5 0 1:29259313:A:G 0 0 0 4.9×10 −5 0 1:29259883:C:G 0 0 0 3.7×10 −5 0 1:29259912:A:C 0 0 0 1.2×10 −5 0 1:29260003:G:C 0 0 0 1.2×10 −5 0 1:29260895:A:C 0.00055 0.0025 0 0.0027 0 1:29275496:G:A 0 0 0 7.3×10 −5 0 1:29275715:G:T 0.0018 0.011 0.0018 0.0092 0.023 1:29279056:G:A 0 0 0 1.2×10 −5 0 1:29279527:G:T 0 6.0×10 −5 0 1.2×10 −5 0 1:29279541:C:A 0 0 0 1.2×10 −5 0 1:29282705:G:A 0.0013 0.0028 0 0.0050 0.0018 1:29282710:C:T 0 0 0 8.5×10 −5 0 1:29284754:C:T 0 0 0 9.7×10 −5 0 1:29291897:G:A 0 0 0 0.00023 0.00045 1:29291929:G:A 0 0.00030 0 2.4×10 −5 0 1:29291942:C:T 0 0.00042 0.00020 0.00052 0 1:29291958:A:G 0.0058 0.00030 0 2.4×10 −5 0 1:29291978:G:A 0 0 0 0.00015 0 1:29303853:A:G 0 0.00036 0 0.00024 0 1:29303863:C:T 0 0.00012 0 0.00032 0 1:29303875:G:T 0.094 0.0065 9.9 × 10–5 0.00078 0.00045 1:29304795:G:A 0 0 0 1.2×10 −5 0 1:29305397:A:G 0.23 0.38 0.48 0.25 0.38 1:29311488:C:T 0.00018 0 0 4.9×10 −5 0.00045 1:29311513:C:A 0 0 0 3.7×10 −5 0 1:29311706:G:C 0 0 0 1.2×10 −5 0 1:29312579:C:T 0 6.0×10 −5 0 9.7×10 −5 0 1:29312615:G:A 0 0.0011 0 0.00015 0 1:29315374:C:T 0 5.99×10 −5 0 2.4×10 −5 0.00045 1:29315448:G:A 0 0 0 1.2×10 −5 0 1:29317767:C:T 0 0.00036 0 0.0014 0 1:29317838:G:A 0 0 0 0.00016 0 1:29320709:A:G 0 0 0 6.1×10 −5 0 1:29323643:C:T 0 0 0.00030 1.2×10 −5 0
Frequency of variants involved in
PTPRU association with weight loss on
semaglutide by ancestry.
Statistical analyses were performed using R, PLINK, Regenie, and
related tools. Logistic regression, linear regression, Firth-penalized
regression, linear mixed models, and Fisher’s exact tests were applied in
the relevant analyses, adjusting for appropriate covariates. Multiple-testing
correction was applied using the FDR or Bonferroni adjustment, as specified in
the relevant analyses. Detailed statistical models and software parameters are
provided in the METHOD DETAILS
section.
An interactive web portal enables users to explore the PheWAS results,
examine fine-scale ancestry associations with clinical phenotypes, and download
supplemental data at atlas-phewas.mednet.ucla.edu .
Results
UCLA Health serves Los Angeles County, one of the world’s most
ancestrally diverse metropolitan areas, with a population of 9.6 million. To
date, the UCLA ATLAS initiative has consented ~250,000 UCLA Health
patients and has collected biomaterials from ~130,000. ATLAS enrollment
largely reflects the composition of UCLA Health patients, who are concentrated
across Los Angeles ( Figure 1b ). The EHR,
initiated in 2013, enables continuous longitudinal stratification of
participants by disease states, with a mean and median of 8.6 and 7.9 years of
participation per individual. As of November 2024, ATLAS genomic data included
WES data from 61,797 participants and custom array genotyping using the Illumina
Global Screening Array from 92,164 participants. Extensive quality control (QC)
indicated high data quality ( Figure S1 ). We leveraged these data to interrogate social and
genetic factors that affect disease risk and health outcomes ( Figure 1a ).
Genotyped cohort demographics are summarized in Table 1 and the Methods . Biobank
participants were older in age and had a higher comorbidity index than
non-biobank patients (2.9 versus 1.7 mean Elixhauser index;
1-year post-collection), consistent with a higher number of clinical visits,
which favors enrollment (controlled for data-completeness, Methods ;
Figure 1c ). EHR-based phenotypes,
including vital signs, disease diagnoses, and lab tests, were defined through
harmonization procedures and subsequently validated ( Methods ; Figure S2a – l ). Common disease
domains were led by endocrine/metabolic, cardiovascular, gastrointestinal
disorders and neoplasms, reflecting known global health challenges ( Figure 1d ; Figure S2m – p ; Methods ) 31 – 33 . The ATLAS EHR data includes a total of
71,739,582 lab tests (considering complete blood count [CBC], lipid, metabolic,
hemoglobin A1c [HbA1c], and 25-hydroxyvitamin D panels), with a mean of 540 lab
test results per participant ( Table S1 ). Detailed prescription data show a total of 5,952,958
prescriptions; the 50 most frequently prescribed medications are listed in Table S1 .
Although we primarily consider genetically determined ancestry in our
analyses, rather than the social constructs of race and ethnicity 34 , 35 , we note substantial diversity based on participant
self-reports. For instance, 62.7% of participants self-identify as White, 12% as
Asian, 4.6% as Black or African American, 2.6% as Middle Eastern or North
African, 0.9% as American Indian or Alaska Native and 0.3% as Pacific Islander.
A substantial proportion (14.8%) report Hispanic or Latino ethnicity ( Table 1 ).
Self-identified race/ancestry, a cultural-societal construct, and
genetic ancestry are conceptually distinct 29 , 34 , 36 and genetic ancestry must be considered
to prevent confounding in genetic association studies 35 , 37 . We classified biobank participants into six broad-scale
ancestry populations: European (EUR), African (AFR), South Asian (SAS), East
Asian (EAS), Admixed American (AMR) and an Unclassifiable (UNC) population,
aligning with the 1000 Genomes Project super-populations 38 ( Figure
1e – g ;
Methods ). ATLAS is diverse relative to most other large
biobanks ( Methods ). Consistent with self-reported race/ethnicity,
almost a third of the biobank participants (32%) were assigned to non-EUR
genetic ancestries. There was significant agreement between self-reported race
and broad-scale ancestry – 99% of self-identified White participants were
assigned to EUR or AMR populations ( Figure S3a ). Of those with
“unknown” self-reported race, 46% were assigned to AMR ancestry
and 45% to EUR.
We next asked how health system usage or disease burden varied by
broad-scale ancestry, observing that the mean number of encounters varied
significantly across ancestries (P-value = 7.5×10 −138 ,
ANCOVA, adjusted for genetic sex, age and Barriers to Accessing Services
[BAS] 39 ), with the
highest numbers of total encounters in participants from AMR ancestry (adjusted
mean = 16.2), followed by AFR (15.8) and SAS (14.2), with the lowest values in
EUR (12.2) and EAS (12.9) ( Figure 1h ). We
also assessed health burden using the comorbidity index score 40 , 41 , which was higher in EAS (adjusted mean = 3.7) and AMR
(3.6) compared to others (all other populations 2.6-2.9; Figure 1i ; P-value =
2.6×10 −5 , ANCOVA, adjusted for genetic sex, age, BAS
and Area Deprivation Index [ADI] 42 ). Similar results for disease distribution across
ancestries were obtained after applying inverse probability weighting (IPW) to
adjust for participation bias relative to the broader UCLA Health population
( Figure S3b ;
Methods ). Increased encounter numbers were only partially
explained by elevated comorbidity index scores, indicated by modest correlations
between the two (mean total encounters vs . Elixhauser
Comorbidity Index: r = 0.32, P-value <2.2×10 −16 ;
mean hospital encounters vs . Elixhauser Comorbidity Index: r =
0.26, P-value <2.2×10 −16 ; Figure S3c – d ).
The population diversity of ATLAS enables the interrogation of the
combined genetic and social/environmental effects on the risk of disease
diagnoses. To illustrate this, we assessed variation in medical conditions
across populations, replicating multiple known associations
( Methods ; Table S2 ; Figure
S3e ). Further, we identified previously unreported associations
( Figure 1j ), including a lower risk of
epilepsy in EAS compared to EUR participants (odds ratio [OR] = 0.46 with 95%
confidence interval [0.32, 0.65], P Bonferroni =
2.3×10 −4 ), not previously detected 43 , 44 . Similarly, we found significantly reduced risk for
bipolar disorder in those with AMR ancestry, clarifying conflicting findings in
previous smaller studies 45 , 46 (OR = 0.47 [0.40, 0.56],
P Bonferroni = 1.5×10 −17 ). We also show a
significantly lower risk of sleep apnea in EAS participants relative to EUR (OR
= 0.67 [0.59, 0.75], PBonferroni = 4.1×10 −10 ), which
has been controversial 47 – 50 , with
a weaker, but significant effect after correcting for body mass index (BMI) (OR
= 0.84 [0.74, 0.95], P-value = 4.6×10 −3 ; Table S2 ). These
associations remained significant after adjusting for socioeconomic status
(SES), using both the BAS and ADI measures, and with IPW ( Table S2 ).
Fine-scale ancestries 37 , 51 – 53 , many of which are
understudied 14 – 16 , reflect recent geographic or
demographic stratification and can reveal important contributions to health
disparities and disease risk 14 . We quantified fine-scale ancestry using identity-by-descent
(IBD) 53 , 54 in a population nearly three times
larger than any previously published study 52 , 53 . This
approach identifies genomic regions shared between individuals due to a common
ancestor and defines fine-scale ancestries, which we refer to as
“clusters” ( Methods ) 52 . We identified 36 fine-scale ancestry
clusters with at least 30 participants ( Figure
2a – b ; Figure S4a – c ; Methods ), labeled
with both a numeric identifier ( e.g. , IBD-01) sorted based on
the cluster sample size, and a corresponding cluster name to facilitate
interpretation ( Methods ). Because PCA captures ancient population
structure and IBD segments reflect recent shared ancestry, full agreement
between these approaches is not expected, as is observed for some of our IBD
clusters, such as IBD-02 ( Figure 2a ). We
replicated previously identified clusters and added unreported ones, such as a
Native Hawaiian cluster (n IBD-29 = 64) and a Bantu cluster
(n IBD-35 = 32). The largest clusters consisted of Northern
Europeans (n IBD-01 = 33,675) and Southern Europeans
(n IBD-02 = 14,841). The remaining clusters represent the
heterogeneity of ancestral origins in Los Angeles, including Ashkenazi Jewish
(n IBD-03 = 14,262), Filipino (n IBD-09 = 1,438),
Iranian Jewish (n IBD-11 = 707), and Armenian
(n IBD-16,IBD-23,IBD-31 = 560) populations. The diversity of ATLAS
is reflected not only by a relatively large portion of non-EUR individuals, but
also by substantial diversity within the EUR broad-scale ancestry population.
Several EUR clusters are large relative to published datasets, including
Southern European, Ashkenazi Jewish, Armenian and Iranian Jewish 3 , 53 , 55 – 57 ( Table S3 ). Similarly, for EAS, we
identified multiple clusters, including Filipino, representing the largest
published Filipino genotyped cohort of which we are aware 58 . Despite this strength, for some
populations, sample sizes are limited, which leads to relatively small
clusters.
For a nuanced understanding of diagnostic variation, we tested the
prevalence of 1,253 phecodes 59 – 61 across
24 fine-scale clusters with at least 100 participants. Table S3 shows all associations and
Figure 2c highlights selected
associations (also available at https://atlas-phewas.mednet.ucla.edu/ancestry ). We replicated known
findings ( Methods ) and uncovered numerous associations between
fine-scale clusters and diseases. For instance, Filipinos, an understudied
population in genetics research, exhibited the lowest risk for vitamin B-complex
deficiency (OR IBD-09 = 0.50 [0.36, 0.67], FDR =
2.1×10 −4 ), and the highest risk for cholesterolosis
of the gallbladder (OR IBD-09 = 4.3 [2.9, 6.2], FDR =
5.6×10 −12 ), both previously unreported. In Iranian
Jewish subjects, another group with low research inclusion, we detected a
previously unreported elevated risk for glaucoma (OR IBD-11 = 1.9
[1.5-2.4], FDR = 2.9×10 −6 ), a condition known to be
more common in individuals of African and East Asian ancestry. We also observed
several shared health risk patterns among Jewish ancestry clusters. Both Iranian
and Ashkenazi Jewish clusters showed the highest risk for bladder cancer
(Iranian Jewish: CI IBD-11 = 1.4-5.6, FDR = 0.008; Ashkenazi Jewish:
OR IBD-03 = 1.5 [1.1, 1.9], FDR = 0.02) and a high hyperplasia of
prostate risk (Iranian Jewish: OR IBD-11 = 2.3 [1.8, 2.8], FDR =
1.1×10 −13 ; Ashkenazi Jewish: OR IBD-03 =
1.7 [1.6, 1.8], FDR = 4.6×10 −66 ), and the lowest risk
for cirrhosis of liver (Iranian Jewish: CI IBD-11 = 0.03-0.4, FDR =
3.1×10 −2 ; Ashkenazi Jewish: OR IBD-03 =
0.4 [0.3, 0.5] , FDR = 6.8×10 −24 ). Mexicans and South
Americans suffered consistently more from hormones’ adverse effects in
therapeutic use, a previously unreported association (Mexican American clusters:
OR IBD-04,-07,and-12 = 2.0-2.8 [1.3, 3.4], FDR =
1.6×10 −2 - 2.1×10 −24 ;
combined South Americans: OR IBD-08 = 1.9 [1.4, 2.5], FDR =
1.3×10 −4 ). These associations persisted after SES
adjustment, using the BAS and ADI ranks. In most cases, associations held after
applying IPW and also adjusting for SES, except for hormone adverse effects,
where only the largest Mexican American group remained highly significant,
likely due to smaller sample sizes resulting from the IPW analyses ( Table S3 ). These findings
demonstrate the utility of fine-scale ancestry in identifying disease risk
profiles in underrepresented populations.
We next focused on cardio-metabolic diseases due to their global impact
on public health. We selected a few targeted cardio-metabolic phenotypes and
compared disease risk across fine-scale clusters within the same broad-scale
ancestry ( Methods ; Figure S4d ). Strikingly, among
Asian clusters, the Filipino cluster had an elevated risk for all tested
cardio-metabolic conditions, which remained after adjusting for BMI. Across
fine-scale EUR ancestries, the larger Armenian cluster showed high
cardiometabolic disease risk. Iranian clusters, both Jewish and non-Jewish, had
a relatively higher risk of coronary atherosclerosis, hyperlipidemia and type 2
diabetes, but not hypertension. Adjusting for SES factors and adding IPW
maintained these significant patterns ( Table S3 ), except abdominal aortic
aneurysm in Filipinos, which lost significance, and hyperlipidemia in Iranian
Jewish participants, which did not remain significant only when both SES
adjustment and IPW were applied- likely due to reduced sample size, although the
OR remained elevated.
To examine the robustness of ATLAS in predicting genetic risk and to
assess the quality of our dataset, we first tested the utility of PGS in
stratifying risk for common disorders. We observed high PGS performance in EUR
individuals–on average, 18.6% of patients diagnosed with major disorders
were in the top PGS decile, rising to 41% for type 1 diabetes ( Figure 3 ; STAR
Methods ). Consistent with prior reports 62 – 65 , this predictive power was diminished in non-European
ancestries ( Figure 3 ;
Methods ).
Next, we performed phenome-wide association studies (PheWAS) using
Regenie 66 to map the
common variant architecture of clinical diagnoses and laboratory measurements,
conducting separate analyses within each of our broad- and fine-scale ancestral
cohorts ( Methods ; n = 84,110 total participants). After linkage
disequilibrium (LD) pruning, we identified 19,431 unique variant-phenotype
associations passing a genome-wide significance threshold of
5×10 −8 ( Table S4 ), with 5,772 passing the
most conservative Bonferroni threshold of 6.4×10 −11
(5×10 −8 conditioned on 776 tested phenotypes). The
majority (83.8%) of associations were detected in only one ancestry at
genome-wide significance, though utilizing a more permissive threshold of
1×10 −4 showed cross-ancestry support for a
substantial portion of associations (54.7%; Figure S6a – b ). Supporting the validity of the
data used for PheWAS, we confirmed ancestry-specific differences in APOE allele
frequencies and attenuated Alzheimer’s risk associated with the ε4
allele in African Americans and replicated known associations such as
PNPLA3 rs738409-G with non-alcoholic fatty liver in AMR
groups 67
( Methods ; Figure S6c – e ). We identified numerous ‘previously unreported
findings’, which we defined as associations that have not been reported
at genome-wide significance in the existing literature, the European Molecular
Biology Laboratory-European Bioinformatics Institute (EMBL-EBI) GWAS
Catalog 68 , or summary
statistics from UKBB 69 , Taiwan
BioBank 7 , and
AoU 5 . Notably, 39.5% of
associations were unreported in the European Molecular Biology
Laboratory-European Bioinformatics Institute EMBL-EBI GWAS Catalog 68 ( Methods ). We
performed replication analyses for the top previously unreported findings in
independent biobanks, including AoU 5 , UKBB 69 ,
Taiwan BioBank 7 , and
BioMe 8 ( Table S4 ;
Methods ). Summary statistics are available at atlas-phewas.mednet.ucla.edu .
We identified a previously unreported association between
FN3K rs7208565-T and increased risk for intestinal
disaccharidase deficiency across EUR populations ( Figure 4A ; P EUR = 3.8 × 10 −41 ;
OR EUR = 1.3; MAF EUR = 33.3%; MAF IBD-01 =
32.3%; MAF IBD-02 = 34.3%; MAF IBD-03 = 35.6%) and the AMR
population (P AMR = 4.9 × 10 −11 ;
OR AMR = 1.3; MAF AMR = 42.3%; Table S4 ). This variant was
additionally associated with increased risk of the “other abnormal
glucose” phecode (P EUR = 1.8 × 10 −42 ;
OR EUR = 1.2), whereas a variant in close LD (rs113373052-T) was
associated with increased HbA1c (P EUR = 5.4 ×
10 −97 ; β EUR = 0.13), reflecting
additional consequences on glucose homeostasis. Intestinal disaccharidase
deficiency is the inability to completely digest sugars such as lactose and
sucrose, leading to irritable bowel syndrome (IBS)-like symptoms 70 . FN3K
phosphorylates glycated proteins, preventing the formation of advanced glycation
end-products 71 .
Supportive of these findings, rs7208565 has been previously associated with type
2 diabetes 72 .
Furthermore, we discovered several previously unreported associations
driven by low-frequency variants (MAF ~1-2%) in non-EUR cohorts. For
instance, we linked rs115750084-G in DPP6 , a gene implicated in
synaptic signaling 73 and
neurological disorders 74 , 75 , to an increased risk of major
depressive disorder in the AFR cohort (P AFR = 4.6 ×
10 −8 ; OR AFR = 4.2; MAF AFR = 1.1%;
Figure S6F ). In
the EAS cohort, we associated rs77742325-G in MAS1 , a key
component of the bone-protective ACE2/Angiotensin-(1-7)/Mas axis, 76 with osteoporosis
(P EAS = 1.5 × 10 −8 ; OR EAS =
3.5; MAF EAS = 1.8%; Figure S6G ). We also linked an
intronic deletion (rs202215133-ATATCATAG>A) in UBR5, an E3 ubiquitin
ligase implicated in neuroinflammatory and neurodevelopmental
syndromes, 77 , 78 with migraine headache in the Mexican
American cluster (P IBD-04 = 1.7 × 10 −8 ;
OR IBD-04 = 5.6; MAF IBD-04 = 1.1%; Figure S6H ). These associations
were independently replicated, albeit with more modest effect sizes in the
replication cohort ( Table
S4 ). Given that the variants are low MAF, further validation through
direct sequencing in large cohorts will be essential to confirm their
generalizability.
Similarly, our large AMR discovery cohort revealed several unreported
associations. For instance, we linked STARD7 rs17419569-C with
increased risk for asthma in the Mexican American cluster ( Figure 4B ; P IBD-04 = 1.9 ×
10 −9 ; OR IBD-04 = 2.6; MAF IBD-04 =
2.5%). Prior work suggests that decreased STARD7 expression is
associated with enhanced allergic responses in the human lung and significant
increases in airway hyperresponsiveness in haploinsufficient
Stard7 mice, 79 whereas rare variants within STARD7 were
linked with asthma (P UKBB = 5.1 × 10 −3 ;
Table S4 ). In the
same cluster, GPX7/SHISAL2A rs74744741-C was associated with
gastrointestinal reflux disease (GERD) ( Figure
4C ; P IBD-04 = 5.8 × 10 −10 ;
OR IBD-04 = 1.6; MAF IBD-04 = 8.9%).
GPX7 has been associated with carcinogenesis in the context
of GERD-associated Barrett’s esophagus, 80 whereas SHISAL2A has an
unknown function, but is highly expressed in the small intestine and in lymphoid
tissues. 81
Additionally, we identified two low-frequency AMR variants associated with
chronic renal failure (CRF; Figure 4D ),
rs112680741-C (P AMR = 2.9 × 10 −11 ;
OR AMR = 3.8; MAF AMR = 1.3%) and rs2744548-C
(P AMR = 9.6 × 10 −11 ; OR AMR =
3.2; MAF AMR = 1.7%), nominating GPLD1 ,
ALDH5A1 , and KIAA0319 as potential CRF
risk genes. These associations did not replicate in available biobank
populations ( Table
S4 ), possibly because the ATLAS AMR populations from the Los Angeles area
represent distinct Mexican and South American ancestries 82 , 83 , whereas those from other biobanks (BioMe, for example) are
comprised of individuals with distinct ancestral cluster assignments to Puerto
Rico and the Dominican Republic. Other cohort-specific characteristics may also
contribute to these differences, so independent replication in additional
cohorts is needed.
This first wave of the ATLAS WES catalog comprises more than 11.5
million autosomal variants, with a median of 9,821 missense and 159
loss-of-function (LOF) variants per individual, closely matching expectations
from other studies 84 . Most
(98%) WES variants were rare (MAF < 1% in any broad-scale ancestry),
including 2,873,731 rare missense and 172,462 rare LOF variants, with a median
of 15 rare LOF and 403 rare missense variants per participant ( Table S5 ). AFR participants showed
the greatest number of rare LOF and missense variants, matching prior
results 85 . To evaluate
our WES, we analyzed a predefined set of clinically relevant rare variants known
to have elevated frequencies in specific populations and demonstrated that ATLAS
captures these established patterns ( Methods ; Figure S7a – e ). Then, we examined differences
in 69 ClinPGx pharmacogenomic variants 86 , 87 across
fine-scale ancestries ( Figure 4e ; Table S4 ). We identified
several unreported enrichments in specific clusters, including a
SLCO1B1 variant, rs4149056, which affects the metabolism
and toxicity of statins 88 in
Ashkenazi and Iranian Jews; a CYP4F2 variant, rs2108622, that
can increase the required dosage of warfarin 89 , enriched in Ashkenazi Jewish, Lebanese, Armenian 1,
Egyptian Christian, and South Asian 2 individuals; and PYD
variants, which affect the toxicity of chemotherapy drugs 90 , abundant in Mexican American, African
American and South American clusters. This suggests the importance of
considering fine-scale ancestry in precision drug prescribing.
Next, we used ATLAS’ ancestral heterogeneity to identify
enrichments of known rare clinically relevant variants, aggregated by gene
across fine-scale clusters to illustrate the utility of fine-scale ancestry in
identifying populations at higher risk of rare, monogenic disease. We queried
the full set of the curated rare pathogenic/likely-pathogenic (P/LP)
ClinGen 91 variants
that have strong clinical and genetic evidence (n = 643), identifying 5,223
unrelated carriers ( Figure 5a ). Alongside
known associations ( STAR Methods ), we
detected a previously unreported elevated carrier frequency for variants in the
Centers for Disease Control and Prevention (CDC) Tier 1 gene,
LDLR , which causes familial hypercholesterolemia (FH) in
the Filipino cluster (OR IBD-09 = 3.9 [1.4, 8.8], FDR = 0.03). This
finding is notable given the high prevalence of dyslipidemia in
Filipinos 92 ( Figure S4D ). Indeed, LDLR
FH variants accounted for 3.1% (population attributable risk [PAR]) of high LDL
cases (LDL > 190 mg/dl) in the Filipino cluster. However, due to the low
number of carriers in this cluster (<10), the PAR CI was wide [0.3,
26.1], and independent replication is required to refine this finding.
Additionally, we identified previously unreported elevated carrier frequency of
variants that cause non-syndromic genetic deafness within two different genes in
two separate clusters: (1) MYO15A in the Mexican American 1
cluster (OR IBD-04 = 6.9 [2.6, 17.4], FDR =
8.44×10 −4 ) and (2) CDH23 in the
Chinese + Korean (OR IBD-05 = 6.0 [2.0, 15.6], FDR =
6.04×10 −3 ) and Japanese (OR IBD-10 = 28.1
[8.9, 77.1], FDR = 7.4 × 10 −6 ) clusters. We also
identified previously unreported elevated carrier frequencies in the Ashkenazi
Jewish cluster for Pendred syndrome variants in SLC264A
(OR IBD-03 = 3.8 [2.9, 5.0], FDR = 8.6 ×
10 −19 ) and recombinase activating gene 2 deficiency in
RAG2 (OR IBD-03 = 2.9 [1.4, 5.8], FDR = 0.02).
Finally, we focused on the clinically actionable American College of Medical
Genetics and Genomics (ACMG) 93
genes, where ClinGen-curated variants within these genes manifested a strong EUR
bias ( Figure 5b ; Figure S7f ; Methods ).
This bias was recently shown in AoU 23 and we provide additional supporting evidence by
demonstrating this bias in a single health system.
As clinical variant databases are primarily ascertained from EUR
populations, we took a conservative approach to call rare predicted deleterious
LOF (dLOF) and missense coding variants (dMIS). High-confidence dLOF were
defined using the loss-of-function transcript effect estimator
(LOFTEE) 94 , which
accounts for biological context beyond protein truncation. dMIS variants were
defined using a consensus of a majority (5 of 9) of state-of-the-art
computational methods based on the Critical Assessment of Genome Interpretation
(CAGI) project 95 for
pathogenicity prediction ( Methods ; Figure S7g – h ). This approach is distinct from
the one taken with ClinGen variants – variants identified in this manner
may not have literature evidence for their impact on gene function but are
potentially less biased by EUR oversampling.
First, we kept our focus on ACMG genes 93 , repeating the ClinGen-based analysis
described above using predicted rare dLOF and dMIS variants. We examined the
frequency of 1,602 rare dLOF and 6,487 rare dMIS variants within ACMG variants
across populations ( Figure 5c ; Figure S7i ). The results
were consistent with previous work showing that AFR participants harbor
substantially more dLOF compared to other ancestries 85 , and highlighted other groups with
previously unreported low or high ACMG putative damaging variants for further
investigation. Overall, these results contrasted in comparison with rare
variants identified in ClinGen ( Figure 5b ),
consistent with the interpretation that ClinGen variant frequencies are biased
toward EUR participants, as is the case for most curated clinical datasets.
Second, to examine the influence of dLOF and dMIS on ATLAS phenotypes,
we performed exome-wide association studies (ExWAS) across 17,537 protein-coding
genes ( Methods ). Using Regenie’s unified gene burden
association strategy 66 , we
identified 1,099 unique Bonferroni-significant gene-trait associations. Within
the EUR population, we replicated numerous known associations; 45 of the top 50
associations had been previously reported 84 ( Table S6 ). These include PKD1 with cystic kidney
disease and TTN with primary intrinsic cardiomyopathy ( Figure 5e ) 84 . We also identified numerous previously
unreported associations ( Table S6 ). In the Northern European cluster, we detected gene-level
associations between HNRNPA1L2 and acquired absence of the
breast (GENE_PIBD-01 = 1.0×10 −25 ), breast cancer
(GENE_PIBD-01 = 2.2× 10-8 ), and malignant neoplasm of the
female breast (GENE_P IBD-01 = 1.5 × 10 −7 ).
Notably, HNRNPA1L2 is on chromosome 13, ~20 megabases
from BRCA2 , representing a different locus. Though
HNRNPA1L2 has not been previously associated with breast
cancer, it is highly expressed in breast invasive carcinoma 81 and a paralog HNRNPA1
was previously implicated in breast cancer progression 96 . In the Ashkenazi Jewish (IBD-03)
cluster, we observed associations between EPG5 and HDL
cholesterol level (P LOF_MIS = 3.6×10 −10 ;
β LOF_MIS = 1.8 [1.2, 2.3]) and triglyceride level
(P LOF_MIS = 1.33×10 −8 ;
β LOF_MIS = −1.81 [−2.4, −1.2]).
EPG5 is an autophagy tethering factor classically
associated with Vici syndrome 97 , a severe developmental disorder, though our findings
suggest an additional role in lipophagy and lipid metabolism 98 .
In the AMR population, we identified an unreported association between
dLOF and dMIS variation in CLN3 and cystic kidney disease
(GENE_P AMR = 8.7×10 −9 ), supported by
experimental evidence that CLN3 is highly expressed in
medullary collecting duct principal cells and plays a role in
osmoregulation 99 . We
additionally identified unreported associations between PPARG
and viral pneumonia (P AMR,MIS = 1.1×10 −6 ,
OR AMR,MIS = 53.0 [13.5, 208.5]), as well as
NADSYN1 and abnormal lung examination findings
(P AMR,MIS = 1.2×10 −5 ,
OR AMR,MIS = 3.7 [2.1, 6.5]). PPARG is known to
regulate macrophage response to pulmonary inflammation 100 , while NADSYN1
knockout in mouse models was shown to cause abnormal lung development 101 , though neither has been
previously associated with respiratory traits in humans.
Within the AFR population and African American cluster, we uncovered
unreported associations between DDHD2 and dysphagia
(P IBD-06,LOF = 1.2×10 −6 ,
OR IBD-06,LOF = 8.3 [3.7, 18.7]) and EFCAB13 and
essential hypertension (P AFR,MIS = 9.3×10 −7 ,
OR AFR,MIS = 13.1 [4.5, 37.7]). These findings align with prior
studies; variants in DDHD2 have been shown to cause hereditary
spastic paraplegia and symptoms of dysphagia 102 , though not in African Americans, and
methylation studies have prioritized EFCAB13 as a risk factor
for heart failure 103 . We
detected an additional association between ANKZF1 and
peripheral vascular disease (P AFR,LOF_MIS =
1.6×10 −6 , OR AFR,LOF_MIS = 58.8 [12.6,
275.1]), which is supported by recent experimental findings 104 .
Next, in the EAS population, we observed a previously unreported
association between AK7 a nd non-rheumatic mitral valve
disorders (P EAS,LOF = 7.6×10 −7 ,
OR EAS,LOF = 85.7 [39.2, 187.2]). While AK7 has
never been associated with cardiovascular phenotypes, it has a well-established
role in primary ciliary dyskinesia 106 and mutations in other ciliary genes have been
associated with mitral valve prolapse 107 . We further identified an unreported association
between KIF2B and nephritis/nephropathy in the Chinese/Korean
cluster (P IBD-05,MIS = 7.2×10 −7 ,
OR IBD-05,MIS = 26.2 [9.0, 76.4]). Though KIF2B
has not been linked to renal traits, other genes in the kinesin family (KIF)
have been implicated in the pathogenesis of renal cancer 108 , 109 . All previously unreported findings described above were
independently replicated in AoU 5 , UKBB 69 or
BioMe 8 at a nominal
P-value < 0.05 ( Table
S6 ).
Finally, we examined known risk associations to identify
ancestry-specific patterns of deleterious rare variation. We tested two known
GBA1 rare missense variants, p.Glu365Lys and p.Thr408Met,
in the AMR and EUR populations; both variants are reported to increase risk for
Parkinson’s disease 110 . Despite similar allele frequencies within each population,
p.Glu365Lys was only associated with increased risk in EUR
(P EUR,E365K = 0.01; P AMR,E365K = 0.8), whereas
p.Thr408Met was only associated with increased risk in AMR participants
(P EUR,T408M = 0.2; P AMR,T408M = 0.003), displaying
evidence of ancestry-specific risk stratification ( Figure S7j ). While we hypothesize
that penetrance is modulated by population-specific genetic backgrounds, further
studies are required to identify the causes of these differential associations.
In total, We performed replication analyses for 21 of our top previously
unreported PheWAS and ExWAS findings in independent biobanks and replicated 16
(76%) at a nominal p value of 0.05.
One of the advantages of EHR data is the ability to consider dynamic
changes over time. To illustrate this, we integrated all main study components
(genetic ancestry, common and rare genetic variants) to study the efficacy of
glucagon-like peptide-1 (GLP-1) receptor agonists (GLP1-RAs) for weight loss. We
focused on semaglutide, which had the greatest number of prescriptions in ATLAS
( Figure
S8a – b ; 7,340 participants; Figure 6a ).
We observed a steady decrease in average weight up to ~60 weeks of
semaglutide treatment, consistent with previous findings 111 – 113 ( Figure S8c ). We tested if the medication dose, route, sex, age, and
initial BMI affected semaglutide efficacy, corroborating that dose and
subcutaneous delivery versus oral delivery were positively correlated with
weight loss ( Figure 6b ; linear
mixed-effects model [LMM]: dose: P Bonferroni =
3.3×10 −97 , semipartial R² (explained
variance) = 1.8%, effect size = −1.0; route oral vs.
subcutaneous: PBonferroni = 7.2×10 −37 , semi-partial
R 2 = 1.1%, effect size = 2.2). We did not detect a significant
effect of age and sex on efficacy.
Next, we asked if semaglutide’s effects varied by broad-scale
ancestry, finding that ancestry significantly influenced weight loss over 60
weeks of treatment ( Figure 6c ; ANOVA on an
LMM; ancestry: P Bonferroni = 0.002; ancestry×time:
P Bonferroni = 0.0006). Comparisons between EUR and the other
populations revealed less weight loss in AMR, and a slower rate of weight loss
in AMR and EAS ( Figure 6d ; LMM;
ancestry AMR : P Bonferroni =
2.9×10 −3 , β = 0.8;
ancestry AMR ×time: P Bonferroni =
5.0×10 −3 , β = 45.2; ancestry EAS
×time: P Bonferroni = 3.2×10 −3 ,
β = 76.4), consistent with recent analysis based on a single time point,
which showed reduced efficacy in AMR and AFR cohorts (here, in AFR, the result
was not significant but the direction of effect was consistent). 113
Given limited evidence on the role of inherited factors in GLP1-RAs
effectiveness 114 , we
used ATLAS to explore the contribution of common genetic variation to
semaglutide-related weight loss. We found no correlation between BMI PGS and
semaglutide efficacy. However, weight loss was negatively correlated with DM2
PGS ( Figure 6e ; LMM, PGS High
vs. Low : P Bonferroni =
2.8×10 −4 , β = 0.96; PGS Med
vs. Low : P Bonferroni =
8.8×10 −3 , β = 0.73). The same relationship
was observed in a simplified model using only maximum weight loss
recorded 114 (linear
regression, P Bonferroni = 1.3×10 −2 , β
= 0.31; Figure S8d ). We
could not find an equally sized cohort for replication, but we used EUR AoU,
which had a sample size 50% smaller than the ATLAS cohort used in this analysis
with array genotyping (1,578 vs . 3,165 in AoU and ATLAS). In
AoU EUR participants, we observed a concordant direction of effect, supporting
replication of the finding, although statistical significance was not reached.
This may reflect the smaller sample size or other population differences, such
as in SES and compliance ( Figure S8e ). Larger samples are needed for more formal replication.
As a second step, we conducted a genome-wide association study (GWAS)
meta-analysis across ancestries, but did not identify any significant loci
( Methods ; Figure S8f ).
Finally, we used WES to test gene-level associations with semaglutide
response (Regenie 66 ;
Methods ). Given the modest sample size, we focused on EUR
individuals and limited our test to the subset of proteins whose plasma
abundance in humans was recently shown to be altered by semaglutide 115 . We identified a significant
association of weight loss on semaglutide with PTPRU
(P Bonferroni = 7.6×10 −3 , β =
−0.87) ( Figure 6f ; Figure S8g ). For replication, we
used AoU, with a cohort that was 21% smaller than the ATLAS cohort used (1,581
vs . 2,012 in AoU and ATLAS with WES). This analysis
supported our original finding by showing a consistent direction of effect,
although significance was not reached, possibly due to the smaller sample size
or differences between the tested populations (AoU: β = −0.27,
P-value = 2×10 −1 ). Therefore, larger sample sizes in
future studies will be needed to formally replicate this finding. However, the
Bonferroni adjusted combined P from ATLAS and AoU remained significant,
supporting this finding (P Bonferroni =
2.3×10 −2 ). This association involves 37 variants,
of which one is common (rs2235937; nominally associated with weight loss; P =
9.7×10 −3 , β = −0.06 in EUR), and the
rest are rare, the frequency of which varied across ancestries ( Figure S8h ; Methods ).
The direction of effect indicates that PTPRU activity is
negatively associated with weight loss in patients taking semaglutide (variants
in PTPRU with a predicted high or moderate functional
consequence contribute to weight loss on semaglutide). Maretty et
al. highlighted PTPRU as a protein whose abundance is elevated in
individuals with higher genetic risk for increased BMI or type 2 diabetes and is
downregulated by semaglutide treatment 115 . Although direct causality has not been established,
these observations, together with our findings, support a model in which
semaglutide-induced downregulation of PTPRU and genetically
reduced PTPRU function promote greater weight loss during semaglutide treatment.
This gene has no previous known functional relationship to weight loss or
related metabolic functions. However, combined with the data from serum
proteomics 115 , our
genetic analysis nominates this protein kinase and its pathways as candidates
for further investigation.
Resource
Questions should be directed to the lead contact, Daniel H. Geschwind
(
[email protected] ).
No materials were generated in this study.
Discussion
Large-scale, longitudinal population studies have accelerated our
understanding of the causes and consequences of a wide variety of biomedical
conditions. The Framingham study 116 , 117 laid the
foundation for population-based cohorts, followed by efforts like the UKBB and
health system biobanks 3 , 8 , 9 , 11 , 53 , 72 , 84 , 118 , 119 . Despite the known sources of
errors and incomplete phenotyping in health system EHRs, multiple studies have shown
that many of these factors can be mitigated, allowing robust analyses 3 , 9 , 10 , 120 . Indeed, we were able to validate dozens of
previously identified genetic associations in ATLAS based on EHR-derived phenotypes
alone. We also leverage the heterogeneity of genetic ancestries in our population to
validate and extend our knowledge of disease burden and genetic risk factors.
Prior work has mostly relied on broad-scale genetic ancestry to estimate
health risk. 5 , 30 , 117
Fine-scale ancestry discerns more recent genetic variation via
shared IBD, 120 but smaller
clusters can limit power. Through PheWAS and ExWAS across broad- and fine-scale
ancestries, we identified numerous unreported genetic associations, many of which
were subsequently replicated. We leveraged fine-scale ancestry definitions to
benefit populations that are underrepresented in genetics and medical
research 122 . For example,
to our knowledge, ATLAS consists of the largest genetic dataset of Filipino
individuals (n = 1,438 vs. 1,028 in the next largest
dataset 58 ), which can be
used to improve precision care for this population and fuel discovery. Similarly,
our Ashkenazi Jewish cluster of 14,261 participants, which is 2.8 times larger than
any prior single-study genetic cohort 57 , enabled previously unreported associations and enhanced risk
quantification. Future work should assess the transferability of our results across
populations 123 .
Discerning genetic variant risk frequency and effect across fine-scale
ancestries is powerful for optimized risk stratification and tailoring therapeutic
intervention. It has been well demonstrated that diagnostic misclassification due to
differences in variant frequencies across ancestries is a serious risk 22 . We emphasize that leveraging
diversity within a single regional biobank can reduce confounding by ancestral
genetic differences and non-genetic variables that differ across countries or
geographically distinct biobanks. We demonstrated the application of longitudinal
EHR to gain insights into health outcomes over time, by linking genetic factors to
weight loss on semaglutide. This involved integrating dynamic EHR changes with
detailed prescription information, including the medication dose and route, which we
identified as major confounders. We considered only a single GLP-1RA drug,
semaglutide, to avoid the complexity introduced by different GLP-1RAs, which have
differing efficacy 124 . Lastly,
our study highlights fine-scale populations with unreported low or high ACMG
putative damaging variants for further investigation. We expect ATLAS to grow and
become an engine for clinical intervention, allowing researchers to rapidly identify
individuals for clinical studies and to implement precision medicine as the field
evolves.
Certain limitations should be considered. First, due to the nature of
EHRs, our study captures a partial view of patient phenotypes and care. Second,
our analysis focuses on the risk of receiving a diagnosis for a disease, which
is related, but not equivalent to the underlying risk of disease development.
This may influence how our findings should be interpreted in the context of true
disease incidence. Third, we relied on computational predictors for comparing
the abundances of predicted damaging variants in ACMG genes across ancestries.
Some predictors may introduce ancestry-related biases 125 , 126 , potentially affecting the accuracy of the results. To
mitigate these potential biases, we applied rigorous filtering and only retained
variants upon concordant agreement in a majority of nine tools. Additionally,
comparing the numbers of predicted damaging variants across ancestries may be
less stable when small ancestry groups are involved. We therefore excluded
groups with fewer than 400 individuals and reflected uncertainty using error
bars. Lastly, while our study introduces multiple previously unreported
associations, findings that could not be replicated should be further evaluated
in independent cohorts as they become available.
Introduction
The integration of electronic health records (EHRs) with genetic data is
transforming biomedical research, offering unprecedented opportunities for
preventing and managing common medical conditions 1 . Longitudinal sampling, linked with genetic and
environmental data, provides advantages over standard cohort-driven
research 2 . This has fueled
the creation of nationwide biobanks, including the UK Biobank (UKBB) 3 , 4 , All of Us (AoU) 5 , FinnGen 6 and
Taiwan Biobank 7 , as well as several
large-scale academic biobanks, such as Mt. Sinai’s BioMe 8 , Vanderbilt’s BioVU 9 , Geissinger’s MyCode 10 , and the Michigan Genomics
Initiative 11 . Integration
of these biobanks has permitted innovative collaborative efforts, such as the eMERGE
consortium 12 , COVID-19
host genomics initiative 13 , and
the Global Biobank Initiative 1 .
Although these efforts have substantially advanced genetic and biomedical
discovery, their concentration on participants of European (EUR) ancestry limits
generalizability 14 – 18 . As PGS are validated for clinical
use, the importance of measuring their accuracy in diverse populations grows; recent
analyses show a continuous relationship between ancestral distance from the
reference population and the utility of PGS 17 . Similarly, rare genetic variation has substantial
ancestry-specific effects and distribution; for example, in African Americans,
APOE4 alleles have reduced impact on Alzheimer’s disease
risk 19 and rare protective
PCSK9 variants are more prevalent 20 . The interpretation of clinically-relevant
rare variation is further hampered by a bias towards EUR variants in genetic
databases 21 , 22 . Including non-EUR populations reveals
substantial disparities in clinically-relevant rare variant frequencies 23 and increases statistical power
for discovery 24 , 25 . Thus, greater ancestral diversity in
biobanks with detailed medical records strengthens efforts to advance precision
health 14 , 26 , 27 .
Here, we analyzed data from the UCLA ATLAS Community Health Initiative,
linking EHR with genomic information for 92,164 participants with array genotyping
and 61,797 with whole-exome sequencing (WES). ATLAS reflects Los Angeles’
ancestral diversity within a single health system 28 – 30 , which reduces confounding from differing clinical practices
and enables robust cross-population comparisons. We performed phenome-wide
associations using common and rare variants within continental (broad-scale) and
sub-continental (fine-scale) ancestral groups, and quantified disease diagnoses and
genetic risk across these groups. Using longitudinal EHR data, we employed
semaglutide as a case study to identify genetic factors that alter drug efficacy. We
identified multiple unreported ancestry-specific risk associations, further
demonstrating the utility of ancestral diversity for more equitable, personalized
medicine research ( Figure 1a , schematic
overview).
Supplementary Material
Figure S1. Whole-exome sequencing quality control, related to
the
STAR Methods . a. Coverage
distribution. The mean and median values were calculated across all samples.
Fold enrichment is the degree to which the baited region is enriched
compared to the background genomic region. b. Percentages of
bases above different coverage thresholds. Every y-axis point shows the
percentage of bases above the corresponding x-axis value. The blue area
shows 95% of the data, while the surrounding gray lines represent the
remaining 5% of the distribution. c. Average base counts by
exome capture regions. d. Excluded bases from coverage
calculation based on Picard 127
e. Insert size distribution. f. Samtools 128 summary statistics
output. g . Genotype number distribution according to
RTGtools 129 .
h . Summary output from RTGtools. Ti/Tv is the ratio of
transition (Ti) to transversion (Tv). This included SNPs from off-target
regions. i. The distribution of variant counts across all
samples after cohort re-genotyping.
Figure S2. Phenotype validation and prevalence, related
to
Figure 1
and the
STAR Methods .
( A–J ) Comparison of vital signs and laboratory test
distributions between case and control groups defined by phecodes, with p
values indicating significance from Wilcoxon tests. ( K and L )
Overlap between cancer cases defined from EHR cancer diagnosis data
(“Cancer Stage Fact”) and phecode-based case-control groups
for breast ( K ) and prostate cancer ( L ).
( M ) Prevalence of frequent phecodes. The red bars indicate
the prevalence of the most frequent phecodes in subset of the UCLA ATLAS
population who had an encounter between 1–2 years after their initial
encounter, at the time of sample collection. The blue bars represent the
prevalence of phecodes in this same subset at their encounters 1–2
years after their initial encounter. ( N ) The change in phecode
prevalence over one year post-collection in the subset of the UCLA ATLAS
population who had an encounter between 1–2 years after their initial
encounter. The triangles represent the number of new patients, and the blue
bars represent the percentage of patients. ( O and P ) The
difference in ( O ) phecode group and ( P ) phecode
prevalence between participants in the UCLA biobank (purple; “UCLA
ATLAS”) at their time of collection compared with all other UCLA
Health patients with EHR data (green; “Rest of Data Discovery
Repository [DDR]”) within one year of the ATLAS launch date. Logistic
regression tests were used to compare the two groups adjusted for sex and
age (ORs with 95% confidence intervals, shown in the right panels). After
Bonferroni multiple testing corrections, ** denotes a significant difference
between the two populations.
Figure S3. Broad-scale genetic ancestry and health
characteristics, related to
Figure 1 . a. Agreement
between self-reported race and genetic ancestry predictions. b .
The Elixhauser comorbidity index varies across genetic ancestries after
applying inverse probability weighting. ANOVA on a linear regression model
tested the overall effect of the categorical ancestry predictor, yielding
its P-value. Adjusted means and 95% CI are presented per ancestry group.
c-d. The relationship between total and hospital encounters
and the comorbidity index. r, Pearson correlation, P, the correlation
P-value. e . Variation in laboratory or vital sign measurements
across broad-scale ancestries. ANCOVA adjusted for genetic sex and age was
performed separately for each tested phenotype.
Figure S4. Fine-scale genetic ancestry supplementary results,
related to
Figure 2 . a-b. Quality
control of identity-by-descent (IBD) segments. Distribution of IBD segment
length per chromosome, before ( A ) and after ( B )
removing human leukocyte antigens (HLA), centromere, and IBD depth outliers.
After quality control, IBD segments display exponential decay of segment
length as expected (most noticeably for chromosomes 6, 15 and 22). The
numbers at the top of each plot represent the chromosome number.
c . The distribution of genetic ancestry between fine- and
broad-scale populations. IBD clusters were sorted by the predominant
broad-scale ancestry of participants in each cluster, rather than cluster
size as shown in Figure 2a , providing
an alternative visualization. D. Cardio-metabolic disease risk
for each fine-scale group within the same broad-scale ancestry.
Representative cardio-metabolic phecodes were selected, and only populations
with at least 100 participants were tested. In cases of small sample sizes,
‘–’ was used instead of numeric values to protect
patient privacy. Filled points represent significant results (FDR ≤
0.05). Firth’s bias-reduced logistic regression adjusted for BMI,
sex, and age was used to obtain OR. Among Asian clusters, Filipino
individuals were at high risk for all tested medical conditions (essential
hypertension: OR IBD-09 =1.6 [1.4, 1.9], FDR = 8.6 ×
10 −11 ; type 2 diabetes: OR IBD-09 = 1.5
[1.3, 1.7], FDR = 7.4 × 10 −6 ; coronary
atherosclerosis: OR IBD-09 = 1.4 [1.1, 1.7], FDR = 5.3 ×
10 −3 ; abdominal aortic aneurysm: OR IBD-09 =
2.8. [1.4, 5.3], FDR = 7.1 × 10 −3 ; hyperlipidemia:
OR IBD-09 = 1.2 [1.0, 1.4], FDR = 2.4 ×
10 −2 ). The largest Armenian cluster and Jewish and
non-Jewish Iranian clusters showed a high risk for type 2 diabetes (Armenian
1: OR IBD-16 = 2.0 [1.5, 2.7], FDR = 2.5 ×
10 −6 ; Iranian Jewish: OR IBD-11 = 2.4 [1.9,
2.9], FDR = 6.0 × 10 −16 ; Iranian:
OR IBD-17 = 2.2 [1.6, 2.9], FDR = 1.6 ×
10 −6 ), hyperlipidemia (Armenian 1: OR IBD-16
= 1.6 [1.2-2.0], FDR = 8.6 × 10 −4 ; Iranian Jewish:
OR IBD-11 = 1.5 [1.2, 1.8], FDR = .1.0 ×
10 −4 ; Iranian: OR IBD-17 = 1.8 [1.4, 2.3],
FDR = 2.9 × 10 −5 ) and coronary atherosclerosis
(Armenian 1: OR IBD-16 = 1.9 [1.4, 2.5], FDR = 8.8 ×
10 −5 ; Iranian Jewish: OR IBD-11 = 1.8 [1.5,
2.2], FDR = 4.2 × 10 −8 ; Iranian:
OR IBD-17 = 1.9 [1.4, 2.5], FDR = 2.9 ×
10 −5 ).
Figure S5. polygenic score association results, related
to
Figure 3
and the
STAR Methods . a-e
Selected associations between polygenic scores (PGS) and diseases in
European (EUR) individuals. f-i The odds ratio (OR) and
prevalence of cases for the top and bottom PGS deciles across non-EUR
ancestries. Logistic regression was used to obtain the OR and p values,
which were adjusted using FDR.
Figure S6. PhWAS supplementary results, related to
Figure 4 .
a-b. Distribution of pruned
variant-trait associations. The number of genome-wide significant pruned gene-trait associations
shared across fine-scale ancestries, using a prioritized gene within 10kb of the pruned variant
and matching effect direction. b presents the same associations described in a, but with a more
permissive threshold for gene-trait support from less powered ancestral clusters. c. APOE
haplotype frequency across fine-scale cohorts. d. Allele frequency for known risk variants across
fine-scale cohorts; shading indicates the level of over- (red) or under- (blue) enrichment of a
haplotype/allele in a given cohort via Fisher’s exact test. In c-d, Fisher’s exact test was utilized to
compare allele frequency. Bold border indicates a significant result (FDR ≤ 0.1). e. Impact of the
non-alcoholic cirrhosis risk variant rs738409-G on cirrhosis and clinical sequalae across finescale
cohorts. Odds ratios and 95% CI were calculated using logistic regression. f-h. Replicated
low-MAF PheWAS associations between rs115750084-G and major depressive disorder (f),
rs77742325-G and osteoporosis with no other symptoms (g) and rs202215133-A and migraine
(h).
Figure S7. Known or putative rare pathogenic variants, related to Figure 5
Figure 5 and the STAR
Methods.
a. Frequency differences in Familial Mediterranean Fever (FMF) known risk alleles. b.
The risk of carrying the HBB:p.E7V variant. c. The risk of carrying loss-of-function variants in
PCSK9. In a-c, the error bars show the 95% Wilson score confidence intervals. d-e. The
frequency of BRCA Ashkenazi Jewish founder alleles across broad-scale (d) and fine-scale
ancestry groups (e). f. Differences across populations in the total numbers of rare ClinGene P/LP
variants in American College of Medical Genetics (ACMG) genes, excluding Ashkenazi Jewish,
as a sensitivity analysis to one presented in Figure 5b. ‘Ref’ is the total number of reference
alleles, and ‘Alt’ of P/LP ClinGen. Fisher’s exact tests were used to produce odds ratios. Filled
points represent significant results at the level of nominal P-value (≤ 0.05). Only fine-scale
ancestry clusters with more than 400 participants were included. g-h. Distributions of missense
variant pathogenicity rank scores for nine computational tools. Each subplot compares the
distribution of all missense variants (purple to green) to known pathogenic variants (red), defined
as either: g. ClinGen curated missense variants. h. ClinVar pathogenic/likely pathogenic (P/LP)
missense variants. The black dashed line indicates the median rankscore across all missense
variants for the given tool. The red dashed line denotes the median rankscore among pathogenic
variants in the corresponding dataset. The purple dashed line marks the likelihood-based
intersection cutoff derived from the point at which the pathogenic and background distributions
cross. i. The number of rare computationally predicted damaging missense and LOF alleles per
individual across ancestries. Mann-Whitney U test with a Bonferroni correction was applied to
test the difference in the distribution of rare LOF and predicted damaging missense counts per
individual between any broad- or fine-scale group compared to all others. Statistically significant
differences are indicated by an asterisk (*). Only fine-scale ancestry clusters with more than 100 participants were included. j. Ancestry-specific carrier counts for two known GBA1 rare LOF
variants (left) and their impact on Parkinson’s disease risk (right) via logistic regression.
Figure S8 . GLP1-RAs ATLAS users and semaglutide
investigation complementary data, related to
Figure 6 . a . Age by sex of
GLP-1 receptor agonist (GLP1-RAs) users. b . Prescription
numbers for GLP1-RAs by simple generic names. c . Weight loss
patterns across time. Presented are smoothed longitudinal data using a
functional boxplot approach, with median values and pointwise intervals
between the 20th and 80th quantiles. Vertical tick marks along the x-axis
show the deciles of the data distribution, indicating where most data points
are concentrated across weeks. d . The relationship between bins
of polygenic scores (PGS) for body mass index (BMI) (left) and type 2
diabetes mellitus (right), and weight loss in response to semaglutide. Shown
are Loess smoothed plots with 95% CI based on the maximum weight loss for
patients across weeks. p values and effect sizes were obtained from a linear
regression model with covariates, and a Bonferroni correction was applied.
e . The relationship between type 2 diabetes mellitus PGS
and weight loss in response to semaglutide in EUR AoU participants. Scaled
PGS were divided into groups and linear mixed model fitted values were
plotted with 95% CI, based on longitudinal data with repeated weight
measurements. P-values were obtained from a linear mixed-effects model that
included covariates. f . GWAS Q-Q plot for weight loss on
semaglutide displays no significant findings. g . Gene-level
test Q-Q plot for weight-loss on semaglutide (genomic inflation factor:
0.95). h. Differences in carrier numbers of semaglutide
efficacy involved alleles in PTPRU across ancestries using
Firth’s bias-reduced logistic regression. Only variants that
exhibited at least one significant difference between EUR and another
ancestry are shown. Filled dots represent significant odds ratios (FDR
≤0.05).
Table S1. Summary of lab test results and prescription numbers for
the top 50 prescribed medications in ATLAS, related to the STAR Methods .
Table S2. Phecodes and broad-scale ancestry associations, related
to Figure 1 and the STAR Methods .
Table S3. Fine-scale ancestry cluster sample sizes compared with
published cohorts and phecode associations, related to Figure 2 and the STAR Methods .
Table S4. PheWAS results and pharmacogenomic variants, related to
Figure 4 .
Table S5. Summary of WES variant counts by category and differences
across broad-scale ancestries, related to Figure 5 and the STAR
Methods .
Table S6. ExWAS results, related to Figure 5 .
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.