Author
A.S. and F.M.‐T. conceived the project and funding. J.P.‐S., A.C.‐M., X.B., A.G.‐C., L.C.‐M., C.M.‐C., and A.S. carried out the analyses. A.C.‐M. and X.B. designed the GALOMICS web resource. J.P.‐S., A.G.‐C., C.M.‐C., and A.S. contributed to a critical discussion of the main findings. ASE wrote the draft of the study. All the authors critically reviewed the final version of the manuscript.
Results
We first removed non‐autosomal variants, indels, and variants with a genotyping rate below 99%. After applying additional filters (e.g., removal of triallelic SNPs and indels), the dataset was reduced to over 17.2 million variants, which were used for downstream analyses unless additional filtering was required, e.g., for Hardy–Weinberg disequilibrium, LD. Approximately 36.4% of the variants were located within intronic regions, while 32.3% were in intergenic regions. Additionally, 15.2% of the variants were identified in regulatory or splice site regions, and 0.8% were in exons. Among the variants located in exons, > 35 K were nonsynonymous, > 29 K were synonymous, and > 79 K were found in 3′/5′‐untranslated regions (UTR). The remaining 500 variants were classified as either stop‐gained or stop‐loss.
Next, we investigated the familial relationship between donors in GALOMICS as done previously [ 42 , 50 ]. Of the initial 94 WGS, three were excluded due to close relationships (Figure 1B ) or their behavior as outliers in the MDS analysis (see below).
Finally, we generated an open‐access web platform, GALOMICS ( https://galomics.genpob.eu ), to facilitate the search for allele variants in the GAL database. It hosts > 18 million biallelic variants from the GAL dataset. Users can query variants by rsID or genomic coordinates, visualize allele distributions within Galicia and globally, and access functional and clinical annotations from Ensembl, ClinVar, OMIM, and UniProtKB. The tool has been designed to be scalable, anticipating the integration of new variants from other datasets. Many of these datasets represent disease conditions in Galician donors, some of which have already been generated, for example, [ 38 , 42 , 43 , 44 , 51 ].
The MDS plot (Figure 1C ) of worldwide population genomic datasets shows most samples clustering around the three main continental groups: sub‐Saharan Africa, Asia, and Europe. The first dimension (12.35% of the variation) primarily separates the sub‐Saharan group (LWK and YRI) from the rest, while Dimension 2 (7.24%) distinguishes East Asians (CHB and JPT) from Europeans (represented by BAS, CEU, FRA, GAL, GBR, IBS, SAR, and TSI) and sub‐Saharan Africa. North African datasets (BED, DRU, MOZ, and PAL) cluster closer to Europeans, but are slightly shifted towards the sub‐Saharan vertex, with a few outliers showing strong sub‐Saharan affinity. In a close‐up of the European/North African cluster (Figure 1C ; left), Galician genomes are embedded but slightly displaced from the European core, showing a subtle shift towards the North African pole. Dimension 3 (0.64%; Figure 1C ; right) separates North Africans from Europeans, with Bedouins (BED) at the opposite pole, while Dimension 4 (0.29%) separates North Africans and Europeans from Mozabites (MOZ; Figure 1C ; right). These two dimensions highlight Bedouins and Mozabites at the extremes of North African variation, with the rest of the North African samples showing more genetic affinity to Europeans.
We used the unsupervised clustering algorithm of ADMIXTURE to estimate ancestry proportions in Galician genomes, influenced by gene flow between African and European populations (Figure 2A ). Our analysis utilized population datasets representing major continental groups, including sub‐Saharan Africa, North Africa, East Asia, and Europe, enriched with additional North African samples from [ 19 ]. We explored K values from 2 to 10 contributors, and focused on K = 3, K = 4, and K = 5 as the most demographically informative. At K = 3, ADMIXTURE identifies ancestry components from sub‐Saharan Africa, East Asia, and North Africa/Europe. At K = 4, the model distinguishes the North African from the European component, suggesting shared ancestry between these regions (Figure 2A ; left). At K = 5 (the model with the lowest cross‐validation value), the algorithm divided the North African/Middle Eastern component into two: one mainly characterizing Mozabites (hereafter referred to as “Mozabites North African”‐like ancestry), and another reflecting Middle Eastern variation (represented by Palestinian, Bedouin and Druze) and the other North African datasets (hereafter referred to as “non‐Mozabites North African”‐like ancestry). All the datasets in North Africa and the Middle Eastern have, however, variable proportions of these two ancestries. These ancestries are also present in several European datasets (FRA, SAR, TSI, IBS, and GAL), while the “non‐Mozabite North African”‐like ancestry is absent in BAS, and both are absent in CEU and GBR.
Ancestry composition, spatial distribution of ancestry components, and shared‐drift signals in the GAL dataset, assessed through ADMIXTURE, interpolated frequency mapping, and f
3 ‐statistics. (A) Bar plot depicting individual ancestry proportions estimated using ADMIXTURE with both unsupervised (left) and supervised (right) clustering approaches, considering K values from 4 to 5. (B) Interpolated frequency map of the Middle Eastern and North African component among GAL donors across the Galician region, based on the ancestry component identified in (A) (see indications for 2B_1; 2B_2, and 2B_3). (C) f
3 ‐statistics in the form f
3 (X EUR , X NAF/SSH ; GAL), measuring shared genetic drift. Each dot represents an f
3 value calculated for a pair of source populations with GAL as the target population. Blue dots correspond to combinations involving sub‐Saharan African populations, red dots indicate combinations with North African or Middle Eastern populations, and filled red dots highlight statistically significant results ( Z ‐score ≥ 3), suggestive of admixture events contributing to the genetic makeup of the GAL population. Negative f
3 values indicate potential evidence of admixture, while positive values reflect shared genetic drift. See Table 1 for population code details. SS‐Africa: Sub‐Saharan Africa. NA/ME: North Africa/Middle Eastern.
To better estimate the North African/Middle Eastern component in GAL, we performed a supervised ADMIXTURE analysis using the following datasets (Figure 2B , right): YRI and LWK (sub‐Saharan Africa), CHB and JPT (East Asia), and GBR and CEU (Europe). For North African and Middle Eastern references, we selected individuals having > 90% (initial run) or > 80% (second run) of their primary ancestry component to represent their most unadmixed lineages and minimize European ancestry. Under the 90% threshold, 18 Bedouins met the criteria, whereas with the 80% threshold 21 Bedouins and 18 Mozabites qualified. In the initial run, GAL shared 84.5% European, 13.5% Middle Eastern, 2.0% sub‐Saharan ancestry, and 0.1% East Asia, compared with IBS (87.0%, 11.6%, 1.2%, and 0.1%), and BAS (92.2%, 7.8%, 0.0%, and 0.0%). In the second run, GAL showed 83.2%, 7.4%, 8.7%, 0.6%, and 0.1% from Europe, the Middle Eastern, North Africa, sub‐Saharan Africa, and East Asia respectively, versus IBS (86.9%, 6.7%, 5.9%, 0.4%, and 0.1%), and BAS (94.9%, 4.0%, 1.1%, 0.0%, and 0.0%). Therefore, the North African/Middle Eastern component in GAL ranged from 13.5% to 16.1% (mean: 14.8%), higher than in IBS (7.8%–12.6%, mean: 8.25%), and BAS (5.1%–7.8%, mean: 6.45%). Differences among all datasets were statistically significant (Supporting Information Data S1 Figure S2 ). A minor yet statistically significant sub‐Saharan ancestry was also detected in GAL (1.0%–2.8%, mean: 1.9%), higher than in IBS (0.3%–1.2%, mean: 0.75%), and absent in BAS. Interestingly, although the North African/Middle Eastern ancestry was largely homogeneous across Galicia, we observed a subtle South‐to‐North gradient, with slightly higher proportions in southern regions (Figure 2B ).
According to fastGLOBETROTTER , Mozabite‐like ancestry contributed 6.9%, while Middle Eastern‐like ancestry contributed 4.6% to GAL, yielding a combined 11.5%, slightly lower than ADMIXTURE estimates.
f
3 ‐statistics confirmed that GAL exhibits admixture signatures consistent with gene flow between European and African sources (Figure 2C ), with several f
3
(X EUR , X NAF/SSH ; GAL) tests yielding significantly negative values.
Finally, mitochondrial DNA (mtDNA) analysis showed that 1.1% of the samples belong to North African haplogroups (U6), with no Sub‐Saharan (L) lineages detected. In contrast, Y‐chromosome data from GAL males ( n = 33) revealed 21.2% (E1b1b1) lineages of North African/Middle Eastern origin.
Time estimates from fastGLOBETROTTER indicate the occurrence of an admixture event in the Galician population that is best explained as gene flow between a predominantly European ancestral group and a more admixed African/Middle Eastern group. The inferred minor source shows a particularly strong North African signature, represented by Mozabites (MOZ), contributing approximately 42%, alongside a smaller Middle Eastern component, represented by Bedouins (BED), accounting for about 19% of the same source. The overall admixture event is dated to approximately 54 generations ago, which translates to a calendar range of around 510–690 ce , assuming a generation time of 26.1 years.
A two‐date admixture model was also tested, but it did not provide a better fit than the one‐date model, supporting the hypothesis of a single, temporally localized admixture event. The admixture signals were further characterized by the presence of a sub‐Saharan African component, with YRI acting as a proxy. While this component was less prominent, it may reflect either a genuine signal of sub‐Saharan input or result from model limitations in distinguishing between shared North African and sub‐Saharan ancestry.
To validate and complement the findings from fastGLOBETROTTER , admixture dating analyses were also performed using ALDER . In agreement with fastGLOBETROTTER , this software estimated the admixture involving North African ancestry (using MOZ as a proxy) to have occurred approximately 50 generations ago, and admixture with Middle Eastern ancestry (using BED as a proxy) around 54 generations ago. In addition, ALDER identified admixture involving sub‐Saharan African ancestry (using YRI and LWK as proxies) at approximately 57 generations ago, which corresponds to around 540 CE .
Taken together, both fastGLOBETROTTER and ALDER yield overlapping time estimates for the introduction of North African and Middle Eastern ancestry into Galicia. For interpretative clarity, we adopt the central point estimates of both methods, 52 generations from ALDER and 54 from fastGLOBETROTTER , to define a reference interval of admixture of 620–670 ce .
We also estimated the timing of the introduction of the North African/Middle Eastern genetic component into various Iberian regions using the NDNAB dataset. While these estimates must be interpreted with caution due to the dataset's relatively low SNP density and limited sample sizes, they nonetheless reveal informative regional patterns. The inferred admixture dates (in generations, with 95% confidence intervals) were as follows: Andalusia/Central Iberia: 33 (95% CI: 31−36), Catalonia: 39 (95% CI: 35−45), Basque Country/Navarra: 55 (95% CI: 31−70), Cantabria: 46 (95% CI: 41−61), and Asturias: 51 (95% CI: 46−55). Notably, Galicia showed an admixture time of 43 generations in the NDNAB dataset, compared to 52 generations when using the higher‐resolution WGS data from the GAL cohort. This discrepancy suggests that the NDNAB estimate may underestimate the true admixture date due to its lower resolution. Despite the inherent limitations, the overall trend across regions supports an earlier admixture in northern Iberia compared to the South and the Center. To further validate this signal, we merged the northern regions (Basque Country/Navarra, Cantabria, Asturias, and Galicia; n = 164) to increase statistical power and compared them with the Andalusia/Central Iberia group ( n = 191). This analysis yielded an admixture estimate of 48 generations (95% CI: 45−51) for northern Iberia, significantly older than the 33 generations estimated for the South–Central regions. These findings suggest that the introduction of North African/Middle Eastern ancestry into Iberia may have occurred in multiple waves or via regionally distinct demographic processes, with earlier influxes into the North than previously assumed.
At the highest level of genome resolution (WGS), fineSTRUCTURE analysis of the GAL dataset identified 10 terminal branches, though only five formed clusters containing at least four individuals (Figure 3A ). The Galician genetic landscape is predominantly shaped by a single major cluster, referred to here as the “Main” cluster, comprising 66 genomes and uniformly distributed across most of the region (Figure 3B ).
FineSTRUCTURE analysis of GAL whole genomes. (A) Dendrogram depicting the highest level of phylogenetic resolution. (B) PCA of fineSTRUCTURE clusters, illustrating genetic differentiation among individuals. (C) Interpolated frequency map showing the geographic distribution of the identified clusters within the GAL dataset.
The remaining four clusters are significantly smaller and more localized. The “Porto do Son” cluster (west Coast, n = 8) shows the most pronounced genetic isolation from the broader Galician population. The “Cuntis/A Estrada” cluster (inner Pontevedra, n = 6) and the “A Mariña/Foz” cluster (northeast Coast, n = 4) are also geographically distinct, with the latter located along the northern coastland. Lastly, a small and scattered “Central/West” cluster ( n = 5) is dispersed throughout central and West Galicia (Figure 3A,B ).
Interpolated frequency maps further emphasize the dominance of the “Main” cluster, while the “Porto do Son” cluster stands out distinctly (Figure 3B ). The “Cuntis/A Estrada” and “A Mariña/Foz” clusters are less visible due to their blending with the “Main” cluster in their respective areas. Meanwhile, the “Central/West” cluster, being both small and dispersed, is further diluted by the influence of the dominant “Main” cluster. This overall pattern highlights the remarkable homogeneity of the Galician genetic landscape, with only limited instances of localized isolation, most notably in the village of Porto do Son.
The PCA of the fineSTRUCTURE results (Figure 3C ) also reveals the five clusters described above. PC1 (21.4% variance explained) places the “Porto do Son” and “A Estrada/Cuntis” clusters at the opposite ends of the axis, yet both align on the same pole of PC2 (15.5%), distinguishing them from the remaining clusters. PC3 (8.0%) positions the “Central/West Galicia” and “A Mariña” at opposite extremes, while PC4 (5.2%) further separates these two clusters from the rest. The clustering pattern shown by PC1 to PC4 does not strictly follow the phylogenetic structure suggested by the phylogenetic dendrogram, where, for example, the deepest split occurs between the “Central/West” cluster and the rest. However, it demonstrates a distinct segregation of individuals into five genetically differentiated clusters.
The low genetic structure revealed by WGS data in Galicia contrasts sharply with the highly stratified landscape reported by Bycroft et al. [ 7 ] in their study of the Iberian Peninsula, particularly Galicia, which was based on lower‐resolution genomic data. These authors reported identifying 145 genetic clusters among Spanish donors, nearly half of which (~70) were concentrated in a small, densely populated area within the Galician province of Pontevedra. Remarkably, a closer examination of their data and maps reveals that Galicia accounts for 107 of the 145 clusters (Pardo‐Seco et al. unpublished). Although this specific area of Pontevedra was more densely sampled in Bycroft et al. [ 7 ] than in the GAL study, the observed discrepancy in stratification levels does not appear to be solely attributable to differences in sampling density (Pardo‐Seco et al. unpublished).
In light of the sharp contrast in clustering patterns, we performed a fineSTRUCTURE analysis on the NDNAB dataset ( n = 453, with 474 761 SNPs after applying the same filters as for the whole‐genome data). This dataset includes independent Galician samples alongside donors from various Spanish regions, enabling a detailed assessment of Galicia's population structure within a national context. NDNAB encompasses individuals from Galicia, Asturias, Navarra, Cantabria, the Basque Country, Catalonia, Andalusia, Castilla‐La Mancha, and Castilla y León, all regions also examined in Bycroft et al.'s study [ 7 ]; Figure 4A . Consistent with our findings in the GAL dataset, fineSTRUCTURE identified only three clusters among Galician donors in the NDNAB dataset. Similarly, one cluster contained most of the Galician samples ( n = 46), mirroring the “Main” GAL cluster, while the remaining two minor clusters comprised 7 and 3 individuals, respectively. In the broader NDNAB Iberian dataset, a total of 23 clusters were detected (including the Galician ones), each closely reflecting the geographic origin of the donors (Figure 4B ; Supporting Information Data S1 Figure S3 ). Unlike Bycroft et al. [ 7 ] the first major split in the phylogenetic tree of our NDNAB Spanish dataset separates clusters exclusively found in the northern and geographically adjacent region of Cantabria (three clusters) and the Basque Country/Navarra (two clusters) from the rest of the Iberian Peninsula. The next phylogenetic subdivision separates Asturias (six clusters, with one dominant cluster dominating most individuals in the region [ n = 36]), followed by Catalonia (with a unique and well‐differentiated cluster; n = 94). Galician genomes emerge at the most recent branching point of the phylogeny, clustering alongside a mixed group of individuals from Andalusia and Central Iberia. This group comprises seven clusters, with three major ones including 128, 35, and 32 individuals, respectively. This clustering pattern differs markedly from that reported by Bycroft et al. [ 7 ], where Galicians from Pontevedra, their hyper‐highly stratified Galician province, exhibited the earliest genetic differentiation and accounted for about half of the total clusters found in Iberia. In their analysis, the next major split separated the Basque Country individuals from the rest of the Iberian Peninsula, followed by a division isolating the remaining Galician samples in their dataset.
Geographic, phylogenetic, and clustering structure of the NDNAB dataset, shown through sampling maps, phylogenetic and dendrogram analyses, PCA, and interpolated frequency maps. (A) Map illustrating the sampling locations of the NDNAB dataset, which includes Iberian donors whose grandparents were born in the same provinces. Circle sizes are proportional to sample sizes and are centered on the main provincial capital cities. (B) Circular phylogeny (top‐left) of the NDNAB dataset (fully expanded in Supporting Information Data S1 Figure S3 ) and a dendrogram (right), where terminal branches are collapsed into nodes representing major geographic regions in Iberia (except Galicia), corresponding to the main autochthonous populations. Unlike other regions, the three Galician clusters were kept separate to facilitate direct comparisons with those observed in the GAL dataset. The proportions indicated to the right of the regional names represent the percentage of samples in each cluster originating from the corresponding autochthonous region, acknowledging that some individuals were sampled in other (often neighboring) regions. Additionally, the total sample size, the number of terminal branches (quantities separated by hyphens), and their respective sample sizes are provided (e.g., fineSTRUCTURE detected three clusters in Cantabria, with sizes 14, 8, and 3 individuals). Symbols in the collapsed nodes of the dendrogram correspond to those used in panels C and D. (C) PCA of NDNAB donors, illustrating the clustering pattern observed in (B). (D) Interpolated frequency maps of the continental Spanish territory, depicting the geographic distribution of clusters identified in (B); the numbers following the region names indicate the maximum number of fineSTRUCTURE clusters represented in the maps. For the specific case of the Andalusia/Central region, we show the distribution of the three main clusters in this area (bottom maps), with sample sizes of 128, 35, and 32, each displaying slightly different geographic distributions.
The PCA analysis of the fineSTRUCTURE clusters found in the NDNAB database (Figure 4C ) further supports the phylogenetic structure observed in the dendrogram (Figure 4B ). The primary separation in PC1 (8.3%) distinguishes the Cantabria/Basque Country/Navarra clusters from the rest of the dataset. PC2 (4.1%) further differentiates the Basque Country/Navarra from the remaining regions. PC3 (2.7%) highlights the genetic distinction between Asturias and Cantabria, while PC4 (2.1%) again emphasizes the differences among these northern Iberian regions. Notably, all four principal components successfully capture and distinguish the clusters identified in the dendrogram, reinforcing the observed population structure.
Analyzing inbreeding levels and ROH segments can reveal events of demographic isolation and offer insights into a population's demographic history and genetic diversity. This study presents the first investigation of ROH patterns in the Galician population, enabling comparisons with other European and non‐European groups at the highest level of genetic resolution possible, namely, that provided by WGS data.
The Galician population (GAL) exhibits ROH patterns comparable to those observed in other European populations (Figure 5A ). However, a more detailed analysis reveals slightly higher levels of inbreeding compared to the national Spanish dataset represented by the Iberian Peninsula (IBS). Specifically, GAL has a larger count of long ROH segments, a longer mean of these segments, and a larger cumulative length compared to IBS ( nROH
L
[GAL] = 9 vs. nROH
L
[IBS] = 8, p ‐value = 1.5 × 10 −1 ; mROH
L
[GAL] = 2.4 vs. mROH
L
[IBS] = 2.1, p ‐value = 5.2 × 10 −8 ; sROH
L
[GAL] = 20.6 vs. sROH
L
[IBS] = 15.2, p ‐value = 9.2 × 10 −4 ); Table S1 . In line with these findings, GAL also has a lower count of short ROH segments, a shorter mean of these segments, and a smaller cumulative length compared to IBS ( nROH
S
[GAL] = 2504 vs. nROH
S
[IBS] = and 2592, p ‐value = 5.8 × 10 −19 ; mROH
S
[GAL] = 0.199 vs. mROH
S
[IBS] = 0.203, p ‐value = 9.2 × 10 −19 ; sROH
S
[GAL] = 499 vs. sROH
S
[IBS] = 525, p ‐value = 3.4 × 10 −26 ); Table S1 . Overall, these metrics support the observations of significantly higher inbreeding levels in GAL relative to IBS and other European datasets, except when comparing long ROH segments in historically isolated populations (Basques and Sardinians), which show higher values. Additionally, total inbreeding coefficient values indicate higher inbreeding in GAL compared to IBS ( F
ROH
[GAL] = 0.007 vs. F
ROH
[IBS] = 0.005, p ‐value = 9.3 × 10 −4 ) and other continental European datasets, although still lower than in Sardinians and Basques; Table S1 .
Analysis of inbreeding in GAL genomes and reference worldwide population datasets. (A) Various ROH estimates indicate that GAL shows patterns consistent with those expected for a dataset of predominantly European ancestry. (B) F ‐statistic analysis conducted on European ancestry datasets. Further details on the different statistical measures are provided in Table S1 and Data S2 . (C) Interpolated frequency maps showing the distribution of different ROH values across GAL genomes.
In agreement with ROH values, Galicia (GAL) has higher values of F
HET
, F
HAT1
and, F
HAT3
compared to IBS ( F
HET
[GAL] = 0.005 vs. F
HET
[IBS] = 0.003, p ‐value = 6.4 × 10 −2 ; F
HAT1
[GAL] = 0.012 vs. F
HAT1
[IBS] =0.0001, p ‐value = 2.7 × 10 −16 ; and F
HAT3
[GAL] = 0.008 vs. F
HAT3
[IBS] =0.0002, p ‐value = 5.2 × 10 −9 ); and equal values of F
HAT2
for both populations ( F
HAT2
[GAL and IBS] = 0.004); Figure 5B , Table S3 . Within Europe, GAL's metrics are intermediate, with Sardinians and Basques showing the highest homozygosity values among these datasets.
Additionally, the “Porto do Son” cluster significantly contributes to the higher inbreeding values in GAL compared to IBS. For instance, metrics for long ROH segments in this cluster are substantially higher (and statistically significant) than those for the general GAL: nROH
L
[Porto do Son] = 11; mROH
L
[Porto do Son] = 3.5; and sROH
L
[Porto do Son] = 45. Conversely, short ROH segment metrics are generally lower: nROH
S
[Porto do Son] = 2442; mROH
S
[Porto do Son] = 0.199; sROH
S
[Porto do Son] = 487; Table S1 . In fact, the inbreeding values for the GAL dataset excluding “Porto do Son,” are similar to those observed in the other clusters (GAL: nROH
L
[GAL without Porto do Son] = 9; mROH
L
[GAL without Porto do Son] = 2.6; and sROH
L
[GAL without Porto do Son] = 26; nROH
S
[GAL without Porto do Son] = 2502; mROH
S
[GAL without Porto do Son] = 0.199; sROH
S
[GAL without Porto do Son] = 497). The frequency‐interpolated maps in Figure 5C make the distintive singularity of higher inbreeding values in “Porto do Son” more visible.
Overall, these results indicate that inbreeding levels in Galicia are broadly comparable to those observed elsewhere on the Iberian Peninsula, with only marginally elevated values in certain localities, such as Porto do Son.
To assess allele differences of potential biomedical relevance, we performed a single‐site association test comparing the GAL cohort to the IBS dataset. This analysis was restricted to SNPs with a MAF > 0.05 in cases and controls, as the focus was solely on common variants. A total of 85 SNP variants (out of > 17 M) surpassed the Bonferroni threshold (Supporting Information Data S1 Figure S4 ; Table S2 ). Among these, only one SNP (rs9917044) was located in an exonic region ( ZNF28 ; Zinc Finger Protein 28 gene) and was synonymous. Additionally, 25 variants were found in intronic regions, 13 were noncoding but classified as regulatory, and 46 were intergenic. When these variants were compared against a merged European cohort (CEU, GBR, IBS, and TSI) and using a MAF > 0.05 in cases and controls, 57 remained significant after applying the Bonferroni threshold. None of the 85 SNP variants appear to be associated with a disease trait in SNPedia. The gene‐based test between GAL and IBS identified only four genes as statistically significant after Bonferroni correction: ATP5F1AP10 (two SNPs), SDR42E1P4 (two SNPs), FRG2EP (17 SNPs), and DUX4L34 (13 SNPs); only the first two genes (supported on only two SNPs) remained significant when compared to the merged European cohort (Supporting Information Data S1 Figure S5 ).
By computing PRSs for eleven common diseases, we investigated whether demographic particularities in Galicia, compared to the broader Iberian (IBS) population, influence the distribution of alleles associated with disease risk in GAL genomes. For most PRS metrics, no statistically significant differences were observed between GAL and IBS. However, despite their small quantitative effect, the differences were statistically significant for PRS ASD (IBS median = −0.22 [interquartile range (IQR): −0.68–0.53], GAL median = 0.15 [IQR: −0.41–0.71]; p ‐value = 0.038) and PRS IBD (IBS median = −0.25 [IQR: −0.89–0.50], GAL median = 0.25 [IQR: −0.62–0.86]; p ‐value = 0.007); Supporting Information Data S1 ; Figure S6 . The wide IQR in these metrics suggests substantial variation among individuals, indicating that some donors have markedly higher or lower polygenic risk relative to others in the group.
To further investigate the microgeographic stratification of disease risk within Galicia, we analyzed the spatial distribution of PRS metrics for the eleven common diseases across the region. The maps in Figure 6A illustrate distinct geographic patterns of risk. Ovarian cancer risk is more pronounced along the western Atlantic coastline, peaking in the southwest. Breast cancer risk varies slightly by PRS but is generally elevated in both eastern and western Galicia and lower in a central North‐to‐East longitudinal band. Coronary artery disease and inflammatory bowel disease exhibit broadly similar distributions. Atrial fibrillation risk increases northward. Type 2 diabetes risk peaks in central Galicia and is lowest in the southeast. Acute lower respiratory infection risk is lowest in the northwest and highest in the northeast. Autism spectrum disorder exhibits intermediate risk from mid‐central to northern Galicia, with the lowest in the southwest and the highest in the southeast. Schizophrenia risk is highest in the southeasternmost corner and remains elevated across most of the region, except in the northwest. Alzheimer's disease follows an approximately inverse pattern.
Geographic distribution and cluster‐level variation of PRSs for common diseases across Galicia, shown through interpolated maps and median PRS comparisons. (A) Interpolated maps showing the distribution of polygenic risk for common diseases across the Galician territory. (B) Median PRS values for various common diseases across the fineSTRUCTURE clusters, ordered from lowest to highest risk. Note that median values do not fully capture the variability of PRSs distributions (which can be seen in detail in Supporting Information Data S1 Figure S7 ) and may not directly correspond to the interpolated maps in (A). These estimates are preliminary and may vary with larger sample sizes in future analyses, but they establish a foundation for further academic exploration. The top panel in (B) provides a graphical representation of the geographic locations of the main Galician clusters, as inferred from Figure 3 .
We further examined disease risk stratification by genetic cluster. Although the sample sizes for most clusters (except for the “Main” cluster) were limited, and the different PRSs encompass a wide range of variant counts (ranging from dozens to millions), we identified statistically significant differences in several comparisons (Supporting Information Data S1 Figure S7 ). The most notable findings include: (a) Coronary artery disease risk: Significant differences were observed between “Porto do Son” and “A Estrada/Cuntis” ( p ‐value = 0.016), “A Estrada/Cuntis” and “A Mariña/Foz” ( p ‐value = 0.019), and “A Estrada/Cuntis” and the rest of Galicia ( p ‐value = 0.038); (b) Inflammatory bowel disease risk: Significant differences were found between “A Estrada/Cuntis” and the “Main” cluster ( p ‐value = 0.026), as well as between “A Estrada/Cuntis” and the rest of Galicia ( p ‐value = 0.040). Additionally, differences were observed between the “Main” cluster and “A Mariña/Foz” ( p ‐value = 0.025), and between “A Mariña/Foz” and the rest of Galicia ( p ‐value = 0.038); (c) Atrial fibrillation risk: Significant differences were detected between the “Main” cluster and “A Mariña/Foz” ( p ‐value = 0.007), and between “A Mariña/Foz” and the rest of Galicia ( p ‐value = 0.008); (d) Alzheimer's disease risk: Differences were observed between the “Main” cluster and “A Mariña/Foz” ( p ‐value = 0.012), the “Main” cluster and the rest of Galicia ( p ‐value = 0.007), and between “A Mariña/Foz” and the rest of Galicia ( p ‐value = 0.027); and (e) Autism spectrum disorder risk: A significant difference was found between “A Estrada/Cuntis” and “A Mariña/Foz” ( p ‐value = 0.019). Therefore, of the 13 statistically significant contrasts identified, 8 involved the “A Mariña/Foz” cluster and 6 involved the “A Estrada/Cuntis” cluster. Despite their relatively small sample sizes, these two clusters account for a substantial proportion of the major disease risk stratification observed in our Galician sample.
In addition, PRS metrics for the different clusters can be ranked according to the diseases considered. For instance, the “Porto do Son” cluster has the highest risk for ovarian cancer and the lowest for schizophrenia, the “A Mariña/Foz” cluster has the highest risk for atrial fibrillation but the lowest for coronary artery disease, while the “A Estrada/Cuntis” cluster has the lowest risk for atrial fibrillation and the highest risk for inflammatory bowel disease (Figure 6B ).
Discussion
Galicia's geographic position at the far western edge of the Iberian Peninsula, combined with its historical isolation, has led some studies to hypothesize that the region may harbor unique genomic features [ 7 , 14 , 15 , 16 , 52 , 53 ]. WGS data from Galicia allow an in‐depth investigation at the highest possible genomic resolution, revealing patterns of genetic variation that are broadly shared with other Iberian populations, yet also display some distinct regional traits. The genetic landscape of the Galician population reflects a demographic history consistent with early settlement, subtle differentiation from the broader Iberian Peninsula, and an early admixture with North African and Middle Eastern ancestries. PCA and fineSTRUCTURE clustering reveal that Galicians form specific but closely related genetic clusters, showing only limited divergence from other Iberian groups.
Galicia demonstrates slightly higher values of nROH
L
, sROH
L
, and mROH
L
, compared to IBS and other European datasets. This pattern extends to shorter ROH segments, resulting in a higher total inbreeding coefficient ( F
ROH
) than in IBS and other Europeans. Accordingly, inbreeding coefficients indicate slightly higher overall homozygosity in Galicia compared to IBS and other European datasets. While statistically significant differences with the rest of the Peninsula exist, they are quantitatively modest and largely driven by the distinct “Porto do Son” cluster (as also inferred from PCA plots and spatial genetic maps). The slightly higher inbreeding coefficients in Porto do Son likely reflect historical geographic isolation and elevated consanguinity. Marriage patterns within small coastal or rural communities, combined with local limited gene flow, can increase genetic homogeneity over generations. Nevertheless, even in Porto do Son, inbreeding levels remain within the Iberian norm, indicating that Galicia's demographic and cultural factors produce autozygosity patterns similar to those of neighboring regions. Therefore, the slightly elevated inbreeding levels observed in Galicia challenge the hypothesis of pronounced population stratification in the region (against Bycroft et al. [ 7 ]). If strong stratification were present, one would expect allele frequency disparities within Galicia to result in numerous genetic clusters and elevated homozygosity levels consistent with a Wahlund effect, a pattern not observed in this study. The minor divergence existing between GAL and the rest of the Iberian Peninsula most likely resulted from genetic drift and small‐scale fluctuations in allele frequencies across the genome. There is no compelling evidence to suggest that long‐standing geographic isolation is the primary factor shaping the genetic landscape of Galicia.
The relatively elevated African ancestry component existing in Galicia, greater than in other parts of Iberia, may have played a role in mitigating levels of inbreeding through increased outbreeding, thereby contributing to the preservation of genetic diversity in Galicia.
The limited number of clusters observed in Galician whole genomes, along with the moderate levels of inbreeding, suggests a demographic history marked by extensive gene flow, both within Galicia and from genetically distinct, non‐Galician populations, including groups from North Africa and the Middle Eastern. Such a pattern aligns with Galicia's geographic and socio‐historical context, characterized by the absence of major geographical barriers and a predominantly rural, dispersed demographic structure. These factors likely promoted both intra‐regional mobility and external genetic exchange, helping to preserve high genetic diversity and limiting inbreeding across the relatively small geographical area of Galicia. While certain areas in Galicia may have experienced localized demographic conditions conducive to increased inbreeding or mild founder effects, as previously documented for mutations causing autosomal recessive congenital ichthyosis [ 52 ] and other inherited disorders, our data reveal a markedly lower degree of fine‐scale population structure than that reported by Bycroft et al. [ 7 ]. Their study, based on low‐density genome‐wide data, identified ~70 genetic clusters recognized by these authors. However, as critically reanalyzed in Pardo‐Seco et al. (unpublished data), the actual number of clusters in their data appears to exceed 107 across the Galician territory out of 145 observed in the whole Iberian Peninsula, accounting for ~74% of the total. These clusters are heavily concentrated in a small southwestern section of Pontevedra Galician province, near the border with northern Portugal. This strikingly high number of clusters is particularly unexpected given the province's relatively small size, its high population density (Pontevedra is among the most densely populated provinces in Spain, with over 200 inhabitants per square kilometer; https://ide.depo.gal ), and its historical, cultural, and religious significance. Notably, the city of Tui in this area has long served as a key historical religious center in both Galicia and northern Portugal. The discrepancy is even more pronounced considering that this southwestern region is one of the most cosmopolitan areas in Galicia, home to Vigo (the province's largest urban center and Galicia's most populous city), and one of the world's busiest fishing ports. Moreover, the region lacks any substantial geographical barriers that might explain deep population substructure. The sharp discrepancies between our results and those in Bycroft et al. [ 7 ] could have been due to technical artifacts (Pardo‐Seco et al. unpublished data), amplified by the high sensitivity of the fine
STRUCTURE tool, as well as flawed misinterpretations of the data.
Several studies (most of them focusing on uniparental markers) suggest that Galicia has a proportion of North African genetic ancestry that is comparable to or even exceeds that of other Iberian regions (Supporting Information Data S6 ), despite its geographic distance from North Africa. Using high‐resolution WGS data, we estimate that the North African/Middle Eastern genomic component in GAL ranges from 13.5% to 16.1%, significantly higher than previously reported for autosomal data and exceeding estimates for other Iberian datasets (IBS: 7.8%–12.6%; BAS: 5.1%–7.8%). Approximately half of this component is shared with “Mozabite‐like” ancestry, with Mozabites (MOZ) representing Berbers from Algeria's M'Zab Valley. The remaining proportion is attributed to Middle Eastern ancestry, primarily represented by the Bedouin (BED) population, nomadic Arab groups historically associated with the Syrian and Arabian deserts, who expanded throughout the Levant, Mesopotamia, and North Africa following the Islamic conquests of the 7th century. This African ancestry is present throughout the entire Galician territory, implying long‐term, extensive gene flow and demographic integration across the region. When examining the spatial distribution of this component using interpolated individual frequencies, a subtle yet consistent South‐to‐North gradient of North African and Middle Eastern ancestry emerges, suggesting a southern entry point into Galicia via maritime or overland routes from the southwest. This pattern warrants further investigation in future studies with larger sample sizes and comparable spatial and genomic resolution. A novel and noteworthy observation is that the North African component is much more pronounced in the Y‐chromosome (21.2%) and autosomal genome (13.5%–16.5%) than in the mitochondrial genome (1.0%), suggesting sex‐biased admixture with a predominant paternal contribution. This male‐skewed pattern is consistent with broader trends observed throughout the Iberian Peninsula, likely reflecting the historical context of migrations before, during, or after the Islamic expansion, where male‐dominated incursions and settlements were common.
The detection of sub‐Saharan ancestry may suggest an even more complex demographic dynamic than the North African/Middle Eastern–European admixture process. This could include either actual sub‐Saharan input or model misattribution due to shared drift or admixture history between “Mozabites‐like” ancestry and sub‐Saharan groups like YRI. We detected a low but statistically significant sub‐Saharan African component in both GAL (1.0%–2.8%; mean: 1.9%) and IBS (0.3%–1.2%; mean: 0.75%), whereas this component was absent in the BAS dataset. Given that the reference populations used in this study are essentially unadmixed, this sub‐Saharan component remains undetectable when analyzed within North African or sub‐Saharan datasets. Furthermore, the minimal representation of this component in Iberian datasets, combined with the limited availability of comprehensive African WGS data, presents a challenge for further investigation into its precise origins. The existence of this sub‐Saharan component in Iberia has been particularly reported in the literature on uniparental markers. A meta‐analysis of mtDNA variation in Iberia [ 54 ] identified the presence of a characteristic sub‐Saharan component in various Iberian regions, with proportions ranging from 0% to 5% for example, Asturias: 3.6%; Galicia: 2.9%; La Rioja: 5.0%; with a mean of 1.4% in mainland Spain. The introduction of typical sub‐Saharan genetic variation into Europe, particularly in Iberia, has been explored by Cerezo et al. [ 55 ]. Phylogeographic analyses of mtDNA haplotypes suggest that ~65% of European L‐lineages (indicative of sub‐Saharan mtDNA ancestry) likely arrived in Iberia during historical periods, including Romanization, Arab rule, and the Atlantic slave trade. The remaining 35% of European‐specific subclades of L mtDNA may have originated from gene flow between sub‐Saharan Africa and Europe as early as 11 000 years ago. Additionally, the presence of sub‐Saharan mtDNA components has been documented in Iberian Romani populations [ 56 ]. More recently, ancient DNA studies further support the introgression of early sub‐Saharan populations into Iberia, providing evidence of an early prehistoric migration route from Africa to Iberia, likely via the Strait of Gibraltar [ 57 ].
The admixture signals detected by GLOBETROTTER and ALDER point to a historically significant episode of gene flow into Galicia, likely occurring between 620 and 670 ce . The predominant African‐related ancestry in this event appears to derive from North African populations, with a smaller but noteworthy Middle Eastern component. The presence of sub‐Saharan genetic signatures, though more modest, may reflect earlier contact or ancestral gene flow mediated through North African intermediary populations, consistent with patterns of trans‐Saharan connectivity. Moreover, our admixture timing estimates reveal a striking pattern within the Iberian Peninsula: the North African and Middle Eastern genetic components appear to have entered northern Iberia significantly earlier than in central and southern regions. Notably, after accounting for the potential underestimation of admixture dates, the pooled northern regions (including the Basque Country/Navarra, Cantabria, Asturias, and Galicia) show an average admixture date around 770 ce (range: 690–880 ce ). In contrast, estimates for Andalusia and central Iberia are more recent, averaging around 1160 ce (range: 1110–1240 ce ).
While the southern admixture signal might align with the well‐documented Arab‐Berber expansion during the Islamic conquest and subsequent centuries of rule, the earlier dates in the North (though requiring cautious interpretation) challenge this conventional narrative and suggest additional independent episodes of contact and gene flow. Overall, these findings highlight a notable North African/Middle Eastern genetic influence on the Iberian Peninsula, particularly in the North, and specifically in the northwestern region of Galicia, predating the Islamic rule. Our data estimate the arrival of this component in a single, concise admixture event occurring between 620 and 670 ce , though statistical limitations, methodological constraints, or, more likely, the demographic complexity of this ancestral signal may mask earlier episodes dating back to the Roman period or even in (pre‐)historic times [ 55 , 57 ]. One could, for instance, hypothesize a steady, small‐scale migration, trade networks, and/or a multifaceted influx of this ancestry into Galicia over several centuries, predating the Islamic conquest, with a peak occurring during the 6th to 7th century, which would give rise to the signal detected in our model. Investigating the pre‐Islamic North African‐Galician is key to understanding this legacy, given Islam's limited northern impact.
Numerous historical links between North Africa and Galicia date to the Roman era. One such link is the Cohors I Celtiberorum , stationed first in the Roman province of Mauretania Tingitana (present‐day northern Morocco) in 132 ce (end of Trajan's reign), then later posted to Cidadela , in Sobrado dos Monxes (A Coruña, Galicia), where a garrison of ~500 members was accommodated. Likely, many soldiers settled locally upon discharge [ 58 , 59 ], and the cohort remained in Cidadela until around the 4th century. Through this movement, some members may have introduced African ancestry to Galicia. More importantly, Roman military units like the Cohors I Celtiberorum rarely moved alone: they were typically accompanied by civilians (families, merchants, craftsmen, and service providers), who supported the soldiers' logistical needs and helped establish semipermanent settlements near camps. Evidence of such settlements (referred to as vicus , insulae , or even vicus ) exists in Cidadela . Curiously, a nearby town still preserves the name “Insua,” possibly reflecting this legacy. It is therefore plausible that accompanying North African civilians contributed to the local demographic and genetic landscape, facilitating not only military but also cultural, demographic, and economic exchanges between North Africa and northwestern Iberia.
Additionally, there is evidence of well‐established Atlantic trade networks, centered on salted fish, garum, and especially cereals transported in amphorae (e.g., present‐day Tunisia), that connected North African ports to the Galician coast via the Balearic Islands, the Iberian Levant, and the Strait of Gibraltar, bringing with them merchants, sailors, and laborers of North African origin [ 60 , 61 , 62 ]. Although much of this commerce likely had a primarily cultural rather than demographic impact, these intertwined military and commercial currents may have contributed to a North African genetic legacy in Galicia centuries before any Islamic political or cultural influence. Another example from Roman times is the presence of the Legio VII Gemina , established in the North of Iberia (part of the Gallaecia province), which included many recruits from Africa [ 63 ].
Historical records from the Late Antique period also show North African and Middle Eastern influence in Galicia, from the end of Roman rule in the northwest of Iberia to the onset of the Islamic era (i.e., 5th to 7th centuries CE). During this time, the Iberian Peninsula underwent significant transformations following the collapse of Roman authority and the emergence of new political and religious structures, particularly under the Visigoth Kingdom. While slavery had long been a feature of Roman society, the slave trade not only persisted but intensified during this period, often taking on new religious and economic dimensions. Churches, monasteries, and civil elites, especially in northwestern Iberia (modern‐day Galicia), became prominent economic and spiritual centers after the 5th century [ 64 ]. These institutions, especially monasteries and churches, relied heavily on enslaved labor for agriculture, construction, and domestic service. Ecclesiastical institutions and the aristocracy were large landowners who needed manpower, most of whom were slaves, designated as servi in the texts of the time. Around 655 CE , Richimirus , Bishop of Dume (predecessor to Fructuosus ) freed numerous slaves in his will and donated > 500 slaves, some of whom came from his personal property [ 64 , 65 ]. This source provides insight into the scale of slave arrivals in the northwest during the 6th and especially the 7th centuries. By this time, the Iberian Peninsula was linked to broader Mediterranean/Atlantic trade and slave routes, including those connecting North Africa and the Eastern Mediterranean. Slaves of African or Middle Eastern origin were often brought into Iberia via maritime and overland routes (through Carthage, Ceuta, or southern Hispania). Historical sources show that the trade flowed in both directions; many slaves were Britons and Saxons sold in markets in Rome and Constantinople [ 62 ]. These enslaved individuals may have been prisoners of war, victims of raids, or traded through expanding networks controlled by Byzantine, Berber, or Visigothic intermediaries. Contemporary texts also describe slaves brought to Gaul to supply churches and major landowners, including individuals of Mauritanian origin [ 66 ]. Evidence suggests that specialized traders, possibly of Eastern origin, such as Syrians, Greeks, and Jews, dominated this market, organizing systematic raids to capture slaves in the Mauritanian territories, beyond the reach of contemporary political and administrative control [ 66 ]. It is also plausible that a significant community of Eastern merchants settled in the area during the 6th and 7th centuries. This community, likely connected to the Visigothic Church and monarchy, oversaw the flow of goods between the Mediterranean and Galicia and may have played a key role in managing this highly profitable slave trade [ 62 , 67 ].
Finally, recent genomic analyses of the remains believed to be those of Bishop Teodomiro of Iria Flavia (Galicia), born in the late 7th century, indicate that approximately one‐fifth of his genome is of North African origin; possibly up to 30% (Ricardo Pérez‐Varela, personal communication). The authors of the study stated that “…their ancestors may have received North African gene flow during the Roman period or, more recently, during or after the Islamic conquest” [ 68 ]; see also [ 67 ]. Also, recent findings from Oteo‐García et al. [ 69 ] support the idea of substantial pre‐Islamic North African genetic influence in Iberia. These authors reconstructed the medieval genetic history of eastern Iberia using ancient DNA from the Valencian region. Their analysis, focused on the effects of the Morisco deportations, shows that North African–related ancestry was already present in medieval Valencia and remained largely unchanged after the Reconquista .
Taken together, this body of evidence challenges the traditional view that North African and Middle Eastern genetic influence in northern Iberia resulted solely from Islamic rule and the Reconquista ; contra [ 7 ]. Instead, it suggests that earlier, independent episodes of gene flow may have contributed to this legacy. Potential factors shaping the demographic landscape of northwestern Iberia include prehistoric contacts, movements during the Roman period, post‐Phoenician trade networks, and trans‐Mediterranean contacts predating the 8th‐century conquest, including the trafficking of enslaved individuals. Religious institutions, particularly monasteries, likely played a central role in this process, acting both as recipients and administrators of enslaved labor.
Despite the limited statistical power, no specific common genetic markers of biomedical significance were identified that distinctly differentiate Galicia from the broader Iberian Peninsula. A gene‐based analysis using SKAT, which accounts for rare variation, identified a small number of genes with statistically significant differences compared to the IBS cohort. Additionally, minor differences in disease risk between GAL and IBS were observed for various common diseases; however, the magnitude of these differences remains relatively low.
Of more interesting consideration is our observation of the disease risk microgeographically stratified within Galicia. We demonstrate that disease risk follows distinct spatial patterns, with certain regions exhibiting elevated risk for specific diseases while showing lower risk for others. Furthermore, for the first time, we illustrate how disease risk stratifies by genetic cluster, an approach that does not necessarily align with the patterns observed in PRS‐interpolated geographical maps. Notably, we show that it is possible to rank disease risk within genetic clusters, offering a novel perspective on the distribution of common disease susceptibility across the region.
While our observation of regional differences in PRS values for various conditions in Galicia is academically relevant, its implications for public health policy remain uncertain. On the one hand, PRS reflects genetic risk alone and constitutes just one of many factors contributing to overall disease susceptibility. For instance, in breast and ovarian cancer, risk stratification algorithms incorporate additional variables such as hormone replacement therapy, parity, endometriosis, contraceptive use, and age of menarche. At the regional level, where Galicia has full autonomy in managing public health policies, the primary focus would likely be on optimizing individual risk stratification through appropriate PRSs, rather than on accounting for microgeographic variation within the region. Incorporating such fine‐scale differences would be particularly challenging for public health authorities due to factors such as population mobility and the complexity of tracing ancestral origins. For example, while our study includes individuals with all four grandparents born within an 80 km radius, real‐world clinical cases are likely to exhibit more diverse regional ancestries. Moreover, the prevailing trend in the scientific community favors the development of PRSs with broad applicability across large populations, as seen with the widely used CanRisk model for breast and ovarian cancer. This trend is likely to hold even in populations with significantly more complex ancestry patterns than Galicia. For example, efforts are underway to develop software capable of improving PRS portability by integrating fine‐scale ancestry components into PRS predictive power, as exemplified by ANCHOR, which has been applied to UK Biobank data [ 70 ].
Moreover, although subtle differences in SNP frequencies and gene markers are observed between Galicians and the broader Iberian population, the magnitude of these differences appears limited, casting doubt on the cost‐effectiveness of implementing large‐scale, region‐specific WGS initiatives that are not explicitly focused on disease‐related applications. If allele frequencies, LD patterns, and genetic risk predictions remain largely consistent between Galicia and the broader Iberian region, and if existing national and international PRS models already demonstrate strong performance in predicting disease risk among Galicians, the rationale for extensive regional WGS projects becomes questionable. Instead, considering the still substantial costs associated with WGS, and even more critically, the challenges related to the storage and management of such voluminous data, a more targeted genomic approach, such as focusing sequencing efforts on coding regions through custom gene panels or WES, may offer a more efficient and cost‐effective alternative. This strategy would be particularly valuable for prioritizing conditions of high public health relevance, such as hereditary breast and colorectal cancers. Supporting this view, recent research from Regeneron Pharmaceuticals [ 71 ] based on UK Biobank data indicates that prioritizing imputed WES over full WGS can substantially improve the discovery rate in genetic association studies, delivering critical insights while optimizing resource allocation.
Conclusions
This study represents the first comprehensive analysis of WGS variation in the Galician population. Despite Galicia's historical reputation as a genetically isolated region within the Iberian Peninsula, our results reveal no strong signatures of long‐term isolation. On the contrary, we observed only subtle fine‐scale population structure, challenging previous claims of extreme substructure in the area.
Importantly, we identified a higher North African and Middle Eastern genetic component in Galician genomes than previously reported. This ancestry is widespread across Galicia, suggesting widespread historical admixture and sustained gene flow. Such connectivity may have mitigated inbreeding otherwise expected in a small, demographically scattered population, particularly compared with other Iberian groups with larger effective sizes. We detected a male‐biased African contribution and a subtle South‐to‐North frequency gradient, possibly reflecting southern entry routes, maritime or overland, though further investigation is needed. The inferred timing predates Islamic rule, challenging the traditional assumption that Galicia's African ancestry stems primarily from Arab‐Berber expansion or the Reconquista .
Finally, PRS analyses show modest variation across Galicia, with microgeographic stratification both regionally and among fine‐scale genetic clusters. However, WGS offers limited advantage over exome sequencing for identifying common, clinically relevant variants, highlighting the need to match sequencing strategies to specific research goals. In many cases, targeted sequencing (exome or gene panels) may be a more cost‐effective strategy for maximizing clinically actionable insights.
Introduction
There is a growing interest in characterizing genome‐wide variability and genetic structure at small or regional geographic scales in different human populations. Several national European examples include extensive GWAS studies performed in Britain [ 1 ], France [ 2 ], Finland [ 3 ], Italy [ 4 ], Ireland [ 5 ], Scotland [ 6 ], and Spain [ 7 ]. This renewed interest in fine‐scale genetic structure is essential for controlling genetic ancestry in common variant association studies [ 8 ], as well as understanding rare variant association studies, since rare variants are often specific to certain geographically clustered populations [ 9 ]. In addition, fine‐scale genetic structure can clarify the relationships between closely related populations and uncover recent historical events. However, there is a scarcity of studies performed with sequencing data, since most of these population studies have been carried out using genome‐wide genotypic data; an early notable exception being the Genome of the Netherlands Project [ 10 ]. Large‐scale sequencing is far more precise and informative in revealing genetic structure, and it also allows for novel variants to be discovered. In the Iberian Peninsula, the shortage of large‐scale genomic studies is even more noticeable, with, to our knowledge, only three studies published, one using Whole Exome Sequencing (WES) data [ 11 ], and the other two using Whole‐Genome Sequencing (WGS), one of them focusing on the northeastern Iberian region of Catalonia [ 12 ], and the other primarily addressing signals of natural selection [ 13 ].
Galicia is a region occupying the northwestern corner of the Iberian Peninsula that possesses several unique characteristics. Its geographical isolation from the rest of the peninsula has, to a certain degree, maintained its identity and distinctive culture, including its own language. Besides, at the genomic level, Galicia seems to show a special differentiation, with examples of possible founder effects related to several recessive Mendelian diseases such as ichthyosis [ 14 ] or Wilson disease [ 15 ], as well as exhibiting signs of local genetic inbreeding [ 14 ] or isolation and low genetic diversity [ 16 ]. Its peripheral location in Europe and its relative isolation make Galicia particularly intriguing for studying population genetic structure. Recent data on the Spanish population by Bycroft et al. [ 7 ] have reported extraordinary fine‐scale population substructure within Galicia, albeit using genotyping data (~693 K SNPs) and exhibiting a geographically biased sampling towards the southwestern region of the Pontevedra province, where the great majority of the Galician samples were recruited.
Previous studies report a significant North African genetic contribution to Iberian populations, especially in Galicia, likely stemming from Arab‐Berber migrations during the 711 ce Muslim conquest. Although Islamic rule lasted until 1492, Galicia (part of the Kingdom of Asturias) became a Christian stronghold and a base for early resistance. By the mid‐8th century, it had largely returned to Christian control, playing a key role in the Reconquista . Despite limited direct Muslim influence in northern Iberia, genetic evidence suggests substantial North African ancestry in Galicia. Today, WGS offers new opportunities to explore this ancestry with unprecedented resolution.
In this study, we generated WGS data from 94 individuals of Galician origin, a sample size comparable to that of other WGS studies conducted in Iberia and Europe. The availability of high‐resolution genomic data enabled us to investigate inbreeding patterns and explore the extent of African and Middle Eastern admixture in Galicia. We analyzed genetic variation in the region at the highest level of resolution to build upon previous remarkable findings suggesting the presence of ultrafine population structure in Galicia [ 7 ]. To further evaluate these findings, we additionally genotyped samples from across Iberia ( n = 453), closely replicating the sampling strategy of Bycroft et al. [ 7 ] to critically assess their reported results in the Galician population and, more broadly, in Spain. Finally, we conducted the first investigation into the biomedical implications of fine‐scale population structure within Galicia, analyzing microspatial patterns of polygenic risk for various common diseases.
Coi Statement
The authors declare no conflicts of interest.
Materials And Methods
We collected 94 DNA samples from donors representing the entire territory of Galicia (northwest Spain), forming the GALOMICS (GAL) dataset. Sampling emphasized geographic homogeneity, with the average distance between the birthplaces of donors and their four grandparents being 6.6 km (standard deviation (SD) 15.8). The sample predominantly represents rural origins (~60% of donors), consistent with the population distribution of the region (Instituto Galego de Estatística, https://www.ige.gal/ ); Supporting Information Data S1 Figure S1 . Provincial representation was balanced: A Coruña (39.6%), Lugo (23.1%), Ourense (20.9%), and Pontevedra (16.5%).
Written informed consent was obtained for all participants. The study was approved by the Galician Ethics Committee (code 2017/399) and complies with Spanish biomedical and data protection regulations.
Saliva was collected using Oragene OG‐500 kits (DNA Genotek). DNA was isolated with prepIT L2P. After quality control (QC), 91 samples were retained for analysis.
To contextualize the Galician genomes within Iberian diversity, we also analyzed genotyping data from 453 donors obtained from the Spanish National DNA Bank (NDNAB), covering all major Spanish regions. These samples were genotyped with the Axiom Genome‐Wide Human Origins 1 Array at CEGEN (Santiago de Compostela), and quality control followed Affymetrix standard guidelines. Samples with dish QC below the default threshold of 0.82 were excluded. Genotyping was performed using the “AxiomGT1_all” algorithm. 599 424 SNPs were successfully retrieved from NDNAB.
Reference populations for comparative analyses were drawn from the 1000 Genomes Project ( 1000G ), the Human Genome Diversity Project ( HGDP ) [ 17 , 18 ], and published North African and Middle Eastern WGS datasets [ 19 ] (Table 1 , Figure 1A ). The combined dataset ( GAL + 1000G + HGDP ) included ~11 million overlapping SNPs . Bioinformatic processing of the 1000G data used previous bioinformatic developments [ 20 , 21 ].
Population WGS datasets used in the present study.
Overview of datasets, relatedness genome filtering, and population structure of the GAL dataset, shown through sampling maps, kinship analysis, and MDS comparisons. (A) Map illustrating the datasets used in this study, with a zoomed‐in section providing details on the Galician (GAL) dataset sampling locations (indicating also the four provinces of the region: A Coruña, Lugo, Ourense, and Pontevedra). (B) Kinship analysis of the GAL dataset, showing that after filtering out closely related individuals (removed from the plot), the remaining samples used in this study exhibit no close familial relationships. (C) MDS analysis comparing the GAL dataset to major continental populations. The left MDS plot includes a zoomed‐in view of the European genomes, highlighting the centroids of the population datasets. See Table 1 for population code details.
WGS of GAL samples was performed by BGI using the BGISEQ‐500 platform. Genomic DNA was fragmented (~350 bp), adapter‐ligated, and amplified by rolling circle amplification to produce DNA nanoballs, which were sequenced in paired‐end mode. The resulting FASTQ data were processed with the BGISEQ‐500 base‐calling software. Quality control was performed using FastQC v0.11.9 and MultiQC v1.10 [ 22 ]. Reads were mapped to the GRCh37 reference genome using BWA v0.7.17 [ 23 ], duplicates were marked with Picard v2.27.0, and variants were called using DeepVariant v1.5.0 [ 24 ] and GLnexus [ 25 ]. Functional annotation employed CADD v1.6 [ 26 ], providing deleteriousness scores for coding and noncoding variants.
Kinship relationships among the GAL samples were studied using pairwise identity‐by‐descent (IBD) patterns and using PLINK v1.9 [ 27 ].
Uniparental markers were extracted from the WGS data: mitochondrial haplogroups were classified with HaploGrep [ 28 ], and Y‐chromosome haplogroups were inferred with Yleaf [ 29 ].
Autosomal population structure was explored through Multidimensional Scaling (MDS) on average pairwise identity‐by‐state distances using R's cmdscale function.
Ancestry components were estimated using ADMIXTURE [ 30 ], applying standard filters (Minor Allele Frequency (MAF) > 0.05, linkage disequilibrium (LD) pruning r
2 0.5). Both unsupervised (general view) and supervised runs were performed, the latter incorporating reference groups with > 80% and > 90% African or Middle Eastern ancestry to minimize confounding by shared Mediterranean background.
Evidence of admixture was formally tested using the f
3 ‐statistic [ 31 ] of the form f
3 (X EUR , X NAF/SSH ; GAL), where X EUR represents a European subset and X NAF/SSH a specific North African or Sub‐Saharan African dataset. Significant negative values indicate admixture between European and African‐related sources.
Admixture timing was first estimated using ALDER [ 32 ], and reference populations representing Sub‐Saharan Africa (YRI), Europe (CEU), North Africa (MOZ), and the Middle Eastern (BED). Input files for ALDER were prepared with EIGENSOFT [ 31 ]. Next, we used fastGLOBETROTTER [ 33 ] to obtain independent estimates for the time of admixture; the files were prepared using CHROMOPAINTER [ 34 ], and the same reference populations were used for ALDER . We report results for the most parsimonious single‐date admixture scenario. To avoid bias, individuals with < 80% North African or Middle Eastern ancestry were excluded. 95% confidence intervals for inferred dates were obtained by 200 bootstrap resamples. We also explored the NDNAB dataset to investigate temporal and regional patterns of North African and Middle Eastern ancestry in Iberia, assessing whether their introduction occurred at different times across the Peninsula. In particular, we examined whether the presence of North African ancestry in northern Iberia may reflect demographic events other than the Islamic rule, such as Roman‐era, pre‐Islamic, or post‐Reconquista migrations. Generation time was assumed to be 26.1 years [ 35 ].
Fine‐scale population structure was assessed using CHROMOPAINTER and fineSTRUCTURE [ 34 ]. Genotyping data were first phased and missing genotypes imputed using Shapeit v2 [ 36 ] with 1000G phased data as reference. The coancestry matrix generated by CHROMOPAINTER was used to infer clusters and phylogenetic relationships within Galicia and across Spain in two steps. First, a subset of 10% of the sample size and four chromosomes (1, 7, 14, 22) was used to estimate the global mutation rate and switch rate with 10 Expectation–Maximization iterations. These parameters were then applied to infer the coancestry matrix. Principal Component Analyses (PCA) based on this matrix were produced using scripts provided by the authors ( http://www.paintmychromosomes.com/finestructure/ ).
Runs of Homozygosity (ROH) and inbreeding coefficients were calculated using the WGS data following [ 19 ] and using PLINK v1.9. ROH were identified using sliding windows across the genomes (segment definition: ≥ 50 SNPs over ≥ 100 kbp, allowing gaps of up to 2000 kbp between consecutive SNPs, with ≤ 1 heterozygous site, and allowing ≤ 5 missing genotypes). To compare GAL with other populations, Wilcoxon tests were applied to several ROH metrics: the total number ( nROH
S
, nROH
L
), cumulative length ( sROH
S
, sROH
L
), and mean length ( mROH
S
, mROH
L
) of short (≤ 1.5 Mb) and long (> 1.5 Mb) ROH segments. We also calculated FROH, an estimator of Wright's FIT inbreeding coefficient, defined as the total length of ROH > 1.5 Mb divided by the genome length (2795 Mb in this study) [ 37 ]. We computed four single‐point inbreeding estimators based on allele frequencies ( F
HET
, F
HAT1
, F
HAT2
, and F
HAT3
) [ 37 ].
To explore the biomedical relevance of regional genomic variation, allele frequency differences between GAL and Iberian (IBS) or pan‐European (CEU + GBR + IBS + TSI) datasets were tested for common variants (MAF > 0.05 in cases and controls, Hardy–Weinberg equilibrium p ‐value > 0.001) following [ 38 ]. Gene‐based rare variant analyses were conducted using the SKAT‐O framework [ 39 ] with CADD annotations, applying Bonferroni correction for multiple testing. The top SNP and gene associations were functionally annotated using SNPedia ( https://www.snpedia.com/ ), focusing on medically relevant traits.
We assessed regional differences in Polygenic Risk Scores (PRS) between the GAL and IBS populations to explore potential variation in genetic predisposition to common diseases. Selected PRSs included Coronary Artery Disease (PRSCAD; PGS000013), Atrial Fibrillation (PRSAF; PGS000016), Type 2 Diabetes (PRST2D; PGS000014), Inflammatory Bowel Disease (PRSIBD; PGS000017), and Breast Cancer (PRSBC1; PGS000015) [ 40 ]; Breast Cancer (PRSBC2; PGS000004) and Ovarian Cancer (PRSOCAC; PGS003394) [ 41 ] (CanRisk calculator; https://www.canrisk.org ); Acute Lower Respiratory Infection (PRSALRI; PGS000925); Schizophrenia (PRSSCH; PGS000133) [ 34 ]; Autism Spectrum Disorder (PRSASD; PGS000327) [ 35 ]; and Alzheimer's Disease (PRSALZ; PGS004146); see [ 42 , 43 , 44 , 45 , 46 , 47 ]. PRSs were calculated with PLINK v1.9 [ 27 ], and scores were standardized to one standard deviation within the combined GAL+IBS dataset.
Apart from the software and indications mentioned above, data curation, analyses, and graphical representations were performed with the R statistical package [ 48 ].
We developed GALOMICS ( https://galomics.genpob.eu ), a searchable web tool providing access to over 18 million biallelic variants from whole‐genome data. Users can query SNPs by rsID or chromosomal position and visualize allele frequencies, genotypes, and biomedical annotations. External data from Ensembl, ClinGen, UniProtKB, OMIM, and dbSNP are integrated in JSON format, with variant‐level details displayed in a ClinVar summary table.
Geographic interpolation of cluster frequencies was performed with SAGA GIS v9.6.1 [ 49 ] using ordinary Kriging.
Further methodological details, including additional QC parameters, software settings, and extended PRS descriptions, are available in Supporting Information Data S5 .
Supplementary Material
Table S1: fsb271253‐sup‐0001‐TableS1.xlsx.
Table S2: fsb271253‐sup‐0002‐TableS2.xls.
Table S3: fsb271253‐sup‐0003‐TableS3.xlsx.
Data S1: fsb271253‐sup‐0004‐DataS1.docx.
Data S2: fsb271253‐sup‐0005‐DataS2.docx.
Data S3: fsb271253‐sup‐0006‐DataS3.docx.
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.