Optimizing Polygenic Scores for Complex Morphological Traits: A Case Study in Nasal Shape Prediction

preprint OA: closed
📄 Open PDF Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-16

This study developed a novel framework using multivariate GWAS, genetically informative phenotypes, and model benchmarking to improve polygenic prediction accuracy for complex morphological traits like nasal shape.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-16 · read from full text

This paper develops a framework to improve polygenic score (PGS) prediction for complex morphological traits by using multivariate GWAS summary statistics to guide SNP selection, while preserving univariate effect sizes for existing PGS tools, and by defining “genetically informative” phenotypes. The authors apply and benchmark their approach for predicting 3D nasal morphology in 52,896 European-ancestry participants from UK Biobank, comparing multivariate-GWAS-guided clumping and thresholding (C+T) against a traditional univariate GWAS-based approach. They report significantly improved prediction for eigen-shapes when SNP selection is based on multivariate GWAS (mean variance explained 3.88% vs 2.02% in the test set), and show that heritability-optimized phenotypes are more predictable than PCA-derived eigen-shapes; LDpred2 performed best among tested PGS methods. A key limitation noted is that well-powered GWAS for facial morphology are not currently feasible with univariate phenotyping, motivating their multivariate strategy. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Polygenic scores (PGS) facilitate the prediction of an individual’s phenotype from their genotype. Typically, PGS methods apply regularization or use a clumping and thresholding (C+T) approach to handle SNP inclusion. To achieve good prediction accuracy, these approaches rely on effect size estimates from well-powered genome-wide association studies (GWAS). However, this is currently not feasible for morphological shape when phenotyped as univariate traits. Here, we introduce a novel framework to enhance polygenic prediction through three key components: (1)leveraging multivariate GWAS summary statistics for improved SNP selection, (2) defining genetically informative phenotypes, and (3) benchmarking PGS methods to select the optimal model. Our approach integrates multivariate GWAS, which performs an omnibus test against all phenotypic variables jointly with increased power. Specifically, our approach leverages P values from multivariate GWAS to improve SNP selection while maintaining the effect size estimates for the univariate trait under investigation, allowing the use of current PGS tools. We evaluated our proposed method for predicting 3D nasal morphology using a dataset of 52,896 individuals of European ancestry from the UK Biobank. Using the C+T method, SNP selection based on multivariate GWAS resulted in significantly improved phenotypic prediction ( P = 9.74e-5) for eigen-shapes, with a mean variance explained of 3.88% (SD = 1.59%) compared to 2.02% (SD = 1.10%) using a traditional univariate approach in the test set (n = 2,896). We also tested whether heritability-optimized phenotypes were more predictable than eigen-shapes derived from principal component analysis (PCA). On average, with SNP selection based on multivariate GWAS using the C+T method, heritability-optimized phenotypes yielded greater predictive performance, with PGS scores explaining 2.72%-10.37% of phenotypic variance, compared to 1.05%-6.84% for eigen-shapes. Furthermore, benchmarking several PGS methods revealed that LDpred2 consistently achieved the best performance for predicting nasal morphology. Our results demonstrate that combining multivariate GWAS P values with optimized phenotypes and advanced PGS models leads to more accurate polygenic prediction for complex morphological traits.
Full text 51,001 characters · extracted from preprint-html · click to expand
Optimizing Polygenic Scores for Complex Morphological Traits: A Case Study in Nasal Shape Prediction | bioRxiv /* */ /* */ <!-- <!-- /*! * yepnope1.5.4 * (c) WTFPL, GPLv2 */ (function(a,b,c){function d(a){return"[object Function]"==o.call(a)}function e(a){return"string"==typeof a}function f(){}function g(a){return!a||"loaded"==a||"complete"==a||"uninitialized"==a}function h(){var a=p.shift();q=1,a?a.t?m(function(){("c"==a.t?B.injectCss:B.injectJs)(a.s,0,a.a,a.x,a.e,1)},0):(a(),h()):q=0}function i(a,c,d,e,f,i,j){function k(b){if(!o&&g(l.readyState)&&(u.r=o=1,!q&&h(),l.onload=l.onreadystatechange=null,b)){"img"!=a&&m(function(){t.removeChild(l)},50);for(var d in y[c])y[c].hasOwnProperty(d)&&y[c][d].onload()}}var j=j||B.errorTimeout,l=b.createElement(a),o=0,r=0,u={t:d,s:c,e:f,a:i,x:j};1===y[c]&&(r=1,y[c]=[]),"object"==a?l.data=c:(l.src=c,l.type=a),l.width=l.height="0",l.onerror=l.onload=l.onreadystatechange=function(){k.call(this,r)},p.splice(e,0,u),"img"!=a&&(r||2===y[c]?(t.insertBefore(l,s?null:n),m(k,j)):y[c].push(l))}function j(a,b,c,d,f){return q=0,b=b||"j",e(a)?i("c"==b?v:u,a,b,this.i++,c,d,f):(p.splice(this.i++,0,a),1==p.length&&h()),this}function k(){var a=B;return a.loader={load:j,i:0},a}var l=b.documentElement,m=a.setTimeout,n=b.getElementsByTagName("script")[0],o={}.toString,p=[],q=0,r="MozAppearance"in l.style,s=r&&!!b.createRange().compareNode,t=s?l:n.parentNode,l=a.opera&&"[object Opera]"==o.call(a.opera),l=!!b.attachEvent&&!l,u=r?"object":l?"script":"img",v=l?"script":u,w=Array.isArray||function(a){return"[object Array]"==o.call(a)},x=[],y={},z={timeout:function(a,b){return b.length&&(a.timeout=b[0]),a}},A,B;B=function(a){function b(a){var a=a.split("!"),b=x.length,c=a.pop(),d=a.length,c={url:c,origUrl:c,prefixes:a},e,f,g;for(f=0;f<d;f++)g=a[f].split("="),(e=z[g.shift()])&&(c=e(c,g));for(f=0;f<b;f++)c=x[f](c);return c}function g(a,e,f,g,h){var i=b(a),j=i.autoCallback;i.url.split(".").pop().split("?").shift(),i.bypass||(e&&(e=d(e)?e:e[a]||e[g]||e[a.split("/").pop().split("?")[0]]),i.instead?i.instead(a,e,f,g,h):(y[i.url]?i.noexec=!0:y[i.url]=1,f.load(i.url,i.forceCSS||!i.forceJS&&"css"==i.url.split(".").pop().split("?").shift()?"c":c,i.noexec,i.attrs,i.timeout),(d(e)||d(j))&&f.load(function(){k(),e&&e(i.origUrl,h,g),j&&j(i.origUrl,h,g),y[i.url]=2})))}function h(a,b){function c(a,c){if(a){if(e(a))c||(j=function(){var a=[].slice.call(arguments);k.apply(this,a),l()}),g(a,j,b,0,h);else if(Object(a)===a)for(n in m=function(){var b=0,c;for(c in a)a.hasOwnProperty(c)&&b++;return b}(),a)a.hasOwnProperty(n)&&(!c&&!--m&&(d(j)?j=function(){var a=[].slice.call(arguments);k.apply(this,a),l()}:j[n]=function(a){return function(){var b=[].slice.call(arguments);a&&a.apply(this,b),l()}}(k[n])),g(a[n],j,b,n,h))}else!c&&l()}var h=!!a.test,i=a.load||a.both,j=a.callback||f,k=j,l=a.complete||f,m,n;c(h?a.yep:a.nope,!!i),i&&c(i)}var i,j,l=this.yepnope.loader;if(e(a))g(a,0,l,0);else if(w(a))for(i=0;i (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0];var j=d.createElement(s);var dl=l!='dataLayer'?'&l='+l:'';j.src='//www.googletagmanager.com/gtm.js?id='+i+dl;j.type='text/javascript';j.async=true;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-M677548'); Skip to main content Home About Submit ALERTS / RSS Search for this keyword Advanced Search New Results Optimizing Polygenic Scores for Complex Morphological Traits: A Case Study in Nasal Shape Prediction View ORCID Profile Meng Yuan , View ORCID Profile Seppe Goovaerts , View ORCID Profile Nina Claessens , Jay Devine , View ORCID Profile Sara Becelaere , View ORCID Profile Isabelle Cleynen , View ORCID Profile Peter Claes doi: https://doi.org/10.1101/2025.09.14.676081 Meng Yuan 1 Department of Electrical Engineering, ESAT/PSI, KU Leuven , Leuven, Belgium 2 Department of Human Genetics, KU Leuven , Leuven, Belgium 3 Medical Imaging Research Center, University Hospitals Leuven , Leuven, Belgium 4 British Heart Foundation Cardiovascular Epidemiology Unit, Department of Public Health and Primary Care, University of Cambridge , Cambridge, UK 5 Victor Phillip Dahdaleh Heart and Lung Research Institute, University of Cambridge , Cambridge, UK Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Meng Yuan For correspondence: meng.yuan{at}kuleuven.be peter.claes{at}kuleuven.be Seppe Goovaerts 2 Department of Human Genetics, KU Leuven , Leuven, Belgium 3 Medical Imaging Research Center, University Hospitals Leuven , Leuven, Belgium Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Seppe Goovaerts Nina Claessens 1 Department of Electrical Engineering, ESAT/PSI, KU Leuven , Leuven, Belgium 3 Medical Imaging Research Center, University Hospitals Leuven , Leuven, Belgium Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Nina Claessens Jay Devine 1 Department of Electrical Engineering, ESAT/PSI, KU Leuven , Leuven, Belgium 2 Department of Human Genetics, KU Leuven , Leuven, Belgium 3 Medical Imaging Research Center, University Hospitals Leuven , Leuven, Belgium Find this author on Google Scholar Find this author on PubMed Search for this author on this site Sara Becelaere 2 Department of Human Genetics, KU Leuven , Leuven, Belgium Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Sara Becelaere Isabelle Cleynen 2 Department of Human Genetics, KU Leuven , Leuven, Belgium Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Isabelle Cleynen Peter Claes 1 Department of Electrical Engineering, ESAT/PSI, KU Leuven , Leuven, Belgium 2 Department of Human Genetics, KU Leuven , Leuven, Belgium 3 Medical Imaging Research Center, University Hospitals Leuven , Leuven, Belgium 6 Murdoch Children’s Research Institute , Melbourne, Victoria, Australia Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Peter Claes For correspondence: meng.yuan{at}kuleuven.be peter.claes{at}kuleuven.be Abstract Full Text Info/History Metrics Data/Code Preview PDF Abstract Polygenic scores (PGS) facilitate the prediction of an individual’s phenotype from their genotype. Typically, PGS methods apply regularization or use a clumping and thresholding (C+T) approach to handle SNP inclusion. To achieve good prediction accuracy, these approaches rely on effect size estimates from well-powered genome-wide association studies (GWAS). However, this is currently not feasible for morphological shape when phenotyped as univariate traits. Here, we introduce a novel framework to enhance polygenic prediction through three key components: (1)leveraging multivariate GWAS summary statistics for improved SNP selection, (2) defining genetically informative phenotypes, and (3) benchmarking PGS methods to select the optimal model. Our approach integrates multivariate GWAS, which performs an omnibus test against all phenotypic variables jointly with increased power. Specifically, our approach leverages P values from multivariate GWAS to improve SNP selection while maintaining the effect size estimates for the univariate trait under investigation, allowing the use of current PGS tools. We evaluated our proposed method for predicting 3D nasal morphology using a dataset of 52,896 individuals of European ancestry from the UK Biobank. Using the C+T method, SNP selection based on multivariate GWAS resulted in significantly improved phenotypic prediction ( P = 9.74e-5) for eigen-shapes, with a mean variance explained of 3.88% (SD = 1.59%) compared to 2.02% (SD = 1.10%) using a traditional univariate approach in the test set (n = 2,896). We also tested whether heritability-optimized phenotypes were more predictable than eigen-shapes derived from principal component analysis (PCA). On average, with SNP selection based on multivariate GWAS using the C+T method, heritability-optimized phenotypes yielded greater predictive performance, with PGS scores explaining 2.72%-10.37% of phenotypic variance, compared to 1.05%-6.84% for eigen-shapes. Furthermore, benchmarking several PGS methods revealed that LDpred2 consistently achieved the best performance for predicting nasal morphology. Our results demonstrate that combining multivariate GWAS P values with optimized phenotypes and advanced PGS models leads to more accurate polygenic prediction for complex morphological traits. Introduction Predicting an individual’s genetic predisposition for a complex disease or trait is a critical task known as polygenic scoring that has far-reaching scientific, ethical, and societal implications. For many genetic disorders, polygenic scores (PGS) show promise in early risk detection, stratification of patient groups, and predicting therapeutic responses [ 1 ], [ 2 ], [ 3 ], [ 4 ], [ 5 ]. PGS are typically calculated as the weighted sum of allele dosages of single nucleotide polymorphisms (SNPs), where trait-associated SNPs are usually identified in genome-wide association studies (GWAS). The weights reflect the magnitude of association between each SNP and the trait. Unfortunately, PGS cannot serve as stand-alone prediction tools, as they capture only a fraction of the genetic contribution, with other genetic and non-genetic factors also playing a significant role [ 6 ]. Complex morphological traits are particularly difficult to predict because their phenotypic variation is multipartite, and no common consensus exists to define phenotypes. Moreover, morphological traits are highly polygenic, shaped by many genetic effects that drive spatially and temporally dependent developmental processes acted on environmental context [ 7 ], [ 8 ], [ 9 ]. To date, the most comprehensive review on facial shape in European individuals estimated that 501 independent SNPs across 303 loci explain only 13.7% of the sample variance [ 10 ]. Another study focused on variants with <1% frequency found seven genes to be enriched for rare variants affecting facial shape [ 11 ]. In general, our current understanding of common and rare genetic variants is unlikely to yield an accurate reconstruction of the entire multivariate face due to this complexity. Optimizing a latent trait that is highly heritable based on a multivariate description and then predicting these distinct univariate facial traits, on the other hand, is a more feasible task that has not been fully explored. Various methods have been proposed to calculate PGS, differing primarily in two key aspects: the selection of SNPs and the weights assigned to them [ 12 ]. The standard approach, clumping and thresholding (C+T), selects a set of approximately independent SNPs by applying linkage disequilibrium (LD) clumping to SNPs that exceed a predefined P value threshold [ 13 ], [ 14 ], [ 15 ]. Although computationally and conceptually simple, the C+T method uses only a subset of SNPs, ignoring many other SNPs along with their LD information, which limits prediction accuracy [ 16 ]. To address this, more sophisticated PGS methods have been developed to incorporate genome-wide SNPs and their LD structure. These methods typically perform shrinkage on all SNP effect sizes, leveraging commonly used regularization techniques (e.g., Lassosum [ 17 ], Lassosum2 [ 18 ]) or Bayesian shrinkage approaches (e.g., LDpred [ 19 ], LDpred2 [ 20 ], and PRS-CS [ 21 ]). Recent comparative studies of PGS methods have shown that no approach is universally optimal, as prediction accuracy depends on trait-specific genetic architecture, the statistical power of GWAS, and ancestry composition [ 22 ], [ 23 ]. While sufficiently powered GWAS are essential for accurate PGS, this remains currently unachievable for facial traits when phenotyped univariately, necessitating alternative strategies. Previous multivariate GWAS of facial shape have uncovered more genomic loci than their univariate GWAS counterparts [ 10 ]. Open-ended multivariate GWAS improve genomic discovery because they leverage cross-trait genetic covariance, while univariate GWAS ignores this. In addition, they reduce the multiple-testing burden compared to analyzing all genetically correlated traits separately [ 24 ]. Here, we investigate the optimization of three distinct components to enhance PGS prediction of complex morphological traits: SNP selection, phenotype definition, and the PGS model. Three-dimensional (3D) nose morphology, an easy to recognize and heritable [ 25 ] part of the human face, is used as a case study. First, we propose the use of multivariate GWAS P values to improve SNP selection, leveraging their ability to capture a variant’s association with multiple traits simultaneously. We retain univariate effect size estimates for PGS construction, ensuring compatibility with existing tools designed for single-trait prediction. Second, rather than predicting the full nasal shape, we focus on distinct nasal features, evaluating how different phenotype definitions influence predictability. To this end, we compare three types of nasal phenotypes: (1) widely used traditional anthropometric measures (i.e., inter-landmark distances);eigen-shapes derived from principal component analysis (PCA); (3) heritability-enriched traits generated through optimization algorithms [ 26 ]. Third, we benchmark several PGS methods, including PRSice-2 [ 15 ], LDpred2 [ 20 ], Lassosum2 [ 18 ] and PRS-CS [ 21 ], and identify the best-performing model for nasal trait prediction. Results In this work, we considered three key components in PGS construction, as illustrated in Fig. 1A : (1) phenotype definition, (2) selection of SNPs, and (3) weight assignment strategies in PGS methods. Nose morphology data (n = 52,896) was sourced from the UK Biobank [ 27 ] and extracted from full head MRI images (see Methods section). From this dataset, 45,000 samples were used as training data to perform discovery GWAS and primary PGS model building. A separate set of 5,000 samples was used as validation data for hyperparameter fine-tuning ( Fig. 1B ). Finally, 2,896 samples were used as independent test data to evaluate prediction performance ( Fig. 1C ). Prediction performance was assessed by fitting a linear model where PGS scores were used to predict the original phenotypes, with R 2 (coefficient of determination) quantifying the proportion of phenotypic variance explained. Higher R 2 indicates better prediction. Download figure Open in new tab Figure 1. PGS workflow. Phenotype definition: The phenotype of interest was defined in three ways: (1) anthropometric measures (10 linear distances among 5 sparse landmarks as described previously in [ 25 ], see Supplementary File 2); (2) 21 eigen-shapes derived from applying PCA to dense landmarks, or (3) 21 heritability-enriched traits obtained by training a genetic algorithm (GA) model to identify directions in feature space with high heritability [ 26 ] (see Methods). SNP inclusion: SNPs associated with the phenotype of interest are typically identified through univariate GWAS (UV-GWAS; a separate GWAS for each trait), where both SNP selection and weight estimation rely on the same discovery GWAS. In this study, we used univariate effect size estimates of UV-GWAS as weights to ensure compatibility with existing PGS tools and explored alternative SNP inclusion strategies. First, we used meta-analyzed univariate GWAS (UV-Meta-GWAS) P values for SNP selection, which aggregates univariate GWAS results within each phenotype category (see Methods). Second, we evaluated P values from multivariate GWAS (MV-GWAS) for SNP selection. The MV-GWAS was conducted on eigen-shapes using canonical correlation analysis (CCA), a well-established method that has successfully identified genomic loci associated with craniofacial morphology by performing an omnibus test against all phenotypic variables jointly with increased power [ 28 ]. Methods benchmarking: Some studies [ 12 ], [ 22 ], [ 23 ] have compared various PGS methods for other types of traits; however, none have benchmarked these methods for complex morphological traits. Therefore, we utilized a PGS toolbox ( https://github.com/SaraBecelaere/STREAM-PRS ) which includes the C+T approach implemented in PRSice-2 [ 15 ], and several effect size shrinkage methods such as LDpred2 [ 20 ], Lassosum [ 17 ], Lassosum2 [ 18 ] and PRS-CS [ 21 ]. Detailed descriptions of training and hyperparameter settings for each method are provided in Methods section. Fig. 2 compares prediction performance under different SNP inclusion strategies using the C+T approach (numerical details in Supplementary File 1, Table S1). A consistent pattern emerged across all phenotype groups: MV-GWAS yielded the best results, followed by UV-Meta-GWAS, while standard UV-GWAS performed the least effectively. A substantial improvement was observed for heritability-enriched traits and eigen-shapes using MV-GWAS for SNP selection (mean variance explained of 5.86% [SD=1.86%] vs. 3.51% [SD=1.55%] and 3.88% [SD=1.59%] vs. 2.02% [SD=1.10%], respectively). Meanwhile, inter-landmark distances exhibited a relatively moderate gain (4.71% [SD=0.98%] vs. 3.15% [SD=1.22%]). When comparing phenotype definitions using the simple C+T approach and MV-GWAS, heritability-enriched traits were more predictable than eigen-shapes (median R 2 : 5.57%; range: 2.72-10.37% vs. median 4.39%; range: 1.05-6.84%; P = 1.3e-3). Inter-landmark distances demonstrated exhibited a moderate level of predictability (median: 4.66%; range: 3.40-6.02%). Download figure Open in new tab Figure 2. Comparison of prediction performance using the C+T approach. After confirming that multivariate GWAS P values result in more optimal SNP selection, we compared PGS approaches under this scheme ( Fig. 3 ). LDpred2 consistently showed relatively higher performance across all phenotypes. For inter-landmark distances, LDpred2 achieved a higher mean explained variance (7.57%) than PRSice-2 (4.71%; P = 8.5e-3) and PRS-CS (5.02%; P Download figure Open in new tab Figure 3. Comparison of prediction performance using different phenotyping methods and PGS approaches with multivariate GWAS P values for SNP inclusion. = 1.85e-2) and was not significantly different from Lassosum2 (5.63%; P = 1.59e-1); all P values are Bonferroni-corrected for three tests. When comparing phenotype definitions using LDpred2, heritability-enriched traits and inter-landmark distances showed similar performance, both surpassing PCA-based eigen-shapes predictions (details in Supplementary File 1, Table S2). Fig. 4 illustrates the prediction performance for 10 representative traits, including linear distances, the first 10 eigen-shapes, and the first 10 heritability-enriched traits (visualizations of all nasal traits are in Supplementary File 2). The highest phenotypic variance explained was observed for the linear distance between the nasion and cheek points (R 2 = 11.5%, Fig. 4A , column 7). Similarly, a heritability-enriched trait ( Fig. 4C , column 5), characterized by a larger, rounded, and pointing nose tip, also exhibited high predictability (R 2 = 10.3%). Notably, another heritability-optimized trait ( Fig. 4C , column 10), which effectively captures fine-scale morphological features (particularly alar crease curvature), achieved high predictive performance (R 2 = 10.9%). In contrast, a similar linear distance measure ( Fig. 4A , column 10) showed relatively low performance (R 2 = 5.3%), demonstrating that linear measures alone cannot adequately represent complex curvature features. Download figure Open in new tab Figure 4. Visualization of nasal traits. (A) The first panel displays the 10 linear distances alongside their corresponding prediction performance. (B) The second panel displays the first 10 principal components (PCs). Shape variation is illustrated in gray, showing the range from the mean to plus and minus 3 times the standard deviation (SD) along each PC, positioned vertically with respect to each other (top + 3SD, and bottom -3D). Colored bars represent differences (in mm) between these two opposite deviations from the mean. (C) Similarly, the third panel presents the first 10 heritability-enriched traits. Discussion In summary, we proposed a framework to optimize PGS construction for complex morphological traits through three key components: (1) leveraging multivariate GWAS summary statistics to enhance SNP selection, (2) defining genetically informative phenotypes by exploiting the inherent multivariate nature of morphological traits, and (3) benchmarking PGS methods to identify the best-performing model. We evaluated the framework using a large dataset from the UK Biobank and found that incorporating MV-GWAS P values significantly improved PGS prediction accuracy compared to standard UV-GWAS. Among phenotype definitions, heritability-enriched traits and inter-landmark distances yielded better performance than eigen-shapes. Furthermore, comparisons across PGS models revealed that LDpred2 consistently outperformed other methods in predicting nasal morphology. A key contribution of our study is the use of open-ended MV-GWAS summary statistics to refine SNP selection in PGS development. MV-GWAS identified more significant associations than UV-GWAS, capturing all genomic loci detected by UV-GWAS while also uncovering additional ones [ 10 ]. This is particularly advantageous for complex traits like the human face, where genetic variants often affect multiple facial traits simultaneously, each with small effect sizes. By accounting for trait dependencies, MV-GWAS enhances statistical power to identify pleiotropic SNPs. As expected, UV-Meta-GWAS, which is another way to boost power, performed intermediately, falling between UV-GWAS and MV-GWAS. From a computational standpoint, integrating MV-GWAS P values into existing PGS toolboxes is straightforward, requiring only separate inputs for SNP selection (i.e., P values) and weight estimation (i.e., allele effect sizes). Given its superior performance in all comparisons, we strongly recommend using MV-GWAS for selecting candidate SNPs in PGS calculation for complex traits, especially for morphological traits, such as facial structure, cranial vault shape, or brain morphology. The multivariate nature of morphological traits provides an opportunity to optimize latent traits for heritability. Accordingly, heritability-enriched traits demonstrated higher prediction performance than PCA-based eigen-shapes across all PGS methods. This aligns with previous findings [ 26 ], [ 29 ], as these phenotypes were explicitly optimized to capture genetically informative features of nasal shape, making them better suited for genotype-phenotype analyses. While PCA remains a valuable tool for extracting features from high-dimensional data, its derived features are driven by statistical variance explained which may not correspond well to heritability. Supporting this, previous studies [ 26 ], [ 29 ] have demonstrated that heritability-enriched phenotypes exhibit higher SNP-heritability and more effectively identify relevant genomic loci compared to eigen-shapes. When using the C+T approach, heritability-enriched traits yielded relatively higher predictive accuracy than inter-landmark distances. However, both approaches achieved comparable performance with optimal PGS models. Inter-landmark distances are defined based on anatomical and biological prior knowledge [ 30 ], [ 31 ], and a previous study [ 25 ] has shown their high SNP-heritability, indicating strong alignment with genetically determined shape variation. A limitation of this approach, however, is that the genetic findings are constrained by the available landmarks. This is particularly evident in nasal shape analysis ( Fig. 4 ), where subtle changes in complex morphology (e.g., curvature) cannot be adequately represented by only five landmarks and their inter-landmark distances. In contrast, heritability-enriched traits are optimized from dense landmarks and can capture highly heritable, localized regions with greater precision. Ultimately, this represents a trade-off: inter-landmark measures are simpler to define but require advanced PGS models (e.g., LDpred2) for accurate prediction, whereas heritability-enriched traits require training sophisticated optimization algorithms to extract features from dense landmarks but achieve strong prediction performance even with basic PGS methods (e.g., C+T approach). A comparison of four PGS methods (PRSice-2 [ 15 ], LDpred2 [ 20 ], Lassosum2 [ 18 ] and PRS-CS [ 21 ]) revealed that LDpred2 achieved the highest prediction accuracy overall. While some studies [ 12 ], [ 20 ] suggest that PGS methods which more formally model trait-specific genetic architecture (e.g., LDpred2, PRS-CS) outperform simpler approaches like C+T approach, we found that the C+T method remained competitive for nasal morphology. This may be due to the highly polygenic nature of nasal traits, and thresholding methods generally outperform PRS-CS for such traits [ 22 ]. Moreover, optimal parameter settings for PGS methods that account for the genetic architecture of specific traits could further enhance the prediction performance. Computationally, the C+T approach (as implemented in PRSice-2) was the most efficient, taking ∼20 minutes per univariate trait on a single CPU to first generate SNP weights based on the discovery sample, then apply them to the validation and testing samples. In contrast, PRS-CS required ∼6 hours, while LDpred2 and Lassosum2 (jointly implemented in [ 32 ]) took ∼3 hours in total. To improve prediction accuracy, we recommend applying multiple PGS methods, selecting the best-performing model, and using a grid search to tune hyperparameters. While our framework advances PGS construction for nasal traits, and by extension facial traits, the practical utility of current models remains limited. Even the best-performing models explained only ∼10% of phenotypic variance, a level unlikely to be meaningful for most real-world applications in forensic, anthropological, and clinical research. This limitation arises because PGS captures only the additive genetic component of a trait, excluding non-genetic factors and gene– environment interactions; thus, its predictive accuracy is inherently bounded by heritability [ 6 ], [ 33 ]. Additionally, GWAS do not account for transcriptomic or proteomic variation further limiting their explanatory power [ 34 ]. Translating PGS into practice also faces significant clinical and social hurdles, as naive implementations risk introducing severe bias and misinterpretation [ 35 ], [ 36 ]. Given these challenges, it may be more practical to focus on predicting specific, distinctive traits rather than attempting to predict the entire morphological shape. A limitation of our study is its focus on individuals of European ancestry. Although increased diversity introduces statistical challenges, methods like PRS-CSx [ 37 ] can effectively couple genetic effects across populations through a shared prior. Studies have demonstrated that training PGS models in ancestrally diverse cohorts improves the weight estimation of genetic variants, particularly for variants with higher frequencies in non-European populations [ 22 ], [ 23 ]. For example, a multi-ancestry GWAS meta-analysis of inter-landmark distance-based facial features explained 1.61%–7.50% of the phenotypic variance [ 38 ]. Their performance is comparable to ours, despite being based on a smaller sample size that included 9,674 East Asians and 10,115 Europeans. This suggests that incorporating diverse ancestries may enhance PGS model performance, or the relatively high phenotypic variance explained is partially driven by population-level (ancestry-related) structure in addition to individual-level differences. Future research could explore multi-ancestry PGS approaches to enhance prediction accuracy and understand the properties of our proposed strategies within and across ancestries. In conclusion, we introduced a framework to optimize PGS construction for complex morphological traits by enhancing SNP selection through the integration of MV-GWAS summary statistics, genetically informative phenotype definitions, and PGS model selection. When applied to the prediction of nasal morphology, this framework identified distinctive traits with notably higher predictability. Our findings offer valuable insights for future PGS-based predictions, particularly for multivariate traits or, more broadly, sets of genetically correlated traits. Materials and methods Dataset Our study utilized data from the UK Biobank [ 27 ], which comprises genetic and head MRI data from approximately 60,000 participants in the UK. We restricted our analysis to unrelated individuals of European ancestry. Ancestry assignment for each participant was based on the Pan-UK Biobank Project ( https://pan.ukbb.broadinstitute.org/ ). Related individuals were identified using the KING-robust [ 39 ] kinship estimator at a threshold of 0.0442 (third-degree relatives) followed by random selection of one individual per related group. Genotype imputation procedures for the UK Biobank have been described in [ 27 ]. After imputation, we applied quality control filters to retain variants with an imputation INFO score > 0.3, minor allele frequency (MAF) > 0.01, genotyping missingness rate 1 × 10-6. Additionally, individuals with more than 5% missing genotype data were excluded from the analysis. Phenotype processing, including 3D image quality control and sparse/dense landmarking, was performed as described in [Goovaerts et al, in preparation]. The final dataset included 8,922,008 SNPs and 52,896 individuals. We partitioned the data into 45,000 samples for discovery GWAS and primary PGS model training, 5,000 for validation and hyperparameter tuning, and 2,896 for evaluating prediction performance. Phenotyping methods The phenotype of interest included anthropometric traits, eigen-shapes, and heritability-enriched traits, with detailed methodologies described in previous work [ 25 ], [ 26 ]. Briefly, we focused on 5 anatomical nasal landmarks and computed 10 inter-landmark Euclidean distances between landmarks (Supplementary File 2). Besides sparse landmarks, we also used dense landmark configurations, represented as a 3D matrix of dimensions N (number of shapes), L (7,160 quasi-landmarks), and 3 (x, y, z coordinates). After mean-centering and reshaping the landmarks into a 2D matrix, we applied low-rank singular value decomposition (SVD). To retain meaningful variation, we combined PCA with parallel analysis [ 40 ], [ 41 ], yielding 21 eigen-shapes that captured 98.21% of nasal shape variation. While eigen-shapes are widely used, they are unsupervised and may not align with genetically relevant phenotypic axes. Therefore, we applied a third approach optimizing phenotype extraction for genetic analyses. As described in [ 26 ], we applied PCA to construct a lower-dimensional feature space encoding complex shape variations. Then, we employed an optimization algorithm, more explicitly a genetic algorithm (GA), to identify directions or traits in this space with high SNP-heritability (training details and hyperparameter settings are in Supplementary File 3). SNP-heritability was computed via GREML [ 42 ], [ 43 ] based on unrelated individuals using SNPlib toolbox [ 44 ] ( https://github.com/jiarui-li/SNPLIB ). Genome-wide association analysis For UV-GWAS, we performed linear regression (function ‘regstats’ from Matlab 2022b) under an additive genetic model (SNP dosages: 0, 1, 2), adjusting for covariates (sex, age, age-squared, height, weight, and the first ten genomic ancestry axes), prior to the regression. To obtain P values for UV-Meta-GWAS, we took the lowest P value for each SNP across all traits within the same phenotype group. For MV-GWAS, we treated the 21 eigen-shapes as a unified representation of multivariate shape variation. Using CCA (function ‘canoncorr’ from Matlab 2022b), we identified the linear combination of these 21 components that maximally correlated with the SNP dosage in the discovery cohort. Prior to CCA, we adjusted for the same set of covariates used in the UV-GWAS. PGS method benchmarking PGS method benchmarking was performed using STREAM-PRS toolbox as described in [Becelaere et al, in revision], which incorporates the C+T approach implemented in PRSice-2 [ 15 ], and shrinkage-based methods LDpred2 [ 20 ], Lassosum2 [ 18 ] and PRS-CS [ 21 ]. For the C+T method, we tested a series of P value thresholds: 5e-8, 1e-7, 1e-6, 1e-5, 1e-4, 0.001, 0.01, 0.02, 0.05, 0.1, 0.2, 0.3, 0.4, 0.5, and 0.6. Clumping was performed using a 250 kb window and an LD R 2 threshold of 0.1. For LDpred2 (jointly implemented with Lassosum2 in [ 32 ]), we applied the LDpred2-inf, LDpred2-grid, and LDpred2-auto models. When running LDpred2-grid, we tested a grid of hyperparameters, with the proportion of causal variants p from a sequence of 10 values from 1e-6 to 1 on a log-scale and the heritability coefficient h 2 = (0.3, 0.7, 1, 1.5). For Lassosum2, we used a grid of hyperparameters: nLambda = 20, lambdaMinRatio = 0.01, and delta = (0.001, 0.05, 1). For PRS-CS, we tested the global shrinkage parameter phi = (1, 0.1, 0.01, 0.001, 1e-4, 1e-6). For all PGS models, optimal hyperparameters were selected using a validation set. Prediction performance was evaluated by fitting a linear model in which the PGS scores were used to predict the original phenotype, with R 2 (coefficient of determination) quantifying the proportion of phenotypic variance explained. We considered three key components in PGS construction: phenotype definition, SNP inclusion, and weight assignment strategies in PGS models. (A) A total of 45,000 samples were used as training data to perform discovery GWAS and to build the PGS model; (B) 5,000 samples were used as validation data for hyperparameter fine-tuning; and (C) 2,896 samples were used as an independent test set to evaluate prediction performance. We evaluated three categories of phenotypes: (1) anthropometric measures (i.e., Distance), (2) eigen-shapes derived from principal component analysis (PCA), and (3) heritability-enriched traits obtained by training a genetic algorithm (GA). Three SNP inclusion strategies were tested: (1) standard univariate GWAS (UV-GWAS); (2) SNPs identified based on meta-analyzed univariate GWAS (UV-Meta-GWAS), which aggregates univariate GWAS results within each phenotype category; and (3) multivariate GWAS (MV-GWAS), where we treated the 21 eigen-shapes as a unified representation of multivariate shape variation and performed GWAS using canonical correlation analysis (CCA). We evaluated three categories of phenotypes: (1) anthropometric measures (i.e., Distance), (2) eigen-shapes derived from principal component analysis (PCA), and (3) heritability-enriched traits obtained by training a genetic algorithm (GA). Four PGS methods were compared: the C+T approach (implemented in PRSice-2), PRS-CS, Lassosum2 and LDpred2. Ethical approval The committee/institutional research board of UK Biobank gave ethical approval for collection of the UK Biobank data ( https://www.ukbiobank.ac.uk/learn-more-about-uk-biobank/about-us/ethics ). Approval to use UK Biobank at an individual level in this work was obtained under application no. 88320. Local ethics approval at the KU Leuven, Belgium, was provided under PRET G-2022-5272. Data availability All the data and detailed information for the UK Biobank, including genetic markers, covariates and MRI images are available through application ( http://www.ukbiobank.ac.uk/register-apply/ ). This research has been conducted using the UKB resource under application no. 88320. We are grateful for all the participants in that resource. This manuscript reflects the views of the authors and may not reflect the opinions or views of the UK Biobank funders and investigators. Code availability Software for the PGS pipeline is available at https://github.com/SaraBecelaere/STREAM-PRS . Code for training the genetic algorithm is available at https://github.com/mm-yuan/optimize_phenotyping . Author contributions M.Y. Formal analysis, Methodology, Investigation, Visualization, Data curation, Software, Writing – original draft, Writing – review & editing S.G. Investigation, Data curation, Writing – review & editing N.C. Investigation, Data curation, Writing – review & editing J.D. Investigation, Data curation, Writing – review & editing S.B. Investigation, Software, Resources, Writing – review & editing I.C. Investigation, Software, Resources, Writing – review & editing P.C. Conceptualization, Methodology, Investigation, Supervision, Funding acquisition, Writing – review & editing, Project administration Funding information This work was supported by the Bijzonder Onderzoeksfonds (BOF) C1, KU Leuven, C14/20/081 and Fonds voor Wetenschappelijk Onderzoek (FWO), Flanders G017225N. Supporting information Supplementary File 1 -Table.xlsx Supplementary Table S1: Source data for Fig. 2 Supplementary Table S2: Source data for Fig. 3 Supplementary File 2 – Additional Materials The facial template, nasal landmark labels, and visualizations of all nasal traits are available online at https://doi.org/10.6084/m9.figshare.29621687 . Supplementary File 3 – Implementation details Funder Information Declared Bijzonder Onderzoeksfonds (BOF) C1, KU Leuven , C14/20/081 Fonds voor Wetenschappelijk Onderzoek (FWO), Flanders , G017225N References [1]. ↵ A. V Khera et al. , “ Genome-wide polygenic scores for common diseases identify individuals with risk equivalent to monogenic mutations ,” Nat Genet , vol. 50 , no. 9 , pp. 1219 – 1224 , 2018 , doi: 10.1038/s41588-018-0183-z . OpenUrl CrossRef [2]. ↵ N. Chatterjee , J. Shi , and M. García-Closas , “ Developing and evaluating polygenic risk prediction models for stratified disease prevention ,” Nat Rev Genet , vol. 17 , no. 7 , pp. 392 – 406 , 2016 , doi: 10.1038/nrg.2016.27 . OpenUrl CrossRef [3]. ↵ A. Torkamani , N. E. Wineinger , and E. J. Topol , “ The personal and clinical utility of polygenic risk scores ,” Nat Rev Genet , vol. 19 , no. 9 , pp. 581 – 590 , 2018 , doi: 10.1038/s41576-018-0018-x . OpenUrl CrossRef PubMed [4]. ↵ S. A. Lambert , G. Abraham , and M. Inouye , “ Towards clinical utility of polygenic risk scores ,” Hum Mol Genet , vol. 28 , no. R2 , pp. R133 – R142 , Nov . 2019 , doi: 10.1093/hmg/ddz187 . OpenUrl CrossRef PubMed [5]. ↵ I. J. Kullo , C. M. Lewis , M. Inouye , A. R. Martin , S. Ripatti , and N. Chatterjee , “ Polygenic scores in biomedical research ,” Nat Rev Genet , vol. 23 , no. 9 , pp. 524 – 532 , 2022 , doi: 10.1038/s41576-022-00470-z . OpenUrl CrossRef PubMed [6]. ↵ N. R. Wray et al. , “ From Basic Science to Clinical Application of Polygenic Risk Scores: A Primer ,” JAMA Psychiatry , vol. 78 , no. 1 , pp. 101 – 109 , Jan . 2021 , doi: 10.1001/jamapsychiatry.2020.3049 . OpenUrl CrossRef PubMed [7]. ↵ D. Sero et al. , “ Facial recognition from DNA using face-to-DNA classifiers ,” Nat Commun , vol. 10 , no. 1 , p. 2557 , 2019 , doi: 10.1038/s41467-019-10617-y . OpenUrl CrossRef PubMed [8]. ↵ B. Hallgrimsson , W. Mio , R. S. Marcucio , and R. Spritz , “ Let’s Face It—Complex Traits Are Just Not That Simple ,” PLoS Genet , vol. 10 , no. 11 , pp. e1004724. , Nov . 2014 , [Online]. Available : doi: 10.1371/journal.pgen.1004724 OpenUrl CrossRef PubMed [9]. ↵ P. Claes and M. D. Shriver , “ Establishing a Multidisciplinary Context for Modeling 3D Facial Shape from DNA ,” PLoS Genet , vol. 10 , no. 11 , pp. e1004725. , Nov . 2014 , [Online]. Available : doi: 10.1371/journal.pgen.1004725 OpenUrl CrossRef [10]. ↵ S. Naqvi et al. , “ Decoding the Human Face: Progress and Challenges in Understanding the Genetics of Craniofacial Morphology ,” Annu Rev Genomics Hum Genet , vol. 23 , no. 1 , pp. 383 – 412 , Aug . 2022 , doi: 10.1146/annurev-genom-120121-102607 . OpenUrl CrossRef [11]. ↵ D. Liu et al. , “ Impact of low-frequency coding variants on human facial shape ,” Sci Rep , vol. 11 , no. 1 , p. 748 , 2021 , doi: 10.1038/s41598-020-80661-y . OpenUrl CrossRef [12]. ↵ G. Ni et al. , “ A Comparison of Ten Polygenic Score Methods for Psychiatric Disorders Applied Across Multiple Cohorts ,” Biol Psychiatry , vol. 90 , no. 9 , pp. 611 – 620 , 2021 , doi: 10.1016/j.biopsych.2021.04.018 . OpenUrl CrossRef [13]. ↵ S. M. Purcell et al. , “ Common polygenic variation contributes to risk of schizophrenia and bipolar disorder ,” Nature , vol. 460 , no. 7256 , pp. 748 – 752 , 2009 , doi: 10.1038/nature08185 . OpenUrl CrossRef PubMed Web of Science [14]. ↵ J. Euesden , C. M. Lewis , and P. F. O’Reilly , “ PRSice: Polygenic Risk Score software ,” Bioinformatics , vol. 31 , no. 9 , pp. 1466 – 1468 , May 2015 , doi: 10.1093/bioinformatics/btu848 . OpenUrl CrossRef PubMed [15]. ↵ S. W. Choi and P. F. O’Reilly , “ PRSice-2: Polygenic Risk Score software for biobank-scale data ,” Gigascience , vol. 8 , no. 7 , p. giz082 , Jul . 2019 , doi: 10.1093/gigascience/giz082 . OpenUrl CrossRef PubMed [16]. ↵ W. Jiang , L. Chen , M. J. Girgenti , and H. Zhao , “ Tuning parameters for polygenic risk score methods using GWAS summary statistics from training data ,” Nat Commun , vol. 15 , no. 1 , p. 24 , 2024 , doi: 10.1038/s41467-023-44009-0 . OpenUrl CrossRef PubMed [17]. ↵ T. S. H. Mak , R. M. Porsch , S. W. Choi , X. Zhou , and P. C. Sham , “ Polygenic scores via penalized regression on summary statistics ,” Genet Epidemiol , vol. 41 , no. 6 , pp. 469 – 480 , Sep . 2017 , doi: 10.1002/gepi.22050 . OpenUrl CrossRef PubMed [18]. ↵ F. Privé , J. Arbel , H. Aschard , and B. J. Vilhjálmsson , “ Identifying and correcting for misspecifications in GWAS summary statistics and polygenic scores ,” Human Genetics and Genomics Advances , vol. 3 , no. 4 , Oct . 2022 , doi: 10.1016/j.xhgg.2022.100136 . OpenUrl CrossRef PubMed [19]. ↵ B. J. Vilhjálmsson et al. , “ Modeling Linkage Disequilibrium Increases Accuracy of Polygenic Risk Scores ,” The American Journal of Human Genetics , vol. 97 , no. 4 , pp. 576 – 592 , Oct . 2015 , doi: 10.1016/j.ajhg.2015.09.001 . OpenUrl CrossRef PubMed [20]. ↵ F. Privé , J. Arbel , and B. J. Vilhjálmsson , “ LDpred2: better, faster, stronger ,” Bioinformatics , vol. 36 , no. 22–23 , pp. 5424 – 5431 , Apr . 2021 , doi: 10.1093/bioinformatics/btaa1029 . OpenUrl CrossRef PubMed [21]. ↵ T. Ge , C.-Y. Chen , Y. Ni , Y.-C. A. Feng , and J. W. Smoller , “ Polygenic prediction via Bayesian regression and continuous shrinkage priors ,” Nat Commun , vol. 10 , no. 1 , p. 1776 , 2019 , doi: 10.1038/s41467-019-09718-5 . OpenUrl CrossRef PubMed [22]. ↵ Y. Wang et al. , “ Polygenic prediction across populations is influenced by ancestry, genetic architecture, and methodology ,” Cell Genomics , vol. 3 , no. 10 , Oct . 2023 , doi: 10.1016/j.xgen.2023.100408 . OpenUrl CrossRef [23]. ↵ S. Gunn et al. , “ Comparison of methods for building polygenic scores for diverse populations ,” Human Genetics and Genomics Advances , vol. 6 , no. 1 , p. 100355 , 2025 , doi: 10.1016/j.xhgg.2024.100355 . OpenUrl CrossRef [24]. ↵ T. E. Galesloot , K. van Steen , L. A. L. M. Kiemeney , L. L. Janss , and S. H. Vermeulen , “ A Comparison of Multivariate Genome-Wide Association Methods ,” PLoS One , vol. 9 , no. 4 , pp. e95923. , Apr . 2014 , [Online].Available : doi: 10.1371/journal.pone.0095923 OpenUrl CrossRef PubMed [25]. ↵ M. Yuan et al. , “ Mapping genes for human face shape: Exploration of univariate phenotyping strategies ,” PLoS Comput Biol , vol. 20 , no. 12 , pp. e1012617-, Dec . 2024 , [Online]. Available : doi: 10.1371/journal.pcbi.1012617 OpenUrl CrossRef [26]. ↵ M. Yuan et al. , “ Optimized phenotyping of complex morphological traits: enhancing discovery of common and rare genetic variants ,” Brief Bioinform , vol. 26 , no. 2 , p. bbaf090 , Mar . 2025 , doi: 10.1093/bib/bbaf090 . OpenUrl CrossRef PubMed [27]. ↵ C. Bycroft et al. , “ The UK Biobank resource with deep phenotyping and genomic data ,” Nature , vol. 562 , no. 7726 , pp. 203 – 209 , 2018 , doi: 10.1038/s41586-018-0579-z . OpenUrl CrossRef PubMed [28]. ↵ J. D. White et al. , “ Insights into the genetic architecture of the human face ,” Nat Genet , vol. 53 , no. 1 , pp. 45 – 53 , 2021 , doi: 10.1038/s41588-020-00741-7 . OpenUrl CrossRef PubMed [29]. ↵ M. Yuan et al. , “ Data-driven trait heritability-based extraction of human facial phenotypes ,” in 2023 IEEE International Conference on Bioinformatics and Biomedicine (BIBM) , 2023 , pp. 312 – 319 . doi: 10.1109/BIBM58861.2023.10385885 . OpenUrl CrossRef [30]. ↵ C. J. Percival et al. , “ The effect of automated landmark identification on morphometric analyses ,” J Anat , vol. 234 , no. 6 , pp. 917 – 935 , Jun . 2019 , doi: 10.1111/joa.12973 . OpenUrl CrossRef PubMed [31]. ↵ S. Katina et al. , “ The definitions of three-dimensional landmarks on the human face: an interdisciplinary view ,” J Anat , vol. 228 , no. 3 , pp. 355 – 365 , Mar . 2016 , doi: 10.1111/joa.12407 . OpenUrl CrossRef PubMed [32]. ↵ F. Privé , H. Aschard , A. Ziyatdinov , and M. G. B. Blum , “ Efficient analysis of large-scale genome-wide data with two R packages: bigstatsr and bigsnpr ,” Bioinformatics , vol. 34 , no. 16 , pp. 2781 – 2787 , Aug . 2018 , doi: 10.1093/bioinformatics/bty185 . OpenUrl CrossRef [33]. ↵ P. M. Visscher , W. G. Hill , and N. R. Wray , “ Heritability in the genomics era — concepts and misconceptions ,” Nat Rev Genet , vol. 9 , no. 4 , pp. 255 – 266 , 2008 , doi: 10.1038/nrg2322 . OpenUrl CrossRef PubMed Web of Science [34]. ↵ C. Manzoni et al. , “ Genome, transcriptome and proteome: the rise of omics data and their integration in biomedical sciences ,” Brief Bioinform , vol. 19 , no. 2 , pp. 286 – 302 , Mar . 2018 , doi: 10.1093/bib/bbw114 . OpenUrl CrossRef PubMed [35]. ↵ A. Torkamani , N. E. Wineinger , and E. J. Topol , “ The personal and clinical utility of polygenic risk scores ,” Nat Rev Genet , vol. 19 , no. 9 , pp. 581 – 590 , 2018 , doi: 10.1038/s41576-018-0018-x . OpenUrl CrossRef PubMed [36]. ↵ N. R. Wray , J. Yang , B. J. Hayes , A. L. Price , M. E. Goddard , and P. M. Visscher , “ Pitfalls of predicting complex traits from SNPs ,” Nat Rev Genet , vol. 14 , no. 7 , pp. 507 – 515 , 2013 , doi: 10.1038/nrg3457 . OpenUrl CrossRef PubMed [37]. ↵ Y. Ruan et al. , “ Improving polygenic prediction in ancestrally diverse populations ,” Nat Genet , vol. 54 , no. 5 , pp. 573 – 580 , 2022 , doi: 10.1038/s41588-022-01054-7 . OpenUrl CrossRef [38]. ↵ S. Du et al. , “ A multi-ancestry GWAS meta-analysis of facial features and its application in predicting archaic human features ,” Journal of Genetics and Genomics , vol. 52 , no. 4 , pp. 513 – 524 , 2025 , doi: 10.1016/j.jgg.2024.07.005 . OpenUrl CrossRef [39]. ↵ A. Manichaikul , J. C. Mychaleckyj , S. S. Rich , K. Daly , M. Sale , and W.-M. Chen , “ Robust relationship inference in genome-wide association studies ,” Bioinformatics , vol. 26 , no. 22 , pp. 2867 – 2873 , Nov . 2010 , doi: 10.1093/bioinformatics/btq559 . OpenUrl CrossRef PubMed Web of Science [40]. ↵ J. C. Hayton , D. G. Allen , and V. Scarpello , “ Factor Retention Decisions in Exploratory Factor Analysis: a Tutorial on Parallel Analysis ,” Organ Res Methods , vol. 7 , no. 2 , pp. 191 – 205 , Apr . 2004 , doi: 10.1177/1094428104263675 . OpenUrl CrossRef Web of Science [41]. ↵ S. B. Franklin , D. J. Gibson , P. A. Robertson , J. T. Pohlmann , and J. S. Fralish , “ Parallel Analysis: a method for determining significant principal components ,” Journal of Vegetation Science , vol. 6 , no. 1 , pp. 99 – 106 , Feb . 1995 , doi: 10.2307/3236261 . OpenUrl CrossRef [42]. ↵ J. Yang et al. , “ Common SNPs explain a large proportion of the heritability for human height ,” Nat Genet , vol. 42 , no. 7 , pp. 565 – 569 , 2010 , doi: 10.1038/ng.608 . OpenUrl CrossRef PubMed Web of Science [43]. ↵ J. Yang , J. Zeng , M. E. Goddard , N. R. Wray , and P. M. Visscher , “ Concepts, estimation and interpretation of SNP-based heritability ,” Nat Genet , vol. 49 , no. 9 , pp. 1304 – 1310 , 2017 , doi: 10.1038/ng.3941 . OpenUrl CrossRef PubMed [44]. ↵ J. Li et al. , “ Robust genome-wide ancestry inference for heterogeneous datasets: illustrated using the 1,000 genome project with 3D facial images ,” Sci Rep , vol. 10 , no. 1 , Dec . 2020 , doi: 10.1038/s41598-020-68259-w . OpenUrl CrossRef View the discussion thread. Back to top Previous Next Posted September 17, 2025. Download PDF Data/Code Email Thank you for your interest in spreading the word about bioRxiv. NOTE: Your email address is requested solely to identify you as the sender of this article. Your Email * Your Name * Send To * Enter multiple addresses on separate lines or separate them with commas. You are going to email the following Optimizing Polygenic Scores for Complex Morphological Traits: A Case Study in Nasal Shape Prediction Message Subject (Your Name) has forwarded a page to you from bioRxiv Message Body (Your Name) thought you would like to see this page from the bioRxiv website. Your Personal Message CAPTCHA This question is for testing whether or not you are a human visitor and to prevent automated spam submissions. Share Optimizing Polygenic Scores for Complex Morphological Traits: A Case Study in Nasal Shape Prediction Meng Yuan , Seppe Goovaerts , Nina Claessens , Jay Devine , Sara Becelaere , Isabelle Cleynen , Peter Claes bioRxiv 2025.09.14.676081; doi: https://doi.org/10.1101/2025.09.14.676081 Share This Article: Copy Citation Tools Optimizing Polygenic Scores for Complex Morphological Traits: A Case Study in Nasal Shape Prediction Meng Yuan , Seppe Goovaerts , Nina Claessens , Jay Devine , Sara Becelaere , Isabelle Cleynen , Peter Claes bioRxiv 2025.09.14.676081; doi: https://doi.org/10.1101/2025.09.14.676081 Citation Manager Formats BibTeX Bookends EasyBib EndNote (tagged) EndNote 8 (xml) Medlars Mendeley Papers RefWorks Tagged Ref Manager RIS Zotero Tweet Widget Facebook Like Google Plus One Subject Area Bioinformatics Subject Areas All Articles Animal Behavior and Cognition (7635) Biochemistry (17697) Bioengineering (13895) Bioinformatics (41951) Biophysics (21456) Cancer Biology (18594) Cell Biology (25520) Clinical Trials (138) Developmental Biology (13381) Ecology (19903) Epidemiology (2067) Evolutionary Biology (24323) Genetics (15612) Genomics (22510) Immunology (17738) Microbiology (40401) Molecular Biology (17184) Neuroscience (88622) Paleontology (667) Pathology (2833) Pharmacology and Toxicology (4825) Physiology (7644) Plant Biology (15158) Scientific Communication and Education (2046) Synthetic Biology (4296) Systems Biology (9825) Zoology (2271)

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00