{"paper_id":"57470a69-8960-4701-b042-c16685254727","body_text":"Choice of phenotype scale is critical in biobank-based G×E tests \n \nManuela Costantino1, Renée Fonseca2, Zhengtong Liu3, Zhenhong Huang1, Sriram \nSankararaman3,4,5, Iain Mathieson6*, Andy Dahl1* \n \n1 Section of Genetic Medicine, University of Chicago, Chicago, IL, USA \n2 Department of Human Genetics, University of Chicago, Chicago, IL, USA \n3 Department of Computer Science, UCLA, Los Angeles, CA, USA \n4 Department of Human Genetics, David Geffen School of Medicine, UCLA, Los Angeles, CA, USA \n5 Department of Computational Medicine, David Geffen School of Medicine, UCLA, Los Angeles, CA, USA \n6 Department of Genetics, Perelman School of Medicine, University of Pennsylvania, Philadelphia, PA, USA \n \n* These authors jointly supervised the work \n \nCorresponding author: Manuela Costantino (mcostantino@uchicago.edu) \n \nAbstract \n \nThe importance of gene-environment interactions (G×E) for complex human traits is heavily \ndebated. Recently, biobank-based GWAS have revealed many statistically significant G ×E \nsignals, though most lack clear evidence of biological significance. Here, we partly explain this \ndiscrepancy by showing that many G ×E signals simplify to additive effects on a different \nphenotype scale, a classical concern that is currently underappreciated. Our results clearly \ndistinguish G×Sex effects on height, which vanish on the log scale, from G×Sex effects on \ntestosterone, where the log scale uncovers biologically meaningful female-specific effects. \nAcross 32 phenotypes in UK Biobank, we find that scaling by a power transformation can \nexplain 46% of PGS×Sex interactions, and that simple log transformation can explain 23%, with \nsimilar results for other environments. We also show that phenotype scale can substantially \nimpact GWAS discovery and the construction and evaluation of polygenic scores. Finally, we \nprovide a set of guidelines to consider and choose phenotype scale in modern genetic studies.  \n \nIntroduction \n \nGene-environment (G×E) interaction models are a classical framework for studying how genetic \narchitecture depends on the environment (1,2). G×E effects are well established in  model \nsystems, such as yeast fitness depending on growth medium or gene regulatory effects \ndepending on chemical exposures (3–6). Examples of biomedically significant G×E have also \nbeen identified in complex human diseases, such as the sex-specific effect of KLF14 variants on \ntype 2 diabetes (7) and the 17q locus effect on asthma that depends on childhood rhinovirus \nillness (8). \n \nNonetheless, the importance of G×E for complex human traits generally remains unclear. Early \nattempts to identify G×E in candidate gene studies produced largely spurious results, but this \ncan be simply attributed to underpowered studies (9,10). With the availability of biobank-based \nGWAS datasets, however, significant G×E signals have become the rule rather than the \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 21, 2026. ; https://doi.org/10.64898/2026.01.20.694695doi: bioRxiv preprint \n\n \nexception (11–15). However, these large studies face the opposite problem because they have \npower to detect subtle technical artifacts (16). For example, heteroskedasticity must be carefully \nmodelled to avoid pervasive false positive G×E in TWAS and heritability estimation (15,17–19). \n \nHere, we explore an underappreciated choice in G×E studies: the scale on which phenotypes \nare analyzed. For example, because males are typically taller than females, a genetic variant \nthat increases height by 1% will have a greater effect in males on the default centimeter scale. It \nhas long been known that phenotype scale is critical in G×E studies (2,20,21). However, current \nG×E studies often ignore this decision, implicitly choosing whatever default scale is provided \nwithout discussion or robustness analyses. \n \nWe establish a framework to distinguish scale-dependent and -independent G×E by identifying \nthe most-additive phenotype scale. We focus on a G×E test based on polygenic scores \n(PGSxE), which has become the predominant model for G×E in biobanks for its simplicity and \npower (22). We use a range of phenotypes in UK Biobank to understand which G×E analyses \nand which phenotypes are liable to scale-dependent bias. We then characterize the impact of \nphenotype scale on GWAS and PRS, which is substantial for some phenotypes. We conclude \nby discussing best practices for phenotype scaling to prioritize biologically meaningful genetic \ndiscoveries when the optimal phenotype scale is unknown. \n \nResults \n \nScale-dependent and -independent statistical interactions in theory \n \nThe existence and magnitude of statistical interactions depend on phenotype measurement \nscale. By “scale,” we refer to monotonic transformations of a nonnegative quantitative \nphenotype (21,23), which preserve the phenotype’s rank order. For a given dataset, if an \ninteraction is statistically significant on all considered scales, we call it scale-independent; \notherwise, we call it scale-dependent. Our definition explicitly depends on power because we \nare focused on the new interactions emerging from biobank-scale data. \n \nAny non-linear scale transformation on an additive phenotype will induce interactions with \nsufficient power (Note S2). But, perhaps surprisingly, scale-independent interactions exist. A \nsimple example is the exclusive or (XOR), where a genetic factor, G, increases the phenotype in \none environment but decreases it in another; regardless of the phenotype scale, the sign of G’s \nexpected effect will flip based on E (Note S1.1). Generalizing this idea, we prove that G×E is \nscale-independent for Gaussian G and E if the sign of E determines the sign of G’s effect (Note \nS1.2); however, even in this simple case, scale-independence is not robust to outliers (Note \nS1.3). The former conclusion is consistent with prior work relating the sign of G’s effect to  \nscale-independent G×E (24,25) or “crossover” in reaction norm models (26).  \n \n \n \n \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 21, 2026. ; https://doi.org/10.64898/2026.01.20.694695doi: bioRxiv preprint \n\n \nAdditive genetic scales can be recovered in simulations \n \nWe focus on a simple approach to study G×E using polygenic scores (PGS), which fits PGS×E \ninteraction to test if the PGS effect depends on E. This models the component of G×E that is \ncorrelated with additive genetic effects, with E systematically amplifying or dampening the \nadditive effects (27). We fit this model for a range of scales defined as power transformations, \ni.e., exponentiating the phenotype by λ (Fig 1). When λ=1, no transformation is applied; an \nimportant special case is λ=0, which corresponds to log transformation. We then profile PGSxE \np-values across λ from -1 to 2, asking if some λ recovers a truly additive scale.  \n \nTo test if our approach could distinguish scale-dependent vs -independent PGS×E, we \nsimulated a latent phenotype and then transformed it to mimic an observed phenotype, varying \nif the latent phenotype is additive (Fig 1A-C) or has G×E (Fig 1D-F ). When the latent scale was \ndirectly observed, PGS×E tests were calibrated on the observed scale, as expected; in fact, in \nthis setting, applying a power transformation creates scale-dependent PGS×E (Fig 1A). When \nthe phenotype is observed on the exponential scale, however, PGS×E tests were severely \ninflated; as expected, this was corrected by log transformation (Fig 1B). Notably, we do not \nobserve power transformations meaningfully reducing power for when PGS×E is truly present \non the latent scale (Fig 1D-F ). Finally, our approach is unable to remove scale-dependent \nPGS×E when the inverse-logit scale is observed, as expected because this is not a power \ntransformation (Fig 1C).  \n \nWe also tested the rank-inverse normal transformation (RINT), which scales the phenotype to \nhave approximately Gaussian quantiles. RINT did not depend on the observed scale, as \nexpected, and always gave modest inflation for PGS×E effects (FPR=0.073), even when the \nadditive scale is observed.  \n \nIn conclusion, separating scale-dependent and -independent G×E is challenging but possible in \ntheory and practice; more importantly, careful consideration of phenotype scale can provide \ninsights into the nature of G×E. \n \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 21, 2026. ; https://doi.org/10.64898/2026.01.20.694695doi: bioRxiv preprint \n\n \n \nFigure 1 - Power transformation removes some forms of scale-dependent G×E in simulations: \nEach blue line shows the results for testing the interaction between a PGS and an environment, E, \naveraged over simulation replicates. The Y axis shows the p-value of the PGS×E interaction term and the \ndotted black line shows p=0.05 after Bonferroni correction. The X axis denotes the different power (or \nBox-Cox) transformations applied to the phenotype with parameter λ (Methods). Importantly, λ=1 \ncorresponds to no transformation and λ=0 to log transformation. \n \nScale-dependent G×Sex effects on height are removed by log transformation \n \nWe first study height, a classic additive trait in human genetics (28,29). We applied our power \ntransformation procedure to evaluate scale-dependent G×E in ~300,000 White British \nindividuals in UK Biobank (30) (Methods). We tested for PGS×E interactions between 5 \nenvironmental variables (E) and 36 polygenic sores that were computed from GWAS that do not \ninclude UKB individuals (Table S3) (31). \n \nWe first study sex as an environment. Surprisingly, despite the limited evidence for sex-specific \ngenetic effects on height (32), we find a highly significant PGS×Sex effect (p=2e-29, Fig 2A). \nNonetheless, this unexpected effect can be eliminated by log transformation (p=0.09), i.e., \npower transformation with λ=0. This suggests that genetic effects on height act multiplicatively \nand that height should be studied on the log scale, which differs from common practice in \nGWAS (28) but is established in cross-species comparisons (33) and has been shown in \nhumans (34). Across the four other E, we find one additional PGS×E that vanishes on the log \nscale (alcohol intake frequency), two that were not significant on any scale (smoking status, \nstatin usage), and one that was significant on all tested scales (age, Fig S1).  \n \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 21, 2026. ; https://doi.org/10.64898/2026.01.20.694695doi: bioRxiv preprint \n\n \nWe then compared the genetic architecture of height and log-height. We found that their GWAS \nloci were nearly identical (97.1% overlap, Fig 2B; Manhattan and QQ plots in Fig S2). Both \nscales also have similar additive heritabilities and near-perfect genetic correlations with each \nother and sex-specific GWAS (Fig 2C, sex-specific Manhattan plots in Fig S5). Furthermore, we \nfound that the PGS constructed from the log and default GWAS gave similar prediction accuracy \n(Fig 2D, Methods). This consistency across scales and sexes is because the log transformation \nhas minimal impact on the distribution of height (Fig 2E), providing an almost linear mapping \n(Fig 2F ). Overall, the PGS×Sex effect on height illustrates how phenotype scale can \ndramatically affect G×E tests even when it has negligible impact on additive genetic \narchitecture. \n \n \n \nFigure 2 - Scale-Dependence in Height: A. Results of the PGS×Sex analysis. Each gray line shows the \np-values of the interaction term between a PGS and sex. The dotted black line shows the Bonferroni \nsignificance threshold. The height PGS is labeled for clarity. B. Number of GWAS hits when performed on \nthe log scale vs. the default scale. C. Genetic correlations (lower triangle) and heritabilities (diagonal) for \nGWAS of height, log-height, and sex-specific variants. D. Prediction accuracy for the PGS constructed \nfrom the default scale GWAS vs the exponentiated PGS constructed from log-scale GWAS; accuracy is \nmeasured by Pearson R2 on the default scale and calculated in either the entire sample or separately in \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 21, 2026. ; https://doi.org/10.64898/2026.01.20.694695doi: bioRxiv preprint \n\n \nfemales and males. E. Distribution of height on the default and the log scale, stratified by sex. F. Effect of \nlog transformation on height phenotype values.  \n \n \nG×Sex effects on testosterone are largely scale-independent \n \nWe next studied an exemplary sex-specific trait to contrast with height: testosterone. Unlike in \nheight, PGS×Sex effects largely remained significant for all tested power transformations (Fig \n3A), suggesting they are scale-independent. Intuitively, scale-independent interactions are \nexpected because the biology of testosterone qualitatively differs between males and females \n(35). Supporting this interpretation, all significant PGS×E interactions for the other four \nenvironments were scale-dependent (Fig S3). \n \nWe then compared the testosterone GWAS on the log vs default (nmol/L) scales. Unlike for \nheight, most GWAS hits for testosterone are only detected on one scale (27.6% overlap, Figs \n3B and S4). We expected this discrepancy arose because males explain the vast majority of \nphenotypic variance on the default scale but not the log scale (Fig 3E). Indeed, we found that \nthe default scale has genetic correlation of 0.99 and 0.15 with the male- and female-specific \nGWAS (Fig3C), respectively, and that its GWAS loci primarily overlap the male-specific GWAS \n(87.2%, sex-specific Manhattan plots in Fig S6). Conversely, the log scale GWAS had much \nmore similar genetic correlations with the sex-specific GWAS (0.70 and 0.74, Fig3C). For \nexample, the female-specific effect associated with FGF9, which has been linked to gonad \ndevelopment (35–38), is significant on the log scale (p=4.29e-9) but not on the default scale \n(p=0.13).  \n \nAs the log scale improves detection of female-specific genetic effects on testosterone, we \nexpected the log-scale PGS would improve prediction. Surprisingly, this reduced Pearson R2 \n(from 2.50% to 1.52%, Fig 3D). However, this is explained by the fact that Pearson R2 is itself \nscale-dependent: while the Pearson R 2 does indeed dramatically increase for females (13-fold, \nfrom 0.05% to 0.65%), it decreases in males (from 5.41% to 2.83%), constituting a net loss \nbecause Pearson R2 weights groups by phenotypic variance (Fig 3E,F ). On the other hand, \nevaluating PGS R2 on the log scale gives the expected result, where the PGS constructed from \nthe log scale performs better (from 1.02% to 1.74%). This demonstrates how the choice of \nphenotype scale affects not only PGS but also the relative prioritization of individuals in \nevaluation metrics. We observe similar results when using Spearman’s rank correlation \ncoefficient to calculate R2 (Fig S8B), which is a scale-independent metric but remains sensitive \nto the GWAS scale used to construct PGS. \n \nOverall, the choice of phenotype scale qualitatively changes inferences on the nature of additive \nand GxSex effects on testosterone. \n \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 21, 2026. ; https://doi.org/10.64898/2026.01.20.694695doi: bioRxiv preprint \n\n \n \nFigure 3 - Scale-Dependence in Testosterone: A. Results of the PGS×Sex analysis. Each gray line \nshows the p-values of the interaction term between a PGS and sex. The dotted black line shows the \nBonferroni significance threshold.  The four most significant PGS are labeled for clarity. The testosterone \nPGS (in orange) was constructed in UK Biobank due to lack of external GWAS (Methods). B. Number of \nGWAS hits when performed on the log scale vs. the default scale. C. Genetic correlations (lower triangle) \nand heritabilities (diagonal) for GWAS of testosterone, log-testosterone, and sex-specific variants. D. \nPrediction accuracy for the PGS constructed from the default scale GWAS vs the exponentiated PGS \nconstructed from log-scale GWAS; accuracy is measured by Pearson R2 on the default scale and \ncalculated in either the entire sample or separately in females and males. PGS are constructed either \nfrom default scale GWAS, or by exponentiating the PGS constructed from log-scale GWAS. E. \nDistribution of testosterone on the default and the log scale, stratified by sex. F. Effect of log \ntransformation on testosterone phenotype values.  \n \n \nScale-dependent and -independent G×E are pervasive in biobank-scale data  \n \nWe expanded our PGS×Sex analysis to 32 quantitative phenotypes (Table S2). Across PGS \nand phenotypes, we find 188 PGS×Sex interactions that are significant on the default scale \n(15.3% of the total tests, p<0.05/36). Of these, 46.3% become non-significant after power \ntransformation with some λ ∈  [-1,2], and 23.4% are eliminated simply by log transformation (i.e. \nλ=0, Fig 4A). Some phenotypes, such as average fat-free arm mass, mainly have \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 21, 2026. ; https://doi.org/10.64898/2026.01.20.694695doi: bioRxiv preprint \n\n \nscale-dependent interactions with sex (10/11 vanish on the log scale), while others, such as \ndiastolic blood pressure, only have scale-independent interactions. We find similar patterns for \nthe other four environments, where 7-21% of PGS×E effects are not significant after log \ntransformation (Fig 4B, S10). The results are robust when we instead define statistical \nsignificance by p<0.05 (Fig S11). \n \nWe also find examples of interactions that cannot be eliminated by power transformation and yet \nare likely scale-dependent (Fig S9). For example, we find complex interactions for biomarkers \nthat are targeted by common drugs, like LDL, which may be eliminated by more sophisticated \ntransformations (17). This highlights how our power transformation approach only provides a \nlower bound for the severity of scale-dependent bias in G×E. \n \nIt is known that scale transformations introduce stronger mean-variance relationships when the \ncoefficient of variation (CV, the ratio of the phenotype’s standard deviation to its mean) is higher \n(23,39), as we observed in our data (Fig S12.A). To ask how CV impacts interactions, we \ncompared cross-sex genetic correlations on the default vs log scale. Across phenotypes, we \nfind a significant relationship between CV and the change in cross-sex genetic correlation after \nlog scaling (p=1.2e7, Fig 4C, Fig S12.B). We also find a significant relationship between CV \nand the change in R2 from constructing the PGS on the default vs log scale (p=1.08e-0.2, Fig \nS12.C). We conclude that the coefficient of variation is one ingredient determining the \nrobustness of a phenotype's genetic architecture to scale transformation. \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 21, 2026. ; https://doi.org/10.64898/2026.01.20.694695doi: bioRxiv preprint \n\n \n \nFigure 4 - Scale-Dependence in 32 complex traits: A. Results of PGS×Sex analysis across 32 traits. \nGrey cells show interactions that were insignificant on the default scale; dark purple cells show those that \nwere significant on both the default and the log scale; and light purple cells show those that were \nsignificant on the default scale but not on the log scale. B. Number of scale-dependent and \nscale-independent PGS×E interactions found across all PGS and all phenotypes for each environment. C. \nRelationship between the coefficient of variation and the absolute value of the change in cross-sex \ngenetic correlation after log scaling the phenotype; relative changes are consistent (Fig S12.B) .  \n \n \nDiscussion \n \nWe found that choice of phenotype scale can qualitatively change inferences about the \nexistence and nature of gene-environment interactions (G×E) in complex human traits. Although \nscale-dependence is a textbook concern, its dramatic impact on modern G×E studies is \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 21, 2026. ; https://doi.org/10.64898/2026.01.20.694695doi: bioRxiv preprint \n\n \nunderappreciated. This especially applies to PGS×E models, which are increasingly common \n(22). Concretely, across 32 phenotypes in UK Biobank, we find that 23.4% of significant \nPGS×Sex vanish after simply log transforming. In some cases, we also find that GWAS loci and \nPGS construction and evaluation depend heavily on scale. In general, we find that phenotype \nscale is more important for phenotypes with larger coefficients of variation (CV). \n \nWe contrast G×Sex effects at two ends of the phenotype spectrum: height and testosterone. In \nheight, which is generally considered additive (40), we found highly significant PGS×Sex effects \nthat vanish after log transformation. Nonetheless, its GWAS and PGS are negligibly impacted by \nchoice of scale, partly because height has a low CV. By contrast, in testosterone, which has \nsex-specific biology (35), we found that G×Sex interactions could not be eliminated by power \ntransformation. Yet still, we found that scale impacts biological discoveries: the log scale GWAS \nfound female-specific effects that are missed by the default scale, such as the FGF9 effect on \ngonad development.  \n \nPhenotype scale can also impact PGS construction and evaluation. When evaluating Pearson \nR2 on the default scale, the optimal scale for PGS construction (log or default) varied across \nphenotypes (Fig S8.A), which also held for the scale-independent Spearman R2 (Fig S8.B). By \ncontrast, the log-constructed PGS was generally superior when evaluating Pearson R2 on the \nlog scale (Fig S8.C). Interestingly, we found that PGS evaluation metrics weigh individuals \ndifferently depending on scale. For example, evaluating Pearson R2 for testosterone on the \ndefault scale is essentially equivalent to ignoring females. Thus, seemingly subtle modifications \nto phenotype scale can dramatically impact PGS accuracy and equity.  \n \nAcross phenotypes, we found that significant PGS×Sex interactions often vanish after power \ntransformation. Many studies have interpreted PGS×E signals as the environment \nmechanistically amplifying or buffering additive genetic effects (27,41–49). However, \namplification signals are expected for log-additive phenotypes; in fact, they are equivalent to first \norder for binary E (Note S2.1). Thus, measurement scale is a major challenge in identifying \nmeaningful forms of amplification.  \n \nOther interaction tests are also affected by phenotype scale. For instance, epistasis (G×G) \nstudies in deep mutational scans historically used statistical tests prone to scale artifacts \n(41,50–52). Phenotype scale is less relevant to epistasis in complex traits because SNP effects \nare typically weak (53) and PGS are themselves additive (49); however, scale is likely relevant \nto epistasis between large-effect variants or burden scores. Phenotype scale also applies to \nbinary phenotypes in the form of link function choice, such as probit vs logit (Fig S13, (54). More \nbroadly, phenotype scale is liable to confound any non-additive statistical test, such as \nquantile-dependent effects (55,56), effects on variance (57,58), or machine learning models \n(57,58). \n \nWhat scale should be used in practice? It depends on the goal. Ideally, the scale is carefully \nchosen based on domain expertise, or prespecified goals like maximizing PRS R2 on a \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 21, 2026. ; https://doi.org/10.64898/2026.01.20.694695doi: bioRxiv preprint \n\n \nclinically-relevant scale (59). Nonetheless, even in this case, prediction may be improved by \nperforming GWAS on a different scale and then transforming back (Fig S8.A).  \n \nStatistically, the scale can be chosen for parsimony. Early work in quantitative genetics suggests \nthat optimal scales could be defined as those that minimize heteroskedasticity, epistasis, \ndominance effects, or non-Gaussianity (2,20,21,23). In the context of G×E studies, we propose \nchoosing phenotype scale to minimize G×E signals, corresponding to the null hypothesis that \nthe phenotype is additive on some scale. This is natural because significant G×E is expected \nwhen an additive phenotype is tested on a non-additive scale. We generally found that RINT \neliminated extreme false positives, though it was always slightly inflated and cannot be easily \ninterpreted. \n \nIt is much more challenging to define the optimal scale for the goal of learning biology. The link \nbetween statistical and biological interaction is tenuous in both directions: biologically \nmeaningful G×E can be statistically insignificant, and biologically-meaningless G×E is expected \nunder many forms of model misspecification, such as an incorrect phenotype scale. Further, \nscale-dependent G×E can be informative: for example, the observation that G×E effects on \nheight disappear on the log scale tells us that effects on height are generally multiplicative  (34), \nand accounting for scale-dependent interactions can reveal more meaningful interactions that \nmerit further study (41–47). Overall, scale-independence of GxE is neither necessary nor \nsufficient for biological insight. \n \nOur study is an early effort to characterize the impact of phenotype scaling on biobank-based \ngenetic studies. We focused on power transformations for simplicity, but richer functional forms \nshould be considered in the future (60). It is likely that scale transformation has the greatest \nimpact on the tails of the distribution, which are not the focus of our study but may be important \nclinically (56,59,61). Further work should also examine the impact of scale on constructed \nphenotypes, such as BMI or waist-to-hip ratio. In summary, phenotype scale is a critical \nparameter in biobank-based genetic studies, especially PGSxE, that should be carefully \nselected, reported, and discussed.  \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 21, 2026. ; https://doi.org/10.64898/2026.01.20.694695doi: bioRxiv preprint \n\n \nMethods \n \nSimulations \n \nWe simulated traits using N=100,000, m=25 and =0.3 under the following model: \n \n \n \n \nHere,  is an x  matrix of polygenic scores,  is a vector of length  of environment values \nand  is a vector of individual-level noise.  is the phenotype on the true scale, \nwhile  is the phenotype on the observed scale, which was generated using one of three \ndifferent scaling functions: ,  and . We \nperformed 100 simulations using =0 and 100 simulations using =0.1. The variance explained \nby the environment was 0.2.  \n \nWe then try to recover the additive scale, while only knowing ,  and . To this end, we fit a \nlinear model with an interaction term to 25 scaled versions of  using Box-Cox transformations \nfor  between -1 and 2. These are defined as \n \n \n \nfor . Additionally, we evaluated rank-inverse normal transformation (RINT) which \ntransforms the quantile of the phenotype to match the quantiles of a Gaussian distribution (62). \nWe determined the significance threshold using Bonferroni correction.  \n \nUK Biobank Data \n \nThe UK Biobank is a publicly available prospective dataset from the United Kingdom that \ncontains genetic and deep phenotypic data for participants aged 40-69 (30). In this study, we \nselected 32 phenotypes to be analyzed, which are listed in supplementary table 2. We manually \nconstructed hip to waist ratio adjusted for BMI, as defined by (42), hip to waist ratio, and \naverage arm fat-free mass (the mean between the right and left arms). \n \nWe excluded individuals who had missing data for the phenotype, environment, or covariates in \nany given analysis. We also excluded outlier individuals whose phenotype value fell in the top or \nbottom 0.01% of the distribution. Analyses were performed on “White British” individuals, except \nPGS R2 evaluation was performed on “White European” individuals (30,63). We also excluded \none individual from each pair of 3rd degree or closer relatives and individuals with missing \ngenotype call rates greater than 0.01.  \n \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 21, 2026. ; https://doi.org/10.64898/2026.01.20.694695doi: bioRxiv preprint \n\n \nEnvironment Data \n \nWe used five different environmental contexts in our analyses, which are described in Table S1. \nStatin usage was defined using labels defined by (17), which classified individuals as either \nusers or non-users of the drug. Smoking status was modelled as a three-level categorical factor: \nnever smoker, previous smoker and current smoker. Alcohol intake frequency was modeled as \ncontinuous, but the variable contains 6 discrete, ordinal values. The fields used for each of \nthese variables are detailed in supplementary table 1.  \n \nPGS×C models in UKB \n \nWe fit the following model, for  from -1 to 2:  \n \n \n \nHere  is the Box-Cox (power) transformed value of the phenotype for  and  is one of \nthe five environmental contexts described above.  is a standard PGS that was trained on \nan external dataset, as described by (31) (Table S3); the sole exception is the testosterone PGS \nin Fig 3.A, which was not available from external data, so we constructed it in “White British” \nand tested the interaction in “White European, as outlined below. This PGS is not generally \ntrained on the trait corresponding to .  is a matrix of covariates, which contains sex, age, \nage2, assessment center and the first 40 principal components. The assessment center is \nencoded as a categorical factor with 26 different levels. The inclusion of an interaction term \nbetween the covariates and the environmental variable (σ) ensures proper adjustment for \nconfounding effects (64). These models were fit in R v4.3.2. \n \nGWAS and Polygenic Score Construction \n \nWe ran GWAS using PLINK2 (v2.00a6LM AVX2 Intel, 4 July 2024) on two versions of each \nphenotype: once on the baseline scale and once on the log scale (65). For this analysis, we \nremoved SNPs with a missing call rate greater than 0.01, a minor allele frequency smaller than \n0.01 or a Hardy-Weinberg equilibrium exact test p-value smaller than 1×10 -6. This resulted in a \ntotal of 6,539,458 SNPs included in the final dataset. We included age, sex and the first ten \ngenetic principal components as covariates. Significant loci were reported with a threshold of \n5×10-8 after LD clumping with an r2 threshold of 0.1 and a window size of 250kb. We used 90% \nof the “White British” sample for this analysis. \n \nThese GWAS were then used to build PGS, using the same clumping parameters, and these \nwere tested on the remaining 10% of the “White British” cohort. We used a set of p-value \nthresholds, ranging from 5×10 -8 to 0.5, and selected the optimal threshold based on prediction \nperformance for the phenotype on the same scale as the GWAS was run, as measured by R2 \n(generally default-scale Pearson, log-scale Pearson for Fig S8.C, and Spearman for Fig S8.B). \n \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 21, 2026. ; https://doi.org/10.64898/2026.01.20.694695doi: bioRxiv preprint \n\n \nWe then evaluated the PGS in a separate sample, 24,950 “White European” individuals, first on \nthe entire group and then separately on males and females. For scores where the GWAS was \ntrained on the baseline scale, we used the following model. \n \n \n \nWhen the GWAS was performed on the log transformed phenotype, we used an alternative \nmodel to adjust for this change in scale. \n \n \n \nIn both models,  is a matrix of covariates, containing age, sex and the first 10 genetic principal \ncomponents. R 2 was used to evaluate each model fit and we compared the two scales using \n.  \n \nwhere we emphasize the calculation of correlations is on the default scale. We also performed \nthis procedure using Spearman’s rank correlation to define R 2, and the inverse procedure when \ntesting Pearson R2  on the log scale. \n \nHeritability and Genetic Correlation \n \nWe calculated the genetic correlation between sex-specific GWAS and GWAS on the default or \nlog scale. We obtained the GWAS summary statistics on males and females using the \nprocedure outlined above. We then used the package LDSC to obtain heritability estimates for \neach analysis and estimates of genetic correlation between each pair of analyses (66,67). We \nfiltered SNPs to the HapMap3 set and used LD scores precalculated from the 1000 Genomes \nEuropean population.  \n \nIdentifying Shared and Specific GWAS Hits \n \nTo determine the number of shared hits between a pair of GWAS, we first created a third set of \np-values, where each SNP was assigned the smallest p-value between those of the two studies \nof interest. We then clumped this third set, using plink and the parameters outlined above and \nsubset the clumped SNPs to those with a p-value below the threshold of 5e-8. Each of these \nsignificant lead hits is then either labeled as shared or only belonging to one of the other two \nGWAS, depending on whether their p-value was below or above the threshold in each study.  \n \n \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 21, 2026. ; https://doi.org/10.64898/2026.01.20.694695doi: bioRxiv preprint \n\n \nCode availability \n \nAll code used for analyses is available on Github: \nhttps://github.com/manuelacostantino1/gxe_scale.  \n \n \nAcknowledgements \n \nThis research has been conducted using the UKB Resource under Application Number 89052. \nA.D. is supported by R35GM150822. M.C. is supported by Fonds de Recherche du Québec \nSanté. We thank Carl Veller for helpful feedback on Proposition 1. We thank the participants in \nUKB for making this study possible. We thank the Center for Research Informatics for providing \nthe computing resources. The Center for Research Informatics is funded by the Biological \nSciences Division at the University of Chicago with additional funding provided by the Institute \nfor Translational Medicine, CTSA grant number 2U54TR002389-06 from the National Institutes \nof Health. \n \n \n \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 21, 2026. ; https://doi.org/10.64898/2026.01.20.694695doi: bioRxiv preprint \n\n \nReferences \n \n1. Tabery J. R. A. Fisher, Lancelot Hogben, and the origin(s) of genotype-environment \ninteraction. J Hist Biol. 2008 Winter;41(4):717–61. \n2. Walsh B, Lynch M. Evolution and selection of quantitative traits. London, England: Oxford \nUniversity Press; 2018. 1496 p. \n3. Lin W, Wall JD, Li G, Newman D, Yang Y, Abney M, et al. Genetic regulatory effects in \nresponse to a high cholesterol, high fat diet in baboons [Internet]. bioRxivorg. 2023. \nAvailable from: http://biorxiv.org/lookup/doi/10.1101/2023.08.01.551489 \n4. Smith EN, Kruglyak L. Gene-environment interaction in yeast gene expression. PLoS Biol. \n2008 Apr 15;6(4):e83. \n5. Knowles DA, Davis JR, Edgington H, Raj A, Favé MJ, Zhu X, et al. Allele-specific \nexpression reveals interactions between genetic variation and environment. Nat Methods. \n2017 Jul;14(7):699–702. \n6. Lea AJ, Peng J, Ayroles JF. Diverse environmental perturbations reveal the evolution and \ncontext-dependency of genetic effects on gene expression levels. Genome Res. 2022 \nOct;32(10):1826–39. \n7. Small KS, Todorčević M, Civelek M, El-Sayed Moustafa JS, Wang X, Simon MM, et al. \nRegulatory variants at KLF14 influence type 2 diabetes risk via a female-specific effect on \nadipocyte size and body composition. Nat Genet. 2018 Apr;50(4):572–80. \n8. Calışkan M, Bochkov YA, Kreiner-Møller E, Bønnelykke K, Stein MM, Du G, et al. \nRhinovirus wheezing illness and genetic risk of childhood-onset asthma. N Engl J Med. \n2013 Apr 11;368(15):1398–407. \n9. Caspi A, McClay J, Moffitt TE, Mill J, Martin J, Craig IW, et al. Role of genotype in the cycle \nof violence in maltreated children. Science. 2002 Aug 2;297(5582):851–4. \n10. Border R, Johnson EC, Evans LM, Smolen A, Berley N, Sullivan PF, et al. No support for \nhistorical candidate gene or candidate gene-by-interaction hypotheses for major depression \nacross multiple large samples. Am J Psychiatry. 2019 May 1;176(5):376–87. \n11. Wang H, Zhang F, Zeng J, Wu Y, Kemper KE, Xue A, et al. Genotype-by-environment \ninteractions inferred from genetic effects on phenotypic variability in the UK Biobank. Sci \nAdv. 2019 Aug;5(8):eaaw3538. \n12. Kerin M, Marchini J. Inferring gene-by-environment interactions with a Bayesian \nwhole-genome regression model. Am J Hum Genet. 2020 Oct 1;107(4):698–713. \n13. Di Scipio M, Khan M, Mao S, Chong M, Judge C, Pathan N, et al. A versatile, fast and \nunbiased method for estimation of gene-by-environment interaction effects on \nbiobank-scale datasets. Nat Commun. 2023 Aug 25;14(1):5196. \n14. Hillary RF, Gadd DA, Kuncheva Z, Mangelis T, Lin T, Ferber K, et al. Systematic discovery \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 21, 2026. ; https://doi.org/10.64898/2026.01.20.694695doi: bioRxiv preprint \n\n \nof gene-environment interactions underlying the human plasma proteome in UK Biobank. \nNat Commun. 2024 Aug 26;15(1):7346. \n15. Pazokitoroudi A, Liu Z, Dahl A, Zaitlen N, Rosset S, Sankararaman S. A scalable and \nrobust variance components method reveals insights into the architecture of \ngene-environment interactions underlying complex traits. Am J Hum Genet. 2024 Jul \n11;111(7):1462–80. \n16. Lindsay B, Liu J. Model assessment tools for a model false world. Stat Sci. 2009 Aug \n1;24(3):303–18. \n17. Sadowski M, Thompson M, Mefford J, Haldar T, Oni-Orisan A, Border R, et al. \nCharacterizing the genetic architecture of drug response using gene-context interaction \nmethods. Cell Genom. 2024 Dec 11;4(12):100722. \n18. Dahl A, Nguyen K, Cai N, Gandal MJ, Flint J, Zaitlen N. A robust method uncovers \nsignificant context-specific heritability in diverse complex traits. Am J Hum Genet. 2020 Jan \n2;106(1):71–91. \n19. Yang J, Lee SH, Goddard ME, Visscher PM. GCTA: a tool for genome-wide complex trait \nanalysis. Am J Hum Genet. 2011 Jan 7;88(1):76–82. \n20. Powers L. Determining scales and the use of transformations in studies on weight per \nlocule of tomato fruit. Biometrics. 1950 Jun;6(2):145–63. \n21. Mather K, Jinks JL. Biometrical Genetics: the study of continuous variation [Internet]. 2nd \ned. London, England: Chapman and Hall; 1971. 416 p. Available from: \nhttp://dx.doi.org/10.1007/978-1-4899-3404-8 \n22. Herrera-Luis E, Benke K, Volk H, Ladd-Acosta C, Wojcik GL. Gene-environment \ninteractions in human health. Nat Rev Genet. 2024 Nov;25(11):768–84. \n23. Falconer DS. Introduction To Quantitative Genetics 4th Edition. 1996 Feb 16; Available \nfrom: https://archive.org/details/IntroductionToQuantitativeGenetics \n24. Sadowski M, Dahl AW, Zaitlen N, Border R. The geometry of G × E: how scaling and \nendogenous treatment effects shape interaction direction [Internet]. bioRxiv. 2025. p. \n2025.07.15.664999. Available from: http://dx.doi.org/10.1101/2025.07.15.664999v1 \n25. Sverdlov S, Thompson EA. The epistasis boundary: Linear vs. Nonlinear \ngenotype-phenotype relationships [Internet]. bioRxiv. bioRxiv; 2018. p. 503466. Available \nfrom: http://dx.doi.org/10.1101/503466v1.full-text \n26. Manuck SB. The reaction norm in gene x environment interaction. Mol Psychiatry. 2010 \nSep;15(9):881–2. \n27. Peterson RE, Cai N, Dahl AW, Bigdeli TB, Edwards AC, Webb BT, et al. Molecular genetic \nanalysis subdivided by adversity exposure suggests etiologic heterogeneity in major \ndepression. Am J Psychiatry. 2018 Jun 1;175(6):545–54. \n28. Yengo L, Vedantam S, Marouli E, Sidorenko J, Bartell E, Sakaue S, et al. A saturated map \nof common genetic variants associated with human height. Nature. 2022 \nOct;610(7933):704–12. \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 21, 2026. ; https://doi.org/10.64898/2026.01.20.694695doi: bioRxiv preprint \n\n \n29. Yang J, Bakshi A, Zhu Z, Hemani G, Vinkhuyzen AAE, Nolte IM, et al. Genome-wide \ngenetic homogeneity between sexes and populations for human height and body mass \nindex. Hum Mol Genet. 2015 Dec 20;24(25):7445–9. \n30. Bycroft C, Freeman C, Petkova D, Band G, Elliott LT, Sharp K, et al. The UK Biobank \nresource with deep phenotyping and genomic data. Nature. 2018 Oct 10;562(7726):203–9. \n31. Thompson DJ, Wells D, Selzam S, Peneva I, Moore R, Sharp K, et al. UK Biobank release \nand systematic evaluation of optimised polygenic risk scores for 53 diseases and \nquantitative traits. medRxiv [Internet]. 2022; Available from: \nhttp://dx.doi.org/10.1101/2022.06.16.22276246 \n32. Randall JC, Winkler TW, Kutalik Z, Berndt SI, Jackson AU, Monda KL, et al. Sex-stratified \ngenome-wide association studies including 270,000 individuals show sexual dimorphism in \ngenetic loci for anthropometric traits. PLoS Genet. 2013 Jun;9(6):e1003500. \n33. Kozłowski J, Gawelczyk AT. Why are species’ body size distributions usually skewed to the \nright?: Species’ body size distribution. Funct Ecol. 2002 Aug;16(4):419–32. \n34. Slavskii SA, Kuznetsov IA, Shashkova TI, Bazykin GA, Axenovich TI, Kondrashov FA, et al. \nThe limits of normal approximation for adult height. Eur J Hum Genet. 2021 \nJul;29(7):1082–91. \n35. Sinnott-Armstrong N, Naqvi S, Rivas M, Pritchard JK. GWAS of three molecular traits \nhighlights core genes and pathways alongside a highly polygenic background. Elife. 2021 \nFeb 15;10:e58615. \n36. Lutzmann M, Grey C, Traver S, Ganier O, Maya-Mendoza A, Ranisavljevic N, et al. MCM8- \nand MCM9-deficient mice reveal gametogenesis defects and genome instability due to \nimpaired homologous recombination. Mol Cell. 2012 Aug 24;47(4):523–34. \n37. Wood-Trageser MA, Gurbuz F, Yatsenko SA, Jeffries EP, Kotan LD, Surti U, et al. MCM9 \nmutations are associated with ovarian failure, short stature, and chromosomal instability. \nAm J Hum Genet. 2014 Dec 4;95(6):754–62. \n38. Colvin JS, Green RP, Schmahl J, Capel B, Ornitz DM. Male-to-female sex reversal in mice \nlacking fibroblast growth factor 9. Cell. 2001 Mar 23;104(6):875–89. \n39. Horner TW, Comstock RE, Robinson HF. Non-allelic Gene Interactions and the \nInterpretation of Quantitative Genetic Data. North Carolina Agricultural Experiment Station; \n1955. \n40. Bicknell LS, Hirschhorn JN, Savarirayan R. The genetic basis of human height. Nat Rev \nGenet. 2025 Sep;26(9):604–19. \n41. Otwinowski J, McCandlish DM, Plotkin JB. Inferring the shape of global epistasis. Proc Natl \nAcad Sci U S A. 2018 Aug 7;115(32):E7550–8. \n42. Zhu C, Ming MJ, Cole JM, Edge MD, Kirkpatrick M, Harpak A. Amplification is the primary \nmode of gene-by-sex interaction in complex human traits. Cell Genom. 2023 May \n10;3(5):100297. \n43. Mostafavi H, Harpak A, Agarwal I, Conley D, Pritchard JK, Przeworski M. Variable \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 21, 2026. ; https://doi.org/10.64898/2026.01.20.694695doi: bioRxiv preprint \n\n \nprediction accuracy of polygenic scores within an ancestry group. Elife. 2020 Jan \n30;9:e48376. \n44. Marderstein AR, Davenport ER, Kulm S, Van Hout CV, Elemento O, Clark AG. Leveraging \nphenotypic variability to identify genetic interactions in human phenotypes. Am J Hum \nGenet. 2021 Jan 7;108(1):49–67. \n45. Marderstein AR, Kulm S, Peng C, Tamimi R, Clark AG, Elemento O. A \npolygenic-score-based approach for identification of gene-drug interactions stratifying \nbreast cancer risk. Am J Hum Genet. 2021 Sep 2;108(9):1752–64. \n46. Durvasula A, Price AL. Distinct explanations underlie gene-environment interactions in the \nUK Biobank. medRxiv [Internet]. 2024 Apr 18; Available from: \nhttp://dx.doi.org/10.1101/2023.09.22.23295969 \n47. Nagpal S, Gibson G, 2024. Dual exposure-by-polygenic score interactions highlight \ndisparities across social groups in the proportion needed to benefit [Internet]. medRxiv. \n2024. Available from: http://medrxiv.org/lookup/doi/10.1101/2024.07.29.24311065 \n48. Marigorta UM, Gibson G. A simulation study of gene-by-environment interactions in GWAS \nimplies ample hidden effects. Front Genet. 2014 Jul 21;5:225. \n49. Sheppard B, Rappoport N, Loh PR, Sanders SJ, Zaitlen N, Dahl A. A model and test for \ncoordinated polygenic epistasis in complex traits. Proc Natl Acad Sci U S A. 2021 Apr \n13;118(15):e1922305118. \n50. Carlson MO, Andrews BL, Simons YB. Robust detection of specific epistasis using rank \nstatistics. bioRxiv [Internet]. 2025 Apr 10; Available from: \nhttp://dx.doi.org/10.1101/2025.04.08.647864v1 \n51. Bergen SE, Ploner A, Howrigan D, CNV Analysis Group and the Schizophrenia Working \nGroup of the Psychiatric Genomics Consortium, O’Donovan MC, Smoller JW, et al. Joint \ncontributions of rare copy number variants and common SNPs to risk for schizophrenia. Am \nJ Psychiatry. 2019 Jan 1;176(1):29–35. \n52. Park Y, Metzger BPH, Thornton JW. The simplicity of protein sequence-function \nrelationships. Nat Commun. 2024 Sep 11;15(1):7953. \n53. Mäki-Tanila A, Hill WG. Influence of gene interaction on complex trait variation with \nmultilocus models. Genetics. 2014 Sep;198(1):355–67. \n54. Kendler KS, Gardner CO. Interpretation of interactions: guide for the perplexed. Br J \nPsychiatry. 2010 Sep;197(3):170–1. \n55. Mefford J, Smullen M, Zhang F, Sadowski M, Border R, Dahl A, et al. Beyond predictive R2: \nQuantile regression and non-equivalence tests reveal complex relationships of traits and \npolygenic scores. Am J Hum Genet. 2025 Jun 5;112(6):1363–75. \n56. Wang C, Wang T, Kiryluk K, Wei Y, Aschard H, Ionita-Laza I. Genome-wide discovery for \nbiomarkers using quantile regression at biobank scale. Nat Commun. 2024 Jul \n31;15(1):6460. \n57. Brown AA, Buil A, Viñuela A, Lappalainen T, Zheng HF, Richards JB, et al. Genetic \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 21, 2026. ; https://doi.org/10.64898/2026.01.20.694695doi: bioRxiv preprint \n\n \ninteractions affecting human gene expression identified by variance association mapping. \nElife. 2014 Apr 25;3:e01381. \n58. Kelemen M, Xu Y, Jiang T, Zhao JH, Anderson CA, Wallace C, et al. Performance of \ndeep-learning-based approaches to improve polygenic scores. Nat Commun. 2025 Jun \n2;16(1):5122. \n59. Khera AV, Chaffin M, Aragam KG, Haas ME, Roselli C, Choi SH, et al. Genome-wide \npolygenic scores for common diseases identify individuals with risk equivalent to \nmonogenic mutations. Nat Genet. 2018 Sep;50(9):1219–24. \n60. Fusi N, Lippert C, Lawrence ND, Stegle O. Warped linear mixed models for the genetic \nanalysis of transformed phenotypes. Nat Commun. 2014 Sep 19;5(1):4890. \n61. Xu C, Ganesh SK, Zhou X. Statistical construction of calibrated prediction intervals for \npolygenic score-based phenotype prediction. Nat Genet. 2025 Nov;57(11):2891–900. \n62. Dahl A, Iotchkova V, Baud A, Johansson Å, Gyllensten U, Soranzo N, et al. A \nmultiple-phenotype imputation method for genetic studies. Nat Genet. 2016 \nApr;48(4):466–72. \n63. Dahl A, Thompson M, An U, Krebs M, Appadurai V, Border R, et al. Phenotype integration \nimproves power and preserves specificity in biobank-based genetic studies of major \ndepressive disorder. Nat Genet. 2023 Dec;55(12):2082–93. \n64. Keller MC. Gene × environment interaction studies have not properly controlled for potential \nconfounders: the problem and the (simple) solution. Biol Psychiatry. 2014 Jan \n1;75(1):18–24. \n65. Chang CC, Chow CC, Tellier LC, Vattikuti S, Purcell SM, Lee JJ. Second-generation PLINK: \nrising to the challenge of larger and richer datasets. Gigascience. 2015 Feb 25;4(1):7. \n66. Bulik-Sullivan B, Finucane HK, Anttila V, Gusev A, Day FR, Loh PR, et al. An atlas of \ngenetic correlations across human diseases and traits. Nat Genet. 2015 \nNov;47(11):1236–41. \n67. Bulik-Sullivan BK, Loh PR, Finucane HK, Ripke S, Yang J, Schizophrenia Working Group of \nthe Psychiatric Genomics Consortium, et al. LD Score regression distinguishes confounding \nfrom polygenicity in genome-wide association studies. Nat Genet. 2015 Mar;47(3):291–5. \n \n \n \n \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 21, 2026. ; https://doi.org/10.64898/2026.01.20.694695doi: bioRxiv preprint \n\n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 21, 2026. ; https://doi.org/10.64898/2026.01.20.694695doi: bioRxiv preprint \n\n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 21, 2026. ; https://doi.org/10.64898/2026.01.20.694695doi: bioRxiv preprint \n\n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 21, 2026. ; https://doi.org/10.64898/2026.01.20.694695doi: bioRxiv preprint \n\n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 21, 2026. ; https://doi.org/10.64898/2026.01.20.694695doi: bioRxiv preprint","source_license":"CC-BY-4.0","license_restricted":false}