A robust and powerful GWAS method for family trios supporting within-family Mendelian randomization analysis | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article A robust and powerful GWAS method for family trios supporting within-family Mendelian randomization analysis Shun Zhang, Hao-Wen Chen, Jia-Hao Mai, Qiu-Wen Zhu, Yuan-Sheng Li, and 2 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6163190/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Effect size estimates in genome-wide association studies (GWAS) and Mendelian randomization (MR) studies for independent individuals may be biased due to dynastic effect (DE) and residual population stratification (RPS). Existing GWAS methods for family trios effectively controlled such biases, while only using parental and offspring’s genotypes and offspring’s phenotype, and not incorporating parental phenotypes, which causes loss in estimation accuracy and test power. Therefore, we proposed a novel GWAS method based on structural equation modelling for family trios, denoted by FT-SEM. FT-SEM simultaneously uses parental and offspring’s genotypes and phenotypes. Simulation results demonstrate that FT-SEM substantially improves estimation accuracy and test power while controlling bias and type I error rate. Using family trios from Minnesota Center for Twin and Family Research (MCTFR), we found that DE and RPS greatly distort the results only based on independent individuals, and FT-SEM effectively corrects such biases. Combining the GWAS results from MCTFR with existing summary data, we performed several two-sample MR analyses. We observed that the effects of BMI on nicotine, alcohol consumption and behavior disorder were due to bias rather than causality. Our findings underscore the necessity of using families to validate the results of GWAS and MR, and highlight FT-SEM’s advantages. Biostatistics dynastic effect residual population stratification structure equation modelling Mendelian randomization family structure Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Introduction Mendelian randomization (MR) is an approach that uses genetic variants (single nucleotide polymorphisms, SNPs) as instrumental variables (IVs) to estimate the causal effect between exposure and outcome 1 – 4 . MR has become increasingly popular due to the widespread application in the field of causal inference with summary data from genome-wide association studies (GWAS). For MR estimates to be valid, the genetic instrument should satisfy three key assumptions: (1) relevance, it must associate with the exposure, (2) independence, ensuring that no confounding factor influences both the instrument and the outcome, and (3) exclusion, the association between the instrument and the outcome must be entirely mediated via the exposure. However, an increasing number of studies have identified biases inherent in the traditional MR approaches when applied in practice. For instance, if the factors like dynastic effect (DE), residual population stratification (RPS), and horizontal pleiotropy, are not taken into account, biased estimates of causal effects may be obtained 5 – 7 . The issues of DE and RPS have been examined in Lawson et al. 8 and Brumpton et al. 9 , 10 , which can cause bias in the MR analysis. Figure 1 illustrates these two mechanisms. In a family trio with both parents and their offspring, DE refers to the association between the parental genotypes and the offspring’s exposure/outcome which can be mediated via the unobserved parental phenotypes or other unknown mechanisms. The presence of DE (the arrows between parental genotypes and offspring’s exposure/outcome) violates the second key MR assumption of independence. Considering the parental genotypes will close these pathways and yield unbiased estimates. RPS represents that there are unobserved confounding factors which simultaneously affect genotype (IV) and outcome, thereby violating the second assumption of IVs. In addition, horizontal pleiotropy denotes a scenario where an IV not only influences the outcome indirectly through the exposure but also affects the outcome directly through other alternative pathways, which violates the exclusion assumption. Fortunately, horizontal pleiotropy could be corrected by the robust MR methods 7 , 11 – 16 . On the other hand, confounding in GWAS estimates and causal estimates, which can be induced by DE and RPS, has been addressed by using family-based study designs 5 , 9 , 17 – 20 . Family-based study designs can generally be categorized into the following three types: (1) family trio, which includes genotypes and phenotypes of both parents and offspring, (2) mother-offspring pair, which includes genotypes and phenotypes of mother and her offspring, and (3) sibling pair, which includes genotypes and phenotypes of two sibs in a family. For family trios, a straightforward approach involves using the offspring’s phenotype as the dependent variable and the offspring’s genotype and parental genotypes as independent variables, and fitting a general linear model for GWAS 9 . This approach can be regarded as fitting a structural equation model (SEM) by only using observed variables. The resulting GWAS summary data can then be used for two-sample MR analysis. However, this approach does not use the phenotypic information of parents in family trios. For mother-offspring pairs, the mother’s genotype and phenotype, and offspring’s phenotype can be treated as observed variables, while the grandmother’s genotype and offspring’s genotype are modeled as latent variables within a SEM for GWAS 18 – 20 . If the grandmother’s or offspring’s genotype is available, it can also be included as observed variables in the model. This method corrects for biases arising from the maternal effect, such as maternal DE, and only requires maternal genotype, which enhances its applicability. However, this method ignores paternal effect, and discards paternal genotype and phenotype, and offspring’s genotype. As such, this method has mainly been used in specific research areas currently, such as studying birth weight. For sibling pairs, the genotypes and the phenotypes of two sibs can be differenced and then analyzed using a general linear model applied to the differenced data for GWAS. This method also corrects for DE and RPS and is widely used due to the relatively higher availability of sibling data in practice 21 , 22 . The two-sample MR methods could be conducted to get causal effect estimate by combining the GWAS summary data from this method. However, it is not applicable to family trios with only one offspring. Therefore, it is essential to further develop GWAS methods based on family trios to enhance data utilization while controlling for DE and RPS, and then facilitate the corresponding two–sample within-family MR analysis. In addition, in this study, for convenience, we used the offspring effect to refer to the effect of offspring’s genotype on his/her own phenotype. The paternal effect denotes the effect of paternal genotype on offspring’s phenotype, while the maternal effect represents the effect of maternal genotype on offspring’s phenotype. Recall the SEM, as a diverse set of methods, which has been widely used to address problems in economics, sociology, psychology and epidemiology currently. Some researchers have applied SEM to analyze the family data for linkage analysis 23 , while others have modeled genes as latent variables, using SEM to perform multi-SNP and multi-phenotype association analyses 24 . In recent years, SEM has also been extended to accommodate the field of GWAS, such as the association analysis of five psychiatric traits, to take account of potential correlations among multiple phenotypes 25 . To address the inherent computational speed limitation of SEM when performing GWAS, an R package GW-SEM 2.0 has been developed 26 . Other researchers have suggested that incorporating the information on family structure into GWAS could give unbiased estimates of the effects between the SNP and phenotypes 5 , 9 . Using SEM to analyze mother-offspring pairs, some researchers divided the effects of SNP on birth weight into maternal and offspring effects 18 . With the widespread application of MR in recent years, concerns have been raised that MR analyses using independent individuals could derive biased causal effect estimates due to the factors such as DE and RPS 5 , 9 . To address these issues, two-sample MR methods based on family data 9 and MR methods based on SEM using sibling data (MR-DoC) have been proposed 27 . The SEM applied to mother-offspring pairs is also considered to be combined with existing two-sample MR methods, providing more robust causal effect estimates 19 . Therefore, in this study, we first constructed a SEM for family trios (denoted by FT-SEM) to conduct the family-based GWAS, which simultaneously uses parental genotypes and phenotypes as well as offspring’s genotypes and phenotypes in each family trio. We carried out extensive simulation studies to demonstrate the accuracy of point estimation and the precision of interval estimation of the genetic effect, and the corresponding hypothesis testing results obtained from GWAS using FT-SEM in the presence of DE and RPS. We also gave how to integrate the results of the GWAS using FT-SEM with existing two-sample MR methods for causal inference. We performed various simulation studies to illustrate the accuracy of point estimation and the precision of interval estimation of the causal effect, and the corresponding hypothesis testing results, even when DE and RPS exist. Utilizing 1,187 family trios from the Minnesota Center for Twin and Family Research (MCTFR) and summary data concerning 99,998 siblings derived from earlier studies 21 , we underscored the necessity of using families to validate the results of GWAS and MR, and highlighted FT-SEM’s advantages. Results Method overview The proposed FT-SEM is a GWAS method for family trios, with its technical details described in the Methods section. FT-SEM considers parental genotypes and phenotypes as well as offspring’s genotypes and phenotypes as observed variables within the SEM framework. Additionally, it treats the genotypes of grandparents as latent variables, enabling a comprehensive analysis of family trios. After performing GWAS, the results obtained using FT-SEM can be organized as summary data and combined with the existing two-sample MR method, the ratio method 28 , for conducting two-sample MR analysis. Note that the ratio method can be considered as a special case of the inverse-variance weighted (IVW) method when there is only a single instrumental variable. Simulation overview We performed simulation studies to evaluate the effectiveness of FT-SEM in GWAS and subsequent two-sample MR analysis, and compared it with existing GWAS methods based on family trios and those only based on independent individuals. In these comparisons, lm_parent represents the existing GWAS method for family trios, which fits a general linear model with the offspring’s phenotype being the dependent variable, the offspring’s genotype being the predictor and the parental genotypes being the covariates. This method is detailed in Brumpton et al. 9 and it should be noted that lm_parent does not incorporate parental phenotypes. Additionally, lm denotes the GWAS method only based on the offspring in all the family trios (i.e. independent individuals), which fits a general linear model regarding the offspring’s phenotype as the dependent variable and the offspring’s genotype as a predictor. Note that lm ignores the information of all the parents in all the family trios. In the GWAS simulation studies, we set the minor allele frequency (MAF) at a SNP to be 0.3 and considered the sample sizes of 1,000, 2,000, and 3,000 under three scenarios (DE, RPS, and there is no bias). Detailed simulation procedures and parameter settings are described in the Methods section. Furthermore, we used phenotypic variance explained (PVE) to specify the effect size. PVE indicates the proportion of variance in one variable explained by another, such as the proportion of variance in offspring’s phenotype explained by offspring’s genotype. The residual correlation between parental phenotypes and offspring’s phenotype was set to be 0.3 and 0.6 to indicate the within-family correlation. In the simulation for two-sample MR analysis, we organized the GWAS results of offspring effect derived from the three methods (FT-SEM, lm_parent, and lm) as summary data and applied the ratio method to estimate the offspring causal effect for two-sample MR analysis. Before analysis, the sample was evenly divided into two subsets: one retaining SNPs and the exposure, and the other keeping SNPs and the outcome as observational data. We fixed the MAF at 0.3 and used the sample sizes of 2,000, 4,000, and 6,000, ensuring that the two samples in the MR analysis have equal sizes of 1,000, 2,000, and 3,000, respectively. Other simulation details are consistent with those in the GWAS simulation. Simulations for GWAS Figure 2 displays the simulated biases and root mean squared errors (RMSEs) of offspring effect estimates, the coverage probabilities (CPs) and the average widths of the corresponding 95% confidence intervals (CIs), and the type I error rates for FT-SEM, lm_parent and lm under the null hypothesis with the PVE of offspring’s phenotype explained by offspring’s genotype being set to be 0, and the residual correlation being set to be 0.6 under DE, RPS and no bias. First, we examined the performances of FT-SEM, lm_parent and lm in association analyses under DE. For point estimation, both FT-SEM and lm_parent provide unbiased estimates (the biases are close to 0), while lm yields biased estimates (the biases are far away from 0). The RMSEs of point estimates decrease as the sample size increases for all the three methods, while FT-SEM achieves the lowest RMSE, and the RMSE of lm_parent based on family trios is smaller than lm only based on independent individuals. For interval estimation, the CPs of FT-SEM and lm_parent are kept around 95% irrespective of sample size, whereas the CP of lm is not maintained close to 95% and decreases when the sample size increases. The widths of the CIs for all the three methods decrease when the sample size is larger, with FT-SEM and lm having similar performance and consistently producing the narrowest intervals. For hypothesis testing, both FT-SEM and lm_parent control the type I error rates well under the null hypothesis (around 0.05). However, lm has the inflated type I error rates. On the other hand, all the results under RPS are similar to those under DE. When no bias (e.g., no DE and no RPS) is present, all the methods obtain unbiased point estimates and precise confidence intervals, and control the type I error rates well. However, FT-SEM and lm have similar performances in RMSE and the average width of CI, which are less than those of lm_parent. Figure 3 illustrates the simulated biases and RMSEs of offspring effect estimates, CPs and average widths of the corresponding 95% CIs, and test powers for FT-SEM, lm_parent and lm under the alternative hypothesis, where the PVE in the offspring’s phenotype explained by the offspring’s genotype is fixed at 0.2% (i.e. the offspring effect is 0.069). The simulation results for point and interval estimations are consistent with those under the null hypothesis. For hypothesis testing, all the methods are more powerful for larger sample sizes. Under DE and RPS, FT-SEM demonstrates much higher power than lm_parent; lm has the largest power while lm cannot maintain its type I error rate. When no bias is present, FT-SEM still exhibits a little larger power than lm, and both FT-SEM and lm are much more powerful than lm_parent. This advantage of FT-SEM in power arises from its utilization of additional parental information, which is not considered in lm. In addition, Supplementary Figs. 1 and 2 show the corresponding results when the residual correlation is 0.3, which are similar to those for 0.6, except that, in the absence of bias, FT-SEM in the RMSE, the width of the 95% CI and the power performs slightly worse than lm. Simulations for two-sample MR Figure 4 presents the simulated biases and RMSEs of causal effect estimates of offspring’s exposure on offspring’s outcome, CPs and average widths of the corresponding 95% CIs, and the type I error rates for FT-SEM, lm_parent and lm. Here, the PVE in offspring’s exposure explained by offspring’s genotype is taken to be 10% (i.e. the offspring effect is 0.488) to ensure that the SNP serves as a strong IV and the ratio estimate produces an accurate causal effect estimate. The PVE in offspring’s outcome explained by offspring’s genotype is set to be 0%, corresponding to a true causal effect value of 0 (i.e. the null hypothesis of no causal effect). From Fig. 4 , we found that almost all the results are similar to those from previous GWAS. Specifically, FT-SEM achieves the lowest RMSE among all the methods while maintaining unbiased point estimates. It provides the narrowest interval widths while ensuring correct CPs and effectively controlling the type I error rates, except that the interval width of lm is the narrowest under DE. Figure 5 gives the corresponding results of estimations and test powers under the alternative hypothesis, where the PVE in offspring’s exposure due to offspring’s genotype is still fixed at 10% and the PVE in offspring’s outcome attributed to offspring’s genotype is assumed to be 0.2% (i.e. the true offspring causal effect is 0.141). It is shown in Fig. 5 that almost all the results are similar to those from previous GWAS under the alternative hypothesis. In detail, FT-SEM has the smallest RMSE among all the methods while obtaining unbiased point estimates. The interval widths of FT-SEM is the narrowest while keeping the CPs around 95%, except that lm has the narrowest interval width under DE. Furthermore, FT-SEM is much more powerful than lm_parent while FT-SEM can effectively control the type I error rate. In addition, Supplementary Figs. 3 and 4 demonstrate the corresponding results when the residual correlation is 0.3, with the conclusions similar to those for 0.6, except that, under no bias, FT-SEM in the RMSE, the width of the 95% CI and the power performs a little worse than lm. Real data applications Firstly, we conducted a GWAS based on the MCTFR data set. This data set includes 6,784 individuals from 2,183 families, and 1,187 family trios with both parents and one offspring are selected for later analysis. We considered the following quality control (QC) criteria—(1) genotype call rate > 99%, (2) MAF > 1%, (3) individual call rate > 99%, and (4) Hardy-Weinberg equilibrium p -value > \(\:1\times\:{10}^{-7}\) . As such, the data set retains 510,275 SNPs on autosomes. Moreover, the data set features five phenotypes: nicotine composite score (NIC), alcohol consumption composite score (CON), alcohol dependence composite score (DEP), illicit drug composite score (DRG), and behavioral disinhibition composite score (BD), which are highly correlated 29 . Hence, for the GWAS of these five phenotypes, a more lenient threshold of significance level, as proposed in the previous study 30 , is adopted as follows: a SNP locus is considered statistically significant if its p -value is less than 1 × \(\:{10}^{-5}\:\) for one phenotype and its p -value is less than 1 × \(\:{10}^{-3}\:\) for at least one additional phenotype. Furthermore, all the five phenotypes are continuous variables but deviate from normality assumption. To address this issue, we applied an inverse normal transformation (INT) as recommended in the previous study 31 to convert them to approximate normality for later analysis. In our real data applications, we compared the results of three methods (FT-SEM, lm_parent and lm). FT-SEM and lm_parent use 1,187 family trios, while lm utilizes only 2,183 independent individuals by randomly selecting one member from each family. Note that both FT-SEM and lm_parent can simultaneously estimate offspring, paternal and maternal effects, and having paternal or maternal effect means that there may be DE or RPS. In contrast, lm can only estimate offspring effect, which may be influenced by unadjusted paternal and maternal effects. Supplementary Tables 1–3 summarize the SNPs having the statistically significant offspring effects on NIC, CON, DEP, DRG and BD identified by FT-SEM, lm_parent and lm, respectively. It is noteworthy that FT-SEM and lm_parent jointly identify SNP rs9349391, which shows a significant association with DEP ( p -value = \(\:4.516\times\:{10}^{-6}\) ) and a potential association with CON ( p -value = \(\:1.187\times\:{10}^{-4}\) ) based on FT-SEM. It is also significantly associated with DEP ( p -value = \(\:8.980\times\:{10}^{-8}\) ) and is potentially associated with CON ( p -value = \(\:2.923\times\:{10}^{-4}\) ) and DRG ( p -value = \(\:9.212\times\:{10}^{-4}\) ) based on lm_parent. However, no loci are identified simultaneously by all the three methods. This discrepancy may result from false positives in lm. Furthermore, Supplementary Tables 4 and 5 report the SNPs showing the statistically significant paternal effects associated with NIC, CON, DEP, DRG and BD revealed by FT-SEM and lm_parent, respectively. Among them, rs11012099, rs7965596 and rs4788008 are jointly detected by both methods. These loci are significantly associated with DEP, DEP and BD ( p -values = \(\:4.500\times\:{10}^{-6}\) , \(\:3.050\times\:{10}^{-6}\) and \(\:6.810\times\:{10}^{-7}\) , respectively) and show potential associations with NIC, CON and DEP ( p -values = \(\:5.120\times\:{10}^{-6}\) , \(\:2.000\times\:{10}^{-4}\) and \(\:4.810\times\:{10}^{-5}\) , respectively) based on FT-SEM. These loci also demonstrate significant associations with DRG, DEP and BD ( p -values = \(\:3.610\times\:{10}^{-7}\) , \(\:7.920\times\:{10}^{-6}\) and \(\:9.660\times\:{10}^{-6}\) , respectively) and are potentially associated with DEP, DRG and DEP ( p -values = \(\:8.600\times\:{10}^{-7}\) , \(\:5.160\times\:{10}^{-4}\) and \(\:7.680\times\:{10}^{-4}\) , respectively) based on lm_parent. Moreover, Supplementary Tables 6 and 7 present significant loci for maternal effects observed by FT-SEM and lm_parent, respectively. Three loci (rs7299766, rs7309476 and rs2334242) are found by both methods. These loci show significant associations with CON, CON and BD ( p -values = \(\:5.950\times\:{10}^{-7}\) , \(\:9.020\times\:{10}^{-7}\) and \(\:6.490\times\:{10}^{-7}\) , respectively) and potential associations with DEP, DEP and DRG ( p -values = \(\:4.990\times\:{10}^{-4}\) , \(\:5.560\times\:{10}^{-4}\) and \(\:8.330\times\:{10}^{-7}\) , respectively) based on FT-SEM. These loci are also significantly associated with CON, CON and DRG ( p -values = \(\:5.750\times\:{10}^{-7}\) , \(\:7.031\times\:{10}^{-7}\) and \(\:5.804\times\:{10}^{-7}\) , respectively), and demonstrate potential association with DEP, DEP and NIC ( p -values = \(\:1.437\times\:{10}^{-4}\) , \(\:1.507\times\:{10}^{-4}\) and \(\:6.434\times\:{10}^{-7}\) , respectively) using lm_parent. After the GWAS for the MCTFR data set, we performed two-sample MR analysis to investigate the causal effects of offspring’s body mass index (BMI) on offspring’s five phenotypes (NIC, CON, DEP, DRG and BD). For the first sample with BMI and genotypes, we used the GWAS summary data derived from the previous study based on sibling data 21 with the sample size of 99,998 and independent individuals 32 with the sample size of 681,275. The second sample consists of the GWAS summary data of offspring effects obtained from our previous analyses for the MCTFR data set. The IVW method serves as the primary approach for the two-sample MR analysis due to the presence of multiple instrumental variables. Specifically, IVW FT-SEM and IVW lm_parent refer to the methods where the sibling-based GWAS summary data are used as the first sample, and the GWAS summary data based on the MCTFR data set using FT-SEM and lm_parent, respectively, are used as the second sample. IVW lm is the approach where the GWAS summary data derived from independent individuals are employed as the first sample, and the GWAS summary data based on the MCTFR data set using lm are utilized as the second sample for MR analysis. Figure 6 presents the results of two-sample MR analysis. IVW lm identifies significant causal effects of offspring’s BMI on offspring’s NIC, CON and BD, with the p -values being 0.006, 0.007, and 0.014, respectively. On the other hand, the results of IVW lm show that causal effects of offspring’s BMI on offspring’s DEP and DRG are not statistically significant ( p -values > 0.05). In contrast, both IVW FT-SEM and IVW lm_parent detect no causal effects of offspring’s BMI on offspring’s NIC, CON, DEP, DRG or BD ( p -values > 0.05). The causal effects identified by IVW lm for NIC, CON and BD may result from DE or RPS, and thus may lead to false positives. In addition, IVW FT-SEM demonstrates the highest estimation precision (the narrowest interval widths), aligning with the findings from our simulation study. Finally, as mentioned above, we only used the IVW method to conduct the two-sample MR analysis. Here, to further investigate whether our two-sample MR results could be attributed to horizontal pleiotropy, we also employed the weighted median, weighted mode and MR-Egger methods to carry out the two-sample MR analysis, which are robust against various forms of horizontal pleiotropy. Supplementary Table 8 summarizes the results of testing horizontal pleiotropy using the MR-Egger method, which shows that horizontal pleiotropy is not statistically significant in all the cases ( p -values > 0.05). Supplementary Figs. 5–7 illustrate the results of the weighted median, weighted mode and MR-Egger methods, where the second sample is the GWAS summary data respectively derived from FT-SEM, lm_parent and lm for the five phenotypes in the MCTFR data set. These three figures show that offspring’s BMI does not exhibit significant causal effects on offspring’s five phenotypes, which is consistent with the conclusions drawn from the IVW method in Fig. 6 (except for the IVW lm results for NIC, CON and BD). Discussion In this paper, we introduced a powerful method FT-SEM for GWAS which is based on SEM using family trios and further illustrated how to perform two-sample MR analysis utilizing the results derived from FT-SEM. FT-SEM leverages the information on both parental genotypes and phenotypes as well as offspring’s genotypes and phenotypes. Compared to previous methods based on family trios, FT-SEM uses a greater amount of information, thereby improving estimation accuracy and statistical power. Additionally, we explained how confounding factors arising from DE and RPS could bias MR studies for independent individuals. The simulation results indicate that when DE and RPS are present, the GWAS and MR methods based on independent individuals (i.e. lm and IVW lm) yield biased estimates and cannot control the type I error rates well. However, the GWAS methods based on family trios (i.e. FT-SEM and lm_parent) as well as the corresponding two-sample MR methods (IVW FT-SEM and IVW lm_parent) respectively based on the results of FT-SEM and lm_parent could provide unbiased point estimates and accurate CPs, and maintain well-controlled type I error rates in the presence of DE and RPS. Compared to lm_parent and IVW lm_parent, FT-SEM and IVW FT-SEM have lower RMSEs, narrower CIs, and greater statistical powers. In the GWAS for the MCTFR data set, we observed that the results of lm based on independent individuals are inconsistent with those of FT-SEM and lm_parent using family trios. This suggests that previous GWAS results only based on independent individuals might have been influenced by DE and RPS. We also conducted the family-based two-sample MR and the two-sample MR based on independent individuals using additional summary data 21 , 32 and the GWAS summary data derived from FT-SEM, lm_parent and lm for the MCTFR data set to estimate the causal effects of BMI on NIC, CON, DEP, DRG and BD. We found that, after incorporating the paternal and maternal effects, there are no causal effects of offspring’s BMI on offspring’s NIC, CON and BD, which is consistent with the conclusions drawn from previous family-based MR analyses 9 , 21 , while IVW lm only based on independent individuals does. This demonstrates that the results of previous MR studies using independent individuals might have been subject to bias. On the other hand, the proposed FT-SEM method can be considered as an extension of SEM based on mother-offspring pairs 18 to accommodate family trios. When the paternal effect is not considered, the model for FT-SEM is reduced to a SEM based on mother-offspring pairs. Furthermore, it should be noted in Brumpton et al. 9 that another bias can affect GWAS analysis and two-sample MR analysis based on independent individuals—cross-trait assortative mating. Cross-trait assortative mating occurs when partners are selected based on different phenotypes. We also conducted the simulations to address this bias. The simulation results demonstrate that the model is robust to this bias (results not shown in the paper for brevity). In addition, our studies on the application of SEM in GWAS and MR further demonstrate the flexibility of SEM. In summary, our findings underscore the necessity of using families to validate the results of GWAS and MR and highlight the advantages of FT-SEM. Sociological research indicated that there was a correlation between BMI and smoking 33 , but this relationship might be due to DE and RPS 21 . These effects operate through complex biological and social mechanisms, where genetic variants may exhibit pleiotropic effects on multiple phenotypes through shared neural pathways, and environmental factors also create shared familial contexts. Moreover, studies have showed that variations at genetic loci associated with alcohol addiction may influence the susceptibility to increased BMI during childhood and adolescence 34 . Our two-sample MR analyses in the real data applications suggest that the causal effects of offspring’s BMI on offspring’s phenotypes such as NIC, CON and BD may result from biases introduced by DE and RPS. For example, parental BMI may be correlated with offspring smoking and alcohol consumption, driven by family-related confounding factors such as a household culture that promotes these behaviors. Such a culture may lead to higher parental BMI and simultaneously increase the likelihood of smoking and alcohol consumption in offspring through multiple mechanisms: direct environmental exposure, behavioral modeling, social learning and the establishment of family norms. These mechanisms can be further modified by epigenetic regulation, where parental lifestyle factors may influence offspring’s phenotypes through DNA methylation and other epigenetic modifications. Since this confounder cannot be directly adjusted in the model, it manifests as an effect of parental BMI on offspring smoking and alcohol consumption. In our two-sample MR analyses based on independent individuals, parental SNPs influence offspring smoking and alcohol consumption via the parental BMI pathway and impact offspring’s SNPs through genetic inheritance. In this scenario, parental SNPs act as confounders, distorting the results of the two-sample MR analysis based on independent individuals—a phenomenon we term "dynastic effect". The IVW FT-SEM and IVW lm_parent methods discussed in this study address this issue by including parental SNPs as covariates in the model. This inclusion blocks the confounding pathway, preventing bias from DE and producing unbiased causal effect estimates. In general, the sample size available for family-based GWAS and two-sample MR analyses is smaller than GWAS and two-sample MR analyses based on independent individuals. Therefore, in practice, the results obtained from FT-SEM and IVW FT-SEM can be considered more robust but less efficient estimates compared to those from GWAS and two-sample MR methods using independent individuals. Consequently, if there is evidence of discrepancies between the estimates from the methods based on independent individuals and those based on family data, the robust and unbiased family-based estimates are generally preferred. While our proposed FT-SEM method is unaffected by DE and RPS when conducting GWAS analyses and using the corresponding GWAS results for two-sample MR analyses, it remains susceptible to horizontal pleiotropy. This issue could be addressed by incorporating robust two-sample MR methods that correct for or model horizontal pleiotropy 7 , 11 – 15 . There are five potential limitations to our proposed FT-SEM method. Firstly, FT-SEM assumes multivariate normality among the observed variables and linearity in the relationship between genotypes and phenotypes. In instance of non-normality, employing an appropriate data transformation could aid in fulfilling the assumption of multivariate normality 31 . Secondly, we assumed additive genetic effects for all the genotypes and the genetic effects across the generations were equal. These assumptions are generally valid in biological contexts. However, if it is not met, the model's estimates might be subject to some bias. Thirdly, in contrast to linear models, FT-SEM needs more computational time. Fourthly, FT-SEM is only applicable to the scenarios where each family has a single offspring. If a family has multiple offspring, the model needs to be extended or alternative methods should be considered. Lastly, the model assumes the presence of only main effects, without considering interactions. In summary, we proposed a novel method in this paper, FT-SEM, based on the SEM using family trios to perform GWAS and further introduced how to use the results from FT-SEM to conduct two-sample MR analysis. Compared to the traditional methods for GWAS and two-sample MR using family trios, FT-SEM produces more precise estimates and more powerful tests, while ensuring robust estimation. Future work should focus on addressing how to enhance the computational speed of the SEM, using our proposed FT-SEM method to provide more robust summary data for subsequent downstream two-sample MR analyses and extending our methods to adapt to different family structures (e.g. nuclear families with multiple offspring and pedigrees). Additionally, efforts could be directed toward refining FT-SEM to address the scenarios involving nonlinear effects and interaction terms between variables in future. Methods FT-SEM method In family trios, the genetic effects at single loci can be partitioned into paternal, maternal and offspring components (respectively denoted by \(\:{\beta\:}_{FE}\) , \(\:{\beta\:}_{ME}\) and \(\:{\beta\:}_{OE}\) ), which is similar to early researches 9 , 19 . Therefore, the SEM model (FT-SEM) for family trios can be represented as: $$\:\begin{array}{c}{\varvec{Z}}_{\varvec{i}}={\varvec{\beta\:}}_{0}+{\varvec{G}}_{\varvec{i}}\beta\:+{\varvec{U}}_{\varvec{i}}+{\varvec{ϵ}}_{\varvec{i}}\:\#\left(1\right)\end{array}$$ $$\:{\varvec{Z}}_{\varvec{i}}=\left(\begin{array}{c}{Z}_{F,i}\\\:{Z}_{M,i}\\\:{Z}_{O,i}\end{array}\right),{\varvec{\beta\:}}_{0}=\left(\begin{array}{c}{\beta\:}_{0F}\\\:{\beta\:}_{0M}\\\:{\beta\:}_{0O}\end{array}\right),{\varvec{G}}_{\varvec{i}}=\left[\begin{array}{ccc}G{F}_{1,i}&\:G{M}_{1,i}&\:{F}_{i}\\\:G{F}_{2,i}&\:G{M}_{2,i}&\:{M}_{i}\\\:{F}_{i}&\:{M}_{i}&\:{O}_{i}\end{array}\right],\varvec{\beta\:}=\left(\begin{array}{c}{\beta\:}_{FE}\\\:{\beta\:}_{ME}\\\:{\beta\:}_{OE}\end{array}\right),\:{\varvec{U}}_{\varvec{i}}=\left[\begin{array}{c}{\varvec{U}}_{\varvec{F},\varvec{i}}^{\varvec{T}}{\varvec{\beta\:}}_{\varvec{U}\varvec{E},\varvec{F}}\\\:{\varvec{U}}_{\varvec{M},\varvec{i}}^{\varvec{T}}{\varvec{\beta\:}}_{\varvec{U}\varvec{E},\varvec{M}}\\\:{\varvec{U}}_{\varvec{O},\varvec{i}}^{\varvec{T}}{\varvec{\beta\:}}_{\varvec{U}\varvec{E},\varvec{O}}\end{array}\right]$$ $$\:{\varvec{ϵ}}_{\varvec{i}}=\left(\begin{array}{c}{ϵ}_{F,i}\\\:{ϵ}_{M,i}\\\:{ϵ}_{O,i}\end{array}\right)\sim\varvec{M}\varvec{V}\varvec{N}\left(0,\left[\begin{array}{ccc}{\sigma\:}_{F}^{2}&\:{\sigma\:}_{FM}&\:{\sigma\:}_{FO}\\\:{\sigma\:}_{FM}&\:{\sigma\:}_{M}^{2}&\:{\sigma\:}_{MO}\\\:{\sigma\:}_{FO}&\:{\sigma\:}_{MO}&\:{\sigma\:}_{O}^{2}\end{array}\right]\right)$$ where \(\:{Z}_{F,i}\) , \(\:{Z}_{M,i}\) and \(\:{Z}_{O,i}\:\) are respectively the phenotypes of the father, the mother and the offspring in family \(\:i\) ( \(\:i=1,\:2,\:\cdots\:,\:N\) ), where \(\:N\) is the number of family trios (i.e. sample size). \(\:{\beta\:}_{0F}\) , \(\:{\beta\:}_{0M}\) and \(\:{\beta\:}_{0O}\) represent the intercept terms for the parents and their offspring, respectively. \(\:{F}_{i}\) , \(\:{M}_{i}\) , \(\:{O}_{i}\) , \(\:G{F}_{1,i}\) , \(\:G{M}_{1,i}\) , \(\:G{F}_{2,i}\) and \(\:G{M}_{2,i}\) are the numbers of the genetic variants respectively corresponding to the father, the mother, the offspring and the grandparents of the offspring in family i , even if \(\:G{F}_{1,i}\) , \(\:G{M}_{1,i}\) , \(\:G{F}_{2,i}\) and \(\:G{M}_{2,i}\) may not be collected in practice and will be regarded as latent variables in SEM later. \(\:{\beta\:}_{FE}\) , \(\:{\beta\:}_{ME}\) and \(\:{\beta\:}_{OE}\) are paternal, maternal and offspring effects. \(\:\:{\varvec{U}}_{\varvec{F},\varvec{i}}\) , \(\:{\varvec{U}}_{\varvec{M},\varvec{i}}\) and \(\:{\varvec{U}}_{\varvec{O},\varvec{i}}\) are p -dimensional vectors representing the confounding variables for the parents and the offspring in family i , respectively. \(\:{\varvec{\beta\:}}_{\varvec{U}\varvec{E},\varvec{F}}\) , \(\:\:{\varvec{\beta\:}}_{\varvec{U}\varvec{E},\varvec{M}}\) and \(\:{\varvec{\beta\:}}_{\varvec{U}\varvec{E},\varvec{O}}\) are p -dimensional vectors representing the corresponding effects of the confounding variables. \(\:{\varvec{ϵ}}_{\varvec{i}}\) represents the residual vector for family i , which follows a multivariate normal distribution, where \(\:{ϵ}_{F,i}\) , \(\:{ϵ}_{M,i}\) , and \(\:{ϵ}_{O,i}\) are respectively the residuals of the father’s, mother’s and offspring’s phenotypes. Here, \(\:{\sigma\:}_{F}^{2}\) , \(\:{\sigma\:}_{M}^{2}\) and \(\:{\sigma\:}_{O}^{2}\) are used to respectively denote the variances of the residuals of the father’s, mother’s and offspring’s phenotypes. Moreover, \(\:{\sigma\:}_{FM}\) , \(\:{\sigma\:}_{FO}\:\) and \(\:{\sigma\:}_{MO}\) are the covariances of the residuals among the father’s, mother’s, and offspring’s phenotypes, respectively. We assumed that the genetic effects in the parental generation are respectively identical to those in the offspring generation (i.e. sharing the same \(\:{\beta\:}_{FE}\) , \(\:{\beta\:}_{ME}\) and \(\:{\beta\:}_{OE}\) ). This assumption has been mentioned in previous study 18 . This model can be fitted using the SEM. Figure 7 presents the path diagram of FT-SEM, which can be fitted using the maximum likelihood estimation with the OpenMx software package 35 . Hypothesis testing for FT-SEM is conducted using the Wald test. In this model, we simultaneously tested the paternal, maternal and offspring effects ( \(\:{\beta\:}_{FE}\) , \(\:{\beta\:}_{ME}\) and \(\:{\beta\:}_{OE}\) ). The results can serve as summary data for various downstream analyses, including two-sample MR analysis. Particular importance is to test for the offspring effect ( \(\:{\beta\:}_{OE}\) ) and to conduct its corresponding two-sample MR analysis. In addition, some robust two-sample MR methods can be employed to control horizontal pleiotropy 7 , 11 – 15 . Furthermore, note that the estimated causal effect may be biased if the associations between SNPs and both the exposure and the outcome are estimated within the same sample 36 . Fortunately, this bias can be eliminated by employing a sample-splitting approach to estimate the associations in different samples. Finally, in the FT-SEM model, we assumed that there is no additional relationship between parental phenotypes and offspring phenotypes. This assumption is based on the fact that the effects of parental phenotypes on offspring’s phenotypes can be partly explained by parental genotypes, while the remaining effects could be reflected through considering the correlations between the residuals of parental phenotypes and that of offspring’s phenotype (i.e. \(\:{\sigma\:}_{FO}\) and \(\:{\sigma\:}_{MO}\) ). In fact, this assumption has also been made in previous structural equation models based on mother-offspring pairs 18 . Simulation settings Firstly, we performed the simulations to investigate the bias and the RMSE of the point estimate of the offspring effect ( \(\:{\beta\:}_{OE}\) ), the CP and the average width of the corresponding 95% CI, and the type I error rate and the test power of FT-SEM in modeling family trios. We generated the genotypes of the grandparents (i.e. \(\:G{F}_{1}\) , \(\:G{M}_{1}\) , \(\:G{F}_{2}\) , and \(\:G{M}_{2}\) ) based on MAF at a single SNP and assuming that Hardy-Weinberg equilibrium holds, and then simulated the genotypes of the parents and the offspring (i.e. \(\:F\) , \(\:M,\:\) and \(\:O\) ) according to Mendelian inheritance. Each of the SNP effect sizes \(\:(\:{\beta\:}_{FE}\) , \(\:{\beta\:}_{ME}\) , and \(\:{\beta\:}_{OE}\) ) and the confounding effects (i.e. \(\:{\varvec{\beta\:}}_{\varvec{U}\varvec{E},\varvec{F}},{\varvec{\beta\:}}_{\varvec{U}\varvec{E},\varvec{M}}\) and \(\:{\varvec{\beta\:}}_{\varvec{U}\varvec{E},\varvec{O}}\) ) was set by using the equation \(\:\beta\:=\:\frac{PVE}{Var\left(SNP\right)}\) , where PVE represents the proportion of the variance in the phenotype explained by the SNP, and \(\:Var\left(SNP\right)\) denotes the variance of the genotype distribution at the SNP. If \(\:PV{E}_{sum}\) was used to represent the sum of the PVE values across all the variables, then the three variances of the residuals (i.e. \(\:{\sigma\:}_{F}^{2}\) , \(\:{\sigma\:}_{M}^{2}\) and \(\:{\sigma\:}_{O}^{2}\) ) were assumed to be the same for simplicity and were set by using the \(\:PV{E}_{sum}\) (i.e. \(\:{\sigma\:}_{F}^{2}\) = \(\:{\sigma\:}_{M}^{2}\) = \(\:{\sigma\:}_{O}^{2}\) = 1 – \(\:PV{E}_{sum}\) ). We used the parameter \(\:\rho\:\:\) to control the correlations among the residuals with \(\:{\sigma\:}_{FM}={\sigma\:}_{FO}={\sigma\:}_{MO}=\rho\:(1-PV{E}_{sum})\) , and \(\:\rho\:\) was set to be 0.3 and 0.6. The simulations for GWAS were divided into three scenarios: the presence of DE, the presence of RPS, and no bias. To simulate DE, we considered a homogenous population with MAF = 0.3 and assigned nonzero values to the parental effects ( \(\:{\beta\:}_{FE}\) and \(\:{\beta\:}_{ME}\) ). Specifically, when \(\:{\beta\:}_{FE}\) or \(\:{\beta\:}_{ME}\) is nonzero, DE is present. In our simulations, we assumed that \(\:PV{E}_{FE}=PV{E}_{ME}=0.5%\) , corresponding to \(\:{\beta\:}_{FE}={\beta\:}_{ME}=0.109\) . To simulate RPS, we considered a population consisting of two subpopulations with MAFs being 0.25 and 0.35, respectively, while keeping other parameters to be the same in these two subpopulations which are both consistent with those under DE. As such, we introduced a variable \(\:L\) (assigned a value of 1 if the family trio belongs to the first subpopulation and 2 if the family trio belongs to the second subpopulation, where a family trio comes from these two subpopulations with equal probabilities) with its effect on the phenotype being denoted as \(\:{\beta\:}_{LE}\) . The PVE of this variable was set to be \(\:PV{E}_{LE}=5%\) , corresponding to \(\:{\beta\:}_{LE}=0.447\) . In this scenario, we assumed no DE (i.e. \(\:{\beta\:}_{FE}={\beta\:}_{ME}=0\) ). To simulate no bias scenario, we considered a homogenous population and set \(\:{\beta\:}_{FE}={\beta\:}_{ME}=0\) . For all the three scenarios simulated, the intercept was taken as 0 and we generated the values of the confounding variable \(\:U\) for the father, the mother and the offspring. These values of \(\:U\:\) for the parents and the offspring were independent of each other but shared the same \(\:PV{E}_{U}=0.3\) , corresponding to identical effects on the phenotypes (i.e. \(\:{\beta\:}_{UE,F}={\beta\:}_{UE,M}={\beta\:}_{UE,O}=0.548\) ). Then, we evaluated all the methods under both the null and alternative hypotheses. Under the null hypothesis, we set \(\:PV{E}_{OE}=0%\) , corresponding to \(\:{\beta\:}_{OE}=0\) . Under the alternative hypothesis, we set \(\:PV{E}_{OE}=0.2%\) , corresponding to \(\:{\beta\:}_{OE}=0.069\) . After setting the parameters, the phenotypes were generated using Eq. (1). Once the phenotypes were generated, the genotypes of the grandparents and the variables \(\:L\) and \(\:U\) were removed and treated as missing values. We considered the sample sizes ( \(\:N\) ) of 1,000, 2,000 and 3,000, and compared the results of three methods (FT-SEM, lm_parent and lm). For each simulation setting, we considered 10,000 replicates. As mentioned previously, after conducting the GWAS using FT-SEM based on family trios, the GWAS results can be applied to various downstream analyses, including two-sample MR. In this context, robust two-sample MR methods can be employed to consider horizontal pleiotropy. Therefore, we further performed the simulations to investigate the bias and the RMSE of the point estimate of the causal effect of offspring’s exposure on offspring’s outcome, the CP and the average width of the corresponding 95% CI, and the type I error rate and the test power of these two-sample MR methods. For simplicity, we selected the ratio estimation approach as the two-sample MR method, which can be considered as a simplified version of the IVW method with only one instrumental variable. In the two-sample MR simulation, the exposure and outcome variables (denoted as X and Y , respectively) are simulated separately. Firstly, we set \(\:\alpha\:=\frac{PV{E}_{gy}}{PV{E}_{gx}}\) , where \(\:PV{E}_{gx}\) represents the PVE of the exposure explained by offspring effect (similar to \(\:PV{E}_{OE}\) in the GWAS) and \(\:PV{E}_{gy}\) denotes the PVE of the outcome attributed to vertical pleiotropic effect in the absence of horizontal pleiotropy. Then, the generation of the exposure can follow Eq. (1), and the simulation of the outcome is also based on Eq. (1) with an additional term \(\:\alpha\:X\) included to represent the causal effect of the exposure on the outcome. For simulating the exposure in all the scenarios, we set \(\:PV{E}_{U}=20\text{%}\) (i.e. \(\:{\beta\:}_{UE,F}={\beta\:}_{UE,M}={\beta\:}_{UE,O}=0.447\) ) and \(\:PV{E}_{gx}\) to be \(\:10\text{%}\:\) for the exposure (i.e. \(\:{\beta\:}_{OE}=0.488\) ), indicating that the SNP is a strong instrumental variable. For simulating the outcome in all the scenarios, we set \(\:PV{E}_{U}=30\text{%}\) (i.e. \(\:{\beta\:}_{UE,F}={\beta\:}_{UE,M}={\beta\:}_{UE,O}=0.548\) ). Afte that, under the null hypothesis, we set \(\:PV{E}_{gy}=0\text{%}\) , corresponding to \(\:\alpha\:=0\) , and we assumed that \(\:PV{E}_{gy}=0.2\text{%}\) , corresponding to \(\:\alpha\:=0.141\) under the alternative hypothesis. Furthermore, when simulating the outcome, the \(\:PV{E}_{OE}\) was set to 0%, indicating that there are no horizontal pleiotropy. When simulating the exposure in the presence of DE, we considered a homogenous population with MAF = 0.3 and specified \(\:PV{E}_{FE}=PV{E}_{ME}=2.5\text{%}\) , which corresponds to \(\:{\beta\:}_{FE}={\beta\:}_{ME}=0.244\) . For simulating the outcome in the presence of DE, we used the same population and set \(\:PV{E}_{FE}=PV{E}_{ME}=0.4\text{%}\) , leading to \(\:{\beta\:}_{FE}={\beta\:}_{ME}=0.098\) . For the simulations involving the exposure in the presence of RPS, we considered two subpopulations with MAFs being 0.25 and 0.35, respectively, which is similar to the simulation for GWAS. And then, we also introduced L in the model to perform the effect of RPS, and set \(\:PV{E}_{LE}=5\text{%}\) , resulting in \(\:{\beta\:}_{LE}=0.224\) to simulate the exposure. After that, when simulating the outcome in the presence of RPS, we assigned \(\:PV{E}_{LE}=10\text{%}\) , which yields \(\:{\beta\:}_{LE}=0.316\) . In this scenario, we assumed no DE (i.e. \(\:{\beta\:}_{FE}={\beta\:}_{ME}=0\) ). In the scenario of no bias, for both the exposure and outcome variables, we considered a homogenous population which is the same as DE, and set \(\:\:{\beta\:}_{FE}={\beta\:}_{ME}=0\) . After generating the exposure and outcome variables, the samples were randomly divided into two parts of equal sizes. In the first part (the first sample), the outcome variable was removed to investigate the association between the offspring’s genotype and the exposure, while in the second part (the second sample), the exposure variable was deleted to study the association between the offspring’s genotype and the outcome. For each simulation setting, we used the sample sizes of 2,000, 4,000, and 6,000, ensuring that the two samples in the MR analysis had equal sizes of 1,000, 2,000, and 3,000, respectively. Finally, we considered 10,000 replicates, and we compared the results of three methods (FT-SEM, lm_parent and lm) with the ratio estimation approach. Real data applications To illustrate the approaches and evaluate the potential biases from DE and RPS, we performed the family-based GWAS by using the MCTFR data set, and combined these GWAS results and the summary data from the sibling-based GWAS 21 to conduct the family-based two-sample MR analyses. The MCTFR data set is accessible via the Database of Genotypes and Phenotypes (dbGaP) with the accession number phs000620.v1.p1 ( https://www.ncbi.nlm.nih.gov/projects/gap/cgi-bin/study.cgi?study_id=phs000620.v1.p1 ). The data set includes the information from 2,183 families and 6,784 individuals. Supplementary Fig. 8 shows the QC process applied to this data set, resulting in a final total of 1,187 family trios used for the FT-SEM and lm_parent analyses. Moreover, lm uses the data of 2,183 independent individuals by randomly selecting one individual from each family. Additionally, the data set contains five phenotypes—NIC, CON, DEP, DRG and BD—all of which were used to carry out the GWAS. The GWAS based on the MCTFR data set was conducted using three methods (FT-SEM, lm_parent and lm). When fitting FT-SEM, the offspring's gender, date of birth and age, and the father's and mother's dates of birth and ages were included as covariates. For lm_parent, we considered the covariates: the offspring’s gender, date of birth and age. lm incorporated the gender, the date of birth, the age and the generation of the individual as the covariates, and also took account of the interaction terms between generation and gender, generation and date of birth, and generation and age. Additionally, the first 10 principal components of all the autosomal SNPs were included in lm to adjust for population stratification. To implement the two-sample MR analysis and investigate the causal effects of offspring’s BMI on offspring’s five phenotypes (NIC, CON, DEP, DRG and BD), we first need to select IVs. Supplementary Fig. 9 details the filtering process of the IVs used in IVW FT-SEM and IVW lm_parent, while Supplementary Fig. 10 outlines the filtering process of the IVs for IVW lm. SNP-exposure results were derived either from the sibling-based GWAS 21 (for IVW FT-SEM and IVW lm_parent) or from the GWAS summary data based on independent individuals 32 (for IVW lm), while SNP-outcome results were obtained from our GWAS based on the MCTFR data set. Moreover, weak instruments with the F-statistic < 10 were excluded according to standard guidelines 37 . For IVW FT-SEM and IVW lm_parent, 27 independent SNPs ( \(\:{r}^{2}\:<\:0.001\) within 10,000 kb) significantly associated with BMI ( \(\:p\:<5\times\:{10}^{-8}\) ) were selected. These SNPs were identified from the combined results of the sibling-based GWAS (“ieu-b-4816”) and the MCTFR-based GWAS using FT-SEM and lm_parent. For IVW lm, 448 independent SNPs meeting the same criteria were selected from the combined results of the independent GWAS (“ieu-b-40”) and the MCTFR-based lm. Additionally, the variants were clumped, and allele effect sizes were harmonized across the data sets using the TwoSampleMR package 38 . The weighted median, the weighted mode, and the MR-Egger are two-sample MR methods that are more robust to horizontal pleiotropy compared to IVW, but they have lower statistical power. Therefore, they can be used in sensitivity analyses to validate the results of our two-sample MR study. The summary data extraction, two-sample MR implementation, and sensitivity analyses (IVW, weighted median, weighted mode and MR-Egger) were performed using the TwoSampleMR package 38 – 40 . Declarations Author contributions S.Z. developed the methods, performed the simulation, conducted real data analysis, and wrote the manuscript. H.C., J.-H.M., Q.-W.Z., J.-Y.Z. validated the models and the results. H.-W.C, J.-H.M, Q.-W.Z., Y.-S.L., X.-B.W. and J.-Y.Z. reviewed and edited the manuscript. X.-B.W., J.-Y.Z. administrated the project and supervised the research. All authors contributed to the article and approved the submitted version. Acknowledgments This work was supported by the National Natural Science Foundation of China (grant no. 82173619), Guangdong Basic and Applied Basic Research Foundation (grant no. 2023A1515011242) and Hainan Province Science and Technology Special Fund (grant no. ZDYF2025SHFZ046). The Minnesota Center for Twin and Family Research (MCTFR) was supported by the National Institute on Drug Abuse (grant no. U01 DA024417). The sample ascertainment and data collection in MCTFR data were supported by the National Institute on Drug Abuse (grant nos. R37 DA05147 and R01 DA13240), the National Institute on Alcohol Abuse and Alcoholism (grant nos. R01 AA09367 and R01 AA11886), and the National Institute of Mental Health (grant no. R01 MH66140). Data availability The Minnesota Center for Twin and Family Research data used for this study can be found on the database of Genotypes and Phenotypes with accession number phs000620.v1.p1 ( https://www.ncbi.nlm.nih.gov/projects/gap/cgi-bin/study.cgi?study_id=phs000620.v1.p1 .) The summary data from within-sibship GWAS are available at https://gwas.mrcieu.ac.uk/datasets/ieu-b-4816/ . The summary data from independent GWAS are available at https://gwas.mrcieu.ac.uk/datasets/ieu-b-40/ . The code using in this paper are available at https://github.com/ShunZhang0816/FT-SEM/ . Competing interests The authors declare no competing interests. References Pingault, J.-B. et al. Using genetic data to strengthen causal inference in observational research. Nat Rev Genet 19 , 566–580 (2018). Davey Smith, G. & Hemani, G. Mendelian randomization: genetic anchors for causal inference in epidemiological studies. Hum Mol Genet 23 , R89–R98 (2014). Smith, G. D. & Ebrahim, S. ‘Mendelian randomization’: can genetic epidemiology contribute to understanding environmental determinants of disease? Int J Epidemiol 32 , 1–22 (2003). Burgess, S., Small, D. S. & Thompson, S. G. A review of instrumental variable estimators for Mendelian randomization. Stat Methods Med Res 26 , 2333–2355 (2017). Hwang, L.-D., Davies, N. M., Warrington, N. M. & Evans, D. M. Integrating Family-Based and Mendelian Randomization Designs. Cold Spring Harb Perspect Med 11 , a039503 (2021). Windmeijer, F., Farbmacher, H., Davies, N. & Davey Smith, G. On the Use of the Lasso for Instrumental Variables Estimation with Some Invalid Instruments. J Am Stat Assoc 114 , 1339–1350 (2019). Hartwig, F. P., Davey Smith, G. & Bowden, J. Robust inference in summary data Mendelian randomization via the zero modal pleiotropy assumption. Int J Epidemiol 46 , 1985–1998 (2017). Lawson, D. J. et al. Is population structure in the genetic biobank era irrelevant, a challenge, or an opportunity? Hum Genet 139 , 23–41 (2020). Brumpton, B. et al. Avoiding dynastic, assortative mating, and population stratification biases in Mendelian randomization through within-family analyses. Nat Commun 11 , 3519 (2020). Kong, A. et al. The nature of nurture: Effects of parental genotypes. Science 359 , 424–428 (2018). Bowden, J., Davey Smith, G. & Burgess, S. Mendelian randomization with invalid instruments: effect estimation and bias detection through Egger regression. Int J Epidemiol 44 , 512–525 (2015). Yuan, Z. et al. Likelihood-based Mendelian randomization analysis with automated instrument selection and horizontal pleiotropic modeling. Sci Adv 8 , eabl5744 (2022). Bowden, J., Davey Smith, G., Haycock, P. C. & Burgess, S. Consistent Estimation in Mendelian Randomization with Some Invalid Instruments Using a Weighted Median Estimator. Genet Epidemiol 40 , 304–314 (2016). Burgess, S., Foley, C. N., Allara, E., Staley, J. R. & Howson, J. M. M. A robust and efficient method for Mendelian randomization with hundreds of genetic variants. Nat Commun 11 , 376 (2020). Qi, G. & Chatterjee, N. Mendelian randomization analysis using mixture models for robust and efficient estimation of causal effects. Nat Commun 10 , 1941 (2019). Yuan, Z. et al. Testing and controlling for horizontal pleiotropy with probabilistic Mendelian randomization in transcriptome-wide association studies. Nat Commun 11 , 3861 (2020). Fulker, D. W., Cherny, S. S., Sham, P. C. & Hewitt, J. K. Combined linkage and association sib-pair analysis for quantitative traits. Am J Hum Genet 64 , 259–267 (1999). Warrington, N. M., Freathy, R. M., Neale, M. C. & Evans, D. M. Using structural equation modelling to jointly estimate maternal and fetal effects on birthweight in the UK Biobank. Int J Epidemiol 47 , 1229–1241 (2018). Evans, D. M., Moen, G.-H., Hwang, L.-D., Lawlor, D. A. & Warrington, N. M. Elucidating the role of maternal environmental exposures on offspring health and disease using two-sample Mendelian randomization. Int J Epidemiol 48 , 861–875 (2019). Wang, G., Warrington, N. M. & Evans, D. M. Partitioning genetic effects on birthweight at classical human leukocyte antigen loci into maternal and fetal components, using structural equation modelling. Int J Epidemiol 53 , dyad142 (2023). Howe, L. J. et al. Within-sibship genome-wide association analyses decrease bias in estimates of direct genetic effects. Nat Genet 54 , 581–592 (2022). Howe, L. J. et al. Educational attainment, health outcomes and mortality: a within-sibship Mendelian randomization study. Int J Epidemiol 52 , 1579–1591 (2023). Morris, N. J., Elston, R. C. & Stein, C. M. A framework for structural equation models in general pedigrees. Hum Hered 70 , 278–286 (2010). Stein, C. M., Morris, N. J., Hall, N. B. & Nock, N. L. Structural Equation Modeling. Methods Mol Biol 1666 , 557–580 (2017). Grotzinger, A. D. et al. Genomic structural equation modelling provides insights into the multivariate genetic architecture of complex traits. Nat Hum Behav 3 , 513–525 (2019). Pritikin, J. N., Neale, M. C., Prom-Wormley, E. C., Clark, S. L. & Verhulst, B. GW-SEM 2.0: Efficient, Flexible, and Accessible Multivariate GWAS. Behav Genet 51 , 343–357 (2021). Minică, C. C., Dolan, C. V., Boomsma, D. I., De Geus, E. & Neale, M. C. Extending Causality Tests with Genetic Instruments: An Integration of Mendelian Randomization with the Classical Twin Design. Behav Genet 48 , 337–349 (2018). Burgess, S. & Thompson, S. G. Mendelian Randomization: Methods for Causal Inference Using Genetic Variants . (CRC Press, Boca Raton London New York, 2021). Kong, Y.-F. et al. Detection of Parent-of-Origin Effects for the Variants Associated With Behavioral Disinhibition in the MCTFR Data. Front Genet 13 , 831685 (2022). McGue, M. et al. A genome-wide association study of behavioral disinhibition. Behav Genet 43 , 363–373 (2013). McCaw, Z. R., Lane, J. M., Saxena, R., Redline, S. & Lin, X. Operating characteristics of the rank-based inverse normal transformation for quantitative trait analysis in genome-wide association studies. Biometrics 76 , 1262–1272 (2020). Yengo, L. et al. Meta-analysis of genome-wide association studies for height and body mass index in ∼700000 individuals of European ancestry. Hum Mol Genet 27 , 3641–3649 (2018). Wang, P. et al. Smoking and Socio-demographic correlates of BMI. BMC Public Health 16 , 500 (2016). Lichenstein, S. D. et al. Familial risk for alcohol dependence and developmental changes in BMI: the moderating influence of addiction and obesity genes. Pharmacogenomics 15 , 1311–1321 (2014). Mc, N. et al. OpenMx 2.0: Extended Structural Equation and Statistical Modeling. Psychometrika 81 , 535–549 (2016). Hartwig, F. P. & Davies, N. M. Why internal weights should be avoided (not only) in MR-Egger regression. Int J Epidemiol 45 , 1676–1678 (2016). Burgess, S., Thompson, S. G., & CRP CHD Genetics Collaboration. Avoiding bias from weak instruments in Mendelian randomization studies. Int J Epidemiol 40 , 755–764 (2011). Hemani, G. et al. The MR-Base platform supports systematic causal inference across the human phenome. Elife 7 , e34408 (2018). Elsworth, B. et al. The MRC IEU OpenGWAS data infrastructure. 2020.08.10.244293 Preprint at https://doi.org/10.1101/2020.08.10.244293 (2020). Hemani, G., Tilling, K. & Davey Smith, G. Orienting the causal relationship between imprecisely measured traits using GWAS summary data. PLoS Genet 13 , e1007081 (2017). Additional Declarations The authors declare no competing interests. Supplementary Files SupplementaryMaterial.docx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6163190","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":424558649,"identity":"cbead5db-1511-47f0-a9cf-79b75941b9c7","order_by":0,"name":"Shun Zhang","email":"","orcid":"https://orcid.org/0009-0008-9008-9008","institution":"Southern Medical University","correspondingAuthor":false,"prefix":"","firstName":"Shun","middleName":"","lastName":"Zhang","suffix":""},{"id":424559209,"identity":"863e3d1a-cd88-4b25-8b8e-eda324530807","order_by":1,"name":"Hao-Wen Chen","email":"","orcid":"https://orcid.org/0009-0001-3410-0582","institution":"Southern Medical University","correspondingAuthor":false,"prefix":"","firstName":"Hao-Wen","middleName":"","lastName":"Chen","suffix":""},{"id":424559210,"identity":"7511640c-9bff-44b4-9f76-bed8de250be6","order_by":2,"name":"Jia-Hao Mai","email":"","orcid":"https://orcid.org/0009-0003-5129-2293","institution":"Southern Medical University","correspondingAuthor":false,"prefix":"","firstName":"Jia-Hao","middleName":"","lastName":"Mai","suffix":""},{"id":424559211,"identity":"1db62858-3a7b-42cd-9efb-6c157fa67bc8","order_by":3,"name":"Qiu-Wen Zhu","email":"","orcid":"https://orcid.org/0009-0000-3398-4833","institution":"Southern Medical University","correspondingAuthor":false,"prefix":"","firstName":"Qiu-Wen","middleName":"","lastName":"Zhu","suffix":""},{"id":424559212,"identity":"3e1f776c-aaad-40c1-a943-aed8ebb6a84f","order_by":4,"name":"Yuan-Sheng Li","email":"","orcid":"https://orcid.org/0009-0007-9106-8623","institution":"Southern Medical University","correspondingAuthor":false,"prefix":"","firstName":"Yuan-Sheng","middleName":"","lastName":"Li","suffix":""},{"id":424559213,"identity":"3e9f08b7-5405-4b1c-9a24-ae05f3e117e1","order_by":5,"name":"Xian-Bo Wu","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAy0lEQVRIiWNgGAWjYHACNoYEAwkefgiHmUgtDyos5CQbSNHC+OBMhbHBAWK1GNzIMXuQ2CaRuPn86TQJhgrrxAb2swcIaElLNwBp2XYjd5sEw5n0xAaevAS8WsxuJB+TgGjh3SbB2HY4sUGCx4CAFpB6kMP6zwK1/CNKC9CWhDMSxgYMQIcxNhChxf7Ms3SDhAoJOYkbuZstEo6lG7fx5ODXItmeY/bwh0EdD3//2Y03PtRYy/azn8GvBRUkMIBidhSMglEwCkYBxQAAWg5GOmxOT28AAAAASUVORK5CYII=","orcid":"https://orcid.org/0000-0002-2706-9599","institution":"Southern Medical University","correspondingAuthor":true,"prefix":"","firstName":"Xian-Bo","middleName":"","lastName":"Wu","suffix":""},{"id":424558648,"identity":"5c28e0f1-342c-4e0d-ad58-ad2ed5784342","order_by":6,"name":"Ji-Yuan Zhou","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAwUlEQVRIiWNgGAWjYDACCTC2gXB4SNCSRqoWBobDJGiRn918TMKy7by97owExgdv2xjkzQlpYZxzLE1Csu124rYbCcyGc9sYDHc2ENDCLJFjBtKSYHYjgU2at40hweAAAS1sEC3n7IFa2H8TpYUHouUAI9BhbMxEaZGQSEu2kDiXnLjtzMNmyTnnJAw3ENIiPyP54G2JMjt7s+PJBz+8KbORJ2gLCDBLsoEoxgYGWDQRBIwf/hCncBSMglEwCkYoAADeOzpYXG6mHAAAAABJRU5ErkJggg==","orcid":"https://orcid.org/0000-0003-0866-4402","institution":"Southern Medical University","correspondingAuthor":true,"prefix":"","firstName":"Ji-Yuan","middleName":"","lastName":"Zhou","suffix":""}],"badges":[],"createdAt":"2025-03-05 13:52:07","currentVersionCode":1,"declarations":{"humanSubjects":false,"vertebrateSubjects":false,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":false,"humanSubjectConsent":false,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":false},"doi":"10.21203/rs.3.rs-6163190/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6163190/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":77864936,"identity":"25b42a08-0d7f-4edd-8083-234fbac64085","added_by":"auto","created_at":"2025-03-06 09:23:00","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":231966,"visible":true,"origin":"","legend":"\u003cp\u003eDirected acyclic graphs showing how dynastic effect and residual population stratification can confound GWAS and MR studies. The GWAS estimate of the genetic effect of the SNP on exposure or outcome and the MR estimate of the causal effect of the exposure on the outcome may be biased due to unobserved confounders among the SNP, the exposure, and the outcome. (a) Dynastic effect can induce a statistically confounding structure in the SNP-outcome association, where F, M and O are respectively the genotypes of the parents and their offspring in a family trio, 𝑿\u003csub\u003e𝒇\u003c/sub\u003e, 𝑿\u003csub\u003e𝒎\u003c/sub\u003e and 𝑿\u003csub\u003e𝒐\u003c/sub\u003e, 𝒀\u003csub\u003e𝒇\u003c/sub\u003e, 𝒀\u003csub\u003e𝒎\u003c/sub\u003e and 𝒀\u003csub\u003e𝒐\u003c/sub\u003e, and 𝑼\u003csub\u003e𝒇\u003c/sub\u003e, 𝑼\u003csub\u003e𝒎\u003c/sub\u003e and 𝑼𝒐 are respectively the corresponding exposures, outcomes and confounders. (b) Residual population stratification can confound the association between the SNP and the outcome, which may bias the estimates of GWAS and MR, where 𝑿, 𝒀 and 𝑼 are respectively the exposure, the outcome and the confounder.\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-6163190/v1/de52e27141801072086f730b.png"},{"id":77866398,"identity":"105b07a3-6459-4c94-892c-43ec34accfab","added_by":"auto","created_at":"2025-03-06 09:31:00","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":312744,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSimulated biases and RMSEs, CPs and average widths of the corresponding 95% CIs and type I error rates of offspring effect estimates for FT-SEM, lm_parent and lm with MAF of 0.3 when the sample size is fixed at 1,000, 2,000 and 3,000, with the PVE of offspring’s phenotype explained by offspring’s genotype being set to be 0, and the residual correlation being set to be 0.6 under dynastic effect, residual population stratification and no bias. a-c\u003c/strong\u003e Biases of offspring effect estimates for FT-SEM, lm_parent and lm under dynastic effect, residual population stratification and no bias, respectively. \u003cstrong\u003ed-f\u003c/strong\u003eRMSEs of offspring effect estimates for FT-SEM, lm_parent and lm under dynastic effect, residual population stratification and no bias, respectively. \u003cstrong\u003eg-i\u003c/strong\u003eCPs of offspring effect estimates for FT-SEM, lm_parent and lm under dynastic effect, residual population stratification and no bias, respectively. \u003cstrong\u003ej-l\u003c/strong\u003eAverage widths of offspring effect estimates for FT-SEM, lm_parent and lm under dynastic effect, residual population stratification and no bias, respectively. \u0026nbsp;\u003cstrong\u003em-o\u003c/strong\u003e Type I error rates of offspring effect estimates for FT-SEM, lm_parent and lm under dynastic effect, residual population stratification and no bias, respectively.\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-6163190/v1/254c6c821890cdaa2ecabc92.png"},{"id":77866393,"identity":"9de3be4f-ba19-49c7-828c-c17127a7d645","added_by":"auto","created_at":"2025-03-06 09:31:00","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":328489,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSimulated biases and RMSEs, CPs and average widths of the corresponding 95% CIs and powers of offspring effect estimates for FT-SEM, lm_parent and lm with MAF of 0.3 when the sample size is fixed at 1,000, 2,000 and 3,000, with the PVE of offspring’s phenotype explained by offspring’s genotype being set to be 0.2%, and the residual correlation being set to be 0.6 under dynastic effect, residual population stratification and no bias\u003c/strong\u003e. \u003cstrong\u003ea-c\u003c/strong\u003e Biases of offspring effect estimates for FT-SEM, lm_parent and lm under dynastic effect, residual population stratification and no bias, respectively. \u003cstrong\u003ed-f\u003c/strong\u003e RMSEs of offspring effect estimates for FT-SEM, lm_parent and lm under dynastic effect, residual population stratification and no bias, respectively. \u003cstrong\u003eg-i\u003c/strong\u003e CPs of offspring effect estimates for FT-SEM, lm_parent and lm under dynastic effect, residual population stratification and no bias, respectively. \u003cstrong\u003ej-l\u003c/strong\u003eAverage widths of offspring effect estimates for FT-SEM, lm_parent and lm under dynastic effect, residual population stratification and no bias, respectively. \u003cstrong\u003em-o\u003c/strong\u003e Test power of offspring effect estimates for FT-SEM, lm_parent and lm under dynastic effect, residual population stratification and no bias, respectively.\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-6163190/v1/2697b8916b257b96e3667092.png"},{"id":77867395,"identity":"00d7ae44-272a-41eb-9877-f8816b58b141","added_by":"auto","created_at":"2025-03-06 09:39:00","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":322411,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSimulated biases and RMSEs, CPs and average widths of the corresponding 95% CIs and type I error rates of causal effect estimates\u003c/strong\u003e \u003cstrong\u003eof offspring’s exposure on offspring’s outcome for FT-SEM, lm_parent and lm with MAF of 0.3 and the residual correlation of 0.6 under the null hypothesis when the sample size is fixed to be 2,000, 4,000 and 6,000. a-c\u003c/strong\u003e Biases of causal effect estimates for FT-SEM, lm_parent and lm using the ratio method under dynastic effect, residual population stratification and no bias, respectively. \u003cstrong\u003ed-f\u003c/strong\u003e RMSEs of causal effect estimates for FT-SEM, lm_parent and lm using the ratio method under dynastic effect, residual population stratification and no bias, respectively. \u003cstrong\u003eg-i\u003c/strong\u003e CPs of causal effect estimates for FT-SEM, lm_parent and lm using the ratio method under dynastic effect, residual population stratification and no bias, respectively. \u003cstrong\u003ej-l\u003c/strong\u003eAverage widths of causal effect estimates for FT-SEM, lm_parent and lm using the ratio method under dynastic effect, residual population stratification and no bias, respectively. \u003cstrong\u003em-o\u003c/strong\u003e Type I error rates of causal effect estimates for FT-SEM, lm_parent and lm using the ratio method under dynastic effect, residual population stratification and no bias, respectively.\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-6163190/v1/93c4c6412ff1d8498caf80d2.png"},{"id":77864939,"identity":"0d0a36b0-51ce-4c16-bc5a-fe7569ca6d1a","added_by":"auto","created_at":"2025-03-06 09:23:00","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":342239,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSimulated biases and RMSEs, CPs and average widths of the corresponding 95% CIs and powers of causal effect estimates of offspring’s exposure on offspring’s outcome for FT-SEM, lm_parent and lm with MAF of 0.3 and the residual correlation of 0.6 under the alternative hypothesis when the sample size is fixed to be 2,000, 4,000 and 6,000. a-c\u003c/strong\u003e Biases of causal effect estimates for FT-SEM, lm_parent and lm using the ratio method under dynastic effect, residual population stratification and no bias, respectively. \u003cstrong\u003ed-f\u003c/strong\u003e RMSEs of causal effect estimates for FT-SEM, lm_parent and lm using the ratio method under dynastic effect, residual population stratification and no bias, respectively. \u003cstrong\u003eg-i\u003c/strong\u003e CPs of causal effect estimates for FT-SEM, lm_parent and lm using the ratio method under dynastic effect, residual population stratification and no bias, respectively. \u003cstrong\u003ej-l\u003c/strong\u003eAverage widths of causal effect estimates for FT-SEM, lm_parent and lm using the ratio method under dynastic effect, residual population stratification and no bias, respectively. \u003cstrong\u003em-o\u003c/strong\u003e Test power of causal effect estimates for FT-SEM, lm_parent and lm using the ratio method under dynastic effect, residual population stratification and no bias, respectively.\u003c/p\u003e","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-6163190/v1/6e47a1720d6b81ff6f46d66f.png"},{"id":77868131,"identity":"16b92695-6d75-46cb-a837-af1dcf65f5c0","added_by":"auto","created_at":"2025-03-06 09:47:00","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":239156,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003ePoint estimates, 95% CIs and \u003c/strong\u003e\u003cem\u003e\u003cstrong\u003ep\u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003e-values for the causal effects of offspring’s BMI on offspring’s NIC, CON, DEP, DRG and BD by using IVW FT-SEM, IVW lm_parent and IVW lm. \u003c/strong\u003eThe sample size for the two-sample MR is denoted as \u003cem\u003eN\u003c/em\u003e, which represents the sum of the sample sizes used in the two stages. IVW FT-SEM utilizes sibling-based GWAS summary data for BMI and GWAS summary data derived from FT-SEM for NIC, CON, DEP, DRG and BD \u0026nbsp;(\u003cem\u003eN\u003c/em\u003e = 101,185); IVW lm_parent is based on sibling-based GWAS summary data for BMI and GWAS summary data derived from lm_parent for the five outcomes (\u003cem\u003eN\u003c/em\u003e = 101,185); IVW lm uses GWAS summary data based on independent individuals for BMI and GWAS summary data derived from lm for the five outcomes (\u003cem\u003eN\u003c/em\u003e = 683,458).\u003c/p\u003e","description":"","filename":"floatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-6163190/v1/9ffd2ac0abdf38169e56a78a.png"},{"id":77864950,"identity":"7cff1f0b-7907-44df-b9c2-69ef6b12442e","added_by":"auto","created_at":"2025-03-06 09:23:00","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":265806,"visible":true,"origin":"","legend":"\u003cp\u003eFT-SEM used to estimate paternal, maternal and offspring genetic effects on phenotypes. One-headed arrows represent causal paths and two-headed arrows indicate correlation relationships. The six observed variables (in squares) denote the father’s, mother’s and offspring’s genotypes (𝐹, 𝑀 and 𝑂 ) and their phenotypes (𝑍\u003csub\u003e𝐹\u003c/sub\u003e , 𝑍\u003csub\u003e𝑀\u003c/sub\u003e and 𝑍\u003csub\u003e𝑜\u003c/sub\u003e ). The latent variables (in circles) represent the genotypes of the grandparents (𝐺𝐹1 , 𝐺𝑀\u003csub\u003e1 \u003c/sub\u003e, 𝐺𝐹\u003csub\u003e2\u003c/sub\u003e and 𝐺𝑀\u003csub\u003e2\u003c/sub\u003e ). The total variances of the latent genotypes and those of the observed genotypes are set to be Ф and are estimated from the data. The path coefficients between the offspring and their parents are both set to be 0.5 according to Mendelian inheritance. The path coefficients 𝛽\u003csub\u003e𝐹𝐸 \u003c/sub\u003e, 𝛽\u003csub\u003e𝑀𝐸\u003c/sub\u003e and 𝛽\u003csub\u003e𝑂𝐸\u003c/sub\u003e refer to paternal, maternal and offspring effects on offspring’s phenotype, respectively.\u003c/p\u003e","description":"","filename":"floatimage7.png","url":"https://assets-eu.researchsquare.com/files/rs-6163190/v1/7fd9025f37371cf9f1bafb91.png"},{"id":77869307,"identity":"abd43663-eede-4165-b09c-8db16e8cdd19","added_by":"auto","created_at":"2025-03-06 09:55:06","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":3063197,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6163190/v1/4491ee60-0646-41a4-8af6-2d8172ca982e.pdf"},{"id":77866400,"identity":"520ce2c1-f15a-46f1-93d3-6505017d19cb","added_by":"auto","created_at":"2025-03-06 09:31:00","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":2377785,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryMaterial.docx","url":"https://assets-eu.researchsquare.com/files/rs-6163190/v1/7e8406259529f791690f6712.docx"}],"financialInterests":"The authors declare no competing interests.","formattedTitle":"\u003cp\u003e\u003cstrong\u003eA robust and powerful GWAS method for family trios supporting within-family Mendelian randomization analysis\u003c/strong\u003e\u003c/p\u003e","fulltext":[{"header":"Introduction","content":"\u003cp\u003eMendelian randomization (MR) is an approach that uses genetic variants (single nucleotide polymorphisms, SNPs) as instrumental variables (IVs) to estimate the causal effect between exposure and outcome\u003csup\u003e\u003cspan additionalcitationids=\"CR2 CR3\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e. MR has become increasingly popular due to the widespread application in the field of causal inference with summary data from genome-wide association studies (GWAS). For MR estimates to be valid, the genetic instrument should satisfy three key assumptions: (1) relevance, it must associate with the exposure, (2) independence, ensuring that no confounding factor influences both the instrument and the outcome, and (3) exclusion, the association between the instrument and the outcome must be entirely mediated via the exposure. However, an increasing number of studies have identified biases inherent in the traditional MR approaches when applied in practice. For instance, if the factors like dynastic effect (DE), residual population stratification (RPS), and horizontal pleiotropy, are not taken into account, biased estimates of causal effects may be obtained\u003csup\u003e\u003cspan additionalcitationids=\"CR6\" citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e. The issues of DE and RPS have been examined in Lawson et al.\u003csup\u003e\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e and Brumpton et al.\u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e,\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u003c/sup\u003e, which can cause bias in the MR analysis. Figure\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e illustrates these two mechanisms. In a family trio with both parents and their offspring, DE refers to the association between the parental genotypes and the offspring\u0026rsquo;s exposure/outcome which can be mediated via the unobserved parental phenotypes or other unknown mechanisms. The presence of DE (the arrows between parental genotypes and offspring\u0026rsquo;s exposure/outcome) violates the second key MR assumption of independence. Considering the parental genotypes will close these pathways and yield unbiased estimates. RPS represents that there are unobserved confounding factors which simultaneously affect genotype (IV) and outcome, thereby violating the second assumption of IVs. In addition, horizontal pleiotropy denotes a scenario where an IV not only influences the outcome indirectly through the exposure but also affects the outcome directly through other alternative pathways, which violates the exclusion assumption. Fortunately, horizontal pleiotropy could be corrected by the robust MR methods\u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e,\u003cspan additionalcitationids=\"CR12 CR13 CR14 CR15\" citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eOn the other hand, confounding in GWAS estimates and causal estimates, which can be induced by DE and RPS, has been addressed by using family-based study designs\u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e,\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e,\u003cspan additionalcitationids=\"CR18 CR19\" citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e. Family-based study designs can generally be categorized into the following three types: (1) family trio, which includes genotypes and phenotypes of both parents and offspring, (2) mother-offspring pair, which includes genotypes and phenotypes of mother and her offspring, and (3) sibling pair, which includes genotypes and phenotypes of two sibs in a family. For family trios, a straightforward approach involves using the offspring\u0026rsquo;s phenotype as the dependent variable and the offspring\u0026rsquo;s genotype and parental genotypes as independent variables, and fitting a general linear model for GWAS\u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e. This approach can be regarded as fitting a structural equation model (SEM) by only using observed variables. The resulting GWAS summary data can then be used for two-sample MR analysis. However, this approach does not use the phenotypic information of parents in family trios. For mother-offspring pairs, the mother\u0026rsquo;s genotype and phenotype, and offspring\u0026rsquo;s phenotype can be treated as observed variables, while the grandmother\u0026rsquo;s genotype and offspring\u0026rsquo;s genotype are modeled as latent variables within a SEM for GWAS\u003csup\u003e\u003cspan additionalcitationids=\"CR19\" citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e. If the grandmother\u0026rsquo;s or offspring\u0026rsquo;s genotype is available, it can also be included as observed variables in the model. This method corrects for biases arising from the maternal effect, such as maternal DE, and only requires maternal genotype, which enhances its applicability. However, this method ignores paternal effect, and discards paternal genotype and phenotype, and offspring\u0026rsquo;s genotype. As such, this method has mainly been used in specific research areas currently, such as studying birth weight. For sibling pairs, the genotypes and the phenotypes of two sibs can be differenced and then analyzed using a general linear model applied to the differenced data for GWAS. This method also corrects for DE and RPS and is widely used due to the relatively higher availability of sibling data in practice\u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e,\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e. The two-sample MR methods could be conducted to get causal effect estimate by combining the GWAS summary data from this method. However, it is not applicable to family trios with only one offspring. Therefore, it is essential to further develop GWAS methods based on family trios to enhance data utilization while controlling for DE and RPS, and then facilitate the corresponding two\u0026ndash;sample within-family MR analysis. In addition, in this study, for convenience, we used the offspring effect to refer to the effect of offspring\u0026rsquo;s genotype on his/her own phenotype. The paternal effect denotes the effect of paternal genotype on offspring\u0026rsquo;s phenotype, while the maternal effect represents the effect of maternal genotype on offspring\u0026rsquo;s phenotype.\u003c/p\u003e \u003cp\u003eRecall the SEM, as a diverse set of methods, which has been widely used to address problems in economics, sociology, psychology and epidemiology currently. Some researchers have applied SEM to analyze the family data for linkage analysis\u003csup\u003e\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e\u003c/sup\u003e, while others have modeled genes as latent variables, using SEM to perform multi-SNP and multi-phenotype association analyses\u003csup\u003e\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/sup\u003e. In recent years, SEM has also been extended to accommodate the field of GWAS, such as the association analysis of five psychiatric traits, to take account of potential correlations among multiple phenotypes\u003csup\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u003c/sup\u003e. To address the inherent computational speed limitation of SEM when performing GWAS, an R package GW-SEM 2.0 has been developed\u003csup\u003e\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e\u003c/sup\u003e. Other researchers have suggested that incorporating the information on family structure into GWAS could give unbiased estimates of the effects between the SNP and phenotypes\u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e,\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e. Using SEM to analyze mother-offspring pairs, some researchers divided the effects of SNP on birth weight into maternal and offspring effects\u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e. With the widespread application of MR in recent years, concerns have been raised that MR analyses using independent individuals could derive biased causal effect estimates due to the factors such as DE and RPS\u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e,\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e. To address these issues, two-sample MR methods based on family data\u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e and MR methods based on SEM using sibling data (MR-DoC) have been proposed\u003csup\u003e\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e\u003c/sup\u003e. The SEM applied to mother-offspring pairs is also considered to be combined with existing two-sample MR methods, providing more robust causal effect estimates\u003csup\u003e\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eTherefore, in this study, we first constructed a SEM for family trios (denoted by FT-SEM) to conduct the family-based GWAS, which simultaneously uses parental genotypes and phenotypes as well as offspring\u0026rsquo;s genotypes and phenotypes in each family trio. We carried out extensive simulation studies to demonstrate the accuracy of point estimation and the precision of interval estimation of the genetic effect, and the corresponding hypothesis testing results obtained from GWAS using FT-SEM in the presence of DE and RPS. We also gave how to integrate the results of the GWAS using FT-SEM with existing two-sample MR methods for causal inference. We performed various simulation studies to illustrate the accuracy of point estimation and the precision of interval estimation of the causal effect, and the corresponding hypothesis testing results, even when DE and RPS exist. Utilizing 1,187 family trios from the Minnesota Center for Twin and Family Research (MCTFR) and summary data concerning 99,998 siblings derived from earlier studies\u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e, we underscored the necessity of using families to validate the results of GWAS and MR, and highlighted FT-SEM\u0026rsquo;s advantages.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003eMethod overview\u003c/p\u003e \u003cp\u003eThe proposed FT-SEM is a GWAS method for family trios, with its technical details described in the Methods section. FT-SEM considers parental genotypes and phenotypes as well as offspring\u0026rsquo;s genotypes and phenotypes as observed variables within the SEM framework. Additionally, it treats the genotypes of grandparents as latent variables, enabling a comprehensive analysis of family trios. After performing GWAS, the results obtained using FT-SEM can be organized as summary data and combined with the existing two-sample MR method, the ratio method\u003csup\u003e\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e\u003c/sup\u003e, for conducting two-sample MR analysis. Note that the ratio method can be considered as a special case of the inverse-variance weighted (IVW) method when there is only a single instrumental variable.\u003c/p\u003e \u003cp\u003eSimulation overview\u003c/p\u003e \u003cp\u003eWe performed simulation studies to evaluate the effectiveness of FT-SEM in GWAS and subsequent two-sample MR analysis, and compared it with existing GWAS methods based on family trios and those only based on independent individuals. In these comparisons, lm_parent represents the existing GWAS method for family trios, which fits a general linear model with the offspring\u0026rsquo;s phenotype being the dependent variable, the offspring\u0026rsquo;s genotype being the predictor and the parental genotypes being the covariates. This method is detailed in Brumpton et al.\u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e and it should be noted that lm_parent does not incorporate parental phenotypes. Additionally, lm denotes the GWAS method only based on the offspring in all the family trios (i.e. independent individuals), which fits a general linear model regarding the offspring\u0026rsquo;s phenotype as the dependent variable and the offspring\u0026rsquo;s genotype as a predictor. Note that lm ignores the information of all the parents in all the family trios.\u003c/p\u003e \u003cp\u003eIn the GWAS simulation studies, we set the minor allele frequency (MAF) at a SNP to be 0.3 and considered the sample sizes of 1,000, 2,000, and 3,000 under three scenarios (DE, RPS, and there is no bias). Detailed simulation procedures and parameter settings are described in the Methods section. Furthermore, we used phenotypic variance explained (PVE) to specify the effect size. PVE indicates the proportion of variance in one variable explained by another, such as the proportion of variance in offspring\u0026rsquo;s phenotype explained by offspring\u0026rsquo;s genotype. The residual correlation between parental phenotypes and offspring\u0026rsquo;s phenotype was set to be 0.3 and 0.6 to indicate the within-family correlation.\u003c/p\u003e \u003cp\u003eIn the simulation for two-sample MR analysis, we organized the GWAS results of offspring effect derived from the three methods (FT-SEM, lm_parent, and lm) as summary data and applied the ratio method to estimate the offspring causal effect for two-sample MR analysis. Before analysis, the sample was evenly divided into two subsets: one retaining SNPs and the exposure, and the other keeping SNPs and the outcome as observational data. We fixed the MAF at 0.3 and used the sample sizes of 2,000, 4,000, and 6,000, ensuring that the two samples in the MR analysis have equal sizes of 1,000, 2,000, and 3,000, respectively. Other simulation details are consistent with those in the GWAS simulation.\u003c/p\u003e \u003cp\u003eSimulations for GWAS\u003c/p\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e displays the simulated biases and root mean squared errors (RMSEs) of offspring effect estimates, the coverage probabilities (CPs) and the average widths of the corresponding 95% confidence intervals (CIs), and the type I error rates for FT-SEM, lm_parent and lm under the null hypothesis with the PVE of offspring\u0026rsquo;s phenotype explained by offspring\u0026rsquo;s genotype being set to be 0, and the residual correlation being set to be 0.6 under DE, RPS and no bias. First, we examined the performances of FT-SEM, lm_parent and lm in association analyses under DE. For point estimation, both FT-SEM and lm_parent provide unbiased estimates (the biases are close to 0), while lm yields biased estimates (the biases are far away from 0). The RMSEs of point estimates decrease as the sample size increases for all the three methods, while FT-SEM achieves the lowest RMSE, and the RMSE of lm_parent based on family trios is smaller than lm only based on independent individuals. For interval estimation, the CPs of FT-SEM and lm_parent are kept around 95% irrespective of sample size, whereas the CP of lm is not maintained close to 95% and decreases when the sample size increases. The widths of the CIs for all the three methods decrease when the sample size is larger, with FT-SEM and lm having similar performance and consistently producing the narrowest intervals. For hypothesis testing, both FT-SEM and lm_parent control the type I error rates well under the null hypothesis (around 0.05). However, lm has the inflated type I error rates. On the other hand, all the results under RPS are similar to those under DE. When no bias (e.g., no DE and no RPS) is present, all the methods obtain unbiased point estimates and precise confidence intervals, and control the type I error rates well. However, FT-SEM and lm have similar performances in RMSE and the average width of CI, which are less than those of lm_parent.\u003c/p\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e illustrates the simulated biases and RMSEs of offspring effect estimates, CPs and average widths of the corresponding 95% CIs, and test powers for FT-SEM, lm_parent and lm under the alternative hypothesis, where the PVE in the offspring\u0026rsquo;s phenotype explained by the offspring\u0026rsquo;s genotype is fixed at 0.2% (i.e. the offspring effect is 0.069). The simulation results for point and interval estimations are consistent with those under the null hypothesis. For hypothesis testing, all the methods are more powerful for larger sample sizes. Under DE and RPS, FT-SEM demonstrates much higher power than lm_parent; lm has the largest power while lm cannot maintain its type I error rate. When no bias is present, FT-SEM still exhibits a little larger power than lm, and both FT-SEM and lm are much more powerful than lm_parent. This advantage of FT-SEM in power arises from its utilization of additional parental information, which is not considered in lm. In addition, Supplementary Figs.\u0026nbsp;1 and 2 show the corresponding results when the residual correlation is 0.3, which are similar to those for 0.6, except that, in the absence of bias, FT-SEM in the RMSE, the width of the 95% CI and the power performs slightly worse than lm.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eSimulations for two-sample MR\u003c/p\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e presents the simulated biases and RMSEs of causal effect estimates of offspring\u0026rsquo;s exposure on offspring\u0026rsquo;s outcome, CPs and average widths of the corresponding 95% CIs, and the type I error rates for FT-SEM, lm_parent and lm. Here, the PVE in offspring\u0026rsquo;s exposure explained by offspring\u0026rsquo;s genotype is taken to be 10% (i.e. the offspring effect is 0.488) to ensure that the SNP serves as a strong IV and the ratio estimate produces an accurate causal effect estimate. The PVE in offspring\u0026rsquo;s outcome explained by offspring\u0026rsquo;s genotype is set to be 0%, corresponding to a true causal effect value of 0 (i.e. the null hypothesis of no causal effect). From Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e, we found that almost all the results are similar to those from previous GWAS. Specifically, FT-SEM achieves the lowest RMSE among all the methods while maintaining unbiased point estimates. It provides the narrowest interval widths while ensuring correct CPs and effectively controlling the type I error rates, except that the interval width of lm is the narrowest under DE. Figure\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e gives the corresponding results of estimations and test powers under the alternative hypothesis, where the PVE in offspring\u0026rsquo;s exposure due to offspring\u0026rsquo;s genotype is still fixed at 10% and the PVE in offspring\u0026rsquo;s outcome attributed to offspring\u0026rsquo;s genotype is assumed to be 0.2% (i.e. the true offspring causal effect is 0.141). It is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e that almost all the results are similar to those from previous GWAS under the alternative hypothesis. In detail, FT-SEM has the smallest RMSE among all the methods while obtaining unbiased point estimates. The interval widths of FT-SEM is the narrowest while keeping the CPs around 95%, except that lm has the narrowest interval width under DE. Furthermore, FT-SEM is much more powerful than lm_parent while FT-SEM can effectively control the type I error rate. In addition, Supplementary Figs.\u0026nbsp;3 and 4 demonstrate the corresponding results when the residual correlation is 0.3, with the conclusions similar to those for 0.6, except that, under no bias, FT-SEM in the RMSE, the width of the 95% CI and the power performs a little worse than lm.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eReal data applications\u003c/p\u003e \u003cp\u003eFirstly, we conducted a GWAS based on the MCTFR data set. This data set includes 6,784 individuals from 2,183 families, and 1,187 family trios with both parents and one offspring are selected for later analysis. We considered the following quality control (QC) criteria\u0026mdash;(1) genotype call rate\u0026thinsp;\u0026gt;\u0026thinsp;99%, (2) MAF\u0026thinsp;\u0026gt;\u0026thinsp;1%, (3) individual call rate\u0026thinsp;\u0026gt;\u0026thinsp;99%, and (4) Hardy-Weinberg equilibrium \u003cem\u003ep\u003c/em\u003e-value \u0026gt; \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:1\\times\\:{10}^{-7}\\)\u003c/span\u003e\u003c/span\u003e. As such, the data set retains 510,275 SNPs on autosomes. Moreover, the data set features five phenotypes: nicotine composite score (NIC), alcohol consumption composite score (CON), alcohol dependence composite score (DEP), illicit drug composite score (DRG), and behavioral disinhibition composite score (BD), which are highly correlated\u003csup\u003e\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e\u003c/sup\u003e. Hence, for the GWAS of these five phenotypes, a more lenient threshold of significance level, as proposed in the previous study\u003csup\u003e\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e, is adopted as follows: a SNP locus is considered statistically significant if its \u003cem\u003ep\u003c/em\u003e-value is less than 1 \u0026times; \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{10}^{-5}\\:\\)\u003c/span\u003e\u003c/span\u003efor one phenotype and its \u003cem\u003ep\u003c/em\u003e-value is less than 1 \u0026times;\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{10}^{-3}\\:\\)\u003c/span\u003e\u003c/span\u003efor at least one additional phenotype. Furthermore, all the five phenotypes are continuous variables but deviate from normality assumption. To address this issue, we applied an inverse normal transformation (INT) as recommended in the previous study\u003csup\u003e\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e\u003c/sup\u003e to convert them to approximate normality for later analysis.\u003c/p\u003e \u003cp\u003eIn our real data applications, we compared the results of three methods (FT-SEM, lm_parent and lm). FT-SEM and lm_parent use 1,187 family trios, while lm utilizes only 2,183 independent individuals by randomly selecting one member from each family. Note that both FT-SEM and lm_parent can simultaneously estimate offspring, paternal and maternal effects, and having paternal or maternal effect means that there may be DE or RPS. In contrast, lm can only estimate offspring effect, which may be influenced by unadjusted paternal and maternal effects. Supplementary Tables\u0026nbsp;1\u0026ndash;3 summarize the SNPs having the statistically significant offspring effects on NIC, CON, DEP, DRG and BD identified by FT-SEM, lm_parent and lm, respectively. It is noteworthy that FT-SEM and lm_parent jointly identify SNP rs9349391, which shows a significant association with DEP (\u003cem\u003ep\u003c/em\u003e-value = \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:4.516\\times\\:{10}^{-6}\\)\u003c/span\u003e\u003c/span\u003e ) and a potential association with CON (\u003cem\u003ep\u003c/em\u003e-value = \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:1.187\\times\\:{10}^{-4}\\)\u003c/span\u003e\u003c/span\u003e) based on FT-SEM. It is also significantly associated with DEP (\u003cem\u003ep\u003c/em\u003e-value = \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:8.980\\times\\:{10}^{-8}\\)\u003c/span\u003e\u003c/span\u003e) and is potentially associated with CON (\u003cem\u003ep\u003c/em\u003e-value = \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:2.923\\times\\:{10}^{-4}\\)\u003c/span\u003e\u003c/span\u003e) and DRG (\u003cem\u003ep\u003c/em\u003e-value = \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:9.212\\times\\:{10}^{-4}\\)\u003c/span\u003e\u003c/span\u003e) based on lm_parent. However, no loci are identified simultaneously by all the three methods. This discrepancy may result from false positives in lm. Furthermore, Supplementary Tables\u0026nbsp;4 and 5 report the SNPs showing the statistically significant paternal effects associated with NIC, CON, DEP, DRG and BD revealed by FT-SEM and lm_parent, respectively. Among them, rs11012099, rs7965596 and rs4788008 are jointly detected by both methods. These loci are significantly associated with DEP, DEP and BD (\u003cem\u003ep\u003c/em\u003e-values = \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:4.500\\times\\:{10}^{-6}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:3.050\\times\\:{10}^{-6}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:6.810\\times\\:{10}^{-7}\\)\u003c/span\u003e\u003c/span\u003e, respectively) and show potential associations with NIC, CON and DEP (\u003cem\u003ep\u003c/em\u003e-values = \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:5.120\\times\\:{10}^{-6}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:2.000\\times\\:{10}^{-4}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:4.810\\times\\:{10}^{-5}\\)\u003c/span\u003e\u003c/span\u003e, respectively) based on FT-SEM. These loci also demonstrate significant associations with DRG, DEP and BD (\u003cem\u003ep\u003c/em\u003e-values = \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:3.610\\times\\:{10}^{-7}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:7.920\\times\\:{10}^{-6}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:9.660\\times\\:{10}^{-6}\\)\u003c/span\u003e\u003c/span\u003e, respectively) and are potentially associated with DEP, DRG and DEP (\u003cem\u003ep\u003c/em\u003e-values = \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:8.600\\times\\:{10}^{-7}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:5.160\\times\\:{10}^{-4}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:7.680\\times\\:{10}^{-4}\\)\u003c/span\u003e\u003c/span\u003e, respectively) based on lm_parent. Moreover, Supplementary Tables\u0026nbsp;6 and 7 present significant loci for maternal effects observed by FT-SEM and lm_parent, respectively. Three loci (rs7299766, rs7309476 and rs2334242) are found by both methods. These loci show significant associations with CON, CON and BD (\u003cem\u003ep\u003c/em\u003e-values = \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:5.950\\times\\:{10}^{-7}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:9.020\\times\\:{10}^{-7}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:6.490\\times\\:{10}^{-7}\\)\u003c/span\u003e\u003c/span\u003e, respectively) and potential associations with DEP, DEP and DRG (\u003cem\u003ep\u003c/em\u003e-values = \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:4.990\\times\\:{10}^{-4}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:5.560\\times\\:{10}^{-4}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:8.330\\times\\:{10}^{-7}\\)\u003c/span\u003e\u003c/span\u003e, respectively) based on FT-SEM. These loci are also significantly associated with CON, CON and DRG (\u003cem\u003ep\u003c/em\u003e-values = \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:5.750\\times\\:{10}^{-7}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:7.031\\times\\:{10}^{-7}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:5.804\\times\\:{10}^{-7}\\)\u003c/span\u003e\u003c/span\u003e, respectively), and demonstrate potential association with DEP, DEP and NIC (\u003cem\u003ep\u003c/em\u003e-values = \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:1.437\\times\\:{10}^{-4}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:1.507\\times\\:{10}^{-4}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:6.434\\times\\:{10}^{-7}\\)\u003c/span\u003e\u003c/span\u003e, respectively) using lm_parent.\u003c/p\u003e \u003cp\u003eAfter the GWAS for the MCTFR data set, we performed two-sample MR analysis to investigate the causal effects of offspring\u0026rsquo;s body mass index (BMI) on offspring\u0026rsquo;s five phenotypes (NIC, CON, DEP, DRG and BD). For the first sample with BMI and genotypes, we used the GWAS summary data derived from the previous study based on sibling data\u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e with the sample size of 99,998 and independent individuals\u003csup\u003e\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e\u003c/sup\u003e with the sample size of 681,275. The second sample consists of the GWAS summary data of offspring effects obtained from our previous analyses for the MCTFR data set. The IVW method serves as the primary approach for the two-sample MR analysis due to the presence of multiple instrumental variables. Specifically, IVW FT-SEM and IVW lm_parent refer to the methods where the sibling-based GWAS summary data are used as the first sample, and the GWAS summary data based on the MCTFR data set using FT-SEM and lm_parent, respectively, are used as the second sample. IVW lm is the approach where the GWAS summary data derived from independent individuals are employed as the first sample, and the GWAS summary data based on the MCTFR data set using lm are utilized as the second sample for MR analysis.\u003c/p\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e presents the results of two-sample MR analysis. IVW lm identifies significant causal effects of offspring\u0026rsquo;s BMI on offspring\u0026rsquo;s NIC, CON and BD, with the \u003cem\u003ep\u003c/em\u003e-values being 0.006, 0.007, and 0.014, respectively. On the other hand, the results of IVW lm show that causal effects of offspring\u0026rsquo;s BMI on offspring\u0026rsquo;s DEP and DRG are not statistically significant (\u003cem\u003ep\u003c/em\u003e-values\u0026thinsp;\u0026gt;\u0026thinsp;0.05). In contrast, both IVW FT-SEM and IVW lm_parent detect no causal effects of offspring\u0026rsquo;s BMI on offspring\u0026rsquo;s NIC, CON, DEP, DRG or BD (\u003cem\u003ep\u003c/em\u003e-values\u0026thinsp;\u0026gt;\u0026thinsp;0.05). The causal effects identified by IVW lm for NIC, CON and BD may result from DE or RPS, and thus may lead to false positives. In addition, IVW FT-SEM demonstrates the highest estimation precision (the narrowest interval widths), aligning with the findings from our simulation study.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eFinally, as mentioned above, we only used the IVW method to conduct the two-sample MR analysis. Here, to further investigate whether our two-sample MR results could be attributed to horizontal pleiotropy, we also employed the weighted median, weighted mode and MR-Egger methods to carry out the two-sample MR analysis, which are robust against various forms of horizontal pleiotropy. Supplementary Table\u0026nbsp;8 summarizes the results of testing horizontal pleiotropy using the MR-Egger method, which shows that horizontal pleiotropy is not statistically significant in all the cases (\u003cem\u003ep\u003c/em\u003e-values\u0026thinsp;\u0026gt;\u0026thinsp;0.05). Supplementary Figs.\u0026nbsp;5\u0026ndash;7 illustrate the results of the weighted median, weighted mode and MR-Egger methods, where the second sample is the GWAS summary data respectively derived from FT-SEM, lm_parent and lm for the five phenotypes in the MCTFR data set. These three figures show that offspring\u0026rsquo;s BMI does not exhibit significant causal effects on offspring\u0026rsquo;s five phenotypes, which is consistent with the conclusions drawn from the IVW method in Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e (except for the IVW lm results for NIC, CON and BD).\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eIn this paper, we introduced a powerful method FT-SEM for GWAS which is based on SEM using family trios and further illustrated how to perform two-sample MR analysis utilizing the results derived from FT-SEM. FT-SEM leverages the information on both parental genotypes and phenotypes as well as offspring’s genotypes and phenotypes. Compared to previous methods based on family trios, FT-SEM uses a greater amount of information, thereby improving estimation accuracy and statistical power. Additionally, we explained how confounding factors arising from DE and RPS could bias MR studies for independent individuals. The simulation results indicate that when DE and RPS are present, the GWAS and MR methods based on independent individuals (i.e. lm and IVW lm) yield biased estimates and cannot control the type I error rates well. However, the GWAS methods based on family trios (i.e. FT-SEM and lm_parent) as well as the corresponding two-sample MR methods (IVW FT-SEM and IVW lm_parent) respectively based on the results of FT-SEM and lm_parent could provide unbiased point estimates and accurate CPs, and maintain well-controlled type I error rates in the presence of DE and RPS. Compared to lm_parent and IVW lm_parent, FT-SEM and IVW FT-SEM have lower RMSEs, narrower CIs, and greater statistical powers. In the GWAS for the MCTFR data set, we observed that the results of lm based on independent individuals are inconsistent with those of FT-SEM and lm_parent using family trios. This suggests that previous GWAS results only based on independent individuals might have been influenced by DE and RPS. We also conducted the family-based two-sample MR and the two-sample MR based on independent individuals using additional summary data\u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e,\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e\u003c/sup\u003e and the GWAS summary data derived from FT-SEM, lm_parent and lm for the MCTFR data set to estimate the causal effects of BMI on NIC, CON, DEP, DRG and BD. We found that, after incorporating the paternal and maternal effects, there are no causal effects of offspring’s BMI on offspring’s NIC, CON and BD, which is consistent with the conclusions drawn from previous family-based MR analyses\u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e,\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e, while IVW lm only based on independent individuals does. This demonstrates that the results of previous MR studies using independent individuals might have been subject to bias. On the other hand, the proposed FT-SEM method can be considered as an extension of SEM based on mother-offspring pairs\u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e to accommodate family trios. When the paternal effect is not considered, the model for FT-SEM is reduced to a SEM based on mother-offspring pairs. Furthermore, it should be noted in Brumpton et al.\u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e that another bias can affect GWAS analysis and two-sample MR analysis based on independent individuals—cross-trait assortative mating. Cross-trait assortative mating occurs when partners are selected based on different phenotypes. We also conducted the simulations to address this bias. The simulation results demonstrate that the model is robust to this bias (results not shown in the paper for brevity). In addition, our studies on the application of SEM in GWAS and MR further demonstrate the flexibility of SEM. In summary, our findings underscore the necessity of using families to validate the results of GWAS and MR and highlight the advantages of FT-SEM.\u003c/p\u003e \u003cp\u003eSociological research indicated that there was a correlation between BMI and smoking\u003csup\u003e\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e\u003c/sup\u003e, but this relationship might be due to DE and RPS\u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e. These effects operate through complex biological and social mechanisms, where genetic variants may exhibit pleiotropic effects on multiple phenotypes through shared neural pathways, and environmental factors also create shared familial contexts. Moreover, studies have showed that variations at genetic loci associated with alcohol addiction may influence the susceptibility to increased BMI during childhood and adolescence\u003csup\u003e\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e\u003c/sup\u003e. Our two-sample MR analyses in the real data applications suggest that the causal effects of offspring’s BMI on offspring’s phenotypes such as NIC, CON and BD may result from biases introduced by DE and RPS. For example, parental BMI may be correlated with offspring smoking and alcohol consumption, driven by family-related confounding factors such as a household culture that promotes these behaviors. Such a culture may lead to higher parental BMI and simultaneously increase the likelihood of smoking and alcohol consumption in offspring through multiple mechanisms: direct environmental exposure, behavioral modeling, social learning and the establishment of family norms. These mechanisms can be further modified by epigenetic regulation, where parental lifestyle factors may influence offspring’s phenotypes through DNA methylation and other epigenetic modifications. Since this confounder cannot be directly adjusted in the model, it manifests as an effect of parental BMI on offspring smoking and alcohol consumption. In our two-sample MR analyses based on independent individuals, parental SNPs influence offspring smoking and alcohol consumption via the parental BMI pathway and impact offspring’s SNPs through genetic inheritance. In this scenario, parental SNPs act as confounders, distorting the results of the two-sample MR analysis based on independent individuals—a phenomenon we term \"dynastic effect\". The IVW FT-SEM and IVW lm_parent methods discussed in this study address this issue by including parental SNPs as covariates in the model. This inclusion blocks the confounding pathway, preventing bias from DE and producing unbiased causal effect estimates.\u003c/p\u003e \u003cp\u003eIn general, the sample size available for family-based GWAS and two-sample MR analyses is smaller than GWAS and two-sample MR analyses based on independent individuals. Therefore, in practice, the results obtained from FT-SEM and IVW FT-SEM can be considered more robust but less efficient estimates compared to those from GWAS and two-sample MR methods using independent individuals. Consequently, if there is evidence of discrepancies between the estimates from the methods based on independent individuals and those based on family data, the robust and unbiased family-based estimates are generally preferred. While our proposed FT-SEM method is unaffected by DE and RPS when conducting GWAS analyses and using the corresponding GWAS results for two-sample MR analyses, it remains susceptible to horizontal pleiotropy. This issue could be addressed by incorporating robust two-sample MR methods that correct for or model horizontal pleiotropy\u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e,\u003cspan additionalcitationids=\"CR12 CR13 CR14\" citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e–\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eThere are five potential limitations to our proposed FT-SEM method. Firstly, FT-SEM assumes multivariate normality among the observed variables and linearity in the relationship between genotypes and phenotypes. In instance of non-normality, employing an appropriate data transformation could aid in fulfilling the assumption of multivariate normality\u003csup\u003e\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e\u003c/sup\u003e. Secondly, we assumed additive genetic effects for all the genotypes and the genetic effects across the generations were equal. These assumptions are generally valid in biological contexts. However, if it is not met, the model's estimates might be subject to some bias. Thirdly, in contrast to linear models, FT-SEM needs more computational time. Fourthly, FT-SEM is only applicable to the scenarios where each family has a single offspring. If a family has multiple offspring, the model needs to be extended or alternative methods should be considered. Lastly, the model assumes the presence of only main effects, without considering interactions.\u003c/p\u003e \u003cp\u003eIn summary, we proposed a novel method in this paper, FT-SEM, based on the SEM using family trios to perform GWAS and further introduced how to use the results from FT-SEM to conduct two-sample MR analysis. Compared to the traditional methods for GWAS and two-sample MR using family trios, FT-SEM produces more precise estimates and more powerful tests, while ensuring robust estimation. Future work should focus on addressing how to enhance the computational speed of the SEM, using our proposed FT-SEM method to provide more robust summary data for subsequent downstream two-sample MR analyses and extending our methods to adapt to different family structures (e.g. nuclear families with multiple offspring and pedigrees). Additionally, efforts could be directed toward refining FT-SEM to address the scenarios involving nonlinear effects and interaction terms between variables in future.\u003c/p\u003e \u003cp\u003e\u003c/p\u003e \u003cp\u003e \u003c/p\u003e "},{"header":"Methods","content":"\u003cp\u003eFT-SEM method\u003c/p\u003e\u003cp\u003eIn family trios, the genetic effects at single loci can be partitioned into paternal, maternal and offspring components (respectively denoted by \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{FE}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{ME}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{OE}\\)\u003c/span\u003e\u003c/span\u003e), which is similar to early researches\u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e,\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e. Therefore, the SEM model (FT-SEM) for family trios can be represented as:\u003c/p\u003e\u003cdiv id=\"Equa\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equa\" name=\"EquationSource\"\u003e\n$$\\:\\begin{array}{c}{\\varvec{Z}}_{\\varvec{i}}={\\varvec{\\beta\\:}}_{0}+{\\varvec{G}}_{\\varvec{i}}\\beta\\:+{\\varvec{U}}_{\\varvec{i}}+{\\varvec{ϵ}}_{\\varvec{i}}\\:\\#\\left(1\\right)\\end{array}$$\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equb\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equb\" name=\"EquationSource\"\u003e\n$$\\:{\\varvec{Z}}_{\\varvec{i}}=\\left(\\begin{array}{c}{Z}_{F,i}\\\\\\:{Z}_{M,i}\\\\\\:{Z}_{O,i}\\end{array}\\right),{\\varvec{\\beta\\:}}_{0}=\\left(\\begin{array}{c}{\\beta\\:}_{0F}\\\\\\:{\\beta\\:}_{0M}\\\\\\:{\\beta\\:}_{0O}\\end{array}\\right),{\\varvec{G}}_{\\varvec{i}}=\\left[\\begin{array}{ccc}G{F}_{1,i}\u0026amp;\\:G{M}_{1,i}\u0026amp;\\:{F}_{i}\\\\\\:G{F}_{2,i}\u0026amp;\\:G{M}_{2,i}\u0026amp;\\:{M}_{i}\\\\\\:{F}_{i}\u0026amp;\\:{M}_{i}\u0026amp;\\:{O}_{i}\\end{array}\\right],\\varvec{\\beta\\:}=\\left(\\begin{array}{c}{\\beta\\:}_{FE}\\\\\\:{\\beta\\:}_{ME}\\\\\\:{\\beta\\:}_{OE}\\end{array}\\right),\\:{\\varvec{U}}_{\\varvec{i}}=\\left[\\begin{array}{c}{\\varvec{U}}_{\\varvec{F},\\varvec{i}}^{\\varvec{T}}{\\varvec{\\beta\\:}}_{\\varvec{U}\\varvec{E},\\varvec{F}}\\\\\\:{\\varvec{U}}_{\\varvec{M},\\varvec{i}}^{\\varvec{T}}{\\varvec{\\beta\\:}}_{\\varvec{U}\\varvec{E},\\varvec{M}}\\\\\\:{\\varvec{U}}_{\\varvec{O},\\varvec{i}}^{\\varvec{T}}{\\varvec{\\beta\\:}}_{\\varvec{U}\\varvec{E},\\varvec{O}}\\end{array}\\right]$$\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equc\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equc\" name=\"EquationSource\"\u003e\n$$\\:{\\varvec{ϵ}}_{\\varvec{i}}=\\left(\\begin{array}{c}{ϵ}_{F,i}\\\\\\:{ϵ}_{M,i}\\\\\\:{ϵ}_{O,i}\\end{array}\\right)\\sim\\varvec{M}\\varvec{V}\\varvec{N}\\left(0,\\left[\\begin{array}{ccc}{\\sigma\\:}_{F}^{2}\u0026amp;\\:{\\sigma\\:}_{FM}\u0026amp;\\:{\\sigma\\:}_{FO}\\\\\\:{\\sigma\\:}_{FM}\u0026amp;\\:{\\sigma\\:}_{M}^{2}\u0026amp;\\:{\\sigma\\:}_{MO}\\\\\\:{\\sigma\\:}_{FO}\u0026amp;\\:{\\sigma\\:}_{MO}\u0026amp;\\:{\\sigma\\:}_{O}^{2}\\end{array}\\right]\\right)$$\u003c/div\u003e\u003c/div\u003e\u003cp\u003ewhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{Z}_{F,i}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{Z}_{M,i}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{Z}_{O,i}\\:\\)\u003c/span\u003e\u003c/span\u003eare respectively the phenotypes of the father, the mother and the offspring in family \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:i\\)\u003c/span\u003e\u003c/span\u003e (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:i=1,\\:2,\\:\\cdots\\:,\\:N\\)\u003c/span\u003e\u003c/span\u003e), where \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:N\\)\u003c/span\u003e\u003c/span\u003e is the number of family trios (i.e. sample size). \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{0F}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{0M}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{0O}\\)\u003c/span\u003e\u003c/span\u003e represent the intercept terms for the parents and their offspring, respectively. \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{F}_{i}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{M}_{i}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{O}_{i}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:G{F}_{1,i}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:G{M}_{1,i}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:G{F}_{2,i}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:G{M}_{2,i}\\)\u003c/span\u003e\u003c/span\u003e are the numbers of the genetic variants respectively corresponding to the father, the mother, the offspring and the grandparents of the offspring in family \u003cem\u003ei\u003c/em\u003e, even if \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:G{F}_{1,i}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:G{M}_{1,i}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:G{F}_{2,i}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:G{M}_{2,i}\\)\u003c/span\u003e\u003c/span\u003e may not be collected in practice and will be regarded as latent variables in SEM later. \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{FE}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{ME}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{OE}\\)\u003c/span\u003e\u003c/span\u003e are paternal, maternal and offspring effects.\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\:{\\varvec{U}}_{\\varvec{F},\\varvec{i}}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\varvec{U}}_{\\varvec{M},\\varvec{i}}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\varvec{U}}_{\\varvec{O},\\varvec{i}}\\)\u003c/span\u003e\u003c/span\u003e are \u003cem\u003ep\u003c/em\u003e-dimensional vectors representing the confounding variables for the parents and the offspring in family \u003cem\u003ei\u003c/em\u003e, respectively. \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\varvec{\\beta\\:}}_{\\varvec{U}\\varvec{E},\\varvec{F}}\\)\u003c/span\u003e\u003c/span\u003e,\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\:{\\varvec{\\beta\\:}}_{\\varvec{U}\\varvec{E},\\varvec{M}}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\varvec{\\beta\\:}}_{\\varvec{U}\\varvec{E},\\varvec{O}}\\)\u003c/span\u003e\u003c/span\u003e are \u003cem\u003ep\u003c/em\u003e-dimensional vectors representing the corresponding effects of the confounding variables. \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\varvec{ϵ}}_{\\varvec{i}}\\)\u003c/span\u003e\u003c/span\u003e represents the residual vector for family \u003cem\u003ei\u003c/em\u003e, which follows a multivariate normal distribution, where \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{ϵ}_{F,i}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{ϵ}_{M,i}\\)\u003c/span\u003e\u003c/span\u003e, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{ϵ}_{O,i}\\)\u003c/span\u003e\u003c/span\u003e are respectively the residuals of the father’s, mother’s and offspring’s phenotypes. Here, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\sigma\\:}_{F}^{2}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\sigma\\:}_{M}^{2}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\sigma\\:}_{O}^{2}\\)\u003c/span\u003e\u003c/span\u003e are used to respectively denote the variances of the residuals of the father’s, mother’s and offspring’s phenotypes. Moreover, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\sigma\\:}_{FM}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\sigma\\:}_{FO}\\:\\)\u003c/span\u003e\u003c/span\u003eand \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\sigma\\:}_{MO}\\)\u003c/span\u003e\u003c/span\u003e are the covariances of the residuals among the father’s, mother’s, and offspring’s phenotypes, respectively. We assumed that the genetic effects in the parental generation are respectively identical to those in the offspring generation (i.e. sharing the same \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{FE}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{ME}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{OE}\\)\u003c/span\u003e\u003c/span\u003e). This assumption has been mentioned in previous study\u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003eThis model can be fitted using the SEM. Figure\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003e presents the path diagram of FT-SEM, which can be fitted using the maximum likelihood estimation with the OpenMx software package\u003csup\u003e\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e\u003c/sup\u003e. Hypothesis testing for FT-SEM is conducted using the Wald test. In this model, we simultaneously tested the paternal, maternal and offspring effects (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{FE}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{ME}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{OE}\\)\u003c/span\u003e\u003c/span\u003e). The results can serve as summary data for various downstream analyses, including two-sample MR analysis. Particular importance is to test for the offspring effect (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{OE}\\)\u003c/span\u003e\u003c/span\u003e) and to conduct its corresponding two-sample MR analysis. In addition, some robust two-sample MR methods can be employed to control horizontal pleiotropy\u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e,\u003cspan additionalcitationids=\"CR12 CR13 CR14\" citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e–\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e. Furthermore, note that the estimated causal effect may be biased if the associations between SNPs and both the exposure and the outcome are estimated within the same sample\u003csup\u003e\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e\u003c/sup\u003e. Fortunately, this bias can be eliminated by employing a sample-splitting approach to estimate the associations in different samples. Finally, in the FT-SEM model, we assumed that there is no additional relationship between parental phenotypes and offspring phenotypes. This assumption is based on the fact that the effects of parental phenotypes on offspring’s phenotypes can be partly explained by parental genotypes, while the remaining effects could be reflected through considering the correlations between the residuals of parental phenotypes and that of offspring’s phenotype (i.e. \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\sigma\\:}_{FO}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\sigma\\:}_{MO}\\)\u003c/span\u003e\u003c/span\u003e). In fact, this assumption has also been made in previous structural equation models based on mother-offspring pairs\u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003eSimulation settings\u003c/p\u003e\u003cp\u003eFirstly, we performed the simulations to investigate the bias and the RMSE of the point estimate of the offspring effect (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{OE}\\)\u003c/span\u003e\u003c/span\u003e), the CP and the average width of the corresponding 95% CI, and the type I error rate and the test power of FT-SEM in modeling family trios. We generated the genotypes of the grandparents (i.e. \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:G{F}_{1}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:G{M}_{1}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:G{F}_{2}\\)\u003c/span\u003e\u003c/span\u003e, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:G{M}_{2}\\)\u003c/span\u003e\u003c/span\u003e) based on MAF at a single SNP and assuming that Hardy-Weinberg equilibrium holds, and then simulated the genotypes of the parents and the offspring (i.e. \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:F\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:M,\\:\\)\u003c/span\u003e\u003c/span\u003eand \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:O\\)\u003c/span\u003e\u003c/span\u003e) according to Mendelian inheritance. Each of the SNP effect sizes \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:(\\:{\\beta\\:}_{FE}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{ME}\\)\u003c/span\u003e\u003c/span\u003e, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{OE}\\)\u003c/span\u003e\u003c/span\u003e) and the confounding effects (i.e. \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\varvec{\\beta\\:}}_{\\varvec{U}\\varvec{E},\\varvec{F}},{\\varvec{\\beta\\:}}_{\\varvec{U}\\varvec{E},\\varvec{M}}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\varvec{\\beta\\:}}_{\\varvec{U}\\varvec{E},\\varvec{O}}\\)\u003c/span\u003e\u003c/span\u003e) was set by using the equation \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\beta\\:=\\:\\frac{PVE}{Var\\left(SNP\\right)}\\)\u003c/span\u003e\u003c/span\u003e, where \u003cem\u003ePVE\u003c/em\u003e represents the proportion of the variance in the phenotype explained by the SNP, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:Var\\left(SNP\\right)\\)\u003c/span\u003e\u003c/span\u003e denotes the variance of the genotype distribution at the SNP. If \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:PV{E}_{sum}\\)\u003c/span\u003e\u003c/span\u003e was used to represent the sum of the \u003cem\u003ePVE\u003c/em\u003e values across all the variables, then the three variances of the residuals (i.e. \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\sigma\\:}_{F}^{2}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\sigma\\:}_{M}^{2}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\sigma\\:}_{O}^{2}\\)\u003c/span\u003e\u003c/span\u003e) were assumed to be the same for simplicity and were set by using the \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:PV{E}_{sum}\\)\u003c/span\u003e\u003c/span\u003e (i.e. \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\sigma\\:}_{F}^{2}\\)\u003c/span\u003e\u003c/span\u003e = \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\sigma\\:}_{M}^{2}\\)\u003c/span\u003e\u003c/span\u003e = \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\sigma\\:}_{O}^{2}\\)\u003c/span\u003e\u003c/span\u003e = 1 – \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:PV{E}_{sum}\\)\u003c/span\u003e\u003c/span\u003e). We used the parameter \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\rho\\:\\:\\)\u003c/span\u003e\u003c/span\u003eto control the correlations among the residuals with \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\sigma\\:}_{FM}={\\sigma\\:}_{FO}={\\sigma\\:}_{MO}=\\rho\\:(1-PV{E}_{sum})\\)\u003c/span\u003e\u003c/span\u003e, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\rho\\:\\)\u003c/span\u003e\u003c/span\u003e was set to be 0.3 and 0.6. The simulations for GWAS were divided into three scenarios: the presence of DE, the presence of RPS, and no bias. To simulate DE, we considered a homogenous population with MAF = 0.3 and assigned nonzero values to the parental effects (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{FE}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{ME}\\)\u003c/span\u003e\u003c/span\u003e). Specifically, when \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{FE}\\)\u003c/span\u003e\u003c/span\u003e or \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{ME}\\)\u003c/span\u003e\u003c/span\u003e is nonzero, DE is present. In our simulations, we assumed that \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:PV{E}_{FE}=PV{E}_{ME}=0.5%\\)\u003c/span\u003e\u003c/span\u003e, corresponding to \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{FE}={\\beta\\:}_{ME}=0.109\\)\u003c/span\u003e\u003c/span\u003e. To simulate RPS, we considered a population consisting of two subpopulations with MAFs being 0.25 and 0.35, respectively, while keeping other parameters to be the same in these two subpopulations which are both consistent with those under DE. As such, we introduced a variable \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:L\\)\u003c/span\u003e\u003c/span\u003e (assigned a value of 1 if the family trio belongs to the first subpopulation and 2 if the family trio belongs to the second subpopulation, where a family trio comes from these two subpopulations with equal probabilities) with its effect on the phenotype being denoted as \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{LE}\\)\u003c/span\u003e\u003c/span\u003e. The \u003cem\u003ePVE\u003c/em\u003e of this variable was set to be \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:PV{E}_{LE}=5%\\)\u003c/span\u003e\u003c/span\u003e, corresponding to \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{LE}=0.447\\)\u003c/span\u003e\u003c/span\u003e. In this scenario, we assumed no DE (i.e. \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{FE}={\\beta\\:}_{ME}=0\\)\u003c/span\u003e\u003c/span\u003e). To simulate no bias scenario, we considered a homogenous population and set \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{FE}={\\beta\\:}_{ME}=0\\)\u003c/span\u003e\u003c/span\u003e. For all the three scenarios simulated, the intercept was taken as 0 and we generated the values of the confounding variable \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:U\\)\u003c/span\u003e\u003c/span\u003e for the father, the mother and the offspring. These values of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:U\\:\\)\u003c/span\u003e\u003c/span\u003efor the parents and the offspring were independent of each other but shared the same \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:PV{E}_{U}=0.3\\)\u003c/span\u003e\u003c/span\u003e, corresponding to identical effects on the phenotypes (i.e. \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{UE,F}={\\beta\\:}_{UE,M}={\\beta\\:}_{UE,O}=0.548\\)\u003c/span\u003e\u003c/span\u003e). Then, we evaluated all the methods under both the null and alternative hypotheses. Under the null hypothesis, we set \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:PV{E}_{OE}=0%\\)\u003c/span\u003e\u003c/span\u003e, corresponding to \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{OE}=0\\)\u003c/span\u003e\u003c/span\u003e. Under the alternative hypothesis, we set \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:PV{E}_{OE}=0.2%\\)\u003c/span\u003e\u003c/span\u003e, corresponding to \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{OE}=0.069\\)\u003c/span\u003e\u003c/span\u003e. After setting the parameters, the phenotypes were generated using Eq.\u0026nbsp;(1). Once the phenotypes were generated, the genotypes of the grandparents and the variables \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:L\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:U\\)\u003c/span\u003e\u003c/span\u003e were removed and treated as missing values. We considered the sample sizes (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:N\\)\u003c/span\u003e\u003c/span\u003e) of 1,000, 2,000 and 3,000, and compared the results of three methods (FT-SEM, lm_parent and lm). For each simulation setting, we considered 10,000 replicates.\u003c/p\u003e\u003cp\u003eAs mentioned previously, after conducting the GWAS using FT-SEM based on family trios, the GWAS results can be applied to various downstream analyses, including two-sample MR. In this context, robust two-sample MR methods can be employed to consider horizontal pleiotropy. Therefore, we further performed the simulations to investigate the bias and the RMSE of the point estimate of the causal effect of offspring’s exposure on offspring’s outcome, the CP and the average width of the corresponding 95% CI, and the type I error rate and the test power of these two-sample MR methods. For simplicity, we selected the ratio estimation approach as the two-sample MR method, which can be considered as a simplified version of the IVW method with only one instrumental variable. In the two-sample MR simulation, the exposure and outcome variables (denoted as \u003cem\u003eX\u003c/em\u003e and \u003cem\u003eY\u003c/em\u003e, respectively) are simulated separately. Firstly, we set \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\alpha\\:=\\frac{PV{E}_{gy}}{PV{E}_{gx}}\\)\u003c/span\u003e\u003c/span\u003e, where \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:PV{E}_{gx}\\)\u003c/span\u003e\u003c/span\u003e represents the \u003cem\u003ePVE\u003c/em\u003e of the exposure explained by offspring effect (similar to \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:PV{E}_{OE}\\)\u003c/span\u003e\u003c/span\u003e in the GWAS) and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:PV{E}_{gy}\\)\u003c/span\u003e\u003c/span\u003e denotes the \u003cem\u003ePVE\u003c/em\u003e of the outcome attributed to vertical pleiotropic effect in the absence of horizontal pleiotropy. Then, the generation of the exposure can follow Eq.\u0026nbsp;(1), and the simulation of the outcome is also based on Eq.\u0026nbsp;(1) with an additional term \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\alpha\\:X\\)\u003c/span\u003e\u003c/span\u003e included to represent the causal effect of the exposure on the outcome. For simulating the exposure in all the scenarios, we set \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:PV{E}_{U}=20\\text{%}\\)\u003c/span\u003e\u003c/span\u003e (i.e. \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{UE,F}={\\beta\\:}_{UE,M}={\\beta\\:}_{UE,O}=0.447\\)\u003c/span\u003e\u003c/span\u003e) and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:PV{E}_{gx}\\)\u003c/span\u003e\u003c/span\u003e to be \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:10\\text{%}\\:\\)\u003c/span\u003e\u003c/span\u003e for the exposure (i.e. \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{OE}=0.488\\)\u003c/span\u003e\u003c/span\u003e), indicating that the SNP is a strong instrumental variable. For simulating the outcome in all the scenarios, we set \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:PV{E}_{U}=30\\text{%}\\)\u003c/span\u003e\u003c/span\u003e (i.e. \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{UE,F}={\\beta\\:}_{UE,M}={\\beta\\:}_{UE,O}=0.548\\)\u003c/span\u003e\u003c/span\u003e). Afte that, under the null hypothesis, we set \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:PV{E}_{gy}=0\\text{%}\\)\u003c/span\u003e\u003c/span\u003e, corresponding to \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\alpha\\:=0\\)\u003c/span\u003e\u003c/span\u003e, and we assumed that \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:PV{E}_{gy}=0.2\\text{%}\\)\u003c/span\u003e\u003c/span\u003e, corresponding to \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\alpha\\:=0.141\\)\u003c/span\u003e\u003c/span\u003e under the alternative hypothesis. Furthermore, when simulating the outcome, the \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:PV{E}_{OE}\\)\u003c/span\u003e\u003c/span\u003e was set to 0%, indicating that there are no horizontal pleiotropy. When simulating the exposure in the presence of DE, we considered a homogenous population with MAF = 0.3 and specified \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:PV{E}_{FE}=PV{E}_{ME}=2.5\\text{%}\\)\u003c/span\u003e\u003c/span\u003e, which corresponds to \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{FE}={\\beta\\:}_{ME}=0.244\\)\u003c/span\u003e\u003c/span\u003e. For simulating the outcome in the presence of DE, we used the same population and set \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:PV{E}_{FE}=PV{E}_{ME}=0.4\\text{%}\\)\u003c/span\u003e\u003c/span\u003e, leading to \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{FE}={\\beta\\:}_{ME}=0.098\\)\u003c/span\u003e\u003c/span\u003e. For the simulations involving the exposure in the presence of RPS, we considered two subpopulations with MAFs being 0.25 and 0.35, respectively, which is similar to the simulation for GWAS. And then, we also introduced \u003cem\u003eL\u003c/em\u003e in the model to perform the effect of RPS, and set \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:PV{E}_{LE}=5\\text{%}\\)\u003c/span\u003e\u003c/span\u003e, resulting in \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{LE}=0.224\\)\u003c/span\u003e\u003c/span\u003e to simulate the exposure. After that, when simulating the outcome in the presence of RPS, we assigned \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:PV{E}_{LE}=10\\text{%}\\)\u003c/span\u003e\u003c/span\u003e, which yields \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{LE}=0.316\\)\u003c/span\u003e\u003c/span\u003e. In this scenario, we assumed no DE (i.e. \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{FE}={\\beta\\:}_{ME}=0\\)\u003c/span\u003e\u003c/span\u003e). In the scenario of no bias, for both the exposure and outcome variables, we considered a homogenous population which is the same as DE, and set\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\:{\\beta\\:}_{FE}={\\beta\\:}_{ME}=0\\)\u003c/span\u003e\u003c/span\u003e. After generating the exposure and outcome variables, the samples were randomly divided into two parts of equal sizes. In the first part (the first sample), the outcome variable was removed to investigate the association between the offspring’s genotype and the exposure, while in the second part (the second sample), the exposure variable was deleted to study the association between the offspring’s genotype and the outcome. For each simulation setting, we used the sample sizes of 2,000, 4,000, and 6,000, ensuring that the two samples in the MR analysis had equal sizes of 1,000, 2,000, and 3,000, respectively. Finally, we considered 10,000 replicates, and we compared the results of three methods (FT-SEM, lm_parent and lm) with the ratio estimation approach.\u003c/p\u003e\u003cp\u003eReal data applications\u003c/p\u003e\u003cp\u003eTo illustrate the approaches and evaluate the potential biases from DE and RPS, we performed the family-based GWAS by using the MCTFR data set, and combined these GWAS results and the summary data from the sibling-based GWAS\u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e to conduct the family-based two-sample MR analyses. The MCTFR data set is accessible via the Database of Genotypes and Phenotypes (dbGaP) with the accession number phs000620.v1.p1 (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ncbi.nlm.nih.gov/projects/gap/cgi-bin/study.cgi?study_id=phs000620.v1.p1\u003c/span\u003e\u003cspan address=\"https://www.ncbi.nlm.nih.gov/projects/gap/cgi-bin/study.cgi?study_id=phs000620.v1.p1\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). The data set includes the information from 2,183 families and 6,784 individuals. Supplementary Fig.\u0026nbsp;8 shows the QC process applied to this data set, resulting in a final total of 1,187 family trios used for the FT-SEM and lm_parent analyses. Moreover, lm uses the data of 2,183 independent individuals by randomly selecting one individual from each family. Additionally, the data set contains five phenotypes—NIC, CON, DEP, DRG and BD—all of which were used to carry out the GWAS.\u003c/p\u003e\u003cp\u003eThe GWAS based on the MCTFR data set was conducted using three methods (FT-SEM, lm_parent and lm). When fitting FT-SEM, the offspring's gender, date of birth and age, and the father's and mother's dates of birth and ages were included as covariates. For lm_parent, we considered the covariates: the offspring’s gender, date of birth and age. lm incorporated the gender, the date of birth, the age and the generation of the individual as the covariates, and also took account of the interaction terms between generation and gender, generation and date of birth, and generation and age. Additionally, the first 10 principal components of all the autosomal SNPs were included in lm to adjust for population stratification.\u003c/p\u003e\u003cp\u003eTo implement the two-sample MR analysis and investigate the causal effects of offspring’s BMI on offspring’s five phenotypes (NIC, CON, DEP, DRG and BD), we first need to select IVs. Supplementary Fig.\u0026nbsp;9 details the filtering process of the IVs used in IVW FT-SEM and IVW lm_parent, while Supplementary Fig.\u0026nbsp;10 outlines the filtering process of the IVs for IVW lm. SNP-exposure results were derived either from the sibling-based GWAS\u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e (for IVW FT-SEM and IVW lm_parent) or from the GWAS summary data based on independent individuals\u003csup\u003e\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e\u003c/sup\u003e (for IVW lm), while SNP-outcome results were obtained from our GWAS based on the MCTFR data set. Moreover, weak instruments with the F-statistic \u0026lt; 10 were excluded according to standard guidelines\u003csup\u003e\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e\u003c/sup\u003e. For IVW FT-SEM and IVW lm_parent, 27 independent SNPs (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{r}^{2}\\:\u0026lt;\\:0.001\\)\u003c/span\u003e\u003c/span\u003e within 10,000 kb) significantly associated with BMI (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:p\\:\u0026lt;5\\times\\:{10}^{-8}\\)\u003c/span\u003e\u003c/span\u003e) were selected. These SNPs were identified from the combined results of the sibling-based GWAS (“ieu-b-4816”) and the MCTFR-based GWAS using FT-SEM and lm_parent. For IVW lm, 448 independent SNPs meeting the same criteria were selected from the combined results of the independent GWAS (“ieu-b-40”) and the MCTFR-based lm. Additionally, the variants were clumped, and allele effect sizes were harmonized across the data sets using the TwoSampleMR package\u003csup\u003e\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e\u003c/sup\u003e. The weighted median, the weighted mode, and the MR-Egger are two-sample MR methods that are more robust to horizontal pleiotropy compared to IVW, but they have lower statistical power. Therefore, they can be used in sensitivity analyses to validate the results of our two-sample MR study. The summary data extraction, two-sample MR implementation, and sensitivity analyses (IVW, weighted median, weighted mode and MR-Egger) were performed using the TwoSampleMR package\u003csup\u003e\u003cspan additionalcitationids=\"CR39\" citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e–\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eAuthor contributions\u003c/h2\u003e\n\u003cp\u003eS.Z. developed the methods, performed the simulation, conducted real data analysis, and wrote the manuscript. H.C., J.-H.M., Q.-W.Z., J.-Y.Z. validated the models and the results. H.-W.C, J.-H.M, Q.-W.Z., Y.-S.L., X.-B.W. and J.-Y.Z. reviewed and edited the manuscript. X.-B.W., J.-Y.Z. administrated the project and supervised the research. All authors contributed to the article and approved the submitted version.\u003c/p\u003e\n\u003ch2\u003eAcknowledgments\u003c/h2\u003e\n\u003cp\u003eThis work was supported by the National Natural Science Foundation of China (grant no. 82173619), Guangdong Basic and Applied Basic Research Foundation (grant no. 2023A1515011242) and Hainan Province Science and Technology Special Fund (grant no. ZDYF2025SHFZ046). The Minnesota Center for Twin and Family Research (MCTFR) was supported by the National Institute on Drug Abuse (grant no. U01 DA024417). The sample ascertainment and data collection in MCTFR data were supported by the National Institute on Drug Abuse (grant nos. R37 DA05147 and R01 DA13240), the National Institute on Alcohol Abuse and Alcoholism (grant nos. R01 AA09367 and R01 AA11886), and the National Institute of Mental Health (grant no. R01 MH66140).\u003c/p\u003e\n\u003ch2\u003eData availability\u003c/h2\u003e\n\u003cp\u003eThe Minnesota Center for Twin and Family Research data used for this study can be found on the database of Genotypes and Phenotypes with accession number phs000620.v1.p1 (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ncbi.nlm.nih.gov/projects/gap/cgi-bin/study.cgi?study_id=phs000620.v1.p1\u003c/span\u003e\u003c/span\u003e.) The summary data from within-sibship GWAS are available at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://gwas.mrcieu.ac.uk/datasets/ieu-b-4816/\u003c/span\u003e\u003c/span\u003e. The summary data from independent GWAS are available at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://gwas.mrcieu.ac.uk/datasets/ieu-b-40/\u003c/span\u003e\u003c/span\u003e. The code using in this paper are available at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/ShunZhang0816/FT-SEM/\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e\n\u003ch2\u003eCompeting interests\u003c/h2\u003e\n\u003cp\u003eThe authors declare no competing interests.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003ePingault, J.-B. \u003cem\u003eet al.\u003c/em\u003e Using genetic data to strengthen causal inference in observational research. \u003cem\u003eNat Rev Genet\u003c/em\u003e \u003cstrong\u003e19\u003c/strong\u003e, 566\u0026ndash;580 (2018).\u003c/li\u003e\n\u003cli\u003eDavey Smith, G. \u0026amp; Hemani, G. Mendelian randomization: genetic anchors for causal inference in epidemiological studies. \u003cem\u003eHum Mol Genet\u003c/em\u003e \u003cstrong\u003e23\u003c/strong\u003e, R89\u0026ndash;R98 (2014).\u003c/li\u003e\n\u003cli\u003eSmith, G. D. \u0026amp; Ebrahim, S. \u0026lsquo;Mendelian randomization\u0026rsquo;: can genetic epidemiology contribute to understanding environmental determinants of disease? \u003cem\u003eInt J Epidemiol\u003c/em\u003e \u003cstrong\u003e32\u003c/strong\u003e, 1\u0026ndash;22 (2003).\u003c/li\u003e\n\u003cli\u003eBurgess, S., Small, D. S. \u0026amp; Thompson, S. G. A review of instrumental variable estimators for Mendelian randomization. \u003cem\u003eStat Methods Med Res\u003c/em\u003e \u003cstrong\u003e26\u003c/strong\u003e, 2333\u0026ndash;2355 (2017).\u003c/li\u003e\n\u003cli\u003eHwang, L.-D., Davies, N. M., Warrington, N. M. \u0026amp; Evans, D. M. Integrating Family-Based and Mendelian Randomization Designs. \u003cem\u003eCold Spring Harb Perspect Med\u003c/em\u003e \u003cstrong\u003e11\u003c/strong\u003e, a039503 (2021).\u003c/li\u003e\n\u003cli\u003eWindmeijer, F., Farbmacher, H., Davies, N. \u0026amp; Davey Smith, G. On the Use of the Lasso for Instrumental Variables Estimation with Some Invalid Instruments. \u003cem\u003eJ Am Stat Assoc\u003c/em\u003e \u003cstrong\u003e114\u003c/strong\u003e, 1339\u0026ndash;1350 (2019).\u003c/li\u003e\n\u003cli\u003eHartwig, F. P., Davey Smith, G. \u0026amp; Bowden, J. Robust inference in summary data Mendelian randomization via the zero modal pleiotropy assumption. \u003cem\u003eInt J Epidemiol\u003c/em\u003e \u003cstrong\u003e46\u003c/strong\u003e, 1985\u0026ndash;1998 (2017).\u003c/li\u003e\n\u003cli\u003eLawson, D. J. \u003cem\u003eet al.\u003c/em\u003e Is population structure in the genetic biobank era irrelevant, a challenge, or an opportunity? \u003cem\u003eHum Genet\u003c/em\u003e \u003cstrong\u003e139\u003c/strong\u003e, 23\u0026ndash;41 (2020).\u003c/li\u003e\n\u003cli\u003eBrumpton, B. \u003cem\u003eet al.\u003c/em\u003e Avoiding dynastic, assortative mating, and population stratification biases in Mendelian randomization through within-family analyses. \u003cem\u003eNat Commun\u003c/em\u003e \u003cstrong\u003e11\u003c/strong\u003e, 3519 (2020).\u003c/li\u003e\n\u003cli\u003eKong, A. \u003cem\u003eet al.\u003c/em\u003e The nature of nurture: Effects of parental genotypes. \u003cem\u003eScience\u003c/em\u003e \u003cstrong\u003e359\u003c/strong\u003e, 424\u0026ndash;428 (2018).\u003c/li\u003e\n\u003cli\u003eBowden, J., Davey Smith, G. \u0026amp; Burgess, S. Mendelian randomization with invalid instruments: effect estimation and bias detection through Egger regression. \u003cem\u003eInt J Epidemiol\u003c/em\u003e \u003cstrong\u003e44\u003c/strong\u003e, 512\u0026ndash;525 (2015).\u003c/li\u003e\n\u003cli\u003eYuan, Z. \u003cem\u003eet al.\u003c/em\u003e Likelihood-based Mendelian randomization analysis with automated instrument selection and horizontal pleiotropic modeling. \u003cem\u003eSci Adv\u003c/em\u003e \u003cstrong\u003e8\u003c/strong\u003e, eabl5744 (2022).\u003c/li\u003e\n\u003cli\u003eBowden, J., Davey Smith, G., Haycock, P. C. \u0026amp; Burgess, S. Consistent Estimation in Mendelian Randomization with Some Invalid Instruments Using a Weighted Median Estimator. \u003cem\u003eGenet Epidemiol\u003c/em\u003e \u003cstrong\u003e40\u003c/strong\u003e, 304\u0026ndash;314 (2016).\u003c/li\u003e\n\u003cli\u003eBurgess, S., Foley, C. N., Allara, E., Staley, J. R. \u0026amp; Howson, J. M. M. A robust and efficient method for Mendelian randomization with hundreds of genetic variants. \u003cem\u003eNat Commun\u003c/em\u003e \u003cstrong\u003e11\u003c/strong\u003e, 376 (2020).\u003c/li\u003e\n\u003cli\u003eQi, G. \u0026amp; Chatterjee, N. Mendelian randomization analysis using mixture models for robust and efficient estimation of causal effects. \u003cem\u003eNat Commun\u003c/em\u003e \u003cstrong\u003e10\u003c/strong\u003e, 1941 (2019).\u003c/li\u003e\n\u003cli\u003eYuan, Z. \u003cem\u003eet al.\u003c/em\u003e Testing and controlling for horizontal pleiotropy with probabilistic Mendelian randomization in transcriptome-wide association studies. \u003cem\u003eNat Commun\u003c/em\u003e \u003cstrong\u003e11\u003c/strong\u003e, 3861 (2020).\u003c/li\u003e\n\u003cli\u003eFulker, D. W., Cherny, S. S., Sham, P. C. \u0026amp; Hewitt, J. K. Combined linkage and association sib-pair analysis for quantitative traits. \u003cem\u003eAm J Hum Genet\u003c/em\u003e \u003cstrong\u003e64\u003c/strong\u003e, 259\u0026ndash;267 (1999).\u003c/li\u003e\n\u003cli\u003eWarrington, N. M., Freathy, R. M., Neale, M. C. \u0026amp; Evans, D. M. Using structural equation modelling to jointly estimate maternal and fetal effects on birthweight in the UK Biobank. \u003cem\u003eInt J Epidemiol\u003c/em\u003e \u003cstrong\u003e47\u003c/strong\u003e, 1229\u0026ndash;1241 (2018).\u003c/li\u003e\n\u003cli\u003eEvans, D. M., Moen, G.-H., Hwang, L.-D., Lawlor, D. A. \u0026amp; Warrington, N. M. Elucidating the role of maternal environmental exposures on offspring health and disease using two-sample Mendelian randomization. \u003cem\u003eInt J Epidemiol\u003c/em\u003e \u003cstrong\u003e48\u003c/strong\u003e, 861\u0026ndash;875 (2019).\u003c/li\u003e\n\u003cli\u003eWang, G., Warrington, N. M. \u0026amp; Evans, D. M. Partitioning genetic effects on birthweight at classical human leukocyte antigen loci into maternal and fetal components, using structural equation modelling. \u003cem\u003eInt J Epidemiol\u003c/em\u003e \u003cstrong\u003e53\u003c/strong\u003e, dyad142 (2023).\u003c/li\u003e\n\u003cli\u003eHowe, L. J. \u003cem\u003eet al.\u003c/em\u003e Within-sibship genome-wide association analyses decrease bias in estimates of direct genetic effects. \u003cem\u003eNat Genet\u003c/em\u003e \u003cstrong\u003e54\u003c/strong\u003e, 581\u0026ndash;592 (2022).\u003c/li\u003e\n\u003cli\u003eHowe, L. J. \u003cem\u003eet al.\u003c/em\u003e Educational attainment, health outcomes and mortality: a within-sibship Mendelian randomization study. \u003cem\u003eInt J Epidemiol\u003c/em\u003e \u003cstrong\u003e52\u003c/strong\u003e, 1579\u0026ndash;1591 (2023).\u003c/li\u003e\n\u003cli\u003eMorris, N. J., Elston, R. C. \u0026amp; Stein, C. M. A framework for structural equation models in general pedigrees. \u003cem\u003eHum Hered\u003c/em\u003e \u003cstrong\u003e70\u003c/strong\u003e, 278\u0026ndash;286 (2010).\u003c/li\u003e\n\u003cli\u003eStein, C. M., Morris, N. J., Hall, N. B. \u0026amp; Nock, N. L. Structural Equation Modeling. \u003cem\u003eMethods Mol Biol\u003c/em\u003e \u003cstrong\u003e1666\u003c/strong\u003e, 557\u0026ndash;580 (2017).\u003c/li\u003e\n\u003cli\u003eGrotzinger, A. D. \u003cem\u003eet al.\u003c/em\u003e Genomic structural equation modelling provides insights into the multivariate genetic architecture of complex traits. \u003cem\u003eNat Hum Behav\u003c/em\u003e \u003cstrong\u003e3\u003c/strong\u003e, 513\u0026ndash;525 (2019).\u003c/li\u003e\n\u003cli\u003ePritikin, J. N., Neale, M. C., Prom-Wormley, E. C., Clark, S. L. \u0026amp; Verhulst, B. GW-SEM 2.0: Efficient, Flexible, and Accessible Multivariate GWAS. \u003cem\u003eBehav Genet\u003c/em\u003e \u003cstrong\u003e51\u003c/strong\u003e, 343\u0026ndash;357 (2021).\u003c/li\u003e\n\u003cli\u003eMinică, C. C., Dolan, C. V., Boomsma, D. I., De Geus, E. \u0026amp; Neale, M. C. Extending Causality Tests with Genetic Instruments: An Integration of Mendelian Randomization with the Classical Twin Design. \u003cem\u003eBehav Genet\u003c/em\u003e \u003cstrong\u003e48\u003c/strong\u003e, 337\u0026ndash;349 (2018).\u003c/li\u003e\n\u003cli\u003eBurgess, S. \u0026amp; Thompson, S. G. \u003cem\u003eMendelian Randomization: Methods for Causal Inference Using Genetic Variants\u003c/em\u003e. (CRC Press, Boca Raton London New York, 2021).\u003c/li\u003e\n\u003cli\u003eKong, Y.-F. \u003cem\u003eet al.\u003c/em\u003e Detection of Parent-of-Origin Effects for the Variants Associated With Behavioral Disinhibition in the MCTFR Data. \u003cem\u003eFront Genet\u003c/em\u003e \u003cstrong\u003e13\u003c/strong\u003e, 831685 (2022).\u003c/li\u003e\n\u003cli\u003eMcGue, M. \u003cem\u003eet al.\u003c/em\u003e A genome-wide association study of behavioral disinhibition. \u003cem\u003eBehav Genet\u003c/em\u003e \u003cstrong\u003e43\u003c/strong\u003e, 363\u0026ndash;373 (2013).\u003c/li\u003e\n\u003cli\u003eMcCaw, Z. R., Lane, J. M., Saxena, R., Redline, S. \u0026amp; Lin, X. Operating characteristics of the rank-based inverse normal transformation for quantitative trait analysis in genome-wide association studies. \u003cem\u003eBiometrics\u003c/em\u003e \u003cstrong\u003e76\u003c/strong\u003e, 1262\u0026ndash;1272 (2020).\u003c/li\u003e\n\u003cli\u003eYengo, L. \u003cem\u003eet al.\u003c/em\u003e Meta-analysis of genome-wide association studies for height and body mass index in \u0026sim;700000 individuals of European ancestry. \u003cem\u003eHum Mol Genet\u003c/em\u003e \u003cstrong\u003e27\u003c/strong\u003e, 3641\u0026ndash;3649 (2018).\u003c/li\u003e\n\u003cli\u003eWang, P. \u003cem\u003eet al.\u003c/em\u003e Smoking and Socio-demographic correlates of BMI. \u003cem\u003eBMC Public Health\u003c/em\u003e \u003cstrong\u003e16\u003c/strong\u003e, 500 (2016).\u003c/li\u003e\n\u003cli\u003eLichenstein, S. D. \u003cem\u003eet al.\u003c/em\u003e Familial risk for alcohol dependence and developmental changes in BMI: the moderating influence of addiction and obesity genes. \u003cem\u003ePharmacogenomics\u003c/em\u003e \u003cstrong\u003e15\u003c/strong\u003e, 1311\u0026ndash;1321 (2014).\u003c/li\u003e\n\u003cli\u003eMc, N. \u003cem\u003eet al.\u003c/em\u003e OpenMx 2.0: Extended Structural Equation and Statistical Modeling. \u003cem\u003ePsychometrika\u003c/em\u003e \u003cstrong\u003e81\u003c/strong\u003e, 535\u0026ndash;549 (2016).\u003c/li\u003e\n\u003cli\u003eHartwig, F. P. \u0026amp; Davies, N. M. Why internal weights should be avoided (not only) in MR-Egger regression. \u003cem\u003eInt J Epidemiol\u003c/em\u003e \u003cstrong\u003e45\u003c/strong\u003e, 1676\u0026ndash;1678 (2016).\u003c/li\u003e\n\u003cli\u003eBurgess, S., Thompson, S. G., \u0026amp; CRP CHD Genetics Collaboration. Avoiding bias from weak instruments in Mendelian randomization studies. \u003cem\u003eInt J Epidemiol\u003c/em\u003e \u003cstrong\u003e40\u003c/strong\u003e, 755\u0026ndash;764 (2011).\u003c/li\u003e\n\u003cli\u003eHemani, G. \u003cem\u003eet al.\u003c/em\u003e The MR-Base platform supports systematic causal inference across the human phenome. \u003cem\u003eElife\u003c/em\u003e \u003cstrong\u003e7\u003c/strong\u003e, e34408 (2018).\u003c/li\u003e\n\u003cli\u003eElsworth, B. \u003cem\u003eet al.\u003c/em\u003e The MRC IEU OpenGWAS data infrastructure. 2020.08.10.244293 Preprint at https://doi.org/10.1101/2020.08.10.244293 (2020).\u003c/li\u003e\n\u003cli\u003eHemani, G., Tilling, K. \u0026amp; Davey Smith, G. Orienting the causal relationship between imprecisely measured traits using GWAS summary data. \u003cem\u003ePLoS Genet\u003c/em\u003e \u003cstrong\u003e13\u003c/strong\u003e, e1007081 (2017).\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[{"identity":"0a313f87-afbb-4349-97ea-8354f166859e","identifier":"10.13039/501100001809","name":"National Natural Science Foundation of China","awardNumber":"grant no. 82173619","order_by":0}],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"Southern Medical University","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"dynastic effect, residual population stratification, structure equation modelling, Mendelian randomization, family structure","lastPublishedDoi":"10.21203/rs.3.rs-6163190/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6163190/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eEffect size estimates in genome-wide association studies (GWAS) and Mendelian randomization (MR) studies for independent individuals may be biased due to dynastic effect (DE) and residual population stratification (RPS). Existing GWAS methods for family trios effectively controlled such biases, while only using parental and offspring\u0026rsquo;s genotypes and offspring\u0026rsquo;s phenotype, and not incorporating parental phenotypes, which causes loss in estimation accuracy and test power. Therefore, we proposed a novel GWAS method based on structural equation modelling for family trios, denoted by FT-SEM. FT-SEM simultaneously uses parental and offspring\u0026rsquo;s genotypes and phenotypes. Simulation results demonstrate that FT-SEM substantially improves estimation accuracy and test power while controlling bias and type I error rate. Using family trios from Minnesota Center for Twin and Family Research (MCTFR), we found that DE and RPS greatly distort the results only based on independent individuals, and FT-SEM effectively corrects such biases. Combining the GWAS results from MCTFR with existing summary data, we performed several two-sample MR analyses. We observed that the effects of BMI on nicotine, alcohol consumption and behavior disorder were due to bias rather than causality. Our findings underscore the necessity of using families to validate the results of GWAS and MR, and highlight FT-SEM\u0026rsquo;s advantages.\u003c/p\u003e","manuscriptTitle":"A robust and powerful GWAS method for family trios supporting within-family Mendelian randomization analysis","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-03-06 09:22:55","doi":"10.21203/rs.3.rs-6163190/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"bac32faf-c5c1-45d5-935a-2ad3a8f69c40","owner":[],"postedDate":"March 6th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":45242968,"name":"Biostatistics"}],"tags":[],"updatedAt":"2025-03-06T09:22:55+00:00","versionOfRecord":[],"versionCreatedAt":"2025-03-06 09:22:55","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6163190","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6163190","identity":"rs-6163190","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.