Efficiency Under Epistasis of Haplotype-Based Genomic Selection for Pure Line Breeding | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Efficiency Under Epistasis of Haplotype-Based Genomic Selection for Pure Line Breeding José Marcelo Viana, Jean Paulo Silva This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7745171/v1 This work is licensed under a CC BY 4.0 License Status: Under Revision Version 1 posted 10 You are reading this latest preprint version Abstract The haplotype-based approach is considered a more efficient alternative for genomic selection under epistasis. In this simulation-based study, we evaluated the efficacy of haplotype-based genomic selection under seven digenic epistasis types and the significance of including additive x additive (AxA) effect in selection, in pure line breeding. We simulated a genome comprised of 10 chromosomes of 100 cM, including 1000 genes and 49825 SNPs. We assumed positive dominance and 30% of epistatic genes. We performed haplotype-based genomic selection, using haplotypes of sizes 3, 6, and 9. The training set included all F2 individuals and 18% of the F3 to F7 individuals. Selection efficacy was assessed using realized total genetic gain. We also evaluated the correlation between prediction accuracy and realized genetic gain and the decrease in the genotypic variance. Regardless of the epistasis type, genomic selection is an efficient procedure to develop superior pure lines. There was no haplotype size effect on the genomic selection efficacy. Under no selection, the AxA genetic value showed no significant negative correlation with the number of favorable genes. Under selection, the inclusion of the predicted AxA value for selection did not increase the selection efficacy. Haplotype- and SNP-based genomic selection were equally efficient. Epistasis can decrease the total genetic gain in pure line breeding, compared to traits determined only by additive and dominance gene effects. Under epistasis, there was a highly positive correlation between genetic gain and prediction accuracy. There is no difference between genomic selection procedures regarding the decrease in the genotypic variance. Biological sciences/Genetics/Plant breeding Biological sciences/Plant sciences/Plant breeding haplotype genomic selection pure lines epistasis Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Introduction Genomic selection has been successfully employed for the improvement of self-pollinated crops such as soybean, wheat, and rice, primarily through SNP-based approach (Tessema et al. , 2020; Baertschi et al. , 2021; Bandillo et al. , 2023). While single nucleotide polymorphisms (SNPs) are widely used, this approach presents limitations. SNPs don't always reveal all variations and allelic combinations of genes controlling a trait, and their effectiveness often decreases with lower relatedness between training and validation populations (Villumsen et al. , 2009; Habier et al. , 2010). Linkage disequilibrium (LD) between markers was initially believed to be the primary source of information for genomic prediction. However, Habier et al. (2007) showed that SNP also capture genetic relationships among individuals. This suggests that the accuracy of SNP-based predictions may depend more on capturing relationships among individuals than on direct LD with causal QTLs (Habier et al. , 2007; Weber et al. , 2023). In contrast, haplotype-based approach offers a promising alternative. They can improve predictive accuracy by potentially increasing LD with causal loci and capturing complex interactions among nearby genes (Meuwissen et al. , 2014; Hess et al. , 2017; Jiang et al. , 2018). Weber et al. (2023) evaluated the accuracy of haplotype-based genomic prediction, reporting inconsistent results across traits, species, and models. Processing a wheat dataset, the higher increases in prediction accuracy from haplotyping, relative to the SNP-based approach, were observed for grain yield and sedimentation value. The accuracies increased from 0.70 to 0.82, from 0.49 to 0.62, and from 0.49 to 0.64, depending on the method of haplotyping and statistical model. Using the soybean, maize, and canola datasets, no differences were observed between haplotype- and SNP-based approaches. Sallam et al. (2020) observed more consistent benefits from haplotype-based genomic predictions. Their investigation evaluated haplotypes of varying sizes under two cross-validation schemes. All haplotype sizes outperformed SNP-based for grain yield and protein content, with the haplotype size 15 providing the greatest prediction accuracy gains. Notably, stratified sampling enhanced predictive performance the most, achieving up to 14.3% improvement for grain yield and 16.8% for protein content. These studies clearly demonstrate that while haplotype-based genomic prediction holds promise, its success is highly contingent upon multiple interacting factors, such as model, trait architecture, species, population, haplotyping method, and training set. The discussion about the best way to model genetic variation extends to the inclusion of epistatic effects in genomic prediction models. However, the literature presents contrasting results regarding the impact of including these interactions. It is noteworthy that, of the limited number of studies including epistasis, the vast majority have solely focused on models accounting for additive x additive effects. In Weber et al. (2023), incorporating additive x additive effects into haplotype-based genomic prediction resulted in only marginal improvements in prediction accuracy for most traits. For example, in wheat, grain yield prediction and its epistatic extension showed a negligible increase, from 0.80 to 0.81. In soybean, the same model improved prediction accuracy for oil content from 0.68 to 0.69, with no increase for protein content. Miller et al. (2023), evaluating parent selection strategies in soybean breeding, reported that including additive x additive effects led to more than a 50% increase in prediction accuracy, depending on the selection method. García-Barrios et al. (2024), studying multi-environment genomic prediction in wheat, found no significant improvement when incorporating additive x additive interactions, compared to purely additive model. Quantifying the impact of epistasis on selection has long been a challenge in plant and animal breeding. With the advancement of genomic data, there has been renewed interest in modeling epistatic interactions; however, assessing the significance of epistasis in empirical datasets remains difficult, since there is no previous knowledge if there is epistasis. In this simulation-based study, we evaluated the efficacy of haplotype-based genomic selection under seven digenic epistasis types and the significance of including additive x additive effect in selection, in pure line breeding. To our knowledge, this is the first simulation-based investigation on the efficiency of haplotype-based genomic selection under epistasis in pure line breeding. Materials and Methods Using the software REALbreeding (available by request), we simulated a genome comprising 10 chromosomes of 100 cM, including 1000 minor genes, and 49825 SNPs. The average densities for genes and SNPs were 1 and 0.02 cM, respectively. The simulated trait was grain yield in soybean, assuming a planting density of 300000 plants per hectare. As an input for the software, we defined, assuming no epistasis, the minimum genotypic and phenotypic values as 7 and 4 g/plant, respectively, while the maximum values were 16 and 20 g/plant, respectively. The degree of dominance ranged from 0 to 1.2, with a mean of 0.6. Broad-sense heritability was set at 30% in each generation. We assumed digenic epistasis for 30% of the genes, resulting in 150 epistatic pairs. The following epistasis types were modeled based on Kempthorne (1955): complementary (9:7 in F₂), duplicate (15:1 in F₂), dominant (12:3:1 in F₂), recessive (9:3:4 in F₂), dominant and recessive (13:3 in F₂), duplicate genes with cumulative effects (9:6:1 in F₂), non-epistatic gene interaction (9:3:3:4 in F₂), and an admixture of types. For the admixture, REALbreeding randomly assigns an epistasis type to each pair of interacting genes. For each epistasis, we kept the same pairs of epistatic genes. The base population consisted of 250 F₂ individuals derived from a cross between two contrasting pure lines, one carrying 70% and the other 30% of the favorable alleles. For details on the simulation procedures and theoretical background, refer to Viana & Garcia (2022). From F 3 to F 8 generations, we performed haplotype-based genomic selection based on predicted sum of the additive and additive x additive values, using haplotype sizes 3, 6, and 9. The training set included all F 2 individuals and 18% of the F 3 to F 7 individuals. Each progeny consisted of 20 plants. The F 3 consisted of 250 progenies (5000 plants). The number of progenies from F 4 to F 8 were 150 (3000 plants), 50 (1000 plants), 20 (400 plants), and 5 (100 plants), respectively. The proportion of selected plants were 3.0, 5.0, 1.7, 2.5, and 1.2%. As reference scenarios, we also evaluated SNP-based selection, haplotype-based selection based only in the predicted additive value, no selection, and selection based on true genotypic value. Under no selection, we advanced 500 individuals in each generation using the single seed descent (SSD) method. For selection based on the true genotypic values, we used the values provided by REALbreeding . We replicated each scenario 10 times. We fitted the model , where the are the incidence matrices associated with the vectors of additive, dominance, additive x additive (AxA), additive x dominance (AxD), and dominance x dominance (DxD) genetic values, respectively. The model was implemented using the BGLR R package (Pérez and De Los Campos, 2014). We used the defaults for burn in (500) and iterations (2000), checking for convergence. The additive and dominance genomic matrices were computed following the methods proposed by VanRaden (2008) (first method) and Su et al. (2012), respectively, using the sommer R package (Covarrubias-Pazaran, 2016). We applied the Hadamard product to the additive and dominance relationship matrices to generate the epistatic genomic matrices. The AxD genetic value corresponds to the sum of AxD and dominance x additive (DxA) values. We built haplotypes based on a fixed number of adjacent SNPs along the chromosome, employing our own R code. In this study, we defined haplotype lengths of 3, 6, and 9, with length 9 applied only to an admixture of types. To evaluate the haplotype pattern in the F 2 population, we used Haploview (Barrett et al. , 2005). Haplotypes were defined using three distinct methods: Gabriel et al. (2002), the four gamete rule (Wang et al. (2002), and the solid spine of LD method. Haplotyping requires phased SNP genotypes, which we conveniently obtained from the REALbreeding ’s output, where the left allele is of maternal origin. We adopted the haplotype model proposed by Da (2015), in which each haplotype allele is treated as a SNP - a procedure referred to as ‘pseudo-SNP’ by Karimi et al. (2018). With a haplotype size , the total number haplotype alleles is . We calculated the genetic gains by the difference between the genotypic means of successive generations. We estimated the additive value prediction accuracy by Pearson’s correlation between the predicted additive values and the true genotypic values. Because of selection, REALbreeding does not compute the additive values. We computed the genotypic variance from the true genotypic values provided by REALbreeding . For testing differences between total genetic gains, we applied regression analysis or t test assuming distinct variances. Aiming to compute the correlations between the number of favorable genes and the additive and additive plus AxA values, we used the genotypes for genes and the true genetic values, assuming no selection from generations F 2 to F 8 , provided by REALbreeding . Results The analysis of the haplotype pattern in the F 2 generation (Figure 1) shows, irrespective of the method, a predominance of long haplotypes (average of 22.2 SNPs). However, only 4.5% of the SNPs were included in a haplotype. The ratio F 2 epistatic variance/genotypic variance ranged from 1.1, assuming complementary epistasis, to 13.9%, assuming duplicate epistasis. The additive variance ranged from 41.8 to 141.7% (due to negative covariances) of the genotypic variance in F 2 , while the AxA genotypic variance ranged from 0.5 to 6.1%. For the admixture of epistasis types, epistatic and AxA variances accounted for 6.4 and 2.0% of the genotypic variance, respectively. The total genetic gains from the haplotype-based genomic selection, based on predicted A + AxA values were 208.4, 212.0, and 218.1 kg ha -1 , for haplotypes sizes 3, 6, and 9, respectively (Figure 2). However, the linear and quadratic regression analyses evidenced no haplotype size effect (R 2 = 2.3 and 2.2%, respectively). Thus, the haplotype-based genomic selection provided an average total genetic gain of 212.8 kg ha -1 . The SNP-based genomic selection using A + AxA value provided a total genetic gain of 211.7 kg ha -1 , showing that both procedures are equivalent (Figure 2). Fitting the complete model but ignoring AxA effects in selection, the total genetic gains with haplotype- and SNP-based analyses were 229.5, 204.1, 204.1, for haplotype sizes 3, 6, and 9, respectively, and 203.0 kg ha -1 , respectively. Even for haplotype size 3, we observed non-significant changes in the selection efficacy by ignoring the predicted AxA effect in selection (-3.9 to 10.1%). For the seven distinct epistasis types we observed essentially the same results: no significant haplotype size effect, no significant difference between haplotype- and SNP-based approaches, and no significant effect by selecting based only in the predicted additive effect (Figure 2). The only exceptions occurred for complementary and duplicate epistasis, where significant decreases in the total genetic gains occurred by increasing the haplotype size from 3 to 6 (P values of 0.04 and 0.005). Comparing SNP with the best haplotype size, there was no significant differences between the total genetic gains. Regarding the impact of ignoring the AxA effect on selection, we observed no significant changes in the total genetic gains. The changes ranged from -14.8 to 17.1 kg ha -1 . The epistasis that minimized and maximized the total genetic gains were, respectively, duplicate, from genomic selection based on haplotype size 6 and predicted A + AxA value (163.6 kg ha -1 ), and non-epistatic gene interaction, from selection based on haplotype of size 6 and predicted additive value (271.8 kg ha -1 ), a statistically significant difference (Figure 2). The correlations between the ratio epistatic variance/genotypic variance with the minimum, average, and maximum total genetic gains were negative of intermediate magnitude (-0.42, -0.31, and -0.23). The scenario of no selection clearly shows how all genomic selection methods efficiently increased the means of the generations. In case of no selection and due to positive dominance, the decreases in the F 8 generation, relative to F 3 , ranged from -122.7, under duplicate epistasis, to -232.2 kg ha -1 , assuming non-epistatic gene interaction, also inversely proportional to the ratio epistatic variance/genotypic variance. Regarding the scenario of selection based on the true genotypic value, the minimum and maximum total genetic gains were 189.4, under recessive epistasis, and 256.0 kg ha -1 , assuming non-epistatic gene interaction. Again, the total genetic gain was inversely proportional to the ratio epistatic variance/genotypic variance (correlation of -0.24). The difference between these two gains is not statistically significant. The correlations of the additive and additive x additive values with the number of favorable genes evidenced that genomic selection based on the additive value was expected to be more effective than selection based on the sum A + AA, because predominantly negative values for the correlation between AxA and the number of favorable genes (Figure 3), irrespective of the epistasis type. Based on the magnitude of the correlations, we can state that the additive x additive genetic values showed no significant correlation with the number of favorable genes. The correlation between the additive value and the number of favorable genes ranged from approximately 0.6 under duplicate epistasis to approximately 0.8 assuming duplicate genes with cumulative effects. As expected, all selection procedures led to a significant decrease in the genotypic variance, irrespective of the epistasis type (99.0% on average) (Figure 4). Excepting dominant epistasis, based on the decrease in the genotypic variance under no selection, due to inbreeding, selection was responsible for 55.0 to 92.0% of the observed decrease. Selection based on the true genotypic value led to comparable decreases, from 52.0 to 92.0%. Regarding the prediction accuracy, the estimated values show clearly how haplotype - and SNP-based genomic selection are effective in pure line breeding, relative to phenotypic selection (full phenotyping), irrespective of the epistasis type (Figure 5). The accuracies ranged from 0.70 to 0.90, regardless of epistasis and selection process, decreasing from F 3 to F 7 . Another important result was the high positive correlation between genetic gains and prediction accuracies. Regardless of the method and epistasis, the correlations ranged from 0.87 to 0.99. Comparing our results for SNP-based genomic selection, including epistasis, with the results assuming no epistasis but the same genome, number and positions of genes and SNPs, parents, degree of dominance etc. (Silva and Viana, 2025), we additionally observed that epistasis increased or decreased the total genetic gains in pure line breeding, depending on the epistasis type. The change in the total genetic gains by including and modelling epistasis ranged from -22.2 to 3.6% (-9.3% on average). Finally, it is important to emphasize that, under epistasis, all genomic selection processes provided F 8 progenies statistically equal to the superior parent. The differences ranged from -1.0 to 2.0%, the same result observed assuming additive-dominance model. Discussion Based on several inheritance studies for qualitative traits (see any standard Genetics book), hundreds of inheritance studies for quantitative traits, using generation mean analysis, QTL mapping, and genome-wide studies, and many transcriptome, proteome, and metabolome investigations, geneticists agree that epistasis is the rule and absence of epistasis is the exception for complex traits (Mackay, 2014). Why, then, quantifying and assessing the significance of epistasis in evolution and selection remains a challenge for geneticists and breeders, even with the help of SNPs? Because, in general, epistatic genetic values have a relative lower significance in determining genotypic value, relative to the additive value. Or better, because the additive variance is the most significant component of the genotypic variance (Hill et al. , 2008). Furthermore, epistatic effects can be difficult to quantify and falsely declared under incomplete LD at low SNP density (De Los Campos et al. , 2019). In this first simulation-based investigation on the efficacy under epistasis of haplotype-based genomic selection in pure line breeding, we provided clear evidence that a genomic analysis can be effective to quantify epistatic effects, if they have a significant contribution in determining the quantitative trait, and to provide substantial genetic gains. Additionally, we showed that the heritable additive x additive effect can contribute for increasing selection efficacy, depending on the correlation between the number of favorable genes and the AxA effect. In our study, the correlations under no selection were negative of low magnitude, regardless of the epistasis type and the ratio epistatic variance/genotypic variance. Our evidence of no significance for selection of the additive x additive epistatic values was observed by García-Barrios et al. (2024) and Weber et al. (2023). Although there are inconclusive results for the efficacy of haplotype- and SNP-based genomic selection, haplotyping presents some expected advantages. Among them, we emphasize that haplotypes can better capture epistatic effects (Jiang et al. , 2018). Applying haplotype-based genomic selection from F 3 to F 7 , the total genetic gains ranged from 188 to 251 kg ha -1 , depending on the epistasis type and the ratio epistatic variance/genotypic variance. This means 38 to 50 kg ha -1 year -1 . The correlation between the gain and the ratio was -0.3. Krause et al. (2023) estimated in 18 to 40 kg ha -1 year -1 the realized genetic gains in soybean public programs. From the assessment of the haplotype size effect, we observed that increasing the haplotype size from 3 to 9, under an admixture of epistasis types, and from 3 to 6 for the seven distinct epistasis, did not increase the genomic selection efficacy. Unfortunately, there is no previous study with self-pollinated crops assessing the haplotype size effect based on realized genetic gains. Only based on breeding value prediction accuracy. Sallam et al. (2020) also did not observe haplotype size effect. The prediction accuracies for haplotype sizes 5, 10, 15, and 20 oscillated from 0.31 to 0.34. Regarding the significance of the predicted AxA effect, because predominantly negative of low magnitude correlations with the number of favorable genes, under no selection, selection based on the additive values was expected to be only slightly efficient than selection based on the additive plus AxA value. And the two processes were equally efficient. The ratio between the realized total genetic gains ranged from 0.9 to 1.2 (1.0 on average). García-Barrios et al. (2024) evaluated multi-environment prediction models and datasets for wheat and found similar performance between purely additive model and models incorporating additive x additive effects. For wheat also, Jiang and Reif (2015) observed slightly improved prediction accuracies when additive x additive effects were included. Raffo et al. (2022) concluded that epistasis can enhance predictive ability, especially under conditions of high relatedness. However, the difference between the prediction accuracies was only 0.06. Weber et al. (2023), one of the few studies in plant breeding to evaluate the impact of including epistatic effects in haplotype-based models, found only marginal improvements in predictive accuracy for grain yield in wheat and for protein and oil contents in soybean, regardless of whether haplotypes or SNPs were used. Comparing haplotype- and SNP-based approaches, we observed that both methods were equivalent, considering total realized genetic gain, additive value prediction accuracy, and decrease in the genotypic variance, irrespective of the epistasis and ratio epistatic variance/genotypic variance. The SNP-based approach provided genetic gains from 184 to 247 kg ha -1 , depending on the epistasis type and the ratio epistatic variance/genotypic variance. Using a wheat inbred line panel, Difabachew et al. (2023) showed that predictive performance varied depending on haplotype definition and trait. Haplotypes defined by LD provided higher accuracy for disease resistance, while fixed-length haplotyping performed better for plant height. However, all haplotype-based approaches were less effective for grain yield. Positive results favoring haplotype-based models are often associated with low marker density. This trend is consistent with findings from Villumsen et al. (2009), who showed that haplotype-based prediction tends to outperform SNP-based models particularly when marker density is low. Regardless of the statistical method, inbred line panel, missing data rate, and trait, He et al. (2023) also observed equivalence in prediction accuracy for haplotypes and SNPs in rice. Conclusions Regardless of the epistasis type, genomic selection is an efficient procedure to develop superior pure lines. There was no haplotype size effect on the genomic selection efficacy. Under no selection, the AxA genetic value showed no significant negative correlation with the number of favorable genes. Under selection, the inclusion of the predicted AxA value for selection did not increase the selection efficacy. Haplotype- and SNP-based genomic selection were equally efficient. Epistasis can decrease the total genetic gain in pure line breeding, compared to traits determined only by additive and dominance gene effects. Under epistasis, there was a highly positive correlation between genetic gain and prediction accuracy. There is no difference between genomic selection procedures regarding the decrease in the genotypic variance. Declarations Acknowledgments Competing Interests The authors declare no conflict of interest. Author Contributions Data Archiving The dataset is available at https://doi.org/10.6084/m9.figshare.29825855.v1. References Baertschi C, Cao T-V, Bartholomé J, Ospina Y, Quintero C, Frouin J, et al. (2021). Impact of early genomic prediction for recurrent selection in an upland rice synthetic population (J Holland, Ed.). G3 GenesGenomesGenetics 11: jkab320. Bandillo NB, Jarquin D, Posadas LG, Lorenz AJ, Graef GL (2023). Genomic selection performs as effectively as phenotypic selection for increasing seed yield in soybean. Plant Genome 16: e20285. Barrett JC, Fry B, Maller J, Daly MJ (2005). Haploview: analysis and visualization of LD and haplotype maps. Bioinformatics 21: 263–265. Covarrubias-Pazaran G (2016). Genome-Assisted Prediction of Quantitative Traits Using the R Package sommer (A Zhang, Ed.). PLOS ONE 11: e0156744. Da Y (2015). Multi-allelic haplotype model based on genetic partition for genomic prediction and variance component estimation using SNP markers. BMC Genet 16: 144. De Los Campos G, Sorensen DA, Toro MA (2019). Imperfect Linkage Disequilibrium Generates Phantom Epistasis (& Perils of Big Data). G3 GenesGenomesGenetics 9: 1429–1436. Difabachew YF, Frisch M, Langstroff AL, Stahl A, Wittkop B, Snowdon RJ, et al. (2023). Genomic prediction with haplotype blocks in wheat. Front Plant Sci 14: 1168547. Gabriel SB, Schaffner SF, Nguyen H, Moore JM, Roy J, Blumenstiel B, et al. (2002). The Structure of Haplotype Blocks in the Human Genome. Science 296: 2225–2229. García-Barrios G, Crespo-Herrera L, Cruz-Izquierdo S, Vitale P, Sandoval-Islas JS, Gerard GS, et al. (2024). Genomic Prediction from Multi-Environment Trials of Wheat Breeding. Genes 15: 417. Habier D, Fernando RL, Dekkers JCM (2007). The Impact of Genetic Relationship Information on Genome-Assisted Breeding Values. Genetics 177: 2389–2397. Habier D, Tetens J, Seefried F-R, Lichtner P, Thaller G (2010). The impact of genetic relationship information on genomic breeding values in German Holstein cattle. Genet Sel Evol 42: 5. He S, Liang S, Meng L, Cao L, Ye G (2023). Sparse Phenotyping and Haplotype-Based Models for Genomic Prediction in Rice. Rice 16: 27. Hess M, Druet T, Hess A, Garrick D (2017). Fixed-length haplotypes can improve genomic prediction accuracy in an admixed dairy cattle population. Genet Sel Evol 49: 54. Hill WG, Goddard ME, Visscher PM (2008). Data and Theory Point to Mainly Additive Genetic Variance for Complex Traits (TFC Mackay, Ed.). PLoS Genet 4: e1000008. Jiang Y, Reif JC (2015). Modeling Epistasis in Genomic Selection. Genetics 201: 759–768. Jiang Y, Schmidt RH, Reif JC (2018). Haplotype-Based Genome-Wide Prediction Models Exploit Local Epistatic Interactions Among Markers. G3 GenesGenomesGenetics 8: 1687–1699. Karimi Z, Sargolzaei M, Robinson JAB, Schenkel FS (2018). Assessing haplotype-based models for genomic evaluation in Holstein cattle (J Plaizier, Ed.). Can J Anim Sci 98: 750–759. Kempthorne O (1955). THE THEORETICAL VALUES OF CORRELATIONS BETWEEN RELATIVES IN RANDOM MATING POPULATIONS. Genetics 40: 153–167. Krause MD, Piepho H-P, Dias KOG, Singh AK, Beavis WD (2023). Models to estimate genetic gain of soybean seed yield from annual multi-environment field trials. Theor Appl Genet 136: 252. Mackay TFC (2014). Epistasis and quantitative traits: using model organisms to study gene–gene interactions. Nat Rev Genet 15: 22–33. Meuwissen TH, Odegard J, Andersen-Ranberg I, Grindflek E (2014). On the distance of genetic relationships and the accuracy of genomic prediction in pig breeding. Genet Sel Evol 46: 49. Miller MJ, Song Q, Fallen B, Li Z (2023). Genomic prediction of optimal cross combinations to accelerate genetic improvement of soybean (Glycine max). Front Plant Sci 14. Pérez P, De Los Campos G (2014). Genome-Wide Regression and Prediction with the BGLR Statistical Package. Genetics 198: 483–495. Raffo MA, Sarup P, Guo X, Liu H, Andersen JR, Orabi J, et al. (2022). Improvement of genomic prediction in advanced wheat breeding lines by including additive-by-additive epistasis. Theor Appl Genet 135: 965–978. Sallam AH, Conley E, Prakapenka D, Da Y, Anderson JA (2020). Improving Prediction Accuracy Using Multi-allelic Haplotype Prediction and Training Population Optimization in Wheat. G3 GenesGenomesGenetics 10: 2265–2273. Silva JPAD, Viana JMS (2025). Efficiency of Genomic Selection for Developing Superior Pure Lines. Agronomy 15: 2247. Tessema BB, Liu H, Sørensen AC, Andersen JR, Jensen J (2020). Strategies Using Genomic Selection to Increase Genetic Gain in Breeding Programs for Wheat. Front Genet 11: 578123. VanRaden PM (2008). Efficient Methods to Compute Genomic Predictions. J Dairy Sci 91: 4414–4423. Viana JMS, Garcia AAF (2022). Significance of linkage disequilibrium and epistasis on genetic variances in noninbred and inbred populations. BMC Genomics 23: 286. Villumsen TM, Janss L, Lund MS (2009). The importance of haplotype length and heritability using genomic selection in dairy cattle. J Anim Breed Genet 126: 3–13. Wang N, Akey JM, Zhang K, Chakraborty R, Jin L (2002). Distribution of Recombination Crossovers and the Origin of Haplotype Blocks: The Interplay of Population History, Recombination, and Mutation. Am J Hum Genet 71: 1227–1234. Weber SE, Frisch M, Snowdon RJ, Voss-Fels KP (2023). Haplotype blocks for genomic prediction: a comparative evaluation in multiple crop datasets. Front Plant Sci 14: 1217589. Additional Declarations There is no duality of interest Cite Share Download PDF Status: Under Revision Version 1 posted Editorial decision: revise 13 Jan, 2026 Review # 3 received at journal 01 Dec, 2025 Review # 2 received at journal 03 Nov, 2025 Review # 1 received at journal 31 Oct, 2025 Reviewer # 3 agreed at journal 28 Oct, 2025 Reviewer # 2 agreed at journal 20 Oct, 2025 Reviewer # 1 agreed at journal 16 Oct, 2025 Reviewers invited by journal 15 Oct, 2025 First submitted to journal 29 Sep, 2025 Editor assigned by journal 29 Sep, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7745171","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":529949972,"identity":"fc635215-12cb-4736-935b-c4d5257a2dd7","order_by":0,"name":"José Marcelo Viana","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAs0lEQVRIiWNgGAWjYBACxgYgkQDE/AyMDyA8orVINjAbEKcFDgwOEKuFedrxZxIPau7IG99IZpNg3HGPCIfNzjGTSDj2zHAbWMuZYqK0MBsksB1m3HYj/7ABY1sCMVrSHxsk/Dtsv3lGMjOxWhIMHyS2HU7cIJHM+IBILTlALX2Hk2ececz4IPEMEVoMZ6c/OPjj22Hb/vZkhgMfdxCjpQGZR4QGBgZ5YhSNglEwCkbBCAcAYdY9e9Jl9T0AAAAASUVORK5CYII=","orcid":"https://orcid.org/0000-0002-5063-4648","institution":"Federal University of Viçosa","correspondingAuthor":true,"prefix":"","firstName":"José","middleName":"Marcelo","lastName":"Viana","suffix":""},{"id":529949973,"identity":"979b58bb-a92c-4e3c-a0bb-a159ad7ddd05","order_by":1,"name":"Jean Paulo Silva","email":"","orcid":"","institution":"Federal University of Viçosa","correspondingAuthor":false,"prefix":"","firstName":"Jean","middleName":"Paulo","lastName":"Silva","suffix":""}],"badges":[],"createdAt":"2025-09-29 20:35:44","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7745171/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7745171/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":94742672,"identity":"5f450fb3-aef2-47eb-b468-867b0ae293ea","added_by":"auto","created_at":"2025-10-30 09:00:45","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":251247,"visible":true,"origin":"","legend":"\u003cp\u003eFrequency distribution of haplotypes across 10 chromosomes, for one simulation and F\u003csub\u003e2\u003c/sub\u003e generation, using Gabriel et al. (2002), four-gamete rule, and solid spine method.\u003c/p\u003e","description":"","filename":"Figure1H.tif.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7745171/v1/1c4824c275094ec4b774dcf6.jpg"},{"id":94742674,"identity":"707c7186-423a-4839-8b75-949233a4b093","added_by":"auto","created_at":"2025-10-30 09:00:45","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":337208,"visible":true,"origin":"","legend":"\u003cp\u003eAverage grain yield (kg ha\u003csup\u003e-1\u003c/sup\u003e) over generations for haplotype- and SNP-based genomic selection using only additive (A) and additive plus additive x additive (A+AA) values, with haplotype sizes (HS) 3, 6, and 9, selection based on the true genotypic value (GV), and no selection (NS), assuming complementary (b), duplicate (c), dominant (d), recessive (e), and dominant and recessive (f) epistasis, duplicate genes with cumulative effects (g), non-epistatic gene interaction (h), and an admixture of types (a). The standard deviations for the means ranged from 2.0 to 47.9 kg ha\u003csup\u003e-1\u003c/sup\u003e, proportional to generation.\u003c/p\u003e","description":"","filename":"Figure2H.tif.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7745171/v1/de3666bfea6ac439b16f3975.jpg"},{"id":94742676,"identity":"24e9077a-bee3-40b0-a65e-46d91022d3b9","added_by":"auto","created_at":"2025-10-30 09:00:45","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":329218,"visible":true,"origin":"","legend":"\u003cp\u003eCorrelations between the number of favorable genes (N) and the additive (A) and additive x additive (AA) values, from F\u003csub\u003e2\u003c/sub\u003e to F\u003csub\u003e8\u003c/sub\u003e, assuming seven epistasis and an admixture of types (Co = complementary, Du = duplicate, Do = dominant, Re = recessive, DR = dominant and recessive, Dg = duplicate genes with cumulative effects, Ne = non-epistatic gene interaction, and All = all types).\u003c/p\u003e","description":"","filename":"Figure3H.tif.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7745171/v1/32075900187164d0ffe55142.jpg"},{"id":94742675,"identity":"6c85a3cc-9e9e-4dee-ba2b-4a39cfe09715","added_by":"auto","created_at":"2025-10-30 09:00:45","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":340048,"visible":true,"origin":"","legend":"\u003cp\u003eAverage genotypic variance ((kg ha\u003csup\u003e-1\u003c/sup\u003e)\u003csup\u003e2\u003c/sup\u003e) for grain yield over generations for haplotype- and SNP-based genomic selection using only additive (A) and additive plus additive x additive (A+AA) values, with haplotype sizes (HS) 3, 6, and 9, selection based on the true genotypic value (GV), and no selection (NS), assuming complementary (b), duplicate (c), dominant (d), recessive (e), and dominant and recessive (f) epistasis, duplicate genes with cumulative effects (g), non-epistatic gene interaction (h), and an admixture of types (a). The standard deviations for the variances ranged from 24.3 to 4340.0 (kg ha\u003csup\u003e-1\u003c/sup\u003e)\u003csup\u003e2\u003c/sup\u003e, inversely proportional to generation.\u003c/p\u003e","description":"","filename":"Figure4H.tif.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7745171/v1/e5722e8d93b55819695ae924.jpg"},{"id":94742677,"identity":"e176f83e-aa97-4a74-94ee-11c914c38d7a","added_by":"auto","created_at":"2025-10-30 09:00:45","extension":"jpg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":349244,"visible":true,"origin":"","legend":"\u003cp\u003eAverage prediction accuracy for grain yield over generations for haplotype - and SNP-based genomic selection using only additive (A) and additive plus additive x additive (A+AA) values, with haplotype sizes (HS) 3, 6, and 9, assuming complementary (b), duplicate (c), dominant (d), recessive (e), and dominant and recessive (f) epistasis, duplicate genes with cumulative effects (g), non-epistatic gene interaction (h), and an admixture of types (a). The standard deviations for the accuracies ranged from 0.00 to 0.17, proportional to generation. P indicates the accuracy of the phenotypic selection.\u003c/p\u003e","description":"","filename":"Figure5H.tif.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7745171/v1/d0ecd4de87eb46193ff41376.jpg"},{"id":94823818,"identity":"c807290b-401c-4f21-a32f-092edd2f4b25","added_by":"auto","created_at":"2025-10-31 06:48:06","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1963032,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7745171/v1/56bf7cc2-7190-4f18-a581-c7b597e81b6b.pdf"}],"financialInterests":"There is no duality of interest","formattedTitle":"\u003cp\u003eEfficiency Under Epistasis of Haplotype-Based Genomic Selection for Pure Line Breeding\u003c/p\u003e","fulltext":[{"header":"Introduction","content":"\u003cp\u003eGenomic selection has been successfully employed for the improvement of self-pollinated crops such as soybean, wheat, and rice, primarily through SNP-based approach (Tessema \u003cem\u003eet al.\u003c/em\u003e, 2020; Baertschi \u003cem\u003eet al.\u003c/em\u003e, 2021; Bandillo \u003cem\u003eet al.\u003c/em\u003e, 2023). While single nucleotide polymorphisms (SNPs) are widely used, this approach presents limitations. SNPs don\u0026apos;t always reveal all variations and allelic combinations of genes controlling a trait, and their effectiveness often decreases with lower relatedness between training and validation populations (Villumsen \u003cem\u003eet al.\u003c/em\u003e, 2009; Habier \u003cem\u003eet al.\u003c/em\u003e, 2010). Linkage disequilibrium (LD) between markers was initially believed to be the primary source of information for genomic prediction. However, Habier et al. (2007) showed that SNP also capture genetic relationships among individuals. This suggests that the accuracy of SNP-based predictions may depend more on capturing relationships among individuals than on direct LD with causal QTLs (Habier \u003cem\u003eet al.\u003c/em\u003e, 2007; Weber \u003cem\u003eet al.\u003c/em\u003e, 2023). In contrast, haplotype-based approach offers a promising alternative. They can improve predictive accuracy by potentially increasing LD with causal loci and capturing complex interactions among nearby genes (Meuwissen \u003cem\u003eet al.\u003c/em\u003e, 2014; Hess \u003cem\u003eet al.\u003c/em\u003e, 2017; Jiang \u003cem\u003eet al.\u003c/em\u003e, 2018).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eWeber et al. (2023) evaluated the accuracy of haplotype-based genomic prediction, reporting inconsistent results across traits, species, and models. Processing a wheat dataset, the higher increases in prediction accuracy from haplotyping, relative to the SNP-based approach, were observed for grain yield and sedimentation value. The accuracies increased from 0.70 to 0.82, from 0.49 to 0.62, and from 0.49 to 0.64, depending on the method of haplotyping and statistical model. Using the soybean, maize, and canola datasets, no differences were observed between haplotype- and SNP-based approaches. Sallam et al. (2020) observed more consistent benefits from haplotype-based genomic predictions. Their investigation evaluated haplotypes of varying sizes under two cross-validation schemes. All haplotype sizes outperformed SNP-based for grain yield and protein content, with the haplotype size 15 providing the greatest prediction accuracy gains. Notably, stratified sampling enhanced predictive performance the most, achieving up to 14.3% improvement for grain yield and 16.8% for protein content. These studies clearly demonstrate that while haplotype-based genomic prediction holds promise, its success is highly contingent upon multiple interacting factors, such as model, trait architecture, species, population, haplotyping method, and training set.\u003c/p\u003e\n\u003cp\u003eThe discussion about the best way to model genetic variation extends to the inclusion of epistatic effects in genomic prediction models. However, the literature presents contrasting results regarding the impact of including these interactions. It is noteworthy that, of the limited number of studies including epistasis, the vast majority have solely focused on models accounting for additive x additive effects. In Weber et al. (2023), incorporating additive x additive effects into haplotype-based genomic prediction resulted in only marginal improvements in prediction accuracy for most traits. For example, in wheat, grain yield prediction and its epistatic extension showed a negligible increase, from 0.80 to 0.81. In soybean, the same model improved prediction accuracy for oil content from 0.68 to 0.69, with no increase for protein content. Miller et al. (2023), evaluating parent selection strategies in soybean breeding, reported that including additive x additive effects led to more than a 50% increase in prediction accuracy, depending on the selection method. Garc\u0026iacute;a-Barrios et al. (2024), studying multi-environment genomic prediction in wheat, found no significant improvement when incorporating additive x additive interactions, compared to purely additive model.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eQuantifying the impact of epistasis on selection has long been a challenge in plant and animal breeding. With the advancement of genomic data, there has been renewed interest in modeling epistatic interactions; however, assessing the significance of epistasis in empirical datasets remains difficult, since there is no previous knowledge if there is epistasis. In this simulation-based study, we evaluated the efficacy of haplotype-based genomic selection under seven digenic epistasis types and the significance of including additive x additive effect in selection, in pure line breeding. To our knowledge, this is the first simulation-based investigation on the efficiency of haplotype-based genomic selection under epistasis in pure line breeding.\u0026nbsp;\u003c/p\u003e"},{"header":"Materials and Methods","content":"\u003cp\u003eUsing the software \u003cem\u003eREALbreeding\u0026nbsp;\u003c/em\u003e(available by request), we simulated a genome comprising 10 chromosomes of 100 cM, including 1000 minor genes, and 49825 SNPs. The average densities for genes and SNPs were 1 and 0.02 cM, respectively. The simulated trait was grain yield in soybean, assuming a planting density of 300000 plants per hectare. As an input for the software, we defined, assuming no epistasis, the minimum genotypic and phenotypic values as 7 and 4 g/plant, respectively, while the maximum values were 16 and 20 g/plant, respectively. The degree of dominance ranged from 0 to 1.2, with a mean of 0.6. Broad-sense heritability was set at 30% in each generation.\u003c/p\u003e\n\u003cp\u003eWe assumed digenic epistasis for 30% of the genes, resulting in 150 epistatic pairs. The following epistasis types were modeled based on Kempthorne (1955): complementary (9:7 in F₂), duplicate (15:1 in F₂), dominant (12:3:1 in F₂), recessive (9:3:4 in F₂), dominant and recessive (13:3 in F₂), duplicate genes with cumulative effects (9:6:1 in F₂), non-epistatic gene interaction (9:3:3:4 in F₂), and an admixture of types. For the admixture, \u003cem\u003eREALbreeding\u003c/em\u003e randomly assigns an epistasis type to each pair of interacting genes. For each epistasis, we kept the same pairs of epistatic genes. The base population consisted of 250 F₂ individuals derived from a cross between two contrasting pure lines, one carrying 70% and the other 30% of the favorable alleles. For details on the simulation procedures and theoretical background, refer to\u0026nbsp;Viana \u0026amp; Garcia (2022).\u003c/p\u003e\n\u003cp\u003eFrom F\u003csub\u003e3\u003c/sub\u003e to F\u003csub\u003e8\u003c/sub\u003e generations, we performed haplotype-based genomic selection based on predicted sum of the additive and additive x additive values, using haplotype sizes 3, 6, and 9.\u0026nbsp;The training set included all F\u003csub\u003e2\u0026nbsp;\u003c/sub\u003eindividuals and 18% of the F\u003csub\u003e3\u003c/sub\u003e to F\u003csub\u003e7\u0026nbsp;\u003c/sub\u003eindividuals.\u0026nbsp;Each progeny consisted of 20 plants. The F\u003csub\u003e3\u003c/sub\u003e consisted of 250 progenies (5000 plants). The number of progenies from F\u003csub\u003e4\u003c/sub\u003e to F\u003csub\u003e8\u003c/sub\u003e were 150 (3000 plants), 50 (1000 plants), 20 (400 plants), and 5 (100 plants), respectively. The proportion of selected plants were 3.0, 5.0, 1.7, 2.5, and 1.2%. As reference scenarios, we also evaluated SNP-based selection, haplotype-based selection based only in the predicted additive value, no selection, and selection based on true genotypic value. Under no selection, we advanced 500 individuals in each generation using the single seed descent (SSD) method. For selection based on the true genotypic values, we used the values provided by \u003cem\u003eREALbreeding\u003c/em\u003e. We replicated each scenario 10 times.\u003c/p\u003e\n\u003cp\u003eWe fitted the model \u0026nbsp;, where the \u0026nbsp; are the incidence matrices associated with the vectors of additive, dominance, additive x additive (AxA), additive x dominance (AxD), and dominance x dominance (DxD) genetic values, respectively. The model was implemented using the BGLR R package\u0026nbsp;(P\u0026eacute;rez and De Los Campos, 2014). We used the defaults for burn in (500) and iterations (2000), checking for convergence. The additive and dominance genomic matrices were computed following the methods proposed by\u0026nbsp;VanRaden (2008)\u0026nbsp;(first method) and Su et al. (2012), respectively, using the sommer R package\u0026nbsp;(Covarrubias-Pazaran, 2016). We applied the Hadamard product to the additive and dominance relationship matrices to generate the epistatic genomic matrices. The AxD genetic value corresponds to the sum of AxD and dominance x additive (DxA) values.\u003c/p\u003e\n\u003cp\u003eWe built haplotypes based on a fixed number of adjacent SNPs along the chromosome, employing our own R code. In this study, we defined haplotype lengths of 3, 6, and 9, with length 9 applied only to an admixture of types. To evaluate the haplotype pattern in the F\u003csub\u003e2\u0026nbsp;\u003c/sub\u003epopulation, we used Haploview (Barrett \u003cem\u003eet al.\u003c/em\u003e, 2005). Haplotypes were defined using three distinct methods: Gabriel et al. (2002), the four gamete rule (Wang et al. (2002), and the solid spine of LD method. Haplotyping requires phased SNP genotypes, which we conveniently obtained from the \u003cem\u003eREALbreeding\u003c/em\u003e\u0026rsquo;s output, where the left allele is of maternal origin. We adopted the haplotype model proposed by Da (2015), in which each haplotype allele is treated as a SNP - a procedure referred to as \u0026lsquo;pseudo-SNP\u0026rsquo; by Karimi et al. (2018). With a haplotype size \u0026nbsp;, the total number haplotype alleles is\u0026nbsp;.\u003c/p\u003e\n\u003cp\u003eWe calculated the genetic gains by the difference between the genotypic means of successive generations. We estimated the additive value prediction accuracy by Pearson\u0026rsquo;s correlation between the predicted additive values and the true genotypic values. Because of selection, \u003cem\u003eREALbreeding\u003c/em\u003e does not compute the additive values. We computed the genotypic variance from the true genotypic values provided by \u003cem\u003eREALbreeding\u003c/em\u003e. For testing differences between total genetic gains, we applied regression analysis or t test assuming distinct variances. Aiming to compute the correlations between the number of favorable genes and the additive and additive plus AxA values, we used the genotypes for genes and the true genetic values, assuming no selection from generations F\u003csub\u003e2\u003c/sub\u003e to F\u003csub\u003e8\u003c/sub\u003e, provided by \u003cem\u003eREALbreeding\u003c/em\u003e.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003eThe analysis of the haplotype pattern in the F\u003csub\u003e2\u003c/sub\u003e generation (Figure 1) shows, irrespective of the method, a predominance of long haplotypes (average of 22.2 SNPs). However, only 4.5% of the SNPs were included in a haplotype. The ratio F\u003csub\u003e2\u003c/sub\u003e epistatic variance/genotypic variance ranged from 1.1, assuming complementary epistasis, to 13.9%, assuming duplicate epistasis. The additive variance ranged from 41.8 to 141.7% (due to negative covariances) of the genotypic variance in F\u003csub\u003e2\u003c/sub\u003e, while the AxA genotypic variance ranged from 0.5 to 6.1%.\u003c/p\u003e\n\u003cp\u003eFor the admixture of epistasis types, epistatic and AxA variances accounted for 6.4 and 2.0% of the genotypic variance, respectively. The total genetic gains from the haplotype-based genomic selection, based on predicted A + AxA values were 208.4, 212.0, and 218.1 kg ha\u003csup\u003e-1\u003c/sup\u003e, for haplotypes sizes 3, 6, and 9, respectively (Figure 2). However, the linear and quadratic regression analyses evidenced no haplotype size effect (R\u003csup\u003e2\u003c/sup\u003e = 2.3 and 2.2%, respectively). Thus, the haplotype-based genomic selection provided an average total genetic gain of 212.8 kg ha\u003csup\u003e-1\u003c/sup\u003e. The SNP-based genomic selection using A + AxA value provided a total genetic gain of 211.7 kg ha\u003csup\u003e-1\u003c/sup\u003e, showing that both procedures are equivalent (Figure 2). Fitting the complete model but ignoring AxA effects in selection, the total genetic gains with haplotype- and SNP-based analyses were 229.5, 204.1, 204.1, for haplotype sizes 3, 6, and 9, respectively, and 203.0 kg ha\u003csup\u003e-1\u003c/sup\u003e, respectively. Even for haplotype size 3, we observed non-significant changes in the selection efficacy by ignoring the predicted AxA effect in selection (-3.9 to 10.1%).\u003c/p\u003e\n\u003cp\u003eFor the seven distinct epistasis types we observed essentially the same results: no significant haplotype size effect, no significant difference between haplotype- and SNP-based approaches, and no significant effect by selecting based only in the predicted additive effect (Figure 2). The only exceptions occurred for complementary and duplicate epistasis, where significant decreases in the total genetic gains occurred by increasing the haplotype size from 3 to 6 (P values of 0.04 and 0.005). Comparing SNP with the best haplotype size, there was no significant differences between the total genetic gains. Regarding the impact of ignoring the AxA effect on selection, we observed no significant changes in the total genetic gains. The changes ranged from\u0026nbsp;-14.8 to 17.1 kg ha\u003csup\u003e-1\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003eThe epistasis that minimized and maximized the total genetic gains were, respectively, duplicate, from genomic selection based on haplotype size 6 and predicted A + AxA value (163.6 kg ha\u003csup\u003e-1\u003c/sup\u003e), and non-epistatic gene interaction, from selection based on haplotype of size 6 and predicted additive value (271.8 kg ha\u003csup\u003e-1\u003c/sup\u003e), a statistically significant difference (Figure 2). The correlations between the ratio epistatic variance/genotypic variance with the minimum, average, and maximum total genetic gains were negative of intermediate magnitude (-0.42,\u0026nbsp;-0.31, and\u0026nbsp;-0.23). The scenario of no selection clearly shows how all genomic selection methods efficiently increased the means of the generations. In case of no selection and due to positive dominance, the decreases in the F\u003csub\u003e8\u003c/sub\u003e generation, relative to F\u003csub\u003e3\u003c/sub\u003e, ranged from\u0026nbsp;-122.7, under duplicate epistasis, to\u0026nbsp;-232.2 kg ha\u003csup\u003e-1\u003c/sup\u003e, assuming non-epistatic gene interaction, also inversely proportional to the ratio epistatic variance/genotypic variance. Regarding the scenario of selection based on the true genotypic value, the minimum and maximum total genetic gains were 189.4, under recessive epistasis, and 256.0 kg ha\u003csup\u003e-1\u003c/sup\u003e, assuming non-epistatic gene interaction. Again, the total genetic gain was inversely proportional to the ratio epistatic variance/genotypic variance (correlation of\u0026nbsp;-0.24). The difference between these two gains is not statistically significant.\u003c/p\u003e\n\u003cp\u003eThe correlations of the additive and additive x additive values with the number of favorable genes evidenced that genomic selection based on the additive value was expected to be more effective than selection based on the sum A + AA, because predominantly negative values for the correlation between AxA and the number of favorable genes (Figure 3), irrespective of the epistasis type. Based on the magnitude of the correlations, we can state that the additive x additive genetic values showed no significant correlation with the number of favorable genes. The correlation between the additive value and the number of favorable genes ranged from approximately 0.6 under duplicate epistasis to approximately 0.8 assuming duplicate genes with cumulative effects.\u003c/p\u003e\n\u003cp\u003eAs expected, all selection procedures led to a significant decrease in the genotypic variance, irrespective of the epistasis type (99.0% on average) (Figure 4). Excepting dominant epistasis, based on the decrease in the genotypic variance under no selection, due to inbreeding, selection was responsible for 55.0 to 92.0% of the observed decrease. Selection based on the true genotypic value led to comparable decreases, from 52.0 to 92.0%. Regarding the prediction accuracy, the estimated values show clearly how haplotype - and SNP-based genomic selection are effective in pure line breeding, relative to phenotypic selection (full phenotyping), irrespective of the epistasis type (Figure 5). The accuracies ranged from 0.70 to 0.90, regardless of epistasis and selection process, decreasing from F\u003csub\u003e3\u003c/sub\u003e to F\u003csub\u003e7\u003c/sub\u003e. Another important result was the high positive correlation between genetic gains and prediction accuracies. Regardless of the method and epistasis, the correlations ranged from 0.87 to 0.99.\u003c/p\u003e\n\u003cp\u003eComparing our results for SNP-based genomic selection, including epistasis, with the results assuming no epistasis but the same genome, number and positions of genes and SNPs, parents, degree of dominance etc. (Silva and Viana, 2025), we additionally observed that epistasis increased or decreased the total genetic gains in pure line breeding, depending on the epistasis type. The change in the total genetic gains by including and modelling epistasis ranged from\u0026nbsp;-22.2 to 3.6% (-9.3% on average). Finally, it is important to emphasize that, under epistasis, all genomic selection processes provided F\u003csub\u003e8\u003c/sub\u003e progenies statistically equal to the superior parent. The differences ranged from -1.0 to 2.0%, the same result observed assuming additive-dominance model.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eBased on several inheritance studies for qualitative traits (see any standard Genetics book), hundreds of inheritance studies for quantitative traits, using generation mean analysis, QTL mapping, and genome-wide studies, and many transcriptome, proteome, and metabolome investigations, geneticists agree that epistasis is the rule and absence of epistasis is the exception for complex traits (Mackay, 2014). Why, then, quantifying and assessing the significance of epistasis in evolution and selection remains a challenge for geneticists and breeders, even with the help of SNPs? Because, in general, epistatic genetic values have a relative lower significance in determining genotypic value, relative to the additive value. Or better, because the additive variance is the most significant component of the genotypic variance (Hill \u003cem\u003eet al.\u003c/em\u003e, 2008). Furthermore, epistatic effects can be difficult to quantify and falsely declared under incomplete LD at low SNP density (De Los Campos \u003cem\u003eet al.\u003c/em\u003e, 2019).\u003c/p\u003e\n\u003cp\u003eIn this first simulation-based investigation on the efficacy under epistasis of haplotype-based genomic selection in pure line breeding, we provided clear evidence that a genomic analysis can be effective to quantify epistatic effects, if they have a significant contribution in determining the quantitative trait, and to provide substantial genetic gains. Additionally, we showed that the heritable additive x additive effect can contribute for increasing selection efficacy, depending on the correlation between the number of favorable genes and the AxA effect. In our study, the correlations under no selection were negative of low magnitude, regardless of the epistasis type and the ratio epistatic variance/genotypic variance. Our evidence of no significance for selection of the additive x additive epistatic values was observed by Garc\u0026iacute;a-Barrios et al. (2024) and Weber et al. (2023). Although there are inconclusive results for the efficacy of haplotype- and SNP-based genomic selection, haplotyping presents some expected advantages. Among them, we emphasize that haplotypes can better capture epistatic effects (Jiang \u003cem\u003eet al.\u003c/em\u003e, 2018). Applying haplotype-based genomic selection from F\u003csub\u003e3\u003c/sub\u003e to F\u003csub\u003e7\u003c/sub\u003e, the total genetic gains ranged from 188 to 251 kg ha\u003csup\u003e-1\u003c/sup\u003e, depending on the epistasis type and the ratio epistatic variance/genotypic variance. This means 38 to 50 kg ha\u003csup\u003e-1\u0026nbsp;\u003c/sup\u003eyear\u003csup\u003e-1\u003c/sup\u003e. The correlation between the gain and the ratio was\u0026nbsp;-0.3.\u0026nbsp;Krause et al. (2023)\u0026nbsp;estimated in 18 to 40 kg ha\u003csup\u003e-1\u0026nbsp;\u003c/sup\u003eyear\u003csup\u003e-1\u003c/sup\u003e the realized genetic gains in soybean public programs.\u003c/p\u003e\n\u003cp\u003eFrom the assessment of the haplotype size effect, we observed that increasing the haplotype size from 3 to 9, under an admixture of epistasis types, and from 3 to 6 for the seven distinct epistasis, did not increase the genomic selection efficacy. Unfortunately, there is no previous study with self-pollinated crops assessing the haplotype size effect based on realized genetic gains. Only based on breeding value prediction accuracy. Sallam et al. (2020) also did not observe haplotype size effect. The prediction accuracies for haplotype sizes 5, 10, 15, and 20 oscillated from 0.31 to 0.34. Regarding the significance of the predicted AxA effect, because predominantly negative of low magnitude correlations with the number of favorable genes, under no selection, selection based on the additive values was expected to be only slightly efficient than selection based on the additive plus AxA value. And the two processes were equally efficient. The ratio between the realized total genetic gains ranged from 0.9 to 1.2 (1.0 on average). Garc\u0026iacute;a-Barrios et al. (2024) evaluated multi-environment prediction models and datasets for wheat and found similar performance between purely additive model and models incorporating additive x additive effects. For wheat also, Jiang and Reif (2015) observed slightly improved prediction accuracies when additive x additive effects were included. Raffo et al. (2022) concluded that epistasis can enhance predictive ability, especially under conditions of high relatedness. However, the difference between the prediction accuracies was only 0.06. Weber et al. (2023), one of the few studies in plant breeding to evaluate the impact of including epistatic effects in haplotype-based models, found only marginal improvements in predictive accuracy for grain yield in wheat and for protein and oil contents in soybean, regardless of whether haplotypes or SNPs were used.\u003c/p\u003e\n\u003cp\u003eComparing haplotype- and SNP-based approaches, we observed that both methods were equivalent, considering total realized genetic gain, additive value prediction accuracy, and decrease in the genotypic variance, irrespective of the epistasis and ratio epistatic variance/genotypic variance. The SNP-based approach provided genetic gains from 184 to 247 kg ha\u003csup\u003e-1\u003c/sup\u003e, depending on the epistasis type and the ratio epistatic variance/genotypic variance. Using a wheat inbred line panel, Difabachew et al. (2023) showed that predictive performance varied depending on haplotype definition and trait. Haplotypes defined by LD provided higher accuracy for disease resistance, while fixed-length haplotyping performed better for plant height. However, all haplotype-based approaches were less effective for grain yield. Positive results favoring haplotype-based models are often associated with low marker density. This trend is consistent with findings from Villumsen et al. (2009), who showed that haplotype-based prediction tends to outperform SNP-based models particularly when marker density is low. Regardless of the statistical method, inbred line panel, missing data rate, and trait, He et al. (2023) also observed equivalence in prediction accuracy for haplotypes and SNPs in rice.\u003c/p\u003e"},{"header":"Conclusions","content":"\u003cp\u003eRegardless of the epistasis type, genomic selection is an efficient procedure to develop superior pure lines. There was no haplotype size effect on the genomic selection efficacy. Under no selection, the AxA genetic value showed no significant negative correlation with the number of favorable genes. Under selection, the inclusion of the predicted AxA value for selection did not increase the selection efficacy. Haplotype- and SNP-based genomic selection were equally efficient. Epistasis can decrease the total genetic gain in pure line breeding, compared to traits determined only by additive and dominance gene effects. Under epistasis, there was a highly positive correlation between genetic gain and prediction accuracy. There is no difference between genomic selection procedures regarding the decrease in the genotypic variance.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eAcknowledgments\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting Interests\u0026nbsp;\u003c/strong\u003eThe authors declare no conflict of interest.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor Contributions\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData Archiving\u0026nbsp;\u003c/strong\u003eThe dataset is available at https://doi.org/10.6084/m9.figshare.29825855.v1.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eBaertschi C, Cao T-V, Bartholom\u0026eacute; J, Ospina Y, Quintero C, Frouin J, et al. (2021). Impact of early genomic prediction for recurrent selection in an upland rice synthetic population (J Holland, Ed.). G3 GenesGenomesGenetics 11: jkab320.\u003c/li\u003e\n\u003cli\u003eBandillo NB, Jarquin D, Posadas LG, Lorenz AJ, Graef GL (2023). Genomic selection performs as effectively as phenotypic selection for increasing seed yield in soybean. Plant Genome 16: e20285.\u003c/li\u003e\n\u003cli\u003eBarrett JC, Fry B, Maller J, Daly MJ (2005). Haploview: analysis and visualization of LD and haplotype maps. Bioinformatics 21: 263\u0026ndash;265.\u003c/li\u003e\n\u003cli\u003eCovarrubias-Pazaran G (2016). Genome-Assisted Prediction of Quantitative Traits Using the R Package sommer (A Zhang, Ed.). PLOS ONE 11: e0156744.\u003c/li\u003e\n\u003cli\u003eDa Y (2015). Multi-allelic haplotype model based on genetic partition for genomic prediction and variance component estimation using SNP markers. BMC Genet 16: 144.\u003c/li\u003e\n\u003cli\u003eDe Los Campos G, Sorensen DA, Toro MA (2019). Imperfect Linkage Disequilibrium Generates Phantom Epistasis (\u0026amp; Perils of Big Data). G3 GenesGenomesGenetics 9: 1429\u0026ndash;1436.\u003c/li\u003e\n\u003cli\u003eDifabachew YF, Frisch M, Langstroff AL, Stahl A, Wittkop B, Snowdon RJ, et al. (2023). Genomic prediction with haplotype blocks in wheat. Front Plant Sci 14: 1168547.\u003c/li\u003e\n\u003cli\u003eGabriel SB, Schaffner SF, Nguyen H, Moore JM, Roy J, Blumenstiel B, et al. (2002). The Structure of Haplotype Blocks in the Human Genome. Science 296: 2225\u0026ndash;2229.\u003c/li\u003e\n\u003cli\u003eGarc\u0026iacute;a-Barrios G, Crespo-Herrera L, Cruz-Izquierdo S, Vitale P, Sandoval-Islas JS, Gerard GS, et al. (2024). Genomic Prediction from Multi-Environment Trials of Wheat Breeding. Genes 15: 417.\u003c/li\u003e\n\u003cli\u003eHabier D, Fernando RL, Dekkers JCM (2007). The Impact of Genetic Relationship Information on Genome-Assisted Breeding Values. Genetics 177: 2389\u0026ndash;2397.\u003c/li\u003e\n\u003cli\u003eHabier D, Tetens J, Seefried F-R, Lichtner P, Thaller G (2010). The impact of genetic relationship information on genomic breeding values in German Holstein cattle. Genet Sel Evol 42: 5.\u003c/li\u003e\n\u003cli\u003eHe S, Liang S, Meng L, Cao L, Ye G (2023). Sparse Phenotyping and Haplotype-Based Models for Genomic Prediction in Rice. Rice 16: 27.\u003c/li\u003e\n\u003cli\u003eHess M, Druet T, Hess A, Garrick D (2017). Fixed-length haplotypes can improve genomic prediction accuracy in an admixed dairy cattle population. Genet Sel Evol 49: 54.\u003c/li\u003e\n\u003cli\u003eHill WG, Goddard ME, Visscher PM (2008). Data and Theory Point to Mainly Additive Genetic Variance for Complex Traits (TFC Mackay, Ed.). PLoS Genet 4: e1000008.\u003c/li\u003e\n\u003cli\u003eJiang Y, Reif JC (2015). Modeling Epistasis in Genomic Selection. Genetics 201: 759\u0026ndash;768.\u003c/li\u003e\n\u003cli\u003eJiang Y, Schmidt RH, Reif JC (2018). Haplotype-Based Genome-Wide Prediction Models Exploit Local Epistatic Interactions Among Markers. G3 GenesGenomesGenetics 8: 1687\u0026ndash;1699.\u003c/li\u003e\n\u003cli\u003eKarimi Z, Sargolzaei M, Robinson JAB, Schenkel FS (2018). Assessing haplotype-based models for genomic evaluation in Holstein cattle (J Plaizier, Ed.). Can J Anim Sci 98: 750\u0026ndash;759.\u003c/li\u003e\n\u003cli\u003eKempthorne O (1955). THE THEORETICAL VALUES OF CORRELATIONS BETWEEN RELATIVES IN RANDOM MATING POPULATIONS. Genetics 40: 153\u0026ndash;167.\u003c/li\u003e\n\u003cli\u003eKrause MD, Piepho H-P, Dias KOG, Singh AK, Beavis WD (2023). Models to estimate genetic gain of soybean seed yield from annual multi-environment field trials. Theor Appl Genet 136: 252.\u003c/li\u003e\n\u003cli\u003eMackay TFC (2014). Epistasis and quantitative traits: using model organisms to study gene\u0026ndash;gene interactions. Nat Rev Genet 15: 22\u0026ndash;33.\u003c/li\u003e\n\u003cli\u003eMeuwissen TH, Odegard J, Andersen-Ranberg I, Grindflek E (2014). On the distance of genetic relationships and the accuracy of genomic prediction in pig breeding. Genet Sel Evol 46: 49.\u003c/li\u003e\n\u003cli\u003eMiller MJ, Song Q, Fallen B, Li Z (2023). Genomic prediction of optimal cross combinations to accelerate genetic improvement of soybean (Glycine max). Front Plant Sci 14.\u003c/li\u003e\n\u003cli\u003eP\u0026eacute;rez P, De Los Campos G (2014). Genome-Wide Regression and Prediction with the BGLR Statistical Package. Genetics 198: 483\u0026ndash;495.\u003c/li\u003e\n\u003cli\u003eRaffo MA, Sarup P, Guo X, Liu H, Andersen JR, Orabi J, et al. (2022). Improvement of genomic prediction in advanced wheat breeding lines by including additive-by-additive epistasis. Theor Appl Genet 135: 965\u0026ndash;978.\u003c/li\u003e\n\u003cli\u003eSallam AH, Conley E, Prakapenka D, Da Y, Anderson JA (2020). Improving Prediction Accuracy Using Multi-allelic Haplotype Prediction and Training Population Optimization in Wheat. G3 GenesGenomesGenetics 10: 2265\u0026ndash;2273.\u003c/li\u003e\n\u003cli\u003eSilva JPAD, Viana JMS (2025). Efficiency of Genomic Selection for Developing Superior Pure Lines. Agronomy 15: 2247.\u003c/li\u003e\n\u003cli\u003eTessema BB, Liu H, S\u0026oslash;rensen AC, Andersen JR, Jensen J (2020). Strategies Using Genomic Selection to Increase Genetic Gain in Breeding Programs for Wheat. Front Genet 11: 578123.\u003c/li\u003e\n\u003cli\u003eVanRaden PM (2008). Efficient Methods to Compute Genomic Predictions. J Dairy Sci 91: 4414\u0026ndash;4423.\u003c/li\u003e\n\u003cli\u003eViana JMS, Garcia AAF (2022). Significance of linkage disequilibrium and epistasis on genetic variances in noninbred and inbred populations. BMC Genomics 23: 286.\u003c/li\u003e\n\u003cli\u003eVillumsen TM, Janss L, Lund MS (2009). The importance of haplotype length and heritability using genomic selection in dairy cattle. J Anim Breed Genet 126: 3\u0026ndash;13.\u003c/li\u003e\n\u003cli\u003eWang N, Akey JM, Zhang K, Chakraborty R, Jin L (2002). Distribution of Recombination Crossovers and the Origin of Haplotype Blocks: The Interplay of Population History, Recombination, and Mutation. Am J Hum Genet 71: 1227\u0026ndash;1234.\u003c/li\u003e\n\u003cli\u003eWeber SE, Frisch M, Snowdon RJ, Voss-Fels KP (2023). Haplotype blocks for genomic prediction: a comparative evaluation in multiple crop datasets. Front Plant Sci 14: 1217589.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"heredity","isNatureJournal":false,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"hdy","sideBox":"Learn more about [Heredity](http://www.nature.com/hdy/)","snPcode":"41437","submissionUrl":"https://mts-hdy.nature.com/cgi-bin/main.plex","title":"Heredity","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"Nature AJ","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"haplotype, genomic selection, pure lines, epistasis","lastPublishedDoi":"10.21203/rs.3.rs-7745171/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7745171/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"The haplotype-based approach is considered a more efficient alternative for genomic selection under epistasis. In this simulation-based study, we evaluated the efficacy of haplotype-based genomic selection under seven digenic epistasis types and the significance of including additive x additive (AxA) effect in selection, in pure line breeding. We simulated a genome comprised of 10 chromosomes of 100 cM, including 1000 genes and 49825 SNPs. We assumed positive dominance and 30% of epistatic genes. We performed haplotype-based genomic selection, using haplotypes of sizes 3, 6, and 9. The training set included all F2 individuals and 18% of the F3 to F7 individuals. Selection efficacy was assessed using realized total genetic gain. We also evaluated the correlation between prediction accuracy and realized genetic gain and the decrease in the genotypic variance. Regardless of the epistasis type, genomic selection is an efficient procedure to develop superior pure lines. There was no haplotype size effect on the genomic selection efficacy. Under no selection, the AxA genetic value showed no significant negative correlation with the number of favorable genes. Under selection, the inclusion of the predicted AxA value for selection did not increase the selection efficacy. Haplotype- and SNP-based genomic selection were equally efficient. Epistasis can decrease the total genetic gain in pure line breeding, compared to traits determined only by additive and dominance gene effects. Under epistasis, there was a highly positive correlation between genetic gain and prediction accuracy. There is no difference between genomic selection procedures regarding the decrease in the genotypic variance.","manuscriptTitle":"Efficiency Under Epistasis of Haplotype-Based Genomic Selection for Pure Line Breeding","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-10-30 09:00:41","doi":"10.21203/rs.3.rs-7745171/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"revise","date":"2026-01-13T14:00:14+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"This content is not available.","date":"2025-12-02T01:27:42+00:00","index":3,"fulltext":"This content is not available."},{"type":"editorInvitedReview","content":"This content is not available.","date":"2025-11-03T08:54:06+00:00","index":2,"fulltext":"This content is not available."},{"type":"editorInvitedReview","content":"This content is not available.","date":"2025-10-31T16:01:34+00:00","index":1,"fulltext":"This content is not available."},{"type":"reviewerAgreed","content":"This content is not available.","date":"2025-10-28T13:36:32+00:00","index":3,"fulltext":"This content is not available."},{"type":"reviewerAgreed","content":"This content is not available.","date":"2025-10-20T07:16:01+00:00","index":2,"fulltext":"This content is not available."},{"type":"reviewerAgreed","content":"This content is not available.","date":"2025-10-16T20:22:03+00:00","index":1,"fulltext":"This content is not available."},{"type":"reviewersInvited","content":"","date":"2025-10-15T09:15:07+00:00","index":"","fulltext":""},{"type":"submitted","content":"Heredity","date":"2025-09-29T20:34:55+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-09-29T20:34:55+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"heredity","isNatureJournal":false,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"hdy","sideBox":"Learn more about [Heredity](http://www.nature.com/hdy/)","snPcode":"41437","submissionUrl":"https://mts-hdy.nature.com/cgi-bin/main.plex","title":"Heredity","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"Nature AJ","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"553c83e9-d67f-48f2-b7d3-b28b225fef7f","owner":[],"postedDate":"October 30th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"in-revision","subjectAreas":[{"id":57146334,"name":"Biological sciences/Genetics/Plant breeding"},{"id":57146335,"name":"Biological sciences/Plant sciences/Plant breeding"}],"tags":[],"updatedAt":"2026-05-05T17:56:36+00:00","versionOfRecord":[],"versionCreatedAt":"2025-10-30 09:00:41","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7745171","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7745171","identity":"rs-7745171","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.