Genotype Imputation from Low-Coverage WGS Using Haplotype Reference Panels in Cultivated Strawberry | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Genotype Imputation from Low-Coverage WGS Using Haplotype Reference Panels in Cultivated Strawberry Tim Koorevaar, Johan H. Willemsen, Richard G.F. Visser, Paul Arens, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7241985/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 19 Nov, 2025 Read the published version in BMC Genomics → Version 1 posted 13 You are reading this latest preprint version Abstract Background To implement high-throughput sequencing-based genotyping in a strawberry ( Fragaria × ananassa ) breeding program, we aimed to construct a haplotype reference panel and explore its utility through genotype dosage imputation of low-coverage (1×) sequencing data. Although genotyping by whole genome sequencing (WGS) provides high SNP density, its cost remains a limitation for large-scale application. Imputation from low coverage data using a reference panel offers a cost effective alternative, but this approach has not yet been optimized for allo-octoploid strawberry. Results To reduce genotyping errors that limit phasing accuracy, we combined high sequencing depth (> 15×) with variant filtering based on average allele balance (AAB), linkage disequilibrium (LD), and Mendelian error rates (MER). Statistical phasing using SHAPEIT5 resulted in a mean switch error rate of 0.9%, with 50% of the genome covered by haplotype blocks of at least 654 kb (QHN50) without phase switches. To evaluate downstream imputation, samples from three genetically distinct populations (California, Florida, and HCFF) were downsampled to 1× and imputed using reference panels of varying size and composition (via GLIMPSE2). Both panel size and genetic diversity influenced imputation accuracy, with concordance rates ranging from 0.87 to 0.97 for the smallest panel and 0.94 to 0.98 for the largest, excluding three outliers. Conclusions These findings demonstrate that constructing a large, genetically diverse haplotype reference panel improves genotype dosage imputation from low-coverage sequencing data. However, high accuracy is still achievable with limited resources, making this a cost-efficient alternative to SNP arrays when adopting WGS-based genotyping in breeding programs. The strategy is broadly applicable to other crops where dense genotyping is needed but resources are limited. In such cases, sequencing approximately 50 genetically representative samples at ≥ 25× depth is recommended for building a reference panel suitable for imputation. WGS whole genome sequencing SHAPEIT5 phasing haplotype reference panel low-coverage imputation strawberry genetic diversity QHN50 Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Introduction In recent years, whole-genome sequencing (WGS) has emerged as a powerful tool for high-throughput genotyping in various crops, including garden strawberry ( Fragaria × ananassa ) [ 1 – 4 ]. WGS offers several significant advantages over traditional Single Nucleotide Polymorphism (SNP) arrays. While SNP arrays are widely used due to their relatively low cost and straightforward analysis pipelines, they present clear limitations. For instance, they suffer from ascertainment bias, as markers are pre-selected based on a small subset of samples [ 5 ]. If this subset is not representative of the whole target population, it might not have captured relevant genetic diversity. Moreover, SNP arrays contain a fixed and limited set of markers that cannot be modified or extended. Therefore, they are unable to detect population-specific or novel variants. In contrast, WGS enables the detection of a much broader spectrum of genetic variation, including rare variants, and provides greater flexibility in both experimental design and downstream analysis. Although WGS remains more expensive per sample, it provides much more genomic information and becomes more accessible in the future as sequencing costs decline. With the decreasing cost of sequencing technologies and the availability of several high-quality reference genomes [ 6 , 7 ], high-throughput genotyping in garden strawberry is increasingly shifting toward WGS-based approaches. For example, single variants identified through WGS have been used to discover and optimize diagnostic markers for red fruit colour intensity (BxANR_5UTR) and perpetual flowering (PFRU) [ 1 , 8 ]. Additionally, WGS has facilitated more comprehensive genetic diversity analyses by identifying and characterizing millions of genome-wide variants [ 2 , 9 ]. These developments illustrate that WGS is not only technically feasible in cultivated strawberry but also offers practical advantages in both breeding and research contexts, particularly where dense genotyping is required or where SNP arrays fall short. In addition to utilizing individual variants, genetic studies using WGS can also allow to obtain haplotypes, which offer additional advantages over single-variant analyses. Haplotypes are groups of variants (SNPs, insertions, deletions) that are inherited together from a single parent because they are located on the same chromosome. These variants will be inherited together until there has been a recombination. Utilizing haplotypes instead of biallelic SNPs after a reference panel has been established has multiple advantages. Firstly, imputation of low coverage data becomes more accurate by using a haplotype reference panel [ 10 ]. This helps reduce sequencing costs, making sequence based genotyping more affordable for larger studies and breeding applications. Secondly, research shows that haplotypes can improve genome-wide association study (GWAS) results. Studies using both real and simulated data have found that haplotypes could lead to higher quantitative trait locus (QTL) mapping accuracy and greater QTL detection power compared to individual SNPs [ 11 – 14 ]. Beyond that, haplotypes reflect identity-by-descent (IBD) rather than identity-by-state (IBS) in contrast to alleles of individual SNPs. Therefore, haplotypes can provide deeper insights into the genetic structure of a population by improving ancestry reconstruction, identifying allele-specific expression, and detecting selective sweeps with greater accuracy [ 14 ]. Haplotypes can be reconstructed from sequencing data using three main approaches. The first relies on sequencing read information to link variants that appear on the same sequencing fragment. This method is particularly effective with long-read technologies such as HiFi PacBio or Nanopore, as their longer reads can connect more SNPs, resulting in longer haplotypes [ 15 ]. A second approach uses linkage information in combination with the knowledge that only one of the two alleles is passed on from a parent (pedigree data, such as parent-offspring trios), to phase genotypes. The third approach is statistical phasing, as implemented in tools like SHAPEIT5 and BEAGLE 5.4 [ 16 , 17 ], that utilize hidden Markov models (HMM) to probabilistically infer haplotype phases. Here, the accuracy of haplotype inference mainly depends on the size of the reference panel and, to a lesser extent, their frequency distribution, as haplotypes that appear more frequently in the dataset are easier to detect and hence more reliable. In general, and for the functionally diploid Fragaria × ananassa , statistical phasing based on short-read sequencing is currently the most convenient haplotyping strategy. Compared to long-read-based phasing and pedigree-based phasing, it offers a practical balance between genotyping cost and data availability. Short-read resequencing is more cost-effective and more available than long-read sequencing, making statistical phasing based on short read resequencing easier to apply in breeding programs. In addition, statistical phasing does not depend on complete pedigree records, which are often incomplete in published datasets and breeding program records. While pedigree-based phasing can provide high accuracy, its application is limited when trios are incomplete or missing. Therefore, pedigree-based phasing is often less scalable and less flexible than statistical phasing, particularly in datasets with incomplete pedigree information or missing genotypes, e.g. , parents that were not genotyped and also not available anymore (so also not possible to genotype). Phased haplotypes can also be used for genotype dosage imputation of low-coverage sequencing data. When phasing is accurate, and the low-coverage samples are genetically related to a well-constructed reference panel, imputation performance is generally high [ 3 , 18 , 19 ]. Conversely, imputation accuracy declines if the reference panel lacks sufficient haplotype diversity, is not population-matched, or if the sequencing depth is too low to resolve ambiguous regions. Therefore, it also allows for assessing phasing quality across samples for which no parental genomes are available [ 16 ]. Rubinacci et al. (2021) demonstrated that statistical phasing methods can lead to high imputation accuracy when large reference panels (> 10,000) are used in human studies [ 20 ]. Similarly, high imputation performance has been achieved in agricultural species such as in cattle and sugar beet, using smaller panels (~ 300–1500), suggesting that statistical phasing remains effective across different reference panel sizes and species [ 19 , 21 ]. However, haplotyping in F. × ananassa , by using statistical phasing software, is more complicated than in humans, cattle or sugar beet because of its allo-octoploid nature (2n = 8x = 56). Allo-octoploidy complicates genotyping and increases error rates, while high quality data is important for accurate haplotype reconstruction. Genotyping errors can lead to switch errors, which occur when the phasing between consecutive heterozygous genotypes is incorrect, causing the haplotype to jump from one homologous chromosome to the other. These errors impact downstream analyses because these can propagate through downstream analyses [ 22 ]. Although F. × ananassa follows disomic inheritance, the four subgenomes per chromosome show high sequence similarity [ 23 ]. This could cause issues with read alignment, variant calling accuracy (genotyping errors) and, subsequently, phasing accuracy. Genotyping errors can be avoided by sequencing at high coverage and additionally, by applying variant filtering methods that improve downstream phasing accuracy [ 9 ]. High coverage sequencing (> 40×) in particular, reduces the likelihood of allelic dropout, which occurs when one of the two alleles at a heterozygous site is not captured during sequencing [ 24 ] However, while high depth reduces allelic dropout ( i.e. detecting too few alleles), it does not correct for genotyping errors caused by subgenome sequence similarity, where reads from different homoeologous regions may be misaligned, sometimes resulting in the detection of too many alleles. These errors are especially common in F. × ananassa [ 9 ]. For this reason, post hoc filtering strategies are also necessary. For example, filtering based on average allele balance (AAB) and linkage disequilibrium (LD) has been shown to reduce switch error rates by up to 44%, primarily by removing variants with elevated genotyping error rates [ 9 ]. Consequently, combining high sequencing depth with robust filtering strategies offers a practical and scalable approach to reduce genotyping noise in complex polyploid genomes. To support the implementation of high-throughput sequencing-based genotyping in a Fragaria × ananassa breeding context, this study focused on constructing varying haplotype reference panels and evaluating their utility for genotype dosage imputation from low-coverage (1×) sequencing data. To achieve this, we first minimized genotyping errors through a combination of high sequencing depth and stringent variant filtering, including filters based on average allele balance (AAB), linkage disequilibrium (LD), and Mendelian error rates (MER). Second, we phased and evaluated performance using SHAPEIT5 by using switch error rate and quality haplotype N50 (QHN50). Third, because haplotype inference accuracy (and consequently genotype dosage imputation) is known to strongly depend on the size of the reference panel and, to a lesser extent, its allele frequency distribution, we assessed how the haplotype reference panel’s composition, both in size and genetic diversity, impacted the imputation accuracy of genotype dosages based on low-coverage (1×) sequencing data. For this, sequencing reads from samples belonging to three genetically distinct populations were downsampled to 1× coverage. Then, the genotype dosages of these downsampled samples were imputed using a range of haplotype reference panel compositions. Materials and Methods A haplotype reference panel for Fragaria × ananassa was constructed using whole genome sequencing (WGS) data from 765 samples, where initial depth filtering (> 10x on chr_1A) resulted in 740 remaining samples. After further analysis of the effect of sample sequencing depth on switch error rates, we filtered even further on high depth (> 15×) across all 28 chromosomes, resulting in a filtered dataset of 653 high-depth samples. To build a reliable panel, genotyping errors were reduced by combining high sequencing depth (for samples) with strict variant filtering, including average allele balance (AAB), linkage disequilibrium (LD), and Mendelian error rate (MER) filters. All retained variants were phased using SHAPEIT5. To evaluate how reference panel composition influences downstream performance, samples from three genetically distinct populations were selected and their genotype dosages were imputed using panels of varying size and composition. This setup allowed us to see how different panel configurations affected the accuracy of imputed low-coverage WGS data. Genotype data Whole-genome sequencing (WGS) data was obtained from four distinct sources, all sequenced using Illumina 150 bp paired-end technology, though with varying sequencing depths. Two out of the four datasets were publicly available and included 137 samples from a breeding program in California [ 25 ] and 80 samples from a breeding program in Florida [ 4 ]. In this study, we refer to these sources collectively as cal_flo. Additionally, two datasets were generated in-house at Fresh Forward Breeding B.V. The first in-house dataset comprised sequenced samples from previous research projects, including a biparental mapping population (Holiday × Korona), while the second was specifically curated to ensure representativeness of the breeding program [ 26 ]. From these two datasets, we identified a representative group of high-chill samples, hereafter referred to as High Chill Fresh Forward (HCFF). Altogether, WGS data for 765 samples was available, predominantly consisting of F. × ananassa , but also including some F. chiloensis and F. virginiana samples. Raw reads were aligned to the F. × ananassa genome (FaRR1), by using minimap2 (v2.24) with the -ax sr preset [ 6 , 27 ]. Sambamba (v1.0.0) was used to convert the SAM output from minimap2 into BAM format while simultaneously filtering out unmapped reads using the filter -F "(not unmapped)". The filtered reads were then sorted with Samtools sort (v1.9). Variant calling was performed using BCFtools (v1.9) mpileup and call, applying the multiallelic caller (-mv) to generate the final VCF files. Biallelic SNPs were retained based on stringent SNP quality thresholds (QUAL > 40 and MAPQ > 30, DP 9575, and missingness rate 10× on chr_1A were removed for downstream filtering analyses, resulting in 740 remaining samples. Filtering erroneous variants due to subgenome similarity To exclude erroneous SNPs resulting from the high subgenome similarity in allo-octoploid strawberry, we applied three different SNP filtering strategies. First, Average Allele Balance (AAB) filtering was used, where only SNPs with 0.45 < AAB < 0.57 were retained [ 9 ]. Second, LD-Based filtering was used, by randomly selecting 140,000 SNPs (5,000 per chromosome) as an anchor set, and all SNPs were assigned to chromosomes based on LD to these anchor SNPs [ 9 ]. SNPs were retained if they showed an LD SNP count > 20 and a sum of squared LD (SSLD) > 0.93. Third, Mendelian error rate (MER) filtering was applied to take advantage of available pedigree information [ 28 ]. This method identifies variants that violate Mendelian expectations in duos and trios. Genotyping errors caused by subgenome sequence similarity also tend to result in elevated Mendelian error rates in SNPs, making MER filtering a useful addition for identifying and removing such problematic variants. As a result, 110 duos and 138 trios were used to filter out SNPs with MER > 0.01, allowing a maximum of two Mendelian errors per SNP ( Error! Reference source not found. ). The final dataset, consisting of 740 samples following these subsequent filtering steps, was designated as the standard dataset. SNP Binning and Genetic Diversity Analysis To reduce the number of SNPs to a manageable level for genetic diversity analyses, we developed a binning script ( Supplementary File S1 ). This script used connected components based on linkage disequilibrium (LD) to form groups of similar SNPs. Connected components are subgraphs where any two vertices are connected by paths, helping identifying clusters of SNPs likely to be inherited together. All SNPs that were in groups smaller than 10 SNPs were removed. For groups consisting of 10 or more co-segregating SNPs, we retained the SNP with the highest QUAL score. In this study, we binned SNPs using windows of 1 Mb, a minimum LD threshold of 0.9, and a minimal group size of 10 SNPs, resulting in the binned dataset. To evaluate genetic diversity, we categorized samples by the origin of their sequencing data (California, Florida, Fresh Forward Breeding B.V.) and performed principal components analysis (PCA) using the prcomp function in R to analyse the genetic diversity among these groups. Phasing optimization using high-coverage data (> 15 × coverage) After initial depth filtering, where samples with more than 10× coverage on chromosome 1A were retained, an analysis was conducted to evaluate the effect of individual sequencing depth on phasing accuracy. Based on this evaluation, a threshold of 15× average sequencing depth was selected to optimize phasing accuracy, as informed by the initial switch error rate analysis. Specifically, the average depth per SNP was calculated for each sample in the standard SNP dataset. Samples with an average depth below 15× were subsequently excluded, resulting in a final set of 653 high depth samples for further analysis. Given the impracticality of pedigree- or long-read-based haplotyping at scale, statistical phasing of short-read resequencing data stands out as the most feasible and effective strategy, especially with the availability of WGS data for 740 samples from both public repositories and in-house collections at Fresh Forward Breeding B.V. [ 2 , 4 ]. Therefore, phasing was performed using SHAPEIT (v5.1.1). To optimize its performance, the influence of three different parameters on phasing accuracy was investigated [ 16 ]. First, 101 parents from duos and trios present in this panel were removed from the standard dataset, so that these offspring are phased without information about their parental haplotypes. Subsequently, the remaining samples were phased using SHAPEIT5 with varying stepwise parameter values for the parameters hmm-ne, hmm-window, and pbwt-window (Table 1 ). Phasing accuracy was assessed by computing the switch error rate (SER) for phased samples belonging to duos or trios, using the unphased genotypic data of their parents [ 29 ]. The total number of switch errors across samples was divided by the total number of evaluated heterozygous SNPs to obtain the overall switch error rate (SER). The SER was computed separately for each of the 28 chromosomes, with each chromosome phased in an independent run. Although the phasing algorithm is stochastic and might introduce variability at the level of individual chromosomes, aggregating results across all chromosomes offsets this variability. Therefore, repeated iterations were not needed, as the conclusions were based on the combined results from all 28 independent runs. This approach allowed us to systematically evaluate the effect of different SHAPEIT5 parameter settings on phasing accuracy. To better understand the effect of genotyping errors and low MAF on phasing quality, we first separated flip errors from the total number of switch errors. Specifically, flip errors are double switch errors, where the haplotype switches and then immediately switches back to the correct state. These are typically caused by genotyping errors or low-frequency variants. Next, we investigated the relationship between flip errors and remaining switch errors. Additionally, both were compared across minor allele frequency bins to assess how allele frequency contributes to phasing accuracy. Table 1 SHAPEIT5 parameters and their tested range for optimization of phasing accuracy. SHAPEIT5 parameter Test range Hmm-ne 500–500,000 (step size: starting with 500 and increasing up to 400,000, default: 15,000) Hmm-window 1–10 (step size: 1, default: 4) Pbwt-window 1–10 (step size: 1, default: 4) In addition to evaluating phasing accuracy using the switch error rate, an alternative metric was proposed by Lo et al. (2011), known as the adjusted N50 (AN50). AN50 is defined as the minimum haplotype block length such that 50 percent of all phased variants are located in blocks of at least this size [ 30 ]. Building on this,Duitama et al. (2012) introduced the quality adjusted N50 (QAN50), which follows the same principle but further divides each haplotype block into the longest sub-blocks without any detected switch errors [ 31 ]. In this study, we computed a similar metric, but with two key differences: we did not adjust for the number of variants per haplotype block, and we did not evaluate the haplotype length based on the number of variants. Instead, we focused on the proportion of the genome that was phased without switch errors. We refer to this measure as the quality haplotype N50 (QHN50). To calculate it, all haplotype blocks were ordered from largest to smallest, and the cumulative proportion of the total phased sequence covered by haplotypes of a given length was computed across all samples and chromosomes. The resulting QHN50 represents the haplotype length at which 50 percent of the phased sequence is covered by blocks of at least that length. Selection of samples for imputation To assess the influence of genetic diversity on phasing accuracy, we evaluated phasing accuracy indirectly through genotype dosage imputation accuracy of downsampled sequencing data (to 1× coverage) from various samples. Although estimating phasing accuracy directly by computing the switch error rate is straightforward and efficient, this approach is limited to populations with known pedigrees, which were only partly available for two datasets in this study. Therefore, we relied on an indirect approach: imputing genotype dosages in downsampled samples using a phased reference panel, where higher accuracy reflects better phasing quality. This approach allows phasing performance to be evaluated beyond pedigreed samples and across genetically diverse backgrounds [ 16 ]. A total of 66 samples were selected for downsampling and imputation from three genetically distinct populations (California, Florida, and High Chill Fresh Forward (HCFF)) each with a clearly defined origin. Within each population, 25% of samples were randomly selected (Table 2 ). To simulate low-coverage whole genome sequencing, the fraction of reads required to reach 1× coverage was first calculated for each sample based on the total sequencing depth across all chromosomes. Subsequently, the corresponding BAM files were downsampled using samtools view -s with the computed read fraction, resulting in BAM files with approximately 1× depth. Table 2 Overview of population sizes and selected samples for imputation. nr Origin population Total population size (> 15x depth) Selected samples for imputation (~ 25% of total population size) Remaining samples for reference panels 1 California 32 8 24 2 Florida 62 16 46 3 HCFF 169 42 127 Phasing and imputation using different reference panels To investigate how haplotype reference panel size and composition affect phasing and low-coverage imputation accuracy, we tested six different haplotype reference panels. In addition to using the entire population (> 15× depth) as a reference panel, various combinations of the three populations (California, Florida, and High Chill Fresh Forward (HCFF)) were tested. Due to SHAPEIT5's minimum sample size requirement of 50 samples, reference panels consisting solely of the California (24) or Florida (46) populations were excluded (Table 3 ). In all cases, the downsampled samples (66) were always excluded from all haplotype reference panels. To further separate the effect of panel diversity from panel size, we randomly selected 70 samples from the HCFF population, matching the size of the combined California + Florida panel. This allowed us to compare panels of equal size but different genetic diversity. Additionally, to address whether it is preferable to maximize the number of samples or prioritize haplotype quality, we also included an extra reference panel composed of the entire population without applying the minimum sample depth filter (15×). Table 3 Overview of reference panels used for imputation, where the downsampled samples were excluded from the reference panels. nr Reference panel Reference panel size (number of samples) 1 California + Florida 70 2 HCFF small 70 3 HCFF 127 4 HCFF + California + Florida 197 5 Whole population 587 Extra Whole population (no 15× DP filter) 674 Each reference panel was phased independently using SHAPEIT5 to prepare for downstream imputation. Default parameters were used except for the hmmne , which was set to 7500. Additionally, the default MCMC iteration scheme was extended (10b,1p,1b,1p,1b,1p,1b,1p,10m) to improve phasing accuracy, following the recommendations by Hofmeister et al. (2023), though at the cost of increased runtime [ 16 ]. The resulting phased reference panels were then used as input for GLIMPSE2 to impute the genotype dosages of downsampled samples. GLIMPSE2, recommended by the SHAPEIT5 authors for low-coverage WGS imputation, uses a haplotype reference panel to refine genotype likelihoods and infer missing genotypes [ 20 , 32 ]. All imputation runs were performed using GLIMPSE2 with default settings. Imputation accuracy evaluation Imputed genotype dosages were compared against the original high-depth genotype dosages from the standard dataset (ground truth). Imputation accuracy was assessed using concordance rates, computed for each imputed sample using each reference panel. The concordance rate measures the proportion of correctly imputed genotype dosages relative to the total number of imputed dosages. Since heterozygous dosages are more difficult to impute than homozygous dosages in low-coverage WGS data, concordance rates were separately calculated for homozygous and heterozygous dosages. Following Zhang et al. (2023), we used a 3×3 cross-classification table to compare the imputed genotype dosages with the ground truth. In this table, homozygous reference genotypes (aa) correspond to dosage 0, heterozygous genotypes (ab) to dosage 1, and homozygous alternative genotypes (bb) to dosage 2 [ 18 ]. Each cell, n ij , represents the number of SNPs where the true dosage is i and the imputed dosage is j (Table 4 ). Using these tables, an overall summary was created by computing concordance rates per sample and per reference panel, aggregated across all chromosomes. Additionally, these concordance values were separated for homozygous and heterozygous ground truth calls to better capture differences in imputation performance between variant types. Table 4 Marginal cross-classification table of true and imputed genotypes from a single sample. Imputed genotypes True genotypes aa ab bb Total aa n 11 n 12 n 13 n 1. ab n 21 n 22 n 23 n 2. bb n 31 n 32 n 33 n 3. Total n .1 n .2 n .3 n.. Results Impact of filtering methods on SNP retention Applying three filtering methods resulted in the retention of 7.22M SNPs over all chromosomes (Fig. 1 A). Individually, strict thresholds for Average Allele Balance (AAB), LD-based filtering, and Mendelian error rates (MER) retained 11.47M, 11.72M, and 13.49M SNPs, respectively. While LD-based filtering and MER retained similar SNP counts, MER retained slightly more overall. Despite this, each method captured distinct SNP sets, with 7.22M SNPs shared across all three. On average, this corresponds to 258K SNPs per chromosome. Notably, AAB filtering retained the fewest unique SNPs compared to the other two methods. To further examine differences between the filtering methods, we analysed the SNPs that were discarded (Fig. 1 B). All three methods agreed on removing 2.69M SNPs, suggesting these were the most evident unreliable SNPs. The largest overlap between two methods was between AAB and LD-based filtering, which together flagged an additional 2.60M SNPs for removal. In contrast, the smallest overlap was between MER and LD-based filtering, with only 0.44M additional SNPs identified. Each method also flagged at least 1.66M SNPs for removal that were not identified by the other two, highlighting the different criteria of each method. Impact of sequencing depth on phasing accuracy To assess the impact of sequencing depth on phasing accuracy, we first phased the whole population (n = 740) and evaluated a biparental population, in which the noise from dataset variability and genetic diversity is minimized (Fig. 2 A). A clear trend is visible between the switch error rate and the sequencing depth of the phased offspring samples, with the lowest switch error rates observed at sequencing depths higher than 25×. However, 356 samples in the whole WGS dataset have depths below 25×. To balance between maximizing the number of samples and ensuring phasing accuracy, we filtered for an average sequencing depth of at least 15× per sample, removing 87 samples and resulting in a final dataset of 640 samples (Fig. 2 B). Linear relationship between flip errors and remaining switch errors To better estimate the switch error rate, flip errors were first identified and removed from the total set of switch errors. Flip errors, defined as double switch events where the haplotype switches for one SNP, were separated to isolate remaining switch errors. A linear relationship was observed between flip error rates and remaining switch error rates across samples (Fig. 2 C). To further examine this pattern, the relationship between switch error rates and minor allele frequency (MAF) was assessed. Switch errors were not enriched in the lowest MAF bin, and their frequency remained relatively consistent across MAF bins (Fig. 2 D). These results suggest that flip errors and remaining switch errors are correlated and that switch errors do not predominantly occur at low MAF sites. Optimized SHAPEIT5 parameters Three parameters were tested stepwise to evaluate their impact on phasing accuracy and determine the optimal settings for phasing F. × ananassa WGS data. First, we assessed the hmm-window parameter, which showed minimal variation in phasing accuracy. The lowest accuracy was observed at values of 1 and 2 cM, but overall, there was no strong justification to deviate from the default setting of 4 cM (Fig. 3 A). Similarly, testing different pbwt-window settings revealed no substantial effect on phasing accuracy, leading to the retention of the default 4 cM value (Fig. 3 B). Moreover, for the effective population size (hmm-ne), accuracy was slightly better at 7,500 compared to the default of 15,000 (Fig. 3 C). Although the difference was small, we opted to set hmm-ne to 7,500 for downstream analyses, as it showed a marginal improvement in some cases, which resulted in a mean switch error rate of 0.92%. Accurate Haplotype Block Length Distribution To evaluate the practical usefulness of the haplotypes in the reference panel, we analysed the distribution of haplotype block lengths without switch errors. Flip errors were first removed to isolate uninterrupted haplotype blocks, separated only by single switch errors. Most blocks were relatively short, with a median length of 30 kb and a mean of 155 kb. To assess how block size contributes to overall phasing, all blocks were ranked by length, and the cumulative proportion of the total phased sequence covered by blocks of at least a given size was computed across all samples and chromosomes (Fig. 4 ). This analysis yielded a mean quality haplotype N50 (QHN50) of 654 kb and a mean QHN90 of 101 kb. These results show that a substantial portion of the phased genome is composed of long, uninterrupted haplotype blocks. Imputation panel composition A principal components analysis (PCA) was conducted using the binned dataset of all 740 samples to visualize genetic relationships ( Figure S1 ). Here, subpopulations were highlighted to define groups based on their genetic origin. Several samples belonging to F. virginiana (dark blue) and F. chiloensis (purple) were present in the dataset and seemed to strongly influence PC2. PC1 appears to reflect chilling requirement, with low-chill populations clustering at the left side of the plot and high-chill populations on the right. Additionally, two clusters of samples belonging to two biparental populations were visible in the lower-right and middle-right areas of the plot. Still, three distinct groups could be defined (based on the data origin). California and Florida formed separate, well-defined clusters, both at the left side of PC1. In contrast, a large group of high chill samples from Fresh Forward (HCFF) was identified as counterpart of these two low chill clusters (Florida and California). The biparental groups from HCFF were excluded from further analysis to avoid bias introduced by many samples with low genetic variation. To validate the distinctness of these three groups, a separate principal components analysis was conducted with just these three groups (Fig. 5 ). The groups remained well separated, although a few samples with Florida origin were in the HCFF group, and one HCFF sample was positioned near the California group. Additionally, the groups varied in size: California (32 samples), Florida (62 samples), and HCFF (169 samples). For each group, 25% of the samples were randomly selected to test imputation accuracy based on different haplotype reference panels (highlighted in Fig. 5 ). The PCA plot confirmed that these samples were representative of their respective groups (Fig. 5 ). Genetic Diversity Across Reference Panels Assessed by Proportion of Polymorphic SNPs To complement insights from principal components analysis (PCA), we quantified genetic diversity by calculating the proportion of polymorphic SNPs within each reference panel. This summary metric provides an informative overview of within-panel variation without compensating for panel size. As expected, the Whole_pop_alldp panel showed the highest diversity, with 100% of SNPs being polymorphic, since SNP discovery was performed across this entire dataset, including low-coverage samples, but excluding the selected samples for imputation tests. In contrast, smaller subsets of this panel showed less polymorphism. Specifically, the cal_flo panel contained only 83% polymorphic SNPs. However, another panel of identical size, HCFF_small, retained 98% polymorphic SNPs, suggesting that population composition, rather than panel size alone, influenced the observed diversity. Additionally, all other panels (HCFF, HCFF_cal_flo, and Whole_pop) displayed consistently high diversity, with the proportion of polymorphic SNPs ranging from 0.99 to 1.00. Imputation accuracy across different haplotype reference panels To assess imputation accuracy, selected samples from different groups were imputed using IMPUTE2 with haplotype reference panels of varying sizes and genetic diversity [ 20 , 32 ]. Overall, imputation accuracy was high, with total concordance rates exceeding 0.92 for most samples (Fig. 6 A). However, an exception was observed when HCFF samples were imputed using the cal_flo reference panel, which resulted in lower concordance rates (mean concordance of 0.90). Additionally, three HCFF outliers showed consistently lower concordance across all reference panels. To better understand these trends with respect to variant types, we analysed concordance rates separately for homozygous (Fig. 6 B) and heterozygous calls (Fig. 6 C). As expected, homozygous calls exhibited higher concordance than heterozygous calls. Moreover, population-level differences in imputation accuracy were apparent. For example, California samples consistently showed the highest homozygous SNP concordance across all panels, followed by Florida and HCFF samples. Genotype dosage imputation performance for homozygous sites also varied depending on the reference panel. Specifically, cal_flo again resulted in the lowest homozygous concordance for HCFF samples. In contrast, when HCFF samples were included in the reference panel, concordance differences across populations were minimal. In contrast to the relatively consistent homozygous results, heterozygous imputation rates showed greater variation across both panels and populations. Among all reference sets, cal_flo again produced the lowest concordance rates across all samples, followed by HCFF_small, HCFF, HCFF_cal_flo, and Whole_pop. The highest accuracy was achieved using the Whole_pop_alldp panel (the panel that also included samples that were removed with the < 15x coverage filtering). Importantly, imputation accuracy of Whole_pop and Whole_pop_alldp remained largely comparable despite the latter including lower-depth samples. Specifically, total concordance increased only slightly from 0.964 to 0.965, a change mainly driven by a small improvement in heterozygous call accuracy (from 0.884 to 0.886), while homozygous concordance remained nearly unchanged. Additionally, the population origin of imputed samples significantly influenced heterozygous performance. When using the Whole_pop panel, Florida samples exhibited the highest median heterozygous concordance, followed by HCFF and California samples. Interestingly, although California samples had the lowest heterozygous accuracy among the three groups, they achieved the highest overall imputation concordance. This outcome can be attributed to the predominance of homozygous reference calls as these calls showed higher concordance rates, with California samples having more homozygous reference calls than Florida, and Florida more than HCFF ( Supplementary File S2 ). Higher imputation accuracy for F. vesca derived (A) subgenomes To investigate subgenome-specific patterns in imputation performance, we analysed heterozygous concordance rates per chromosome across all imputed samples by using Whole_pop as reference panel. This chromosome-level resolution provided insight into how different subgenomes contribute to overall accuracy. Interestingly, the F. vesca derived A subgenomes consistently demonstrated the highest concordance rates across all chromosomes. Moreover, their median heterozygous concordance exceeded 0.90 in every case, clearly distinguishing them from other subgenomes, where median concordance values were ranging between 0.85 and 0.90, with no specific chromosomes deviating substantially from this pattern (Fig. 7 A). To explore whether higher imputation accuracy also corresponds to improved phasing quality, we computed the median QHN90 score for each chromosome using the standard dataset where samples with < 15× coverage were removed (Fig. 7 B). If phasing were more accurate for the A subgenomes, one would expect to observe higher QHN90 values for these chromosomes. However, the pattern was less consistent. Although chromosomes 2A, 3A, 4A, 5A, and 7A showed relatively high QHN90 scores, this trend did not hold across all A subgenomes. In particular, chromosomes 1A and 6A did not have the highest QHN90 scores among their respective subgenome groups, suggesting that phasing accuracy may not fully explain the observed differences in imputation performance. Discussion This study aimed to construct a high-quality haplotype reference panel for Fragaria × ananassa using whole-genome sequencing (WGS) data and to evaluate its utility for accurate dosage imputation from low-coverage sequencing. We found that combining high sequencing depth with rigorous variant filtering (based on allele balance, linkage disequilibrium, and Mendelian errors) enabled statistical phasing using SHAPEIT5 with switch error rates around 1%, corresponding to an average QHN50 of 654 kb. The results also showed that once a reference panel (using > 15× sequencing data) capturing most, if not all, possible alleles is established, low coverage (1×) sequenced varieties can be imputed with high accuracy for downstream applications. This accuracy was maintained with only minimal loss, even when using smaller or less genetically diverse haplotype reference panels. These findings demonstrate that high-throughput WGS-based genotyping is a viable strategy for allo-polyploid crops like cultivated strawberry, where a single, well-constructed haplotype panel can support robust dosage imputation across genetically diverse breeding varieties. This work extends previous findings in diploid species and simpler genomes by showing that statistical phasing and imputation are also effective under the genomic complexity of allo-octoploids. Genotyping errors are main cause of switch errors Although extensive filtering was applied to the WGS dataset in F. × ananassa , switch errors remained. First, three filtering methods were used to obtain a reliable SNP dataset across all samples, including those without pedigree data. Second, samples with insufficient sequencing depth (< 15×) were removed. However, a switch error rate of 0.008–0.011 was still observed across all chromosomes, meaning an error occurred roughly every 100 heterozygous SNPs. One possible explanation could be that population size and diversity influenced switch error rates. However, increasing the reference panel from HCFF (n = 142) to the Whole Population (n = 640) led to only a marginal improvement in imputation accuracy. This suggests that population size and genetic diversity were only small factors in haplotyping accuracy in our dataset. Instead, the presence of a linear relationship between flip ( i.e. double switch errors) and remaining switch errors ( i.e. errors spanning at least two consecutive heterozygous SNPs) points to a shared underlying cause. Since flip errors are typically associated with low MAF SNPs and genotyping errors, we examined whether switch errors were enriched in low MAF variants. Although the MAF < 0.05 bin contained the highest number of SNPs, switch errors were not disproportionately concentrated in this bin. Consequently, genotyping errors, rather than MAF, appear to be the main factor driving switch errors. Moreover, switch error rates could only marginally be improved by optimizing three phasing parameters, which aligns with expectations: phasing parameter adjustments cannot correct switch errors that are caused by genotyping errors. Given these findings, an approach to further reduce genotyping errors could be to focus on variant call accuracy. While filtering had already been applied at the whole SNP and sample levels, no additional filtering was done after the default variant calling with bcftools . Prioritizing missing variant calls over unreliable ones may be more effective when performing phasing. One approach could be to filter based on a minimum depth threshold for variant calls, but this threshold would need to be adjusted for samples that were sequenced with lower depth. Without such an adjustment, samples with lower sequencing depths could end up with highly skewed missingness rates. Another option is to filter based on unbalanced allele frequencies at heterozygous sites (Allelic Bias) [ 9 , 33 ]. This is similar to the Average Allele Bias filtering used in this study but applied at the individual call level rather than the population level. Additionally, setting thresholds for Phred-scaled likelihoods (PL) or genotype quality (GQ) on the genotype dosage calls could further reduce variant call errors, helping to reduce switch errors and ultimately improve imputation accuracy. However, when combining multiple variant call filtering strategies, care must be taken not to remove too many calls, as retaining a sufficient number of informative sites is important for effective phasing and imputation. Since SHAPEIT5 includes internal imputation for missing variant calls during phasing, some missing data is tolerated [ 34 ]. Still, it remains unclear at what level of variant call missingness phasing performance begins to decline. Potential Bias from Mendelian Filtering in Switch Error Rate Estimation To filter SNPs, we applied Mendelian error rates (MER), which may have introduced bias in the evaluation of switch error rates. This is because switch error rates were calculated downstream using duos and trios, the same data that is used for MER filtering. However, we still considered SNP-level MER filtering useful, because if erroneous (based on MER) genotype dosages cluster in one SNP, other genotype dosages for the same SNP that could not be evaluated due to missing pedigree data are also more likely to be erroneous. To minimize the potential MER bias on SER, we set a maximum Mendelian error rate of 1%, ensuring that the effect on switch error rate estimation remained limited. Despite this, some degree of bias is inevitable. Therefore, we used imputation accuracy as an additional proxy for haplotype accuracy, as errors in haplotypes directly reduce imputation performance [ 34 ]. Reference genome relatedness drives imputation accuracy differences Imputation accuracy varied across samples from different origins, with clear differences observed between homozygous and heterozygous calls, which showed mean concordance rates of 0.99 and 0.88, respectively, when using the Whole_pop reference panel. California samples showed the highest homozygosity, followed by Florida, while HCFF samples were the most heterozygous. Consequently, total concordance rates were highest for California, followed by Florida and then HCFF samples. Increased homozygosity in Californian breeding material was also reported by Feldmann et al. (2024) who attributed it to selection, random genetic drift, and selective sweeps [ 35 ]. Interestingly, homozygous calls were more accurately imputed for California samples than for Florida or HCFF samples, with the largest gap observed between California and HCFF. This is likely due to the reference genome used for alignment (FaRR1), which originates from California and is genetically closer to California samples than to those from Florida or HCFF [ 6 ]. As a result, reads containing reference alleles are more likely to align correctly, improving variant calling, phasing, and ultimately imputation, an effect known as reference genome bias. Indeed, imputed California samples had a much higher proportion of homozygous reference calls than homozygous alternative calls, compared to both Florida and HCFF samples. Moreover, homozygous reference calls generally had higher imputation accuracy than homozygous alternative calls ( Supplementary File S1 ). Similarly, Florida samples had more homozygous reference calls than HCFF, suggesting they are also more closely related to the reference genome. These results highlight the significant impact of both homozygosity and reference genome relatedness on imputation accuracy, as well as on haplotype reconstruction, since homozygous regions do not require phasing and therefore do not introduce switch errors. In contrast to the trend observed for homozygous calls, where California samples showed the highest imputation accuracies, heterozygous imputation rates followed a different pattern. On average, Florida samples had higher imputation rates for heterozygous sites than California samples regardless of the haplotype reference panel used. Given that heterozygous calls always include one reference allele (matching the reference genome), one might expect that phasing or imputing the other allele would be easier when the reference genome is closely related. Yet, our results did not support this assumption. One possible explanation is that the California reference population included in the haplotype reference panel was relatively small (n = 24) compared to the Florida (n = 42) and HCFF (n = 142) populations. If certain alternative alleles from the to-be-imputed California samples were not represented in the reference panel, this could have limited imputation accuracy. Since a shared haplotype can only be found if it is present in the haplotype reference panel [ 36 , 37 ], this limitation may have contributed to the lower heterozygous imputation rates observed for California samples. Another possible explanation is that relevant haplotypes were present in the reference panel but contained more switch errors. Notably, the average sequencing depth of the ground truth dataset for California samples (20.7×) was lower than that of Florida (32.5×) and HCFF (28.0×) samples. Since lower sequencing depth is associated with elevated switch error rates [ 36 ], this may have introduced more phasing errors in the California-specific haplotypes. Therefore, these haplotypes may have been less reliable, ultimately leading to lower imputation accuracy for rare or population specific haplotypes in California samples. Genetic diversity and size of reference panel drive imputation accuracy To investigate how panel size and genetic diversity influence imputation accuracy, the downsampled samples were imputed using five different haplotype reference panels. Panel size was consistently associated with improved imputation accuracy. Even adding 87 previously filtered samples with lower sequencing depth (< 15×) to the Whole_pop reference panel led to a slight increase in imputation performance. Therefore, while sequencing at a depth greater than 15× is recommended, including some samples with lower coverage does not negatively impact overall accuracy. To assess the effect of genetic diversity, we constructed a smaller HCFF panel with the same sample size as the cal_flo panel. The HCFF_small panel had a much higher proportion of polymorphic SNPs (98%) than the cal_flo panel (83%), indicating greater genetic diversity. This difference likely reflects the number of breeding objectives in each region: while California and Florida breeding programs mainly target low-chill, open-field cultivation, the HCFF panel includes high-chill selections from multiple cultivation systems such as open field, greenhouse (spring), and greenhouse (winter). In terms of imputation outcomes, HCFF samples imputed with the HCFF_small panel showed notably higher accuracy compared to those imputed with the cal_flo panel. In contrast, for California and Florida samples, imputation performance was similar between the two panels, with slightly higher imputation rates for HCFF_small. This suggests that the genetic diversity required for accurate imputation of California and Florida samples is sufficiently captured in both panels. However, for accurate imputation of HCFF samples, the required haplotypes are better represented in the HCFF_small panel. Nevertheless, it is worth noting that the cal_flo panel still achieved a median total imputation rate of nearly 0.97 for HCFF samples, with heterozygous concordance rates exceeding 0.75. Together, these results highlight the importance of both genetic diversity and panel size for accurate imputation. At the same time, they also show the robustness of this approach, as accurate imputation of low-coverage WGS data is achievable across various panel configurations. Concordance outliers Three outliers were identified, each showing lower total concordance rates compared to the other samples. All three outliers were from HCFF, which also had the highest number of imputed samples (42). For one of these, the pedigree data had been removed during a prior pedigree check due to elevated Mendelian errors. Elevated Mendelian errors can arise from various factors, such as errors in pedigree records or sample swaps. While the exact cause of this outlier remains uncertain, its combination with a very high heterozygosity rate suggests the possibility of severe contamination. The second outlier did not exhibit elevated Mendelian errors (with both parents known) but still displayed a high heterozygosity percentage. This could simply be due to high heterozygosity affecting the overall concordance rates. The third outlier was a parent in multiple other trios, but no elevated pedigree errors were detected. This sample also showed a slightly higher heterozygosity rate. In this case, contamination with another strawberry sample seems the most likely explanation, as it would not cause additional pedigree errors but could result in erroneous haplotypes. Higher imputation accuracy for the F. vesca -derived A subgenomes Chromosome-level imputation results revealed a consistent pattern in which the A subgenomes, derived from F. vesca , showed higher imputation accuracy than the B, C, and D subgenomes. This elevated accuracy may be explained by the A subgenome being the most recently incorporated to the octoploid strawberry, and therefore showing less genetic variation (and therefore less SNPs) and lower haplotype diversity than the B, C and D subgenomes [ 9 , 38 , 39 ]. These factors may have contributed to more accurate variant calls and, combined with lower haplotype diversity, resulted in higher quality haplotypes. Supporting this, several, but not all, A subgenomes also showed higher QHN90 scores (Fig. 7 B), suggesting longer contiguous haplotype blocks. Moreover, lower haplotype diversity on the A subgenomes may increase the chance that haplotypes to be imputed are already represented in the haplotype reference panel. In addition, these characteristics may also have contributed to a higher assembly quality of the A subgenomes in the reference genome, further improving variant calling and imputation performance. Haplotype reference panel beyond imputation In this study, we demonstrated that haplotype reference panels can be effectively used to impute low-coverage (1×) WGS data in allo-octoploid strawberry. However, besides genotype dosage imputation, such haplotype reference panels are also valuable in breeding applications that require accurate haplotypes. For instance, they can support haplotype-based parental checks and pedigree reconstruction, as shown in fruit crops like apple, where shared haplotype blocks were used to resolve unknown parentage and trace ancestry across generations [ 40 ]. Haplotype structure has also been used to inform the selection of elite haplotypes across germplasm pools, as shown in wheat through a haplotype-led breeding framework that improved mapping resolution and selection precision [ 41 ]. Additionally, haplotypes can improve genomic prediction by capturing local linkage patterns and multi-allelic effects that are often missed by single SNPs. Studies in canola, wheat, soybean, and dairy cattle have shown that using haplotypes as predictors can lead to higher genomic prediction accuracy, especially for complex traits [ 42 , 43 ]. Moreover, haplotypes were useful for the development of diagnostic markers. For example, in wheat, LD-defined haploblocks were used to fine-map a disease resistance QTL and design a minimal marker panel that could trace the resistant haplotype with high specificity [ 44 ]. For many of these applications, long and accurate haplotype blocks are important. In our study, we achieved a mean QHN50 of 654 kb, indicating that half of the phased genome is represented in blocks of at least this length without switch errors. This level of haplotype contiguity is likely sufficient for tasks such as haplotype tracing, genomic prediction, or diagnostic marker development, where full chromosome-level phasing is not strictly necessary. Conclusions In this study, we constructed a haplotype reference panel for Fragaria × ananassa using whole-genome sequencing data and a statistical phasing algorithm (SHAPEIT5). High sequencing depth (> 15×), combined with strict variant filtering based on average allele balance, linkage disequilibrium, and Mendelian error rate, was important to reduce genotyping errors and improve phasing accuracy. Although optimization of phasing parameters had limited effect, these filtering strategies helped reduce switch errors and enabled accurate phasing, with switch error rates of approximately 0.01. This corresponded to haplotype blocks with a mean QHN50 of 654 kb, meaning that half of the phased genome was covered by blocks of at least this length without switch errors. If even longer haplotype blocks are required, phasing accuracy can still be improved by further reducing genotyping errors. For constructing haplotype reference panels, sequencing at a depth of at least 25× is recommended (Fig. 2 A). However, available samples with sequencing depth lower than 25× can still be valuable and should be included. Furthermore, we constructed haplotype reference panels with varying genetic diversity and size, to investigate their influence on low-coverage WGS genotype dosage imputation. As expected, larger panels yielded better results. However, genetic diversity only had a limited effect, where less diverse panels had lower imputation accuracies for samples whose haplotypes were not represented well in the panel. In addition, the A subgenome consistently showed higher imputation accuracy than the B, C, and D subgenomes, which may relate to its dominant role in the genome. Together, these findings provide practical guidelines for building haplotype reference panels and show that statistical phasing and imputation of low-coverage WGS data is robust across different panel configurations. The approach is also applicable to other crops, where sequencing around 50 genetically representative samples at ≥ 25× depth can enable accurate dosage imputation from low-coverage data. Declarations Competing interests The authors declare no competing interests. Funding This project was supported by Fresh Forward Breeding B.V. and the TKI project ‘LWV20.112 Application of sequence-based multi-allelic markers in genetics and breeding of polyploids’ (BO-68–001-042-WPR). Author Contribution T.K. led the conceptualization, investigation, software implementation, visualization, and drafting of the manuscript. T.K., J.H.W., P.A., R.G.F.V., and C.M. contributed to the conceptualization, methodology and reviewed and provided feedback on manuscript drafts. All authors approved the final version of the manuscript and are accountable for their contributions. Acknowledgement We would like to thank Fresh Forward Breeding B.V. for providing the sequencing data used in this study, and Rian Peters for DNA isolation of the Fresh Forward core collection. Data Availability The California and Florida sequencing data are publicly available from their respective studies [4, 25]. The Fresh Forward Breeding B.V. datasets generated and/or analysed during the current study are not publicly available because they are owned by private company Fresh Forward Breeding B.V. but can be requested from Fresh Forward Breeding B.V. Scripts for computing average allele balance and LD-based filtering are available in a previous publication [9]. SHAPEIT5 is available at https://odelaneau.github.io/shapeit5/. GLIMPSE2 is available at https://odelaneau.github.io/GLIMPSE/. Script for binning is available in Supplementary File S1. Script for QHN computation is available in Supplementary File S3. References Saiga S, Tada M, Segawa T, Sugihara Y, Nishikawa M, Makita N, et al. NGS-based genome wide association study helps to develop co-dominant marker for the physical map-based locus of PFRU controlling flowering in cultivated octoploid strawberry. Euphytica. 2022;219:6. https://doi.org/10.1007/s10681-022-03132-7 . Hardigan MA, Lorant A, Pincot DDA, Feldmann MJ, Famula RA, Acharya CB, et al. Unraveling the Complex Hybrid Ancestry and Domestication History of Cultivated Strawberry. Mol Biol Evol. 2021;38:2285–305. https://doi.org/10.1093/molbev/msab024 . Liu S, Martin KE, Snelling WM, Long R, Leeds TD, Vallejo RL et al. Accurate genotype imputation from low-coverage whole-genome sequencing data of rainbow trout. G3 Genes|Genomes|Genetics. 2024;14:jkae168. https://doi.org/10.1093/g3journal/jkae168 Fan Z, Whitaker VM. Genomic signatures of strawberry domestication and diversification. Plant Cell. 2024;36:1622–36. https://doi.org/10.1093/plcell/koad314 . Geibel J, Reimer C, Pook T, Weigend S, Weigend A, Simianer H. How imputation can mitigate SNP ascertainment Bias. BMC Genomics. 2021;22:340. https://doi.org/10.1186/s12864-021-07663-6 . Hardigan MA, Feldmann MJ, Pincot DDA, Famula RA, Vachev MV, Madera MA et al. Blueprint for Phasing and Assembling the Genomes of Heterozygous Polyploids: Application to the Octoploid Genome of Strawberry. bioRxiv. 2021;:2021.11.03.467115. https://doi.org/10.1101/2021.11.03.467115 Han H, Barbey CR, Fan Z, Verma S, Whitaker VM, Lee S. Telomere-to-Telomere and Haplotype-Phased Genome Assemblies of the Heterozygous Octoploid ‘Florida Brilliance’ Strawberry (Fragaria × ananassa). bioRxiv. 2022;:2022.10.05.509768. https://doi.org/10.1101/2022.10.05.509768 Labadie M, Vallin G, Potier A, Petit A, Ring L, Hoffmann T, et al. High Resolution Quantitative Trait Locus Mapping and Whole Genome Sequencing Enable the Design of an Anthocyanidin Reductase-Specific Homoeo-Allelic Marker for Fruit Colour Improvement in Octoploid Strawberry (Fragaria × ananassa). Front Plant Sci. 2022;13. https://doi.org/10.3389/fpls.2022.869655 . Koorevaar T, Willemsen JH, Hildebrand D, Visser RGF, Arens P, Maliepaard C. How to handle high subgenome sequence similarity in allopolyploid Fragaria x ananassa: linkage disequilibrium based variant filtering. BMC Genomics. 2024;25:1150. https://doi.org/10.1186/s12864-024-10987-8 . Das S, Abecasis GR, Browning BL. Genotype Imputation from Large Reference Panels. 2018. https://doi.org/10.1146/annurev-genom-083117 Gawenda I, Thorwarth P, Günther T, Ordon F, Schmid KJ. Genome-wide association studies in elite varieties of German winter barley using single-marker and haplotype-based methods. Plant Breeding. 2015;134:28–39. https://doi.org/10.1111/pbr.12237 . Qian L, Hickey LT, Stahl A, Werner CR, Hayes B, Snowdon RJ et al. Exploring and Harnessing Haplotype Diversity to Improve Yield Stability in Crops. Front Plant Sci. 2017;8-2017. https://doi.org/10.3389/fpls.2017.01534 Helal MMU, Gill RA, Tang M, Yang L, Hu M, Yang L, et al. SNP- and Haplotype-Based GWAS of Flowering-Related Traits in Brassica napus. Plants. 2021;10. https://doi.org/10.3390/plants10112475 . Bhat JA, Yu D, Bohra A, Ganie SA, Varshney RK. Features and applications of haplotypes in crop breeding. Commun Biol. 2021;4:1266. https://doi.org/10.1038/s42003-021-02782-y . Martin M, Patterson M, Garg S, O Fischer S, Pisanti N, Klau GW, et al. WhatsHap: fast and accurate read-based phasing. bioRxiv. 2016;085050. https://doi.org/10.1101/085050 . Hofmeister RJ, Ribeiro DM, Rubinacci S, Delaneau O. Accurate rare variant phasing of whole-genome and whole-exome sequencing data in the UK Biobank. Nat Genet. 2023;55:1243–9. https://doi.org/10.1038/s41588-023-01415-w . Browning BL, Tian X, Zhou Y, Browning SR. Fast two-stage phasing of large-scale sequence data. Am J Hum Genet. 2021;108:1880–90. https://doi.org/10.1016/j.ajhg.2021.08.005 . Zhang Z, Wang A, Hu H, Wang L, Gong M, Yang Q, et al. Anim Res One Health. 2023;1:4–16. https://doi.org/10.1002/aro2.8 . The efficient phasing and imputation pipeline of low-coverage whole genome sequencing data using a high-quality and publicly available reference panel in cattle. Niehoff T, Pook T, Gholami M, Beissinger T. Imputation of low-density marker chip data in plant breeding: Evaluation of methods based on sugar beet. Plant Genome. 2022;15:e20257. https://doi.org/10.1002/tpg2.20257 . Rubinacci S, Ribeiro DM, Hofmeister RJ, Delaneau O. Efficient phasing and imputation of low-coverage sequencing data using large reference panels. Nat Genet. 2021;53:120–6. https://doi.org/10.1038/s41588-020-00756-0 . Oget-Ebrad C, Kadri NK, Moreira GCM, Karim L, Coppieters W, Georges M, et al. Benchmarking phasing software with a whole-genome sequenced cattle pedigree. BMC Genomics. 2022;23:130. https://doi.org/10.1186/s12864-022-08354-6 . Browning BL, Browning SR. Genotype error biases trio-based estimates of haplotype phase accuracy. Am J Hum Genet. 2022;109:1016–25. https://doi.org/10.1016/j.ajhg.2022.04.019 . Rousseau-Gueutin M, Lerceteau-Köhler E, Barrot L, Sargent DJ, Monfort A, Simpson D, et al. Comparative Genetic Mapping Between Octoploid and Diploid Fragaria Species Reveals a High Level of Colinearity Between Their Genomes and the Essentially Disomic Behavior of the Cultivated Octoploid Strawberry. Genetics. 2008;179:2045–60. https://doi.org/10.1534/genetics.107.083840 . Cooke TF, Yee M-C, Muzzio M, Sockell A, Bell R, Cornejo OE, et al. GBStools: A Statistical Method for Estimating Allelic Dropout in Reduced Representation Sequencing Data. PLoS Genet. 2016;12:e1005631. https://doi.org/10.1371/journal.pgen.1005631 . Hardigan MA, Feldmann MJ, Lorant A, Bird KA, Famula R, Acharya C, et al. Genome Synteny Has Been Conserved Among the Octoploid Progenitors of Cultivated Strawberry Over Millions of Years of Evolution. Front Plant Sci. 2020;10. https://doi.org/10.3389/fpls.2019.01789 . Koorevaar T, Willemsen JH, Visser RGF, Arens P, Maliepaard C. Construction of a strawberry breeding core collection to capture and exploit genetic variation. BMC Genomics. 2023;24:740. https://doi.org/10.1186/s12864-023-09824-1 . Li H. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics. 2018;34:3094–100. https://doi.org/10.1093/bioinformatics/bty191 . Pincot DDA, Ledda M, Feldmann MJ, Hardigan MA, Poorten TJ, Runcie DE et al. Social network analysis of the genealogy of strawberry: retracing the wild roots of heirloom and modern cultivars. G3 Genes|Genomes|Genetics. 2021;11:jkab015. https://doi.org/10.1093/g3journal/jkab015 Delaneau O, Coulonges C, Zagury J-F. Shape-IT: new rapid and accurate algorithm for haplotype inference. BMC Bioinformatics. 2008;9:540. https://doi.org/10.1186/1471-2105-9-540 . Lo C, Bashir A, Bansal V, Bafna V. Strobe sequence design for haplotype assembly. BMC Bioinformatics. 2011;12:S24. https://doi.org/10.1186/1471-2105-12-S1-S24 . Duitama J, McEwen GK, Huebsch T, Palczewski S, Schulz S, Verstrepen K, et al. Fosmid-based whole genome haplotyping of a HapMap trio child: evaluation of Single Individual Haplotyping techniques. Nucleic Acids Res. 2012;40:2041–53. https://doi.org/10.1093/nar/gkr1042 . Rubinacci S, Hofmeister RJ, Sousa da Mota B, Delaneau O. Imputation of low-coverage sequencing data from 150,119 UK Biobank genomes. Nat Genet. 2023;55:1088–90. https://doi.org/10.1038/s41588-023-01438-3 . Muyas F, Bosio M, Puig A, Susak H, Domènech L, Escaramis G, et al. Allele balance bias identifies systematic genotyping errors and false disease associations. Hum Mutat. 2019;40:115–26. https://doi.org/10.1002/humu.23674 . Delaneau O, Zagury J-F, Robinson MR, Marchini JL, Dermitzakis ET. Accurate, scalable and integrative haplotype estimation. Nat Commun. 2019;10:5436. https://doi.org/10.1038/s41467-019-13225-y . Feldmann MJ, Pincot DDA, Seymour DK, Famula RA, Jiménez NP, López CM, et al. A dominance hypothesis argument for historical genetic gains and the fixation of heterosis in octoploid strawberry. Genetics. 2024;228:iyae159. https://doi.org/10.1093/genetics/iyae159 . Auton A, Abecasis GR, Altshuler DM, Durbin RM, Bentley DR, Chakravarti A, et al. A global reference for human genetic variation. Nature. 2015;526:68. https://doi.org/10.1038/NATURE15393 . Zhou W, Fritsche LG, Das S, Zhang H, Nielsen JB, Holmen OL, et al. Improving power of association tests using multiple sets of imputed genotypes from distributed reference panels. Genet Epidemiol. 2017;41:744–55. https://doi.org/10.1002/GEPI.22067 . Edger PP, Poorten TJ, VanBuren R, Hardigan MA, Colle M, McKain MR, et al. Origin and evolution of the octoploid strawberry genome. Nat Genet. 2019;51:541–7. https://doi.org/10.1038/s41588-019-0356-4 . Fan Z, Liston A, Soltis DE, Soltis PS, Ashman T-L, Hummer KE et al. Homoploid hybridization adds clarity to the origins of octoploid strawberries. Proceedings of the National Academy of Sciences. 2025;122:e2502814122. https://doi.org/10.1073/pnas.2502814122 Howard NP, van de Weg E, Bedford DS, Peace CP, Vanderzande S, Clark MD, et al. Elucidation of the ‘Honeycrisp’ pedigree through haplotype analysis with a multi-family integrated SNP linkage map and a large apple (Malus×domestica) pedigree-connected SNP data set. Hortic Res. 2017;4:17003. https://doi.org/10.1038/hortres.2017.3 . Brinton J, Ramirez-Gonzalez RH, Simmonds J, Wingen L, Orford S, Griffiths S, et al. A haplotype-led approach to increase the precision of wheat breeding. Commun Biol. 2020;3:1–11. https://doi.org/10.1038/s42003-020-01413-2 . Weber SE, Frisch M, Snowdon RJ, Voss-Fels KP. Haplotype blocks for genomic prediction: a comparative evaluation in multiple crop datasets. Front Plant Sci. 2023;14:1217589. https://doi.org/10.3389/FPLS.2023.1217589 . Hess M, Druet T, Hess A, Garrick D. Fixed-length haplotypes can improve genomic prediction accuracy in an admixed dairy cattle population. Genet Selection Evol. 2017;49:1–14. https://doi.org/10.1186/s12711-017-0329-Y . Bruschi M, Bozzoli M, Ratti C, Sciara G, Goudemand E, Devaux P, et al. Dissecting the genetic basis of resistance to Soil-borne cereal mosaic virus (SBCMV) in durum wheat by bi-parental mapping and GWAS. Theor Appl Genet. 2024;137:1–23. https://doi.org/10.1007/s00122-024-04709-7 . Additional Declarations No competing interests reported. Supplementary Files S1Binning.zip S2impconc.xlsx S3QHNcomputation.zip SupplementaryFigures.docx Cite Share Download PDF Status: Published Journal Publication published 19 Nov, 2025 Read the published version in BMC Genomics → Version 1 posted Editorial decision: Revision requested 03 Oct, 2025 Reviews received at journal 02 Oct, 2025 Reviews received at journal 11 Sep, 2025 Reviewers agreed at journal 03 Sep, 2025 Reviews received at journal 30 Aug, 2025 Reviewers agreed at journal 29 Aug, 2025 Reviewers agreed at journal 28 Aug, 2025 Reviewers agreed at journal 27 Aug, 2025 Reviewers invited by journal 27 Aug, 2025 Editor invited by journal 01 Aug, 2025 Editor assigned by journal 31 Jul, 2025 Submission checks completed at journal 31 Jul, 2025 First submitted to journal 29 Jul, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7241985","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":507941395,"identity":"41421807-6837-402f-9b2f-4a02f98158b1","order_by":0,"name":"Tim Koorevaar","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAyUlEQVRIiWNgGAWjYDACZgYGxgY4rwJMMh7Ap4MHVcsZCIVfCwOyFsY2IrTYs3MnMM6osMnnl0h++Lly3uHE7Q28Bwg4jHcD44YzaZYzZ6QZS57ddjhxzgG+BMJaHrYdNjA4c8BAsnHb7cQZDDwGxGmxP3P888/GOcRq2Qiyhb3HTLKxgRgth3k3HJxxJs1A4nhPmWXDsf/GM5gJaGHvP7vxYU+FjQF/M/vmmw01abIz2HsMH+DTAgJoZjITUj8KRsEoGAWjgCAAAMbaSstPjob2AAAAAElFTkSuQmCC","orcid":"","institution":"Fresh Forward Breeding B.V.","correspondingAuthor":true,"prefix":"","firstName":"Tim","middleName":"","lastName":"Koorevaar","suffix":""},{"id":507941397,"identity":"5227cfe9-6f3f-4b32-aaa2-024e76143b80","order_by":1,"name":"Johan H. Willemsen","email":"","orcid":"","institution":"Fresh Forward Breeding B.V.","correspondingAuthor":false,"prefix":"","firstName":"Johan","middleName":"H.","lastName":"Willemsen","suffix":""},{"id":507941399,"identity":"02493691-0130-4a93-9d5b-0bc88d468bd5","order_by":2,"name":"Richard G.F. Visser","email":"","orcid":"","institution":"Wageningen University and Research Plant Breeding","correspondingAuthor":false,"prefix":"","firstName":"Richard","middleName":"G.F.","lastName":"Visser","suffix":""},{"id":507941402,"identity":"5cad367b-4632-422d-8242-d3e6f39bbafd","order_by":3,"name":"Paul Arens","email":"","orcid":"","institution":"Wageningen University and Research Plant Breeding","correspondingAuthor":false,"prefix":"","firstName":"Paul","middleName":"","lastName":"Arens","suffix":""},{"id":507941403,"identity":"297f55fa-6e13-45e5-89d5-4324f72d028e","order_by":4,"name":"Chris Maliepaard","email":"","orcid":"","institution":"Wageningen University and Research Plant Breeding","correspondingAuthor":false,"prefix":"","firstName":"Chris","middleName":"","lastName":"Maliepaard","suffix":""}],"badges":[],"createdAt":"2025-07-29 10:08:31","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7241985/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7241985/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1186/s12864-025-12270-w","type":"published","date":"2025-11-19T15:58:27+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":90541904,"identity":"a98bbbdb-6ad3-4199-8ac7-4db53db2e48c","added_by":"auto","created_at":"2025-09-03 23:58:21","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":122269,"visible":true,"origin":"","legend":"\u003cp\u003eVenn diagrams for retained (A) and discarded (B) SNPs over three post hoc filtering methods, AAB, LD and MER based.\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-7241985/v1/1868d85780c2f5c0db322669.png"},{"id":90542664,"identity":"c6f8cc77-0ec0-4f12-909b-dc7fdde2772c","added_by":"auto","created_at":"2025-09-04 00:06:21","extension":"jpeg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":673720,"visible":true,"origin":"","legend":"\u003cp\u003eA). Switch error rates compared to sequencing depth (×) for a biparental population (n=46) with Holiday and Korona as parents. B). Average depth per sample of all 740 available samples (without downsampled samples). Red lines indicate depth of 15× in both figures, which was used as filtering criterium. C). Flip errors vs. remaining switch errors over all chromosomes and all samples (for which pedigree data was available). D). Top figure: frequency of SNPs in minor allele frequency bins (x axis). Bottom figure: switch error rate counts per minor allele frequency bin.\u003c/p\u003e","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7241985/v1/3abd003f57717cf9a937f3ff.jpeg"},{"id":90541911,"identity":"e64007d1-0cb5-4030-8a84-4210b9fd97fe","added_by":"auto","created_at":"2025-09-03 23:58:21","extension":"jpeg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":36975,"visible":true,"origin":"","legend":"\u003cp\u003eSwitch error rates with varying parameter values for hmm-window (A), pbwt-window (B) and hmm-ne (C).\u003c/p\u003e","description":"","filename":"groupimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7241985/v1/d02a11170d4f0bb98729b133.jpeg"},{"id":90541924,"identity":"26c0c117-118e-422d-b273-2529ecdc8565","added_by":"auto","created_at":"2025-09-03 23:58:21","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":287352,"visible":true,"origin":"","legend":"\u003cp\u003eCumulative distribution of haplotype block lengths. The plot shows the cumulative percentage of the total phased block length as a function of block size (in Mb) across all samples and chromosomes. For each block length bin (0.01 Mb intervals), the cumulative percentage represents the fraction of the total phased sequence contained in blocks of at least that length.\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-7241985/v1/58f574187d901c3e77dc1370.png"},{"id":90541921,"identity":"285b162e-9c9d-4add-aa75-b3adba8facf0","added_by":"auto","created_at":"2025-09-03 23:58:21","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":64202,"visible":true,"origin":"","legend":"\u003cp\u003ePrincipal Components Analysis (PCA) of three groups with distinct origin (California, Florida, HCFF (High Chill)). For each group, the 25% of selected samples are highlighted in a lighter shade of their respective colour.\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-7241985/v1/8dabac2c7a121fea04917b8d.png"},{"id":90541915,"identity":"eecdd283-31d1-4174-bdc7-3644090e7a97","added_by":"auto","created_at":"2025-09-03 23:58:21","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":164283,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eConcordance rates between imputed and ground truth SNPs for samples from three population origins using five different haplotype reference panels, shown separately for A) all SNPs, B) homozygous SNPs, and C) heterozygous SNPs.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-7241985/v1/e6a77f1f70ca2cf41087eec6.png"},{"id":90541919,"identity":"87897214-38a2-4f31-b3d1-ab6f1f5bf2f8","added_by":"auto","created_at":"2025-09-03 23:58:21","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":346402,"visible":true,"origin":"","legend":"\u003cp\u003eA) Boxplots showing the distribution of heterozygous SNP concordance rates across all chromosomes, based on imputation using the Whole_pop reference panel. Each boxplot represents concordance values across all imputed samples for a given chromosome. (B) Bar plots showing the median QHN90 over all phased samples with the y-axis denoting block lengths in kilobases (kb).\u003c/p\u003e","description":"","filename":"floatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-7241985/v1/071a0fc98abf185d4857ed8b.png"},{"id":96650375,"identity":"22b6cbd4-70b1-4cc6-896e-743969d27cfa","added_by":"auto","created_at":"2025-11-24 16:11:38","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2650625,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7241985/v1/e611b2ce-ef01-4b79-a038-bd845e4cf191.pdf"},{"id":90542662,"identity":"25bea5d0-c308-4764-bd34-7a1ea84efb25","added_by":"auto","created_at":"2025-09-04 00:06:21","extension":"zip","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":3979,"visible":true,"origin":"","legend":"","description":"","filename":"S1Binning.zip","url":"https://assets-eu.researchsquare.com/files/rs-7241985/v1/c5b048ca3a90777aa3b56077.zip"},{"id":90541906,"identity":"7350e98d-8f3e-4d3f-ba3e-0449aa0b37ea","added_by":"auto","created_at":"2025-09-03 23:58:21","extension":"xlsx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":23074,"visible":true,"origin":"","legend":"","description":"","filename":"S2impconc.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-7241985/v1/ff9fe06540a05a2f0681d5ff.xlsx"},{"id":90541925,"identity":"2a5c35fd-5bda-45eb-ad72-d793857389be","added_by":"auto","created_at":"2025-09-03 23:58:21","extension":"zip","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":171186,"visible":true,"origin":"","legend":"","description":"","filename":"S3QHNcomputation.zip","url":"https://assets-eu.researchsquare.com/files/rs-7241985/v1/2c29bc58b0aa7775c4fcb78d.zip"},{"id":90541920,"identity":"4d56d286-b4ae-4808-85e8-43ac99369b4c","added_by":"auto","created_at":"2025-09-03 23:58:21","extension":"docx","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":105626,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryFigures.docx","url":"https://assets-eu.researchsquare.com/files/rs-7241985/v1/87cdaf1e303c35e69b94c9cb.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Genotype Imputation from Low-Coverage WGS Using Haplotype Reference Panels in Cultivated Strawberry","fulltext":[{"header":"Introduction","content":"\u003cp\u003eIn recent years, whole-genome sequencing (WGS) has emerged as a powerful tool for high-throughput genotyping in various crops, including garden strawberry (\u003cem\u003eFragaria \u0026times; ananassa\u003c/em\u003e) [\u003cspan additionalcitationids=\"CR2 CR3\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. WGS offers several significant advantages over traditional Single Nucleotide Polymorphism (SNP) arrays. While SNP arrays are widely used due to their relatively low cost and straightforward analysis pipelines, they present clear limitations. For instance, they suffer from ascertainment bias, as markers are pre-selected based on a small subset of samples [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. If this subset is not representative of the whole target population, it might not have captured relevant genetic diversity. Moreover, SNP arrays contain a fixed and limited set of markers that cannot be modified or extended. Therefore, they are unable to detect population-specific or novel variants.\u003c/p\u003e\u003cp\u003eIn contrast, WGS enables the detection of a much broader spectrum of genetic variation, including rare variants, and provides greater flexibility in both experimental design and downstream analysis. Although WGS remains more expensive per sample, it provides much more genomic information and becomes more accessible in the future as sequencing costs decline. With the decreasing cost of sequencing technologies and the availability of several high-quality reference genomes [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e], high-throughput genotyping in garden strawberry is increasingly shifting toward WGS-based approaches. For example, single variants identified through WGS have been used to discover and optimize diagnostic markers for red fruit colour intensity (BxANR_5UTR) and perpetual flowering (PFRU) [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. Additionally, WGS has facilitated more comprehensive genetic diversity analyses by identifying and characterizing millions of genome-wide variants [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e, \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. These developments illustrate that WGS is not only technically feasible in cultivated strawberry but also offers practical advantages in both breeding and research contexts, particularly where dense genotyping is required or where SNP arrays fall short.\u003c/p\u003e\u003cp\u003eIn addition to utilizing individual variants, genetic studies using WGS can also allow to obtain haplotypes, which offer additional advantages over single-variant analyses. Haplotypes are groups of variants (SNPs, insertions, deletions) that are inherited together from a single parent because they are located on the same chromosome. These variants will be inherited together until there has been a recombination. Utilizing haplotypes instead of biallelic SNPs after a reference panel has been established has multiple advantages. Firstly, imputation of low coverage data becomes more accurate by using a haplotype reference panel [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. This helps reduce sequencing costs, making sequence based genotyping more affordable for larger studies and breeding applications. Secondly, research shows that haplotypes can improve genome-wide association study (GWAS) results. Studies using both real and simulated data have found that haplotypes could lead to higher quantitative trait locus (QTL) mapping accuracy and greater QTL detection power compared to individual SNPs [\u003cspan additionalcitationids=\"CR12 CR13\" citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. Beyond that, haplotypes reflect identity-by-descent (IBD) rather than identity-by-state (IBS) in contrast to alleles of individual SNPs. Therefore, haplotypes can provide deeper insights into the genetic structure of a population by improving ancestry reconstruction, identifying allele-specific expression, and detecting selective sweeps with greater accuracy [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e].\u003c/p\u003e\u003cp\u003eHaplotypes can be reconstructed from sequencing data using three main approaches. The first relies on sequencing read information to link variants that appear on the same sequencing fragment. This method is particularly effective with long-read technologies such as HiFi PacBio or Nanopore, as their longer reads can connect more SNPs, resulting in longer haplotypes [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]. A second approach uses linkage information in combination with the knowledge that only one of the two alleles is passed on from a parent (pedigree data, such as parent-offspring trios), to phase genotypes. The third approach is statistical phasing, as implemented in tools like SHAPEIT5 and BEAGLE 5.4 [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e, \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e], that utilize hidden Markov models (HMM) to probabilistically infer haplotype phases. Here, the accuracy of haplotype inference mainly depends on the size of the reference panel and, to a lesser extent, their frequency distribution, as haplotypes that appear more frequently in the dataset are easier to detect and hence more reliable.\u003c/p\u003e\u003cp\u003eIn general, and for the functionally diploid \u003cem\u003eFragaria \u0026times; ananassa\u003c/em\u003e, statistical phasing based on short-read sequencing is currently the most convenient haplotyping strategy. Compared to long-read-based phasing and pedigree-based phasing, it offers a practical balance between genotyping cost and data availability. Short-read resequencing is more cost-effective and more available than long-read sequencing, making statistical phasing based on short read resequencing easier to apply in breeding programs. In addition, statistical phasing does not depend on complete pedigree records, which are often incomplete in published datasets and breeding program records. While pedigree-based phasing can provide high accuracy, its application is limited when trios are incomplete or missing. Therefore, pedigree-based phasing is often less scalable and less flexible than statistical phasing, particularly in datasets with incomplete pedigree information or missing genotypes, \u003cem\u003ee.g.\u003c/em\u003e, parents that were not genotyped and also not available anymore (so also not possible to genotype).\u003c/p\u003e\u003cp\u003ePhased haplotypes can also be used for genotype dosage imputation of low-coverage sequencing data. When phasing is accurate, and the low-coverage samples are genetically related to a well-constructed reference panel, imputation performance is generally high [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e, \u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. Conversely, imputation accuracy declines if the reference panel lacks sufficient haplotype diversity, is not population-matched, or if the sequencing depth is too low to resolve ambiguous regions. Therefore, it also allows for assessing phasing quality across samples for which no parental genomes are available [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]. Rubinacci et al. (2021) demonstrated that statistical phasing methods can lead to high imputation accuracy when large reference panels (\u0026gt;\u0026thinsp;10,000) are used in human studies [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. Similarly, high imputation performance has been achieved in agricultural species such as in cattle and sugar beet, using smaller panels (~\u0026thinsp;300\u0026ndash;1500), suggesting that statistical phasing remains effective across different reference panel sizes and species [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e, \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e].\u003c/p\u003e\u003cp\u003eHowever, haplotyping in \u003cem\u003eF. \u0026times; ananassa\u003c/em\u003e, by using statistical phasing software, is more complicated than in humans, cattle or sugar beet because of its allo-octoploid nature (2n\u0026thinsp;=\u0026thinsp;8x\u0026thinsp;=\u0026thinsp;56). Allo-octoploidy complicates genotyping and increases error rates, while high quality data is important for accurate haplotype reconstruction. Genotyping errors can lead to switch errors, which occur when the phasing between consecutive heterozygous genotypes is incorrect, causing the haplotype to jump from one homologous chromosome to the other. These errors impact downstream analyses because these can propagate through downstream analyses [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e]. Although \u003cem\u003eF. \u0026times; ananassa\u003c/em\u003e follows disomic inheritance, the four subgenomes per chromosome show high sequence similarity [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e]. This could cause issues with read alignment, variant calling accuracy (genotyping errors) and, subsequently, phasing accuracy.\u003c/p\u003e\u003cp\u003eGenotyping errors can be avoided by sequencing at high coverage and additionally, by applying variant filtering methods that improve downstream phasing accuracy [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. High coverage sequencing (\u0026gt;\u0026thinsp;40\u0026times;) in particular, reduces the likelihood of allelic dropout, which occurs when one of the two alleles at a heterozygous site is not captured during sequencing [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e] However, while high depth reduces allelic dropout (\u003cem\u003ei.e.\u003c/em\u003e detecting too few alleles), it does not correct for genotyping errors caused by subgenome sequence similarity, where reads from different homoeologous regions may be misaligned, sometimes resulting in the detection of too many alleles. These errors are especially common in \u003cem\u003eF. \u0026times; ananassa\u003c/em\u003e [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. For this reason, post hoc filtering strategies are also necessary. For example, filtering based on average allele balance (AAB) and linkage disequilibrium (LD) has been shown to reduce switch error rates by up to 44%, primarily by removing variants with elevated genotyping error rates [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. Consequently, combining high sequencing depth with robust filtering strategies offers a practical and scalable approach to reduce genotyping noise in complex polyploid genomes.\u003c/p\u003e\u003cp\u003eTo support the implementation of high-throughput sequencing-based genotyping in a \u003cem\u003eFragaria \u0026times; ananassa\u003c/em\u003e breeding context, this study focused on constructing varying haplotype reference panels and evaluating their utility for genotype dosage imputation from low-coverage (1\u0026times;) sequencing data. To achieve this, we first minimized genotyping errors through a combination of high sequencing depth and stringent variant filtering, including filters based on average allele balance (AAB), linkage disequilibrium (LD), and Mendelian error rates (MER). Second, we phased and evaluated performance using SHAPEIT5 by using switch error rate and quality haplotype N50 (QHN50). Third, because haplotype inference accuracy (and consequently genotype dosage imputation) is known to strongly depend on the size of the reference panel and, to a lesser extent, its allele frequency distribution, we assessed how the haplotype reference panel\u0026rsquo;s composition, both in size and genetic diversity, impacted the imputation accuracy of genotype dosages based on low-coverage (1\u0026times;) sequencing data. For this, sequencing reads from samples belonging to three genetically distinct populations were downsampled to 1\u0026times; coverage. Then, the genotype dosages of these downsampled samples were imputed using a range of haplotype reference panel compositions.\u003c/p\u003e"},{"header":"Materials and Methods","content":"\u003cp\u003eA haplotype reference panel for \u003cem\u003eFragaria \u0026times; ananassa\u003c/em\u003e was constructed using whole genome sequencing (WGS) data from 765 samples, where initial depth filtering (\u0026gt;\u0026thinsp;10x on chr_1A) resulted in 740 remaining samples. After further analysis of the effect of sample sequencing depth on switch error rates, we filtered even further on high depth (\u0026gt;\u0026thinsp;15\u0026times;) across all 28 chromosomes, resulting in a filtered dataset of 653 high-depth samples. To build a reliable panel, genotyping errors were reduced by combining high sequencing depth (for samples) with strict variant filtering, including average allele balance (AAB), linkage disequilibrium (LD), and Mendelian error rate (MER) filters. All retained variants were phased using SHAPEIT5. To evaluate how reference panel composition influences downstream performance, samples from three genetically distinct populations were selected and their genotype dosages were imputed using panels of varying size and composition. This setup allowed us to see how different panel configurations affected the accuracy of imputed low-coverage WGS data.\u003c/p\u003e\u003cp\u003e\u003cb\u003eGenotype data\u003c/b\u003e\u003c/p\u003e\u003cp\u003eWhole-genome sequencing (WGS) data was obtained from four distinct sources, all sequenced using Illumina 150 bp paired-end technology, though with varying sequencing depths. Two out of the four datasets were publicly available and included 137 samples from a breeding program in California [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e] and 80 samples from a breeding program in Florida [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. In this study, we refer to these sources collectively as cal_flo. Additionally, two datasets were generated in-house at Fresh Forward Breeding B.V. The first in-house dataset comprised sequenced samples from previous research projects, including a biparental mapping population (Holiday \u0026times; Korona), while the second was specifically curated to ensure representativeness of the breeding program [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e]. From these two datasets, we identified a representative group of high-chill samples, hereafter referred to as High Chill Fresh Forward (HCFF). Altogether, WGS data for 765 samples was available, predominantly consisting of \u003cem\u003eF. \u0026times; ananassa\u003c/em\u003e, but also including some \u003cem\u003eF. chiloensis\u003c/em\u003e and \u003cem\u003eF. virginiana\u003c/em\u003e samples.\u003cdiv class=\"BlockQuote\"\u003e\u003cp\u003eRaw reads were aligned to the \u003cem\u003eF. \u0026times; ananassa\u003c/em\u003e genome (FaRR1), by using \u003cem\u003eminimap2\u003c/em\u003e (v2.24) with the -ax sr preset [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]. Sambamba (v1.0.0) was used to convert the SAM output from minimap2 into BAM format while simultaneously filtering out unmapped reads using the filter -F \"(not unmapped)\". The filtered reads were then sorted with Samtools sort (v1.9). Variant calling was performed using BCFtools (v1.9) mpileup and call, applying the multiallelic caller (-mv) to generate the final VCF files. Biallelic SNPs were retained based on stringent SNP quality thresholds (QUAL\u0026thinsp;\u0026gt;\u0026thinsp;40 and MAPQ\u0026thinsp;\u0026gt;\u0026thinsp;30, DP\u0026thinsp;\u0026lt;\u0026thinsp;29555, DP\u0026thinsp;\u0026gt;\u0026thinsp;9575, and missingness rate\u0026thinsp;\u0026lt;\u0026thinsp;10%). Samples with an average depth\u0026thinsp;\u0026gt;\u0026thinsp;10\u0026times; on chr_1A were removed for downstream filtering analyses, resulting in 740 remaining samples.\u003c/p\u003e\u003cp\u003e\u003cb\u003eFiltering erroneous variants due to subgenome similarity\u003c/b\u003e\u003c/p\u003e\u003cp\u003eTo exclude erroneous SNPs resulting from the high subgenome similarity in allo-octoploid strawberry, we applied three different SNP filtering strategies. First, Average Allele Balance (AAB) filtering was used, where only SNPs with 0.45\u0026thinsp;\u0026lt;\u0026thinsp;AAB\u0026thinsp;\u0026lt;\u0026thinsp;0.57 were retained [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. Second, LD-Based filtering was used, by randomly selecting 140,000 SNPs (5,000 per chromosome) as an anchor set, and all SNPs were assigned to chromosomes based on LD to these anchor SNPs [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. SNPs were retained if they showed an LD SNP count\u0026thinsp;\u0026gt;\u0026thinsp;20 and a sum of squared LD (SSLD)\u0026thinsp;\u0026gt;\u0026thinsp;0.93. Third, Mendelian error rate (MER) filtering was applied to take advantage of available pedigree information [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e]. This method identifies variants that violate Mendelian expectations in duos and trios. Genotyping errors caused by subgenome sequence similarity also tend to result in elevated Mendelian error rates in SNPs, making MER filtering a useful addition for identifying and removing such problematic variants. As a result, 110 duos and 138 trios were used to filter out SNPs with MER\u0026thinsp;\u0026gt;\u0026thinsp;0.01, allowing a maximum of two Mendelian errors per SNP (\u003cb\u003eError! Reference source not found.\u003c/b\u003e). The final dataset, consisting of 740 samples following these subsequent filtering steps, was designated as the standard dataset.\u003c/p\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e\u003cb\u003eSNP Binning and Genetic Diversity Analysis\u003c/b\u003e\u003c/p\u003e\u003cp\u003eTo reduce the number of SNPs to a manageable level for genetic diversity analyses, we developed a binning script (\u003cem\u003eSupplementary File S1\u003c/em\u003e). This script used connected components based on linkage disequilibrium (LD) to form groups of similar SNPs. Connected components are subgraphs where any two vertices are connected by paths, helping identifying clusters of SNPs likely to be inherited together. All SNPs that were in groups smaller than 10 SNPs were removed. For groups consisting of 10 or more co-segregating SNPs, we retained the SNP with the highest QUAL score. In this study, we binned SNPs using windows of 1 Mb, a minimum LD threshold of 0.9, and a minimal group size of 10 SNPs, resulting in the binned dataset. To evaluate genetic diversity, we categorized samples by the origin of their sequencing data (California, Florida, Fresh Forward Breeding B.V.) and performed principal components analysis (PCA) using the \u003cem\u003eprcomp\u003c/em\u003e function in R to analyse the genetic diversity among these groups.\u003c/p\u003e\u003cp\u003e\u003cb\u003ePhasing optimization using high-coverage data (\u0026gt;\u0026thinsp;15 \u0026times; coverage)\u003c/b\u003e\u003c/p\u003e\u003cp\u003eAfter initial depth filtering, where samples with more than 10\u0026times; coverage on chromosome 1A were retained, an analysis was conducted to evaluate the effect of individual sequencing depth on phasing accuracy. Based on this evaluation, a threshold of 15\u0026times; average sequencing depth was selected to optimize phasing accuracy, as informed by the initial switch error rate analysis. Specifically, the average depth per SNP was calculated for each sample in the standard SNP dataset. Samples with an average depth below 15\u0026times; were subsequently excluded, resulting in a final set of 653 high depth samples for further analysis.\u003c/p\u003e\u003cp\u003eGiven the impracticality of pedigree- or long-read-based haplotyping at scale, statistical phasing of short-read resequencing data stands out as the most feasible and effective strategy, especially with the availability of WGS data for 740 samples from both public repositories and in-house collections at Fresh Forward Breeding B.V. [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e, \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. Therefore, phasing was performed using SHAPEIT (v5.1.1). To optimize its performance, the influence of three different parameters on phasing accuracy was investigated [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]. First, 101 parents from duos and trios present in this panel were removed from the standard dataset, so that these offspring are phased without information about their parental haplotypes. Subsequently, the remaining samples were phased using SHAPEIT5 with varying stepwise parameter values for the parameters hmm-ne, hmm-window, and pbwt-window (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e).\u003c/p\u003e\u003cp\u003ePhasing accuracy was assessed by computing the switch error rate (SER) for phased samples belonging to duos or trios, using the unphased genotypic data of their parents [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e]. The total number of switch errors across samples was divided by the total number of evaluated heterozygous SNPs to obtain the overall switch error rate (SER). The SER was computed separately for each of the 28 chromosomes, with each chromosome phased in an independent run. Although the phasing algorithm is stochastic and might introduce variability at the level of individual chromosomes, aggregating results across all chromosomes offsets this variability. Therefore, repeated iterations were not needed, as the conclusions were based on the combined results from all 28 independent runs. This approach allowed us to systematically evaluate the effect of different SHAPEIT5 parameter settings on phasing accuracy.\u003c/p\u003e\u003cp\u003eTo better understand the effect of genotyping errors and low MAF on phasing quality, we first separated flip errors from the total number of switch errors. Specifically, flip errors are double switch errors, where the haplotype switches and then immediately switches back to the correct state. These are typically caused by genotyping errors or low-frequency variants. Next, we investigated the relationship between flip errors and remaining switch errors. Additionally, both were compared across minor allele frequency bins to assess how allele frequency contributes to phasing accuracy.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eSHAPEIT5 parameters and their tested range for optimization of phasing accuracy.\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"2\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eSHAPEIT5 parameter\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eTest range\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eHmm-ne\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e500\u0026ndash;500,000 (step size: starting with 500 and increasing up to 400,000, default: 15,000)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eHmm-window\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e1\u0026ndash;10 (step size: 1, default: 4)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ePbwt-window\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e1\u0026ndash;10 (step size: 1, default: 4)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eIn addition to evaluating phasing accuracy using the switch error rate, an alternative metric was proposed by Lo et al. (2011), known as the adjusted N50 (AN50). AN50 is defined as the minimum haplotype block length such that 50 percent of all phased variants are located in blocks of at least this size [\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e]. Building on this,Duitama et al. (2012) introduced the quality adjusted N50 (QAN50), which follows the same principle but further divides each haplotype block into the longest sub-blocks without any detected switch errors [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]. In this study, we computed a similar metric, but with two key differences: we did not adjust for the number of variants per haplotype block, and we did not evaluate the haplotype length based on the number of variants. Instead, we focused on the proportion of the genome that was phased without switch errors. We refer to this measure as the quality haplotype N50 (QHN50). To calculate it, all haplotype blocks were ordered from largest to smallest, and the cumulative proportion of the total phased sequence covered by haplotypes of a given length was computed across all samples and chromosomes. The resulting QHN50 represents the haplotype length at which 50 percent of the phased sequence is covered by blocks of at least that length.\u003c/p\u003e\u003cp\u003e\u003cb\u003eSelection of samples for imputation\u003c/b\u003e\u003c/p\u003e\u003cp\u003eTo assess the influence of genetic diversity on phasing accuracy, we evaluated phasing accuracy indirectly through genotype dosage imputation accuracy of downsampled sequencing data (to 1\u0026times; coverage) from various samples. Although estimating phasing accuracy directly by computing the switch error rate is straightforward and efficient, this approach is limited to populations with known pedigrees, which were only partly available for two datasets in this study. Therefore, we relied on an indirect approach: imputing genotype dosages in downsampled samples using a phased reference panel, where higher accuracy reflects better phasing quality. This approach allows phasing performance to be evaluated beyond pedigreed samples and across genetically diverse backgrounds [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e].\u003c/p\u003e\u003cp\u003eA total of 66 samples were selected for downsampling and imputation from three genetically distinct populations (California, Florida, and High Chill Fresh Forward (HCFF)) each with a clearly defined origin. Within each population, 25% of samples were randomly selected (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). To simulate low-coverage whole genome sequencing, the fraction of reads required to reach 1\u0026times; coverage was first calculated for each sample based on the total sequencing depth across all chromosomes. Subsequently, the corresponding BAM files were downsampled using \u003cem\u003esamtools view -s\u003c/em\u003e with the computed read fraction, resulting in BAM files with approximately 1\u0026times; depth.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eOverview of population sizes and selected samples for imputation.\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"5\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003enr\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eOrigin population\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eTotal population size (\u0026gt;\u0026thinsp;15x depth)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eSelected samples for imputation (~\u0026thinsp;25% of total population size)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eRemaining samples for reference panels\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eCalifornia\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e32\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e8\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e24\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eFlorida\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e62\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e16\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e46\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e3\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eHCFF\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e169\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e42\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e127\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e\u003cb\u003ePhasing and imputation using different reference panels\u003c/b\u003e\u003c/p\u003e\u003cp\u003eTo investigate how haplotype reference panel size and composition affect phasing and low-coverage imputation accuracy, we tested six different haplotype reference panels. In addition to using the entire population (\u0026gt;\u0026thinsp;15\u0026times; depth) as a reference panel, various combinations of the three populations (California, Florida, and High Chill Fresh Forward (HCFF)) were tested. Due to SHAPEIT5's minimum sample size requirement of 50 samples, reference panels consisting solely of the California (24) or Florida (46) populations were excluded (Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). In all cases, the downsampled samples (66) were always excluded from all haplotype reference panels.\u003c/p\u003e\u003cp\u003eTo further separate the effect of panel diversity from panel size, we randomly selected 70 samples from the HCFF population, matching the size of the combined California\u0026thinsp;+\u0026thinsp;Florida panel. This allowed us to compare panels of equal size but different genetic diversity. Additionally, to address whether it is preferable to maximize the number of samples or prioritize haplotype quality, we also included an extra reference panel composed of the entire population without applying the minimum sample depth filter (15\u0026times;).\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eOverview of reference panels used for imputation, where the downsampled samples were excluded from the reference panels.\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"3\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003enr\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eReference panel\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eReference panel size (number of samples)\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eCalifornia\u0026thinsp;+\u0026thinsp;Florida\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e70\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eHCFF small\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e70\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e3\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eHCFF\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e127\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e4\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eHCFF\u0026thinsp;+\u0026thinsp;California\u0026thinsp;+\u0026thinsp;Florida\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e197\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e5\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eWhole population\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e587\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eExtra\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eWhole population (no 15\u0026times; DP filter)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e674\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eEach reference panel was phased independently using SHAPEIT5 to prepare for downstream imputation. Default parameters were used except for the \u003cem\u003ehmmne\u003c/em\u003e, which was set to 7500. Additionally, the default MCMC iteration scheme was extended (10b,1p,1b,1p,1b,1p,1b,1p,10m) to improve phasing accuracy, following the recommendations by Hofmeister et al. (2023), though at the cost of increased runtime [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]. The resulting phased reference panels were then used as input for GLIMPSE2 to impute the genotype dosages of downsampled samples. GLIMPSE2, recommended by the SHAPEIT5 authors for low-coverage WGS imputation, uses a haplotype reference panel to refine genotype likelihoods and infer missing genotypes [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e, \u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e]. All imputation runs were performed using GLIMPSE2 with default settings.\u003c/p\u003e\u003cp\u003e\u003cb\u003eImputation accuracy evaluation\u003c/b\u003e\u003c/p\u003e\u003cp\u003eImputed genotype dosages were compared against the original high-depth genotype dosages from the standard dataset (ground truth). Imputation accuracy was assessed using concordance rates, computed for each imputed sample using each reference panel. The concordance rate measures the proportion of correctly imputed genotype dosages relative to the total number of imputed dosages.\u003c/p\u003e\u003cp\u003eSince heterozygous dosages are more difficult to impute than homozygous dosages in low-coverage WGS data, concordance rates were separately calculated for homozygous and heterozygous dosages. Following Zhang et al. (2023), we used a 3\u0026times;3 cross-classification table to compare the imputed genotype dosages with the ground truth. In this table, homozygous reference genotypes (aa) correspond to dosage 0, heterozygous genotypes (ab) to dosage 1, and homozygous alternative genotypes (bb) to dosage 2 [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. Each cell, \u003cem\u003en\u003c/em\u003e\u003csub\u003e\u003cem\u003eij\u003c/em\u003e\u003c/sub\u003e, represents the number of SNPs where the true dosage is \u003cem\u003ei\u003c/em\u003e and the imputed dosage is \u003cem\u003ej\u003c/em\u003e (Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e). Using these tables, an overall summary was created by computing concordance rates per sample and per reference panel, aggregated across all chromosomes. Additionally, these concordance values were separated for homozygous and heterozygous ground truth calls to better capture differences in imputation performance between variant types.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eMarginal cross-classification table of true and imputed genotypes from a single sample.\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"5\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eImputed genotypes\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colspan=\"4\" nameend=\"c5\" namest=\"c2\"\u003e\u003cp\u003eTrue genotypes\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eaa\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eab\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003ebb\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eTotal\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eaa\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cem\u003en\u003c/em\u003e\u003csub\u003e\u003cem\u003e11\u003c/em\u003e\u003c/sub\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e\u003cem\u003en\u003c/em\u003e\u003csub\u003e\u003cem\u003e12\u003c/em\u003e\u003c/sub\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e\u003cem\u003en\u003c/em\u003e\u003csub\u003e\u003cem\u003e13\u003c/em\u003e\u003c/sub\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e\u003cem\u003en\u003c/em\u003e\u003csub\u003e\u003cem\u003e1.\u003c/em\u003e\u003c/sub\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eab\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cem\u003en\u003c/em\u003e\u003csub\u003e\u003cem\u003e21\u003c/em\u003e\u003c/sub\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e\u003cem\u003en\u003c/em\u003e\u003csub\u003e\u003cem\u003e22\u003c/em\u003e\u003c/sub\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e\u003cem\u003en\u003c/em\u003e\u003csub\u003e\u003cem\u003e23\u003c/em\u003e\u003c/sub\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e\u003cem\u003en\u003c/em\u003e\u003csub\u003e\u003cem\u003e2.\u003c/em\u003e\u003c/sub\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ebb\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cem\u003en\u003c/em\u003e\u003csub\u003e\u003cem\u003e31\u003c/em\u003e\u003c/sub\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e\u003cem\u003en\u003c/em\u003e\u003csub\u003e\u003cem\u003e32\u003c/em\u003e\u003c/sub\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e\u003cem\u003en\u003c/em\u003e\u003csub\u003e\u003cem\u003e33\u003c/em\u003e\u003c/sub\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e\u003cem\u003en\u003c/em\u003e\u003csub\u003e\u003cem\u003e3.\u003c/em\u003e\u003c/sub\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eTotal\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cem\u003en\u003c/em\u003e\u003csub\u003e\u003cem\u003e.1\u003c/em\u003e\u003c/sub\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e\u003cem\u003en\u003c/em\u003e\u003csub\u003e\u003cem\u003e.2\u003c/em\u003e\u003c/sub\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e\u003cem\u003en\u003c/em\u003e\u003csub\u003e\u003cem\u003e.3\u003c/em\u003e\u003c/sub\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e\u003cem\u003en..\u003c/em\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003e\u003cb\u003eImpact of filtering methods on SNP retention\u003c/b\u003e\u003c/p\u003e\u003cp\u003eApplying three filtering methods resulted in the retention of 7.22M SNPs over all chromosomes (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eA). Individually, strict thresholds for Average Allele Balance (AAB), LD-based filtering, and Mendelian error rates (MER) retained 11.47M, 11.72M, and 13.49M SNPs, respectively. While LD-based filtering and MER retained similar SNP counts, MER retained slightly more overall. Despite this, each method captured distinct SNP sets, with 7.22M SNPs shared across all three. On average, this corresponds to 258K SNPs per chromosome. Notably, AAB filtering retained the fewest unique SNPs compared to the other two methods.\u003c/p\u003e\u003cp\u003eTo further examine differences between the filtering methods, we analysed the SNPs that were discarded (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eB). All three methods agreed on removing 2.69M SNPs, suggesting these were the most evident unreliable SNPs. The largest overlap between two methods was between AAB and LD-based filtering, which together flagged an additional 2.60M SNPs for removal. In contrast, the smallest overlap was between MER and LD-based filtering, with only 0.44M additional SNPs identified. Each method also flagged at least 1.66M SNPs for removal that were not identified by the other two, highlighting the different criteria of each method.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003e\u003cb\u003eImpact of sequencing depth on phasing accuracy\u003c/b\u003e\u003c/p\u003e\u003cp\u003eTo assess the impact of sequencing depth on phasing accuracy, we first phased the whole population (n\u0026thinsp;=\u0026thinsp;740) and evaluated a biparental population, in which the noise from dataset variability and genetic diversity is minimized (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA). A clear trend is visible between the switch error rate and the sequencing depth of the phased offspring samples, with the lowest switch error rates observed at sequencing depths higher than 25\u0026times;. However, 356 samples in the whole WGS dataset have depths below 25\u0026times;. To balance between maximizing the number of samples and ensuring phasing accuracy, we filtered for an average sequencing depth of at least 15\u0026times; per sample, removing 87 samples and resulting in a final dataset of 640 samples (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eB).\u003c/p\u003e\u003cp\u003e\u003cb\u003eLinear relationship between flip errors and remaining switch errors\u003c/b\u003e\u003c/p\u003e\u003cp\u003eTo better estimate the switch error rate, flip errors were first identified and removed from the total set of switch errors. Flip errors, defined as double switch events where the haplotype switches for one SNP, were separated to isolate remaining switch errors. A linear relationship was observed between flip error rates and remaining switch error rates across samples (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eC). To further examine this pattern, the relationship between switch error rates and minor allele frequency (MAF) was assessed. Switch errors were not enriched in the lowest MAF bin, and their frequency remained relatively consistent across MAF bins (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eD). These results suggest that flip errors and remaining switch errors are correlated and that switch errors do not predominantly occur at low MAF sites.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003e\u003cb\u003eOptimized SHAPEIT5 parameters\u003c/b\u003e\u003c/p\u003e\u003cp\u003eThree parameters were tested stepwise to evaluate their impact on phasing accuracy and determine the optimal settings for phasing \u003cem\u003eF. \u0026times; ananassa\u003c/em\u003e WGS data. First, we assessed the hmm-window parameter, which showed minimal variation in phasing accuracy. The lowest accuracy was observed at values of 1 and 2 cM, but overall, there was no strong justification to deviate from the default setting of 4 cM (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eA). Similarly, testing different pbwt-window settings revealed no substantial effect on phasing accuracy, leading to the retention of the default 4 cM value (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eB). Moreover, for the effective population size (hmm-ne), accuracy was slightly better at 7,500 compared to the default of 15,000 (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eC). Although the difference was small, we opted to set hmm-ne to 7,500 for downstream analyses, as it showed a marginal improvement in some cases, which resulted in a mean switch error rate of 0.92%.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003e\u003cb\u003eAccurate Haplotype Block Length Distribution\u003c/b\u003e\u003c/p\u003e\u003cp\u003eTo evaluate the practical usefulness of the haplotypes in the reference panel, we analysed the distribution of haplotype block lengths without switch errors. Flip errors were first removed to isolate uninterrupted haplotype blocks, separated only by single switch errors. Most blocks were relatively short, with a median length of 30 kb and a mean of 155 kb. To assess how block size contributes to overall phasing, all blocks were ranked by length, and the cumulative proportion of the total phased sequence covered by blocks of at least a given size was computed across all samples and chromosomes (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e). This analysis yielded a mean quality haplotype N50 (QHN50) of 654 kb and a mean QHN90 of 101 kb. These results show that a substantial portion of the phased genome is composed of long, uninterrupted haplotype blocks.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003e\u003cb\u003eImputation panel composition\u003c/b\u003e\u003c/p\u003e\u003cp\u003eA principal components analysis (PCA) was conducted using the binned dataset of all 740 samples to visualize genetic relationships (\u003cem\u003eFigure \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e\u003c/em\u003e). Here, subpopulations were highlighted to define groups based on their genetic origin. Several samples belonging to \u003cem\u003eF. virginiana\u003c/em\u003e (dark blue) and \u003cem\u003eF. chiloensis\u003c/em\u003e (purple) were present in the dataset and seemed to strongly influence PC2. PC1 appears to reflect chilling requirement, with low-chill populations clustering at the left side of the plot and high-chill populations on the right. Additionally, two clusters of samples belonging to two biparental populations were visible in the lower-right and middle-right areas of the plot.\u003c/p\u003e\u003cp\u003eStill, three distinct groups could be defined (based on the data origin). California and Florida formed separate, well-defined clusters, both at the left side of PC1. In contrast, a large group of high chill samples from Fresh Forward (HCFF) was identified as counterpart of these two low chill clusters (Florida and California). The biparental groups from HCFF were excluded from further analysis to avoid bias introduced by many samples with low genetic variation.\u003c/p\u003e\u003cp\u003eTo validate the distinctness of these three groups, a separate principal components analysis was conducted with just these three groups (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e). The groups remained well separated, although a few samples with Florida origin were in the HCFF group, and one HCFF sample was positioned near the California group. Additionally, the groups varied in size: California (32 samples), Florida (62 samples), and HCFF (169 samples). For each group, 25% of the samples were randomly selected to test imputation accuracy based on different haplotype reference panels (highlighted in Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e). The PCA plot confirmed that these samples were representative of their respective groups (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e).\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003e\u003cb\u003eGenetic Diversity Across Reference Panels Assessed by Proportion of Polymorphic SNPs\u003c/b\u003e\u003c/p\u003e\u003cp\u003eTo complement insights from principal components analysis (PCA), we quantified genetic diversity by calculating the proportion of polymorphic SNPs within each reference panel. This summary metric provides an informative overview of within-panel variation without compensating for panel size. As expected, the Whole_pop_alldp panel showed the highest diversity, with 100% of SNPs being polymorphic, since SNP discovery was performed across this entire dataset, including low-coverage samples, but excluding the selected samples for imputation tests. In contrast, smaller subsets of this panel showed less polymorphism. Specifically, the cal_flo panel contained only 83% polymorphic SNPs. However, another panel of identical size, HCFF_small, retained 98% polymorphic SNPs, suggesting that population composition, rather than panel size alone, influenced the observed diversity. Additionally, all other panels (HCFF, HCFF_cal_flo, and Whole_pop) displayed consistently high diversity, with the proportion of polymorphic SNPs ranging from 0.99 to 1.00.\u003c/p\u003e\u003cp\u003e\u003cb\u003eImputation accuracy across different haplotype reference panels\u003c/b\u003e\u003c/p\u003e\u003cp\u003eTo assess imputation accuracy, selected samples from different groups were imputed using IMPUTE2 with haplotype reference panels of varying sizes and genetic diversity [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e, \u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e]. Overall, imputation accuracy was high, with total concordance rates exceeding 0.92 for most samples (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eA). However, an exception was observed when HCFF samples were imputed using the cal_flo reference panel, which resulted in lower concordance rates (mean concordance of 0.90). Additionally, three HCFF outliers showed consistently lower concordance across all reference panels.\u003c/p\u003e\u003cp\u003eTo better understand these trends with respect to variant types, we analysed concordance rates separately for homozygous (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eB) and heterozygous calls (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eC). As expected, homozygous calls exhibited higher concordance than heterozygous calls. Moreover, population-level differences in imputation accuracy were apparent. For example, California samples consistently showed the highest homozygous SNP concordance across all panels, followed by Florida and HCFF samples. Genotype dosage imputation performance for homozygous sites also varied depending on the reference panel. Specifically, cal_flo again resulted in the lowest homozygous concordance for HCFF samples. In contrast, when HCFF samples were included in the reference panel, concordance differences across populations were minimal.\u003c/p\u003e\u003cp\u003eIn contrast to the relatively consistent homozygous results, heterozygous imputation rates showed greater variation across both panels and populations. Among all reference sets, cal_flo again produced the lowest concordance rates across all samples, followed by HCFF_small, HCFF, HCFF_cal_flo, and Whole_pop. The highest accuracy was achieved using the Whole_pop_alldp panel (the panel that also included samples that were removed with the \u0026lt;\u0026thinsp;15x coverage filtering). Importantly, imputation accuracy of Whole_pop and Whole_pop_alldp remained largely comparable despite the latter including lower-depth samples. Specifically, total concordance increased only slightly from 0.964 to 0.965, a change mainly driven by a small improvement in heterozygous call accuracy (from 0.884 to 0.886), while homozygous concordance remained nearly unchanged.\u003c/p\u003e\u003cp\u003eAdditionally, the population origin of imputed samples significantly influenced heterozygous performance. When using the Whole_pop panel, Florida samples exhibited the highest median heterozygous concordance, followed by HCFF and California samples. Interestingly, although California samples had the lowest heterozygous accuracy among the three groups, they achieved the highest overall imputation concordance. This outcome can be attributed to the predominance of homozygous reference calls as these calls showed higher concordance rates, with California samples having more homozygous reference calls than Florida, and Florida more than HCFF (\u003cem\u003eSupplementary File S2\u003c/em\u003e).\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003e\u003cb\u003eHigher imputation accuracy for\u003c/b\u003e \u003cb\u003eF. vesca\u003c/b\u003e \u003cb\u003ederived (A) subgenomes\u003c/b\u003e\u003c/p\u003e\u003cp\u003eTo investigate subgenome-specific patterns in imputation performance, we analysed heterozygous concordance rates per chromosome across all imputed samples by using Whole_pop as reference panel. This chromosome-level resolution provided insight into how different subgenomes contribute to overall accuracy. Interestingly, the \u003cem\u003eF. vesca\u003c/em\u003e derived A subgenomes consistently demonstrated the highest concordance rates across all chromosomes. Moreover, their median heterozygous concordance exceeded 0.90 in every case, clearly distinguishing them from other subgenomes, where median concordance values were ranging between 0.85 and 0.90, with no specific chromosomes deviating substantially from this pattern (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eA).\u003c/p\u003e\u003cp\u003eTo explore whether higher imputation accuracy also corresponds to improved phasing quality, we computed the median QHN90 score for each chromosome using the standard dataset where samples with \u0026lt;\u0026thinsp;15\u0026times; coverage were removed (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eB). If phasing were more accurate for the A subgenomes, one would expect to observe higher QHN90 values for these chromosomes. However, the pattern was less consistent. Although chromosomes 2A, 3A, 4A, 5A, and 7A showed relatively high QHN90 scores, this trend did not hold across all A subgenomes. In particular, chromosomes 1A and 6A did not have the highest QHN90 scores among their respective subgenome groups, suggesting that phasing accuracy may not fully explain the observed differences in imputation performance.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eThis study aimed to construct a high-quality haplotype reference panel for \u003cem\u003eFragaria \u0026times; ananassa\u003c/em\u003e using whole-genome sequencing (WGS) data and to evaluate its utility for accurate dosage imputation from low-coverage sequencing. We found that combining high sequencing depth with rigorous variant filtering (based on allele balance, linkage disequilibrium, and Mendelian errors) enabled statistical phasing using SHAPEIT5 with switch error rates around 1%, corresponding to an average QHN50 of 654 kb. The results also showed that once a reference panel (using\u0026thinsp;\u0026gt;\u0026thinsp;15\u0026times; sequencing data) capturing most, if not all, possible alleles is established, low coverage (1\u0026times;) sequenced varieties can be imputed with high accuracy for downstream applications. This accuracy was maintained with only minimal loss, even when using smaller or less genetically diverse haplotype reference panels. These findings demonstrate that high-throughput WGS-based genotyping is a viable strategy for allo-polyploid crops like cultivated strawberry, where a single, well-constructed haplotype panel can support robust dosage imputation across genetically diverse breeding varieties. This work extends previous findings in diploid species and simpler genomes by showing that statistical phasing and imputation are also effective under the genomic complexity of allo-octoploids.\u003c/p\u003e\u003cp\u003e\u003cb\u003eGenotyping errors are main cause of switch errors\u003c/b\u003e\u003c/p\u003e\u003cp\u003eAlthough extensive filtering was applied to the WGS dataset in \u003cem\u003eF. \u0026times; ananassa\u003c/em\u003e, switch errors remained. First, three filtering methods were used to obtain a reliable SNP dataset across all samples, including those without pedigree data. Second, samples with insufficient sequencing depth (\u0026lt;\u0026thinsp;15\u0026times;) were removed. However, a switch error rate of 0.008\u0026ndash;0.011 was still observed across all chromosomes, meaning an error occurred roughly every 100 heterozygous SNPs. One possible explanation could be that population size and diversity influenced switch error rates. However, increasing the reference panel from HCFF (n\u0026thinsp;=\u0026thinsp;142) to the Whole Population (n\u0026thinsp;=\u0026thinsp;640) led to only a marginal improvement in imputation accuracy. This suggests that population size and genetic diversity were only small factors in haplotyping accuracy in our dataset.\u003c/p\u003e\u003cp\u003eInstead, the presence of a linear relationship between flip (\u003cem\u003ei.e.\u003c/em\u003e double switch errors) and remaining switch errors (\u003cem\u003ei.e.\u003c/em\u003e errors spanning at least two consecutive heterozygous SNPs) points to a shared underlying cause. Since flip errors are typically associated with low MAF SNPs and genotyping errors, we examined whether switch errors were enriched in low MAF variants. Although the MAF\u0026thinsp;\u0026lt;\u0026thinsp;0.05 bin contained the highest number of SNPs, switch errors were not disproportionately concentrated in this bin. Consequently, genotyping errors, rather than MAF, appear to be the main factor driving switch errors. Moreover, switch error rates could only marginally be improved by optimizing three phasing parameters, which aligns with expectations: phasing parameter adjustments cannot correct switch errors that are caused by genotyping errors.\u003c/p\u003e\u003cp\u003eGiven these findings, an approach to further reduce genotyping errors could be to focus on variant call accuracy. While filtering had already been applied at the whole SNP and sample levels, no additional filtering was done after the default variant calling with \u003cem\u003ebcftools\u003c/em\u003e. Prioritizing missing variant calls over unreliable ones may be more effective when performing phasing. One approach could be to filter based on a minimum depth threshold for variant calls, but this threshold would need to be adjusted for samples that were sequenced with lower depth. Without such an adjustment, samples with lower sequencing depths could end up with highly skewed missingness rates. Another option is to filter based on unbalanced allele frequencies at heterozygous sites (Allelic Bias) [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e, \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e]. This is similar to the Average Allele Bias filtering used in this study but applied at the individual call level rather than the population level. Additionally, setting thresholds for Phred-scaled likelihoods (PL) or genotype quality (GQ) on the genotype dosage calls could further reduce variant call errors, helping to reduce switch errors and ultimately improve imputation accuracy. However, when combining multiple variant call filtering strategies, care must be taken not to remove too many calls, as retaining a sufficient number of informative sites is important for effective phasing and imputation. Since SHAPEIT5 includes internal imputation for missing variant calls during phasing, some missing data is tolerated [\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e]. Still, it remains unclear at what level of variant call missingness phasing performance begins to decline.\u003c/p\u003e\u003cp\u003e\u003cb\u003ePotential Bias from Mendelian Filtering in Switch Error Rate Estimation\u003c/b\u003e\u003c/p\u003e\u003cp\u003eTo filter SNPs, we applied Mendelian error rates (MER), which may have introduced bias in the evaluation of switch error rates. This is because switch error rates were calculated downstream using duos and trios, the same data that is used for MER filtering. However, we still considered SNP-level MER filtering useful, because if erroneous (based on MER) genotype dosages cluster in one SNP, other genotype dosages for the same SNP that could not be evaluated due to missing pedigree data are also more likely to be erroneous. To minimize the potential MER bias on SER, we set a maximum Mendelian error rate of 1%, ensuring that the effect on switch error rate estimation remained limited. Despite this, some degree of bias is inevitable. Therefore, we used imputation accuracy as an additional proxy for haplotype accuracy, as errors in haplotypes directly reduce imputation performance [\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e].\u003c/p\u003e\u003cp\u003e\u003cb\u003eReference genome relatedness drives imputation accuracy differences\u003c/b\u003e\u003c/p\u003e\u003cp\u003eImputation accuracy varied across samples from different origins, with clear differences observed between homozygous and heterozygous calls, which showed mean concordance rates of 0.99 and 0.88, respectively, when using the Whole_pop reference panel. California samples showed the highest homozygosity, followed by Florida, while HCFF samples were the most heterozygous. Consequently, total concordance rates were highest for California, followed by Florida and then HCFF samples. Increased homozygosity in Californian breeding material was also reported by Feldmann et al. (2024) who attributed it to selection, random genetic drift, and selective sweeps [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e].\u003c/p\u003e\u003cp\u003eInterestingly, homozygous calls were more accurately imputed for California samples than for Florida or HCFF samples, with the largest gap observed between California and HCFF. This is likely due to the reference genome used for alignment (FaRR1), which originates from California and is genetically closer to California samples than to those from Florida or HCFF [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. As a result, reads containing reference alleles are more likely to align correctly, improving variant calling, phasing, and ultimately imputation, an effect known as reference genome bias. Indeed, imputed California samples had a much higher proportion of homozygous reference calls than homozygous alternative calls, compared to both Florida and HCFF samples. Moreover, homozygous reference calls generally had higher imputation accuracy than homozygous alternative calls (\u003cem\u003eSupplementary File S1\u003c/em\u003e). Similarly, Florida samples had more homozygous reference calls than HCFF, suggesting they are also more closely related to the reference genome. These results highlight the significant impact of both homozygosity and reference genome relatedness on imputation accuracy, as well as on haplotype reconstruction, since homozygous regions do not require phasing and therefore do not introduce switch errors.\u003c/p\u003e\u003cp\u003eIn contrast to the trend observed for homozygous calls, where California samples showed the highest imputation accuracies, heterozygous imputation rates followed a different pattern. On average, Florida samples had higher imputation rates for heterozygous sites than California samples regardless of the haplotype reference panel used. Given that heterozygous calls always include one reference allele (matching the reference genome), one might expect that phasing or imputing the other allele would be easier when the reference genome is closely related. Yet, our results did not support this assumption. One possible explanation is that the California reference population included in the haplotype reference panel was relatively small (n\u0026thinsp;=\u0026thinsp;24) compared to the Florida (n\u0026thinsp;=\u0026thinsp;42) and HCFF (n\u0026thinsp;=\u0026thinsp;142) populations. If certain alternative alleles from the to-be-imputed California samples were not represented in the reference panel, this could have limited imputation accuracy. Since a shared haplotype can only be found if it is present in the haplotype reference panel [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e, \u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e], this limitation may have contributed to the lower heterozygous imputation rates observed for California samples. Another possible explanation is that relevant haplotypes were present in the reference panel but contained more switch errors. Notably, the average sequencing depth of the ground truth dataset for California samples (20.7\u0026times;) was lower than that of Florida (32.5\u0026times;) and HCFF (28.0\u0026times;) samples. Since lower sequencing depth is associated with elevated switch error rates [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e], this may have introduced more phasing errors in the California-specific haplotypes. Therefore, these haplotypes may have been less reliable, ultimately leading to lower imputation accuracy for rare or population specific haplotypes in California samples.\u003c/p\u003e\u003cp\u003e\u003cb\u003eGenetic diversity and size of reference panel drive imputation accuracy\u003c/b\u003e\u003c/p\u003e\u003cp\u003eTo investigate how panel size and genetic diversity influence imputation accuracy, the downsampled samples were imputed using five different haplotype reference panels. Panel size was consistently associated with improved imputation accuracy. Even adding 87 previously filtered samples with lower sequencing depth (\u0026lt;\u0026thinsp;15\u0026times;) to the Whole_pop reference panel led to a slight increase in imputation performance. Therefore, while sequencing at a depth greater than 15\u0026times; is recommended, including some samples with lower coverage does not negatively impact overall accuracy.\u003c/p\u003e\u003cp\u003eTo assess the effect of genetic diversity, we constructed a smaller HCFF panel with the same sample size as the cal_flo panel. The HCFF_small panel had a much higher proportion of polymorphic SNPs (98%) than the cal_flo panel (83%), indicating greater genetic diversity. This difference likely reflects the number of breeding objectives in each region: while California and Florida breeding programs mainly target low-chill, open-field cultivation, the HCFF panel includes high-chill selections from multiple cultivation systems such as open field, greenhouse (spring), and greenhouse (winter).\u003c/p\u003e\u003cp\u003eIn terms of imputation outcomes, HCFF samples imputed with the HCFF_small panel showed notably higher accuracy compared to those imputed with the cal_flo panel. In contrast, for California and Florida samples, imputation performance was similar between the two panels, with slightly higher imputation rates for HCFF_small. This suggests that the genetic diversity required for accurate imputation of California and Florida samples is sufficiently captured in both panels. However, for accurate imputation of HCFF samples, the required haplotypes are better represented in the HCFF_small panel. Nevertheless, it is worth noting that the cal_flo panel still achieved a median total imputation rate of nearly 0.97 for HCFF samples, with heterozygous concordance rates exceeding 0.75. Together, these results highlight the importance of both genetic diversity and panel size for accurate imputation. At the same time, they also show the robustness of this approach, as accurate imputation of low-coverage WGS data is achievable across various panel configurations.\u003c/p\u003e\u003cp\u003e\u003cb\u003eConcordance outliers\u003c/b\u003e\u003c/p\u003e\u003cp\u003eThree outliers were identified, each showing lower total concordance rates compared to the other samples. All three outliers were from HCFF, which also had the highest number of imputed samples (42). For one of these, the pedigree data had been removed during a prior pedigree check due to elevated Mendelian errors. Elevated Mendelian errors can arise from various factors, such as errors in pedigree records or sample swaps. While the exact cause of this outlier remains uncertain, its combination with a very high heterozygosity rate suggests the possibility of severe contamination. The second outlier did not exhibit elevated Mendelian errors (with both parents known) but still displayed a high heterozygosity percentage. This could simply be due to high heterozygosity affecting the overall concordance rates. The third outlier was a parent in multiple other trios, but no elevated pedigree errors were detected. This sample also showed a slightly higher heterozygosity rate. In this case, contamination with another strawberry sample seems the most likely explanation, as it would not cause additional pedigree errors but could result in erroneous haplotypes.\u003c/p\u003e\u003cp\u003e\u003cb\u003eHigher imputation accuracy for the\u003c/b\u003e \u003cb\u003eF. vesca\u003c/b\u003e\u003cb\u003e-derived A subgenomes\u003c/b\u003e\u003c/p\u003e\u003cp\u003eChromosome-level imputation results revealed a consistent pattern in which the A subgenomes, derived from \u003cem\u003eF. vesca\u003c/em\u003e, showed higher imputation accuracy than the B, C, and D subgenomes. This elevated accuracy may be explained by the A subgenome being the most recently incorporated to the octoploid strawberry, and therefore showing less genetic variation (and therefore less SNPs) and lower haplotype diversity than the B, C and D subgenomes [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e, \u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e, \u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e]. These factors may have contributed to more accurate variant calls and, combined with lower haplotype diversity, resulted in higher quality haplotypes. Supporting this, several, but not all, A subgenomes also showed higher QHN90 scores (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eB), suggesting longer contiguous haplotype blocks. Moreover, lower haplotype diversity on the A subgenomes may increase the chance that haplotypes to be imputed are already represented in the haplotype reference panel. In addition, these characteristics may also have contributed to a higher assembly quality of the A subgenomes in the reference genome, further improving variant calling and imputation performance.\u003c/p\u003e\u003cp\u003e\u003cb\u003eHaplotype reference panel beyond imputation\u003c/b\u003e\u003c/p\u003e\u003cp\u003eIn this study, we demonstrated that haplotype reference panels can be effectively used to impute low-coverage (1\u0026times;) WGS data in allo-octoploid strawberry. However, besides genotype dosage imputation, such haplotype reference panels are also valuable in breeding applications that require accurate haplotypes. For instance, they can support haplotype-based parental checks and pedigree reconstruction, as shown in fruit crops like apple, where shared haplotype blocks were used to resolve unknown parentage and trace ancestry across generations [\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e]. Haplotype structure has also been used to inform the selection of elite haplotypes across germplasm pools, as shown in wheat through a haplotype-led breeding framework that improved mapping resolution and selection precision [\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e]. Additionally, haplotypes can improve genomic prediction by capturing local linkage patterns and multi-allelic effects that are often missed by single SNPs. Studies in canola, wheat, soybean, and dairy cattle have shown that using haplotypes as predictors can lead to higher genomic prediction accuracy, especially for complex traits [\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e, \u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e]. Moreover, haplotypes were useful for the development of diagnostic markers. For example, in wheat, LD-defined haploblocks were used to fine-map a disease resistance QTL and design a minimal marker panel that could trace the resistant haplotype with high specificity [\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e].\u003c/p\u003e\u003cp\u003eFor many of these applications, long and accurate haplotype blocks are important. In our study, we achieved a mean QHN50 of 654 kb, indicating that half of the phased genome is represented in blocks of at least this length without switch errors. This level of haplotype contiguity is likely sufficient for tasks such as haplotype tracing, genomic prediction, or diagnostic marker development, where full chromosome-level phasing is not strictly necessary.\u003c/p\u003e"},{"header":"Conclusions","content":"\u003cp\u003eIn this study, we constructed a haplotype reference panel for \u003cem\u003eFragaria \u0026times; ananassa\u003c/em\u003e using whole-genome sequencing data and a statistical phasing algorithm (SHAPEIT5). High sequencing depth (\u0026gt;\u0026thinsp;15\u0026times;), combined with strict variant filtering based on average allele balance, linkage disequilibrium, and Mendelian error rate, was important to reduce genotyping errors and improve phasing accuracy. Although optimization of phasing parameters had limited effect, these filtering strategies helped reduce switch errors and enabled accurate phasing, with switch error rates of approximately 0.01. This corresponded to haplotype blocks with a mean QHN50 of 654 kb, meaning that half of the phased genome was covered by blocks of at least this length without switch errors. If even longer haplotype blocks are required, phasing accuracy can still be improved by further reducing genotyping errors. For constructing haplotype reference panels, sequencing at a depth of at least 25\u0026times; is recommended (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA). However, available samples with sequencing depth lower than 25\u0026times; can still be valuable and should be included.\u003c/p\u003e\u003cp\u003eFurthermore, we constructed haplotype reference panels with varying genetic diversity and size, to investigate their influence on low-coverage WGS genotype dosage imputation. As expected, larger panels yielded better results. However, genetic diversity only had a limited effect, where less diverse panels had lower imputation accuracies for samples whose haplotypes were not represented well in the panel. In addition, the A subgenome consistently showed higher imputation accuracy than the B, C, and D subgenomes, which may relate to its dominant role in the genome. Together, these findings provide practical guidelines for building haplotype reference panels and show that statistical phasing and imputation of low-coverage WGS data is robust across different panel configurations. The approach is also applicable to other crops, where sequencing around 50 genetically representative samples at \u0026ge;\u0026thinsp;25\u0026times; depth can enable accurate dosage imputation from low-coverage data.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eCompeting interests\u003c/h2\u003e\n\u003cp\u003eThe authors declare no competing interests.\u003c/p\u003e\n\u003ch2\u003eFunding\u003c/h2\u003e\n\u003cp\u003eThis project was supported by Fresh Forward Breeding B.V. and the TKI project \u0026lsquo;LWV20.112 Application of sequence-based multi-allelic markers in genetics and breeding of polyploids\u0026rsquo; (BO-68\u0026ndash;001-042-WPR).\u003c/p\u003e\n\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\n\u003cp\u003eT.K. led the conceptualization, investigation, software implementation, visualization, and drafting of the manuscript. T.K., J.H.W., P.A., R.G.F.V., and C.M. contributed to the conceptualization, methodology and reviewed and provided feedback on manuscript drafts. All authors approved the final version of the manuscript and are accountable for their contributions.\u003c/p\u003e\n\u003ch2\u003eAcknowledgement\u003c/h2\u003e\n\u003cp\u003eWe would like to thank Fresh Forward Breeding B.V. for providing the sequencing data used in this study, and Rian Peters for DNA isolation of the Fresh Forward core collection.\u003c/p\u003e\n\u003ch2\u003eData Availability\u003c/h2\u003e\n\u003cp\u003eThe California and Florida sequencing data are publicly available from their respective studies [4, 25]. The Fresh Forward Breeding B.V. datasets generated and/or analysed during the current study are not publicly available because they are owned by private company Fresh Forward Breeding B.V. but can be requested from Fresh Forward Breeding B.V. Scripts for computing average allele balance and LD-based filtering are available in a previous publication [9]. SHAPEIT5 is available at https://odelaneau.github.io/shapeit5/. GLIMPSE2 is available at https://odelaneau.github.io/GLIMPSE/. Script for binning is available in Supplementary File S1. Script for QHN computation is available in Supplementary File S3.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eSaiga S, Tada M, Segawa T, Sugihara Y, Nishikawa M, Makita N, et al. NGS-based genome wide association study helps to develop co-dominant marker for the physical map-based locus of PFRU controlling flowering in cultivated octoploid strawberry. Euphytica. 2022;219:6. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s10681-022-03132-7\u003c/span\u003e\u003cspan address=\"10.1007/s10681-022-03132-7\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHardigan MA, Lorant A, Pincot DDA, Feldmann MJ, Famula RA, Acharya CB, et al. Unraveling the Complex Hybrid Ancestry and Domestication History of Cultivated Strawberry. Mol Biol Evol. 2021;38:2285\u0026ndash;305. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1093/molbev/msab024\u003c/span\u003e\u003cspan address=\"10.1093/molbev/msab024\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLiu S, Martin KE, Snelling WM, Long R, Leeds TD, Vallejo RL et al. Accurate genotype imputation from low-coverage whole-genome sequencing data of rainbow trout. G3 Genes|Genomes|Genetics. 2024;14:jkae168. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1093/g3journal/jkae168\u003c/span\u003e\u003cspan address=\"10.1093/g3journal/jkae168\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eFan Z, Whitaker VM. Genomic signatures of strawberry domestication and diversification. Plant Cell. 2024;36:1622\u0026ndash;36. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1093/plcell/koad314\u003c/span\u003e\u003cspan address=\"10.1093/plcell/koad314\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGeibel J, Reimer C, Pook T, Weigend S, Weigend A, Simianer H. How imputation can mitigate SNP ascertainment Bias. BMC Genomics. 2021;22:340. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/s12864-021-07663-6\u003c/span\u003e\u003cspan address=\"10.1186/s12864-021-07663-6\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHardigan MA, Feldmann MJ, Pincot DDA, Famula RA, Vachev MV, Madera MA et al. Blueprint for Phasing and Assembling the Genomes of Heterozygous Polyploids: Application to the Octoploid Genome of Strawberry. bioRxiv. 2021;:2021.11.03.467115. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1101/2021.11.03.467115\u003c/span\u003e\u003cspan address=\"10.1101/2021.11.03.467115\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHan H, Barbey CR, Fan Z, Verma S, Whitaker VM, Lee S. Telomere-to-Telomere and Haplotype-Phased Genome Assemblies of the Heterozygous Octoploid \u0026lsquo;Florida Brilliance\u0026rsquo; Strawberry (Fragaria \u0026times; ananassa). bioRxiv. 2022;:2022.10.05.509768. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1101/2022.10.05.509768\u003c/span\u003e\u003cspan address=\"10.1101/2022.10.05.509768\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLabadie M, Vallin G, Potier A, Petit A, Ring L, Hoffmann T, et al. High Resolution Quantitative Trait Locus Mapping and Whole Genome Sequencing Enable the Design of an Anthocyanidin Reductase-Specific Homoeo-Allelic Marker for Fruit Colour Improvement in Octoploid Strawberry (Fragaria \u0026times; ananassa). Front Plant Sci. 2022;13. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3389/fpls.2022.869655\u003c/span\u003e\u003cspan address=\"10.3389/fpls.2022.869655\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKoorevaar T, Willemsen JH, Hildebrand D, Visser RGF, Arens P, Maliepaard C. How to handle high subgenome sequence similarity in allopolyploid Fragaria x ananassa: linkage disequilibrium based variant filtering. BMC Genomics. 2024;25:1150. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/s12864-024-10987-8\u003c/span\u003e\u003cspan address=\"10.1186/s12864-024-10987-8\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDas S, Abecasis GR, Browning BL. Genotype Imputation from Large Reference Panels. 2018. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1146/annurev-genom-083117\u003c/span\u003e\u003cspan address=\"10.1146/annurev-genom-083117\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGawenda I, Thorwarth P, G\u0026uuml;nther T, Ordon F, Schmid KJ. Genome-wide association studies in elite varieties of German winter barley using single-marker and haplotype-based methods. Plant Breeding. 2015;134:28\u0026ndash;39. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1111/pbr.12237\u003c/span\u003e\u003cspan address=\"10.1111/pbr.12237\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eQian L, Hickey LT, Stahl A, Werner CR, Hayes B, Snowdon RJ et al. Exploring and Harnessing Haplotype Diversity to Improve Yield Stability in Crops. Front Plant Sci. 2017;8-2017. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3389/fpls.2017.01534\u003c/span\u003e\u003cspan address=\"10.3389/fpls.2017.01534\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHelal MMU, Gill RA, Tang M, Yang L, Hu M, Yang L, et al. SNP- and Haplotype-Based GWAS of Flowering-Related Traits in Brassica napus. Plants. 2021;10. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3390/plants10112475\u003c/span\u003e\u003cspan address=\"10.3390/plants10112475\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBhat JA, Yu D, Bohra A, Ganie SA, Varshney RK. Features and applications of haplotypes in crop breeding. Commun Biol. 2021;4:1266. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/s42003-021-02782-y\u003c/span\u003e\u003cspan address=\"10.1038/s42003-021-02782-y\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMartin M, Patterson M, Garg S, O Fischer S, Pisanti N, Klau GW, et al. WhatsHap: fast and accurate read-based phasing. bioRxiv. 2016;085050. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1101/085050\u003c/span\u003e\u003cspan address=\"10.1101/085050\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHofmeister RJ, Ribeiro DM, Rubinacci S, Delaneau O. Accurate rare variant phasing of whole-genome and whole-exome sequencing data in the UK Biobank. Nat Genet. 2023;55:1243\u0026ndash;9. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/s41588-023-01415-w\u003c/span\u003e\u003cspan address=\"10.1038/s41588-023-01415-w\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBrowning BL, Tian X, Zhou Y, Browning SR. Fast two-stage phasing of large-scale sequence data. Am J Hum Genet. 2021;108:1880\u0026ndash;90. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.ajhg.2021.08.005\u003c/span\u003e\u003cspan address=\"10.1016/j.ajhg.2021.08.005\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eZhang Z, Wang A, Hu H, Wang L, Gong M, Yang Q, et al. Anim Res One Health. 2023;1:4\u0026ndash;16. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1002/aro2.8\u003c/span\u003e\u003cspan address=\"10.1002/aro2.8\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. The efficient phasing and imputation pipeline of low-coverage whole genome sequencing data using a high-quality and publicly available reference panel in cattle.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eNiehoff T, Pook T, Gholami M, Beissinger T. Imputation of low-density marker chip data in plant breeding: Evaluation of methods based on sugar beet. Plant Genome. 2022;15:e20257. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1002/tpg2.20257\u003c/span\u003e\u003cspan address=\"10.1002/tpg2.20257\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRubinacci S, Ribeiro DM, Hofmeister RJ, Delaneau O. Efficient phasing and imputation of low-coverage sequencing data using large reference panels. Nat Genet. 2021;53:120\u0026ndash;6. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/s41588-020-00756-0\u003c/span\u003e\u003cspan address=\"10.1038/s41588-020-00756-0\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eOget-Ebrad C, Kadri NK, Moreira GCM, Karim L, Coppieters W, Georges M, et al. Benchmarking phasing software with a whole-genome sequenced cattle pedigree. BMC Genomics. 2022;23:130. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/s12864-022-08354-6\u003c/span\u003e\u003cspan address=\"10.1186/s12864-022-08354-6\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBrowning BL, Browning SR. Genotype error biases trio-based estimates of haplotype phase accuracy. Am J Hum Genet. 2022;109:1016\u0026ndash;25. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.ajhg.2022.04.019\u003c/span\u003e\u003cspan address=\"10.1016/j.ajhg.2022.04.019\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRousseau-Gueutin M, Lerceteau-K\u0026ouml;hler E, Barrot L, Sargent DJ, Monfort A, Simpson D, et al. Comparative Genetic Mapping Between Octoploid and Diploid Fragaria Species Reveals a High Level of Colinearity Between Their Genomes and the Essentially Disomic Behavior of the Cultivated Octoploid Strawberry. Genetics. 2008;179:2045\u0026ndash;60. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1534/genetics.107.083840\u003c/span\u003e\u003cspan address=\"10.1534/genetics.107.083840\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eCooke TF, Yee M-C, Muzzio M, Sockell A, Bell R, Cornejo OE, et al. GBStools: A Statistical Method for Estimating Allelic Dropout in Reduced Representation Sequencing Data. PLoS Genet. 2016;12:e1005631. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1371/journal.pgen.1005631\u003c/span\u003e\u003cspan address=\"10.1371/journal.pgen.1005631\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHardigan MA, Feldmann MJ, Lorant A, Bird KA, Famula R, Acharya C, et al. Genome Synteny Has Been Conserved Among the Octoploid Progenitors of Cultivated Strawberry Over Millions of Years of Evolution. Front Plant Sci. 2020;10. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3389/fpls.2019.01789\u003c/span\u003e\u003cspan address=\"10.3389/fpls.2019.01789\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKoorevaar T, Willemsen JH, Visser RGF, Arens P, Maliepaard C. Construction of a strawberry breeding core collection to capture and exploit genetic variation. BMC Genomics. 2023;24:740. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/s12864-023-09824-1\u003c/span\u003e\u003cspan address=\"10.1186/s12864-023-09824-1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLi H. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics. 2018;34:3094\u0026ndash;100. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1093/bioinformatics/bty191\u003c/span\u003e\u003cspan address=\"10.1093/bioinformatics/bty191\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePincot DDA, Ledda M, Feldmann MJ, Hardigan MA, Poorten TJ, Runcie DE et al. Social network analysis of the genealogy of strawberry: retracing the wild roots of heirloom and modern cultivars. G3 Genes|Genomes|Genetics. 2021;11:jkab015. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1093/g3journal/jkab015\u003c/span\u003e\u003cspan address=\"10.1093/g3journal/jkab015\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDelaneau O, Coulonges C, Zagury J-F. Shape-IT: new rapid and accurate algorithm for haplotype inference. BMC Bioinformatics. 2008;9:540. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/1471-2105-9-540\u003c/span\u003e\u003cspan address=\"10.1186/1471-2105-9-540\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLo C, Bashir A, Bansal V, Bafna V. Strobe sequence design for haplotype assembly. BMC Bioinformatics. 2011;12:S24. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/1471-2105-12-S1-S24\u003c/span\u003e\u003cspan address=\"10.1186/1471-2105-12-S1-S24\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDuitama J, McEwen GK, Huebsch T, Palczewski S, Schulz S, Verstrepen K, et al. Fosmid-based whole genome haplotyping of a HapMap trio child: evaluation of Single Individual Haplotyping techniques. Nucleic Acids Res. 2012;40:2041\u0026ndash;53. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1093/nar/gkr1042\u003c/span\u003e\u003cspan address=\"10.1093/nar/gkr1042\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRubinacci S, Hofmeister RJ, Sousa da Mota B, Delaneau O. Imputation of low-coverage sequencing data from 150,119 UK Biobank genomes. Nat Genet. 2023;55:1088\u0026ndash;90. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/s41588-023-01438-3\u003c/span\u003e\u003cspan address=\"10.1038/s41588-023-01438-3\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMuyas F, Bosio M, Puig A, Susak H, Dom\u0026egrave;nech L, Escaramis G, et al. Allele balance bias identifies systematic genotyping errors and false disease associations. Hum Mutat. 2019;40:115\u0026ndash;26. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1002/humu.23674\u003c/span\u003e\u003cspan address=\"10.1002/humu.23674\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDelaneau O, Zagury J-F, Robinson MR, Marchini JL, Dermitzakis ET. Accurate, scalable and integrative haplotype estimation. Nat Commun. 2019;10:5436. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/s41467-019-13225-y\u003c/span\u003e\u003cspan address=\"10.1038/s41467-019-13225-y\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eFeldmann MJ, Pincot DDA, Seymour DK, Famula RA, Jim\u0026eacute;nez NP, L\u0026oacute;pez CM, et al. A dominance hypothesis argument for historical genetic gains and the fixation of heterosis in octoploid strawberry. Genetics. 2024;228:iyae159. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1093/genetics/iyae159\u003c/span\u003e\u003cspan address=\"10.1093/genetics/iyae159\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAuton A, Abecasis GR, Altshuler DM, Durbin RM, Bentley DR, Chakravarti A, et al. A global reference for human genetic variation. Nature. 2015;526:68. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/NATURE15393\u003c/span\u003e\u003cspan address=\"10.1038/NATURE15393\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eZhou W, Fritsche LG, Das S, Zhang H, Nielsen JB, Holmen OL, et al. Improving power of association tests using multiple sets of imputed genotypes from distributed reference panels. Genet Epidemiol. 2017;41:744\u0026ndash;55. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1002/GEPI.22067\u003c/span\u003e\u003cspan address=\"10.1002/GEPI.22067\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eEdger PP, Poorten TJ, VanBuren R, Hardigan MA, Colle M, McKain MR, et al. Origin and evolution of the octoploid strawberry genome. Nat Genet. 2019;51:541\u0026ndash;7. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/s41588-019-0356-4\u003c/span\u003e\u003cspan address=\"10.1038/s41588-019-0356-4\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eFan Z, Liston A, Soltis DE, Soltis PS, Ashman T-L, Hummer KE et al. Homoploid hybridization adds clarity to the origins of octoploid strawberries. Proceedings of the National Academy of Sciences. 2025;122:e2502814122. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1073/pnas.2502814122\u003c/span\u003e\u003cspan address=\"10.1073/pnas.2502814122\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHoward NP, van de Weg E, Bedford DS, Peace CP, Vanderzande S, Clark MD, et al. Elucidation of the \u0026lsquo;Honeycrisp\u0026rsquo; pedigree through haplotype analysis with a multi-family integrated SNP linkage map and a large apple (Malus\u0026times;domestica) pedigree-connected SNP data set. Hortic Res. 2017;4:17003. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/hortres.2017.3\u003c/span\u003e\u003cspan address=\"10.1038/hortres.2017.3\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBrinton J, Ramirez-Gonzalez RH, Simmonds J, Wingen L, Orford S, Griffiths S, et al. A haplotype-led approach to increase the precision of wheat breeding. Commun Biol. 2020;3:1\u0026ndash;11. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/s42003-020-01413-2\u003c/span\u003e\u003cspan address=\"10.1038/s42003-020-01413-2\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWeber SE, Frisch M, Snowdon RJ, Voss-Fels KP. Haplotype blocks for genomic prediction: a comparative evaluation in multiple crop datasets. Front Plant Sci. 2023;14:1217589. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3389/FPLS.2023.1217589\u003c/span\u003e\u003cspan address=\"10.3389/FPLS.2023.1217589\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHess M, Druet T, Hess A, Garrick D. Fixed-length haplotypes can improve genomic prediction accuracy in an admixed dairy cattle population. Genet Selection Evol. 2017;49:1\u0026ndash;14. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/s12711-017-0329-Y\u003c/span\u003e\u003cspan address=\"10.1186/s12711-017-0329-Y\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBruschi M, Bozzoli M, Ratti C, Sciara G, Goudemand E, Devaux P, et al. Dissecting the genetic basis of resistance to Soil-borne cereal mosaic virus (SBCMV) in durum wheat by bi-parental mapping and GWAS. Theor Appl Genet. 2024;137:1\u0026ndash;23. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s00122-024-04709-7\u003c/span\u003e\u003cspan address=\"10.1007/s00122-024-04709-7\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"bmc-genomics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"gics","sideBox":"Learn more about [BMC Genomics](http://bmcgenomics.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/gics","title":"BMC Genomics","twitterHandle":"#BMCGenomics","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"WGS, whole genome sequencing, SHAPEIT5, phasing, haplotype reference panel, low-coverage, imputation, strawberry, genetic diversity, QHN50","lastPublishedDoi":"10.21203/rs.3.rs-7241985/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7241985/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e\u003cp\u003eTo implement high-throughput sequencing-based genotyping in a strawberry (\u003cem\u003eFragaria \u0026times; ananassa\u003c/em\u003e) breeding program, we aimed to construct a haplotype reference panel and explore its utility through genotype dosage imputation of low-coverage (1\u0026times;) sequencing data. Although genotyping by whole genome sequencing (WGS) provides high SNP density, its cost remains a limitation for large-scale application. Imputation from low coverage data using a reference panel offers a cost effective alternative, but this approach has not yet been optimized for allo-octoploid strawberry.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e\u003cp\u003eTo reduce genotyping errors that limit phasing accuracy, we combined high sequencing depth (\u0026gt;\u0026thinsp;15\u0026times;) with variant filtering based on average allele balance (AAB), linkage disequilibrium (LD), and Mendelian error rates (MER). Statistical phasing using SHAPEIT5 resulted in a mean switch error rate of 0.9%, with 50% of the genome covered by haplotype blocks of at least 654 kb (QHN50) without phase switches. To evaluate downstream imputation, samples from three genetically distinct populations (California, Florida, and HCFF) were downsampled to 1\u0026times; and imputed using reference panels of varying size and composition (via GLIMPSE2). Both panel size and genetic diversity influenced imputation accuracy, with concordance rates ranging from 0.87 to 0.97 for the smallest panel and 0.94 to 0.98 for the largest, excluding three outliers.\u003c/p\u003e\u003ch2\u003eConclusions\u003c/h2\u003e\u003cp\u003eThese findings demonstrate that constructing a large, genetically diverse haplotype reference panel improves genotype dosage imputation from low-coverage sequencing data. However, high accuracy is still achievable with limited resources, making this a cost-efficient alternative to SNP arrays when adopting WGS-based genotyping in breeding programs. The strategy is broadly applicable to other crops where dense genotyping is needed but resources are limited. In such cases, sequencing approximately 50 genetically representative samples at \u0026ge;\u0026thinsp;25\u0026times; depth is recommended for building a reference panel suitable for imputation.\u003c/p\u003e","manuscriptTitle":"Genotype Imputation from Low-Coverage WGS Using Haplotype Reference Panels in Cultivated Strawberry","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-09-03 23:58:16","doi":"10.21203/rs.3.rs-7241985/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2025-10-03T07:16:03+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-10-02T09:09:44+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-09-12T03:54:44+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"144058817887852867592946689523549292201","date":"2025-09-03T15:00:40+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-08-31T02:30:45+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"110841300817169767876896696327202210472","date":"2025-08-29T23:37:37+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"230079875934608095285934985986196293118","date":"2025-08-29T03:33:44+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"188682200943021703891734723701475649900","date":"2025-08-27T05:42:32+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-08-27T04:58:49+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2025-08-01T06:58:27+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-08-01T00:15:56+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-08-01T00:15:23+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Genomics","date":"2025-07-29T10:01:23+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"bmc-genomics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"gics","sideBox":"Learn more about [BMC Genomics](http://bmcgenomics.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/gics","title":"BMC Genomics","twitterHandle":"#BMCGenomics","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"0cd74009-06e8-46e8-abfc-f4dfb2923f0a","owner":[],"postedDate":"September 3rd, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2025-11-24T16:06:47+00:00","versionOfRecord":{"articleIdentity":"rs-7241985","link":"https://doi.org/10.1186/s12864-025-12270-w","journal":{"identity":"bmc-genomics","isVorOnly":false,"title":"BMC Genomics"},"publishedOn":"2025-11-19 15:58:27","publishedOnDateReadable":"November 19th, 2025"},"versionCreatedAt":"2025-09-03 23:58:16","video":"","vorDoi":"10.1186/s12864-025-12270-w","vorDoiUrl":"https://doi.org/10.1186/s12864-025-12270-w","workflowStages":[]},"version":"v1","identity":"rs-7241985","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7241985","identity":"rs-7241985","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.