Beyond SNPs: Scalable Detection of Structural Variants Unlocks Hidden Genetic Diversity in Tomato | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Beyond SNPs: Scalable Detection of Structural Variants Unlocks Hidden Genetic Diversity in Tomato Reza Shekasteband This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8628781/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Structural variants (SVs)-large genomic alterations such as insertions, deletions, duplications, and translocations-are widespread in tomato genomes and play a critical role in phenotypic diversity. However, their detection has traditionally depended on expensive long-read sequencing technologies. Consistent with previous studies on structural variation, this study demonstrates a cost-effective approach for SV discovery using repurposed short-read sequencing data (150 bp), enabling integration of SVs into breeding workflows without additional sequencing investment. Using Illumina whole-genome data from 60 diverse tomato lines, including wild accessions, landraces, transgenic lines, and modern breeding lines, we identified over 71,000 high-confidence SVs, including a significant number of private doubletons with the Manta caller, as well as 10.9 million short genetic variants. SVs were unevenly distributed across chromosomes, clustering in subtelomeric regions and near disease-resistance loci, with chromosomes 6, 7, and 9 showing the highest densities. Wild accessions harbored nearly twice as many SVs as cultivated lines, with deletions dominating wild genomes and insertions more prevalent in cultivated tomatoes than in the reference genome. Comparative phylogenetic analysis revealed strong concordance between SV-based and SNP/InDel-based trees (Baker’s γ = 0.95), while SV data improved pedigree-consistent clustering and resolved ambiguous lineage relationships. These findings highlight SVs as hidden drivers of tomato diversity and valuable resources for marker-assisted selection, trait mapping, and genomic studies. Mining archived short-read datasets could offer breeding programs a scalable, low-cost strategy to unlock latent SV information, accelerate genetic improvement, and enhance genome-to-phenome insights. Limitations include under-detection of complex rearrangements, warranting targeted validation for critical loci. Structural variants (SVs) Tomato genomics Short-read sequencing Marker-assisted breeding Genetic diversity Manta SV caller Phylogenetic analysis Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 1. Introduction Tomato ( Solanum lycopersicum L.), one of the most important vegetable crops in the world, has been scrutinized with next-generation sequencing (NGS) technologies and other genomic tools soon after the genome of the inbred processing tomato line ‘Heinz1706’ was sequenced (Sato et al., 2012 ). In recent years, hundreds of wild, landrace, heirloom, and modern tomato genomes have been re-sequenced, and millions of short genetic variants (single-nucleotide variants, SNVs, and short insertions and deletions, InDels) were identified, underlining the vast genetic variation in the tomato clade within the genus Solanum (Gao et al., 2019 ). Ease of generating ample short genetic variation genotypic data from mapping populations and genome-wide association studies (GWAS) led to genetic mapping of numerous traits and disease resistance, QTL mapping, and consequently, functional gene discoveries in tomatoes (Pascual et al., 2016 ; Prasanna et al., 2023 ). Subsequently, breeding tomatoes for traits of interest and disease resistance has been accelerated significantly by exploiting DNA markers developed during the process (Prasanna et al., 2023 ). Structural variants (SVs), which occur at high frequencies, are also an important part of the genome and have gained rapid research attention in the next-generation sequencing (NGS) era (Jobson & Roberts, 2022 ). Sizable variations in the physical architecture of the genome, SVs, are pervasive features that significantly impact an organism's phenotype. Any genomic variations larger than 30 base pairs fall into this category. Large deletions (DEL) and insertions (INS) are the predominant types of SVs, followed by duplication (DUP), breakends (BND), and other genome modifications (Alonge et al., 2020 ; Fuentes et al., 2019 ). Identification and precise physical mapping of SVs have been difficult due to their size and complex nature. Thus, modern genomic studies rely heavily on short genetic variants (SNPs and short InDels) for population genetics, marker-trait association studies, GWAS, and QTL mapping. Recent advances in whole-genome resequencing technologies capable of capturing long reads and new computational tools have enabled SV discovery, opening a new frontier in genetic studies. According to multiple studies, SVs contribute significantly to genomic variation in humans. It has been shown that the impact of SVs on human genomic variation is 3.4 times that of SNVs (Huddleston et al., 2017). The average genomic variation between two humans due to SVs is estimated to be 15 times that due to SNVs, which are 1.5% and 0.1%, respectively (Pang et al., 2010). Sizable genomic variations associated with phenotypes are also common in plants. For example, Fuentes et al. ( 2019 ) identified more than 63 million SVs in 3,000 rice genomes. Furthermore, this study found that long SVs are prevalent in promoter regions, whereas short SVs are prevalent in the 5′ UTR (Fuentes et al., 2019 ). Structural variants are present throughout the genome; however, studies show that they occur at high frequency in centromeric and subtelomeric regions, underscoring the need to map their physical locations (Turner et al., 2007 ). Thousands of SVs have been identified in tomatoes, emphasizing the importance of the considerable genomic variation in the genetic architecture of this vegetable crop (Alonge et al., 2020 ; N. Li et al., 2023 ; Liu et al., 2022 ). Additionally, multiple structural variants associated with plant growth habits, fruit size, flavor, disease resistance, and pathogen recognition have been identified in tomatoes (Gao et al., 2019 ; Jobson & Roberts, 2022 ; M. B. Lee et al., 2022 ; Wang et al., 2020 ). Present among modern tomato cultivars with low frequencies, SVs play a substantial role in the enormous genetic diversity present between wild and modern tomato accessions (Alonge et al., 2020 ). Structural variants can extend our understanding of genome-to-phenome interactions, explaining the genetics behind phenotypes that are not possible to elucidate based on short genetic variants (Zhou et al., 2022 ). To facilitate the deployment of widespread SVs in tomato genomics and breeding, it is crucial to adopt affordable, straightforward yet reliable approaches to identify, map, and utilize SVs in public and private plant breeding programs. This study used short sequencing reads (150 base pairs) generated through common whole-genome sequencing approaches to identify a large number of SVs, map them to the tomato genome, and explore their potential for marker-assisted tomato breeding and genomic studies. This study also demonstrates that the contribution of structural variants to genetic diversity is significant enough to serve as the sole source for phylogenetic studies, highlighting the importance of SVs in the genetic architecture of modern tomatoes. 2. Materials and Methods 2.1. Plant material A core collection of 60 tomato lines, comprising wild accessions, landraces, specialty types, transgenic, and modern tomato lines, was used in this study. The wild tomato accessions and landrace tomatoes were obtained from the Tomato Genetics Resource Center (Davis, CA, USA) and the USDA's National Plant Germplasm System (NPGS). Modern and specialty-type tomato inbred lines have been developed and released by the tomato breeding program at North Carolina State University (NC State University, 2024 ). The tomato breeding program at NCSU is located in the Mountain Horticultural Crops Research and Extension Center (MHCREC) in Mills River, NC. In the last 5 decades, more than 70 elite tomato breeding lines and 35 F 1 fresh-market tomato cultivars have been developed, commercialized (e.g.,Gardner, 1982 ; Gardner & Panthee, 2012 ), and transferred to public and commercial tomato research programs (NC State University, 2024 ). The core tomato collection in this study encompasses a diverse range of horticultural characteristics (Supplementary Table 1). 2.2. Whole genome sequencing and short read alignment Leaf samples were bulked from six seedlings at the 4-week stage for genomic DNA extraction. Genomic DNA was extracted with a DNeasy Plant Mini Kit (Qiagen, Germantown, MD). It was sent to a Novogene Co., Inc. (Beijing, China) facility in California, USA, for sequencing library preparation and whole-genome sequencing. Standard Illumina DNA libraries were prepared with an insert size of 350 bp and then sequenced on an Illumina HiSeq instrument (Illumina, San Diego, CA), producing 150-bp paired-end reads with, on average, 22 Gb of raw data per sample. Raw sequencing reads were aligned using Bowtie2 (Langmead & Salzberg, 2012 ) to the published ‘Heinz1706’ reference genome sequence (Sato et al., 2012 ) and then sorted and indexed with SAMtools [version 1.9, (Li et al., 2009 )] 2.3. Genetic variant calling Single-nucleotide polymorphisms (SNPs) and short insertions and deletions (InDels) were called from aligned reads individually and soft-filtered using Linux system commands, BCFtools [version 1.9, (Danecek et al., 2021 )]. The merge option in BCFtool was used to merge individual files into a single VCF file. Additionally, BCFtools and VCFtools [version 0.1.13, (Danecek et al., 2011 )]were used to extract variant statistics, filter VCF files, and identify private doubletons. 2.4. Structural variant calling Four independent rounds of workflow were executed to call SVs from aligned short reads (150 bp) using the Manta structural variant caller with internal default settings (Chen et al., 2016 ). Manta can call SV on individual or small sets of samples. In the first round (R1), SV calling was performed on individual BAM files, generating a single VCF file per sample. Small sets of samples were batched for the next three rounds (R2, R3, and R4) to run a single Manta workflow. Tomato lines with a similar pedigree, fruit shapes, and sizes were grouped to run a single round of SV calling (R2). During the third round (R3), lines were randomly assigned to a batch of 15 samples. The sample assignment for the fourth round (R4) was more complex. For this round, seven lines (round, plum, and grape tomato lines) with and without common disease resistance genes were kept fixed in each batch. Then, eight random samples were assigned to the batch to call and use SVs from. To call SVs from the three wild tomato accessions in this round, the seven fixed lines were batched only with the wild tomato accessions. Any genomic variations larger than 50 base pairs were classified as SVs and categorized into four major categories. Large deletions (DEL), insertions (INS), transversions (BND), and duplication (DUP). Chromosome translocations, BND-type SVs, were categorized into two subcategories based on the relocation of the segment within the same chromosome (intrachromosomal translocations, BNDa) or between nonhomologous chromosomes (interchromosomal translocations, BNDb). At the final step, the ‘merge’ option in BCFtools was used to merge individual VCF files, generating a single merged VCF file per round. 2.5. Comparative genomics visualization To generate the short genetic variant and SV density histograms, each tomato chromosome was divided into 1 Mbp sliding windows (778 windows on the entire genome) with single-nucleotide overlap using BEDtools [version 2.31.1 9, (Quinlan & Hall, 2010 )] based on the most recent reference genome of tomato [SL4.00, (Hosmani et al., 2019 )]. Then, the number of short genetic variants and SVs in each window was calculated using BEDtools. Density histograms for individual SV types (DEL, INS, BNDa, BNDb, and DUP), overall SVs, and short genetic variants were generated and displayed in circular graphs by using Circos software (Krzywinski et al., 2009 ). 2.6. Genetic distance and phylogenetic trees The VCF file containing short genetic variants, along with the output from the R4-round SV calling, was used to calculate genetic distances among the 60 tomato lines separately. Tassel v5.0 (Bradbury et al., 2007 ) was then used to calculate the genetic distance matrix. TASSEL calculates genetic distance as IBS (identity by state) similarity, where IBS is defined as two fragments of DNA that are identical for reasons other than a recent shared common ancestor. In clustering, lower IBS values indicate greater genetic similarity; a value of 0 indicates an individual's distance from itself (Bradbury et al., 2007 ). 2.7. Similarity clustering of the results The results of the genetic distance matrix were imported into the RStudio (version 1.2.5042) environment, hierarchically clustered (complete linkage method), and plotted against each other as tanglegrams using the “dendextend” package (Galili, 2015 ). For the tanglegrams, the association of the branches in the two opposing dendrograms was statistically measured by calculating Baker’s gamma index (BGI) (Baker, 1974 ). Cophenetic correlation matrices of these dendrograms were calculated and visualized by the package “corrplot” [version 0.95, (Wei & Simko, 2024 )] 3. Results 3.1. SNP and short Indel Calling Over six years, a whole-genome sequencing database of diverse tomato lines was established at NCSU and used in this study. Using Illumina next-generation sequencing (NGS) technology, 150-bp paired-end reads with an average of 26.81 Gb of raw data (35× sequencing depth) per sample were generated. Raw sequencing reads were aligned to the tomato reference genome (version SL4.0, ‘Hienz1706’), yielding an average of 179 million high-quality reads per sample, with 97.96% of reads mapping primarily to the reference genome (Supplementary Table 1) . Short genetic variants were called for each sample individually and merged into a single file, generating more than 10.9 million (10,953,613) genetic variants. Of these, 10,886,815 genetic sites (10,044,343 SNPs and 842,472 short InDels) were polymorphic among the lines in the collection (Supplementary Table 4a) . In contrast to the reference genome, 66,798 monomorphic genetic variants (59,379 homozygous and 7,419 heterozygous) were present in all 57 tomato lines in the collection. Among the identified polymorphic genetic variants, 5,674,878 sites (5,200,573 SNPs and 474,305 short InDels) showed homozygosity across the lines. A high percentage of polymorphic genetic variants (47.9%) existed at the heterozygous stage in at least one and a maximum of 56 lines of the collection. In contrast, heterozygous monomeric genetic variants accounted for a small percentage of the variation (0.07%), at only 7,419 genetic sites. We aligned sequencing reads from three wild tomato accessions (WTAs) to the cultivated tomato reference genome (‘Heinz1706’) to identify genetic variants. We identified 46,414,599 genetic variants (44,110,425 SNPs and 2,304,174 short InDels). Of those, 40,317,231 genetic sites showed polymorphism among the three WTAs, with 22,065,066 (20,477,542 SNPs and 1,587,524 short InDels) sites being homozygous and 18,252,165 (17,595,414 SNPs and 656,751 InDels) sites being heterozygous. Approximately 45.3% of the polymorphic genetic variants were heterozygous in at least one wild tomato accession. There were 6,097,368 monomorphic genetic variants among the three accessions compared to the reference genome. Monomorphic variants at 5,876,446 sites (5,818,629 SNPs and 57,817 short InDels) showed homozygosity in all three wild tomato accessions. Meanwhile, the heterozygous monomorphic variants were observed at 220,922 genetic sites (218,840 SNPs and 2,082 short InDels) across the three accessions (Supplementary Table 4b). 3.2. Structural variant calling The Manta structural variant caller was used to call structural variants (SVs) from the same sequencing reads, aligned to the reference genome described in the previous section. We identified five major structural variants: deletions (DEL), insertions (INS), tandem duplications (DUP), and two Inversions reported as break ends (BNDa and BNDb). Four approaches (R1, R2, R3, and R4) were implemented, resulting in increased SV calls in most samples ( Fig. 1 A, 1 B, and 1 C). A total of 78,427 SVs were called when individual samples were subjected to independent variant calling by Manta (R1). Of those, 56,150 sites were flagged with the ‘PASS’ in the filter column, signifying that the variants have met Manta's thresholds for quality and reliability. There were 80,633, 87,328, and 89,459 SVs called in the next three rounds, R2, R3, and R4, respectively. Of which, 62,779, 68,768, and 71,262 SVs were designated high quality with ‘PASS’ flag in the filter column, respectively ( Fig. 1 A ). The SVs called from the fourth round (R4) with the ‘PASS’ flag (71,262 SVs) were used in the downstream data analysis and visualization. Further data analysis revealed that 25,325 and 45,937 SVs were present in the genomes of the cultivated and three wild tomato accessions, respectively. Of those, 2,494 SVs (2088 homozygous and 406 heterozygous) were common across all lines in the collection (wild tomato accessions and cultivated tomatoes), leaving 22,832 SVs present only in the cultivated tomatoes but not in any of the three WTAs in this study (Supplementary Table 4c and 4d) . Among SVs identified in the cultivated tomato collection, 25,301 SVs (16,785 homozygous and 8516 heterozygous) were polymorphic across 57 lines. There were only 24 monomeric SVs among the lines (Supplementary Table 4c) . Among the three wild tomato accessions, 41,405 polymorphic SVs were identified, of which 31,372 and 10,033 SVs were homogeneous and heterogeneous, respectively (Supplementary Table 4d) . The investigation of allele frequencies revealed that a significant proportion of the SVs were present only at the heterogeneous stage, accounting for 33% (8,522 SVs) in cultivated tomatoes, compared with 10,247 SVs (22%) in wild tomato accessions. 3.3. Statistics of SV data In the cultivated tomato collection, a total of 25,325 SVs with ‘PASS’ flag were identified as present in at least one line. Of those, 3.018 sites (12%) were private doubleton SVs uniquely present in a single line. The remaining SVs (22,307 SVs) were present in two or more lines in the collection, with some SVs being present in all 57 lines (Supplementary Table 2) . The number of SVs identified in the grape tomato breeding lines developed at the NC tomato breeding program (NC 1–6 grape lines) was significantly high. Among them, ‘NC 4grape’ with 5,280 homozygous SVs harbors the highest number of SVs, and ‘NC 5grape’ carries the lowest number of SVs in grape lines (4,462 SVs). Nevertheless, SVs in the ‘NC 5grape’ lines were still significantly high compared to the rest of the cultivated tomatoes in the collection ( Fig. 1 B ) . Old open-pollinated cultivars and landrace tomatoes, such as ‘Thessaloniki’ and ‘Pearl Harbor’, had the fewest SVs in their genomes, 522 and 326, respectively ( Fig. 1 B ) . The distribution of the SVs on the tomato chromosomes differed from a uniform distribution of an average of 2,111 SVs per chromosome. The number of SVs on chromosome 9 was almost double the average (4,050 SVs), and chromosome 6 harbored 3,278 SVs, 55% higher than the average. The lowest number of SVs, 794, was observed on chromosome 8 ( Fig. 3 A ) . Additionally, each chromosome harbored significant numbers of private doubleton SVs, each present in a single line; the number per chromosome ranged from 119 to 624. Chromosome 6 contains the most private doubleton SVs (624), accounting for 19% of the SVs on this chromosome. Chromosomes 11 and 8, with 119 and 120 counts, showed the fewest private doubleton appearances among all 12 chromosomes ( Fig. 3 A ) . Furthermore, the number of private doubleton SVs varied significantly among the cultivated tomato lines. A modern, large-fruited NC breeding line, ‘NC 233253’, possessed 317, followed by four lines with more than 200 doubletons in their genome (NC 6 Grape, Mogeor, NC 3 Grape, and LA4442). About half of the lines contained four or fewer private doubleton SVs, with six lines exhibiting none ( Fig. 2 A ). The investigation of the SV type revealed that 12,647 sites harbored insertion-type SVs in cultivated tomato genomes, accounting for 50% of the total SVs when aligned to the ‘Heinz1706’ genome; meanwhile, deletions accounted for 41% of the SVs (DEL = 10,321 SVs). We also identified 2,004 and 353 genetic sites with BND (8%) and DUP (1) type SVs in the collection, respectively ( Fig. 3 B and 3 C ). In cultivated tomatoes, 682 BNDa were identified; the remaining BNDs (BNDb = 1,322 SVs) were relocated onto nonhomologous chromosomes (Supplementary Table 4c and Fig. 4 A ) . In the three wild tomato accessions, 45,937 SVs were called with the ‘PASS’ flag using a single Manta workflow from the short read (150 bp) aligned to the tomato reference genome, of which 23,242 SVs (51%) were private doubletons, present only in one wild tomato accession. The remaining SVs (22,695) were common across two or all three wild accessions in the study (Supplementary Table 3) . The highest number of private doubleton SVs, 16,386, was observed in LA2582 (4n); meanwhile, LA1938 (2n) had the lowest number of private doubleton SVs among the three lines (2,944). Similarly, the LA1930 (2n) wild tomato accession shows 3,912 private doubleton SVs when aligned to the reference genome ( Fig. 2 B ) . The distribution of the total SVs on the chromosomes of the three wild tomato accessions was more uniform, with the following outliers. The number of SVs on chromosome 1 was 30% higher than average per chromosome (4,962 SVs). Similarly, chromosomes 3 and 4, with 4,614 and 4,143 SVs, had approximately 21% and 8% higher numbers than the average, respectively. In contrast, chromosomes 5, 11, and 12 had 17%, 18%, and 11% fewer SVs than the average, with 3,170, 3,125, and 3,390 SVs, respectively. The number of SVs on the remaining six chromosomes was close to the average number of 3,828 SVs ( Fig. 3 D ) . Each chromosome in the three wild tomato accessions displayed a substantial density of private doubleton structural variants, ranging from 1,253 on chromosome 5 to pronounced peaks on chromosomes 3 and 4 with 2,375 and 2,342 doubleton SVs, respectively ( Fig. 3 D ) . The deletion-type SVs were dominant in the genomes of the three wild tomato accessions, accounting for 25,626 sites and 56% of the genome architecture. Insertion-type SVs were present at 16,766 genetic sites and were the second dominant type (37%) of SV in the genomes of wild tomato accessions. The other two SV types, BND and DUP, were observed at 2,951 and 594 genetic sites, accounting for 6% and 1% of the total SVs, respectively ( Fig. 3 E and 3 F ) . Further classification of BND-type SVs revealed 1,104 BNDa and 1,847 BNDb-type SVs across the three wild tomato accessions (Supplementary Table 4d and Fig. 4 B ) . 3.4. Distribution of SV on Chromosomes The physical locations of the identified short genetic variants, total SVs, and different SV types were mapped on the tomato genome, revealing distinct patterns of distribution and density across different chromosomes. On average, there were 37 SVs per 1 Mbp span of the cultivated tomato genome; however, the distribution of SVs across chromosomes is highly uneven and inconsistent, with multiple hot spots exhibiting significantly higher numbers of SVs ( Fig. 4 A ) . In general, high densities of SVs were observed in the distal and subtelomeric regions of the chromosomes. Nevertheless, low-density SV coverage was present on the proximal and centromeric regions of all the chromosomes. On chromosomes 4, 9, 11, and 12, SV density increases significantly when moving towards both subtelomeric regions; however, the rest of the chromosomes show a different distribution pattern with SVs clustered towards one of the distal regions. The distal part of the long arms of Chromosomes 1, 2, 3, 5, and 7 contains many SVs, and, in contrast, the distal part of the short arms of Chromosomes 6, 8, and 10 shows highly dense SV presence ( Fig. 4 A ) . Three regions, two on chromosome 6 and one on chromosome 7, exhibited the highest density of SVs in the genome of cultivated tomatoes. Nevertheless, there is a high density of SVs on the entire span of chromosome 6 (3278 SVs total), with 1128 SVs packed together in the very first 7 Mbp of the chromosome and 839 SVs clustered 20 Mbp away on another 7 Mbp span located at 20-27Mbp of the physical map. Chromosome 7, in contrast, harbors very low SV density compared to other chromosomes, except for 1,045 SVs clustered at the end of the chromosome at 57–64 Mbp (7 MbP). Other than the high peaks evident on the SV density histogram of the chromosomes, the density of the SV coverage on the entire span of chromosomes 6 and 9 was significantly higher than that of any other chromosome on the cultivated tomato genome. In contrast, chromosomes 8 and 10 exhibit low SV distribution across their entire length ( Fig. 4 A ). Except for chromosome 2, the distribution of SVs called from the alignment of sequencing reads from 3 wild tomato accessions onto the cultivated tomato genome showed uniform patterns across all 12 chromosomes. A high distribution of SVs across the entire chromosome span was flanked by high-density peaks at the two ends of the chromosomes, creating a pronounced U-shape on 11 chromosomes. Chromosome 2 shows a different distribution pattern, with the major SV distribution confined to the end of the long arm ( Fig. 4 B ). In both cultivated and three wild tomato accessions, there is a strong correlation between the four SV types present in the chromosomes. The number of all four SV types increases proportionally to their abundance as the overall SV distribution increases ( Figs. 4 A and 4 B ). The distribution pattern of the SVs aligns with that of the short genetic variants in cultivated tomato in this study. In the three wild tomato accessions, however, short genetic variants are uniformly distributed across all 12 chromosomes, which is inconsistent with the U-shaped distribution of SVs across chromosomes ( Fig. 4 C ) . 3.5. Cophylogenetic Trees Two separate genetic distances and, consequently, two phylogenetic trees were generated based on short genetic variants and SV variation across all 60 lines. A tanglogram was generated to visualize and quantify similarities between two phylogenetic trees ( Fig. 5 ) . There was a robust correlation between the two phylogenetic trees, with Cophenetic and Baker correlation coefficients of 0.7472186 and 0.9511537, respectively. The tomato lines in this study clustered into seven distinct clades in both phylogenetic trees; however, a single line, ‘NC 233253’, formed a separate, eighth cluster (SV-C8) in the SV phylogenetic tree. One of the wild tomato accessions, LA2582, grouped with NC grape tomatoes in the short genetic variant phylogenetic tree, but clustered with two other wild tomato accessions in the SV tree (SNV-C6 and SV-C7). The tomato lines in the remaining clades show minimal shuffling within each clade in two comparative dendrograms ( Fig. 5 ) . 4. Discussion 4.1. A scalable path to SV-detection in tomato Consistent with previous studies on structural variants in the genome od tomatoes and other crops, this study demonstrates that the Manta SV caller can detect and map a substantial number of SVs using short-read sequencing (150 bp), with sufficient fidelity for downstream breeding decisions, significantly lowering barriers relative to long-read sequencing approaches in cost, turnaround time, and analytical complexity. It has been shown that by extracting SV information from existing whole-genome datasets, programs can convert sunk sequencing costs into new genetic insights that immediately inform trait-mapping, selection, and product-development pipelines (Zhou et al., 2022 ). The result is a practical, budget-conscious approach to integrating SVs alongside SNPs/InDels into routine marker-assisted workflows without new sequencing campaigns. In this study, we successfully called SVs in the genomes of a collection of tomato lines using short sequencing reads (150 bp). We also investigated their prevalence by SV type in different tomato lines, their distribution across chromosomes, and their importance in the tomato genome. The identification and study of short genetic variants have been the central focus of whole-genome sequencing with short-read data. The investigation of these short genetic variants was not the focus of this study; however, we explored the correlation between short genetic variants and SVs to estimate and validate the contribution of the SV to the genetic architecture and, consequently, to the genetic diversity of the tomato lines in this study. This finding will also encourage researchers to revisit and delve deeper into previously generated whole-genome sequencing data, potentially identifying the leading genetic cause of the phenotypes of interest in gene discovery studies and improving our understanding of functional genome-phenome interactions, especially in plant breeding and gene discovery. 4.2. Workflow effect on SV calling To use Manta for SV calling, the initial alignment of raw sequencing reads by Bowtie2 requires a critical adjustment. Bowtie2 operates end-to-end read alignment by default, but the BAM file generated by this approach did not produce any SV calls when subjected to the Manta workflow. Realignment of the raw reads using the –local option in Bowtie2 was crucial for SV calling with Manta in the next step. According to the user manual, Manta can be used for batch SV calling of individual and small sets of samples; however, the number of samples is limited. In this study, 15 samples were processed in a single Manta workflow without error. To investigate the efficacy of SV calling with Manta, four approaches (R1, R2, R3, and R4) were implemented, resulting in increased SV calls across most samples. In general, significantly more SVs were called when multiple samples were batched and subjected to a single Manta workflow. It is not clear to us why Manta calls more SVs when multiple samples are included in a single workflow; however, it has been shown previously that an ensemble approach to call SVs, using multiple SV callers, can also improve the chance of SV identification and precision without substantially decreasing sensitivity (Zarate et al., 2018). 4.3. Prevalence of SV in the genome of tomato A significant and satisfactory number of SVs were detected in the genomes of 60 tomato lines in the collection using short-read sequencing data. Compared with cultivated tomatoes, significantly more SVs were identified in wild tomato accessions; furthermore, LA2582, a tetraploid wild tomato accession, has significantly more SVs than the other two wild accessions. The total number of SVs identified in this study was consistent with those reported in other studies (Alonge et al., 2020 ; Li et al., 2023 ; Liu et al., 2022 ). Alonge et al. ( 2020 ) identified 238,490 structural variants across 100 tomato genomes; however, they used long reads and included 100 tomato lines, with a greater proportion of wild tomato accessions. Deletion and insertion SV types were slightly higher in wild tomato accessions and cultivated tomatoes, respectively, compared to the reference genome. This is evidence, but not proof, that the genome size of the fresh-market tomatoes in the collection is larger than that of the reference genome. Meanwhile, the reference genome is slightly larger than those of the wild tomato accessions used in this study. Given that most of the cultivated breeding lines in this study are modern tomato lines, the presence of a significant number of SVs unique to and homozygous in each line (private doubletons) is noteworthy and warrants further exploration in integrated genome-wide studies for gene discovery and quantitative genetic studies. The presence of large numbers of private doubletons was independent of fruit size and shape, as NC 233253 and NC 9BS produce the largest fruit among lines in the collection, yet one has the highest number of doubletons and the other contains none. Additionally, grape tomato lines and lines with nugget-sized fruits (Mogeor and LA4442) contain almost the same number of doubletons as large-fruited lines. Different distribution patterns emerged when mapping the physical location of the SVs onto tomato chromosomes. In the cultivated tomato collection, the high density of SVs on chromosomes is strongly associated with short genetic-variant distributions, which, in turn, are associated with known disease-resistance loci and with genetic introgression from wild tomatoes that carry them. For instance, the concentration of SVs on chromosome 9 aligns well with multiple disease resistance genes on the chromosome, including the Frl gene (Fusarium crown and root rot), the Tm-2a gene (tomato and tobacco mosaic virus), the Ph-3 gene (late blight caused by Phytophthora infestans ), and the Sw-5b gene (tomato spotted wilt virus), with the Tm-2a gene being associated with the largest introgression in the lines with resistance allele. Meanwhile, an extensive introgression associated with the Mi gene on chromosome 6 is a source of high concentration of SVs on the short arm of the chromosome. The high peak of SVs on the distal part of chromosome 7 is linked to the I3 gene (Fusarium wilt, race 3) (Du et al., 2025 ; Foolad, 2007 ; Simko et al., 2021 ). Beyond variations in fruit shape, size, and other horticultural traits, the cultivated tomatoes examined in this study also possess the disease-resistance genes referenced above. The association of genomic regions with low SV density on chromosomes 8 and 10, and of large segments on other chromosomes, with fruit and horticultural characteristics of cultivated tomatoes (fresh market and processing) is evident in this study but requires further investigation. Random and scattered structural variants in the tomato genome led to variation in plant growth habits, fruit size, flavor, disease resistance, and pathogen recognition (Gao et al., 2019 ; Jobson & Roberts, 2022 ; M. B. Lee et al., 2022 ; Wang et al., 2020 ). However, introgression of disease resistance from wild relatives of cultivated tomato appears to be the primary source of high-density SV clusters in the tomato genome. 4.4. Cophylogenetic relationship A cophylogenetic study using tanglegrams provided a visual and statistical phylogenomic comparison of two main genotyping approaches based on genetic variants. Baker correlation coefficients of 0.95 indicate statistically similar clustering between phylogenetic trees generated from short genetic variants and from major SVs. However, meaningful differences were observed among the clusters generated by the two different genetic variants. Interestingly, the phylogenetic tree generated from SV polymorphism among the 60 lines was more accurate given the known pedigree information for the tomato lines (NC State University, 2024 ). For instance, one of the wild tomato accessions is clustered with grape tomato lines based on short genetic variants, but realigned with other wild tomato accessions when SVs were used to generate the tree. Additionally, a new breeding line from the NCSU tomato breeding program, ‘NC 233253’, with minimal NCSU genetics in its pedigree, clustered separately from all other lines in the SV-based phylogenetic tree. This finding implies that integrating traditional short variant polymorphisms with SV data can significantly improve genotyping resolution and, consequently, increase the accuracy of phylogenetic and marker-trait association studies. Given that both genotypic data can be generated from the same short sequencing reads. Integrating SV-derived distances with short variant genotypes tightens phylogenetic clustering and improves pedigree-consistent grouping, increasing confidence when making selection and crossing decisions. The tanglegram analysis, supported by high Baker’s gamma and strong cophenetic correlation, demonstrates that SV data enhances discriminatory power, uncovering genomic distinctions often missed by SNPs—such as correctly grouping a wild accession with other wild types and isolating a breeding line outside the program pedigree. For commercial breeding teams, this improved resolution enables more accurate segment tracking, reduces misclassification, and ensures greater predictability in advanced lines. Evidently, the frequency and physical map of SVs in tomatoes are highly associated with and clustered around the distribution of short genetic variants in the genome and can be used in combination with short variants to improve the mapping resolution of genomic studies. Additionally, structural variants are excellent targets for developing DNA markers using high-resolution melting (HRM) curve-based methods. Genotyping by HRM curve provides the most affordable genotyping platform for small breeding programs; however, DNA marker development based on SNVs is challenging because most cases show insufficient shift in melting temperature to discriminate heterozygosity (Lee et al., 2018 ). 4.5. Limitations While short-read SV calling (Manta) performed well across workflows, it may under-detect complex rearrangements or transgene insertion architectures, as evidenced by missed events in transformed lines with complicated breakpoints. In four transgenic tomato lines with a “Moneymaker” background and two related backcrossed lines, we confirmed seven unique SVs (private doubletons) flagged by Manta as potential transgene insertion sites. Yet, Manta missed additional insertion sites, likely due to the complexity of these regions in the host genome. We employed an alternative method to identify six additional insertion sites (unpublished data), illustrating both the strengths and limitations of the Manta-based approach. Additionally, batching improves discovery but does not guarantee capture of all SV classes; breakpoint-only detection can leave multi-break rearrangements partially unresolved. To mitigate risk in product development, tiered validation (visual inspection, orthogonal assays, targeted long-read confirmation for critical loci) should be embedded before large-scale deployment. It has been pointed out that Manta and other SV callers only detect breakpoints, whereas multiple breakpoints are characteristic of complex genomic rearrangements in an organism (Shale et al., 2020 ). This study leveraged short-read sequencing to identify widespread structural variants (SVs) and map their positions on the tomato genome, highlighting the need for deeper investigations into complex genomic rearrangements and the genetic architecture of cultivated and wild tomatoes in genome–phenome interaction research. Conclusion Short-read SV detection offers a scalable, cost-effective strategy for integrating structural variation into tomato genomics and breeding pipelines, unlocking hidden genetic diversity and improving the accuracy of phylogenetic and marker-trait association studies. Declarations Acknowledgements We thank Dr. Dorith Rotenberg and Anna Whitefield (Professors at the Entomology and Plant Pathology Department, NCSU) for supplying the transgenic lines and Dr. Randy Gardner (NC State University, Emeritus Professor) for providing seeds from NC GEM tomato breeding lines. Funding This research was supported by multiple sources. An internal startup grant from the Department of Horticultural Science at North Carolina State University provided initial funding for project development. Additional support for next-generation sequencing (NGS) data generation was provided through federal and state programs, including: USDA National Institute of Food and Agriculture (USDA-NIFA), Grant No. 2022-67013-37076 US Department of Agriculture Agricultural Marketing Service (USDA/AMS) and North Carolina Department of Agriculture and Consumer Services (NCDACS), Grant Nos. 23SCBPNC1202-00 and AM200100XXXXG061 California Tomato Research Institute (CTRI), Grant No. 2021 - 294 Conflicts of Interest The author declares no conflicts of interest. Data Availability Statement Supporting data are provided within the article and supplementary files or available from the corresponding author upon reasonable request. Whole-genome sequencing datasets remain restricted due to ongoing research and ethical constraints. References Alonge, M., Wang, X., Benoit, M., Soyk, S., Pereira, L., Zhang, L., Suresh, H., Ramakrishnan, S., Maumus, F., Ciren, D., Levy, Y., Harel, T. H., Shalev-Schlosser, G., Amsellem, Z., Razifard, H., Caicedo, A. L., Tieman, D. M., Klee, H., Kirsche, M., … Lippman, Z. B. (2020). Major Impacts of Widespread Structural Variation on Gene Expression and Crop Improvement in Tomato. Cell , 182 (1). https://doi.org/10.1016/j.cell.2020.05.021 Baker, F. B. (1974). Stability of Two Hierarchical Grouping Techniques Case 1: Sensitivity to Data Errors. In Source: Journal of the American Statistical Association (Vol. 69, Issue 346). Bradbury, P. J., Zhang, Z., Kroon, D. E., Casstevens, T. M., Ramdoss, Y., & Buckler, E. S. (2007). TASSEL: Software for association mapping of complex traits in diverse samples. Bioinformatics , 23 (19), 2633–2635. https://doi.org/10.1093/bioinformatics/btm308 Chen, X., Schulz-Trieglaff, O., Shaw, R., Barnes, B., Schlesinger, F., Källberg, M., Cox, A. J., Kruglyak, S., & Saunders, C. T. (2016). Manta: Rapid detection of structural variants and indels for germline and cancer sequencing applications. Bioinformatics , 32 (8), 1220–1222. https://doi.org/10.1093/bioinformatics/btv710 Danecek, P., Auton, A., Abecasis, G., Albers, C. A., Banks, E., DePristo, M. A., Handsaker, R. E., Lunter, G., Marth, G. T., Sherry, S. T., McVean, G., & Durbin, R. (2011). The variant call format and VCFtools. Bioinformatics , 27 (15), 2156–2158. https://doi.org/10.1093/bioinformatics/btr330 Danecek, P., Bonfield, J. K., Liddle, J., Marshall, J., Ohan, V., Pollard, M. O., Whitwham, A., Keane, T., McCarthy, S. A., & Davies, R. M. (2021). Twelve years of SAMtools and BCFtools. GigaScience , 10 (2). https://doi.org/10.1093/gigascience/giab008 Du, M., Sun, C., Deng, L., Zhou, M., Li, J., Du, Y., Ye, Z., Huang, S., Li, T., Yu, J., Li, C. B., & Li, C. (2025). Molecular breeding of tomato: Advances and challenges. In Journal of Integrative Plant Biology (Vol. 67, Issue 3, pp. 669–721). John Wiley and Sons Inc. https://doi.org/10.1111/jipb.13879 Foolad, M. R. (2007). Genome mapping and molecular breeding of tomato. International Journal of Plant Genomics , 2007 . https://doi.org/10.1155/2007/64358 Fuentes, R. R., Chebotarov, D., Duitama, J., Smith, S., Hoz, J. F. D. La, Mohiyuddin, M., Wing, R. A., Mcnally, K. L., Tatarinova, T., Grigoriev, A., Mauleon, R., & Alexandrov, N. (2019). Structural variants in 3000 rice genomes. In Genome Research (Vol. 29, Issue 5, p. 870). https://doi.org/10.1101/gr.241240.118 Galili, T. (2015). dendextend: An R package for visualizing, adjusting and comparing trees of hierarchical clustering. Bioinformatics , 31 (22), 3718–3720. https://doi.org/10.1093/bioinformatics/btv428 Gao, L., Gonda, I., Sun, H., Ma, Q., Bao, K., Tieman, D. M., Burzynski-Chang, E. A., Fish, T. L., Stromberg, K. A., Sacks, G. L., Thannhauser, T. W., Foolad, M. R., Diez, M. J., Blanca, J., Canizares, J., Xu, Y., Knaap, E. Van Der, Huang, S., Klee, H. J., … Fei, Z. (2019). The tomato pan-genome uncovers new genes and a rare allele regulating fruit flavor. In Nature Genetics (Vol. 51, Issue 6, p. 1044). https://doi.org/10.1038/s41588-019-0410-2 Gardner, R. G. (1982). NC50-7 Breeding Line, “Cherokee”, and “Mountain Pride” Tomato . Gardner, R. G., & Panthee, D. R. (2012). Tomato Spotted Wilt Virus-resistant Fresh-market Tomato Breeding Lines: NC 58S, NC 123S, NC 127S, and NC 132S. In HORTSCIENCE (Vol. 47, Issue 4). http://www.ces.ncsu.edu/ Hosmani, P. S., Flores-Gonzalez, M., van de Geest, H., Maumus, F., Bakker, L. V., Schijlen, E., van Haarst, J., Cordewener, J., Sanchez-Perez, G., Peters, S., Fei, Z., Giovannoni, J. J., Mueller, L. A., & Saha, S. (2019). An improved de novo assembly and annotation of the tomato reference genome using single-molecule sequencing, Hi-C proximity ligation and optical maps . https://doi.org/10.1101/767764 Jobson, E., & Roberts, R. (2022). Genomic structural variation in tomato and its role in plant immunity. In Molecular Horticulture (Vol. 2, Issue 1). https://doi.org/10.1186/s43897-022-00029-w Krzywinski, M., Schein, J., Birol, I., Connors, J., Gascoyne, R., Horsman, D., Jones, S. J., & Marra, M. A. (2009). Circos: An information aesthetic for comparative genomics. Genome Research , 19 (9), 1639–1645. https://doi.org/10.1101/gr.092759.109 Langmead, B., & Salzberg, S. L. (2012). Fast gapped-read alignment with Bowtie 2. Nature Methods , 9 (4), 357–359. https://doi.org/10.1038/nmeth.1923 Lee, M. B., Shekasteband, R., Hutton, S. F., & Lee, T. G. (2022). A mutant allele of the flowering promoting factor 1 gene at the tomato BRACHYTIC locus reduces plant height with high quality fruit. Plant Direct , 6 (8). https://doi.org/10.1002/pld3.422 Lee, T. G., Shekasteband, R., Menda, N., Mueller, L. A., & Hutton, S. F. (2018). Molecular markers to select for the j-2–mediated jointless pedicel in tomato. HortScience , 53 (2). https://doi.org/10.21273/HORTSCI12628-17 Li, H., Handsaker, B., Wysoker, A., Fennell, T., Ruan, J., Homer, N., Marth, G., Abecasis, G., & Durbin, R. (2009). The Sequence Alignment/Map format and SAMtools. Bioinformatics , 25 (16), 2078–2079. https://doi.org/10.1093/bioinformatics/btp352 Li, N., He, Q., Wang, J., Wang, B., Zhao, J., Huang, S., Yang, T., Tang, Y., Yang, S., Aisimutuola, P., Xu, R., Hu, J., Jia, C., Ma, K., Li, Z., Jiang, F., Gao, J., Lan, H., Zhou, Y., … Yu, Q. (2023). Super-pangenome analyses highlight genomic diversity and structural variation across wild and cultivated tomato species. Nature Genetics , 55 (5). https://doi.org/10.1038/s41588-023-01340-y Liu, L., Zhang, K., Bai, J., Lu, J., Lu, X., Hu, J., Pan, C., He, S., Yuan, J., Zhang, Y., Zhang, M., Guo, Y., Wang, X., Huang, Z., Du, Y., Cheng, F., & Li, J. (2022). All-flesh fruit in tomato is controlled by reduced expression dosage of AFF through a structural variant mutation in the promoter. Journal of Experimental Botany , 73 (1). https://doi.org/10.1093/jxb/erab401 NC State University. (2024, April). https://mountainhort.ces.ncsu.edu/fresh-market-tomato-breeding/ . Pascual, L., Albert, E., Sauvage, C., Duangjit, J., Bouchet, J. P., Bitton, F., Desplat, N., Brunel, D., Le Paslier, M. C., Ranc, N., Bruguier, L., Chauchard, B., Verschave, P., & Causse, M. (2016). Dissecting quantitative trait variation in the resequencing era: Complementarity of bi-parental, multi-parental and association panels. Plant Science , 242 , 120–130. https://doi.org/10.1016/j.plantsci.2015.06.017 Prasanna, H., Rai, N., Hussain, Z., Yerasu, S. R., & Tiwari, J. K. (2023). Tomato: Breeding and Genomics. Vegetable Science , 50 (Special). https://doi.org/10.61180/vegsci.2023.v50.spl.02 Quinlan, A. R., & Hall, I. M. (2010). BEDTools: A flexible suite of utilities for comparing genomic features. Bioinformatics , 26 (6), 841–842. https://doi.org/10.1093/bioinformatics/btq033 Sato, S., Tabata, S., Hirakawa, H., Asamizu, E., Shirasawa, K., Isobe, S., Kaneko, T., Nakamura, Y., Shibata, D., Aoki, K., Egholm, M., Knight, J., Bogden, R., Li, C., Shuang, Y., Xu, X., Pan, S., Cheng, S., Liu, X., … Gianese, G. (2012). The tomato genome sequence provides insights into fleshy fruit evolution. Nature , 485 (7400), 635–641. https://doi.org/10.1038/nature11119 Shale, C., Baber, J., Cameron, D. L., Wong, M., Cowley, M. J., Papenfuss, A. T., Cuppen, E., & Priestley, P. (2020). Unscrambling cancer genomes via integrated analysis of structural variation and copy number . https://doi.org/10.1101/2020.12.03.410860 Simko, I., Jia, M., Venkatesh, J., Kang, B. C., Weng, Y., Barcaccia, G., Lanteri, S., Bhattarai, G., & Foolad, M. R. (2021). Genomics and Marker-Assisted Improvement of Vegetable Crops. Critical Reviews in Plant Sciences , 40 (4), 303–365. https://doi.org/10.1080/07352689.2021.1941605 Turner, D. J., Miretti, M., Rajan, D., Fiegler, H., Carter, N. P., Blayney, M. L., Beck, S., & Hurles, M. E. (2007). Germline rates of de novo meiotic deletions and duplications causing several genomic disorders. In Nature Genetics (Vol. 40, Issue 1, p. 90). https://doi.org/10.1038/ng.2007.40 Wang, X., Gao, L., Jiao, C., Stravoravdis, S., Hosmani, P. S., Saha, S., Zhang, J., Mainiero, S., Strickler, S. R., Catala, C., Martin, G. B., Mueller, L. A., Vrebalov, J., Giovannoni, J. J., Wu, S., & Fei, Z. (2020). Genome of Solanum pimpinellifolium provides insights into structural variants during tomato breeding. In Nature Communications (Vol. 11, Issue 1). https://doi.org/10.1038/s41467-020-19682-0 Wei, T., & Simko, V. (2024). R package “corrplot”: Visualization of a Correlation Matrix (Version 0.95). Available from https://github.com/taiyun/corrplot . Zhou, Y., Zhang, Z., Bao, Z., Li, H., Lyu, Y., Zan, Y., Wu, Y., Cheng, L., Fang, Y., Wu, K., Zhang, J., Lyu, H., Lin, T., Gao, Q., Saha, S., Mueller, L., Fei, Z., Städler, T., Xu, S., … Huang, S. (2022). Graph pangenome captures missing heritability and empowers tomato breeding. Nature , 606 (7914). https://doi.org/10.1038/s41586-022-04808-9 Tables Table 1 . The number of monomorphic and polymorphic structural variants present in 60 lines is in the collection. Cultivate tomatoes INS DEL BNDa BNDb DUP Total Monomophic Homozygous 6 12 0 0 0 18 Heterozygous 1 4 0 0 1 6 Polymorphic Homozygous 8,571 7,787 314 81 32 16,785 Heterozygous 4,069 2,518 368 1241 320 8,516 Total 12,647 10,321 682 1322 353 25,325 Wild tomato accessions Monomorphic Homozygous 2,663 1,392 69 164 30 4,318 Heterozygous 25 116 22 38 13 214 Polymorphic Homozygous 12,794 18,110 291 157 20 31,372 Heterozygous 1,284 6,008 722 1,488 531 10,033 Total 16,766 25,626 1,104 1,847 594 45,937 Table 2 . Number of monomorphic and polymorphic short genetic variants (SNPs and short InDels) present in 60 lines in the collection Cultivated tomatoes SNPs InDels Total Monomorphic Homozygous 48,715 10,664 59,379 Heterozygous 7,173 246 7,419 Polymorphic Homozygous 5,200,573 474,305 5,674,878 Heterozygous 4,843,770 368,167 5,211,937 Total 10,100,231 853,382 10,953,613 Wild tomato accessions Monomorphic Homozygous 5,818,629 57,817 5,876,446 Heterozygous 218,840 2,082 220,922 Polymorphic Homozygous 20,477,542 1,587,524 22,065,066 Heterozygous 17,595,414 656,751 18,252,165 Total 44,110,425 2,304,174 46,414,599 Additional Declarations No competing interests reported. Supplementary Files SupplementaryTables.xlsx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8628781","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":585429595,"identity":"363c2c24-51f2-4c92-bba8-7f813a6f2c4d","order_by":0,"name":"Reza Shekasteband","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAuklEQVRIiWNgGAWjYFACHgbGBgYGOSiPmXgtxqRrSWwgWot5/9mDD2fU1KVvZz/+TIKhwhqmFzeQuZGXbLjh2OHcnT05ZhIMZ9IJa5GQ4DGTfMB2IHfDDR42Cca2w0Ro4T9j/vPBv7p0gxvszyQY/xGjhSHHjHFjG3OCwQ0GMwnGBmK0SOQYS87sO2y44UyOsUXCsXRjYhxm+LHnW528wfHjD298qLGWJagFFSSQpnwUjIJRMApGAS4AAGLTPJ3kWjhjAAAAAElFTkSuQmCC","orcid":"","institution":"North Carolina State University","correspondingAuthor":true,"prefix":"","firstName":"Reza","middleName":"","lastName":"Shekasteband","suffix":""}],"badges":[],"createdAt":"2026-01-18 02:23:14","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8628781/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8628781/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":104336757,"identity":"170e9668-56ba-4077-97b4-dce096618486","added_by":"auto","created_at":"2026-03-10 16:07:58","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":209816,"visible":true,"origin":"","legend":"\u003cp\u003eStructural variant analysis in 60 tomato lines using four Manta workflows. (A) Total SVs (purple) vs. filtered SVs (red). (B-C) Homozygous SVs with ‘PASS’ flag in cultivated and wild accessions\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-8628781/v1/579ecbe7441258c7bef8f29e.png"},{"id":104336756,"identity":"d3779120-6db5-4948-88cc-b8ad9ca953e3","added_by":"auto","created_at":"2026-03-10 16:07:57","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":155546,"visible":true,"origin":"","legend":"\u003cp\u003eAnalysis of private doubleton structural variants (SVs) across 60 tomato lines. (A-B) Homozygous private doubleton SVs with ‘PASS’ flag in cultivated and wild accessions\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-8628781/v1/5c7a94061104160e9e8dac20.png"},{"id":104336759,"identity":"6fedbd60-d04a-4d63-8cef-45fb86b8c3ac","added_by":"auto","created_at":"2026-03-10 16:07:58","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":179974,"visible":true,"origin":"","legend":"\u003cp\u003eDistribution of structural variants (SVs) and SV types in tomato chromosomes. (A) Total SVs (red) and private doubletons (blue) in 57 cultivated tomatoes. (B-C) SV type counts and percentage distribution in cultivated tomatoes. (D-F) Same for three wild accessions\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-8628781/v1/df8dd1087b4d87f22a086c13.png"},{"id":104779867,"identity":"bd0dc19a-131f-4dea-918d-733ab052c524","added_by":"auto","created_at":"2026-03-17 07:47:07","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":669575,"visible":true,"origin":"","legend":"\u003cp\u003eCircular histogram of total structural variant counts and types. (A) SV map in cultivated tomatoes; (B) SV map in wild accessions. Rings: BNDb, BNDa (red = homozygous, blue = heterozygous), chromosomes, total BND, DUP, DEL, INS, overall SV count. (C) Comparative SV and short variant distribution in cultivated vs. wild tomatoes\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-8628781/v1/5b0bc05309aaabe8152fefe2.png"},{"id":104779630,"identity":"1fa26454-02c6-4785-91d9-67b76e4e4238","added_by":"auto","created_at":"2026-03-17 07:43:39","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":201353,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eCophylogenetic trees based on SVs and short genetic variants. \u003c/strong\u003ePhylogenetic trees generated using structural variants compared to short genetic variants (SNPs and short InDels)\u003c/p\u003e","description":"","filename":"5.png","url":"https://assets-eu.researchsquare.com/files/rs-8628781/v1/e667b52d458a6c193fdc23d0.png"},{"id":109315348,"identity":"5ec55160-9a5b-4879-afb2-80d6198e12aa","added_by":"auto","created_at":"2026-05-15 12:10:46","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1507871,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8628781/v1/d1d1253b-4f36-49b8-a804-b9e5f00110c2.pdf"},{"id":104336761,"identity":"b956d2b3-945f-427e-a8b7-3fec5f1cf090","added_by":"auto","created_at":"2026-03-10 16:07:58","extension":"xlsx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":7429125,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryTables.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-8628781/v1/193a797f2d760161e0f177b3.xlsx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Beyond SNPs: Scalable Detection of Structural Variants Unlocks Hidden Genetic Diversity in Tomato","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eTomato (\u003cem\u003eSolanum lycopersicum\u003c/em\u003e L.), one of the most important vegetable crops in the world, has been scrutinized with next-generation sequencing (NGS) technologies and other genomic tools soon after the genome of the inbred processing tomato line \u0026lsquo;Heinz1706\u0026rsquo; was sequenced (Sato et al., \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e2012\u003c/span\u003e). In recent years, hundreds of wild, landrace, heirloom, and modern tomato genomes have been re-sequenced, and millions of short genetic variants (single-nucleotide variants, SNVs, and short insertions and deletions, InDels) were identified, underlining the vast genetic variation in the tomato clade within the genus Solanum (Gao et al., \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). Ease of generating ample short genetic variation genotypic data from mapping populations and genome-wide association studies (GWAS) led to genetic mapping of numerous traits and disease resistance, QTL mapping, and consequently, functional gene discoveries in tomatoes (Pascual et al., \u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e2016\u003c/span\u003e; Prasanna et al., \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). Subsequently, breeding tomatoes for traits of interest and disease resistance has been accelerated significantly by exploiting DNA markers developed during the process (Prasanna et al., \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e2023\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eStructural variants (SVs), which occur at high frequencies, are also an important part of the genome and have gained rapid research attention in the next-generation sequencing (NGS) era (Jobson \u0026amp; Roberts, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e2022\u003c/span\u003e). Sizable variations in the physical architecture of the genome, SVs, are pervasive features that significantly impact an organism's phenotype. Any genomic variations larger than 30 base pairs fall into this category. Large deletions (DEL) and insertions (INS) are the predominant types of SVs, followed by duplication (DUP), breakends (BND), and other genome modifications (Alonge et al., \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2020\u003c/span\u003e; Fuentes et al., \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). Identification and precise physical mapping of SVs have been difficult due to their size and complex nature. Thus, modern genomic studies rely heavily on short genetic variants (SNPs and short InDels) for population genetics, marker-trait association studies, GWAS, and QTL mapping. Recent advances in whole-genome resequencing technologies capable of capturing long reads and new computational tools have enabled SV discovery, opening a new frontier in genetic studies.\u003c/p\u003e \u003cp\u003eAccording to multiple studies, SVs contribute significantly to genomic variation in humans. It has been shown that the impact of SVs on human genomic variation is 3.4 times that of SNVs (Huddleston et al., 2017). The average genomic variation between two humans due to SVs is estimated to be 15 times that due to SNVs, which are 1.5% and 0.1%, respectively (Pang et al., 2010). Sizable genomic variations associated with phenotypes are also common in plants. For example, Fuentes et al. (\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e2019\u003c/span\u003e) identified more than 63\u0026nbsp;million SVs in 3,000 rice genomes. Furthermore, this study found that long SVs are prevalent in promoter regions, whereas short SVs are prevalent in the 5\u0026prime; UTR (Fuentes et al., \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). Structural variants are present throughout the genome; however, studies show that they occur at high frequency in centromeric and subtelomeric regions, underscoring the need to map their physical locations (Turner et al., \u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e2007\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eThousands of SVs have been identified in tomatoes, emphasizing the importance of the considerable genomic variation in the genetic architecture of this vegetable crop (Alonge et al., \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2020\u003c/span\u003e; N. Li et al., \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Liu et al., \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e2022\u003c/span\u003e). Additionally, multiple structural variants associated with plant growth habits, fruit size, flavor, disease resistance, and pathogen recognition have been identified in tomatoes (Gao et al., \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e2019\u003c/span\u003e; Jobson \u0026amp; Roberts, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e2022\u003c/span\u003e; M. B. Lee et al., \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e2022\u003c/span\u003e; Wang et al., \u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). Present among modern tomato cultivars with low frequencies, SVs play a substantial role in the enormous genetic diversity present between wild and modern tomato accessions (Alonge et al., \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). Structural variants can extend our understanding of genome-to-phenome interactions, explaining the genetics behind phenotypes that are not possible to elucidate based on short genetic variants (Zhou et al., \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e2022\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eTo facilitate the deployment of widespread SVs in tomato genomics and breeding, it is crucial to adopt affordable, straightforward yet reliable approaches to identify, map, and utilize SVs in public and private plant breeding programs. This study used short sequencing reads (150 base pairs) generated through common whole-genome sequencing approaches to identify a large number of SVs, map them to the tomato genome, and explore their potential for marker-assisted tomato breeding and genomic studies. This study also demonstrates that the contribution of structural variants to genetic diversity is significant enough to serve as the sole source for phylogenetic studies, highlighting the importance of SVs in the genetic architecture of modern tomatoes.\u003c/p\u003e"},{"header":"2. Materials and Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e2.1. Plant material\u003c/h2\u003e \u003cp\u003eA core collection of 60 tomato lines, comprising wild accessions, landraces, specialty types, transgenic, and modern tomato lines, was used in this study. The wild tomato accessions and landrace tomatoes were obtained from the Tomato Genetics Resource Center (Davis, CA, USA) and the USDA's National Plant Germplasm System (NPGS). Modern and specialty-type tomato inbred lines have been developed and released by the tomato breeding program at North Carolina State University (NC State University, \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). The tomato breeding program at NCSU is located in the Mountain Horticultural Crops Research and Extension Center (MHCREC) in Mills River, NC. In the last 5 decades, more than 70 elite tomato breeding lines and 35 F\u003csub\u003e1\u003c/sub\u003e fresh-market tomato cultivars have been developed, commercialized (e.g.,Gardner, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e1982\u003c/span\u003e; Gardner \u0026amp; Panthee, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e2012\u003c/span\u003e), and transferred to public and commercial tomato research programs (NC State University, \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). The core tomato collection in this study encompasses a diverse range of horticultural characteristics \u003cb\u003e(Supplementary Table\u0026nbsp;1).\u003c/b\u003e\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e2.2. Whole genome sequencing and short read alignment\u003c/h2\u003e \u003cp\u003eLeaf samples were bulked from six seedlings at the 4-week stage for genomic DNA extraction. Genomic DNA was extracted with a DNeasy Plant Mini Kit (Qiagen, Germantown, MD). It was sent to a Novogene Co., Inc. (Beijing, China) facility in California, USA, for sequencing library preparation and whole-genome sequencing. Standard Illumina DNA libraries were prepared with an insert size of 350 bp and then sequenced on an Illumina HiSeq instrument (Illumina, San Diego, CA), producing 150-bp paired-end reads with, on average, 22 Gb of raw data per sample. Raw sequencing reads were aligned using Bowtie2 (Langmead \u0026amp; Salzberg, \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e2012\u003c/span\u003e) to the published \u0026lsquo;Heinz1706\u0026rsquo; reference genome sequence (Sato et al., \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e2012\u003c/span\u003e) and then sorted and indexed with SAMtools [version 1.9, (Li et al., \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e2009\u003c/span\u003e)]\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e2.3. Genetic variant calling\u003c/h2\u003e \u003cp\u003eSingle-nucleotide polymorphisms (SNPs) and short insertions and deletions (InDels) were called from aligned reads individually and soft-filtered using Linux system commands, BCFtools [version 1.9, (Danecek et al., \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e2021\u003c/span\u003e)]. The merge option in BCFtool was used to merge individual files into a single VCF file. Additionally, BCFtools and VCFtools [version 0.1.13, (Danecek et al., \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e2011\u003c/span\u003e)]were used to extract variant statistics, filter VCF files, and identify private doubletons.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e2.4. Structural variant calling\u003c/h2\u003e \u003cp\u003eFour independent rounds of workflow were executed to call SVs from aligned short reads (150 bp) using the Manta structural variant caller with internal default settings (Chen et al., \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e2016\u003c/span\u003e). Manta can call SV on individual or small sets of samples. In the first round (R1), SV calling was performed on individual BAM files, generating a single VCF file per sample. Small sets of samples were batched for the next three rounds (R2, R3, and R4) to run a single Manta workflow. Tomato lines with a similar pedigree, fruit shapes, and sizes were grouped to run a single round of SV calling (R2). During the third round (R3), lines were randomly assigned to a batch of 15 samples. The sample assignment for the fourth round (R4) was more complex. For this round, seven lines (round, plum, and grape tomato lines) with and without common disease resistance genes were kept fixed in each batch. Then, eight random samples were assigned to the batch to call and use SVs from. To call SVs from the three wild tomato accessions in this round, the seven fixed lines were batched only with the wild tomato accessions. Any genomic variations larger than 50 base pairs were classified as SVs and categorized into four major categories. Large deletions (DEL), insertions (INS), transversions (BND), and duplication (DUP). Chromosome translocations, BND-type SVs, were categorized into two subcategories based on the relocation of the segment within the same chromosome (intrachromosomal translocations, BNDa) or between nonhomologous chromosomes (interchromosomal translocations, BNDb). At the final step, the \u0026lsquo;merge\u0026rsquo; option in BCFtools was used to merge individual VCF files, generating a single merged VCF file per round.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003e2.5. Comparative genomics visualization\u003c/h2\u003e \u003cp\u003eTo generate the short genetic variant and SV density histograms, each tomato chromosome was divided into 1 Mbp sliding windows (778 windows on the entire genome) with single-nucleotide overlap using BEDtools [version 2.31.1 9, (Quinlan \u0026amp; Hall, \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e2010\u003c/span\u003e)] based on the most recent reference genome of tomato [SL4.00, (Hosmani et al., \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e2019\u003c/span\u003e)]. Then, the number of short genetic variants and SVs in each window was calculated using BEDtools. Density histograms for individual SV types (DEL, INS, BNDa, BNDb, and DUP), overall SVs, and short genetic variants were generated and displayed in circular graphs by using Circos software (Krzywinski et al., \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e2009\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003e2.6. Genetic distance and phylogenetic trees\u003c/h2\u003e \u003cp\u003eThe VCF file containing short genetic variants, along with the output from the R4-round SV calling, was used to calculate genetic distances among the 60 tomato lines separately. Tassel v5.0 (Bradbury et al., \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e2007\u003c/span\u003e) was then used to calculate the genetic distance matrix. TASSEL calculates genetic distance as IBS (identity by state) similarity, where IBS is defined as two fragments of DNA that are identical for reasons other than a recent shared common ancestor. In clustering, lower IBS values indicate greater genetic similarity; a value of 0 indicates an individual's distance from itself (Bradbury et al., \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e2007\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003e2.7. Similarity clustering of the results\u003c/h2\u003e \u003cp\u003eThe results of the genetic distance matrix were imported into the RStudio (version 1.2.5042) environment, hierarchically clustered (complete linkage method), and plotted against each other as tanglegrams using the \u0026ldquo;dendextend\u0026rdquo; package (Galili, \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e2015\u003c/span\u003e). For the tanglegrams, the association of the branches in the two opposing dendrograms was statistically measured by calculating Baker\u0026rsquo;s gamma index (BGI) (Baker, \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e1974\u003c/span\u003e). Cophenetic correlation matrices of these dendrograms were calculated and visualized by the package \u0026ldquo;corrplot\u0026rdquo; [version 0.95, (Wei \u0026amp; Simko, \u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e2024\u003c/span\u003e)]\u003c/p\u003e \u003c/div\u003e"},{"header":"3. Results","content":"\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003e3.1. SNP and short Indel Calling\u003c/h2\u003e \u003cp\u003eOver six years, a whole-genome sequencing database of diverse tomato lines was established at NCSU and used in this study. Using Illumina next-generation sequencing (NGS) technology, 150-bp paired-end reads with an average of 26.81 Gb of raw data (35\u0026times; sequencing depth) per sample were generated. Raw sequencing reads were aligned to the tomato reference genome (version SL4.0, \u0026lsquo;Hienz1706\u0026rsquo;), yielding an average of 179\u0026nbsp;million high-quality reads per sample, with 97.96% of reads mapping primarily to the reference genome \u003cb\u003e(Supplementary Table\u0026nbsp;1)\u003c/b\u003e. Short genetic variants were called for each sample individually and merged into a single file, generating more than 10.9\u0026nbsp;million (10,953,613) genetic variants. Of these, 10,886,815 genetic sites (10,044,343 SNPs and 842,472 short InDels) were polymorphic among the lines in the collection \u003cb\u003e(Supplementary Table\u0026nbsp;4a)\u003c/b\u003e. In contrast to the reference genome, 66,798 monomorphic genetic variants (59,379 homozygous and 7,419 heterozygous) were present in all 57 tomato lines in the collection. Among the identified polymorphic genetic variants, 5,674,878 sites (5,200,573 SNPs and 474,305 short InDels) showed homozygosity across the lines. A high percentage of polymorphic genetic variants (47.9%) existed at the heterozygous stage in at least one and a maximum of 56 lines of the collection. In contrast, heterozygous monomeric genetic variants accounted for a small percentage of the variation (0.07%), at only 7,419 genetic sites.\u003c/p\u003e \u003cp\u003eWe aligned sequencing reads from three wild tomato accessions (WTAs) to the cultivated tomato reference genome (\u0026lsquo;Heinz1706\u0026rsquo;) to identify genetic variants. We identified 46,414,599 genetic variants (44,110,425 SNPs and 2,304,174 short InDels). Of those, 40,317,231 genetic sites showed polymorphism among the three WTAs, with 22,065,066 (20,477,542 SNPs and 1,587,524 short InDels) sites being homozygous and 18,252,165 (17,595,414 SNPs and 656,751 InDels) sites being heterozygous. Approximately 45.3% of the polymorphic genetic variants were heterozygous in at least one wild tomato accession. There were 6,097,368 monomorphic genetic variants among the three accessions compared to the reference genome. Monomorphic variants at 5,876,446 sites (5,818,629 SNPs and 57,817 short InDels) showed homozygosity in all three wild tomato accessions. Meanwhile, the heterozygous monomorphic variants were observed at 220,922 genetic sites (218,840 SNPs and 2,082 short InDels) across the three accessions \u003cb\u003e(Supplementary Table\u0026nbsp;4b).\u003c/b\u003e\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003e3.2. Structural variant calling\u003c/h2\u003e \u003cp\u003eThe Manta structural variant caller was used to call structural variants (SVs) from the same sequencing reads, aligned to the reference genome described in the previous section. We identified five major structural variants: deletions (DEL), insertions (INS), tandem duplications (DUP), and two Inversions reported as break ends (BNDa and BNDb). Four approaches (R1, R2, R3, and R4) were implemented, resulting in increased SV calls in most samples \u003cb\u003e(\u003c/b\u003eFig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eA, \u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eB, and \u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eC). A total of 78,427 SVs were called when individual samples were subjected to independent variant calling by Manta (R1). Of those, 56,150 sites were flagged with the \u0026lsquo;PASS\u0026rsquo; in the filter column, signifying that the variants have met Manta's thresholds for quality and reliability. There were 80,633, 87,328, and 89,459 SVs called in the next three rounds, R2, R3, and R4, respectively. Of which, 62,779, 68,768, and 71,262 SVs were designated high quality with \u0026lsquo;PASS\u0026rsquo; flag in the filter column, respectively \u003cb\u003e(\u003c/b\u003eFig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eA\u003cb\u003e).\u003c/b\u003e The SVs called from the fourth round (R4) with the \u0026lsquo;PASS\u0026rsquo; flag (71,262 SVs) were used in the downstream data analysis and visualization.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eFurther data analysis revealed that 25,325 and 45,937 SVs were present in the genomes of the cultivated and three wild tomato accessions, respectively. Of those, 2,494 SVs (2088 homozygous and 406 heterozygous) were common across all lines in the collection (wild tomato accessions and cultivated tomatoes), leaving 22,832 SVs present only in the cultivated tomatoes but not in any of the three WTAs in this study \u003cb\u003e(Supplementary Table\u0026nbsp;4c and 4d)\u003c/b\u003e. Among SVs identified in the cultivated tomato collection, 25,301 SVs (16,785 homozygous and 8516 heterozygous) were polymorphic across 57 lines. There were only 24 monomeric SVs among the lines \u003cb\u003e(Supplementary Table\u0026nbsp;4c)\u003c/b\u003e. Among the three wild tomato accessions, 41,405 polymorphic SVs were identified, of which 31,372 and 10,033 SVs were homogeneous and heterogeneous, respectively \u003cb\u003e(Supplementary Table\u0026nbsp;4d)\u003c/b\u003e. The investigation of allele frequencies revealed that a significant proportion of the SVs were present only at the heterogeneous stage, accounting for 33% (8,522 SVs) in cultivated tomatoes, compared with 10,247 SVs (22%) in wild tomato accessions.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003e3.3. Statistics of SV data\u003c/h2\u003e \u003cp\u003eIn the cultivated tomato collection, a total of 25,325 SVs with \u0026lsquo;PASS\u0026rsquo; flag were identified as present in at least one line. Of those, 3.018 sites (12%) were private doubleton SVs uniquely present in a single line. The remaining SVs (22,307 SVs) were present in two or more lines in the collection, with some SVs being present in all 57 lines \u003cb\u003e(Supplementary Table\u0026nbsp;2)\u003c/b\u003e. The number of SVs identified in the grape tomato breeding lines developed at the NC tomato breeding program (NC 1\u0026ndash;6 grape lines) was significantly high. Among them, \u0026lsquo;NC 4grape\u0026rsquo; with 5,280 homozygous SVs harbors the highest number of SVs, and \u0026lsquo;NC 5grape\u0026rsquo; carries the lowest number of SVs in grape lines (4,462 SVs). Nevertheless, SVs in the \u0026lsquo;NC 5grape\u0026rsquo; lines were still significantly high compared to the rest of the cultivated tomatoes in the collection \u003cb\u003e(\u003c/b\u003eFig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eB\u003cb\u003e)\u003c/b\u003e. Old open-pollinated cultivars and landrace tomatoes, such as \u0026lsquo;Thessaloniki\u0026rsquo; and \u0026lsquo;Pearl Harbor\u0026rsquo;, had the fewest SVs in their genomes, 522 and 326, respectively \u003cb\u003e(\u003c/b\u003eFig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eB\u003cb\u003e)\u003c/b\u003e. The distribution of the SVs on the tomato chromosomes differed from a uniform distribution of an average of 2,111 SVs per chromosome. The number of SVs on chromosome 9 was almost double the average (4,050 SVs), and chromosome 6 harbored 3,278 SVs, 55% higher than the average. The lowest number of SVs, 794, was observed on chromosome 8 \u003cb\u003e(\u003c/b\u003eFig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eA\u003cb\u003e)\u003c/b\u003e. Additionally, each chromosome harbored significant numbers of private doubleton SVs, each present in a single line; the number per chromosome ranged from 119 to 624. Chromosome 6 contains the most private doubleton SVs (624), accounting for 19% of the SVs on this chromosome. Chromosomes 11 and 8, with 119 and 120 counts, showed the fewest private doubleton appearances among all 12 chromosomes \u003cb\u003e(\u003c/b\u003eFig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eA\u003cb\u003e)\u003c/b\u003e. Furthermore, the number of private doubleton SVs varied significantly among the cultivated tomato lines. A modern, large-fruited NC breeding line, \u0026lsquo;NC 233253\u0026rsquo;, possessed 317, followed by four lines with more than 200 doubletons in their genome (NC 6 Grape, Mogeor, NC 3 Grape, and LA4442). About half of the lines contained four or fewer private doubleton SVs, with six lines exhibiting none \u003cb\u003e(\u003c/b\u003eFig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA\u003cb\u003e).\u003c/b\u003e The investigation of the SV type revealed that 12,647 sites harbored insertion-type SVs in cultivated tomato genomes, accounting for 50% of the total SVs when aligned to the \u0026lsquo;Heinz1706\u0026rsquo; genome; meanwhile, deletions accounted for 41% of the SVs (DEL\u0026thinsp;=\u0026thinsp;10,321 SVs). We also identified 2,004 and 353 genetic sites with BND (8%) and DUP (1) type SVs in the collection, respectively \u003cb\u003e(\u003c/b\u003eFig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eB and \u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eC\u003cb\u003e).\u003c/b\u003e In cultivated tomatoes, 682 BNDa were identified; the remaining BNDs (BNDb\u0026thinsp;=\u0026thinsp;1,322 SVs) were relocated onto nonhomologous chromosomes \u003cb\u003e(Supplementary Table\u0026nbsp;4c and\u003c/b\u003e Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eA\u003cb\u003e)\u003c/b\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eIn the three wild tomato accessions, 45,937 SVs were called with the \u0026lsquo;PASS\u0026rsquo; flag using a single Manta workflow from the short read (150 bp) aligned to the tomato reference genome, of which 23,242 SVs (51%) were private doubletons, present only in one wild tomato accession. The remaining SVs (22,695) were common across two or all three wild accessions in the study \u003cb\u003e(Supplementary Table\u0026nbsp;3)\u003c/b\u003e. The highest number of private doubleton SVs, 16,386, was observed in LA2582 (4n); meanwhile, LA1938 (2n) had the lowest number of private doubleton SVs among the three lines (2,944). Similarly, the LA1930 (2n) wild tomato accession shows 3,912 private doubleton SVs when aligned to the reference genome \u003cb\u003e(\u003c/b\u003eFig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eB\u003cb\u003e)\u003c/b\u003e. The distribution of the total SVs on the chromosomes of the three wild tomato accessions was more uniform, with the following outliers. The number of SVs on chromosome 1 was 30% higher than average per chromosome (4,962 SVs). Similarly, chromosomes 3 and 4, with 4,614 and 4,143 SVs, had approximately 21% and 8% higher numbers than the average, respectively. In contrast, chromosomes 5, 11, and 12 had 17%, 18%, and 11% fewer SVs than the average, with 3,170, 3,125, and 3,390 SVs, respectively. The number of SVs on the remaining six chromosomes was close to the average number of 3,828 SVs \u003cb\u003e(\u003c/b\u003eFig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eD\u003cb\u003e)\u003c/b\u003e. Each chromosome in the three wild tomato accessions displayed a substantial density of private doubleton structural variants, ranging from 1,253 on chromosome 5 to pronounced peaks on chromosomes 3 and 4 with 2,375 and 2,342 doubleton SVs, respectively \u003cb\u003e(\u003c/b\u003eFig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eD\u003cb\u003e)\u003c/b\u003e. The deletion-type SVs were dominant in the genomes of the three wild tomato accessions, accounting for 25,626 sites and 56% of the genome architecture. Insertion-type SVs were present at 16,766 genetic sites and were the second dominant type (37%) of SV in the genomes of wild tomato accessions. The other two SV types, BND and DUP, were observed at 2,951 and 594 genetic sites, accounting for 6% and 1% of the total SVs, respectively \u003cb\u003e(\u003c/b\u003eFig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eE and \u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eF\u003cb\u003e)\u003c/b\u003e. Further classification of BND-type SVs revealed 1,104 BNDa and 1,847 BNDb-type SVs across the three wild tomato accessions \u003cb\u003e(Supplementary Table\u0026nbsp;4d and\u003c/b\u003e Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eB\u003cb\u003e)\u003c/b\u003e.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003e3.4. Distribution of SV on Chromosomes\u003c/h2\u003e \u003cp\u003eThe physical locations of the identified short genetic variants, total SVs, and different SV types were mapped on the tomato genome, revealing distinct patterns of distribution and density across different chromosomes. On average, there were 37 SVs per 1 Mbp span of the cultivated tomato genome; however, the distribution of SVs across chromosomes is highly uneven and inconsistent, with multiple hot spots exhibiting significantly higher numbers of SVs \u003cb\u003e(\u003c/b\u003eFig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eA\u003cb\u003e)\u003c/b\u003e. In general, high densities of SVs were observed in the distal and subtelomeric regions of the chromosomes. Nevertheless, low-density SV coverage was present on the proximal and centromeric regions of all the chromosomes. On chromosomes 4, 9, 11, and 12, SV density increases significantly when moving towards both subtelomeric regions; however, the rest of the chromosomes show a different distribution pattern with SVs clustered towards one of the distal regions. The distal part of the long arms of Chromosomes 1, 2, 3, 5, and 7 contains many SVs, and, in contrast, the distal part of the short arms of Chromosomes 6, 8, and 10 shows highly dense SV presence \u003cb\u003e(\u003c/b\u003eFig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eA\u003cb\u003e)\u003c/b\u003e.\u003c/p\u003e \u003cp\u003eThree regions, two on chromosome 6 and one on chromosome 7, exhibited the highest density of SVs in the genome of cultivated tomatoes. Nevertheless, there is a high density of SVs on the entire span of chromosome 6 (3278 SVs total), with 1128 SVs packed together in the very first 7 Mbp of the chromosome and 839 SVs clustered 20 Mbp away on another 7 Mbp span located at 20-27Mbp of the physical map. Chromosome 7, in contrast, harbors very low SV density compared to other chromosomes, except for 1,045 SVs clustered at the end of the chromosome at 57\u0026ndash;64 Mbp (7 MbP). Other than the high peaks evident on the SV density histogram of the chromosomes, the density of the SV coverage on the entire span of chromosomes 6 and 9 was significantly higher than that of any other chromosome on the cultivated tomato genome. In contrast, chromosomes 8 and 10 exhibit low SV distribution across their entire length \u003cb\u003e(\u003c/b\u003eFig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eA\u003cb\u003e).\u003c/b\u003e\u003c/p\u003e \u003cp\u003eExcept for chromosome 2, the distribution of SVs called from the alignment of sequencing reads from 3 wild tomato accessions onto the cultivated tomato genome showed uniform patterns across all 12 chromosomes. A high distribution of SVs across the entire chromosome span was flanked by high-density peaks at the two ends of the chromosomes, creating a pronounced U-shape on 11 chromosomes. Chromosome 2 shows a different distribution pattern, with the major SV distribution confined to the end of the long arm \u003cb\u003e(\u003c/b\u003eFig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eB\u003cb\u003e).\u003c/b\u003e\u003c/p\u003e \u003cp\u003eIn both cultivated and three wild tomato accessions, there is a strong correlation between the four SV types present in the chromosomes. The number of all four SV types increases proportionally to their abundance as the overall SV distribution increases \u003cb\u003e(\u003c/b\u003eFigs.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eA and \u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eB\u003cb\u003e).\u003c/b\u003e The distribution pattern of the SVs aligns with that of the short genetic variants in cultivated tomato in this study. In the three wild tomato accessions, however, short genetic variants are uniformly distributed across all 12 chromosomes, which is inconsistent with the U-shaped distribution of SVs across chromosomes \u003cb\u003e(\u003c/b\u003eFig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eC\u003cb\u003e)\u003c/b\u003e.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003e3.5. Cophylogenetic Trees\u003c/h2\u003e \u003cp\u003eTwo separate genetic distances and, consequently, two phylogenetic trees were generated based on short genetic variants and SV variation across all 60 lines. A tanglogram was generated to visualize and quantify similarities between two phylogenetic trees \u003cb\u003e(\u003c/b\u003eFig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e\u003cb\u003e)\u003c/b\u003e. There was a robust correlation between the two phylogenetic trees, with Cophenetic and Baker correlation coefficients of 0.7472186 and 0.9511537, respectively. The tomato lines in this study clustered into seven distinct clades in both phylogenetic trees; however, a single line, \u0026lsquo;NC 233253\u0026rsquo;, formed a separate, eighth cluster (SV-C8) in the SV phylogenetic tree. One of the wild tomato accessions, LA2582, grouped with NC grape tomatoes in the short genetic variant phylogenetic tree, but clustered with two other wild tomato accessions in the SV tree (SNV-C6 and SV-C7). The tomato lines in the remaining clades show minimal shuffling within each clade in two comparative dendrograms \u003cb\u003e(\u003c/b\u003eFig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e\u003cb\u003e)\u003c/b\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"4. Discussion","content":"\u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003e4.1. A scalable path to SV-detection in tomato\u003c/h2\u003e \u003cp\u003eConsistent with previous studies on structural variants in the genome od tomatoes and other crops, this study demonstrates that the Manta SV caller can detect and map a substantial number of SVs using short-read sequencing (150 bp), with sufficient fidelity for downstream breeding decisions, significantly lowering barriers relative to long-read sequencing approaches in cost, turnaround time, and analytical complexity. It has been shown that by extracting SV information from existing whole-genome datasets, programs can convert sunk sequencing costs into new genetic insights that immediately inform trait-mapping, selection, and product-development pipelines (Zhou et al., \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e2022\u003c/span\u003e). The result is a practical, budget-conscious approach to integrating SVs alongside SNPs/InDels into routine marker-assisted workflows without new sequencing campaigns.\u003c/p\u003e \u003cp\u003eIn this study, we successfully called SVs in the genomes of a collection of tomato lines using short sequencing reads (150 bp). We also investigated their prevalence by SV type in different tomato lines, their distribution across chromosomes, and their importance in the tomato genome. The identification and study of short genetic variants have been the central focus of whole-genome sequencing with short-read data. The investigation of these short genetic variants was not the focus of this study; however, we explored the correlation between short genetic variants and SVs to estimate and validate the contribution of the SV to the genetic architecture and, consequently, to the genetic diversity of the tomato lines in this study. This finding will also encourage researchers to revisit and delve deeper into previously generated whole-genome sequencing data, potentially identifying the leading genetic cause of the phenotypes of interest in gene discovery studies and improving our understanding of functional genome-phenome interactions, especially in plant breeding and gene discovery.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec18\" class=\"Section2\"\u003e \u003ch2\u003e4.2. Workflow effect on SV calling\u003c/h2\u003e \u003cp\u003eTo use Manta for SV calling, the initial alignment of raw sequencing reads by Bowtie2 requires a critical adjustment. Bowtie2 operates end-to-end read alignment by default, but the BAM file generated by this approach did not produce any SV calls when subjected to the Manta workflow. Realignment of the raw reads using the \u0026ndash;local option in Bowtie2 was crucial for SV calling with Manta in the next step.\u003c/p\u003e \u003cp\u003eAccording to the user manual, Manta can be used for batch SV calling of individual and small sets of samples; however, the number of samples is limited. In this study, 15 samples were processed in a single Manta workflow without error. To investigate the efficacy of SV calling with Manta, four approaches (R1, R2, R3, and R4) were implemented, resulting in increased SV calls across most samples. In general, significantly more SVs were called when multiple samples were batched and subjected to a single Manta workflow. It is not clear to us why Manta calls more SVs when multiple samples are included in a single workflow; however, it has been shown previously that an ensemble approach to call SVs, using multiple SV callers, can also improve the chance of SV identification and precision without substantially decreasing sensitivity (Zarate et al., 2018).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec19\" class=\"Section2\"\u003e \u003ch2\u003e4.3. Prevalence of SV in the genome of tomato\u003c/h2\u003e \u003cp\u003eA significant and satisfactory number of SVs were detected in the genomes of 60 tomato lines in the collection using short-read sequencing data. Compared with cultivated tomatoes, significantly more SVs were identified in wild tomato accessions; furthermore, LA2582, a tetraploid wild tomato accession, has significantly more SVs than the other two wild accessions. The total number of SVs identified in this study was consistent with those reported in other studies (Alonge et al., \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2020\u003c/span\u003e; Li et al., \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Liu et al., \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e2022\u003c/span\u003e). Alonge et al. (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2020\u003c/span\u003e) identified 238,490 structural variants across 100 tomato genomes; however, they used long reads and included 100 tomato lines, with a greater proportion of wild tomato accessions. Deletion and insertion SV types were slightly higher in wild tomato accessions and cultivated tomatoes, respectively, compared to the reference genome. This is evidence, but not proof, that the genome size of the fresh-market tomatoes in the collection is larger than that of the reference genome. Meanwhile, the reference genome is slightly larger than those of the wild tomato accessions used in this study.\u003c/p\u003e \u003cp\u003eGiven that most of the cultivated breeding lines in this study are modern tomato lines, the presence of a significant number of SVs unique to and homozygous in each line (private doubletons) is noteworthy and warrants further exploration in integrated genome-wide studies for gene discovery and quantitative genetic studies. The presence of large numbers of private doubletons was independent of fruit size and shape, as NC 233253 and NC 9BS produce the largest fruit among lines in the collection, yet one has the highest number of doubletons and the other contains none. Additionally, grape tomato lines and lines with nugget-sized fruits (Mogeor and LA4442) contain almost the same number of doubletons as large-fruited lines.\u003c/p\u003e \u003cp\u003eDifferent distribution patterns emerged when mapping the physical location of the SVs onto tomato chromosomes. In the cultivated tomato collection, the high density of SVs on chromosomes is strongly associated with short genetic-variant distributions, which, in turn, are associated with known disease-resistance loci and with genetic introgression from wild tomatoes that carry them. For instance, the concentration of SVs on chromosome 9 aligns well with multiple disease resistance genes on the chromosome, including the \u003cem\u003eFrl\u003c/em\u003e gene (Fusarium crown and root rot), the \u003cem\u003eTm-2a\u003c/em\u003e gene (tomato and tobacco mosaic virus), the \u003cem\u003ePh-3\u003c/em\u003e gene (late blight caused by \u003cem\u003ePhytophthora infestans\u003c/em\u003e), and the \u003cem\u003eSw-5b\u003c/em\u003e gene (tomato spotted wilt virus), with the \u003cem\u003eTm-2a\u003c/em\u003e gene being associated with the largest introgression in the lines with resistance allele. Meanwhile, an extensive introgression associated with the \u003cem\u003eMi\u003c/em\u003e gene on chromosome 6 is a source of high concentration of SVs on the short arm of the chromosome. The high peak of SVs on the distal part of chromosome 7 is linked to the \u003cem\u003eI3\u003c/em\u003e gene (Fusarium wilt, race 3) (Du et al., \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e2025\u003c/span\u003e; Foolad, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e2007\u003c/span\u003e; Simko et al., \u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e2021\u003c/span\u003e). Beyond variations in fruit shape, size, and other horticultural traits, the cultivated tomatoes examined in this study also possess the disease-resistance genes referenced above. The association of genomic regions with low SV density on chromosomes 8 and 10, and of large segments on other chromosomes, with fruit and horticultural characteristics of cultivated tomatoes (fresh market and processing) is evident in this study but requires further investigation. Random and scattered structural variants in the tomato genome led to variation in plant growth habits, fruit size, flavor, disease resistance, and pathogen recognition (Gao et al., \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e2019\u003c/span\u003e; Jobson \u0026amp; Roberts, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e2022\u003c/span\u003e; M. B. Lee et al., \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e2022\u003c/span\u003e; Wang et al., \u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). However, introgression of disease resistance from wild relatives of cultivated tomato appears to be the primary source of high-density SV clusters in the tomato genome.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec20\" class=\"Section2\"\u003e \u003ch2\u003e4.4. Cophylogenetic relationship\u003c/h2\u003e \u003cp\u003eA cophylogenetic study using tanglegrams provided a visual and statistical phylogenomic comparison of two main genotyping approaches based on genetic variants. Baker correlation coefficients of 0.95 indicate statistically similar clustering between phylogenetic trees generated from short genetic variants and from major SVs. However, meaningful differences were observed among the clusters generated by the two different genetic variants. Interestingly, the phylogenetic tree generated from SV polymorphism among the 60 lines was more accurate given the known pedigree information for the tomato lines (NC State University, \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). For instance, one of the wild tomato accessions is clustered with grape tomato lines based on short genetic variants, but realigned with other wild tomato accessions when SVs were used to generate the tree. Additionally, a new breeding line from the NCSU tomato breeding program, \u0026lsquo;NC 233253\u0026rsquo;, with minimal NCSU genetics in its pedigree, clustered separately from all other lines in the SV-based phylogenetic tree. This finding implies that integrating traditional short variant polymorphisms with SV data can significantly improve genotyping resolution and, consequently, increase the accuracy of phylogenetic and marker-trait association studies. Given that both genotypic data can be generated from the same short sequencing reads.\u003c/p\u003e \u003cp\u003eIntegrating SV-derived distances with short variant genotypes tightens phylogenetic clustering and improves pedigree-consistent grouping, increasing confidence when making selection and crossing decisions. The tanglegram analysis, supported by high Baker\u0026rsquo;s gamma and strong cophenetic correlation, demonstrates that SV data enhances discriminatory power, uncovering genomic distinctions often missed by SNPs\u0026mdash;such as correctly grouping a wild accession with other wild types and isolating a breeding line outside the program pedigree. For commercial breeding teams, this improved resolution enables more accurate segment tracking, reduces misclassification, and ensures greater predictability in advanced lines.\u003c/p\u003e \u003cp\u003eEvidently, the frequency and physical map of SVs in tomatoes are highly associated with and clustered around the distribution of short genetic variants in the genome and can be used in combination with short variants to improve the mapping resolution of genomic studies. Additionally, structural variants are excellent targets for developing DNA markers using high-resolution melting (HRM) curve-based methods. Genotyping by HRM curve provides the most affordable genotyping platform for small breeding programs; however, DNA marker development based on SNVs is challenging because most cases show insufficient shift in melting temperature to discriminate heterozygosity (Lee et al., \u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e2018\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec21\" class=\"Section2\"\u003e \u003ch2\u003e4.5. Limitations\u003c/h2\u003e \u003cp\u003eWhile short-read SV calling (Manta) performed well across workflows, it may under-detect complex rearrangements or transgene insertion architectures, as evidenced by missed events in transformed lines with complicated breakpoints. In four transgenic tomato lines with a \u0026ldquo;Moneymaker\u0026rdquo; background and two related backcrossed lines, we confirmed seven unique SVs (private doubletons) flagged by Manta as potential transgene insertion sites. Yet, Manta missed additional insertion sites, likely due to the complexity of these regions in the host genome. We employed an alternative method to identify six additional insertion sites (unpublished data), illustrating both the strengths and limitations of the Manta-based approach. Additionally, batching improves discovery but does not guarantee capture of all SV classes; breakpoint-only detection can leave multi-break rearrangements partially unresolved. To mitigate risk in product development, tiered validation (visual inspection, orthogonal assays, targeted long-read confirmation for critical loci) should be embedded before large-scale deployment. It has been pointed out that Manta and other SV callers only detect breakpoints, whereas multiple breakpoints are characteristic of complex genomic rearrangements in an organism (Shale et al., \u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). This study leveraged short-read sequencing to identify widespread structural variants (SVs) and map their positions on the tomato genome, highlighting the need for deeper investigations into complex genomic rearrangements and the genetic architecture of cultivated and wild tomatoes in genome\u0026ndash;phenome interaction research.\u003c/p\u003e \u003c/div\u003e"},{"header":"Conclusion","content":"\u003cp\u003eShort-read SV detection offers a scalable, cost-effective strategy for integrating structural variation into tomato genomics and breeding pipelines, unlocking hidden genetic diversity and improving the accuracy of phylogenetic and marker-trait association studies.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe thank Dr. Dorith Rotenberg and Anna Whitefield (Professors at the Entomology and Plant Pathology Department, NCSU) for supplying the transgenic lines and Dr. Randy Gardner (NC State University, Emeritus Professor) for providing seeds from NC GEM tomato breeding lines.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis research was supported by multiple sources. An internal startup grant from the Department of Horticultural Science at North Carolina State University provided initial funding for project development. Additional support for next-generation sequencing (NGS) data generation was provided through federal and state programs, including:\u003c/p\u003e\n\u003cul class=\"decimal_type\"\u003e\n \u003cli\u003eUSDA National Institute of Food and Agriculture (USDA-NIFA), Grant No. 2022-67013-37076\u003c/li\u003e\n \u003cli\u003eUS Department of Agriculture Agricultural Marketing Service (USDA/AMS) and North Carolina Department of Agriculture and Consumer Services (NCDACS), Grant Nos. 23SCBPNC1202-00 and AM200100XXXXG061\u003c/li\u003e\n \u003cli\u003eCalifornia Tomato Research Institute (CTRI), Grant No. 2021 - 294\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003e\u003cstrong\u003eConflicts of Interest\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe author declares no conflicts of interest.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData Availability Statement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eSupporting data are provided within the article and supplementary files or available from the corresponding author upon reasonable request. Whole-genome sequencing datasets remain restricted due to ongoing research and ethical constraints.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eAlonge, M., Wang, X., Benoit, M., Soyk, S., Pereira, L., Zhang, L., Suresh, H., Ramakrishnan, S., Maumus, F., Ciren, D., Levy, Y., Harel, T. H., Shalev-Schlosser, G., Amsellem, Z., Razifard, H., Caicedo, A. L., Tieman, D. M., Klee, H., Kirsche, M., \u0026hellip; Lippman, Z. B. (2020). Major Impacts of Widespread Structural Variation on Gene Expression and Crop Improvement in Tomato. \u003cem\u003eCell\u003c/em\u003e, \u003cem\u003e182\u003c/em\u003e(1). https://doi.org/10.1016/j.cell.2020.05.021\u003c/li\u003e\n\u003cli\u003eBaker, F. B. (1974). Stability of Two Hierarchical Grouping Techniques Case 1: Sensitivity to Data Errors. In \u003cem\u003eSource: Journal of the American Statistical Association\u003c/em\u003e (Vol. 69, Issue 346).\u003c/li\u003e\n\u003cli\u003eBradbury, P. J., Zhang, Z., Kroon, D. E., Casstevens, T. M., Ramdoss, Y., \u0026amp; Buckler, E. S. (2007). TASSEL: Software for association mapping of complex traits in diverse samples. \u003cem\u003eBioinformatics\u003c/em\u003e, \u003cem\u003e23\u003c/em\u003e(19), 2633\u0026ndash;2635. https://doi.org/10.1093/bioinformatics/btm308\u003c/li\u003e\n\u003cli\u003eChen, X., Schulz-Trieglaff, O., Shaw, R., Barnes, B., Schlesinger, F., K\u0026auml;llberg, M., Cox, A. J., Kruglyak, S., \u0026amp; Saunders, C. T. (2016). Manta: Rapid detection of structural variants and indels for germline and cancer sequencing applications. \u003cem\u003eBioinformatics\u003c/em\u003e, \u003cem\u003e32\u003c/em\u003e(8), 1220\u0026ndash;1222. https://doi.org/10.1093/bioinformatics/btv710\u003c/li\u003e\n\u003cli\u003eDanecek, P., Auton, A., Abecasis, G., Albers, C. A., Banks, E., DePristo, M. A., Handsaker, R. E., Lunter, G., Marth, G. T., Sherry, S. T., McVean, G., \u0026amp; Durbin, R. (2011). The variant call format and VCFtools. \u003cem\u003eBioinformatics\u003c/em\u003e, \u003cem\u003e27\u003c/em\u003e(15), 2156\u0026ndash;2158. https://doi.org/10.1093/bioinformatics/btr330\u003c/li\u003e\n\u003cli\u003eDanecek, P., Bonfield, J. K., Liddle, J., Marshall, J., Ohan, V., Pollard, M. O., Whitwham, A., Keane, T., McCarthy, S. A., \u0026amp; Davies, R. M. (2021). Twelve years of SAMtools and BCFtools. \u003cem\u003eGigaScience\u003c/em\u003e, \u003cem\u003e10\u003c/em\u003e(2). https://doi.org/10.1093/gigascience/giab008\u003c/li\u003e\n\u003cli\u003eDu, M., Sun, C., Deng, L., Zhou, M., Li, J., Du, Y., Ye, Z., Huang, S., Li, T., Yu, J., Li, C. B., \u0026amp; Li, C. (2025). Molecular breeding of tomato: Advances and challenges. In \u003cem\u003eJournal of Integrative Plant Biology\u003c/em\u003e (Vol. 67, Issue 3, pp. 669\u0026ndash;721). John Wiley and Sons Inc. https://doi.org/10.1111/jipb.13879\u003c/li\u003e\n\u003cli\u003eFoolad, M. R. (2007). Genome mapping and molecular breeding of tomato. \u003cem\u003eInternational Journal of Plant Genomics\u003c/em\u003e, \u003cem\u003e2007\u003c/em\u003e. https://doi.org/10.1155/2007/64358\u003c/li\u003e\n\u003cli\u003eFuentes, R. R., Chebotarov, D., Duitama, J., Smith, S., Hoz, J. F. D. La, Mohiyuddin, M., Wing, R. A., Mcnally, K. L., Tatarinova, T., Grigoriev, A., Mauleon, R., \u0026amp; Alexandrov, N. (2019). Structural variants in 3000 rice genomes. In \u003cem\u003eGenome Research\u003c/em\u003e (Vol. 29, Issue 5, p. 870). https://doi.org/10.1101/gr.241240.118\u003c/li\u003e\n\u003cli\u003eGalili, T. (2015). dendextend: An R package for visualizing, adjusting and comparing trees of hierarchical clustering. \u003cem\u003eBioinformatics\u003c/em\u003e, \u003cem\u003e31\u003c/em\u003e(22), 3718\u0026ndash;3720. https://doi.org/10.1093/bioinformatics/btv428\u003c/li\u003e\n\u003cli\u003eGao, L., Gonda, I., Sun, H., Ma, Q., Bao, K., Tieman, D. M., Burzynski-Chang, E. A., Fish, T. L., Stromberg, K. A., Sacks, G. L., Thannhauser, T. W., Foolad, M. R., Diez, M. J., Blanca, J., Canizares, J., Xu, Y., Knaap, E. Van Der, Huang, S., Klee, H. J., \u0026hellip; Fei, Z. (2019). The tomato pan-genome uncovers new genes and a rare allele regulating fruit flavor. In \u003cem\u003eNature Genetics\u003c/em\u003e (Vol. 51, Issue 6, p. 1044). https://doi.org/10.1038/s41588-019-0410-2\u003c/li\u003e\n\u003cli\u003eGardner, R. G. (1982). \u003cem\u003eNC50-7 Breeding Line, \u0026ldquo;Cherokee\u0026rdquo;, and \u0026ldquo;Mountain Pride\u0026rdquo; Tomato\u003c/em\u003e.\u003c/li\u003e\n\u003cli\u003eGardner, R. G., \u0026amp; Panthee, D. R. (2012). Tomato Spotted Wilt Virus-resistant Fresh-market Tomato Breeding Lines: NC 58S, NC 123S, NC 127S, and NC 132S. In \u003cem\u003eHORTSCIENCE\u003c/em\u003e (Vol. 47, Issue 4). http://www.ces.ncsu.edu/\u003c/li\u003e\n\u003cli\u003eHosmani, P. S., Flores-Gonzalez, M., van de Geest, H., Maumus, F., Bakker, L. V., Schijlen, E., van Haarst, J., Cordewener, J., Sanchez-Perez, G., Peters, S., Fei, Z., Giovannoni, J. J., Mueller, L. A., \u0026amp; Saha, S. (2019). \u003cem\u003eAn improved de novo assembly and annotation of the tomato reference genome using single-molecule sequencing, Hi-C proximity ligation and optical maps\u003c/em\u003e. https://doi.org/10.1101/767764\u003c/li\u003e\n\u003cli\u003eJobson, E., \u0026amp; Roberts, R. (2022). Genomic structural variation in tomato and its role in plant immunity. In \u003cem\u003eMolecular Horticulture\u003c/em\u003e (Vol. 2, Issue 1). https://doi.org/10.1186/s43897-022-00029-w\u003c/li\u003e\n\u003cli\u003eKrzywinski, M., Schein, J., Birol, I., Connors, J., Gascoyne, R., Horsman, D., Jones, S. J., \u0026amp; Marra, M. A. (2009). Circos: An information aesthetic for comparative genomics. \u003cem\u003eGenome Research\u003c/em\u003e, \u003cem\u003e19\u003c/em\u003e(9), 1639\u0026ndash;1645. https://doi.org/10.1101/gr.092759.109\u003c/li\u003e\n\u003cli\u003eLangmead, B., \u0026amp; Salzberg, S. L. (2012). Fast gapped-read alignment with Bowtie 2. \u003cem\u003eNature Methods\u003c/em\u003e, \u003cem\u003e9\u003c/em\u003e(4), 357\u0026ndash;359. https://doi.org/10.1038/nmeth.1923\u003c/li\u003e\n\u003cli\u003eLee, M. B., Shekasteband, R., Hutton, S. F., \u0026amp; Lee, T. G. (2022). A mutant allele of the flowering promoting factor 1 gene at the tomato BRACHYTIC locus reduces plant height with high quality fruit. \u003cem\u003ePlant Direct\u003c/em\u003e, \u003cem\u003e6\u003c/em\u003e(8). https://doi.org/10.1002/pld3.422\u003c/li\u003e\n\u003cli\u003eLee, T. G., Shekasteband, R., Menda, N., Mueller, L. A., \u0026amp; Hutton, S. F. (2018). Molecular markers to select for the j-2\u0026ndash;mediated jointless pedicel in tomato. \u003cem\u003eHortScience\u003c/em\u003e, \u003cem\u003e53\u003c/em\u003e(2). https://doi.org/10.21273/HORTSCI12628-17\u003c/li\u003e\n\u003cli\u003eLi, H., Handsaker, B., Wysoker, A., Fennell, T., Ruan, J., Homer, N., Marth, G., Abecasis, G., \u0026amp; Durbin, R. (2009). The Sequence Alignment/Map format and SAMtools. \u003cem\u003eBioinformatics\u003c/em\u003e, \u003cem\u003e25\u003c/em\u003e(16), 2078\u0026ndash;2079. https://doi.org/10.1093/bioinformatics/btp352\u003c/li\u003e\n\u003cli\u003eLi, N., He, Q., Wang, J., Wang, B., Zhao, J., Huang, S., Yang, T., Tang, Y., Yang, S., Aisimutuola, P., Xu, R., Hu, J., Jia, C., Ma, K., Li, Z., Jiang, F., Gao, J., Lan, H., Zhou, Y., \u0026hellip; Yu, Q. (2023). Super-pangenome analyses highlight genomic diversity and structural variation across wild and cultivated tomato species. \u003cem\u003eNature Genetics\u003c/em\u003e, \u003cem\u003e55\u003c/em\u003e(5). https://doi.org/10.1038/s41588-023-01340-y\u003c/li\u003e\n\u003cli\u003eLiu, L., Zhang, K., Bai, J., Lu, J., Lu, X., Hu, J., Pan, C., He, S., Yuan, J., Zhang, Y., Zhang, M., Guo, Y., Wang, X., Huang, Z., Du, Y., Cheng, F., \u0026amp; Li, J. (2022). All-flesh fruit in tomato is controlled by reduced expression dosage of AFF through a structural variant mutation in the promoter. \u003cem\u003eJournal of Experimental Botany\u003c/em\u003e, \u003cem\u003e73\u003c/em\u003e(1). https://doi.org/10.1093/jxb/erab401\u003c/li\u003e\n\u003cli\u003eNC State University. (2024, April). \u003cem\u003ehttps://mountainhort.ces.ncsu.edu/fresh-market-tomato-breeding/\u003c/em\u003e.\u003c/li\u003e\n\u003cli\u003ePascual, L., Albert, E., Sauvage, C., Duangjit, J., Bouchet, J. P., Bitton, F., Desplat, N., Brunel, D., Le Paslier, M. C., Ranc, N., Bruguier, L., Chauchard, B., Verschave, P., \u0026amp; Causse, M. (2016). Dissecting quantitative trait variation in the resequencing era: Complementarity of bi-parental, multi-parental and association panels. \u003cem\u003ePlant Science\u003c/em\u003e, \u003cem\u003e242\u003c/em\u003e, 120\u0026ndash;130. https://doi.org/10.1016/j.plantsci.2015.06.017\u003c/li\u003e\n\u003cli\u003ePrasanna, H., Rai, N., Hussain, Z., Yerasu, S. R., \u0026amp; Tiwari, J. K. (2023). Tomato: Breeding and Genomics. \u003cem\u003eVegetable Science\u003c/em\u003e, \u003cem\u003e50\u003c/em\u003e(Special). https://doi.org/10.61180/vegsci.2023.v50.spl.02\u003c/li\u003e\n\u003cli\u003eQuinlan, A. R., \u0026amp; Hall, I. M. (2010). BEDTools: A flexible suite of utilities for comparing genomic features. \u003cem\u003eBioinformatics\u003c/em\u003e, \u003cem\u003e26\u003c/em\u003e(6), 841\u0026ndash;842. https://doi.org/10.1093/bioinformatics/btq033\u003c/li\u003e\n\u003cli\u003eSato, S., Tabata, S., Hirakawa, H., Asamizu, E., Shirasawa, K., Isobe, S., Kaneko, T., Nakamura, Y., Shibata, D., Aoki, K., Egholm, M., Knight, J., Bogden, R., Li, C., Shuang, Y., Xu, X., Pan, S., Cheng, S., Liu, X., \u0026hellip; Gianese, G. (2012). The tomato genome sequence provides insights into fleshy fruit evolution. \u003cem\u003eNature\u003c/em\u003e, \u003cem\u003e485\u003c/em\u003e(7400), 635\u0026ndash;641. https://doi.org/10.1038/nature11119\u003c/li\u003e\n\u003cli\u003eShale, C., Baber, J., Cameron, D. L., Wong, M., Cowley, M. J., Papenfuss, A. T., Cuppen, E., \u0026amp; Priestley, P. (2020). \u003cem\u003eUnscrambling cancer genomes via integrated analysis of structural variation and copy number\u003c/em\u003e. https://doi.org/10.1101/2020.12.03.410860\u003c/li\u003e\n\u003cli\u003eSimko, I., Jia, M., Venkatesh, J., Kang, B. C., Weng, Y., Barcaccia, G., Lanteri, S., Bhattarai, G., \u0026amp; Foolad, M. R. (2021). Genomics and Marker-Assisted Improvement of Vegetable Crops. \u003cem\u003eCritical Reviews in Plant Sciences\u003c/em\u003e, \u003cem\u003e40\u003c/em\u003e(4), 303\u0026ndash;365. https://doi.org/10.1080/07352689.2021.1941605\u003c/li\u003e\n\u003cli\u003eTurner, D. J., Miretti, M., Rajan, D., Fiegler, H., Carter, N. P., Blayney, M. L., Beck, S., \u0026amp; Hurles, M. E. (2007). Germline rates of de novo meiotic deletions and duplications causing several genomic disorders. In \u003cem\u003eNature Genetics\u003c/em\u003e (Vol. 40, Issue 1, p. 90). https://doi.org/10.1038/ng.2007.40\u003c/li\u003e\n\u003cli\u003eWang, X., Gao, L., Jiao, C., Stravoravdis, S., Hosmani, P. S., Saha, S., Zhang, J., Mainiero, S., Strickler, S. R., Catala, C., Martin, G. B., Mueller, L. A., Vrebalov, J., Giovannoni, J. J., Wu, S., \u0026amp; Fei, Z. (2020). Genome of Solanum pimpinellifolium provides insights into structural variants during tomato breeding. In \u003cem\u003eNature Communications\u003c/em\u003e (Vol. 11, Issue 1). https://doi.org/10.1038/s41467-020-19682-0\u003c/li\u003e\n\u003cli\u003eWei, T., \u0026amp; Simko, V. (2024). \u003cem\u003eR package \u0026ldquo;corrplot\u0026rdquo;: Visualization of a Correlation Matrix (Version 0.95). Available from https://github.com/taiyun/corrplot\u003c/em\u003e.\u003c/li\u003e\n\u003cli\u003eZhou, Y., Zhang, Z., Bao, Z., Li, H., Lyu, Y., Zan, Y., Wu, Y., Cheng, L., Fang, Y., Wu, K., Zhang, J., Lyu, H., Lin, T., Gao, Q., Saha, S., Mueller, L., Fei, Z., St\u0026auml;dler, T., Xu, S., \u0026hellip; Huang, S. (2022). Graph pangenome captures missing heritability and empowers tomato breeding. \u003cem\u003eNature\u003c/em\u003e, \u003cem\u003e606\u003c/em\u003e(7914). https://doi.org/10.1038/s41586-022-04808-9\u003c/li\u003e\n\u003c/ol\u003e"},{"header":"Tables","content":"\u003cp\u003e\u003cstrong\u003eTable 1\u003c/strong\u003e. The number of monomorphic and polymorphic structural variants present in 60 lines is in the collection.\u003c/p\u003e\n\u003ctable border=\"0\" cellspacing=\"0\" cellpadding=\"0\" width=\"100%\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"2\" valign=\"bottom\" style=\"width: 33px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eCultivate tomatoes\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 11px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eINS\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 9px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eDEL\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 11px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eBNDa\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 11px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eBNDb\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 11px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eDUP\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 11px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eTotal\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"2\" style=\"width: 17px;\"\u003e\n \u003cp\u003eMonomophic\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 16px;\"\u003e\n \u003cp\u003eHomozygous\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 9px;\"\u003e\n \u003cp\u003e12\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e18\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 16px;\"\u003e\n \u003cp\u003eHeterozygous\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 9px;\"\u003e\n \u003cp\u003e4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e6\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"2\" style=\"width: 17px;\"\u003e\n \u003cp\u003ePolymorphic\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 16px;\"\u003e\n \u003cp\u003eHomozygous\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e8,571\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 9px;\"\u003e\n \u003cp\u003e7,787\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e314\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e81\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e32\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e16,785\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 16px;\"\u003e\n \u003cp\u003eHeterozygous\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e4,069\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 9px;\"\u003e\n \u003cp\u003e2,518\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e368\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e1241\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e320\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e8,516\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 17px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eTotal\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 16px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e12,647\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 9px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e10,321\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e682\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e1322\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e353\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e25,325\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"8\" valign=\"bottom\" style=\"width: 100px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eWild tomato accessions\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 17px;\"\u003e\n \u003cp\u003eMonomorphic\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 16px;\"\u003e\n \u003cp\u003eHomozygous\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e2,663\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 9px;\"\u003e\n \u003cp\u003e1,392\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e69\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e164\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e30\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e4,318\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 17px;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 16px;\"\u003e\n \u003cp\u003eHeterozygous\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e25\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 9px;\"\u003e\n \u003cp\u003e116\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e22\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e38\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e13\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e214\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 17px;\"\u003e\n \u003cp\u003ePolymorphic\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 16px;\"\u003e\n \u003cp\u003eHomozygous\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e12,794\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 9px;\"\u003e\n \u003cp\u003e18,110\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e291\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e157\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e20\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e31,372\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 17px;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 16px;\"\u003e\n \u003cp\u003eHeterozygous\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e1,284\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 9px;\"\u003e\n \u003cp\u003e6,008\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e722\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e1,488\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e531\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e10,033\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 17px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eTotal\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 16px;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e16,766\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 9px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e25,626\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e1,104\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e1,847\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e594\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e45,937\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 2\u003c/strong\u003e. Number of monomorphic and polymorphic short genetic variants (SNPs and short InDels) present in 60 lines in the collection\u003c/p\u003e\n\u003ctable border=\"0\" cellspacing=\"0\" cellpadding=\"0\" width=\"100%\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"2\" valign=\"bottom\" style=\"width: 47px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eCultivated tomatoes\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eSNPs\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eInDels\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eTotal\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"2\" valign=\"top\" style=\"width: 24px;\"\u003e\n \u003cp\u003eMonomorphic\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 23px;\"\u003e\n \u003cp\u003eHomozygous\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18px;\"\u003e\n \u003cp\u003e48,715\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16px;\"\u003e\n \u003cp\u003e10,664\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18px;\"\u003e\n \u003cp\u003e59,379\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 23px;\"\u003e\n \u003cp\u003eHeterozygous\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18px;\"\u003e\n \u003cp\u003e7,173\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16px;\"\u003e\n \u003cp\u003e246\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18px;\"\u003e\n \u003cp\u003e7,419\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"2\" valign=\"top\" style=\"width: 24px;\"\u003e\n \u003cp\u003ePolymorphic\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 23px;\"\u003e\n \u003cp\u003eHomozygous\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18px;\"\u003e\n \u003cp\u003e5,200,573\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16px;\"\u003e\n \u003cp\u003e474,305\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18px;\"\u003e\n \u003cp\u003e5,674,878\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 23px;\"\u003e\n \u003cp\u003eHeterozygous\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18px;\"\u003e\n \u003cp\u003e4,843,770\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16px;\"\u003e\n \u003cp\u003e368,167\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18px;\"\u003e\n \u003cp\u003e5,211,937\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 24px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eTotal\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 23px;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e10,100,231\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e853,382\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e10,953,613\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"5\" valign=\"bottom\" style=\"width: 100px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eWild tomato accessions\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 24px;\"\u003e\n \u003cp\u003eMonomorphic\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 23px;\"\u003e\n \u003cp\u003eHomozygous\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18px;\"\u003e\n \u003cp\u003e5,818,629\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16px;\"\u003e\n \u003cp\u003e57,817\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18px;\"\u003e\n \u003cp\u003e5,876,446\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 24px;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 23px;\"\u003e\n \u003cp\u003eHeterozygous\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18px;\"\u003e\n \u003cp\u003e218,840\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16px;\"\u003e\n \u003cp\u003e2,082\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18px;\"\u003e\n \u003cp\u003e220,922\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 24px;\"\u003e\n \u003cp\u003ePolymorphic\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 23px;\"\u003e\n \u003cp\u003eHomozygous\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18px;\"\u003e\n \u003cp\u003e20,477,542\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16px;\"\u003e\n \u003cp\u003e1,587,524\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18px;\"\u003e\n \u003cp\u003e22,065,066\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 24px;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 23px;\"\u003e\n \u003cp\u003eHeterozygous\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18px;\"\u003e\n \u003cp\u003e17,595,414\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16px;\"\u003e\n \u003cp\u003e656,751\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18px;\"\u003e\n \u003cp\u003e18,252,165\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 24px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eTotal\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 23px;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e44,110,425\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 16px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e2,304,174\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e46,414,599\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Structural variants (SVs), Tomato genomics, Short-read sequencing, Marker-assisted breeding, Genetic diversity, Manta SV caller, Phylogenetic analysis","lastPublishedDoi":"10.21203/rs.3.rs-8628781/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8628781/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eStructural variants (SVs)-large genomic alterations such as insertions, deletions, duplications, and translocations-are widespread in tomato genomes and play a critical role in phenotypic diversity. However, their detection has traditionally depended on expensive long-read sequencing technologies. Consistent with previous studies on structural variation, this study demonstrates a cost-effective approach for SV discovery using repurposed short-read sequencing data (150 bp), enabling integration of SVs into breeding workflows without additional sequencing investment. Using Illumina whole-genome data from 60 diverse tomato lines, including wild accessions, landraces, transgenic lines, and modern breeding lines, we identified over 71,000 high-confidence SVs, including a significant number of private doubletons with the Manta caller, as well as 10.9\u0026nbsp;million short genetic variants. SVs were unevenly distributed across chromosomes, clustering in subtelomeric regions and near disease-resistance loci, with chromosomes 6, 7, and 9 showing the highest densities. Wild accessions harbored nearly twice as many SVs as cultivated lines, with deletions dominating wild genomes and insertions more prevalent in cultivated tomatoes than in the reference genome. Comparative phylogenetic analysis revealed strong concordance between SV-based and SNP/InDel-based trees (Baker\u0026rsquo;s γ\u0026thinsp;=\u0026thinsp;0.95), while SV data improved pedigree-consistent clustering and resolved ambiguous lineage relationships. These findings highlight SVs as hidden drivers of tomato diversity and valuable resources for marker-assisted selection, trait mapping, and genomic studies. Mining archived short-read datasets could offer breeding programs a scalable, low-cost strategy to unlock latent SV information, accelerate genetic improvement, and enhance genome-to-phenome insights. Limitations include under-detection of complex rearrangements, warranting targeted validation for critical loci.\u003c/p\u003e","manuscriptTitle":"Beyond SNPs: Scalable Detection of Structural Variants Unlocks Hidden Genetic Diversity in Tomato","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-03-10 16:07:53","doi":"10.21203/rs.3.rs-8628781/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"99d686b0-07b7-4a64-b7f2-618482e54d4f","owner":[],"postedDate":"March 10th, 2026","published":true,"recentEditorialEvents":[{"type":"decision","content":"Rejected","date":"2026-05-15T12:00:30+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-05-01T17:18:22+00:00","index":23,"fulltext":""}],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2026-05-15T12:10:22+00:00","versionOfRecord":[],"versionCreatedAt":"2026-03-10 16:07:53","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8628781","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8628781","identity":"rs-8628781","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.