Transposable element-mediated structural variations drive gene expression and agronomic trait diversity in Brassica napus

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract Brassica napus is a globally significant polyploid oilseed crop with a complex genome, harbors substantial yet largely unexplored genetic diversity. In this study, we upgraded the genomes of the elite rapeseed cultivar Zhongshuang 11 (ZS11) and the spring cultivar Westar and their genome annotations. Additionally, we assembled the genomes of six other representative varieties, namely ZS2, Bugle, 352, 862, Ribenyoucai (RB), and Baihua (BH), and constructed a comprehensive pan-genome comprising 101,903 non-redundant pan-gene clusters derived from 22 genomes. Analysis of the pan-genome led to the identification of 124,682 transposable element- mediated structural variations (TEMSVs), which are unevenly distributed among subgenomes and have a significant impact on gene expression. Notably, this uneven distribution correlates with unbalanced expression of homoeologous gene pairs (uHGPs), with the degree of imbalance increasing alongside the number of TEMSVs. These uHGPs consistently exhibit biased expression patterns across different tissues. Furthermore, using structural variation-based genome-wide association studies (SV-GWAS), we identified two candidate genes, BnaA06.PC78 and BnaC03.UCU1, which are associated with seed oil content and fatty acid composition. This study offers valuable insights and genetic resources for elucidating gene expression dynamics among subgenomes and advancing B. napus breeding.
Full text 152,189 characters · extracted from preprint-html · click to expand
Transposable element-mediated structural variations drive gene expression and agronomic trait diversity in Brassica napus | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Transposable element-mediated structural variations drive gene expression and agronomic trait diversity in Brassica napus Hu Zhao, Zengdong Tan, Yuyu Zheng, Zhilin Guan, Xuqing Wang, Jianshun Yang, and 5 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6452497/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted You are reading this latest preprint version Abstract Brassica napus is a globally significant polyploid oilseed crop with a complex genome, harbors substantial yet largely unexplored genetic diversity. In this study, we upgraded the genomes of the elite rapeseed cultivar Zhongshuang 11 (ZS11) and the spring cultivar Westar and their genome annotations. Additionally, we assembled the genomes of six other representative varieties, namely ZS2, Bugle, 352, 862, Ribenyoucai (RB), and Baihua (BH), and constructed a comprehensive pan-genome comprising 101,903 non-redundant pan-gene clusters derived from 22 genomes. Analysis of the pan-genome led to the identification of 124,682 transposable element- mediated structural variations (TEMSVs), which are unevenly distributed among subgenomes and have a significant impact on gene expression. Notably, this uneven distribution correlates with unbalanced expression of homoeologous gene pairs (uHGPs), with the degree of imbalance increasing alongside the number of TEMSVs. These uHGPs consistently exhibit biased expression patterns across different tissues. Furthermore, using structural variation-based genome-wide association studies (SV-GWAS), we identified two candidate genes, BnaA06.PC78 and BnaC03.UCU1 , which are associated with seed oil content and fatty acid composition. This study offers valuable insights and genetic resources for elucidating gene expression dynamics among subgenomes and advancing B. napus breeding. Biological sciences/Plant sciences/Natural variation in plants Biological sciences/Genetics/Genome/Genetic variation Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Introduction Brassica napus (AACC, 2n = 38) is a leading oilseed crop cultivated globally for edible vegetable oil, animal feed and vegetables [ 1 ]. In 2022, the total production of rapeseed was 87.2 million tonnes, with an average yield of 2.18 tonnes per hectare (FAO, 2022; http://faostat.fao.org ). As an allotetraploid species, B. napus originated approximately 7,500 years ago from an interspecific hybridization between two diploid progenitors B. rapa (AA, 2n = 20) and B. oleracea (CC, 2n = 18) [ 2 , 3 ]. This unique genomic composition has spurred extensive research into its genetic architecture, particularly the divergent evolutionary trajectories of its subgenomes [ 4 , 5 ]. Previous studies on polyploid subgenomes have predominantly focused on genomic genetic variation [ 6 ], epigenetic modification and gene expression divergence [ 7 – 9 ], selected genes [ 10 ] and trait formation [ 11 ] and evolution [ 12 , 13 ]. Investigations have revealed imbalances in gene expression, regulation, and levels of active epigenetic signal modifications between the subgenomes of B. napus [ 4 , 9 ]. Additionally, SNP-eQTL mapping has uncovered feedback regulation between homoeologous gene pairs, suggesting mechanisms that maintain partial expression dosage across subgenomes [ 3 ]. Despite these advances, structural variations (SVs) are increasingly recognized as a major source of genetic diversity, often explaining more phenotypic variation than single nucleotide polymorphisms (SNPs) [ 14 ]. Graph-based pan-genomics studies, along with cataloging genetic variations, are instrumental in dissecting crop domestication, environmental adaptation, and phenotypic diversification [ 15 , 16 ]. For instance, analysis of SV effects on glucosinolate biosynthesis and transport pathways in rapeseed demonstrated the importance of SVs in remodeling key traits [ 17 ]. Among the various types of SVs, transposable element (TE) insertions are particularly significant because TEs tend to be highly enriched near genes and can exert profound influences on gene function and phenotype [ 18 , 19 ]. A notable case is the 3.9-kb CACTA-like TE insertion in the promoter of BnaA9.CYP78A9 , which significantly correlates with variations in silique length and seed weight [ 20 ]. However, most research on agronomically important traits in B. napus has focused on SNP-level variations, leaving the impact of large-scale SVs on trait domestication and improvement largely unexplored. Recent advances in long-read sequencing technologies [ 21 , 22 ] have transformed plant genome research by enabling a deeper exploration of intraspecific genetic diversity, particularly in economically important crops [ 23 – 25 ]. The pan-genome has emerged as a revolutionary approach that captures the full spectrum of genetic diversity, thus facilitating the comprehensive characterization of SVs [ 22 , 25 – 27 ]. Despite these advances in crops like maize and rice, the number of high-quality reference genomes available for B. napus remains limited [ 3 ]. To address this gap, we refined near-complete genomes of two representative B. napus accessions (ZS11 and Westar) and newly generated chromosome-level assemblies for six additional varieties (ZS2, Bugle, 352, 862, RB, and BH). Combined with published genomes, we constructed a graph-based pan-genome containing 22 high-quality B. napus genomes. Analysis of seed RNA-seq data from 505 B. napus accessions revealed that transposable element–mediated structural variations (TEMSVs) can significantly alter gene expression patterns across subgenomes. Moreover, SV-based genome-wide association studies (SV-GWAS) identified two SVs associated with seed oil content (SOC) and seed fatty acid content (SFAC). Our findings suggest that the expression imbalance of homoeologous gene pairs among subgenomes may be driven by TEMSVs, highlighting the impact of SVs on key traits in B. napus . This study introduces a graph-based pan-genome of B. napus as a valuable tool for understanding SVs and advancing rapeseed genetic improvement. Results Two near-complete genome assemblies in B. napus To upgrade the reference genomes of ZS11 and Westar, we sequenced the two genomes with Oxford Nanopore Technology (ONT) and obtained 154.13 GB (~ 154 × coverage) and 81.31 GB (~ 81 × coverage) of ultra-long reads, respectively (Supplementary Table 1). Combined with the PacBio data generated in previous study [ 20 ], each genome was independently assembled and error-corrected using three methods, including NextDenovo, Flye and NECAT. Subsequently, the preliminary assemblies (ZS11 v0 and Westar v0) were used as integration backbones, leading to the production of two near-complete genomes, designated ZS11 v1 and Westar v1, which contained 15 and 1 gap-free chromosomes, respectively. The final ZS11 v1 assembly reached 1,014 Mb with only 38 gaps (Contig N50: 58.66 Mb), marking a 38.9-fold improvement in contiguity over ZS11 v0 assembly (999 Mb, 4,470 gaps, Contig N50: 1.51 Mb). Similarly, Westar v1 achieved 984 Mb with 94 gaps (Contig N50: 23.71 Mb), representing a 7.6-fold enhancement compared to Westar v0 (982.41 Mb, 2,587 gaps, Contig N50: 3.13 Mb) (Fig. 1 A). By searching for telomere-specific tandem repeats, we identified telomeric sequences at the ends of 16 chromosomes in ZS11 v1 and Westar v1 (Fig. 1 B, Supplementary Fig. 1). Given the high abundance of repetitive sequences in centromeric regions and the limited coverage of genetic markers, centromeres are notoriously challenging to assemble. Here, we mapped Brassicaceae centromere-specific repetitive sequences to the ZS11 v1 and Westar v1 genomes (Fig. 1 B, Supplementary Fig. 1). In the upgraded ZS11 and Westar reference genomes, the identified centromeric region spanned 153.28 Mb and 136.71 Mb (Supplementary Table 2–3). Among them, ZS11 v1 with the number of annotated genes (N v1 = 1,099) within this region increasing by 78.89% compared to the ZS11 v0. Furthermore, the proportion of expressed genes in the centromeres rose to 48.32% in ZS11 v1, compared to 34.34% in the v0 assembly (Supplementary Fig. 2). Gene Ontology (GO) enrichment analysis further revealed that the genes within the newly assembled centromeric regions were significantly enriched in pathways related to energy metabolism and photosynthetic electron transport (Supplementary Table 4). Overall, these findings indicate that our improved assembly offers a more complete and functionally informative centromeric region, providing a valuable resource for future functional genomic studies. To further validate the completeness of the ZS11 v1 and Westar v1 genomes, we aligned 273 RNA-seq datasets from 91 distinct tissues, including root, seedling, stem, leaves, flower buds, flower tissues, silique walls, and seeds [ 29 ], to the new assemblies. The average mapping rates for the ZS11 and Westar v1 genomes were 76.1% and 75.2%, respectively, representing improvements of 2.2% and 1.4% over their respective v0 genome assemblies (Supplementary Fig. 3). Additionally, the benchmarking universal single-copy orthologs (BUSCO) analysis revealed that the new assemblies captured 99.75% and 99.69% of the expected BUSCO genes for ZS11 v1 and Westar v1, respectively (Fig. 1 A). Collectively, these findings demonstrate a substantial improvement in genome quality and completeness. In the newly assembled ZS11 and Westar genomes, we annotated 111,715 and 114,412 genes, respectively, using a combination of multi-transcript integration, homologous species protein prediction, and de novo prediction methods (Fig. 1 A). The average gene lengths increased to 2,498.39 bp and 2,229.65 bp, as compared to their respective v0 assemblies, which had average gene lengths of 2,061.34 bp and 2,110.54 bp. As eukaryotic genes typically consist of multiple exons, the proportion of multi-exon genes increased from 78.79% and 78.86% in the v0 genomes (ZS11 and Westar) to 82.08% and 84.03% in the upgraded assemblies, indicating enhanced resolution in gene structure annotation. In addition, annotation of untranslated regions (UTRs) showed substantial improvement. Prior to the upgradation, 23,575 and 29,694 genes were annotated with UTRs in ZS11 and Westar, accounting for 30.70% and 31.14% of the total annotated genes, respectively. After the upgradation, these numbers increased to 34,785 and 35,130 genes, with a total of 69,762 and 73,725 UTRs annotated in the ZS11 v1 and Westar v1 genomes (Supplementary Table 5). Comparisons of transcript length, coding sequence (CDS) length, intron length, exon number, and exon length between the upgraded gene sets and those of B. oleracea and B. rapa revealed similar distributions, thereby supporting the reliability of our improved annotation (Supplementary Fig. 4). The ZS11 v0 genome has underpinned numerous functional genomics studies in B. napus [ 20 , 28 , 29 ]. However, the gene annotation in the ZS11 v0 genome remains suboptimal and warrants improvement. For example, only two copies of TRANSPARENT TESTA 4 ( TT4 ), a gene that known to affect seed coat content (SCC) and SOC [ 30 ], were previously annotated on chromosomes C01 and C02 in the ZS11 v0 genome. In contrast, 13 TT4 copies were identified in the ZS11 v1 genome, which is consistent with the copy number reported in the Darmor-bzh v10 genome [ 31 ]. Additionally, an unassigned copy in the Darmor-bzh v10 genome was successfully assigned to chromosome C09 in the ZS11 v1, and gene expression analysis confirmed that 10 of these TT4 copies were actively expressed in ZS11 developing seeds (Fig. 1 C). Two GLUCOSINOLATE TRANSPORTER-2 ( GTR2 ) genes on A09 and C02 encoding key transporter proteins were identified to control seed glucosinolates content (SGC) using the Darmor-bzh v10 genome[ 32 ] and the ZS11 v1 genome. However, these two genes were not annotated in the ZS11 v0 genome (Fig. 1 D-E). Together, these findings underscore the importance of comprehensive and accurate gene annotation for advancing genomic research and for improving our understanding of key agronomic traits in B. napus . Chromosome-level assemblies of six diverse accessions of Brassica napus To comprehensively capture genome variations of B. napus , six additional representative varieties, namely ZS2, Bugle, 352, 862, RB, and BH, were selected for de novo assembly based on the population structure analysis and the phenotypic characterization from previous study [ 33 ]. PacBio sequencing yield 94.39 to 146.29 Gb of data for each variety. The assembled genome sizes ranged from 824.77 to 949.56 Mb, with Contig N50 values ranging from 1.44 to 4.62 Mb. A total of 90,960 to 98,777 protein-coding genes were annotated, and 98.60 to 99.09% of the conserved BUSCO genes were identified, demonstrating the high integrity of the assembled genomes (Supplementary Table 6). Gene-based pan-genome and core-genome analyses of 20 B. napus accessions In addition to the eight genomes assembled in this study, an additional 12 previously published genome assemblies with gene annotations (including Da-Ae, Tapior, Darmor, Express617, Gangan, No2127, P130, P202, Quinta, Shengli, Xiaoyun and Zheyou7) for subsequent analysis [ 20 , 31 , 34 – 37 ]. By clustering 1,911,393 predicted gene models from these 20 accessions, we constructed a protein-coding gene-based pan-genome with 101,903 non-redundant pan-gene clusters (i.e., the cumulative set of all genes), which is close to the number reported in a previous rapeseed pan-genome analysis [ 20 ]. These clusters were classified into four categories based on their frequency of occurrence: 17,877 core (present in all 20 accessions) and 13,096 softcore (present in at least 18 accessions), 66,748 dispensable clusters (present in 2 to 17 accessions) and 4,182 private clusters (present in only one accession). The private clusters represent 4.02% of the total gene sets (Fig. 2 A-B). Pan-genome size simulations indicated that the number of gene cluster approached saturation when more than 14 accessions were included, and the addition of the 20th accession contributed only 110 new clusters, demonstrating the representativeness of our selected accessions (Fig. 2 A-B). Notably, semi-winter and spring rapeseed exhibited a greater prevalence of diverse gene families (Supplementary Fig. 5A). Dispensable gene families accounted for 65.6% of the total clusters, contributing an average of 34.7% of a single accession’s genome, whereas core and softcore genes collectively accounted for 64.6% (Fig. 2 C-D). Compared with core/softcore genes, dispensable and private genes had fewer exons (Supplementary Fig. 5B) but displayed significantly higher non-synonymous and synonymous substitution ratio (Ka/Ks). In particular, 811 dispensable gene families enriched in nucleotide binding and ATP binding had a Ka/Ks ratio greater than 1, suggesting positive selection pressure in functional categories such as nucleotide binding and ATP binding (Fig. 2 E). Conversely, 53,129 gene families with Ka/Ks ratio less than 1 indicated purifying selection, particularly in functions related to the nucleus and metal ion binding (Supplementary Fig. 5C). GO analysis indicated that core genes were enriched in several essential biological processes including macromolecule modification, glycosylation, and phosphorus metabolic process, while dispensable genes were associated with processes like DNA integration, hormone responses, and telomere maintenance (Fig. 2 F; Supplementary Table 7). Finally, we constructed a comprehensive database BrassicaPanDB ( http://rgmi.hzau.edu.cn/pan/ ), to provide researchers with access to this pan-genome. The database includes functions for data download, genome browsing, gene search, and BLAST analysis, enabling users to explore the genetic diversity of these 20 B. napus accessions in depth. Extensive genetic variations and the variant-integrated graph-based pan-genome The constructed rapeseed graph-based pan-genome integrates extensive genetic variations, providing essential insights into sequence divergence. The analysis included the 20 genomes and two additional genomes (GH06 and ZY821) [ 38 ], enabling accurate identification of single nucleotide polymorphisms (SNPs), small insertions and deletions (InDels) and SVs (InDels and inversions ≥ 50 bp). Using Minigraph-Cactus [ 39 ], the sequences of 21 B. napus genomes were aligned against the ZS11 v1 genome. A total of 11,120,890 SNPs and 5,583,367 InDels (≤ 50 bp), as well as 221,062 non-redundant SVs, including 124,836 insertions (ranging from 50 bp to 839 kb), 95,685 deletions (ranging from 50 bp to 159 kb) and 541 inversions (ranging from 606 bp to 1.22 Mb) were detected (Fig. 3 A-B; Supplementary Table 8; Supplementary Fig. 5D-E). Of these SVs, 744 are present in all 21 query genomes, 4,479 present in 19–20 genomes, 130,975 present in 2–18 genomes, and 84,864 unique to individual genomes (Fig. 3 A; Supplementary Table 9). To verify the accuracy, a subset of SVs in known flowering-time genes (e.g., FLC and FT ) was validated using PacBio long reads, confirming high assembly quality (Supplementary Fig. 6). Presence-absence variants (PAVs) were the most abundant type of SV (Fig. 2 ; Supplementary Fig. 5D). A total of 220,521 non-redundant PAVs were included in subsequent analyses. To examine the distribution and specificity of PAVs among different genomes, we analyzed the patterns of shared and unique SVs across the rapeseed accessions. The number of shared SVs sharply declined for the first three genomes and slowly decreased thereafter (Fig. 3 B). The number of PAVs per genome ranged from 39,690 to 87,518, with the variety no2127 exhibiting the highest number of private PAVs, probably due to its synthetic origin (Fig. 3 C; Supplementary Table 9–10). We also found that the number of PAVs gradually declined as the distance of PAVs from the nearest gene increased. Notably, 49.6% SVs were located in genic and flanking regions (within 5 kb of the transcription start site), suggesting these SVs may affect gene expression (Fig. 3 D). Moreover, PAVs were enriched in intergenic repetitive regions, with 71.9% (N = 158,486) of PAVs being TEs, suggesting that TEs may have driven the formation of most PAVs in the rapeseed genome. We defined these types of SV as TE-mediated SVs (TEMSVs). Specifically, most TEMSVs were DNA/Helitron (most abundant in number) and long terminal repeat (LTR) transposons (longest in length) (Supplementary Fig. 7A). A set of 4,064 TEMSVs are located in the promoter regions (2 kb upstream of the translation-start sites) or gene bodies, with SVs being more prevalent in genes with low expression levels (Supplementary Fig. 7B). Spring and semi-winter type rapeseed varieties are widely cultivated throughout the world. By mapping reads of 505 B. napus accessions (comprising winter, semi-winter, and spring type) to this graph-based genome, a total of 431,765 SVs were identified in the population. A neighbor-joining tree constructed from these SVs distinctly grouped the different ecotypes, indicating that SVs are associated with ecotype differentiation (Supplementary Fig. 7B-C). Notably, the distribution of structural variants (SVs) was asymmetric between ecotypes, with the An subgenome containing a greater number of ecotype-specific SVs (i.e., structural variants that are present, absent, or differ in size in over 70% of individuals within a single ecotype). We also observed certain chromosomal regions enriched with SVs showing subpopulation imbalance. For example, ecotype-specific SVs on chromosomes A09 (18.5–32.7 Mb), C02 (2.1–2.9 Mb, 33.8–37.9 Mb), and C08 (3.4–6.5 Mb) suggest these regions may harbor loci involved in ecotype differentiation (Fig. 3 E). The interval from 18.5 to 32.7 Mb on chromosome C02 includes the extensively studied flowering time inhibitor gene BnaC02.FLC , hypothesized to mediate ecological differentiation through negative regulation of flowering time. Two known TEMSVs affect BnaC02.FLC expression [ 20 ], leading to variation in flowering time among individuals carrying these mutations. Additional TEMSVs were also found at 2.1–2.9 Mb on chromosome C02 and 3.4–6.5 Mb on chromosome C08, near genes involved in rapeseed development. For instance, BnaC02.SDP1 , which participates in the biosynthesis of linolenic acid, linoleic acid, and arachidonic acid, could potentially influence rapeseed SOC by regulating related pathways. A TEMSV located about 4 kb upstream of the BnaC02.SDP1 in semi-winter rapeseed might cause developmental differences among ecotypes. At the same time, a TEMSV was identified in the pericentromeric region of A09. This SV is located downstream of the BnaA09.CPK gene of semi-winter rapeseed, which encodes peroxidase and calcium-dependent protein kinase (CPK) and may have the effect of improving stress resistance (Supplementary Fig. 7D). Collectively, these findings highlight the significant role of SVs in shaping the genetic architecture of B. napus , driving ecotype differentiation, modulating gene expression, and enhancing key agronomic traits. Effect of TEMSVs on Gene Expression SVs can affect the expression of nearby genes by altering coding sequences or disrupting regulatory elements. In particular, TEMSVs play a significant role in regulating gene expression [ 40 , 41 ]. Using the graph-based pan-genome of B. napus , we identified 124,682 high-confidence TEMSVs, with more than 80% sequence length overlap between SVs and transposable elements. 49.46% TEMSVs were localized in the distant genomic regions (beyond 3 kb upstream of promoter, Fig. 4 A). The distribution of various types of TEMSVs was consistent in the subgenomes (Supplementary Fig. 8). We defined a “TEMSV-gene” as any gene containing TEMSVs within its coding region or within 3 kb upstream of its transcription start site. Notably, the proportion of TEMSV-genes in the An subgenome (61.48%) was substantially higher than in the Cn subgenome (43.78%) (Fig. 4 B). Moreover, the average number of TEMSVs per gene in the An subgenome was significantly greater than that in the Cn subgenome (Fig. 4 C). We compared the expression levels of each TEMSV-gene in the 505 B. napus accessions with and without TEMSV. The differences in gene expression between these two groups were considerably greater than what would be expected by random sampling (Fig. 4 D). A 592 bp TE-mediated insertion was identified 247 bp upstream of ULTRACURVATA 1 ( BnaC03.UCU1 ). In the surveyed B. napus lines, 81.35% carried this insertion (N = 410) (Fig. 4 E). Gene expression analysis at 20 days after flowering (DAF) of population revealed that the insertion significantly reduced the expression of BnaC03.UCU1 ( P exp = 1.2 \(\:\times\:\) 10 −5 ; Fig. 4 F). In addition, SV-based GWAS analysis identified a significant association between the 592 bp insertion and seed oleic acid (C18:1) content (Supplementary Fig. 9A). Varieties with the insertion exhibited significantly lower C18:1 content than those without the insertion ( P C18:1 = 1.0 \(\:\times\:\) 10 −6 ; Fig. 4 G). To further explore the effect of cis -acting variants in this gene on C18:1, eGWAS of BnaC03.UCU1 was performed in 20 DAF, and the results showed co-localization of the eQTL for this gene with the QTL for C18:1 content (Supplementary Fig. 9B-C). These findings suggest that the insertion in promoter of BnaC03.UCU1 may alters the seed C18:1 content by modulating gene expression. TEMSV effects on the expression of homoeologous gene pairs between subgenomes To investigate the effects of TEMSVs on inter-subgenomic gene expression, we examined its impact on the expression of homoeologous gene pairs (HGPs). An HGP was defined as a pair of homoeologous copies with highest sequence similarity between subgenomes, with one copy from An and one from Cn. A total of 19,856 HGPs were identified (Supplementary Table 11), among which 15,902 (80.09%) contained TEMSVs. Notably, in HGPs harboring TEMSVs, the distribution was unbalanced: 9,159 (46.13%) HGPs exhibited a higher number of TEMSVs in the An copy, nearly twice as many as in the Cn copy (Supplementary Fig. 10). Based on the ZS11 whole-period transcriptome data and seed transcriptome data of the population [ 28 , 33 ], we constructed a comprehensive gene expression profile of HGPs. Overall, homoeologous genes in HGPs exhibited significantly higher expression correlation than randomly paired genes (RGPs) (Fig. 5 A, Supplementary Fig. 11). Additionally, we categorized HGPs into three types based on the position of TEMSVs: (i) TEMSVs in the coding region, (ii) TEMSVs in the non-coding region, and (iii) no TEMSVs. Analysis of the expression correlations revealed that HGPs without TEMSVs had a higher proportion of highly correlated expression than the other two categories (Fig. 5 B-C), suggesting that the presence of TEMSVs partially disrupts the expression correlation of homoeologous genes. To further explore inter-subgenomic relationships, we performed differential expression analysis of HGPs across various tissues of ZS11. Among the identified HGPs, 51.25% of HGPs exhibit unbalanced expression (uHGP, defined as | log 2 (fold change)| > 1 and P adj < 0.05) in at least one tissue (Supplementary Fig. 12). We compared expression fold changes between HGPs with TEMSVs and non-TEMSV HGPs (i.e., genes whose regions do not overlap with any TEMSV) across the seven tissues (silique, seed, silique wall, bud, flower, leaf and stem) over multiple time points, we found that the proportion HGPs with TEMSVs exhibiting over a one-fold change in expression was consistently higher than that of non-TEMSV HGPs (Fig. 5 D; Supplementary Table 12). Furthermore, we examined how the proportion of uHGPs varied with the number of TEMSVs present in HGPs. Our analysis revealed that as the number of TEMSVs increased, the proportion of uHGPs also increased (Fig. 5 E), a trend consistently observed across all tissues (Supplementary Fig. 13). These findings suggest that TEMSVs contribute to unbalanced expression of HGPs among subgenomes, with the degree of imbalance intensifying as more TEMSVs accumulate. Silique wall samples contained the highest number of TEMSV-associated uHGPs (478). We defined HGPs with unbalanced expression in all examined tissues as “shared uHGPs”. Notably, the number of shared uHGPs decreased as the criterion was extended to include more tissues. However, this decline was significantly slower for TEMSV-HGPs than for randomly sampled RGPs (Fig. 5 F). In addition, the expression fold changes of shared TEMSV-HGPs were found to be more stable across different tissues (Fig. 5 G). For example, the gene CDC5 , which regulate the cell division cycle, contains a TEMSV in the An copy (Supplementary Fig. 14). Analysis of the expression profile of this HGP revealed that the An copy consistently exhibited 3- to 4-fold higher expression than the Cn copy across all tissues (Fig. 5 H). These results suggest that uHGPs containing TEMSVs maintain a stable pattern of unbalanced expression across tissues. Identification of HGPs with dosage compensation effects Among 19,856 HGPs, 43,762 TEMSVs were identified. Of these TEMSVs, 1,405 were associated with significant changes in the expression of nearby genes at both 20 and 40 DAF in seed (Student's t -test; P < 0.05). In addition, among the 292 HGPs, when the expression level of homoeologous copy near TEMSV changes, the expression level of distant homoeologous copy also shows significant changes (Student's t -test; P < 0.05). Notably, the magnitude of expression change was significantly greater in the homoeologous copy proximal to TEMSVs compared with its distal counterpart (Fig. 6 A, Supplementary Fig. 15A). Further analysis revealed that in 32.30% of HGPs with TEMSVs, the two homoeologous copies exhibited opposite expression changes (Fig. 6 B, Supplementary Fig. 15B). For example, in the HGP of NUCLEOSIDE DIPHOSPHATE KINASE 3 ( NDPK3 ), the absence of a TE insertion corresponded with significantly higher expression of BnaA02.NDPK3 relative to BnaC02.NDPK3 . However, following the insertion of an 8,141 bp sequence at the 3,209 bp upstream the TSS of BnaA02.NDPK3 (Fig. 6 C), the expression level of BnaA02.NDPK3 significantly decreased, whereas that of BnaC02.NDPK3 markedly increased. Consequently, the overall expression level of the NDPK3 homoeologous pair was reduced after TE insertion (Fig. 6 D). These results suggest that some HGPs are subject to feedback regulation. When TEMSVs cause changes in one gene copy, the other copy adjusts its expression, thereby exhibiting a dosage compensation effect. This compensatory mechanism may help stabilize overall gene expression in the face of perturbations caused by TEMSVs. A 104-bp TE-mediated insertion influenced seed oil content in B. napus TEMSV can affect not only gene expression but also important agronomic traits. In B. napus , SOC is one of the important traits. Using the SV map constructed by pan-genome, we performed an SV-based genome-wide association study (SV-GWAS) to analyze SOC. We identified a 104-bp TE-mediated insertion (TEMI) located in the promoter of the peptidase C78 ( BnaA06 . PC78 ) on chromosome A06 that is significantly associated with SOC (Fig. 6 E-F). Among 505 accessions examined, 63.98% carried the TEMI allele. Importantly, accessions harboring the TEMI exhibited significantly lower SOC than those lacking the insertion ( P SOC = 1.41 \(\:\times\:\) 10 −5 ; Fig. 6 G-H). Expression analysis further revealed that at 40 DAF in seed, the TEMI allele significantly decreased the expression of BnaA06.PC78 , whereas its homolog in the Cn subgenome ( BnaC07.PC78 ) was significantly increased. At 20 DAF, although the decrease in BnaA06.PC78 expression was not statistically significant, a downward trend was observed alongside a significant upregulation of BnaC07.PC78 (Fig. 6 I-J). These findings suggest that the HGP of Bna.PC78 is subject to feedback regulation, wherein the perturbation of one copy by TEMI leads to compensatory expression changes in its homolog, thereby influencing SOC. Discussion The genomes of ZS11 and Westar have become indispensable resources in B. napus research. ZS11 is frequently used as a reference genome for diverse analyses [ 3 ], whereas Westar’s high transformation efficiency makes it a preferred recipient for gene editing experiments [ 42 ]. Our comprehensive mapping of nearly complete genomes using ONT represents a significant advance. With only 38 gaps in ZS11 and 94 in Westar, and notably 15 gap-free chromosomes in ZS11 and one in Westar (Fig. 1 A), these assemblies markedly improve upon previous versions. The number of genes annotated in the centromere region of ZS11 v1 increased by 78.89% compared to the earlier ZS11 v0 assembly (Supplementary Fig. 2). Future efforts should integrate Hi-C, PacBio HiFi, and sequence-specific gap-filling approaches to resolve these challenging regions. Moreover, by extending our analysis to six additional representative B. napus lines, whose assemblies ranged from 824.77 to 949.56 Mb with consistently high BUSCO scores (98.60-99.09%; Supplementary Table 6), we have broadened the genomic framework to encompass both semi-wintering and spring ecotypes. These high-quality assemblies not only enhance our understanding of B. napus genetic diversity but also provide valuable resources for future breeding programs. SVs have emerged as a significant contributor to genetic diversity and phenotypic variation in crops [ 14 – 16 ]. Pan-genome analyses, which integrate multiple high-quality assemblies, offer an effective platform for accurately identifying these SVs [ 20 ]. In this study, we constructed a comprehensive pan-genome using 22 high-quality B. napus genomes, enabling the identification of 221,062 unique SVs (Fig. 3 A). Analysis of resequencing data from 505 accessions further highlighted the value of SVs for differentiating B. napus ecotypes (Supplementary Fig. 7B-C). This extensive reservoir of genetic variation constitutes a valuable resource for elucidating genotype–phenotype relationships and accelerating molecular breeding in B. napus . Natural hybridization and genome doubling in B. napus have resulted in significant asymmetry between the An and Cn subgenomes, especially in terms of gene expression and epigenetic modifications [ 9 ]. Previous eQTL mapping studies have shown that asymmetric transcriptional regulation underpins expression imbalances between subgenomes [ 29 ]. Here, we extend these findings by showing that TEMSVs are more abundant in the An subgenome, thereby contributing to unbalanced expression of HGPs (Fig. 4 ). Moreover, our results indicate that as the number of TEMSVs increases, the expression imbalance between the two subgenomic copies increasingly pronounced, with these effects consistently observed across multiple tissues (Fig. 5 ). These findings deepen our theoretical understanding of regulatory interactions within polyploid genomes. Current investigations into the effects of TEMSVs on HGPs have primarily focused on transcriptional regulation. As chromatin accessibility is closely linked to gene expression [ 43 ], further research is warranted to explore the epistatic effects of TEMSVs on HGP regulation. Integrating pan-genomics with other omics technologies is a powerful strategy for dissecting trait variation and accelerating crop improvement. By incorporating SVs and other genetic variants into a graph-based pan-genome framework, we improved the resolution of GWAS and identified robust associations between specific SVs and key agronomic traits. Notably, two SVs with significant effects on seed traits were identified (Fig. 4 E-G, Fig. 6 E-J). Future studies should focus on functional characterization of these SVs and their associated genes using CRISPR-based genome editing. Additionally, expanding trait association analyses to encompass a broader range of phenotypes may facilitate the discovery of further functionally relevant SVs. These efforts will underscore the functional importance of structural variants and their potential application in breeding programs. Methods Genome sequencing and de novo assembly In this study, the collection of young leaves of ZS11 and Westar were subjected to Nanopore library construction and sequencing. PacBio and Illumina sequencing was performed for six genomes of Zhongshuang 2 (ZS2), Bugle, 352, 862, Ribenyoucai (RB) and Baihua (BH). Canu software was applied to filter out reads less than 5,000 bp in length from Nanopore and PacBio sequencing results and perform self-correction and trimming [ 44 ]. For Illumina data, removal of reads with consecutive mass values ≤ 20 bases up to 40%; removal of N-containing reads and removing duplication contamination. The genome assembly process was divided into two steps: (1) De-novo assembly of Nanopore or PacBio data using NextDenovo software [ 45 ]. (2) Correction of single-base errors and small indel errors in the assembly results with Illumina data using NextPolish software [ 46 ]. Repetitive sequence annotation Tandem repeat sequences were first annotated using GMATA and Tandem Repeats Finder (TRF) software[ 47 , 48 ]. The non-initial repeat libraries for each genome were then predicted using MITE-hunter and RepeatModeler2 (using default parameters), respectively [ 49 , 50 ]. The obtained libraries were then compared with the TE class Repbase ( http://www.girinst.org/repbase ) to categorize the type of each repeat family. To further characterize repetitive sequences across the genome, RepeatMasker was applied to match the sequences against new repetitive sequence libraries and Repbase TE libraries in order to search for known and novel transposable elements [ 51 ]. Overlapping transposable elements belonging to the same repeat class were organized and merged. Gene prediction Three independent methods were used for gene prediction, including ab initio prediction, homology search, and reference-guided transcriptome assembly. Firstly, the homoeologous peptides of Arabidopsis thaliana , B. oleracea , B. rapa and B.napus Darmor were compared with each genome using GeMoMa [ 52 ], and then the gene structure information was obtained. Next gene prediction based on RNA-seq was performed using STAR (default) to compare the filtered RNA-seq reads to the reference genome [ 53 ]. Further transcripts are assembled using StringTie and open reading frames (ORFs) are predicted using PASA to generate a training set, and non-initial gene prediction is performed on the training set using AUGUSTUS with default parameters[ 54 – 56 ]. Finally, different sets of predicted genes obtained from head prediction, homology annotation and transcriptome prediction were integrated by EVidenceModeler (EVM)[ 55 ], besides that, genes with TEs were removed using TransposonPSI package ( http://transposonpsi.sourceforge.net/ ) and filtered the heterozygous genes. Based on the RNA-seq assembly results, untranslated regions (UTRs) and alternative splicing regions were identified using PASA, the longest transcripts of each locus were retained, and regions other than ORFs were designated as UTRs. Functional annotation of gene models Gene functional information as well as protein motifs and structural domains were annotated by comparison with public databases (SwissProt, NR, KEGG, KOG and Gene Ontology). Putative structural domains and GO Term of genes were determined using InterProScan software and default parameters. And then the EvidenceModeler integrated protein sequences were compared with four public protein databases using BLASTP with a threshold of 1e-05 for E-value [ 55 , 57 ], and the results with the lowest E-value were retained. The results from the five databases were finally integrated. Annotation of non-coding RNA (ncRNA) Two strategies were used for the study: searching against databases and utilizing model predictions to obtain ncRNAs (non-coding RNAs). Transfer RNA (tRNA) was predicted using tRNAscan-SE and eukaryotic parameters [ 58 ]. Rfam database was searched using Infernal cmscan for MicroRNA, rRNA, small nuclear RNA and small nucleolar RNA[ 59 ]. rRNA and its subunits were predicted using RNAmmer [ 60 ]. Mitosis and telomere annotation Mitotic-specific repetitive sequences (CentBr1, CentBr2, CRB, PCRBr, TR238, and TR805) were aligned to the reference genome using blastn software [ 61 , 62 ], and mitotic regions were identified by counting the density of associated repeat sequences and gene density of each chromosome region in a 100 Kb window with a sliding window of 100 Kb step size. The annotation of telomeres was done by searching for telomere-specific repeats 5'- (TTTAGGG)n -3' [ 63 , 64 ]. Phylogenetic analysis In order to construct phylogenetic relationships among 20 varieties of Brassica napus , we did an all-to-all blastp with peptide sequences of protein-coding genes annotated from these 20 genomes by diamond [ 65 ] with cut-off evalues of 1e-5 and then input the results into OrthoFinder. The single-copy orthologous genes were further extracted from OrthoFinder results, protein sequences were aligned by muscle v3.8.31 [ 66 ] and conserved sites from multiple sequence alignment were extracted by Gblocks v0.91b [ 67 ]. The phylogenetic tree was constructed by RAxML v8.2.12 [ 68 ] to automatically find the best model and perform 1000 ultrafast bootstrap analyzes to test the robustness of each branch. The online tool iTOL ( http://itol.embl.de ) was used to visualize the constructed tree. Construction of the protein-coding gene-based pan-genome We did an all-to-all alignment with 1,911,393 protein sequences of the 20 B. napus varieties and input the result file into OrthoFinder for gene family construction. Finally, we obtained 101,903 non-redundant gene clusters and 48,233 clusters with only one gene. We then classified non-redundant gene clusters into 4 categories: core gene clusters that were conserved in all 20 varieties; soft-core gene clusters, which were present in 18–19 varieties; dispensable gene clusters, which were found in 2–17 genomes; and private gene clusters, which contained genes from only 1 sample (not including unassigned genes). The longest transcript was chosen to represent each gene. Ka/Ks calculation of different types of pan-genes Non-synonymous substitution rates (Ka), synonymous substitution rates (Ks), and Ka/Ks in core, softcore, and dispensable gene clusters were computed using the WGDI [ 69 ] with default parameters. We conducted amino acid alignment for gene pairs in each cluster first and then converted the results into the coding sequence (CDS) alignment using PAL2NAL [ 70 ]. The alignments were further passed to the PAML [ 71 ] to obtain Ka/Ks values. Construction of the graph-based pan-genome and SVs calling First, we used Minigraph-Cactus to construct the graph pangenome of the 22 high-quality B. napus genome assemblies based on sequence alignment. The ZS11 genome assembled in this study was set as the reference and the other 21 genomes were added into the multi-assembly graph successively. The fragments differing from the reference genome are displayed as different paths in the generated graphical fragment assembly (GFA) file. After obtaining the first version of the pan-genome, we used MUMmer4 [ 72 ] to align the other 21 genomes to the ZS11 reference genome, and used SyRI [ 73 ] to obtain the structural variation between the genome and ZS11. Finally, we integrated the structural variants produced by both methods and used vg toolkit v1.34.0 [ 74 ] to construct the final complete pan-genome. To genotype the SVs in the 505 B. napus accessions we mapped the short reads from each individual to the graph-based pan-genome via vg using default parameters. After filtering individuals with a missing rate above 0.5 or minor allele frequency (MAF) above 0.05 using Plink [ 75 ]. Declarations Acknowledgements Computations in this study were conducted on the bioinformatics computing platform of the National Key Laboratory of Crop Genetic Improvement, Huazhong Agricultural University. This work is supported by National Natural Science Foundation of China (32200501), Project X2662024ZKPY001 supported by the Fundamental Research Funds for the Central Universities, Knowledge Innovation Program of Wuhan-Basic Research (2022020801010221), Project supported by the Fundamental Research Funds for the Central Universities (2662022YJ016), Chongqing Technical Innovation and Application Development Special Project (CSTB2024TIAD-KPX0025), Basic Research Project in 2023 of Yazhouwan National Laboratory (GL23YCKY01) and the fellowship of China National Postdoctoral Program for Innovation Talents (BX20240299). Author contributions H.Z., K.L., J.S. and L.G. designed and supervised this study. Z.T., Y.Z., J.Y. and Z.Z. performed the bioinformatics analysis. Z.T., Y.Z. and Z.G. prepared the manuscript. H.Z., K.L., J.S. and Q.Y. revised the manuscript. All the authors read and approved the manuscript. Competing interests The authors declare no competing interests. References Gu J, Guan Z, Jiao Y, Liu K, Hong D (2024) The story of a decade: genomics, functional genomics and molecular breeding in Brassica napus. Plant Commun Chalhoub B, Denoeud F, Liu S, Parkin IA, Tang H, Wang X, Chiquet J, Belcram H, Tong C, Samans B (2014) Early allopolyploid evolution in the post-Neolithic Brassica napus oilseed genome. science 345:950–953 Tan Z, Han X, Dai C, Lu S, He H, Yao X, Chen P, Yang C, Zhao L, Yang QY (2024) Functional genomics of Brassica napus: Progress, challenges, and perspectives. J Integr Plant Biol 66:484–509 Lu K, Wei L, Li X, Wang Y, Wu J, Liu M, Zhang C, Chen Z, Xiao Z, Jian H (2019) Whole-genome resequencing reveals Brassica napus origin and genetic loci involved in its improvement. Nat Commun 10:1154 Qian L, Qian W, Snowdon RJ (2014) Sub-genomic selection patterns as a signature of breeding in the allopolyploid Brassica napus genome. BMC Genomics 15:1–17 Rong J, Feltus FA, Waghmare VN, Pierce GJ, Chee PW, Draye X, Saranga Y, Wright RJ, Wilkins TA, May OL (2007) Meta-analysis of polyploid cotton QTL shows unequal contributions of subgenomes to a complex network of genes and gene clusters implicated in lint fiber development. Genetics 176:2577–2588 Lv Z, Li Z, Wang M, Zhao F, Zhang W, Li C, Gong L, Zhang Y, Mason AS, Liu B (2021) Conservation and trans-regulation of histone modification in the A and B subgenomes of polyploid wheat during domestication and ploidy transition. BMC Biol 19:1–16 Wang M, Li Z, Zhang Ye, Zhang Y, Xie Y, Ye L, Zhuang Y, Lin K, Zhao F, Guo J (2021) An atlas of wheat epigenetic regulatory elements reveals subgenome divergence in the regulation of development and stress responses. Plant Cell 33:865–881 Zhang Q, Guan P, Zhao L, Ma M, Xie L, Li Y, Zheng R, Ouyang W, Wang S, Li H (2021) Asymmetric epigenome maps of subgenomes reveal imbalanced transcription and distinct evolutionary trends in Brassica napus. Mol Plant 14:604–619 Zhu H, Han X, Lv J, Zhao L, Xu X, Zhang T, Guo W (2011) Structure, expression differentiation and evolution of duplicated fiber developmental genes in Gossypium barbadense and G. hirsutum. BMC Plant Biol 11:1–15 Zhuang W, Chen H, Yang M, Wang J, Pandey MK, Zhang C, Chang W-C, Zhang L, Zhang X, Tang R (2019) The genome of cultivated peanut provides insight into legume karyotypes, polyploid evolution and crop domestication. Nat Genet 51:865–876 Cheng F, Wu J, Cai X, Liang J, Freeling M, Wang X (2018) Gene retention, fractionation and subgenome differences in polyploid plants. Nat plants 4:258–268 Zhang K, Wang X, Cheng F (2019) Plant polyploidy: origin, evolution, and its influence on crop domestication. Hortic Plant J 5:231–239 Alonge M, Wang X, Benoit M, Soyk S, Pereira L, Zhang L, Suresh H, Ramakrishnan S, Maumus F, Ciren D (2020) Major impacts of widespread structural variation on gene expression and crop improvement in tomato. Cell 182:145–161 e123 Zhou Y, Zhang Z, Bao Z, Li H, Lyu Y, Zan Y, Wu Y, Cheng L, Fang Y, Wu K (2022) Graph pangenome captures missing heritability and empowers tomato breeding. Nature 606:527–534 Guo N, Wang S, Wang T, Duan M, Zong M, Miao L, Han S, Wang G, Liu X, Zhang D (2024) A graph-based pan-genome of Brassica oleracea provides new insights into its domestication and morphotype diversification. Plant Commun 5 Zhang Y, Yang Z, He Y, Liu D, Liu Y, Liang C, Xie M, Jia Y, Ke Q, Zhou Y Structural variation reshapes population gene expression and trait variation in 2,105 Brassica napus accessions. Nat Genet 2024:1–13 Niu X-M, Xu Y-C, Li Z-W, Bian Y-T, Hou X-H, Chen J-F, Zou Y-P, Jiang J, Wu Q, Ge S (2019) Transposable elements drive rapid phenotypic variation in Capsella rubella. Proceedings of the National Academy of Sciences 116:6908–6913 Domínguez M, Dugas E, Benchouaia M, Leduque B, Jiménez-Gómez JM, Colot V, Quadrana L (2020) The impact of transposable elements on tomato diversity. Nat Commun 11:4058 Song J-M, Guan Z, Hu J, Guo C, Yang Z, Wang S, Liu D, Wang B, Lu S, Zhou R (2020) Eight high-quality genomes reveal pan-genome architecture and ecotype differentiation of Brassica napus. Nat plants 6:34–45 Bayer PE, Golicz AA, Scheben A, Batley J, Edwards D (2020) Plant pan-genomes are the new reference. Nat plants 6:914–920 Schreiber M, Jayakodi M, Stein N, Mascher M Plant pangenomes for crop improvement, biodiversity and evolution. Nat Rev Genet 2024:1–15 Liu Y, Du H, Li P, Shen Y, Peng H, Liu S, Zhou G-A, Zhang H, Liu Z, Shi M (2020) Pan-genome of wild and cultivated soybeans. Cell 182:162–176 e113 Shang L, Li X, He H, Yuan Q, Song Y, Wei Z, Lin H, Hu M, Zhao F, Zhang C (2022) A super pan-genomic landscape of rice. Cell Res 32:878–896 Gui S, Wei W, Jiang C, Luo J, Chen L, Wu S, Li W, Wang Y, Li S, Yang N (2022) A pan-Zea genome map for enhancing maize improvement. Genome Biol 23:178 Wang M, Li J, Qi Z, Long Y, Pei L, Huang X, Grover CE, Du X, Xia C, Wang P (2022) Genomic innovation and regulatory rewiring during evolution of the cotton genus Gossypium. Nat Genet 54:1959–1971 Huang Y, He J, Xu Y, Zheng W, Wang S, Chen P, Zeng B, Yang S, Jiang X, Liu Z (2023) Pangenome analysis provides insight into the evolution of the orange subfamily and a key gene for citric acid accumulation in citrus fruits. Nat Genet 55:1964–1975 Liu D, Yu L, Wei L, Yu P, Wang J, Zhao H, Zhang Y, Zhang S, Yang Z, Chen G (2021) BnTIR: an online transcriptome platform for exploring RNA-seq libraries for oil crop Brassica napus. Plant Biotechnol J 19:1895 Tan Z, Peng Y, Xiong Y, Xiong F, Zhang Y, Guo N, Tu Z, Zong Z, Wu X, Ye J (2022) Comprehensive transcriptional variability analysis reveals gene networks regulating seed oil content of Brassica napus. Genome Biol 23:233 Li L, Tian Z, Chen J, Tan Z, Zhang Y, Zhao H, Wu X, Yao X, Wen W, Chen W (2023) Characterization of novel loci controlling seed oil content in Brassica napus by marker metabolite-based multi-omics analysis. Genome Biol 24:141 Rousseau-Gueutin M, Belser C, Da Silva C, Richard G, Istace B, Cruaud C, Falentin C, Boideau F, Boutte J, Delourme R (2020) Long-read assembly of the Brassica napus reference genome Darmor-bzh. GigaScience 9:giaa137 Tan Z, Xie Z, Dai L, Zhang Y, Zhao H, Tang S, Wan L, Yao X, Guo L, Hong D (2022) Genome-and transcriptome‐wide association studies reveal the genetic basis and the breeding history of seed glucosinolate content in Brassica napus. Plant Biotechnol J 20:211–225 Tang S, Zhao H, Lu S, Yu L, Zhang G, Zhang Y, Yang Q-Y, Zhou Y, Wang X, Ma W (2021) Genome-and transcriptome-wide association studies provide insights into the genetic basis of natural variation of seed oil content in Brassica napus. Mol Plant 14:470–487 Wang P, Song Y, Xie Z, Wan M, Xia R, Jiao Y, Zhang H, Yang G, Fan Z, Yang QY, Hong D (2023) Xiaoyun, a model accession for functional genomics research in Brassica napus . Plant Commun 5:100727 Yang Z, Wang S, Wei L, Huang Y, Liu D, Jia Y, Luo C, Lin Y, Liang C, Hu Y (2023) BnIR: A multi-omics database with various tools for Brassica napus research and breeding. Mol Plant 16:775–789 Lee H, Chawla HS, Obermeier C, Dreyer F, Abbadi A, Snowdon R (2020) Chromosome-scale assembly of winter oilseed rape Brassica napus. Front Plant Sci 11:496 Davis JT, Li R, Kim S, Michelmore R, Kim S, Maloof JN (2023) Whole-genome sequence of synthetically derived Brassica napus inbred cultivar Da-Ae. G3: Genes, Genomes, Genetics 13:jkad026 Qu C, Zhu M, Hu R, Niu Y, Chen S, Zhao H, Li C, Wang Z, Yin N, Sun F (2023) Comparative genomic analyses reveal the genetic basis of the yellow-seed trait in Brassica napus. Nat Commun 14:5194 Hickey G, Monlong J, Ebler J, Novak AM, Eizenga JM, Gao Y, Human Pangenome Reference C, Marschall T, Li H, Paten B (2024) Pangenome graph construction from genome alignments with Minigraph-Cactus. Nat Biotechnol 42:663–673 Qin P, Lu H, Du H, Wang H, Chen W, Chen Z, He Q, Ou S, Zhang H, Li X (2021) Pan-genome analysis of 33 genetically diverse rice accessions reveals hidden genomic variations. Cell 184:3542–3558 e3516 Gill RA, Scossa F, King GJ, Golicz AA, Tong C, Snowdon RJ, Fernie AR, Liu S (2021) On the role of transposable elements in the regulation of gene expression and subgenomic interactions in crop genomes. CRC Crit Rev Plant Sci 40:157–189 Dai C, Li Y, Li L, Du Z, Lin S, Tian X, Li S, Yang B, Yao W, Wang J (2020) An efficient Agrobacterium-mediated transformation method using hypocotyl as explants for Brassica napus. Mol Breeding 40:1–13 Voss TC, Hager GL (2014) Dynamic regulation of transcriptional states by chromatin and transcription factors. Nat Rev Genet 15:69–81 Koren S, Walenz BP, Berlin K, Miller JR, Bergman NH, Phillippy AM (2017) Canu: scalable and accurate long-read assembly via adaptive k-mer weighting and repeat separation. Genome Res 27:722–736 Hu J, Wang Z, Sun Z, Hu B, Ayoola AO, Liang F, Li J, Sandoval JR, Cooper DN, Ye K (2024) NextDenovo: an efficient error correction and accurate assembly tool for noisy long reads. Genome Biol 25:107 Hu J, Fan J, Sun Z, Liu S (2020) NextPolish: a fast and efficient genome polishing tool for long-read assembly. Bioinformatics 36:2253–2255 Wang X, Wang L (2016) GMATA: an integrated software package for genome-scale SSR mining, marker development and viewing. Front Plant Sci 7:215951 Benson G (1999) Tandem repeats finder: a program to analyze DNA sequences. Nucleic Acids Res 27:573–580 Han Y, Wessler SR (2010) MITE-Hunter: a program for discovering miniature inverted-repeat transposable elements from genomic sequences. Nucleic Acids Res 38:e199–e199 Flynn JM, Hubley R, Goubert C, Rosen J, Clark AG, Feschotte C, Smit AF (2020) RepeatModeler2 for automated genomic discovery of transposable element families. Proceedings of the National Academy of Sciences 117:9451–9457 Chen N (2004) Using Repeat Masker to identify repetitive elements in genomic sequences. Current protocols in bioinformatics 5:4.10. 11-14.10. 14 Keilwagen J, Hartung F, Grau J GeMoMa: homology-based gene prediction utilizing intron position conservation and RNA-seq data. Gene prediction: Methods protocols 2019:161–177 Dobin A, Davis CA, Schlesinger F, Drenkow J, Zaleski C, Jha S, Batut P, Chaisson M, Gingeras TR (2013) STAR: ultrafast universal RNA-seq aligner. Bioinformatics 29:15–21 Pertea M, Pertea GM, Antonescu CM, Chang T-C, Mendell JT, Salzberg SL (2015) StringTie enables improved reconstruction of a transcriptome from RNA-seq reads. Nat Biotechnol 33:290–295 Haas BJ, Salzberg SL, Zhu W, Pertea M, Allen JE, Orvis J, White O, Buell CR, Wortman JR (2008) Automated eukaryotic gene structure annotation using EVidenceModeler and the Program to Assemble Spliced Alignments. Genome Biol 9:1–22 Stanke M, Diekhans M, Baertsch R, Haussler D (2008) Using native and syntenically mapped cDNA alignments to improve de novo gene finding. Bioinformatics 24:637–644 Kent WJ (2002) BLAT—the BLAST-like alignment tool. Genome Res 12:656–664 Chan PP, Lin BY, Mak AJ, Lowe TM (2021) tRNAscan-SE 2.0: improved detection and functional classification of transfer RNA genes. Nucleic Acids Res 49:9077–9096 Nawrocki EP, Eddy SR (2013) Infernal 1.1: 100-fold faster RNA homology searches. Bioinformatics 29:2933–2935 Lagesen K, Hallin P, Rødland EA, Stærfeldt H-H, Rognes T, Ussery DW (2007) RNAmmer: consistent and rapid annotation of ribosomal RNA genes. Nucleic Acids Res 35:3100–3108 Lim K-B, De Jong H, Yang T-J, Park J-Y, Kwon S-J, Kim JS, Lim M-H, Kim JA, Jin M, Jin Y-M (2005) Characterization of rDNAs and tandem repeats in the heterochromatin of Brassica rapa. Mol Cells 19:436–444 Lim KB, Yang TJ, Hwang YJ, Kim JS, Park JY, Kwon SJ, Kim J, Choi BS, Lim MH, Jin M (2007) Characterization of the centromere and peri-centromere retrotransposons in Brassica rapa and their distribution in related Brassica species. Plant J 49:173–183 Cox AV, Bennett ST, Parokonny AS, Kenton A, Callimassia MA, Bennett MD (1993) Comparison of plant telomere locations using a PCR-generated synthetic probe. Ann Botany 72:239–247 Koo D-H, Hong CP, Batley J, Chung YS, Edwards D, Bang J-W, Hur Y, Lim YP (2011) Rapid divergence of repetitive DNAs in Brassica relatives. Genomics 97:173–185 Buchfink B, Reuter K, Drost H-G (2021) Sensitive protein alignments at tree-of-life scale using DIAMOND. Nat Methods 18:366–368 Edgar RC (2010) Search and clustering orders of magnitude faster than BLAST. Bioinformatics 26:2460–2461 Talavera G, Castresana J (2007) Improvement of phylogenies after removing divergent and ambiguously aligned blocks from protein sequence alignments. Syst Biol 56:564–577 Stamatakis A (2014) RAxML version 8: a tool for phylogenetic analysis and post-analysis of large phylogenies. Bioinformatics 30:1312–1313 Sun P, Jiao B, Yang Y, Shan L, Li T, Li X, Xi Z, Wang X, Liu J (2022) WGDI: A user-friendly toolkit for evolutionary analyses of whole-genome duplications and ancestral karyotypes. Mol Plant 15:1841–1851 Suyama M, Torrents D, Bork P (2006) PAL2NAL: robust conversion of protein sequence alignments into the corresponding codon alignments. Nucleic Acids Res 34:W609–W612 Yang Z (2007) PAML 4: phylogenetic analysis by maximum likelihood. Mol Biol Evol 24:1586–1591 Delcher AL, Phillippy A, Carlton J, Salzberg SL (2002) Fast algorithms for large-scale genome alignment and comparison. Nucleic Acids Res 30:2478–2483 Goel M, Sun H, Jiao WB, Schneeberger K (2019) SyRI: finding genomic rearrangements and local sequence differences from whole-genome assemblies. Genome Biol 20:277 Garrison E, Siren J, Novak AM, Hickey G, Eizenga JM, Dawson ET, Jones W, Garg S, Markello C, Lin MF, Paten B, Durbin R (2018) Variation graph toolkit improves read mapping by representing genetic variation in the reference. Nat Biotechnol 36:875–879 Purcell S, Neale B, Todd-Brown K, Thomas L, Ferreira MA, Bender D, Maller J, Sklar P, de Bakker PI, Daly MJ, Sham PC (2007) PLINK: a tool set for whole-genome association and population-based linkage analyses. Am J Hum Genet 81:559–575 Additional Declarations There is NO Competing Interest. Supplementary Files 20250414SupTable.xlsx Supplementary Table 1-12 20250410supfigure.pdf Supplementary Figure 1-15 SupplementaryFiguresandTableslegends.docx Cite Share Download PDF Status: Under Review Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6452497","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":448944125,"identity":"1131fe1d-a093-4f35-ad0a-acea1e4a3670","order_by":0,"name":"Hu Zhao","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA70lEQVRIiWNgGAWjYDACCQYGZgYGGyDJ3MAMFTMgRksakGQkTcthIItYLfKzmx8+Lmw7H83fDtRS8McusYG9eZsEQ80dnFoY5xwzNp7Zdjt3xmGglhk8yYkNPMfKJBiOPcOphVkiwUyaF6ilAaSFR4I5sUEix0yCseEwTi1sEunfgFrO5c4HazGoT2yQf4NfCw/QTKCWA7kbwFoSDgNt4cGvRUIip9iY51xy7kaglsM8B44bt/GkFVskHMOtRX5G+sbHPGV2ufPOHz74mOdPtWw/++GNNz7U4NaCAg6AfQciEojTMApGwSgYBaMABwAA0bVOQve6mqkAAAAASUVORK5CYII=","orcid":"https://orcid.org/0000-0001-5046-6632","institution":"Huazhong Agricultural University","correspondingAuthor":true,"prefix":"","firstName":"Hu","middleName":"","lastName":"Zhao","suffix":""},{"id":448944126,"identity":"01f861aa-4019-4c7f-a4fa-fe1987467e21","order_by":1,"name":"Zengdong Tan","email":"","orcid":"","institution":"Yazhouwan National Laboratory","correspondingAuthor":false,"prefix":"","firstName":"Zengdong","middleName":"","lastName":"Tan","suffix":""},{"id":448944127,"identity":"f234b9d8-b731-4484-aeb4-f21eed13dc2e","order_by":2,"name":"Yuyu Zheng","email":"","orcid":"","institution":"Huazhong Agricultural University","correspondingAuthor":false,"prefix":"","firstName":"Yuyu","middleName":"","lastName":"Zheng","suffix":""},{"id":448944128,"identity":"7b76e50a-be32-4b77-a65d-38dc132c5af7","order_by":3,"name":"Zhilin Guan","email":"","orcid":"","institution":"Huazhong Agricultural University","correspondingAuthor":false,"prefix":"","firstName":"Zhilin","middleName":"","lastName":"Guan","suffix":""},{"id":448944129,"identity":"7749b07d-a743-4e26-90e9-bf18a26be4a3","order_by":4,"name":"Xuqing Wang","email":"","orcid":"","institution":"Huazhong Agricultural University","correspondingAuthor":false,"prefix":"","firstName":"Xuqing","middleName":"","lastName":"Wang","suffix":""},{"id":448944130,"identity":"e6d1a17a-fb3d-4ffb-8a5c-e7923da29284","order_by":5,"name":"Jianshun Yang","email":"","orcid":"","institution":"Huazhong Agricultural University","correspondingAuthor":false,"prefix":"","firstName":"Jianshun","middleName":"","lastName":"Yang","suffix":""},{"id":448944131,"identity":"525dac5e-b912-41ca-a5db-82f31ff87ea1","order_by":6,"name":"Zhanxiang Zong","email":"","orcid":"https://orcid.org/0000-0001-5287-420X","institution":"National Key Laboratory of Crop Genetic Improvement","correspondingAuthor":false,"prefix":"","firstName":"Zhanxiang","middleName":"","lastName":"Zong","suffix":""},{"id":448944132,"identity":"43af14ec-317e-4c4e-bd5e-56740a1d3d8a","order_by":7,"name":"Qing-Yong Yang","email":"","orcid":"https://orcid.org/0000-0002-3510-8906","institution":"Huazhong Agricultural University","correspondingAuthor":false,"prefix":"","firstName":"Qing-Yong","middleName":"","lastName":"Yang","suffix":""},{"id":448944133,"identity":"a8c12fdf-ce83-436a-ae0e-5d0c9177f4d9","order_by":8,"name":"Liang Guo","email":"","orcid":"https://orcid.org/0000-0001-7191-5062","institution":"Huazhong Agricultural University","correspondingAuthor":false,"prefix":"","firstName":"Liang","middleName":"","lastName":"Guo","suffix":""},{"id":448944134,"identity":"d767c673-3f3a-4ed0-8559-f040466d3924","order_by":9,"name":"Jiaming Song","email":"","orcid":"https://orcid.org/0000-0002-6636-4152","institution":"Southwest University","correspondingAuthor":false,"prefix":"","firstName":"Jiaming","middleName":"","lastName":"Song","suffix":""},{"id":448944135,"identity":"427bad75-7092-4e4e-9cb5-60f6ffbf9af0","order_by":10,"name":"Kede Liu","email":"","orcid":"https://orcid.org/0000-0002-1395-2049","institution":"Huazhong Agricultural University","correspondingAuthor":false,"prefix":"","firstName":"Kede","middleName":"","lastName":"Liu","suffix":""}],"badges":[],"createdAt":"2025-04-15 08:20:17","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6452497/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6452497/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":81613546,"identity":"e9b8995a-ee2d-4188-a9bf-a8d2a26fab5c","added_by":"auto","created_at":"2025-04-29 07:51:59","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":666233,"visible":true,"origin":"","legend":"\u003cp\u003eSee image above for figure legend.\u003c/p\u003e","description":"","filename":"figure1.png","url":"https://assets-eu.researchsquare.com/files/rs-6452497/v1/1b1fefb9c382ffe9b87a3e1a.png"},{"id":81613049,"identity":"c473c7b5-7cd7-41b7-9907-3fdec30d7335","added_by":"auto","created_at":"2025-04-29 07:43:59","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":218270,"visible":true,"origin":"","legend":"\u003cp\u003eSee image above for figure legend.\u003c/p\u003e","description":"","filename":"figure2.png","url":"https://assets-eu.researchsquare.com/files/rs-6452497/v1/e8b8c79d144c3a2db360b20c.png"},{"id":81613048,"identity":"77577a2b-8958-45bf-a44b-84a6e16bd896","added_by":"auto","created_at":"2025-04-29 07:43:59","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":587373,"visible":true,"origin":"","legend":"\u003cp\u003eSee image above for figure legend.\u003c/p\u003e","description":"","filename":"figure3.png","url":"https://assets-eu.researchsquare.com/files/rs-6452497/v1/a7be63f4cfecffeb418d4424.png"},{"id":81613047,"identity":"247d55f8-e692-49eb-bc41-69544ca05942","added_by":"auto","created_at":"2025-04-29 07:43:59","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":72525,"visible":true,"origin":"","legend":"\u003cp\u003eSee image above for figure legend.\u003c/p\u003e","description":"","filename":"figure4.png","url":"https://assets-eu.researchsquare.com/files/rs-6452497/v1/470879bc8b4a5ce10ba5e2ad.png"},{"id":81613052,"identity":"10638d5c-4996-4ea8-ae8e-0af384b7bb46","added_by":"auto","created_at":"2025-04-29 07:43:59","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":296107,"visible":true,"origin":"","legend":"\u003cp\u003eSee image above for figure legend.\u003c/p\u003e","description":"","filename":"figure5.png","url":"https://assets-eu.researchsquare.com/files/rs-6452497/v1/2d72970cdbcfd723e1627010.png"},{"id":81613050,"identity":"26ff7710-586c-4422-b5af-d3521777602f","added_by":"auto","created_at":"2025-04-29 07:43:59","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":189056,"visible":true,"origin":"","legend":"\u003cp\u003eSee image above for figure legend.\u003c/p\u003e","description":"","filename":"figure6.png","url":"https://assets-eu.researchsquare.com/files/rs-6452497/v1/428e9371777f7897da032b50.png"},{"id":81614831,"identity":"3a0005f4-2e01-4d48-a02c-9524a913497b","added_by":"auto","created_at":"2025-04-29 08:00:05","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2666839,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6452497/v1/0f7cb2df-81a4-4710-b596-de5dee81a7ca.pdf"},{"id":81613078,"identity":"d15399a7-bf7a-4461-b1d0-482f27a626b2","added_by":"auto","created_at":"2025-04-29 07:44:02","extension":"xlsx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":1104765,"visible":true,"origin":"","legend":"Supplementary Table 1-12","description":"","filename":"20250414SupTable.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-6452497/v1/45efdbfda75213d853d2e0c1.xlsx"},{"id":81613547,"identity":"6ecb7ba8-4948-498e-9b28-e81e0df2efc2","added_by":"auto","created_at":"2025-04-29 07:51:59","extension":"pdf","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":6191546,"visible":true,"origin":"","legend":"Supplementary Figure 1-15","description":"","filename":"20250410supfigure.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6452497/v1/9df966f4fb38c21fac304c1c.pdf"},{"id":81613044,"identity":"bd52a106-758f-44c6-8cd9-0f77e8aa1414","added_by":"auto","created_at":"2025-04-29 07:43:59","extension":"docx","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":16938,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryFiguresandTableslegends.docx","url":"https://assets-eu.researchsquare.com/files/rs-6452497/v1/124d34299ee0b8c8bea6f383.docx"}],"financialInterests":"There is \u003cb\u003eNO\u003c/b\u003e Competing Interest.","formattedTitle":"Transposable element-mediated structural variations drive gene expression and agronomic trait diversity in Brassica napus","fulltext":[{"header":"Introduction","content":"\u003cp\u003e \u003cem\u003eBrassica napus\u003c/em\u003e (AACC, 2n\u0026thinsp;=\u0026thinsp;38) is a leading oilseed crop cultivated globally for edible vegetable oil, animal feed and vegetables [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. In 2022, the total production of rapeseed was 87.2\u0026nbsp;million tonnes, with an average yield of 2.18 tonnes per hectare (FAO, 2022; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://faostat.fao.org\u003c/span\u003e\u003cspan address=\"http://faostat.fao.org\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). As an allotetraploid species, \u003cem\u003eB. napus\u003c/em\u003e originated approximately 7,500 years ago from an interspecific hybridization between two diploid progenitors \u003cem\u003eB. rapa\u003c/em\u003e (AA, 2n\u0026thinsp;=\u0026thinsp;20) and \u003cem\u003eB. oleracea\u003c/em\u003e (CC, 2n\u0026thinsp;=\u0026thinsp;18) [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e, \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. This unique genomic composition has spurred extensive research into its genetic architecture, particularly the divergent evolutionary trajectories of its subgenomes [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. Previous studies on polyploid subgenomes have predominantly focused on genomic genetic variation [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e], epigenetic modification and gene expression divergence [\u003cspan additionalcitationids=\"CR8\" citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e], selected genes [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e] and trait formation [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e] and evolution [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. Investigations have revealed imbalances in gene expression, regulation, and levels of active epigenetic signal modifications between the subgenomes of \u003cem\u003eB. napus\u003c/em\u003e [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. Additionally, SNP-eQTL mapping has uncovered feedback regulation between homoeologous gene pairs, suggesting mechanisms that maintain partial expression dosage across subgenomes [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eDespite these advances, structural variations (SVs) are increasingly recognized as a major source of genetic diversity, often explaining more phenotypic variation than single nucleotide polymorphisms (SNPs) [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. Graph-based pan-genomics studies, along with cataloging genetic variations, are instrumental in dissecting crop domestication, environmental adaptation, and phenotypic diversification [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e, \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]. For instance, analysis of SV effects on glucosinolate biosynthesis and transport pathways in rapeseed demonstrated the importance of SVs in remodeling key traits [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]. Among the various types of SVs, transposable element (TE) insertions are particularly significant because TEs tend to be highly enriched near genes and can exert profound influences on gene function and phenotype [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e, \u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. A notable case is the 3.9-kb CACTA-like TE insertion in the promoter of \u003cem\u003eBnaA9.CYP78A9\u003c/em\u003e, which significantly correlates with variations in silique length and seed weight [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. However, most research on agronomically important traits in \u003cem\u003eB. napus\u003c/em\u003e has focused on SNP-level variations, leaving the impact of large-scale SVs on trait domestication and improvement largely unexplored.\u003c/p\u003e \u003cp\u003eRecent advances in long-read sequencing technologies [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e, \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e] have transformed plant genome research by enabling a deeper exploration of intraspecific genetic diversity, particularly in economically important crops [\u003cspan additionalcitationids=\"CR24\" citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]. The pan-genome has emerged as a revolutionary approach that captures the full spectrum of genetic diversity, thus facilitating the comprehensive characterization of SVs [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e, \u003cspan additionalcitationids=\"CR26\" citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]. Despite these advances in crops like maize and rice, the number of high-quality reference genomes available for \u003cem\u003eB. napus\u003c/em\u003e remains limited [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. To address this gap, we refined near-complete genomes of two representative \u003cem\u003eB. napus\u003c/em\u003e accessions (ZS11 and Westar) and newly generated chromosome-level assemblies for six additional varieties (ZS2, Bugle, 352, 862, RB, and BH). Combined with published genomes, we constructed a graph-based pan-genome containing 22 high-quality \u003cem\u003eB. napus\u003c/em\u003e genomes. Analysis of seed RNA-seq data from 505 \u003cem\u003eB. napus\u003c/em\u003e accessions revealed that transposable element\u0026ndash;mediated structural variations (TEMSVs) can significantly alter gene expression patterns across subgenomes. Moreover, SV-based genome-wide association studies (SV-GWAS) identified two SVs associated with seed oil content (SOC) and seed fatty acid content (SFAC). Our findings suggest that the expression imbalance of homoeologous gene pairs among subgenomes may be driven by TEMSVs, highlighting the impact of SVs on key traits in \u003cem\u003eB. napus\u003c/em\u003e. This study introduces a graph-based pan-genome of \u003cem\u003eB. napus\u003c/em\u003e as a valuable tool for understanding SVs and advancing rapeseed genetic improvement.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003e \u003cb\u003eTwo near-complete genome assemblies in\u003c/b\u003e \u003cb\u003eB. napus\u003c/b\u003e\u003c/p\u003e \u003cp\u003eTo upgrade the reference genomes of ZS11 and Westar, we sequenced the two genomes with Oxford Nanopore Technology (ONT) and obtained 154.13 GB (~\u0026thinsp;154 \u0026times; coverage) and 81.31 GB (~\u0026thinsp;81 \u0026times; coverage) of ultra-long reads, respectively (Supplementary Table\u0026nbsp;1). Combined with the PacBio data generated in previous study [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e], each genome was independently assembled and error-corrected using three methods, including NextDenovo, Flye and NECAT. Subsequently, the preliminary assemblies (ZS11 v0 and Westar v0) were used as integration backbones, leading to the production of two near-complete genomes, designated ZS11 v1 and Westar v1, which contained 15 and 1 gap-free chromosomes, respectively. The final ZS11 v1 assembly reached 1,014 Mb with only 38 gaps (Contig N50: 58.66 Mb), marking a 38.9-fold improvement in contiguity over ZS11 v0 assembly (999 Mb, 4,470 gaps, Contig N50: 1.51 Mb). Similarly, Westar v1 achieved 984 Mb with 94 gaps (Contig N50: 23.71 Mb), representing a 7.6-fold enhancement compared to Westar v0 (982.41 Mb, 2,587 gaps, Contig N50: 3.13 Mb) (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eA). By searching for telomere-specific tandem repeats, we identified telomeric sequences at the ends of 16 chromosomes in ZS11 v1 and Westar v1 (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eB, Supplementary Fig.\u0026nbsp;1).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eGiven the high abundance of repetitive sequences in centromeric regions and the limited coverage of genetic markers, centromeres are notoriously challenging to assemble. Here, we mapped Brassicaceae centromere-specific repetitive sequences to the ZS11 v1 and Westar v1 genomes (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eB, Supplementary Fig.\u0026nbsp;1). In the upgraded ZS11 and Westar reference genomes, the identified centromeric region spanned 153.28 Mb and 136.71 Mb (Supplementary Table\u0026nbsp;2\u0026ndash;3). Among them, ZS11 v1 with the number of annotated genes (N\u003csub\u003ev1\u003c/sub\u003e = 1,099) within this region increasing by 78.89% compared to the ZS11 v0. Furthermore, the proportion of expressed genes in the centromeres rose to 48.32% in ZS11 v1, compared to 34.34% in the v0 assembly (Supplementary Fig.\u0026nbsp;2). Gene Ontology (GO) enrichment analysis further revealed that the genes within the newly assembled centromeric regions were significantly enriched in pathways related to energy metabolism and photosynthetic electron transport (Supplementary Table\u0026nbsp;4). Overall, these findings indicate that our improved assembly offers a more complete and functionally informative centromeric region, providing a valuable resource for future functional genomic studies.\u003c/p\u003e \u003cp\u003eTo further validate the completeness of the ZS11 v1 and Westar v1 genomes, we aligned 273 RNA-seq datasets from 91 distinct tissues, including root, seedling, stem, leaves, flower buds, flower tissues, silique walls, and seeds [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e], to the new assemblies. The average mapping rates for the ZS11 and Westar v1 genomes were 76.1% and 75.2%, respectively, representing improvements of 2.2% and 1.4% over their respective v0 genome assemblies (Supplementary Fig.\u0026nbsp;3). Additionally, the benchmarking universal single-copy orthologs (BUSCO) analysis revealed that the new assemblies captured 99.75% and 99.69% of the expected BUSCO genes for ZS11 v1 and Westar v1, respectively (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eA). Collectively, these findings demonstrate a substantial improvement in genome quality and completeness.\u003c/p\u003e \u003cp\u003eIn the newly assembled ZS11 and Westar genomes, we annotated 111,715 and 114,412 genes, respectively, using a combination of multi-transcript integration, homologous species protein prediction, and \u003cem\u003ede novo\u003c/em\u003e prediction methods (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eA). The average gene lengths increased to 2,498.39 bp and 2,229.65 bp, as compared to their respective v0 assemblies, which had average gene lengths of 2,061.34 bp and 2,110.54 bp. As eukaryotic genes typically consist of multiple exons, the proportion of multi-exon genes increased from 78.79% and 78.86% in the v0 genomes (ZS11 and Westar) to 82.08% and 84.03% in the upgraded assemblies, indicating enhanced resolution in gene structure annotation. In addition, annotation of untranslated regions (UTRs) showed substantial improvement. Prior to the upgradation, 23,575 and 29,694 genes were annotated with UTRs in ZS11 and Westar, accounting for 30.70% and 31.14% of the total annotated genes, respectively. After the upgradation, these numbers increased to 34,785 and 35,130 genes, with a total of 69,762 and 73,725 UTRs annotated in the ZS11 v1 and Westar v1 genomes (Supplementary Table\u0026nbsp;5). Comparisons of transcript length, coding sequence (CDS) length, intron length, exon number, and exon length between the upgraded gene sets and those of \u003cem\u003eB. oleracea\u003c/em\u003e and \u003cem\u003eB. rapa\u003c/em\u003e revealed similar distributions, thereby supporting the reliability of our improved annotation (Supplementary Fig.\u0026nbsp;4).\u003c/p\u003e \u003cp\u003eThe ZS11 v0 genome has underpinned numerous functional genomics studies in \u003cem\u003eB. napus\u003c/em\u003e [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e, \u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e, \u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e]. However, the gene annotation in the ZS11 v0 genome remains suboptimal and warrants improvement. For example, only two copies of \u003cem\u003eTRANSPARENT TESTA 4\u003c/em\u003e (\u003cem\u003eTT4\u003c/em\u003e), a gene that known to affect seed coat content (SCC) and SOC [\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e], were previously annotated on chromosomes C01 and C02 in the ZS11 v0 genome. In contrast, 13 \u003cem\u003eTT4\u003c/em\u003e copies were identified in the ZS11 v1 genome, which is consistent with the copy number reported in the Darmor-bzh v10 genome [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]. Additionally, an unassigned copy in the Darmor-bzh v10 genome was successfully assigned to chromosome C09 in the ZS11 v1, and gene expression analysis confirmed that 10 of these \u003cem\u003eTT4\u003c/em\u003e copies were actively expressed in ZS11 developing seeds (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eC). Two \u003cem\u003eGLUCOSINOLATE TRANSPORTER-2\u003c/em\u003e (\u003cem\u003eGTR2\u003c/em\u003e) genes on A09 and C02 encoding key transporter proteins were identified to control seed glucosinolates content (SGC) using the Darmor-bzh v10 genome[\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e] and the ZS11 v1 genome. However, these two genes were not annotated in the ZS11 v0 genome (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eD-E). Together, these findings underscore the importance of comprehensive and accurate gene annotation for advancing genomic research and for improving our understanding of key agronomic traits in \u003cem\u003eB. napus\u003c/em\u003e.\u003c/p\u003e \u003cp\u003e \u003cb\u003eChromosome-level assemblies of six diverse accessions of\u003c/b\u003e \u003cb\u003eBrassica napus\u003c/b\u003e\u003c/p\u003e \u003cp\u003eTo comprehensively capture genome variations of \u003cem\u003eB. napus\u003c/em\u003e, six additional representative varieties, namely ZS2, Bugle, 352, 862, RB, and BH, were selected for de novo assembly based on the population structure analysis and the phenotypic characterization from previous study [\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e]. PacBio sequencing yield 94.39 to 146.29 Gb of data for each variety. The assembled genome sizes ranged from 824.77 to 949.56 Mb, with Contig N50 values ranging from 1.44 to 4.62 Mb. A total of 90,960 to 98,777 protein-coding genes were annotated, and 98.60 to 99.09% of the conserved BUSCO genes were identified, demonstrating the high integrity of the assembled genomes (Supplementary Table\u0026nbsp;6).\u003c/p\u003e \u003cp\u003e \u003cb\u003eGene-based pan-genome and core-genome analyses of 20\u003c/b\u003e \u003cb\u003eB. napus\u003c/b\u003e \u003cb\u003eaccessions\u003c/b\u003e\u003c/p\u003e \u003cp\u003eIn addition to the eight genomes assembled in this study, an additional 12 previously published genome assemblies with gene annotations (including Da-Ae, Tapior, Darmor, Express617, Gangan, No2127, P130, P202, Quinta, Shengli, Xiaoyun and Zheyou7) for subsequent analysis [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e, \u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e, \u003cspan additionalcitationids=\"CR35 CR36\" citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e]. By clustering 1,911,393 predicted gene models from these 20 accessions, we constructed a protein-coding gene-based pan-genome with 101,903 non-redundant pan-gene clusters (i.e., the cumulative set of all genes), which is close to the number reported in a previous rapeseed pan-genome analysis [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. These clusters were classified into four categories based on their frequency of occurrence: 17,877 core (present in all 20 accessions) and 13,096 softcore (present in at least 18 accessions), 66,748 dispensable clusters (present in 2 to 17 accessions) and 4,182 private clusters (present in only one accession). The private clusters represent 4.02% of the total gene sets (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA-B). Pan-genome size simulations indicated that the number of gene cluster approached saturation when more than 14 accessions were included, and the addition of the 20th accession contributed only 110 new clusters, demonstrating the representativeness of our selected accessions (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA-B). Notably, semi-winter and spring rapeseed exhibited a greater prevalence of diverse gene families (Supplementary Fig.\u0026nbsp;5A). Dispensable gene families accounted for 65.6% of the total clusters, contributing an average of 34.7% of a single accession\u0026rsquo;s genome, whereas core and softcore genes collectively accounted for 64.6% (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eC-D). Compared with core/softcore genes, dispensable and private genes had fewer exons (Supplementary Fig.\u0026nbsp;5B) but displayed significantly higher non-synonymous and synonymous substitution ratio (Ka/Ks). In particular, 811 dispensable gene families enriched in nucleotide binding and ATP binding had a Ka/Ks ratio greater than 1, suggesting positive selection pressure in functional categories such as nucleotide binding and ATP binding (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eE). Conversely, 53,129 gene families with Ka/Ks ratio less than 1 indicated purifying selection, particularly in functions related to the nucleus and metal ion binding (Supplementary Fig.\u0026nbsp;5C). GO analysis indicated that core genes were enriched in several essential biological processes including macromolecule modification, glycosylation, and phosphorus metabolic process, while dispensable genes were associated with processes like DNA integration, hormone responses, and telomere maintenance (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eF; Supplementary Table\u0026nbsp;7). Finally, we constructed a comprehensive database BrassicaPanDB (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://rgmi.hzau.edu.cn/pan/\u003c/span\u003e\u003cspan address=\"http://rgmi.hzau.edu.cn/pan/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e), to provide researchers with access to this pan-genome. The database includes functions for data download, genome browsing, gene search, and BLAST analysis, enabling users to explore the genetic diversity of these 20 \u003cem\u003eB. napus\u003c/em\u003e accessions in depth.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eExtensive genetic variations and the variant-integrated graph-based pan-genome\u003c/h2\u003e \u003cp\u003eThe constructed rapeseed graph-based pan-genome integrates extensive genetic variations, providing essential insights into sequence divergence. The analysis included the 20 genomes and two additional genomes (GH06 and ZY821) [\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e], enabling accurate identification of single nucleotide polymorphisms (SNPs), small insertions and deletions (InDels) and SVs (InDels and inversions\u0026thinsp;\u0026ge;\u0026thinsp;50 bp). Using Minigraph-Cactus [\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e], the sequences of 21 \u003cem\u003eB. napus\u003c/em\u003e genomes were aligned against the ZS11 v1 genome. A total of 11,120,890 SNPs and 5,583,367 InDels (\u0026le;\u0026thinsp;50 bp), as well as 221,062 non-redundant SVs, including 124,836 insertions (ranging from 50 bp to 839 kb), 95,685 deletions (ranging from 50 bp to 159 kb) and 541 inversions (ranging from 606 bp to 1.22 Mb) were detected (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eA-B; Supplementary Table\u0026nbsp;8; Supplementary Fig.\u0026nbsp;5D-E). Of these SVs, 744 are present in all 21 query genomes, 4,479 present in 19\u0026ndash;20 genomes, 130,975 present in 2\u0026ndash;18 genomes, and 84,864 unique to individual genomes (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eA; Supplementary Table\u0026nbsp;9). To verify the accuracy, a subset of SVs in known flowering-time genes (e.g., \u003cem\u003eFLC\u003c/em\u003e and \u003cem\u003eFT\u003c/em\u003e) was validated using PacBio long reads, confirming high assembly quality (Supplementary Fig.\u0026nbsp;6).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003ePresence-absence variants (PAVs) were the most abundant type of SV (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e; Supplementary Fig.\u0026nbsp;5D). A total of 220,521 non-redundant PAVs were included in subsequent analyses. To examine the distribution and specificity of PAVs among different genomes, we analyzed the patterns of shared and unique SVs across the rapeseed accessions. The number of shared SVs sharply declined for the first three genomes and slowly decreased thereafter (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eB). The number of PAVs per genome ranged from 39,690 to 87,518, with the variety no2127 exhibiting the highest number of private PAVs, probably due to its synthetic origin (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eC; Supplementary Table\u0026nbsp;9\u0026ndash;10). We also found that the number of PAVs gradually declined as the distance of PAVs from the nearest gene increased. Notably, 49.6% SVs were located in genic and flanking regions (within 5 kb of the transcription start site), suggesting these SVs may affect gene expression (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eD). Moreover, PAVs were enriched in intergenic repetitive regions, with 71.9% (N\u0026thinsp;=\u0026thinsp;158,486) of PAVs being TEs, suggesting that TEs may have driven the formation of most PAVs in the rapeseed genome. We defined these types of SV as TE-mediated SVs (TEMSVs). Specifically, most TEMSVs were DNA/Helitron (most abundant in number) and long terminal repeat (LTR) transposons (longest in length) (Supplementary Fig.\u0026nbsp;7A). A set of 4,064 TEMSVs are located in the promoter regions (2 kb upstream of the translation-start sites) or gene bodies, with SVs being more prevalent in genes with low expression levels (Supplementary Fig.\u0026nbsp;7B).\u003c/p\u003e \u003cp\u003eSpring and semi-winter type rapeseed varieties are widely cultivated throughout the world. By mapping reads of 505 \u003cem\u003eB. napus\u003c/em\u003e accessions (comprising winter, semi-winter, and spring type) to this graph-based genome, a total of 431,765 SVs were identified in the population. A neighbor-joining tree constructed from these SVs distinctly grouped the different ecotypes, indicating that SVs are associated with ecotype differentiation (Supplementary Fig.\u0026nbsp;7B-C). Notably, the distribution of structural variants (SVs) was asymmetric between ecotypes, with the An subgenome containing a greater number of ecotype-specific SVs (i.e., structural variants that are present, absent, or differ in size in over 70% of individuals within a single ecotype). We also observed certain chromosomal regions enriched with SVs showing subpopulation imbalance. For example, ecotype-specific SVs on chromosomes A09 (18.5\u0026ndash;32.7 Mb), C02 (2.1\u0026ndash;2.9 Mb, 33.8\u0026ndash;37.9 Mb), and C08 (3.4\u0026ndash;6.5 Mb) suggest these regions may harbor loci involved in ecotype differentiation (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eE). The interval from 18.5 to 32.7 Mb on chromosome C02 includes the extensively studied flowering time inhibitor gene \u003cem\u003eBnaC02.FLC\u003c/em\u003e, hypothesized to mediate ecological differentiation through negative regulation of flowering time. Two known TEMSVs affect \u003cem\u003eBnaC02.FLC\u003c/em\u003e expression [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e], leading to variation in flowering time among individuals carrying these mutations. Additional TEMSVs were also found at 2.1\u0026ndash;2.9 Mb on chromosome C02 and 3.4\u0026ndash;6.5 Mb on chromosome C08, near genes involved in rapeseed development. For instance, \u003cem\u003eBnaC02.SDP1\u003c/em\u003e, which participates in the biosynthesis of linolenic acid, linoleic acid, and arachidonic acid, could potentially influence rapeseed SOC by regulating related pathways. A TEMSV located about 4 kb upstream of the \u003cem\u003eBnaC02.SDP1\u003c/em\u003e in semi-winter rapeseed might cause developmental differences among ecotypes. At the same time, a TEMSV was identified in the pericentromeric region of A09. This SV is located downstream of the \u003cem\u003eBnaA09.CPK\u003c/em\u003e gene of semi-winter rapeseed, which encodes peroxidase and calcium-dependent protein kinase (CPK) and may have the effect of improving stress resistance (Supplementary Fig.\u0026nbsp;7D). Collectively, these findings highlight the significant role of SVs in shaping the genetic architecture of \u003cem\u003eB. napus\u003c/em\u003e, driving ecotype differentiation, modulating gene expression, and enhancing key agronomic traits.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eEffect of TEMSVs on Gene Expression\u003c/h3\u003e\n\u003cp\u003eSVs can affect the expression of nearby genes by altering coding sequences or disrupting regulatory elements. In particular, TEMSVs play a significant role in regulating gene expression [\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e, \u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e]. Using the graph-based pan-genome of \u003cem\u003eB. napus\u003c/em\u003e, we identified 124,682 high-confidence TEMSVs, with more than 80% sequence length overlap between SVs and transposable elements. 49.46% TEMSVs were localized in the distant genomic regions (beyond 3 kb upstream of promoter, Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eA). The distribution of various types of TEMSVs was consistent in the subgenomes (Supplementary Fig.\u0026nbsp;8). We defined a \u0026ldquo;TEMSV-gene\u0026rdquo; as any gene containing TEMSVs within its coding region or within 3 kb upstream of its transcription start site. Notably, the proportion of TEMSV-genes in the An subgenome (61.48%) was substantially higher than in the Cn subgenome (43.78%) (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eB). Moreover, the average number of TEMSVs per gene in the An subgenome was significantly greater than that in the Cn subgenome (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eC). We compared the expression levels of each TEMSV-gene in the 505 \u003cem\u003eB. napus\u003c/em\u003e accessions with and without TEMSV. The differences in gene expression between these two groups were considerably greater than what would be expected by random sampling (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eD).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eA 592 bp TE-mediated insertion was identified 247 bp upstream of \u003cem\u003eULTRACURVATA 1\u003c/em\u003e (\u003cem\u003eBnaC03.UCU1\u003c/em\u003e). In the surveyed \u003cem\u003eB. napus\u003c/em\u003e lines, 81.35% carried this insertion (N\u0026thinsp;=\u0026thinsp;410) (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eE). Gene expression analysis at 20 days after flowering (DAF) of population revealed that the insertion significantly reduced the expression of \u003cem\u003eBnaC03.UCU1\u003c/em\u003e (\u003cem\u003eP\u003c/em\u003e\u003csub\u003eexp\u003c/sub\u003e = 1.2\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\times\\:\\)\u003c/span\u003e\u003c/span\u003e10\u003csup\u003e\u0026minus;5\u003c/sup\u003e; Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eF). In addition, SV-based GWAS analysis identified a significant association between the 592 bp insertion and seed oleic acid (C18:1) content (Supplementary Fig.\u0026nbsp;9A). Varieties with the insertion exhibited significantly lower C18:1 content than those without the insertion (\u003cem\u003eP\u003c/em\u003e\u003csub\u003eC18:1\u003c/sub\u003e = 1.0\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\times\\:\\)\u003c/span\u003e\u003c/span\u003e10\u003csup\u003e\u0026minus;6\u003c/sup\u003e; Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eG). To further explore the effect of \u003cem\u003ecis\u003c/em\u003e-acting variants in this gene on C18:1, eGWAS of \u003cem\u003eBnaC03.UCU1\u003c/em\u003e was performed in 20 DAF, and the results showed co-localization of the eQTL for this gene with the QTL for C18:1 content (Supplementary Fig.\u0026nbsp;9B-C). These findings suggest that the insertion in promoter of \u003cem\u003eBnaC03.UCU1\u003c/em\u003e may alters the seed C18:1 content by modulating gene expression.\u003c/p\u003e\n\u003ch3\u003eTEMSV effects on the expression of homoeologous gene pairs between subgenomes\u003c/h3\u003e\n\u003cp\u003eTo investigate the effects of TEMSVs on inter-subgenomic gene expression, we examined its impact on the expression of homoeologous gene pairs (HGPs). An HGP was defined as a pair of homoeologous copies with highest sequence similarity between subgenomes, with one copy from An and one from Cn. A total of 19,856 HGPs were identified (Supplementary Table\u0026nbsp;11), among which 15,902 (80.09%) contained TEMSVs. Notably, in HGPs harboring TEMSVs, the distribution was unbalanced: 9,159 (46.13%) HGPs exhibited a higher number of TEMSVs in the An copy, nearly twice as many as in the Cn copy (Supplementary Fig.\u0026nbsp;10).\u003c/p\u003e \u003cp\u003eBased on the ZS11 whole-period transcriptome data and seed transcriptome data of the population [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e, \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e], we constructed a comprehensive gene expression profile of HGPs. Overall, homoeologous genes in HGPs exhibited significantly higher expression correlation than randomly paired genes (RGPs) (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eA, Supplementary Fig.\u0026nbsp;11). Additionally, we categorized HGPs into three types based on the position of TEMSVs: (i) TEMSVs in the coding region, (ii) TEMSVs in the non-coding region, and (iii) no TEMSVs. Analysis of the expression correlations revealed that HGPs without TEMSVs had a higher proportion of highly correlated expression than the other two categories (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eB-C), suggesting that the presence of TEMSVs partially disrupts the expression correlation of homoeologous genes.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eTo further explore inter-subgenomic relationships, we performed differential expression analysis of HGPs across various tissues of ZS11. Among the identified HGPs, 51.25% of HGPs exhibit unbalanced expression (uHGP, defined as |\u003cem\u003elog\u003c/em\u003e\u003csub\u003e\u003cem\u003e2\u003c/em\u003e\u003c/sub\u003e(fold change)| \u0026gt; 1 and \u003cem\u003eP\u003c/em\u003e\u003csub\u003eadj\u003c/sub\u003e \u0026lt; 0.05) in at least one tissue (Supplementary Fig.\u0026nbsp;12). We compared expression fold changes between HGPs with TEMSVs and non-TEMSV HGPs (i.e., genes whose regions do not overlap with any TEMSV) across the seven tissues (silique, seed, silique wall, bud, flower, leaf and stem) over multiple time points, we found that the proportion HGPs with TEMSVs exhibiting over a one-fold change in expression was consistently higher than that of non-TEMSV HGPs (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eD; Supplementary Table\u0026nbsp;12). Furthermore, we examined how the proportion of uHGPs varied with the number of TEMSVs present in HGPs. Our analysis revealed that as the number of TEMSVs increased, the proportion of uHGPs also increased (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eE), a trend consistently observed across all tissues (Supplementary Fig.\u0026nbsp;13). These findings suggest that TEMSVs contribute to unbalanced expression of HGPs among subgenomes, with the degree of imbalance intensifying as more TEMSVs accumulate.\u003c/p\u003e \u003cp\u003eSilique wall samples contained the highest number of TEMSV-associated uHGPs (478). We defined HGPs with unbalanced expression in all examined tissues as \u0026ldquo;shared uHGPs\u0026rdquo;. Notably, the number of shared uHGPs decreased as the criterion was extended to include more tissues. However, this decline was significantly slower for TEMSV-HGPs than for randomly sampled RGPs (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eF). In addition, the expression fold changes of shared TEMSV-HGPs were found to be more stable across different tissues (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eG). For example, the gene \u003cem\u003eCDC5\u003c/em\u003e, which regulate the cell division cycle, contains a TEMSV in the An copy (Supplementary Fig.\u0026nbsp;14). Analysis of the expression profile of this HGP revealed that the An copy consistently exhibited 3- to 4-fold higher expression than the Cn copy across all tissues (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eH). These results suggest that uHGPs containing TEMSVs maintain a stable pattern of unbalanced expression across tissues.\u003c/p\u003e\n\u003ch3\u003eIdentification of HGPs with dosage compensation effects\u003c/h3\u003e\n\u003cp\u003eAmong 19,856 HGPs, 43,762 TEMSVs were identified. Of these TEMSVs, 1,405 were associated with significant changes in the expression of nearby genes at both 20 and 40 DAF in seed (Student's \u003cem\u003et\u003c/em\u003e-test; \u003cem\u003eP\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.05). In addition, among the 292 HGPs, when the expression level of homoeologous copy near TEMSV changes, the expression level of distant homoeologous copy also shows significant changes (Student's \u003cem\u003et\u003c/em\u003e-test; \u003cem\u003eP\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.05). Notably, the magnitude of expression change was significantly greater in the homoeologous copy proximal to TEMSVs compared with its distal counterpart (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eA, Supplementary Fig.\u0026nbsp;15A). Further analysis revealed that in 32.30% of HGPs with TEMSVs, the two homoeologous copies exhibited opposite expression changes (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eB, Supplementary Fig.\u0026nbsp;15B). For example, in the HGP of \u003cem\u003eNUCLEOSIDE DIPHOSPHATE KINASE 3\u003c/em\u003e (\u003cem\u003eNDPK3\u003c/em\u003e), the absence of a TE insertion corresponded with significantly higher expression of \u003cem\u003eBnaA02.NDPK3\u003c/em\u003e relative to \u003cem\u003eBnaC02.NDPK3\u003c/em\u003e. However, following the insertion of an 8,141 bp sequence at the 3,209 bp upstream the TSS of \u003cem\u003eBnaA02.NDPK3\u003c/em\u003e (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eC), the expression level of \u003cem\u003eBnaA02.NDPK3\u003c/em\u003e significantly decreased, whereas that of \u003cem\u003eBnaC02.NDPK3\u003c/em\u003e markedly increased. Consequently, the overall expression level of the \u003cem\u003eNDPK3\u003c/em\u003e homoeologous pair was reduced after TE insertion (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eD). These results suggest that some HGPs are subject to feedback regulation. When TEMSVs cause changes in one gene copy, the other copy adjusts its expression, thereby exhibiting a dosage compensation effect. This compensatory mechanism may help stabilize overall gene expression in the face of perturbations caused by TEMSVs.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cb\u003eA 104-bp TE-mediated insertion influenced seed oil content in\u003c/b\u003e \u003cb\u003eB. napus\u003c/b\u003e\u003c/p\u003e \u003cp\u003eTEMSV can affect not only gene expression but also important agronomic traits. In \u003cem\u003eB. napus\u003c/em\u003e, SOC is one of the important traits. Using the SV map constructed by pan-genome, we performed an SV-based genome-wide association study (SV-GWAS) to analyze SOC. We identified a 104-bp TE-mediated insertion (TEMI) located in the promoter of the \u003cem\u003epeptidase C78\u003c/em\u003e (\u003cem\u003eBnaA06\u003c/em\u003e.\u003cem\u003ePC78\u003c/em\u003e) on chromosome A06 that is significantly associated with SOC (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eE-F). Among 505 accessions examined, 63.98% carried the TEMI allele. Importantly, accessions harboring the TEMI exhibited significantly lower SOC than those lacking the insertion (\u003cem\u003eP\u003c/em\u003e\u003csub\u003eSOC\u003c/sub\u003e = 1.41\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\times\\:\\)\u003c/span\u003e\u003c/span\u003e10\u003csup\u003e\u0026minus;5\u003c/sup\u003e; Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eG-H).\u003c/p\u003e \u003cp\u003eExpression analysis further revealed that at 40 DAF in seed, the TEMI allele significantly decreased the expression of \u003cem\u003eBnaA06.PC78\u003c/em\u003e, whereas its homolog in the Cn subgenome (\u003cem\u003eBnaC07.PC78\u003c/em\u003e) was significantly increased. At 20 DAF, although the decrease in \u003cem\u003eBnaA06.PC78\u003c/em\u003e expression was not statistically significant, a downward trend was observed alongside a significant upregulation of \u003cem\u003eBnaC07.PC78\u003c/em\u003e (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eI-J). These findings suggest that the HGP of \u003cem\u003eBna.PC78\u003c/em\u003e is subject to feedback regulation, wherein the perturbation of one copy by TEMI leads to compensatory expression changes in its homolog, thereby influencing SOC.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eThe genomes of ZS11 and Westar have become indispensable resources in \u003cem\u003eB. napus\u003c/em\u003e research. ZS11 is frequently used as a reference genome for diverse analyses [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e], whereas Westar\u0026rsquo;s high transformation efficiency makes it a preferred recipient for gene editing experiments [\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e]. Our comprehensive mapping of nearly complete genomes using ONT represents a significant advance. With only 38 gaps in ZS11 and 94 in Westar, and notably 15 gap-free chromosomes in ZS11 and one in Westar (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eA), these assemblies markedly improve upon previous versions. The number of genes annotated in the centromere region of ZS11 v1 increased by 78.89% compared to the earlier ZS11 v0 assembly (Supplementary Fig.\u0026nbsp;2). Future efforts should integrate Hi-C, PacBio HiFi, and sequence-specific gap-filling approaches to resolve these challenging regions. Moreover, by extending our analysis to six additional representative \u003cem\u003eB. napus\u003c/em\u003e lines, whose assemblies ranged from 824.77 to 949.56 Mb with consistently high BUSCO scores (98.60-99.09%; Supplementary Table\u0026nbsp;6), we have broadened the genomic framework to encompass both semi-wintering and spring ecotypes. These high-quality assemblies not only enhance our understanding of \u003cem\u003eB. napus\u003c/em\u003e genetic diversity but also provide valuable resources for future breeding programs.\u003c/p\u003e \u003cp\u003eSVs have emerged as a significant contributor to genetic diversity and phenotypic variation in crops [\u003cspan additionalcitationids=\"CR15\" citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]. Pan-genome analyses, which integrate multiple high-quality assemblies, offer an effective platform for accurately identifying these SVs [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. In this study, we constructed a comprehensive pan-genome using 22 high-quality \u003cem\u003eB. napus\u003c/em\u003e genomes, enabling the identification of 221,062 unique SVs (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eA). Analysis of resequencing data from 505 accessions further highlighted the value of SVs for differentiating \u003cem\u003eB. napus\u003c/em\u003e ecotypes (Supplementary Fig.\u0026nbsp;7B-C). This extensive reservoir of genetic variation constitutes a valuable resource for elucidating genotype\u0026ndash;phenotype relationships and accelerating molecular breeding in \u003cem\u003eB. napus\u003c/em\u003e.\u003c/p\u003e \u003cp\u003eNatural hybridization and genome doubling in \u003cem\u003eB. napus\u003c/em\u003e have resulted in significant asymmetry between the An and Cn subgenomes, especially in terms of gene expression and epigenetic modifications [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. Previous eQTL mapping studies have shown that asymmetric transcriptional regulation underpins expression imbalances between subgenomes [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e]. Here, we extend these findings by showing that TEMSVs are more abundant in the An subgenome, thereby contributing to unbalanced expression of HGPs (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e). Moreover, our results indicate that as the number of TEMSVs increases, the expression imbalance between the two subgenomic copies increasingly pronounced, with these effects consistently observed across multiple tissues (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e). These findings deepen our theoretical understanding of regulatory interactions within polyploid genomes. Current investigations into the effects of TEMSVs on HGPs have primarily focused on transcriptional regulation. As chromatin accessibility is closely linked to gene expression [\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e], further research is warranted to explore the epistatic effects of TEMSVs on HGP regulation.\u003c/p\u003e \u003cp\u003eIntegrating pan-genomics with other omics technologies is a powerful strategy for dissecting trait variation and accelerating crop improvement. By incorporating SVs and other genetic variants into a graph-based pan-genome framework, we improved the resolution of GWAS and identified robust associations between specific SVs and key agronomic traits. Notably, two SVs with significant effects on seed traits were identified (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eE-G, Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eE-J). Future studies should focus on functional characterization of these SVs and their associated genes using CRISPR-based genome editing. Additionally, expanding trait association analyses to encompass a broader range of phenotypes may facilitate the discovery of further functionally relevant SVs. These efforts will underscore the functional importance of structural variants and their potential application in breeding programs.\u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eMethods\u003c/h2\u003e \u003cdiv id=\"Sec9\" class=\"Section3\"\u003e \u003ch2\u003eGenome sequencing and de novo assembly\u003c/h2\u003e \u003cp\u003eIn this study, the collection of young leaves of ZS11 and Westar were subjected to Nanopore library construction and sequencing. PacBio and Illumina sequencing was performed for six genomes of Zhongshuang 2 (ZS2), Bugle, 352, 862, Ribenyoucai (RB) and Baihua (BH). Canu software was applied to filter out reads less than 5,000 bp in length from Nanopore and PacBio sequencing results and perform self-correction and trimming [\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e]. For Illumina data, removal of reads with consecutive mass values\u0026thinsp;\u0026le;\u0026thinsp;20 bases up to 40%; removal of N-containing reads and removing duplication contamination.\u003c/p\u003e \u003cp\u003eThe genome assembly process was divided into two steps: (1) De-novo assembly of Nanopore or PacBio data using NextDenovo software [\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e]. (2) Correction of single-base errors and small indel errors in the assembly results with Illumina data using NextPolish software [\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e\n\u003ch3\u003eRepetitive sequence annotation\u003c/h3\u003e\n\u003cp\u003eTandem repeat sequences were first annotated using GMATA and Tandem Repeats Finder (TRF) software[\u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e, \u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e]. The non-initial repeat libraries for each genome were then predicted using MITE-hunter and RepeatModeler2 (using default parameters), respectively [\u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e, \u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e]. The obtained libraries were then compared with the TE class Repbase (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://www.girinst.org/repbase\u003c/span\u003e\u003cspan address=\"http://www.girinst.org/repbase\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) to categorize the type of each repeat family. To further characterize repetitive sequences across the genome, RepeatMasker was applied to match the sequences against new repetitive sequence libraries and Repbase TE libraries in order to search for known and novel transposable elements [\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e]. Overlapping transposable elements belonging to the same repeat class were organized and merged.\u003c/p\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eGene prediction\u003c/h2\u003e \u003cp\u003eThree independent methods were used for gene prediction, including ab initio prediction, homology search, and reference-guided transcriptome assembly. Firstly, the homoeologous peptides of \u003cem\u003eArabidopsis thaliana\u003c/em\u003e, \u003cem\u003eB. oleracea\u003c/em\u003e, \u003cem\u003eB. rapa\u003c/em\u003e and \u003cem\u003eB.napus\u003c/em\u003e Darmor were compared with each genome using GeMoMa [\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e], and then the gene structure information was obtained. Next gene prediction based on RNA-seq was performed using STAR (default) to compare the filtered RNA-seq reads to the reference genome [\u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e]. Further transcripts are assembled using StringTie and open reading frames (ORFs) are predicted using PASA to generate a training set, and non-initial gene prediction is performed on the training set using AUGUSTUS with default parameters[\u003cspan additionalcitationids=\"CR55\" citationid=\"CR54\" class=\"CitationRef\"\u003e54\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e]. Finally, different sets of predicted genes obtained from head prediction, homology annotation and transcriptome prediction were integrated by EVidenceModeler (EVM)[\u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e55\u003c/span\u003e], besides that, genes with TEs were removed using TransposonPSI package (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://transposonpsi.sourceforge.net/\u003c/span\u003e\u003cspan address=\"http://transposonpsi.sourceforge.net/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) and filtered the heterozygous genes. Based on the RNA-seq assembly results, untranslated regions (UTRs) and alternative splicing regions were identified using PASA, the longest transcripts of each locus were retained, and regions other than ORFs were designated as UTRs.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eFunctional annotation of gene models\u003c/h2\u003e \u003cp\u003eGene functional information as well as protein motifs and structural domains were annotated by comparison with public databases (SwissProt, NR, KEGG, KOG and Gene Ontology). Putative structural domains and GO Term of genes were determined using InterProScan software and default parameters. And then the EvidenceModeler integrated protein sequences were compared with four public protein databases using BLASTP with a threshold of 1e-05 for E-value [\u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e55\u003c/span\u003e, \u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e57\u003c/span\u003e], and the results with the lowest E-value were retained. The results from the five databases were finally integrated.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eAnnotation of non-coding RNA (ncRNA)\u003c/h2\u003e \u003cp\u003eTwo strategies were used for the study: searching against databases and utilizing model predictions to obtain ncRNAs (non-coding RNAs). Transfer RNA (tRNA) was predicted using tRNAscan-SE and eukaryotic parameters [\u003cspan citationid=\"CR58\" class=\"CitationRef\"\u003e58\u003c/span\u003e]. Rfam database was searched using Infernal cmscan for MicroRNA, rRNA, small nuclear RNA and small nucleolar RNA[\u003cspan citationid=\"CR59\" class=\"CitationRef\"\u003e59\u003c/span\u003e]. rRNA and its subunits were predicted using RNAmmer [\u003cspan citationid=\"CR60\" class=\"CitationRef\"\u003e60\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003eMitosis and telomere annotation\u003c/h2\u003e \u003cp\u003eMitotic-specific repetitive sequences (CentBr1, CentBr2, CRB, PCRBr, TR238, and TR805) were aligned to the reference genome using blastn software [\u003cspan citationid=\"CR61\" class=\"CitationRef\"\u003e61\u003c/span\u003e, \u003cspan citationid=\"CR62\" class=\"CitationRef\"\u003e62\u003c/span\u003e], and mitotic regions were identified by counting the density of associated repeat sequences and gene density of each chromosome region in a 100 Kb window with a sliding window of 100 Kb step size. The annotation of telomeres was done by searching for telomere-specific repeats 5'- (TTTAGGG)n -3' [\u003cspan citationid=\"CR63\" class=\"CitationRef\"\u003e63\u003c/span\u003e, \u003cspan citationid=\"CR64\" class=\"CitationRef\"\u003e64\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003ePhylogenetic analysis\u003c/h2\u003e \u003cp\u003eIn order to construct phylogenetic relationships among 20 varieties of \u003cem\u003eBrassica napus\u003c/em\u003e, we did an all-to-all blastp with peptide sequences of protein-coding genes annotated from these 20 genomes by diamond [\u003cspan citationid=\"CR65\" class=\"CitationRef\"\u003e65\u003c/span\u003e] with cut-off evalues of 1e-5 and then input the results into OrthoFinder. The single-copy orthologous genes were further extracted from OrthoFinder results, protein sequences were aligned by muscle v3.8.31 [\u003cspan citationid=\"CR66\" class=\"CitationRef\"\u003e66\u003c/span\u003e] and conserved sites from multiple sequence alignment were extracted by Gblocks v0.91b [\u003cspan citationid=\"CR67\" class=\"CitationRef\"\u003e67\u003c/span\u003e]. The phylogenetic tree was constructed by RAxML v8.2.12 [\u003cspan citationid=\"CR68\" class=\"CitationRef\"\u003e68\u003c/span\u003e] to automatically find the best model and perform 1000 ultrafast bootstrap analyzes to test the robustness of each branch. The online tool iTOL (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://itol.embl.de\u003c/span\u003e\u003cspan address=\"http://itol.embl.de\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) was used to visualize the constructed tree.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003eConstruction of the protein-coding gene-based pan-genome\u003c/h2\u003e \u003cp\u003eWe did an all-to-all alignment with 1,911,393 protein sequences of the 20 \u003cem\u003eB. napus\u003c/em\u003e varieties and input the result file into OrthoFinder for gene family construction. Finally, we obtained 101,903 non-redundant gene clusters and 48,233 clusters with only one gene. We then classified non-redundant gene clusters into 4 categories: core gene clusters that were conserved in all 20 varieties; soft-core gene clusters, which were present in 18\u0026ndash;19 varieties; dispensable gene clusters, which were found in 2\u0026ndash;17 genomes; and private gene clusters, which contained genes from only 1 sample (not including unassigned genes). The longest transcript was chosen to represent each gene.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003eKa/Ks calculation of different types of pan-genes\u003c/h2\u003e \u003cp\u003eNon-synonymous substitution rates (Ka), synonymous substitution rates (Ks), and Ka/Ks in core, softcore, and dispensable gene clusters were computed using the WGDI [\u003cspan citationid=\"CR69\" class=\"CitationRef\"\u003e69\u003c/span\u003e] with default parameters. We conducted amino acid alignment for gene pairs in each cluster first and then converted the results into the coding sequence (CDS) alignment using PAL2NAL [\u003cspan citationid=\"CR70\" class=\"CitationRef\"\u003e70\u003c/span\u003e]. The alignments were further passed to the PAML [\u003cspan citationid=\"CR71\" class=\"CitationRef\"\u003e71\u003c/span\u003e] to obtain Ka/Ks values.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec18\" class=\"Section2\"\u003e \u003ch2\u003eConstruction of the graph-based pan-genome and SVs calling\u003c/h2\u003e \u003cp\u003eFirst, we used Minigraph-Cactus to construct the graph pangenome of the 22 high-quality \u003cem\u003eB. napus\u003c/em\u003e genome assemblies based on sequence alignment. The ZS11 genome assembled in this study was set as the reference and the other 21 genomes were added into the multi-assembly graph successively. The fragments differing from the reference genome are displayed as different paths in the generated graphical fragment assembly (GFA) file. After obtaining the first version of the pan-genome, we used MUMmer4 [\u003cspan citationid=\"CR72\" class=\"CitationRef\"\u003e72\u003c/span\u003e] to align the other 21 genomes to the ZS11 reference genome, and used SyRI [\u003cspan citationid=\"CR73\" class=\"CitationRef\"\u003e73\u003c/span\u003e] to obtain the structural variation between the genome and ZS11. Finally, we integrated the structural variants produced by both methods and used vg toolkit v1.34.0 [\u003cspan citationid=\"CR74\" class=\"CitationRef\"\u003e74\u003c/span\u003e] to construct the final complete pan-genome.\u003c/p\u003e \u003cp\u003eTo genotype the SVs in the 505 \u003cem\u003eB. napus\u003c/em\u003e accessions we mapped the short reads from each individual to the graph-based pan-genome via vg using default parameters. After filtering individuals with a missing rate above 0.5 or minor allele frequency (MAF) above 0.05 using Plink [\u003cspan citationid=\"CR75\" class=\"CitationRef\"\u003e75\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eComputations in this study were conducted on the bioinformatics computing platform of the National Key Laboratory of Crop Genetic Improvement, Huazhong Agricultural University. This work is supported by National Natural Science Foundation of China (32200501),\u0026nbsp;Project X2662024ZKPY001 supported by the Fundamental Research Funds for the Central Universities, Knowledge Innovation Program of Wuhan-Basic Research (2022020801010221), Project supported by the Fundamental Research Funds for the Central Universities (2662022YJ016), Chongqing Technical Innovation and Application Development Special Project (CSTB2024TIAD-KPX0025),\u0026nbsp;Basic Research Project in 2023 of Yazhouwan National Laboratory (GL23YCKY01) and the fellowship of China National Postdoctoral Program for Innovation Talents (BX20240299).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eH.Z., K.L., J.S. and L.G. designed and supervised this study. Z.T., Y.Z., J.Y. and Z.Z. performed the bioinformatics analysis. Z.T., Y.Z. and Z.G. prepared the manuscript. H.Z., K.L., J.S. and Q.Y. revised the manuscript. All the authors read and approved the manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare no competing interests.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eGu J, Guan Z, Jiao Y, Liu K, Hong D (2024) The story of a decade: genomics, functional genomics and molecular breeding in Brassica napus. Plant Commun\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChalhoub B, Denoeud F, Liu S, Parkin IA, Tang H, Wang X, Chiquet J, Belcram H, Tong C, Samans B (2014) Early allopolyploid evolution in the post-Neolithic Brassica napus oilseed genome. \u003cem\u003escience\u003c/em\u003e 345:950\u0026ndash;953\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTan Z, Han X, Dai C, Lu S, He H, Yao X, Chen P, Yang C, Zhao L, Yang QY (2024) Functional genomics of Brassica napus: Progress, challenges, and perspectives. J Integr Plant Biol 66:484\u0026ndash;509\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLu K, Wei L, Li X, Wang Y, Wu J, Liu M, Zhang C, Chen Z, Xiao Z, Jian H (2019) Whole-genome resequencing reveals Brassica napus origin and genetic loci involved in its improvement. Nat Commun 10:1154\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eQian L, Qian W, Snowdon RJ (2014) Sub-genomic selection patterns as a signature of breeding in the allopolyploid Brassica napus genome. BMC Genomics 15:1\u0026ndash;17\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRong J, Feltus FA, Waghmare VN, Pierce GJ, Chee PW, Draye X, Saranga Y, Wright RJ, Wilkins TA, May OL (2007) Meta-analysis of polyploid cotton QTL shows unequal contributions of subgenomes to a complex network of genes and gene clusters implicated in lint fiber development. Genetics 176:2577\u0026ndash;2588\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLv Z, Li Z, Wang M, Zhao F, Zhang W, Li C, Gong L, Zhang Y, Mason AS, Liu B (2021) Conservation and trans-regulation of histone modification in the A and B subgenomes of polyploid wheat during domestication and ploidy transition. BMC Biol 19:1\u0026ndash;16\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang M, Li Z, Zhang Ye, Zhang Y, Xie Y, Ye L, Zhuang Y, Lin K, Zhao F, Guo J (2021) An atlas of wheat epigenetic regulatory elements reveals subgenome divergence in the regulation of development and stress responses. Plant Cell 33:865\u0026ndash;881\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang Q, Guan P, Zhao L, Ma M, Xie L, Li Y, Zheng R, Ouyang W, Wang S, Li H (2021) Asymmetric epigenome maps of subgenomes reveal imbalanced transcription and distinct evolutionary trends in Brassica napus. Mol Plant 14:604\u0026ndash;619\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhu H, Han X, Lv J, Zhao L, Xu X, Zhang T, Guo W (2011) Structure, expression differentiation and evolution of duplicated fiber developmental genes in Gossypium barbadense and G. hirsutum. BMC Plant Biol 11:1\u0026ndash;15\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhuang W, Chen H, Yang M, Wang J, Pandey MK, Zhang C, Chang W-C, Zhang L, Zhang X, Tang R (2019) The genome of cultivated peanut provides insight into legume karyotypes, polyploid evolution and crop domestication. Nat Genet 51:865\u0026ndash;876\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCheng F, Wu J, Cai X, Liang J, Freeling M, Wang X (2018) Gene retention, fractionation and subgenome differences in polyploid plants. Nat plants 4:258\u0026ndash;268\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang K, Wang X, Cheng F (2019) Plant polyploidy: origin, evolution, and its influence on crop domestication. Hortic Plant J 5:231\u0026ndash;239\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAlonge M, Wang X, Benoit M, Soyk S, Pereira L, Zhang L, Suresh H, Ramakrishnan S, Maumus F, Ciren D (2020) Major impacts of widespread structural variation on gene expression and crop improvement in tomato. Cell 182:145\u0026ndash;161 e123\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhou Y, Zhang Z, Bao Z, Li H, Lyu Y, Zan Y, Wu Y, Cheng L, Fang Y, Wu K (2022) Graph pangenome captures missing heritability and empowers tomato breeding. Nature 606:527\u0026ndash;534\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGuo N, Wang S, Wang T, Duan M, Zong M, Miao L, Han S, Wang G, Liu X, Zhang D (2024) A graph-based pan-genome of Brassica oleracea provides new insights into its domestication and morphotype diversification. Plant Commun 5\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang Y, Yang Z, He Y, Liu D, Liu Y, Liang C, Xie M, Jia Y, Ke Q, Zhou Y Structural variation reshapes population gene expression and trait variation in 2,105 Brassica napus accessions. Nat Genet 2024:1\u0026ndash;13\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNiu X-M, Xu Y-C, Li Z-W, Bian Y-T, Hou X-H, Chen J-F, Zou Y-P, Jiang J, Wu Q, Ge S (2019) Transposable elements drive rapid phenotypic variation in Capsella rubella. \u003cem\u003eProceedings of the National Academy of Sciences\u003c/em\u003e 116:6908\u0026ndash;6913\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDom\u0026iacute;nguez M, Dugas E, Benchouaia M, Leduque B, Jim\u0026eacute;nez-G\u0026oacute;mez JM, Colot V, Quadrana L (2020) The impact of transposable elements on tomato diversity. Nat Commun 11:4058\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSong J-M, Guan Z, Hu J, Guo C, Yang Z, Wang S, Liu D, Wang B, Lu S, Zhou R (2020) Eight high-quality genomes reveal pan-genome architecture and ecotype differentiation of Brassica napus. Nat plants 6:34\u0026ndash;45\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBayer PE, Golicz AA, Scheben A, Batley J, Edwards D (2020) Plant pan-genomes are the new reference. Nat plants 6:914\u0026ndash;920\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSchreiber M, Jayakodi M, Stein N, Mascher M Plant pangenomes for crop improvement, biodiversity and evolution. Nat Rev Genet 2024:1\u0026ndash;15\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu Y, Du H, Li P, Shen Y, Peng H, Liu S, Zhou G-A, Zhang H, Liu Z, Shi M (2020) Pan-genome of wild and cultivated soybeans. Cell 182:162\u0026ndash;176 e113\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShang L, Li X, He H, Yuan Q, Song Y, Wei Z, Lin H, Hu M, Zhao F, Zhang C (2022) A super pan-genomic landscape of rice. Cell Res 32:878\u0026ndash;896\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGui S, Wei W, Jiang C, Luo J, Chen L, Wu S, Li W, Wang Y, Li S, Yang N (2022) A pan-Zea genome map for enhancing maize improvement. Genome Biol 23:178\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang M, Li J, Qi Z, Long Y, Pei L, Huang X, Grover CE, Du X, Xia C, Wang P (2022) Genomic innovation and regulatory rewiring during evolution of the cotton genus Gossypium. Nat Genet 54:1959\u0026ndash;1971\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHuang Y, He J, Xu Y, Zheng W, Wang S, Chen P, Zeng B, Yang S, Jiang X, Liu Z (2023) Pangenome analysis provides insight into the evolution of the orange subfamily and a key gene for citric acid accumulation in citrus fruits. Nat Genet 55:1964\u0026ndash;1975\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu D, Yu L, Wei L, Yu P, Wang J, Zhao H, Zhang Y, Zhang S, Yang Z, Chen G (2021) BnTIR: an online transcriptome platform for exploring RNA-seq libraries for oil crop Brassica napus. Plant Biotechnol J 19:1895\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTan Z, Peng Y, Xiong Y, Xiong F, Zhang Y, Guo N, Tu Z, Zong Z, Wu X, Ye J (2022) Comprehensive transcriptional variability analysis reveals gene networks regulating seed oil content of Brassica napus. Genome Biol 23:233\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi L, Tian Z, Chen J, Tan Z, Zhang Y, Zhao H, Wu X, Yao X, Wen W, Chen W (2023) Characterization of novel loci controlling seed oil content in Brassica napus by marker metabolite-based multi-omics analysis. Genome Biol 24:141\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRousseau-Gueutin M, Belser C, Da Silva C, Richard G, Istace B, Cruaud C, Falentin C, Boideau F, Boutte J, Delourme R (2020) Long-read assembly of the Brassica napus reference genome Darmor-bzh. GigaScience 9:giaa137\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTan Z, Xie Z, Dai L, Zhang Y, Zhao H, Tang S, Wan L, Yao X, Guo L, Hong D (2022) Genome-and transcriptome‐wide association studies reveal the genetic basis and the breeding history of seed glucosinolate content in Brassica napus. Plant Biotechnol J 20:211\u0026ndash;225\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTang S, Zhao H, Lu S, Yu L, Zhang G, Zhang Y, Yang Q-Y, Zhou Y, Wang X, Ma W (2021) Genome-and transcriptome-wide association studies provide insights into the genetic basis of natural variation of seed oil content in Brassica napus. Mol Plant 14:470\u0026ndash;487\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang P, Song Y, Xie Z, Wan M, Xia R, Jiao Y, Zhang H, Yang G, Fan Z, Yang QY, Hong D (2023) Xiaoyun, a model accession for functional genomics research in \u003cem\u003eBrassica napus\u003c/em\u003e. Plant Commun 5:100727\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang Z, Wang S, Wei L, Huang Y, Liu D, Jia Y, Luo C, Lin Y, Liang C, Hu Y (2023) BnIR: A multi-omics database with various tools for Brassica napus research and breeding. Mol Plant 16:775\u0026ndash;789\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLee H, Chawla HS, Obermeier C, Dreyer F, Abbadi A, Snowdon R (2020) Chromosome-scale assembly of winter oilseed rape Brassica napus. Front Plant Sci 11:496\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDavis JT, Li R, Kim S, Michelmore R, Kim S, Maloof JN (2023) Whole-genome sequence of synthetically derived Brassica napus inbred cultivar Da-Ae. \u003cem\u003eG3: Genes, Genomes, Genetics\u003c/em\u003e 13:jkad026\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eQu C, Zhu M, Hu R, Niu Y, Chen S, Zhao H, Li C, Wang Z, Yin N, Sun F (2023) Comparative genomic analyses reveal the genetic basis of the yellow-seed trait in Brassica napus. Nat Commun 14:5194\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHickey G, Monlong J, Ebler J, Novak AM, Eizenga JM, Gao Y, Human Pangenome Reference C, Marschall T, Li H, Paten B (2024) Pangenome graph construction from genome alignments with Minigraph-Cactus. Nat Biotechnol 42:663\u0026ndash;673\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eQin P, Lu H, Du H, Wang H, Chen W, Chen Z, He Q, Ou S, Zhang H, Li X (2021) Pan-genome analysis of 33 genetically diverse rice accessions reveals hidden genomic variations. Cell 184:3542\u0026ndash;3558 e3516\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGill RA, Scossa F, King GJ, Golicz AA, Tong C, Snowdon RJ, Fernie AR, Liu S (2021) On the role of transposable elements in the regulation of gene expression and subgenomic interactions in crop genomes. CRC Crit Rev Plant Sci 40:157\u0026ndash;189\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDai C, Li Y, Li L, Du Z, Lin S, Tian X, Li S, Yang B, Yao W, Wang J (2020) An efficient Agrobacterium-mediated transformation method using hypocotyl as explants for Brassica napus. Mol Breeding 40:1\u0026ndash;13\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVoss TC, Hager GL (2014) Dynamic regulation of transcriptional states by chromatin and transcription factors. Nat Rev Genet 15:69\u0026ndash;81\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKoren S, Walenz BP, Berlin K, Miller JR, Bergman NH, Phillippy AM (2017) Canu: scalable and accurate long-read assembly via adaptive k-mer weighting and repeat separation. Genome Res 27:722\u0026ndash;736\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHu J, Wang Z, Sun Z, Hu B, Ayoola AO, Liang F, Li J, Sandoval JR, Cooper DN, Ye K (2024) NextDenovo: an efficient error correction and accurate assembly tool for noisy long reads. Genome Biol 25:107\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHu J, Fan J, Sun Z, Liu S (2020) NextPolish: a fast and efficient genome polishing tool for long-read assembly. Bioinformatics 36:2253\u0026ndash;2255\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang X, Wang L (2016) GMATA: an integrated software package for genome-scale SSR mining, marker development and viewing. Front Plant Sci 7:215951\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBenson G (1999) Tandem repeats finder: a program to analyze DNA sequences. Nucleic Acids Res 27:573\u0026ndash;580\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHan Y, Wessler SR (2010) MITE-Hunter: a program for discovering miniature inverted-repeat transposable elements from genomic sequences. Nucleic Acids Res 38:e199\u0026ndash;e199\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFlynn JM, Hubley R, Goubert C, Rosen J, Clark AG, Feschotte C, Smit AF (2020) RepeatModeler2 for automated genomic discovery of transposable element families. \u003cem\u003eProceedings of the National Academy of Sciences\u003c/em\u003e 117:9451\u0026ndash;9457\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen N (2004) Using Repeat Masker to identify repetitive elements in genomic sequences. \u003cem\u003eCurrent protocols in bioinformatics\u003c/em\u003e 5:4.10. 11-14.10. 14\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKeilwagen J, Hartung F, Grau J GeMoMa: homology-based gene prediction utilizing intron position conservation and RNA-seq data. Gene prediction: Methods protocols 2019:161\u0026ndash;177\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDobin A, Davis CA, Schlesinger F, Drenkow J, Zaleski C, Jha S, Batut P, Chaisson M, Gingeras TR (2013) STAR: ultrafast universal RNA-seq aligner. Bioinformatics 29:15\u0026ndash;21\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePertea M, Pertea GM, Antonescu CM, Chang T-C, Mendell JT, Salzberg SL (2015) StringTie enables improved reconstruction of a transcriptome from RNA-seq reads. Nat Biotechnol 33:290\u0026ndash;295\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHaas BJ, Salzberg SL, Zhu W, Pertea M, Allen JE, Orvis J, White O, Buell CR, Wortman JR (2008) Automated eukaryotic gene structure annotation using EVidenceModeler and the Program to Assemble Spliced Alignments. Genome Biol 9:1\u0026ndash;22\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eStanke M, Diekhans M, Baertsch R, Haussler D (2008) Using native and syntenically mapped cDNA alignments to improve de novo gene finding. Bioinformatics 24:637\u0026ndash;644\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKent WJ (2002) BLAT\u0026mdash;the BLAST-like alignment tool. Genome Res 12:656\u0026ndash;664\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChan PP, Lin BY, Mak AJ, Lowe TM (2021) tRNAscan-SE 2.0: improved detection and functional classification of transfer RNA genes. Nucleic Acids Res 49:9077\u0026ndash;9096\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNawrocki EP, Eddy SR (2013) Infernal 1.1: 100-fold faster RNA homology searches. Bioinformatics 29:2933\u0026ndash;2935\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLagesen K, Hallin P, R\u0026oslash;dland EA, St\u0026aelig;rfeldt H-H, Rognes T, Ussery DW (2007) RNAmmer: consistent and rapid annotation of ribosomal RNA genes. Nucleic Acids Res 35:3100\u0026ndash;3108\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLim K-B, De Jong H, Yang T-J, Park J-Y, Kwon S-J, Kim JS, Lim M-H, Kim JA, Jin M, Jin Y-M (2005) Characterization of rDNAs and tandem repeats in the heterochromatin of Brassica rapa. Mol Cells 19:436\u0026ndash;444\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLim KB, Yang TJ, Hwang YJ, Kim JS, Park JY, Kwon SJ, Kim J, Choi BS, Lim MH, Jin M (2007) Characterization of the centromere and peri-centromere retrotransposons in Brassica rapa and their distribution in related Brassica species. Plant J 49:173\u0026ndash;183\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCox AV, Bennett ST, Parokonny AS, Kenton A, Callimassia MA, Bennett MD (1993) Comparison of plant telomere locations using a PCR-generated synthetic probe. Ann Botany 72:239\u0026ndash;247\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKoo D-H, Hong CP, Batley J, Chung YS, Edwards D, Bang J-W, Hur Y, Lim YP (2011) Rapid divergence of repetitive DNAs in Brassica relatives. Genomics 97:173\u0026ndash;185\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBuchfink B, Reuter K, Drost H-G (2021) Sensitive protein alignments at tree-of-life scale using DIAMOND. Nat Methods 18:366\u0026ndash;368\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEdgar RC (2010) Search and clustering orders of magnitude faster than BLAST. Bioinformatics 26:2460\u0026ndash;2461\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTalavera G, Castresana J (2007) Improvement of phylogenies after removing divergent and ambiguously aligned blocks from protein sequence alignments. Syst Biol 56:564\u0026ndash;577\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eStamatakis A (2014) RAxML version 8: a tool for phylogenetic analysis and post-analysis of large phylogenies. Bioinformatics 30:1312\u0026ndash;1313\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSun P, Jiao B, Yang Y, Shan L, Li T, Li X, Xi Z, Wang X, Liu J (2022) WGDI: A user-friendly toolkit for evolutionary analyses of whole-genome duplications and ancestral karyotypes. Mol Plant 15:1841\u0026ndash;1851\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSuyama M, Torrents D, Bork P (2006) PAL2NAL: robust conversion of protein sequence alignments into the corresponding codon alignments. Nucleic Acids Res 34:W609\u0026ndash;W612\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang Z (2007) PAML 4: phylogenetic analysis by maximum likelihood. Mol Biol Evol 24:1586\u0026ndash;1591\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDelcher AL, Phillippy A, Carlton J, Salzberg SL (2002) Fast algorithms for large-scale genome alignment and comparison. Nucleic Acids Res 30:2478\u0026ndash;2483\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGoel M, Sun H, Jiao WB, Schneeberger K (2019) SyRI: finding genomic rearrangements and local sequence differences from whole-genome assemblies. Genome Biol 20:277\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGarrison E, Siren J, Novak AM, Hickey G, Eizenga JM, Dawson ET, Jones W, Garg S, Markello C, Lin MF, Paten B, Durbin R (2018) Variation graph toolkit improves read mapping by representing genetic variation in the reference. Nat Biotechnol 36:875\u0026ndash;879\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePurcell S, Neale B, Todd-Brown K, Thomas L, Ferreira MA, Bender D, Maller J, Sklar P, de Bakker PI, Daly MJ, Sham PC (2007) PLINK: a tool set for whole-genome association and population-based linkage analyses. Am J Hum Genet 81:559\u0026ndash;575\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"nature-portfolio","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"","title":"Nature Portfolio","twitterHandle":"","acdcEnabled":false,"dfaEnabled":false,"editorialSystem":"ejp","reportingPortfolio":"","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-6452497/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6452497/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e \u003cem\u003eBrassica napus\u003c/em\u003e is a globally significant polyploid oilseed crop with a complex genome, harbors substantial yet largely unexplored genetic diversity. In this study, we upgraded the genomes of the elite rapeseed cultivar Zhongshuang 11 (ZS11) and the spring cultivar Westar and their genome annotations. Additionally, we assembled the genomes of six other representative varieties, namely ZS2, Bugle, 352, 862, Ribenyoucai (RB), and Baihua (BH), and constructed a comprehensive pan-genome comprising 101,903 non-redundant pan-gene clusters derived from 22 genomes. Analysis of the pan-genome led to the identification of 124,682 transposable element- mediated structural variations (TEMSVs), which are unevenly distributed among subgenomes and have a significant impact on gene expression. Notably, this uneven distribution correlates with unbalanced expression of homoeologous gene pairs (uHGPs), with the degree of imbalance increasing alongside the number of TEMSVs. These uHGPs consistently exhibit biased expression patterns across different tissues. Furthermore, using structural variation-based genome-wide association studies (SV-GWAS), we identified two candidate genes, \u003cem\u003eBnaA06.PC78\u003c/em\u003e and \u003cem\u003eBnaC03.UCU1\u003c/em\u003e, which are associated with seed oil content and fatty acid composition. This study offers valuable insights and genetic resources for elucidating gene expression dynamics among subgenomes and advancing \u003cem\u003eB. napus\u003c/em\u003e breeding.\u003c/p\u003e","manuscriptTitle":"Transposable element-mediated structural variations drive gene expression and agronomic trait diversity in Brassica napus","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-04-29 07:43:54","doi":"10.21203/rs.3.rs-6452497/v1","editorialEvents":[],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"nature-communications","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"NCOMMS","sideBox":"Learn more about [Nature Communications](http://www.nature.com/ncomms/)","snPcode":"","submissionUrl":"https://mts-ncomms.nature.com/","title":"Nature Communications","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"Nature Communications","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"32d45144-cdf0-4d32-967f-e247d124fda1","owner":[],"postedDate":"April 29th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[{"id":47770269,"name":"Biological sciences/Plant sciences/Natural variation in plants"},{"id":47770270,"name":"Biological sciences/Genetics/Genome/Genetic variation"}],"tags":[],"updatedAt":"2026-03-24T05:10:59+00:00","versionOfRecord":[],"versionCreatedAt":"2025-04-29 07:43:54","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6452497","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6452497","identity":"rs-6452497","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-26T02:00:01.498150+00:00
License: CC-BY-4.0