A telomere-to-telomere genome assembly of Pythium species

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract Pythium spp. include numerous important pathogens that affect both plants and animals, posing serious threats to agricultural production and the health of humans and livestock. Currently, high-quality genomic resources for these species are scarce, impeding the understanding of population evolution and the development of prevention and control strategies. Here, the first complete telomere-to-telomere (T2T) genome of Pythium sp. was produced by integrating NGS, PacBio HiFi,ONT and Hi-C data. The genome size is 59.9 Mb in total, with a scaffold N50 length of 3.5 Mb. BUSCO assessment indicated that the genome assembly achieved 94.91% (n=255) completeness. A total of 14,111 protein-coding genes were predicted, with 12,711 genes (90.08% of the total) annotated using functional databases. Additionally, we predicted 382 effectors. This T2T genome is an excellent resource for furthering the understanding of the evolutionary path and pathogenesis mechanisms of Pythium spp.
Full text 90,247 characters · extracted from preprint-html · click to expand
A telomere-to-telomere genome assembly of Pythium species | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF data-descriptor A telomere-to-telomere genome assembly of Pythium species YUHANG ZHANG, YAXUAN GUO, SHUAI LI, CHAOSONG HUANG, YUHAN GAO, and 6 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8634265/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Pythium spp. include numerous important pathogens that affect both plants and animals, posing serious threats to agricultural production and the health of humans and livestock. Currently, high-quality genomic resources for these species are scarce, impeding the understanding of population evolution and the development of prevention and control strategies. Here, the first complete telomere-to-telomere (T2T) genome of Pythium sp. was produced by integrating NGS, PacBio HiFi,ONT and Hi-C data. The genome size is 59.9 Mb in total, with a scaffold N50 length of 3.5 Mb. BUSCO assessment indicated that the genome assembly achieved 94.91% (n=255) completeness. A total of 14,111 protein-coding genes were predicted, with 12,711 genes (90.08% of the total) annotated using functional databases. Additionally, we predicted 382 effectors. This T2T genome is an excellent resource for furthering the understanding of the evolutionary path and pathogenesis mechanisms of Pythium spp. Figures Figure 1 Figure 2 Figure 3 Figure 4 Background & Summary Pythium species belonging to the phylum Oomycota, which are transmitted through soil and water and have a broad host range worldwide, can infect important crops such as maize 1 (Fig. 1 a–d), soybean 2 , kidney bean 3 , wheat 4 , and tomato 5 at both the seedling and adult plant stages, causing symptoms such as stalk rot, root rot, and seedling blight and severely affecting crop yield. Furthermore, Pythium species can also cause infections in humans, horses, dogs, and other mammals, involving the eyes, arteries, skin, and gastrointestinal tract; in severe cases, it can even be life-threatening 6 , 7 . Therefore, elucidating the pathogenic mechanism of Pythium is highly important for agricultural production as well as human and livestock health. We isolated strain P16-1 from maize plants infected with stalk rot disease. P16-1 is strongly pathogenic to maize across multiple niches. Although several contig-level Pythium genomes have been assembled using Illumina sequencing 8 , many gaps remain, and the numbers of centromeres and telomeres have not been determined. A telomere-to-telomere (T2T) genome sequence that enables accurate characterization is highly important for understanding Pythium ’s evolutionary relationships and pathogenesis mechanisms and for developing control measures. To produce a gap-free reference genome, we sequenced the P16-1 genome to generate 4.37 Gb (~ 62× coverage) of raw MGI (MGI Tech Co., Ltd., Shenzhen) data. The raw data were filtered using fastp (v0.23.4) 9 , and k-mer analysis was performed (Fig. 1 e), which revealed a heterozygosity of 2.62%. We generated 4.5 Gb (~ 60× coverage) of high-fidelity (HiFi) and 75 Gb (~ 1071× coverage) of ONT (Oxford Nanopore Technologies, Oxford) raw data. Hifiasm (v0.24.0) 10 was used for hybrid assembly based on ONT data exceeding 50 kb and HiFi data, yielding the preliminarily assembled HiFiasmMix version of the genome. We also generated 17.9 Gb (~ 585× coverage) of Hi-C (high-throughput chromatin conformation capture) raw data, which was utilized for scaffolding based on chromatin interaction signals, resulting in a scaffolding rate of 99.26% (Fig. 1 f). Subsequently, gap closing and genome polishing were performed, ultimately resulting in a T2T genome with a size of 59.9 Mb and an N50 of 3.5 Mb (Fig. 1 g). We aligned this genome against the Nonredundant Protein Database (NR) and constructed a phylogenetic tree using high-quality genomes of Pythium species and their close relatives (Fig. 2). The results indicated that this genome represents a new species, which we propose to name Pythium maydis sp. nov. We next assessed the obtained genome. Using minimap2 (v2.26)¹¹ and PanDepth (https://github.com/HuiyangYu/PanDept h), we calculated the average coverage depth of regions within each bin. Using Merfin (v1.0)¹² and Merqury (v1.3)¹³, the genome was evaluated based on k-mers, yielding the consensus quality (QV) of the genome assembly. Using SeqKit (v2.5.1)¹⁴, statistical analysis was performed on telomeric intervals. To predict centromeric intervals (Fig. 3a), we used TRF (v4.09.1)¹⁵ to analyse tandem repeats in the genome, and RepeatMasker (Revision1.331)¹⁶ was used to predict repeat sequences in the genome. SRF (latest version)¹⁷ was subsequently employed to identify the units of the candidate centromeric repeat sequences. Finally, using centromeric interval information obtained from quarTeT (v1.1.6)¹⁸ as a reference, we integrated the above information to jointly determine the centromeric repeat units and intervals (Fig. 3a). We used BISER (v1.4) 19 to identify segmental duplication (SD) sequences in the genome (Fig. 3b) and minUniqueKmer (v1.0) 20 to calculate the length of the minimum unique k-mer (MUK) in the genome (Fig. 3c). RNAmmer (https://github.com/fujch7/rnamme r) was used to predict rRNA positions. HiFi and ONT data were aligned against the T2T genome using pbmm2 (https://github.com/PacificBiosciences/pbmm2 /) and modkit (https://github.com/nanoporetech/modki t), respectively. Methylation statistics were obtained by analysing the alignment results using pb-CpG-tools (https://github.com/PacificBiosciences/pb-CpG-tool s). Finally, the above information was visualized using karyoploteR (v4.0.4) (Fig. 3d). Figure 2 Phylogenetic tree construction of Pythium maydis with high-quality genome species of Pythium and their closely related species. To perform genome annotation, we generated 3.28 Gb of HiFi transcriptome data and 9.7 Gb of MGI transcriptome data. A total of 14,111 genes were predicted, with an average gene length of 1,967.03 bp and an average CDS length of 1,631.05 bp (Table 1 ). Statistical analysis of gene functional annotations from different databases revealed 12,711 genes annotated to functional databases (Fig. 4b), accounting for 90.08% of the total number of genes, with a transcriptome support rate of 90.06%. We also predicted genomic repeat sequences and noncoding RNAs (ncRNAs) (Table 1 ). Because Pythium maydis is a pathogen, we focused on predicting effectors Figure 3 Assessment of the Pythium maydis strain P16-1 genome. ( a ) Graphical representation of the positions of centromeres and telomeres on eighteen chromosomes. ( b ) Circos plot of the SD chromosomal distribution. ( c ) Statistical plot of MUK lengths in genomes and chromosomes. ( d ) Distribution map of various types of chromosomal information. and identified 382 effectors (Fig. 4a). We annotated the encoded protein sequences by alignment against the Pathogen–Host Interactions database (PHI) (Fig. 4c). Methods Mycelial culture Mycelia were cultured on potato dextrose agar (PDA) plates at 25°C for 7 days for mycelial growth. Cellophane sheets were laid on the plates to facilitate the collection of mycelia. Figure 4 Overview of the telomere-to-telomere genome and annotation of Pythium maydis strain P16-1. ( a ) Starting from the innermost circle moving outwards, ① chromosome ideogram, ② GC content, ③ gene density, ④ exon density, ⑤ TE (transposable element) density, ⑥ simple repeat density, ⑦ t-RNA density, ⑧ methylation density, and ⑨ effector protein gene density. ( b ) Overview of the number of genes in the P16-1 genome annotated by different databases. ( c ) Results of alignment and annotation of encoded protein sequences with the PHI database. DNA and RNA sequencing High-molecular-weight genomic DNA was extracted from Pythium maydis using the CTAB method. NGS Platform Sequencing: Libraries were constructed with an MGIEasy Universal DNA Library Prep Kit V1.0 (CAT#1000005250, MGI) following the standard protocol. The qualified libraries were subsequently sequenced on the MGI DNBSEQ-T7 platform. PacBio HiFi sequencing: Libraries were constructed using a SMRTbell® Prep Kit 3.0. The qualified libraries were subsequently sequenced on the PacBio Revio platform. Nanopore PromethION Ultralong Sequencing: DNA was extracted using a BAC-long DNA Kit. For the ultralong Nanopore library, genomic DNA was selected (> 50 kb) with the SageHLS HMW library system and then processed using the Ligation Sequencing 1D Kit. DNA libraries were constructed and sequenced on the PromethION platform. Hi-C: Library preparation was performed according to the established method 21 . The final library was subsequently sequenced on an Illumina NovaSeq X Plus/MIGI DNBSEQ-T7 platform. Table 1 Genome assembly and annotation statistics. Characteristics Pythium maydis Repetitive elements Number of elements Length of sequence (bp) Percentage of sequence (%) LINEs 5,282 2,251,583 3.76 SINEs 1,119 815,977 1.36 LTR 9,206 6,698,165 11.19 DNA 6,954 3,396,470 5.67 RC 382 220,031 0.37 MITE 43 8,195 0.01 Total TEs 22,986 13,390,421 22.36 Total TRs 10,654 1,076,610 1.8 Simple repeats 586 85,043 0.14 Low complexity 20 2,391 0 ncRNA Copy Number Total Length (bp) Percentage of sequence (%) rRNA 18S 49 89,154 0.1489 28S 50 221,644 0.3702 5.8S 0 0 0 5S 50 5,850 0.0098 sRNA snRNA 8 1412 0.0024 miRNA 0 0 0 spliceosomal 59 7415 0.0124 other 5 1517 0.0025 tRNA tRNA 1,487 111644 0.1865 Protein-coding genes number Average gene length (bp) Average CDS length (bp) BUSCO completeness (%) 14,111 1967.03 1631.05 94.9 Total RNA was extracted by grinding tissue in an RNeasy Mini Kit on dry ice and processed following the protocol provided by the manufacturer. NGS platform sequencing: Qualified RNA samples were reverse transcribed using an Iso-Seq Express 2.0 kit and purified with SMRTbell magnetic beads. After library construction, sequencing was conducted on the MGI DNBSEQ-T7 platform. PacBio Iso-Seq: Total RNA was used to construct a Kinnex full-length RNA library using an Iso-Seq express 2.0 kit (PacBio® 103-071-500) and a Kinnex full-length RNA kit (PacBio®103-072-000). The Kinnex full-length RNA library was subsequently sequenced on the PacBio Revio platform. Genome size and heterozygosity estimation To understand the genomic characteristics of Pythium maydis , k-mer analysis was performed using Illumina DNA data prior to genome assembly to estimate genome size and heterozygosity. Briefly, quality-filtered reads were subjected to 21-mer frequency distribution analysis using the KMC program. By analysing the 21-mer depth distribution from the 350-bp library-cleaned sequencing reads in GenomeScope software, we estimated the genome size of Pythium maydis with the following equation: G = K-num/K-depth (where K-num is the total number of 21-mers, K-depth denotes the k-mer depth, and G represents the genome size). Furthermore, GenomeScope was used to estimate the heterozygosity of the Pythium maydis genome. Genome assembly Based on the ONT data (filtered using seqkit (v2.5.1) with the parameters “seq --min-len 50000”, retaining ONT reads longer than 50 kb) and HiFi data, a preliminary assembled HiFiasmMix version of the genome was obtained through hybrid assembly using hifiasm (v0.24.0) (parameter: --ul ONT50kreads). All quality-controlled data were aligned to the reference genome using bwa (v0.7.17-r1198-dirty) 22 (parameters: mem − 5SP) and filtered using filter_bam (v2.0.0; https://github.com/genomewalker/bam-filter ), followed by analysis with HapHiC (v1.0.5)21 for clustering, reassigning, ordering and orientation, and building pseudomolecules. Prior to clustering, preprocessing was performed using HapHiC (v1.0.5) 23 to correct misassemblies. The contigs were subsequently clustered into groups using the Markov clustering algorithm. Reassignment was performed using HapHiC (v1.0.5) 23 . Following manual curation, a BUSCO assessment was conducted using Compleasm (v0.2.5) 24 . Nanopore Ultralong reads that were unmapped to the reference genome, had poor mapping quality, and were located at both ends of gaps were extracted. These reads were subsequently used for local iterative assembly to identify the optimal gap-filling path, generating sequences overlapping with both sides of the gap. These sequences were remapped to the gap-containing genome to replace the corresponding gap regions. HiFi reads unmapped to the genome, HiFi reads containing telomeric sequences, and reads mapped to chromosome ends without assembled telomeres were extracted for local assembly. The optimal gap-filling paths were identified to generate sequences overlapping the gap side, after which the corresponding gap regions were replaced. K-mer libraries of 21 k-mers and 31 k-mers were generated separately using yak (v1.0; https://github.com/lh3/yak ). HiFi reads were then mapped to the gap-filled genome using minimap2 (v2.26)¹¹ to generate BAM files. The genome was subsequently polished using Nextpolish2 (v0.2.0) 25 to obtain the final polished genome. Genome annotation Transcriptome sequences were mapped back to the genome using GMAP 26 to obtain transcript location information. Afterwards, the reassembled transcripts were realigned to the genome using PASA (v2.3.3) 27 , and optimal open reading frames (ORFs) were identified using GeneMark-ST (v5.1) 28 , resulting in the generation of transcriptome-based gene position predictions. Then, the protein sequences of the target species were obtained. These protein sequences were aligned back to the genome using GeMoMa (v1.6.1) 29 to determine the alignment position information of the corresponding proteins, thereby obtaining gene information annotated based on homologous proteins. Using gene information annotated from transcriptome sequences, gene models were predicted using GeneMark-ST (v5.1) 28 . The predicted genes were aligned against the Swiss-Prot Database 30 using BLASTP (v2.7.1). The alignment results were then filtered (identity ≥ 95%), and based on the filtered results, genes predicted by GeneMark-ST(v5.1) 28 were selected as the training set for AUGUSTUS (v3.3.1) 31 model training. Finally, genes in the genome were predicted using AUGUSTUS (v3.3.1) 31 with the predicted model. The results from de novo prediction, homology-based annotation, and transcriptome-based prediction were integrated using EVM (v1.1.1) 27 to obtain a nonredundant exon set, followed by the removal of invalid sequences using TransposonPSI 32 . The genomic protein sequences were aligned against the NR, Kyoto Encyclopedia of Genes and Genomes (KEGG) 3 ³, Eukaryotic Orthologous Groups of proteins (KOG)³ 4 , Gene Ontology (GO)³ 5 , and Swiss-Prot Database²⁹ using BLASTP (v2.7.1) and InterProScan (5.32-71.0) 36 . The annotated information of these protein sequences was categorized and classified. We used GMATA (v2.2) 37 and TRF (v4.07b) 15 to identify tandem repeats genome wide. For interspersed repeats, we first used MITE-hunter 38 to identify miniature inverted-repeat transposable elements (MITEs). The results from LTR_finder 39 and LTR_harvest 40 were subsequently used to construct an LTR sequence library with LTR_retriever 41 . After the two libraries were integrated, RepeatModeler (v1.0.11) 42 was used to perform de novo searches for repetitive sequences, which were then classified using TEclass 43 , after which the above data were integrated with Repbase 44 . Using the integrated data and RepeatMasker (v1.331) 42 , we searched for repetitive sequences genome wide. We used Infernal (v1.1.2) 45 to search against the Rfam 46 database for predicting rRNAs, snRNAs, and miRNAs. Additionally, tRNAscan-SE (v2.0) 47 was used to predict genomic tRNA sequences. Technical validation Compared with the assembled Pythium genomes by Christoffel et al. 8 , our genome assembly demonstrated enhanced continuity and completeness attributed to the utilization of long reads and Hi-C sequencing. This assembled genome has no gaps, and the telomeres and centromeres of each chromosome have been captured. Furthermore, the BUSCO assessment indicated an overall completeness of 94.9%, including 90.2% single-copy BUSCOs, 4.71% duplicated BUSCOs, 1.57% fragmented BUSCOs, and 3.53% missing BUSCOs. Declarations Competing interests The authors declare that they have no competing interests. Author Contribution C.D., Y.C. and Z.C. conceived and supervised the study. Y.Z. collected the samples and conducted the DNA extraction and sequencing. Z.C. and Y.Z. performed the bioinformatics analysis and prepared the figures and tables. C.D., Y.C., Z.C., Y.Z., Y.G., S.L., C.H., W.W., Z.Z. and S.S. wrote and revised the manuscript. All the authors read and approved the final manuscript. Acknowledgements This work was supported by grants from the National Key Research & Development Program of China (2021YFD1200700), the Program of Discovery and Application of Resistance Genes to Stalk Rot and Ear Rot in Maize (BLCZK20240920), and the Agricultural Science and Technology Innovation Program of the Chinese Academy of Agricultural Sciences (01-ICS-02). Correspondence and requests for materials should be addressed to C.D., Y.C. and Z.C. Data Availability The telomere-to-telomere (T2T) genome assembly of *Pythium maydis* . has been deposited at the National Center for Biotechnology Information (NCBI) GenBank under the accession number GCA_054083355.1 48 . The raw sequencing data, including Nanopore Ultralong reads, PacBio HiFi reads, MGI short reads, Hi-C data, and RNA-Seq data, have been deposited in the NCBI Sequence Read Archive (SRA) under the accession numbers SRR37082792- SRR37082796 49 . All datasets are publicly available and can be accessed via the NCBI database. Code availability All the software used in this study is in the public domain, with the parameters clearly described in the Methods section and this section. If no detailed parameters were mentioned for the software, default parameters were used as suggested by the developer. References Duan C., Song F., Sun S., Guo C., Zhu Z., Wang X. Characterization and molecular mapping of two novel genes resistant to Pythium stalk rot in maize. Phytopathology . 109, 804–809 (2019). Broders KD., Lipps PE., Paul PA., Dorrance AE. Characterization of Pythium spp. associated with corn and soybean seed and seedling disease in ohio. Plant Dis . 91, 727–735 (2007). Rossman DR., Rojas A., Jacobs JL., Mukankusi C., Kelly JD., Chilvers MI. Pathogenicity and virulence of soilborne oomycetes on phaseolus vulgaris. Plant Dis . 101, 1851–1859 (2017). Reeves ER., Kerns JP, Cowger C, Shew BB. Pythium spp. Associated with root rot and stunting of winter wheat in north carolina. Plant Dis . 105, 986–996 (2021). Ma M., Taylor PWJ., Chen D., Vaghefi N., He JZ. Major soilborne pathogens of field processing tomatoes and management strategies. Microorganisms . 11, 263 (2023). Yolanda H., Krajaejun T. Review of methods and antimicrobial agents for susceptibility testing against Pythium insidiosum . Heliyon. 6, e03737 (2020). Gurnani B., Kaur K., Venugopal A., Srinivasan B., Bagga B., Iyer G., Christy J., Prajna L., Vanathi M., Garg P., Narayana S., Agarwal S., Sahu S. Pythium insidiosum keratitis - A review. Indian J Ophthalmol . 70, 1107–1120 (2022). Nguyen HDT., Dodge A., Dadej K., Rintoul TL., Ponomareva E., Martin FN., de Cock AWAM., Lévesque CA., Redhead SA., Spies CFJ. Whole genome sequencing and phylogenomic analysis show support for the splitting of genus Pythium . Mycologia . 114, 501–515 (2022). Chen S., Zhou Y., Chen Y., Gu J. fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics , 34, p. i884-i890 (2018). Cheng H., Jarvis ED., Fedrigo O., Koepfli KP., Urban L., Gemmell NJ., Li H. Haplotype-resolved assembly of diploid genomes without parental data. Nat Biotechnol . 40, 1332–1335 (2022). Li, H. New strategies to improve minimap2 alignment accuracy. Bioinformatics. 37, 4572–4574 (2021). Formenti G., Rhie A., Walenz BP., Thibaud-Nissen F., Shafin K., Koren S., Myers EW., Jarvis ED., Phillippy AM. Merfin: improved variant filtering, assembly evaluation and polishing via k-mer validation. Nat Methods . 19, 696–704 (2022). Rhie A., Walenz BP., Koren S., Phillippy AM. Merqury: reference-free quality, completeness, and phasing assessment for genome assemblies. Genome Biol . 21, 245 (2020). Shen W., Le S., Li Y., Hu F. SeqKit: A cross-platform and ultrafast toolkit for FASTA/Q file manipulation. PLoS ONE . 11, e0163962 (2016). Benson G. Tandem repeats finder: a program to analyze DNA sequences. Nucleic Acids Research. 27, 573–580 (1999). Bedell JA., Korf I., Gish W. MaskerAid: a performance enhancement to RepeatMasker. Bioinformatics. 16, 1040–1 (2000). Zhang Y., Chu J., Cheng H., Li H. De novo reconstruction of satellite repeat units from sequence data. Genome Res . 33, 1994–2001 (2023). Lin Y., Ye C., Li X., Chen Q., Wu Y., Zhang F., Pan R., Zhang S., Chen S., Wang X., Cao S., Wang Y., Yue Y., Liu Y., Yue J. quarTeT: a telomere-to-telomere toolkit for gap-free genome assembly and centromeric repeat identification. Horticulture Research . 10, uhad127 (2023). Išerić H., Alkan C., Hach F., Numanagić I. Fast characterization of segmental duplication structure in multiple genome assemblies. Algorithms Mol Biol. 17, 4 (2022). Kirsche M., Das A., Schatz MC. Sapling: accelerating suffix array queries with learned data models, Bioinformatics , 37, 744–749 (2021). Belton JM., McCord RP., Gibcus JH., Naumova N., Zhan Y., Dekker J. Hi-C: A comprehensive technique to capture the conformation of genomes. Methods . 58, 268–76 (2012). Li H. Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM. Genomics . 1303.3997v2. (2013) Zeng X., Yi Z., Zhang X., Du Y., Li Y., Zhou Z., Chen S., Zhao H., Yang S., Wang Y., Chen G. Chromosome-level scaffolding of haplotype-resolved assemblies using Hi-C data without reference genomes. Nature Plants . 10, 1184–1200 (2024). Huang N., Li H. compleasm: a faster and more accurate reimplementation of BUSCO. Bioinformatics , 39, btad595 (2023). Hu J., Wang Z., Liang F., Liu SL., Ye K., Wang DP. NextPolish2: A repeat-aware polishing tool for genomes assembled using HiFi long reads. Genomics, Proteomics & Bioinformatics . 22, qzad009 (2024). Wu TD., Watanabe CK. GMAP: a genomic mapping and alignment program for mRNA and EST sequences. Bioinformatics. Bioinformatics. 21, 1859–75 (2005). Haas BJ., Salzberg SL., Zhu W., Pertea M., Allen JE., Orvis J., White O., Buell CR., Wortman JR. Automated eukaryotic gene structure annotation using EVidenceModeler and the program to assemble spliced alignments. Genome Biol . 9, R7 (2008). Tang S., Lomsadze A., Borodovsky M. Identification of protein coding regions in RNA transcripts. Nucleic Acids Res . 43, e78 (2015). Keilwagen J., Wenk M., Erickson JL., Schattat MH., Grau J., Hartung F. Using intron position conservation for homology-based gene prediction. Nucleic Acids Res . 44, e89 (2016). Bairoch A., Apweiler R., Wu CH., Barker WC., Boeckmann B., Ferro S., Gasteiger E., Huang H., Lopez R., Magrane M., Martin MJ., Natale DA., O'Donovan C., Redaschi N., Yeh LS. The universal protein resource (UniProt). Nucleic Acids Res . 33, D154-9 (2005). Stanke M., Diekhans M., Baertsch R., Haussler D. Using native and syntenically mapped cDNA alignments to improve de novo gene finding. Bioinformatics. 24, 637–44 (2008). Urasaki N., Takagi H., Natsume S., Uemura A., Taniai N., Miyagi N., Fukushima M., Suzuki S., Tarora K., Tamaki M., Sakamoto M., Terauchi R., Matsumura H. Draft genome sequence of bitter gourd ( Momordica charantia ), a vegetable and medicinal plant in tropical and subtropical regions. DNA Res . 24, 51–58 (2017). Ogata H., Goto S., Sato K., Fujibuchi W., Bono H., Kanehisa M. KEGG: kyoto encyclopedia of genes and genomes. Nucleic Acids Res . 28, 27–30 (2000). Galperin MY., Makarova KS., Wolf YI., Koonin EV. Expanded microbial genome coverage and improved protein family annotation in the COG database. Nucleic Acids Res . 43, D261-9 (2015). Ashburner M., Ball CA., Blake JA., Botstein D., Butler H., Cherry JM., Davis AP., Dolinski K., Dwight SS., Eppig JT., Harris MA., Hill DP., Issel-Tarver L., Kasarskis A., Lewis S., Matese JC., Richardson JE., Ringwald M., Rubin GM., Sherlock G. Gene ontology: tool for the unification of biology. The Gene Ontology Consortium. Nat Genet . 25, 25–9 (2000). Zdobnov EM., Apweiler R. InterProScan–an integration platform for the signature-recognition methods in InterPro. Bioinformatics. 17, 847–8 (2001). Wang, X., L. Wang. GMATA: An integrated software package for genome-scale SSR Mining, Marker Development and Viewing. Front Plant Sci . 7, 1350 (2016). Han, Y., S.R. Wessler, MITE-Hunter: a program for discovering miniature inverted-repeat transposable elements from genomic sequences. Nucleic Acids Res . 38, 199 (2010). Xu, Z., H. Wang. LTR_FINDER: an efficient tool for the prediction of full-length LTR retrotransposons. Nucleic Acids Res . 35, 265–8 (2007). Ellinghaus D., Kurtz S., Willhoeft U. LTRharvest, an efficient and flexible software for de novo detection of LTR retrotransposons. BMC Bioinformatics . 9, 18 (2008). Ou S., Jiang N. LTR_retriever: A highly accurate and sensitive program for identification of long terminal repeat retrotransposons. Plant Physiol . 176, 1410–1422 (2018). Bedell JA., Korf I., Gish W. MaskerAid: a performance enhancement to RepeatMasker. Bioinformatics. 16, 1040–1 (2000). Abrusán G., Grundmann N., DeMester L., Makalowski W. TEclass–a tool for automated classification of unknown eukaryotic transposable elements. Bioinformatics. 25, 1329–30 (2009). Jurka J., Kapitonov VV., Pavlicek A., Klonowski P., Kohany O., Walichiewicz J. Repbase Update, a database of eukaryotic repetitive elements. Cytogenet Genome Res . 110, 462–7 (2005). Nawrocki EP., Eddy SR. Eddy. Infernal 1.1: 100-fold faster RNA homology searches. Bioinformatics . 29, 2933–5 (2013). Griffiths-Jones S., Moxon S., Marshall M., Khanna A., Eddy SR., Bateman A. Rfam: annotating non-coding RNAs in complete genomes. Nucleic Acids Res . 33, D121-4 (2005). Lowe TM., Eddy SR. tRNAscan-SE: a program for improved detection of transfer RNA genes in genomic sequence. Nucleic Acids Res . 25, 955–64 (1997). NCBI GenBank https://www.ncbi.nlm.nih.gov/assembly/GCA_054083355.1 (2025). NCBI Sequence Read Archive https://www.ncbi.nlm.nih.gov /sra?linkname=bioproject_sra_all&from_uid=1358918 (2026). Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8634265","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"data-descriptor","associatedPublications":[],"authors":[{"id":600085692,"identity":"d3f1dbad-b587-47ed-b110-0f6dcaa21ac2","order_by":0,"name":"YUHANG ZHANG","email":"","orcid":"","institution":"Chinese Academy of Agricultural Sciences","correspondingAuthor":false,"prefix":"","firstName":"YUHANG","middleName":"","lastName":"ZHANG","suffix":""},{"id":600085693,"identity":"ed2d707f-3f93-4f7b-b430-d8e99bb6309d","order_by":1,"name":"YAXUAN GUO","email":"","orcid":"","institution":"Chinese Academy of Agricultural Sciences","correspondingAuthor":false,"prefix":"","firstName":"YAXUAN","middleName":"","lastName":"GUO","suffix":""},{"id":600085694,"identity":"b26397d0-85f8-4482-b9d2-c73a1bcc7d5c","order_by":2,"name":"SHUAI LI","email":"","orcid":"","institution":"Chinese Academy of Agricultural Sciences","correspondingAuthor":false,"prefix":"","firstName":"SHUAI","middleName":"","lastName":"LI","suffix":""},{"id":600085695,"identity":"c9b8b142-c5a2-4cdd-ada6-a62a7d945104","order_by":3,"name":"CHAOSONG HUANG","email":"","orcid":"","institution":"Chinese Academy of Agricultural Sciences","correspondingAuthor":false,"prefix":"","firstName":"CHAOSONG","middleName":"","lastName":"HUANG","suffix":""},{"id":600085696,"identity":"7333296f-feb2-4682-836e-ad032711d613","order_by":4,"name":"YUHAN GAO","email":"","orcid":"","institution":"Chinese Academy of Agricultural Sciences","correspondingAuthor":false,"prefix":"","firstName":"YUHAN","middleName":"","lastName":"GAO","suffix":""},{"id":600085697,"identity":"0e48126f-0b76-4cab-80ee-f6a5359a5079","order_by":5,"name":"WENQI WU","email":"","orcid":"","institution":"Chinese Academy of Agricultural Sciences","correspondingAuthor":false,"prefix":"","firstName":"WENQI","middleName":"","lastName":"WU","suffix":""},{"id":600085698,"identity":"5788fec5-2b96-4b38-9ad1-b4f5a280b67d","order_by":6,"name":"SULI SUN","email":"","orcid":"","institution":"Chinese Academy of Agricultural Sciences","correspondingAuthor":false,"prefix":"","firstName":"SULI","middleName":"","lastName":"SUN","suffix":""},{"id":600085699,"identity":"237bb9fd-c78f-47e2-a2ad-de5685570638","order_by":7,"name":"ZHENDONG ZHU","email":"","orcid":"","institution":"Chinese Academy of Agricultural Sciences","correspondingAuthor":false,"prefix":"","firstName":"ZHENDONG","middleName":"","lastName":"ZHU","suffix":""},{"id":600085700,"identity":"3a4f47e4-b20e-481e-b332-0676d9d1614f","order_by":8,"name":"YANYONG CAO","email":"","orcid":"","institution":"Henan Academy of Agricultural Sciences","correspondingAuthor":false,"prefix":"","firstName":"YANYONG","middleName":"","lastName":"CAO","suffix":""},{"id":600085701,"identity":"55f472d2-8a40-44cb-b108-b22ec26d013e","order_by":9,"name":"ZIXIANG CHENG","email":"","orcid":"","institution":"Chinese Academy of Agricultural Sciences","correspondingAuthor":false,"prefix":"","firstName":"ZIXIANG","middleName":"","lastName":"CHENG","suffix":""},{"id":600085702,"identity":"7ed8c276-30a6-436f-9ca1-25ee8392411e","order_by":10,"name":"CANXING DUAN","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA1klEQVRIiWNgGAWjYDCCA1CajYGB8QHJWpgNSNMC0iVBlA6+G8nPHvyouZPYJ91+reJHjZ09A/vZA3i1SN5IMzfsOfbMmE3mTNnNnmPJiQ08eQl4tRjcSDCTZmA7LMcmkZN2m4GNOYFBgge/nwxupH+TZvh3mAekpZjhX709EVpyzKQZ20C2pB9jBjIYGwhpkTzzpkyyt++wMdAWZiDjeGIbTw5+LXzH07dJ/Ph2OHH+jPSHH358q7bnZz9DZPwwMEDdw0aseiBgf0CC4lEwCkbBKBhJAACxv0KRsWTSsgAAAABJRU5ErkJggg==","orcid":"","institution":"Chinese Academy of Agricultural Sciences","correspondingAuthor":true,"prefix":"","firstName":"CANXING","middleName":"","lastName":"DUAN","suffix":""}],"badges":[],"createdAt":"2026-01-19 02:23:16","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8634265/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8634265/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":103985886,"identity":"a3eff099-ff9f-4ea4-a630-ddd16a1e986c","added_by":"auto","created_at":"2026-03-05 10:27:04","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":1144587,"visible":true,"origin":"","legend":"\u003cp\u003ePhenotype of maize infected by P16-1 and its telomere-to-telomere genome assembly. (\u003cstrong\u003ea\u003c/strong\u003e) Field disease phenotype of maize infected by P16-1. (\u003cstrong\u003eb\u003c/strong\u003e) Internal stalk phenotype of healthy and P16-1-infected maize plants. (\u003cstrong\u003ec\u003c/strong\u003e) Internal stalk phenotype of maize artificially inoculated with P16-1. (\u003cstrong\u003ed\u003c/strong\u003e) Colony morphology after 7 days of incubation at 25°C. (\u003cstrong\u003ee\u003c/strong\u003e) A genome survey of P16-1 was performed using GenomeScope. (\u003cstrong\u003ef\u003c/strong\u003e) Contact map. Chromosome-level heatmap of P16-1. (\u003cstrong\u003eg\u003c/strong\u003e) T2T genome snail plot.\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-8634265/v1/3ea7b4a3da2b54d46d582af3.png"},{"id":103985890,"identity":"3ddd77a9-54fc-42c7-9118-acc9d9df45aa","added_by":"auto","created_at":"2026-03-05 10:27:07","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":191750,"visible":true,"origin":"","legend":"\u003cp\u003ePhylogenetic tree construction of \u003cem\u003ePythium maydis\u003c/em\u003e with high-quality genome species of \u003cem\u003ePythium\u003c/em\u003e and their closely related species.\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-8634265/v1/a47b843f50c083fcf4ed21a3.png"},{"id":103985888,"identity":"612a95d3-ce55-4e4f-8eec-1abeec4de0ad","added_by":"auto","created_at":"2026-03-05 10:27:05","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":372388,"visible":true,"origin":"","legend":"\u003cp\u003eAssessment of the \u003cem\u003ePythium maydis \u003c/em\u003estrain P16-1\u003cem\u003e \u003c/em\u003egenome. (\u003cstrong\u003ea\u003c/strong\u003e) Graphical representation of the positions of centromeres and telomeres on eighteen chromosomes. (\u003cstrong\u003eb\u003c/strong\u003e) Circos plot of the SD chromosomal distribution. (\u003cstrong\u003ec\u003c/strong\u003e) Statistical plot of MUK lengths in genomes and chromosomes. (\u003cstrong\u003ed\u003c/strong\u003e) Distribution map of various types of chromosomal information.\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-8634265/v1/7bfd4df5f5371d00312a0188.png"},{"id":103985889,"identity":"866423e4-a372-487b-8586-9d421e668d17","added_by":"auto","created_at":"2026-03-05 10:27:05","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":383325,"visible":true,"origin":"","legend":"\u003cp\u003eOverview of the telomere-to-telomere genome and annotation of \u003cem\u003ePythium maydis \u003c/em\u003estrain P16-1. (\u003cstrong\u003ea\u003c/strong\u003e) Starting from the innermost circle moving outwards, ① chromosome ideogram, ② GC content, ③ gene density, ④ exon density, ⑤ TE (transposable element) density, ⑥ simple repeat density, ⑦ t-RNA density, ⑧ methylation density, and ⑨ effector protein gene density. (\u003cstrong\u003eb\u003c/strong\u003e) Overview of the number of genes in the P16-1 genome annotated by different databases. (\u003cstrong\u003ec\u003c/strong\u003e) Results of alignment and annotation of encoded protein sequences with the PHI database.\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-8634265/v1/99171f5f9b327d895690e44e.png"},{"id":103985910,"identity":"d7ef0b35-a7e0-40ab-aea1-04184bd0041c","added_by":"auto","created_at":"2026-03-05 10:27:16","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2582640,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8634265/v1/124de4e4-57dd-49a2-804a-e01abd658058.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"A telomere-to-telomere genome assembly of Pythium species","fulltext":[{"header":"Background \u0026 Summary","content":"\u003cp\u003e \u003cem\u003ePythium\u003c/em\u003e species belonging to the phylum Oomycota, which are transmitted through soil and water and have a broad host range worldwide, can infect important crops such as maize\u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u003c/sup\u003e (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ea\u0026ndash;d), soybean\u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e, kidney bean\u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e, wheat\u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e, and tomato\u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e at both the seedling and adult plant stages, causing symptoms such as stalk rot, root rot, and seedling blight and severely affecting crop yield. Furthermore, \u003cem\u003ePythium\u003c/em\u003e species can also cause infections in humans, horses, dogs, and other mammals, involving the eyes, arteries, skin, and gastrointestinal tract; in severe cases, it can even be life-threatening\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e,\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e. Therefore, elucidating the pathogenic mechanism of \u003cem\u003ePythium\u003c/em\u003e is highly important for agricultural production as well as human and livestock health.\u003c/p\u003e \u003cp\u003eWe isolated strain P16-1 from maize plants infected with stalk rot disease. P16-1 is strongly pathogenic to maize across multiple niches. Although several contig-level \u003cem\u003ePythium\u003c/em\u003e genomes have been assembled using Illumina sequencing\u003csup\u003e\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e, many gaps remain, and the numbers of centromeres and telomeres have not been determined. A telomere-to-telomere (T2T) genome sequence that enables accurate characterization is highly important for understanding \u003cem\u003ePythium\u003c/em\u003e\u0026rsquo;s evolutionary relationships and pathogenesis mechanisms and for developing control measures.\u003c/p\u003e \u003cp\u003eTo produce a gap-free reference genome, we sequenced the P16-1 genome to generate 4.37 Gb (~\u0026thinsp;62\u0026times; coverage) of raw MGI (MGI Tech Co., Ltd., Shenzhen) data. The raw data were filtered using fastp (v0.23.4)\u003csup\u003e9\u003c/sup\u003e, and k-mer analysis was performed (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ee), which revealed a heterozygosity of 2.62%. We generated 4.5 Gb (~\u0026thinsp;60\u0026times; coverage) of high-fidelity (HiFi) and 75 Gb (~\u0026thinsp;1071\u0026times; coverage) of ONT (Oxford Nanopore Technologies, Oxford) raw data. Hifiasm (v0.24.0)\u003csup\u003e10\u003c/sup\u003e was used for hybrid assembly based on ONT data exceeding 50 kb and HiFi data, yielding the preliminarily assembled HiFiasmMix version of the genome. We also generated 17.9 Gb (~\u0026thinsp;585\u0026times; coverage) of Hi-C (high-throughput chromatin conformation capture) raw data, which was utilized for scaffolding based on chromatin interaction signals, resulting in a scaffolding rate of 99.26% (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ef). Subsequently, gap closing and genome polishing were performed, ultimately resulting in a T2T genome with a size of 59.9 Mb and an N50 of 3.5 Mb (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eg). We aligned this genome against the Nonredundant Protein Database (NR) and constructed a phylogenetic tree using high-quality genomes of \u003cem\u003ePythium\u003c/em\u003e species and their close relatives (Fig.\u0026nbsp;2). The results indicated that this genome represents a new species, which we propose to name \u003cem\u003ePythium maydis\u003c/em\u003e sp. nov.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e We next assessed the obtained genome. Using minimap2 (v2.26)\u0026sup1;\u0026sup1; and PanDepth \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e(https://github.com/HuiyangYu/PanDept\u003c/span\u003e\u003cspan address=\"http://(https://github.com/HuiyangYu/PanDept\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003eh), we calculated the average coverage depth of regions within each bin. Using Merfin (v1.0)\u0026sup1;\u0026sup2; and Merqury (v1.3)\u0026sup1;\u0026sup3;, the genome was evaluated based on k-mers, yielding the consensus quality (QV) of the genome assembly. Using SeqKit (v2.5.1)\u0026sup1;⁴, statistical analysis was performed on telomeric intervals. To predict centromeric intervals (Fig.\u0026nbsp;3a), we used TRF (v4.09.1)\u0026sup1;⁵ to analyse tandem repeats in the genome, and RepeatMasker (Revision1.331)\u0026sup1;⁶ was used to predict repeat sequences in the genome. SRF (latest version)\u0026sup1;⁷ was subsequently employed to identify the units of the candidate centromeric repeat sequences. Finally, using centromeric interval information obtained from quarTeT (v1.1.6)\u0026sup1;⁸ as a reference, we integrated the above information to jointly determine the centromeric repeat units and intervals (Fig.\u0026nbsp;3a). We used BISER (v1.4)\u003csup\u003e19\u003c/sup\u003e to identify segmental duplication (SD) sequences in the genome (Fig.\u0026nbsp;3b) and minUniqueKmer (v1.0)\u003csup\u003e20\u003c/sup\u003e to calculate the length of the minimum unique k-mer (MUK) in the genome (Fig.\u0026nbsp;3c). RNAmmer \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e(https://github.com/fujch7/rnamme\u003c/span\u003e\u003cspan address=\"http://(https://github.com/fujch7/rnamme\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003er) was used to predict rRNA positions. HiFi and ONT data were aligned against the T2T genome using pbmm2 \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e(https://github.com/PacificBiosciences/pbmm2\u003c/span\u003e\u003cspan address=\"http://(https://github.com/PacificBiosciences/pbmm2\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e/) and modkit \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e(https://github.com/nanoporetech/modki\u003c/span\u003e\u003cspan address=\"http://(https://github.com/nanoporetech/modki\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003et), respectively. Methylation statistics were obtained by analysing the alignment results using pb-CpG-tools \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e(https://github.com/PacificBiosciences/pb-CpG-tool\u003c/span\u003e\u003cspan address=\"http://(https://github.com/PacificBiosciences/pb-CpG-tool\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003es). Finally, the above information was visualized using karyoploteR (v4.0.4) (Fig.\u0026nbsp;3d).\u003c/p\u003e \u003cp\u003eFigure\u0026nbsp;2 Phylogenetic tree construction of \u003cem\u003ePythium maydis\u003c/em\u003e with high-quality genome species of \u003cem\u003ePythium\u003c/em\u003e and their closely related species.\u003c/p\u003e \u003cp\u003e To perform genome annotation, we generated 3.28 Gb of HiFi transcriptome data and 9.7 Gb of MGI transcriptome data. A total of 14,111 genes were predicted, with an average gene length of 1,967.03 bp and an average CDS length of 1,631.05 bp (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). Statistical analysis of gene functional annotations from different databases revealed 12,711 genes annotated to functional databases (Fig.\u0026nbsp;4b), accounting for 90.08% of the total number of genes, with a transcriptome support rate of 90.06%. We also predicted genomic repeat sequences and noncoding RNAs (ncRNAs) (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). Because \u003cem\u003ePythium maydis\u003c/em\u003e is a pathogen, we focused on predicting effectors\u003c/p\u003e \u003cp\u003eFigure\u0026nbsp;3 Assessment of the \u003cem\u003ePythium maydis\u003c/em\u003e strain P16-1 genome. (\u003cb\u003ea\u003c/b\u003e) Graphical representation of the positions of centromeres and telomeres on eighteen chromosomes. (\u003cb\u003eb\u003c/b\u003e) Circos plot of the SD chromosomal distribution. (\u003cb\u003ec\u003c/b\u003e) Statistical plot of MUK lengths in genomes and chromosomes. (\u003cb\u003ed\u003c/b\u003e) Distribution map of various types of chromosomal information.\u003c/p\u003e \u003cp\u003eand identified 382 effectors (Fig.\u0026nbsp;4a). We annotated the encoded protein sequences by alignment against the Pathogen\u0026ndash;Host Interactions database (PHI) (Fig.\u0026nbsp;4c).\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eMycelial culture\u003c/h2\u003e \u003cp\u003e Mycelia were cultured on potato dextrose agar (PDA) plates at 25\u0026deg;C for 7 days for mycelial growth. Cellophane sheets were laid on the plates to facilitate the collection of mycelia.\u003c/p\u003e \u003cp\u003eFigure\u0026nbsp;4 Overview of the telomere-to-telomere genome and annotation of \u003cem\u003ePythium maydis\u003c/em\u003e strain P16-1. (\u003cb\u003ea\u003c/b\u003e) Starting from the innermost circle moving outwards, ① chromosome ideogram, ② GC content, ③ gene density, ④ exon density, ⑤ TE (transposable element) density, ⑥ simple repeat density, ⑦ t-RNA density, ⑧ methylation density, and ⑨ effector protein gene density. (\u003cb\u003eb\u003c/b\u003e) Overview of the number of genes in the P16-1 genome annotated by different databases. (\u003cb\u003ec\u003c/b\u003e) Results of alignment and annotation of encoded protein sequences with the PHI database.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eDNA and RNA sequencing\u003c/h3\u003e\n\u003cp\u003eHigh-molecular-weight genomic DNA was extracted from \u003cem\u003ePythium maydis\u003c/em\u003e using the CTAB method. NGS Platform Sequencing: Libraries were constructed with an MGIEasy Universal DNA Library Prep Kit V1.0 (CAT#1000005250, MGI) following the standard protocol. The qualified libraries were subsequently sequenced on the MGI DNBSEQ-T7 platform. PacBio HiFi sequencing: Libraries were constructed using a SMRTbell\u0026reg; Prep Kit 3.0. The qualified libraries were subsequently sequenced on the PacBio Revio platform. Nanopore PromethION Ultralong Sequencing: DNA was extracted using a BAC-long DNA Kit. For the ultralong Nanopore library, genomic DNA was selected (\u0026gt;\u0026thinsp;50 kb) with the SageHLS HMW library system and then processed using the Ligation Sequencing 1D Kit. DNA libraries were constructed and sequenced on the PromethION platform. Hi-C: Library preparation was performed according to the established method\u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e. The final library was subsequently sequenced on an Illumina NovaSeq X Plus/MIGI DNBSEQ-T7 platform.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eGenome assembly and annotation statistics.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colspan=\"3\" nameend=\"c3\" namest=\"c1\"\u003e \u003cp\u003eCharacteristics\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"3\" nameend=\"c6\" namest=\"c4\"\u003e \u003cp\u003e\u003cem\u003ePythium maydis\u003c/em\u003e\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"10\" rowspan=\"11\"\u003e \u003cp\u003eRepetitive elements\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eNumber of elements\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eLength of sequence (bp)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003ePercentage of sequence (%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eLINEs\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e5,282\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e2,251,583\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e3.76\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eSINEs\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1,119\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e815,977\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e1.36\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eLTR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e9,206\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e6,698,165\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e11.19\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eDNA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e6,954\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e3,396,470\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e5.67\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eRC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e382\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e220,031\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.37\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eMITE\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e43\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e8,195\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.01\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eTotal TEs\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e22,986\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e13,390,421\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e22.36\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eTotal TRs\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e10,654\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1,076,610\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e1.8\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eSimple repeats\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e586\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e85,043\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.14\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eLow complexity\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e20\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e2,391\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"9\" rowspan=\"10\"\u003e \u003cp\u003encRNA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eCopy Number\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eTotal Length (bp)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003ePercentage of sequence (%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"3\" rowspan=\"4\"\u003e \u003cp\u003erRNA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e18S\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e49\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e89,154\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.1489\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e28S\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e50\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e221,644\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.3702\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e5.8S\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e5S\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e50\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e5,850\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.0098\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"3\" rowspan=\"4\"\u003e \u003cp\u003esRNA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003esnRNA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1412\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.0024\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003emiRNA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003espliceosomal\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e59\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e7415\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.0124\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eother\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1517\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.0025\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003etRNA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003etRNA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1,487\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e111644\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.1865\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eProtein-coding genes\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003enumber\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eAverage gene length (bp)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eAverage CDS length (bp)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eBUSCO completeness (%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003e14,111\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1967.03\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1631.05\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e94.9\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eTotal RNA was extracted by grinding tissue in an RNeasy Mini Kit on dry ice and processed following the protocol provided by the manufacturer. NGS platform sequencing: Qualified RNA samples were reverse transcribed using an Iso-Seq Express 2.0 kit and purified with SMRTbell magnetic beads. After library construction, sequencing was conducted on the MGI DNBSEQ-T7 platform. PacBio Iso-Seq: Total RNA was used to construct a Kinnex full-length RNA library using an Iso-Seq express 2.0 kit (PacBio\u0026reg; 103-071-500) and a Kinnex full-length RNA kit (PacBio\u0026reg;103-072-000). The Kinnex full-length RNA library was subsequently sequenced on the PacBio Revio platform.\u003c/p\u003e\n\u003ch3\u003eGenome size and heterozygosity estimation\u003c/h3\u003e\n\u003cp\u003eTo understand the genomic characteristics of \u003cem\u003ePythium maydis\u003c/em\u003e, k-mer analysis was performed using Illumina DNA data prior to genome assembly to estimate genome size and heterozygosity. Briefly, quality-filtered reads were subjected to 21-mer frequency distribution analysis using the KMC program. By analysing the 21-mer depth distribution from the 350-bp library-cleaned sequencing reads in GenomeScope software, we estimated the genome size of \u003cem\u003ePythium maydis\u003c/em\u003e with the following equation: G\u0026thinsp;=\u0026thinsp;K-num/K-depth (where K-num is the total number of 21-mers, K-depth denotes the k-mer depth, and G represents the genome size). Furthermore, GenomeScope was used to estimate the heterozygosity of the \u003cem\u003ePythium maydis\u003c/em\u003e genome.\u003c/p\u003e\n\u003ch3\u003eGenome assembly\u003c/h3\u003e\n\u003cp\u003eBased on the ONT data (filtered using seqkit (v2.5.1) with the parameters \u0026ldquo;seq --min-len 50000\u0026rdquo;, retaining ONT reads longer than 50 kb) and HiFi data, a preliminary assembled HiFiasmMix version of the genome was obtained through hybrid assembly using hifiasm (v0.24.0) (parameter: --ul ONT50kreads). All quality-controlled data were aligned to the reference genome using bwa (v0.7.17-r1198-dirty)\u003csup\u003e\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e (parameters: mem \u0026minus;\u0026thinsp;5SP) and filtered using filter_bam (v2.0.0; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/genomewalker/bam-filter\u003c/span\u003e\u003cspan address=\"https://github.com/genomewalker/bam-filter\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e), followed by analysis with HapHiC (v1.0.5)21 for clustering, reassigning, ordering and orientation, and building pseudomolecules. Prior to clustering, preprocessing was performed using HapHiC (v1.0.5)\u003csup\u003e23\u003c/sup\u003e to correct misassemblies. The contigs were subsequently clustered into groups using the Markov clustering algorithm. Reassignment was performed using HapHiC (v1.0.5)\u003csup\u003e23\u003c/sup\u003e. Following manual curation, a BUSCO assessment was conducted using Compleasm (v0.2.5)\u003csup\u003e24\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eNanopore Ultralong reads that were unmapped to the reference genome, had poor mapping quality, and were located at both ends of gaps were extracted. These reads were subsequently used for local iterative assembly to identify the optimal gap-filling path, generating sequences overlapping with both sides of the gap. These sequences were remapped to the gap-containing genome to replace the corresponding gap regions. HiFi reads unmapped to the genome, HiFi reads containing telomeric sequences, and reads mapped to chromosome ends without assembled telomeres were extracted for local assembly. The optimal gap-filling paths were identified to generate sequences overlapping the gap side, after which the corresponding gap regions were replaced. K-mer libraries of 21 k-mers and 31 k-mers were generated separately using yak (v1.0; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/lh3/yak\u003c/span\u003e\u003cspan address=\"https://github.com/lh3/yak\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). HiFi reads were then mapped to the gap-filled genome using minimap2 (v2.26)\u0026sup1;\u0026sup1; to generate BAM files. The genome was subsequently polished using Nextpolish2 (v0.2.0)\u003csup\u003e25\u003c/sup\u003e to obtain the final polished genome.\u003c/p\u003e\n\u003ch3\u003eGenome annotation\u003c/h3\u003e\n\u003cp\u003eTranscriptome sequences were mapped back to the genome using GMAP\u003csup\u003e\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e\u003c/sup\u003e to obtain transcript location information. Afterwards, the reassembled transcripts were realigned to the genome using PASA (v2.3.3)\u003csup\u003e27\u003c/sup\u003e, and optimal open reading frames (ORFs) were identified using GeneMark-ST (v5.1)\u003csup\u003e28\u003c/sup\u003e, resulting in the generation of transcriptome-based gene position predictions. Then, the protein sequences of the target species were obtained. These protein sequences were aligned back to the genome using GeMoMa (v1.6.1)\u003csup\u003e29\u003c/sup\u003e to determine the alignment position information of the corresponding proteins, thereby obtaining gene information annotated based on homologous proteins. Using gene information annotated from transcriptome sequences, gene models were predicted using GeneMark-ST (v5.1)\u003csup\u003e28\u003c/sup\u003e. The predicted genes were aligned against the Swiss-Prot Database\u003csup\u003e\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e using BLASTP (v2.7.1). The alignment results were then filtered (identity\u0026thinsp;\u0026ge;\u0026thinsp;95%), and based on the filtered results, genes predicted by GeneMark-ST(v5.1)\u003csup\u003e28\u003c/sup\u003e were selected as the training set for AUGUSTUS (v3.3.1)\u003csup\u003e31\u003c/sup\u003e model training. Finally, genes in the genome were predicted using AUGUSTUS (v3.3.1)\u003csup\u003e31\u003c/sup\u003e with the predicted model. The results from de novo prediction, homology-based annotation, and transcriptome-based prediction were integrated using EVM (v1.1.1)\u003csup\u003e27\u003c/sup\u003e to obtain a nonredundant exon set, followed by the removal of invalid sequences using TransposonPSI\u003csup\u003e\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e\u003c/sup\u003e. The genomic protein sequences were aligned against the NR, Kyoto Encyclopedia of Genes and Genomes (KEGG)\u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e\u0026sup3;, Eukaryotic Orthologous Groups of proteins (KOG)\u0026sup3;\u003csup\u003e4\u003c/sup\u003e, Gene Ontology (GO)\u0026sup3;\u003csup\u003e5\u003c/sup\u003e, and Swiss-Prot Database\u0026sup2;⁹ using BLASTP (v2.7.1) and InterProScan (5.32-71.0)\u003csup\u003e36\u003c/sup\u003e. The annotated information of these protein sequences was categorized and classified.\u003c/p\u003e \u003cp\u003eWe used GMATA (v2.2)\u003csup\u003e37\u003c/sup\u003e and TRF (v4.07b)\u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e to identify tandem repeats genome wide. For interspersed repeats, we first used MITE-hunter\u003csup\u003e\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e\u003c/sup\u003e to identify miniature inverted-repeat transposable elements (MITEs). The results from LTR_finder\u003csup\u003e\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e\u003c/sup\u003e and LTR_harvest\u003csup\u003e\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e\u003c/sup\u003e were subsequently used to construct an LTR sequence library with LTR_retriever\u003csup\u003e\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e\u003c/sup\u003e. After the two libraries were integrated, RepeatModeler (v1.0.11)\u003csup\u003e\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e\u003c/sup\u003e was used to perform de novo searches for repetitive sequences, which were then classified using TEclass\u003csup\u003e\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e\u003c/sup\u003e, after which the above data were integrated with Repbase\u003csup\u003e\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e\u003c/sup\u003e. Using the integrated data and RepeatMasker (v1.331)\u003csup\u003e\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e\u003c/sup\u003e, we searched for repetitive sequences genome wide.\u003c/p\u003e \u003cp\u003eWe used Infernal (v1.1.2)\u003csup\u003e45\u003c/sup\u003e to search against the Rfam\u003csup\u003e\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e\u003c/sup\u003e database for predicting rRNAs, snRNAs, and miRNAs. Additionally, tRNAscan-SE (v2.0)\u003csup\u003e47\u003c/sup\u003e was used to predict genomic tRNA sequences.\u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eTechnical validation\u003c/h2\u003e \u003cp\u003eCompared with the assembled \u003cem\u003ePythium\u003c/em\u003e genomes by Christoffel et al.\u003csup\u003e8\u003c/sup\u003e, our genome assembly demonstrated enhanced continuity and completeness attributed to the utilization of long reads and Hi-C sequencing. This assembled genome has no gaps, and the telomeres and centromeres of each chromosome have been captured. Furthermore, the BUSCO assessment indicated an overall completeness of 94.9%, including 90.2% single-copy BUSCOs, 4.71% duplicated BUSCOs, 1.57% fragmented BUSCOs, and 3.53% missing BUSCOs.\u003c/p\u003e \u003c/div\u003e"},{"header":"Declarations","content":"\u003cp\u003e \u003ch2\u003eCompeting interests\u003c/h2\u003e \u003cp\u003eThe authors declare that they have no competing interests.\u003c/p\u003e \u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eC.D., Y.C. and Z.C. conceived and supervised the study. Y.Z. collected the samples and conducted the DNA extraction and sequencing. Z.C. and Y.Z. performed the bioinformatics analysis and prepared the figures and tables. C.D., Y.C., Z.C., Y.Z., Y.G., S.L., C.H., W.W., Z.Z. and S.S. wrote and revised the manuscript. All the authors read and approved the final manuscript.\u003c/p\u003e\u003ch2\u003eAcknowledgements\u003c/h2\u003e \u003cp\u003eThis work was supported by grants from the National Key Research \u0026amp; Development Program of China (2021YFD1200700), the Program of Discovery and Application of Resistance Genes to Stalk Rot and Ear Rot in Maize (BLCZK20240920), and the Agricultural Science and Technology Innovation Program of the Chinese Academy of Agricultural Sciences (01-ICS-02).\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eCorrespondence\u003c/strong\u003e and requests for materials should be addressed to C.D., Y.C. and Z.C.\u003c/p\u003e\n\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eThe telomere-to-telomere (T2T) genome assembly of *Pythium maydis* . has been deposited at the National Center for Biotechnology Information (NCBI) GenBank under the accession number GCA_054083355.1 48 . The raw sequencing data, including Nanopore Ultralong reads, PacBio HiFi reads, MGI short reads, Hi-C data, and RNA-Seq data, have been deposited in the NCBI Sequence Read Archive (SRA) under the accession numbers SRR37082792- SRR37082796 49 . All datasets are publicly available and can be accessed via the NCBI database.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCode availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll the software used in this study is in the public domain, with the parameters clearly described in the Methods section and this section. If no detailed parameters were mentioned for the software, default parameters were used as suggested by the developer.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eDuan C., Song F., Sun S., Guo C., Zhu Z., Wang X. Characterization and molecular mapping of two novel genes resistant to \u003cem\u003ePythium\u003c/em\u003e stalk rot in maize. \u003cem\u003ePhytopathology\u003c/em\u003e. 109, 804\u0026ndash;809 (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBroders KD., Lipps PE., Paul PA., Dorrance AE. Characterization of \u003cem\u003ePythium\u003c/em\u003e spp. associated with corn and soybean seed and seedling disease in ohio. \u003cem\u003ePlant Dis\u003c/em\u003e. 91, 727\u0026ndash;735 (2007).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRossman DR., Rojas A., Jacobs JL., Mukankusi C., Kelly JD., Chilvers MI. Pathogenicity and virulence of soilborne oomycetes on phaseolus vulgaris. \u003cem\u003ePlant Dis\u003c/em\u003e. 101, 1851\u0026ndash;1859 (2017).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eReeves ER., Kerns JP, Cowger C, Shew BB. \u003cem\u003ePythium\u003c/em\u003e spp. Associated with root rot and stunting of winter wheat in north carolina. \u003cem\u003ePlant Dis\u003c/em\u003e. 105, 986\u0026ndash;996 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMa M., Taylor PWJ., Chen D., Vaghefi N., He JZ. Major soilborne pathogens of field processing tomatoes and management strategies. \u003cem\u003eMicroorganisms\u003c/em\u003e. 11, 263 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYolanda H., Krajaejun T. Review of methods and antimicrobial agents for susceptibility testing against \u003cem\u003ePythium insidiosum\u003c/em\u003e. \u003cem\u003eHeliyon.\u003c/em\u003e 6, e03737 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGurnani B., Kaur K., Venugopal A., Srinivasan B., Bagga B., Iyer G., Christy J., Prajna L., Vanathi M., Garg P., Narayana S., Agarwal S., Sahu S. \u003cem\u003ePythium insidiosum\u003c/em\u003e keratitis - A review. \u003cem\u003eIndian J Ophthalmol\u003c/em\u003e. 70, 1107\u0026ndash;1120 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNguyen HDT., Dodge A., Dadej K., Rintoul TL., Ponomareva E., Martin FN., de Cock AWAM., L\u0026eacute;vesque CA., Redhead SA., Spies CFJ. Whole genome sequencing and phylogenomic analysis show support for the splitting of genus \u003cem\u003ePythium\u003c/em\u003e. \u003cem\u003eMycologia\u003c/em\u003e. 114, 501\u0026ndash;515 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen S., Zhou Y., Chen Y., Gu J. fastp: an ultra-fast all-in-one FASTQ preprocessor. \u003cem\u003eBioinformatics\u003c/em\u003e, 34, p. i884-i890 (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCheng H., Jarvis ED., Fedrigo O., Koepfli KP., Urban L., Gemmell NJ., Li H. Haplotype-resolved assembly of diploid genomes without parental data. \u003cem\u003eNat Biotechnol\u003c/em\u003e. 40, 1332\u0026ndash;1335 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi, H. New strategies to improve minimap2 alignment accuracy. \u003cem\u003eBioinformatics.\u003c/em\u003e 37, 4572\u0026ndash;4574 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFormenti G., Rhie A., Walenz BP., Thibaud-Nissen F., Shafin K., Koren S., Myers EW., Jarvis ED., Phillippy AM. Merfin: improved variant filtering, assembly evaluation and polishing via k-mer validation. \u003cem\u003eNat Methods\u003c/em\u003e. 19, 696\u0026ndash;704 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRhie A., Walenz BP., Koren S., Phillippy AM. Merqury: reference-free quality, completeness, and phasing assessment for genome assemblies. \u003cem\u003eGenome Biol\u003c/em\u003e. 21, 245 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShen W., Le S., Li Y., Hu F. SeqKit: A cross-platform and ultrafast toolkit for FASTA/Q file manipulation. \u003cem\u003ePLoS ONE\u003c/em\u003e. 11, e0163962 (2016).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBenson G. Tandem repeats finder: a program to analyze DNA sequences. \u003cem\u003eNucleic Acids Research.\u003c/em\u003e 27, 573\u0026ndash;580 (1999).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBedell JA., Korf I., Gish W. MaskerAid: a performance enhancement to RepeatMasker. \u003cem\u003eBioinformatics.\u003c/em\u003e 16, 1040\u0026ndash;1 (2000).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang Y., Chu J., Cheng H., Li H. De novo reconstruction of satellite repeat units from sequence data. \u003cem\u003eGenome Res\u003c/em\u003e. 33, 1994\u0026ndash;2001 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLin Y., Ye C., Li X., Chen Q., Wu Y., Zhang F., Pan R., Zhang S., Chen S., Wang X., Cao S., Wang Y., Yue Y., Liu Y., Yue J. quarTeT: a telomere-to-telomere toolkit for gap-free genome assembly and centromeric repeat identification. \u003cem\u003eHorticulture Research\u003c/em\u003e. 10, uhad127 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eIšerić H., Alkan C., Hach F., Numanagić I. Fast characterization of segmental duplication structure in multiple genome assemblies. \u003cem\u003eAlgorithms Mol Biol.\u003c/em\u003e 17, 4 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKirsche M., Das A., Schatz MC. Sapling: accelerating suffix array queries with learned data models, \u003cem\u003eBioinformatics\u003c/em\u003e, 37, 744\u0026ndash;749 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBelton JM., McCord RP., Gibcus JH., Naumova N., Zhan Y., Dekker J. Hi-C: A comprehensive technique to capture the conformation of genomes. \u003cem\u003eMethods\u003c/em\u003e. 58, 268\u0026ndash;76 (2012).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi H. Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM. \u003cem\u003eGenomics\u003c/em\u003e. 1303.3997v2. (2013)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZeng X., Yi Z., Zhang X., Du Y., Li Y., Zhou Z., Chen S., Zhao H., Yang S., Wang Y., Chen G. Chromosome-level scaffolding of haplotype-resolved assemblies using Hi-C data without reference genomes. \u003cem\u003eNature Plants\u003c/em\u003e. 10, 1184\u0026ndash;1200 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHuang N., Li H. compleasm: a faster and more accurate reimplementation of BUSCO. \u003cem\u003eBioinformatics\u003c/em\u003e, 39, btad595 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHu J., Wang Z., Liang F., Liu SL., Ye K., Wang DP. NextPolish2: A repeat-aware polishing tool for genomes assembled using HiFi long reads. \u003cem\u003eGenomics, Proteomics \u0026amp; Bioinformatics\u003c/em\u003e. 22, qzad009 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWu TD., Watanabe CK. GMAP: a genomic mapping and alignment program for mRNA and EST sequences. Bioinformatics. \u003cem\u003eBioinformatics.\u003c/em\u003e 21, 1859\u0026ndash;75 (2005).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHaas BJ., Salzberg SL., Zhu W., Pertea M., Allen JE., Orvis J., White O., Buell CR., Wortman JR. Automated eukaryotic gene structure annotation using EVidenceModeler and the program to assemble spliced alignments. \u003cem\u003eGenome Biol\u003c/em\u003e. 9, R7 (2008).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTang S., Lomsadze A., Borodovsky M. Identification of protein coding regions in RNA transcripts. \u003cem\u003eNucleic Acids Res\u003c/em\u003e. 43, e78 (2015).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKeilwagen J., Wenk M., Erickson JL., Schattat MH., Grau J., Hartung F. Using intron position conservation for homology-based gene prediction. \u003cem\u003eNucleic Acids Res\u003c/em\u003e. 44, e89 (2016).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBairoch A., Apweiler R., Wu CH., Barker WC., Boeckmann B., Ferro S., Gasteiger E., Huang H., Lopez R., Magrane M., Martin MJ., Natale DA., O'Donovan C., Redaschi N., Yeh LS. The universal protein resource (UniProt). \u003cem\u003eNucleic Acids Res\u003c/em\u003e. 33, D154-9 (2005).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eStanke M., Diekhans M., Baertsch R., Haussler D. Using native and syntenically mapped cDNA alignments to improve de novo gene finding. \u003cem\u003eBioinformatics.\u003c/em\u003e 24, 637\u0026ndash;44 (2008).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eUrasaki N., Takagi H., Natsume S., Uemura A., Taniai N., Miyagi N., Fukushima M., Suzuki S., Tarora K., Tamaki M., Sakamoto M., Terauchi R., Matsumura H. Draft genome sequence of bitter gourd (\u003cem\u003eMomordica charantia\u003c/em\u003e), a vegetable and medicinal plant in tropical and subtropical regions. \u003cem\u003eDNA Res\u003c/em\u003e. 24, 51\u0026ndash;58 (2017).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOgata H., Goto S., Sato K., Fujibuchi W., Bono H., Kanehisa M. KEGG: kyoto encyclopedia of genes and genomes. \u003cem\u003eNucleic Acids Res\u003c/em\u003e. 28, 27\u0026ndash;30 (2000).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGalperin MY., Makarova KS., Wolf YI., Koonin EV. Expanded microbial genome coverage and improved protein family annotation in the COG database. \u003cem\u003eNucleic Acids Res\u003c/em\u003e. 43, D261-9 (2015).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAshburner M., Ball CA., Blake JA., Botstein D., Butler H., Cherry JM., Davis AP., Dolinski K., Dwight SS., Eppig JT., Harris MA., Hill DP., Issel-Tarver L., Kasarskis A., Lewis S., Matese JC., Richardson JE., Ringwald M., Rubin GM., Sherlock G. Gene ontology: tool for the unification of biology. The Gene Ontology Consortium. \u003cem\u003eNat Genet\u003c/em\u003e. 25, 25\u0026ndash;9 (2000).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZdobnov EM., Apweiler R. InterProScan\u0026ndash;an integration platform for the signature-recognition methods in InterPro. \u003cem\u003eBioinformatics.\u003c/em\u003e 17, 847\u0026ndash;8 (2001).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang, X., L. Wang. GMATA: An integrated software package for genome-scale SSR Mining, Marker Development and Viewing. \u003cem\u003eFront Plant Sci\u003c/em\u003e. 7, 1350 (2016).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHan, Y., S.R. Wessler, MITE-Hunter: a program for discovering miniature inverted-repeat transposable elements from genomic sequences. \u003cem\u003eNucleic Acids Res\u003c/em\u003e. 38, 199 (2010).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXu, Z., H. Wang. LTR_FINDER: an efficient tool for the prediction of full-length LTR retrotransposons. \u003cem\u003eNucleic Acids Res\u003c/em\u003e. 35, 265\u0026ndash;8 (2007).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEllinghaus D., Kurtz S., Willhoeft U. LTRharvest, an efficient and flexible software for de novo detection of LTR retrotransposons. \u003cem\u003eBMC Bioinformatics\u003c/em\u003e. 9, 18 (2008).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOu S., Jiang N. LTR_retriever: A highly accurate and sensitive program for identification of long terminal repeat retrotransposons. \u003cem\u003ePlant Physiol\u003c/em\u003e. 176, 1410\u0026ndash;1422 (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBedell JA., Korf I., Gish W. MaskerAid: a performance enhancement to RepeatMasker. \u003cem\u003eBioinformatics.\u003c/em\u003e 16, 1040\u0026ndash;1 (2000).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAbrus\u0026aacute;n G., Grundmann N., DeMester L., Makalowski W. TEclass\u0026ndash;a tool for automated classification of unknown eukaryotic transposable elements. \u003cem\u003eBioinformatics.\u003c/em\u003e 25, 1329\u0026ndash;30 (2009).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJurka J., Kapitonov VV., Pavlicek A., Klonowski P., Kohany O., Walichiewicz J. Repbase Update, a database of eukaryotic repetitive elements. \u003cem\u003eCytogenet Genome Res\u003c/em\u003e. 110, 462\u0026ndash;7 (2005).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNawrocki EP., Eddy SR. Eddy. Infernal 1.1: 100-fold faster RNA homology searches. \u003cem\u003eBioinformatics\u003c/em\u003e. 29, 2933\u0026ndash;5 (2013).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGriffiths-Jones S., Moxon S., Marshall M., Khanna A., Eddy SR., Bateman A. Rfam: annotating non-coding RNAs in complete genomes. \u003cem\u003eNucleic Acids Res\u003c/em\u003e. 33, D121-4 (2005).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLowe TM., Eddy SR. tRNAscan-SE: a program for improved detection of transfer RNA genes in genomic sequence. \u003cem\u003eNucleic Acids Res\u003c/em\u003e. 25, 955\u0026ndash;64 (1997).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNCBI GenBank \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ncbi.nlm.nih.gov/assembly/GCA_054083355.1\u003c/span\u003e\u003cspan address=\"https://www.ncbi.nlm.nih.gov/assembly/GCA_054083355.1\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNCBI Sequence Read Archive \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ncbi.nlm.nih.gov\u003c/span\u003e\u003cspan address=\"https://www.ncbi.nlm.nih.gov\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e/sra?linkname=bioproject_sra_all\u0026amp;from_uid=1358918 (2026).\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-8634265/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8634265/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cem\u003ePythium\u003c/em\u003e spp. include numerous important pathogens that affect both plants and animals, posing serious threats to agricultural production and the health of humans and livestock. Currently, high-quality genomic resources for these species are scarce, impeding the understanding of population evolution and the development of prevention and control strategies. Here, the first complete telomere-to-telomere (T2T) genome of \u003cem\u003ePythium \u003c/em\u003esp. was produced by integrating NGS, PacBio HiFi,ONT and Hi-C data. The genome size is 59.9 Mb in total, with a scaffold N50 length of 3.5 Mb. BUSCO assessment indicated that the genome assembly achieved 94.91% (n=255) completeness. A total of 14,111 protein-coding genes were predicted, with 12,711 genes (90.08% of the total) annotated using functional databases. Additionally, we predicted 382 effectors. This T2T genome is an excellent resource for furthering the understanding of the evolutionary path and pathogenesis mechanisms of \u003cem\u003ePythium \u003c/em\u003espp.\u003c/p\u003e","manuscriptTitle":"A telomere-to-telomere genome assembly of Pythium species","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-03-05 10:26:40","doi":"10.21203/rs.3.rs-8634265/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"6a8a5f82-ec40-4d3f-bdf5-3090b4dcc779","owner":[],"postedDate":"March 5th, 2026","published":true,"recentEditorialEvents":[{"type":"editorInvitedReview","content":"","date":"2026-05-08T14:39:15+00:00","index":53,"fulltext":""}],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2026-03-05T10:26:41+00:00","versionOfRecord":[],"versionCreatedAt":"2026-03-05 10:26:40","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8634265","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8634265","identity":"rs-8634265","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-28T02:00:01.590549+00:00
License: CC-BY-4.0