PacBio long read-assembled draft genome of Pythium insidiosum strain Pi-S isolated from a Thai patient with pythiosis

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Objectives: Pythium insidiosum is the causative agent of pythiosis, a difficult-to-treat condition, in humans and animals worldwide. Biological information about this filamentous microorganism is sparse. Genomes of several P. insidiosum strains were sequenced using the Illumina short-read NGS platform, producing incomplete genome sequence data. PacBio long-read platform was employed to obtain a better-quality genome of Pythium insidiosum . The obtained genome data could promote basic research on the pathogen’s biology and pathogenicity. Data description: gDNA sample was extracted from the P. insidiosum strain Pi-S for whole-genome sequencing by PacBio long-read NGS platform. Raw reads were assembled using CANU (v2.1), polished using ARROW (SMRT link version 5.0.1), aligned with the original raw PacBio reads using pbmm2 (v1.2.1), consensus sequence checked using ARROW, and gene predicted using Funannotate pipeline (v1.7.4). The genome completion was assessed using BUSCO (v4.0.2). As the result, 840 contigs (maximum length: 1.3 Mb; N 50 229.9 Kb; L 50 70) were obtained. Sequence assembly showed a genome size of 66.7 Mb (178x coverage; 57.2% G-C content) that contained 20,375 ORFs. A BUSCO-based assessment revealed 85.5% genome completion. All assembled contigs sequences have been deposited in the DDBJ database under the accession numbers BBXB02000001 - BBXB02000840.
Full text 43,458 characters · extracted from preprint-html · click to expand
PacBio long read-assembled draft genome of Pythium insidiosum strain Pi-S isolated from a Thai patient with pythiosis | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Data Note PacBio long read-assembled draft genome of Pythium insidiosum strain Pi-S isolated from a Thai patient with pythiosis Theerapong Krajaejun, Preecha Patumcharoenpol, Thidarat Rujirawat, and 4 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-2124426/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 13 Oct, 2023 Read the published version in BMC Research Notes → Version 1 posted 6 You are reading this latest preprint version Abstract Objectives: Pythium insidiosum is the causative agent of pythiosis, a difficult-to-treat condition, in humans and animals worldwide. Biological information about this filamentous microorganism is sparse. Genomes of several P. insidiosum strains were sequenced using the Illumina short-read NGS platform, producing incomplete genome sequence data. PacBio long-read platform was employed to obtain a better-quality genome of Pythium insidiosum . The obtained genome data could promote basic research on the pathogen’s biology and pathogenicity. Data description: gDNA sample was extracted from the P. insidiosum strain Pi-S for whole-genome sequencing by PacBio long-read NGS platform. Raw reads were assembled using CANU (v2.1), polished using ARROW (SMRT link version 5.0.1), aligned with the original raw PacBio reads using pbmm2 (v1.2.1), consensus sequence checked using ARROW, and gene predicted using Funannotate pipeline (v1.7.4). The genome completion was assessed using BUSCO (v4.0.2). As the result, 840 contigs (maximum length: 1.3 Mb; N 50 : 229.9 Kb; L 50 : 70) were obtained. Sequence assembly showed a genome size of 66.7 Mb (178x coverage; 57.2% G-C content) that contained 20,375 ORFs. A BUSCO-based assessment revealed 85.5% genome completion. All assembled contigs sequences have been deposited in the DDBJ database under the accession numbers BBXB02000001 - BBXB02000840. Pythium insidiosum Pythiosis Draft genome Next-Generation Sequencing Objective Short- and long-read genome sequencing technologies are now widely available for generating genome data of various organisms [1]. The major difference between the two is the maximum length of generated sequence reads: 150-400 bases for the short-read platforms (i.e., Illumina and Ion Torrent) [2, 3] and many kilobases for the long-read platforms (i.e., Pacific BioSciences (PacBio) [4] and Oxford Nanopore [5]). Another difference is the genome sequence coverage, in which the short-read platforms tend to produce a relatively-higher coverage (although likely incomplete) genome of an organism of interest. In contrast, the long-read platforms generate a relatively more complete genome (although at a higher cost). The genomes of 10 strains of Pythium insidiosum (the causative agent of the difficult-to-treat infectious disease, namely pythiosis, in humans and various animals worldwide [6, 7]) were sequenced using the Illumina short-read platform, resulting in up to sixty thousand contigs and thus incompletely assembled genome data [8–13]. This study aims to obtain a better-quality genome of P. insidiosum using the PacBio long-read. The P. insidiosum strain Pi-S was selected because it has been widely referred to in recent immunological, genomic, transcriptomic, and proteomic studies [8, 14–20]. The short-read version of the strain Pi-S genome data shows 53.2 Mb in size and comprises 1,192 contigs (average length: 44,664 bases; N 50 : 146,252 bases), 52% G-C content, 10% N composition, and 14,962 open reading frames (ORFs) [8]. When it becomes available, better-quality genome data of P. insidiosum can promote an in-depth genetic exploration and a better understanding of this filamentous microorganism's biology, evolution, and pathogenicity. Such advances could lead to the development of a clinical application for preventing, diagnosing, and treating the disease caused by this devastating pathogen. Data description Genomic DNA (gDNA) was obtained from 7-day-old liquid culture of P. insidiosum (strain Pi-S) using our previously described protocol [21]. Harvested hyphae (1,000 mg wet weight) were ground in a mortar after adding liquid nitrogen. Resulting hyphal power was moved to a sterile 50-mL tube containing 5 ml of the extraction solution [250 mM NaCl, 100 mM Tris–HCl (pH 8.0), 100 mM ethylenediaminetetraacetic acid (pH 8.0), and 1% sodium dodecyl sulfate] and RNase A (final concentration: 50 µg/ml). The cell lysate was incubated (with gentle inversion) at 37 °C for 2 hr before adding proteinase K (final concentration: 50 mg/mL). After the sample was incubated at 56 °C overnight, the supernatant was collected by centrifugation (10,000 xg) at room temperature for 30 min, mixed with an equal volume of phenol:chloroform:isoamyl alcohol (25:24:1) solution, gently inverted for 15 min, and centrifuged (2,000 xg) at 25 °C (room temperature) for 30 min. The resulting aqueous phase was moved to a new 50-ml tube, mixed with the same volume of isopropanol, gently inverted 10 times, and centrifuged (2,000 xg) at 25 °C for 5 min. A resulting gDNA pellet was collected, washed with 70% ethanol, air dried, and resuspended in 5 mM Tris-HCl (pH 8.0). Extracted gDNA was kept at -20 °C and transported the following day to the National Omics Center, National Science and Technology Development Agency, Pathum Thani, Thailand, for genome sequencing. Genome sequencing was done following the method of Shearman et al. [22]. In brief, a P. insidiosum gDNA sample was purified using the Ampure PB beads (Pacific Biosciences, Menlo Park, USA) and checked for DNA integrity using the Pippin Pulse Electrophoresis System (Sage Science, Beverly, USA). The gDNA sample (10 ng) was torn using a Covaris gTube (4,500 rpm for 2 min) for library preparation (using a ‘0.75%DF Marker S1 high-pass 15–20 kb’ BluePippin cassette with 12-50kb selection) according to the Pacific Biosciences protocol (20 kb Template Preparation Using BluePippin Size-Selection). Genome sequencing was conducted on the PacBio RSII. Raw reads with a length of at least 20 kb were defined as seed reads, which were corrected by shorter reads (sizes < 20 kb) using the Pacific Biosciences SMRT analysis software v2.3.0 (the RS_PreAssembler.1 protocol with default settings). PacBio-derived corrected reads (2,632.0 Mb) were assembled using CANU (v2.1) [23], a software that shows an ability to produce highly contiguous assemblies. Three rounds of polishing were carried out using ARROW [24] (SMRT link version 5.0.1). The first initial assemblies were aligned with the original raw PacBio reads using pbmm2 (v1.2.1) (https://github.com/PacificBiosciences/pbmm2). The resulting alignment was then used for calling the consensus sequence using ARROW. Gene prediction was made with the Funannotate pipeline (v1.7.4) (https://github.com/nextgenusfs/funannotate). As a result, a total of 840 contigs (maximum length: ~1.3 Mb; minimum length: 1,294 bases; N 50 : 229.9 Kb; L 50 : 70) were obtained. The sequence assembly showed a total genome length of 66.7 Mb (178x coverage) containing 57.2% G-C content and 20,375 ORFs. The genome completion estimated using BUSCO (v4.0.2) [25] and the Fungi Ogb10 dataset (containing 250 conserve eukaryotic genes) was 85.5%. All contig sequences have been deposited in the NCBI database under the accession numbers BBXB02000001 - BBXB02000840 via the project PRJDB3817 ( Table 1 ). Table 1 : Overview of data files/data sets. Label Name of data file/data set File types (file extension) Data repository and identifier (DOI or accession number) Data file 1 P. insidiosum strain Pi-S, whole genome shotgun sequencing project FASTA BBXB02000001 - BBXB02000840 Limitations We used the PacBio long-read NGS platform to improve the genome quality of P. insidiosum strain Pi-S. Compared with the short-read NGS platforms (such as Illumina), genome sequencing employing the long-read platforms (such as PacBio) provides a more complete genome sequence but is a relatively more expensive technique and shows a higher base error rate. Abbreviations DNA, Deoxyribonucleic acid gDNA, Genomic deoxyribonucleic acid NCBI, National Center for Biotechnology Information NGS, Next-generation sequencing OD, Optical density ORF, Open reading frame Declarations Ethics approval and consent to participate This study was approved by the Human Research Ethics Committee, Faculty of Medicine, Ramathibodi Hospital, Mahidol University (MURA2022/274). Consent for publication Not applicable. Availability of data and materials The draft genome sequence of the P. insidiosum strain Pi-S comprising 840 contigs (accession numbers BBXB02000001 - BBXB02000840) is available in the DDBJ database. Competing interests None. Funding This study obtained financial support from the Faculty of Medicine, Ramathibodi Hospital, Mahidol University [Grant number CF_65003] and the National Research Council of Thailand (NRCT) and Mahidol University [Grant number N42A650339]. The funders had no role in the design of the study and collection, analysis, and interpretation of data and in writing the manuscript. Authors’ contributions T.K. conceived the project. S.T., T.R., T.L., and W.Y. performed the experiments. W.K., P.P., S.T., T.R., and T.K. analyzed the data. W.K., P.P., T.R., S.T., and T.K. wrote the manuscript. All authors reviewed the manuscript. T.K. acquired the research funds. Acknowledgments Not applicable. References Kittichotirat W, Krajaejun T. Application of Genome Sequencing to Study Infectious Diseases. J Infect Dis Antimicrob Agents. 2019;36:47–58. Bentley DR, Balasubramanian S, Swerdlow HP, Smith GP, Milton J, Brown CG, et al. Accurate whole human genome sequencing using reversible terminator chemistry. Nature. 2008;456:53–9. Rothberg J, Myers J. Semiconductor sequencing for life. J Biomol Tech. 2011;22:S41–2. Rhoads A, Au KF. PacBio Sequencing and Its Applications. Genomics Proteomics Bioinformatics. 2015;13:278–89. Lu H, Giordano F, Ning Z. Oxford Nanopore MinION Sequencing and Genome Assembly. Genomics Proteomics Bioinformatics. 2016;14:265–79. Chitasombat MN, Jongkhajornpong P, Lekhanont K, Krajaejun T. Recent update in diagnosis and treatment of human pythiosis. PeerJ. 2020;8:e8555. Yolanda H, Krajaejun T. Global Distribution and Clinical Features of Pythiosis in Humans and Animals. J Fungi Basel Switz. 2022;8:182. Rujirawat T, Patumcharoenpol P, Lohnoo T, Yingyong W, Lerksuthirat T, Tangphatsornruang S, et al. Draft Genome Sequence of the Pathogenic Oomycete Pythium insidiosum Strain Pi-S, Isolated from a Patient with Pythiosis. Genome Announc. 2015;3:e00574-15. Kittichotirat W, Patumcharoenpol P, Rujirawat T, Lohnoo T, Yingyong W, Krajaejun T. Draft genome and sequence variant data of the oomycete Pythium insidiosum strain Pi45 from the phylogenetically-distinct Clade-III. Data Brief. 2017;15:896–900. Krajaejun T, Kittichotirat W, Patumcharoenpol P, Rujirawat T, Lohnoo T, Yingyong W. Data on whole genome sequencing of the oomycete Pythium insidiosum strain CBS 101555 from a horse with pythiosis in Brazil. BMC Res Notes. 2018;11:880. Krajaejun T, Kittichotirat W, Patumcharoenpol P, Rujirawat T, Lohnoo T, Yingyong W. Draft genome sequence of the oomycete Pythium destruens strain ATCC 64221 from a horse with pythiosis in Australia. BMC Res Notes. 2020;13:329. Patumcharoenpol P, Rujirawat T, Lohnoo T, Yingyong W, Vanittanakom N, Kittichotirat W, et al. Draft genome sequences of the oomycete Pythium insidiosum strain CBS 573.85 from a horse with pythiosis and strain CR02 from the environment. Data Brief. 2018;16:47–50. Krajaejun T, Kittichotirat W, Patumcharoenpol P, Rujirawat T, Lohnoo T, Yingyong W. Genome data of four Pythium insidiosum strains from the phylogenetically-distinct clades I, II, and III. BMC Res Notes. 2021;14:197. Rujirawat T, Patumcharoenpol P, Lohnoo T, Yingyong W, Kumsang Y, Payattikul P, et al. Probing the Phylogenomics and Putative Pathogenicity Genes of Pythium insidiosum by Oomycete Genome Analyses. Sci Rep. 2018;8:4135. Krajaejun T, Rujirawat T, Lohnoo T, Yingyong W, Sae-Chew P, Reamtong O, et al. Secretome Profiling by Proteogenomic Analysis Shows Species-Specific, Temperature-Dependent, and Putative Virulence Proteins of Pythium insidiosum. J Fungi Basel Switz. 2022;8:527. Sae-Chew P, Rujirawat T, Kumsang Y, Payattikul P, Lohnoo T, Yingyong W, et al. Automated Cell-Free Multiprotein Synthesis Facilitates the Identification of a Secretory, Oligopeptide Elicitor-Like, Immunoreactive Protein of the Oomycete Pythium insidiosum. mSystems. 2020;5. Rotchanapreeda T, Sae-Chew P, Lohnoo T, Yingyong W, Rujirawat T, Kumsang Y, et al. Immunological Cross-Reactivity of Proteins Extracted from the Oomycete Pythium insidiosum and the Fungus Basidiobolus ranarum Compromises the Detection Specificity of Immunodiagnostic Assays for Pythiosis. J Fungi Basel Switz. 2021;7:474. Krajaejun T, Lerksuthirat T, Garg G, Lowhnoo T, Yingyong W, Khositnithikul R, et al. Transcriptome analysis reveals pathogenicity and evolutionary history of the pathogenic oomycete Pythium insidiosum. Fungal Biol. 2014;118:640–53. Krajaejun T, Lohnoo T, Jittorntam P, Srimongkol A, Kumsang Y, Yingyong W, et al. Assessment of matrix-assisted laser desorption ionization-time of flight mass spectrometry for identification and biotyping of the pathogenic oomycete Pythium insidiosum. Int J Infect Dis. 2018;77:61–7. Rujirawat T, Sridapan T, Lohnoo T, Yingyong W, Kumsang Y, Sae-Chew P, et al. Single nucleotide polymorphism-based multiplex PCR for identification and genotyping of the oomycete Pythium insidiosum from humans, animals and the environment. Infect Genet Evol. 2017;54:429–36. Lohnoo T, Jongruja N, Rujirawat T, Yingyon W, Lerksuthirat T, Nampoon U, et al. Efficiency comparison of three methods for extracting genomic DNA of the pathogenic oomycete Pythium insidiosum. J Med Assoc Thai. 2014;97:342–8. Shearman JR, Sonthirod C, Naktang C, Sangsrakru D, Yoocha T, Chatbanyong R, et al. Assembly of the durian chloroplast genome using long PacBio reads. Sci Rep. 2020;10:15980. Koren S, Walenz BP, Berlin K, Miller JR, Bergman NH, Phillippy AM. Canu: scalable and accurate long-read assembly via adaptive k-mer weighting and repeat separation. Genome Res. 2017;27:722–36. Chin C-S, Alexander DH, Marks P, Klammer AA, Drake J, Heiner C, et al. Nonhybrid, finished microbial genome assemblies from long-read SMRT sequencing data. Nat Methods. 2013;10:563–9. Waterhouse RM, Seppey M, Simão FA, Manni M, Ioannidis P, Klioutchnikov G, et al. BUSCO Applications from Quality Assessments to Gene Prediction and Phylogenomics. Mol Biol Evol. 2018;35:543–8. Additional Declarations No competing interests reported. Cite Share Download PDF Status: Published Journal Publication published 13 Oct, 2023 Read the published version in BMC Research Notes → Version 1 posted Editorial decision: Major revision 26 Oct, 2022 Reviews received at journal 25 Oct, 2022 Reviewers agreed at journal 24 Oct, 2022 Reviewers invited by journal 20 Oct, 2022 Submission checks completed at journal 20 Oct, 2022 First submitted to journal 20 Oct, 2022 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-2124426","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Data Note","associatedPublications":[],"authors":[{"id":145660036,"identity":"eff7d5f0-15df-4b5b-87a3-17e7041edb6e","order_by":0,"name":"Theerapong Krajaejun","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAsUlEQVRIiWNgGAWjYNCCigNg6sAD4rWcOcDAA9KSQLQOxjaIFgaitOjOPsD4uHDeHTl7scMPgbbYyek2ENBidi6B2XjmtmfGPNJpBkAtycZmBwhpOcPAJs277XBij3QCSMuBxG1EaGH/zTsHpCX9A9Fa2Jh5G0Bacoi3hVma59hhY57bOQUHEgyI8wvjZ56aw3Lss9M3f/hQYSdHUAsDA/8HJI4BQeWjYBSMglEwCogBAMidQfXvwAq8AAAAAElFTkSuQmCC","orcid":"","institution":"Mahidol University","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Theerapong","middleName":"","lastName":"Krajaejun","suffix":""},{"id":145660037,"identity":"c6037bce-b916-42d0-a57b-f64addbf54ca","order_by":1,"name":"Preecha Patumcharoenpol","email":"","orcid":"","institution":"Kasetsart University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Preecha","middleName":"","lastName":"Patumcharoenpol","suffix":""},{"id":145660038,"identity":"d941dd84-bd38-4299-82c2-a3d4bfa3c606","order_by":2,"name":"Thidarat Rujirawat","email":"","orcid":"","institution":"Mahidol University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Thidarat","middleName":"","lastName":"Rujirawat","suffix":""},{"id":145660039,"identity":"87fd2329-4cb8-400b-a84d-16cf35605bf4","order_by":3,"name":"Weerayuth Kittichotirat","email":"","orcid":"","institution":"King Mongkut’s University of Technology Thonburi","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Weerayuth","middleName":"","lastName":"Kittichotirat","suffix":""},{"id":145660040,"identity":"56005cb4-033e-410a-9cfd-1f98eba69b8c","order_by":4,"name":"Sithichoke Tangphatsornruang","email":"","orcid":"","institution":"National Science and Technology Development Agency","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Sithichoke","middleName":"","lastName":"Tangphatsornruang","suffix":""},{"id":145660041,"identity":"d727b4ad-34c6-4699-a150-fe06a15deab6","order_by":5,"name":"Tassanee Lohnoo","email":"","orcid":"","institution":"Mahidol University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Tassanee","middleName":"","lastName":"Lohnoo","suffix":""},{"id":145660042,"identity":"96bf661f-8b8d-4072-a974-9270f5373e78","order_by":6,"name":"Wanta Yingyong","email":"","orcid":"","institution":"Mahidol University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Wanta","middleName":"","lastName":"Yingyong","suffix":""}],"badges":[],"createdAt":"2022-10-01 22:14:12","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-2124426/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-2124426/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1186/s13104-023-06532-7","type":"published","date":"2023-10-13T15:02:57+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":44793295,"identity":"472905fb-81b6-4023-b2d6-920f68a9f163","added_by":"auto","created_at":"2023-10-17 14:57:49","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":251066,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-2124426/v1/526c527b-974d-49dd-818b-2dfef8cbb9b6.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"PacBio long read-assembled draft genome of Pythium insidiosum strain Pi-S isolated from a Thai patient with pythiosis","fulltext":[{"header":"Objective ","content":"\u003cp\u003eShort- and long-read genome sequencing technologies are now widely available for generating genome data of various organisms [1]. The major difference between the two is the maximum length of generated sequence reads: 150-400 bases for the short-read platforms (i.e., Illumina and Ion Torrent) [2, 3] and many kilobases for the long-read platforms (i.e., Pacific BioSciences (PacBio) [4] and Oxford Nanopore [5]). Another difference is the genome sequence coverage, in which the short-read platforms tend to produce a relatively-higher coverage (although likely incomplete) genome of an organism of interest. In contrast, the long-read platforms generate a relatively more complete genome (although at a higher cost). The genomes of 10 strains of \u003cem\u003ePythium insidiosum\u003c/em\u003e (the causative agent of the difficult-to-treat infectious disease, namely pythiosis, in humans and various animals worldwide [6, 7]) were sequenced using the Illumina short-read platform, resulting in up to sixty thousand contigs and thus incompletely assembled genome data [8\u0026ndash;13]. This study aims to obtain a better-quality genome of \u003cem\u003eP. insidiosum\u003c/em\u003e using the PacBio long-read. The \u003cem\u003eP. insidiosum\u0026nbsp;\u003c/em\u003estrain Pi-S was selected because it has been widely referred to in recent immunological, genomic, transcriptomic, and proteomic studies [8, 14\u0026ndash;20]. The short-read version of the strain Pi-S genome data shows 53.2 Mb in size and comprises 1,192 contigs (average length: 44,664 bases; \u003cem\u003eN\u003c/em\u003e\u003csub\u003e50\u003c/sub\u003e: 146,252 bases), 52% G-C content, 10% N composition, and 14,962 open reading frames (ORFs) [8]. When it becomes available, better-quality genome data of \u003cem\u003eP. insidiosum\u003c/em\u003e can promote an in-depth genetic exploration and a better understanding of this filamentous microorganism\u0026apos;s biology, evolution, and pathogenicity. Such advances could lead to the development of a clinical application for preventing, diagnosing, and treating the disease caused by this devastating pathogen.\u003c/p\u003e"},{"header":"Data description ","content":"\u003cp\u003eGenomic DNA (gDNA) was obtained from 7-day-old liquid culture of \u003cem\u003eP. insidiosum\u0026nbsp;\u003c/em\u003e(strain Pi-S) using our previously described protocol [21]. Harvested hyphae (1,000 mg wet weight) were ground in a mortar after adding liquid nitrogen. Resulting hyphal power was moved to a sterile 50-mL tube containing 5 ml of the extraction solution [250 mM NaCl, 100 mM Tris\u0026ndash;HCl (pH 8.0), 100 mM ethylenediaminetetraacetic acid (pH 8.0), and 1% sodium dodecyl sulfate] and RNase A (final concentration: 50 \u0026micro;g/ml). The cell lysate was incubated (with gentle inversion) at 37 \u0026deg;C for 2 hr before adding proteinase K (final concentration: 50 mg/mL). After the sample was incubated at 56 \u0026deg;C overnight, the supernatant was collected by centrifugation (10,000 xg) at room temperature for 30 min, mixed with an equal volume of phenol:chloroform:isoamyl alcohol (25:24:1) solution, gently inverted for 15 min, and centrifuged (2,000 xg) at 25 \u0026deg;C (room temperature) for 30 min. The resulting aqueous phase was moved to a new 50-ml tube, mixed with the same volume of isopropanol, gently inverted 10 times, and centrifuged (2,000 xg) at 25 \u0026deg;C for 5 min. A resulting gDNA pellet was collected, washed with 70% ethanol, air dried, and resuspended in 5 mM Tris-HCl (pH 8.0). Extracted gDNA was kept at -20 \u0026deg;C and transported the following day to the National Omics Center, National Science and Technology Development Agency, Pathum Thani, Thailand, for genome sequencing.\u003c/p\u003e\n\u003cp\u003eGenome sequencing was done following the method of Shearman et al. [22]. In brief, a \u003cem\u003eP. insidiosum\u003c/em\u003e gDNA sample was purified using the Ampure PB beads (Pacific Biosciences, Menlo Park, USA) and checked for DNA integrity using the Pippin Pulse Electrophoresis System (Sage Science, Beverly, USA). The gDNA sample (10 ng) was torn using a Covaris gTube (4,500 rpm for 2 min) for library preparation (using a \u0026lsquo;0.75%DF Marker S1 high-pass 15\u0026ndash;20 kb\u0026rsquo; BluePippin cassette with 12-50kb selection) according to the Pacific Biosciences protocol (20 kb Template Preparation Using BluePippin Size-Selection). Genome sequencing was conducted on the PacBio RSII. Raw reads with a length of at least 20 kb were defined as seed reads, which were corrected by shorter reads (sizes \u0026lt; 20 kb) using the Pacific Biosciences SMRT analysis software v2.3.0 (the RS_PreAssembler.1 protocol with default settings).\u003c/p\u003e\n\u003cp\u003ePacBio-derived corrected reads (2,632.0 Mb) were assembled using CANU (v2.1) [23], a software that shows an ability to produce highly contiguous assemblies. Three rounds of polishing were carried out using ARROW [24] (SMRT link version 5.0.1). The first initial assemblies were aligned with the original raw PacBio reads using pbmm2 (v1.2.1) (https://github.com/PacificBiosciences/pbmm2). The resulting alignment was then used for calling the consensus sequence using ARROW. Gene prediction was made with the Funannotate pipeline (v1.7.4) (https://github.com/nextgenusfs/funannotate). As a result, a total of 840 contigs (maximum length: ~1.3 Mb; minimum length: 1,294 bases; \u003cem\u003eN\u003c/em\u003e\u003csub\u003e50\u003c/sub\u003e: 229.9 Kb; \u003cem\u003eL\u003c/em\u003e\u003csub\u003e50\u003c/sub\u003e: 70) were obtained. The sequence assembly showed a total genome length of 66.7 Mb (178x coverage) containing 57.2% G-C content and 20,375 ORFs. The genome completion estimated using BUSCO (v4.0.2) [25] and the Fungi Ogb10 dataset (containing 250 conserve eukaryotic genes) was 85.5%. All contig sequences have been deposited in the NCBI database under the accession numbers BBXB02000001 - BBXB02000840 via the project PRJDB3817 (\u003cstrong\u003eTable 1\u003c/strong\u003e).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 1\u003c/strong\u003e: Overview of data files/data sets.\u003c/p\u003e\n\u003ctable border=\"1\" cellpadding=\"0\" cellspacing=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"11.538461538461538%\"\u003e\n \u003cp\u003e\u003cstrong\u003eLabel\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"19.23076923076923%\"\u003e\n \u003cp\u003e\u003cstrong\u003eName of data file/data set\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"16.346153846153847%\"\u003e\n \u003cp\u003e\u003cstrong\u003eFile types\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003e(file extension)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"52.88461538461539%\"\u003e\n \u003cp\u003e\u003cstrong\u003eData repository and identifier (DOI or accession number)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"11.538461538461538%\"\u003e\n \u003cp\u003eData file 1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"19.23076923076923%\"\u003e\n \u003cp\u003e\u003cem\u003eP. insidiosum\u003c/em\u003e strain Pi-S,\u0026nbsp;whole genome shotgun sequencing project\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"16.346153846153847%\"\u003e\n \u003cp\u003eFASTA\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"52.88461538461539%\"\u003e\n \u003cp\u003eBBXB02000001 - BBXB02000840\u003c/p\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e"},{"header":"Limitations","content":"\u003cp\u003eWe used the PacBio long-read NGS platform to improve the genome quality of \u003cem\u003eP. insidiosum\u0026nbsp;\u003c/em\u003estrain Pi-S. Compared with the short-read NGS platforms (such as Illumina), genome sequencing employing the long-read platforms (such as PacBio) provides a more complete genome sequence but is a relatively more expensive technique and shows a higher base error rate.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cp\u003eDNA, Deoxyribonucleic acid\u003c/p\u003e\n\u003cp\u003egDNA, Genomic deoxyribonucleic acid\u003c/p\u003e\n\u003cp\u003eNCBI, National Center for Biotechnology Information\u003c/p\u003e\n\u003cp\u003eNGS, Next-generation sequencing\u003c/p\u003e\n\u003cp\u003eOD, Optical density\u003c/p\u003e\n\u003cp\u003eORF, Open reading frame\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study was approved by the Human Research Ethics Committee, Faculty of Medicine, Ramathibodi Hospital, Mahidol University (MURA2022/274).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe draft genome sequence of the \u003cem\u003eP. insidiosum\u003c/em\u003e strain Pi-S comprising 840 contigs (accession numbers BBXB02000001 - BBXB02000840) is available in the DDBJ database.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNone.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study obtained financial support from the Faculty of Medicine, Ramathibodi Hospital, Mahidol University [Grant number CF_65003] and the National Research Council of Thailand (NRCT) and Mahidol University [Grant number N42A650339]. The funders had no role in the design of the study and collection, analysis, and interpretation of data and in writing the manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026rsquo; contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eT.K. conceived the project. S.T., T.R., T.L., and W.Y. performed the experiments. W.K., P.P., S.T., T.R., and T.K. analyzed the data. W.K., P.P., T.R., S.T., and T.K. wrote the manuscript. All authors reviewed the manuscript. T.K. acquired the research funds.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgments\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e"},{"header":"References ","content":"\u003col\u003e\n\u003cli\u003eKittichotirat W, Krajaejun T. Application of Genome Sequencing to Study Infectious Diseases. J Infect Dis Antimicrob Agents. 2019;36:47\u0026ndash;58.\u003c/li\u003e\n\u003cli\u003eBentley DR, Balasubramanian S, Swerdlow HP, Smith GP, Milton J, Brown CG, et al. Accurate whole human genome sequencing using reversible terminator chemistry. Nature. 2008;456:53\u0026ndash;9.\u003c/li\u003e\n\u003cli\u003eRothberg J, Myers J. Semiconductor sequencing for life. J Biomol Tech. 2011;22:S41\u0026ndash;2.\u003c/li\u003e\n\u003cli\u003eRhoads A, Au KF. PacBio Sequencing and Its Applications. Genomics Proteomics Bioinformatics. 2015;13:278\u0026ndash;89.\u003c/li\u003e\n\u003cli\u003eLu H, Giordano F, Ning Z. Oxford Nanopore MinION Sequencing and Genome Assembly. Genomics Proteomics Bioinformatics. 2016;14:265\u0026ndash;79.\u003c/li\u003e\n\u003cli\u003eChitasombat MN, Jongkhajornpong P, Lekhanont K, Krajaejun T. Recent update in diagnosis and treatment of human pythiosis. PeerJ. 2020;8:e8555.\u003c/li\u003e\n\u003cli\u003eYolanda H, Krajaejun T. Global Distribution and Clinical Features of Pythiosis in Humans and Animals. J Fungi Basel Switz. 2022;8:182.\u003c/li\u003e\n\u003cli\u003eRujirawat T, Patumcharoenpol P, Lohnoo T, Yingyong W, Lerksuthirat T, Tangphatsornruang S, et al. Draft Genome Sequence of the Pathogenic Oomycete Pythium insidiosum Strain Pi-S, Isolated from a Patient with Pythiosis. Genome Announc. 2015;3:e00574-15.\u003c/li\u003e\n\u003cli\u003eKittichotirat W, Patumcharoenpol P, Rujirawat T, Lohnoo T, Yingyong W, Krajaejun T. Draft genome and sequence variant data of the oomycete Pythium insidiosum strain Pi45 from the phylogenetically-distinct Clade-III. Data Brief. 2017;15:896\u0026ndash;900.\u003c/li\u003e\n\u003cli\u003eKrajaejun T, Kittichotirat W, Patumcharoenpol P, Rujirawat T, Lohnoo T, Yingyong W. Data on whole genome sequencing of the oomycete Pythium insidiosum strain CBS 101555 from a horse with pythiosis in Brazil. BMC Res Notes. 2018;11:880.\u003c/li\u003e\n\u003cli\u003eKrajaejun T, Kittichotirat W, Patumcharoenpol P, Rujirawat T, Lohnoo T, Yingyong W. Draft genome sequence of the oomycete Pythium destruens strain ATCC 64221 from a horse with pythiosis in Australia. BMC Res Notes. 2020;13:329.\u003c/li\u003e\n\u003cli\u003ePatumcharoenpol P, Rujirawat T, Lohnoo T, Yingyong W, Vanittanakom N, Kittichotirat W, et al. Draft genome sequences of the oomycete Pythium insidiosum strain CBS 573.85 from a horse with pythiosis and strain CR02 from the environment. Data Brief. 2018;16:47\u0026ndash;50.\u003c/li\u003e\n\u003cli\u003eKrajaejun T, Kittichotirat W, Patumcharoenpol P, Rujirawat T, Lohnoo T, Yingyong W. Genome data of four Pythium insidiosum strains from the phylogenetically-distinct clades I, II, and III. BMC Res Notes. 2021;14:197.\u003c/li\u003e\n\u003cli\u003eRujirawat T, Patumcharoenpol P, Lohnoo T, Yingyong W, Kumsang Y, Payattikul P, et al. Probing the Phylogenomics and Putative Pathogenicity Genes of Pythium insidiosum by Oomycete Genome Analyses. Sci Rep. 2018;8:4135.\u003c/li\u003e\n\u003cli\u003eKrajaejun T, Rujirawat T, Lohnoo T, Yingyong W, Sae-Chew P, Reamtong O, et al. Secretome Profiling by Proteogenomic Analysis Shows Species-Specific, Temperature-Dependent, and Putative Virulence Proteins of Pythium insidiosum. J Fungi Basel Switz. 2022;8:527.\u003c/li\u003e\n\u003cli\u003eSae-Chew P, Rujirawat T, Kumsang Y, Payattikul P, Lohnoo T, Yingyong W, et al. Automated Cell-Free Multiprotein Synthesis Facilitates the Identification of a Secretory, Oligopeptide Elicitor-Like, Immunoreactive Protein of the Oomycete Pythium insidiosum. mSystems. 2020;5.\u003c/li\u003e\n\u003cli\u003eRotchanapreeda T, Sae-Chew P, Lohnoo T, Yingyong W, Rujirawat T, Kumsang Y, et al. Immunological Cross-Reactivity of Proteins Extracted from the Oomycete Pythium insidiosum and the Fungus Basidiobolus ranarum Compromises the Detection Specificity of Immunodiagnostic Assays for Pythiosis. J Fungi Basel Switz. 2021;7:474.\u003c/li\u003e\n\u003cli\u003eKrajaejun T, Lerksuthirat T, Garg G, Lowhnoo T, Yingyong W, Khositnithikul R, et al. Transcriptome analysis reveals pathogenicity and evolutionary history of the pathogenic oomycete Pythium insidiosum. Fungal Biol. 2014;118:640\u0026ndash;53.\u003c/li\u003e\n\u003cli\u003eKrajaejun T, Lohnoo T, Jittorntam P, Srimongkol A, Kumsang Y, Yingyong W, et al. Assessment of matrix-assisted laser desorption ionization-time of flight mass spectrometry for identification and biotyping of the pathogenic oomycete Pythium insidiosum. Int J Infect Dis. 2018;77:61\u0026ndash;7.\u003c/li\u003e\n\u003cli\u003eRujirawat T, Sridapan T, Lohnoo T, Yingyong W, Kumsang Y, Sae-Chew P, et al. Single nucleotide polymorphism-based multiplex PCR for identification and genotyping of the oomycete Pythium insidiosum from humans, animals and the environment. Infect Genet Evol. 2017;54:429\u0026ndash;36.\u003c/li\u003e\n\u003cli\u003eLohnoo T, Jongruja N, Rujirawat T, Yingyon W, Lerksuthirat T, Nampoon U, et al. Efficiency comparison of three methods for extracting genomic DNA of the pathogenic oomycete Pythium insidiosum. J Med Assoc Thai. 2014;97:342\u0026ndash;8.\u003c/li\u003e\n\u003cli\u003eShearman JR, Sonthirod C, Naktang C, Sangsrakru D, Yoocha T, Chatbanyong R, et al. Assembly of the durian chloroplast genome using long PacBio reads. Sci Rep. 2020;10:15980.\u003c/li\u003e\n\u003cli\u003eKoren S, Walenz BP, Berlin K, Miller JR, Bergman NH, Phillippy AM. Canu: scalable and accurate long-read assembly via adaptive k-mer weighting and repeat separation. Genome Res. 2017;27:722\u0026ndash;36.\u003c/li\u003e\n\u003cli\u003eChin C-S, Alexander DH, Marks P, Klammer AA, Drake J, Heiner C, et al. Nonhybrid, finished microbial genome assemblies from long-read SMRT sequencing data. Nat Methods. 2013;10:563\u0026ndash;9.\u003c/li\u003e\n\u003cli\u003eWaterhouse RM, Seppey M, Sim\u0026atilde;o FA, Manni M, Ioannidis P, Klioutchnikov G, et al. BUSCO Applications from Quality Assessments to Gene Prediction and Phylogenomics. Mol Biol Evol. 2018;35:543\u0026ndash;8.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"bmc-research-notes","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"resn","sideBox":"Learn more about [BMC Research Notes](http://bmcresnotes.biomedcentral.com)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/resn/default.aspx","title":"BMC Research Notes","twitterHandle":"@BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Pythium insidiosum, Pythiosis, Draft genome, Next-Generation Sequencing","lastPublishedDoi":"10.21203/rs.3.rs-2124426/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-2124426/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eObjectives:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cem\u003ePythium insidiosum\u003c/em\u003e\u0026nbsp;is the causative agent of pythiosis, a difficult-to-treat condition, in humans and animals worldwide. Biological information about this filamentous microorganism is sparse. Genomes of several \u003cem\u003eP. insidiosum\u003c/em\u003e strains were sequenced using the Illumina short-read NGS platform, producing incomplete genome sequence data. PacBio long-read platform was employed to obtain a better-quality genome of\u0026nbsp;\u003cem\u003ePythium insidiosum\u003c/em\u003e. The obtained genome data could promote basic research on the pathogen’s biology and pathogenicity.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData description:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003egDNA sample was extracted from the \u003cem\u003eP. insidiosum\u003c/em\u003e strain Pi-S for whole-genome sequencing by PacBio long-read NGS platform. Raw reads were assembled using CANU (v2.1), polished using ARROW (SMRT link version 5.0.1), aligned with the original raw PacBio reads using pbmm2 (v1.2.1), consensus sequence checked using ARROW, and gene predicted using Funannotate pipeline (v1.7.4). The genome completion was assessed using BUSCO (v4.0.2). As the result, 840 contigs (maximum length: 1.3 Mb; \u003cem\u003eN\u003c/em\u003e\u003csub\u003e50\u003c/sub\u003e: 229.9 Kb; \u003cem\u003eL\u003c/em\u003e\u003csub\u003e50\u003c/sub\u003e: 70) were obtained. Sequence assembly showed a genome size of 66.7 Mb (178x coverage; 57.2% G-C content) that contained 20,375 ORFs. A BUSCO-based assessment revealed 85.5% genome completion. All assembled contigs sequences have been deposited in the DDBJ database under the accession numbers BBXB02000001 - BBXB02000840.\u003c/p\u003e","manuscriptTitle":"PacBio long read-assembled draft genome of Pythium insidiosum strain Pi-S isolated from a Thai patient with pythiosis","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2023-03-28 10:05:14","doi":"10.21203/rs.3.rs-2124426/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Major revision","date":"2022-10-26T10:47:22+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2022-10-25T17:01:44+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"590382fa-966b-4656-a410-377681b21836","date":"2022-10-24T14:27:11+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2022-10-20T13:20:16+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2022-10-20T06:52:09+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Research Notes","date":"2022-10-20T04:57:10+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"bmc-research-notes","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"resn","sideBox":"Learn more about [BMC Research Notes](http://bmcresnotes.biomedcentral.com)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/resn/default.aspx","title":"BMC Research Notes","twitterHandle":"@BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"650e6ba5-2d12-4f3f-b72f-4622a383eb43","owner":[],"postedDate":"March 28th, 2023","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2023-10-17T14:57:40+00:00","versionOfRecord":{"articleIdentity":"rs-2124426","link":"https://doi.org/10.1186/s13104-023-06532-7","journal":{"identity":"bmc-research-notes","isVorOnly":false,"title":"BMC Research Notes"},"publishedOn":"2023-10-13 15:02:57","publishedOnDateReadable":"October 13th, 2023"},"versionCreatedAt":"2023-03-28 10:05:14","video":"","vorDoi":"10.1186/s13104-023-06532-7","vorDoiUrl":"https://doi.org/10.1186/s13104-023-06532-7","workflowStages":[]},"version":"v1","identity":"rs-2124426","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-2124426","identity":"rs-2124426","version":["v1"]},"buildId":"7rjqhiLT3MXkJMwkYKINL","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00
unpaywall
last seen: 2026-05-29T02:00:03.542394+00:00
License: CC-BY-4.0