{"paper_id":"1bbb83dd-1d40-4c9b-8df2-f5dd8cfcf437","body_text":"NanoPlasmiQC: Full plasmid sequencing with \nONT long-reads and automatic data analysis \n \nJulie Anne V. S. de Oliveira1,+,Vinson Ng1,+, Katharina Wolff1, Boas Pucker1,* \n1 Plant Biotechnology and Bioinformatics, Institute for Cellular and Molecular Botany (IZMB), \nUniversity of Bonn, Kirschallee 1, 53115 Bonn, Germany \n \n+ contributed equally \n* corresponding author:pucker@uni-bonn.de \n \nORCIDs: \nJAVSdO: 0000-0001-9930-6373 \nVN: 0009-0006-3941-9628 \nKW: 0000-0002-3616-0482 \nBP: 0000-0002-3321-7471 \n \nAbstract \nLong-read sequencing has shown a rapid technological development during the last years. It \nhas been established as the standard method for the sequencing of plant genomes and has \nalso gained importance for full plasmid sequencing. As Sanger sequencing has a limited read \nlength of about 1 kb, long read sequencing offers a great advantage, as the full plasmid can be \nsequenced in one read. \nHere, we present a cost-effective workflow to sequence full plasmids and compare the results \nagainst an expectation. The per plasmid cost of this workflow is determined by the number of \nplasmids investigated simultaneously, but can be lower than the price of a single Sanger \nsequencing reaction. We developed a workflow for automatic data processing, which allows us \nto complete sequencing and data analysis within a day. \n \nIntroduction \nSince the invention of chain termination sequencing (Sanger et al., 1977), often referred to as \n‘Sanger sequencing’ to honor the inventor, researchers have benefited greatly from unraveling \nthe sequence of DNA molecules. For almost 50 years, sequencing technologies have shown \nrapid development with massive parallel sequencing of the second generation and long reads of \nthe third generation representing the most important breakthroughs (Margulies et al., 2005; \nMetzker, 2010; Mardis, 2013, 2017; Pucker et al., 2022; de Oliveira et al., 2026). In plant \ngenomics, the long-read technologies offered by Pacific Biosciences (PacBio) and Oxford \nNanopore Technologies (ONT) are the most important sequencing methods (Marks et al., 2021; \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 3, 2026. ; https://doi.org/10.64898/2026.04.01.715842doi: bioRxiv preprint \n\n \n \nSchwacke et al., 2025; de Oliveira et al., 2026). In addition to numerous applications in \nagriculture (Zhou et al., 2020; Jayakodi et al., 2020; Walkowiak et al., 2020; Pucker et al., 2022), \nplant genome sequences are fundamental to unravel complex biosynthesis pathways leading to \nproducts with biotechnological or biomedical value (Cheng et al., 2025; Hakim et al., 2025). \n \nNanopore has also frequently been used to characterize the products of mutagenesis or \ngenome editing experiments to identify any off-target effects. T-DNA insertion lines, generated \nby randomly inserting a T-DNA into a plant genome, were fundamental for plant biology \nresearch, because homologous recombination does not work effectively in plants. While these \nlines form an important resource for the discovery of gene functions, a comprehensive genetic \ncharacterization is necessary to rule out additional events besides the desired gene knock-out. \nAn efficient way to characterize such lines is long-read sequencing combined with a local \nassembly for all T-DNA insertion loci (Pucker et al., 2021). A recently released pipeline allows \nthe automatic analysis of CRISPR/Cas editing sites based on long-reads (Chen et al., 2026). \n \nAssessing gene functions and biotechnological studies often involve the construction of \nplasmids. For many years, plasmids were validated by Sanger sequencing. It is believed that up \nto 40% of certain plasmid types with inverted tandem repeats might contain an unexpected \nmutation (Bai et al., 2025). Circuit-seq was developed to analyze plasmids with nanopore \nsequencing without prior knowledge about the plasmid sequence (Emiliani et al., 2022). Until \nrecently, plasmid sequencing relying only on long-reads was not recommended due to small \ninsertions/deletions caused by noisy long reads (Hernandez et al., 2024). However, with recent \ntechnological improvements leading to a raw read accuracy of about 99 %, full plasmid \nsequencing utilizing nanopore long-reads could be feasible now. Especially ambitious projects \ninvolving large plasmids adopted full plasmid sequencing as an attractive alternative to \ncombining numerous Sanger sequencing reactions. Full plasmid sequencing removes the need \nto design sequencing primers and ensures that the complete plasmid and not just the newly \ninserted DNA construct is checked. Costs of commercial offers for full plasmid sequencing often \nexceed 10 € per plasmid. Given that many biotechnological projects involve the generation and \nvalidation of numerous plasmids, costs for the sequencing add up. Sophisticated pooling \nsystems have been proposed to reduce these costs by mixing multiple plasmids into one \nsample without barcoding (Uematsu & Baskin, 2025). \n \nHere, we present a workflow for cost-effective and sustainable full plasmid analysis based on \nONT long-read sequencing. The steps of sample preparation, library construction, and \nnanopore sequencing are described. An automatic data analysis workflow implemented in \nPython generates user-friendly outputs that enable plasmid validation by life scientists. \n \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 3, 2026. ; https://doi.org/10.64898/2026.04.01.715842doi: bioRxiv preprint \n\n \n \nMaterials and Methods \nDNA preparations \nAll plasmids submitted for sequencing should have a concentration of about 50 ng/µL \n(NanoDrop measurement) and should be resuspended in TE buffer. Although differences in \nplasmid size are going to cause different sequencing depth, this approach makes it convenient \nfor users and the total coverage should be sufficiently high to resolve even plasmids with a \nlower concentration. Plasmids are pooled together, taking 1µL of each plasmid. While an \noptimization would be possible based on the size of the different plasmids, this approach \nensures efficient handling of large numbers of samples. The plasmid mix concentration is \nmeasured and diluted to a concentration of 10-15 ng/µL, as a total of 150 ng is to be used in the \nlibrary preparation. \n \nLong-read sequencing \nThe library preparation is done with 10 µL of pooled plasmids using the Rapid Sequencing Kit \nfrom ONT (SQK-RAD114). Sequencing is conducted on R10 flow cells on a PromethION after \ncompletion of plant genome sequencing projects. The nanopores recovered through a \nDNase-based washing step (EXP-WSH004) still offer sufficient capacity for the sequencing of \ndozens of plasmids. Sequencing is typically performed over night to ensure sufficient data are \navailable for the following analysis. Basecalling is done with dorado v1.4.0 (ONT) in high \naccuracy mode (HAC).  \n \nData analysis \nExpected sequences of all plasmids were collected in a FASTA file. Headers in this file were \ncleaned from invalid characters as part of the data analysis process. A read mapping with \nminimap2 v2.26-r1175 (Li, 2018) against all expected sequences was conducted with the flag \nsecondary=no in addition to default parameters. The resulting mapping was split per plasmid \nreference using samtools v1.19.2 (Li et al., 2009) and the reads per plasmid were extracted and \nconverted into a FASTQ file with samtools. A variant calling per plasmid reference was \nperformed with bcftools v1.19 (Danecek et al., 2021) using default parameters. Filtering of the \nextracted reads was done with seqkit v2.3.0 (Shen et al., 2024) to ensure that no redundant \nreads are contained. Due to very high coverage, a subsampling was required to reduce the \namount of reads subjected to the following assembly step. Miniasm v0.3-r179 (Li, 2016) was \nused for a de novo assembly per plasmid using default parameters. Racon v1.5.0 (Vaser et al., \n2017) was applied to polish the assembled plasmid sequence. Integrative Genomics Viewer \nv2.7.2 (Robinson et al., 2023) was used for manual inspection of results. The entire data \nanalysis workflow was wrapped into a Python script which is available via GitHub \n(https://github.com/bpucker/NanoPlasmiQC). \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 3, 2026. ; https://doi.org/10.64898/2026.04.01.715842doi: bioRxiv preprint \n\n \n \n \nResults & Discussion \nWe designed a workflow to enable cost-effective and sustainable plasmid sequencing with ONT \nlong-reads (Fig. 1). \n \nFig. 1: Workflow designed for full plasmid sequencing with ONT long-reads. \nSequencing output \nSequencing a mixture of plasmids on a re-used PromethION flow cell is expected to return low \noutput when compared against typical plant sequencing runs. However, we managed to get a \ntotal of 12.7 Gbp and 0.58 Gbp, respectively, in two test runs (Table 1). Plasmids typically have \na size of 5-20 kbp and can therefore be fully covered by individual reads. The limited size of \nplasmids results in a low N50 of 7.2 kbp and 2.6 kbp, respectively, in the two test runs. An \naverage sequencing depth of 10x per plasmid should be sufficient to validate the sequence of \nmost plasmids. Since the pooling of plasmids with different concentrations and sizes will result \nin biases, we aim for a coverage of >100x to ensure that even large plasmids are well captured \nin the resulting dataset. \n \nTable 1: Plasmid sequencing output statistics. \nRunID Output [Gbp] N50 [bp] Number of reads Average read length [bp] \nRP080 12.75 7219 3405632 3745 \nRP094 0.58 2584 487472 1199 \n \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 3, 2026. ; https://doi.org/10.64898/2026.04.01.715842doi: bioRxiv preprint \n\n \n \n \nImplementation of the plasmid validation tool \nProcessing of the sequencing reads in order to validate the accuracy of constructed plasmids \nwas automated with a Python script serving as wrapper for numerous analysis steps (Fig. 2). \nThe provided reference FASTA file is cleaned to ensure no problematic characters in the \nheaders can break downstream tools (1). Statistics of the provided FASTQ file are calculated for \ndocumentation purposes (2). Minimap2 is run to align ONT long-reads to the expected plasmid \nsequences with the secondary=no flag to ensure specific read mapping (3). Resulting SAM files \nare converted into BAM, sorted, and indexed (4). The BAM file is split by plasmid reference \nsequence with samtools to enable independent analysis of all plasmids in the downstream steps \n(5). The provided FASTA file with all plasmid reference sequences is also split accordingly and \nindexed with samtools (6). For each BAM file, the reads are extracted in FASTQ format with \nsamtools using the -F2308 flag and checked for duplicates with seqkit (7). The coverage per \nplasmid is calculated to enable subsampling to reach a target coverage (8). Variant calling is \nconducted with bcftools using the “-Q 10” and “-q 10” flags as well as downstream filtering with \nQUAL>20 and DP> a coverage cutoff (9). Based on the subsampled reads per plasmid, \nminimap2 and miniasm are used to de novo assemble the plasmid sequence (10). These \nassemblies are polished with racon (11). \n \n \n \nFig. 2: Plasmid sequencing data analysis workflow. \n \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 3, 2026. ; https://doi.org/10.64898/2026.04.01.715842doi: bioRxiv preprint \n\n \n \n \nProof of concept \nAs a proof of concept the previously published plasmid pBF3038 (Fang et al., 2011) was \nsequenced and assessed with the presented workflow. To illustrate the results obtained through \nthe automatic data analysis, the results of pBF3038 were visualized with IGV(Fig. 3). \n \n \nFig. 3: Screenshots of different results for pBF3038. A: Visualization of read mapping against \nthe expected sequence of pBF3038 in IGV. Additionally, the variant calling is included. B: \nPlasmid map of the pBF3038 sequence assembled with miniasm. \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 3, 2026. ; https://doi.org/10.64898/2026.04.01.715842doi: bioRxiv preprint \n\n \n \n \nThe mapping of the assigned long-reads shows that the entire length of the plasmid can be \ncovered in single reads at a high coverage. Additionally, it also shows a proof of concept, that \nthe reads can be assigned to the individual plasmids. Further, the resequencing and variant \ncalling revealed several point mutations within our sequenced copy of pBF3038. An annotation \nof the assembled plasmid was able to validate the assembly process and shows all key genetic \nelements of the plasmid were sequenced and assembled. \n \nDeclarations \nEthics approval and consent to participate \nNot applicable \n \nConsent for publication \nNot applicable \n \nAvailability of data and materials \nThe scripts developed for this study are available via GitHub (https://github.com/bpucker/NanoPlasmiQC). \nCompeting interests \nThe authors declare that they have no competing interests. \n \nFunding \nNot applicable \n \nAuthors' contributions \nJAVSdO, VN, KW, and BP planned the project. VN and KW prepared the plasmid mixtures prior \nto library preparation. JAVSdO conducted library preparation, sequencing, and bioinformatic \nanalyses. BP supervised the work, developed the tool for automatic data analysis, conducted \nbioinformatic analyses, and wrote the manuscript. All authors approved the final version of the \nmanuscript and agreed to its submission. \n \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 3, 2026. ; https://doi.org/10.64898/2026.04.01.715842doi: bioRxiv preprint \n\n \n \nAcknowledgements \nThis work was supported by the de.NBI Cloud within the German Network for Bioinformatics \nInfrastructure (de.NBI) and ELIXIR-DE (Forschungszentrum Jülich and W-de.NBI-001, \nW-de.NBI-004, W-de.NBI-008, W-de.NBI-010, W-de.NBI-013, W-de.NBI-014, W-de.NBI-016, \nW-de.NBI-022). We thank all members of the Plant Biotechnology and Bioinformatics group for \ntheir support and feedback during the process. \nReferences \nBai X, Hong JF, Yu S, Hu DY, Chen AY, Rich CA, Shi SJ, Xu SY, Croucher DM, Müssar KJ, et al. \n2025. Prevalence of errors in lab-made plasmids across the globe. : 2024.06.17.596931. doi: \n10.1101/2024.06.17.596931.  \nChen Y, Gao X-H, Vichas A, Wang J, Golhar R, Neuhaus I. 2026. ALPINE: A Scalable Pipeline \nfor Comprehensive Classification of Gene-Editing Outcomes from Long-Read Amplicon \nSequencing. : 2026.03.27.714831. doi: 10.64898/2026.03.27.714831.  \nCheng L-T, Wang Z-L, Zhu Q-H, Ye M, Ye C-Y. 2025. A long road ahead to reliable and \ncomplete medicinal plant genomes. Nature Communications 16: 2150. doi: \n10.1038/s41467-025-57448-8.  \nDanecek P, Bonfield JK, Liddle J, Marshall J, Ohan V, Pollard MO, Whitwham A, Keane T, \nMcCarthy SA, Davies RM, et al. 2021. Twelve years of SAMtools and BCFtools. GigaScience \n10: giab008. doi: 10.1093/gigascience/giab008.  \nEmiliani FE, Hsu I, McKenna A. 2022. Multiplexed Assembly and Annotation of Synthetic \nBiology Constructs Using Long-Read Nanopore Sequencing. ACS Synthetic Biology 11: \n2238–2246. doi: 10.1021/acssynbio.2c00126.  \nFang F, Salmon K, Shen MWY, Aeling KA, Ito E, Irwin B, Tran UPC, Hatfield GW, Da Silva NA, \nSandmeyer S. 2011. A vector set for systematic metabolic engineering in Saccharomyces \ncerevisiae. Yeast 28: 123–136. doi: 10.1002/yea.1824.  \nHakim SE, Choudhary N, Malhotra K, Peng J, Bültemeier A, Arafa A, Friedhoff R, Bauer M, \nEikenberg J, Witte C-P, et al. 2025. Phylogenomics and metabolic engineering reveal a \nconserved gene cluster in Solanaceae plants for withanolide biosynthesis. Nature \nCommunications 16: 6367. doi: 10.1038/s41467-025-61686-1.  \nHernandez SI, Berezin C-T, Miller KM, Peccoud SJ, Peccoud J. 2024. Sequencing Strategy to \nEnsure Accurate Plasmid Assembly. ACS Synthetic Biology 13: 4099–4109. doi: \n10.1021/acssynbio.4c00539.  \nJayakodi M, Padmarasu S, Haberer G, Bonthala VS, Gundlach H, Monat C, Lux T, Kamal N, \nLang D, Himmelbach A, et al. 2020. The barley pan-genome reveals the hidden legacy of \nmutation breeding. Nature 588: 284–289. doi: 10.1038/s41586-020-2947-8.  \nLi H. 2016. Minimap and miniasm: fast mapping and de novo assembly for noisy long \nsequences. Bioinformatics 32: 2103–2110. doi: 10.1093/bioinformatics/btw152.  \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 3, 2026. ; https://doi.org/10.64898/2026.04.01.715842doi: bioRxiv preprint \n\n \n \nLi H. 2018. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics 34: \n3094–3100. doi: 10.1093/bioinformatics/bty191.  \nLi H, Handsaker B, Wysoker A, Fennell T, Ruan J, Homer N, Marth G, Abecasis G, Durbin R, \n1000 Genome Project Data Processing Subgroup. 2009. The Sequence Alignment/Map format \nand SAMtools. Bioinformatics 25: 2078–2079. doi: 10.1093/bioinformatics/btp352.  \nMardis ER. 2013. Next-generation sequencing platforms. Annual Review of Analytical Chemistry \n6: 287–303. doi: 10.1146/annurev-anchem-062012-092628.  \nMardis ER. 2017. DNA sequencing technologies: 2006–2016. Nature Protocols 12: 213–218. \ndoi: 10.1038/nprot.2016.182.  \nMargulies M, Egholm M, Altman WE, Attiya S, Bader JS, Bemben LA, Berka J, Braverman MS, \nChen Y-J, Chen Z, et al. 2005. Genome sequencing in microfabricated high-density picolitre \nreactors. Nature 437: 376–380. doi: 10.1038/nature03959.  \nMarks RA, Hotaling S, Frandsen PB, VanBuren R. 2021. Representation and participation \nacross 20 years of plant genome sequencing. Nature Plants 7: 1571–1578. doi: \n10.1038/s41477-021-01031-8.  \nMetzker ML. 2010. Sequencing technologies — the next generation. Nature Reviews Genetics \n11: 31–46. doi: 10.1038/nrg2626.  \nde Oliveira JAVS, Choudhary N, Meckoni SN, Nowak MS, Hagedorn M, Pucker B. 2026. \nCookbook for plant genome sequences. BMC Genomics. doi: 10.1186/s12864-026-12623-z.  \nPucker B, Irisarri I, Vries J de, Xu B. 2022. Plant genome sequence assembly in the era of long \nreads: Progress, challenges and future directions. Quantitative Plant Biology 3: e5. doi: \n10.1017/qpb.2021.18.  \nPucker B, Kleinbölting N, Weisshaar B. 2021. Large scale genomic rearrangements in selected \nArabidopsis thaliana T-DNA lines are caused by T-DNA insertion mutagenesis. BMC Genomics \n22: 599. doi: 10.1186/s12864-021-07877-8.  \nRobinson JT, Thorvaldsdottir H, Turner D, Mesirov JP. 2023. igv.js: an embeddable JavaScript \nimplementation of the Integrative Genomics Viewer (IGV). Bioinformatics 39: btac830. doi: \n10.1093/bioinformatics/btac830.  \nSanger F, Nicklen S, Coulson AR. 1977. DNA sequencing with chain-terminating inhibitors. \nProceedings of the National Academy of Sciences of the United States of America 74: \n5463–5467. doi: 10.1073/pnas.74.12.5463.  \nSchwacke R, Bolger ME, Usadel B. 2025. PubPlant – a continuously updated online resource \nfor sequenced and published plant genomes. Frontiers in Plant Science 16. doi: \n10.3389/fpls.2025.1603547.  \nShen W, Sipos B, Zhao L. 2024. SeqKit2: A Swiss army knife for sequence and alignment \nprocessing. iMeta 3: e191. doi: 10.1002/imt2.191.  \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 3, 2026. ; https://doi.org/10.64898/2026.04.01.715842doi: bioRxiv preprint \n\n \n \nUematsu M, Baskin JM. 2025. Barcode-free multiplex plasmid sequencing using Bayesian \nanalysis and nanopore sequencing (MM Zambrano and D Weigel, Eds). eLife 12: RP88794. doi: \n10.7554/eLife.88794.  \nVaser R, Sović I, Nagarajan N, Šikić M. 2017. Fast and accurate de novo genome assembly \nfrom long uncorrected reads. Genome Research 27: 737–746. doi: 10.1101/gr.214270.116.  \nWalkowiak S, Gao L, Monat C, Haberer G, Kassa MT, Brinton J, Ramirez-Gonzalez RH, \nKolodziej MC, Delorean E, Thambugala D, et al. 2020. Multiple wheat genomes reveal global \nvariation in modern breeding. Nature 588: 277–283. doi: 10.1038/s41586-020-2961-x.  \nZhou Y, Chebotarov D, Kudrna D, Llaca V, Lee S, Rajasekar S, Mohammed N, Al-Bader N, \nSobel-Sorenson C, Parakkal P, et al. 2020. A platinum standard pan-genome resource that \nrepresents the population structure of Asian rice. Scientific Data 7: 113. doi: \n10.1038/s41597-020-0438-2.  \n \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 3, 2026. ; https://doi.org/10.64898/2026.04.01.715842doi: bioRxiv preprint","source_license":"CC-BY-4.0","license_restricted":false}