{"paper_id":"800490f1-9c43-4d67-a23f-7c69993ae2cc","body_text":"Genome Report: Pseudomolecule-scale genome assemblies of\nDrepanocaryum sewerzowii and Marmoritis complanata\nSamuel J. Smit 1,*\nCaragh Whitehead 1\nSally R. James 2\nDaniel C. Jeffares 3\nGrant Godden 4\nDeli Peng 5, 6\nHang Sun 7\nBenjamin R. Lichman 1,*\n1 Centre for Novel Agricultural Products, Department of Biology, University of York, York, UK\n2 Bioscience T echnology Facility, Department of Biology, University of York, York, UK\n3 York Biomedical Research Institute, Department of Biology, University of York, York, UK\n4 Florida Museum of Natural History, University of Florida, Gainesville, FL, USA\n5 School of Life Science, Yunnan Normal University, Kunming, Yunnan, China\n6 Key Laboratory of Yunnan for Biomass Energy and Biotechnology of Environment, Yunnan\nNormal University, Kunming, China\n7 Key Laboratory for Plant Diversity and Biogeography of East Asia/Yunnan Key Laboratory\nfor Integrative Conservation of Plant Species with Extremely Small Populations, Kunming\nInstitute of Botany, Chinese Academy of Sciences, Kunming, China\nKeywords: Lamiaceae, Nepetinae, chromosome-level assembly, Nanopore sequencing,\nHi-C sequencing\nAbstract\nThe Nepetoideae, a subfamily of Lamiaceae (mint family), is rich in aromatic plants, many of\nwhich are sought after for their use as flavours and fragrances or for their medicinal\nproperties. Here we present genome assemblies for two species in Nepetiodeae:\nDrepanocaruym sewerzowii and Marmoritis complanata. Both assemblies were generated\nusing Oxford Nanopore Q20+ reads with contigs anchored to nine pseudomolecules that\nresulted in 335 Mb and 305 Mb assemblies, respectively, and BUSCO scores above 95% for\nboth the assembly and annotation. We furthermore provide a species tree for the Lamiaceae\nusing only genome derived gene models, complementing existing transcriptome and\nmarker-based phylogenies.\n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted April 28, 2024. ; https://doi.org/10.1101/2024.04.23.590777doi: bioRxiv preprint \n\nIntroduction\nThe mint family (Lamiaceae) is the sixth largest plant family with a number of species\nregarded as important for medicinal, aromatic and ornamental properties (Harley, R, M et al.\n2004; Zhao et al. 2021; Rose et al. 2022). Within the Lamiaceae, species from the\nNepetoideae are renowned for the accumulation of terpenoids, with tissues used for the\nextraction of essential oils or as traditional herbal medicines (Wink 2003; Frezza et al. 2019).\nThe clade includes widely recognised aromatic species such as mint, lavender, lemon balm\nand catnip; the volatile terpenoids produced by these plants are responsible for their\ncharacteristic fragrances. The ethnobotanical and commercial relevance of this plant family\nhas resulted in considerable scientific interest, including genome assemblies for 36 species\nat the time of writing (“Published Plant Genomes”).\nHere we present the genome assemblies for two Nepetoideae species, namely\nDrepanocaryum sewerzowii (Regel) Pojark. and Marmoritis complanata (Dunn)\nA.L.Budantzev. M. complanata is endemic to the subnival band of the Himalaya-Hengduan\nMountains, a unique arctic-alpine region recognised as a biodiversity hotspot (Myers et al.\n2000; Sun et al. 2017). This unique habitat necessitates careful control of seed germination\nto ensure survival (Peng et al. 2018). M. complanata and other species of the genus are also\nused as traditional herbal medicines to treat a variety of ailments that include digestive,\nreproductive, musculoskeletal and skin disorders (Zaman et al. 2022). D. sewerzowii is\nnative to a region that ranges from Iran to Central Asia and Pakistan and is the sole\nrepresentative of this genus (Serpooshan et al. 2018).\nThese two species are part of the Nepetinae, a subtribe of the mint family (Lamiaceae,\nsubfamily Nepetoideae, tribe Mentheae) that consists of 375 species and 9-12 genera of\nwhich Nepeta L. is considered the type genus encompassing 200-300 species. Other genera\nin this subfamily include Dracocephalum L., Hymenocrater Fisch. & C.A. Mey., Lophanthus\nAdans., Agastache Clayton ex Gronov. and Schizonepeta (Benth.) Briq. (Serpooshan et al.\n2018; Rose et al. 2023). The phylogenetic relationship of M. complanata and D. sewerzowii\nrelative to N. cataria L., N. racemosa Lam., A. rugosa (Fisch. & C.A. Mey.) Kuntze and S.\ntenuifolia (Benth.) Briq. is what prompted our efforts to assemble these genomes. We have\nbeen exploring the evolutionary, genomic and enzymatic innovations of monoterpenoid\nbiosynthesis in these species (Lichman et al. 2019, 2020; Hernández Lozada et al. 2022; Liu\net al. 2023). However, the available genomic resources provide limited taxonomic coverage.\nThe genome assemblies presented here will allow us to further explore the evolutionary\ninnovations that have impacted terpenoid biosynthesis in the mint family.\n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted April 28, 2024. ; https://doi.org/10.1101/2024.04.23.590777doi: bioRxiv preprint \n\nMethods and Materials\nPlant growth conditions\nD. sewerzowii seeds were obtained from the Millennium Seed Bank at the Royal Botanic\nGardens, Kew (serial no. 0694027). M. complanata seeds were collected from Puyong Pass\nShangri-la County, Yunnan Province, SW China (99°55′E, 28°24′N), 4620 m a.s.l (Peng et al.\n2018). Seeds were germinated on 1% water agar in a growth room set to 16 h day length,\ntemperature of 20 (±2) °C, relative humidity of 60% (±10%) and a NS12 light spectrum at\n120 µmol m -2s-1 PPFD using Valoya L28 LED lights (Helsinki, Finland). Once a radical\nemerged, the seedlings were transferred to 7 cm square pots containing Levington Advance\nSeed and Modular FS2 (ICL Professional Horticulture) seedling soil that was pre-treated with\nCalypso (Bayer). Once established, a single individual was selected and maintained as a\nclonal population by propagation using cuttings.\nGenome size estimation by FCM\nGenome size estimations were performed through flow cytometry (FCM) using the method of\n(Dolezel et al. 2007). Briefly, the LB01 buffer was used together with N. cataria tissue to\nprepare a reference standard with a previously reported genome size (Mint Evolutionary\nGenomics Consortium 2018). A CytoFLEX LX (Beckman Coulter) flow cytometer with a 561\nnm excitation laser, 610/20 emission filter and a flow rate of 30 µL/min was used. The\nthreshold was set to 488 nm forward scatter to exclude instrument noise and background\nsignal from the buffer.\nNucleic acid isolation\nHMW DNA isolation and sequencing\nHigh molecular weight (HMW) DNA was extracted in duplicate from ~1 g of young leaf tissue\nusing the Nucleobond HMW DNA Extraction kit (Macherey-Nagel, Germany). HMW DNA\npurity and concentration was assessed by Nanodrop and Qubit, whereafter the extractions\nwere combined. Small fragment DNA elimination was performed with the Circulomics short\nread eliminator kit (PacBio). Briefly, an equal volume of SRE reagent was added to the\nsample, and this was centrifuged for 1 h at 12,000 x g. The pellet was washed with 70%\nethanol before resuspending in TE buffer with low EDTA. DNA quality and quantity was\nassessed with a nanodrop spectrophotometer (Thermo Fischer Scientific), Agilent\nT apestation (running genomic DNA screentape) and Qubit fluorimeter (Invitrogen).\nSequencing was performed with the ligation sequencing kit SQK-LSK114 (Oxford Nanopore\nT echnologies), as per the manufacturer's guidelines, with limited modifications; namely\nextending the reaction times for end preparation to 30 min at each temperature, and\n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted April 28, 2024. ; https://doi.org/10.1101/2024.04.23.590777doi: bioRxiv preprint \n\nextending adapter ligation steps to an hour). Sequencing was performed on a single\npromethION FLO-PRO114 flowcell (Oxford Nanopore T echnologies) per species, with\nnuclease flush and sample reload steps performed every 24 h through the run time. For M.\ncomplanata two additional runs using the SQK-LSK112 ligation sequencing kit (Oxford\nNanopore T echnologies) and FLO-MIN112 minION flowcells (Oxford Nanopore\nT echnologies) were performed.\nSuper accuracy base calling was performed using guppy (Oxford Nanopore T echnologies)\nversion 6.1.5 for D. sewerzowii and version 6.3.9 for M. complanata. Read length and quality\nwas assessed using Nanoplot (De Coster and Rademakers 2023). D. sewerzowii reads were\nfiltered for a 10 kb minimum length using Nanofilt (De Coster et al. 2018). For M. complanata\nwe combined all reads from the promethION and minION runs and then filtered using\nNanofilt (De Coster et al. 2018) with a 3 kb length and Q15 quality cutoff .\nGenomic DNA isolation and Illumina sequencing\nGenomic DNA (gDNA) was extracted from 100 mg of young leaf tissue, in duplicate, using a\nCTAB extraction method (Doyle and Doyle 1990) and treated with RNAse A. Removal of\nRNA was confirmed through gel electrophoresis followed by gDNA quality and quantity\nassessment with a nanodrop spectrophotometer and a Qubit fluorometer (Invitrogen). A total\nof 508 ng and 752 ng of gDNA for D. sewerzowii and M. complanata, respectively, was sent\nfor library preparation and paired-end Illumina sequencing with Novogene (Cambridge, UK).\nRNA isolation and sequencing\nRNA was extracted from 80-100 mg of tissue with the Direct-Zol RNA extraction kit (Zymo\nResearch, CA, USA) as per the manufacturer guidelines. For D. sewerzowii young and\nmature leaves, closed and open flowers and stems were used. For M. complanata root,\nyoung and mature leaf and stem tissues were used. RNA quality was assessed with an\nAgilent bioanalyzer. Library preparation and paired-end Illumina sequencing was performed\nby Novogene (Cambridge, UK).\nHi-C sequencing\nFreshly harvested young leaf leaf tissue was fixed in 1% formaldehyde and washed as per\nthe Phase Genomics (Seattle, WA, USA) sample preparation protocol. Following fixation, the\ntissue was flash frozen in liquid nitrogen and homogenised using a tissue lyser. The Hi-C\nlibraries were prepared and sequenced by Phase Genomics (Seattle, WA, USA).\n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted April 28, 2024. ; https://doi.org/10.1101/2024.04.23.590777doi: bioRxiv preprint \n\nGenome Assembly\nFiltered nanopore reads for the respective genomes were used for assembly and error\ncorrection. Both species were first assembled using Flye (Lin et al. 2016; Kolmogorov et al.\n2019) (--iterations 0 and --nano-hq flags). M. complanata was also assembled with NECAT\n(Chen et al. 2021) using the default configuration file settings. Our error correction pipeline\nentailed polishing with long reads by two rounds of RACON (Vaser et al. 2017), with reads\nmapped using minimap2 (Li 2018), followed by two rounds of MEDAKA (medaka: Sequence\ncorrection provided by ONT Research 2018) polishing. Short reads were mapped using\nbwa-mem (Li 2013) and duplicate reads marked using Picard (Picard toolkit 2019) prior to\ntwo iterative rounds of polishing with Pilon (Walker et al. 2014).\nThe M. complanata Flye and NECAT assemblies were merged with Quickmerge\n(Chakraborty et al. 2016; Solares et al. 2018) due to the low N50 scores. The overlap cutoff\n(-c flag) was five and the length cutoff (-l) was 100,000 with the NECAT assembly used as\nthe query. The NECAT-Flye merged assembly underwent another two rounds of short read\nerror correction using Pilon. For M. complanata we purged the merged assembly of haplotigs\nprior to HiC scaffolding while D. serwerzowii was purged after HiC scaffolding. Haplotig\npurging was performed using the purge haplotigs pipeline (Roach et al. 2018). Contigs were\nscaffolded into pseudomolecules by Phase Genomics (Seattle, WA, USA) using the Proximo\nGenome Scaffolding Platform. Contiguity and completeness was assessed throughout the\nassembly pipeline using BUSCO (Benchmarking for University Single Copy Orthologs)\nv5.4.2 with the embryophyta_odb10 dataset (Manni et al. 2021).\nGenome annotation\nRepeats and transposable elements were annotated using the Earl Grey v3.2 (Baril et al.\n2023, 2024) pipeline with default settings followed by softmasking of the repeats using the\nmaskfasta function of bedtools. The BRAKER3 pipeline (v3.0.6) (Stanke et al. 2006, 2008;\nGotoh 2008; Iwata and Gotoh 2012; Buchfink et al. 2015; Hoff et al. 2016, 2019; Kovaka et\nal. 2019; Pertea and Pertea 2020; Brůna et al. 2021; Bruna et al. 2024) was used to predict\ngene models using mRNA and protein evidence. For protein evidence we generated a\nrepresentative database from 52 Mint species (48 Lamiaceae and four from Lamiales\nfamilies) using the transcriptomes from (Mint Evolutionary Genomics Consortium 2018).\nMMseqs2 (Steinegger and Söding 2017) was used to remove identical sequences from the\ndatabase. For mRNA evidence we aligned RNAseq reads from the different tissues using\nSTAR (Dobin et al. 2013) with default settings and the “--outSAMstrandField intronMotif” flag.\nThe respective bam outputs were merged using samtools (Danecek et al. 2021) and used as\ninput for BRAKER3. The BRAKER annotation output was reformatted to GFF3 using AGAT\n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted April 28, 2024. ; https://doi.org/10.1101/2024.04.23.590777doi: bioRxiv preprint \n\n(Dainat et al. 2023) followed by extraction and translation of the longest open-reading for\neach predicted coding sequence. Annotation completeness was assessed using BUSCO\n(Manni et al. 2021) in protein mode with the embryophyta_odb10 dataset.\nSpecies tree and macrosynteny analysis\nMarkerminer (Chamala et al. 2015) was used to identify single-copy genes using predicted\ncoding genes from representative Lamiaceae genomes (Sup. T able 1) and Paulownia\nfortunei (Seem.) Hemsl. as an outgroup. Genes present in 26 of the 27 species were\nincluded. The MAFFT alignments generated as part of the Markerminer pipeline were\ntrimmed for gaps using using the gappyout algorithm of trimAl v1.4.1 (Capella-Gutiérrez et\nal. 2009) and concatenated into a supermatrix with partitions using the catfasta2phyml script\n(https://github.com/nylander/catfasta2phyml). A species-tree was inferred by maximum\nlikelihood with partition models (Chernomor et al. 2016) using IQ-TREE 2 (Minh et al. 2020)\nwith ModelFinder (Kalyaanamoorthy et al. 2017), ultrafast bootstraps (UFBoot2, X1000)\n(Hoang et al. 2018), and SH-aLRT supports (X1000) (Guindon et al. 2010). In addition, a\nspecies tree using protein sequences was inferred using the STAG (Species Tree inference\nfrom All Genes) method of Orthofinder (Emms and Kelly 2015, 2017, 2018, 2019). Pairwise\nmacrosynteny analyses were performed against A. rugosa (Park et al. 2023) and S.\ntenuifolia (Liu et al. 2023) using the JCVI (T ang et al. 2015) implementation of MCScan\n(T ang et al. 2008). MCScan orthologs were identified in full mode with predicted protein\nsequences and default settings.\nExpression analysis\nRNAseq read alignments were evaluated with STAR (Dobin et al. 2013) and assed with\nqualimap (García-Alcalde et al. 2012; Okonechnikov et al. 2016). Qualimap reports were\naggregated with MultiQC (Ewels et al. 2016). Expression counts as transcripts per million\n(TPM) were generated using Salmon (Patro et al. 2017). The transcript index for Salmon\nwas generated using the full set of predicted coding sequences from BRAKER3.\nResults and Discussion\nChromosome level assemblies\nWe sequenced the genomes for D. sewerzowii and M. complanata using Oxford Nanopore\nlong reads and Proximo HiC scaffolding (Phase Genomics) resulting in two\nchromosome-level assemblies. A total of 99.24 Gb of super accurate nanopore reads were\ngenerated for D. sewerzowii with 80 Gb of reads being greater than 10 kb at a mean read\nquality (Q-score) of 16.6. The size filtered reads provided 242X coverage at an estimated\n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted April 28, 2024. ; https://doi.org/10.1101/2024.04.23.590777doi: bioRxiv preprint \n\ngenome size of 330 Mb, as determined through FCM (Supl. Fig. 1). The initial Flye assembly\nresulted in 472 contigs, an N50 of 17 Mb, a total assembly length of 333.75 Mb and a\nBUSCO score of 98.7%. Polishing with long and short reads reduced the number of contigs\nto 134 and assembly size to 332.85 Mb while maintaining a N50 of 17 Mb. The BUSCO\nscore increased slightly to 98.8% after polishing. HiC scaffolding orientated the assembly to\nnine pseudomolecules (Supl. Fig 2A), which is in agreement with the chromosome counts\nreported by Bordbar (2023). The nine pseudomolecules contained 97.6% of the contigs,\nrepresenting 324.87 Mb of the total assembly at a N50 of 35.2 Mb and L50 of 5 (T able 1).\nFigure 1. Circos plots for the genome assemblies of D. sewerzowii and M. complanata\ndepicting density (1 Mb bins) of genes, total repeats, gypsy and copia elements along the 9\npseudomolecules.\nThe M. complanata genome size was estimated at 337 Mb using FCM (Supl. Fig. 1). We\nobtained 59.3 Gb of reads after length and quality filtering, providing 176X coverage with a\nmean Q-score of 18. We tried various different read filtering cutoffs for both length and\nquality with all attempts using Flye failing to reach a N50 greater than ~335 kb. After\npolishing the best Flye assembly was 420 Mb in size with a N50 of 335 kb, 3001 contigs and\na BUSCO score of 98.5%, of which 17.1% were duplicated. NECAT resulted in a more\ncontiguous genome assembly of 457 Mb with a N50 of 1 Mb, 869 contigs and 98.6%\nBUSCO, of which 36.7% were duplicated. The inflated genome size and high number of\nduplicate BUSCO genes suggested that the fragmented assemblies contained a high\nnumber of haplotigs (contigs of a single haplotype), that would artificially inflate genome size.\n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted April 28, 2024. ; https://doi.org/10.1101/2024.04.23.590777doi: bioRxiv preprint \n\nTable 1. Assembly and annotation metrics.\nD. sewerzowii M. complanata\nAssembly Statistics\nAssembly size (Mb) 332.85 305.55\nNumber of\npseudomolecules\n9 9\nN50 (Mb) 35.13 27.69\nL50 5 5\nL90 9 28\nGC% 38.56 37.45\nNumber of Ns 2600 18 800\nAnnotation Statistics\nAssembly BUSCO*\nn=1440\nC: 99.0 %\nS: 96.3 %\nD: 2.7 %\nC: 95.7 %\nS: 85.7 %\nD: 10.0 %\nAnnotation BUSCO*\nn=1440\nC: 95.0 %\nS: 92.1 %\nD: 2.9 %\nC: 95.3 %\nS: 86.1 %\nD: 9.2 %\nPredicted coding genes 24 221 25 080\nPredicted proteins 26 989 28 384\nPercentage repeats T otal: 62%\nDNA: 2.38 %\nLINE: 0.53%\nLTR: 40.61%\nT otal: 53%\nDNA: 4.43%\nLINE: 2.92%\nLTR: 25.83%\n* Complete ( C), Single (S), Duplicated (D)\nIn an attempt to increase the continuity of the assembly (N50 score) we merged the Flye and\nNECAT assemblies. The NECAT assembly had fewer contigs and greater N50 and was\ntherefore selected to be the query genome with the Flye assembly used to improve the\nquery genome. We evaluated the impact of haplotig purging before and after merging. Each\nassembly was purged of haplotigs prior to merging and compared to a merged assembly\nthat was purged as the final step. In each iteration we polished twice with short reads after\nmerging. The merging increased the N50 to 3 Mb regardless of when we purged the\nhaplotigs. The timing of the purging step had a large impact on the number of contigs\ntogether with a minor impact on the duplicated BUSCOs. Merging, polishing and then\n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted April 28, 2024. ; https://doi.org/10.1101/2024.04.23.590777doi: bioRxiv preprint \n\npurging the haplotigs resulted in the most contiguous assembly (305.6 Mb) with the fewest\nnumber of contigs (338) and a BUSCO score of 95.6 %. HiC scaffolding assembled the\ncontigs into nine pseudomolecules (Supl. Fig 2B), which is in agreement with karyotype\ninformation (Sun 2016), totaling 258 Mb (85% of the total assembly). The pseudomolecules\nhad a BUSCO score of 91.3% with the total assembly having a BUSCO of 95.7% (T able 1).\nPseudomolecule termini were manually inspected for presence of the TTTAGGG telomeric\nrepeat. Seven of the D. sewerzowii pseudomolecules contained this repeat on at least one\nend with chr. 4 and 7 having it on both ends. For M. complanata we found this repeat on six\npseudomolecules with chr. 7 and 8 having it on both ends. The presence of this repeat on\nboth ends indicates a telomere to telomere assembly for these chromosomes.\nRepeat and genome annotations\nRepeat annotation revealed that 62% of the D. sewerzowii genome and 53% of the M.\ncomplanata genome are repeats (T able 1). The largest portion of the repeats were long\nterminal repeats (LTR), occupying 40.6% and 25.8% of the respective genomes.\nSubsequent to repeat masking our gene annotation, using ab initio, protein and mRNA\npredictions, resulted in 24,221 and 25,080 gene regions that encode for 26,989 and 28,384\nproteins for the respective genomes. BUSCO analysis of the primary isoforms was 95% for\nboth genomes. Gene and repeat density showed an inverse relationship along the\nchromosomes (Figure 1).\nThe RNAseq data we produced found evidence for expression of the majority of genes.\nRNAseq reads mapped to gene models showed that 85% (22,794/26,815) of the genes were\nexpressed in at least one tissue type for D. sewerzowii and 88% (24,919/28,384) of the\ngenes in M. complanata. Expression matrices as transcripts per million (TPM) are available\nin Supplementary T ables 3 and 4.\nPairwise macrosynteny and species tree\nWe compared our assemblies to the closest relatives with pseudomolecule assemblies,\nnamely A. rugosa (Park et al. 2023) and S. tenuifolia (Liu et al. 2023) (Figure 2). A. rugosa\nhas a 9 chromosome assembly with macrosynteny revealing that the overall genome\nstructures for both D. sewerzowii and M. complanata are similar to this species.\nMacrosynteny revealed a number of chromosome fusion events in S. tenuifolia relative to the\nother three genomes. Although only 85% of the M. complantum contigs were anchored to\npseudomolecules the overall structure of the chromosomes (relative to A. rugosa and D.\nsewerzowii) and presence of large syntenic blocks indicate a reasonably complete assembly.\n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted April 28, 2024. ; https://doi.org/10.1101/2024.04.23.590777doi: bioRxiv preprint \n\nFigure 2. Pairwise macrosynteny analysis of the assembled genomes relative to closely\nrelated species with chromosome level assemblies. Conserved collinear blocks are linked by\nthe grey lines.\nThe phylogenetic relationships of the Lamiaceae have been reported using plastid, nuclear\nand transcriptome approaches. The species trees presented in Figure 3 used genome\nderived gene models for phylogenomic inference, complementing existing species trees\n(Serpooshan et al. 2018; Mint Evolutionary Genomics Consortium 2018; Rose et al. 2022,\n2023). The STAG species-tree used multi-copy gene families (i.e. orthogroups) predicted by\nOrthofinder using protein sequences. The consensus tree in Fig 3A shows internal bipartition\nsupport for 5,296 orthogroups in which all species are present. The ML tree (Fig. 3B) was\ninferred from a single-copy gene supermatrix totaling 340,706 nucleotide sites with all but\ntwo branches showing above 98% support for both ultrafast bootstraps and SH-aLRT . The\ntwo branches indicated by the asterisk were not well supported, bootstrap and SH-aLRT\n<85%.\n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted April 28, 2024. ; https://doi.org/10.1101/2024.04.23.590777doi: bioRxiv preprint \n\nFigure 3. Species-trees inferred with Lamiaceae genome derived gene models. (A) STAG\nspecies-tree inferred with Orthofinder protein orthogroups. Support values show the\nproportion of trees at which the internal bipartitions occur for all species. (B)\nMaximum-likelihood species tree using single copy nucleotide sequences. Branches with\ncircles are fully supported (>98%) as judged by ultrafast bootstraps and SH-aLRT . Branches\nindicated by the asterisk are less well supported (<85%).\n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted April 28, 2024. ; https://doi.org/10.1101/2024.04.23.590777doi: bioRxiv preprint \n\nIn both the ML and STAG topologies, D. sewerzowii is recovered as a sister to a clade that\nincludes M. complanata and Nepeta, which are sister to each other. While our phylogenomic\nresults corroborate existing hypotheses regarding the close relationships among these\ngenera, our trees are incongruent with previously reported topologies (Supl. Fig. 3). For\nexample, nuclear phylogenetic results by Rose et al. (2023) report D. sewerzowii as sister to\nNepeta, which together are sister to the sister lineages Hymenocrater and (Lophanthus +\nMarmoritis). This contrasts with plastid-based phylogenetic results reported in the same\nstudy, which recover Nepeta as sister to a clade comprising the sister taxa, Drepanocaryum\nand Hymenocrater, and their sister, (Lophanthus + Marmoritis), and with results by\nSerrpooshan et al. (2018), which recover D. sewerzowii as sister to a mixed and partially\nunresolved clade of Hymenocrater, Lophanthus, Marmoritis, and Nepeta. T opological\ndiscordances among trees reported in this and previous studies likely reflect differences in\ntaxonomic and molecular sampling, but they also highlight the complexity of resolving\nintergeneric relationships within Nepetinae. The species-tree presented here (Fig. 3)\nprovides necessary context for comparative genomics, although interpretations should be\nconsidered alongside available transcriptome- and marker-based phylogenies until additional\nNepetinae genomes and phylogenomic results become available. Nevertheless, the\ngenomes presented here provide a valuable resource to explore the evolutionary trajectories\nunderpinning the remarkable innovations in specialised metabolism within the Lamiaceae.\nConclusion\nPlant genome assemblies are being generated at a remarkable rate, with two-thirds of\navailable plant genome assemblies generated within the last 3 years (Xie et al. 2024). Here\nwe present the chromosome-level genome assemblies of D. sewerzowii and M. complanata,\nrepresenting the first assemblies from these genera. The gene and repeat annotations,\nalong with expression matrices, present a comprehensive resource for comparative\ngenomics. The species-tree using gene models from available Lamiaceae genome\nassemblies provides a reference point that celebrates the number of sequenced species.\nThese genome assemblies will allow us to decipher the evolutionary innovations that\nresulted in the remarkably diverse number of specialised metabolites found in the\nLamiaceae.\nData Availability Statement\nThe raw reads for whole genome and transcriptome sequencing are available in the National\nCenter for Biotechnology Information Sequence Read Archive BioProject PRJNA1097548\n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted April 28, 2024. ; https://doi.org/10.1101/2024.04.23.590777doi: bioRxiv preprint \n\nand PRJNA1095452. The genome assembly, annotation files and gene expression\nabundance datasets are available through Figshare as supplementary data.\nAcknowledgments\nWe would like to thank Prof. C. Robin Buell for her advice and on assembling mint genomes.\nThe Viking cluster was used during this project, which is a high performance compute facility\nprovided by the University of York. We are grateful for computational support from the\nUniversity of York, IT Services and the Research IT team. We are grateful to the University\nof York Horticulture T eam for the propagation and care of our plant material. We would like to\nthank Karen Hobb from the University of York Imaging and Cytometry Laboratory for\nassistance with FCM.\nConflict of Interest\nThe authors declare no competing interests.\nFunder Information\nThis work was financially supported by the BBSRC (BB/V006452/1) and UKRI\n(MR/S01862X/1).\n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted April 28, 2024. ; https://doi.org/10.1101/2024.04.23.590777doi: bioRxiv preprint \n\nSupplementary Data\nSupplementary Figure 1. FCM ungated histograms for D. sewerzowii (A) and M.\ncomplanata (B). N. cataria tetraploid (G1 NECA) and N. racemosa diploid (G1 NEMU)\nreferences are shown relative to that of D. sewerzowii and M. complanata, labelled as G1\nsample.\nSupplementary Figure 2. HiC contact maps for D. sewerzowii (A) and M. complanata (B).\n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted April 28, 2024. ; https://doi.org/10.1101/2024.04.23.590777doi: bioRxiv preprint \n\nSupplementary Figure 3. Summary of tree incongruence between the genome derived\nspecies tree (this study) and trees reported using Bayesian Inference (BI), maximum clade\ncredibility (MCC) or maximum parsimony (MP) for tree inference with plastid markers,\nnuclear markers or nuclear ribosomal internal transcribed spacer regions (NRITS).\n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted April 28, 2024. ; https://doi.org/10.1101/2024.04.23.590777doi: bioRxiv preprint \n\nSupplementary Table 1 - Genome assemblies used for comparative genomics and\nphylogenomics.\nSpecies Reference\nAgastache rugosa (Fisch. & C.A.Mey.) Kuntze Park et al. 2023\nCallicarpa americana L. Hamilton et al. 2020\nDrepanocaryum sewerzowii (Regel) Pojark. This work\nHyssopus officinalis L. Lichman et al. 2020\nIsodon rubescens (Hemsl.) H.Hara Sun et al. 2023\nLavandula angustifolia Mill. Hamilton et al. 2023\nMarmoritis complanata (Dunn) A.L.Budantzev This work\nMentha longifolia (L.) L. Vining et al. 2022\nNepeta cataria L. Lichman et al. 2020\nNepeta racemosa Lam. Lichman et al. 2020\nOcimum basilicum L. Bornowski et al. 2020\nOriganum majorana L. Bornowski et al. 2020\nOriganum vulgare L. Bornowski et al. 2020\nPaulownia fortunei (Seem.) Hemsl. Cao et al. 2021\nPerilla citriodora (Makino) Nakai Zhang et al. 2021\nPerilla frutescens (L.) Britton Zhang et al. 2021\nPogostemon cablin (Blanco) Benth. Shen et al. 2022\nSalvia bowleyana Dunn Zheng et al. 2021\nSalvia hispanica L. Wang et al. 2022\nSalvia miltiorrhiza Bunge Pan et al. 2023\nSalvia rosmarinus Spenn. Han et al. 2023\nSalvia splendens Sellow ex Nees Jia et al. 2021\nSchizonepeta tenuifolia (Benth.) Briq. Liu et al. 2023\nScutellaria baicalensis Georgi Xu et al. 2020\nScutellaria barbata D.Don Xu et al. 2020\nTeucrium marum L. Smit et al. 2024\nThymus quinquecostatus Čelak. Sun et al. 2022\n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted April 28, 2024. ; https://doi.org/10.1101/2024.04.23.590777doi: bioRxiv preprint \n\nLiterature Cited\nBaril, T ., J. Galbraith, and A. Hayward, 2023 Earl Grey. Zenodo.\nBaril, T ., J. Galbraith, and A. Hayward, 2024 Earl Grey: A fully automated user-friendly\ntransposable element annotation and analysis pipeline. Mol. Biol. Evol. 41:\n2022.06.30.498289.\nBordbar, F ., 2023 New chromosome counts in Lamiaceae from flora of Iran - II. JABS 17:\n298–305.\nBornowski, N., J. P . Hamilton, P . Liao, J. C. Wood, N. Dudareva et al., 2020 Genome\nsequencing of four culinary herbs reveals terpenoid genes underlying chemodiversity in\nthe Nepetoideae. DNA Res. 27.:\nBrůna, T ., K. J. Hoff, A. Lomsadze, M. Stanke, and M. Borodovsky, 2021 BRAKER2:\nAutomatic eukaryotic genome annotation with GeneMark-EP+ and AUGUSTUS\nsupported by a protein database. NAR Genom Bioinform 3: lqaa108.\nBruna, T ., A. Lomsadze, and M. Borodovsky, 2024 GeneMark-ETP: Automatic gene finding\nin eukaryotic genomes in consistency with extrinsic data. bioRxiv.\nBuchfink, B., C. Xie, and D. H. Huson, 2015 Fast and sensitive protein alignment using\nDIAMOND. Nat. Methods 12: 59–60.\nCao, Y ., G. Sun, X. Zhai, P . Xu, L. Ma et al., 2021 Genomic insights into the fast growth of\npaulownias and the formation of Paulownia witches’ broom. Mol. Plant 14: 1668–1682.\nCapella-Gutiérrez, S., J. M. Silla-Martínez, and T . Gabaldón, 2009 trimAl: A tool for\nautomated alignment trimming in large-scale phylogenetic analyses. Bioinformatics 25:\n1972–1973.\nChakraborty, M., J. G. Baldwin-Brown, A. D. Long, and J. J. Emerson, 2016 Contiguous and\naccurate de novo assembly of metazoan genomes with modest long read coverage.\nNucleic Acids Res. 44: e147.\nChamala, S., N. García, G. T . Godden, V. Krishnakumar, I. E. Jordon-Thaden et al., 2015\nMarkerMiner 1.0: A new application for phylogenetic marker development using\n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted April 28, 2024. ; https://doi.org/10.1101/2024.04.23.590777doi: bioRxiv preprint \n\nangiosperm transcriptomes. Appl. Plant Sci. 3.:\nChen, Y ., F . Nie, S.-Q. Xie, Y .-F . Zheng, Q. Daiet al., 2021 Efficient assembly of nanopore\nreads via highly accurate and intact error correction. Nat. Commun. 12: 60.\nChernomor, O., A. von Haeseler, and B. Q. Minh, 2016 T errace aware data structure for\nphylogenomic inference from supermatrices. Syst. Biol. 65: 997–1008.\nDainat, J., D. Hereñú, Murray, K, D, E. Davis, K. Crouch et al., 2023 NBISweden/AGAT:\nAGAT-v1.2.0.\nDanecek, P ., J. K. Bonfield, J. Liddle, J. Marshall, V. Ohan et al., 2021 Twelve years of\nSAMtools and BCFtools. Gigascience 10.:\nDe Coster, W., S. D’Hert, D. T . Schultz, M. Cruts, and C. Van Broeckhoven, 2018 NanoPack:\nVisualizing and processing long-read sequencing data. Bioinformatics 34: 2666–2669.\nDe Coster, W., and R. Rademakers, 2023 NanoPack2: Population-scale evaluation of\nlong-read sequencing data. Bioinformatics 39.:\nDobin, A., C. A. Davis, F . Schlesinger, J. Drenkow, C. Zaleski et al., 2013 STAR: Ultrafast\nuniversal RNA-seq aligner. Bioinformatics 29: 15–21.\nDolezel, J., J. Greilhuber, and J. Suda, 2007 Estimation of nuclear DNA content in plants\nusing flow cytometry. Nat. Protoc. 2: 2233–2244.\nDoyle, J. J., and J. L. Doyle, 1990 Isolation of plant DNA from fresh tissue. Focus 12: 13–15.\nEmms, D. M., and S. Kelly, 2019 OrthoFinder: Phylogenetic orthology inference for\ncomparative genomics. Genome Biol. 20: 238.\nEmms, D. M., and S. Kelly, 2015 OrthoFinder: Solving fundamental biases in whole genome\ncomparisons dramatically improves orthogroup inference accuracy. Genome Biol. 16:\n157.\nEmms, D. M., and S. Kelly, 2018 STAG: Species Tree Inference from All Genes. bioRxiv\n267914.\nEmms, D. M., and S. Kelly, 2017 STRIDE: Species Tree Root Inference from Gene\nDuplication Events. Mol. Biol. Evol. 34: 3267–3278.\nEwels, P ., M. Magnusson, S. Lundin, and M. Käller, 2016 MultiQC: Summarize analysis\n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted April 28, 2024. ; https://doi.org/10.1101/2024.04.23.590777doi: bioRxiv preprint \n\nresults for multiple tools and samples in a single report. Bioinformatics 32: 3047–3048.\nFrezza, C., A. Venditti, M. Serafini, and A. Bianco, 2019 Chapter 4 - Phytochemistry,\nchemotaxonomy, ethnopharmacology, and nutraceutics of Lamiaceae, pp. 125–178 in\nStudies in Natural Products Chemistry, edited by Atta-ur-Rahman. Elsevier.\nGarcía-Alcalde, F ., K. Okonechnikov, J. Carbonell, L. M. Cruz, S. Götz et al., 2012 Qualimap:\nEvaluating next-generation sequencing alignment data. Bioinformatics 28: 2678–2679.\nGotoh, O., 2008 A space-efficient and accurate method for mapping and aligning cDNA\nsequences onto genomic sequence. Nucleic Acids Res. 36: 2630–2638.\nGuindon, S., J.-F . Dufayard, V. Lefort, M. Anisimova, W. Hordijk et al., 2010 New algorithms\nand methods to estimate maximum-likelihood phylogenies: assessing the performance\nof PhyML 3.0. Syst. Biol. 59: 307–321.\nHamilton, J. P ., G. T . Godden, E. Lanier, W. W. Bhat, T . J. Kinser et al., 2020 Generation of a\nchromosome-scale genome assembly of the insect-repellent terpenoid-producing\nLamiaceae species, Callicarpa americana. Gigascience 9.:\nHamilton, J. P ., B. Vaillancourt, J. C. Wood, H. Wang, J. Jiang et al., 2023\nChromosome-scale genome assembly of the “Munstead” cultivar of Lavandula\nangustifolia. BMC Genom Data 24: 75.\nHan, D., W. Li, Z. Hou, C. Lin, Y . Xie et al., 2023 The chromosome-scale assembly of the\nSalvia rosmarinus genome provides insight into carnosic acid biosynthesis. Plant J. 113:\n819–832.\nHarley, R, M, S. Atkins, Budantsev, A, L, Cantino, P , D, Conn, B, J et al., 2004 Labiateae, pp.\n167–275 in The Families and Genera of Vascular Plants, edited by W. Kadereit J.\nSpringer, Berlin, Heidelberg.\nHernández Lozada, N. J., B. Hong, J. C. Wood, L. Caputi, J. Basquin et al., 2022\nBiocatalytic routes to stereo-divergent iridoids. Nat. Commun. 13: 4718.\nHoang, D. T ., O. Chernomor, A. von Haeseler, B. Q. Minh, and L. S. Vinh, 2018 UFBoot2:\nImproving the ultrafast bootstrap approximation. Mol. Biol. Evol. 35: 518–522.\nHoff, K. J., S. Lange, A. Lomsadze, M. Borodovsky, and M. Stanke, 2016 BRAKER1:\n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted April 28, 2024. ; https://doi.org/10.1101/2024.04.23.590777doi: bioRxiv preprint \n\nUnsupervised RNA-Seq-based genome annotation with GeneMark-ET and\nAUGUSTUS. Bioinformatics 32: 767–769.\nHoff, K. J., A. Lomsadze, M. Borodovsky, and M. Stanke, 2019 Whole-Genome annotation\nwith BRAKER. Methods Mol. Biol. 1962: 65–95.\nIwata, H., and O. Gotoh, 2012 Benchmarking spliced alignment programs including Spaln2,\nan extended version of Spaln that incorporates additional species-specific features.\nNucleic Acids Res. 40: e161.\nJia, K.-H., H. Liu, R.-G. Zhang, J. Xu, S.-S. Zhou et al., 2021 Chromosome-scale assembly\nand evolution of the tetraploid Salvia splendens (Lamiaceae) genome. Hortic Res 8:\n177.\nKalyaanamoorthy, S., B. Q. Minh, T . K. F . Wong, A. von Haeseler, and L. S. Jermiin, 2017\nModelFinder: fast model selection for accurate phylogenetic estimates. Nat. Methods\n14: 587–589.\nKolmogorov, M., J. Yuan, Y . Lin, and P . A. Pevzner, 2019 Assembly of long, error-prone\nreads using repeat graphs. Nat. Biotechnol. 37: 540–546.\nKovaka, S., A. V. Zimin, G. M. Pertea, R. Razaghi, S. L. Salzberg et al., 2019 Transcriptome\nassembly from long-read RNA-seq alignments with StringTie2. Genome Biol. 20: 278.\nLi, H., 2013 Aligning sequence reads, clone sequences and assembly contigs with\nBWA-MEM. arXiv [q-bio.GN].\nLi, H., 2018 Minimap2: Pairwise alignment for nucleotide sequences. Bioinformatics 34:\n3094–3100.\nLichman, B. R., G. T . Godden, J. P . Hamilton, L. Palmer, M. O. Kamileen et al., 2020 The\nevolutionary origins of the cat attractant nepetalactone in catnip. Sci Adv 6: eaba0721.\nLichman, B. R., M. O. Kamileen, G. R. Titchiner, G. Saalbach, C. E. M. Stevenson et al.,\n2019 Uncoupled activation and cyclization in catmint reductive terpenoid biosynthesis.\nNat. Chem. Biol. 15: 71–79.\nLin, Y ., J. Yuan, M. Kolmogorov, M. W. Shen, M. Chaisson et al., 2016 Assembly of long\nerror-prone reads using de Bruijn graphs. Proc. Natl. Acad. Sci. U. S. A. 113:\n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted April 28, 2024. ; https://doi.org/10.1101/2024.04.23.590777doi: bioRxiv preprint \n\nE8396–E8405.\nLiu, C., S. J. Smit, J. Dang, P . Zhou, G. T . Godden et al., 2023 A chromosome-level genome\nassembly reveals that a bipartite gene cluster formed via an inverted duplication\ncontrols monoterpenoid biosynthesis in Schizonepeta tenuifolia. Mol. Plant 16: 533–548.\nManni, M., M. R. Berkeley, M. Seppey, F . A. Simão, and E. M. Zdobnov, 2021 BUSCO\nUpdate: Novel and streamlined workflows along with broader and deeper phylogenetic\ncoverage for scoring of eukaryotic, prokaryotic, and viral genomes. Mol. Biol. Evol. 38:\n4647–4654.\nmedaka: Sequence correction provided by ONT Research, 2018.\nMinh, B. Q., H. A. Schmidt, O. Chernomor, D. Schrempf, M. D. Woodhams et al., 2020\nIQ-TREE 2: New models and ffficient methods for phylogenetic inference in the genomic\nera. Mol. Biol. Evol. 37: 1530–1534.\nMint Evolutionary Genomics Consortium, 2018 Phylogenomic mining of the mints reveals\nmultiple mechanisms contributing to the evolution of chemical diversity in Lamiaceae.\nMol. Plant 11: 1084–1096.\nMyers, N., R. A. Mittermeier, C. G. Mittermeier, G. A. da Fonseca, and J. Kent, 2000\nBiodiversity hotspots for conservation priorities. Nature 403: 853–858.\nOkonechnikov, K., A. Conesa, and F . García-Alcalde, 2016 Qualimap 2: Advanced\nmulti-sample quality control for high-throughput sequencing data. Bioinformatics 32:\n292–294.\nPan, X., Y . Chang, C. Li, X. Qiu, X. Cui et al., 2023 Chromosome-level genome assembly of\nSalvia miltiorrhiza with orange roots uncovers the role of Sm2OGD3 in catalyzing\n15,16-dehydrogenation of tanshinones. Hortic Res 10: uhad069.\nPark, H.-S., I. H. Jo, S. Raveendar, N.-H. Kim, J. Gil et al., 2023 A chromosome-level\ngenome assembly of Korean mint (Agastache rugosa). Sci Data 10: 792.\nPatro, R., G. Duggal, M. I. Love, R. A. Irizarry, and C. Kingsford, 2017 Salmon provides fast\nand bias-aware quantification of transcript expression. Nat. Methods 14: 417–419.\nPeng, D.-L., X.-J. Hu, J. Yang, and H. Sun, 2018 Seed dormancy, germination and soil seed\n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted April 28, 2024. ; https://doi.org/10.1101/2024.04.23.590777doi: bioRxiv preprint \n\nbank of Lamiophlomis rotata and Marmoritis complanatum (Labiatae), two endemic\nspecies from Himalaya–Hengduan Mountains. Plant Biosystems - An International\nJournal Dealing with all Aspects of Plant Biology 152: 642–648.\nPertea, G., and M. Pertea, 2020 GFF Utilities: GffRead and GffCompare. F1000Res. 9.:\nPicard toolkit, 2019 Broad Institute.\nPublished Plant Genomes.\nRoach, M. J., S. A. Schmidt, and A. R. Borneman, 2018 Purge Haplotigs: Allelic contig\nreassignment for third-gen diploid genome assemblies. BMC Bioinformatics 19: 460.\nRose, J. P ., J. Wiese, N. Pauley, T . Dirmenci, F . Celepet al., 2023 East Asian-North\nAmerican disjunctions and phylogenetic relationships within subtribe Nepetinae\n(Lamiaceae). Mol. Phylogenet. Evol. 187: 107873.\nRose, J. P ., C.-L. Xiang, K. J. Sytsma, and B. T . Drew, 2022 A timeframe for mint evolution:\ntowards a better understanding of trait evolution and historical biogeography in\nLamiaceae. Bot. J. Linn. Soc. 200: 15–38.\nSerpooshan, F ., Z. Jamzad, T . Nejadsattari, and I. Mehregan, 2018 Molecular phylogenetics\nof Hymenocrater and allies (Lamiaceae): new insights from nrITS, plastid trnL intron and\ntrnL-F intergenic spacer DNA sequences. Nord. J. Bot. 36: njb–01600.\nShen, Y ., W. Li, Y . Zeng, Z. Li, Y . Chenet al., 2022 Chromosome-level and\nhaplotype-resolved genome provides insight into the tetraploid hybrid origin of patchouli.\nNat. Commun. 13: 3511.\nSmit, S. J., S. Ayten, B. A. Radzikowska, J. P . Hamilton, S. Langer et al., 2024 The genomic\nand enzymatic basis for iridoid biosynthesis in cat thyme (T eucrium marum). Plant J.\nSolares, E. A., M. Chakraborty, D. E. Miller, S. Kalsow, K. Hall et al., 2018 Rapid low-cost\nassembly of the Drosophila melanogaster reference genome using low-coverage,\nlong-read sequencing. G3 8: 3143–3154.\nStanke, M., M. Diekhans, R. Baertsch, and D. Haussler, 2008 Using native and syntenically\nmapped cDNA alignments to improve de novo gene finding. Bioinformatics 24:\n637–644.\n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted April 28, 2024. ; https://doi.org/10.1101/2024.04.23.590777doi: bioRxiv preprint \n\nStanke, M., O. Schöffmann, B. Morgenstern, and S. Waack, 2006 Gene prediction in\neukaryotes with a generalized hidden Markov model that uses hints from external\nsources. BMC Bioinformatics 7: 62.\nSteinegger, M., and J. Söding, 2017 MMseqs2 enables sensitive protein sequence searching\nfor the analysis of massive data sets. Nat. Biotechnol. 35: 1026–1028.\nSun, W.-G., 2016 Karyotype of nine endemic species from alpine subnival belt in the\nHengduan mountains, SW China. J. Jpn. Bot. 91: 242–249.\nSun, Y ., J. Shao, H. Liu, H. Wang, G. Wang et al., 2023 A chromosome-level genome\nassembly reveals that tandem-duplicated CYP706V oxidase genes control oridonin\nbiosynthesis in the shoot apex of Isodon rubescens. Mol. Plant 16: 517–532.\nSun, H., J. Zhang, T . Deng, and D. E. Boufford, 2017 Origins and evolution of plant diversity\nin the Hengduan Mountains, China. Plant Divers 39: 161–166.\nSun, M., Y . Zhang, L. Zhu, N. Liu, H. Bai et al., 2022 Chromosome-level assembly and\nanalysis of the Thymus genome provide insights into glandular secretory trichome\nformation and monoterpenoid biosynthesis in thyme. Plant Commun 3: 100413.\nT ang, H., J. E. Bowers, X. Wang, R. Ming, M. Alam et al., 2008 Synteny and collinearity in\nplant genomes. Science 320: 486–488.\nT ang, H., V. Krishnakumar, and J. Li, 2015 jcvi: JCVI utility libraries.\nVaser, R., I. Sović, N. Nagarajan, and M. Šikić, 2017 Fast and accurate de novo genome\nassembly from long uncorrected reads. Genome Res. 27: 737–746.\nVining, K. J., I. Pandelova, I. Lange, A. N. Parrish, A. Lefors et al., 2022 Chromosome-level\ngenome assembly of Mentha longifolia L. reveals gene organization underlying disease\nresistance and essential oil traits. G3 12.:\nWalker, B. J., T . Abeel, T . Shea, M. Priest, A. Abouelliel et al., 2014 Pilon: An integrated tool\nfor comprehensive microbial variant detection and genome assembly improvement.\nPLoS One 9: e112963.\nWang, L., M. Lee, F . Sun, Z. Song, Z. Yang et al., 2022 A chromosome-level genome\nassembly of chia provides insights into high omega-3 content and coat color variation of\n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted April 28, 2024. ; https://doi.org/10.1101/2024.04.23.590777doi: bioRxiv preprint \n\nits seeds. Plant Commun 3: 100326.\nWink, M., 2003 Evolution of secondary metabolites from an ecological and molecular\nphylogenetic perspective. Phytochemistry 64: 3–19.\nXie, L., X. Gong, K. Yang, Y . Huang, S. Zhang et al., 2024 T echnology-enabled great leap in\ndeciphering plant genomes. Nat Plants.\nXu, Z., R. Gao, X. Pu, R. Xu, J. Wang et al., 2020 Comparative genome analysis of\nScutellaria baicalensis and Scutellaria barbata reveals the evolution of active flavonoid\nbiosynthesis. Genomics Proteomics Bioinformatics 18: 230–240.\nZaman, W., J. Ye, M. Ahmad, S. Saqib, Z. K. Shinwari et al., 2022 Phylogenetic exploration\nof traditional Chinese medicinal plants: A case study on Lamiaceae. Pak. J. Bot. 54:\n1033–1040.\nZhang, Y ., Q. Shen, L. Leng, D. Zhang, S. Chen et al., 2021 Incipient diploidization of the\nmedicinal plant Perilla within 10,000 years. Nat. Commun. 12: 5508.\nZhao, F ., Y .-P . Chen, Y . Salmaki, B. T . Drew, T . C. Wilsonet al., 2021 An updated tribal\nclassification of Lamiaceae based on plastome phylogenomics. BMC Biol. 19: 2.\nZheng, X., D. Chen, B. Chen, L. Liang, Z. Huang et al., 2021 Insights into salvianolic acid B\nbiosynthesis from chromosome-scale assembly of the Salvia bowleyana genome. J.\nIntegr. Plant Biol. 63: 1309–1323.\n.CC-BY 4.0 International licensemade available under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is \nThe copyright holder for this preprintthis version posted April 28, 2024. ; https://doi.org/10.1101/2024.04.23.590777doi: bioRxiv preprint","source_license":"CC-BY-4.0","license_restricted":false}