Chloroplast Genome of Rambutan and Comparative Analyses in Sapindaceae | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Chloroplast Genome of Rambutan and Comparative Analyses in Sapindaceae Fei Dong, Zhicong Lin, Jing Lin, Ray Ming, Wenping Zhang This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-128918/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 02 Feb, 2021 Read the published version in Plants → Version 1 posted You are reading this latest preprint version Abstract Background: Rambutan ( Nephelium lappaceum L.) is an important fruit tree belongs to the family Sapindaceae and widely cultivated in Southeast Asia. The chloroplast of plants, as a photosynthetic organelle plays an important role in the photosynthesis and secondary metabolic activities. The chloroplast genome sequencing has become an integral part in understanding the genomic machinery and the phylogenetic histories of rambutan organelles. Results: We sequenced its chloroplast genome and assembled 161,321 bp circular DNA. It is characterized by a typical quadripartite structure composed of a large (86,068 bp) and small (18,153 bp) single-copy region interspersed by two identical inverted repeats (IRs) (28,550 bp). We identified 132 genes including 78 protein-coding, 29 tRNA and 4 rRNA genes, with 21 genes duplicated in the IRs. Sixty-three simple sequence repeats (SSRs) and 98 repetitive sequences were detected. Twenty-nine codons showed biased usage and 49 potential RNA editing sites were predicted across 18 protein-coding genes in the rambutan chloroplast genome. In addition, coding gene sequence divergence analysis of N. lappaceum suggested that ccsA, clpP, rpoA, rps12, psbJ and rps19 were under positive selection, which might reflect specific adaptations of N. lappaceum to its particular living environment. Comparative chloroplast genome analyses from five species in Sapindaceae revealed that a higher similarity was conserved in the IR regions than in the LSC and SSC regions. The phylogenetic analysis showed that N. lappaceum chloroplast genome has the closest relationship with that of Pometia tomentosa. Conclusions: The understanding of the chloroplast genomics of rambutan and comparative analysis of Sapindaceae species would provide insight into future research on the breeding of rambutan and Sapindaceae evolutionary studies. Epigenetics & Genomics Rambutan Nephelium lappaceum Chloroplast genome Sapindaceae RNA editing Phylogeny Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Background Rambutan ( Nephelium lappaceum L.) is an important tropical fruit in the family Sapindaceae and originated in Indonesia and Malay Peninsula [ 1 ]. It is widely cultivated in Southeast Asia and the coastal areas of South China. Malaysians refer to it as “rambutan”, because of the fruit surface covered with thick and elongated spines. The fruits of rambutan are popular in the general population due to its rich nutrients, delicate and characteristic flavor and delicious taste. Rambutan peel extract is rich in phenolic content and exhibited antibacterial activity against many pathogenic bacteria, suggesting its antioxidant and/or antimicrobial properties[ 2 ]. Rambutan has the potential to be used as natural antioxidants and anti-aging agent in pharmaceutical and food industries to replace synthetic ones[ 3 , 4 ]. The Sapindaceae family contains over 150 genera and 2000 species with several economically important crops widely distributed in tropical and subtropical regions[ 5 ]. However, genomic research on Sapindaceae family, especially in the N. lappaceum has been relatively scarce. This lack of genetic information making it difficult to meet the need for improving the quality and agronomic characteristics of rambutan through breeding and gene editing. Chloroplast (cp) are photosynthetic organelles that provide energy to green plants, it plays an important role in the photosynthesis and secondary metabolic activities[ 6 , 7 ]. The chloroplast genomes are maternally inherited in most plants and are highly conserved with its composition and sequence. The typical chloroplast genomes of angiosperms are circular DNA molecule which has a characteristic quadripartite structure with a large single-copy (LSC) region, a small single-copy (SSC) region, and two inverse repeats (IRs) regions[ 6 ]. The length of the genome is between 120 and 170 kb and usually encode 110 to 130 genes, and about 40 genes are specialized participating in photosynthesis, transcription and translation[ 8 , 9 ]. The first chloroplast genome from tobacco ( Nicotiana tabacum ) was sequenced in 1986[ 10 ], With the rapid development of next-generation sequencing technologies, the cost of whole genome sequencing is dropping rapidly[ 11 ]. Complete chloroplast genome sequences now could be easily acquired with relatively low cost. It has been an explosion in the number of available chloroplast genome sequences. Over 4,300 complete chloroplast genome sequences have been submitted in the National Center for Biotechnology Information (NCBI) organelle genome database. Within the Sapindaceae family, the complete chloroplast genomes of eight plant species have been sequenced and were available from the NCBI database. Nevertheless, no chloroplast genome in the genus Nephelium Linn has been reported. In this study, we report the first complete chloroplast genome of N. lappaceum , exploring its general features, SSRs and long repeats, codon usage and analysis of IR contraction and expansion. In addition, nine chloroplast genome sequences were used for analysis of molecular evolution in the Sapindaceae family. We constructed a phylogenetic tree to understand the phylogenetic relationship of Sapindaceae plants. The chloroplast genome sequence and the comprehensive chloroplast genomic analysis of N. lappaceum would provide a theoretical basis for molecular identification and further understanding of the evolutionary history of Sapindaceae family. Results Chloroplast Genome Features of N. lappaceum The structure of N. lappaceum chloroplast genome was analogous to most chloroplast genomes of plants with a typical quadripartite structure. We assemble a closed circular chloroplast genome with 161,321 bp in N. lappaceum . The chloroplast genome contains a pair of inverted repeat regions (IRs) of 28,550 bp, a large single-copy region (LSC) of 86,068 bp and a small single-copy region (SSC) of 18,153 bp (Fig. 1 ). The size of N. lappaceum chloroplast genome was slightly larger than that in S. mukorossi (160,481 bp), P. tomentosa (160,818 bp), D. Longan (160,833 bp) and shorter than that in L. chinensis chloroplast genome (162,524 bp) of Sapindoideae (Table 1 ). The number of chloroplast genes in N. lappaceum was 132, the same with those in D. Longan and L. chinensis (Table 1 ). In addition, there was no significant difference in GC content among the five analytical genomes in Sapindoideae . Table 1 Comparison of the general features of the five Sapindoideae chloroplast genomes. Genome feature Dimocarpus longan Litchi chinensis Pometia tomentosa Sapindus mukorossi Nephelium lappaceum GenBank MG214255 KY635881 MN106254 KM454982 MT936934 Size (bp) 160833 162524 160818 160481 161321 LSC (bp) 85707 85750 85666 85650 86068 SSC (bp) 18270 16568 18360 18873 18153 IR (bp) 28428 30103 28396 27979 28550 Total genes 132 132 133 135 132 Protein genes 87 87 88 88 87 tRNA genes 37 37 37 39 37 rRNA genes 8 8 8 8 8 GC (%) 37.79% 37.80% 37.87% 37.66% 37.77% The overall nucleotide composition of rambutan is: 30.79% A, 31.44% T, 19.27% C, and 18.50% G, with a total GC content of 37.77%. In total, 132 genes were annotated on this chloroplast genome, including 78 protein-coding genes, 29 transfer RNA genes (tRNA) and 4 ribosomal RNA genes (rRNA). Among them, a total of 21 genes were found duplicated in the IR regions, including nine protein-coding genes ( rps3, rps7, rps12, rps19, rpl2, rpl22, rpl23, ndhB and ycf2 ), eight tRNA genes ( trnA-UGC, trnI-CAU, trnI-GAU, trnL-CAA, trnM-CAU, trnN-GUU, trnR-ACG and trnV-GAC ) and four rRNA genes ( rrn4.5 s, rrn5s, rrn16s and rrn23s ) (Additional file 1: Table S2). The genes structure analysis showed that 21 genes contains introns, and 19 of them (11 protein-coding genes and 8 tRNA genes) have one intron, while two genes ( ycf3 and clpP ) have two introns (Additional file 1: Table S3). Characterization of SSRs and repeat sequences A total of 63 SSRs were detected from rambutan chloroplast genome, of which 45 were mononucleotide, 3 dinucleotide, 8 trinucleotide, 5 tetranucleotide and two pentanucleotide (Additional file 1: Table S4). Moreover, we compared the distribution pattern and number of SSRs with eight other chloroplast genomes in Sapindaceae family (Additional file 1: Table S5). The number of mononucleotide repeats is more than the sum of other types (Fig. 2 A), and the number and types of chloroplast SSRs vary in different species. S. mukorossi (91 SSRs) possess the highest number of SSRs while E. cavaleriei (62 SSRs) possesses the lowest. Furthermore, the chloroplast genome of D. longan , L. chinensis , P. tomentosa , D. viscosa , K. paniculate and X. sorbifolium contained 79, 75, 74, 77, 87 and 83 SSRs, respectively (Fig. 2 B). In this study, a total of 98 larger repeats (> 10 bp) were identified in N. lappaceum chloroplast genome composed of 42 forward, 11 reverse, 41 palindromic and 4 complement repeats (Additional file 1: Table S6) using REPuter[ 12 ]. Among them, the largest repeat was a palindromic repeat with a size of 48 bp. Codon usage analysis and RNA editing sites prediction We used 53 protein coding sequences from rambutan chloroplast genome for calculate codon usage frequency and relative synonymous codon usage (RSCU) frequency (Additional file 1: Table S7). All protein coding sequences contain 21,434 codons. In detail, leucine and cysteine are the highest and lowest number of amino acids, they have 2,232 codons (approximately 10.41% of the total) and 236 codons (approximately 1.10% of the total), respectively. While Met (ATG) and Trp (TGG) are encoded by only one codon showed no biased usage (RSCU = 1). 30 codons with RSCU values more than 1, indicating they showed biased usage (Fig. 3 ). Among them, excluding the leucine (UUG) codon was G-ending, the remaining 29 biased usage codons of N. lappaceum were all A/T-ending in the third codon. In addition, there were 49 potential RNA editing sites were found across 18 protein-coding genes in N. lappaceum chloroplast genome and the ndhB gene contained the most RNA editing sites (9) (Additional file 1: Table S8). We also observed that RNA editing sites were all C to U conversion, and took place at the first (30.6%) or second (69.4%) positions of the codons, indicating that editing in the third codon position disappeared quicker than that in the second or first codon position. Furthermore, serine codons were more frequently edited than codons of other amino acids and the conversion from serine to leucine occurred most frequently. Comparative genomes analysis The comparative analysis based on mVISTA was performed between the chloroplast genomes of rambutan with other four Sapindoideae species with the annotated D. longan chloroplast genome as a reference. The five Sapindoideae subfamily chloroplast genomes length between the confines of 160,481 to 162,524 bp. The chloroplast genome of L. chinensis has the largest size, whereas S. mukorossi has the smallest size. Interestingly, the SSC region (16,568 bp) of L. chinensis is the shortest, whereas the SSC region (18,873 bp) of S. mukorossi chloroplast genome is the longest (Fig. 4 ). The IR (A/B) regions exhibited less divergence than the SSC and LSC regions. In addition, the coding regions were more highly conserved than the non-coding regions. Among the five chloroplast genomes, four rRNA genes ( rrn16S, rrn23S, rrn5S, rrn4.5S ) were the most conserved, while 7 genes ( matK, rpoC2, psbB, rpoA, ndhF, ndhD and ycf1 ) showed the most diversity in the coding regions. The highly divergent regions were found in the intergenic spacers and introns, including trnH-GUG-psbA, trnR-UCU-atpA, petN-psbM, psbZ-trnG-GCC, ndhC-trnV-UAC, psbE-petL, rpl16-rps3 and rpl32-trnL-UAG . Expansion and contraction of IR regions We compared the IR regions and the junction sites of the LSC and SSC regions of nine Sapindaceae family chloroplast genomes (including N. lappaceum ) (Fig. 5 ). The IR regions vary in different chloroplast genomes, ranging from 26,923 bp in E. cavaleriei to 30,103 bp in L. chinensis . In our study, the ycf1 gene was located at the SSC/IRA junction in all of the nine chloroplast genomes and the fragment located at the IRa region ranged from 962 bp to 3,183 bp. Moreover, most junctions between LSC and IRa in this study was located downstream of the trnH-GUG , except the S. mukorossi. In addition, the LSC/IRb junction of three species D. viscosa , E. cavaleriei and K. paniculate was located within the coding region of rpl22 and created a location of 110, 40 or 63 bp at the LSC/IRb border. The remaining chloroplast genomes share a similar pattern, the LSC/IRb junction was located in intergenic regions of rpl16 and rps3 , and the IRb/SSC junction between IRb and SSC region (JSB) of five species ( S. mukorossi, X. sorbifoliun, D. viscosa , E. cavaleriei and K. paniculate ) was located between the gene of ycf1 and ndhF . However, other four chloroplast genomes only have ndhF located or near the JSB. Synonymous (Ks) and non-synonymous (Ka) substitution rate analysis To explore molecular evolution of orthologous genes shared by nine Sapindaceae species, particularly genes undergoing purifying or positive selection, we calculated the Ka/Ks ratio of 622 orthologous pairs with 78 protein coding genes (Additional file 1: Table S9). Overall, the average Ka/Ks ratio of the nine chloroplast genomes was 0.20. Total 612 orthologous pairs had a Ka/Ks ratio less than 1 in the nine comparison groups, out of which 546 orthologs had a Ka/Ks ratio less than 0.5 (Fig. 6 ), suggesting that most genes are undergoing strong purifying selection pressures. Moreover, 66 orthologs of 31 genes with a Ka/Ks ratio between 0.5 and 1, 10 orthologous pairs of 6 genes ( ccsA, rpoA, rps12, psbJ, clpPc and rps19 ) with a Ka/Ks ratio greater than 1 were detected in this study, suggesting that these genes might have experienced positive selection in the procedure of evolution. Among them, the Ka/Ks ratio of the ycf1 gene was greater than 0.5 in eight comparison groups, the rpoA and ycf2 gene with Ka/Ks ratio greater than 0.5 was also observed in the comparison of seven and six groups, respectively. Besides, clpP, matK and rps15 gene with Ka/Ks ratio > 0.5 were founded in four out of the eight comparison groups. Phylogenetic analysis We performed multiple sequence alignments using the whole chloroplast genome sequences of nine Sapindaceae species and two Anacardiaceae species as outgroups (Fig. 7 ). All nodes in the ML trees have 100% bootstrap support values, and these 11 chloroplast genome sequences are clustered into three groups. In detail, the five species ( D. longan, L. chinensis, P. tomentosa, N. lappaceum and S. mukorossi ) from Sapindoideae clustered into one group, four species ( K. paniculata, D. viscosa, E. cavaleriei and X. sorbifolium ) from Dodonaeoide ae are in one group, and the two species ( A. occidentale and M. indica ) in Anacardiaceae are cluster into one group. In the Sapindoideae group, the N. lappaceum chloroplast genome sequence showed the closest relationship with P. tomentosa , followed by D. longan and L. chinesis , as far as S. mukorossi . The three groups of this phylogenetic tree of the 11 chloroplast genome sequences are consistent with traditional taxonomy, suggesting that the chloroplast genome could effectively resolve the phylogenetic positions and relationships of species. Disscussion We assembled N. lappaceum complete chloroplast genome sequence and deposited it to GenBank under accession number: MT936934, N. lappaceum chloroplast genome was consistent with the characteristics of most angiosperm species in structure and gene content. Although there are some differences in the sizes of the overall genome, LSC, SSC and IR regions, the numbers of genes and GC content are similar among the five Sapindoideae chloroplast genomes, which to some extent reflects the high conservation of angiosperm chloroplast genomes[ 6 ]. Intron plays an important role in RNA stability, regulation of gene expression and alternative splicing which have been reported in many other species[ 13 , 14 ]. There were two genes ( ycf3 and clpP ) included two introns in the N. lappaceum chloroplast genome. It has been reported that ycf3 gene is essential for the accumulation of the photosystem I (PSI) complex and acts a chaperone that interacts with the PSI subunits at a post-translational level[ 15 , 16 ]. Besides, clpP gene functions as the proteolytic subunit of the ATP-dependent Clp protease in plant chloroplasts and is essential for the development and/or function of plastids with active gene expression in previous studies[ 17 , 18 ]. Thus, study of ycf3 and clpP gene will contribute to further investigation of chloroplast in N. lappaceum . Simple sequence repeats (SSRs), also known as microsatellites, are tandem repeats (1 ~ 6 bp units repeated multiple times) distributed across the entire genome which have been widely applied as molecular markers for determining genetic variations across species and evolutionary studies because of its unique uniparental in inheritance[ 19 – 21 ]. we identified 63 SSRs in N. lappaceum chloroplast genome and most of SSRs were distributed in IGS regions. Mononucleotide SSRs were identified most frequently (68% on average) among the nine analyzed chloroplast genomes of Sapindaceae species, and vast majority of mononucleotide repeats consist of short polyA or polyT repeats sharing a similar pattern in most angiosperm chloroplast genomes[ 22 , 23 ]. Moreover, repetitive sequences are helpful in phylogenetic study and play a vital role in genome rearrangement[ 24 ]. These results can provide chloroplast molecular markers that can be used to quickly identify species, confirm hybrid progeny when breeding. Codon usage biases are found in all eukaryotic and prokaryotic genomes and have been proposed to regulate different aspects of translation process[ 25 ]. High RSCU values of the codons are probably attributed to amino acid functions or peptide structures that avoid transcriptional errors in chloroplast genomes[ 26 , 27 ]. There are 30 codons showed biased usage and most of them were A/T-ending in the third codon. This phenomenon also exists in previous studies[ 22 ]. Codon usage bias of chloroplast genome reflects a selective pressure to increase translation efficiency[ 28 ], and research on codon preferences can help us to better understand gene expression and molecular evolution mechanisms of N. lappaceum. We observed ndhB gene contained the most RNA editing sites within the 49 potential RNA editing sites, and 16 editing sites were U_A type, indicating there was a U_A bias for the distribution of RNA editing sites that was in accordance with previous reports in other species[ 23 , 29 ]. RNA editing is a post-transcriptional regulation pattern involved in the insertion, deletion, or modification of nucleotides that widely exists in land plants[ 30 ]. The first chloroplast RNA editing event of land plant was discovered in the mRNA transcript of rpl2 gene in maize chloroplast genome in 1991[ 31 ]. The most frequent editing events in plants are C-to-U changes, however, U-to-C editing has also been observed[ 32 , 33 ]. Additionally, RNA editing usually occurs in the first or second base of codons, resulting in the conversion of hydrophilic amino acid to hydrophobic[ 34 ]. Comparative analysis of chloroplast genomes is an essential step in genomics which can provide insight into complex evolutionary relationships. The mVISTA analysis showed that five Sapindoideae chloroplast genomes were conserved, with a high degree of synteny and gene order conservation, and the coding region was more conserved than the non-coding region, which is consistent with reports on other angiosperms[ 35 ], suggesting an evolutionary conservation of these genomes at the genome-scale level. In addition, ycf1 gene showed the greatest degree of differentiation. Previous study reported that ycf1 is helpful to provide phylogenetic information at the species level and more variable than matK in Orchidaceae [ 36 ]. Furthermore, ycf1 performed better to identify DNA barcodes of high resolution at species level than any of the matK , rbcL and trnH-psbA [ 35 ]. These variable genic regions found in our study can be regarded as molecular markers for DNA barcoding and phylogenetic studies in Sapindaceae . Although most land plants have relatively conserved cp genomes, the end of the inverted repeats (IRa and IRb) regions differs among various plant species. The expansion or contraction of the IR regions represent important evolutionary events often results in size variation of different chloroplast genomes and is helpful to studying the chloroplast genome evolution history[ 37 , 38 ]. In this study, our results suggested that the boundary of IR/LSC and IR/SSC might be conversed among chloroplast genomes of closely related family species but some differences also occurs between relatively distantly related family species, such as gene overlap length, duplicate of the ycf1 and rps3 genes, even the distance of trnH-GUG from the border near the LSC/IRB junctions, indicating that the expansion and contraction of the IR region led to length and structure changes of chloroplast genomes. The ratio between nonsynonymous (Ka) and synonymous (Ks) nucleotide substitution has been widely used as important markers in genome or gene evolution studies[ 8 ]. Ka/Ks = 1 signifies neutral evolution, Ka/Ks > 1 indicates that the gene is affected by positive selection, whereas Ka/Ks < 1 indicates that the gene is affected by purifying selection[ 39 ]. Additionally, a Ka/Ks ratio of 0.5 was considered as a useful cut-off value to identify genes under positive selection in previous studies[ 40 ]. In our study, the ccsA, rpoA, rps12, psbJ, clpP and rps19 gene with Ka/Ks > 1. It is noteworthy that the ycf1 gene also exhibited high Ka/Ks ratios with Ka/Ks > 0.5 in eight comparison groups. This result is in keeping with the previous observations that the ycf1 gene was more variable than the matK and rbcL genes in most plant, and could be using as an effective biological tool for plant phylogeny study[ 35 ]. The positive selection of genes in N. lappaceum possibly provided help for adaptations to its particular living environment. Numerous studies have shown that chloroplast genome sequences have been successfully used in taxonomic and phylogenetic studies[ 41 ], and contribute to describe the evolutionary relationships between species[ 42 ]. In this study, the topology of the trees consists of two main branches: Dodonaeoideae and evolutionary younger Sapindoideae. And generic relationships of the two subfamilies are basically congruent with the taxonomy of these families. The availability of the completed N. lappaceum chloroplast genome provided us with sequence information that can be used to confirm the phylogenetic position of N. lappaceum and understand the phylogenetic relationships among Sapindaceae . However, we had used only a small number of species in Sapindaceae , further research on other chloroplast genome as well as nuclear genome sequences of Sapindaceae should be sequenced to provide more sufficient evidence to accurately illustrate the evolution of family Sapindaceae. Conclusions We assembled the first complete chloroplast genome of rambutan using Illumina sequencing technology and compared its structure with other Sapindaceae species. The chloroplast genome of N. lappaceum exhibits similar quadripartite structure, gene order, G + C content when compared with other Sapindaceae chloroplast genomes. A total of 63 SSRs and 98 repeat sequences were identified in N. lappaceum chloroplast genome. The research on codon usage of N. lappaceum shows that some amino acids have obvious codon usage bias and the codon preferences may help us understand the evolution mechanisms of N. lappaceum . With PREP prediction, we detected 49 RNA editing loci in 18 protein-coding genes in N. lappaceum. Moreover, the expansion and contraction of the IR regions, leading to the variations in nine Sapindaceae chloroplast genome size. There are 6 genes ( ccsA, rpoA, rps12, psbJ, clpP and rps19 ) were detected with a Ka/Ks ratio > 1, suggesting that these genes experienced positive selection in the evolution. Additionally, phylogenetic analysis using 9 complete chloroplast genome sequences in Sapindaceae strongly supports the close relationship of N. lappaceum and P. tomentosa among sequenced chloroplast genomes in Sapindaceae . Methods Plant material, DNA extraction, and sequencing Young, healthy leaves of the major cultivar of rambutan, Baoyan7, were collected from Baoting (N18°23′, E109°21′) in Hainan Province, China. The leaves were frozen in liquid nitrogen, and maintained at − 80 ◦ C. The total genomic DNA was extracted by 2X cetyltrimethylammonium bromide (CTAB) method [43]. And a library with insert sizes of 300–500 bp was constructed and then sequenced on an Illumina HiSeq2500 platform (Illumina, San Diego, CA, USA) double terminal sequencing method (150 pair-ends). Chloroplast genome assembly and annotation First, FastQC software was performed to evaluate the quality of Illumina paired-end raw reads[44], and low-quality reads were filtered. The remaining clean reads were used for assembly via NOVOPlasty[45] using Dimocarpus longan chloroplast genome(GenBank: MG214255)[46] as the reference genome to generate the first version of rambutan genome. Next, all clean reads were mapped onto the first version genome and the mapped reads were assembled using SPAdes3.14.1[47] and assembled contigs were corrected using the pair-end short reads from HiSeq2500 by Pilon version 1.23 (https://github.com/broadinstitute/pilon)[48] to generate the second version of rambutan chloroplast genome. These two versions were compared and mutually corrected to get the final complete rambutan chloroplast genome. The chloroplast genome was annotated by online program GeSeq (https://chlorobox.mpimp-golm.mpg.de/geseq.html)[49] and CPGAVAS2[50]. Genome features like start/stop codons and intron/exon borders were manually corrected through the comparison of other reported Sapindaceae family chloroplast genomes. In addition, tRNA genes were identified by tRNAscan-SE 2.0 (http://lowelab.ucsc.edu/tRNAscan-SE/)[51]. A circular map of the revised annotated rambutan chloroplast genome was illustrated by using Organellar Genome DRAW (OGDRAW) (https://chlorobox.mpimpgolm.mpg.de/OGDraw.html)[52]. Chloroplast genome analysis The simple sequence repeats (SSR) in nine chloroplast genome sequences of Sapindaceae (including N. lappaceum ) (Additional file 1: Table S1) were identified by online tool MISA (https://webblast.ipk-gatersleben.de/misa/)[53] and the threshold settings were as follows: ten was applied to mononucleotide repeats, five to dinucleotide repeats and four to trinucleotide repeats, three for tetra-, penta-, and hexanucleotide repeats[54]. Repetitive sequences including forward, reverse, palindrome, and complement sequences were analyzed by REPuter program[12], and the parameter was set with a minimum length of repeat region set to 10 bp and the minimum sequence identity set to 90%. The expansion and contraction of the inverted repeat (IR) regions at junction sites from eight Sapindaceae family chloroplast genome sequences, including Dimocarpus longan (MG214255), Litchi chinensis (KY635881), Pometia tomentosa (MN106254), Sapindus mukorossi (KM454982), Dodonaea viscosa (MF155892), Eurycorymbus cavaleriei (MG813997), Koelreuteria paniculate (KY859413) and Xanthoceras sorbifolium (KY779850), were examined and plotted using IRscope online program (https://irscope.shinyapps.io/irapp/)[55]. Codon usage of the N. lappaceum chloroplast genome was analyzed via GALAXY platform (https://galaxy.pasteur.fr) [56] with CodonW online tool. The length of protein-coding gene less than 300 nucleotide and the repetitive genes sequences were removed to reduce deviation of the results[57]. Finally, 53 CDS in N. lappaceum were selected for further codon usage analysis. Besides, putative RNA editing sites were predicted using PREP-Cp web server (http://prep.unl.edu/cgi-bin/cp-input.pl)[58] with a cutoff value of 0.8. Genome Comparison We downloaded four whole chloroplast genome sequences of Sapindoideae subfamily from the National Center for Biotechnology Information (NCBI) Organelle Genome and Nucleotide Resources database, including D. longan [46], L. chinensis , P. tomentosa [59] and S. mukorossi [60]. The mVISTA online program (Shuffle-LAGAN mode)[61, 62] was used to compare chloroplast genome sequence of rambutan with the other species from Sapindoideae subfamily, in which the annotation of D. longan as the reference. Positive selection analysis of protein sequence We analyzed synonymous (Ks) and non-synonymous (Ka) substitution rates to investigate the molecular evolutionary process of Sapindaceae family, The protein-coding genes of N. lappaceum were separately compared with eight closely related species in Sapindaceae family: D. longan , L. chinensis , P. tomentosa , S. mukorossi, D. viscosa [63], E. cavaleriei [64], K. paniculate [65] and X. sorbifolium [66] using ParaAT 2.0[67], then the Ka/Ks value was calculated by KaKs_calculator 2.0[68] with NG method[69]. Phylogenetic analysis In order to deeply detect the evolutionary relationship of Sapindaceae family, we aligned 9 complete chloroplast genomes (including N. lappaceum ) with MAFFT version 7[70]. The best fitting nucleotide substitution model (TVM + I + G ) was chosen by jModelTest v2.1.7[71]. Phylogenetic analysis was then inferred by ML (maximum-likelihood) method based on the TVM + I + G substitution model in PAUP* 4.0[72] with 1000 bootstrap replicates. Anacardium occidentale (KY635877) and Mangifera indica (KY635882)[73] in Anacardiaceae family were set as the outgroup. Abbreviations Cp Chloroplast; SSRs:Simple sequence repeats; IRs:Inverted repeats; LSC:Large single-copy; SSC:Small single-copy; ML:Maximum-likelihood Declarations Acknowledgements This work was supported by startup fund from Fujian Agriculture and Forestry University. Author contributions F.D. performed most of the data analysis, and wrote the manuscript. W.Z. collected experiment materials and data. Z. L. and J. L. helped in genome assembly strategy design. R.M., W. Z. and F. D. designed the project and revised the manuscript. Availability of Data and Materials The chloroplast genome assembly and annotation were deposited in GenBank under the accession number of MT936934. All other data and material generated in this manuscript are available from the corresponding author upon reasonable request. Ethics approval and consent to participate Not applicable. Consent for publication Not applicable. Competing interests The authors declare no competing interests. Funding Not applicable. References Lim TK: Edible Medicinal and Non Medicinal Plants 2015: Springer; 2012. Palanisamy UD, Cheng HM, Masilamani T, Subramaniam T, Ling LT, Radhakrishnan AK: Rind of the rambutan, Nephelium lappaceum, a potential source of natural antioxidants . Food Chem 2008, 109 (1):54–63. Zhuang Y, Ma Q, Guo Y, Sun L: Protective effects of rambutan (Nephelium lappaceum) peel phenolics on H2O2-induced oxidative damages in HepG2 cells and d-galactose-induced aging mice . Food Chemical Toxicology 2017, 108 (Pt B):554–562. c NNMPab, B TTL, A JVC, A KR: Evaluation of antimicrobial activity of rambutan (Nephelium lappaceum L.) peel extracts . Int J Food Microbiol 2020, 321 . Harrington MG, Edwards KJ, Johnson SA, Chase MW, Gadek PA: Phylogenetic Inference in Sapindaceae sensu lato Using Plastid matK and rbcL DNA Sequences . Syst Bot 2005, 30 (2):366–382. Wicke S, Schneeweiss GM, Depamphilis CW, Kai FM, Quandt D: The evolution of the plastid chromosome in land plants: gene content, gene order, gene function . Plant MolBiol 2011, 76 (3–5):273–297. Bobik K, Burch-Smith TM: Chloroplast signaling within, between and beyond cells . Front Plant Sci 2015, 6 :781. Wolfe KH, Li W, Sharp PM: Rates of nucleotide substitution vary greatly among plant mitochondrial, chloroplast, and nuclear DNAs . Proc Natl Acad Sci U S A 1987, 84 (24):9054–9058. Palmer JD: Comparative Organization of Chloroplast Genomes . Annu Rev Genet 1985, 19 (1):325–354. Shinozaki K, Ohme M, Tanaka M, Wakasugi T, Sugiura M: The complete nucleotide sequence of the tobacco chloroplast genome: its gene organization and expression . Plant Mol Biol Rep 1986, 5 (9):2043–2049. Li C, Lin F, An D, Wang W, Huang R: Genome Sequencing and Assembly by Long Reads in Plants . Genes 2017, 9 (1):6. Kurtz S, Choudhuri JV, Ohlebusch E, Schleiermacher C, Stoye J, Giegerich R: REPuter: the manifold applications of repeat analysis on a genomic scale . Nucleic Acids Res 2001, 29 (22):4633–4642. Nguyen Dinh S, Sai TZT, Nawaz G, Lee K, Kang H: Abiotic stresses affect differently the intron splicing and expression of chloroplast genes in coffee plants (Coffea arabica) and rice (Oryza sativa) . J Plant Physiol 2016, 201 :85–94. Mirzaei S, Mansouri M, Mohammadi-Nejad G, Sablok G: Comparative assessment of chloroplast transcriptional responses highlights conserved and unique patterns across Triticeae members under salt stress . Photosynth Res 2017, 136 (3):357–369. Naver H, Boudreau E, Rochaix JD: Functional studies of YCF3: Its role in assembly of photosystem I and Interactions with some of its subunits . Plant Cell 2002, 13 (12):2731–2745. Boudreau E, Takahashi Y, Lemieux C, Turmel M, Rochaix JD: The chloroplast ycf3 and ycf4 open reading frames of Chlamydomonas reinhardtii are required for the accumulation of the photosystem I complex . Embo J 1997, 16 (20):6095–6104. Clarke AK, Schelin J, Porankiewicz J: Inactivation of the clpP1 gene for the proteolytic subunit of the ATP-dependent Clp protease in the cyanobacterium Synechococcus limits growth and light acclimation . Plant MolBiol 1998, 37 (5):791–801. Bruce CA, Cunningham KA, Stern DB: The plastid clpP gene may not be essential for plant cell viability . Plant Cell Physiol 2003(1):93–95. Varshney RK, Sigmund R, Borner A, Korzun V, Stein N, Sorrells ME, Langridge P, Graner A: Interspecific transferability and comparative mapping of barley EST-SSR markers in wheat, rye and rice . Plant Sci 2005, 168 (1):195–202. Yang A, Zhang J, Tian H, Yao X: Characterization of 39 novel EST-SSR markers for Liriodendron tulipifera and cross-species amplification in L. chinense (Magnoliaceae) . Am J Bot 2012, 99 (11):e460-464. Li B, Lin F, Huang P, Guo W, Zheng Y: Development of nuclear SSR and chloroplast genome markers in diverse Liriodendron chinense germplasm based on low-coverage whole genome sequencing . Biol Res 2020, 53 (1):21. Gao B, Yuan L, Tang T, Hou J, Pan K, Wei N: The complete chloroplast genome sequence of Alpinia oxyphylla Miq. and comparison analysis within the Zingiberaceae family . PLoS One 2019, 14 (6). Yan C, Du J, Gao L, Li Y, Hou X: The complete chloroplast genome sequence of watercress (Nasturtium officinale R. Br.): Genome organization, adaptive evolution and phylogenetic relationships in Cardamineae . Gene 2019, 699 :24–36. Cavalier-Smith T: Chloroplast Evolution: Secondary Symbiogenesis and Multiple Losses . Curr Biol 2002, 12 (2):R62-R64. Hershberg R, Petrov DA: Selection on codon bias . Annu Rev Genet 2008, 42 (1):287–299. Raman G, Park S, Lee EM, Park SJ: Evidence of mitochondrial DNA in the chloroplast genome of Convallaria keiskei and its subsequent evolution in the Asparagales . Sci Rep 2019, 9 (1):5028. Purabi M, Rofinayasmin BO, Katharina M, Ramakrishnan N, Jennifer AH: Codon usage and codon pair patterns in non-grass monocot genomes . Ann Bot 2017(6):1–17. Morton BR, Wright SI: Selective Constraints on Codon Usage of Nuclear Genes from Arabidopsis thaliana . Mol Biol Evol 2006, 24 (1):122–129. Wang W, Yu H, Wang J, Lei W, Gao J, Qiu X, Wang J: The Complete Chloroplast Genome Sequences of the Medicinal Plant Forsythia suspensa (Oleaceae) . Int J Mol Sci 2017, 18 (11):2288. Smith HC, Gott JM, Hanson MR: A guide to RNA editing . RNA-Publ RNA Soc 1997, 3 (10):1105–1123. Hoch B: Editing of a chloroplast mRNA by creation of an initiation codon . Nature 1991, 353 (6340):178–180. Maier RM, Zeltz P, Kossel H, Bonnard G, Gualberto JM, Grienenberger JM: RNA editing in plant mitochondria and chloroplasts . Plant MolBiol 1996, 32 (1):343–365. Schmitzlinneweber C, Barkan A: RNA splicing and RNA editing in chloroplasts . Topics in Current Genetics 2007, 19 :213–248. Shikanai T: RNA editing in plant organelles: machinery, physiological function and evolution . Cellular Molecular Life Sciences 2006, 63 (6):698–708. Dong W, Xu C, Li C, Sun J, Zuo Y, Shi S, Cheng T, Guo J, Zhou S: ycf1, the most promising plastid DNA barcode of land plants . Sci Rep 2015, 5 :8348. Neubig KM, Whitten WM, Carlsward BS, Blanco MA, Endara L, Williams NH, Moore M: Phylogenetic utility of ycf1 in orchids: a plastid gene more variable than matK . Plant Syst Evol 2009, 277 (1–2):75–84. Dugas DV, Hernandez D, Koenen EJM, Schwarz E, Straub S, Hughes CE, Jansen RK, Nageswara-Rao M, Staats M, Trujillo JT: Mimosoid legume plastome evolution: IR expansion, tandem repeat expansions, and accelerated rate of evolution in clpP . Sci Rep 2015, 5 :16958. Yu X, Tan W, Zhang H, Gao H, Tian X: Complete Chloroplast Genomes of Ampelopsis humulifolia and Ampelopsis japonica : Molecular Structure , Comparative Analysis , and Phylogenetic Analysis . Plants-Basel 2019, 8 (10):410. Yang Z, Bielawski JP: Statistical methods for detecting molecular adaptation . Trends Ecol Evol 2000, 15 (12):496–503. Swanson WJ, Wong A, Wolfner MF, Aquadro CF: Evolutionary Expressed Sequence Tag Analysis of Drosophila Female Reproductive Tracts Identifies Genes Subjected to Positive Selection . Genetics 2004, 168 (3):1457–1465. Gitzendanner MA, Soltis PS, Wong GK, Ruhfel BR, Soltis DE: Plastid phylogenomic analysis of green plants: A billion years of evolutionary history . American Journal of Botany 2018, 105 (3):291–301. Du YP, Bi Y, Yang FP, Zhang MF, Zhang XH: Complete chloroplast genome sequences of Lilium: Insights into evolutionary dynamics and phylogenetic analyses . Sci Rep 2017, 7 (1). Porebski S, Bailey LG, Baum BR: Modification of a CTAB DNA extraction protocol for plants containing high polysaccharide and polyphenol components . Plant Mol Biol Rep 1997, 15 (1):8–15. Andrews S: FastQC A Quality Control tool for High Throughput Sequence Data . Babraham Institute 2015. Nicolas D, Patrick M, Guillaume S: NOVOPlasty: de novo assembly of organelle genomes from whole genome data . Nucleic Acids Res 2017(4):4. Wang K, Li L, Zhao M, Li S, Sun H, Lv Y, Wang Y: Characterization of the complete chloroplast genome of longan (Dimocarpus longan Lour.) using illumina paired-end sequencing . Mitochondrial DNA Part B 2017, 2 (2):904–906. Bankevich A, Nurk S, Antipov D, Gurevich AA, Dvorkin M, Kulikov AS, Lesin VM, Nikolenko SI, Pham S, Prjibelski AD: SPAdes: a new genome assembly algorithm and its applications to single-cell sequencing . J Comput Biol 2012, 19 (5):455–477. Walker BJ, Abeel T, Shea T, Priest M, Abouelliel A, Sakthikumar S, Cuomo CA, Zeng Q, Wortman JR, Young S: Pilon: an integrated tool for comprehensive microbial variant detection and genome assembly improvement . PLoS One 2014, 9 (11):e112963. Michael T, Pascal L, Tommaso P, Ulbricht-Jones ES, Axel F, Ralph B, Stephan G: GeSeq – versatile and accurate annotation of organelle genomes . Nucleic Acids Res 2017(W1):W1. Shi L, Chen H, Jiang M, Wang L, Wu X, Huang L, Liu C: CPGAVAS2, an integrated plastome sequence annotator and analyzer . Nucleic Acids Res 2019, 47 (W1):W65-W73. Chan PP, Lowe TMJMoMB: tRNAscan-SE: Searching for tRNA Genes in Genomic Sequences . Methods Mol Biol 2019, 1962 :1–14. Lohse M, Drechsel O, Bock R: OrganellarGenomeDRAW (OGDRAW): a tool for the easy generation of high-quality custom graphical maps of plastid and mitochondrial genomes . Curr Genet 2007, 52 (5):267–274. Beier S, Thiel T, Munch T, Scholz U, Mascher M: MISA-web: a web server for microsatellite prediction . Bioinformatics 2017, 33 (16):2583–2585. Li Q, Wan JM: [SSRHunter: development of a local searching software for SSR sites] . Yi Chuan 2005, 27 (5):808–810. Amiryousefi A, Hyvonen J, Poczai P: IRscope: an online program to visualize the junction sites of chloroplast genomes . Bioinformatics 2018, 34 (17):3030–3031. Afgan E, Baker D, Den Beek MV, Blankenberg D, Bouvier D, Cech M, Chilton J, Clements D, Coraor N, Eberhard C: The Galaxy platform for accessible, reproducible and collaborative biomedical analyses: 2016 update . Nucleic Acids Res 2016, 44 (W1):W3-W10. Wright F: The effective number of codons used in a gene . Gene 1990, 87 (1):23–29. Mower JP: The PREP suite: predictive RNA editors for plant mitochondrial genes, chloroplast genes and user-defined alignments . Nucleic Acids Res 2009, 37 :253–259. Wang Y, Yuan X, Zhang J: The complete chloroplast genome sequence of Pometia tomentosa . Mitochondrial DNA Part B 2019, 4 (2):3950–3951. Yang B, Li M, Ma J, Fu Z, Tian J: The complete chloroplast genome sequence of Sapindus mukorossi . Mitochondrial DNA Part A 2016, 27 (3):1825–1826. Frazer KA, Pachter L, Poliakov A, Rubin EM, Dubchak IJNAR: VISTA: computational tools for comparative genomics . Nucleic Acids Res 2004, 32 (Web Server issue):W273-279. Brudno M, Malde S, Poliakov A, Do CB, Couronne O, Dubchak I, Batzoglou S: Glocal alignment: finding rearrangements during alignment . Bioinformatics 2003, 19 :54–62. Saina JK, Gichira AW, Li ZZ, Hu GW, Wang QF, Liao K: The complete chloroplast genome sequence of Dodonaea viscosa: comparative and phylogenetic analyses . Genetica 2017, 146 (1):101–113. Du X, Xin G, Ren X, Liu H, Hao N, Jia G, Liu W: The complete chloroplast genome of Eurycorymbus cavaleriei (Sapindaceae), a Tertiary relic species endemic to China . Conserv Genet Resour 2018. Kim SC, Baek SH, Hong KN, Lee JW: Characterization of the complete chloroplast genome of Koelreuteria paniculata (Sapindaceae) . Conserv Genet Resour 2018, 10 (4):69–72. Chen SY, Zhang XZ: Characterization of the complete chloroplast genome of Xanthoceras sorbifolium, an endangered oil tree . Conserv Genet Resour 2017, 9 (4):595–598. Zhang Z, Xiao J, Wu J, Zhang H, Liu G, Wang X, Dai L: ParaAT: a parallel tool for constructing multiple protein-coding DNA alignments . Biochem Biophys Res Commun 2012, 419 (4):779–781. Wang D, Zhang Y, Zhang Z, Zhu J, Yu J: KaKs_Calculator 2.0: A Toolkit Incorporating Gamma-Series Methods and Sliding Window Strategies . Genomics,Proteomics & Bioinformatics 2010, 8 (1):77–80. Nei M, Gojobori T: Simple methods for estimating the numbers of synonymous and nonsynonymous nucleotide substitutions . Mol Biol Evol 1986, 3 (5):418–426. Katoh K, Rozewicki J, Yamada KD: MAFFT online service: multiple sequence alignment, interactive sequence choice and visualization . Brief Bioinform 2017. Darriba D, Taboada GL, Doallo R, Posada D: jModelTest 2: more models, new heuristics and parallel computing . Nat Methods 2012, 9 (8):772–772. Cummings MP: PAUP* (Phylogenetic Analysis Using Parsimony (and Other Methods)) . Dictionary of Bioinformatics Computational Biology 2004. Azim MK, Khan IA, Zhang Y: Characterization of mango (Mangifera indicaL.) transcriptome and chloroplast genome . Plant MolBiol 2014, 85 (1–2):193–208. Supplementary Files TableS1TableS8.xlsx Additional file 1: Table S1. The reference chloroplast genomes used in this study. Table S2. Gene composition in N. lappaceum chloroplast genome. Table S3. The introns and exons length of intron-containing genes in N. lappaceum chloroplast genome. Table S4. Total number of perfect simple sequence repeats (SSRs) identified within the chloroplast genome of Nephelium lappaceum. Table S5. The distribution pattern and number of simple sequence repeats (SSRs) identified within the chloroplast genome of Sapindaceae. Table S6. Long repeat sequences in the N.lappaceum chloroplast genome. Table S7. Codon usage for N. lappaceum chloroplast genome. Total 21434 codons, using 53 CDS. Table S8. RNA editing sites in the Nephelium lappaceum chloroplast genome. TableS9.xlsx Table S9. Ka/Ks ratio between pairwise of species protein coding sequences in nine Sapindaceae species. Cite Share Download PDF Status: Published Journal Publication published 02 Feb, 2021 Read the published version in Plants → Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-128918","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":6914341,"identity":"8da358c8-262c-4e8c-9ca0-fe91ffb91323","order_by":0,"name":"Fei Dong","email":"","orcid":"","institution":"College of Life Sciences, Fujian Agriculture and Forestry University, Fuzhou 350002, Fujian, China","correspondingAuthor":false,"prefix":"","firstName":"Fei","middleName":"","lastName":"Dong","suffix":""},{"id":6914342,"identity":"be2276ef-5d4b-40ba-9c9d-39208a7c4aa9","order_by":1,"name":"Zhicong Lin","email":"","orcid":"","institution":"College of Agriculture, Fujian Agriculture and Forestry University, Fuzhou, 350002, Fujian, China","correspondingAuthor":false,"prefix":"","firstName":"Zhicong","middleName":"","lastName":"Lin","suffix":""},{"id":6914343,"identity":"30caf2c2-8a85-4280-b3f0-12399c8a7905","order_by":2,"name":"Jing Lin","email":"","orcid":"","institution":"Center for Genomics and Biotechnology, Fujian Provincial Key Laboratory of Haixia Applied Plant Systems Biology, Key Laboratory of Genetics, Fujian Agriculture and Forestry University, Fuzhou, 350002","correspondingAuthor":false,"prefix":"","firstName":"Jing","middleName":"","lastName":"Lin","suffix":""},{"id":6914344,"identity":"963e4cdf-de65-401b-ba8d-e6f6ba57d50c","order_by":3,"name":"Ray Ming","email":"","orcid":"","institution":"Department of Plant Biology, University of Illinois at Urbana-Champaign, Urbana, IL 61801, USA","correspondingAuthor":false,"prefix":"","firstName":"Ray","middleName":"","lastName":"Ming","suffix":""},{"id":6914345,"identity":"c32126bf-037f-42fc-b018-68b502f0ef2b","order_by":4,"name":"Wenping Zhang","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABB0lEQVRIie3RPUvEMBjA8RyBuARvTfFLPBDIedzRfpWUwrnEW1wEtR4IdTl3Bz+IY6RwXaq3Bhy0ONqhk5ybab01paNg/pCX4fmRIQj5fH81aRe0l2aOxu3JBpPR/QIFq0EE7QmmQwgUd5v36vEqmhw8559zGTLQ+OmVonDpJOXLCcRlgafr5WKmZGIJSWYUJWcuIowSLM42BLQS/HSHU9BUHFGk45WLvNUdobCtBT+W1/aV8Vc/MbQllwyM4h9I5pZQ0kuiUk0gzjSAqcVoLQsW5IRPHyBxkuC2FNV3lkawVbzZyQt2WNxUpj4PncRG7Mq7y+934HYD9/yepN1s0zvn8/l8/7YfePVXFLZz5V8AAAAASUVORK5CYII=","orcid":"","institution":"Center for Genomics and Biotechnology, Fujian Provincial Key Laboratory of Haixia Applied Plant Systems Biology, Key Laboratory of Genetics, Fujian Agriculture and Forestry University, Fuzhou, 350002","correspondingAuthor":true,"prefix":"","firstName":"Wenping","middleName":"","lastName":"Zhang","suffix":""}],"badges":[],"createdAt":"2020-12-15 09:14:13","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-128918/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-128918/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.3390/plants10020283","type":"published","date":"2021-02-02T22:32:08+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":4547830,"identity":"dff6d0b6-b673-4dfa-8632-f75253b727d0","added_by":"auto","created_at":"2020-12-28 17:41:41","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":1127322,"visible":true,"origin":"","legend":"Gene map of N. lappaceum chloroplast genome. Genes drawn outside and inside of the circle are transcribed clockwise and counterclockwise, respectively. Genes belonging to different functional groups are color coded. The darker gray in the inner circle corresponds to GC content. SSC region, LSC region, and inverted repeats (IRA and IRB) are indicated.","description":"","filename":"Figure1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-128918/v1/86272523481b8af8362d94e6.jpg"},{"id":4547831,"identity":"aa6714d6-c05e-4be2-9bbb-82d4d9f0b422","added_by":"auto","created_at":"2020-12-28 17:41:42","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":367970,"visible":true,"origin":"","legend":"Analysis of simple sequence repeats (SSRs) in nine Sapindaceae (including N. lappaceum) chloroplast genomes. (A) Number of different SSRs types detected in nine Sapindaceae (including N. lappaceum) chloroplast genomes. (B) Presence of different SSRs types in all SSRs of nine Sapindaceae (including N. lappaceum) chloroplast genomes. ","description":"","filename":"Figure2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-128918/v1/d22672c3e5d6c3375d846c03.jpg"},{"id":4547834,"identity":"5911123d-3908-40c9-94cf-bdf1326c774a","added_by":"auto","created_at":"2020-12-28 17:41:42","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":389966,"visible":true,"origin":"","legend":"Codon content of 20 amino acids and stop codons in all protein-coding genes of N. lappaceum chloroplast genome. The colour of the histogram corresponds to the colour of codons.","description":"","filename":"Figure3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-128918/v1/cac340de0d9c2d99111d00e6.jpg"},{"id":4547996,"identity":"9c347b30-a09f-4dcd-a85b-c3e065974e70","added_by":"auto","created_at":"2020-12-28 17:44:42","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":1281390,"visible":true,"origin":"","legend":"Comparison of five Sapindoideae subfamily chloroplast genomes (including N. lappaceum), with D. longan as a reference. Gray arrows and thick black lines above the alignment indicate the direction of the gene. Purple bars represent exons, blue bars represent untranslated regions (UTRs), pink bars represent conserved non-coding sequences (CNS), and gray bars represent mRNA. The y-axis indicates the percent identity between 50% and 100%.","description":"","filename":"Figure4.jpg","url":"https://assets-eu.researchsquare.com/files/rs-128918/v1/cd5bd3f1371d55fcf3b4dc23.jpg"},{"id":4548053,"identity":"02dd032b-9572-4222-9fcf-96f7cf7fb578","added_by":"auto","created_at":"2020-12-28 17:47:42","extension":"jpg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":642330,"visible":true,"origin":"","legend":"Comparison of the borders of the LSC, SSC and IR regions among nine Sapindaceae chloroplast genomes. For each species, genes transcribed in positive strand are depicted on the top of their corresponding track from right to left direction, while the genes on the negative strand are depicted below from left to right. The numbers at arrows refer to the distance of the start or end position of a given gene from the corresponding junction site. The T bar above or below the genes indicate the extent of their parts with their corresponding values in base pair. The plotted genes and distances in the vicinity of the junction sites are the scaled projection of the genome. JLB (IRb /LSC), JSB (IRb/SSC), JSA (SSC/IRa) and JLA (IRa/LSC) denote the junction sites between each corresponding two regions of the genome.","description":"","filename":"Figure5.jpg","url":"https://assets-eu.researchsquare.com/files/rs-128918/v1/ebcbe7263d10a56b069662ea.jpg"},{"id":4547995,"identity":"fe7e240d-e9b3-45b2-9d95-b99a86ec1b66","added_by":"auto","created_at":"2020-12-28 17:44:42","extension":"jpg","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":893004,"visible":true,"origin":"","legend":"The Ka/Ks ratios of 78 protein-coding genes of the N. lappaceum chloroplast genome versus eight closely related species of Sapindaceae.","description":"","filename":"Figure6.jpg","url":"https://assets-eu.researchsquare.com/files/rs-128918/v1/b7568cd7a6430f55f0c67ece.jpg"},{"id":4548054,"identity":"ee850ec2-5ffe-491a-ab46-299be5cbf6bc","added_by":"auto","created_at":"2020-12-28 17:47:42","extension":"jpg","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":629445,"visible":true,"origin":"","legend":"The Maximum likelihood (ML) phylogenetic tree of the Sapindaceae family based on chloroplast genome sequences. The numbers in each node were tested by bootstrap analysis with 1000 replicates. Anacardium occidentale and Mangifera indica were set as the outgroups. The position of N. lappaceum is indicated in red text.","description":"","filename":"Figure7.jpg","url":"https://assets-eu.researchsquare.com/files/rs-128918/v1/755f96cb4a1504827e98a676.jpg"},{"id":59812742,"identity":"21923c9d-9103-4146-b37d-394ad9220a13","added_by":"auto","created_at":"2024-07-07 22:32:19","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":7698789,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-128918/v1/189af972-4e5d-4457-953b-92323ad12042.pdf"},{"id":4547994,"identity":"e45243e7-776d-41f1-8d37-7140c35fcb0b","added_by":"auto","created_at":"2020-12-28 17:44:42","extension":"xlsx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":106040,"visible":true,"origin":"","legend":"Additional file 1: Table S1. The reference chloroplast genomes used in this study. Table S2. Gene composition in N. lappaceum chloroplast genome. Table S3. The introns and exons length of intron-containing genes in N. lappaceum chloroplast genome. Table S4. Total number of perfect simple sequence repeats (SSRs) identified within the chloroplast genome of Nephelium lappaceum. Table S5. The distribution pattern and number of simple sequence repeats (SSRs) identified within the chloroplast genome of Sapindaceae. Table S6. Long repeat sequences in the N.lappaceum chloroplast genome. Table S7. Codon usage for N. lappaceum chloroplast genome. Total 21434 codons, using 53 CDS. Table S8. RNA editing sites in the Nephelium lappaceum chloroplast genome. ","description":"","filename":"TableS1TableS8.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-128918/v1/d9f0ec62a513cce78b2d5e8f.xlsx"},{"id":4547835,"identity":"dc300c47-3f15-4672-af8f-3f0056698c7a","added_by":"auto","created_at":"2020-12-28 17:41:42","extension":"xlsx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":118786,"visible":true,"origin":"","legend":"Table S9. Ka/Ks ratio between pairwise of species protein coding sequences in nine Sapindaceae species.","description":"","filename":"TableS9.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-128918/v1/d265c7baa3730d18329d48bd.xlsx"}],"financialInterests":"","formattedTitle":"\u003cp\u003eChloroplast Genome of Rambutan and Comparative Analyses in \u003cem\u003eSapindaceae\u003c/em\u003e\u003c/p\u003e","fulltext":[{"header":"Background","content":" \u003cp\u003eRambutan (\u003cem\u003eNephelium lappaceum\u003c/em\u003e L.) is an important tropical fruit in the family \u003cem\u003eSapindaceae\u003c/em\u003e and originated in Indonesia and Malay Peninsula [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. It is widely cultivated in Southeast Asia and the coastal areas of South China. Malaysians refer to it as \u0026ldquo;rambutan\u0026rdquo;, because of the fruit surface covered with thick and elongated spines. The fruits of rambutan are popular in the general population due to its rich nutrients, delicate and characteristic flavor and delicious taste. Rambutan peel extract is rich in phenolic content and exhibited antibacterial activity against many pathogenic bacteria, suggesting its antioxidant and/or antimicrobial properties[\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. Rambutan has the potential to be used as natural antioxidants and anti-aging agent in pharmaceutical and food industries to replace synthetic ones[\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThe \u003cem\u003eSapindaceae\u003c/em\u003e family contains over 150 genera and 2000 species with several economically important crops widely distributed in tropical and subtropical regions[\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. However, genomic research on \u003cem\u003eSapindaceae\u003c/em\u003e family, especially in the \u003cem\u003eN. lappaceum\u003c/em\u003e has been relatively scarce. This lack of genetic information making it difficult to meet the need for improving the quality and agronomic characteristics of rambutan through breeding and gene editing.\u003c/p\u003e \u003cp\u003eChloroplast (cp) are photosynthetic organelles that provide energy to green plants, it plays an important role in the photosynthesis and secondary metabolic activities[\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. The chloroplast genomes are maternally inherited in most plants and are highly conserved with its composition and sequence. The typical chloroplast genomes of angiosperms are circular DNA molecule which has a characteristic quadripartite structure with a large single-copy (LSC) region, a small single-copy (SSC) region, and two inverse repeats (IRs) regions[\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. The length of the genome is between 120 and 170\u0026nbsp;kb and usually encode 110 to 130 genes, and about 40 genes are specialized participating in photosynthesis, transcription and translation[\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e, \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. The first chloroplast genome from tobacco (\u003cem\u003eNicotiana tabacum\u003c/em\u003e) was sequenced in 1986[\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e], With the rapid development of next-generation sequencing technologies, the cost of whole genome sequencing is dropping rapidly[\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e]. Complete chloroplast genome sequences now could be easily acquired with relatively low cost. It has been an explosion in the number of available chloroplast genome sequences. Over 4,300 complete chloroplast genome sequences have been submitted in the National Center for Biotechnology Information (NCBI) organelle genome database.\u003c/p\u003e \u003cp\u003eWithin the \u003cem\u003eSapindaceae\u003c/em\u003e family, the complete chloroplast genomes of eight plant species have been sequenced and were available from the NCBI database. Nevertheless, no chloroplast genome in the genus \u003cem\u003eNephelium\u003c/em\u003e Linn has been reported. In this study, we report the first complete chloroplast genome of \u003cem\u003eN. lappaceum\u003c/em\u003e, exploring its general features, SSRs and long repeats, codon usage and analysis of IR contraction and expansion. In addition, nine chloroplast genome sequences were used for analysis of molecular evolution in the \u003cem\u003eSapindaceae\u003c/em\u003e family. We constructed a phylogenetic tree to understand the phylogenetic relationship of \u003cem\u003eSapindaceae\u003c/em\u003e plants. The chloroplast genome sequence and the comprehensive chloroplast genomic analysis of \u003cem\u003eN. lappaceum\u003c/em\u003e would provide a theoretical basis for molecular identification and further understanding of the evolutionary history of \u003cem\u003eSapindaceae\u003c/em\u003e family.\u003c/p\u003e "},{"header":"Results","content":"\u003cp\u003e\u003cstrong\u003eChloroplast Genome Features of \u003cspan class=\"BoldItalic\"\u003eN. lappaceum\u003c/span\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe structure of \u003cem\u003eN. lappaceum\u003c/em\u003e chloroplast genome was analogous to most chloroplast genomes of plants with a typical quadripartite structure. We assemble a closed circular chloroplast genome with 161,321\u0026nbsp;bp in \u003cem\u003eN. lappaceum\u003c/em\u003e. The chloroplast genome contains a pair of inverted repeat regions (IRs) of 28,550\u0026nbsp;bp, a large single-copy region (LSC) of 86,068\u0026nbsp;bp and a small single-copy region (SSC) of 18,153\u0026nbsp;bp (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e). The size of \u003cem\u003eN. lappaceum\u003c/em\u003e chloroplast genome was slightly larger than that in \u003cem\u003eS. mukorossi\u003c/em\u003e (160,481\u0026nbsp;bp), \u003cem\u003eP. tomentosa\u003c/em\u003e (160,818\u0026nbsp;bp), \u003cem\u003eD. Longan\u003c/em\u003e (160,833\u0026nbsp;bp) and shorter than that in \u003cem\u003eL. chinensis\u003c/em\u003e chloroplast genome (162,524\u0026nbsp;bp) of \u003cem\u003eSapindoideae\u003c/em\u003e (Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e). The number of chloroplast genes in \u003cem\u003eN. lappaceum\u003c/em\u003e was 132, the same with those in \u003cem\u003eD. Longan\u003c/em\u003e and \u003cem\u003eL. chinensis\u003c/em\u003e (Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e). In addition, there was no significant difference in GC content among the five analytical genomes in \u003cem\u003eSapindoideae\u003c/em\u003e.\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n\u003ctable id=\"Tab1\" border=\"1\"\u003e\u003ccaption\u003e\n\u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\n\u003cdiv class=\"CaptionContent\"\u003e\n\u003cp\u003eComparison of the general features of the five \u003cem\u003eSapindoideae\u003c/em\u003e chloroplast genomes.\u003c/p\u003e\n\u003c/div\u003e\n\u003c/caption\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eGenome feature\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003e\u003cem\u003eDimocarpus longan\u003c/em\u003e\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003e\u003cem\u003eLitchi chinensis\u003c/em\u003e\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003e\u003cem\u003ePometia tomentosa\u003c/em\u003e\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003e\u003cem\u003eSapindus mukorossi\u003c/em\u003e\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003e\u003cem\u003eNephelium lappaceum\u003c/em\u003e\u003c/p\u003e\n\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eGenBank\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eMG214255\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eKY635881\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eMN106254\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eKM454982\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eMT936934\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eSize (bp)\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e160833\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e162524\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e160818\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e160481\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e161321\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eLSC (bp)\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e85707\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e85750\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e85666\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e85650\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e86068\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eSSC (bp)\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e18270\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e16568\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e18360\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e18873\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e18153\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eIR (bp)\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e28428\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e30103\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e28396\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e27979\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e28550\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eTotal genes\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e132\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e132\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e133\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e135\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e132\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eProtein genes\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e87\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e87\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e88\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e88\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e87\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003etRNA genes\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e37\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e37\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e37\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e39\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e37\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003erRNA genes\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e8\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e8\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e8\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e8\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e8\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eGC (%)\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e37.79%\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e37.80%\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e37.87%\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e37.66%\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e37.77%\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003eThe overall nucleotide composition of rambutan is: 30.79% A, 31.44% T, 19.27% C, and 18.50% G, with a total GC content of 37.77%. In total, 132 genes were annotated on this chloroplast genome, including 78 protein-coding genes, 29 transfer RNA genes (tRNA) and 4 ribosomal RNA genes (rRNA). Among them, a total of 21 genes were found duplicated in the IR regions, including nine protein-coding genes (\u003cem\u003erps3, rps7, rps12, rps19, rpl2, rpl22, rpl23, ndhB\u003c/em\u003e and \u003cem\u003eycf2\u003c/em\u003e), eight tRNA genes (\u003cem\u003etrnA-UGC, trnI-CAU, trnI-GAU, trnL-CAA, trnM-CAU, trnN-GUU, trnR-ACG\u003c/em\u003e and \u003cem\u003etrnV-GAC\u003c/em\u003e) and four rRNA genes (\u003cem\u003errn4.5\u0026nbsp;s, rrn5s, rrn16s\u003c/em\u003e and \u003cem\u003errn23s\u003c/em\u003e) (Additional file 1: Table S2). The genes structure analysis showed that 21 genes contains introns, and 19 of them (11 protein-coding genes and 8 tRNA genes) have one intron, while two genes (\u003cem\u003eycf3\u003c/em\u003e and \u003cem\u003eclpP\u003c/em\u003e) have two introns (Additional file 1: Table S3).\u003c/p\u003e\n\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\n\u003ch2\u003eCharacterization of SSRs and repeat sequences\u003c/h2\u003e\n\u003cp\u003eA total of 63 SSRs were detected from rambutan chloroplast genome, of which 45 were mononucleotide, 3 dinucleotide, 8 trinucleotide, 5 tetranucleotide and two pentanucleotide (Additional file 1: Table S4). Moreover, we compared the distribution pattern and number of SSRs with eight other chloroplast genomes in \u003cem\u003eSapindaceae\u003c/em\u003e family (Additional file 1: Table S5). The number of mononucleotide repeats is more than the sum of other types (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003eA), and the number and types of chloroplast SSRs vary in different species. \u003cem\u003eS. mukorossi\u003c/em\u003e (91 SSRs) possess the highest number of SSRs while \u003cem\u003eE. cavaleriei\u003c/em\u003e (62 SSRs) possesses the lowest. Furthermore, the chloroplast genome of \u003cem\u003eD. longan\u003c/em\u003e, \u003cem\u003eL. chinensis\u003c/em\u003e, \u003cem\u003eP. tomentosa\u003c/em\u003e, \u003cem\u003eD. viscosa\u003c/em\u003e, \u003cem\u003eK. paniculate\u003c/em\u003e and \u003cem\u003eX. sorbifolium\u003c/em\u003e contained 79, 75, 74, 77, 87 and 83 SSRs, respectively (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003eB). In this study, a total of 98 larger repeats (\u0026gt;\u0026thinsp;10\u0026nbsp;bp) were identified in \u003cem\u003eN. lappaceum\u003c/em\u003e chloroplast genome composed of 42 forward, 11 reverse, 41 palindromic and 4 complement repeats (Additional file 1: Table S6) using REPuter[\u003cspan class=\"CitationRef\"\u003e12\u003c/span\u003e]. Among them, the largest repeat was a palindromic repeat with a size of 48\u0026nbsp;bp.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec4\" class=\"Section2\"\u003e\n\u003ch2\u003eCodon usage analysis and RNA editing sites prediction\u003c/h2\u003e\n\u003cp\u003eWe used 53 protein coding sequences from rambutan chloroplast genome for calculate codon usage frequency and relative synonymous codon usage (RSCU) frequency (Additional file 1: Table S7). All protein coding sequences contain 21,434 codons. In detail, leucine and cysteine are the highest and lowest number of amino acids, they have 2,232 codons (approximately 10.41% of the total) and 236 codons (approximately 1.10% of the total), respectively. While Met (ATG) and Trp (TGG) are encoded by only one codon showed no biased usage (RSCU\u0026thinsp;=\u0026thinsp;1). 30 codons with RSCU values more than 1, indicating they showed biased usage (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e). Among them, excluding the leucine (UUG) codon was G-ending, the remaining 29 biased usage codons of \u003cem\u003eN. lappaceum\u003c/em\u003e were all A/T-ending in the third codon. In addition, there were 49 potential RNA editing sites were found across 18 protein-coding genes in \u003cem\u003eN. lappaceum\u003c/em\u003e chloroplast genome and the \u003cem\u003endhB\u003c/em\u003e gene contained the most RNA editing sites (9) (Additional file 1: Table S8). We also observed that RNA editing sites were all C to U conversion, and took place at the first (30.6%) or second (69.4%) positions of the codons, indicating that editing in the third codon position disappeared quicker than that in the second or first codon position. Furthermore, serine codons were more frequently edited than codons of other amino acids and the conversion from serine to leucine occurred most frequently.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec5\" class=\"Section2\"\u003e\n\u003ch2\u003eComparative genomes analysis\u003c/h2\u003e\n\u003cp\u003eThe comparative analysis based on mVISTA was performed between the chloroplast genomes of rambutan with other four \u003cem\u003eSapindoideae\u003c/em\u003e species with the annotated \u003cem\u003eD. longan\u003c/em\u003e chloroplast genome as a reference. The five \u003cem\u003eSapindoideae\u003c/em\u003e subfamily chloroplast genomes length between the confines of 160,481 to 162,524\u0026nbsp;bp. The chloroplast genome of \u003cem\u003eL. chinensis\u003c/em\u003e has the largest size, whereas \u003cem\u003eS. mukorossi\u003c/em\u003e has the smallest size. Interestingly, the SSC region (16,568\u0026nbsp;bp) of \u003cem\u003eL. chinensis\u003c/em\u003e is the shortest, whereas the SSC region (18,873\u0026nbsp;bp) of \u003cem\u003eS. mukorossi\u003c/em\u003e chloroplast genome is the longest (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003e). The IR (A/B) regions exhibited less divergence than the SSC and LSC regions. In addition, the coding regions were more highly conserved than the non-coding regions. Among the five chloroplast genomes, four rRNA genes (\u003cem\u003errn16S, rrn23S, rrn5S, rrn4.5S\u003c/em\u003e) were the most conserved, while 7 genes (\u003cem\u003ematK, rpoC2, psbB, rpoA, ndhF, ndhD\u003c/em\u003e and \u003cem\u003eycf1\u003c/em\u003e) showed the most diversity in the coding regions. The highly divergent regions were found in the intergenic spacers and introns, including \u003cem\u003etrnH-GUG-psbA, trnR-UCU-atpA, petN-psbM, psbZ-trnG-GCC, ndhC-trnV-UAC, psbE-petL, rpl16-rps3\u003c/em\u003e and \u003cem\u003erpl32-trnL-UAG\u003c/em\u003e.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec6\" class=\"Section2\"\u003e\n\u003ch2\u003eExpansion and contraction of IR regions\u003c/h2\u003e\n\u003cp\u003eWe compared the IR regions and the junction sites of the LSC and SSC regions of nine \u003cem\u003eSapindaceae\u003c/em\u003e family chloroplast genomes (including \u003cem\u003eN. lappaceum\u003c/em\u003e) (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003e). The IR regions vary in different chloroplast genomes, ranging from 26,923\u0026nbsp;bp in \u003cem\u003eE. cavaleriei\u003c/em\u003e to 30,103\u0026nbsp;bp in \u003cem\u003eL. chinensis\u003c/em\u003e. In our study, the \u003cem\u003eycf1\u003c/em\u003e gene was located at the SSC/IRA junction in all of the nine chloroplast genomes and the fragment located at the IRa region ranged from 962\u0026nbsp;bp to 3,183\u0026nbsp;bp. Moreover, most junctions between LSC and IRa in this study was located downstream of the \u003cem\u003etrnH-GUG\u003c/em\u003e, except the \u003cem\u003eS. mukorossi.\u003c/em\u003e In addition, the LSC/IRb junction of three species \u003cem\u003eD. viscosa\u003c/em\u003e, \u003cem\u003eE. cavaleriei\u003c/em\u003e and \u003cem\u003eK. paniculate\u003c/em\u003e was located within the coding region of \u003cem\u003erpl22\u003c/em\u003e and created a location of 110, 40 or 63\u0026nbsp;bp at the LSC/IRb border. The remaining chloroplast genomes share a similar pattern, the LSC/IRb junction was located in intergenic regions of \u003cem\u003erpl16\u003c/em\u003e and \u003cem\u003erps3\u003c/em\u003e, and the IRb/SSC junction between IRb and SSC region (JSB) of five species (\u003cem\u003eS. mukorossi, X. sorbifoliun, D. viscosa\u003c/em\u003e, \u003cem\u003eE. cavaleriei\u003c/em\u003e and \u003cem\u003eK. paniculate\u003c/em\u003e) was located between the gene of ycf1 and \u003cem\u003endhF\u003c/em\u003e. However, other four chloroplast genomes only have \u003cem\u003endhF\u003c/em\u003e located or near the JSB.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec7\" class=\"Section2\"\u003e\n\u003ch2\u003eSynonymous (Ks) and non-synonymous (Ka) substitution rate analysis\u003c/h2\u003e\n\u003cp\u003eTo explore molecular evolution of orthologous genes shared by nine \u003cem\u003eSapindaceae\u003c/em\u003e species, particularly genes undergoing purifying or positive selection, we calculated the Ka/Ks ratio of 622 orthologous pairs with 78 protein coding genes (Additional file 1: Table S9). Overall, the average Ka/Ks ratio of the nine chloroplast genomes was 0.20. Total 612 orthologous pairs had a Ka/Ks ratio less than 1 in the nine comparison groups, out of which 546 orthologs had a Ka/Ks ratio less than 0.5 (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e6\u003c/span\u003e), suggesting that most genes are undergoing strong purifying selection pressures. Moreover, 66 orthologs of 31 genes with a Ka/Ks ratio between 0.5 and 1, 10 orthologous pairs of 6 genes (\u003cem\u003eccsA, rpoA, rps12, psbJ, clpPc\u003c/em\u003e and \u003cem\u003erps19\u003c/em\u003e) with a Ka/Ks ratio greater than 1 were detected in this study, suggesting that these genes might have experienced positive selection in the procedure of evolution. Among them, the Ka/Ks ratio of the \u003cem\u003eycf1\u003c/em\u003e gene was greater than 0.5 in eight comparison groups, the \u003cem\u003erpoA\u003c/em\u003e and \u003cem\u003eycf2\u003c/em\u003e gene with Ka/Ks ratio greater than 0.5 was also observed in the comparison of seven and six groups, respectively. Besides, \u003cem\u003eclpP, matK\u003c/em\u003e and \u003cem\u003erps15\u003c/em\u003e gene with Ka/Ks ratio\u0026thinsp;\u0026gt;\u0026thinsp;0.5 were founded in four out of the eight comparison groups.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec8\" class=\"Section2\"\u003e\n\u003ch2\u003ePhylogenetic analysis\u003c/h2\u003e\n\u003cp\u003eWe performed multiple sequence alignments using the whole chloroplast genome sequences of nine \u003cem\u003eSapindaceae\u003c/em\u003e species and two \u003cem\u003eAnacardiaceae\u003c/em\u003e species as outgroups (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e7\u003c/span\u003e). All nodes in the ML trees have 100% bootstrap support values, and these 11 chloroplast genome sequences are clustered into three groups. In detail, the five species (\u003cem\u003eD. longan, L. chinensis, P. tomentosa, N. lappaceum\u003c/em\u003e and \u003cem\u003eS. mukorossi\u003c/em\u003e) from \u003cem\u003eSapindoideae\u003c/em\u003e clustered into one group, four species (\u003cem\u003eK. paniculata, D. viscosa, E. cavaleriei\u003c/em\u003e and \u003cem\u003eX. sorbifolium\u003c/em\u003e) from \u003cem\u003eDodonaeoide\u003c/em\u003eae are in one group, and the two species (\u003cem\u003eA. occidentale\u003c/em\u003e and \u003cem\u003eM. indica\u003c/em\u003e) in \u003cem\u003eAnacardiaceae\u003c/em\u003e are cluster into one group. In the \u003cem\u003eSapindoideae\u003c/em\u003e group, the \u003cem\u003eN. lappaceum\u003c/em\u003e chloroplast genome sequence showed the closest relationship with \u003cem\u003eP. tomentosa\u003c/em\u003e, followed by \u003cem\u003eD. longan\u003c/em\u003e and \u003cem\u003eL. chinesis\u003c/em\u003e, as far as \u003cem\u003eS. mukorossi\u003c/em\u003e. The three groups of this phylogenetic tree of the 11 chloroplast genome sequences are consistent with traditional taxonomy, suggesting that the chloroplast genome could effectively resolve the phylogenetic positions and relationships of species.\u003c/p\u003e\n\u003c/div\u003e"},{"header":"Disscussion","content":" \u003cp\u003eWe assembled \u003cem\u003eN. lappaceum\u003c/em\u003e complete chloroplast genome sequence and deposited it to GenBank under accession number: MT936934, \u003cem\u003eN. lappaceum\u003c/em\u003e chloroplast genome was consistent with the characteristics of most angiosperm species in structure and gene content. Although there are some differences in the sizes of the overall genome, LSC, SSC and IR regions, the numbers of genes and GC content are similar among the five \u003cem\u003eSapindoideae\u003c/em\u003e chloroplast genomes, which to some extent reflects the high conservation of angiosperm chloroplast genomes[\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. Intron plays an important role in RNA stability, regulation of gene expression and alternative splicing which have been reported in many other species[\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e, \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. There were two genes (\u003cem\u003eycf3\u003c/em\u003e and \u003cem\u003eclpP\u003c/em\u003e) included two introns in the \u003cem\u003eN. lappaceum\u003c/em\u003e chloroplast genome. It has been reported that \u003cem\u003eycf3\u003c/em\u003e gene is essential for the accumulation of the photosystem I (PSI) complex and acts a chaperone that interacts with the PSI subunits at a post-translational level[\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e, \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]. Besides, \u003cem\u003eclpP\u003c/em\u003e gene functions as the proteolytic subunit of the ATP-dependent Clp protease in plant chloroplasts and is essential for the development and/or function of plastids with active gene expression in previous studies[\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e, \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. Thus, study of \u003cem\u003eycf3\u003c/em\u003e and \u003cem\u003eclpP\u003c/em\u003e gene will contribute to further investigation of chloroplast in \u003cem\u003eN. lappaceum\u003c/em\u003e.\u003c/p\u003e \u003cp\u003eSimple sequence repeats (SSRs), also known as microsatellites, are tandem repeats (1\u0026thinsp;~\u0026thinsp;6\u0026nbsp;bp units repeated multiple times) distributed across the entire genome which have been widely applied as molecular markers for determining genetic variations across species and evolutionary studies because of its unique uniparental in inheritance[\u003cspan additionalcitationids=\"CR20\" citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]. we identified 63 SSRs in \u003cem\u003eN. lappaceum\u003c/em\u003e chloroplast genome and most of SSRs were distributed in IGS regions. Mononucleotide SSRs were identified most frequently (68% on average) among the nine analyzed chloroplast genomes of \u003cem\u003eSapindaceae\u003c/em\u003e species, and vast majority of mononucleotide repeats consist of short polyA or polyT repeats sharing a similar pattern in most angiosperm chloroplast genomes[\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e, \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e]. Moreover, repetitive sequences are helpful in phylogenetic study and play a vital role in genome rearrangement[\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e]. These results can provide chloroplast molecular markers that can be used to quickly identify species, confirm hybrid progeny when breeding.\u003c/p\u003e \u003cp\u003eCodon usage biases are found in all eukaryotic and prokaryotic genomes and have been proposed to regulate different aspects of translation process[\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]. High RSCU values of the codons are probably attributed to amino acid functions or peptide structures that avoid transcriptional errors in chloroplast genomes[\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e, \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]. There are 30 codons showed biased usage and most of them were A/T-ending in the third codon. This phenomenon also exists in previous studies[\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e]. Codon usage bias of chloroplast genome reflects a selective pressure to increase translation efficiency[\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e], and research on codon preferences can help us to better understand gene expression and molecular evolution mechanisms of \u003cem\u003eN. lappaceum.\u003c/em\u003e\u003c/p\u003e \u003cp\u003eWe observed \u003cem\u003endhB\u003c/em\u003e gene contained the most RNA editing sites within the 49 potential RNA editing sites, and 16 editing sites were U_A type, indicating there was a U_A bias for the distribution of RNA editing sites that was in accordance with previous reports in other species[\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e, \u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e]. RNA editing is a post-transcriptional regulation pattern involved in the insertion, deletion, or modification of nucleotides that widely exists in land plants[\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e]. The first chloroplast RNA editing event of land plant was discovered in the mRNA transcript of \u003cem\u003erpl2\u003c/em\u003e gene in maize chloroplast genome in 1991[\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]. The most frequent editing events in plants are C-to-U changes, however, U-to-C editing has also been observed[\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e, \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e]. Additionally, RNA editing usually occurs in the first or second base of codons, resulting in the conversion of hydrophilic amino acid to hydrophobic[\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eComparative analysis of chloroplast genomes is an essential step in genomics which can provide insight into complex evolutionary relationships. The mVISTA analysis showed that five \u003cem\u003eSapindoideae\u003c/em\u003e chloroplast genomes were conserved, with a high degree of synteny and gene order conservation, and the coding region was more conserved than the non-coding region, which is consistent with reports on other angiosperms[\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e], suggesting an evolutionary conservation of these genomes at the genome-scale level. In addition, \u003cem\u003eycf1\u003c/em\u003e gene showed the greatest degree of differentiation. Previous study reported that \u003cem\u003eycf1\u003c/em\u003e is helpful to provide phylogenetic information at the species level and more variable than \u003cem\u003ematK\u003c/em\u003e in \u003cem\u003eOrchidaceae\u003c/em\u003e[\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e]. Furthermore, \u003cem\u003eycf1\u003c/em\u003e performed better to identify DNA barcodes of high resolution at species level than any of the \u003cem\u003ematK\u003c/em\u003e, \u003cem\u003erbcL\u003c/em\u003e and \u003cem\u003etrnH-psbA\u003c/em\u003e[\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e]. These variable genic regions found in our study can be regarded as molecular markers for DNA barcoding and phylogenetic studies in \u003cem\u003eSapindaceae\u003c/em\u003e.\u003c/p\u003e \u003cp\u003eAlthough most land plants have relatively conserved cp genomes, the end of the inverted repeats (IRa and IRb) regions differs among various plant species. The expansion or contraction of the IR regions represent important evolutionary events often results in size variation of different chloroplast genomes and is helpful to studying the chloroplast genome evolution history[\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e, \u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e]. In this study, our results suggested that the boundary of IR/LSC and IR/SSC might be conversed among chloroplast genomes of closely related family species but some differences also occurs between relatively distantly related family species, such as gene overlap length, duplicate of the \u003cem\u003eycf1\u003c/em\u003e and \u003cem\u003erps3\u003c/em\u003e genes, even the distance of \u003cem\u003etrnH-GUG\u003c/em\u003e from the border near the LSC/IRB junctions, indicating that the expansion and contraction of the IR region led to length and structure changes of chloroplast genomes.\u003c/p\u003e \u003cp\u003eThe ratio between nonsynonymous (Ka) and synonymous (Ks) nucleotide substitution has been widely used as important markers in genome or gene evolution studies[\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. Ka/Ks\u0026thinsp;=\u0026thinsp;1 signifies neutral evolution, Ka/Ks\u0026thinsp;\u0026gt;\u0026thinsp;1 indicates that the gene is affected by positive selection, whereas Ka/Ks\u0026thinsp;\u0026lt;\u0026thinsp;1 indicates that the gene is affected by purifying selection[\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e]. Additionally, a Ka/Ks ratio of 0.5 was considered as a useful cut-off value to identify genes under positive selection in previous studies[\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e]. In our study, the \u003cem\u003eccsA, rpoA, rps12, psbJ, clpP\u003c/em\u003e and \u003cem\u003erps19\u003c/em\u003e gene with Ka/Ks\u0026thinsp;\u0026gt;\u0026thinsp;1. It is noteworthy that the \u003cem\u003eycf1\u003c/em\u003e gene also exhibited high Ka/Ks ratios with Ka/Ks\u0026thinsp;\u0026gt;\u0026thinsp;0.5 in eight comparison groups. This result is in keeping with the previous observations that the \u003cem\u003eycf1\u003c/em\u003e gene was more variable than the \u003cem\u003ematK\u003c/em\u003e and \u003cem\u003erbcL\u003c/em\u003e genes in most plant, and could be using as an effective biological tool for plant phylogeny study[\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e]. The positive selection of genes in \u003cem\u003eN. lappaceum\u003c/em\u003e possibly provided help for adaptations to its particular living environment.\u003c/p\u003e \u003cp\u003eNumerous studies have shown that chloroplast genome sequences have been successfully used in taxonomic and phylogenetic studies[\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e], and contribute to describe the evolutionary relationships between species[\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e]. In this study, the topology of the trees consists of two main branches: \u003cem\u003eDodonaeoideae\u003c/em\u003e and evolutionary younger \u003cem\u003eSapindoideae.\u003c/em\u003e And generic relationships of the two subfamilies are basically congruent with the taxonomy of these families. The availability of the completed \u003cem\u003eN. lappaceum\u003c/em\u003e chloroplast genome provided us with sequence information that can be used to confirm the phylogenetic position of \u003cem\u003eN. lappaceum\u003c/em\u003e and understand the phylogenetic relationships among \u003cem\u003eSapindaceae\u003c/em\u003e. However, we had used only a small number of species in \u003cem\u003eSapindaceae\u003c/em\u003e, further research on other chloroplast genome as well as nuclear genome sequences of \u003cem\u003eSapindaceae\u003c/em\u003e should be sequenced to provide more sufficient evidence to accurately illustrate the evolution of family \u003cem\u003eSapindaceae.\u003c/em\u003e\u003c/p\u003e "},{"header":"Conclusions","content":" \u003cp\u003eWe assembled the first complete chloroplast genome of rambutan using Illumina sequencing technology and compared its structure with other \u003cem\u003eSapindaceae\u003c/em\u003e species. The chloroplast genome of \u003cem\u003eN. lappaceum\u003c/em\u003e exhibits similar quadripartite structure, gene order, G\u0026thinsp;+\u0026thinsp;C content when compared with other \u003cem\u003eSapindaceae\u003c/em\u003e chloroplast genomes. A total of 63 SSRs and 98 repeat sequences were identified in \u003cem\u003eN. lappaceum\u003c/em\u003e chloroplast genome. The research on codon usage of \u003cem\u003eN. lappaceum\u003c/em\u003e shows that some amino acids have obvious codon usage bias and the codon preferences may help us understand the evolution mechanisms of \u003cem\u003eN. lappaceum\u003c/em\u003e. With PREP prediction, we detected 49 RNA editing loci in 18 protein-coding genes in \u003cem\u003eN. lappaceum.\u003c/em\u003e Moreover, the expansion and contraction of the IR regions, leading to the variations in nine \u003cem\u003eSapindaceae\u003c/em\u003e chloroplast genome size. There are 6 genes (\u003cem\u003eccsA, rpoA, rps12, psbJ, clpP\u003c/em\u003e and \u003cem\u003erps19\u003c/em\u003e) were detected with a Ka/Ks ratio\u0026thinsp;\u0026gt;\u0026thinsp;1, suggesting that these genes experienced positive selection in the evolution. Additionally, phylogenetic analysis using 9 complete chloroplast genome sequences in \u003cem\u003eSapindaceae\u003c/em\u003e strongly supports the close relationship of \u003cem\u003eN. lappaceum\u003c/em\u003e and \u003cem\u003eP. tomentosa\u003c/em\u003e among sequenced chloroplast genomes in \u003cem\u003eSapindaceae\u003c/em\u003e.\u003c/p\u003e "},{"header":"Methods","content":"\u003cp\u003e\u003cstrong\u003ePlant material, DNA extraction, and sequencing\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eYoung, healthy leaves of the major cultivar of rambutan, Baoyan7, were collected from Baoting (N18\u0026deg;23\u0026prime;, E109\u0026deg;21\u0026prime;) in Hainan Province, China. The leaves were frozen in liquid nitrogen, and maintained at \u0026minus;\u0026thinsp;80 \u003csup\u003e◦\u003c/sup\u003eC. The total genomic DNA was extracted by 2X cetyltrimethylammonium bromide (CTAB) method [43]. And a library with insert sizes of 300\u0026ndash;500\u0026nbsp;bp was constructed and then sequenced on an Illumina HiSeq2500 platform (Illumina, San Diego, CA, USA) double terminal sequencing method (150 pair-ends).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eChloroplast genome assembly and annotation\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eFirst, FastQC software was performed to evaluate the quality of Illumina paired-end raw reads[44], and low-quality reads were filtered. The remaining clean reads were used for assembly via NOVOPlasty[45] using \u003cem\u003eDimocarpus longan\u003c/em\u003e chloroplast genome(GenBank: MG214255)[46] as the reference genome to generate the first version of rambutan genome. Next, all clean reads were mapped onto the first version genome and the mapped reads were assembled using SPAdes3.14.1[47] and assembled contigs were corrected using the pair-end short reads from HiSeq2500 by Pilon version 1.23 (https://github.com/broadinstitute/pilon)[48] to generate the second version of rambutan chloroplast genome. These two versions were compared and mutually corrected to get the final complete rambutan chloroplast genome.\u003c/p\u003e\n\u003cp\u003eThe chloroplast genome was annotated by online program GeSeq (https://chlorobox.mpimp-golm.mpg.de/geseq.html)[49] and CPGAVAS2[50]. Genome features like start/stop codons and intron/exon borders were manually corrected through the comparison of other reported \u003cem\u003eSapindaceae\u003c/em\u003e family chloroplast genomes. In addition, tRNA genes were identified by tRNAscan-SE 2.0 (http://lowelab.ucsc.edu/tRNAscan-SE/)[51]. A circular map of the revised annotated rambutan chloroplast genome was illustrated by using Organellar Genome DRAW (OGDRAW) (https://chlorobox.mpimpgolm.mpg.de/OGDraw.html)[52].\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eChloroplast genome analysis\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe simple sequence repeats (SSR) in nine chloroplast genome sequences of \u003cem\u003eSapindaceae\u003c/em\u003e (including \u003cem\u003eN. lappaceum\u003c/em\u003e) (Additional file 1: Table S1) were identified by online tool MISA (https://webblast.ipk-gatersleben.de/misa/)[53] and the threshold settings were as follows: ten was applied to mononucleotide repeats, five to dinucleotide repeats and four to trinucleotide repeats, three for tetra-, penta-, and hexanucleotide repeats[54]. Repetitive sequences including forward, reverse, palindrome, and complement sequences were analyzed by REPuter program[12], and the parameter was set with a minimum length of repeat region set to 10\u0026nbsp;bp and the minimum sequence identity set to 90%.\u003c/p\u003e\n\u003cp\u003eThe expansion and contraction of the inverted repeat (IR) regions at junction sites from eight \u003cem\u003eSapindaceae\u003c/em\u003e family chloroplast genome sequences, including \u003cem\u003eDimocarpus longan\u003c/em\u003e (MG214255), \u003cem\u003eLitchi chinensis\u003c/em\u003e (KY635881), \u003cem\u003ePometia tomentosa\u003c/em\u003e (MN106254), \u003cem\u003eSapindus mukorossi\u003c/em\u003e (KM454982), \u003cem\u003eDodonaea viscosa\u003c/em\u003e(MF155892), \u003cem\u003eEurycorymbus cavaleriei\u003c/em\u003e (MG813997), \u003cem\u003eKoelreuteria paniculate\u003c/em\u003e (KY859413) and \u003cem\u003eXanthoceras sorbifolium\u003c/em\u003e (KY779850), were examined and plotted using IRscope online program (https://irscope.shinyapps.io/irapp/)[55]. Codon usage of the \u003cem\u003eN. lappaceum\u003c/em\u003e chloroplast genome was analyzed via GALAXY platform (https://galaxy.pasteur.fr) [56] with CodonW online tool. The length of protein-coding gene less than 300 nucleotide and the repetitive genes sequences were removed to reduce deviation of the results[57]. Finally, 53 CDS in \u003cem\u003eN. lappaceum\u003c/em\u003e were selected for further codon usage analysis. Besides, putative RNA editing sites were predicted using PREP-Cp web server (http://prep.unl.edu/cgi-bin/cp-input.pl)[58] with a cutoff value of 0.8.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eGenome Comparison\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe downloaded four whole chloroplast genome sequences of \u003cem\u003eSapindoideae\u003c/em\u003e subfamily from the National Center for Biotechnology Information (NCBI) Organelle Genome and Nucleotide Resources database, including \u003cem\u003eD. longan\u003c/em\u003e[46], \u003cem\u003eL. chinensis\u003c/em\u003e, \u003cem\u003eP. tomentosa\u003c/em\u003e[59] and \u003cem\u003eS. mukorossi\u003c/em\u003e[60]. The mVISTA online program (Shuffle-LAGAN mode)[61, 62] was used to compare chloroplast genome sequence of rambutan with the other species from \u003cem\u003eSapindoideae\u003c/em\u003e subfamily, in which the annotation of \u003cem\u003eD. longan\u003c/em\u003e as the reference.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ePositive selection analysis of protein sequence\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe analyzed synonymous (Ks) and non-synonymous (Ka) substitution rates to investigate the molecular evolutionary process of \u003cem\u003eSapindaceae\u003c/em\u003e family, The protein-coding genes of \u003cem\u003eN. lappaceum\u003c/em\u003e were separately compared with eight closely related species in \u003cem\u003eSapindaceae\u003c/em\u003e family: \u003cem\u003eD. longan\u003c/em\u003e, \u003cem\u003eL. chinensis\u003c/em\u003e, \u003cem\u003eP. tomentosa\u003c/em\u003e, \u003cem\u003eS. mukorossi, D. viscosa\u003c/em\u003e[63], \u003cem\u003eE. cavaleriei\u003c/em\u003e[64], \u003cem\u003eK. paniculate\u003c/em\u003e[65] and \u003cem\u003eX. sorbifolium\u003c/em\u003e[66] using ParaAT 2.0[67], then the Ka/Ks value was calculated by KaKs_calculator 2.0[68] with NG method[69].\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ePhylogenetic analysis\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eIn order to deeply detect the evolutionary relationship of \u003cem\u003eSapindaceae\u003c/em\u003e family, we aligned 9 complete chloroplast genomes (including \u003cem\u003eN. lappaceum\u003c/em\u003e) with MAFFT version 7[70]. The best fitting nucleotide substitution model (TVM\u0026thinsp;+\u0026thinsp;I\u0026thinsp;+\u0026thinsp;G ) was chosen by jModelTest v2.1.7[71]. Phylogenetic analysis was then inferred by ML (maximum-likelihood) method based on the TVM\u0026thinsp;+\u0026thinsp;I\u0026thinsp;+\u0026thinsp;G substitution model in PAUP* 4.0[72] with 1000 bootstrap replicates. \u003cem\u003eAnacardium occidentale\u003c/em\u003e (KY635877) and \u003cem\u003eMangifera indica\u003c/em\u003e (KY635882)[73] in \u003cem\u003eAnacardiaceae\u003c/em\u003e family were set as the outgroup.\u003c/p\u003e"},{"header":"Abbreviations","content":" \u003cdiv class=\"DefinitionList\"\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eCp\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eChloroplast; SSRs:Simple sequence repeats; IRs:Inverted repeats; LSC:Large single-copy; SSC:Small single-copy; ML:Maximum-likelihood\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003c/div\u003e "},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis work was supported by startup fund from Fujian Agriculture and Forestry University.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eF.D.\u0026nbsp;performed\u0026nbsp;most\u0026nbsp;of\u0026nbsp;the\u0026nbsp;data\u0026nbsp;analysis,\u0026nbsp;and\u0026nbsp;wrote\u0026nbsp;the\u0026nbsp;manuscript.\u0026nbsp;W.Z.\u0026nbsp;collected\u0026nbsp;experiment\u0026nbsp;materials\u0026nbsp;and\u0026nbsp;data.\u0026nbsp;\u003cbr /\u003eZ.\u0026nbsp;L.\u0026nbsp;and\u0026nbsp;J.\u0026nbsp;L.\u0026nbsp;helped\u0026nbsp;in\u0026nbsp;genome\u0026nbsp;assembly\u0026nbsp;strategy\u0026nbsp;design.\u0026nbsp;R.M.,\u0026nbsp;W.\u0026nbsp;Z.\u0026nbsp;and\u0026nbsp;F.\u0026nbsp;D.\u0026nbsp;designed\u0026nbsp;the\u0026nbsp;project\u0026nbsp;and\u0026nbsp;revised\u0026nbsp;the\u0026nbsp;manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of Data and Materials\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe chloroplast genome assembly and annotation were deposited in GenBank under the accession number of MT936934. All other data and material generated in this manuscript are available from the corresponding author upon reasonable request.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eLim TK: \u003cb\u003eEdible Medicinal and Non Medicinal Plants\u003c/b\u003e 2015: Springer; 2012.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePalanisamy UD, Cheng HM, Masilamani T, Subramaniam T, Ling LT, Radhakrishnan AK: \u003cb\u003eRind of the rambutan, Nephelium lappaceum, a potential source of natural antioxidants\u003c/b\u003e. \u003cem\u003eFood Chem\u003c/em\u003e 2008, \u003cb\u003e109\u003c/b\u003e(1):54\u0026ndash;63.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhuang Y, Ma Q, Guo Y, Sun L: \u003cb\u003eProtective effects of rambutan (Nephelium lappaceum) peel phenolics on H2O2-induced oxidative damages in HepG2 cells and d-galactose-induced aging mice\u003c/b\u003e. \u003cem\u003eFood Chemical Toxicology\u003c/em\u003e 2017, \u003cb\u003e108\u003c/b\u003e(Pt B):554\u0026ndash;562.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ec NNMPab, B TTL, A JVC, A KR: \u003cb\u003eEvaluation of antimicrobial activity of rambutan (Nephelium lappaceum L.) peel extracts\u003c/b\u003e. \u003cem\u003eInt J Food Microbiol\u003c/em\u003e 2020, \u003cb\u003e321\u003c/b\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHarrington MG, Edwards KJ, Johnson SA, Chase MW, Gadek PA: \u003cb\u003ePhylogenetic Inference in Sapindaceae sensu lato Using Plastid matK and rbcL DNA Sequences\u003c/b\u003e. \u003cem\u003eSyst Bot\u003c/em\u003e 2005, \u003cb\u003e30\u003c/b\u003e(2):366\u0026ndash;382.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWicke S, Schneeweiss GM, Depamphilis CW, Kai FM, Quandt D: \u003cb\u003eThe evolution of the plastid chromosome in land plants: gene content, gene order, gene function\u003c/b\u003e. \u003cem\u003ePlant MolBiol\u003c/em\u003e 2011, \u003cb\u003e76\u003c/b\u003e(3\u0026ndash;5):273\u0026ndash;297.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBobik K, Burch-Smith TM: \u003cb\u003eChloroplast signaling within, between and beyond cells\u003c/b\u003e. \u003cem\u003eFront Plant Sci\u003c/em\u003e 2015, \u003cb\u003e6\u003c/b\u003e:781.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWolfe KH, Li W, Sharp PM: \u003cb\u003eRates of nucleotide substitution vary greatly among plant mitochondrial, chloroplast, and nuclear DNAs\u003c/b\u003e. \u003cem\u003eProc Natl Acad Sci U S A\u003c/em\u003e 1987, \u003cb\u003e84\u003c/b\u003e(24):9054\u0026ndash;9058.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePalmer JD: \u003cb\u003eComparative Organization of Chloroplast Genomes\u003c/b\u003e. \u003cem\u003eAnnu Rev Genet\u003c/em\u003e 1985, \u003cb\u003e19\u003c/b\u003e(1):325\u0026ndash;354.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShinozaki K, Ohme M, Tanaka M, Wakasugi T, Sugiura M: \u003cb\u003eThe complete nucleotide sequence of the tobacco chloroplast genome: its gene organization and expression\u003c/b\u003e. \u003cem\u003ePlant Mol Biol Rep\u003c/em\u003e 1986, \u003cb\u003e5\u003c/b\u003e(9):2043\u0026ndash;2049.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi C, Lin F, An D, Wang W, Huang R: \u003cb\u003eGenome Sequencing and Assembly by Long Reads in Plants\u003c/b\u003e. \u003cem\u003eGenes\u003c/em\u003e 2017, \u003cb\u003e9\u003c/b\u003e(1):6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKurtz S, Choudhuri JV, Ohlebusch E, Schleiermacher C, Stoye J, Giegerich R: \u003cb\u003eREPuter: the manifold applications of repeat analysis on a genomic scale\u003c/b\u003e. \u003cem\u003eNucleic Acids Res\u003c/em\u003e 2001, \u003cb\u003e29\u003c/b\u003e(22):4633\u0026ndash;4642.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNguyen Dinh S, Sai TZT, Nawaz G, Lee K, Kang H: \u003cb\u003eAbiotic stresses affect differently the intron splicing and expression of chloroplast genes in coffee plants (Coffea arabica) and rice (Oryza sativa)\u003c/b\u003e. \u003cem\u003eJ Plant Physiol\u003c/em\u003e 2016, \u003cb\u003e201\u003c/b\u003e:85\u0026ndash;94.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMirzaei S, Mansouri M, Mohammadi-Nejad G, Sablok G: \u003cb\u003eComparative assessment of chloroplast transcriptional responses highlights conserved and unique patterns across Triticeae members under salt stress\u003c/b\u003e. \u003cem\u003ePhotosynth Res\u003c/em\u003e 2017, \u003cb\u003e136\u003c/b\u003e(3):357\u0026ndash;369.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNaver H, Boudreau E, Rochaix JD: \u003cb\u003eFunctional studies of YCF3: Its role in assembly of photosystem I and Interactions with some of its subunits\u003c/b\u003e. \u003cem\u003ePlant Cell\u003c/em\u003e 2002, \u003cb\u003e13\u003c/b\u003e(12):2731\u0026ndash;2745.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBoudreau E, Takahashi Y, Lemieux C, Turmel M, Rochaix JD: \u003cb\u003eThe chloroplast ycf3 and ycf4 open reading frames of Chlamydomonas reinhardtii are required for the accumulation of the photosystem I complex\u003c/b\u003e. \u003cem\u003eEmbo J\u003c/em\u003e 1997, \u003cb\u003e16\u003c/b\u003e(20):6095\u0026ndash;6104.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eClarke AK, Schelin J, Porankiewicz J: \u003cb\u003eInactivation of the clpP1 gene for the proteolytic subunit of the ATP-dependent Clp protease in the cyanobacterium Synechococcus limits growth and light acclimation\u003c/b\u003e. \u003cem\u003ePlant MolBiol\u003c/em\u003e 1998, \u003cb\u003e37\u003c/b\u003e(5):791\u0026ndash;801.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBruce CA, Cunningham KA, Stern DB: \u003cb\u003eThe plastid clpP gene may not be essential for plant cell viability\u003c/b\u003e. \u003cem\u003ePlant Cell Physiol\u003c/em\u003e 2003(1):93\u0026ndash;95.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVarshney RK, Sigmund R, Borner A, Korzun V, Stein N, Sorrells ME, Langridge P, Graner A: \u003cb\u003eInterspecific transferability and comparative mapping of barley EST-SSR markers in wheat, rye and rice\u003c/b\u003e. \u003cem\u003ePlant Sci\u003c/em\u003e 2005, \u003cb\u003e168\u003c/b\u003e(1):195\u0026ndash;202.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang A, Zhang J, Tian H, Yao X: \u003cb\u003eCharacterization of 39 novel EST-SSR markers for Liriodendron tulipifera and cross-species amplification in L. chinense (Magnoliaceae)\u003c/b\u003e. \u003cem\u003eAm J Bot\u003c/em\u003e 2012, \u003cb\u003e99\u003c/b\u003e(11):e460-464.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi B, Lin F, Huang P, Guo W, Zheng Y: \u003cb\u003eDevelopment of nuclear SSR and chloroplast genome markers in diverse Liriodendron chinense germplasm based on low-coverage whole genome sequencing\u003c/b\u003e. \u003cem\u003eBiol Res\u003c/em\u003e 2020, \u003cb\u003e53\u003c/b\u003e(1):21.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGao B, Yuan L, Tang T, Hou J, Pan K, Wei N: \u003cb\u003eThe complete chloroplast genome sequence of Alpinia oxyphylla Miq. and comparison analysis within the Zingiberaceae family\u003c/b\u003e. \u003cem\u003ePLoS One\u003c/em\u003e 2019, \u003cb\u003e14\u003c/b\u003e(6).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYan C, Du J, Gao L, Li Y, Hou X: \u003cb\u003eThe complete chloroplast genome sequence of watercress (Nasturtium officinale R. Br.): Genome organization, adaptive evolution and phylogenetic relationships in Cardamineae\u003c/b\u003e. \u003cem\u003eGene\u003c/em\u003e 2019, \u003cb\u003e699\u003c/b\u003e:24\u0026ndash;36.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCavalier-Smith T: \u003cb\u003eChloroplast Evolution: Secondary Symbiogenesis and Multiple Losses\u003c/b\u003e. \u003cem\u003eCurr Biol\u003c/em\u003e 2002, \u003cb\u003e12\u003c/b\u003e(2):R62-R64.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHershberg R, Petrov DA: \u003cb\u003eSelection on codon bias\u003c/b\u003e. \u003cem\u003eAnnu Rev Genet\u003c/em\u003e 2008, \u003cb\u003e42\u003c/b\u003e(1):287\u0026ndash;299.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRaman G, Park S, Lee EM, Park SJ: \u003cb\u003eEvidence of mitochondrial DNA in the chloroplast genome of Convallaria keiskei and its subsequent evolution in the Asparagales\u003c/b\u003e. \u003cem\u003eSci Rep\u003c/em\u003e 2019, \u003cb\u003e9\u003c/b\u003e(1):5028.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePurabi M, Rofinayasmin BO, Katharina M, Ramakrishnan N, Jennifer AH: \u003cb\u003eCodon usage and codon pair patterns in non-grass monocot genomes\u003c/b\u003e. \u003cem\u003eAnn Bot\u003c/em\u003e 2017(6):1\u0026ndash;17.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMorton BR, Wright SI: \u003cb\u003eSelective Constraints on Codon Usage of Nuclear Genes from Arabidopsis thaliana\u003c/b\u003e. \u003cem\u003eMol Biol Evol\u003c/em\u003e 2006, \u003cb\u003e24\u003c/b\u003e(1):122\u0026ndash;129.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang W, Yu H, Wang J, Lei W, Gao J, Qiu X, Wang J: \u003cb\u003eThe Complete Chloroplast Genome Sequences of the Medicinal Plant Forsythia suspensa (Oleaceae)\u003c/b\u003e. \u003cem\u003eInt J Mol Sci\u003c/em\u003e 2017, \u003cb\u003e18\u003c/b\u003e(11):2288.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSmith HC, Gott JM, Hanson MR: \u003cb\u003eA guide to RNA editing\u003c/b\u003e. \u003cem\u003eRNA-Publ RNA Soc\u003c/em\u003e 1997, \u003cb\u003e3\u003c/b\u003e(10):1105\u0026ndash;1123.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHoch B: \u003cb\u003eEditing of a chloroplast mRNA by creation of an initiation codon\u003c/b\u003e. \u003cem\u003eNature\u003c/em\u003e 1991, \u003cb\u003e353\u003c/b\u003e(6340):178\u0026ndash;180.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMaier RM, Zeltz P, Kossel H, Bonnard G, Gualberto JM, Grienenberger JM: \u003cb\u003eRNA editing in plant mitochondria and chloroplasts\u003c/b\u003e. \u003cem\u003ePlant MolBiol\u003c/em\u003e 1996, \u003cb\u003e32\u003c/b\u003e(1):343\u0026ndash;365.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSchmitzlinneweber C, Barkan A: \u003cb\u003eRNA splicing and RNA editing in chloroplasts\u003c/b\u003e. \u003cem\u003eTopics in Current Genetics\u003c/em\u003e 2007, \u003cb\u003e19\u003c/b\u003e:213\u0026ndash;248.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShikanai T: \u003cb\u003eRNA editing in plant organelles: machinery, physiological function and evolution\u003c/b\u003e. \u003cem\u003eCellular Molecular Life Sciences\u003c/em\u003e 2006, \u003cb\u003e63\u003c/b\u003e(6):698\u0026ndash;708.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDong W, Xu C, Li C, Sun J, Zuo Y, Shi S, Cheng T, Guo J, Zhou S: \u003cb\u003eycf1, the most promising plastid DNA barcode of land plants\u003c/b\u003e. \u003cem\u003eSci Rep\u003c/em\u003e 2015, \u003cb\u003e5\u003c/b\u003e:8348.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNeubig KM, Whitten WM, Carlsward BS, Blanco MA, Endara L, Williams NH, Moore M: \u003cb\u003ePhylogenetic utility of ycf1 in orchids: a plastid gene more variable than matK\u003c/b\u003e. \u003cem\u003ePlant Syst Evol\u003c/em\u003e 2009, \u003cb\u003e277\u003c/b\u003e(1\u0026ndash;2):75\u0026ndash;84.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDugas DV, Hernandez D, Koenen EJM, Schwarz E, Straub S, Hughes CE, Jansen RK, Nageswara-Rao M, Staats M, Trujillo JT: \u003cb\u003eMimosoid legume plastome evolution: IR expansion, tandem repeat expansions, and accelerated rate of evolution in clpP\u003c/b\u003e. \u003cem\u003eSci Rep\u003c/em\u003e 2015, \u003cb\u003e5\u003c/b\u003e:16958.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYu X, Tan W, Zhang H, Gao H, Tian X: \u003cb\u003eComplete Chloroplast Genomes of Ampelopsis humulifolia and Ampelopsis japonica\u003c/b\u003e: \u003cb\u003eMolecular Structure\u003c/b\u003e, \u003cb\u003eComparative Analysis\u003c/b\u003e, \u003cb\u003eand Phylogenetic Analysis\u003c/b\u003e. \u003cem\u003ePlants-Basel\u003c/em\u003e 2019, \u003cb\u003e8\u003c/b\u003e(10):410.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang Z, Bielawski JP: \u003cb\u003eStatistical methods for detecting molecular adaptation\u003c/b\u003e. \u003cem\u003eTrends Ecol Evol\u003c/em\u003e 2000, \u003cb\u003e15\u003c/b\u003e(12):496\u0026ndash;503.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSwanson WJ, Wong A, Wolfner MF, Aquadro CF: \u003cb\u003eEvolutionary Expressed Sequence Tag Analysis of Drosophila Female Reproductive Tracts Identifies Genes Subjected to Positive Selection\u003c/b\u003e. \u003cem\u003eGenetics\u003c/em\u003e 2004, \u003cb\u003e168\u003c/b\u003e(3):1457\u0026ndash;1465.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGitzendanner MA, Soltis PS, Wong GK, Ruhfel BR, Soltis DE: \u003cb\u003ePlastid phylogenomic analysis of green plants: A billion years of evolutionary history\u003c/b\u003e. \u003cem\u003eAmerican Journal of Botany\u003c/em\u003e 2018, \u003cb\u003e105\u003c/b\u003e(3):291\u0026ndash;301.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDu YP, Bi Y, Yang FP, Zhang MF, Zhang XH: \u003cb\u003eComplete chloroplast genome sequences of Lilium: Insights into evolutionary dynamics and phylogenetic analyses\u003c/b\u003e. \u003cem\u003eSci Rep\u003c/em\u003e 2017, \u003cb\u003e7\u003c/b\u003e(1).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePorebski S, Bailey LG, Baum BR: \u003cb\u003eModification of a CTAB DNA extraction protocol for plants containing high polysaccharide and polyphenol components\u003c/b\u003e. \u003cem\u003ePlant Mol Biol Rep\u003c/em\u003e 1997, \u003cb\u003e15\u003c/b\u003e(1):8\u0026ndash;15.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAndrews S: \u003cb\u003eFastQC A Quality Control tool for High Throughput Sequence Data\u003c/b\u003e. \u003cem\u003eBabraham Institute\u003c/em\u003e 2015.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNicolas D, Patrick M, Guillaume S: \u003cb\u003eNOVOPlasty: de novo assembly of organelle genomes from whole genome data\u003c/b\u003e. \u003cem\u003eNucleic Acids Res\u003c/em\u003e 2017(4):4.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang K, Li L, Zhao M, Li S, Sun H, Lv Y, Wang Y: \u003cb\u003eCharacterization of the complete chloroplast genome of longan (Dimocarpus longan Lour.) using illumina paired-end sequencing\u003c/b\u003e. \u003cem\u003eMitochondrial DNA Part B\u003c/em\u003e 2017, \u003cb\u003e2\u003c/b\u003e(2):904\u0026ndash;906.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBankevich A, Nurk S, Antipov D, Gurevich AA, Dvorkin M, Kulikov AS, Lesin VM, Nikolenko SI, Pham S, Prjibelski AD: \u003cb\u003eSPAdes: a new genome assembly algorithm and its applications to single-cell sequencing\u003c/b\u003e. \u003cem\u003eJ Comput Biol\u003c/em\u003e 2012, \u003cb\u003e19\u003c/b\u003e(5):455\u0026ndash;477.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWalker BJ, Abeel T, Shea T, Priest M, Abouelliel A, Sakthikumar S, Cuomo CA, Zeng Q, Wortman JR, Young S: \u003cb\u003ePilon: an integrated tool for comprehensive microbial variant detection and genome assembly improvement\u003c/b\u003e. \u003cem\u003ePLoS One\u003c/em\u003e 2014, \u003cb\u003e9\u003c/b\u003e(11):e112963.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMichael T, Pascal L, Tommaso P, Ulbricht-Jones ES, Axel F, Ralph B, Stephan G: \u003cb\u003eGeSeq \u0026ndash; versatile and accurate annotation of organelle genomes\u003c/b\u003e. \u003cem\u003eNucleic Acids Res\u003c/em\u003e 2017(W1):W1.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShi L, Chen H, Jiang M, Wang L, Wu X, Huang L, Liu C: \u003cb\u003eCPGAVAS2, an integrated plastome sequence annotator and analyzer\u003c/b\u003e. \u003cem\u003eNucleic Acids Res\u003c/em\u003e 2019, \u003cb\u003e47\u003c/b\u003e(W1):W65-W73.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChan PP, Lowe TMJMoMB: \u003cb\u003etRNAscan-SE: Searching for tRNA Genes in Genomic Sequences\u003c/b\u003e. \u003cem\u003eMethods Mol Biol\u003c/em\u003e 2019, \u003cb\u003e1962\u003c/b\u003e:1\u0026ndash;14.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLohse M, Drechsel O, Bock R: \u003cb\u003eOrganellarGenomeDRAW (OGDRAW): a tool for the easy generation of high-quality custom graphical maps of plastid and mitochondrial genomes\u003c/b\u003e. \u003cem\u003eCurr Genet\u003c/em\u003e 2007, \u003cb\u003e52\u003c/b\u003e(5):267\u0026ndash;274.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBeier S, Thiel T, Munch T, Scholz U, Mascher M: \u003cb\u003eMISA-web: a web server for microsatellite prediction\u003c/b\u003e. \u003cem\u003eBioinformatics\u003c/em\u003e 2017, \u003cb\u003e33\u003c/b\u003e(16):2583\u0026ndash;2585.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi Q, Wan JM: \u003cb\u003e[SSRHunter: development of a local searching software for SSR sites]\u003c/b\u003e. \u003cem\u003eYi Chuan\u003c/em\u003e 2005, \u003cb\u003e27\u003c/b\u003e(5):808\u0026ndash;810.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAmiryousefi A, Hyvonen J, Poczai P: \u003cb\u003eIRscope: an online program to visualize the junction sites of chloroplast genomes\u003c/b\u003e. \u003cem\u003eBioinformatics\u003c/em\u003e 2018, \u003cb\u003e34\u003c/b\u003e(17):3030\u0026ndash;3031.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAfgan E, Baker D, Den Beek MV, Blankenberg D, Bouvier D, Cech M, Chilton J, Clements D, Coraor N, Eberhard C: \u003cb\u003eThe Galaxy platform for accessible, reproducible and collaborative biomedical analyses: 2016 update\u003c/b\u003e. \u003cem\u003eNucleic Acids Res\u003c/em\u003e 2016, \u003cb\u003e44\u003c/b\u003e(W1):W3-W10.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWright F: \u003cb\u003eThe effective number of codons used in a gene\u003c/b\u003e. \u003cem\u003eGene\u003c/em\u003e 1990, \u003cb\u003e87\u003c/b\u003e(1):23\u0026ndash;29.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMower JP: \u003cb\u003eThe PREP suite: predictive RNA editors for plant mitochondrial genes, chloroplast genes and user-defined alignments\u003c/b\u003e. \u003cem\u003eNucleic Acids Res\u003c/em\u003e 2009, \u003cb\u003e37\u003c/b\u003e:253\u0026ndash;259.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang Y, Yuan X, Zhang J: \u003cb\u003eThe complete chloroplast genome sequence of Pometia tomentosa\u003c/b\u003e. \u003cem\u003eMitochondrial DNA Part B\u003c/em\u003e 2019, \u003cb\u003e4\u003c/b\u003e(2):3950\u0026ndash;3951.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang B, Li M, Ma J, Fu Z, Tian J: \u003cb\u003eThe complete chloroplast genome sequence of Sapindus mukorossi\u003c/b\u003e. \u003cem\u003eMitochondrial DNA Part A\u003c/em\u003e 2016, \u003cb\u003e27\u003c/b\u003e(3):1825\u0026ndash;1826.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFrazer KA, Pachter L, Poliakov A, Rubin EM, Dubchak IJNAR: \u003cb\u003eVISTA: computational tools for comparative genomics\u003c/b\u003e. \u003cem\u003eNucleic Acids Res\u003c/em\u003e 2004, \u003cb\u003e32\u003c/b\u003e(Web Server issue):W273-279.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBrudno M, Malde S, Poliakov A, Do CB, Couronne O, Dubchak I, Batzoglou S: \u003cb\u003eGlocal alignment: finding rearrangements during alignment\u003c/b\u003e. \u003cem\u003eBioinformatics\u003c/em\u003e 2003, \u003cb\u003e19\u003c/b\u003e:54\u0026ndash;62.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSaina JK, Gichira AW, Li ZZ, Hu GW, Wang QF, Liao K: \u003cb\u003eThe complete chloroplast genome sequence of Dodonaea viscosa: comparative and phylogenetic analyses\u003c/b\u003e. \u003cem\u003eGenetica\u003c/em\u003e 2017, \u003cb\u003e146\u003c/b\u003e(1):101\u0026ndash;113.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDu X, Xin G, Ren X, Liu H, Hao N, Jia G, Liu W: \u003cb\u003eThe complete chloroplast genome of Eurycorymbus cavaleriei (Sapindaceae), a Tertiary relic species endemic to China\u003c/b\u003e. \u003cem\u003eConserv Genet Resour\u003c/em\u003e 2018.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKim SC, Baek SH, Hong KN, Lee JW: \u003cb\u003eCharacterization of the complete chloroplast genome of Koelreuteria paniculata (Sapindaceae)\u003c/b\u003e. \u003cem\u003eConserv Genet Resour\u003c/em\u003e 2018, \u003cb\u003e10\u003c/b\u003e(4):69\u0026ndash;72.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen SY, Zhang XZ: \u003cb\u003eCharacterization of the complete chloroplast genome of Xanthoceras sorbifolium, an endangered oil tree\u003c/b\u003e. \u003cem\u003eConserv Genet Resour\u003c/em\u003e 2017, \u003cb\u003e9\u003c/b\u003e(4):595\u0026ndash;598.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang Z, Xiao J, Wu J, Zhang H, Liu G, Wang X, Dai L: \u003cb\u003eParaAT: a parallel tool for constructing multiple protein-coding DNA alignments\u003c/b\u003e. \u003cem\u003eBiochem Biophys Res Commun\u003c/em\u003e 2012, \u003cb\u003e419\u003c/b\u003e(4):779\u0026ndash;781.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang D, Zhang Y, Zhang Z, Zhu J, Yu J: \u003cb\u003eKaKs_Calculator 2.0: A Toolkit Incorporating Gamma-Series Methods and Sliding Window Strategies\u003c/b\u003e. \u003cem\u003eGenomics,Proteomics \u0026amp; Bioinformatics\u003c/em\u003e 2010, \u003cb\u003e8\u003c/b\u003e(1):77\u0026ndash;80.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNei M, Gojobori T: \u003cb\u003eSimple methods for estimating the numbers of synonymous and nonsynonymous nucleotide substitutions\u003c/b\u003e. \u003cem\u003eMol Biol Evol\u003c/em\u003e 1986, \u003cb\u003e3\u003c/b\u003e(5):418\u0026ndash;426.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKatoh K, Rozewicki J, Yamada KD: \u003cb\u003eMAFFT online service: multiple sequence alignment, interactive sequence choice and visualization\u003c/b\u003e. \u003cem\u003eBrief Bioinform\u003c/em\u003e 2017.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDarriba D, Taboada GL, Doallo R, Posada D: \u003cb\u003ejModelTest 2: more models, new heuristics and parallel computing\u003c/b\u003e. \u003cem\u003eNat Methods\u003c/em\u003e 2012, \u003cb\u003e9\u003c/b\u003e(8):772\u0026ndash;772.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCummings MP: \u003cb\u003ePAUP* (Phylogenetic Analysis Using Parsimony (and Other Methods))\u003c/b\u003e. \u003cem\u003eDictionary of Bioinformatics Computational Biology\u003c/em\u003e 2004.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAzim MK, Khan IA, Zhang Y: \u003cb\u003eCharacterization of mango (Mangifera indicaL.) transcriptome and chloroplast genome\u003c/b\u003e. \u003cem\u003ePlant MolBiol\u003c/em\u003e 2014, \u003cb\u003e85\u003c/b\u003e(1\u0026ndash;2):193\u0026ndash;208.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Rambutan, Nephelium lappaceum, Chloroplast genome, Sapindaceae, RNA editing, Phylogeny","lastPublishedDoi":"10.21203/rs.3.rs-128918/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-128918/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eBackground:\u003c/strong\u003e Rambutan (\u003cem\u003eNephelium lappaceum\u003c/em\u003e L.) is an important fruit tree belongs to the family \u003cem\u003eSapindaceae\u003c/em\u003e and widely cultivated in Southeast Asia. The chloroplast of plants, as a photosynthetic organelle plays an important role in the photosynthesis and secondary metabolic activities. The chloroplast genome sequencing has become an integral part in understanding the genomic machinery and the phylogenetic histories of rambutan organelles.\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eResults:\u003c/strong\u003e We sequenced its chloroplast genome and assembled 161,321 bp circular DNA. It is characterized by a typical quadripartite structure composed of a large (86,068 bp) and small (18,153 bp) single-copy region interspersed by two identical inverted repeats (IRs) (28,550 bp). We identified 132 genes including 78 protein-coding, 29 tRNA and 4 rRNA genes, with 21 genes duplicated in the IRs. Sixty-three simple sequence repeats (SSRs) and 98 repetitive sequences were detected. Twenty-nine codons showed biased usage and 49 potential RNA editing sites were predicted across 18 protein-coding genes in the\u003cem\u003e \u003c/em\u003erambutan chloroplast genome. In addition, coding gene sequence divergence analysis of \u003cem\u003eN. lappaceum\u003c/em\u003e suggested that \u003cem\u003eccsA, clpP, rpoA, rps12, psbJ \u003c/em\u003eand\u003cem\u003e rps19\u003c/em\u003e were under positive selection, which might reflect specific adaptations of \u003cem\u003eN. lappaceum\u003c/em\u003e to its particular living environment. Comparative chloroplast genome analyses from five species in \u003cem\u003eSapindaceae\u003c/em\u003e revealed that a higher similarity was conserved in the IR regions than in the LSC and SSC regions. The phylogenetic analysis showed that \u003cem\u003eN. lappaceum\u003c/em\u003e chloroplast genome has the closest relationship with that of \u003cem\u003ePometia tomentosa. \u003c/em\u003e\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eConclusions:\u003c/strong\u003e The understanding of the chloroplast genomics of rambutan and comparative analysis of \u003cem\u003eSapindaceae\u003c/em\u003e species would provide insight into future research on the breeding of rambutan and \u003cem\u003eSapindaceae\u003c/em\u003e evolutionary studies.\u003c/p\u003e","manuscriptTitle":"Chloroplast Genome of Rambutan and Comparative Analyses in Sapindaceae","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2020-12-28 17:41:40","doi":"10.21203/rs.3.rs-128918/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"fdeb8c39-2521-4456-aabd-390c0b274e25","owner":[],"postedDate":"December 28th, 2020","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":1616781,"name":"Epigenetics \u0026 Genomics"}],"tags":[],"updatedAt":"2024-07-07T22:32:08+00:00","versionOfRecord":{"articleIdentity":"rs-128918","link":"https://doi.org/10.3390/plants10020283","journal":{"identity":"plants","isVorOnly":true,"title":"Plants"},"publishedOn":"2021-02-02 22:32:08","publishedOnDateReadable":"February 2nd, 2021"},"versionCreatedAt":"2020-12-28 17:41:40","video":"","vorDoi":"10.3390/plants10020283","vorDoiUrl":"https://doi.org/10.3390/plants10020283","workflowStages":[]},"version":"v1","identity":"rs-128918","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-128918","identity":"rs-128918","version":["v1"]},"buildId":"_2-kVJe1T_tPrBINL-cwx","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.