Genomic Resources of Broomcorn Millet: Demonstration and Application of a High-throughput BAC Mapping Pipeline

preprint OA: closed
Full text JSON View at publisher

Abstract

Background: With high-efficient water-use and drought tolerance, broomcorn millet has emerged as a candidate for food security. To promote its research process for molecular breeding and functional research, a comprehensive genome resource is of great importance. Results: Herein, we constructed a BAC library for broomcorn millet, generated BAC end sequences based on the clone-array pooled shotgun sequencing strategy and Illumina sequencing technology, and integrated BAC clones into genome by a novel pipeline for BAC end profiling. The BAC library is consisted of 76,023 clones with an average insert length of 123.48 Kb, covering about 9.9-fold of the 850 Mb genome. Of 9,216 clones tested using our pipeline, 8,262 clones were mapped on the broomcorn millet cultivar longmi4 genome. These mapped clones covered 308 of the 829 gaps left by the genome. To our knowledge, this is the only BAC resource for broomcorn millet. Conclusions: We constructed a high-quality BAC libraray for broomcorn millet and designed a novel pipeline for BAC end profiling. BAC clones can be browsed and obtained from our website (http://eightstarsbio.com/gresource/JBrowse-1.16.5/index.html). The high-quality BAC clones mapped on genome in this study will provide a powerful genomic resource for genome gap filling, complex segment sequencing, FISH, functional research, and genetic engineering of broomcorn millet.
Full text 95,962 characters · extracted from preprint-html · click to expand
Genomic Resources of Broomcorn Millet: Demonstration and Application of a High-throughput BAC Mapping Pipeline | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Genomic Resources of Broomcorn Millet: Demonstration and Application of a High-throughput BAC Mapping Pipeline Wei Xu, Mengjie Liang, Xue Yang, Hao Wang, Meizhong Luo This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-536711/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background : With high-efficient water-use and drought tolerance, broomcorn millet has emerged as a candidate for food security. To promote its research process for molecular breeding and functional research, a comprehensive genome resource is of great importance. Results : Herein, we constructed a BAC library for broomcorn millet, generated BAC end sequences based on the clone-array pooled shotgun sequencing strategy and Illumina sequencing technology, and integrated BAC clones into genome by a novel pipeline for BAC end profiling. The BAC library is consisted of 76,023 clones with an average insert length of 123.48 Kb, covering about 9.9-fold of the 850 Mb genome. Of 9,216 clones tested using our pipeline, 8,262 clones were mapped on the broomcorn millet cultivar longmi4 genome. These mapped clones covered 308 of the 829 gaps left by the genome. To our knowledge, this is the only BAC resource for broomcorn millet. Conclusions : We constructed a high-quality BAC libraray for broomcorn millet and designed a novel pipeline for BAC end profiling. BAC clones can be browsed and obtained from our website ( http://eightstarsbio.com/gresource/JBrowse-1.16.5/index.html ). The high-quality BAC clones mapped on genome in this study will provide a powerful genomic resource for genome gap filling, complex segment sequencing, FISH, functional research, and genetic engineering of broomcorn millet. Epigenetics & Genomics Broomcorn millet BAC BES Genomics resources Gap filling CAPSS Jbrowse Figures Figure 1 Figure 2 Figure 3 Figure 4 1 Introduction With the increased global water scarcity caused by climate change and population growth, it is of great importance to exploit the high-efficient water-use crop for the food security of the human in the future. Broomcorn millet ( Panicum miliaceum ), also known as proso millet, panic millet, and wild millet, is one of the traditional five-grain crops in the north of China [ 1 ]. It is a typical C4 plant with high photosynthetic efficiency. It also has a high-efficient utilization ratio of water resource and a capacity for drought resistance, exquisitely adapting to semi-drought or drought conditions [ 2 ]. Furthermore, its growing cycle, 60–90 days from sowing to maturity, is shorter than other cereals [ 3 ]. Broomcorn millet contains more protein than most grains, and a relatively balance array of trace elements and vitamins. More than 8700 accessions of broomcorn millet including landraces and cultivars have been conserved in the National Gene Bank of the Institute of Crop Science, Chinese Academy of Agricultural Sciences, thus providing an abundance of resources for genetic improvement of broomcorn millet. At present, high-quality chromosome-scale genome assemblies of two allotetraploid (2n = 4x = 36) broomcorn millet varieties decoded by Chinese researchers are available [ 4 , 5 ]. These genome assemblies provide the foundation for the molecular breeding of broomcorn millet. However, other gemonic resources are still required to complete and make the full use of the genome assemblies. The tranditional bacterial artificial chromosome (BAC) libraries with genomic DNA inserts of 50 kbp – 300 kbp [ 6 – 8 ] are also important resources for genomic research. They provide natural DNA materials for a variety of experiments, such as intact gene cluster cloning, map-based cloning, whole genome sequencing, comparative genomics analysis and fluorescence in situ hybridization that aim to understand functional elements in the genome [ 9 , 10 ]. The utility of the BAC libraries can be greatly enhanced by mapping the BAC clones on the genome assemblies. Hence, we constructed a high-quality BAC library for broomcorn millet, developed a new pipeline for cost-effectively decoding BAC end seqeunces generated by clone-array pooled shotgun sequencing strategy (abbr. CAPSS) and Illumina sequencing technology, and mapped the BAC clones on the genome assemblies of the broomcorn millet [ 11 ]. 2 Results 2.1 BAC library construction A BAC library of broomcorn millet was constructed with the restriction enzyme Hin dIII using high-molecular-weight genomic DNA prepared from etiolated seedlings. In total, the library consists of 76,032 BAC clones, that were arrayed into 198 independent 384-well microtiter plates. Insert sizing of randomly picked clones showed that the majority of genomic BAC inserts fell into the length range of 97–145.5 kb with an average insert size of 132 kb (Fig. 1 ). 2.2 Construction and sequencing of DNA pools In order to obtain BAC end sequences with high efficiency and low cost, we designed and performed a pipeline based on the clone array pooled shotgun sequencing strategy (Fig. 2 ). Randomly selected 24 384-plates (each plate consists of 16 x 24 clones) from the broomcorn millet BAC library were arranged to a square superpool with 6 row plates and 4 column plates. Therefore, the superpool consisted of 96 row pools (6 plates x 16) and 96 column pools (4 plates x 24). These row pools and column pools were called as the secondary pools, and each secondary pool also consisted of 96 BAC clones. In total, one superpool consisted of 192 secondary pools and 9,216 individual BAC clones. DNA of the secondary pools were extracted and Illumina sequencing libraries with individual index sequences for each secondary pool were prepared. Finally, 96 row pool libraries were mixed in equal amounts as the library X, and 96 column pool libraries were also mixed in equal amounts as the library Y. The average insert sizes of the library X and Y analyzed by the Agilent 2100 Bioanalyzer (Additional file 1: Fig. S1) were 491 bp and 443 bp, respectively. Sequencing was accomplished using the Illumina HiSeq 2000 sequencing platform with paired-end protocol (PE150). The Library X and Y generated 130.82 Gb and 133.50 Gb raw data, respectively. Trimmomatic was employed for trimming adaptors and filtering low-quality or shorter reads. FastQC was employed for evaluating the quality of the preprocessing reads. Then, the valid reads were obtained by filtering the reads traced to E. coli DH10B reference genome using Bowtie2. Finally, we obtained valid yields of 60.40 Gb and 102.37 Gb for libraries X and Y, respectively (Table 1 ). Demultiplexer was employed for demultiplexing and generating each secondary pool reads according to the 7-bp index (Additional file 2). The average valid sequencing depths of row pools and column pools were 45x and 78x, respectively (Additional file 1: Fig. S2). Table 1 The summary of NGS data Library Total size (Gb) Size after QC (Gb) Valid size (Gb) X 130.82 103.78 60.40 Y 133.50 107.92 102.37 2.3 Parsing BAC end sequences We focused our interest on the two special contigs in a BAC: the forward and reverse BAC end seqeunces. We designed two pathways to find them. On one hand, the paired-end reads overlapping with the border sequences of the vector harboring the Hin dIII restriction enzyme site were extracted from the valid data of each secondary pool, and then assembled using Cap3 based on the overlap-layout-consensus method. The vector sequence parts in the consensus were truncated to obtain short BAC end sequences (abbr. BES) starting with AAGCTT ( Hin dIII site). These short BAC end sequences were assigned to the corresponding wells according to the clone array pooled shotgun sequencing strategy. Of the 9,216 wells in 384-plates tested in the pipeline, 8,183 (88.79%) wells were assigned one forward BES and 7,897 (85.69 %) wells were assigned one reverse BES (Additional file 3: Table S1). All assigned BAC end sequences have an average length of 358 bp (Additional file 1: Fig. S3), which is consistent with the insert size of the Illumina libraries (Additional file 1: Fig. S1). On the other hand, the valid NGS data of each secondary pool were firstly assembled using SPAdes, and the contig N50 sizes of all row pools and column pools were counted (Additional file 1: Fig. S4). The average N50 of row pools and column pools were 13.97 kb and 11.45 kb, respectively. Likewise, the vector sequence parts at the ends of the contigs were removed to retain the long BESs starting with AAGCTT. Finally, these long BESs were assigned to the corresponding wells in 384-plates. Of the 9,216 wells, 5,454 (59.18%) wells obtained one forward BES and 5,108 (55.43%) wells obtained one reverse BES (Additional file 3: Table S1). The N50 of the long BESs is 60.87 kb (Additional file 1: Fig. S5). 2.4 Determination of BAC locations on the genome In order to determine the locations of the broomcorn millet BAC clones on the genome, we aligned the short and long BAC end sequences onto the cultivar longmi4 genome with Blastn. Table 2 listed the alignment results. With BAC end sequences with repeats, 14,862 (83.44%) short BAC end sequences including 12,971 single-hit and 1,891 multi-hit sequences were mapped to the genome, and 12,120 (95.28%) long BAC end sequences (all single-hit) were mapped to the genome. With the BAC end sequences without repeats, 7,760 (43.57%) short BAC end sequences including 7,295 single-hit and 465 multi-hit sequences were mapped to the genome, and 10,626 (83.53%) long BAC end sequences (all single-hit) were mapped to the genome. In order to map as many as possible BACs to the genome, the alignment results of the end sequences with repeats were adopted. We wrote a python script to extract the Blast results from short and long BESs. As a result, 5,795 BACs were mapped to the genome using short BAC end sequences, and 6,973 BACs were mapped to the genome using long BAC end sequences. Finally, the Blast results generated by short and long BAC end sequences were integrated, and in total 8,262 BACs (89.65%) were mapped to the genome. Table 2 A summary of the Broomcorn millet BESs and the anchoring results of the Broomcorn millet BAC clones to the longmi 4 genome using the BESs Categories Short BESs Long BESs Integrity BAC end sequences Clones in superpool 9216 9216 Clones with successful BESs 9126 7725 with paired BESs 7790 3890 with single-end BESs 1336 3835 only forward 921 2189 only reverse 415 1646 Total successful BESs 17811 (96.63%) 12721 (66.17%) Alignment Aligned BESs with repeats unmasked 14862 (83.44%) 12120 (95.28%) Single-hit BESs 12971 12120 Multi-hit BESs 1891 0 Aligned BESs with repeats masked 7760 (43.57%) 10626 (83.53%) Single-hit BESs 7295 10626 Multi-hit BESs 465 0 Anchoring to reference sequences Clones anchored to single sites 5795 6973 8262 Clones anchored with single BES 2507 3907 2871 Clones anchored with paired BESs 3288 3066 5391 To verify the accuracy of BAC locations on the genome, 55 BAC clones were randomly picked from 384-plates for BAC end sequencing using Sanger method. After quality control, the 55 paired Sanger end sequences were blasted with the above short and long BESs, and the genome. All BAC clones but one (32G16) were consistent. By checking the 32G16 BAC end sequences, we found that this clone (or rather 32G16 well) was assigned two long forward BESs, one short forward BES, and no reverse BES. Only one long forward BES and the short forward BES were perfectly identical to the Sanger end sequence. Another long forward BES confused the mapping. However, it can be solved by weighting. In summary, the accuracy of the BAC mapping approach in this study was extremely high. The distribution of the 5,391 BACs mapped to the genome by paired BAC end sequences were counted (Table 3 ). These clones covered a total of 432.47 Mb of chromosomes with a total coverage of 50.97%. Among the 18 chromosome sequences of the cultivar longmi4 genome, there left 829 gaps. Our BACs covered 308 of them. The insert sizes of these BACs presented the Gaussian distribution, with an average insert size of 123.48 kb (Fig. 3 ), which is lower than that predicted by pulse field electrophoresis. These BACs are valuable resources for further improvement of the genome. Table 3 The location result of BACs on broomcorn millet chromosomes Chr. Length (Mb) No. gaps No. BACs BACs Covered length (Mb) No. gaps chr 1 69.18 85 439 35.25 34 chr 2 61.15 59 379 31.04 19 chr 3 57.97 50 394 31.03 20 chr 4 56.29 34 359 27.90 14 chr 5 54.13 52 361 29.86 19 chr 6 52.84 46 358 28.24 16 chr 7 51.23 67 286 23.67 27 chr 8 48.26 29 339 26.49 8 chr 9 45.11 70 240 20.96 21 chr10 44.65 53 308 25.44 28 chr11 43.18 30 267 21.90 14 chr12 42.47 30 254 20.79 11 chr13 40.72 50 261 21.11 17 chr14 38.49 32 269 20.23 10 chr15 34.36 34 213 18.14 8 chr16 33.61 45 212 16.72 15 chr17 32.99 25 226 18.26 10 chr18 32.24 38 199 15.32 17 unplaced 9.48 0 27 - 0 Total 848.47 829 5391 432.47 308 2.5 Presentation of BAC locations by JBrowse In order to effectively use the BAC resource, quickly and easily retrieve BAC clones and view other annotaion information, we established a resource website employing a lightweight JBrowse to display the BAC resource information on the broomcorn millet genome (Fig. 4 ). 2.6 Organelle genomes of the broomcorn millet We aligned all BESs in this study to the chloroplast (CM009689) genome of broomcorn millet by local blast. A total of 65 BAC clones were mapped to the chloroplast with more than 99% of similarity, of which 16 BAC clones were determined by paired BESs (25C07, 25N09, 27F16, 27O19, 29B02, 29K11, 32J09, 32N16, 33D23, 33K24, 36A05, 39C22, 42C10, 46C16, 48B15, 48C03). In the absence of mitochrondrial sequence in broomcorn millet, we aligned BESs to all mitochondrial genomes in Gramineae . Of 16 BAC clones with homologous BES sequences, 13 BACs were simultaneously aligned to the chloroplast and 3 BACs were simultaneously aligned to the nuclear genome. 3 Discussion BAC library is still a powerful resource for genome assembly, functional genomics research, and long-term genetic resoure storage for the endangered species. BAC seqeunces, espcially paired BESs, are usally used to detect assembly errors or assist assembly [ 12 ]. The most conventional and convenient approach to decode BESs is Sanger sequencing method. However, its one-by-one style is laborious, time-consuming and expensive. Fortunately, a few of efficient approaches had developped, based on next generation sequencing technologies, such as pBACcode and BAC-anchor[ 13 , 14 ]. pBACode determines paired BESs by a pair of random barcodes flanking the cloning site. BAC-anchor determines general paired BESs by using specific restriction enzyme sites and searching for utlra-long paired-end subreads containing large internal gaps. We previously also developed a high-throughput approach for long accurate BES profiling with PacBio sequencing technology [ 15 ]. However, these approaches focused on paired BES profiles, and they are hard to trace BESs of specific clones in 384-plate wells. In this study, we generated BESs and assigned them to physical wells in 384-plates by applying the characteristics of row-cross-column in two-dimension arrays and cost-effective Illumina platform. These BAC end sequences in average are longer and accurater than that generated by Sanger method, for assembling by SPAdes assembler[ 16 ]. After alignment, 89.65% of clones were successfully associated with the broomcorn millet longmi 4 genome. These clones covered 308 of the 829 gaps left by the genome and can be uesd to close the genome gaps. In conventional whole genome sequencing project, sequencing coverage is usallly at least 30x, and not more than 200x. Higher coverage will result in more sequences that are generated by PCR mutations or sequencing errors and lead to contigs with shorter N50. We added index into each secondary pool and mixed 96 secondary pools as a sequencing library to run in a lane of Illumina flow cell for greatly reducing the cost of sequencing. As a result, after removing the contaminated E. coli genomic DNA reads the average valid sequencing depths of row pools and column pools were 45x and 78x, respectively. In the process of BESs extraction, we designed two pathways: short BES pathway and long BES pathway. Short BES pathway searched all reads overhanging with verctor end sequences before assembling by Cap3; consequently, it generated all potenital BESs. Long BES pathway identified all contigs overhanging with vector end seqeunces after assembling by SPAdes; consequently, it generated longer but less BESs than the first pathway. The assignment of BAC end sequences at the intersection sites is affected by many factors, such as the overlapping rate of the BAC clones and the correct rate of the sequences. The overlapping rate of BAC clones in the superpool is the most important factor. If overlapping BAC clones appear in the same row or column pool, we can assign them to wells easily and correctly. Also If two overlapping BAC clones appear in a different row or column pool, we can rectify them by our previous method [ 17 ]. However, if more than two overlapping BAC clones appear in a different row or column pool, our method will filter out potential BESs, so that the BESs of the intersection well will be absent. In the process of BES assignment, the flow of the forward and reverse BESs are completely independent. When more than one forward and/or reverse BESs were assigned to a well, we cannot determine which BESs are a pair of BESs. If a high-qualty genome is available, it is easy to assess which pair of BESs are derived from the same BAC by mapping. However, if the variety used for BAC library construction is not the same as that the reference genome stands for, the alignment results that do not satisfy the location requirement of BES pairs will be discarded. The clones that may contain a large structural variation can be picked out for further analysis from the 384-plate wells. Plant cells contain an abundance of chloroplasts and mitochondria. The chloroplast genome is generally around 150 kb, while the mitochondrial genome size varies widely, typically between 200 kb and 750 kb [ 18 ]. Although nuclei are extracted for BAC library construction, a trace of organelle DNA cotamination is inevitable. In this study, a small number of clones of chloroplast genome was found by BESs, while clones of mitochondrial genome were almost absent. By high-throughput sequencing and mapping clones in the secondary pool of BAC libraries, we can find the coordinates of genome sequences in the broomcorn millet BAC libraries. Therefore, if we find genes that play an important role in biology, we can quickly locate the BAC clones containing this gene, and obtain the experimental materials for further research and analysis. At the same time, because BAC contains a long DNA fragment (about 120 kb), it is also convenient for us to quickly analyze the upstream and/or downstream DNA elements of interested genes or adjacent genes. 4 Conclusions We constructed a high-quality BAC library for broomcorn millet, developed a high-efficient and low-cost pipeline which can generate and parse ten thousand of paired BAC end sequences at each time, and mapped a total of 8,262 broomcorn millet BACs to the chromosomes. These clones covered 308 of the 829 gaps left by the genome. The high-quality BAC clones mapped on genome in this study will provide a powerful genomic resource for gap filling, complex segment sequencing, FISH, functional research, and genetic engineering of broomcorn millet. 5 Methods 5.1 Plant materials, growth conditions and BAC library construction The seeds of broomcorn millet were provided by professor Mingsheng Chen of the Institute of Genetics and Developmental Biology, Chinese Academy of Sciences and grown at dark conditions under 25°C. The young leaves of seedlings were harvested and mixed, frozen immediately in liquid nitrogen for the extraction of nuclear DNA. High molecular weight nuclear DNA was extracted and BAC library was constructed following our previous protocol [ 6 ]. Partial digestions of DNA plugs with dilution Hin dIII were performed. DNA fragments ranging from 100 kb to 200 kb were recovered from pulse field gel and ligated with pIndigoBAC536-S vector [ 7 ]. The ligation product was used to transform DH10B cells by electroporation. White colonies were picked up and stored in 384-plates at -80 ºC. 5.2 Pool construction Twenty-four 384-plates were chosen form the BAC library of broomcorn millet. These plates were arranged in a 2-dimension superpool. The superpool contains 96 row pools and 96 column pools. A total of 192 pools were processed independently for high quality DNA extraction with the AxyGen AxyPrep Easy-96 plasmid kit. Then, these BAC DNAs were completely digested with ATP-Dependent DNase (Epicentre) to remove the host E. coli DNA. After digestion, these BAC DNAs were sheared in the Bioruptor to an average of 500 bp. During blunt-end repair, overhanging 5’ and 3’ ends were filled in or removed by T4 DNA polymerase. 5’-phosphates were attached using T4 polynucleotide kinase. Tail-A were added using Taq DNA polymerase. Then, adapters were ligated to both ends of the molecules using T4 DNA ligase. The ligation products were cleaned using MagBead DNA Purification Kit (Sangon Biotech, shanghai, CN). Sequencing adaptors were added using PCR. Finally, the products of KOD PCR were cleaned again. 5.3 Illumina sequencing and BAC end sequence analysis NGS sequencing of 2 mixed DNA libraries was performed via the Illumina HiSeq 2000 with 150-bp paired-end protocol (Genewiz, suzhou, CN). Trimmotatic was employed to filter and trim raw reads. FastQC was employed to assess data quality. After adapter filtering and quality assessment, BBMAP/demuxbyname script was employed for deconvolution depending on unique index sequence in each pool. A python script was employed to extract target paired-end reads that cover BES-VES site. Cap3 was employed to assemble those reads to consensuses, and consensuses were then trimmed to remove the part of vector to generate BESs called “short BES”. SPAdes was employed to directly assemble pool reads to contigs. Then, the contigs containing BAC vector sequences were extracted and trimmed using python script to generate the trimmed contigs called “long BES”. Blastn was used to align BESs from raw and column pools, and then the shared BESs were assigned to the wells at the intersection. 5.4 Validation of BAC end sequences The analysis results of the BAC end sequences were validated by Sanger sequencing. Fifty-five BAC clones were randomly selected and their DNAs were extracted using an improved alkaline lysis protocol. Sanger sequencing was accomplished using BAC-F (5’-AACGACGGCCAGTGAATTG-3’) and BAC-R (5’-GATAACAATTTCACACAGG-3’) primers from pIndigoBAC536-S vector backbone. The BAC end sequences from Sanger sequencing were aligned to BES from the analysis results of illumina data using Blastn. 5.5 BAC Mapping on broomcorn millet longmi4 genome The genome sequence of broomcorn millet longmi4 was downloaded from NCBI genome database under GCA_002895445 accession, which was submitted by researchers from China Agricultural University. The local blastn was employed to map all BESs to the genome with the following options: qcov_hsp_perc = 99, perc_identity = 99, outfmt = 6, culling_limit = 1. The results were further converted to a GFF3 format file using a Python script. In this script, following conditions were set: if both forward and reverse BESs in each clone were mapped to the same chromosome, and their orientations were opposite and their interval lengths were less than 250 kb, such clones were recorded in GFF3 format file with three lines; if either forward or reverse BES was uniquely mapped to chromosome, such clones also were recorded in GFF3 format file with two lines; other conditions would be discarded. The GFF3 file was sorted using GFF3sort and presented with JBrowse in our website ( http://eightstarsbio.com/gresource/JBrowse-1.16.5/index.html ). 5.6 Organelle genome analysis The chloroplast genome (CM009689) of brromcorn millet and all mitochondrial genomes (NC_008331, NC_007982, NC_036024, NC_031164, NC_029816, NC_022714, NC_022666, NC_013816, NC_007886, NC_011033, NC_008362, NC_008360, NC_008332, NC_008333) in Gramineae were downloaded form NCBI nucleotide database. All potential BACs of organelle genome were identified by the local blastn with following options: qcov_hsp_perc = 99, perc_identity = 99, outfmt = 6. 6 Abbreviations BAC: bacterial artificial chromosome; BES: BAC end sequence; CAPSS: clone array pooled shotgun sequencing; CM: chloramphenicol; NGS: next generation sequencing; PCR: polymerase chain reaction; VES: vector end sequence. 7 Declarations 7.1 Ethics approval and consent to participate Not applicable 7.2 Consent for publication Not applicable 7.3 Availability of data and materials The source codes are openly available in a GitHub repository ( https://github.com/xuweixw/broomcorn-millet-BAC-library ). Illumina sequencing data are available at Sequence Read Archive (SRA) under the accession PRJNA576359. BAC clones can be browsed and obtained through our website ( http://eightstarsbio.com/gresource/JBrowse-1.16.5/index.html ). 7.4 Competing interests The authors declare that they have no competing interests. 7.5 Funding This work was supported by a grant from the National Natural Science Foundation of China (Grant no. 31671268). 7.6 Authors' contributions WX and MLuo conceived and designed the research framework; WX, MLiang, XY and HW performed the experiments; WX analyzed the data and wrote the manuscript; MLuo supervised the work and finalized this manuscript. All authors read and approved the manuscript. 7.7 Acknowledgements We are grateful to Dr. Mingsheng Chen (the Institute of Genetics and Developmental Biology, Chinese Academy of Sciences) for providing the seeds of broomcorn millet. References Kalinova J, Moudry J. Content and Quality of Protein in Proso Millet ( Panicum miliaceum L .) Varieties. Plant Foods Hum Nutr. 2006;61:45–9. Washburn JD, Schnable JC, Davidse G, Pires JC. Phylogeny and photosynthesis of the grass tribe Paniceae. Am J Bot. 2015;102:1493–505. Baltensperger DD. Progress with proso, pearl and other millets. 2002. Shi J, Ma X, Zhang J, Zhou Y, Liu M, Huang L, et al. Chromosome conformation capture resolved near complete genome assembly of broomcorn millet. Nat Commun. 2019;10:1–9. doi:10.1038/s41467-018-07876-6. Zou C, Li L, Miki D, Li D, Tang Q, Xiao L, et al. The genome of broomcorn millet. Nat Commun. 2019;10:1–12. doi:10.1038/s41467-019-08409-5. Luo M, Wing RA. An improved method for plant BAC library construction. In: Plant Functional Genomics. 2003. p. 3–19. Shi X, Zeng H, Xue Y, Luo M. A pair of new BAC and BIBAC vectors that facilitate BAC/BIBAC library construction and intact large genomic DNA insert exchange. Plant Methods. 2011;7:33. doi:10.1186/1746-4811-7-33. Shizuya H, Birren B, Kim UJ, Mancino V, Slepak T, Tachiiri Y, et al. Cloning and stable maintenance of 300-kilobase-pair fragments of human DNA in Escherichia coli using an F-factor-based vector. Proc Natl Acad Sci U S A. 1992;89:8794–7. Pan Y, Deng Y, Lin H, Kudrna DA, Wing RA, Li L, et al. Comparative BAC-based physical mapping of Oryza sativa ssp. indica var. 93-11 and evaluation of the two rice reference sequence assemblies. Plant J. 2014;77:795–805. Dong G, Shen J, Zhang Q, Wang J, Yu Q, Ming R, et al. Development and Applications of Chromosome-Specific Cytogenetic BAC-FISH Probes in S. spontaneum. Front Plant Sci. 2018;9:218. doi:10.3389/fpls.2018.00218. Cai WW, Chen R, Gibbs RA, Bradley A. A clone-array pooled shotgun strategy for sequencing large genomes. Genome Res. 2001;11:1619–23. Deng Y, Pan Y, Luo M. Detection and correction of assembly errors of rice Nipponbare reference sequence. Plant Biol. 2014;16:643–50. Wei X, Xu Z, Wang G, Hou J, Ma X, Liu H, et al. PBACode: A random-barcode-based high-throughput approach for BAC paired-end sequencing and physical clone mapping. Nucleic Acids Res. 2017;45(7):e52. Yang X, Yang Y, Ling J, Guan J, Guo X, Dong D, et al. A high‐throughput BAC end analysis protocol ( BAC ‐anchor) for profiling genome assembly and physical mapping. Plant Biotechnol J. 2019;18(2):364–72. Zhaozhao D, Tong L, Jiadong L, Zhifei H. High-throughput long paired-end sequencing of a Fosmid library by Pacbio. Plant Methods. 2019;15:142. Bankevich A, Nurk S, Antipov D, Gurevich AA, Dvorkin M, Kulikov AS, et al. SPAdes: A New Genome Assembly Algorithm and Its Applications to Single-Cell Sequencing. J Comput Biol. 2012;19:455–77. Pan Y, Wang X, Liu L, Wang H, Luo M. Whole Genome Mapping with Feature Sets from High-throughput Sequencing Data. PLoS One. 2016;11:1–17. Gualberto JM, Newton KJ. Plant Mitochondrial Genomes: Dynamics and Mechanisms of Mutation. Annu Rev Plant Biol. 2017;68:225–52. doi:10.1146/annurev-arplant-043015-112232. Additional Declarations No competing interests reported. Supplementary Files Addtionalfile1.docx Additionalfile2.xlsx Additionalfile3.docx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-536711","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":27732027,"identity":"4994d9e7-0710-4a4f-b3c4-4c9a9605fa13","order_by":0,"name":"Wei Xu","email":"","orcid":"","institution":"Huazhong Agricultural University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Wei","middleName":"","lastName":"Xu","suffix":""},{"id":27732028,"identity":"8df50771-8c96-43a5-b639-6d8968718a7c","order_by":1,"name":"Mengjie Liang","email":"","orcid":"","institution":"Huazhong Agricultural University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Mengjie","middleName":"","lastName":"Liang","suffix":""},{"id":27732029,"identity":"5e08c36e-d9b8-4878-8290-909de3c60f07","order_by":2,"name":"Xue Yang","email":"","orcid":"","institution":"Huazhong Agricultural University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Xue","middleName":"","lastName":"Yang","suffix":""},{"id":27732030,"identity":"1e85e1b3-ebd0-4828-89bc-339728b65b4a","order_by":3,"name":"Hao Wang","email":"","orcid":"","institution":"Huazhong Agricultural University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Hao","middleName":"","lastName":"Wang","suffix":""},{"id":27732031,"identity":"70939001-c24e-4311-97f4-f91c7ab26bd1","order_by":4,"name":"Meizhong Luo","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA7ElEQVRIiWNgGAWjYFAC5gYGhgoGBjZ2BjaoSAIhLYxALWeAWphJ0sLYBrKNWC26/QfbJD7O2ybPB9Ty8EfNYQZ+9hwDhp87cGsxO3CwTXLmttuGbcwM7AYSxw4zSPa8MWDsPYNHy8HGNmnebbcZgVrYJAzYDjMY3MgxYAY7FZeWw4xALXNu24O1JPw7zGBPUMsxkJaG24lgLQfbgLZIENJyhrHZcsax28ltQGWSjX3pPBJnnhUc7MWn5fzhgzc+1Ny2nd/efEzyxzdrOf725I0PfuLRggRAccrAwAMiDhClYRSMglEwCkYBTgAAMCtLZ4yPGQcAAAAASUVORK5CYII=","orcid":"","institution":"Huazhong Agricultural University","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Meizhong","middleName":"","lastName":"Luo","suffix":""}],"badges":[],"createdAt":"2021-05-18 02:43:59","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-536711/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-536711/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":9403712,"identity":"47136991-4d7f-4591-ac4f-4fa7e365d6c5","added_by":"auto","created_at":"2021-05-20 19:24:59","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":49061,"visible":true,"origin":"","legend":"The insert sizes of randomly selected BAC clones determined by PFGE. The maker in the middle is λ DNA ladder.","description":"","filename":"Fig1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-536711/v1/0b512e9e0bf4c491c5d869e8.jpg"},{"id":9403887,"identity":"ed8c3fb4-a7da-4050-b2a0-4fd8ad74fff7","added_by":"auto","created_at":"2021-05-20 19:27:59","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":174657,"visible":true,"origin":"","legend":"The strategy of BAC end sequence localization based on CAPSS. Twenty-four 384-plates are arranged to a square superpool, and then 96 row pools and 96 column pools are prepared and sequenced by NGS platform. The sequences of each pool are assembled into contigs. If a contig (especially, BAC end sequences, abbr. BES) is shared in a row pool and a column pool, it will be assigned to the well at the intersection of the row pool and the column pool. BAC clones will be further mapped to the reference genome according to assigned contigs.","description":"","filename":"Fig2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-536711/v1/fbf63db3a8aa1124071eec39.jpg"},{"id":9403885,"identity":"53388f4a-1c90-4932-953d-506bb5e32101","added_by":"auto","created_at":"2021-05-20 19:27:59","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":152716,"visible":true,"origin":"","legend":"the statistics of the BAC inserts mapped by paired BESs","description":"","filename":"Fig3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-536711/v1/292c67ad5bd7ed017e231086.jpg"},{"id":9403947,"identity":"5f1ba78b-07df-434f-a15b-73d1f311f950","added_by":"auto","created_at":"2021-05-20 19:30:59","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":212998,"visible":true,"origin":"","legend":"Presentation of broomcorn millet BAC locations by JBrowse. A barbell icon idicates a BAC mapped by forward and reverse BESs (yellow rectangle). An icon with arrow indicates a BAC mapped by only single BES.","description":"","filename":"Fig4.jpg","url":"https://assets-eu.researchsquare.com/files/rs-536711/v1/39cc55bb07fd193c3f2ca803.jpg"},{"id":13693971,"identity":"4f2cb3fc-ce36-427b-ac24-7a34aaa46310","added_by":"auto","created_at":"2021-09-17 12:50:56","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":652329,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-536711/v1/40da190a-6123-4253-a828-201b7f5c4f8e.pdf"},{"id":9403888,"identity":"ac923514-7d33-4295-ae93-db320b003dd4","added_by":"auto","created_at":"2021-05-20 19:27:59","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":2168353,"visible":true,"origin":"","legend":"","description":"","filename":"Addtionalfile1.docx","url":"https://assets-eu.researchsquare.com/files/rs-536711/v1/afeeaeec09192a4b70d402f9.docx"},{"id":9403716,"identity":"c49f1da0-e965-4436-a223-90cd3b2659bc","added_by":"auto","created_at":"2021-05-20 19:24:59","extension":"xlsx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":12822,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile2.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-536711/v1/82c5b39124a2eec200355285.xlsx"},{"id":9403946,"identity":"c8e1fde1-c300-4726-86c4-ddc80e056ecd","added_by":"auto","created_at":"2021-05-20 19:30:59","extension":"docx","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":17867,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile3.docx","url":"https://assets-eu.researchsquare.com/files/rs-536711/v1/c7e0291e0d79d5622d0bfd4e.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Genomic Resources of Broomcorn Millet: Demonstration and Application of a High-throughput BAC Mapping Pipeline","fulltext":[{"header":"1 Introduction","content":" \u003cp\u003eWith the increased global water scarcity caused by climate change and population growth, it is of great importance to exploit the high-efficient water-use crop for the food security of the human in the future. Broomcorn millet (\u003cem\u003ePanicum miliaceum\u003c/em\u003e), also known as proso millet, panic millet, and wild millet, is one of the traditional five-grain crops in the north of China [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. It is a typical C4 plant with high photosynthetic efficiency. It also has a high-efficient utilization ratio of water resource and a capacity for drought resistance, exquisitely adapting to semi-drought or drought conditions [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. Furthermore, its growing cycle, 60\u0026ndash;90 days from sowing to maturity, is shorter than other cereals [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. Broomcorn millet contains more protein than most grains, and a relatively balance array of trace elements and vitamins. More than 8700 accessions of broomcorn millet including landraces and cultivars have been conserved in the National Gene Bank of the Institute of Crop Science, Chinese Academy of Agricultural Sciences, thus providing an abundance of resources for genetic improvement of broomcorn millet.\u003c/p\u003e \u003cp\u003eAt present, high-quality chromosome-scale genome assemblies of two allotetraploid (2n\u0026thinsp;=\u0026thinsp;4x\u0026thinsp;=\u0026thinsp;36) broomcorn millet varieties decoded by Chinese researchers are available [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. These genome assemblies provide the foundation for the molecular breeding of broomcorn millet. However, other gemonic resources are still required to complete and make the full use of the genome assemblies.\u003c/p\u003e \u003cp\u003eThe tranditional bacterial artificial chromosome (BAC) libraries with genomic DNA inserts of 50 kbp \u0026ndash; 300 kbp [\u003cspan additionalcitationids=\"CR7\" citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e] are also important resources for genomic research. They provide natural DNA materials for a variety of experiments, such as intact gene cluster cloning, map-based cloning, whole genome sequencing, comparative genomics analysis and fluorescence \u003cem\u003ein situ\u003c/em\u003e hybridization that aim to understand functional elements in the genome [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e, \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. The utility of the BAC libraries can be greatly enhanced by mapping the BAC clones on the genome assemblies.\u003c/p\u003e \u003cp\u003eHence, we constructed a high-quality BAC library for broomcorn millet, developed a new pipeline for cost-effectively decoding BAC end seqeunces generated by clone-array pooled shotgun sequencing strategy (abbr. CAPSS) and Illumina sequencing technology, and mapped the BAC clones on the genome assemblies of the broomcorn millet [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e].\u003c/p\u003e "},{"header":"2 Results","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\n\u003ch2\u003e2.1 BAC library construction\u003c/h2\u003e\n\u003cp\u003eA BAC library of broomcorn millet was constructed with the restriction enzyme \u003cem\u003eHin\u003c/em\u003edIII using high-molecular-weight genomic DNA prepared from etiolated seedlings. In total, the library consists of 76,032 BAC clones, that were arrayed into 198 independent 384-well microtiter plates. Insert sizing of randomly picked clones showed that the majority of genomic BAC inserts fell into the length range of 97\u0026ndash;145.5 kb with an average insert size of 132 kb (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e).\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec4\" class=\"Section2\"\u003e\n\u003ch2\u003e2.2 Construction and sequencing of DNA pools\u003c/h2\u003e\n\u003cp\u003eIn order to obtain BAC end sequences with high efficiency and low cost, we designed and performed a pipeline based on the clone array pooled shotgun sequencing strategy (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e). Randomly selected 24 384-plates (each plate consists of 16 x 24 clones) from the broomcorn millet BAC library were arranged to a square superpool with 6 row plates and 4 column plates. Therefore, the superpool consisted of 96 row pools (6 plates x 16) and 96 column pools (4 plates x 24). These row pools and column pools were called as the secondary pools, and each secondary pool also consisted of 96 BAC clones. In total, one superpool consisted of 192 secondary pools and 9,216 individual BAC clones.\u003c/p\u003e\n\u003cp\u003eDNA of the secondary pools were extracted and Illumina sequencing libraries with individual index sequences for each secondary pool were prepared. Finally, 96 row pool libraries were mixed in equal amounts as the library X, and 96 column pool libraries were also mixed in equal amounts as the library Y. The average insert sizes of the library X and Y analyzed by the Agilent 2100 Bioanalyzer (Additional file 1: Fig. S1) were 491 bp and 443 bp, respectively. Sequencing was accomplished using the Illumina HiSeq 2000 sequencing platform with paired-end protocol (PE150). The Library X and Y generated 130.82 Gb and 133.50 Gb raw data, respectively. Trimmomatic was employed for trimming adaptors and filtering low-quality or shorter reads. FastQC was employed for evaluating the quality of the preprocessing reads. Then, the valid reads were obtained by filtering the reads traced to \u003cem\u003eE. coli\u003c/em\u003e DH10B reference genome using Bowtie2. Finally, we obtained valid yields of 60.40 Gb and 102.37 Gb for libraries X and Y, respectively (Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e). Demultiplexer was employed for demultiplexing and generating each secondary pool reads according to the 7-bp index (Additional file 2). The average valid sequencing depths of row pools and column pools were 45x and 78x, respectively (Additional file 1: Fig. S2).\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n\u003ctable id=\"Tab1\" border=\"1\"\u003e\u003ccaption\u003e\n\u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\n\u003cdiv class=\"CaptionContent\"\u003e\n\u003cp\u003eThe summary of NGS data\u003c/p\u003e\n\u003c/div\u003e\n\u003c/caption\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eLibrary\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eTotal size (Gb)\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eSize after QC (Gb)\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eValid size (Gb)\u003c/p\u003e\n\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eX\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e130.82\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e103.78\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e60.40\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eY\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e133.50\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e107.92\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e102.37\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec5\" class=\"Section2\"\u003e\n\u003ch2\u003e2.3 Parsing BAC end sequences\u003c/h2\u003e\n\u003cp\u003eWe focused our interest on the two special contigs in a BAC: the forward and reverse BAC end seqeunces. We designed two pathways to find them. On one hand, the paired-end reads overlapping with the border sequences of the vector harboring the \u003cem\u003eHin\u003c/em\u003edIII restriction enzyme site were extracted from the valid data of each secondary pool, and then assembled using Cap3 based on the overlap-layout-consensus method. The vector sequence parts in the consensus were truncated to obtain short BAC end sequences (abbr. BES) starting with AAGCTT (\u003cem\u003eHin\u003c/em\u003edIII site). These short BAC end sequences were assigned to the corresponding wells according to the clone array pooled shotgun sequencing strategy. Of the 9,216 wells in 384-plates tested in the pipeline, 8,183 (88.79%) wells were assigned one forward BES and 7,897 (85.69 %) wells were assigned one reverse BES (Additional file 3: Table S1). All assigned BAC end sequences have an average length of 358 bp (Additional file 1: Fig. S3), which is consistent with the insert size of the Illumina libraries (Additional file 1: Fig. S1).\u003c/p\u003e\n\u003cp\u003eOn the other hand, the valid NGS data of each secondary pool were firstly assembled using SPAdes, and the contig N50 sizes of all row pools and column pools were counted (Additional file 1: Fig. S4). The average N50 of row pools and column pools were 13.97 kb and 11.45 kb, respectively. Likewise, the vector sequence parts at the ends of the contigs were removed to retain the long BESs starting with AAGCTT. Finally, these long BESs were assigned to the corresponding wells in 384-plates. Of the 9,216 wells, 5,454 (59.18%) wells obtained one forward BES and 5,108 (55.43%) wells obtained one reverse BES (Additional file 3: Table S1). The N50 of the long BESs is 60.87 kb (Additional file 1: Fig. S5).\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec6\" class=\"Section2\"\u003e\n\u003ch2\u003e2.4 Determination of BAC locations on the genome\u003c/h2\u003e\n\u003cp\u003eIn order to determine the locations of the broomcorn millet BAC clones on the genome, we aligned the short and long BAC end sequences onto the cultivar longmi4 genome with Blastn. Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e listed the alignment results. With BAC end sequences with repeats, 14,862 (83.44%) short BAC end sequences including 12,971 single-hit and 1,891 multi-hit sequences were mapped to the genome, and 12,120 (95.28%) long BAC end sequences (all single-hit) were mapped to the genome. With the BAC end sequences without repeats, 7,760 (43.57%) short BAC end sequences including 7,295 single-hit and 465 multi-hit sequences were mapped to the genome, and 10,626 (83.53%) long BAC end sequences (all single-hit) were mapped to the genome. In order to map as many as possible BACs to the genome, the alignment results of the end sequences with repeats were adopted. We wrote a python script to extract the Blast results from short and long BESs. As a result, 5,795 BACs were mapped to the genome using short BAC end sequences, and 6,973 BACs were mapped to the genome using long BAC end sequences. Finally, the Blast results generated by short and long BAC end sequences were integrated, and in total 8,262 BACs (89.65%) were mapped to the genome.\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n\u003ctable id=\"Tab2\" border=\"1\"\u003e\u003ccaption\u003e\n\u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\n\u003cdiv class=\"CaptionContent\"\u003e\n\u003cp\u003eA summary of the Broomcorn millet BESs and the anchoring results of the Broomcorn millet BAC clones to the longmi 4 genome using the BESs\u003c/p\u003e\n\u003c/div\u003e\n\u003c/caption\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eCategories\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eShort BESs\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eLong BESs\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eIntegrity\u003c/p\u003e\n\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eBAC end sequences\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eClones in superpool\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e9216\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e9216\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eClones with successful BESs\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e9126\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e7725\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003ewith paired BESs\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e7790\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e3890\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003ewith single-end BESs\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1336\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e3835\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eonly forward\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e921\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e2189\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eonly reverse\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e415\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1646\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eTotal successful BESs\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e17811 (96.63%)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e12721 (66.17%)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eAlignment\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eAligned BESs with repeats unmasked\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e14862 (83.44%)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e12120 (95.28%)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eSingle-hit BESs\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e12971\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e12120\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eMulti-hit BESs\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1891\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eAligned BESs with repeats masked\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e7760 (43.57%)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e10626 (83.53%)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eSingle-hit BESs\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e7295\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e10626\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eMulti-hit BESs\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e465\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eAnchoring to reference sequences\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eClones anchored to single sites\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e5795\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e6973\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e8262\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eClones anchored with single BES\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e2507\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e3907\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e2871\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eClones anchored with paired BESs\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e3288\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e3066\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e5391\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eTo verify the accuracy of BAC locations on the genome, 55 BAC clones were randomly picked from 384-plates for BAC end sequencing using Sanger method. After quality control, the 55 paired Sanger end sequences were blasted with the above short and long BESs, and the genome. All BAC clones but one (32G16) were consistent. By checking the 32G16 BAC end sequences, we found that this clone (or rather 32G16 well) was assigned two long forward BESs, one short forward BES, and no reverse BES. Only one long forward BES and the short forward BES were perfectly identical to the Sanger end sequence. Another long forward BES confused the mapping. However, it can be solved by weighting. In summary, the accuracy of the BAC mapping approach in this study was extremely high.\u003c/p\u003e\n\u003cp\u003eThe distribution of the 5,391 BACs mapped to the genome by paired BAC end sequences were counted (Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e). These clones covered a total of 432.47 Mb of chromosomes with a total coverage of 50.97%. Among the 18 chromosome sequences of the cultivar longmi4 genome, there left 829 gaps. Our BACs covered 308 of them. The insert sizes of these BACs presented the Gaussian distribution, with an average insert size of 123.48 kb (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e), which is lower than that predicted by pulse field electrophoresis. These BACs are valuable resources for further improvement of the genome.\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n\u003ctable id=\"Tab3\" border=\"1\"\u003e\u003ccaption\u003e\n\u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e\n\u003cdiv class=\"CaptionContent\"\u003e\n\u003cp\u003eThe location result of BACs on broomcorn millet chromosomes\u003c/p\u003e\n\u003c/div\u003e\n\u003c/caption\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth rowspan=\"2\" align=\"left\"\u003e\n\u003cp\u003eChr.\u003c/p\u003e\n\u003c/th\u003e\n\u003cth rowspan=\"2\" align=\"left\"\u003e\n\u003cp\u003eLength (Mb)\u003c/p\u003e\n\u003c/th\u003e\n\u003cth rowspan=\"2\" align=\"left\"\u003e\n\u003cp\u003eNo. gaps\u003c/p\u003e\n\u003c/th\u003e\n\u003cth rowspan=\"2\" align=\"left\"\u003e\n\u003cp\u003eNo. BACs\u003c/p\u003e\n\u003c/th\u003e\n\u003cth colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003eBACs Covered\u003c/p\u003e\n\u003c/th\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003elength (Mb)\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eNo. gaps\u003c/p\u003e\n\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003echr 1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e69.18\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e85\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e439\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e35.25\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e34\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003echr 2\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e61.15\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e59\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e379\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e31.04\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e19\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003echr 3\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e57.97\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e50\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e394\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e31.03\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e20\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003echr 4\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e56.29\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e34\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e359\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e27.90\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e14\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003echr 5\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e54.13\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e52\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e361\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e29.86\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e19\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003echr 6\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e52.84\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e46\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e358\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e28.24\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e16\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003echr 7\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e51.23\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e67\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e286\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e23.67\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e27\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003echr 8\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e48.26\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e29\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e339\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e26.49\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e8\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003echr 9\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e45.11\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e70\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e240\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e20.96\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e21\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003echr10\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e44.65\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e53\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e308\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e25.44\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e28\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003echr11\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e43.18\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e30\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e267\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e21.90\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e14\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003echr12\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e42.47\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e30\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e254\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e20.79\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e11\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003echr13\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e40.72\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e50\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e261\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e21.11\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e17\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003echr14\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e38.49\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e32\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e269\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e20.23\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e10\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003echr15\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e34.36\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e34\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e213\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e18.14\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e8\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003echr16\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e33.61\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e45\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e212\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e16.72\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e15\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003echr17\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e32.99\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e25\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e226\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e18.26\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e10\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003echr18\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e32.24\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e38\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e199\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e15.32\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e17\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eunplaced\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e9.48\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e27\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e-\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eTotal\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e848.47\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e829\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e5391\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e432.47\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e308\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec7\" class=\"Section2\"\u003e\n\u003ch2\u003e2.5 Presentation of BAC locations by JBrowse\u003c/h2\u003e\n\u003cp\u003eIn order to effectively use the BAC resource, quickly and easily retrieve BAC clones and view other annotaion information, we established a resource website employing a lightweight JBrowse to display the BAC resource information on the broomcorn millet genome (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003e).\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec8\" class=\"Section2\"\u003e\n\u003ch2\u003e2.6 Organelle genomes of the broomcorn millet\u003c/h2\u003e\n\u003cp\u003eWe aligned all BESs in this study to the chloroplast (CM009689) genome of broomcorn millet by local blast. A total of 65 BAC clones were mapped to the chloroplast with more than 99% of similarity, of which 16 BAC clones were determined by paired BESs (25C07, 25N09, 27F16, 27O19, 29B02, 29K11, 32J09, 32N16, 33D23, 33K24, 36A05, 39C22, 42C10, 46C16, 48B15, 48C03). In the absence of mitochrondrial sequence in broomcorn millet, we aligned BESs to all mitochondrial genomes in \u003cem\u003eGramineae\u003c/em\u003e. Of 16 BAC clones with homologous BES sequences, 13 BACs were simultaneously aligned to the chloroplast and 3 BACs were simultaneously aligned to the nuclear genome.\u003c/p\u003e\n\u003c/div\u003e"},{"header":"3 Discussion","content":" \u003cp\u003eBAC library is still a powerful resource for genome assembly, functional genomics research, and long-term genetic resoure storage for the endangered species. BAC seqeunces, espcially paired BESs, are usally used to detect assembly errors or assist assembly [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. The most conventional and convenient approach to decode BESs is Sanger sequencing method. However, its one-by-one style is laborious, time-consuming and expensive. Fortunately, a few of efficient approaches had developped, based on next generation sequencing technologies, such as pBACcode and BAC-anchor[\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e, \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. pBACode determines paired BESs by a pair of random barcodes flanking the cloning site. BAC-anchor determines general paired BESs by using specific restriction enzyme sites and searching for utlra-long paired-end subreads containing large internal gaps. We previously also developed a high-throughput approach for long accurate BES profiling with PacBio sequencing technology [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]. However, these approaches focused on paired BES profiles, and they are hard to trace BESs of specific clones in 384-plate wells. In this study, we generated BESs and assigned them to physical wells in 384-plates by applying the characteristics of row-cross-column in two-dimension arrays and cost-effective Illumina platform. These BAC end sequences in average are longer and accurater than that generated by Sanger method, for assembling by SPAdes assembler[\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]. After alignment, 89.65% of clones were successfully associated with the broomcorn millet longmi 4 genome. These clones covered 308 of the 829 gaps left by the genome and can be uesd to close the genome gaps.\u003c/p\u003e \u003cp\u003eIn conventional whole genome sequencing project, sequencing coverage is usallly at least 30x, and not more than 200x. Higher coverage will result in more sequences that are generated by PCR mutations or sequencing errors and lead to contigs with shorter N50. We added index into each secondary pool and mixed 96 secondary pools as a sequencing library to run in a lane of Illumina flow cell for greatly reducing the cost of sequencing. As a result, after removing the contaminated \u003cem\u003eE. coli\u003c/em\u003e genomic DNA reads the average valid sequencing depths of row pools and column pools were 45x and 78x, respectively.\u003c/p\u003e \u003cp\u003eIn the process of BESs extraction, we designed two pathways: short BES pathway and long BES pathway. Short BES pathway searched all reads overhanging with verctor end sequences before assembling by Cap3; consequently, it generated all potenital BESs. Long BES pathway identified all contigs overhanging with vector end seqeunces after assembling by SPAdes; consequently, it generated longer but less BESs than the first pathway. The assignment of BAC end sequences at the intersection sites is affected by many factors, such as the overlapping rate of the BAC clones and the correct rate of the sequences. The overlapping rate of BAC clones in the superpool is the most important factor. If overlapping BAC clones appear in the same row or column pool, we can assign them to wells easily and correctly. Also If two overlapping BAC clones appear in a different row or column pool, we can rectify them by our previous method [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]. However, if more than two overlapping BAC clones appear in a different row or column pool, our method will filter out potential BESs, so that the BESs of the intersection well will be absent. In the process of BES assignment, the flow of the forward and reverse BESs are completely independent. When more than one forward and/or reverse BESs were assigned to a well, we cannot determine which BESs are a pair of BESs. If a high-qualty genome is available, it is easy to assess which pair of BESs are derived from the same BAC by mapping. However, if the variety used for BAC library construction is not the same as that the reference genome stands for, the alignment results that do not satisfy the location requirement of BES pairs will be discarded. The clones that may contain a large structural variation can be picked out for further analysis from the 384-plate wells.\u003c/p\u003e \u003cp\u003ePlant cells contain an abundance of chloroplasts and mitochondria. The chloroplast genome is generally around 150 kb, while the mitochondrial genome size varies widely, typically between 200 kb and 750 kb [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. Although nuclei are extracted for BAC library construction, a trace of organelle DNA cotamination is inevitable. In this study, a small number of clones of chloroplast genome was found by BESs, while clones of mitochondrial genome were almost absent.\u003c/p\u003e \u003cp\u003eBy high-throughput sequencing and mapping clones in the secondary pool of BAC libraries, we can find the coordinates of genome sequences in the broomcorn millet BAC libraries. Therefore, if we find genes that play an important role in biology, we can quickly locate the BAC clones containing this gene, and obtain the experimental materials for further research and analysis. At the same time, because BAC contains a long DNA fragment (about 120 kb), it is also convenient for us to quickly analyze the upstream and/or downstream DNA elements of interested genes or adjacent genes.\u003c/p\u003e "},{"header":"4 Conclusions","content":" \u003cp\u003eWe constructed a high-quality BAC library for broomcorn millet, developed a high-efficient and low-cost pipeline which can generate and parse ten thousand of paired BAC end sequences at each time, and mapped a total of 8,262 broomcorn millet BACs to the chromosomes. These clones covered 308 of the 829 gaps left by the genome. The high-quality BAC clones mapped on genome in this study will provide a powerful genomic resource for gap filling, complex segment sequencing, FISH, functional research, and genetic engineering of broomcorn millet.\u003c/p\u003e "},{"header":"5 Methods","content":"\u003cdiv id=\"Sec12\" class=\"Section2\"\u003e\n\u003ch2\u003e5.1 Plant materials, growth conditions and BAC library construction\u003c/h2\u003e\n\u003cp\u003eThe seeds of broomcorn millet were provided by professor Mingsheng Chen of the Institute of Genetics and Developmental Biology, Chinese Academy of Sciences and grown at dark conditions under 25\u0026deg;C. The young leaves of seedlings were harvested and mixed, frozen immediately in liquid nitrogen for the extraction of nuclear DNA. High molecular weight nuclear DNA was extracted and BAC library was constructed following our previous protocol [\u003cspan class=\"CitationRef\"\u003e6\u003c/span\u003e]. Partial digestions of DNA plugs with dilution \u003cem\u003eHin\u003c/em\u003edIII were performed. DNA fragments ranging from 100 kb to 200 kb were recovered from pulse field gel and ligated with pIndigoBAC536-S vector [\u003cspan class=\"CitationRef\"\u003e7\u003c/span\u003e]. The ligation product was used to transform DH10B cells by electroporation. White colonies were picked up and stored in 384-plates at -80 \u0026ordm;C.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec13\" class=\"Section2\"\u003e\n\u003ch2\u003e5.2 Pool construction\u003c/h2\u003e\n\u003cp\u003eTwenty-four 384-plates were chosen form the BAC library of broomcorn millet. These plates were arranged in a 2-dimension superpool. The superpool contains 96 row pools and 96 column pools. A total of 192 pools were processed independently for high quality DNA extraction with the AxyGen AxyPrep Easy-96 plasmid kit. Then, these BAC DNAs were completely digested with ATP-Dependent DNase (Epicentre) to remove the host \u003cem\u003eE. coli\u003c/em\u003e DNA. After digestion, these BAC DNAs were sheared in the Bioruptor to an average of 500 bp. During blunt-end repair, overhanging 5\u0026rsquo; and 3\u0026rsquo; ends were filled in or removed by T4 DNA polymerase. 5\u0026rsquo;-phosphates were attached using T4 polynucleotide kinase. Tail-A were added using Taq DNA polymerase. Then, adapters were ligated to both ends of the molecules using T4 DNA ligase. The ligation products were cleaned using MagBead DNA Purification Kit (Sangon Biotech, shanghai, CN). Sequencing adaptors were added using PCR. Finally, the products of KOD PCR were cleaned again.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec14\" class=\"Section2\"\u003e\n\u003ch2\u003e5.3 Illumina sequencing and BAC end sequence analysis\u003c/h2\u003e\n\u003cp\u003eNGS sequencing of 2 mixed DNA libraries was performed via the Illumina HiSeq 2000 with 150-bp paired-end protocol (Genewiz, suzhou, CN). Trimmotatic was employed to filter and trim raw reads. FastQC was employed to assess data quality. After adapter filtering and quality assessment, BBMAP/demuxbyname script was employed for deconvolution depending on unique index sequence in each pool. A python script was employed to extract target paired-end reads that cover BES-VES site. Cap3 was employed to assemble those reads to consensuses, and consensuses were then trimmed to remove the part of vector to generate BESs called \u0026ldquo;short BES\u0026rdquo;. SPAdes was employed to directly assemble pool reads to contigs. Then, the contigs containing BAC vector sequences were extracted and trimmed using python script to generate the trimmed contigs called \u0026ldquo;long BES\u0026rdquo;. Blastn was used to align BESs from raw and column pools, and then the shared BESs were assigned to the wells at the intersection.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec15\" class=\"Section2\"\u003e\n\u003ch2\u003e5.4 Validation of BAC end sequences\u003c/h2\u003e\n\u003cp\u003eThe analysis results of the BAC end sequences were validated by Sanger sequencing. Fifty-five BAC clones were randomly selected and their DNAs were extracted using an improved alkaline lysis protocol. Sanger sequencing was accomplished using BAC-F (5\u0026rsquo;-AACGACGGCCAGTGAATTG-3\u0026rsquo;) and BAC-R (5\u0026rsquo;-GATAACAATTTCACACAGG-3\u0026rsquo;) primers from pIndigoBAC536-S vector backbone. The BAC end sequences from Sanger sequencing were aligned to BES from the analysis results of illumina data using Blastn.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec16\" class=\"Section2\"\u003e\n\u003ch2\u003e5.5 BAC Mapping on broomcorn millet longmi4 genome\u003c/h2\u003e\n\u003cp\u003eThe genome sequence of broomcorn millet longmi4 was downloaded from NCBI genome database under GCA_002895445 accession, which was submitted by researchers from China Agricultural University. The local blastn was employed to map all BESs to the genome with the following options: qcov_hsp_perc\u0026thinsp;=\u0026thinsp;99, perc_identity\u0026thinsp;=\u0026thinsp;99, outfmt\u0026thinsp;=\u0026thinsp;6, culling_limit\u0026thinsp;=\u0026thinsp;1. The results were further converted to a GFF3 format file using a Python script. In this script, following conditions were set: if both forward and reverse BESs in each clone were mapped to the same chromosome, and their orientations were opposite and their interval lengths were less than 250 kb, such clones were recorded in GFF3 format file with three lines; if either forward or reverse BES was uniquely mapped to chromosome, such clones also were recorded in GFF3 format file with two lines; other conditions would be discarded. The GFF3 file was sorted using GFF3sort and presented with JBrowse in our website (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://eightstarsbio.com/gresource/JBrowse-1.16.5/index.html\u003c/span\u003e\u003c/span\u003e).\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec17\" class=\"Section2\"\u003e\n\u003ch2\u003e5.6 Organelle genome analysis\u003c/h2\u003e\n\u003cp\u003eThe chloroplast genome (CM009689) of brromcorn millet and all mitochondrial genomes (NC_008331, NC_007982, NC_036024, NC_031164, NC_029816, NC_022714, NC_022666, NC_013816, NC_007886, NC_011033, NC_008362, NC_008360, NC_008332, NC_008333) in \u003cem\u003eGramineae\u003c/em\u003e were downloaded form NCBI nucleotide database. All potential BACs of organelle genome were identified by the local blastn with following options: qcov_hsp_perc\u0026thinsp;=\u0026thinsp;99, perc_identity\u0026thinsp;=\u0026thinsp;99, outfmt\u0026thinsp;=\u0026thinsp;6.\u003c/p\u003e\n\u003c/div\u003e"},{"header":"6 Abbreviations","content":"\u003cp\u003eBAC: bacterial artificial chromosome;\u003c/p\u003e\n\u003cp\u003eBES: BAC end sequence;\u003c/p\u003e\n\u003cp\u003eCAPSS: clone array pooled shotgun sequencing;\u003c/p\u003e\n\u003cp\u003eCM: chloramphenicol;\u003c/p\u003e\n\u003cp\u003eNGS: next generation sequencing;\u003c/p\u003e\n\u003cp\u003ePCR: polymerase chain reaction;\u003c/p\u003e\n\u003cp\u003eVES: vector end sequence.\u003c/p\u003e"},{"header":"7 Declarations","content":"\u003ch2\u003e7.1 Ethics approval and consent to participate\u003c/h2\u003e\n\u003cp\u003eNot applicable\u003c/p\u003e\n\u003ch2\u003e7.2 Consent for publication\u003c/h2\u003e\n\u003cp\u003eNot applicable\u003c/p\u003e\n\u003ch2\u003e7.3 Availability of data and materials\u003c/h2\u003e\n\u003cp\u003eThe source codes are openly available in a GitHub repository (\u003ca href=\"https://github.com/xuweixw/broomcorn-millet-BAC-library\"\u003ehttps://github.com/xuweixw/broomcorn-millet-BAC-library\u003c/a\u003e). Illumina sequencing data are available at Sequence Read Archive (SRA) under the accession PRJNA576359. BAC clones can be browsed and obtained through our website (\u003ca href=\"http://eightstarsbio.com/gresource/JBrowse-1.16.5/index.html\"\u003ehttp://eightstarsbio.com/gresource/JBrowse-1.16.5/index.html\u003c/a\u003e).\u003c/p\u003e\n\u003ch2\u003e7.4 Competing interests\u003c/h2\u003e\n\u003cp\u003eThe authors declare that they have no competing interests.\u003c/p\u003e\n\u003ch2\u003e7.5 Funding\u003c/h2\u003e\n\u003cp\u003eThis work was supported by a grant from the National Natural Science Foundation of China (Grant no. 31671268).\u003c/p\u003e\n\u003ch2\u003e7.6 Authors' contributions\u003c/h2\u003e\n\u003cp\u003eWX and MLuo conceived and designed the research framework; WX, MLiang, XY and HW performed the experiments; WX analyzed the data and wrote the manuscript; MLuo supervised the work and finalized this manuscript. All authors read and approved the manuscript.\u003c/p\u003e\n\u003ch2\u003e7.7 Acknowledgements\u003c/h2\u003e\n\u003cp\u003eWe are grateful to Dr. Mingsheng Chen (the Institute of Genetics and Developmental Biology, Chinese Academy of Sciences) for providing the seeds of broomcorn millet.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eKalinova J, Moudry J. Content and Quality of Protein in Proso Millet ( Panicum miliaceum L .) Varieties. Plant Foods Hum Nutr. 2006;61:45\u0026ndash;9.\u003c/li\u003e\n\u003cli\u003eWashburn JD, Schnable JC, Davidse G, Pires JC. Phylogeny and photosynthesis of the grass tribe Paniceae. Am J Bot. 2015;102:1493\u0026ndash;505.\u003c/li\u003e\n\u003cli\u003eBaltensperger DD. Progress with proso, pearl and other millets. 2002.\u003c/li\u003e\n\u003cli\u003eShi J, Ma X, Zhang J, Zhou Y, Liu M, Huang L, et al. Chromosome conformation capture resolved near complete genome assembly of broomcorn millet. Nat Commun. 2019;10:1\u0026ndash;9. doi:10.1038/s41467-018-07876-6.\u003c/li\u003e\n\u003cli\u003eZou C, Li L, Miki D, Li D, Tang Q, Xiao L, et al. The genome of broomcorn millet. Nat Commun. 2019;10:1\u0026ndash;12. doi:10.1038/s41467-019-08409-5.\u003c/li\u003e\n\u003cli\u003eLuo M, Wing RA. An improved method for plant BAC library construction. In: Plant Functional Genomics. 2003. p. 3\u0026ndash;19.\u003c/li\u003e\n\u003cli\u003eShi X, Zeng H, Xue Y, Luo M. A pair of new BAC and BIBAC vectors that facilitate BAC/BIBAC library construction and intact large genomic DNA insert exchange. Plant Methods. 2011;7:33. doi:10.1186/1746-4811-7-33.\u003c/li\u003e\n\u003cli\u003eShizuya H, Birren B, Kim UJ, Mancino V, Slepak T, Tachiiri Y, et al. Cloning and stable maintenance of 300-kilobase-pair fragments of human DNA in Escherichia coli using an F-factor-based vector. Proc Natl Acad Sci U S A. 1992;89:8794\u0026ndash;7.\u003c/li\u003e\n\u003cli\u003ePan Y, Deng Y, Lin H, Kudrna DA, Wing RA, Li L, et al. Comparative BAC-based physical mapping of Oryza sativa ssp. indica var. 93-11 and evaluation of the two rice reference sequence assemblies. Plant J. 2014;77:795\u0026ndash;805.\u003c/li\u003e\n\u003cli\u003eDong G, Shen J, Zhang Q, Wang J, Yu Q, Ming R, et al. Development and Applications of Chromosome-Specific Cytogenetic BAC-FISH Probes in S. spontaneum. Front Plant Sci. 2018;9:218. doi:10.3389/fpls.2018.00218.\u003c/li\u003e\n\u003cli\u003eCai WW, Chen R, Gibbs RA, Bradley A. A clone-array pooled shotgun strategy for sequencing large genomes. Genome Res. 2001;11:1619\u0026ndash;23.\u003c/li\u003e\n\u003cli\u003eDeng Y, Pan Y, Luo M. Detection and correction of assembly errors of rice Nipponbare reference sequence. Plant Biol. 2014;16:643\u0026ndash;50.\u003c/li\u003e\n\u003cli\u003eWei X, Xu Z, Wang G, Hou J, Ma X, Liu H, et al. PBACode: A random-barcode-based high-throughput approach for BAC paired-end sequencing and physical clone mapping. Nucleic Acids Res. 2017;45(7):e52.\u003c/li\u003e\n\u003cli\u003eYang X, Yang Y, Ling J, Guan J, Guo X, Dong D, et al. A high‐throughput BAC end analysis protocol ( BAC ‐anchor) for profiling genome assembly and physical mapping. Plant Biotechnol J. 2019;18(2):364\u0026ndash;72.\u003c/li\u003e\n\u003cli\u003eZhaozhao D, Tong L, Jiadong L, Zhifei H. High-throughput long paired-end sequencing of a Fosmid library by Pacbio. Plant Methods. 2019;15:142.\u003c/li\u003e\n\u003cli\u003eBankevich A, Nurk S, Antipov D, Gurevich AA, Dvorkin M, Kulikov AS, et al. SPAdes: A New Genome Assembly Algorithm and Its Applications to Single-Cell Sequencing. J Comput Biol. 2012;19:455\u0026ndash;77.\u003c/li\u003e\n\u003cli\u003ePan Y, Wang X, Liu L, Wang H, Luo M. Whole Genome Mapping with Feature Sets from High-throughput Sequencing Data. PLoS One. 2016;11:1\u0026ndash;17.\u003c/li\u003e\n\u003cli\u003eGualberto JM, Newton KJ. Plant Mitochondrial Genomes: Dynamics and Mechanisms of Mutation. Annu Rev Plant Biol. 2017;68:225\u0026ndash;52. doi:10.1146/annurev-arplant-043015-112232.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Broomcorn millet, BAC, BES, Genomics resources, Gap filling, CAPSS, Jbrowse","lastPublishedDoi":"10.21203/rs.3.rs-536711/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-536711/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eBackground\u003c/strong\u003e: With high-efficient water-use and drought tolerance, broomcorn millet has emerged as a candidate for food security. To promote its research process for molecular breeding and functional research, a comprehensive genome resource is of great importance. \u003c/p\u003e\u003cp\u003e\u003cstrong\u003eResults\u003c/strong\u003e: Herein, we constructed a BAC library for broomcorn millet, generated BAC end sequences based on the clone-array pooled shotgun sequencing strategy and Illumina sequencing technology, and integrated BAC clones into genome by a novel pipeline for BAC end profiling. The BAC library is consisted of 76,023 clones with an average insert length of 123.48 Kb, covering about 9.9-fold of the 850 Mb genome. Of 9,216 clones tested using our pipeline, 8,262 clones were mapped on the broomcorn millet cultivar longmi4 genome. These mapped clones covered 308 of the 829 gaps left by the genome. To our knowledge, this is the only BAC resource for broomcorn millet.\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eConclusions\u003c/strong\u003e: We constructed a high-quality BAC libraray for broomcorn millet and designed a novel pipeline for BAC end profiling. BAC clones can be browsed and obtained from our website (\u003ca href=\"http://eightstarsbio.com/gresource/JBrowse-1.16.5/index.html\" rel=\"noopener noreferrer\" target=\"_blank\"\u003ehttp://eightstarsbio.com/gresource/JBrowse-1.16.5/index.html\u003c/a\u003e). The high-quality BAC clones mapped on genome in this study will provide a powerful genomic resource for genome gap filling, complex segment sequencing, FISH, functional research, and genetic engineering of broomcorn millet.\u003c/p\u003e","manuscriptTitle":"Genomic Resources of Broomcorn Millet: Demonstration and Application of a High-throughput BAC Mapping Pipeline","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2021-05-20 19:24:57","doi":"10.21203/rs.3.rs-536711/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"907ff6e5-98bd-4aa7-b423-c00e058f4eef","owner":[],"postedDate":"May 20th, 2021","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":4448360,"name":"Epigenetics \u0026 Genomics"}],"tags":[],"updatedAt":"2021-06-02T13:14:16+00:00","versionOfRecord":[],"versionCreatedAt":"2021-05-20 19:24:57","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-536711","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-536711","identity":"rs-536711","version":["v1"]},"buildId":"WrCJVZZCHTDjtuVLN7oU0","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00