Chromosome-level genome assembly of the big-footed bat (Myotis pilosus)

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Some bat species in the genus Myotis have evolved longevity-associated mechanisms and exhibit remarkable resistance to cancer. Among them, the big-footed bat (Myotis pilosus) has been confirmed as a cancer-resistant species. Here, we assembled a chromosome-level genome of the big-footed bat, utilizing a combination of ONT long reads and Hi-C technologies. The size of this genome is 1968.27 Mb, with a contig N50 of 41.29 Mb. All assembled sequences were anchored onto 21 autosomes and X chromosome. We identified 739.02 Mb (37.55%) of repetitive sequences in the genome and predicted 21,368 protein-coding genes. Assessment of the genome assembly quality indicated that the assembled genome of the big-footed bat exhibits excellent continuity, completeness, and accuracy. Taken together, we have generated the first chromosome-level genome assembly of the big-footed bat, providing an important reference resource for genetic and genomic studies of long-lived bat species.
Full text 81,415 characters · extracted from preprint-html · click to expand
Chromosome-level genome assembly of the big-footed bat (Myotis pilosus) | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Data Note Chromosome-level genome assembly of the big-footed bat ( Myotis pilosus ) Shilin Tian, Libiao Zhang, Huabin Zhao This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7561642/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Some bat species in the genus Myotis have evolved longevity-associated mechanisms and exhibit remarkable resistance to cancer. Among them, the big-footed bat ( Myotis pilosus ) has been confirmed as a cancer-resistant species. Here, we assembled a chromosome-level genome of the big-footed bat, utilizing a combination of ONT long reads and Hi-C technologies. The size of this genome is 1968.27 Mb, with a contig N50 of 41.29 Mb. All assembled sequences were anchored onto 21 autosomes and X chromosome. We identified 739.02 Mb (37.55%) of repetitive sequences in the genome and predicted 21,368 protein-coding genes. Assessment of the genome assembly quality indicated that the assembled genome of the big-footed bat exhibits excellent continuity, completeness, and accuracy. Taken together, we have generated the first chromosome-level genome assembly of the big-footed bat, providing an important reference resource for genetic and genomic studies of long-lived bat species. Animal Science big-footed bat genome assembly Figures Figure 1 Background & Summary Bats (Chiroptera) comprise over 1,400 extant species, accounting for approximately 20% of all living mammals. They have successfully evolved a series of unique adaptations, including powered flight, echolocation, a distinctive immune system, and exceptional longevity 1 . Most bat species live more than three times longer than other mammals of similar body size 2 , 3 . Among all bats, species of the genus Myotis exhibit notably long lifespans 4 , with at least 13 Myotis species documented to survive at least 20 years in the wild, including the longest-lived recaptured bat to date—the Brandt’s bat ( Myotis brandtii ), which has a lifespan exceeding 41 years. Molecular and genomic studies have increasingly provided evidence for the longevity mechanisms in bats. Comparative Genomic analyses of M. brandtii indicate that an altered growth hormone/insulin-like growth factor 1 (GH/IGF1) axis, together with adaptations such as hibernation and low reproductive rate, collectively contribute to the exceptional lifespan of Brandt’s bat 5 . Long-lived mammals typically exhibit pronounced cancer resistance; in the long-lived bat species Myotis myotis , oncogenic microRNAs are downregulated, whereas cancer-associated microRNAs are upregulated 6 . Recent studies have also revealed that another Myotis species, M. pilosus , is potentially long-lived, as its primary fibroblast cell lines exhibit significantly higher β-galactosidase activity—a marker of cellular senescence—compared to the known long-lived bat Rhinolophus ferrumequinum (> 30.5 years, data from the AnAge Database 7 ). Further transcriptomic and functional analyses indicate that the downregulation of three genes ( HIF1A , COPS5 , and RPS3 ) plays a key role in the cancer resistance of M. pilosus . In this study, we generated a de novo assembly of the female big-footed bat genome by integrating 158.49 Gb of Oxford Nanopore Technologies (ONT) long reads, 212.83 Gb of Hi-C data, and 60.74 Gb of Illumina paired-end reads (Table 1 ). Using an improved four-step assembly strategy 8 , we obtained a 1968.27 Mb genome with a contig N50 of 41.29 Mb (Table 2 ), anchored to 22 chromosomes (Fig. 1 ). We successfully identified the X chromosome, which spans 127.95 Mb and exhibits a high degree of conservation. The genome contains 739.02 Mb of repetitive sequences (37.55% of the genome) and encodes 21,368 protein-coding genes, 10,757 microRNAs, 166,812 tRNAs, 648 rRNAs, and 3,317 snRNAs. Quality assessment demonstrates that this assembly is comparable to previously published bat genomes in terms of continuity, completeness, and accuracy (Table 2 ). As the first chromosome-level genome of the big-footed bat, it provides a valuable resource for exploring the genetic basis of longevity and cancer resistance in the genus Myotis . Table 1 Summary of paired-end data, Hi-C data and long-read sequencing data Illumina sequences Library Raw data (Gb) High Quality Data Base (Gb) Q20 (%) Q30 (%) Paired-end 60.74 57.29 97.01 91.91 Hi-C 212.83 176.71 97.15 92.18 Long-read sequences Platform Total read bases (Gb) Read number Read N50 (bp) Mean read length (bp) ONT 158.49 6,478,552 32,339 24,463 Table 2 Summary of genome survey, assembly and assessment Type Key Metrics Values Genome Survey Mer-length 17 Number of K mers 51,175,072,948 Estimated genome size (Mb) 2,105.94 Heterozygous rate (%) 0.25 Genome Assembly Assembly length (Mb) 1,968.27 Contig Number 154 Contig N50 length (Mb) 41.29 Anchored onto chromosomes (Mb) 1,966.68 Scaffold N50 length (Mb) 107.20 Repetitive sequences (Length / Percentage) DNA transposon 90.07 Mb / 4.58% LINE 584.93 Mb / 29.72% SINE 9.22 Mb / 0.47% LTR retrotransposon 98.39 Mb / 5.00% Satellite 1.08 Mb / 0.06% Simple repeat 2.11 Mb / 0.11% Others 12.11 Mb / 0.62% Total 739.02 Mb / 37.55% Protein-coding Genes Number 21,368 Average gene length (bp) 32,130.11 Average CDS length per gene (bp) 1,496.34 Average exon number per gene 8.44 Average exon length (bp) 177.37 Average intron length (bp) 3,738.94 ncRNAs (Copy / Length) microRNAs (miRNA) 10,757 / 1.26 Mb transfer RNA (tRNA) 166,812 / 12.14 Mb ribosomal RNAs (rRNA) 648 / 0.11 Mb small nuclear RNAs (snRNA) 3,317 / 0.37 Mb Methods Sampling and sequencing An adult female big-footed bat was collected from Guangzhou City, Guangdong Province, China. The specimen was captured using mist nets at cave sites, placed in a clean cloth bag, and transported to a temporary laboratory for processing. Field sampling was approved by the Guangdong Institute of Applied Biological Resources, Institute of Zoology, Guangdong Academy of Sciences (approval number: GIABR20200810). The sampling activities did not require any additional specific permissions and did not involve any endangered or protected species. DNA was isolated from liver tissue using Qiagen Genomic DNA extraction kits and was tested for quality before use in genomic libraries. Short-insert libraries (~ 350 bp) were prepared from ~ 1.5 µg of genomic DNA using the TruSeq Nano DNA HT Sample Preparation Kit (Illumina) and were sequenced on an Illumina NovaSeq platform to generate 150 bp paired-end reads. Raw reads were quality-filtered using fastp (v0.23.4) 9 . For long-read sequencing, genomic DNA was treated with the NEBNext FFPE DNA Repair Mix (M6630) and the NEBNext End Repair/dA-Tailing Module (E7546) according to the manufacturer’s instructions. Libraries were sequenced on a PromethION platform (Oxford Nanopore Technologies), and reads with Phred quality scores below 7 were removed prior to assembly. Hi-C libraries were constructed from liver cells to assist chromosome-level scaffolding. Cells were cross-linked with formaldehyde, lysed, and digested with DpnII . The resulting sticky ends were biotinylated and proximity-ligated to generate chimeric junctions. Fragments of 300–500 bp were enriched and converted into paired-end Hi-C libraries, which were sequenced on an Illumina NovaSeq platform (150 bp reads). Hi-C reads were filtered using fastp (v0.23.4) 9 and subsequently employed for chromosome anchoring. Genome survey, assembly and assessment Prior to genome assembly, genome size and heterozygosity were estimated through k-mer analysis of high-quality Illumina paired-end reads. K -mers were counted using jellyfish with the mer-length of 17 10 , yielding 51,175,072,948 k -mers ( Table S4 ). These data were subsequently analyzed with the GCE (Genomic Character Estimator) tool 11 to calculate genome characteristics. Genome size (G) was estimated using the formula G = K num / K depth , where K num is the total number of 17-mers and K depth is the 17-mer depth. This analysis produced an estimated genome size of 2,105.94 Mb and a heterozygosity rate of 0.25% ( Table S2 ). We adapted a previously optimized assembly strategy 8 to generate a high-quality genome assembly for the big-footed bat, implementing a four-step pipeline. First, ONT long reads were assembled into preliminary contigs using the “correct-then-assemble” approach implemented in NextDenovo (v2.5.2; https://github.com/Nextomics/NextDenovo ) with parameters read_cutoff = 1k, seed_cutoff = 32k, and blocksize = 3g. These contigs were then polished with Illumina paired-end reads using the “best” algorithm in NextPolish v1.4.1 12 , producing the ContigV1 assembly (243 contigs, contig N50 = 24.78 Mb). Second, Hi-C reads were mapped to ContigV1 using Bowtie2 (single-ended mode) 13 . Invalid self-ligated and unligated fragments within uniquely mapped pairs were removed with the HiCUP pipeline (v0.8.0) 14 , and valid interaction pairs were used to compute linkage frequencies among contigs via an agglomerative hierarchical clustering algorithm 15 , yielding 22 linkage groups. Third, ONT long reads were remapped to ContigV1 with Minimap2 16 , and the best-aligned reads for each linkage group were extracted and locally reassembled using NextDenovo (v2.5.2) with the same parameters as in step one 17 . This local assembly step, designed to reduce false overlaps caused by repetitive sequences during string graph construction 18 , produced the ContigV2 assembly (154 contigs, contig N50 = 41.29 Mb; Table S5 ). Finally, chromosome-scale scaffolds were constructed using linkage information, restriction enzyme sites, and string graph data with HapHiC (v1.0.7) 19 , and any scaffolding or orientation errors indicated by abnormal chromatin interaction patterns were manually corrected. We evaluated the quality of the assembled genome using three complementary approaches. Assembly accuracy was estimated through a K mer–based analysis with Merqury 20 to compute the quality value (QV). Assembly completeness was assessed using BUSCO (v5.4.2) 21 against the mammalia_odb10 dataset, comprising 9,226 conserved single-copy orthologs. In addition, Illumina paired-end reads were realigned to the assembly using BWA 22 to determine the mapping rate and coverage depth, providing further validation of assembly completeness. Identification of repetitive sequences, protein-coding gene and noncoding RNA gene Repetitive elements in the big-footed bat genome were annotated using a combination of de novo prediction and homology-based searches. Candidate repeat libraries were first generated de novo with LTR_FINDER (v1.0.7) 23 and RepeatModeler (v1.0.8) 24 . These libraries, together with the Repbase database, were then used to identify repetitive sequences in the genome via RepeatMasker (v4.0.5) 25 . In addition, transposable elements (TEs) were predicted with RepeatProteinMask (v4.0.5) 25 using default parameters and the Repbase database as a reference. We utilized homologous-, de novo -, and transcriptome-based approaches to predict the protein-coding genes within the big-footed bat genome. For homology-based gene prediction, sequences from six mammalian reference genomes— Homo sapiens (GCF_000001405.39), Mus musculus (GCF_000001635.27), Myotis myotis (GCA_014108235.1) 26 , Phyllostomus discolor (GCA_004126475.3) 26 , Rhinolophus sinicus (GWHFDMV00000000.1) 27 , and Cynopterus sphinx (GWHFDMV00000000.1) 8 —were aligned to the big-footed bat genome using LASTZ (v1.04.15) 28 with parameters T = 2 (no transition), Y = 15,000 (ydrop), L = 3,000 (gappedthresh), and K = 4,500 (hspthresh). Alignment chains were generated with axtChain and processed in TOGA (v1.1.7) 29 , a machine learning–based orthology inference framework, to identify orthologous genes. The resulting six orthologous gene sets were merged to form the final homology-based prediction set. For transcriptome-based gene prediction, RNA-seq datasets from three tissues (brain, liver, and kidney) of the big-footed bat were retrieved from the Genome Sequence Archive at the National Genomics Data Center (NGDC; accession numbers CRR584602–CRR584604 and CRR620862–CRR620867) 30 . Transcript assemblies were generated using Trinity (v2.1.1) 31 and aligned to the genome with PASA 32 , which clustered effective alignments according to genomic coordinates and assembled them into gene models (PASA Trinity set). In parallel, RNA-seq reads were directly mapped to the genome using TopHat (v2.0.13) 33 , and assembled into gene models (Cufflinks set) with Cufflinks (v2.1.1) 34 . For de novo prediction, the repeat-masked genome was analyzed using Augustus (v2.5.5) 35 , GeneID (v1.4) 36 , GeneScan (v1.0) 37 , GlimmerHMM (v3.0.1) 38 , and SNAP (version 2013-11-29) 39 , with Augustus, SNAP, and GlimmerHMM trained on the PASA Trinity set. Gene models derived from transcriptome-based, homology-based, and ab initio predictions were integrated using EVidenceModeler (v1.1.1), applying the following weighting scheme: PASA Trinity set > Homology set = Cufflinks set > Augustus > GeneID = SNAP = GlimmerHMM = GeneScan. Predicted genes shorter than 50 amino acids, supported solely by ab initio evidence, and exhibiting an expression value < 1 were excluded from the final annotation. Noncoding RNA genes in the big-footed bat genome were annotated using a combination of specialized prediction tools and homology-based searches. Transfer RNAs (tRNAs) were identified with tRNAscan-SE (v1.3.1) 40 , while ribosomal RNAs (rRNAs) were detected by BLAST searches against the invertebrate rRNA database with an E-value threshold of 1e-10 41 . Small nuclear RNAs (snRNAs), small nucleolar RNAs, and microRNAs (miRNAs) were annotated using Infernal (v1.1rc4) in conjunction with the Rfam database 42 . Technical Validation The final genome size is closely aligned with the estimated result (2105.94 Mb) from K -mer analysis (Table 1 ). Our assembled genome of the big-footed bat exhibits excellent completeness, as evidenced by the coverage of 99.53% Illumina short-reads across 99.29% of the genome, and recovery of 95.45% of BUSCOs (Benchmarking Universal Single-Copy Orthologs) 21 in 9,226 conserved mammalian genes from the mammalia_odb10 database (Table 2 ). Our BUSCOs metric results surpass the average BUSCOs values of the genomes of the most recently published approximately 60 vertebrates 8 , 26 , 43 , 44 . Furthermore, we used a reference-free and k -mer based approach and estimated a high assembly quality value (QV) of 41.28 (Table 2 ), exceeding the Vertebrate Genome Project (VGP) standard of QV40 20,45 , suggesting a superior accuracy in our assembly. Table 3 Assembly quality evaluation Type Key Metrics Values Assembly completeness: Paired-end reads mapped to the assembly High-quality pairs 190,955,425 Mapping rate (%) 99.53 Depth (X) 28.43 Coverage (%) 99.29 Assembly completeness: BUSCO summary Complete (%) 95.45 Complete and single-copy (%) 93.33 Complete and duplicated (%) 2.11 Fragmented (%) 1.09 Missing (%) 3.46 Assembly accuracy: Quality value 41.28 Declarations Code availability All commands and pipelines used in data processing were executed according to the manual and protocols of the corresponding bioinformatics software. Data Records The raw data, including Illumina, ONT, and HIC sequencing data of the whole genome, was submitted to the NCBI SRA with accession number: SRR35148957-SRR35148964. The genome assembly has been deposited in the GeneBank database under the accession number: JBQVXG000000000. Acknowledgments This study was supported in the National Natural Science Foundation of China (32471689) to S.T. H.Z. was supported in part by the Hubei Provincial Natural Science Foundation (2023AFA015) and Fundamental Research Funds for the Central Universities (2042022dx0003). Author contributions H.Z. conceived and designed research. S.T. designed and performed analyses. L.Z collected the samples. S.T. and H.Z. discussed results and wrote the manuscript. All authors have read and approved the paper. Competing Interest Statement All authors declare no conflict of interest. References Teeling, E.C. et al. Bat Biology, Genomes, and the Bat1K Project: To Generate Chromosome-Level Genomes for All Living Bat Species. Annu Rev Anim Biosci 6 , 23-46 (2018). Munshi-South, J. & Wilkinson, G.S. Bats and birds: Exceptional longevity despite high metabolic rates. Ageing Res Rev 9 , 12-9 (2010). Healy, K. et al. Ecology and mode-of-life explain lifespan variation in birds and mammals. Proceedings of the Royal Society B: Biological Sciences 281 , 20140298 (2014). Podlutsky, A.J., Khritankov, A.M., Ovodov, N.D. & Austad, S.N. A new field record for bat longevity. The Journals of Gerontology Series A: Biological Sciences and Medical Sciences 60 , 1366-1368 (2005). Seim, I. et al. Genome analysis reveals insights into physiology and longevity of the Brandt's bat Myotis brandtii. Nat Commun 4 , 2212 (2013). Kim, E.B. et al. Genome sequencing reveals insights into physiology and longevity of the naked mole rat. Nature 479 , 223-227 (2011). Budovsky, A. et al. LongevityMap: a database of human genetic variants associated with longevity. Trends in Genetics 29 , 559-560 (2013). Tian, S. et al. Comparative analyses of bat genomes identify distinct evolution of immunity in Old World fruit bats. Sci Adv 9 , eadd0141 (2023). Chen, S., Zhou, Y., Chen, Y. & Gu, J. fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics 34 , i884-i890 (2018). Marcais, G. & Kingsford, C. A fast, lock-free approach for efficient parallel counting of occurrences of k-mers. Bioinformatics 27 , 764-70 (2011). Liu, B. et al. Estimation of genomic characteristics by analyzing k-mer frequency in de novo genome projects. arXiv preprint arXiv:1308.2012 (2013). Hu, J., Fan, J., Sun, Z. & Liu, S. NextPolish: a fast and efficient genome polishing tool for long-read assembly. Bioinformatics 36 , 2253-2255 (2020). Langmead, B. & Salzberg, S.L. Fast gapped-read alignment with Bowtie 2. Nat Methods 9 , 357-9 (2012). Wingett, S. et al. HiCUP: pipeline for mapping and processing Hi-C data. F1000Res 4 , 1310 (2015). Li, D. et al. Population genomics identifies patterns of genetic diversity and selection in chicken. BMC Genomics 20 , 263 (2019). Li, H. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics 34 , 3094-3100 (2018). Cheng, H., Concepcion, G.T., Feng, X., Zhang, H. & Li, H. Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm. Nature Methods 18 , 170-175 (2021). Myers, E.W. The fragment assembly string graph. Bioinformatics 21 Suppl 2 , ii79-85 (2005). Zeng, X. et al. Chromosome-level scaffolding of haplotype-resolved assemblies using Hi-C data without reference genomes. Nature plants 10 , 1184-1200 (2024). Rhie, A., Walenz, B.P., Koren, S. & Phillippy, A.M. Merqury: reference-free quality, completeness, and phasing assessment for genome assemblies. Genome Biol 21 , 245 (2020). Simao, F.A., Waterhouse, R.M., Ioannidis, P., Kriventseva, E.V. & Zdobnov, E.M. BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs. Bioinformatics 31 , 3210-2 (2015). Li, H. & Durbin, R. Fast and accurate long-read alignment with Burrows-Wheeler transform. Bioinformatics 26 , 589-95 (2010). Xu, Z. & Wang, H. LTR_FINDER: an efficient tool for the prediction of full-length LTR retrotransposons. Nucleic Acids Res 35 , W265-8 (2007). Smit, A. & Hubley, R.R. Open-1.0. Available from http://www.repeatmasker.org (2008). Smit, A., Hubley, R. & Green, P. RepeatMasker Open-4.0. 2013–2015. (2015). Jebb, D. et al. Six reference-quality genomes reveal evolution of bat adaptations. Nature 583 , 578-584 (2020). Tian, S. et al. Comparative genomics provides insights into chromosomal evolution and immunological adaptation in horseshoe bats. Nat Ecol Evol 9 , 705-720 (2025). Harris, R.S. Improved pairwise Alignmnet of genomic DNA. (2007). Kirilenko, B.M. et al. Integrating gene annotation with orthology inference at scale. Science 380 , eabn3107 (2023). Liu, W. et al. Large-scale across species transcriptomic analysis identifies genetic selection signatures associated with longevity in mammals. EMBO J 42 , e112740 (2023). Grabherr, M.G. et al. Full-length transcriptome assembly from RNA-Seq data without a reference genome. Nat Biotechnol 29 , 644-52 (2011). Haas, B.J. et al. Improving the Arabidopsis genome annotation using maximal transcript alignment assemblies. Nucleic Acids Res 31 , 5654-5666 (2003). Kim, D. et al. TopHat2: accurate alignment of transcriptomes in the presence of insertions, deletions and gene fusions. Genome Biol 14 , R36 (2013). Trapnell, C. et al. Differential gene and transcript expression analysis of RNA-seq experiments with TopHat and Cufflinks. Nat Protoc 7 , 562-78 (2012). Stanke, M. & Waack, S. Gene prediction with a hidden Markov model and a new intron submodel. Bioinformatics 19 Suppl 2 , ii215-25 (2003). Guigo, R. Assembling genes from predicted exons in linear time with dynamic programming. J Comput Biol 5 , 681-702 (1998). Burge, C. & Karlin, S. Prediction of complete gene structures in human genomic DNA. J Mol Biol 268 , 78-94 (1997). Majoros, W.H., Pertea, M. & Salzberg, S.L. TigrScan and GlimmerHMM: two open source ab initio eukaryotic gene-finders. Bioinformatics 20 , 2878-9 (2004). Korf, I. Gene finding in novel genomes. BMC Bioinformatics 5 , 59 (2004). Schattner, P., Brooks, A.N. & Lowe, T.M. The tRNAscan-SE, snoscan and snoGPS web servers for the detection of tRNAs and snoRNAs. Nucleic Acids Res 33 , W686-9 (2005). Altschul, S.F., Gish, W., Miller, W., Myers, E.W. & Lipman, D.J. Basic local alignment search tool. J Mol Biol 215 , 403-10 (1990). Nawrocki, E.P. & Eddy, S.R. Infernal 1.1: 100-fold faster RNA homology searches. Bioinformatics 29 , 2933-5 (2013). Shao, Y. et al. Phylogenomic analyses provide insights into primate evolution. Science 380 , 913-924 (2023). Peng, C. et al. Large-scale snake genome analyses provide insights into vertebrate development. Cell 186 , 2959-2976 e22 (2023). Editorial, N.B. A reference standard for genome biology. Nat Biotechnol 36 , 1121 (2018). Additional Declarations The authors declare no competing interests. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7561642","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Data Note","associatedPublications":[],"authors":[{"id":511751263,"identity":"ad0184fb-746e-418e-a846-fc5bb0802455","order_by":0,"name":"Shilin Tian","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA10lEQVRIiWNgGAWjYDACZiBmbJCQYWNgPsAMFjlAQAcPVAsPGwNbApFaGMBaQDSPAXFa7NmZnz38ucOCh0+655t0YRuDHN+NBMbPBXgdxmZuzHsG6DCZs9ukZ7YxGEveSGCWnoHfL2bSjG1ALRK5227ztjEkbriRwMbMg1cL+zfJn2AtOc9AWuqJ0MJjJsEL0cIG0pJgQFDLYZ4yaYiWNPPfPOckDGeeedgsjU8Le//xbUCH1cnJz0h+bMxTZiPPdzz54Gd8WtCBBAM4mkbBKBgFo2AUUAYAOmQ75UiI23cAAAAASUVORK5CYII=","orcid":"https://orcid.org/0000-0001-8958-1806","institution":"State Key Laboratory of Virology and Biosafety, Key Laboratory of Biodiversity and Environment on the Qinghai-Tibetan Plateau, Ministry of Education, Frontier Science Center for Immunology and Metabolism, Hubei Key Laboratory of Cell Homeostasis, College of Life Sciences, Wuhan University, Wuhan, 430072, China","correspondingAuthor":true,"prefix":"","firstName":"Shilin","middleName":"","lastName":"Tian","suffix":""},{"id":511751265,"identity":"80ebe3b6-190b-46f8-991e-111f84651f8e","order_by":1,"name":"Libiao Zhang","email":"","orcid":"","institution":"Guangdong Key Laboratory of Animal Conservation and Resource Utilization, Guangdong, Public Laboratory of Wild Animal Conservation and Utilization, Institute of Zoology, Guangdong Academy of Sciences, Guangzhou, China.","correspondingAuthor":false,"prefix":"","firstName":"Libiao","middleName":"","lastName":"Zhang","suffix":""},{"id":511751268,"identity":"9a3e3d38-3db6-4ca4-98fe-b22860f4f5c6","order_by":2,"name":"Huabin Zhao","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA4ElEQVRIiWNgGAWjYBACAwYGNoYEBhBiPsAMFjpAvBa2BBK0MIC18BgQp8VcIvnZg4c76vL423u+SRe2Mcjx3Uhg/FyAR4vljDRzg8QzbMUSZ85uk57ZxmAseSOBWXoGPofdSDCTSGzjSWy4kbvtNm8bQ+KGGwlszDx4taR/A2qRSJx/I+cZSEs9EVpyQLYYAA3PYQNpSTAgqOXMm3KDxLaExI1njpn/5jknYTjzzMNmabxajqdve/izrS5x3vHmx8Y8ZTbyfMeTD37GpwUdSAAxYwMJGkbBKBgFo2AUYAMA/y1PHewD0f8AAAAASUVORK5CYII=","orcid":"https://orcid.org/0000-0002-7848-6392","institution":"State Key Laboratory of Virology and Biosafety, Key Laboratory of Biodiversity and Environment on the Qinghai-Tibetan Plateau, Ministry of Education, Frontier Science Center for Immunology and Metabolism, Hubei Key Laboratory of Cell Homeostasis, College of Life Sciences, Wuhan University, Wuhan, 430072, China","correspondingAuthor":true,"prefix":"","firstName":"Huabin","middleName":"","lastName":"Zhao","suffix":""}],"badges":[],"createdAt":"2025-09-08 08:20:47","currentVersionCode":1,"declarations":{"humanSubjects":false,"vertebrateSubjects":true,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":false,"humanSubjectConsent":false,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":true},"doi":"10.21203/rs.3.rs-7561642/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7561642/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":90925582,"identity":"c94fbec7-a612-46c7-868e-3527a2ba4cea","added_by":"auto","created_at":"2025-09-09 15:28:49","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":257980,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eChromatin interactions in each chromosome of the big-footed bat genome at a resolution of 500 Kb.\u003c/strong\u003e The dark\u003cstrong\u003e \u003c/strong\u003ered dots show the high probability of interaction, and the light dots show the low probability of interaction.\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-7561642/v1/acffa8eb7cfafd13e50d1e0a.png"},{"id":90927311,"identity":"64e93a1e-3f1c-4c10-b2f5-de5f399a1e9e","added_by":"auto","created_at":"2025-09-09 15:44:49","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":899520,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7561642/v1/007a9241-a87b-4b62-887d-92cdb5747a0b.pdf"}],"financialInterests":"The authors declare no competing interests.","formattedTitle":"\u003cp\u003e\u003cstrong\u003eChromosome-level genome assembly of the big-footed bat (\u003c/strong\u003e\u003cem\u003e\u003cstrong\u003eMyotis pilosus\u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003e)\u003c/strong\u003e\u003c/p\u003e","fulltext":[{"header":"Background \u0026 Summary","content":"\u003cp\u003eBats (Chiroptera) comprise over 1,400 extant species, accounting for approximately 20% of all living mammals. They have successfully evolved a series of unique adaptations, including powered flight, echolocation, a distinctive immune system, and exceptional longevity\u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u003c/sup\u003e. Most bat species live more than three times longer than other mammals of similar body size\u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e,\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e. Among all bats, species of the genus \u003cem\u003eMyotis\u003c/em\u003e exhibit notably long lifespans\u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e, with at least 13 \u003cem\u003eMyotis\u003c/em\u003e species documented to survive at least 20 years in the wild, including the longest-lived recaptured bat to date\u0026mdash;the Brandt\u0026rsquo;s bat (\u003cem\u003eMyotis brandtii\u003c/em\u003e), which has a lifespan exceeding 41 years.\u003c/p\u003e\u003cp\u003eMolecular and genomic studies have increasingly provided evidence for the longevity mechanisms in bats. Comparative Genomic analyses of \u003cem\u003eM. brandtii\u003c/em\u003e indicate that an altered growth hormone/insulin-like growth factor 1 (GH/IGF1) axis, together with adaptations such as hibernation and low reproductive rate, collectively contribute to the exceptional lifespan of Brandt\u0026rsquo;s bat\u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e. Long-lived mammals typically exhibit pronounced cancer resistance; in the long-lived bat species \u003cem\u003eMyotis myotis\u003c/em\u003e, oncogenic microRNAs are downregulated, whereas cancer-associated microRNAs are upregulated\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e. Recent studies have also revealed that another \u003cem\u003eMyotis\u003c/em\u003e species, \u003cem\u003eM. pilosus\u003c/em\u003e, is potentially long-lived, as its primary fibroblast cell lines exhibit significantly higher β-galactosidase activity\u0026mdash;a marker of cellular senescence\u0026mdash;compared to the known long-lived bat \u003cem\u003eRhinolophus ferrumequinum\u003c/em\u003e (\u0026gt;\u0026thinsp;30.5 years, data from the AnAge Database\u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e). Further transcriptomic and functional analyses indicate that the downregulation of three genes (\u003cem\u003eHIF1A\u003c/em\u003e, \u003cem\u003eCOPS5\u003c/em\u003e, and \u003cem\u003eRPS3\u003c/em\u003e) plays a key role in the cancer resistance of \u003cem\u003eM. pilosus\u003c/em\u003e.\u003c/p\u003e\u003cp\u003eIn this study, we generated a de novo assembly of the female big-footed bat genome by integrating 158.49 Gb of Oxford Nanopore Technologies (ONT) long reads, 212.83 Gb of Hi-C data, and 60.74 Gb of Illumina paired-end reads (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). Using an improved four-step assembly strategy\u003csup\u003e\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e, we obtained a 1968.27 Mb genome with a contig N50 of 41.29 Mb (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e), anchored to 22 chromosomes (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). We successfully identified the X chromosome, which spans 127.95 Mb and exhibits a high degree of conservation. The genome contains 739.02 Mb of repetitive sequences (37.55% of the genome) and encodes 21,368 protein-coding genes, 10,757 microRNAs, 166,812 tRNAs, 648 rRNAs, and 3,317 snRNAs. Quality assessment demonstrates that this assembly is comparable to previously published bat genomes in terms of continuity, completeness, and accuracy (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). As the first chromosome-level genome of the big-footed bat, it provides a valuable resource for exploring the genetic basis of longevity and cancer resistance in the genus \u003cem\u003eMyotis\u003c/em\u003e.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eSummary of paired-end data, Hi-C data and long-read sequencing data\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"6\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"3\" rowspan=\"4\"\u003e\u003cp\u003eIllumina sequences\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003eLibrary\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003eRaw data\u003c/p\u003e\u003cp\u003e(Gb)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"3\" nameend=\"c6\" namest=\"c4\"\u003e\u003cp\u003eHigh Quality Data\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eBase (Gb)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eQ20 (%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eQ30 (%)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003ePaired-end\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e60.74\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e57.29\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e97.01\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e91.91\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eHi-C\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e212.83\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e176.71\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e97.15\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e92.18\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003eLong-read sequences\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003ePlatform\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eTotal read\u003c/p\u003e\u003cp\u003ebases (Gb)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eRead number\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eRead N50\u003c/p\u003e\u003cp\u003e(bp)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eMean read\u003c/p\u003e\u003cp\u003elength (bp)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eONT\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e158.49\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e6,478,552\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e32,339\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e24,463\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eSummary of genome survey, assembly and assessment\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"3\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eType\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eKey Metrics\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eValues\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"3\" rowspan=\"4\"\u003e\u003cp\u003eGenome Survey\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eMer-length\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e17\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eNumber of \u003cem\u003eK\u003c/em\u003emers\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e51,175,072,948\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eEstimated genome size (Mb)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e2,105.94\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eHeterozygous rate (%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.25\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"4\" rowspan=\"5\"\u003e\u003cp\u003eGenome Assembly\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eAssembly length (Mb)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e1,968.27\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eContig Number\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e154\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eContig N50 length (Mb)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e41.29\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eAnchored onto chromosomes (Mb)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e1,966.68\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eScaffold N50 length (Mb)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e107.20\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"7\" rowspan=\"8\"\u003e\u003cp\u003eRepetitive sequences\u003c/p\u003e\u003cp\u003e(Length / Percentage)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eDNA transposon\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e90.07 Mb / 4.58%\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eLINE\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e584.93 Mb / 29.72%\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eSINE\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e9.22 Mb / 0.47%\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eLTR retrotransposon\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e98.39 Mb / 5.00%\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eSatellite\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e1.08 Mb / 0.06%\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eSimple repeat\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e2.11 Mb / 0.11%\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eOthers\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e12.11 Mb / 0.62%\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eTotal\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e739.02 Mb / 37.55%\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"5\" rowspan=\"6\"\u003e\u003cp\u003eProtein-coding\u003c/p\u003e\u003cp\u003eGenes\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eNumber\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e21,368\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eAverage gene length (bp)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e32,130.11\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eAverage CDS length per gene (bp)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e1,496.34\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eAverage exon number per gene\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e8.44\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eAverage exon length (bp)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e177.37\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eAverage intron length (bp)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e3,738.94\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"3\" rowspan=\"4\"\u003e\u003cp\u003encRNAs\u003c/p\u003e\u003cp\u003e(Copy / Length)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003emicroRNAs (miRNA)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e10,757 / 1.26 Mb\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003etransfer RNA (tRNA)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e166,812 / 12.14 Mb\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eribosomal RNAs (rRNA)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e648 / 0.11 Mb\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003esmall nuclear RNAs (snRNA)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e3,317 / 0.37 Mb\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\u003ch2\u003eSampling and sequencing\u003c/h2\u003e\u003cp\u003eAn adult female big-footed bat was collected from Guangzhou City, Guangdong Province, China. The specimen was captured using mist nets at cave sites, placed in a clean cloth bag, and transported to a temporary laboratory for processing. Field sampling was approved by the Guangdong Institute of Applied Biological Resources, Institute of Zoology, Guangdong Academy of Sciences (approval number: GIABR20200810). The sampling activities did not require any additional specific permissions and did not involve any endangered or protected species.\u003c/p\u003e\u003cp\u003eDNA was isolated from liver tissue using Qiagen Genomic DNA extraction kits and was tested for quality before use in genomic libraries. Short-insert libraries (~ 350 bp) were prepared from ~ 1.5 µg of genomic DNA using the TruSeq Nano DNA HT Sample Preparation Kit (Illumina) and were sequenced on an Illumina NovaSeq platform to generate 150 bp paired-end reads. Raw reads were quality-filtered using \u003cem\u003efastp\u003c/em\u003e (v0.23.4)\u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e. For long-read sequencing, genomic DNA was treated with the NEBNext FFPE DNA Repair Mix (M6630) and the NEBNext End Repair/dA-Tailing Module (E7546) according to the manufacturer’s instructions. Libraries were sequenced on a PromethION platform (Oxford Nanopore Technologies), and reads with Phred quality scores below 7 were removed prior to assembly. Hi-C libraries were constructed from liver cells to assist chromosome-level scaffolding. Cells were cross-linked with formaldehyde, lysed, and digested with \u003cem\u003eDpnII\u003c/em\u003e. The resulting sticky ends were biotinylated and proximity-ligated to generate chimeric junctions. Fragments of 300–500 bp were enriched and converted into paired-end Hi-C libraries, which were sequenced on an Illumina NovaSeq platform (150 bp reads). Hi-C reads were filtered using \u003cem\u003efastp\u003c/em\u003e (v0.23.4)\u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e and subsequently employed for chromosome anchoring.\u003c/p\u003e\u003c/div\u003e\n\u003cp\u003e\u003c/p\u003e\n\u003ch3\u003eGenome survey, assembly and assessment\u003c/h3\u003e\n\u003cp\u003ePrior to genome assembly, genome size and heterozygosity were estimated through k-mer analysis of high-quality Illumina paired-end reads. \u003cem\u003eK\u003c/em\u003e-mers were counted using \u003cem\u003ejellyfish\u003c/em\u003e with the mer-length of 17\u003csup\u003e10\u003c/sup\u003e, yielding 51,175,072,948 \u003cem\u003ek\u003c/em\u003e-mers (\u003cb\u003eTable S4\u003c/b\u003e). These data were subsequently analyzed with the GCE (Genomic Character Estimator) tool\u003csup\u003e\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u003c/sup\u003e to calculate genome characteristics. Genome size (G) was estimated using the formula G = \u003cem\u003eK\u003c/em\u003e\u003csub\u003enum\u003c/sub\u003e/\u003cem\u003eK\u003c/em\u003e\u003csub\u003edepth\u003c/sub\u003e, where \u003cem\u003eK\u003c/em\u003e\u003csub\u003enum\u003c/sub\u003e is the total number of 17-mers and \u003cem\u003eK\u003c/em\u003e\u003csub\u003edepth\u003c/sub\u003e is the 17-mer depth. This analysis produced an estimated genome size of 2,105.94 Mb and a heterozygosity rate of 0.25% (\u003cb\u003eTable S2\u003c/b\u003e).\u003c/p\u003e\u003cp\u003eWe adapted a previously optimized assembly strategy\u003csup\u003e\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e to generate a high-quality genome assembly for the big-footed bat, implementing a four-step pipeline. First, ONT long reads were assembled into preliminary contigs using the “correct-then-assemble” approach implemented in NextDenovo (v2.5.2; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/Nextomics/NextDenovo\u003c/span\u003e\u003cspan address=\"https://github.com/Nextomics/NextDenovo\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) with parameters read_cutoff = 1k, seed_cutoff = 32k, and blocksize = 3g. These contigs were then polished with Illumina paired-end reads using the “best” algorithm in NextPolish v1.4.1\u003csup\u003e12\u003c/sup\u003e, producing the ContigV1 assembly (243 contigs, contig N50 = 24.78 Mb). Second, Hi-C reads were mapped to ContigV1 using Bowtie2 (single-ended mode)\u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e. Invalid self-ligated and unligated fragments within uniquely mapped pairs were removed with the HiCUP pipeline (v0.8.0)\u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e, and valid interaction pairs were used to compute linkage frequencies among contigs via an agglomerative hierarchical clustering algorithm\u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e, yielding 22 linkage groups. Third, ONT long reads were remapped to ContigV1 with Minimap2\u003csup\u003e16\u003c/sup\u003e, and the best-aligned reads for each linkage group were extracted and locally reassembled using NextDenovo (v2.5.2) with the same parameters as in step one\u003csup\u003e\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e. This local assembly step, designed to reduce false overlaps caused by repetitive sequences during string graph construction\u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e, produced the ContigV2 assembly (154 contigs, contig N50 = 41.29 Mb; \u003cb\u003eTable S5\u003c/b\u003e). Finally, chromosome-scale scaffolds were constructed using linkage information, restriction enzyme sites, and string graph data with HapHiC (v1.0.7)\u003csup\u003e\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e, and any scaffolding or orientation errors indicated by abnormal chromatin interaction patterns were manually corrected.\u003c/p\u003e\u003cp\u003eWe evaluated the quality of the assembled genome using three complementary approaches. Assembly accuracy was estimated through a \u003cem\u003eK\u003c/em\u003emer–based analysis with Merqury\u003csup\u003e\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e to compute the quality value (QV). Assembly completeness was assessed using BUSCO (v5.4.2)\u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e against the mammalia_odb10 dataset, comprising 9,226 conserved single-copy orthologs. In addition, Illumina paired-end reads were realigned to the assembly using BWA\u003csup\u003e\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e to determine the mapping rate and coverage depth, providing further validation of assembly completeness.\u003c/p\u003e\n\u003ch3\u003eIdentification of repetitive sequences, protein-coding gene and noncoding RNA gene\u003c/h3\u003e\n\u003cp\u003eRepetitive elements in the big-footed bat genome were annotated using a combination of \u003cem\u003ede novo\u003c/em\u003e prediction and homology-based searches. Candidate repeat libraries were first generated \u003cem\u003ede novo\u003c/em\u003e with LTR_FINDER (v1.0.7)\u003csup\u003e\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e\u003c/sup\u003e and RepeatModeler (v1.0.8)\u003csup\u003e\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/sup\u003e. These libraries, together with the Repbase database, were then used to identify repetitive sequences in the genome via RepeatMasker (v4.0.5)\u003csup\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u003c/sup\u003e. In addition, transposable elements (TEs) were predicted with RepeatProteinMask (v4.0.5)\u003csup\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u003c/sup\u003e using default parameters and the Repbase database as a reference.\u003c/p\u003e\u003cp\u003eWe utilized homologous-, \u003cem\u003ede novo\u003c/em\u003e-, and transcriptome-based approaches to predict the protein-coding genes within the big-footed bat genome. For homology-based gene prediction, sequences from six mammalian reference genomes—\u003cem\u003eHomo sapiens\u003c/em\u003e (GCF_000001405.39), \u003cem\u003eMus musculus\u003c/em\u003e (GCF_000001635.27), \u003cem\u003eMyotis myotis\u003c/em\u003e (GCA_014108235.1)\u003csup\u003e\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e\u003c/sup\u003e, \u003cem\u003ePhyllostomus discolor\u003c/em\u003e (GCA_004126475.3)\u003csup\u003e\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e\u003c/sup\u003e, \u003cem\u003eRhinolophus sinicus\u003c/em\u003e (GWHFDMV00000000.1)\u003csup\u003e\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e\u003c/sup\u003e, and \u003cem\u003eCynopterus sphinx\u003c/em\u003e (GWHFDMV00000000.1)\u003csup\u003e\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e—were aligned to the big-footed bat genome using LASTZ (v1.04.15)\u003csup\u003e\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e\u003c/sup\u003e with parameters T = 2 (no transition), Y = 15,000 (ydrop), L = 3,000 (gappedthresh), and K = 4,500 (hspthresh). Alignment chains were generated with axtChain and processed in TOGA (v1.1.7)\u003csup\u003e\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e\u003c/sup\u003e, a machine learning–based orthology inference framework, to identify orthologous genes. The resulting six orthologous gene sets were merged to form the final homology-based prediction set.\u003c/p\u003e\u003cp\u003eFor transcriptome-based gene prediction, RNA-seq datasets from three tissues (brain, liver, and kidney) of the big-footed bat were retrieved from the Genome Sequence Archive at the National Genomics Data Center (NGDC; accession numbers CRR584602–CRR584604 and CRR620862–CRR620867)\u003csup\u003e\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e. Transcript assemblies were generated using Trinity (v2.1.1)\u003csup\u003e\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e\u003c/sup\u003e and aligned to the genome with PASA\u003csup\u003e\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e\u003c/sup\u003e, which clustered effective alignments according to genomic coordinates and assembled them into gene models (PASA Trinity set). In parallel, RNA-seq reads were directly mapped to the genome using TopHat (v2.0.13)\u003csup\u003e\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e\u003c/sup\u003e, and assembled into gene models (Cufflinks set) with Cufflinks (v2.1.1)\u003csup\u003e\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003eFor \u003cem\u003ede novo\u003c/em\u003e prediction, the repeat-masked genome was analyzed using Augustus (v2.5.5)\u003csup\u003e\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e\u003c/sup\u003e, GeneID (v1.4)\u003csup\u003e\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e\u003c/sup\u003e, GeneScan (v1.0)\u003csup\u003e\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e\u003c/sup\u003e, GlimmerHMM (v3.0.1)\u003csup\u003e\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e\u003c/sup\u003e, and SNAP (version 2013-11-29)\u003csup\u003e\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e\u003c/sup\u003e, with Augustus, SNAP, and GlimmerHMM trained on the PASA Trinity set. Gene models derived from transcriptome-based, homology-based, and ab initio predictions were integrated using EVidenceModeler (v1.1.1), applying the following weighting scheme: PASA Trinity set \u0026gt; Homology set = Cufflinks set \u0026gt; Augustus \u0026gt; GeneID = SNAP = GlimmerHMM = GeneScan. Predicted genes shorter than 50 amino acids, supported solely by ab initio evidence, and exhibiting an expression value \u0026lt; 1 were excluded from the final annotation.\u003c/p\u003e\u003cp\u003eNoncoding RNA genes in the big-footed bat genome were annotated using a combination of specialized prediction tools and homology-based searches. Transfer RNAs (tRNAs) were identified with tRNAscan-SE (v1.3.1)\u003csup\u003e\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e\u003c/sup\u003e, while ribosomal RNAs (rRNAs) were detected by BLAST searches against the invertebrate rRNA database with an E-value threshold of 1e-10\u003csup\u003e41\u003c/sup\u003e. Small nuclear RNAs (snRNAs), small nucleolar RNAs, and microRNAs (miRNAs) were annotated using Infernal (v1.1rc4) in conjunction with the Rfam database\u003csup\u003e\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e"},{"header":"Technical Validation","content":"\u003cp\u003eThe final genome size is closely aligned with the estimated result (2105.94 Mb) from \u003cem\u003eK\u003c/em\u003e-mer analysis (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). Our assembled genome of the big-footed bat exhibits excellent completeness, as evidenced by the coverage of 99.53% Illumina short-reads across 99.29% of the genome, and recovery of 95.45% of BUSCOs (Benchmarking Universal Single-Copy Orthologs)\u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e in 9,226 conserved mammalian genes from the mammalia_odb10 database (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). Our BUSCOs metric results surpass the average BUSCOs values of the genomes of the most recently published approximately 60 vertebrates\u003csup\u003e\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e,\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e,\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e,\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e\u003c/sup\u003e. Furthermore, we used a reference-free and \u003cem\u003ek\u003c/em\u003e-mer based approach and estimated a high assembly quality value (QV) of 41.28 (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e), exceeding the Vertebrate Genome Project (VGP) standard of QV40\u003csup\u003e20,45\u003c/sup\u003e, suggesting a superior accuracy in our assembly.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cdiv class=\"gridtable\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eAssembly quality evaluation\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"3\"\u003e\u003c/colgroup\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eType\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eKey Metrics\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eValues\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"3\" rowspan=\"4\"\u003e\u003cp\u003eAssembly completeness: Paired-end reads mapped to the assembly\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eHigh-quality pairs\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e190,955,425\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eMapping rate (%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e99.53\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eDepth (X)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e28.43\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eCoverage (%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e99.29\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"4\" rowspan=\"5\"\u003e\u003cp\u003eAssembly completeness: BUSCO summary\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eComplete (%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e95.45\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eComplete and single-copy (%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e93.33\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eComplete and duplicated (%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e2.11\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eFragmented (%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e1.09\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eMissing (%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e3.46\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eAssembly accuracy: Quality value\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e\u003cp\u003e41.28\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/table\u003e\u003c/div\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eCode\u0026nbsp;availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll commands and pipelines used in data processing were executed according to the manual and protocols of the corresponding bioinformatics software.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData Records\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe raw data, including Illumina, ONT, and HIC sequencing data of the whole genome, was submitted to the NCBI SRA with accession number: SRR35148957-SRR35148964. The genome assembly has been deposited in the GeneBank database under the accession number: JBQVXG000000000.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgments\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study was supported in the National Natural Science Foundation of China (32471689) to S.T. H.Z. was supported in part by the Hubei Provincial Natural Science Foundation (2023AFA015) and Fundamental Research Funds for the Central Universities (2042022dx0003).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eH.Z. conceived and designed research. S.T. designed and performed analyses. L.Z collected the samples. S.T. and H.Z. discussed results and wrote the manuscript. All authors have read and approved the paper.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting Interest Statement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll authors declare no conflict of interest.\u0026nbsp;\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eTeeling, E.C.\u003cem\u003e et al.\u003c/em\u003e Bat Biology, Genomes, and the Bat1K Project: To Generate Chromosome-Level Genomes for All Living Bat Species. \u003cem\u003eAnnu Rev Anim Biosci\u003c/em\u003e \u003cstrong\u003e6\u003c/strong\u003e, 23-46 (2018).\u003c/li\u003e\n\u003cli\u003eMunshi-South, J. \u0026amp; Wilkinson, G.S. Bats and birds: Exceptional longevity despite high metabolic rates. \u003cem\u003eAgeing Res Rev\u003c/em\u003e \u003cstrong\u003e9\u003c/strong\u003e, 12-9 (2010).\u003c/li\u003e\n\u003cli\u003eHealy, K.\u003cem\u003e et al.\u003c/em\u003e Ecology and mode-of-life explain lifespan variation in birds and mammals. \u003cem\u003eProceedings of the Royal Society B: Biological Sciences\u003c/em\u003e \u003cstrong\u003e281\u003c/strong\u003e, 20140298 (2014).\u003c/li\u003e\n\u003cli\u003ePodlutsky, A.J., Khritankov, A.M., Ovodov, N.D. \u0026amp; Austad, S.N. A new field record for bat longevity. \u003cem\u003eThe Journals of Gerontology Series A: Biological Sciences and Medical Sciences\u003c/em\u003e \u003cstrong\u003e60\u003c/strong\u003e, 1366-1368 (2005).\u003c/li\u003e\n\u003cli\u003eSeim, I.\u003cem\u003e et al.\u003c/em\u003e Genome analysis reveals insights into physiology and longevity of the Brandt\u0026apos;s bat Myotis brandtii. \u003cem\u003eNat Commun\u003c/em\u003e \u003cstrong\u003e4\u003c/strong\u003e, 2212 (2013).\u003c/li\u003e\n\u003cli\u003eKim, E.B.\u003cem\u003e et al.\u003c/em\u003e Genome sequencing reveals insights into physiology and longevity of the naked mole rat. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e479\u003c/strong\u003e, 223-227 (2011).\u003c/li\u003e\n\u003cli\u003eBudovsky, A.\u003cem\u003e et al.\u003c/em\u003e LongevityMap: a database of human genetic variants associated with longevity. \u003cem\u003eTrends in Genetics\u003c/em\u003e \u003cstrong\u003e29\u003c/strong\u003e, 559-560 (2013).\u003c/li\u003e\n\u003cli\u003eTian, S.\u003cem\u003e et al.\u003c/em\u003e Comparative analyses of bat genomes identify distinct evolution of immunity in Old World fruit bats. \u003cem\u003eSci Adv\u003c/em\u003e \u003cstrong\u003e9\u003c/strong\u003e, eadd0141 (2023).\u003c/li\u003e\n\u003cli\u003eChen, S., Zhou, Y., Chen, Y. \u0026amp; Gu, J. fastp: an ultra-fast all-in-one FASTQ preprocessor. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cstrong\u003e34\u003c/strong\u003e, i884-i890 (2018).\u003c/li\u003e\n\u003cli\u003eMarcais, G. \u0026amp; Kingsford, C. A fast, lock-free approach for efficient parallel counting of occurrences of k-mers. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cstrong\u003e27\u003c/strong\u003e, 764-70 (2011).\u003c/li\u003e\n\u003cli\u003eLiu, B.\u003cem\u003e et al.\u003c/em\u003e Estimation of genomic characteristics by analyzing k-mer frequency in de novo genome projects. \u003cem\u003earXiv preprint arXiv:1308.2012\u003c/em\u003e (2013).\u003c/li\u003e\n\u003cli\u003eHu, J., Fan, J., Sun, Z. \u0026amp; Liu, S. NextPolish: a fast and efficient genome polishing tool for long-read assembly. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cstrong\u003e36\u003c/strong\u003e, 2253-2255 (2020).\u003c/li\u003e\n\u003cli\u003eLangmead, B. \u0026amp; Salzberg, S.L. Fast gapped-read alignment with Bowtie 2. \u003cem\u003eNat Methods\u003c/em\u003e \u003cstrong\u003e9\u003c/strong\u003e, 357-9 (2012).\u003c/li\u003e\n\u003cli\u003eWingett, S.\u003cem\u003e et al.\u003c/em\u003e HiCUP: pipeline for mapping and processing Hi-C data. \u003cem\u003eF1000Res\u003c/em\u003e \u003cstrong\u003e4\u003c/strong\u003e, 1310 (2015).\u003c/li\u003e\n\u003cli\u003eLi, D.\u003cem\u003e et al.\u003c/em\u003e Population genomics identifies patterns of genetic diversity and selection in chicken. \u003cem\u003eBMC Genomics\u003c/em\u003e \u003cstrong\u003e20\u003c/strong\u003e, 263 (2019).\u003c/li\u003e\n\u003cli\u003eLi, H. Minimap2: pairwise alignment for nucleotide sequences. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cstrong\u003e34\u003c/strong\u003e, 3094-3100 (2018).\u003c/li\u003e\n\u003cli\u003eCheng, H., Concepcion, G.T., Feng, X., Zhang, H. \u0026amp; Li, H. Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm. \u003cem\u003eNature Methods\u003c/em\u003e \u003cstrong\u003e18\u003c/strong\u003e, 170-175 (2021).\u003c/li\u003e\n\u003cli\u003eMyers, E.W. The fragment assembly string graph. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cstrong\u003e21 Suppl 2\u003c/strong\u003e, ii79-85 (2005).\u003c/li\u003e\n\u003cli\u003eZeng, X.\u003cem\u003e et al.\u003c/em\u003e Chromosome-level scaffolding of haplotype-resolved assemblies using Hi-C data without reference genomes. \u003cem\u003eNature plants\u003c/em\u003e \u003cstrong\u003e10\u003c/strong\u003e, 1184-1200 (2024).\u003c/li\u003e\n\u003cli\u003eRhie, A., Walenz, B.P., Koren, S. \u0026amp; Phillippy, A.M. Merqury: reference-free quality, completeness, and phasing assessment for genome assemblies. \u003cem\u003eGenome Biol\u003c/em\u003e \u003cstrong\u003e21\u003c/strong\u003e, 245 (2020).\u003c/li\u003e\n\u003cli\u003eSimao, F.A., Waterhouse, R.M., Ioannidis, P., Kriventseva, E.V. \u0026amp; Zdobnov, E.M. BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cstrong\u003e31\u003c/strong\u003e, 3210-2 (2015).\u003c/li\u003e\n\u003cli\u003eLi, H. \u0026amp; Durbin, R. Fast and accurate long-read alignment with Burrows-Wheeler transform. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cstrong\u003e26\u003c/strong\u003e, 589-95 (2010).\u003c/li\u003e\n\u003cli\u003eXu, Z. \u0026amp; Wang, H. LTR_FINDER: an efficient tool for the prediction of full-length LTR retrotransposons. \u003cem\u003eNucleic Acids Res\u003c/em\u003e \u003cstrong\u003e35\u003c/strong\u003e, W265-8 (2007).\u003c/li\u003e\n\u003cli\u003eSmit, A. \u0026amp; Hubley, R.R. Open-1.0. Available from \u003cem\u003ehttp://www.repeatmasker.org\u003c/em\u003e (2008).\u003c/li\u003e\n\u003cli\u003eSmit, A., Hubley, R. \u0026amp; Green, P. RepeatMasker Open-4.0. 2013\u0026ndash;2015. (2015).\u003c/li\u003e\n\u003cli\u003eJebb, D.\u003cem\u003e et al.\u003c/em\u003e Six reference-quality genomes reveal evolution of bat adaptations. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e583\u003c/strong\u003e, 578-584 (2020).\u003c/li\u003e\n\u003cli\u003eTian, S.\u003cem\u003e et al.\u003c/em\u003e Comparative genomics provides insights into chromosomal evolution and immunological adaptation in horseshoe bats. \u003cem\u003eNat Ecol Evol\u003c/em\u003e \u003cstrong\u003e9\u003c/strong\u003e, 705-720 (2025).\u003c/li\u003e\n\u003cli\u003eHarris, R.S. Improved pairwise Alignmnet of genomic DNA. (2007).\u003c/li\u003e\n\u003cli\u003eKirilenko, B.M.\u003cem\u003e et al.\u003c/em\u003e Integrating gene annotation with orthology inference at scale. \u003cem\u003eScience\u003c/em\u003e \u003cstrong\u003e380\u003c/strong\u003e, eabn3107 (2023).\u003c/li\u003e\n\u003cli\u003eLiu, W.\u003cem\u003e et al.\u003c/em\u003e Large-scale across species transcriptomic analysis identifies genetic selection signatures associated with longevity in mammals. \u003cem\u003eEMBO J\u003c/em\u003e \u003cstrong\u003e42\u003c/strong\u003e, e112740 (2023).\u003c/li\u003e\n\u003cli\u003eGrabherr, M.G.\u003cem\u003e et al.\u003c/em\u003e Full-length transcriptome assembly from RNA-Seq data without a reference genome. \u003cem\u003eNat Biotechnol\u003c/em\u003e \u003cstrong\u003e29\u003c/strong\u003e, 644-52 (2011).\u003c/li\u003e\n\u003cli\u003eHaas, B.J.\u003cem\u003e et al.\u003c/em\u003e Improving the Arabidopsis genome annotation using maximal transcript alignment assemblies. \u003cem\u003eNucleic Acids Res\u003c/em\u003e \u003cstrong\u003e31\u003c/strong\u003e, 5654-5666 (2003).\u003c/li\u003e\n\u003cli\u003eKim, D.\u003cem\u003e et al.\u003c/em\u003e TopHat2: accurate alignment of transcriptomes in the presence of insertions, deletions and gene fusions. \u003cem\u003eGenome Biol\u003c/em\u003e \u003cstrong\u003e14\u003c/strong\u003e, R36 (2013).\u003c/li\u003e\n\u003cli\u003eTrapnell, C.\u003cem\u003e et al.\u003c/em\u003e Differential gene and transcript expression analysis of RNA-seq experiments with TopHat and Cufflinks. \u003cem\u003eNat Protoc\u003c/em\u003e \u003cstrong\u003e7\u003c/strong\u003e, 562-78 (2012).\u003c/li\u003e\n\u003cli\u003eStanke, M. \u0026amp; Waack, S. Gene prediction with a hidden Markov model and a new intron submodel. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cstrong\u003e19 Suppl 2\u003c/strong\u003e, ii215-25 (2003).\u003c/li\u003e\n\u003cli\u003eGuigo, R. Assembling genes from predicted exons in linear time with dynamic programming. \u003cem\u003eJ Comput Biol\u003c/em\u003e \u003cstrong\u003e5\u003c/strong\u003e, 681-702 (1998).\u003c/li\u003e\n\u003cli\u003eBurge, C. \u0026amp; Karlin, S. Prediction of complete gene structures in human genomic DNA. \u003cem\u003eJ Mol Biol\u003c/em\u003e \u003cstrong\u003e268\u003c/strong\u003e, 78-94 (1997).\u003c/li\u003e\n\u003cli\u003eMajoros, W.H., Pertea, M. \u0026amp; Salzberg, S.L. TigrScan and GlimmerHMM: two open source ab initio eukaryotic gene-finders. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cstrong\u003e20\u003c/strong\u003e, 2878-9 (2004).\u003c/li\u003e\n\u003cli\u003eKorf, I. Gene finding in novel genomes. \u003cem\u003eBMC Bioinformatics\u003c/em\u003e \u003cstrong\u003e5\u003c/strong\u003e, 59 (2004).\u003c/li\u003e\n\u003cli\u003eSchattner, P., Brooks, A.N. \u0026amp; Lowe, T.M. The tRNAscan-SE, snoscan and snoGPS web servers for the detection of tRNAs and snoRNAs. \u003cem\u003eNucleic Acids Res\u003c/em\u003e \u003cstrong\u003e33\u003c/strong\u003e, W686-9 (2005).\u003c/li\u003e\n\u003cli\u003eAltschul, S.F., Gish, W., Miller, W., Myers, E.W. \u0026amp; Lipman, D.J. Basic local alignment search tool. \u003cem\u003eJ Mol Biol\u003c/em\u003e \u003cstrong\u003e215\u003c/strong\u003e, 403-10 (1990).\u003c/li\u003e\n\u003cli\u003eNawrocki, E.P. \u0026amp; Eddy, S.R. Infernal 1.1: 100-fold faster RNA homology searches. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cstrong\u003e29\u003c/strong\u003e, 2933-5 (2013).\u003c/li\u003e\n\u003cli\u003eShao, Y.\u003cem\u003e et al.\u003c/em\u003e Phylogenomic analyses provide insights into primate evolution. \u003cem\u003eScience\u003c/em\u003e \u003cstrong\u003e380\u003c/strong\u003e, 913-924 (2023).\u003c/li\u003e\n\u003cli\u003ePeng, C.\u003cem\u003e et al.\u003c/em\u003e Large-scale snake genome analyses provide insights into vertebrate development. \u003cem\u003eCell\u003c/em\u003e \u003cstrong\u003e186\u003c/strong\u003e, 2959-2976 e22 (2023).\u003c/li\u003e\n\u003cli\u003eEditorial, N.B. A reference standard for genome biology. \u003cem\u003eNat Biotechnol\u003c/em\u003e\u003cstrong\u003e36\u003c/strong\u003e, 1121 (2018).\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[{"identity":"83d13e4a-558e-48ac-a8a3-66c94b96f95b","identifier":"10.13039/501100001809","name":"National Natural Science Foundation of China","awardNumber":"32471689","order_by":0},{"identity":"40b8769a-dd6d-483c-8d94-4bf4128460b1","identifier":"10.13039/501100003819","name":"Natural Science Foundation of Hubei Province","awardNumber":"2023AFA015","order_by":1},{"identity":"38c00839-13da-4233-8f9d-16555b050c5e","identifier":"10.13039/501100012226","name":"Fundamental Research Funds for the Central Universities","awardNumber":"2042022dx0003","order_by":2}],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"Wuhan University","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"big-footed bat, genome assembly","lastPublishedDoi":"10.21203/rs.3.rs-7561642/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7561642/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eSome bat species in the genus \u003cem\u003eMyotis\u003c/em\u003e have evolved longevity-associated mechanisms and exhibit remarkable resistance to cancer. Among them, the big-footed bat (\u003cem\u003eMyotis pilosus\u003c/em\u003e) has been confirmed as a cancer-resistant species. Here, we assembled a chromosome-level genome of the big-footed bat, utilizing a combination of ONT long reads and Hi-C technologies. The size of this genome is 1968.27 Mb, with a contig N50 of 41.29 Mb. All assembled sequences were anchored onto 21 autosomes and X chromosome. We identified 739.02 Mb (37.55%) of repetitive sequences in the genome and predicted 21,368 protein-coding genes. Assessment of the genome assembly quality indicated that the assembled genome of the big-footed bat exhibits excellent continuity, completeness, and accuracy. Taken together, we have generated the first chromosome-level genome assembly of the big-footed bat, providing an important reference resource for genetic and genomic studies of long-lived bat species.\u003c/p\u003e","manuscriptTitle":"Chromosome-level genome assembly of the big-footed bat (Myotis pilosus)","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-09-09 15:28:44","doi":"10.21203/rs.3.rs-7561642/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"bf625a14-bb33-40ce-a1fc-09f4b1ae5b27","owner":[],"postedDate":"September 9th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":54346142,"name":"Animal Science"}],"tags":[],"updatedAt":"2025-09-09T15:28:45+00:00","versionOfRecord":[],"versionCreatedAt":"2025-09-09 15:28:44","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7561642","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7561642","identity":"rs-7561642","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00