Full text
94,482 characters
· extracted from
preprint-html
· click to expand
Evolutionary Consequences of Unusually Large Pericentric TE-rich Regions in the Genome of a Neotropical Fig Wasp | bioRxiv /* */ /* */ <!-- <!-- /*! * yepnope1.5.4 * (c) WTFPL, GPLv2 */ (function(a,b,c){function d(a){return"[object Function]"==o.call(a)}function e(a){return"string"==typeof a}function f(){}function g(a){return!a||"loaded"==a||"complete"==a||"uninitialized"==a}function h(){var a=p.shift();q=1,a?a.t?m(function(){("c"==a.t?B.injectCss:B.injectJs)(a.s,0,a.a,a.x,a.e,1)},0):(a(),h()):q=0}function i(a,c,d,e,f,i,j){function k(b){if(!o&&g(l.readyState)&&(u.r=o=1,!q&&h(),l.onload=l.onreadystatechange=null,b)){"img"!=a&&m(function(){t.removeChild(l)},50);for(var d in y[c])y[c].hasOwnProperty(d)&&y[c][d].onload()}}var j=j||B.errorTimeout,l=b.createElement(a),o=0,r=0,u={t:d,s:c,e:f,a:i,x:j};1===y[c]&&(r=1,y[c]=[]),"object"==a?l.data=c:(l.src=c,l.type=a),l.width=l.height="0",l.onerror=l.onload=l.onreadystatechange=function(){k.call(this,r)},p.splice(e,0,u),"img"!=a&&(r||2===y[c]?(t.insertBefore(l,s?null:n),m(k,j)):y[c].push(l))}function j(a,b,c,d,f){return q=0,b=b||"j",e(a)?i("c"==b?v:u,a,b,this.i++,c,d,f):(p.splice(this.i++,0,a),1==p.length&&h()),this}function k(){var a=B;return a.loader={load:j,i:0},a}var l=b.documentElement,m=a.setTimeout,n=b.getElementsByTagName("script")[0],o={}.toString,p=[],q=0,r="MozAppearance"in l.style,s=r&&!!b.createRange().compareNode,t=s?l:n.parentNode,l=a.opera&&"[object Opera]"==o.call(a.opera),l=!!b.attachEvent&&!l,u=r?"object":l?"script":"img",v=l?"script":u,w=Array.isArray||function(a){return"[object Array]"==o.call(a)},x=[],y={},z={timeout:function(a,b){return b.length&&(a.timeout=b[0]),a}},A,B;B=function(a){function b(a){var a=a.split("!"),b=x.length,c=a.pop(),d=a.length,c={url:c,origUrl:c,prefixes:a},e,f,g;for(f=0;f<d;f++)g=a[f].split("="),(e=z[g.shift()])&&(c=e(c,g));for(f=0;f<b;f++)c=x[f](c);return c}function g(a,e,f,g,h){var i=b(a),j=i.autoCallback;i.url.split(".").pop().split("?").shift(),i.bypass||(e&&(e=d(e)?e:e[a]||e[g]||e[a.split("/").pop().split("?")[0]]),i.instead?i.instead(a,e,f,g,h):(y[i.url]?i.noexec=!0:y[i.url]=1,f.load(i.url,i.forceCSS||!i.forceJS&&"css"==i.url.split(".").pop().split("?").shift()?"c":c,i.noexec,i.attrs,i.timeout),(d(e)||d(j))&&f.load(function(){k(),e&&e(i.origUrl,h,g),j&&j(i.origUrl,h,g),y[i.url]=2})))}function h(a,b){function c(a,c){if(a){if(e(a))c||(j=function(){var a=[].slice.call(arguments);k.apply(this,a),l()}),g(a,j,b,0,h);else if(Object(a)===a)for(n in m=function(){var b=0,c;for(c in a)a.hasOwnProperty(c)&&b++;return b}(),a)a.hasOwnProperty(n)&&(!c&&!--m&&(d(j)?j=function(){var a=[].slice.call(arguments);k.apply(this,a),l()}:j[n]=function(a){return function(){var b=[].slice.call(arguments);a&&a.apply(this,b),l()}}(k[n])),g(a[n],j,b,n,h))}else!c&&l()}var h=!!a.test,i=a.load||a.both,j=a.callback||f,k=j,l=a.complete||f,m,n;c(h?a.yep:a.nope,!!i),i&&c(i)}var i,j,l=this.yepnope.loader;if(e(a))g(a,0,l,0);else if(w(a))for(i=0;i (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0];var j=d.createElement(s);var dl=l!='dataLayer'?'&l='+l:'';j.src='//www.googletagmanager.com/gtm.js?id='+i+dl;j.type='text/javascript';j.async=true;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-M677548'); Skip to main content Home About Submit ALERTS / RSS Search for this keyword Advanced Search New Results Evolutionary Consequences of Unusually Large Pericentric TE-rich Regions in the Genome of a Neotropical Fig Wasp View ORCID Profile Zexuan Zhao , View ORCID Profile Kevin Quinteros , View ORCID Profile Carlos A. Machado doi: https://doi.org/10.1101/2025.04.20.649723 Zexuan Zhao 1 Department of Biology, University of Maryland, College Park , MD, USA Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Zexuan Zhao Kevin Quinteros 1 Department of Biology, University of Maryland, College Park , MD, USA Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Kevin Quinteros Carlos A. Machado 1 Department of Biology, University of Maryland, College Park , MD, USA Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Carlos A. Machado For correspondence: machado{at}umd.edu Abstract Full Text Info/History Metrics Supplementary material Preview PDF Abstract Transposable elements (TEs), despite generally being considered deleterious, represent a substantial portion of most eukaryotic genomes. Specific genomic regions, such as telomeres and pericentromeres, are often densely populated with TEs. In these regions, which tend to be gene-poor, reduced recombination shelters the genome from the deleterious effects of TEs. Here, we describe unusually large and continuous pericentromeric TE-rich regions in all chromosomes of the genome assembly of Pegoscapus hoffmeyeri Sp. A (511.79 Mbp), a Neotropical fig wasp that is the obligate pollinator of Ficus obtusifolia . The identified pericentromeric TE-rich regions span nearly half (46%) of the genome, and harbor over 40% of all annotated genes, including 30% of conserved BUSCO genes. We present evidence that low recombination in these TE-rich regions generates strong bimodal molecular evolution patterns genome-wide. Patterns of nucleotide diversity and protein-coding gene evolution in TE-rich regions are consistent with a reduced efficiency of selection and suggestive of strong Hill–Robertson effects. A significant reduction in third codon position GC content (GC3) in TE-rich regions emerged as the most distinctive gene feature differentiating genes in TE-rich regions from those in the rest of the genome, a pattern that likely results from the absence of GC-biased gene conversion. This remarkable binary compartmental genome organization in the genome of P . hoffmeyeri provides a unique example of how genome organization with compartmental TE distribution can lead to context-dependent gene evolution shaped by common evolutionary forces. Introduction Haploid genome sizes of eukaryote species show little correlation with organismal complexity ( Lynch and Conery 2003 ; Wright 2017 ; Blommaert 2020 ). In insects, haploid genome sizes can vary by two orders of magnitude, ranging from 89.6 Mbp in the Antarctic midge, Belgica antarctica ( Kelley et al. 2014 ) to 21.48 Gbp in the grasshopper Bryodemella tuberculata ( Hawlitschek et al. 2023 ). Given that the number of genes remain relatively conserved across taxa, discrepancies in haploid genome sizes are predominantly the result of differences in the extent and composition of intronic and intergenic sequences ( Lynch and Conery 2003 ; Elliott and Gregory 2015a ). As transposable elements (TEs) are major components within these regions ( Bourque et al. 2018 ), they have become a major focus of research into the mechanisms underlying the evolution of genome structure ( Chalopin et al. 2015 ; Elliott and Gregory 2015b ; Sessegolo et al. 2016 ; Kapusta et al. 2017 ; Blommaert et al. 2019 ; Cong et al. 2022 ; Zuo et al. 2023 ; Betancourt et al. 2024 ). Although some studies have reported beneficial effects of TEs through their effects on regulatory evolution, chromosome structure, or horizontal gene transfer ( Klein and O’Neill 2018 ; Choudhary et al. 2020 ; Widen et al. 2023 ), most TEs are thought to be deleterious and to be under purifying selection ( Cridland et al. 2013 ; Blumenstiel et al. 2014 ; Stritt et al. 2018 ; Oggenfuss et al. 2021 ; Dazenière et al. 2022 ; Betancourt et al. 2024 ). Newly inserted TEs can disrupt coding sequences ( Charlesworth and Langley 1989 ; Hancks and Kazazian 2016 ), and remnant TE copies can trigger ectopic recombination and chromosomal rearrangement between non-homologous sequences across chromosomes ( Montgomery et al. 1987 ; Langley et al. 1988 ), dramatically increasing the risk of deleterious effects as TE copy number increases ( Bennetzen and Wang 2014 ). Even when TEs are suppressed through epigenetic modification, the inadvertent spreading of these modifications can negatively affect the expression of neighboring genes ( Choi and Lee 2020 ). Despite these potentially harmful effects associated with TEs, their prevalence across genomes remains a puzzling paradox that underscores their complex evolutionary and functional significance. One explanation for the prevalence of TEs is the existence of “TE refuges” in genomes ( Bertocchi et al. 2018 ). For instance, the Mbp-scale pericentric heterochromatic regions flanking centromeres in Drosophila melanogaster are populated with satellite DNA and TEs and even contain genes of exceptionally large size, such as Myo81F , which spans over 2.5 Mbp ( Hoskins et al. 2007 ; Hoskins et al. 2015 ). Other examples include non-recombining regions of sex chromosomes ( Erlandsson et al. 2000 ; Duhamel et al. 2023 ), B chromosomes ( Bertocchi et al. 2018 ), heterochromatic knobs ( Quesneville 2020 ) and chromosomal inversions ( Carpinteyro-Ponce and Machado 2024 ). These regions, characterized by suppressed recombination, limit the deleterious effects of ectopic recombination and/or slow down the removal of TEs through recombination ( Kent et al. 2017 ). Consequently, these genomic environments often harbor TE families with higher copy numbers, while other gene-rich regions accumulate younger TE families ( Baucom et al. 2009 ). While TEs can accumulate in regions with low recombination, TE silencing could also further reduce recombination rates by introducing repressive chromatin modifications ( Mirouze et al. 2012 ; Yelina et al. 2012 ; Zamudio et al. 2015 ; Huang et al. 2025 ). In addition, the coevolution between TEs, recombination and chromatin structure can also influence gene evolution in direct or indirect ways. TEs can directly influence gene evolution in TE rich regions through their accumulation in introns, resulting in increased gene size ( Yasuhara et al. 2005 ; Corradini et al. 2007 ). Further, in regions of low recombination the efficacy of selection is weakened due to increased Hill–Robertson interference ( Hill and Robertson 1966 ), leading to higher sequence divergence of protein coding genes in TE-rich regions ( Campos et al. 2012 ). The intricate interplay between TEs and chromatin structure, as well as their compound effects on heterochromatic genes, need to be further investigated. In this study, we present a comprehensive analysis of the evolutionary effects of TEs in the genome of one pollinating fig wasp (Superfamily Chalcidoidea, family Agaonidae). The karyotype of Agaonids is relatively stable at 5-6 chromosomes across the family, and most chromosomes are metacentric ( Liu et al. 2011 ; Gokhman et al. 2019 ). The conserved chromosome structure and lack of large-scale chromosome fusions or translocations provide a structurally stable genomic landscape to detect and track the evolution of TE refuges. Fig pollinating wasps are highly specialized organisms that spend most of their life cycle (3-4 weeks) developing in fig tissue, and only 1-2 days as adults searching for a new receptive fig host ( Janzen 1979 ; Weiblen 2002 ). Based on their special ecological niche, fig wasps were used to support the hypothesis that TEs should experience limited accumulation in genomes of specialists whose ecological contacts with other organisms are rare ( Gilbert et al. 2021 ). In fig wasps, this hypothesis was initially supported by the short-read based genome assembly of Ceratosolen solmsi which showed very low (6.4%) TE content ( Xiao et al. 2013 ). However, as genome assemblies for fig wasps have become more refined, the estimated TE content has approached 20% ( Cooper et al. 2020 ; Zhang et al. 2020 ), challenging this original hypothesis. In this study, we further challenge this idea by generating a chromosome-level genome assembly of Pegoscapus hoffmeyeri sp. A ( Grandi 1934 ), one of the two pollinator species associated with Ficus obtusifolia in Panama ( Molbo et al. 2003 ; Molbo et al. 2004 ). The assembled genome is 30-85% larger than published genomes from other fig wasp species due a remarkable increase in TE content. We show that TE content in this species is effectively bimodal, with extremely large pericentric TE-rich regions encompassing almost half of the genome and protein coding genes, and showing patterns of molecular evolution markedly differently than the rest of the genome. Materials and Methods High Molecular DNA Extraction and Reference Genome Sequencing We collected a brood from a fig pollinated by a single foundress on Barro Colorado Island, managed by the Smithsonian Tropical Research Institute. High molecular weight DNA was extracted using the Qiagen Blood and Cell culture Midi Kit following a previously described protocol ( Chakraborty et al. 2016 ). The Pacbio HiFi library was constructed and sequenced in a Sequel II instrument at the Institute for Genome Sciences (University of Maryland). Genome Assembly and Quality Assessment The raw subreads were aligned using CCS algorithm v 6.4.0 ( https://github.com/PacificBiosciences/ccs ) to call high-fidelity (HiFi) long reads. Contigs were assembled by Hifiasm v0.15 ( Cheng et al. 2021 ) and the taxonomical match of each contig was identified through Basic Local Alignment Search Tool (BLAST) against the nt database ( Altschul et al. 1990 ; Sayers et al. 2022 ). Contigs taxonomized as insects or unclassified were scaffolded using RagTag v1.0.2 ( Alonge et al. 2022 ) using the genome assembly of Eupristina verticillata (Accession # GWHFQFG00000000.1) as a reference. Assembly duplication was checked by self-alignment dot plots using D-Genies v1.5.0 ( Cabanettes and Klopp 2018 ). HiFi reads were mapped to the assembly using minimap2 v2.26 ( Li 2018 ; Li 2021 ) and mapping depths were calculated using samtools v1.17 ( Li et al. 2009 ; Danecek et al. 2021 ). The quality of the genome assembly was evaluated using the Genome Evaluation Pipeline ( https://git.imp.fu-berlin.de/cmazzoni/GEP ). It used the Python script ‘assembly_stats.py’ to calculate scaffold and contig statistics from the assembled genome ( Trizna 2020 ). GenomeScope v2.0 estimated genome size, repeat content, and heterozygosity from HiFi reads ( Vurture et al. 2017 ; Ranallo-Benavidez et al. 2020 ). Merqury v1.3 ( Rhie et al. 2020 ) evaluated kmer completeness, accuracy, and heterozygosity of the assembly, while BUSCO v5.4.7 ( Manni et al. 2021 ; Alonge et al. 2022 ) assessed gene completeness against the hymenoptera_odb10.2019-11-20 database. To examine whether the assembly reached pseudochromosome level, telomeric repeats were identified using tidk v0.2.3 ( Brown et al. 2023 ). Genome Annotation De novo transposable element annotation was performed using Extensive de novo TE Annotator (EDTA) v2.1.0 ( Ou et al. 2019 ). EDTA incorporates a variety of repeat annotators into its pipeline: GenomeTools v1.6.2 ( Gremme et al. 2013 ), LTR_FINDER v1.07 ( Xu and Wang 2007 ), LTR_retriever v2.9.0 ( Ou and Jiang 2018 ), Generic Repeat Finder v1.0 ( Shi and Liang 2019 ), TIR-Learner v2.5 ( Su et al. 2019 ), HelitronScanner v1.1 ( Xiong et al. 2014 ) and TEsorter v1.3 ( Zhang et al. 2022 ). It also uses RepeatModeler v2.0.3 ( Flynn et al. 2020 ) to identify remaining TEs and masks the genome using RepeatMasker v4.1.2 ( http://www.repeatmasker.org ). For comparison, we also annotated the genome assemblies of V. javana , W. pumilae , C. solmsi marchali and E. verticillata to control the different performance of TE annotation pipelines. We classified TEs into superfamilies according to ( Wicker et al. 2007 ), and assigned TEs that were classified only into order or higher classification level or unclassified as unclassified TEs. The createRepeatLandscape.pl script from RepeatMasker was used to generate a TE landscape. We employed homology-based and de novo gene prediction on the genome assembly with only long TEs (length >=1 Kb) masked. Homology-based gene prediction was done using Gene Model Mapper (GeMoMA) v1.9 ( Keilwagen et al. 2019 ) with 10 Hymenoptera genome annotations as references with a variety of genetic distances (Supplementary table 5). We turned off static intron length ( GeMoMa.sil=false ) in GeneModelMapper step to infer intron length from reference annotation individually for each gene. It is worth noting that GeMoMa tends to overestimate genes when multiple references are provided. To eliminate bias towards genes with extremely long sequences and control false discovery rate, we tested score/aa ratio from 0.8 to 2.0 with increment = 0.1 in GeMoMa’s Annotation Filter step and evaluated the quality of annotation using BUSCO completeness. We picked score/aa=1.5 as it generated the highest BUSCO completeness (Supplementary table 6). Using this optimized filtering parameter we predicted 12,285 genes, compared to 16,332 genes predicted using the default filter. De novo gene prediction was done using Augustus v3.5.0 ( Stanke et al. 2006 ). We used 1000 genes predicted by GeMoMa as a validation set and the rest as the training set to train and validate a de novo gene model. In total 23,777 genes were predicted by Augustus, of which 13,208 genes do not overlap with GeMoMa predicted genes. We used eggNOG-mapper v2.1.9 ( Cantalapiedra et al. 2021 ) with eggNOG database v5.0 ( Huerta-Cepas et al. 2019 ) for functional annotation of the predicted genes. A total of 10,264 genes from the 12,285 GeMoMA predicted genes and 2,872 from the 13,208 de novo predicted genes were functionally annotated. In total, we annotated 13,136 functional genes using homology-based and de novo approaches. Based on the longest isoforms for each genes, we also annotated the non-degenerate and four-fold degenerate sites using degenotate v1.3 ( Mirchandani et al. 2024 ). Classification of TE-rich Regions We calculated TE densities as the proportion of base pairs covered by homology-based TE annotations in 200-Kbp genomic windows after merging overlapping TE annotations using the GenomicRanges package v3.17 ( Lawrence et al. 2013 ) in R v4.3.0 ( R Core Team 2023 ). As the frequency of TE densities suggests a bimodal distribution (Supplementary figure 8), we fitted mixture distributions of k subpopulations with k from 1 to 3 to the TE densities, and calculated Akaike information criterion (AIC) and Bayesian Information Criterion (BIC) values to select the best mixture model. One assumption to classify genomic windows is local correlation. If the TE density of a window falls between two mean TE density peaks, the classification should rely on the classification of neighboring regions. To employ this idea, we fitted the TE densities along the five chromosomes to a hidden Markov model (HMM) with two states using depmixS4 v1.5.0 ( Visser and Speekenbrink 2010 ) package in R. HMM utilizes both the magnitude of TE densities and the transition of the densities to infer the state of each window. Two states with different mean TE densities can switch along the genome according to a transition probability matrix. The inter-state transition probability represents the probability of switching to a TE-rich window from a background window or vice versa. The lower the transition probability, the bigger the expected TE-rich islands, and a higher level of clustering pattern. We estimated the significance of the clustering pattern by permuting TE densities on the chromosomes and fitting the HMM to the permuted datasets. The p-value of observed transition probability is calculated from the distribution of permuted transition probabilities. Whole Genome Resequencing and Estimation of Nucleotide Diversity We collected eight female offspring from independent fig syconia from Panama’s Barro Colorado Monument. DNA was extracted using Omega Biotek’s E.Z.N.A. Mollusc & Insect DNA Kit. Illumina libraries were prepared using the NEBNext Ultra II FS DNA Library Prep Kit for Illumina (New England Biolabs) and libraries were sequenced on an Illumina HiSeq 3000 instrument at the Institute for Genome Sciences (University of Maryland). We compiled a customized workflow to detect SNPs from the Illumina short read population genomic dataset. Low-quality bases and adaptor sequences were trimmed by fastp v0.23.2 ( Chen et al. 2018 ; Chen 2023 ). The quality of reads was assessed using fastQC v0.23.2 ( Andrews et al. 2012 ) before and after trimming. Using the ‘fq2bam’ command within clara-parabricks v4.0.0 ( https://docs.nvidia.com/clara/parabricks/4.0.0/index.html ), each sample was aligned to the reference genome and duplicates were marked. Average read depth was estimated as the number of reads overlapping each base pair, and average read coverage was estimated as the proportion of base pairs covered by at least one read are examined across the genome and samples (Supplementary figure 9). Nucleotide diversity (π) is the average pairwise difference between all possible pairs of individuals ( Nei and Li 1979 ) and is normalized by the length of the sequences being compared. In the traditional variant calling pipeline, missing data is treated as reference calls, and when calculating π, the denominator used is the length of the sliding windows rather than the actual length of sequences observed, which can potentially lead to an underestimation of π. To correct the bias, we called genotypes at all sites in the reference genome. Variants were generated by DeepVariant 1.4.0-gpu ( Poplin et al. 2018 ), and jointly genotyped by GLnexus v1.4.1 ( Yun et al. 2021 ), and biallelic SNPs were selected by BCFtools v1.17 ( Danecek et al. 2021 ). We used DeepVariant for its high accuracy in genotype calling, while it does not support calling reference loci. Therefore, we generated reference genotype calls by BCFtools v1.17, and merged variant calls and reference calls using bedtools v2.31.0 ( Quinlan and Hall 2010 ) and GATK4 v4.4.0.0 ( Quinlan and Hall 2010 ). We filtered variant and reference genotype calls by vcftools v0.1.16 ( Danecek et al. 2011 ) with parameter --max-missing 0.8 --min-meanDP 15 --max-meanDP 50 -- minDP 10 --maxDP 100 . We further separated genotype calls by gene annotation and TE annotation using bedtools. Nucleotide diversity was calculated using pixy v1.2.7.beta1 ( Korunes and Samuk 2021 ) per 200-Kbp non-overlapping sliding windows. We added the total number of reference calls and variant calls of respective categories in each window as a confounding variable. For comparison, we also calculated uncorrected π from variant calls straight from DeepVariant using vcftools. To assess whether our sequencing depths are sufficient for accurate genotyping, we down-sampled reads to 10x and 20x coverage using SeqKit v2.5.1 ( Shen et al. 2016 ) and calculated recall, precision and F1 scores against variant calls derived from the full set of reads using Haplotype Comparison Tools v0.3.15 ( https://github.com/Illumina/hap.py ). CpG Methylation Detection 5-Methylcytosine (5mC) at CpG sites was first detected on the PacBio HiFi reads from subreads using jasmine v2.0.0 ( https://github.com/pacificbiosciences/jasmine ). Then these reads were aligned to the genome using pbmm2 v1.13.1 ( https://github.com/PacificBiosciences/pbmm2 ). For each CpG in the reference genome, the per-site methylation probability was calculated using pb-CpG-tools v2.3.2 ( https://github.com/PacificBiosciences/pb-CpG-tools ). To establish an optimal probability threshold to classify methylated CpGs, we plotted candidate probability thresholds against the detected number of CpG methylation sites. The optimal threshold (0.65) was identified by locating the threshold value where the absolute value of its local slope is at its minimum (Supplementary figure 13). We assigned CpGs to their nearest genes or TEs to avoid redundancy, and then counted the number of methylated CpGs in 500-bp windows. We then performed principal component analysis (PCA) and determined windows with the highest 5% loadings of PC1 as regions with differentially methylated CpGs. To validate the CpG methylation detection, we tested if lower observed-to-expected CpG ratios of methylated genes were observed than expected (Supplementary figure 16). Observed-to-expected CpG ratios in coding sequences were calculated using the formula from ( Elango et al. 2009 ). Summary of Gene Features The following additional features were tabulated as the input for the machine learning model: number of exons, number of introns, intron GC content, third codon position GC content (GC3), first and second codon position GC content (GC12), coding sequence lengths, total intron lengths, and number of TE-insertions in introns. We also included maximum likelihood codon bias (MCB) estimated using coRdon package v1.20.0 in R ( Elek et al. 2024 ). Synonymous divergence (dS), non-synonymous divergence (dN) and non-synonymous to synonymous divergence ratio (dN/dS) was calculated against protein sequences from another fig wasp E. verticillata using orthologr package v0.4.2 in R ( Drost et al. 2015 ), in which reciprocal BLAST hits were used to detect orthologs and the Comeron method was used to estimate dN and dS ( Comeron 1995 ). We also included gene annotation method (source) as a gene feature to avoid bias in annotation accuracy. Random Forest Analysis and Importance Analysis of Gene Features Random forest, a naive machine learning algorithm, implemented in randomForest R package v4.7-1.1 ( Liaw and Wiener 2002 ) was used to predict whether genes were in TE-rich regions or other regions from aforementioned gene features, and to determine which features were most important. It uses an ensemble of decision trees to predict the response, each of which was built upon a random selection of features. To avoid overfitting, we only trained shallow trees (maximum number of terminal nodes = 5). We randomly selected 80% of genes to train the model and used the rest to test its performance. To avoid bias, the dataset was resampled to balance the number of genes in the TE-rich regions and other regions. The importance of gene features was measured as increased error rate after permuting each feature. Results Pegoscapus hoffmeyeri sp. A has an unexpectedly large genome with high TE content From an extremely rare all-male brood from a single P. hoffmeyeri sp. A foundress, we generated 11.9 Gbp from 1.34M PacBio HiFi reads, expecting 40x coverage according to the sizes of genome assemblies of other fig wasps (around 300 Mbp) ( Table 1 ). However, k-mer analysis (k=31) of the HiFi reads suggests that the haploid genome size is 500.96 Mbp or 40% larger than other published fig wasp genomes, with extremely low heterozygosity (0.001%) (Supplementary figure 1). The low heterozygosity observed in P. hoffmeyeri sp. A is consistent with the highly inbred nature of P. hoffmeyeri sp. A ( Molbo et al. 2004 ). However, the large genome size was unexpected and was further investigated to determine if it was overestimated. View this table: View inline View popup Download powerpoint Table 1: Statistics of the P. hoffmeyeri sp.A genome assembly, and comparison with other fig wasp genome assemblies. The initial contig-level assembly had a length of 511.79 Mbp with a contig N50 of 29.59 Mbp. These findings were consistent with our assembly-free k-mer estimations. Using an assembly from Eupristina verticillata as a reference (Jie et al, unpublished data), we scaffolded 65 contigs onto 5 chromosomes with a scaffold N50 of 103.83 Mbp while only 1.5% of the initial assembly remained unplaced (Supplementary figure 2A, Supplementary figure 3, Table 1 ). Self-alignment of the genome assembly showed no significant duplicated segments in the assembly (Supplementary figure 4A). Benchmarking Universal Single-Copy Orthologs (BUSCO) analysis showed that the assembly contained 5489 out of 5991 (91.6%) complete Hymenoptera orthologs, with only 36 identified (0.6%) duplicates. The BUSCO completeness is consistent with those observed in other fig wasp genome assemblies ( Table 1 ). K-mer analysis indicated that 99.89% of 31-mers present in the HiFi reads are in the assembly, indicating a low level of misassembly such as switch errors. Mapping depths were consistent across the genome (Supplementary figure 4B), except for some sparse regions on the 5’ end of chromosome 1. Altogether these analyses suggest that the estimation of genome size was not inflated and thus the genome of P. hoffmeyeri sp. A is 30-85% larger than previously published fig wasp genomes ( Table 1 ). We conducted further assessment of assembly quality through whole genome annotation. Utilizing a homology-based annotation method (GeMoMA) ( Keilwagen et al. 2019 ), we identified 12,285 genes using a database compiled from mosquitoes, bees, flies, and other fig wasps (Supplementary table 5) ( Friedrich and Muqim 2003 ; Xiao et al. 2013 ; Hoskins et al. 2015 ; Wallberg et al. 2019 ; Dalla Benetta et al. 2020 ; Zhang et al. 2020 ; Wang et al. 2021 ; Habtewold et al. 2023 ; Toga et al. 2024 ; Lee et al. 2025 ). In addition, we predicted 23,777 genes using a de novo approach. We removed de novo predicted genes that overlapped with homology-based genes and annotated the function of the remaining set of genes. The final gene set comprises 13,136 functionally annotated protein-coding genes, which is comparable to other fig wasp genomes (Supplementary figure 2b, Table 1 ). Additionally, we annotated 89 complete rRNA loci (29 5S, 19 5.8S, 20 18S and 21 28S), and 166 tRNA loci (Supplementary figure 2b). We also identified AACCCAAT/ATTGGGTT telomeric repeats at both ends on 4 chromosomal scaffolds (Supplementary figure 2b). We annotated transposable elements (TEs) to investigate their contribution to the observed genome size expansion in Pegoscapus hoffmeyeri sp. A. We found that 33.11% of the genome is occupied by TEs ( Figure 1A ), the highest percentage among all sequenced fig wasp genomes to date. The genome size after excluding TEs is 340 Mbp. This number aligns more closely with genome sizes after TE removal that have been reported in other fig wasp species, which range from 267.78 to 317.79 Mbp ( Table 1 ). The five most abundant classified TE superfamilies—Helitron, Gypsy, Mutator, Copia, and Tc1_Mariner—constitute 6.22%, 5.33%, 3.31%, 2.29%, and 2.29% of the genome, respectively. When compared to E. verticillata (Supplementary figure 5A), P. hoffmeyeri sp. A possesses a larger quantity of DNA transposons, with a notable increase in the Helitron superfamily. Download figure Open in new tab Figure 1 TE composition, landscape and distribution associated with nucleotide diversity. (A) Genome composition of TE and non-TE elements. (B) The TE landscape; the X-axis represents the pairwise CpG-corrected Kimura 2-parameter substitution level of the TEs relative to the consensus sequence of their respective families, analogous to the age of the TEs, while the Y-axis is the percentage of the genome occupied by TEs. (C) Patterns of nucleotide diversity and TE density, with classified TE-rich regions in orange on the X-axis. To estimate the timeline of TE accumulation, we plotted the TE landscape, displaying the distribution of Kimura distances across TE superfamilies. A bimodal distribution was observed ( Figure 1B ), with a majority of TEs in the older part of the bell-shaped distribution. Using a mutation rate of μ = 6.8 × 10 −9 per site per generation estimated from honeybees ( Yang et al. 2015 ) and 6 generations per year ( Weiblen 2002 ), the expected Kimura substitution level is 4.08 per million years. The peak of the bell-shaped distribution is estimated to have occurred approximately 5.6 mya. There is another distinct peak near zero on the TE landscape, which suggests an ongoing transposition activity. The genome of E. verticillata shows a consistently steady and moderate TE activity and also an ongoing transposition activity on the TE landscape (Supplementary figure 5), but both processes have been or are less intense than those observed in P. hoffmeyeri sp. A. One important consequence of the observed TE accumulation is its effect on gene size, specifically intron size. While only 17.54 Mbp of the genome are exons, 170.27 Mbp are introns (Supplementary table 1). Within these introns, 54.91 Mbp are classified as TEs, with 12.18 Mbp specifically identified as Helitrons. Genes with intronic TE insertions (median size = 10,825 bp, n = 6058) are an order of magnitude larger than genes without TE insertions (median size = 1,653 bp, n = 7078) (p<0.001, Wilcoxon test) (Supplementary figure 6). A Binary Genome-wide Pattern of TE Density, Nucleotide Diversity and Protein Coding Gene Evolution To understand the organization of TEs across the genome, we calculated TE densities (see Methods) and explored their spatial pattern ( Figure 1C ). A binary pattern of TE densities was observed, and the rank orders of TE densities per superfamily were similar across the genome (Supplementary figure 7). A modality test supported a bimodal distribution (Supplementary figure 8, Supplementary table 2). A 2-state hidden Markov model trained solely by TE densities identified five major regions with elevated TE densities, amounting to 46.3 % of the genome (237 Mbp) ( Figure 1C ). Around 80% of TEs were located in these TE-rich regions. The average TE density is much higher in TE-rich regions (mean = 56.54%) compared to the other regions (mean = 11.60%). For each chromosome, TE-rich regions comprise 37.6 Mbp to 54.6 Mbp, or 38.56% to 51.02% of the chromosomes (Supplementary table 3). These regions are significantly clustered (p<0.001, permutation test) and major TE-rich regions are located close to the putative center of the chromosomes. Remarkably, these TE-rich regions also encompass 40.8% of all annotated genes, and are thus not gene-poor regions. In fact, the average gene density is only slightly decreased in TE-rich regions (4.26 genes per 200 Kbp) compared to other regions (5.53 genes per 200 Kbp) (p=0.0059, Wilcoxon test) (Supplementary figure 6). To understand the effect of these TE-rich regions on molecular evolutionary processes, we first estimated nucleotide diversity (π) using whole genome short read sequencing from 8 individuals and compared π within and outside these TE-rich regions. 83.2M to 118.2M Illumina reads per individual were generated from 8 female wasps, providing 21x to 30x coverage of the genome (Supplementary table 4, Supplementary figure 9). In total, we called 239,435 SNPs using DeepVariant. Down-sampling analysis shows consistent genotype calls even when coverage was reduced to 10x, and thus the sequencing depth was sufficient for reliable SNP calling (Supplementary figure 10). In repetitive regions, short-read alignment is more challenging, resulting in a higher false negative rate in variant calling. To address this potential source of underestimation of π, we called 361.51M reference genotypes, and calculated π exclusively from these reference and variant positions with similar quality (Supplementary figure 11). Similar to TE densities, nucleotide diversity (π) showed a binary pattern across the genome ( Figures 1C ). π is significantly lower in TE-rich regions (p < 2 x 10 -16 , Wilcoxon test), with a median π of 7.08 × 10 −5 , compared to a median π of 2.84 × 10 −4 in other regions. This pattern is also consistent across exons, intronic TEs, TE-free introns, intergenic TEs and intergenic non-TE sequences ( Figure 2 ). We found that effective sequence size, the total number of reference and variant sites when calculating π, heavily influences π estimation. In intronic and intergenic regions, π stabilized when effective sequence size was above 20 Kbp and was consistently lower in TE-rich regions. Although in exonic regions π hardly stabilized because exons are too sparse, exons in TE-rich regions still exhibited the lowest level of π compared to other sequences. Download figure Open in new tab Figure 2 Reduced nucleotide diversity in TE-rich regions. Genomic sequences were first separated into 200-Kbp windows and then separated by gene features (exons, introns and intergenic) and TE annotations (TEs and non-TEs). π was estimated in each window and plotted against their effective sequence sizes, colored by whether the windows were in TE-rich regions. (A) Exons; (B) Intergenic and intronic TEs and non-TEs. Smoothed lines represent conditional means, and the shaded area around the lines correspond to the 95% confidence intervals. To explore if there were significant differences in large scale patterns of molecular evolution between the two genomic compartments, we identified orthologous genes in E. verticillata and calculated synonymous divergence (dS), non-synonymous divergence (dN) and the dN/dS ratio. We found that genes located in TE-rich regions have lower dS (p < 0.0001, Wilcoxon test), but higher dN (p < 0.0001, Wilcoxon test) and dN/dS ratios (p < 0.0001, Wilcoxon test) when compared to genes outside these regions (Supplementary figure 20). This pattern is suggestive of significant differences in selection efficiency between the compartments. To further explore that idea we calculated π at 0-fold (π 0 ) and 4-fold (π 4 ) degenerate sites. We found that genome-wide π 0 was higher relative to π 4 in TE-rich regions (π 0 = 3.56 × 10⁻⁵; π 4 = 5.53 × 10⁻⁵) compared to other regions (π 0 = 7.76 × 10⁻⁵; π 4 = 3.34 × 10⁻⁴) (Supplementary figure 12A). Further, the contrast of π 0 and π 4 is not significant in TE-rich regions (p=0.2744) but was significant in the other region (p<0.0001) (Supplementary figure 12B). Lower GC3 is the Major Signature of Genes in TE-rich Regions, not DNA Methylation To maintain genome integrity, TEs in plants are often silenced through DNA methylation and histone modification ( Liu and Zhao 2023 ). However, the reduced abundance of DNA methylation in insects suggests that DNA methylation might have lost their role in TE silencing ( Bewick et al. 2017 ; Provataris et al. 2018 ). To explore the potential association between TEs and DNA methylation in fig wasps, we annotated methylated CpG dinucleotides in the genome assembly using the HiFi reads. In total 81K CpG 5mC methylation sites were detected, comprising only 0.6% of the total CpGs in the genome. We found methylation to be significantly lower in the TE-rich region (p<0.001, Wilcoxon test) (Supplementary figure 14). We next examined the distance between each methylated CpG to their nearest genes or TEs. We generated two null distributions by permutating methylated CpGs among all CpGs (labeled “permuted”) and randomly assigning new positions to each methylated CpGs across the whole genome (labeled “random”). Genes were enriched in CpG methylation ( Figure 3A ) when compared with null distributions, especially from 500 bp upstream and 5000 bp downstream of the start codon (Supplementary figure 15). On the other hand, the average distance between TEs and methylated CpGs is greater than the average distance observed in the null distributions ( Figure 3A ), likely resulting from the fact that methylation is restricted to genes. Further analyses show that evolutionarily conserved genes are more heavily methylated (Supplementary table 7, Supplementary figure 17), and methylated CpGs were also more enriched in exons than in introns (Supplementary table 8). Download figure Open in new tab Figure 3 Comparison between genes in TE-rich regions and other regions: (A) Probability density of distances between methylated CpGs to genes and TEs with upstream regions represented by negative coordinates. (B) Importance of gene features in predicting genes in TE-rich regions. Gene features were grouped by biological significance. (C) Comparison of GC3 and intron GC content. To identify the most critical genomic features associated with genes in TE-rich vs other regions we selected a high confidence set of 6,615 protein coding genes by filtering out genes without orthologs in E. verticillata or without introns. A dataset of 18 gene features, organized into five categories was compiled to explore differences between genes located in TE-rich regions and those in other regions. A random forest model was trained to predict whether a gene is in a TE-rich region based on these gene features. The training and testing accuracy reached 82.06% and 78.58%, respectively. Importance analysis revealed that features related to GC content were the main contributors to model accuracy, with GC content at the third codon position (GC3) being the most influential factor ( Figure 3B ). GC3 content in TE-rich regions was significantly lower than that of genes located on the arms of chromosomes (p < 0.01, Wilcoxon test) ( Figure 3C ). Additionally, GC3 content in these TE-rich regions was even more reduced compared to the genome-wide GC context. Intron GC content in TE-rich regions showed a bimodal distribution. Specifically, genes with a higher number of TE insertions have a higher intron GC content close to the genome-wide average GC content (∼30%) (Supplementary figure 19), while introns in genes with fewer TE insertions are extremely AT-biased. Furthermore, features associated with gene size, specifically the number of TE insertions and intron length, are also significant features that distinguish genes in TE-rich and other regions. Maximum likelihood codon bias (MCB) was identified as a relevant feature, albeit less impactful. As MCB and GC3 content are highly correlated (Supplementary figure 21), the pattern of MCB is likely driven by GC3 content. Discussion Fig wasps have become a remarkable system to study plant-pollinator mutualisms due to their obligate association with figs ( Machado et al. 2001 ; Herre et al. 2008 ; Cruaud et al. 2012 ). However, we have only recently started uncovering the genomic underpinnings of this mutualism, albeit with still limited genomic resources available for most fig wasp genera ( Xiao et al. 2013 ; Cooper et al. 2020 ; Zhang et al. 2020 ; Wang et al. 2021 ; Xiao et al. 2021 ; Chen et al. 2022 ). We present the first study exploring the genome organization and evolution of a fig wasp from the Americas, Pegoscapus hoffmeyeri Sp. A, the pollinator of Ficus obtusifolia in Panama. The genome is 1.32 to 1.86 times larger than the genomes of other fig wasp species sequenced so far and the genome size expansion is the result of increased TE content. The abundance of TEs uncovered here challenges the notion that ecological isolation inhibits TE accumulation by limiting horizontal gene transfer ( Gilbert et al. 2021 ). The Genome of P. hoffmeyeri sp.A Exhibits Two Distinct Compartments with Large-scale Bimodal TE Enrichment Patterns TEs exhibit specific distribution patterns throughout eukaryotic genomes ( Sigman and Slotkin 2016 ; Bourque et al. 2018 ). Intrinsic insertion preferences shape TE-specific localization within or close to some genomic features during the initial integration phase, while selective pressures bias overall TE retention in the post-insertion phase ( Sultana et al. 2017 ). For example, the LTR retrotransposon Ty1 in yeast preferentially inserts itself upstream of genes transcribed by RNA Polymerase III ( Guo et al. 2015 ; Cheung et al. 2016 ; Cheung et al. 2018 ). On the other hand, TEs close to genes or TEs that disrupt exons are more rapidly purged ( Cridland et al. 2013 ; Cridland et al. 2013 ; Stritt et al. 2018 ). Given that sequence characteristics only explain the enrichment of few TE families ( Liao et al. 2000 ; Liu et al. 2005 ; Linheiro and Bergman 2012 ), Mbp-scale TE enrichment patterns that apply to most TEs in general more likely result from location-based differential selection pressure. In gene-rich euchromatic regions, a stronger selective pressure exists to remove TEs that disrupt exons and regulatory regions, or that introduce repressive epigenetic marks into neighboring genes ( Sigman and Slotkin 2016 ; Huang et al. 2022 ) and consequently TEs are more likely to be retained in intergenic regions, where other inactive TEs can also serve as a buffer between them and neighboring genes. Given that TEs are also recognized as sources of structural polymorphisms ( Gray 2000 ; Maumus et al. 2015 ; Mun et al. 2021 ; Balachandran et al. 2022 ; Munasinghe et al. 2023 ; Carpinteyro-Ponce and Machado 2024 ), purifying selection against TE insertions can also be strong due to their potential to generate deleterious chromosome rearrangements by ectopic recombination between non-allelic TE insertions, posing major risks to genome stability ( Kent et al. 2017 ). However, reduced recombination rates in certain genomic compartments can shelter TEs from purifying selection, promoting their accumulation in these regions ( Dolgin and Charlesworth 2008 ; Sigman and Slotkin 2016 ). Thus, the interplay between insertion preferences, selection pressures and recombination dynamics can influence the distribution and evolution of TEs within a genome. In the P. hoffmeyeri sp.A genome, we observed a large-scale binary enrichment pattern of TEs, where nearly half of each chromosome, specifically the pericentric region, was populated with TEs. This pattern is consistent with the pericentric heterochromatin observed in Drosophila in terms of location and genomic element composition ( Hoskins et al. 2015 ). However, although these regions in the P. hoffmeyeri sp.A genome are full of TEs, as TEs are also located in introns, they are still gene-rich as they encompass about 40% of all the annotated protein coding genes in this genome. The TE-rich regions reported here are thus substantially different from known gene-poor and heterochromatic pericentric regions. Since recombination is suppressed in centromeres and pericentromeres ( Mancera et al. 2008 ; Comeron et al. 2012 ; Nambiar and Smith 2016 ; Kent et al. 2017 ), the significant enrichment of TEs in these large pericentric regions is likely the result of reduced recombination (see below). GC3 Patterns Provide Indirect but Strong Evidence of Lower Recombination Rates in TE-rich Regions GC3 content was the most significant feature (of 18 gene features) distinguishing TE-rich regions from other regions ( Figure 3B ). GC3 is significantly lower in TE-rich regions and this effect is more pronounced than metrics that are intuitively correlated with TEs, such as the number of TE insertions in introns (Supplementary figure 20). It is well-established that lower GC3 content is associated with lower recombination rates through GC-biased gene conversion (gBGC) in mammals ( Eyre-Walker 1997 ; Galtier et al. 2001 ; Duret and Galtier 2009 ), other eukaryotes ( Pessia et al. 2012 ), and, more relevantly, in honey bees ( Kent et al. 2012 ). This is because in addition to crossover, recombination can also lead to gene conversion, which could favor G/C alleles over A/T alleles, resulting in an increase of GC content following conversion. Therefore, the observed lower GC3 content in TE-rich regions could be a result of the absence of gBGC that recovers the GC content induced by AT biased mutations, as recombination rates are lower in these regions. Patterns of Nucleotide Diversity and Gene Evolution in TE-rich regions Nucleotide diversity is 75% lower in pericentric TE-rich regions ( Figures 1C , 2 ). This pattern was consistent across exonic, intronic and intergenic TE and non-TE sequences, and is also robust against artifacts such as mapping rates and effective sequence sizes. In essence, we found that TE enrichment is associated with reduced nucleotide diversity in neighboring sequences. This large-scale pattern can only be explained by reduction in recombination rates and/or mutation rates as opposed to demographic histories, which are expected to affect the whole genome. Although lower mutation rates are consistent with the lower dS observed in genes from TE-rich regions, this parameter is not a consistent predictor of these genes in our machine learning model ( Figure 3B ) and thus provides very weak support for any mutation rate variation among the two genome compartments. Instead, the observed differences in GC3 content support the conclusion that low nucleotide diversity in TE-rich regions is likely the result of enhanced Hill–Robertson interference (HRI) ( Hill and Robertson 1966 ) caused by reduced recombination rates. Given that TE-rich regions harbor a substantial number of genes, and that most of those genes should be under some level of purifying selection, HRI effects are likely to be strong in these regions due to reduced recombination. Genes located in TE-rich regions had higher dN/dS ratios when compared to genes outside these regions (Supplementary figure 20). At first glance, this is consistent with patterns previously observed in TE islands from inbreeding ants where genes in TE-rich regions exhibit higher adaptive divergence rates ( Schrader et al. 2014 ). In this genome, however, we do not interpret higher dN/dS as an indicator of faster adaptation. As shown in avian and mammalian species, the absence of recombination and gBGC is associated with an increase in dN/dS ratios ( Rousselle et al. 2018 ) due to the reduced efficiency of purifying selection when HRI effects are strong ( Rousselle et al. 2018 ), leading to an inflation of dN values. Instead of adaptive evolution, genes in these low recombination regions are more likely accumulating deleterious mutations, an effect known as Muller’s ratchet ( Muller 1964 ; Felsenstein 1974 ). This is consistent with our observation that the ratio of π at zero-fold degenerate sites to four-fold degenerate sites is higher in TE-rich regions compared to other regions (Supplementary figure 12). Furthermore, a decrease in gBGC that results in lower GC content may also slightly decrease mutation rates because cytosine and guanine exhibit higher inherent mutability (Supplementary figure 22), which consequently contributes to lower rates of synonymous divergence. Therefore, the higher dN/dS is likely a combination of direct and indirect consequences of low recombination rates in the TE-rich regions. Genes in TE-rich Regions may Evolve at the Epigenetic Level If DNA methylation suppresses TE activity, as observed in plants ( Sigman and Slotkin 2016 ), we should see enrichment of methylated CpGs in TEs. Instead, our analysis on the distance between methylated CpGs to genes and TEs found enrichment of methylated CpGs only in gene bodies ( Figure 3A ) and there is no direct association between DNA methylation and TEs. We also found methylated genes are more likely to be evolutionarily conserved BUSCO genes (Supplementary figure 17, Supplementary table 7), consistent with the idea that the function of DNA methylation in insects is to regulate the expression of genes involved in basic biological processes ( Elango et al. 2009 ; Foret et al. 2009 ; Sarda et al. 2012 ; Glastad et al. 2014 ). Given that genes located in TE-rich regions are also more heavily methylated (Supplementary figure 20), it is likely that they could be evolving at the epigenetic level ( Schmitz et al. 2011 ; Hagmann et al. 2015 ; Torres-Garcia et al. 2020 ; Ashe et al. 2021 ; Sarkies 2023 ; Wilson et al. 2023 ). In the P. hoffmeyeri sp. A genome, nearly half of the annotated genes, including 30% of the BUSCO genes, are located in these TE-rich regions, suggesting that gene expression does occur in TE-rich regions. As many of these genes have gone through dramatic intron sequence elongation due to TE insertions (Supplementary figure 6), selection may favor increased methylation as a compensatory mechanism necessary for proper gene expression ( Jeong et al. 2018 ; Wu et al. 2022 ). Given that DNA methylation and histone modifications can be highly related ( Hunt et al. 2013 ), these genes could also adopt histone modifications to facilitate expression within the TE-rich context ( Saha and Mishra 2019 ). For example, although H3K9me3 is a conserved repressive epigenetic modification that induces chromatin compaction by recruiting heterochromatin protein 1 (HP1), it is also required for the expression of heterochromatic genes and suppression of aberrant genes ( Ninova et al. 2020 ). Similar molecular mechanisms are likely being adopted during evolution to ensure gene expression. Implications of Genome Size Expansion on Fig Wasp Evolution The genome sizes of most hymenopteran species sequenced so far range from 180 to 340 Mbp, with approximately 12,000 to 20,000 genes ( Branstetter et al. 2018 ). Furthermore, they have fairly low GC content (30-45%) ( Branstetter et al. 2018 ) and low TE content relative to other arthropods ( Petersen et al. 2019 ). Although P. hoffmeyeri sp. A has a very large genome compared to other fig wasps, the TE-free size of its genome is similar to that of other fig wasps ( Table 1 ). This suggests that TEs play a major role in the evolution of genome size in this group of insects. In contrast to the consistent and moderate TE activity inferred from the TE landscape of E. verticillata , P. hoffmeyeri sp. A has experienced an intense burst of TE activity that peaked approximately 5.6 mya. As a consequence of the absence of alignable TE copies when they are too degraded to be recognized, the TE landscape can only show TE activities up 10 to 15 mya (40-60 Kimura divergence), which is more recent than the crown ages of every pollinating fig wasp genera ( Cruaud et al. 2012 ). Thus, the TE landscapes shown for P. hoffmeyeri sp. A and E. verticillata are independent and represent a partial TE activity history of each genus and some species-specific TE activities. To gain a comprehensive understanding of when the TE burst started in Pegoscapus and what TEs contributed to the genome size expansion, a more extensive sampling of high-quality fig wasp genomes across multiple genera will be necessary. There could be ecological reasons behind the observed genome size expansion. The adaptive hypothesis of genome size evolution posits that a larger genome size may provide evolutionary advantages, since organisms may benefit from larger cells and body sizes ( Gregory et al. 2000 ). However, comparative studies have failed to establish a correlation between body size and genome size ( Xu et al. 2021 ; Yuan et al. 2021 ). In the case of fig wasps, an increase in body size may enhance dispersal capabilities that facilitate access to receptive syconia. However, it can also interfere with their capacity to enter the fig syconium through the narrow ostiole ( Liu et al. 2013 ). Notably, even though P. hoffmeyeri sp.A is one of the largest pollinator species in Panama (Herre 1989), its female body size (2mm) ( Wiebes 1995 ) is still similar to that of E. verticillata ( Waterston 1921 ). Although fig wasp body size is likely the result of coevolution with the size of their host fig ostiole ( Weiblen 2002 ), it will be interesting to determine if there is a significant correlation between body size and genome size in genus Pegoscapus . Data Availability All scripts were deposited at the project github repository ( https://github.com/ZexuanZhao/Pegoscapus-hoffmeyeri-sp.A-genome-paper ). PacBio HiFi reads and the genome assembly of P. hoffmeyeri sp. A are available in NCBI under project number PRJNA1196887. Illumina short reads were deposited under project number PRJNA1198996. Acknowledgements This study was funded by National Science Foundation grants MCB-1716532 and DEB-2225083 to C.A.M. We thank Thomas Kocher, Phillip Johnson and Steve Mount for helpful discussions and comments on the manuscript. We are grateful to the Institute for Genome Sciences (University of Maryland) for generating the sequencing data. Funding National Science Foundation, https://ror.org/021nxhr62 , DEB-2225083 , DEB-1754572 References ↵ Alonge M , Lebeigle L , Kirsche M , Jenike K , Ou S , Aganezov S , Wang X , Lippman ZB , Schatz MC , Soyk S . 2022 . Automated assembly scaffolding using RagTag elevates a new tomato system for high-throughput genome editing . Genome Biol . 23 : 258 . OpenUrl CrossRef PubMed ↵ Altschul SF , Gish W , Miller W , Myers EW , Lipman DJ . 1990 . Basic local alignment search tool . J. Mol. Biol . 215 : 403 – 410 . OpenUrl CrossRef PubMed Web of Science ↵ Andrews S , Krueger F , Segonds-Pichon A , Biggins L , Krueger C , Wingett S . 2012 . FastQC . ↵ Ashe A , Colot V , Oldroyd BP . 2021 . How does epigenetics influence the course of evolution? Philos. Trans. R. Soc. B Biol. Sci . 376 : 20200111 . OpenUrl CrossRef PubMed ↵ Balachandran P , Walawalkar IA , Flores JI , Dayton JN , Audano PA , Beck CR . 2022 . Transposable element-mediated rearrangements are prevalent in human genomes . Nat. Commun . 13 : 7115 . OpenUrl CrossRef PubMed ↵ Baucom RS , Estill JC , Chaparro C , Upshaw N , Jogi A , Deragon J-M , Westerman RP , Sanmiguel PJ , Bennetzen JL . 2009 . Exceptional diversity, non-random distribution, and rapid evolution of retroelements in the B73 maize genome . PLoS Genet . 5 : e1000732 . OpenUrl CrossRef PubMed ↵ Bennetzen JL , Wang H . 2014 . The contributions of transposable elements to the structure, function, and evolution of plant genomes . Annu. Rev. Plant Biol . 65 : 505 – 530 . OpenUrl CrossRef PubMed Web of Science ↵ Bertocchi NA , de Oliveira TD , del Valle Garnero A , Coan RLB , Gunski RJ , Martins C , Torres FP . 2018 . Distribution of CR1-like transposable element in woodpeckers (Aves Piciformes): Z sex chromosomes can act as a refuge for transposable elements . Chromosome Res . 26 : 333 – 343 . OpenUrl CrossRef PubMed ↵ Betancourt AJ , Wei KH-C , Huang Y , Lee YCG . 2024 . Causes and Consequences of Varying Transposable Element Activity: An Evolutionary Perspective . Annu. Rev. Genomics Hum. Genet . 25 : 1 – 25 . OpenUrl CrossRef PubMed ↵ Bewick AJ , Vogel KJ , Moore AJ , Schmitz RJ . 2017 . Evolution of DNA Methylation across Insects . Mol. Biol. Evol . 34 : 654 – 665 . OpenUrl CrossRef PubMed ↵ Blommaert J . 2020 . Genome size evolution: towards new model systems for old questions . Proc. R. Soc. B Biol. Sci . 287 : 20201441 . OpenUrl PubMed ↵ Blommaert J , Riss S , Hecox-Lea B , Mark Welch DB , Stelzer CP . 2019 . Small, but surprisingly repetitive genomes: transposon expansion and not polyploidy has driven a doubling in genome size in a metazoan species complex . BMC Genomics 20 : 466 . OpenUrl CrossRef PubMed ↵ Blumenstiel JP , Chen X , He M , Bergman CM . 2014 . An age-of-allele test of neutrality for transposable element insertions . Genetics 196 : 523 – 538 . OpenUrl Abstract / FREE Full Text ↵ Bourque G , Burns KH , Gehring M , Gorbunova V , Seluanov A , Hammell M , Imbeault M , Izsvák Z , Levin HL , Macfarlan TS , et al. 2018 . Ten things you should know about transposable elements . Genome Biol . 19 : 199 . OpenUrl CrossRef PubMed ↵ Branstetter MG , Childers AK , Cox-Foster D , Hopper KR , Kapheim KM , Toth AL , Worley KC . 2018 . Genomes of the Hymenoptera . Curr. Opin. Insect Sci . 25 : 65 – 75 . OpenUrl CrossRef PubMed ↵ Brown M , González De la Rosa PM , Mark B. 2023 . A Telomere Identification Toolkit . Available from : doi: 10.5281/zenodo.10091385 OpenUrl CrossRef ↵ Cabanettes F , Klopp C . 2018 . D-GENIES: dot plot large genomes in an interactive, efficient and simple way . PeerJ 6 : e4958 . OpenUrl CrossRef PubMed ↵ Campos JL , Charlesworth B , Haddrill PR . 2012 . Molecular Evolution in Nonrecombining Regions of the Drosophila melanogaster Genome . Genome Biol. Evol . 4 : 278 – 288 . OpenUrl CrossRef PubMed ↵ Cantalapiedra CP , Hernández-Plaza A , Letunic I , Bork P , Huerta-Cepas J . 2021 . eggNOG-mapper v2: Functional Annotation, Orthology Assignments, and Domain Prediction at the Metagenomic Scale . Mol. Biol. Evol . 38 : 5825 – 5829 . OpenUrl CrossRef PubMed ↵ Carpinteyro-Ponce J , Machado CA . 2024 . The Complex Landscape of Structural Divergence Between the Drosophila pseudoobscura and D. persimilis Genomes . Genome Biol. Evol . 16 : evae047 . OpenUrl CrossRef PubMed ↵ Chakraborty M , Baldwin-Brown JG , Long AD , Emerson JJ . 2016 . Contiguous and accurate de novo assembly of metazoan genomes with modest long read coverage . Nucleic Acids Res . 44 : e147 . OpenUrl CrossRef PubMed ↵ Chalopin D , Naville M , Plard F , Galiana D , Volff J-N . 2015 . Comparative analysis of transposable elements highlights mobilome diversity and evolution in vertebrates . Genome Biol. Evol . 7 : 567 – 580 . OpenUrl CrossRef PubMed ↵ Charlesworth B , Langley CH . 1989 . The population genetics of Drosophila transposable elements . Annu. Rev. Genet . 23 : 251 – 287 . OpenUrl CrossRef PubMed Web of Science ↵ Chen L , Feng C , Wang R , Nong X , Deng X , Chen X , Yu H . 2022 . A chromosome-level genome assembly of the pollinating fig wasp Valisia javana . DNA Res . 29 : dsac014 . OpenUrl CrossRef PubMed ↵ Chen S. 2023 . Ultrafast one-pass FASTQ data preprocessing, quality control, and deduplication using fastp . iMeta 2 : e107 . OpenUrl CrossRef ↵ Chen S , Zhou Y , Chen Y , Gu J . 2018 . fastp: an ultra-fast all-in-one FASTQ preprocessor . Bioinformatics 34 : i884 – i890 . OpenUrl CrossRef PubMed ↵ Cheng H , Concepcion GT , Feng X , Zhang H , Li H . 2021 . Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm . Nat. Methods 18 : 170 – 175 . OpenUrl CrossRef PubMed ↵ Cheung S , Ma L , Chan PHW , Hu H-L , Mayor T , Chen H-T , Measday V . 2016 . Ty1 Integrase Interacts with RNA Polymerase III-specific Subcomplexes to Promote Insertion of Ty1 Elements Upstream of Polymerase (Pol) III-transcribed Genes . J. Biol. Chem . 291 : 6396 . OpenUrl Abstract / FREE Full Text ↵ Cheung S , Manhas S , Measday V . 2018 . Retrotransposon targeting to RNA polymerase III-transcribed genes . Mob. DNA 9 : 14 . OpenUrl CrossRef PubMed ↵ Choi JY , Lee YCG . 2020 . Double-edged sword: The evolutionary consequences of the epigenetic silencing of transposable elements . PLOS Genet . 16 : e1008872 . OpenUrl CrossRef PubMed ↵ Choudhary MN , Friedman RZ , Wang JT , Jang HS , Zhuo X , Wang T . 2020 . Co-opted transposons help perpetuate conserved higher-order chromosomal structures . Genome Biol . 21 : 16 . OpenUrl CrossRef PubMed ↵ Comeron JM . 1995 . A method for estimating the numbers of synonymous and nonsynonymous substitutions per site . J. Mol. Evol . 41 : 1152 – 1159 . OpenUrl CrossRef PubMed Web of Science ↵ Comeron JM , Ratnappan R , Bailin S . 2012 . The Many Landscapes of Recombination in Drosophila melanogaster . PLOS Genet . 8 : e1002905 . OpenUrl CrossRef PubMed ↵ Cong Y , Ye X , Mei Y , He K , Li F. 2022 . Transposons and non-coding regions drive the intrafamily differences of genome size in insects . iScience 25 : 104873 . OpenUrl CrossRef PubMed ↵ Cooper L , Bunnefeld L , Hearn J , Cook JM , Lohse K , Stone GN . 2020 . Low-coverage genomic data resolve the population divergence and gene flow history of an Australian rain forest fig wasp . Mol. Ecol . 29 : 3649 – 3666 . OpenUrl CrossRef ↵ Corradini N , Rossi F , Giordano E , Caizzi R , Verní F , Dimitri P . 2007 . Drosophila melanogaster as a model for studying protein-encoding genes that are resident in constitutive heterochromatin . Heredity 98 : 3 – 12 . OpenUrl CrossRef PubMed ↵ Cridland JM , Macdonald SJ , Long AD , Thornton KR . 2013 . Abundance and Distribution of Transposable Elements in Two Drosophila QTL Mapping Resources . Mol. Biol. Evol . 30 : 2311 – 2327 . OpenUrl CrossRef PubMed Web of Science ↵ Cruaud A , Rønsted N , Chantarasuwan B , Chou LS , Clement WL , Couloux A , Cousins B , Genson G , Harrison RD , Hanson PE , et al. 2012 . An Extreme Case of Plant–Insect Codiversification: Figs and Fig-Pollinating Wasps . Syst. Biol . 61 : 1029 – 1047 . OpenUrl CrossRef PubMed ↵ Dalla Benetta E , Antoshechkin I , Yang T , Nguyen HQM , Ferree PM , Akbari OS . 2020 . Genome elimination mediated by gene expression from a selfish chromosome . Sci. Adv . 6 : eaaz9808 . OpenUrl FREE Full Text ↵ Danecek P , Auton A , Abecasis G , Albers CA , Banks E , DePristo MA , Handsaker RE , Lunter G , Marth GT , Sherry ST , et al. 2011 . The variant call format and VCFtools . Bioinformatics 27 : 2156 – 2158 . OpenUrl CrossRef PubMed Web of Science ↵ Danecek P , Bonfield JK , Liddle J , Marshall J , Ohan V , Pollard MO , Whitwham A , Keane T , McCarthy SA , Davies RM , et al. 2021 . Twelve years of SAMtools and BCFtools . GigaScience 10 : giab008 . OpenUrl CrossRef PubMed ↵ Dazenière J , Bousios A , Eyre-Walker A . 2022 . Patterns of selection in the evolution of a transposable element . G3 GenesGenomesGenetics 12 : jkac056 . OpenUrl ↵ Dolgin ES , Charlesworth B . 2008 . The Effects of Recombination Rate on the Distribution and Abundance of Transposable Elements . Genetics 178 : 2169 – 2177 . OpenUrl Abstract / FREE Full Text ↵ Drost H-G , Gabel A , Grosse I , Quint M . 2015 . Evidence for Active Maintenance of Phylotranscriptomic Hourglass Patterns in Animal and Plant Embryogenesis . Mol. Biol. Evol . 32 : 1221 – 1231 . OpenUrl CrossRef PubMed ↵ Duhamel M , Hood ME , Rodríguez de la Vega RC , Giraud T. 2023 . Dynamics of transposable element accumulation in the non-recombining regions of mating-type chromosomes in anther-smut fungi . Nat. Commun . 14 : 5692 . OpenUrl CrossRef PubMed ↵ Duret L , Galtier N . 2009 . Biased Gene Conversion and the Evolution of Mammalian Genomic Landscapes . Annu. Rev. Genomics Hum. Genet . 10 : 285 – 311 . OpenUrl CrossRef PubMed Web of Science ↵ Elango N , Hunt BG , Goodisman MAD , Yi SV . 2009 . DNA methylation is widespread and associated with differential gene expression in castes of the honeybee, Apis mellifera . Proc. Natl. Acad. Sci . 106 : 11206 – 11211 . OpenUrl Abstract / FREE Full Text ↵ Elek A , Kuzman M , Vlahovicek K. 2024 . coRdon: Codon Usage Analysis and Prediction of Gene Expressivity . Available from: https://bioconductor.org/packages/coRdon ↵ Elliott TA , Gregory TR . 2015a . What’s in a genome? The C-value enigma and the evolution of eukaryotic genome content . Philos. Trans. R. Soc. B Biol. Sci . 370 : 20140331 . OpenUrl CrossRef PubMed ↵ Elliott TA , Gregory TR . 2015b . Do larger genomes contain more diverse transposable elements? BMC Evol. Biol . 15 : 69 . OpenUrl CrossRef PubMed ↵ Erlandsson R , Wilson JF , Pääbo S . 2000 . Sex Chromosomal Transposable Element Accumulation and Male-Driven Substitutional Evolution in Humans . Mol. Biol. Evol . 17 : 804 – 812 . OpenUrl CrossRef PubMed Web of Science ↵ Eyre-Walker A . 1997 . Recombination and mammalian genome evolution . Proc. R. Soc. Lond. B Biol. Sci . 252 : 237 – 243 . OpenUrl ↵ Felsenstein J . 1974 . The evolutionary advantage of recombination . Genetics 78 : 737 – 756 . OpenUrl Abstract / FREE Full Text ↵ Flynn JM , Hubley R , Goubert C , Rosen J , Clark AG , Feschotte C , Smit AF. 2020 . RepeatModeler2 for automated genomic discovery of transposable element families . Proc. Natl. Acad. Sci . 117 : 9451 – 9457 . OpenUrl Abstract / FREE Full Text ↵ Foret S , Kucharski R , Pittelkow Y , Lockett GA , Maleszka R . 2009 . Epigenetic regulation of the honey bee transcriptome: unravelling the nature of methylated genes . BMC Genomics 10 : 472 . OpenUrl CrossRef PubMed ↵ Friedrich M , Muqim N . 2003 . Sequence and phylogenetic analysis of the complete mitochondrial genome of the flour beetle Tribolium castanaeum . Mol. Phylogenet. Evol . 26 : 502 – 512 . OpenUrl CrossRef PubMed Web of Science ↵ Galtier N , Piganeau G , Mouchiroud D , Duret L . 2001 . GC-Content Evolution in Mammalian Genomes: The Biased Gene Conversion Hypothesis . Genetics 159 : 907 – 911 . OpenUrl FREE Full Text ↵ Gilbert C , Peccoud J , Cordaux R . 2021 . Transposable Elements and the Evolution of Insects . Annu. Rev. Entomol . 66 : 355 – 372 . OpenUrl CrossRef PubMed ↵ Glastad KM , Hunt BG , Goodisman MA . 2014 . Evolutionary insights into DNA methylation in insects . Curr. Opin. Insect Sci . 1 : 25 – 30 . OpenUrl CrossRef PubMed ↵ Gokhman VE , Nugnes F , Bernardo U . 2019 . Chromosomes of Eupristina verticillata Waterston, 1921 and an overview of known karyotypes of chalcid wasps of the family Agaonidae (Hymenoptera) . J. Hymenopt. Res . 71 : 157 – 161 . OpenUrl CrossRef ↵ Grandi G . 1934 . Nuovi Agaonidi’(Hymenoptera Chalcididae) della fauna neotropica . Boll Lab Ent Bologna 7 : 186 – 197 . OpenUrl ↵ Gray YHM . 2000 . It takes two transposons to tango:transposable-element-mediated chromosomal rearrangements . Trends Genet . 16 : 461 – 468 . OpenUrl CrossRef PubMed Web of Science ↵ Gregory TR , Hebert PDN , Kolasa J . 2000 . Evolutionary implications of the relationship between genome size and body size in flatworms and copepods . Heredity 84 : 201 – 208 . OpenUrl CrossRef PubMed Web of Science ↵ Gremme G , Steinbiss S , Kurtz S . 2013 . GenomeTools: A Comprehensive Software Library for Efficient Processing of Structured Genome Annotations . IEEE/ACM Trans. Comput. Biol. Bioinform . 10 : 645 – 656 . OpenUrl CrossRef ↵ Guo Y , Singh PK , Levin HL . 2015 . A long terminal repeat retrotransposon of Schizosaccharomyces japonicus integrates upstream of RNA pol III transcribed genes . Mob. DNA 6 : 19 . OpenUrl CrossRef PubMed ↵ Habtewold T , Wagah M , Tambwe MM , Moore S , Windbichler N , Christophides G , Johnson H , Heaton H , Collins J , Krasheninnikova K , et al. 2023 . A chromosomal reference genome sequence for the malaria mosquito, Anopheles gambiae, Giles, 1902, Ifakara strain . Wellcome Open Res . 8 : 74 . ↵ Hagmann J , Becker C , Müller J , Stegle O , Meyer RC , Wang G , Schneeberger K , Fitz J , Altmann T , Bergelson J , et al. 2015 . Century-scale Methylome Stability in a Recently Diverged Arabidopsis thaliana Lineage . PLOS Genet . 11 : e1004920 . OpenUrl CrossRef PubMed ↵ Hancks DC , Kazazian HH . 2016 . Roles for retrotransposon insertions in human disease . Mob. DNA 7 : 9 . OpenUrl CrossRef PubMed ↵ Hawlitschek O , Sadílek D , Dey L-S , Buchholz K , Noori S , Baez IL , Wehrt T , Brozio J , Trávníček P , Seidel M , et al. 2023 . New estimates of genome size in Orthoptera and their evolutionary implications . PloS One 18 : e0275551 . OpenUrl CrossRef PubMed ↵ Herre EA , Jandér KC , Machado CA . 2008 . Evolutionary Ecology of Figs and Their Associates: Recent Progress and Outstanding Puzzles . Annu. Rev. Ecol. Evol. Syst . 39 : 439 – 458 . OpenUrl CrossRef ↵ Hill WG , Robertson A . 1966 . The effect of linkage on limits to artificial selection . Genet. Res . 8 : 269 – 294 . OpenUrl CrossRef PubMed Web of Science ↵ Hoskins RA , Carlson JW , Kennedy C , Acevedo D , Evans-Holm M , Frise E , Wan KH , Park S , Mendez-Lago M , Rossi F , et al. 2007 . Sequence Finishing and Mapping of Drosophila melanogaster Heterochromatin . Science 316 : 1625 – 1628 . OpenUrl Abstract / FREE Full Text ↵ Hoskins RA , Carlson JW , Wan KH , Park S , Mendez I , Galle SE , Booth BW , Pfeiffer BD , George RA , Svirskas R , et al. 2015 . The Release 6 reference sequence of the Drosophila melanogaster genome . Genome Res . 25 : 445 – 458 . OpenUrl Abstract / FREE Full Text ↵ Huang Y , Gao ZY , Ly K , Lin L , Lambooij J-P , King EG , Janssen A , Wei KH-C , Lee YCG . 2025 . Polymorphic transposable elements contribute to variation in recombination landscapes . Proc. Natl. Acad. Sci . 122 : e2427312122 . OpenUrl CrossRef PubMed ↵ Huang Y , Shukla H , Lee YCG. 2022 . Species-specific chromatin landscape determines how transposable elements shape genome evolution . Nordborg M , Przeworski M , editors. eLife 11 : e81567 . OpenUrl CrossRef PubMed ↵ Huerta-Cepas J , Szklarczyk D , Heller D , Hernández-Plaza A , Forslund SK , Cook H , Mende DR , Letunic I , Rattei T , Jensen LJ , et al. 2019 . eggNOG 5.0: a hierarchical, functionally and phylogenetically annotated orthology resource based on 5090 organisms and 2502 viruses . Nucleic Acids Res . 47 : D309 – D314 . OpenUrl CrossRef PubMed ↵ Hunt BG , Glastad KM , Yi SV , Goodisman MAD . 2013 . Patterning and Regulatory Associations of DNA Methylation Are Mirrored by Histone Modifications in Insects . Genome Biol. Evol . 5 : 591 – 598 . OpenUrl CrossRef PubMed ↵ Janzen DH . 1979 . How to be a Fig . Annu. Rev. Ecol. Syst . 10 : 13 – 51 . OpenUrl CrossRef Web of Science ↵ Jeong H , Wu X , Smith B , Yi SV . 2018 . Genomic Landscape of Methylation Islands in Hymenopteran Insects . Genome Biol. Evol . 10 : 2766 – 2776 . OpenUrl CrossRef PubMed ↵ Kapusta A , Suh A , Feschotte C . 2017 . Dynamics of genome size evolution in birds and mammals . Proc. Natl. Acad. Sci. U. S. A . 114 : E1460 – E1469 . OpenUrl Abstract / FREE Full Text ↵ Kollmar M Keilwagen J , Hartung F , Grau J . 2019 . GeMoMa: Homology-Based Gene Prediction Utilizing Intron Position Conservation and RNA-seq Data . In: Kollmar M , editor. Gene Prediction: Methods and Protocols . New York, NY : Springer . p. 161 – 177 . Available from : doi: 10.1007/978-1-4939-9173-0_9 OpenUrl CrossRef PubMed ↵ Kelley JL , Peyton JT , Fiston-Lavier A-S , Teets NM , Yee M-C , Johnston JS , Bustamante CD , Lee RE , Denlinger DL . 2014 . Compact genome of the Antarctic midge is likely an adaptation to an extreme environment . Nat. Commun . 5 : 4611 . OpenUrl CrossRef PubMed ↵ Kent CF , Minaei S , Harpur BA , Zayed A . 2012 . Recombination is associated with the evolution of genome structure and worker behavior in honey bees . Proc. Natl. Acad. Sci . 109 : 18012 – 18017 . OpenUrl Abstract / FREE Full Text ↵ Kent TV , Uzunović J , Wright SI . 2017 . Coevolution between transposable elements and recombination . Philos. Trans. R. Soc. B Biol. Sci . 372 : 20160458 . OpenUrl CrossRef PubMed ↵ Klein SJ , O’Neill RJ . 2018 . Transposable elements: genome innovation, chromosome diversity, and centromere conflict . Chromosome Res . 26 : 5 . OpenUrl CrossRef PubMed ↵ Korunes KL , Samuk K . 2021 . pixy: Unbiased estimation of nucleotide diversity and divergence in the presence of missing data . Mol. Ecol. Resour . 21 : 1359 – 1368 . OpenUrl CrossRef PubMed ↵ Langley CH , Montgomery E , Hudson R , Kaplan N , Charlesworth B . 1988 . On the role of unequal exchange in the containment of transposable element copy number . Genet. Res . 52 : 223 – 235 . OpenUrl CrossRef PubMed Web of Science ↵ Lawrence M , Huber W , Pagès H , Aboyoun P , Carlson M , Gentleman R , Morgan MT , Carey VJ . 2013 . Software for Computing and Annotating Genomic Ranges . PLOS Comput. Biol . 9 : e1003118 . OpenUrl CrossRef PubMed ↵ Lee J , Kiuchi T , Yamaguchi K , Shigenobu S , Toyoda A , Shimada T . 2025 . A chromosome-level genome assembly of wild silkmoth, Bombyx mandarina . Sci. Data 12 : 27 . OpenUrl CrossRef PubMed ↵ Li H . 2018 . Minimap2: pairwise alignment for nucleotide sequences . Bioinformatics 34 : 3094 – 3100 . OpenUrl CrossRef PubMed ↵ Li H . 2021 . New strategies to improve minimap2 alignment accuracy . Bioinformatics 37 : 4572 – 4574 . OpenUrl CrossRef PubMed ↵ Li H , Handsaker B , Wysoker A , Fennell T , Ruan J , Homer N , Marth G , Abecasis G , Durbin R , 1000 Genome Project Data Processing Subgroup . 2009 . The Sequence Alignment/Map format and SAMtools . Bioinformatics 25 : 2078 – 2079 . OpenUrl CrossRef PubMed Web of Science ↵ Liao G , Rehm EJ , Rubin GM . 2000 . Insertion site preferences of the P transposable element in Drosophila melanogaster . Proc. Natl. Acad. Sci . 97 : 3347 – 3351 . OpenUrl Abstract / FREE Full Text ↵ Liaw A , Wiener M . 2002 . Classification and Regression by randomForest . R News 2 : 18 – 22 . OpenUrl CrossRef ↵ Linheiro RS , Bergman CM. 2012 . Whole Genome Resequencing Reveals Natural Target Site Preferences of Transposable Elements in Drosophila melanogaster . PLOS ONE 7 : e30008 . OpenUrl CrossRef PubMed ↵ Liu B , Zhao M . 2023 . How transposable elements are recognized and epigenetically silenced in plants? Curr. Opin. Plant Biol . 75 : 102428 . OpenUrl CrossRef PubMed ↵ Liu C , Yang D-R , Compton SG , Peng Y-Q . 2013 . Larger Fig Wasps Are More Careful About Which Figs to Enter – With Good Reason . PLOS ONE 8 : e74117 . OpenUrl CrossRef PubMed ↵ Liu G , Geurts AM , Yae K , Srinivasan AR , Fahrenkrug SC , Largaespada DA , Takeda J , Horie K , Olson WK , Hackett PB . 2005 . Target-site Preferences of Sleeping Beauty Transposons . J. Mol. Biol . 346 : 161 – 173 . OpenUrl CrossRef PubMed Web of Science ↵ Liu Q , Ou X hong , Compton SG , Yang D rong . 2011 . Chromosome numbers are not fixed in Agaonidae (Hymenoptera: Chalcidoidea) . Symbiosis 53 : 131 – 137 . OpenUrl CrossRef ↵ Lynch M , Conery JS . 2003 . The Origins of Genome Complexity . Science 302 : 1401 – 1404 . OpenUrl Abstract / FREE Full Text ↵ Machado CA , Jousselin E , Kjellberg F , Compton SG , Herre EA . 2001 . Phylogenetic relationships, historical biogeography and character evolution of fig-pollinating wasps . Proc. R. Soc. Lond. B Biol. Sci . 268 : 685 – 694 . OpenUrl CrossRef GeoRef PubMed ↵ Mancera E , Bourgon R , Brozzi A , Huber W , Steinmetz LM . 2008 . High-resolution mapping of meiotic crossovers and noncrossovers in yeast . Nature 454 : 479 – 485 . OpenUrl CrossRef PubMed Web of Science ↵ Manni M , Berkeley MR , Seppey M , Zdobnov EM . 2021 . BUSCO: Assessing Genomic Data Quality and Beyond . Curr. Protoc . 1 : e323 . OpenUrl CrossRef ↵ Maumus F , Fiston-Lavier A-S , Quesneville H . 2015 . Impact of transposable elements on insect genomes and biology . Curr. Opin. Insect Sci . 7 : 30 – 36 . OpenUrl CrossRef PubMed ↵ Mirchandani CD , Shultz AJ , Thomas GWC , Smith SJ , Baylis M , Arnold B , Corbett-Detig R , Enbody E , Sackton TB . 2024 . A Fast, Reproducible, High-throughput Variant Calling Workflow for Population Genomics . Mol. Biol. Evol . 41 :msad270. ↵ Mirouze M , Lieberman-Lazarovich M , Aversano R , Bucher E , Nicolet J , Reinders J , Paszkowski J . 2012 . Loss of DNA methylation affects the recombination landscape in Arabidopsis . Proc. Natl. Acad. Sci . 109 : 5880 – 5885 . OpenUrl Abstract / FREE Full Text ↵ Molbo D , Machado CA , Herre EA , Keller L . 2004 . Inbreeding and population structure in two pairs of cryptic fig wasp species . Mol. Ecol . 13 : 1613 – 1623 . OpenUrl CrossRef PubMed ↵ Molbo D , Machado CA , Sevenster JG , Keller L , Herre EA . 2003 . Cryptic species of fig-pollinating wasps: Implications for the evolution of the fig–wasp mutualism, sex allocation, and precision of adaptation . Proc. Natl. Acad. Sci . 100 : 5867 – 5872 . OpenUrl Abstract / FREE Full Text ↵ Montgomery E , Charlesworth B , Langley CH . 1987 . A test for the role of natural selection in the stabilization of transposable element copy number in a population of Drosophila melanogaster . Genet. Res . 49 : 31 – 41 . OpenUrl CrossRef PubMed Web of Science ↵ Muller HJ . 1964 . The relation of recombination to mutational advance . Mutat. Res. Mol. Mech. Mutagen . 1 : 2 – 9 . OpenUrl CrossRef ↵ Mun S , Kim S , Lee W , Kang K , Meyer TJ , Han B-G , Han K , Kim H-S . 2021 . A study of transposable element-associated structural variations (TASVs) using a de novo-assembled Korean genome . Exp. Mol. Med . 53 : 615 – 630 . OpenUrl CrossRef PubMed ↵ Munasinghe M , Read A , Stitzer MC , Song B , Menard CC , Ma KY , Brandvain Y , Hirsch CN , Springer N . 2023 . Combined analysis of transposable elements and structural variation in maize genomes reveals genome contraction outpaces expansion . PLOS Genet . 19 : e1011086 . OpenUrl CrossRef PubMed ↵ Nambiar M , Smith GR . 2016 . Repression of harmful meiotic recombination in centromeric regions . Semin. Cell Dev. Biol . 54 : 188 – 197 . OpenUrl CrossRef PubMed ↵ Nei M , Li WH . 1979 . Mathematical model for studying genetic variation in terms of restriction endonucleases . Proc. Natl. Acad. Sci . 76 : 5269 – 5273 . OpenUrl Abstract / FREE Full Text ↵ Ninova M , Godneeva B , Chen Y-CA , Luo Y , Prakash SJ , Jankovics F , Erdélyi M , Aravin AA , Tóth KF . 2020 . The SUMO Ligase Su(var)2-10 Controls Hetero- and Euchromatic Gene Expression via Establishing H3K9 Trimethylation and Negative Feedback Regulation . Mol. Cell 77 : 571 – 585 .e4. OpenUrl CrossRef PubMed ↵ Oggenfuss U , Badet T , Wicker T , Hartmann FE , Singh NK , Abraham L , Karisto P , Vonlanthen T , Mundt C , McDonald BA , et al. 2021 . A population-level invasion by transposable elements triggers genome expansion in a fungal pathogen . Weigel D , Mirouze M , Joly-Lopez Z , Quadrana L , editors. eLife 10 : e69249 . OpenUrl CrossRef PubMed ↵ Ou S , Jiang N . 2018 . LTR_retriever: A Highly Accurate and Sensitive Program for Identification of Long Terminal Repeat Retrotransposons . Plant Physiol . 176 : 1410 – 1422 . OpenUrl Abstract / FREE Full Text ↵ Ou S , Su W , Liao Y , Chougule K , Agda JRA , Hellinga AJ , Lugo CSB , Elliott TA , Ware D , Peterson T , et al. 2019 . Benchmarking transposable element annotation methods for creation of a streamlined, comprehensive pipeline . Genome Biol . 20 : 275 . OpenUrl CrossRef PubMed ↵ Pessia E , Popa A , Mousset S , Rezvoy C , Duret L , Marais GAB . 2012 . Evidence for Widespread GC-biased Gene Conversion in Eukaryotes . Genome Biol. Evol . 4 : 675 – 682 . OpenUrl CrossRef PubMed ↵ Petersen M , Armisén D , Gibbs RA , Hering L , Khila A , Mayer G , Richards S , Niehuis O , Misof B . 2019 . Diversity and evolution of the transposable element repertoire in arthropods with particular reference to insects . BMC Ecol. Evol . 19 : 11 . OpenUrl CrossRef ↵ Poplin R , Chang P-C , Alexander D , Schwartz S , Colthurst T , Ku A , Newburger D , Dijamco J , Nguyen N , Afshar PT , et al. 2018 . A universal SNP and small-indel variant caller using deep neural networks . Nat. Biotechnol . 36 : 983 – 987 . OpenUrl CrossRef PubMed ↵ Provataris P , Meusemann K , Niehuis O , Grath S , Misof B . 2018 . Signatures of DNA Methylation across Insects Suggest Reduced DNA Methylation Levels in Holometabola . Genome Biol. Evol . 10 : 1185 – 1197 . OpenUrl CrossRef PubMed ↵ Quesneville H . 2020 . Twenty years of transposable element analysis in the Arabidopsis thaliana genome . Mob. DNA 11 : 28 . OpenUrl CrossRef PubMed ↵ Quinlan AR , Hall IM . 2010 . BEDTools: a flexible suite of utilities for comparing genomic features . Bioinformatics 26 : 841 – 842 . OpenUrl CrossRef PubMed Web of Science ↵ R Core Team . 2023 . R: A Language and Environment for Statistical Computing . Vienna, Austria : R Foundation for Statistical Computing Available from: https://www.R-project.org/ ↵ Ranallo-Benavidez TR , Jaron KS , Schatz MC . 2020 . GenomeScope 2.0 and Smudgeplot for reference-free profiling of polyploid genomes . Nat. Commun . 11 : 1432 . OpenUrl CrossRef PubMed ↵ Rhie A , Walenz BP , Koren S , Phillippy AM . 2020 . Merqury: reference-free quality, completeness, and phasing assessment for genome assemblies . Genome Biol . 21 : 245 . OpenUrl CrossRef PubMed ↵ Rousselle M , Laverré A , Figuet E , Nabholz B , Galtier N . 2018 . Influence of Recombination and GC-biased Gene Conversion on the Adaptive and Nonadaptive Substitution Rate in Mammals versus Birds . Mol. Biol. Evol . 36 : 458 . OpenUrl ↵ Saha P , Mishra RK . 2019 . Heterochromatic hues of transcription-the diverse roles of noncoding transcripts from constitutive heterochromatin . FEBS J . 286 : 4626 – 4641 . OpenUrl CrossRef PubMed ↵ Sarda S , Zeng J , Hunt BG , Yi SV . 2012 . The Evolution of Invertebrate Gene Body Methylation . Mol. Biol. Evol . 29 : 1907 – 1916 . OpenUrl CrossRef PubMed Web of Science ↵ Sarkies P . 2023 . Evolution beyond DNA: epigenetic drivers for evolutionary change? BMC Biol . 21 : 272 . OpenUrl CrossRef PubMed ↵ Sayers EW , Bolton EE , Brister JR , Canese K , Chan J , Comeau DC , Connor R , Funk K , Kelly C , Kim S , et al. 2022 . Database resources of the national center for biotechnology information . Nucleic Acids Res . 50 : D20 – D26 . OpenUrl CrossRef PubMed ↵ Schmitz RJ , Schultz MD , Lewsey MG , O’Malley RC , Urich MA , Libiger O , Schork NJ , Ecker JR . 2011 . Transgenerational Epigenetic Instability Is a Source of Novel Methylation Variants . Science 334 : 369 – 373 . OpenUrl Abstract / FREE Full Text ↵ Schrader L , Kim JW , Ence D , Zimin A , Klein A , Wyschetzki K , Weichselgartner T , Kemena C , Stökl J , Schultner E , et al. 2014 . Transposable element islands facilitate adaptation to novel environments in an invasive species . Nat. Commun . 5 : 5495 . OpenUrl CrossRef PubMed ↵ Sessegolo C , Burlet N , Haudry A . 2016 . Strong phylogenetic inertia on genome size and transposable element content among 26 species of flies . Biol. Lett . 12 : 20160407 . OpenUrl CrossRef PubMed ↵ Shen W , Le S , Li Y , Hu F. 2016 . SeqKit: A Cross-Platform and Ultrafast Toolkit for FASTA/Q File Manipulation . PLOS ONE 11 : e0163962 . OpenUrl CrossRef PubMed ↵ Shi J , Liang C . 2019 . Generic Repeat Finder: A High-Sensitivity Tool for Genome-Wide De Novo Repeat Detection . Plant Physiol . 180 : 1803 . OpenUrl Abstract / FREE Full Text ↵ Sigman MJ , Slotkin RK . 2016 . The First Rule of Plant Transposable Element Silencing: Location, Location, Location . Plant Cell 28 : 304 – 313 . OpenUrl Abstract / FREE Full Text ↵ Stanke M , Keller O , Gunduz I , Hayes A , Waack S , Morgenstern B . 2006 . AUGUSTUS: ab initio prediction of alternative transcripts . Nucleic Acids Res . 34 : W435 – W439 . OpenUrl CrossRef PubMed Web of Science ↵ Stritt C , Gordon SP , Wicker T , Vogel JP , Roulin AC . 2018 . Recent Activity in Expanding Populations and Purifying Selection Have Shaped Transposable Element Landscapes across Natural Accessions of the Mediterranean Grass Brachypodium distachyon . Genome Biol. Evol . 10 : 304 – 318 . OpenUrl CrossRef PubMed ↵ Su W , Gu X , Peterson T . 2019 . TIR-Learner, a New Ensemble Method for TIR Transposable Element Annotation, Provides Evidence for Abundant New Transposable Elements in the Maize Genome . Mol. Plant 12 : 447 – 460 . OpenUrl CrossRef PubMed ↵ Sultana T , Zamborlini A , Cristofari G , Lesage P . 2017 . Integration site selection by retroviruses and transposable elements in eukaryotes . Nat. Rev. Genet . 18 : 292 – 308 . OpenUrl CrossRef PubMed ↵ Toga K , Sakamoto T , Kanda M , Tamura K , Okuhara K , Tabunoki H , Bono H . 2024 . Long-read genome assembly of the Japanese parasitic wasp Copidosoma floridanum (Hymenoptera: Encyrtidae) . G3 Bethesda Md 14 :jkae127. ↵ Torres-Garcia S , Yaseen I , Shukla M , Audergon PNCB , White SA , Pidoux AL , Allshire RC . 2020 . Epigenetic gene silencing by heterochromatin primes fungal resistance . Nature 585 : 453 – 458 . OpenUrl CrossRef PubMed ↵ Trizna M . 2020 . assembly_stats 0.1.4 . Available from : doi: 10.5281/zenodo.3968775 OpenUrl CrossRef ↵ Visser I , Speekenbrink M . 2010 . depmixS4: An R Package for Hidden Markov Models . J. Stat. Softw . 36 : 1 – 21 . OpenUrl CrossRef PubMed ↵ Vurture GW , Sedlazeck FJ , Nattestad M , Underwood CJ , Fang H , Gurtowski J , Schatz MC . 2017 . GenomeScope: fast reference-free genome profiling from short reads . Bioinformatics 33 : 2202 – 2204 . OpenUrl CrossRef PubMed ↵ Wallberg A , Bunikis I , Pettersson OV , Mosbech M-B , Childers AK , Evans JD , Mikheyev AS , Robertson HM , Robinson GE , Webster MT . 2019 . A hybrid de novo genome assembly of the honeybee, Apis mellifera, with chromosome-length scaffolds . BMC Genomics 20 : 275 . OpenUrl CrossRef PubMed ↵ Wang R , Yang Y , Jing Y , Segar ST , Zhang Y , Wang G , Chen J , Liu Q-F , Chen S , Chen Y , et al. 2021 . Molecular mechanisms of mutualistic and antagonistic interactions in a plant– pollinator association. Nat . Ecol. Evol . 5 : 974 – 986 . OpenUrl ↵ Waterston J . 1921 . On some Bornean Fig-insects (Agaonidae–Hymenoptera Chalcidoidea) . Bull. Entomol. Res . 12 : 35 – 40 . OpenUrl CrossRef ↵ Weiblen GD . 2002 . How to be a fig wasp . Annu. Rev. Entomol . 47 : 299 – 330 . OpenUrl CrossRef PubMed Web of Science ↵ Wicker T , Sabot F , Hua-Van A , Bennetzen JL , Capy P , Chalhoub B , Flavell A , Leroy P , Morgante M , Panaud O , et al. 2007 . A unified classification system for eukaryotic transposable elements . Nat. Rev. Genet . 8 : 973 – 982 . OpenUrl CrossRef PubMed ↵ Widen SA , Bes IC , Koreshova A , Pliota P , Krogull D , Burga A . 2023 . Virus-like transposons cross the species barrier and drive the evolution of genetic incompatibilities . Science 380 : eade0705 . OpenUrl CrossRef PubMed ↵ Wiebes J . 1995 . Agaonidae (Hymenoptera Chalcidoidea) and Ficus (Moraceae): fig wasps and their figs, XV (Meso-American Pegoscapus) . In: Vol. 98 . p. 167 – 183 . OpenUrl ↵ Wilson R , Bourgeois ML , Perez M , Sarkies P . 2023 . Fluctuations in chromatin state at regulatory loci occur spontaneously under relaxed selection and are associated with epigenetically inherited variation in C. elegans gene expression . PLOS Genet . 19 : e1010647 . OpenUrl CrossRef PubMed ↵ Wright SI. 2017 . Evolution of Genome Size . In: eLS . John Wiley & Sons, Ltd . p. 1 – 6 . Available from: https://onlinelibrary.wiley.com/doi/abs/10.1002/9780470015902.a0023983 ↵ Wu X , Bhatia N , Grozinger CM , Yi SV . 2022 . Comparative studies of genomic and epigenetic factors influencing transcriptional variation in two insect species . G3 GenesGenomesGenetics 12 :jkac230. ↵ Xiao J , Wei Xianqin , Zhou Y , Xin Z , Miao Y , Hou H , Li J , Zhao D , Liu J , Chen R , et al. 2021 . Genomes of 12 fig wasps provide insights into the adaptation of pollinators to fig syconia . J. Genet. Genomics 48 : 225 – 236 . OpenUrl CrossRef PubMed ↵ Xiao J-H , Yue Z , Jia L-Y , Yang X-H , Niu L-H , Wang Z , Zhang P , Sun B-F , He S-M , Li Z , et al. 2013 . Obligate mutualism within a host drives the extreme specialization of a fig wasp genome . Genome Biol . 14 : R141 . OpenUrl CrossRef PubMed ↵ Xiong W , He L , Lai J , Dooner HK , Du C. 2014 . HelitronScanner uncovers a large overlooked cache of Helitron transposons in many plant genomes . Proc. Natl. Acad. Sci . 111 : 10263 – 10268 . OpenUrl Abstract / FREE Full Text ↵ Xu H , Ye X , Yang Yajun , Yang Yi , Sun YH , Mei Y , Xiong S , He K , Xu L , Fang Q , et al. 2021 . Comparative Genomics Sheds Light on the Convergent Evolution of Miniaturized Wasps . Mol. Biol. Evol . 38 : 5539 – 5554 . OpenUrl CrossRef PubMed ↵ Xu Z , Wang H . 2007 . LTR_FINDER: an efficient tool for the prediction of full-length LTR retrotransposons . Nucleic Acids Res . 35 : W265 – W268 . OpenUrl CrossRef PubMed Web of Science ↵ Yang S , Wang L , Huang J , Zhang X , Yuan Y , Chen J-Q , Hurst LD , Tian D . 2015 . Parent– progeny sequencing indicates higher mutation rates in heterozygotes . Nature 523 : 463 – 467 . OpenUrl CrossRef PubMed ↵ Yasuhara JC , DeCrease CH , Wakimoto BT . 2005 . Evolution of heterochromatic genes of Drosophila . Proc. Natl. Acad. Sci . 102 : 10958 – 10963 . OpenUrl Abstract / FREE Full Text ↵ Yelina NE , Choi K , Chelysheva L , Macaulay M , de Snoo B , Wijnker E , Miller N , Drouaud J , Grelon M , Copenhaver GP , et al. 2012 . Epigenetic remodeling of meiotic crossover frequency in Arabidopsis thaliana DNA methyltransferase mutants . PLoS Genet . 8 : e1002844 . OpenUrl CrossRef PubMed ↵ Yuan H , Huang Y , Mao Y , Zhang N , Nie Y , Zhang X , Zhou Y , Mao S. 2021 . The Evolutionary Patterns of Genome Size in Ensifera (Insecta: Orthoptera) . Front. Genet . [Internet] 12 . Available from: https://www.frontiersin.org/journals/genetics/articles/10.3389/fgene.2021.693541/full ↵ Yun T , Li H , Chang P-C , Lin MF , Carroll A , McLean CY . 2021 . Accurate, scalable cohort variant calls using DeepVariant and GLnexus . Bioinformatics 36 : 5582 – 5589 . OpenUrl CrossRef PubMed ↵ Zamudio N , Barau J , Teissandier A , Walter M , Borsos M , Servant N , Bourc’his D . 2015 . DNA methylation restrains transposons from adopting a chromatin signature permissive for meiotic recombination . Genes Dev . 29 : 1256 – 1270 . OpenUrl Abstract / FREE Full Text ↵ Zhang R-G , Li G-Y , Wang X-L , Dainat J , Wang Z-X , Ou S , Ma Y . 2022 . TEsorter: An accurate and fast method to classify LTR-retrotransposons in plant genomes . Hortic. Res . 9 :uhac017. ↵ Zhang X , Wang G , Zhang S , Chen S , Wang Y , Wen P , Ma X , Shi Y , Qi R , Yang Y , et al. 2020 . Genomes of the Banyan Tree and Pollinator Wasp Provide Insights into Fig-Wasp Coevolution . Cell 183 : 875 – 889 .e17. OpenUrl CrossRef PubMed ↵ Zuo B , Nneji LM , Sun Y-B . 2023 . Comparative genomics reveals insights into anuran genome size evolution . BMC Genomics 24 : 379 . OpenUrl CrossRef PubMed View the discussion thread. Back to top Previous Next Posted April 24, 2025. Download PDF Supplementary Material Email Thank you for your interest in spreading the word about bioRxiv. NOTE: Your email address is requested solely to identify you as the sender of this article. Your Email * Your Name * Send To * Enter multiple addresses on separate lines or separate them with commas. You are going to email the following Evolutionary Consequences of Unusually Large Pericentric TE-rich Regions in the Genome of a Neotropical Fig Wasp Message Subject (Your Name) has forwarded a page to you from bioRxiv Message Body (Your Name) thought you would like to see this page from the bioRxiv website. Your Personal Message CAPTCHA This question is for testing whether or not you are a human visitor and to prevent automated spam submissions. Share Evolutionary Consequences of Unusually Large Pericentric TE-rich Regions in the Genome of a Neotropical Fig Wasp Zexuan Zhao , Kevin Quinteros , Carlos A. Machado bioRxiv 2025.04.20.649723; doi: https://doi.org/10.1101/2025.04.20.649723 Share This Article: Copy Citation Tools Evolutionary Consequences of Unusually Large Pericentric TE-rich Regions in the Genome of a Neotropical Fig Wasp Zexuan Zhao , Kevin Quinteros , Carlos A. Machado bioRxiv 2025.04.20.649723; doi: https://doi.org/10.1101/2025.04.20.649723 Citation Manager Formats BibTeX Bookends EasyBib EndNote (tagged) EndNote 8 (xml) Medlars Mendeley Papers RefWorks Tagged Ref Manager RIS Zotero Tweet Widget Facebook Like Google Plus One Subject Area Genomics Subject Areas All Articles Animal Behavior and Cognition (7622) Biochemistry (17650) Bioengineering (13871) Bioinformatics (41881) Biophysics (21424) Cancer Biology (18566) Cell Biology (25461) Clinical Trials (138) Developmental Biology (13365) Ecology (19866) Epidemiology (2067) Evolutionary Biology (24290) Genetics (15590) Genomics (22476) Immunology (17713) Microbiology (40331) Molecular Biology (17148) Neuroscience (88473) Paleontology (666) Pathology (2827) Pharmacology and Toxicology (4816) Physiology (7635) Plant Biology (15114) Scientific Communication and Education (2044) Synthetic Biology (4286) Systems Biology (9815) Zoology (2268)
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.