TEgenomeSimulator: A Flexible Framework for Simulating Genomes with Configurable Transposable Element Landscapes

preprint OA: closed CC-BY-NC-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Transposable elements (TEs) are major contributors to genome structure and evolution. However, our ability to study them is limited by the difficulty of annotating and curating them, especially in non-model organisms. This challenge is compounded by a lack of ground-truth datasets for benchmarking, which are nearly impossible to generate through manual curation alone. To overcome this critical limitation, we developed TEgenomeSimulator, a flexible framework for generating synthetic genomes with configurable TE landscapes. TEgenomeSimulator supports both randomly generated and biologically derived backbone sequences, enabling the modeling of TE insertions under diverse structural and evolutionary contexts. Benchmarking against existing simulators demonstrated that TEgenomeSimulator reproduces realistic TE composition, sequence divergence, and integrity distributions while offering greater flexibility in modeling chromosome-level structure. Its modular design offers a tuneable continuum between biological fidelity and experimental control, facilitating systematic benchmarking, algorithm development, and evolutionary modeling of TE dynamics in silico , filling a major gap in the field. The source code of TEgenomeSimulator and the scripts for this paper are available at https://github.com/Plant-Food-Research-Open/TEgenomeSimulator .
Full text 56,598 characters · extracted from oa-pdf · 9 sections · click to expand

Abstract

15 Transposable elements (TEs) are major contributors to genome structure and evolution. 16 However, our ability to study them is limited by the difficulty of annotating and curating them, 17 especially in non-model organisms. This challenge is compounded by a lack of ground-truth 18 datasets for benchmarking, which are nearly impossible to generate through manual curation 19 alone. To overcome this critical limitation, we developed TEgenomeSimulator, a flexible 20 framework for generating synthetic genomes with configurable TE landscapes. 21 TEgenomeSimulator supports both randomly generated and biologically derived backbone 22 sequences, enabling the modeling of TE insertions under diverse structural and evolutionary 23 contexts. Benchmarking against existing simulators demonstrated that TEgenomeSimulator 24 reproduces realistic TE composition, sequence divergence, and integrity distributions while 25 offering greater flexibility in modeling chromosome-level structure. Its modular design offers 26 a tuneable continuum between biological fidelity and experimental control, facilitating 27 systematic benchmarking, algorithm development, and evolutionary modeling of TE 28 dynamics in silico, filling a major gap in the field. The source code of TEgenomeSimulator 29 and the scripts for this paper are available at https://github.com/Plant-Food-Research-30 Open/TEgenomeSimulator. 31 .CC-BY-NC 4.0 International licenseperpetuity. It is made available under a preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint

Introduction

32 Transposable elements (TEs) are abundant components in eukaryotic genomes. Once 33 regarded as parasitic DNA, TEs are now recognized as major contributors to gene 34 regulation, genome architecture, and evolutionary processes [1,2]. TEs also impact 35 agronomic traits, such as plant-pathogen interactions and environmental responses [3], and 36 may act as “silent” regulators of chromosomal architecture and post-zygotic barriers that 37 drive speciation [4]. Consistent with this view, cross-taxon surveys demonstrate that genome 38 size variation can be largely explained by TE accumulation [5-9]. TEs have also been 39 implicated in centromere structure evolution [6,7] and higher-order chromatin architecture 40 [12]. Together, these discoveries establish TEs not only as regulators of genes but also as 41 major catalysts of genome evolution, underscoring the need for high-quality assemblies to 42 accurately resolve repetitive DNA. 43 Advances in long-read sequencing technologies and assembly algorithms have markedly 44 enhanced genome assemblies and the resolution of repetitive regions, leading to increased 45 genome and pan-genome assemblies across taxa. Recent pangenome studies in Brassica, 46 grapevine and potato highlight TE enrichment in dispensable genomic regions linked to 47 resistance gene family expansion, regulatory variation, and reproductive isolation [13–15], 48 further solidifying the structural and functional impact of TEs and their critical role in genome 49 evolution. Despite these advances, TE annotation and interpretation remain underprioritized 50 in many genome projects, in only a few cases have domain experts filled this knowledge gap 51 [11]. 52 The detection and annotation of TEs within assembled genomes has benefited from 53 bioinformatic tool developments [12]; however, there are still major challenges to generating 54 quality annotation, especially for non-model organisms. TEs vary greatly in repetitiveness 55 and have undergone mutation, nested TE insertion and ectopic recombination, complicating 56 the differentiation of aged TE copies from non-TE sequences and hindering accurate 57 reconstruction of the consensus TE sequences [12]. Besides, ground–truth datasets are 58 especially scarce for non-model organisms, and only a small number have had any 59 validation through other ‘omics’ data sources [13,14]. Dfam is a TE database with an open 60 framework for community contribution and transparency in curated and non-curated datasets 61 [16]. As of March 2025, only 0.6% of the TE families in Dfam (release 3.9) have been 62 manually curated, predominantly from mammals. Most curated libraries rely on labor-63 intensive manual curation [17]. Consequently, simulated genomes with traceable TE 64 insertions can address this bottleneck by providing benchmarks for evaluating TE detection 65 and annotation. 66 Simulating TE sequence at scale that reflects the breadth of complexity and diversity of TE 67 composition is essential for benchmarking TE detection tools and modeling TE activity over 68 evolutionary time. Synthetic genomes with deliberately designed TE architectures, including 69 nested insertions, target site duplications (TSDs) and extreme cases of sequence 70 divergence and polymorphisms, provide a controlled framework to probe the limits of TE 71 detection and classification methods. This is particularly crucial in the era of telomere-to-72 telomere assemblies and pangenome-based annotations. By generating such 73 comprehensive and customizable datasets, researchers can systematically evaluate 74 performance under realistic and challenging conditions while also exploring the evolutionary 75 processes shaping genome architecture. A variety of TE simulators have been developed to 76 address some of these needs. For example, denovoTE-eval [18] synthesizes random 77 .CC-BY-NC 4.0 International licenseperpetuity. It is made available under a preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint sequences with user-specified lengths and GC content, requiring an input table containing 78 TE families and parameters to model mutagenesis, sequence decay, and random insertions. 79 However, it can simulate only a single random sequence at a time and offers limited flexibility 80 for adjusting mutagenesis parameters when working with large TE libraries. By relying on 81 predefined TE architectures, SimulaTE [19] can recapitulate intricate TE landscapes and 82 simulate sequencing reads, but it remains limited in its capability to simulate novel 83 evolutionary events that extend beyond the user-provided genome architecture. SLiM [20] 84 excels at modeling selection and demography but represents TEs as single units, each like a 85 single nucleotide variant, restricting exploration of sequence-level architectures and genome 86 size changes. GARLIC [21] generates realistic intergenic sequences based on k-mer 87 composition, GC content, and repeat profiles but lacks flexibility in TE mutagenesis profiles 88 and does not capture the coexistence of coding and non-coding regions. While PrinTE is 89 designed for forward-time simulation on TE mobilization, mutation, and removal, it requires 90 multiple iterative attempts to infer and reproduce the dynamics of a specific TE type, with a 91 focus on LTR retrotransposon [22]. Despite the breadth of existing simulation tools in design, 92 few offer the flexibility to customize genome content while also modeling TE sequence 93 mutagenesis under diverse decay schemes. Even fewer support fine-grained, family-specific 94 controls or allow combinations of multiple modes for complex TE composition simulations. 95 To address these limitations, we developed TEgenomeSimulator (Fig. 1), which was built on 96 concepts introduced in denovoTE-eval. Extending beyond the capability of denovoTE-eval in 97 synthesizing single random sequences and performing random and nested TE insertions, 98 TEgenomeSimulator can synthesize genomes with multiple chromosomes of varying 99 complexity and offers fine-grained control over TE sequence populations for a more realistic 100 fragmentation pattern. TEgenomeSimulator departs from the uniform TSD range approach of 101 denovo-TEeval by modeling TSD length at the TE-superfamily level, a key feature for 102 structure-based TE detection and classification. It can also incorporate real genome 103 assemblies and their TE compositions for more realistic modeling. Furthermore, it allows 104 users to combine different simulation modes to introduce new TE activity beyond known TE 105 compositions, enabling the exploration of diverse evolutionary scenarios. A comparison with 106 existing simulators is available in Table S1. 107 In this paper, we demonstrate that TEgenomeSimulator generates more realistic synthetic 108 genomes than denovoTE-eval and GARLIC. We systematically illustrate how 109 parameterization of TE copy number, fragmentation, mean sequence identity, and standard 110 deviation (SD) influences simulation outcomes and their biological interpretation. The 111 resulting outputs enable robust benchmarking of TE detection and annotation methods. As a 112 proof of concept, we use TEgenomeSimulator to facilitate the evaluation of TE detection by 113 RepeatMasker [23]. TEgenomeSimulator is implemented as a reproducible and portable 114 Python package for Linux, installable via ‘pip’, Docker, or Apptainer. 115 116 117 .CC-BY-NC 4.0 International licenseperpetuity. It is made available under a preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint 118 Figure 1. TEgenomeSimulator flowchart. TEgenomeSimulator operates in three modes—119 Random Synthesized Genome (mode 0), Custom Genome (mode 1), and TE 120 Composition Approximation (mode 2)—shown in blue, grey, and sage, respectively. 121 Workflow components (rounded pink rectangles) are organized into three conceptual layers: 122 Inputs, Processes, and Outputs. Connections between components indicate the data flow 123 and processing steps taken by each mode, highlighting shared and mode-specific elements. 124 See Materials and Methods for details. 125 126

Methods

127 Implementation 128 Simulation Modes 129 Random Synthesized Genome mode (mode 0) synthesizes artificial chromosomes using 130 user-defined chromosome lengths and GC content, then inserts simulated TEs at random 131 positions. A TE library is used to produce TE copies subjected to random mutagenesis, 132 including nucleotide substitutions, single-base insertions and deletions (InDels), TSDs, and 133 fragmentation. 134 .CC-BY-NC 4.0 International licenseperpetuity. It is made available under a preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint Custom Genome mode (mode 1) accepts a user-provided ‘backbone’ genome with TEs 135 removed, or masks and removes TEs using RepeatMasker [23] (using ‘--to_mask’ option) 136 with a user-supplied TE library. Mutation and insertion proceed as in mode 0. 137 TE Composition Approximation mode (mode 2) also utilizes RepeatMasker to remove TEs 138 but additionally infers TE family abundance, nucleotide substitution and InDel rates, and 139 sequence integrity from the source genome, which are then used to parameterize TE 140 simulations. 141 Simulating TE Mutagenesis 142 From the TE library, TEgenomeSimulator extracts family classification and sequence 143 metadata and assigns per-family mutagenesis parameters. In modes 0 and 1, per-family TE 144 copy numbers are sampled from user-defined ranges. Sequence identity (I) is drawn from a 145 Gaussian distribution with user-specified mean and SD (modes 0 and 1) or estimated from 146 RepeatMasker output (mode 2), then adjusted following denovoTE-eval. TSD lengths 147 follows published observations [24], with random TSD sequences synthesized. Sequence 148 integrity in modes 0 and 1 is modeled using a beta distribution controlled by alpha and beta 149 (default α=0.5, β=0.7), producing predominantly fragmented elements, with 0.1% of copies 150 remaining intact by default. Sequence integrity in mode 2 is inferred from the source 151 genome. TE copies are randomly oriented and integrated. See Supplementary Methods for 152 detailed implementation. 153 Random and Nested TE Insertion 154 Nested insertions are generated in 0-30% of LTR retrotransposon copies, with insertion 155 targets and locations selected randomly. 156 Input and Output 157 TE libraries must follow RepeatMasker naming conventions (e.g, >ATCOPIA10#LTR/Copia) 158 and ideally contain full-length sequences (e.g. LTR-INT-LTR for retrotransposons). Split LTR 159 entries are automatically reassembled (see Supplementary Methods). Modes 1 and 2 160 require a genome FASTA file, whereas mode 0 requires a genome index CSV file specifying 161 chromosome name, length, and GC%. Outputs include the simulated genome, all TE copy 162 sequences, and a GFF file recording TE locations, identity, integrity, and nesting 163 relationships. 164 165 Simulator comparison analysis 166 Comparison 1: TEgenomeSimulator vs denovoTE-eval 167 TEgenomeSimulator’s mode 0 was used to synthesize a 10 Mb non-TE random sequence 168 with 35% GC content, followed by TE insertion using a composition defined by 169 TEgenomeSimulator (denoted as TEGS in Fig. 2). Copy numbers per family were set 170 between 1 and 10, with other mutagenesis parameters left at default values. To ensure 171 comparability, settings for size of non-TE backbone, GC%, TE copy number, and sequence 172 identity from TEgenomeSimulator were used for denovoTE-eval (denoted as dnvTEe in Fig. 173 2), while maintaining direct comparisons of other features. 174 Comparison 2: TEgenomeSimulator vs GARLIC 175 .CC-BY-NC 4.0 International licenseperpetuity. It is made available under a preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint Chromosome 1 of the Arabidopsis thaliana (TAIR10, NCBI accession ID GCF_000001735.4) 176 and the curated TE library from the EDTA GitHub repository 177 (https://github.com/oushujun/EDTA/blob/master/database/) were used as inputs. Two 178 genomes were simulated with TEgenomeSimulator: (i) a digital replica of TAIR10 179 chromosome 1 using mode 2 (m2 in Fig. 2); and (ii) a hybrid simulation in which the TE 180 composition inferred by mode 2 was imported into mode 0 to generate a randomly 181 synthesized backbone (m2+m0 in Fig. 2; see Supplementary Methods). 182 GARLIC simulation required four input files: (i) tandem repeat profiles generated by Tandem 183 Repeat Finder (TRF, v4.07) [25] on TAIR10 genome; (ii) gene annotations from Ensembl 184 release 61 (GFF3) converted to GARLIC’s tabular format using a custom script; (iii) 185 RepeatMasker’s alignments of TAIR10 genome using the same TE library; and (iv) the TE 186 library converted from FASTA to EMBL format. GARLIC’s ‘createModel.pl’ and 187 ‘createFakeSequence.pl’ scripts were used to generate a synthetic sequence matching 188 chromosome 1 length. 189 K-mer spectra (k = 6) were analyzed using a custom Python script treating reverse 190 complements as canonical. Shannon entropy (H) on nucleotide probabilities (p) was 191 calculated in 1 Kb sliding window with a step size of 1 bp according to the formula: 192 𝐻 = − $ 𝑝! log"(𝑝!)   !∈{&,(,),*} 193 194 Use-cases demonstration 195 A multi-species TE library was constructed by merging the curated TE library of A. thaliana, 196 Oryza sativa, and Zea mays obtained from EDTA, followed by redundancy removal using cd-197 hit-est with settings following Goubert et al. [17]. The resulting library, along with the TE-198 depleted TAIR10 genome, was used for simulation demonstrations using mode 1. The 199 TAIR10 genome was pre-processed using RepeatMasker and the curated A. thaliana TE 200 library to remove endogenous TEs. Simulation settings corresponding to Fig. 3 are 201 described below. 202 Impact of TE copy number on genome size (mode 1) 203 TE copy number ranges per family were set to 1-10, 5-100, 5-500, 5-1000 and 5-2000 using 204 the TE-depleted A. thaliana genome as input. 205 Impact of alpha and beta on TE integrity (mode 1) 206 For demonstration, TE integrity was modeled with a beta distribution controlled by five [α, β] 207 pairs: [0.5, 1], [0.5, 0.75], [0.5, 0.5], [0.75, 0.5], and [1, 0.5]. 208 Impact of mean sequence identity (mode 1) 209 Identity ranges were set to 90-100, 80-90, 70-80, 70-85, 70-90, 70-95, and 70-100. 210 Combined effects of sequence identity and standard deviation (mode 1) 211 Four scenarios were simulated (Fig. 3D) by varying mean identity ranges (10-80 or 90-100) 212 and SD (1-5 and 15-20) to represent different histories of TE activity (see Supplementary 213 Methods) 214 .CC-BY-NC 4.0 International licenseperpetuity. It is made available under a preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint Digital replicas (mode 2) 215 Genomes of A. thaliana (TAIR10; 120 Mb), O. sativa (AGIS1.0; NCBI accession ID 216 GCF_034140825.1; 386 Mb), Z. mays (Zm-B73-REFERENCE-NAM-5.0; NCBI accession ID 217 GCF_902167145.1; 2.183 Gb), Danio rerio (GRCz12tu; NCBI accession ID 218 GCF_049306965.1; 1.449 Gb) and Drosophila melanogaster (dm6; NCBI accession ID 219 GCF_000001215.4; 138 Mb) were directly used as input with default settings. Acquisition of 220 TE libraries for the mentioned plant species has been described previously. TE libraries for 221 D. rerio and D. melanogaster were downloaded from the RepeatModeler2 GitHub repository 222 (https://github.com/jmf422/TE_annotation/tree/master/benchmark_libraries/RM2). 223 Computational resource tests 224 Mode 0 was benchmarked using five configurations that varied the length of randomly 225 synthesized sequences with 35% GC prior to TE insertion. Four tests, generated a single 226 sequence of 100 Kb, 1 Mb, 10 Mb, or 100 Mb, and a fifth generated ten 100 Mb sequences 227 (1Gb in total). TE copy numbers per family were fixed at 1-10. 228 To test mode 1, the TE-depleted TAIR10 genome and curated A. thaliana TE library were 229 used, with TE copy number ranges of 1-10, 5-100, 5-500, 5-1000, and 5-2000. 230 Mode 2 resource usage was recorded for simulations of A. thaliana, O. sativa, Z. mays, D. 231 melanogaster, and D. rerio. 232 All tests were executed on the same SLURM-managed cluster. Each mode 0 and mode 1 233 test used a single CPU core. Each mode 2 test was assigned 20 CPU cores due to the 234 computational demands of running RepeatMasker. 235 TE recovery rate analysis 236 RepeatMasker was applied to simulated genomes from the four scenarios Fig. 3D using the 237 same TE library as in the simulations. TEgenomeSimulator’s TE annotation files (GFF) 238 served as ground truth and were compared to the BED files converted from RepeatMasker 239 output using Bedtools Intersect [26], applying an 80% reciprocal overlap cutoff. Recovery 240 rate (R) was calculated per integrity and identity bin (10 bins, ranging from 0 to 1) according 241 to the formula: 242 𝑅  =   𝑁𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑞𝑢𝑎𝑙𝑖𝑓𝑖𝑒𝑑 𝑜𝑣𝑒𝑟𝑙𝑎𝑝𝑠 𝑖𝑛 𝑏𝑖𝑛 𝑇𝑜𝑡𝑎𝑙 𝑔𝑟𝑜𝑢𝑛𝑑 − 𝑡𝑟𝑢𝑡ℎ 𝑇𝐸𝑠 𝑖𝑛 𝑏𝑖𝑛 243 244

Results

245 Simulator Comparisons 246 We compared the key features of TEgenomeSimulator to those of denovoTE-eval, GARLIC, 247 SimulaTE, and SLiM (Table S1). Unlike denovoTE-eval, TEgenomeSimulator enables the 248 synthesis of multi-chromosomal genomes within a single simulation run. It also supports 249 modeling of superfamily-specific TSD lengths, allows beta distribution or empirical 250 distribution to model fragmentation, and enables both empirical and combinatorial modeling 251 of TE composition through the sequential application of multiple simulation modes 252 (Supplementary Methods). In contrast to GARLIC, which analyzes genic and non-genic 253 profiles but only synthesizes non-genic regions, TEgenomeSimulator’s mode 2 preserves 254 gene models whose exons or introns are not identified and masked as TEs by 255 .CC-BY-NC 4.0 International licenseperpetuity. It is made available under a preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint RepeatMasker. Moreover, TEgenomeSimulator is uniquely capable of utilizing either 256 randomly generated or empirically derived non-TE backbones for TE insertion, providing 257 enhanced flexibility for the simulation of complex and biologically realistic TE landscapes 258 (Table S1). 259 We assessed TEgenomeSimulator against its closest peers, denovoTE-eval and GARLIC, 260 using two comparison designs (Fig. 2A-B). Comparison 1 (Fig. 2A) utilizes a 10 Mb TE-free 261 backbone, synthesized by TEgenomeSimulator, that is used by both TEgenomeSimulaor 262 and denovoTE-eval for subsequent TE insertion (see Methods). To maintain consistency, 263 the TE mutagenesis parameters of TEgenomeSimulator were reformatted and imported into 264 denovoTE-eval. Both tools yielded comparable TE superfamily composition and divergence 265 distributions (Fig. S1), validating the consistency in the transferred TE composition and 266 sequence identity parameters. We did, however, observed considerable variance in the 267 resulting genome size and composition. denovoTE-eval produced a genome approximately 268 5 Mb larger than TEgenomeSimulator, due to incorporating 7.6% more TE bases(Fig. 2C). 269 We also observed a marked difference in the integrity distributions of simulated TEs. 270 denovoTE-eval displayed flattened TE loci distributions at integrity ranges 0.4-0.7 and 0.7-271 0.9, with no fragmented element exceeding 0.9 (Fig. 2E). This gap between 0.9 and 1 arose 272 because denovoTE-eval caps integrity at 0.9 following fragmentation, while the even 273 patterns reflected its programmed two-tier sampling scheme: integrities 0.7-0.9 for sequence 274 <500 bp and 0.4-0.9 for sequences ≥500 bp. TEgenomeSimulator, however, models integrity 275 using a beta-distribution combined with a fixed fraction of intact TEs. This results in a 276 broader spectrum, spanning degraded to nearly intact elements, better reflecting the 277 sequence integrity observed in natural genomes (Fig. 2F). 278 For Comparison 2 (Fig. 2B), we used TEgenomeSimulator’s mode 2 to build a model of the 279 TE content of TAIR10 chromsome1 (chr1), which was then used as input for TE insertion into 280 TE-depleted chr1 (m2) or a simulated backbone matching the length of chr1 (m2+m0). The 281 resulting genomes were compared to simulations derived from GARLIC. GARLIC produced 282 genomes with a markedly higher proportion of TE sequence (25%) than the original chr1 283 (10.6%) and nearly three times as many TE insertions (Fig. S2A). Both the m2 (9.7%) and 284 m2+m0 (10.1%) genomes more closely matched the balance of TE and non-TE occupancy 285 observed in chr1 (Fig. 2D). GARLIC’s integrity profile was heavily skewed toward values of 286 0.9-1, in sharp contrast to the more realistic integrity ranges captured by 287 TEgenomeSimulator m2 and m2+m0 (Fig. S3). Sequence complexity analyses reinforced 288 these differences (Fig. 2G-I). GARLIC's k-mer profile displayed an extended tail in TE-only 289 regions, which is absent from both chr1 (Ori) and TEgenomeSimulator-synthesized 290 sequence (Fig. 2G). In non-TE regions, GARLIC and TEgenomeSimulator’s m2 closely 291 matched the reference k-mer profile, whereas the distortion in m2+m0 reflected the use of a 292 randomly synthesized non-TE backbone (Fig. 2G). We used Shannon entropy analysis to 293 determine variability within the data, we observed a central dip within chr1 (Ori) that was 294 successfully reproduced by TEgenomeSimulator’s m2 , while GARLIC and m2+m0 failed to 295 capture this feature (Fig. 2H). GARLIC’s outputs exhibited a narrower, higher entropy range 296 in non-TE regions and a lower entropy range in TE-only regions relative to the original 297 genome (Fig. 2I). For TEgenomeSimulator, the TE-only entropy range of m2 was shifted 298 compared to the reference, but its non-TE profile was preserved, while m2+m0 showed 299 markedly narrower non-TE distribution, again reflecting the random backbone synthesis 300 (Fig. 2I). 301 .CC-BY-NC 4.0 International licenseperpetuity. It is made available under a preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint 302 Figure 2. Simulator Comparisons. (A) Comparison 1: mode 0 of TEgenomeSimulator 303 (TEGS) and denovoTE-eval (dnvTEe) were compared using a random synthesized 10 Mb 304 sequence. The TE mutagenesis table (TE mut-table) generated by TEGS was imported to 305 dnvTEe to match TE composition for comparison. (B) Comparison 2: simulated TAIR10 306 chromosome 1 (chr1) generated by GARLIC, TEGS full mode 2 (m2), TEGS partial mode 2 307 followed by full mode 0 (m2+m0), and the original TAIR10 chr1. For m2+m0, the TE mut-308 table, TE-depleted chr1 length and GC content were acquired from m2 and used as inputs 309 .CC-BY-NC 4.0 International licenseperpetuity. It is made available under a preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint for m0. (C, D) Proportions of TE and non-TE sequences from Comparison 1 (C) and 2 (D). 310 (E, F) TE integrity distributions for dnvTEe (E) and TEGS (F) simulations. (G, H) K-mer (G) 311 and Shannon entropy (H) spectra of whole genomes, and the non-TE and TE only 312 sequences, colored by simulation method. (I) Shannon entropy ranges by sequence 313 category. Asterisks indicate p < 0.0001 (Wilcoxon Rank-Sum Test versus the original 314 genome). 315 316 Use cases 317 TE composition is shaped by multiple evolutionary processes, each leaving distinct genomic 318 signatures. To illustrate how TEgenomeSimulator captures these dynamics, we demonstrate 319 key configurations that reproduce major features of empirical TE landscapes. 320 Modeling basic TE features 321 TE accumulation is a primary driver of genome expansion [27–29]. Using the Custom 322 Genome Mode together with the TE-depleted A. thaliana (TAIR10) genome and the curated 323 Arabidopsis TE library (see Methods), we simulated genome expansions ranging from 324 under 200 Mb to over 1 Gb by increasing the maximum copy number per TE family from 10 325 to 2000 (Fig. 3A). This configuration directly models the impact of TE accumulation on 326 genome size. 327 TE integrity reflects the balance between recent insertions and long-term decay. In both 328 Random Synthesized Genome and Custom Genome modes, TE integrity is modelled using 329 a beta distribution across families. When α < β (both between 0 and 1), the distribution takes 330 on an asymmetrical U-shape with a peak near low integrity values (Fig. 3B), reflecting the 331 predominance of TE relics observed in empirical datasets [30]. Conversely, when β < α, the 332 distribution shifts toward high-integrity copies (Fig. 3B), reflecting recent evolutionary TE 333 activity. 334 Alongside modifying TE integrity, we can vary the range of mean sequence identity to reflect 335 mutation accumulation over time. By varying the range of mean sequence identity (70%-336 100%) while constraining copy numbers (5-100 per family), we generated divergence 337 distributions with progressively flattened peaks as identity ranges widened (Fig. 3C), 338 reproducing the expected signatures of sustained or overlapping TE amplification events 339 [31,32]. 340 341 Modeling TE composition under diverse evolutionary scenarios 342 The duration of TE bursts, which can be measured by the standard deviation (SD) of 343 sequence identity, also influences divergence distribution. A short-lived TE burst results in a 344 rapid deposition of nearly identical insertions. These insertions then gradually diverge 345 through accumulated mutation. In contrast, a long-lasting burst over evolutionary time leads 346 to the continuous deposition of similar TE sequences, causing a less synchronized 347 divergence pattern owing to overlapping degradation timelines. Moreover, ancient TE bursts 348 tend to show greater sequence deterioration due to prolonged exposure to mutational decay. 349 We can simulate genomes that represent different evolutionary phases of TE activity (Fig. 350 3D), offering genomic templates with distinct base-level TE composition that can serve as 351 burn-in states for more advanced simulations. As shown in Fig. 3D, the representation of TE 352 burst timeline and duration on the x-axis and y-axis, respectively, allows four distinct TE 353 divergence landscapes to emerge: I) high sequence identity with high SD, modeling recent 354 .CC-BY-NC 4.0 International licenseperpetuity. It is made available under a preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint TE bursts with ongoing or prolonged activity over time; II) low sequence identity with high 355 SD, modeling ancient and prolonged activity, or multiple independent bursts occurring over 356 an extended period; III) low sequence identity with low SD, reflecting ancient TE bursts that 357 were quickly inactivated, with TE activity remaining silenced since the initial event; IV) high 358 sequence identity with low SD, inferring very recent and rapid TE bursts, resulting in a 359 synchronized insertion of highly similar sequences. 360 We modelled these four scenarios using the Custom Genome mode of TEgenomeSimulator. 361 The identity and SD ranges were implemented as shown in Fig. 3D, with identity ranges of 362 90%–100% and 70%–80%, and SD ranges of 15–20 and 1–5. The TE copy number range 363 was fixed between 5 and 1000. The resulting sequence divergence profiles demonstrated a 364 narrowing of the peak as the SD range decreased from 15–20 to 1–5, indicating more 365 instantaneous bursts (Fig. 3E). Furthermore, shifting the identity range from 70–80% to 90–366 100% led to a leftward shift of the divergence peak (from ~0.25 to ~0.1) and an increase in 367 peak height. These results demonstrate that TEgenomeSimulator is well-suited for modeling 368 TE composition under diverse evolutionary scenarios. 369 370 371 .CC-BY-NC 4.0 International licenseperpetuity. It is made available under a preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint 372 Figure 3. Example use-case of TEgenomeSimulator. (A) Impact of altering TE copy 373 number ranges on simulated genome size (mode 0). (B) TE sequence integrity distribution 374 under different alpha and beta settings (mode 0). (C) TE sequence divergence distribution 375 under varying ranges of mean identity across TE families, with a fixed TE copy number 376 range of 5–100 (mode 0). (D) Four historical TE activity scenarios modeled by adjusting 377 mean identity (activity timing) and its standard deviation (SD; activity duration). (E) Resulting 378 TE divergence patterns associated with the settings in (C). (F) TE/nonTE sequence 379 proportions in original and simulated Arabidopsis thaliana genomes (mode 2). (G,H) 380 Sequence divergence (G) and integrity (H) in original and simulated A. thaliana genomes. 381 382 Digital replicas 383 TE composition in real genomes varies substantially across TE families and species 384 because of their complex and diverse evolutionary histories. This variability makes it 385 challenging to simulate realistic genome-wide TE landscapes. To address this, we used 386 TEgenomeSimulator’s TE Composition Approximation Mode to model the TE profiles of A. 387 thaliana (TAIR10), O. sativa, Z. mays, D. rerio, and D. melanogaster and synthesis replicas. 388 We observed the final genome size and proportion of TE occupancy in the digital replicas 389 .CC-BY-NC 4.0 International licenseperpetuity. It is made available under a preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint closely matched those observed in the original genomes (Fig. 3F, Fig. S4-S7). In terms of 390 TE loci and TE bases per superfamily, the simulated TE composition also aligned well with 391 the original genomes (Fig. S4-S8). 392 While the simulated divergence profiles of the replicas resembled those of the original 393 genomes, they were not identical: the fine-scale fluctuations present in the original genomes 394 were smoothed in the replicas due to the use of normal distributions in the simulator. This 395 approach, however, still provides a generalized distribution pattern that captures the major 396 trends in TE sequence evolution (Fig. 3G, Fig. S4-S7). 397 To replicate a real TE integrity pattern for each TE family, the empirical distribution of TE 398 length ratios (relative to consensus sequences) is extracted from the input genome. This 399 enables the simulator to reproduce the integrity distributions more authentically in the 400 resulting simulated genome (Fig. 3H, Fig. S4-S7). 401 These digital replicas provide a robust foundation for simulating future TE-related events, 402 such as new TE bloating or purging episodes, effectively serving as burn-in phases for 403 modeling genome evolution from the current state. 404 Computational resource usage 405 To evaluate the computational resource usage of TEgenomeSimulator, we conducted tests 406 across all three simulation modes (Fig. 4, Table S2). In mode 0, we observed that the total 407 length of randomly synthesized sequences significantly impacts memory usage, particularly 408 for sequences longer than 100 Mb. Synthesizing a large genome over 1 Gb, comprising 10 409 chromosomes of approximately 100 Mb each, required over 12 GB of memory (Fig. 4, Table 410 S2). In mode 1, increasing the TE copy number range per family had a more pronounced 411 impact on runtime than on memory usage (Fig. 4). Simulations with broader copy number 412 ranges took several hours to complete and required multiple gigabytes of memory, 413 depending on the settings. Notably, a simulation with copy number ranges of 5-2000 per 414 family consumed over 6 GB of memory and required nearly 10 hours to complete, resulting 415 in a simulated genome exceeding 1 Gb in size (Fig. 4, Table S2). In contrast, tests varying 416 TE sequence identity and integrity finished in under 3.5 minutes and used less than 1 GB of 417 memory (Table S2). In mode 2, both runtime and memory usage were substantially 418 increased in relation to the size and complexity of the input genomes used for TE 419 composition approximation (Fig. 4, Table S2). The increasing computational demand with 420 genome size appears to correlate with native TE content (Fig. S2-S5). 421 422 423 .CC-BY-NC 4.0 International licenseperpetuity. It is made available under a preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint 424 Figure 4. Computational resource usage of TEgenomeSimulator. Each bubble 425 represents a test run, with size proportional to the simulated genome. Mode 0 varied 426

Reference

sequence length (100 Kb-100 Mb). Mode 1 varied TE copy numbers ranges. 427 Mode 2 approximated TE compositions from five genomes: Arabidopsis thaliana (TAIR10), 428 Oryza sativa (O_sativa), Zea mays (Z_mays), Drosophila melanogaster (dm6), and Danio 429 rerio (D_rerio). 430 431 TE recovery rate test 432 TEgenomeSimulator can generate diverse TE insertions ideal for benchmarking TE 433 detection and annotation tools. We assessed the TE detection capability of RepeatMasker 434 using simulated TE insertions from the four scenarios described in Fig. 3D as ground truth 435 (Fig. 5A). TE recovery rate was measured as the proportion of simulated insertions correctly 436 identified by RepeatMasker, stratified by sequence identity and integrity (Fig. 5A). The 437 resulting heatmaps revealed a strong dependence on sequence identity, with higher 438 recovery for insertions of greater similarity to the consensus (Fig. 5B). In contrast, recovery 439 showed little or no clear dependency on integrity. Bar plots summarizing recovery rate 440 across identity and integrity slices confirmed these trends across the four scenarios: 441 recovery rates remained high at identity above 0.7 but declined progressively with greater 442 divergence, whereas variation across integrity levels lacked a consistent pattern. Although 443 the precise recovery values may be influenced by RepeatMasker’s parameter settings, 444 particularly the divergence and alignment thresholds, the results clearly demonstrate that TE 445 insertions with higher sequence divergence are more difficult to detect. 446 .CC-BY-NC 4.0 International licenseperpetuity. It is made available under a preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint 447 Fig. 5. TE annotation recovery rate of RepeatMasker. (A) TE annotation recovery 448 workflow, with ground-truth (green) and test (purple) annotations; numbered blue rectangles 449 indicate analysis steps (see Methods). (B) Heatmaps of recovery rate by sequence identity 450 and integrity for the four scenarios in Fig. 3D, with marginal bar graphs showing mean 451 recovery. Empty bins are shown in white. 452 453

Discussion

454 TEgenomeSimulator introduces a flexible, mode-based architecture that addresses critical 455

Limitations

in existing TE simulation tools by enabling both randomized simulations and 456 simulations based on real genomic architecture, thereby capturing structural and 457 compositional complexity characteristics of natural TE landscapes. 458 TEgenomeSimulator-generated TE populations display a broad and continuous range of 459 sequence integrity values by modeling either a beta or an empirical distribution, providing a 460 more biologically realistic material for testing annotation tools. While denovoTE-eval served 461 as a foundational model for the development of TEgenomeSimulator, it lacked flexibility in 462 .CC-BY-NC 4.0 International licenseperpetuity. It is made available under a preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint defining superfamily-specific TSDs and exhibited discontinuous integrity distributions due to 463 rigid random sampling. TEgenomeSimulator overcomes this limitation by introducing TE 464 superfamily modeling of TSDs and sequence divergence and integrity, producing outputs 465 that better reflect natural TE diversity and evolutionary behaviors of TE lineages observed in 466 real genomes. By comparison, GARLIC’s emphasis on modeling non-genic contexts is 467 useful for certain studies but results in overrepresentation of high-integrity TE copies and 468 reduced sequence entropy, limiting its value for representing complex genomic structure and 469 simulating evolutionary variation. Other simulation tools, such as SimulaTE and SLiM, 470 emphasize population dynamics or evolutionary events but lack fine control over structural 471 variation or sequence-level mutagenesis parameters that TEgenomeSimulator provides. 472 TEgenomeSimulator’s capability in combining different modes provides a tunable continuum 473 across a spectrum of evolutionary scenarios and genomic realism. For example, a mode 0 474 simulation is ideal for controlled benchmarking under a uniform backbone condition, where 475 the goal is to assess TE detection sensitivity without confounding genomic context effects. 476 Mode 2, approximation to realistic TE composition, best supports validation of annotation 477 tools with authentic genomic contexts, where chromosomal architecture, including 478 centromere, intergenic and intragenic regions, is preserved as the non-TE backbone. Mode 479 1 combines controllable TE composition and mutagenesis with a realistic non-TE backbone. 480 Further flexibility is introduced through the combination of modes (e.g. mode 2 + mode 0 or 481 mode 2 + mode 1; see examples in Supplementary Methods) uniquely enable hybrid 482 simulation approaches that leveraging the TE compositional realism of real genomes (mode 483 2) while expanding diversity through random synthetic non-TE backbones (mode 0) and 484 introducing controllable novel TE activity to a real genome (mode 1), mimicking TE bursts in 485 horizontal transfer or TE reactivation. Besides, combining TEgenomeSimulator’s digital 486 replicas with PrinTE’s forward evolution simulation facilitates investigation in complex 487 evolutionary history [22]. 488 The customizable mutagenesis parameters and compositional realism of 489 TEgenomeSimulator make it ideal for systematically evaluating tool sensitivity across identity 490 and integrity gradients, as demonstrated in our RepeatMasker studies. Hence, 491 TEgenomeSimulator could serve as a reference framework for benchmarking studies, 492 establishing standardized simulation conditions for reproducible TE tool comparisons. 493 Future development could integrate the positional bias of TE insertion, TE dynamics over 494 time, and epigenetic influences, bridging the gap between structural simulation and 495 evolutionary modeling. 496 Key points 497 - TEgenomeSimulator addresses the critical gap in large-scale genomic analysis by 498 providing traceable, simulated genomes that serve as essential “ground-truth” 499 benchmarks for evaluating the accuracy of TE annotation tools, particularly for non-500 model organisms where curated data is scarce. 501 - TEgenomeSimulator offers a unique methodological structure through three 502 interoperable modes that span a spectrum of complexity: from entirely synthetic 503 genomes (mode 0) and applying TEs to custom genome backbones (mode 1) to the 504 high-fidelity approximation of empirical TE landscapes derived from source genomes 505 (mode 2). 506 .CC-BY-NC 4.0 International licenseperpetuity. It is made available under a preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint - TEgenomeSimulator allows for versatile parameterization of TE mutagenesis, 507 including sequence identity, integrity, and copy number. The tool captures greater 508 evolutionary complexity and genomic realism than existing simulators such as 509 denovoTE-eval and GARLIC. 510 - While this manuscript demonstrates TEgenomeSimulator’s utility via RepeatMasker, 511 the framework is built to evaluate a wide range of TE detection and annotation 512 methods, uncovering how specific gradients of sequence conservation and structural 513 integrity impact performance across different algorithms. 514 - TEgenomeSimulator’s ability to combine simulation modes provides a tunable 515 continuum between evolutionary abstraction and genomic realism. This flexibility, 516 paired with its open-source design and compatibility as a “burn-in” input for forward-517 evolution simulators like PrinTE, ensures that it is a comprehensive resource for the 518 genomic research community. 519 520

Acknowledgements

521 We want to thank Drs Chen Wu, Julie Blommaert, Jason Shiller, and David Chagné for 522 reviewing this manuscript and providing insightful comments. 523 524 Author Contributions 525 Ting-Hsuan Chen (Conceptualization [lead], Software [lead], Analysis [lead], Writing — 526 original draft [lead]), Olivia Angelin-Bonnet (Software [supporting], Writing — review & editing 527 [supporting], James Bristow (Software [supporting], Writing — Review & editing 528 [supporting]), Christopher Benson (Writing — Review & editing [supporting]), Shujun Ou 529 (Writing — review & editing [supporting]), Cecilia Deng (Conceptualization [supporting], 530 Writing — review & editing [supporting], Supervision [supporting]), Susan Thomson 531 (Conceptualization [supporting], Funding acquisition [lead], Supervision [lead], Writing — 532 review & editing [lead]) 533 534 Conflict of interest 535 The authors declare that they have no competing interests. 536 Use of AI in this study 537 ChatGPT o4 and 5 were used to help draft and improve scripts during the development of 538 TEgenomeSimulator, and to suggest edits for clarity and grammar in the manuscript. 539 Funding 540 This research was supported by Plant & Food Research through the Kiwifruit Royalty 541 Investment Programme. 542 Data availability 543 TEgenomeSimulator is available at https://github.com/Plant-Food-Research-544 Open/TEgenomeSimulator under the GPL-3.0 license, and all codes for the analyses in this 545 .CC-BY-NC 4.0 International licenseperpetuity. It is made available under a preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint paper are available at https://github.com/Plant-Food-Research-546 Open/TEgenomeSimulator_publication_code under the MIT license. 547

Reference

548 1. Tossolini I, Mencia R, Arce AL, Manavella PA. The genome awakens: transposon-549 mediated gene regulation. Trends Plant Sci. Elsevier; 2025;30:857–71. 550 https://doi.org/10.1016/j.tplants.2025.02.005 551 2. Hayward A, Gilbert C. Transposable elements. Curr Biol. Elsevier; 2022;32:R904–9. 552 https://doi.org/10.1016/j.cub.2022.07.044 553 3. Zanini SF, Bayer PE, Wells R, Snowdon RJ, Batley J, Varshney RK, et al. Pangenomics in 554 crop improvement—from coding structural variations to finding regulatory variants with 555 pangenome graphs. Plant Genome. 2022;15:e20177. https://doi.org/10.1002/tpg2.20177 556 4. Cabral-de-Mello DC, Palacios-Gimenez OM. Repetitive DNAs: the ‘invisible’ regulators of 557 insect adaptation and speciation. Curr Opin Insect Sci. 2025;67:101295. 558 https://doi.org/10.1016/j.cois.2024.101295 559 5. Elliott TA, Gregory TR. Do larger genomes contain more diverse transposable elements? 560 BMC Evol Biol. 2015;15:69. https://doi.org/10.1186/s12862-015-0339-8 561 6. Xie L, Huang Y, Huang W, Shang L, Sun Y, Chen Q, et al. Genetic diversity and evolution 562 of rice centromeres. bioRxiv. 2024;2024.07.28.605524. 563 https://doi.org/10.1101/2024.07.28.605524 564 7. Naish M, Henderson IR. The structure, function, and evolution of plant centromeres. 565 Genome Res. 2024;34:161–78. https://doi.org/10.1101/gr.278409.123 566 8. Bayer PE, Golicz AA, Tirnaz S, Chan CK, Edwards D, Batley J. Variation in abundance of 567 predicted resistance genes in the Brassica oleracea pangenome. Plant Biotechnol J. 568 2019;17:789–800. https://doi.org/10.1111/pbi.13015 569 9. Cochetel N, Minio A, Guarracino A, Garcia JF, Figueroa-Balderas R, Massonnet M, et al. A 570 super-pangenome of the North American wild grape species. Genome Biol. 2023;24:290. 571 https://doi.org/10.1186/s13059-023-03133-2 572 10. Bozan I, Achakkagari SR, Anglin NL, Ellis D, Tai HH, Strömvik MV. Pangenome analyses 573 reveal impact of transposable elements and ploidy on the evolution of potato species. Proc 574 Natl Acad Sci. 2023;120:e2211117120. https://doi.org/10.1073/pnas.2211117120 575 11. Osmanski AB, Paulat NS, Korstian J, Grimshaw JR, Halsey M, Sullivan KAM, et al. 576 Insights into mammalian TE diversity through the curation of 248 genome assemblies. 577 Science. 2023;380:eabn1430. https://doi.org/10.1126/science.abn1430 578 12. Storer JM, Hubley R, Rosen J, Smit AFA. Methodologies for the De novo Discovery of 579 Transposable Element Families. Genes. Multidisciplinary Digital Publishing Institute; 580 2022;13:709. https://doi.org/10.3390/genes13040709 581 13. Bell EA, Butler CL, Oliveira C, Marburger S, Yant L, Taylor MI. Transposable element 582 annotation in non-model species: The benefits of species-specific repeat libraries using 583 semi-automated EDTA and DeepTE de novo pipelines. Mol Ecol Resour [Internet]. 2021 584 [cited 2021 Nov 23];n/a. https://doi.org/10.1111/1755-0998.13489 585 .CC-BY-NC 4.0 International licenseperpetuity. It is made available under a preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint 14. Berthelier J, Casse N, Daccord N, Jamilloux V, Saint-Jean B, Carrier G. A transposable 586 element annotation pipeline and expression analysis reveal potentially active elements in the 587 microalga Tisochrysis lutea. BMC Genomics. 2018;19:378. https://doi.org/10.1186/s12864-588 018-4763-1 589 15. Goerner-Potvin P, Bourque G. Computational tools to unmask transposable elements. 590 Nat Rev Genet. Nature Publishing Group; 2018;19:688–704. https://doi.org/10.1038/s41576-591 018-0050-x 592 16. Storer J, Hubley R, Rosen J, Wheeler TJ, Smit AF. The Dfam community resource of 593 transposable element families, sequence models, and genome annotations. Mob DNA. 594 2021;12:2. https://doi.org/10.1186/s13100-020-00230-y 595 17. Goubert C, Craig RJ, Bilat AF, Peona V, Vogan AA, Protasio AV. A beginner’s guide to 596 manual curation of transposable elements. Mob DNA. 2022;13:7. 597 https://doi.org/10.1186/s13100-021-00259-7 598 18. Rodriguez M, Makalowski W. Software evaluation for de novo detection of transposons. 599 Mob DNA. 2022;13:14. https://doi.org/10.1186/s13100-022-00266-2 600 19. Kofler R. SimulaTE: simulating complex landscapes of transposable elements of 601 populations. Bioinformatics. 2017;34:1419–20. https://doi.org/10.1093/bioinformatics/btx772 602 20. Haller BC, Messer PW. SLiM 4: Multispecies Eco-Evolutionary Modeling. Am Nat. 603 2023;201:E127–39. https://doi.org/10.1086/723601 604 21. Caballero J, Smit AFA, Hood L, Glusman G. Realistic artificial DNA sequences as 605 negative controls for computational genomics. Nucleic Acids Res. 2014;42:e99. 606 https://doi.org/10.1093/nar/gku356 607 22. Benson CW, Chen T-H, Thomson SJ, Deng CH, Ou S. PrinTE: A Forward Simulation 608 Framework for Studying the Role of Transposable Elements in Genome Expansion and 609 Contraction [Internet]. bioRxiv; 2025 [cited 2025 Nov 24]. p. 2025.06.04.657780. 610 https://doi.org/10.1101/2025.06.04.657780 611 23. Smit, AFA, Hubley, R, Green, P. RepeatMasker Open-4.0 [Internet]. 2013. 612 http://www.repeatmasker.org 613 24. Zhao D, Ferguson AA, Jiang N. What makes up plant genomes: The vanishing line 614 between transposable elements and genes. Biochim Biophys Acta BBA - Gene Regul Mech. 615 2016;1859:366–80. https://doi.org/10.1016/j.bbagrm.2015.12.005 616 25. Benson G. Tandem repeats finder: a program to analyze DNA sequences. Nucleic Acids 617 Res. 1999;27:573–80. https://doi.org/10.1093/nar/27.2.573 618 26. Quinlan AR, Hall IM. BEDTools: a flexible suite of utilities for comparing genomic 619 features. Bioinformatics. 2010;26:841–2. https://doi.org/10.1093/bioinformatics/btq033 620 27. Wells JN, Feschotte C. A Field Guide to Eukaryotic Transposable Elements. Annu Rev 621 Genet. Annual Reviews; 2020;54:539–61. https://doi.org/10.1146/annurev-genet-040620-622 022145 623 28. Sotero-Caio CG, Platt RN II, Suh A, Ray DA. Evolution and Diversity of Transposable 624 Elements in Vertebrate Genomes. Genome Biol Evol. 2017;9:161–77. 625 https://doi.org/10.1093/gbe/evw264 626 .CC-BY-NC 4.0 International licenseperpetuity. It is made available under a preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint 29. Tenaillon MI, Hollister JD, Gaut BS. A triptych of the evolution of plant transposable 627 elements. Trends Plant Sci. 2010;15:471–8. https://doi.org/10.1016/j.tplants.2010.05.003 628 30. Chen T-H, Winefield C. Comprehensive analysis of both long and short read 629 transcriptomes of a clonal and a seed-propagated model species reveal the prerequisites for 630 transcriptional activation of autonomous and non-autonomous transposons in plants. Mob 631 DNA. 2022;13:16. https://doi.org/10.1186/s13100-022-00271-5 632 31. Willing E-M, Rawat V, Mandáková T, Maumus F, James GV, Nordström KJV, et al. 633 Genome expansion of Arabis alpina linked with retrotransposition and reduced symmetric 634 DNA methylation. Nat Plants. Nature Publishing Group; 2015;1:14023. 635 https://doi.org/10.1038/nplants.2014.23 636 32. Wang J, Zhang G, Sun C, Chang L, Wang Y, Yang X, et al. DNA gains and losses in 637 gigantic genomes do not track differences in transposable element-host silencing 638 interactions. Commun Biol. Nature Publishing Group; 2025;8:704. 639 https://doi.org/10.1038/s42003-025-08127-3 640 641 642 .CC-BY-NC 4.0 International licenseperpetuity. It is made available under a preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-pdf

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-06-06T02:00:05.402940+00:00
License: CC-BY-NC-4.0