Abstract
15
Transposable elements (TEs) are major contributors to genome structure and evolution. 16
However, our ability to study them is limited by the difficulty of annotating and curating them, 17
especially in non-model organisms. This challenge is compounded by a lack of ground-truth 18
datasets for benchmarking, which are nearly impossible to generate through manual curation 19
alone. To overcome this critical limitation, we developed TEgenomeSimulator, a flexible 20
framework for generating synthetic genomes with configurable TE landscapes. 21
TEgenomeSimulator supports both randomly generated and biologically derived backbone 22
sequences, enabling the modeling of TE insertions under diverse structural and evolutionary 23
contexts. Benchmarking against existing simulators demonstrated that TEgenomeSimulator 24
reproduces realistic TE composition, sequence divergence, and integrity distributions while 25
offering greater flexibility in modeling chromosome-level structure. Its modular design offers 26
a tuneable continuum between biological fidelity and experimental control, facilitating 27
systematic benchmarking, algorithm development, and evolutionary modeling of TE 28
dynamics in silico, filling a major gap in the field. The source code of TEgenomeSimulator 29
and the scripts for this paper are available at https://github.com/Plant-Food-Research-30
Open/TEgenomeSimulator. 31
.CC-BY-NC 4.0 International licenseperpetuity. It is made available under a
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint
Introduction
32
Transposable elements (TEs) are abundant components in eukaryotic genomes. Once 33
regarded as parasitic DNA, TEs are now recognized as major contributors to gene 34
regulation, genome architecture, and evolutionary processes [1,2]. TEs also impact 35
agronomic traits, such as plant-pathogen interactions and environmental responses [3], and 36
may act as “silent” regulators of chromosomal architecture and post-zygotic barriers that 37
drive speciation [4]. Consistent with this view, cross-taxon surveys demonstrate that genome 38
size variation can be largely explained by TE accumulation [5-9]. TEs have also been 39
implicated in centromere structure evolution [6,7] and higher-order chromatin architecture 40
[12]. Together, these discoveries establish TEs not only as regulators of genes but also as 41
major catalysts of genome evolution, underscoring the need for high-quality assemblies to 42
accurately resolve repetitive DNA. 43
Advances in long-read sequencing technologies and assembly algorithms have markedly 44
enhanced genome assemblies and the resolution of repetitive regions, leading to increased 45
genome and pan-genome assemblies across taxa. Recent pangenome studies in Brassica, 46
grapevine and potato highlight TE enrichment in dispensable genomic regions linked to 47
resistance gene family expansion, regulatory variation, and reproductive isolation [13–15], 48
further solidifying the structural and functional impact of TEs and their critical role in genome 49
evolution. Despite these advances, TE annotation and interpretation remain underprioritized 50
in many genome projects, in only a few cases have domain experts filled this knowledge gap 51
[11]. 52
The detection and annotation of TEs within assembled genomes has benefited from 53
bioinformatic tool developments [12]; however, there are still major challenges to generating 54
quality annotation, especially for non-model organisms. TEs vary greatly in repetitiveness 55
and have undergone mutation, nested TE insertion and ectopic recombination, complicating 56
the differentiation of aged TE copies from non-TE sequences and hindering accurate 57
reconstruction of the consensus TE sequences [12]. Besides, ground–truth datasets are 58
especially scarce for non-model organisms, and only a small number have had any 59
validation through other ‘omics’ data sources [13,14]. Dfam is a TE database with an open 60
framework for community contribution and transparency in curated and non-curated datasets 61
[16]. As of March 2025, only 0.6% of the TE families in Dfam (release 3.9) have been 62
manually curated, predominantly from mammals. Most curated libraries rely on labor-63
intensive manual curation [17]. Consequently, simulated genomes with traceable TE 64
insertions can address this bottleneck by providing benchmarks for evaluating TE detection 65
and annotation. 66
Simulating TE sequence at scale that reflects the breadth of complexity and diversity of TE 67
composition is essential for benchmarking TE detection tools and modeling TE activity over 68
evolutionary time. Synthetic genomes with deliberately designed TE architectures, including 69
nested insertions, target site duplications (TSDs) and extreme cases of sequence 70
divergence and polymorphisms, provide a controlled framework to probe the limits of TE 71
detection and classification methods. This is particularly crucial in the era of telomere-to-72
telomere assemblies and pangenome-based annotations. By generating such 73
comprehensive and customizable datasets, researchers can systematically evaluate 74
performance under realistic and challenging conditions while also exploring the evolutionary 75
processes shaping genome architecture. A variety of TE simulators have been developed to 76
address some of these needs. For example, denovoTE-eval [18] synthesizes random 77
.CC-BY-NC 4.0 International licenseperpetuity. It is made available under a
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint
sequences with user-specified lengths and GC content, requiring an input table containing 78
TE families and parameters to model mutagenesis, sequence decay, and random insertions. 79
However, it can simulate only a single random sequence at a time and offers limited flexibility 80
for adjusting mutagenesis parameters when working with large TE libraries. By relying on 81
predefined TE architectures, SimulaTE [19] can recapitulate intricate TE landscapes and 82
simulate sequencing reads, but it remains limited in its capability to simulate novel 83
evolutionary events that extend beyond the user-provided genome architecture. SLiM [20] 84
excels at modeling selection and demography but represents TEs as single units, each like a 85
single nucleotide variant, restricting exploration of sequence-level architectures and genome 86
size changes. GARLIC [21] generates realistic intergenic sequences based on k-mer 87
composition, GC content, and repeat profiles but lacks flexibility in TE mutagenesis profiles 88
and does not capture the coexistence of coding and non-coding regions. While PrinTE is 89
designed for forward-time simulation on TE mobilization, mutation, and removal, it requires 90
multiple iterative attempts to infer and reproduce the dynamics of a specific TE type, with a 91
focus on LTR retrotransposon [22]. Despite the breadth of existing simulation tools in design, 92
few offer the flexibility to customize genome content while also modeling TE sequence 93
mutagenesis under diverse decay schemes. Even fewer support fine-grained, family-specific 94
controls or allow combinations of multiple modes for complex TE composition simulations. 95
To address these limitations, we developed TEgenomeSimulator (Fig. 1), which was built on 96
concepts introduced in denovoTE-eval. Extending beyond the capability of denovoTE-eval in 97
synthesizing single random sequences and performing random and nested TE insertions, 98
TEgenomeSimulator can synthesize genomes with multiple chromosomes of varying 99
complexity and offers fine-grained control over TE sequence populations for a more realistic 100
fragmentation pattern. TEgenomeSimulator departs from the uniform TSD range approach of 101
denovo-TEeval by modeling TSD length at the TE-superfamily level, a key feature for 102
structure-based TE detection and classification. It can also incorporate real genome 103
assemblies and their TE compositions for more realistic modeling. Furthermore, it allows 104
users to combine different simulation modes to introduce new TE activity beyond known TE 105
compositions, enabling the exploration of diverse evolutionary scenarios. A comparison with 106
existing simulators is available in Table S1. 107
In this paper, we demonstrate that TEgenomeSimulator generates more realistic synthetic 108
genomes than denovoTE-eval and GARLIC. We systematically illustrate how 109
parameterization of TE copy number, fragmentation, mean sequence identity, and standard 110
deviation (SD) influences simulation outcomes and their biological interpretation. The 111
resulting outputs enable robust benchmarking of TE detection and annotation methods. As a 112
proof of concept, we use TEgenomeSimulator to facilitate the evaluation of TE detection by 113
RepeatMasker [23]. TEgenomeSimulator is implemented as a reproducible and portable 114
Python package for Linux, installable via ‘pip’, Docker, or Apptainer. 115
116
117
.CC-BY-NC 4.0 International licenseperpetuity. It is made available under a
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint
118
Figure 1. TEgenomeSimulator flowchart. TEgenomeSimulator operates in three modes—119
Random Synthesized Genome (mode 0), Custom Genome (mode 1), and TE 120
Composition Approximation (mode 2)—shown in blue, grey, and sage, respectively. 121
Workflow components (rounded pink rectangles) are organized into three conceptual layers: 122
Inputs, Processes, and Outputs. Connections between components indicate the data flow 123
and processing steps taken by each mode, highlighting shared and mode-specific elements. 124
See Materials and Methods for details. 125
126
Methods
127
Implementation 128
Simulation Modes 129
Random Synthesized Genome mode (mode 0) synthesizes artificial chromosomes using 130
user-defined chromosome lengths and GC content, then inserts simulated TEs at random 131
positions. A TE library is used to produce TE copies subjected to random mutagenesis, 132
including nucleotide substitutions, single-base insertions and deletions (InDels), TSDs, and 133
fragmentation. 134
.CC-BY-NC 4.0 International licenseperpetuity. It is made available under a
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint
Custom Genome mode (mode 1) accepts a user-provided ‘backbone’ genome with TEs 135
removed, or masks and removes TEs using RepeatMasker [23] (using ‘--to_mask’ option) 136
with a user-supplied TE library. Mutation and insertion proceed as in mode 0. 137
TE Composition Approximation mode (mode 2) also utilizes RepeatMasker to remove TEs 138
but additionally infers TE family abundance, nucleotide substitution and InDel rates, and 139
sequence integrity from the source genome, which are then used to parameterize TE 140
simulations. 141
Simulating TE Mutagenesis 142
From the TE library, TEgenomeSimulator extracts family classification and sequence 143
metadata and assigns per-family mutagenesis parameters. In modes 0 and 1, per-family TE 144
copy numbers are sampled from user-defined ranges. Sequence identity (I) is drawn from a 145
Gaussian distribution with user-specified mean and SD (modes 0 and 1) or estimated from 146
RepeatMasker output (mode 2), then adjusted following denovoTE-eval. TSD lengths 147
follows published observations [24], with random TSD sequences synthesized. Sequence 148
integrity in modes 0 and 1 is modeled using a beta distribution controlled by alpha and beta 149
(default α=0.5, β=0.7), producing predominantly fragmented elements, with 0.1% of copies 150
remaining intact by default. Sequence integrity in mode 2 is inferred from the source 151
genome. TE copies are randomly oriented and integrated. See Supplementary Methods for 152
detailed implementation. 153
Random and Nested TE Insertion 154
Nested insertions are generated in 0-30% of LTR retrotransposon copies, with insertion 155
targets and locations selected randomly. 156
Input and Output 157
TE libraries must follow RepeatMasker naming conventions (e.g, >ATCOPIA10#LTR/Copia) 158
and ideally contain full-length sequences (e.g. LTR-INT-LTR for retrotransposons). Split LTR 159
entries are automatically reassembled (see Supplementary Methods). Modes 1 and 2 160
require a genome FASTA file, whereas mode 0 requires a genome index CSV file specifying 161
chromosome name, length, and GC%. Outputs include the simulated genome, all TE copy 162
sequences, and a GFF file recording TE locations, identity, integrity, and nesting 163
relationships. 164
165
Simulator comparison analysis 166
Comparison 1: TEgenomeSimulator vs denovoTE-eval 167
TEgenomeSimulator’s mode 0 was used to synthesize a 10 Mb non-TE random sequence 168
with 35% GC content, followed by TE insertion using a composition defined by 169
TEgenomeSimulator (denoted as TEGS in Fig. 2). Copy numbers per family were set 170
between 1 and 10, with other mutagenesis parameters left at default values. To ensure 171
comparability, settings for size of non-TE backbone, GC%, TE copy number, and sequence 172
identity from TEgenomeSimulator were used for denovoTE-eval (denoted as dnvTEe in Fig. 173
2), while maintaining direct comparisons of other features. 174
Comparison 2: TEgenomeSimulator vs GARLIC 175
.CC-BY-NC 4.0 International licenseperpetuity. It is made available under a
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint
Chromosome 1 of the Arabidopsis thaliana (TAIR10, NCBI accession ID GCF_000001735.4) 176
and the curated TE library from the EDTA GitHub repository 177
(https://github.com/oushujun/EDTA/blob/master/database/) were used as inputs. Two 178
genomes were simulated with TEgenomeSimulator: (i) a digital replica of TAIR10 179
chromosome 1 using mode 2 (m2 in Fig. 2); and (ii) a hybrid simulation in which the TE 180
composition inferred by mode 2 was imported into mode 0 to generate a randomly 181
synthesized backbone (m2+m0 in Fig. 2; see Supplementary Methods). 182
GARLIC simulation required four input files: (i) tandem repeat profiles generated by Tandem 183
Repeat Finder (TRF, v4.07) [25] on TAIR10 genome; (ii) gene annotations from Ensembl 184
release 61 (GFF3) converted to GARLIC’s tabular format using a custom script; (iii) 185
RepeatMasker’s alignments of TAIR10 genome using the same TE library; and (iv) the TE 186
library converted from FASTA to EMBL format. GARLIC’s ‘createModel.pl’ and 187
‘createFakeSequence.pl’ scripts were used to generate a synthetic sequence matching 188
chromosome 1 length. 189
K-mer spectra (k = 6) were analyzed using a custom Python script treating reverse 190
complements as canonical. Shannon entropy (H) on nucleotide probabilities (p) was 191
calculated in 1 Kb sliding window with a step size of 1 bp according to the formula: 192
𝐻 = − $ 𝑝! log"(𝑝!)
!∈{&,(,),*}
193
194
Use-cases demonstration 195
A multi-species TE library was constructed by merging the curated TE library of A. thaliana, 196
Oryza sativa, and Zea mays obtained from EDTA, followed by redundancy removal using cd-197
hit-est with settings following Goubert et al. [17]. The resulting library, along with the TE-198
depleted TAIR10 genome, was used for simulation demonstrations using mode 1. The 199
TAIR10 genome was pre-processed using RepeatMasker and the curated A. thaliana TE 200
library to remove endogenous TEs. Simulation settings corresponding to Fig. 3 are 201
described below. 202
Impact of TE copy number on genome size (mode 1) 203
TE copy number ranges per family were set to 1-10, 5-100, 5-500, 5-1000 and 5-2000 using 204
the TE-depleted A. thaliana genome as input. 205
Impact of alpha and beta on TE integrity (mode 1) 206
For demonstration, TE integrity was modeled with a beta distribution controlled by five [α, β] 207
pairs: [0.5, 1], [0.5, 0.75], [0.5, 0.5], [0.75, 0.5], and [1, 0.5]. 208
Impact of mean sequence identity (mode 1) 209
Identity ranges were set to 90-100, 80-90, 70-80, 70-85, 70-90, 70-95, and 70-100. 210
Combined effects of sequence identity and standard deviation (mode 1) 211
Four scenarios were simulated (Fig. 3D) by varying mean identity ranges (10-80 or 90-100) 212
and SD (1-5 and 15-20) to represent different histories of TE activity (see Supplementary 213
Methods) 214
.CC-BY-NC 4.0 International licenseperpetuity. It is made available under a
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint
Digital replicas (mode 2) 215
Genomes of A. thaliana (TAIR10; 120 Mb), O. sativa (AGIS1.0; NCBI accession ID 216
GCF_034140825.1; 386 Mb), Z. mays (Zm-B73-REFERENCE-NAM-5.0; NCBI accession ID 217
GCF_902167145.1; 2.183 Gb), Danio rerio (GRCz12tu; NCBI accession ID 218
GCF_049306965.1; 1.449 Gb) and Drosophila melanogaster (dm6; NCBI accession ID 219
GCF_000001215.4; 138 Mb) were directly used as input with default settings. Acquisition of 220
TE libraries for the mentioned plant species has been described previously. TE libraries for 221
D. rerio and D. melanogaster were downloaded from the RepeatModeler2 GitHub repository 222
(https://github.com/jmf422/TE_annotation/tree/master/benchmark_libraries/RM2). 223
Computational resource tests 224
Mode 0 was benchmarked using five configurations that varied the length of randomly 225
synthesized sequences with 35% GC prior to TE insertion. Four tests, generated a single 226
sequence of 100 Kb, 1 Mb, 10 Mb, or 100 Mb, and a fifth generated ten 100 Mb sequences 227
(1Gb in total). TE copy numbers per family were fixed at 1-10. 228
To test mode 1, the TE-depleted TAIR10 genome and curated A. thaliana TE library were 229
used, with TE copy number ranges of 1-10, 5-100, 5-500, 5-1000, and 5-2000. 230
Mode 2 resource usage was recorded for simulations of A. thaliana, O. sativa, Z. mays, D. 231
melanogaster, and D. rerio. 232
All tests were executed on the same SLURM-managed cluster. Each mode 0 and mode 1 233
test used a single CPU core. Each mode 2 test was assigned 20 CPU cores due to the 234
computational demands of running RepeatMasker. 235
TE recovery rate analysis 236
RepeatMasker was applied to simulated genomes from the four scenarios Fig. 3D using the 237
same TE library as in the simulations. TEgenomeSimulator’s TE annotation files (GFF) 238
served as ground truth and were compared to the BED files converted from RepeatMasker 239
output using Bedtools Intersect [26], applying an 80% reciprocal overlap cutoff. Recovery 240
rate (R) was calculated per integrity and identity bin (10 bins, ranging from 0 to 1) according 241
to the formula: 242
𝑅 = 𝑁𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑞𝑢𝑎𝑙𝑖𝑓𝑖𝑒𝑑 𝑜𝑣𝑒𝑟𝑙𝑎𝑝𝑠 𝑖𝑛 𝑏𝑖𝑛
𝑇𝑜𝑡𝑎𝑙 𝑔𝑟𝑜𝑢𝑛𝑑 − 𝑡𝑟𝑢𝑡ℎ 𝑇𝐸𝑠 𝑖𝑛 𝑏𝑖𝑛 243
244
Results
245
Simulator Comparisons 246
We compared the key features of TEgenomeSimulator to those of denovoTE-eval, GARLIC, 247
SimulaTE, and SLiM (Table S1). Unlike denovoTE-eval, TEgenomeSimulator enables the 248
synthesis of multi-chromosomal genomes within a single simulation run. It also supports 249
modeling of superfamily-specific TSD lengths, allows beta distribution or empirical 250
distribution to model fragmentation, and enables both empirical and combinatorial modeling 251
of TE composition through the sequential application of multiple simulation modes 252
(Supplementary Methods). In contrast to GARLIC, which analyzes genic and non-genic 253
profiles but only synthesizes non-genic regions, TEgenomeSimulator’s mode 2 preserves 254
gene models whose exons or introns are not identified and masked as TEs by 255
.CC-BY-NC 4.0 International licenseperpetuity. It is made available under a
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint
RepeatMasker. Moreover, TEgenomeSimulator is uniquely capable of utilizing either 256
randomly generated or empirically derived non-TE backbones for TE insertion, providing 257
enhanced flexibility for the simulation of complex and biologically realistic TE landscapes 258
(Table S1). 259
We assessed TEgenomeSimulator against its closest peers, denovoTE-eval and GARLIC, 260
using two comparison designs (Fig. 2A-B). Comparison 1 (Fig. 2A) utilizes a 10 Mb TE-free 261
backbone, synthesized by TEgenomeSimulator, that is used by both TEgenomeSimulaor 262
and denovoTE-eval for subsequent TE insertion (see Methods). To maintain consistency, 263
the TE mutagenesis parameters of TEgenomeSimulator were reformatted and imported into 264
denovoTE-eval. Both tools yielded comparable TE superfamily composition and divergence 265
distributions (Fig. S1), validating the consistency in the transferred TE composition and 266
sequence identity parameters. We did, however, observed considerable variance in the 267
resulting genome size and composition. denovoTE-eval produced a genome approximately 268
5 Mb larger than TEgenomeSimulator, due to incorporating 7.6% more TE bases(Fig. 2C). 269
We also observed a marked difference in the integrity distributions of simulated TEs. 270
denovoTE-eval displayed flattened TE loci distributions at integrity ranges 0.4-0.7 and 0.7-271
0.9, with no fragmented element exceeding 0.9 (Fig. 2E). This gap between 0.9 and 1 arose 272
because denovoTE-eval caps integrity at 0.9 following fragmentation, while the even 273
patterns reflected its programmed two-tier sampling scheme: integrities 0.7-0.9 for sequence 274
<500 bp and 0.4-0.9 for sequences ≥500 bp. TEgenomeSimulator, however, models integrity 275
using a beta-distribution combined with a fixed fraction of intact TEs. This results in a 276
broader spectrum, spanning degraded to nearly intact elements, better reflecting the 277
sequence integrity observed in natural genomes (Fig. 2F). 278
For Comparison 2 (Fig. 2B), we used TEgenomeSimulator’s mode 2 to build a model of the 279
TE content of TAIR10 chromsome1 (chr1), which was then used as input for TE insertion into 280
TE-depleted chr1 (m2) or a simulated backbone matching the length of chr1 (m2+m0). The 281
resulting genomes were compared to simulations derived from GARLIC. GARLIC produced 282
genomes with a markedly higher proportion of TE sequence (25%) than the original chr1 283
(10.6%) and nearly three times as many TE insertions (Fig. S2A). Both the m2 (9.7%) and 284
m2+m0 (10.1%) genomes more closely matched the balance of TE and non-TE occupancy 285
observed in chr1 (Fig. 2D). GARLIC’s integrity profile was heavily skewed toward values of 286
0.9-1, in sharp contrast to the more realistic integrity ranges captured by 287
TEgenomeSimulator m2 and m2+m0 (Fig. S3). Sequence complexity analyses reinforced 288
these differences (Fig. 2G-I). GARLIC's k-mer profile displayed an extended tail in TE-only 289
regions, which is absent from both chr1 (Ori) and TEgenomeSimulator-synthesized 290
sequence (Fig. 2G). In non-TE regions, GARLIC and TEgenomeSimulator’s m2 closely 291
matched the reference k-mer profile, whereas the distortion in m2+m0 reflected the use of a 292
randomly synthesized non-TE backbone (Fig. 2G). We used Shannon entropy analysis to 293
determine variability within the data, we observed a central dip within chr1 (Ori) that was 294
successfully reproduced by TEgenomeSimulator’s m2 , while GARLIC and m2+m0 failed to 295
capture this feature (Fig. 2H). GARLIC’s outputs exhibited a narrower, higher entropy range 296
in non-TE regions and a lower entropy range in TE-only regions relative to the original 297
genome (Fig. 2I). For TEgenomeSimulator, the TE-only entropy range of m2 was shifted 298
compared to the reference, but its non-TE profile was preserved, while m2+m0 showed 299
markedly narrower non-TE distribution, again reflecting the random backbone synthesis 300
(Fig. 2I). 301
.CC-BY-NC 4.0 International licenseperpetuity. It is made available under a
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint
302
Figure 2. Simulator Comparisons. (A) Comparison 1: mode 0 of TEgenomeSimulator 303
(TEGS) and denovoTE-eval (dnvTEe) were compared using a random synthesized 10 Mb 304
sequence. The TE mutagenesis table (TE mut-table) generated by TEGS was imported to 305
dnvTEe to match TE composition for comparison. (B) Comparison 2: simulated TAIR10 306
chromosome 1 (chr1) generated by GARLIC, TEGS full mode 2 (m2), TEGS partial mode 2 307
followed by full mode 0 (m2+m0), and the original TAIR10 chr1. For m2+m0, the TE mut-308
table, TE-depleted chr1 length and GC content were acquired from m2 and used as inputs 309
.CC-BY-NC 4.0 International licenseperpetuity. It is made available under a
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint
for m0. (C, D) Proportions of TE and non-TE sequences from Comparison 1 (C) and 2 (D). 310
(E, F) TE integrity distributions for dnvTEe (E) and TEGS (F) simulations. (G, H) K-mer (G) 311
and Shannon entropy (H) spectra of whole genomes, and the non-TE and TE only 312
sequences, colored by simulation method. (I) Shannon entropy ranges by sequence 313
category. Asterisks indicate p < 0.0001 (Wilcoxon Rank-Sum Test versus the original 314
genome). 315
316
Use cases 317
TE composition is shaped by multiple evolutionary processes, each leaving distinct genomic 318
signatures. To illustrate how TEgenomeSimulator captures these dynamics, we demonstrate 319
key configurations that reproduce major features of empirical TE landscapes. 320
Modeling basic TE features 321
TE accumulation is a primary driver of genome expansion [27–29]. Using the Custom 322
Genome Mode together with the TE-depleted A. thaliana (TAIR10) genome and the curated 323
Arabidopsis TE library (see Methods), we simulated genome expansions ranging from 324
under 200 Mb to over 1 Gb by increasing the maximum copy number per TE family from 10 325
to 2000 (Fig. 3A). This configuration directly models the impact of TE accumulation on 326
genome size. 327
TE integrity reflects the balance between recent insertions and long-term decay. In both 328
Random Synthesized Genome and Custom Genome modes, TE integrity is modelled using 329
a beta distribution across families. When α < β (both between 0 and 1), the distribution takes 330
on an asymmetrical U-shape with a peak near low integrity values (Fig. 3B), reflecting the 331
predominance of TE relics observed in empirical datasets [30]. Conversely, when β < α, the 332
distribution shifts toward high-integrity copies (Fig. 3B), reflecting recent evolutionary TE 333
activity. 334
Alongside modifying TE integrity, we can vary the range of mean sequence identity to reflect 335
mutation accumulation over time. By varying the range of mean sequence identity (70%-336
100%) while constraining copy numbers (5-100 per family), we generated divergence 337
distributions with progressively flattened peaks as identity ranges widened (Fig. 3C), 338
reproducing the expected signatures of sustained or overlapping TE amplification events 339
[31,32]. 340
341
Modeling TE composition under diverse evolutionary scenarios 342
The duration of TE bursts, which can be measured by the standard deviation (SD) of 343
sequence identity, also influences divergence distribution. A short-lived TE burst results in a 344
rapid deposition of nearly identical insertions. These insertions then gradually diverge 345
through accumulated mutation. In contrast, a long-lasting burst over evolutionary time leads 346
to the continuous deposition of similar TE sequences, causing a less synchronized 347
divergence pattern owing to overlapping degradation timelines. Moreover, ancient TE bursts 348
tend to show greater sequence deterioration due to prolonged exposure to mutational decay. 349
We can simulate genomes that represent different evolutionary phases of TE activity (Fig. 350
3D), offering genomic templates with distinct base-level TE composition that can serve as 351
burn-in states for more advanced simulations. As shown in Fig. 3D, the representation of TE 352
burst timeline and duration on the x-axis and y-axis, respectively, allows four distinct TE 353
divergence landscapes to emerge: I) high sequence identity with high SD, modeling recent 354
.CC-BY-NC 4.0 International licenseperpetuity. It is made available under a
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint
TE bursts with ongoing or prolonged activity over time; II) low sequence identity with high 355
SD, modeling ancient and prolonged activity, or multiple independent bursts occurring over 356
an extended period; III) low sequence identity with low SD, reflecting ancient TE bursts that 357
were quickly inactivated, with TE activity remaining silenced since the initial event; IV) high 358
sequence identity with low SD, inferring very recent and rapid TE bursts, resulting in a 359
synchronized insertion of highly similar sequences. 360
We modelled these four scenarios using the Custom Genome mode of TEgenomeSimulator. 361
The identity and SD ranges were implemented as shown in Fig. 3D, with identity ranges of 362
90%–100% and 70%–80%, and SD ranges of 15–20 and 1–5. The TE copy number range 363
was fixed between 5 and 1000. The resulting sequence divergence profiles demonstrated a 364
narrowing of the peak as the SD range decreased from 15–20 to 1–5, indicating more 365
instantaneous bursts (Fig. 3E). Furthermore, shifting the identity range from 70–80% to 90–366
100% led to a leftward shift of the divergence peak (from ~0.25 to ~0.1) and an increase in 367
peak height. These results demonstrate that TEgenomeSimulator is well-suited for modeling 368
TE composition under diverse evolutionary scenarios. 369
370
371
.CC-BY-NC 4.0 International licenseperpetuity. It is made available under a
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint
372
Figure 3. Example use-case of TEgenomeSimulator. (A) Impact of altering TE copy 373
number ranges on simulated genome size (mode 0). (B) TE sequence integrity distribution 374
under different alpha and beta settings (mode 0). (C) TE sequence divergence distribution 375
under varying ranges of mean identity across TE families, with a fixed TE copy number 376
range of 5–100 (mode 0). (D) Four historical TE activity scenarios modeled by adjusting 377
mean identity (activity timing) and its standard deviation (SD; activity duration). (E) Resulting 378
TE divergence patterns associated with the settings in (C). (F) TE/nonTE sequence 379
proportions in original and simulated Arabidopsis thaliana genomes (mode 2). (G,H) 380
Sequence divergence (G) and integrity (H) in original and simulated A. thaliana genomes. 381
382
Digital replicas 383
TE composition in real genomes varies substantially across TE families and species 384
because of their complex and diverse evolutionary histories. This variability makes it 385
challenging to simulate realistic genome-wide TE landscapes. To address this, we used 386
TEgenomeSimulator’s TE Composition Approximation Mode to model the TE profiles of A. 387
thaliana (TAIR10), O. sativa, Z. mays, D. rerio, and D. melanogaster and synthesis replicas. 388
We observed the final genome size and proportion of TE occupancy in the digital replicas 389
.CC-BY-NC 4.0 International licenseperpetuity. It is made available under a
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint
closely matched those observed in the original genomes (Fig. 3F, Fig. S4-S7). In terms of 390
TE loci and TE bases per superfamily, the simulated TE composition also aligned well with 391
the original genomes (Fig. S4-S8). 392
While the simulated divergence profiles of the replicas resembled those of the original 393
genomes, they were not identical: the fine-scale fluctuations present in the original genomes 394
were smoothed in the replicas due to the use of normal distributions in the simulator. This 395
approach, however, still provides a generalized distribution pattern that captures the major 396
trends in TE sequence evolution (Fig. 3G, Fig. S4-S7). 397
To replicate a real TE integrity pattern for each TE family, the empirical distribution of TE 398
length ratios (relative to consensus sequences) is extracted from the input genome. This 399
enables the simulator to reproduce the integrity distributions more authentically in the 400
resulting simulated genome (Fig. 3H, Fig. S4-S7). 401
These digital replicas provide a robust foundation for simulating future TE-related events, 402
such as new TE bloating or purging episodes, effectively serving as burn-in phases for 403
modeling genome evolution from the current state. 404
Computational resource usage 405
To evaluate the computational resource usage of TEgenomeSimulator, we conducted tests 406
across all three simulation modes (Fig. 4, Table S2). In mode 0, we observed that the total 407
length of randomly synthesized sequences significantly impacts memory usage, particularly 408
for sequences longer than 100 Mb. Synthesizing a large genome over 1 Gb, comprising 10 409
chromosomes of approximately 100 Mb each, required over 12 GB of memory (Fig. 4, Table 410
S2). In mode 1, increasing the TE copy number range per family had a more pronounced 411
impact on runtime than on memory usage (Fig. 4). Simulations with broader copy number 412
ranges took several hours to complete and required multiple gigabytes of memory, 413
depending on the settings. Notably, a simulation with copy number ranges of 5-2000 per 414
family consumed over 6 GB of memory and required nearly 10 hours to complete, resulting 415
in a simulated genome exceeding 1 Gb in size (Fig. 4, Table S2). In contrast, tests varying 416
TE sequence identity and integrity finished in under 3.5 minutes and used less than 1 GB of 417
memory (Table S2). In mode 2, both runtime and memory usage were substantially 418
increased in relation to the size and complexity of the input genomes used for TE 419
composition approximation (Fig. 4, Table S2). The increasing computational demand with 420
genome size appears to correlate with native TE content (Fig. S2-S5). 421
422
423
.CC-BY-NC 4.0 International licenseperpetuity. It is made available under a
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint
424
Figure 4. Computational resource usage of TEgenomeSimulator. Each bubble 425
represents a test run, with size proportional to the simulated genome. Mode 0 varied 426
Reference
sequence length (100 Kb-100 Mb). Mode 1 varied TE copy numbers ranges. 427
Mode 2 approximated TE compositions from five genomes: Arabidopsis thaliana (TAIR10), 428
Oryza sativa (O_sativa), Zea mays (Z_mays), Drosophila melanogaster (dm6), and Danio 429
rerio (D_rerio). 430
431
TE recovery rate test 432
TEgenomeSimulator can generate diverse TE insertions ideal for benchmarking TE 433
detection and annotation tools. We assessed the TE detection capability of RepeatMasker 434
using simulated TE insertions from the four scenarios described in Fig. 3D as ground truth 435
(Fig. 5A). TE recovery rate was measured as the proportion of simulated insertions correctly 436
identified by RepeatMasker, stratified by sequence identity and integrity (Fig. 5A). The 437
resulting heatmaps revealed a strong dependence on sequence identity, with higher 438
recovery for insertions of greater similarity to the consensus (Fig. 5B). In contrast, recovery 439
showed little or no clear dependency on integrity. Bar plots summarizing recovery rate 440
across identity and integrity slices confirmed these trends across the four scenarios: 441
recovery rates remained high at identity above 0.7 but declined progressively with greater 442
divergence, whereas variation across integrity levels lacked a consistent pattern. Although 443
the precise recovery values may be influenced by RepeatMasker’s parameter settings, 444
particularly the divergence and alignment thresholds, the results clearly demonstrate that TE 445
insertions with higher sequence divergence are more difficult to detect. 446
.CC-BY-NC 4.0 International licenseperpetuity. It is made available under a
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint
447
Fig. 5. TE annotation recovery rate of RepeatMasker. (A) TE annotation recovery 448
workflow, with ground-truth (green) and test (purple) annotations; numbered blue rectangles 449
indicate analysis steps (see Methods). (B) Heatmaps of recovery rate by sequence identity 450
and integrity for the four scenarios in Fig. 3D, with marginal bar graphs showing mean 451
recovery. Empty bins are shown in white. 452
453
Discussion
454
TEgenomeSimulator introduces a flexible, mode-based architecture that addresses critical 455
Limitations
in existing TE simulation tools by enabling both randomized simulations and 456
simulations based on real genomic architecture, thereby capturing structural and 457
compositional complexity characteristics of natural TE landscapes. 458
TEgenomeSimulator-generated TE populations display a broad and continuous range of 459
sequence integrity values by modeling either a beta or an empirical distribution, providing a 460
more biologically realistic material for testing annotation tools. While denovoTE-eval served 461
as a foundational model for the development of TEgenomeSimulator, it lacked flexibility in 462
.CC-BY-NC 4.0 International licenseperpetuity. It is made available under a
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint
defining superfamily-specific TSDs and exhibited discontinuous integrity distributions due to 463
rigid random sampling. TEgenomeSimulator overcomes this limitation by introducing TE 464
superfamily modeling of TSDs and sequence divergence and integrity, producing outputs 465
that better reflect natural TE diversity and evolutionary behaviors of TE lineages observed in 466
real genomes. By comparison, GARLIC’s emphasis on modeling non-genic contexts is 467
useful for certain studies but results in overrepresentation of high-integrity TE copies and 468
reduced sequence entropy, limiting its value for representing complex genomic structure and 469
simulating evolutionary variation. Other simulation tools, such as SimulaTE and SLiM, 470
emphasize population dynamics or evolutionary events but lack fine control over structural 471
variation or sequence-level mutagenesis parameters that TEgenomeSimulator provides. 472
TEgenomeSimulator’s capability in combining different modes provides a tunable continuum 473
across a spectrum of evolutionary scenarios and genomic realism. For example, a mode 0 474
simulation is ideal for controlled benchmarking under a uniform backbone condition, where 475
the goal is to assess TE detection sensitivity without confounding genomic context effects. 476
Mode 2, approximation to realistic TE composition, best supports validation of annotation 477
tools with authentic genomic contexts, where chromosomal architecture, including 478
centromere, intergenic and intragenic regions, is preserved as the non-TE backbone. Mode 479
1 combines controllable TE composition and mutagenesis with a realistic non-TE backbone. 480
Further flexibility is introduced through the combination of modes (e.g. mode 2 + mode 0 or 481
mode 2 + mode 1; see examples in Supplementary Methods) uniquely enable hybrid 482
simulation approaches that leveraging the TE compositional realism of real genomes (mode 483
2) while expanding diversity through random synthetic non-TE backbones (mode 0) and 484
introducing controllable novel TE activity to a real genome (mode 1), mimicking TE bursts in 485
horizontal transfer or TE reactivation. Besides, combining TEgenomeSimulator’s digital 486
replicas with PrinTE’s forward evolution simulation facilitates investigation in complex 487
evolutionary history [22]. 488
The customizable mutagenesis parameters and compositional realism of 489
TEgenomeSimulator make it ideal for systematically evaluating tool sensitivity across identity 490
and integrity gradients, as demonstrated in our RepeatMasker studies. Hence, 491
TEgenomeSimulator could serve as a reference framework for benchmarking studies, 492
establishing standardized simulation conditions for reproducible TE tool comparisons. 493
Future development could integrate the positional bias of TE insertion, TE dynamics over 494
time, and epigenetic influences, bridging the gap between structural simulation and 495
evolutionary modeling. 496
Key points 497
- TEgenomeSimulator addresses the critical gap in large-scale genomic analysis by 498
providing traceable, simulated genomes that serve as essential “ground-truth” 499
benchmarks for evaluating the accuracy of TE annotation tools, particularly for non-500
model organisms where curated data is scarce. 501
- TEgenomeSimulator offers a unique methodological structure through three 502
interoperable modes that span a spectrum of complexity: from entirely synthetic 503
genomes (mode 0) and applying TEs to custom genome backbones (mode 1) to the 504
high-fidelity approximation of empirical TE landscapes derived from source genomes 505
(mode 2). 506
.CC-BY-NC 4.0 International licenseperpetuity. It is made available under a
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint
- TEgenomeSimulator allows for versatile parameterization of TE mutagenesis, 507
including sequence identity, integrity, and copy number. The tool captures greater 508
evolutionary complexity and genomic realism than existing simulators such as 509
denovoTE-eval and GARLIC. 510
- While this manuscript demonstrates TEgenomeSimulator’s utility via RepeatMasker, 511
the framework is built to evaluate a wide range of TE detection and annotation 512
methods, uncovering how specific gradients of sequence conservation and structural 513
integrity impact performance across different algorithms. 514
- TEgenomeSimulator’s ability to combine simulation modes provides a tunable 515
continuum between evolutionary abstraction and genomic realism. This flexibility, 516
paired with its open-source design and compatibility as a “burn-in” input for forward-517
evolution simulators like PrinTE, ensures that it is a comprehensive resource for the 518
genomic research community. 519
520
Acknowledgements
521
We want to thank Drs Chen Wu, Julie Blommaert, Jason Shiller, and David Chagné for 522
reviewing this manuscript and providing insightful comments. 523
524
Author Contributions 525
Ting-Hsuan Chen (Conceptualization [lead], Software [lead], Analysis [lead], Writing — 526
original draft [lead]), Olivia Angelin-Bonnet (Software [supporting], Writing — review & editing 527
[supporting], James Bristow (Software [supporting], Writing — Review & editing 528
[supporting]), Christopher Benson (Writing — Review & editing [supporting]), Shujun Ou 529
(Writing — review & editing [supporting]), Cecilia Deng (Conceptualization [supporting], 530
Writing — review & editing [supporting], Supervision [supporting]), Susan Thomson 531
(Conceptualization [supporting], Funding acquisition [lead], Supervision [lead], Writing — 532
review & editing [lead]) 533
534
Conflict of interest 535
The authors declare that they have no competing interests. 536
Use of AI in this study 537
ChatGPT o4 and 5 were used to help draft and improve scripts during the development of 538
TEgenomeSimulator, and to suggest edits for clarity and grammar in the manuscript. 539
Funding 540
This research was supported by Plant & Food Research through the Kiwifruit Royalty 541
Investment Programme. 542
Data availability 543
TEgenomeSimulator is available at https://github.com/Plant-Food-Research-544
Open/TEgenomeSimulator under the GPL-3.0 license, and all codes for the analyses in this 545
.CC-BY-NC 4.0 International licenseperpetuity. It is made available under a
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint
paper are available at https://github.com/Plant-Food-Research-546
Open/TEgenomeSimulator_publication_code under the MIT license. 547
Reference
548
1. Tossolini I, Mencia R, Arce AL, Manavella PA. The genome awakens: transposon-549
mediated gene regulation. Trends Plant Sci. Elsevier; 2025;30:857–71. 550
https://doi.org/10.1016/j.tplants.2025.02.005 551
2. Hayward A, Gilbert C. Transposable elements. Curr Biol. Elsevier; 2022;32:R904–9. 552
https://doi.org/10.1016/j.cub.2022.07.044 553
3. Zanini SF, Bayer PE, Wells R, Snowdon RJ, Batley J, Varshney RK, et al. Pangenomics in 554
crop improvement—from coding structural variations to finding regulatory variants with 555
pangenome graphs. Plant Genome. 2022;15:e20177. https://doi.org/10.1002/tpg2.20177 556
4. Cabral-de-Mello DC, Palacios-Gimenez OM. Repetitive DNAs: the ‘invisible’ regulators of 557
insect adaptation and speciation. Curr Opin Insect Sci. 2025;67:101295. 558
https://doi.org/10.1016/j.cois.2024.101295 559
5. Elliott TA, Gregory TR. Do larger genomes contain more diverse transposable elements? 560
BMC Evol Biol. 2015;15:69. https://doi.org/10.1186/s12862-015-0339-8 561
6. Xie L, Huang Y, Huang W, Shang L, Sun Y, Chen Q, et al. Genetic diversity and evolution 562
of rice centromeres. bioRxiv. 2024;2024.07.28.605524. 563
https://doi.org/10.1101/2024.07.28.605524 564
7. Naish M, Henderson IR. The structure, function, and evolution of plant centromeres. 565
Genome Res. 2024;34:161–78. https://doi.org/10.1101/gr.278409.123 566
8. Bayer PE, Golicz AA, Tirnaz S, Chan CK, Edwards D, Batley J. Variation in abundance of 567
predicted resistance genes in the Brassica oleracea pangenome. Plant Biotechnol J. 568
2019;17:789–800. https://doi.org/10.1111/pbi.13015 569
9. Cochetel N, Minio A, Guarracino A, Garcia JF, Figueroa-Balderas R, Massonnet M, et al. A 570
super-pangenome of the North American wild grape species. Genome Biol. 2023;24:290. 571
https://doi.org/10.1186/s13059-023-03133-2 572
10. Bozan I, Achakkagari SR, Anglin NL, Ellis D, Tai HH, Strömvik MV. Pangenome analyses 573
reveal impact of transposable elements and ploidy on the evolution of potato species. Proc 574
Natl Acad Sci. 2023;120:e2211117120. https://doi.org/10.1073/pnas.2211117120 575
11. Osmanski AB, Paulat NS, Korstian J, Grimshaw JR, Halsey M, Sullivan KAM, et al. 576
Insights into mammalian TE diversity through the curation of 248 genome assemblies. 577
Science. 2023;380:eabn1430. https://doi.org/10.1126/science.abn1430 578
12. Storer JM, Hubley R, Rosen J, Smit AFA. Methodologies for the De novo Discovery of 579
Transposable Element Families. Genes. Multidisciplinary Digital Publishing Institute; 580
2022;13:709. https://doi.org/10.3390/genes13040709 581
13. Bell EA, Butler CL, Oliveira C, Marburger S, Yant L, Taylor MI. Transposable element 582
annotation in non-model species: The benefits of species-specific repeat libraries using 583
semi-automated EDTA and DeepTE de novo pipelines. Mol Ecol Resour [Internet]. 2021 584
[cited 2021 Nov 23];n/a. https://doi.org/10.1111/1755-0998.13489 585
.CC-BY-NC 4.0 International licenseperpetuity. It is made available under a
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint
14. Berthelier J, Casse N, Daccord N, Jamilloux V, Saint-Jean B, Carrier G. A transposable 586
element annotation pipeline and expression analysis reveal potentially active elements in the 587
microalga Tisochrysis lutea. BMC Genomics. 2018;19:378. https://doi.org/10.1186/s12864-588
018-4763-1 589
15. Goerner-Potvin P, Bourque G. Computational tools to unmask transposable elements. 590
Nat Rev Genet. Nature Publishing Group; 2018;19:688–704. https://doi.org/10.1038/s41576-591
018-0050-x 592
16. Storer J, Hubley R, Rosen J, Wheeler TJ, Smit AF. The Dfam community resource of 593
transposable element families, sequence models, and genome annotations. Mob DNA. 594
2021;12:2. https://doi.org/10.1186/s13100-020-00230-y 595
17. Goubert C, Craig RJ, Bilat AF, Peona V, Vogan AA, Protasio AV. A beginner’s guide to 596
manual curation of transposable elements. Mob DNA. 2022;13:7. 597
https://doi.org/10.1186/s13100-021-00259-7 598
18. Rodriguez M, Makalowski W. Software evaluation for de novo detection of transposons. 599
Mob DNA. 2022;13:14. https://doi.org/10.1186/s13100-022-00266-2 600
19. Kofler R. SimulaTE: simulating complex landscapes of transposable elements of 601
populations. Bioinformatics. 2017;34:1419–20. https://doi.org/10.1093/bioinformatics/btx772 602
20. Haller BC, Messer PW. SLiM 4: Multispecies Eco-Evolutionary Modeling. Am Nat. 603
2023;201:E127–39. https://doi.org/10.1086/723601 604
21. Caballero J, Smit AFA, Hood L, Glusman G. Realistic artificial DNA sequences as 605
negative controls for computational genomics. Nucleic Acids Res. 2014;42:e99. 606
https://doi.org/10.1093/nar/gku356 607
22. Benson CW, Chen T-H, Thomson SJ, Deng CH, Ou S. PrinTE: A Forward Simulation 608
Framework for Studying the Role of Transposable Elements in Genome Expansion and 609
Contraction [Internet]. bioRxiv; 2025 [cited 2025 Nov 24]. p. 2025.06.04.657780. 610
https://doi.org/10.1101/2025.06.04.657780 611
23. Smit, AFA, Hubley, R, Green, P. RepeatMasker Open-4.0 [Internet]. 2013. 612
http://www.repeatmasker.org 613
24. Zhao D, Ferguson AA, Jiang N. What makes up plant genomes: The vanishing line 614
between transposable elements and genes. Biochim Biophys Acta BBA - Gene Regul Mech. 615
2016;1859:366–80. https://doi.org/10.1016/j.bbagrm.2015.12.005 616
25. Benson G. Tandem repeats finder: a program to analyze DNA sequences. Nucleic Acids 617
Res. 1999;27:573–80. https://doi.org/10.1093/nar/27.2.573 618
26. Quinlan AR, Hall IM. BEDTools: a flexible suite of utilities for comparing genomic 619
features. Bioinformatics. 2010;26:841–2. https://doi.org/10.1093/bioinformatics/btq033 620
27. Wells JN, Feschotte C. A Field Guide to Eukaryotic Transposable Elements. Annu Rev 621
Genet. Annual Reviews; 2020;54:539–61. https://doi.org/10.1146/annurev-genet-040620-622
022145 623
28. Sotero-Caio CG, Platt RN II, Suh A, Ray DA. Evolution and Diversity of Transposable 624
Elements in Vertebrate Genomes. Genome Biol Evol. 2017;9:161–77. 625
https://doi.org/10.1093/gbe/evw264 626
.CC-BY-NC 4.0 International licenseperpetuity. It is made available under a
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint
29. Tenaillon MI, Hollister JD, Gaut BS. A triptych of the evolution of plant transposable 627
elements. Trends Plant Sci. 2010;15:471–8. https://doi.org/10.1016/j.tplants.2010.05.003 628
30. Chen T-H, Winefield C. Comprehensive analysis of both long and short read 629
transcriptomes of a clonal and a seed-propagated model species reveal the prerequisites for 630
transcriptional activation of autonomous and non-autonomous transposons in plants. Mob 631
DNA. 2022;13:16. https://doi.org/10.1186/s13100-022-00271-5 632
31. Willing E-M, Rawat V, Mandáková T, Maumus F, James GV, Nordström KJV, et al. 633
Genome expansion of Arabis alpina linked with retrotransposition and reduced symmetric 634
DNA methylation. Nat Plants. Nature Publishing Group; 2015;1:14023. 635
https://doi.org/10.1038/nplants.2014.23 636
32. Wang J, Zhang G, Sun C, Chang L, Wang Y, Yang X, et al. DNA gains and losses in 637
gigantic genomes do not track differences in transposable element-host silencing 638
interactions. Commun Biol. Nature Publishing Group; 2025;8:704. 639
https://doi.org/10.1038/s42003-025-08127-3 640
641
642
.CC-BY-NC 4.0 International licenseperpetuity. It is made available under a
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
The copyright holder for thisthis version posted March 11, 2026. ; https://doi.org/10.64898/2026.03.09.710711doi: bioRxiv preprint
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.