Lateral gene transfer leaves lasting traces in Rhizaria | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Lateral gene transfer leaves lasting traces in Rhizaria Laura Eme, Jolien van Hooff This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4176859/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted You are reading this latest preprint version Abstract Lateral gene transfer (LGT) is a fundamental process that has contributed to the genetic makeup of various eukaryotic lineages. Yet our comprehension of its prevalence and significance across the eukaryotic domain remains incomplete, particularly when it comes to eukaryote-to-eukaryote transfers. The Rhizaria forms an expansive, ancient, and morphologically diverse clade of mostly free-living, single-celled phagotrophs, whose genetic diversity reflects their adaptability and evolutionary success across a range of habitats. Here, we undertake a comprehensive investigation of LGT within Rhizaria, tracing its role from the clade's emergence to the present day, employing advanced phylogenetic analyses and machine learning techniques. Contrary to previous assertions that LGT in eukaryotes is rare, our findings suggest that at least 8% of genes in contemporary rhizarian genomes were acquired through LGT at various points during their billion-year history. Our analysis reveals that although gene duplications exceed the number of LGT events, the duplicated genes that were originally acquired through LGT have a visible effect on the evolutionary trajectory of the recipient organisms. Our results also indicate that LGTs originating from other eukaryotes are more common than those from prokaryotes and exhibit unique patterns. By offering both quantitative and qualitative insights into the role of LGT in shaping the evolution of a major eukaryotic lineage, this work further demystifies lateral gene transfer in eukaryotes. Biological sciences/Microbiology/Microbial genetics Biological sciences/Evolution/Phylogenetics Biological sciences/Computational biology and bioinformatics/High-throughput screening Biological sciences/Evolution/Molecular evolution Biological sciences/Evolution/Evolutionary genetics Figures Figure 1 Figure 2 Figure 3 Figure 4 Introduction Lateral gene transfer (LGT) refers to the non-inheritance transfer of genetic material between unrelated species, and is well-documented among bacteria and archaea. Its role and prevalence in eukaryotes, however, have been more controversial and less clear, although increasing evidence suggests that LGT has played a significant role in the evolution of many eukaryotic lineages 1–4 . Questions have been raised about the molecular and cellular mechanisms underlying these transfers, their relative importance compared to gene duplication, and their long-term effects on the recipient genomes 5–7 . Myriad studies indicate that LGT has helped eukaryotes acquire new capacities, such as establishing stable relationships with their endosymbionts, adapting to or reverting from parasitism, or thriving in low-oxygen environments 8–12 . Most recently, LGT-derived genes were estimated to constitute 1% of all eukaryotic gene inventories 5 . Regardless, this estimate hints at a significant role of LGT in eukaryotic evolution and contrasts the claim that LGT does not have a long-lasting, cumulative impact 1 . Several factors complicate our understanding of LGT in eukaryotes. Primarily, research has often been directed at pinpointing LGTs that are unique to specific organisms, focusing largely on recent events. This significantly limits our ability to gauge the long-term significance of LGT and hinders a consistent comparison of LGT frequencies across different organisms using a standard set of criteria. Consequently, it remains unclear how the frequency of LGT compares to other evolutionary processes for gene acquisition, such as gene duplication or de novo gene invention. Additionally, the study of LGT between eukaryotic organisms has been relatively neglected compared to LGT from prokaryotes to eukaryotes, largely due to the inherent challenges in detection. State-of-the-art methodologies, which rely on discrepancies between gene and species trees, often fall short when analyzing closely related lineages. This is because single gene or protein phylogenies, derived from a limited phylogenetic signal in a small number of analyzed sites, may not be sufficiently resolved. In addition, independent gene loss across the tree can erroneously suggest LGT events, by grouping sequences from “distant” eukaryotes. As a result, single gene phylogenies frequently lack the necessary information to differentiate LGT between eukaryotes from vertical gene transmission confidently. In this work, we investigated the contribution of LGT to Rhizaria genomes by identifying and characterizing gene transfers from prokaryotic and eukaryotic donors, and by evaluating the relative impact of LGT, gene duplication and de novo gene invention to genome dynamics. Furthermore, we characterized how genes evolved after LGT and showed that they duplicate and get lost frequently. We assessed how LGT genes differed from vertically inherited genes, then leveraged these differences to train a machine learning model capable of identifying LGT-derived genes without phylogenies. In summary, our study offers a detailed, both quantitative and qualitative, exploration of LGT within the relatively unexplored Rhizaria, emphasizing the significant contribution of foreign genes to the genetic makeup of this largely microbial group. Results Extensive gene acquisition in Rhizaria through LGT from prokaryotes and eukaryotes To investigate LGTs in Rhizaria, we aimed to develop a sophisticated phylogeny-based approach. We first collected protein sequences from 29 rhizarians (Supplementary Table 1) and collected their homologs from other eukaryotes, prokaryotes and viruses. After discarding sets of homologs displaying potentially spurious or weak signal (see Methods), we inferred phylogenies for 40,951 protein families. In each tree, we identified the monophyletic clades of rhizarian sequences (hereafter called 'rhizarian clades') and determined their evolutionary origin. Briefly, and taking into account statistical support and after removing potential contaminants, we designated each rhizarian clade as 1) having a vertical origin, if it forms a monophyletic group with stramenopiles or alveolates (the closest relatives of rhizarians, all together forming the SAR group 13–15 , 2), having a lateral origin, if it is nested in a non-SAR clade, or 3), representing a gene invention if the tree only contains rhizarian sequences. We then inferred a species phylogeny of Rhizaria and mapped the LGTs, gene duplications and inventions onto it (Methods). We detected 13,282 LGT events into Rhizaria, at any time point after they diverged from stramenopiles and alveolates. These gave rise, on average, to 30% of the examined proteins of modern Rhizaria (Figure 1, Supplementary Table 2), a significant departure from the previously estimated 1% 5 (but see discussion below). This estimate is relatively consistent among species, whether they are represented by transcriptomic data, or by good-quality genomic data ( Bigelowiella natans : 24%, Plasmodiophora brassicae : 27%). If we ignore all rhizarian “clades” containing sequences from only a single species—thereby rigorously avoiding the misidentification of potential sequence contamination as LGT, yet underestimating real species-specific LGT events—we still identify 1,992 instances of LGT. These events account for approximately 20% of the proteins analyzed in current rhizarian species (Supplementary Figure 1). Notably, 9% of examined proteins result from an LGT into a relatively deep branch of the rhizarian tree, such as in the ancestors of Reticulofilosa, Monadofilosa or Rotaliida, respectively (labeled inner branches, Figure 1). This indicates that LGT has long-term and likely adaptive consequences in Rhizaria. However, and very importantly, this phylogeny-based approach only examined 12% of the proteins found across Rhizaria, due to the numerous quality checks and strict decisions we made to avoid false positives (Methods). Consequently, to investigate a higher proportion of the proteomes, we developed an original prediction approach based on a machine learning classifier. Briefly, we used features of LGT and non-LGT proteins, as identified by our phylogenetic approach, to train a Gradient Boosting classifier, an ‘ensemble’ method based on many decision trees. As input features, we used sequence characteristics such as the calculated HGT index, protein length, protein disorder, predicted subcellular localization and functional category (Methods, Supplementary Table 3). Using this approach, we analyzed another 71% of rhizarian proteins, in addition to those already analyzed using phylogenetics. The remaining proteins were excluded due to either the low probability associated with any type of origin, or the risk of contamination. We inferred that, on average, 8% of the classified proteins originated through LGT (Supplementary Figure 2). This percentage is significantly lower than the estimate obtained solely through the phylogenetic approach. The discrepancy arises mainly because the dataset included a large number of rhizaria-specific genes, which were classified as inventions and often not investigated by phylogenetics in the previous step (see Supplementary Text, 'Machine learning classifier predicts at least 8% of proteins derived from LGT'). Nonetheless, each of our approaches for tracking LGT suggests an estimate much greater than 1% 5 , leading us to argue that the actual percentage of genes in Rhizaria acquired through LGT is likely several folds higher than the eukaryote-wide estimate reported in this review. Given that we consider the phylogenetic approach the most reliable, we restricted all subsequent analyses to characterize LGTs to the results thereof. In order to estimate the LGT rate across Rhizaria, we summed all LGTs inferred between the last common ancestor of Rhizaria and an example modern-day species for which we have trustworthy genomic data, P. brassicae . We found that 726 LGT events contributed to its proteome (Figure 1). Using the SAR divergence time estimates (1,673-1,986 mya) 16 , this comes down to 0.37-0.43 LGTs per million years. If we consider the LGTs inferred by machine learning (Supplementary Figure 2), a total of 1118 LGT events corresponds to 0.56-0.67 LGTs per million years. For B. natans , these rates come down to 0.65-0.77 and 0.89-1.06 LGTs per million years, respectively. These rates are close to the 1 LGT per million years that was previously contested and considered a highly unlikely rate in eukaryotes 2 . It is important to highlight that these rates reflect the outcomes of both LGT and subsequent gene losses, leading to an underestimation of the actual LGT rate, constrained by what is detectable in the present day. We then aimed to assess the relative visible imprint of LGT onto modern genomes, compared to gene duplications (Figure 1) – the mechanism generally assumed to be the main source of genomic novelty in eukaryotes. Overall, gene duplications occurred three times more frequently than LGTs. However, this distribution was not uniform across all branches; in only 36 of the 57 lineages did we identify more duplications than LGTs. Notable exceptions were observed in the ancestors of Rhizaria (181 LGTs vs. 7 duplications) and Cercozoa (241 LGTs vs. 117 duplications). Interestingly, the prevalence of LGT over duplications did not correlate with any specific level of evolutionary depth but rather varied by clade. For instance, the branches within Monadofilosea predominantly experienced more LGTs than duplications, whereas the opposite was true for Foraminifera. This trend in foraminiferans might be attributed to their frequent endoreplication processes, which could lead to high rates of gene duplication as a side effect. These processes might also necessitate the separation of germline from somatic genomes 17 , potentially making LGT fixation more challenging, especially if foreign genes are inserted into somatic rather than germline DNA. We determined the source of each lateral gene transfer (LGT) by identifying the clade within which the Rhizaria are nested. Remarkably, we found that 48% of the LGTs originated from non-SAR eukaryotes, 39% from bacteria, and less than 1% from archaea. The remaining transfers were of unclear origin. A significant majority of the individual branches (46 out of 56) received a higher number of LGTs from eukaryotic sources compared to prokaryotic ones (Supplementary Figure 3, Supplementary Table 4). The few exceptions were all observed at the tips, rather than deeper within the tree, either indicating residual contamination in the sequence data or suggesting that eukaryote-derived genes are more likely to be preserved over the long term compared to those from prokaryotes. Supporting this, when excluding LGTs limited to a single species, the majority of LGTs were eukaryotic in origin: 51%, in contrast to 26% from bacteria, less than 1% from archaea, and 3% from unidentified prokaryotes. A notable tip where numerous LGTs from prokaryotes appear trustworthy is Paulinella chromatophora. It is a unique species that acquired a primary photosynthetic organelle independently of Archaeplastida, in which we confirmed a significant number of bacterial LGTs (402/790, 51%)—a figure higher than previously reported 18 . Interestingly, this species also received many eukaryotic LGTs (299 out of 790, 38%), a finding that has not been previously reported. The fate of laterally transferred genes We first sought to determine how frequently genes of prokaryotic origin assimilated to the host genome through spliceosomal intron acquisition 12 . In the species for which we have genomic data, most prokaryote-derived LGT genes contained at least one intron (Figure 2A, B. natans : 81% of prokaryote-derived LGT genes, P. brassicae : 63%, Reticulomyxa filosa : 69%; Supplementary Figure 4A). The detection of these introns serves two purposes: firstly, it confirms that these prokaryote-origin genes are not contaminants; and secondly, the canonical lengths of these introns indicate that they might be expressed by the recipient genome, a hypothesis that awaits validation from transcriptomic data. Noticeably, the youngest, lineage-specific LGT genes lack introns more often than older LGTs, suggesting a gradual intron acquisition over time (for more details, see Supplementary Text: ‘Introns in prokaryotic-derived LGTs’). To gauge the impact of these newly incorporated genes on the host organisms, we examined their evolutionary trajectories after transfer. If LGT genes duplicate frequently, they might confer a beneficial function, possibly benefiting from increased gene dosage. Although duplications of LGT genes have been documented (Vancaester et al. 2020; Eme et al. 2017; Xu et al. 2016; Shin, Doucet, and Pauchet 2022; Siddique et al. 2022; Sheikh et al. 2023; Sahu et al. 2023), a systematic comparison with native genes is pending. Here, we found that in most lineages (38/56), LGT genes underwent duplication more often than vertically inherited genes (Figure 2C), with significant differences observed in 16 lineages. This suggests these genes not only fulfill crucial adaptive roles, but also that their duplications also contribute to the large proportion of LGT-derived genes in modern-day rhizarian genomes. The potential for LGT genes to be lost more frequently than vertically inherited genes was also examined, as this would indicate their transient impact on lineage evolution. In the limited lineages with statistically significant differences, seven showed a higher loss rate for LGT genes, compared to five for vertically inherited genes, showing no distinct loss tendencies between the two gene categories (Figure 2D). Beyond duplication and loss, we explored the sequence evolutionary rates of LGTs and their propensity for protein domain gain and loss. Our findings indicate that LGT proteins often evolve significantly faster than vertically inherited genes in certain lineages (Supplementary Text, 'LGT sequence divergence'). Moreover, they tend to gain and lose domains more frequently across most lineages (Supplementary Text, 'Domain loss and gain’). In conclusion, our analysis suggests that LGT genes exhibit a more dynamic evolutionary process than vertically inherited genes. Mechanistics insights into LGTs Given that viruses are known to facilitate lateral gene transfer (LGT) in eukaryotes 19 , we searched all gene trees for viral sequences. We found that rhizarian clades originating from LGT were more likely to contain viral sequences than those inherited vertically (18% versus 13%, respectively, P<0.001). Although this finding does not provide definitive evidence, it confirms the prevalence of viral–eukaryotic gene transfer, regardless of directionality 20 , and suggests that viruses could have played a role in mediating some LGT events. We also investigated the genomic context of laterally acquired genes. It has been proposed that LGTs preferentially localize to gene-sparse, potentially heterochromatic areas of the genome to avoid disrupting the genetic integrity of the host 12 . We examined the gene density surrounding LGTs in B. natans and P. brassicae by analyzing the flanking intergenic regions (FIRs), which represent the distances to neighboring genes. For B. natans , we observed no notable difference in FIR lengths between LGTs and vertically inherited genes (Figure 2B). In contrast, in P. brassicae , LGTs were found to have significantly larger FIRs compared to vertically inherited genes (median FIR for LGTs: 628 bp, median FIR for vertically inherited genes: 505 bp, P=0.012), indicating a preference for integration into less gene-dense areas. Furthermore, we investigated whether LGTs are initially integrated into sparse genomic regions and subsequently relocate to denser, potentially more transcriptionally active areas over evolutionary time, as suggested by Husnik and McCutcheon (2018) 12 . Our comparison of FIR lengths for LGTs acquired at various time points within the Rhizaria tree against those of native genes revealed that, in P. brassicae , the most recent, species-specific LGTs are situated in the sparsest genomic regions (median FIR: 777 bp, Supplementary Figure 4E). This pattern supports the hypothesis that LGTs initially insert into gene-poor regions before possibly transitioning to gene-richer and presumably more active areas. Finally, to shed light on putative structural constraints on transfer mechanisms, we comprehensively analyzed LGT genes and the proteins they encode, drawing on features predicted for both modern-day sequences and their ancestors (Methods). Several studies have observed that LGT genes are typically shorter 21–23 , possibly because shorter genes are more likely to be fully integrated and functional upon transfer, even if only a small fragment of foreign DNA is incorporated. Our analysis indeed shows that LGT proteins are generally shorter than their native counterparts (Figure 3A, median LGT protein length: 226 amino acids vs. median native protein length: 248 amino acids, P<0.001). However, breaking it down by origin, proteins from eukaryotic LGTs are slightly longer than native proteins (median length: 258, P<0.001), hence the observed shortness primarily pertains to prokaryotic LGTs. This difference likely reflects the intrinsic length disparity between eukaryotic and prokaryotic proteins 24 , challenging the assumption that the integration of larger DNA segments poses a major barrier to LGT (also see Supplementary Text: ‘Protein lengths across datasets’). Transferred proteins harbor donor-specific signatures To explore the roles LGT genes play within their new hosts, we utilized DeepLoc 25 for localization predictions and EggNOG-mapper 26 for functional annotations. Similar to protein length, the functional signatures of LGTs differ between those donated by prokaryotes and eukaryotes. Overall, the predicted localizations of genes of prokaryotic origin depart more strongly from those of native genes, than those of eukaryotic origin. Prokaryotic LGTs are predicted to localize extracellularly at a rate 2.6 times higher than vertically-inherited genes (P<0.001; Figure 3B, 803 LGTs in total; Supplementary Figure 5E), suggesting a role in host-environment interactions. This is supported by significant enrichments in functional categories such as ‘Cell motility’ and ‘Defense mechanisms’ for prokaryotic LGTs (4.0 and 2.4-fold enrichment, respectively; P<0.001; Figure 3C), despite low absolute numbers of LGTs in these categories (19 and 36, respectively, Supplementary Figure 5G). Conversely, eukaryotic LGTs are most enriched in the nucleus (1.4-fold higher average localization probability compared to vertically inherited genes, P<0.001; Figure 3B, 1515 LGTs in total; Supplementary Figure 5F), a surprising and significant finding without an immediate explanation in terms of adaptive function. Notably, eukaryotic LGTs also show strong overrepresentation in ‘Defense mechanisms’ (3.1-fold enrichment, P<0.001, 53 LGTs in total; Supplementary Figure 4H), and more surprisingly, in informational processes like ‘Transcription’ and ‘Cell cycle control’ (2.2 and 2.1-fold enrichment, P<0.001; Supplementary Figure 5H), the latter two aligning with an enriched predicted nuclear localization. This suggests that LGTs can significantly influence key cellular and informational processes, challenging previous characterizations that emphasized extracellular localization and metabolic functions, which may have stemmed from analyzing predominantly prokaryote-to-eukaryote transfers. We highlight the remodeling of essential processes by foreign-origin genes, such as transcription factors and epigenetic regulators. For instance, several rhizarian lineages have acquired genes encoding Chromo domains (associated with histone modification and transcription regulation) from Opisthokonta. Similarly, Reticulofilosa acquired a translational regulator, Impact, from red algae, predicted to localize to plastids or mitochondria and play a role in stress response translation. These findings underscore the multifaceted contributions of LGTs to the evolution and functional diversification of recipient lineages. Highlighted cases of lateral gene transfers Here, we highlight several instances that exemplify the characteristics more commonly found in LGTs, as opposed to genes inherited vertically. In Foraminifera, our analysis reveals that LGTs frequently fall into the COG category of 'Nucleotide metabolism and transport,' and often originated from prokaryotes. A particularly interesting case is that of a Glutamine amidotransferase class-I enzyme (Figure 4A), which seems to have been independently transferred to other eukaryotic groups, such as alveolates and Discoba (Kinetoplastida), with the possibility of subsequent transfers among these groups. In Rhizaria, specifically within R. filosa , this gene has undergone multiple duplications post-transfer, underscoring its biological significance. The availability of a genomic assembly for R. filosa bolsters our confidence that these are genuine paralogs, rather than assembly artifacts common in transcriptomic data. Although the enzymatic activity of this protein—specifically, its role in catalyzing the removal of the ammonia group from glutamine—is well documented, its specific cellular functions within foraminiferans remain elusive. Intriguingly, our investigation also reveals that some sequences closely match those found in giant viral genes (GVOGs), suggesting the possibility that this gene might have been introduced into the foraminiferan lineage through a viral intermediary. Our other two cases involve the protoplast feeder Leptophrys vorax , which carries many LGTs predicted to be secreted or membrane proteins facing the extracellular environment. These typically come from either plants (Figure 4B) or bacteria (Figure 4C). Most sequences that descended from these LGTs are predicted to be secreted or in the cell membrane (Figure 4B,C; hexagons and triangles). The case highlighted in Figure 4B exemplifies the frequent duplication events that were previously discussed, showcasing their impact on gene families following their transfer. This LGT involves a protein presumptively originating from plants, annotated as 'Late embryogenesis abundant protein 2' (LEA-2, PF03168). Although this function appears to be specific to land plants, homologs can also be found in organisms such as the alga Klebsormidium nitens and various prokaryotes 27 , where it has been hypothesized to have a role in stress tolerance. We can speculate that this transferred gene fulfills a similar role in L. vorax . The second L. vorax signal-peptide containing LGT (Figure 4C) had no functional annotation by EggNOG-mapper. Its bacterial homologs are described as Hlyd-like secretion proteins or Cobalt-zinc-cadmium resistance proteins (CzcB), and they seem otherwise specific to bacteria (PF16576). This protein has been described as involved in resistance to metals like cadmium, zinc and cobalt 28 and might help L. vorax handle varying metal concentrations that can be found in aquatic systems. Discussion This study showcases the substantial contribution of lateral gene transfer to the gene content of eukaryotes, as exemplified by the large Rhizaria phylum. Our phylogenetic screen reveals that on average, LGTs account for approximately 30% of the proteins, which corresponds to individual species percentages ranging from 2.9% ( Sorites sp.) to 64.5% ( Mikrocytos mackini , see also Supplementary Text, ‘High percentage of LGT in Mikrocytos mackini ’) and a median estimate of 26.1%. These numbers significantly surpass the previously estimated 1% 5 and are comparable to the LGT content found in prokaryotic genomes 29 . Such high percentages underscore the scope and depth of our phylogenetic analysis, combined with the inclusion of eukaryote-derived LGTs, a component often overlooked. First and foremost, our findings emphasize the importance of considering LGT as a key evolutionary force shaping eukaryotic genomes. Although gene duplications are more common, LGTs represent a crucial source of genetic novelty, evidenced by the identification of over 13,000 LGT events across the Rhizarian tree of life. Ignoring LGT risks overestimating the number of species-specific genes and misunderstanding the evolutionary history of eukaryotic lineages. Moreover, LGT challenges conventional interpretations of patchy gene distributions in eukaryotes. Instead, it suggests a more dynamic evolutionary landscape than previously thought, marked by gene acquisition by LGT rather than mere vertical inheritance from a Last Eukaryotic Common Ancestor with an inconceivably large gene repertoire, followed by myriad losses. One of the primary challenges in analyzing transcriptomic data from organisms that cannot be grown in pure culture is the difficulty in completely removing potential contamination by foreign sequences, which might be mistaken for lateral gene transfers (LGTs). However, our most rigorous approach to filter these out yielded a conservative estimate, which suggested that 20% of the genes analyzed by phylogeny have originated from LGT. This estimate represents a lower bound due to its exclusion of species-specific LGTs, which are most likely numerous (Supplementary Figure 6) given the considerable phylogenetic distances among many species in our dataset. To fully appreciate the impact of LGT in this group, a deeper and more extensive sampling of the Rhizaria clade is necessary 30,31 . The very same factors that make Rhizaria susceptible to sequence contamination –complex feeding behaviors and (endo)symbiotic relationships– may promote the integration of foreign genes, leading to many LGTs in their genomes. This process of symbiosis-mediated LGT could, in turn, enhance the symbiotic relationships themselves, a phenomenon observed in other evolutionary lineages 32 . Our findings highlight the numerical significance and unique characteristics of eukaryote-to-eukaryote LGTs, distinguishing them from the more commonly studied prokaryote-to-eukaryote ones. A pivotal question centers on if and how these numerous genes derived from eukaryotes integrate into the recipient's cellular machinery. This question particularly applies to genes predicted to be involved in nuclear and cell cycle processes, given that these involve numerous components and interactions, with which such an LGT might need to interact. Eukaryote-derived LGT genes, however, may offer distinct advantages for integration. For instance, the presence of introns in these genes could potentially reduce their susceptibility to gene silencing mechanisms, like those mediated by the HUSH complex 33 . Furthermore, gene exchange via LGT may occur more readily between closely related species, as evidenced in grasses 34,35 , suggesting a propensity for such transfers within similar biological contexts. This potential bias for LGTs among more closely related lineages may imply that our current estimates of LGT events are conservative, since we did not consider LGTs from stramenopiles, alveolates, or other rhizarians, i.e. the closest relatives to the acceptor species. Importantly, our survey of LGTs within the “deep” parts of the Rhizaria phylogeny enabled us to map the trajectory of gene transfers over time, shedding light on how these genes evolved and potentially adapted within their new environments. Notably, we found that older LGT genes tend to have introns more frequently than their more recently acquired counterparts, aligning with expectations regarding gene evolution post-transfer, in particular from prokaryotic donors. Interestingly, in the case of P. brassicae , we also observed that older LGT genes are situated in regions of the genome with higher gene density compared to newer LGTs, suggesting that these genes may migrate from areas of lower to higher gene density over time. This phenomenon of intragenomic migration, coupled with the acquisition of introns, could potentially facilitate increased gene expression. Consequently, these observations raise an intriguing question about the expression levels of older LGT genes compared to newer ones, suggesting that time might play a crucial role in the integration and functional significance of transferred genes within their new genomic contexts. Through our innovative machine learning approach, we ventured into uncharted territory to identify lateral gene transfer (LGT) candidates where traditional phylogenetic methods fall short. Our initial efforts focused on analyzing limited characteristics inherent to the protein products of genes, relying on predictive rather than experimental data. Incorporating additional information, such as transcriptional activity levels or specific genomic attributes of the genes in question, could significantly improve the accuracy of our model. For example, the surrounding gene density—an aspect highlighted earlier—may serve as an additional indicator for identifying LGT events. Ideally, adding such comprehensive data would enable our model not only to determine LGT likelihood, but also to predict a gene's origin, be it prokaryotic or eukaryotic, thereby offering deeper insights into the evolutionary pathways of gene transfer. In conclusion, our investigation into lateral gene transfer (LGT) within Rhizaria illuminates the evolutionary dynamics of this enigmatic clade, providing valuable insights and methodologies for the functional characterization of LGT-derived genes. The ability to systematically assess LGT patterns, facilitated by our detection approaches, extends beyond Rhizaria to a broad spectrum of genomically characterized eukaryotic clades. This comprehensive analysis not only promises a more complete understanding of LGT's quantitative and qualitative impacts across the eukaryotic tree but also highlights the intricate interplay between biological interactions and genetic exchange in these organisms. Our study underscores the nuanced nature of LGT, revealing the challenges in distinguishing LGT from vertical inheritance, particularly in closely related lineages where gene trees may not offer clear differentiation. This complexity necessitates refined analytical techniques to accurately identify genuine LGTs, which would advance our understanding of evolutionary processes in Rhizaria and across the diverse landscape of eukaryotic life. Methods Rhizaria species phylogeny reconstruction We curated a dataset comprising predicted protein sequences from 29 Rhizaria lineages, sourced from both genomic (four) and transcriptomic data (25) (Supplementary Table 1). (Note here that the genomic assembly of R. filosa exhibited considerable fragmentation, with, for instance, a median of one gene per scaffold. Consequently, we predominantly treated it as a 'transcriptomic' dataset in most subsequent analyses, except when specifically examining introns.) To infer the species phylogeny of these 29 lineages, we employed PhyloFisher, an advanced phylogenetic pipeline designed for eukaryotic species tree inference 36 . Briefly, this pipeline facilitated the systematic search and manual curation of orthologs across 240 marker genes within our focal species, culminating in a concatenated alignment utilized for phylogenetic inference. We initiated the PhyloFisher process using its provided scripts, including config.py, fisher.py, informant.py, working_dataset_constructor.py, sgt_constructor.py, and forest.py, all with their default settings. Given that the PhyloFisher v.1.0 dataset encompasses diverse Rhizaria species, we leveraged the 'phylogenetically-informed' approach within the 'fisher.py' step by specifying the rhizarian species in that dataset as the 'Blast Seed.' To classify the Rhizaria sequences from our dataset and, if necessary, those from the PhyloFisher dataset, as ortholog, paralog or contamination, we utilized ParaSorter, a graphical user interface tool integrated into PhyloFisher. Furthermore, we employed 'apply_to_db.py' to incorporate our taxa into the PhyloFisher dataset and 'select_orthologs.py' to exclude two marker genes (H2A and PYGB) due to the complexity of selecting orthologs using the single gene trees. Subsequently, we used 'prep_final_dataset.py' and 'matrix_constructor.py' to build the concatenated alignment. Details regarding the coverages of the 29 focal Rhizaria lineages in this alignment are provided in Supplementary Table 1 ('Supermatrix coverage'). The eukaryotic species tree encompassing these 29 Rhizaria lineages was inferred using IQ-TREE v.2.0.3 37 under the LG+G4+C60+F model, with ultrafast bootstrapping (1000 replicates). We arbitrarily rooted the tree on the branch uniting Obazoa+Amoebozoa+CRuMs and Metamonada+Discoba. Within Rhizaria, all branches but one were well-supported (>80%). However, M. mackini clustered within Metamonada, probably due to a long-branch artifact 38 . Consequently, in the Rhizaria phylogeny utilized for mapping LGT events (as illustrated in Figure 1), we pruned the M. mackini branch and grafted it as a sister to Gromia sphaerica , in accordance with its position reported previously 38 . The original phylogeny is available in Supplementary Figure 7. “Rhizaria-and-sister-clades” dataset assembly In addition to Rhizaria, we extensively collected sequence data from members of Stramenopila and Alveolata (collectively referred to as Halvaria), given their status as the closest relatives of Rhizaria. As detailed below, any gene shared between Rhizaria and Halvaria was deemed ancestral to the SAR supergroup in our analyses. This comprehensive sampling of Rhizaria's sister clades was undertaken to mitigate the risk of erroneous claims regarding lateral gene transfers (LGTs). Protein families construction We utilized OrthoFinder (version 2.3.8) to identify orthogroups within a dataset of 418 SAR proteomes 39 . From this comprehensive set of orthogroups, encompassing singleton 'orthogroups' as well, we retained only those containing at least one sequence from our 29 Rhizaria representative proteomes (n=341,035). Subsequently, we compiled sequences from non-SAR eukaryotes and viruses available in the NR database (as of May 20th, 2020) and downsampled these sequences to 90% identity using CD-hit (version 4.8.1) 40 , in order to reduce the computational load associated with homology searches and subsequent analyses. For Bacteria and Archaea, we sourced data from GTDB release 89 (Parks et al., 2020), downsampling these sequences with CD-hit at a 70% identity threshold. We conducted homolog searches using the rhizarian sequences within the orthogroups as queries against all the data listed above (non-SAR eukaryotes, viruses, archaea and bacteria); we used DIAMOND v2.0.9 41 in 'ultra-sensitive' mode, applying an e-value cutoff of 0.001. To mitigate biases stemming from uneven taxonomic representation in databases, we performed these searches separately against seven distinct databases: Archaea, Bacteria, Metazoa, Viridiplantae, Fungi, all other eukaryotes, and viruses. In each search, the maximum number of target sequences was restricted to 2000. To ensure that the rhizarian sequences serving as queries were not outliers within their respective orthogroups, those longer than twice the median length of the orthogroup were ignored. This precautionary step aimed to prevent potential artifacts arising from abnormally long sequences, such as additional protein domains, which could lead to spurious hits compared to other members within the orthogroup. Finally, we incorporated outgroup hits from all seven outgroup datasets into the SAR orthogroups corresponding to the rhizarian queries. Single protein tree inference For the reconstruction of individual protein trees, we only selected the expanded orthogroups comprising a minimum of 15 and a maximum of 6000 members. Groups with fewer than 15 sequences were excluded due to the challenge of reliably inferring transfer directionality. Conversely, groups exceeding 6000 sequences were deemed computationally intensive with a corresponding phylogeny likely unresolved. Nevertheless, to accommodate these large orthogroups exceeding 6000 members, we ‘recreated’ these by retrieving fewer homologues from each database (500 instead 2000). Any resulting orthogroups that still contained more than 6000 members at this step were not retained for further analysis. Subsequently, sequences within each orthogroup were aligned using MAFFT v7.407 42 in 'auto' mode, followed by alignment trimming with BMGE v1.12 43 using specific parameters (-m BLOSUM30 -b 3 -g 0.7 -h 0.5). Alignments with fewer than 50 positions post-trimming were discarded to avoid spurious topologies arising from weak phylogenetic signal. Additionally, individual sequences displaying over 80% gaps in the trimmed alignment were removed. Alignments containing fewer than 15 sequences were also excluded from subsequent phylogenetic inference. Finally, phylogenies for the remaining orthogroups were inferred using IQ-TREE v2.0.3, under the LG+F+R5 model. Branch support values were obtained via the SH-like approximate likelihood ratio test with 1000 replicates (-aLRT 1000) (Guindon et al., 2010). All 40,951 resulting trees are available as Newick files in Supplementary Dataset 1. LGT detection based on single gene trees Summary of the workflow. We comprehensively screened 40,951 single protein trees to identify potential lateral gene transfers (LGTs) using the ETE3 toolkit 44 . Specifically, we inferred the evolutionary origins of rhizarian sequences within these trees. These origins were categorized as either 'vertical' (indicating parent-offspring inheritance), 'lateral' (signifying an LGT event at the base of the rhizarian clade or taxon of interest), or 'invention' (implying a novel gene, unique to Rhizaria). This categorization was based on the taxonomic affiliation of the clade within which the rhizarian ones were nested. To elaborate, a vertical origin was inferred if the closest sister clade, or second closest sister clade (sister to group formed by the rhizarian clade plus its closest outgroup) consisted of SAR sequences. Conversely, a lateral origin was suggested if these two sister clades were primarily composed of prokaryotes, or eukaryotes from non-SAR groups. Lastly, a clade was classified as a rhizarian invention if the entirety of the tree was predominantly composed of rhizarian sequences. Detailed workflow. The analysis process is outlined in Supplementary Figure 8. Initially, we annotated each leaf in the single gene trees with a taxonomic classification using NCBI Taxonomy 45 for eukaryotes and viruses (as of October 4th, 2021) and GTDB for prokaryotes (release 89). Leaves were categorized into groups: prokaryotes, non-SAR eukaryotes, halvarians (Stramenopila and Alveolata), rhizarians, viruses, or unknown (Supplementary Figure 8B). Trees without rhizarian sequences were excluded from the analysis. Recognizing viruses primarily as gene transfer agents rather than original gene sources, we removed viral sequences, while noting their proportion within each orthogroup. Sequences of ‘unknown’ sources were also removed. For trees with 80% or more rhizarian sequences, we classified the tree as a rhizarian invention, attributing non-rhizarian sequences to either contamination or lateral gene transfers (LGTs) from Rhizaria. In other cases, to discern whether rhizarian sequences derived from vertical or lateral origins, we first identified all rhizarian clades in a tree after rooting the tree on a random non-rhizarian leaf with a non-rhizarian sister clade (Supplementary Figure 8C). Acknowledging potential inaccuracies due to single-gene tree limitations, we merged rhizarian clades that appeared artificially paraphyletic. Two criteria guided this merging: clades were considered monophyletic if only one non-rhizarian sequence was breaking this monophyly, presumed to be contamination or LGT (Supplementary Figure 8D); and clades were merged if fewer than two strongly supported branches (with SH-aLRT support values ≥ 0.8) separated them, thereby correcting for the possible poor resolution of the tree. We merged such rhizarian clades by pruning one and grafting it at the base of the other, keeping the non-rhizarian sister clades intact (Supplementary Figure 8E). After merging, the total number of rhizarian clades was tallied (Supplementary Figure 8F). Trees containing more than five rhizarian clades were deemed as likely too ambiguous for origin analysis due to insufficient phylogenetic signal. We aimed to then distinguish authentic "rhizarian" clades from those representing contaminant sequences. If a clade contained multiple rhizarian taxa, we considered it less likely to be the result of contamination. However, for clades consisting of sequences from a single species, we discarded the entire clade if at least one of its sequences displayed high identity to any stramenopile or alveolate sequence (>90%), or to any other eukaryotic or prokaryotic sequence (>80%), as determined by a DIAMOND BLASTP search with parameters ultra-sensitive -k 1 --max-hsps 1 --query-cover 70 (Supplementary Figures 8G, 9). To ensure a comprehensive search, we used our large SAR database for stramenopile and alveolate searches and NR database for other sequences. Subsequently, we determined the evolutionary origin of the remaining rhizarian clades by identifying their sister lineages. We first annotated internal nodes and leaves in the tree using NCBI Taxonomy and GTDB’s taxonomy (Supplementary Figure 8H). For a node to be annotated, we require its parent branch to be well-supported (>0.8 SH-like support). We annotated nodes at the lowest taxonomic level possible if it represented the last common ancestor of at least 80% of their descendant leaves. For example, we annotated a node that contained 20 Bilateria and three Euglenozoa leaves, as ‘Bilateria’. Allowing for such ‘interspersing’ sequences takes into account the possibility that the node’s composition was affected by LGT, contamination or other artifacts, while likely reflecting the actual ancestral nature of the clade. After annotation, we examined the identities of sister clades of the rhizarian clade in an unrooted configuration. If both sister clades comprised Stramenopila and/or Alveolata, we rooted the tree on the ancestral rhizarian node (i.e. the last common ancestor of all rhizarian sequences in the clade of interest), in line with the generally accepted position of Rhizaria in the eukaryotic tree of life, as sister-group to both Stramenopila and Alveolata. Alternatively, we selected the sister clade with fewer SAR sequences and rooted on the leaf with the longest branch to the rhizarian clade. If neither sister clade contained SAR sequences, we chose the leaf with the longest branch to the rhizarian clade. By parsing the tree from leaves to root, we then identified the two first parent nodes with support values above 0.8, upward from the rhizarian clade (Supplementary Figure 8I). In some cases, only a single parent node was found. Each parent was annotated broadly as 'SAR', 'eukaryotes', 'prokaryotes', or 'mix' (if it contained both prokaryotic and eukaryotic sequences), and more specifically when possible (e.g., 'Metazoa'). If either the first or second parent was identified as 'SAR', we inferred a vertical origin for the ancestral rhizarian sequence (Supplementary Figure 8J). Similarly, if the clade descending from the parent contained a diverse range of eukaryotes but lacked Stramenopile and Alveolata representation, we also inferred a vertical origin, suggesting potential gene loss from stramenopiles and alveolates. Conversely, if the clades descending from the parent nodes included (i) prokaryotes, (ii) specific lower-level eukaryotic groups (e.g., 'Metazoa'), or (iii) a mix of prokaryotes and eukaryotes, we inferred lateral gene transfer (LGT) as the origin of the rhizarian homologs. Detailed configurations and associated inferred origins for the ancestral rhizarian sequences are provided in Supplementary Table 5. Subsequently, we inferred the taxonomic identity of the putative donor based on the identity of the first or second parent (see Supplementary Table 5 for details). To prevent erroneous origin inferences of highly divergent sequences (which are often erroneously placed in single gene trees), we excluded rhizarian clades from subsequent analyses if the median tip-to-tip distance between the rhizarian sequences and those from the first sister clade exceeded four expected substitutions per site (Supplementary Figure 8K). Following the analysis of all single-gene trees, we validated the lateral origins identified in B. natans and P. brassicae using available genome data to ensure that the identified LGTs were not due to contamination (Supplementary Figure 8L). For this, we stipulated that a gene corresponding to an inferred LGT event must be located on a scaffold containing at least one gene inferred to have been inherited vertically. It is important to note that while the proteomes of R. filosa and Globobulimina sp. were derived from genomic data, we could not use them for this validation step. The assembly of R. filosa primarily consisted of very short scaffolds, typically containing only a single gene, and the annotated genomic sequence for Globobulimina sp. was not publicly available. Overall, our analysis of the single-gene trees enabled us to determine the origins of 12% of the proteins within our 29 rhizarian lineages. Supplementary Figure 10 illustrates the percentage of sequences remaining after each step in our workflow. Predicting LGT using protein sequence features A substantial fraction (69%) of all initial rhizarian sequences were not assigned an evolutionary origin through our phylogenetic approach because they belonged to either small orthogroups (6000 sequences, even after downsampling of the outgroup). To get a more complete picture of evolutionary dynamics in Rhizaria, we explored several approaches to predict the origins of these sequences by non-phylogenetic means. All approaches operate at the level of individual sequences. The individual results were then combined to infer the origin of the cluster they belong to. As a mean to benchmark these approaches, we first tested them on orthogroups that we could analyze through phylogenetics, and for which rhizarian sequences formed a single monophyletic clade in the tree; we regarded them as representing the most straightforward evolutionary history, and thus trusted that we correctly inferred their origins. This constituted a test set of 35,876 sequences which we inferred to belong to rhizarian clade of either vertical or lateral origin (21,224 and 14,652 sequences, respectively). Note that here, the assigned origin of a sequence corresponds to that of its ancestral rhizarian node (i.e., to the entire clade it belongs to; see ‘Detecting LGT with single gene trees’). We first tested HGT index, which uses the ratio of the DIAMOND search bit score of the best ‘alien’ hit (non-SAR eukaryotes and prokaryotes) divided by the bit score of the best ‘native’ hit (stramenopile or alveolate). Considering the large overlap in HGT indices for sequences from vertical and lateral origins, we did not expect an excellent performance from this approach (Supplementary Figure 11A). Indeed, using the default threshold value of 1 46 as indicative of an HGT, this method showed a 70% accuracy rate (Supplementary Figure 11B), with a relatively high 31% rate of false positives (i.e., vertically inherited genes classified as HGT). We also tested a linear discriminant analysis, a machine learning technique, leveraging the scikit-learn software library 47 to distinguish between two types of genetic origins: LGT and vertical inheritance. This approach aimed to identify a boundary that best separates these two groups. Despite achieving an overall accuracy of only 59%, this method showed a significant improvement in reducing the misclassification of vertically inherited genes as LGT to just 6%, a notable decrease compared to the 31% misclassification rate using the default threshold in the HGT index method (Supplementary Figure 11C). However, the moderate success rate of this and similar HGT index-based approaches was anticipated, given the substantial overlap in HGT index values between sequences of vertical and lateral origins, indicating a limited ability of the HGT index alone to accurately predict genetic origin. Recognizing the limitations of relying solely on the HGT index, we explored alternative strategies to enhance origin prediction by utilizing distinguishing features of LGTs (see ‘Transferred proteins harbor demarcating signatures’). Here, our approach involved analyzing individual protein sequences for diagnostic characteristics (Supplementary Table 3), rather than assessing gene families evolutionary histories. We meticulously compiled these features and prepared our data for machine learning analysis, employing one-hot encoding to manage categorical variables. Our evaluation of various classifiers available in scikit-learn focused on those capable of handling our dataset's mix of categorical and numeric features. Among the techniques we tested, Gradient Boosting emerged as the most effective. This ‘ensemble’ method, which combines multiple decision trees, outperformed others in accuracy, the area under the curve (AUC), and the F1 score. We refined our Gradient Boosting model through hyperparameter tuning with a five-fold cross-validation on the training set. Based on this, we trained the classifier with 1,000 decision trees (n_estimator) and a maximum tree depth of 20 for each tree, to prevent overfitting. Additionally, we addressed potential biases from imbalanced data (disproportionate numbers of lateral vs. vertical genes) by balancing these two categories in our training sample. Training on 80% of this balanced dataset and testing on the remaining 20%, our optimized Gradient Boosting classifier achieved a 78% accuracy rate. This performance significantly surpassed that of methods relying solely on the HGT index, also reducing the false positive rate to 26%, although this remains high (Supplementary Figure 11D). Further analysis of our classifier's effectiveness included examining the significance of various features in the model, using the "feature_importances_" attribute to assess their roles based on Gini impurity (Supplementary Table 3). Further building on our observations that eukaryotic lateral gene transfers (LGTs) differ significantly from prokaryotic LGTs, we explored a more complex model capable of distinguishing among three categories: 'vertical' (inherited within the species), 'LGT eukaryotic' (gene transfers from eukaryotes), and 'LGT prokaryotic' (gene transfers from prokaryotes). Using a gradient boosting model adapted for this multi-class classification, we found that the accuracy dropped to 70% from the 78% achieved with the simpler two-class model. This decrease in accuracy primarily stemmed from difficulties in distinguishing between vertically inherited genes and those acquired from eukaryotic LGTs (Supplementary Figure 11E). Specifically, 27% of the genes we identified as vertically inherited were misclassified as eukaryotic LGTs, and 7% as prokaryotic LGTs. This supports our earlier findings that eukaryotic LGTs often closely resemble genes inherited vertically, more so than prokaryotic LGTs do. Given our goal was to estimate the overall rate of LGT rather than to precisely categorize eukaryotic and prokaryotic LGTs, we decided to continue using the more straightforward two-class model (Supplementary Figure 11D). We also sought to identify the origins of orthogroups that were not amenable to analysis with our previous phylogenetic methods. For OGs consisting of more than 80% rhizarian sequences (excluding viruses), we inferred an origin through a process known as 'invention' in a manner similar to our phylogenetic analysis. For the remaining orthogroups, we applied our gradient boosting classifier. Recognizing that some orthogroups could contain genes from multiple origins within Rhizaria, we first divided these orthogroups into smaller clusters using MCL 48 with an inflation factor of 1.2. This approach resulted in most orthogroups forming a single cluster, with only a small percentage showing more complex patterns indicative of multiple origins. For each cluster resulting from MCL, we then used our classifier to predict the origin of each gene as 'vertical' or 'lateral'. We also calculated the average probability of belonging to either of these two categories for all genes within a cluster. A cluster's origin was determined to be the category with the highest average probability, but only if this probability exceeded a 70% confidence threshold, discarding the most dubious cases. Additionally, to ensure the reliability of our predictions and avoid mislabeling due to potential contamination, we applied the same criteria used in our phylogenetic analysis. This involved checking for unusually high sequence similarities to non-rhizarian sequences and, in specific cases, confirming the presence of native genes on the same genetic scaffold when possible. The native nature of genes was deducted from the earlier phylogenetics-based detection strategy. Comprehensive reconstruction of gene evolutionary histories After identifying the origins of rhizarian genes, we carried out a comprehensive reconstruction of their evolution across the Rhizaria tree of life. This reconstruction mapped out gene duplications, transfers, and losses, utilizing two sets of data: one derived solely from the phylogenetic LGT detection (Figure 1) and another that combined both phylogeny-based detection and machine learning or distribution-based predictions (Supplementary Figure 2). We started by compiling protein sequences from rhizarian clades or clusters. When an origin was determined based on species distribution, clusters corresponded directly to orthogroups. However, in some cases, orthogroups were subdivided during the MCL stage. For rhizarian clades containing at least three sequences, we aligned them using the MAFFT ‘linsi’ tool and refined the alignment with BMGE (settings: -m BLOSUM45 -b 3 -g 0.7 -h 0.5). We then employed GeneRax v2.0.1 49 to infer gene trees containing only rhizarian sequences and reconcile these with the overall Rhizaria species tree, using the LG+G substitution model and the UndatedDL reconciliation model. For nodes or clusters with fewer sequences, we used Notung v2.9.1.5 50 for pairs of sequences, or created a hypothetical phylogenetic branch for single sequences, ensuring each scenario was accommodated in our analysis. As input for Notung, we either used the pruned subtree corresponding to the rhizarian clade, or, for the results from the prediction approach, inferred a tree using the alignment and trimming procedure as described above (‘Large-scale inference of single gene trees’) and subsequently FastTree v2.0 51 (settings: -lg -gamma). We annotated all resulting reconciled single gene trees with the information collected during the detection or prediction, most importantly the type of origin (vertical, lateral or invention) and, in case of a lateral origin, the donor clade. To ensure a comprehensive reconstruction, we updated the reconciled gene trees to include inferred branches representing either missing ancestral nodes or gene losses, reflecting our analysis findings. These additions were crucial for accurately mapping gene evolution. For example, when our analysis (either detection or prediction) indicated that the origin of a rhizarian clade was vertical (i.e. inherited from an older eukaryotic ancestor), but that the root of the reconciled gene tree was mapped onto a more recent node than the last common ancestor of Rhizaria (e.g., 'Foraminifera'), we addressed this discrepancy by adding the missing branches at the base of the tree. These additions, creating previously missing internal nodes such as 'Retaria' and 'Rhizaria', effectively adjusted the ancestral rhizarian node back to the common ancestor of Rhizaria. These newly added sister branches were identified as representing 'lost' branches. We also annotated them accordingly, including the clade we hypothesized they were lost from, like 'Radiolaria' and 'Cercozoa' in our example. By annotating these newly added branches and internal nodes, we captured a more detailed picture of rhizarian gene history, including duplications, transfers, and losses, and ancestral gene content. For all the branches we added to the gene trees, we assigned a length of zero. After making these adjustments, we proceeded to gather or infer specific information for each node (both internal and external, except for those representing lost genes) within the updated reconciled gene trees. This information included: (1) the origin of the gene, categorized as vertical, lateral, invention, or duplication. This categorization was based on the outcomes of our detection or prediction analyses or whether the node was identified as a result of duplication, or represented a duplication itself; (2) the source clade, in cases where the gene's origin was lateral; (3) the median distance from the node to the tree leaves; (4) the total count of gene duplications and losses among the node's descendants, providing insight into the gene's evolutionary dynamics. We introduced the concept of a 'residence index' of a gene as a measure to capture its retention across the species tree. This metric is derived by first pruning the species tree for the species that retained this gene, which therefore reflects the branches in which this gene resided, and then summing the branch length of that pruned species tree. The 'residence' metric is crucial for normalizing the estimates of gene duplication and loss. Indeed, if a gene is heavily retained, we have more chances of observing gene duplication and loss events than if it got lost early on in the tree (or is inferred as such due to missing data). Similarly, an older gene will have had more time to duplicate than a younger one. This metric thus allows us to more accurately compare duplication and loss frequencies between gene families regardless of their distribution across the species tree.Finally, we aggregated the data from all reconciled single gene trees to tally LGTs, gene duplications, and losses across all branches of the species tree. This comprehensive approach also enables us to infer which genes existed in the ancestral lineages leading to the current members of the Rhizaria included in our study. Characterizing LGT evolution and function using gene evolutionary histories We used reconciled single-gene trees to explore whether genes acquired through lateral gene transfer (LGT) exhibit distinctive evolutionary dynamics, functional, and structural characteristics compared to genes inherited vertically. For the most reliable insights, we exclusively employed data derived from a phylogeny-based detection approach. We used the set of genes deduced to have been present in a particular lineage within the rhizarian species tree based on the inferences described above. We compared genes with a presumed 'vertical' origin in Rhizaria against those acquired via LGT directly in the same lineage. Thus, we excluded genes that, despite being vertically inherited from the parent, were originally acquired in the Rhizaria through LGT in an older ancestor. This exclusion was based on the premise that such genes might blur the distinction between vertical and lateral transfer, as they may have already adapted to the rhizarian host and exhibit characteristics more akin to vertical inheritance. We investigated whether LGTs show distinct patterns of gene duplication and loss compared to vertically inherited genes. For each branch, this involved tallying the occurrences of gene duplications and losses among their descendants and normalizing these counts by the ‘residence index’ described above. This aimed to account for missing data and to facilitate comparisons across taxonomic levels. For calculating normalized losses, we only considered losses detected in the initial reconciled single gene tree, excluding losses attributed to the reassignment of vertically inherited nodes within Rhizaria (see ‘Comprehensive reconstruction of gene evolutionary histories’). To determine if the origin of genes (vertical vs lateral) influenced the normalized rates of duplication and loss, we conducted a two-sided Mann-Whitney U test, assuming both categories had variable distributions (i.e., more than one unique normalized count). We compared the mean normalized counts for each gene origin type and analyzed the differences between these means. Similarly, for LGTs, we calculated the median branch length from root-to-tip and normalized these lengths using the root-to-tip median branch length in the species tree. Additionally, we examined the dynamics of protein domain acquisition and loss in genes that were either inherited vertically or acquired laterally. This analysis was based on Pfam domain annotations of genes at each node of the rhizarian tree, and the node right above it. For this, we used HMMER's hmmscan against Pfam-A database version 3.1b2 52 to analyze sequences from the sister clade, results which we combined to the domain annotation already done for the rhizarian sequences. By tracking the domain changes (gains and losses) at each node within the rhizarian gene tree, relative to its ancestral state and throughout its descendants, we quantitatively assessed these changes. Like for gene duplications and losses, the counts of domain dynamics were normalized considering the age and retention of the gene in the rhizarian clade (using the ‘residence index’ described above). For the structural and functional analysis of lateral gene transfers (LGTs), we predicted various attributes of rhizarian protein sequences (Supplementary Table 6). These included intrinsic protein disorder with MobiDB-lite v3.10.0 53 , coiled-coil regions and many protein family annotations from InterProScan v5.48-83.0 54 , protein subcellular localization predictions with SignalP v5.0b 55 , TargetP v2.0 56 and DeepLoc v1.0 25 , transmembrane domains with Phobius v1.01 57 and TMHMM v2.0c 58 , viral and NCLDV affinities using VOGDB and GVOG datasets from ViralRecall v2.0 59 , carbohydrate enzyme activities with dbCAN2 60 and COG functional categories with eggNOG-mapper v2.1.5 26 . For each ancestral node identified as having acquired genes via LGT, we annotated its descendants. This included collecting data on all measurable features from the leaf nodes. Depending on the feature's nature (qualitative or quantitative), we applied appropriate methodologies for projection onto internal nodes. For instance, Pfam annotations were treated with Dollo parsimony, requiring presence in at least 10% of leaf nodes for consideration. Numerical attributes, like protein length, internal nodes were annotated by using median values across its descending leaves. We compared these annotated features between LGTs and vertically inherited genes, looking for statistically significant differences. This comparison was conducted either branch-specifically within the species tree or across all examined branches, employing chi-square tests for boolean data and Mann-Whitney U tests for numeric data (Figure 3). We reported either the presence or absence of certain features (for boolean data) or median values (for numeric data), ensuring comprehensive analysis of structural and functional gene properties. This extensive examination allows us to discern any significant distinctions between the two gene origins across various biological and molecular aspects. Characterizing LGT intron content and genomic context using individual sequences To analyze intron acquisition, we extracted gene information from the GFF files of the genome assemblies for B. natans , P. brassicae , and R. filosa . This data included the count of introns per gene, the sum of introns lengths, the intronic proportion of each gene, and the average length of introns within a gene. We calculated the median values for these four metrics and determined the percentage of genes containing at least one intron. These calculations were performed separately for both vertically inherited and laterally acquired proteins within each species. Moreover, we categorized laterally acquired proteins based on their acquisition timeline, allowing for a comparative analysis of intron properties between 'young' and 'old' LGTs. We employed the Mann-Whitney U test or the chi-square test to examine if the intron characteristics differed between laterally acquired and vertically inherited genes. Additionally, we specifically analyzed LGTs believed to have been transferred from eukaryotes and those from prokaryotes, comparing their intron features with those of vertically inherited genes to discern any distinct patterns. In our examination of the genomic context surrounding LGTs, we were unable to include R. filosa due to its highly fragmented genome. Consequently, we investigated gene densities, viral gene abundance, giant virus gene abundance, and transposable element (TE) abundance in B. natans and P. brassicae . To operationalize these genomic features, we calculated the distances to the nearest gene, virus-annotated gene (identified using VOG HMM profiles), giant virus-annotated gene (identified using GVOG HMM profiles), and TE for each species, examining both strands of DNA for a comprehensive assessment. TE annotations of the genomes were obtained with RepeatModeler v2.0.1 with the Run the LTR structural discovery pipeline 61 and RepeatMasker v4.1.1 62 with the Dfam TE Tools Container v1.2 63 . We determined the median distances of each genomic feature from both vertically inherited and laterally acquired proteins, further categorizing the latter by their time of acquisition to discern patterns between 'new' and 'old' LGTs. To assess whether the spatial genomic contexts of these proteins differed based on their origin or acquisition timing, we utilized the Mann-Whitney U test to compare the distribution of distances. Selection and gene tree inference of illustrative LGTs We analyzed the enrichment of specific COG categories and subcellular localizations in LGTs compared to vertically inherited genes, and determined common donor groups for these LGTs. Focusing on lineages that exhibited both functional overrepresentation and a high frequency of specific donors, we meticulously examined the gene phylogenies of these LGTs for confirmation. We inferred a phylogeny using IQ-TREE under a more sophisticated evolutionary model (LG+C60+G) with 1000 ultrafast bootstraps. For large orthogroups, we streamlined the dataset by including only the 1000 non-rhizarian sequences closest to the rhizarian sequences, based on the shortest branch lengths in the original gene tree. Following thorough validation and careful review of the phylogenies, we selected three representative cases for detailed presentation in Figure 4. These selected examples showcase the variety of functions and evolutionary histories of LGTs within our study. Statistics and data visualization We used the Python toolkit SciPy 64 for all statistical analyses performed in this study. If applicable, such as in the case of the Mann-Whitney U test, we used the ‘two-tailed’ version of that test. Across all statistical tests, we performed a P-value correction for multiple testing errors, namely the Benjamini-Hochberg procedure, which decreases the false discovery rate (FDR). For this purpose, we used the statsmodels python package 65 . The thereafter applied significance threshold was P<0.05. All figures were generated with R's ggplot2 package and with the ggtree package 66 , except Figure 4, which was made using iTOL 67 . For visualization of the phylogenies in Figure 4, and to increase their interpretability, we rerooted them in a way that emphasized the recipient lineage and its sister groups. Declarations Code availability Scripts for detecting and predicting the origins of Rhizaria genes are available at https://github.com/jolienvanhooff/lgtcallrhizaria Acknowledgements This work was supported by the European Research Council under the European Union’s Horizon 2020 research and innovation programme (ERC Starting Grant to L.E. for Macro-EpiK, 803151). Bioinformatic analyses were run on a local cluster with the help of Philippe Deschamps, on the IFB Core Cluster, and the ABiMS Cluster. We thank the following people for sharing sequence data of various SAR species: Christian Woehle and Alexandra-Sophie Roy (Kiel University), Rebecca Gast (Woods Hole Oceanographic Institution), Nick Irwin and Varsha Mathur (University of Oxford), Fabien Burki (Uppsala University), Denis Tikhonenkov (Papanin Institute for Biology of Inland Waters), Jürgen Strassert (Leibniz Institute of Freshwater Ecology and Inland Fisheries), Tonje Marita Bjerkan Heggeset (SINTEF Materials and Chemistry), Kristina Terpis and Chris Lane (The University of Rhode Island), and Matthew Brown (Mississippi State University). We thank the PhyloFisher team (Mississippi State University) for providing support in inferring the eukaryotic species phylogeny, Edward Susko (Dalhousie University) for advice on implementing multiple testing correction, and Michael Seidl (Utrecht University) for advice on genomic density estimates. We thank Michelle Leger, David Moreira, Andrew Roger, Purificación López-García, and Courtney Stairs for fruitful discussions. Author information Authors and affiliations Ecologie Systématique Evolution, CNRS, Université Paris-Saclay, AgroParisTech, Gif-sur-Yvette, France Jolien J.E. van Hooff, Laura Eme Laboratory of Microbiology, Wageningen University and Research, Wageningen, The Netherlands Jolien J.E. van Hooff (present address) Contributions Conceptualization: J.J.E.v.H. and L.E.; investigation and software development: J.J.E.v.H.; result analysis and interpretation: J.J.E.v.H. and L.E. supervision: L.E.; writing: J.J.E.v.H. and L.E. Funding acquisition: L.E. Corresponding authors Correspondence to Laura Eme and Jolien van Hooff. Competing interests The authors declare no competing interests. Supplementary Material Supplementary Figures, Supplementary Tables and Text can be found as Supplementary Material. Supplementary Datasets 1-4 can be found at Figshare (https://figshare.com/projects/Lateral_gene_transfers_LGTs_in_Rhizaria/158240). Figure and table captions, dataset descriptions and Supplementary Text can be found in the file ‘Supplementary Information’. References Ku, C. et al. Endosymbiotic origin and differential loss of eukaryotic genes. Nature 524 , 427–432 (2015). Martin, W. F. Too Much Eukaryote LGT. Bioessays 39 , 1700115 (2017). Leger, M. M., Eme, L., Stairs, C. W. & Roger, A. J. Demystifying Eukaryote Lateral Gene Transfer (Response to Martin 2017 DOI: 10.1002/bies.201700115). Bioessays (2018) doi:10.1002/bies.201700242. Roger, A. J. Reply to ’ Eukaryote lateral gene transfer is Lamarckian ’. Nature Ecology & Evolution 2018 (2018). Van Etten, J. & Bhattacharya, D. Horizontal Gene Transfer in Eukaryotes: Not if, but How Much? Trends Genet. 36 , 915–925 (2020). Sibbald, S. J., Eme, L., Archibald, J. M. & Roger, A. J. Lateral Gene Transfer Mechanisms and Pan-genomes in Eukaryotes. Trends Parasitol. (2020) doi:10.1016/j.pt.2020.07.014. Tria, F. D. K. et al. Gene Duplications Trace Mitochondria to the Onset of Eukaryote Complexity. Genome Biol. Evol. 13 , (2021). Husnik, F. et al. Horizontal gene transfer from diverse bacteria to an insect genome enables a tripartite nested mealybug symbiosis. Cell 153 , 1567–1578 (2013). Shi-Kunne, X., van Kooten, M., Depotter, J. R. L., Thomma, B. P. H. J. & Seidl, M. F. The Genome of the Fungal Pathogen Verticillium dahliae Reveals Extensive Bacterial to Fungal Gene Transfer. Genome Biol. Evol. 11 , 855–868 (2019). Eme, L., Gentekaki, E., Curtis, B., Archibald, J. M. & Roger, A. J. Lateral Gene Transfer in the Adaptation of the Anaerobic Parasite Blastocystis to the Gut. Curr. Biol. 27 , 807–820 (2017). Xu, F. et al. On the reversibility of parasitism: adaptation to a free-living lifestyle via gene acquisitions in the diplomonad Trepomonas sp. PC1. BMC Biol. 14 , 62 (2016). Husnik, F. & McCutcheon, J. P. Functional horizontal gene transfer from bacteria to eukaryotes. Nat. Rev. Microbiol. 16 , 67–79 (2018). Burki, F. et al. Phylogenomics Reshuffles the Eukaryotic Supergroups. PLoS One 2 , e790 (2007). Hackett, J. D. et al. Phylogenomic analysis supports the monophyly of cryptophytes and haptophytes and the association of rhizaria with chromalveolates. Mol. Biol. Evol. 24 , 1702–1713 (2007). Burki, F., Roger, A. J., Brown, M. W. & Simpson, A. G. B. The New Tree of Eukaryotes. Trends Ecol. Evol. 35 , 43–55 (2020). Strassert, J. F. H., Irisarri, I., Williams, T. A. & Burki, F. A molecular timescale for eukaryote evolution with implications for the origin of red algal-derived plastids. Nat. Commun. 12 , 1–13 (2021). Goetz, E. J. et al. Foraminifera as a model of the extensive variability in genome dynamics among eukaryotes. Bioessays 44 , e2100267 (2022). Nowack, E. C. M. et al. Gene transfers from diverse bacteria compensate for reductive genome evolution in the chromatophore of Paulinella chromatophora. Proc. Natl. Acad. Sci. U. S. A. 113 , 12214–12219 (2016). Gilbert, C. & Cordaux, R. Viruses as vectors of horizontal transfer of genetic material in eukaryotes. Curr. Opin. Virol. 25 , 16–22 (2017). Irwin, N. A. T., Pittis, A. A., Richards, T. A. & Keeling, P. J. Systematic evaluation of horizontal gene transfer between eukaryotes and viruses. Nat Microbiol (2021) doi:10.1038/s41564-021-01026-3. Fan, X. et al. Phytoplankton pangenome reveals extensive prokaryotic horizontal gene transfer of diverse functions. Science Advances 6 , eaba0111 (2020). Ciach, M. A., Pawłowska, J. & Muszewska, A. Horizontal gene transfer in 44 early diverging fungi favors short, metabolic, extracellular proteins from associated bacteria. bioRxiv 2021.12.02.471044 (2021) doi:10.1101/2021.12.02.471044. Vancaester, E., Depuydt, T., Osuna-Cruz, C. M. & Vandepoele, K. Comprehensive and Functional Analysis of Horizontal Gene Transfer Events in Diatoms. Mol. Biol. Evol. 37 , 3243–3257 (2020). Nevers, Y., Glover, N. M., Dessimoz, C. & Lecompte, O. Protein length distribution is remarkably uniform across the tree of life. Genome Biol. 24 , 135 (2023). Almagro Armenteros, J. J., Sønderby, C. K., Sønderby, S. K., Nielsen, H. & Winther, O. DeepLoc: prediction of protein subcellular localization using deep learning. Bioinformatics 33 , 3387–3395 (2017). Cantalapiedra, C. P., Hernández-Plaza, A., Letunic, I., Bork, P. & Huerta-Cepas, J. eggNOG-mapper v2: Functional Annotation, Orthology Assignments, and Domain Prediction at the Metagenomic Scale. Mol. Biol. Evol. (2021) doi:10.1093/molbev/msab293. Mertens, J., Aliyu, H. & Cowan, D. A. LEA Proteins and the Evolution of the WHy Domain. Appl. Environ. Microbiol. 84 , (2018). Bouzat, J. L. & Hoostal, M. J. Evolutionary analysis and lateral gene transfer of two-component regulatory systems associated with heavy-metal tolerance in bacteria. J. Mol. Evol. 76 , 267–279 (2013). Corel, E. et al. Bipartite Network Analysis of Gene Sharings in the Microbial World. Mol. Biol. Evol. 35 , 899–913 (2018). Sibbald, S. J. & Archibald, J. M. More protist genomes needed. Nature Ecology &Amp; Evolution 1 , 0145 (2017). Burki, F. & Keeling, P. J. Rhizaria. Curr. Biol. 24 , R103–7 (2014). Gilbert, C. & Maumus, F. Multiple horizontal acquisitions of plant genes in the whitefly Bemisia tabaci. bioRxiv 2022.01.12.476015 (2022) doi:10.1101/2022.01.12.476015. Seczynska, M., Bloor, S., Cuesta, S. M. & Lehner, P. J. Genome surveillance by HUSH-mediated silencing of intronless mobile elements. Nature (2021) doi:10.1038/s41586-021-04228-1. Hibdige, S. G. S., Raimondeau, P., Christin, P.-A. & Dunning, L. T. Widespread lateral gene transfer among grasses. New Phytol. (2021) doi:10.1111/nph.17328. Dunning, L. T. et al. Lateral transfers of large DNA fragments spread functional genes among grasses. Proc. Natl. Acad. Sci. U. S. A. 116 , 4416–4425 (2019). Tice, A. K. et al. PhyloFisher: A phylogenomic package for resolving eukaryotic relationships. PLoS Biol. 19 , e3001365 (2021). Minh, B. Q. et al. IQ-TREE 2: New Models and Efficient Methods for Phylogenetic Inference in the Genomic Era. Mol. Biol. Evol. 37 , 1530–1534 (2020). Burki, F. et al. Phylogenomics of the Intracellular Parasite Mikrocytos mackini Reveals Evidence for a Mitosome in Rhizaria. Curr. Biol. 23 , 1541–1547 (2013). Emms, D. M. & Kelly, S. OrthoFinder: phylogenetic orthology inference for comparative genomics. Genome Biol. 20 , 1–14 (2019). Fu, L., Niu, B., Zhu, Z., Wu, S. & Li, W. CD-HIT: accelerated for clustering the next-generation sequencing data. Bioinformatics 28 , 3150–3152 (2012). Buchfink, B., Reuter, K. & Drost, H.-G. Sensitive protein alignments at tree-of-life scale using DIAMOND. Nat. Methods 18 , 366–368 (2021). Katoh, K. & Standley, D. M. MAFFT multiple sequence alignment software version 7: improvements in performance and usability. Mol. Biol. Evol. 30 , 772–780 (2013). Criscuolo, A. & Gribaldo, S. BMGE (Block Mapping and Gathering with Entropy): a new software for selection of phylogenetic informative regions from multiple sequence alignments. BMC Evol. Biol. 10 , 210 (2010). Huerta-Cepas, J., Serra, F. & Bork, P. ETE 3: Reconstruction, Analysis, and Visualization of Phylogenomic Data. Mol. Biol. Evol. 33 , 1635–1638 (2016). Schoch, C. L. et al. NCBI Taxonomy: a comprehensive update on curation, resources and tools. Database 2020 , (2020). Boschetti, C. et al. Biochemical diversification through foreign gene expression in bdelloid rotifers. PLoS Genet. 8 , e1003035 (2012). Pedregosa, F. et al. Scikit-learn: Machine learning in Python. the Journal of machine Learning research 12 , 2825–2830 (2011). Enright, A. J., Van Dongen, S. & Ouzounis, C. A. An efficient algorithm for large-scale detection of protein families. Nucleic Acids Res. 30 , 1575–1584 (2002). Morel, B., Kozlov, A. M., Stamatakis, A. & Szöllősi, G. J. GeneRax: A Tool for Species-Tree-Aware Maximum Likelihood-Based Gene Family Tree Inference under Gene Duplication, Transfer, and Loss. Mol. Biol. Evol. 37 , 2763–2774 (2020). Stolzer, M. et al. Inferring duplications, losses, transfers and incomplete lineage sorting with nonbinary species trees. Bioinformatics 28 , i409–i415 (2012). Price, M. N., Dehal, P. S. & Arkin, A. P. FastTree 2--approximately maximum-likelihood trees for large alignments. PLoS One 5 , e9490 (2010). Mistry, J. et al. Pfam: The protein families database in 2021. Nucleic Acids Res. 49 , D412–D419 (2021). Necci, M., Piovesan, D., Clementel, D., Dosztányi, Z. & Tosatto, S. C. E. MobiDB-lite 3.0: fast consensus annotation of intrinsic disorder flavours in proteins. Bioinformatics (2020) doi:10.1093/bioinformatics/btaa1045. Jones, P. et al. InterProScan 5: genome-scale protein function classification. Bioinformatics 30 , 1236–1240 (2014). Almagro Armenteros, J. J. et al. SignalP 5.0 improves signal peptide predictions using deep neural networks. Nat. Biotechnol. 37 , 420–423 (2019). Almagro Armenteros, J. J. et al. Detecting sequence signals in targeting peptides using deep learning. Life Science Alliance 2 , e201900429 (2019). Käll, L., Krogh, A. & Sonnhammer, E. L. L. A combined transmembrane topology and signal peptide prediction method. J. Mol. Biol. 338 , 1027–1036 (2004). Krogh, A., Larsson, B., von Heijne, G. & Sonnhammer, E. L. L. Predicting transmembrane protein topology with a hidden markov model: application to complete genomes11Edited by F. Cohen. J. Mol. Biol. 305 , 567–580 (2001). Aylward, F. O. & Moniruzzaman, M. ViralRecall-A Flexible Command-Line Tool for the Detection of Giant Virus Signatures in ’Omic Data. Viruses 13 , (2021). Zhang, H. et al. dbCAN2: a meta server for automated carbohydrate-active enzyme annotation. Nucleic Acids Res. 46 , W95–W101 (2018). Flynn, J. M. et al. RepeatModeler2 for automated genomic discovery of transposable element families. Proc. Natl. Acad. Sci. U. S. A. 117 , 9451–9457 (2020). RepeatMasker Home Page. http://www.repeatmasker.org/. TETools: Dfam Transposable Element Tools Docker Container . (Github). Virtanen, P. et al. SciPy 1.0: fundamental algorithms for scientific computing in Python. Nat. Methods 17 , 261–272 (2020). Seabold, S. & Perktold, J. Statsmodels: Econometric and statistical modeling with python. in Proceedings of the 9th Python in Science Conference (SciPy, 2010). doi:10.25080/majora-92bf1922-011. Yu, G., Smith, D. K., Zhu, H., Guan, Y. & Lam, T. T.-Y. Ggtree : An r package for visualization and annotation of phylogenetic trees with their covariates and other associated data. Methods Ecol. Evol. 8 , 28–36 (2017). Letunic, I. & Bork, P. Interactive Tree Of Life (iTOL) v5: an online tool for phylogenetic tree display and annotation. Nucleic Acids Res. 49 , W293–W296 (2021). Additional Declarations There is NO Competing Interest. Supplementary Files SupplementaryFigure1.pdf Supplementary Figure 1 SupplementaryFigure2.pdf Supplementary Figure 2 SupplementaryFigure3.pdf Supplementary Figure 3 SupplementaryFigure4.pdf Supplementary Figure 4 SupplementaryFigure5.pdf Supplementary Figure 5 SupplementaryFigure6.pdf Supplementary Figure 6 SupplementaryFigure7.pdf Supplementary Figure 7 SupplementaryFigure8.pdf Supplementary Figure 8 SupplementaryFigure9.pdf Supplementary Figure 9 SupplementaryFigure10.pdf Supplementary Figure 10 SupplementaryFigure11.pdf Supplementary Figure 11 SupplementaryFigure12.pdf Supplementary Figure 12 VanHooffRhizariaLGTManuscriptv6SupplementaryMaterial.docx Supplementary Material Cite Share Download PDF Status: Under Review Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4176859","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":284935872,"identity":"d5675d60-ab36-41c3-9c48-94a2176ccc27","order_by":0,"name":"Laura Eme","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABY0lEQVRIie2QMWvCQBiG7/jgXFJvPVFIf8JXAraCtT+kSyTglOBQKEKFBoS4SF2V/olAobg1cmCX/IHSRZdOGdKlOFTshVZJ1O6F5oHj3nyXh3s5QnJy/iDFbSq4hJhJgJ8B+97O1KLzlMK2SQs2CmSORDLCg4owNwlImn2lEFbjJalf8lJUnS+8elvv8+mcThrtIg8gXnaFjgGwOKVotjEekJYzvrdPsem1ar4EQBpaNU+YrDSYiRM/ABili9kG0Yh0/Fe7KsxQIgKw8soDZIJ/kCNXUD/gMl2MRwb9JGvn6SVMlDXqPSgsqXeLTBV7X7niIrklrQjbAI0Eji80pXQCJBIYoZ5Eph6wrG5p7ilvV1BByxmFrWulWIgSDEHDZ1XMZOXKTFhjmVE4tx5o1Gk4w758LC2xgfpwuojp5EYFVSzq1s/vnnvZd0/A3QF1Ux/Hcvf4EBlFd3/5KycnJ+ef8AV5v2wzfw+/9QAAAABJRU5ErkJggg==","orcid":"https://orcid.org/0000-0002-0510-8868","institution":"CNRS, Université Paris-Saclay","correspondingAuthor":true,"prefix":"","firstName":"Laura","middleName":"","lastName":"Eme","suffix":""},{"id":284935873,"identity":"454f4613-42c6-4df4-b371-5d41cfb920cd","order_by":1,"name":"Jolien van Hooff","email":"","orcid":"https://orcid.org/0000-0001-8754-1894","institution":"CNRS, Université Paris-Saclay","correspondingAuthor":false,"prefix":"","firstName":"Jolien","middleName":"van","lastName":"Hooff","suffix":""}],"badges":[],"createdAt":"2024-03-27 14:38:09","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4176859/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4176859/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":53701533,"identity":"0740cc29-8ab9-49c1-a97f-c425c2288059","added_by":"auto","created_at":"2024-03-29 05:30:10","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":648021,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eLGT, duplication and invention events across the rhizarian tree of life.\u003c/strong\u003e Absolute LGT event counts are displayed on each branch, while bar charts atop the branches show the relative frequencies of LGT, duplication, and invention events. Adjacent to the tree, stacked bar charts illustrate the contributions of vertical inheritance, LGT, and gene invention to the gene inventories of each existing species. A blue gradient dot to the right signifies the proportion of proteins analyzed through our phylogenetic process relative to the total protein count. An inset details the Rhizarian phylogeny with accurate branch lengths, arranging species in the same sequence as the schematic tree, and displays the four major clades within Rhizaria. The scale bar denotes the expected number of substitutions per site.\u003c/p\u003e","description":"","filename":"Figure1.png","url":"https://assets-eu.researchsquare.com/files/rs-4176859/v1/bc6db3b3b1ac892c651ac965.png"},{"id":53701531,"identity":"e762ce65-fe80-4650-80ec-e17259187fdb","added_by":"auto","created_at":"2024-03-29 05:30:10","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":197135,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eEvolutionary fates and genomic contexts of genes acquired by LGT.\u003c/strong\u003e A. Fraction of vertically inherited and LGT-derived genes displaying at least one intron. Intron presence was only assessed for three species with appropriate genomic data: \u003cem\u003eB. natans\u003c/em\u003e, \u003cem\u003eP. brassicae\u003c/em\u003eand \u003cem\u003eR. filosa\u003c/em\u003e. P-values were obtained with the chi-square test of independence. B. Distributions of the flanking intergenic regions (FIRs) lengths for LGTs and vertically inherited genes in \u003cem\u003eB. natans\u003c/em\u003e and \u003cem\u003eP. brassicae\u003c/em\u003e. P-values were obtained with the Mann-Whitney U-test. C,D. Gene duplications (C) and losses (D) (mean of normalized values, see Methods) for vertically inherited genes and genes acquired by lateral transfer across lineages in the rhizarian tree of life. Note that no LGT was inferred at the base of ‘Ascetosporea’ (Figure 1). Losses could only be inferred for non-terminal branches. P-values were obtained with the Mann-Whitney U-test.\u003c/p\u003e","description":"","filename":"Figure2.png","url":"https://assets-eu.researchsquare.com/files/rs-4176859/v1/85cbebb985d29f2f451cfb0f.png"},{"id":53701529,"identity":"237cb749-c46a-4379-b39f-1ed1fd78e659","added_by":"auto","created_at":"2024-03-29 05:30:09","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":240550,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eFeatures of proteins acquired by LGT\u003c/strong\u003e. A. Distributions of protein lengths of vertically inherited genes, all proteins acquired by LGT, proteins acquired by LGT from prokaryotes only, and from eukaryotes only. P-values were derived from a Mann-Whitney U test on the distributions of the probabilities for LGT-derived versus vertically inherited proteins. B. Localization propensities of LGT proteins compared to vertically inherited proteins based on DeepLoc calculations. If relative localization propensity \u0026gt;1, the LGTs have on average a stronger predicted probability to localize to a specific cellular compartment than vertically inherited genes. LGTs are divided based on prokaryotic or eukaryotic donor sources. C. Fold-enrichment of LGT proteins in COG functional categories compared to vertically inherited genes, split by donor. This analysis considers the predicted annotations at ancestral nodes in single gene trees (Methods). Chi-square tests assess the significance of differences in the proportion of LGT and vertically inherited proteins within each category.\u003c/p\u003e","description":"","filename":"Figure3.png","url":"https://assets-eu.researchsquare.com/files/rs-4176859/v1/20b9e2f59255a4bfbb4eea4c.png"},{"id":53701541,"identity":"5c39b9d2-1ad4-4406-b638-fd6187c476b2","added_by":"auto","created_at":"2024-03-29 05:30:11","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":739660,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eLGT cases illustrating some of the main patterns identified across Rhizaria\u003c/strong\u003e. A-C. Maximum-likelihood phylogenies inferred for selected LGTs with IQ-TREE under the LG+C60+G evolutionary model. The rhizarian sequences descending from LGT events are highlighted in light blue. The description comprises the most informative annotation from the rhizarian sequences themselves or from their homologs in other taxa. The ‘COG category’ summarizes the EggNOG-mapper annotations of all sequences descending from the LGT. Signal peptides were predicted with TargetP and cell membrane/extracellular localization of proteins with DeepLoc. Putative homologs in giant viruses were identified using hmmscan of the rhizarian proteins against a database of HMM profiles from GVOG.\u003c/p\u003e","description":"","filename":"Figure4.png","url":"https://assets-eu.researchsquare.com/files/rs-4176859/v1/4e77eb27e6924ec8ca6ef4eb.png"},{"id":55288962,"identity":"5849ab0d-6473-4bbd-9426-19195cf92ed7","added_by":"auto","created_at":"2024-04-25 08:44:06","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1652435,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4176859/v1/76703cac-f605-4fd7-8737-eda3b7a5348e.pdf"},{"id":53701530,"identity":"09c8a4af-5878-4905-b4d2-88eef68a6721","added_by":"auto","created_at":"2024-03-29 05:30:09","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":246783,"visible":true,"origin":"","legend":"Supplementary Figure 1","description":"","filename":"SupplementaryFigure1.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4176859/v1/c0d1bc448cb54d80523676fa.pdf"},{"id":53701532,"identity":"29445bee-b655-4da1-9ddf-be2f4e92f6c0","added_by":"auto","created_at":"2024-03-29 05:30:10","extension":"pdf","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":264336,"visible":true,"origin":"","legend":"Supplementary Figure 2","description":"","filename":"SupplementaryFigure2.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4176859/v1/0dc49ced4833be19aed7f630.pdf"},{"id":53701892,"identity":"e3690a44-3221-4a45-978c-67511266cceb","added_by":"auto","created_at":"2024-03-29 05:38:10","extension":"pdf","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":170898,"visible":true,"origin":"","legend":"Supplementary Figure 3","description":"","filename":"SupplementaryFigure3.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4176859/v1/364e54173cefe069d4fd314b.pdf"},{"id":53701535,"identity":"70704000-af27-479a-9a8c-dc6f122dc547","added_by":"auto","created_at":"2024-03-29 05:30:10","extension":"pdf","order_by":4,"title":"","display":"","copyAsset":false,"role":"supplement","size":698508,"visible":true,"origin":"","legend":"\u003cp\u003eSupplementary Figure 4\u003c/p\u003e","description":"","filename":"SupplementaryFigure4.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4176859/v1/b43b3017433441caeefff31e.pdf"},{"id":53701543,"identity":"5734e416-ecb6-4c91-9b95-af73fe72b2e2","added_by":"auto","created_at":"2024-03-29 05:30:11","extension":"pdf","order_by":5,"title":"","display":"","copyAsset":false,"role":"supplement","size":49455,"visible":true,"origin":"","legend":"\u003cp\u003eSupplementary Figure 5\u003c/p\u003e","description":"","filename":"SupplementaryFigure5.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4176859/v1/ea8e5cc0498125b308bf066e.pdf"},{"id":53701537,"identity":"1529a6f6-eb9f-443d-8fcc-983a91c675e6","added_by":"auto","created_at":"2024-03-29 05:30:10","extension":"pdf","order_by":6,"title":"","display":"","copyAsset":false,"role":"supplement","size":75384,"visible":true,"origin":"","legend":"\u003cp\u003eSupplementary Figure 6\u003c/p\u003e","description":"","filename":"SupplementaryFigure6.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4176859/v1/a556e447e081829e0b77ed06.pdf"},{"id":53701547,"identity":"2d519c5d-1ce0-44cc-a6b1-c5683a8237ef","added_by":"auto","created_at":"2024-03-29 05:30:11","extension":"pdf","order_by":7,"title":"","display":"","copyAsset":false,"role":"supplement","size":212582,"visible":true,"origin":"","legend":"Supplementary Figure 7","description":"","filename":"SupplementaryFigure7.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4176859/v1/e56a1a5539de805821ceae08.pdf"},{"id":53701893,"identity":"2e5d2754-c003-476b-aa71-61f48a39c906","added_by":"auto","created_at":"2024-03-29 05:38:11","extension":"pdf","order_by":8,"title":"","display":"","copyAsset":false,"role":"supplement","size":594969,"visible":true,"origin":"","legend":"\u003cp\u003eSupplementary Figure 8\u003c/p\u003e","description":"","filename":"SupplementaryFigure8.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4176859/v1/8650b186caefab9bab5d16f3.pdf"},{"id":53701538,"identity":"d7e4474f-1913-40a1-be2f-15d01f0af75d","added_by":"auto","created_at":"2024-03-29 05:30:11","extension":"pdf","order_by":9,"title":"","display":"","copyAsset":false,"role":"supplement","size":15102,"visible":true,"origin":"","legend":"Supplementary Figure 9","description":"","filename":"SupplementaryFigure9.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4176859/v1/12704fb96d7be5b60f137468.pdf"},{"id":53701544,"identity":"ab9fc711-36ee-4c4a-a3b9-f48a94e15adf","added_by":"auto","created_at":"2024-03-29 05:30:11","extension":"pdf","order_by":10,"title":"","display":"","copyAsset":false,"role":"supplement","size":149614,"visible":true,"origin":"","legend":"\u003cp\u003eSupplementary Figure 10\u003c/p\u003e","description":"","filename":"SupplementaryFigure10.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4176859/v1/603944be74b6c3f63a3f7da7.pdf"},{"id":53701542,"identity":"f2633163-21cf-4a76-9389-bb8a3827710c","added_by":"auto","created_at":"2024-03-29 05:30:11","extension":"pdf","order_by":11,"title":"","display":"","copyAsset":false,"role":"supplement","size":36474,"visible":true,"origin":"","legend":"Supplementary Figure 11","description":"","filename":"SupplementaryFigure11.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4176859/v1/74d8c19300fa562be03603c8.pdf"},{"id":53701545,"identity":"fc7cb87c-8f72-4246-8b19-6f680d450602","added_by":"auto","created_at":"2024-03-29 05:30:11","extension":"pdf","order_by":12,"title":"","display":"","copyAsset":false,"role":"supplement","size":25507,"visible":true,"origin":"","legend":"Supplementary Figure 12","description":"","filename":"SupplementaryFigure12.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4176859/v1/af7e62065336ea0e9b2727c3.pdf"},{"id":53701540,"identity":"8e9a0a6d-6645-44c7-93b4-c550d1b383e9","added_by":"auto","created_at":"2024-03-29 05:30:11","extension":"docx","order_by":13,"title":"","display":"","copyAsset":false,"role":"supplement","size":39442,"visible":true,"origin":"","legend":"Supplementary Material","description":"","filename":"VanHooffRhizariaLGTManuscriptv6SupplementaryMaterial.docx","url":"https://assets-eu.researchsquare.com/files/rs-4176859/v1/8628f3eb832743abfa7cef34.docx"}],"financialInterests":"There is \u003cb\u003eNO\u003c/b\u003e Competing Interest.","formattedTitle":"Lateral gene transfer leaves lasting traces in Rhizaria","fulltext":[{"header":"Introduction","content":"\u003cp\u003eLateral gene transfer (LGT) refers to the non-inheritance transfer of genetic material between unrelated species, and is well-documented among bacteria and archaea. Its role and prevalence in eukaryotes, however, have been more controversial and less clear, although increasing evidence suggests that LGT has played a significant role in the evolution of many eukaryotic lineages\u003csup\u003e1\u0026ndash;4\u003c/sup\u003e. Questions have been raised about the molecular and cellular mechanisms underlying these transfers, their relative importance compared to gene duplication, and their long-term effects on the recipient genomes\u003csup\u003e5\u0026ndash;7\u003c/sup\u003e. Myriad studies indicate that LGT has helped eukaryotes acquire new capacities, such as establishing stable relationships with their endosymbionts, adapting to or reverting from parasitism, or thriving in low-oxygen environments\u003csup\u003e8\u0026ndash;12\u003c/sup\u003e. Most recently, LGT-derived genes were estimated to constitute 1% of all eukaryotic gene inventories\u003csup\u003e5\u003c/sup\u003e. Regardless, this estimate hints at a significant role of LGT in eukaryotic evolution and contrasts the claim that LGT does not have a long-lasting, cumulative impact\u003csup\u003e1\u003c/sup\u003e.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eSeveral factors complicate our understanding of LGT in eukaryotes. Primarily, research has often been directed at pinpointing LGTs that are unique to specific organisms, focusing largely on recent events. This significantly limits our ability to gauge the long-term significance of LGT and hinders a consistent comparison of LGT frequencies across different organisms using a standard set of criteria. Consequently, it remains unclear how the frequency of LGT compares to other evolutionary processes for gene acquisition, such as gene duplication or de novo gene invention. Additionally, the study of LGT between eukaryotic organisms has been relatively neglected compared to LGT from prokaryotes to eukaryotes, largely due to the inherent challenges in detection. State-of-the-art methodologies, which rely on discrepancies between gene and species trees, often fall short when analyzing closely related lineages. This is because single gene or protein phylogenies, derived from a limited phylogenetic signal in a small number of analyzed sites, may not be sufficiently resolved. In addition, independent gene loss across the tree can erroneously suggest LGT events, by grouping sequences from \u0026ldquo;distant\u0026rdquo; eukaryotes. As a result, single gene phylogenies frequently lack the necessary information to differentiate LGT between eukaryotes from vertical gene transmission confidently.\u003c/p\u003e\n\u003cp\u003eIn this work, we investigated the contribution of LGT to Rhizaria genomes by identifying and characterizing gene transfers from prokaryotic and eukaryotic donors, and by evaluating the relative impact of LGT, gene duplication and \u003cem\u003ede novo\u003c/em\u003e gene invention to genome dynamics. Furthermore, we characterized how genes evolved after LGT and showed that they duplicate and get lost frequently. We assessed how LGT genes differed from vertically inherited genes, then leveraged these differences to train a machine learning model capable of identifying LGT-derived genes without phylogenies. In summary, our study offers a detailed, both quantitative and qualitative, exploration of LGT within the relatively unexplored Rhizaria, emphasizing the significant contribution of foreign genes to the genetic makeup of this largely microbial group.\u003c/p\u003e"},{"header":"Results","content":"\u003ch2\u003eExtensive gene acquisition in Rhizaria through LGT from prokaryotes and eukaryotes\u003c/h2\u003e\n\u003cp\u003eTo investigate LGTs in Rhizaria, we aimed to develop a sophisticated phylogeny-based approach. We first collected protein sequences from 29 rhizarians (Supplementary Table 1) and collected their homologs from other eukaryotes, prokaryotes and viruses. After discarding sets of homologs displaying potentially spurious or weak signal (see Methods), we inferred phylogenies for 40,951 protein families. In each tree, we identified the monophyletic clades of rhizarian sequences (hereafter called \u0026apos;rhizarian clades\u0026apos;) and determined their evolutionary origin. Briefly, and taking into account statistical support and after removing potential contaminants, we designated each rhizarian clade as 1) having a vertical origin, if it forms a monophyletic group with stramenopiles or alveolates (the closest relatives of rhizarians, all together forming the SAR group\u003csup\u003e13\u0026ndash;15\u003c/sup\u003e, 2), having a lateral origin, if it is nested in a non-SAR clade, or 3), representing a gene invention if the tree only contains rhizarian sequences. We then inferred a species phylogeny of Rhizaria and mapped the LGTs, gene duplications and inventions onto it (Methods). We detected 13,282 LGT events into Rhizaria, at any time point after they diverged from stramenopiles and alveolates. These gave rise, on average, to \u0026nbsp; 30% of the examined proteins of modern Rhizaria (Figure 1, Supplementary Table 2), a significant departure from the previously estimated 1%\u003csup\u003e5\u003c/sup\u003e (but see discussion below). This estimate is relatively consistent among species, whether they are represented by transcriptomic data, or by good-quality genomic data (\u003cem\u003eBigelowiella natans\u003c/em\u003e: 24%, \u003cem\u003ePlasmodiophora brassicae\u003c/em\u003e: 27%). If we ignore all rhizarian \u0026ldquo;clades\u0026rdquo; containing sequences from only a single species\u0026mdash;thereby rigorously avoiding the misidentification of potential sequence contamination as LGT, yet underestimating real species-specific LGT events\u0026mdash;we still identify 1,992 instances of LGT. These events account for approximately 20% of the proteins analyzed in current rhizarian species (Supplementary Figure 1). Notably, 9% of examined proteins result from an LGT into a relatively deep branch of the rhizarian tree, such as in the ancestors of Reticulofilosa, Monadofilosa or Rotaliida, respectively (labeled inner branches, Figure 1). This indicates that LGT has long-term and likely adaptive consequences in Rhizaria. However, and very importantly, this phylogeny-based approach only examined 12% of the proteins found across Rhizaria, due to the numerous quality checks and strict decisions we made to avoid false positives (Methods).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eConsequently, to investigate a higher proportion of the proteomes, we developed an original prediction approach based on a machine learning classifier. Briefly, we used features of LGT and non-LGT proteins, as identified by our phylogenetic approach, to train a Gradient Boosting classifier, an \u0026lsquo;ensemble\u0026rsquo; method based on many decision trees. As input features, we used sequence characteristics such as the calculated HGT index, protein length, protein disorder, predicted subcellular localization and functional category (Methods, Supplementary Table 3). Using this approach, we analyzed another 71% of rhizarian proteins, in addition to those already analyzed using phylogenetics. The remaining proteins were excluded due to either the low probability associated with any type of origin, or the risk of contamination. We inferred that, on average, 8% of the classified proteins originated through LGT (Supplementary Figure 2). This percentage is significantly lower than the estimate obtained solely through the phylogenetic approach. The discrepancy arises mainly because the dataset included a large number of rhizaria-specific genes, which were classified as inventions and often not investigated by phylogenetics in the previous step (see Supplementary Text, \u0026apos;Machine learning classifier predicts at least 8% of proteins derived from LGT\u0026apos;). Nonetheless, each of our approaches for tracking LGT suggests an estimate much greater than 1%\u003csup\u003e5\u003c/sup\u003e, leading us to argue that the actual percentage of genes in Rhizaria acquired through LGT is likely several folds higher than the eukaryote-wide estimate reported in this review. Given that we consider the phylogenetic approach the most reliable, we restricted all subsequent analyses to characterize LGTs to the results thereof. \u0026nbsp;\u003c/p\u003e\n\u003cp\u003eIn order to estimate the LGT rate across Rhizaria, we summed all LGTs inferred between the last common ancestor of Rhizaria and an example modern-day species for which we have trustworthy genomic data, \u003cem\u003eP. brassicae\u003c/em\u003e. We found that 726 LGT events contributed to its\u003cem\u003e\u0026nbsp;\u003c/em\u003eproteome (Figure 1). Using the SAR divergence time estimates (1,673-1,986 mya)\u003csup\u003e16\u003c/sup\u003e, this comes down to 0.37-0.43 LGTs per million years. If we consider the LGTs inferred by machine learning (Supplementary Figure 2), a total of 1118 LGT events corresponds to 0.56-0.67 LGTs per million years. For \u003cem\u003eB. natans\u003c/em\u003e, these rates come down to 0.65-0.77 and 0.89-1.06 LGTs per million years, respectively. These rates are close to the 1 LGT per million years that was previously contested and considered a highly unlikely rate in eukaryotes\u003csup\u003e2\u003c/sup\u003e. It is important to highlight that these rates reflect the outcomes of both LGT and subsequent gene losses, leading to an underestimation of the actual LGT rate, constrained by what is detectable in the present day.\u003c/p\u003e\n\u003cp\u003eWe then aimed to assess the relative visible imprint of LGT onto modern genomes, compared to gene duplications (Figure 1) \u0026ndash; the mechanism generally assumed to be the main source of genomic novelty in eukaryotes. Overall, gene duplications occurred three times more frequently than LGTs. However, this distribution was not uniform across all branches; in only 36 of the 57 lineages did we identify more duplications than LGTs. Notable exceptions were observed in the ancestors of Rhizaria (181 LGTs vs. 7 duplications) and Cercozoa (241 LGTs vs. 117 duplications). Interestingly, the prevalence of LGT over duplications did not correlate with any specific level of evolutionary depth but rather varied by clade. For instance, the branches within Monadofilosea predominantly experienced more LGTs than duplications, whereas the opposite was true for Foraminifera. This trend in foraminiferans might be attributed to their frequent endoreplication processes, which could lead to high rates of gene duplication as a side effect. These processes might also necessitate the separation of germline from somatic genomes\u003csup\u003e17\u003c/sup\u003e, potentially making LGT fixation more challenging, especially if foreign genes are inserted into somatic rather than germline DNA.\u003c/p\u003e\n\u003cp\u003eWe determined the source of each lateral gene transfer (LGT) by identifying the clade within which the Rhizaria are nested. Remarkably, we found that 48% of the LGTs originated from non-SAR eukaryotes, 39% from bacteria, and less than 1% from archaea. The remaining transfers were of unclear origin. A significant majority of the individual branches (46 out of 56) received a higher number of LGTs from eukaryotic sources compared to prokaryotic ones (Supplementary Figure 3, Supplementary Table 4). The few exceptions were all observed at the tips, rather than deeper within the tree, either indicating residual contamination in the sequence data or suggesting that eukaryote-derived genes are more likely to be preserved over the long term compared to those from prokaryotes. Supporting this, when excluding LGTs limited to a single species, the majority of LGTs were eukaryotic in origin: 51%, in contrast to 26% from bacteria, less than 1% from archaea, and 3% from unidentified prokaryotes. A notable tip where numerous LGTs from prokaryotes appear trustworthy is \u003cem\u003ePaulinella chromatophora.\u0026nbsp;\u003c/em\u003eIt is\u003cem\u003e\u0026nbsp;\u003c/em\u003ea unique species that acquired a primary photosynthetic organelle independently of Archaeplastida, in which we confirmed a significant number of bacterial LGTs (402/790, 51%)\u0026mdash;a figure higher than previously reported\u003csup\u003e18\u003c/sup\u003e. Interestingly, this species also received many eukaryotic LGTs (299 out of 790, 38%), a finding that has not been previously reported.\u003c/p\u003e\n\u003ch2\u003eThe fate of laterally transferred genes\u003c/h2\u003e\n\u003cp\u003eWe first sought to determine how frequently genes of prokaryotic origin assimilated to the host genome through spliceosomal intron acquisition\u003csup\u003e12\u003c/sup\u003e. In the species for which we have genomic data, most prokaryote-derived LGT genes contained at least one intron (Figure 2A, \u003cem\u003eB. natans\u003c/em\u003e: 81% of prokaryote-derived LGT genes, \u003cem\u003eP. brassicae\u003c/em\u003e: 63%, \u003cem\u003eReticulomyxa filosa\u003c/em\u003e: 69%; Supplementary Figure 4A). The detection of these introns serves two purposes: firstly, it confirms that these prokaryote-origin genes are not contaminants; and secondly, the canonical lengths of these introns indicate that they might be expressed by the recipient genome, a hypothesis that awaits validation from transcriptomic data. Noticeably, the youngest, lineage-specific LGT genes lack introns more often than older LGTs, suggesting a gradual intron acquisition over time (for more details, see Supplementary Text: \u0026lsquo;Introns in prokaryotic-derived LGTs\u0026rsquo;).\u003c/p\u003e\n\u003cp\u003eTo gauge the impact of these newly incorporated genes on the host organisms, we examined their evolutionary trajectories after transfer. If LGT genes duplicate frequently, they might confer a beneficial function, possibly benefiting from increased gene dosage. Although duplications of LGT genes have been documented (Vancaester et al. 2020; Eme et al. 2017; Xu et al. 2016; Shin, Doucet, and Pauchet 2022; Siddique et al. 2022; Sheikh et al. 2023; Sahu et al. 2023), a systematic comparison with native genes is pending. Here, we found that in most lineages (38/56), LGT genes underwent duplication more often than vertically inherited genes (Figure 2C), with significant differences observed in 16 lineages. This suggests these genes not only fulfill crucial adaptive roles, but also that their duplications also contribute to the large proportion of LGT-derived genes in modern-day rhizarian genomes. The potential for LGT genes to be lost more frequently than vertically inherited genes was also examined, as this would indicate their transient impact on lineage evolution. In the limited lineages with statistically significant differences, seven showed a higher loss rate for LGT genes, compared to five for vertically inherited genes, showing no distinct loss tendencies between the two gene categories (Figure 2D).\u003c/p\u003e\n\u003cp\u003eBeyond duplication and loss, we explored the sequence evolutionary rates of LGTs and their propensity for protein domain gain and loss. Our findings indicate that LGT proteins often evolve significantly faster than vertically inherited genes in certain lineages (Supplementary Text, \u0026apos;LGT sequence divergence\u0026apos;). Moreover, they tend to gain and lose domains more frequently across most lineages (Supplementary Text, \u0026apos;Domain loss and gain\u0026rsquo;). In conclusion, our analysis suggests that LGT genes exhibit a more dynamic evolutionary process than vertically inherited genes.\u003c/p\u003e\n\u003ch2\u003eMechanistics insights into LGTs\u003c/h2\u003e\n\u003cp\u003eGiven that viruses are known to facilitate lateral gene transfer (LGT) in eukaryotes\u003csup\u003e19\u003c/sup\u003e, we searched all gene trees for viral sequences. We found that rhizarian clades originating from LGT were more likely to contain viral sequences than those inherited vertically (18% \u003cem\u003eversus\u003c/em\u003e 13%, respectively, P\u0026lt;0.001). Although this finding does not provide definitive evidence, it confirms the prevalence of viral\u0026ndash;eukaryotic gene transfer, regardless of directionality\u003csup\u003e20\u003c/sup\u003e, and suggests that viruses could have played a role in mediating some LGT events.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eWe also investigated the genomic context of laterally acquired genes. It has been proposed that LGTs preferentially localize to gene-sparse, potentially heterochromatic areas of the genome to avoid disrupting the genetic integrity of the host\u003csup\u003e12\u003c/sup\u003e. We examined the gene density surrounding LGTs in \u003cem\u003eB. natans\u003c/em\u003e and \u003cem\u003eP. brassicae\u003c/em\u003e by analyzing the flanking intergenic regions (FIRs), which represent the distances to neighboring genes. For \u003cem\u003eB. natans\u003c/em\u003e, we observed no notable difference in FIR lengths between LGTs and vertically inherited genes (Figure 2B). In contrast, in \u003cem\u003eP. brassicae\u003c/em\u003e, LGTs were found to have significantly larger FIRs compared to vertically inherited genes (median FIR for LGTs: 628 bp, median FIR for vertically inherited genes: 505 bp, P=0.012), indicating a preference for integration into less gene-dense areas. Furthermore, we investigated whether LGTs are initially integrated into sparse genomic regions and subsequently relocate to denser, potentially more transcriptionally active areas over evolutionary time, as suggested by Husnik and McCutcheon (2018)\u003csup\u003e12\u003c/sup\u003e. Our comparison of FIR lengths for LGTs acquired at various time points within the Rhizaria tree against those of native genes revealed that, in \u003cem\u003eP. brassicae\u003c/em\u003e, the most recent, species-specific LGTs are situated in the sparsest genomic regions (median FIR: 777 bp, Supplementary Figure 4E). This pattern supports the hypothesis that LGTs initially insert into gene-poor regions before possibly transitioning to gene-richer and presumably more active areas.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eFinally, to shed light on putative structural constraints on transfer mechanisms, we comprehensively analyzed LGT genes and the proteins they encode, drawing on features predicted for both modern-day sequences and their ancestors (Methods). Several studies have observed that LGT genes are typically shorter\u003csup\u003e21\u0026ndash;23\u003c/sup\u003e, possibly because shorter genes are more likely to be fully integrated and functional upon transfer, even if only a small fragment of foreign DNA is incorporated. Our analysis indeed shows that LGT proteins are generally shorter than their native counterparts (Figure 3A, median LGT protein length: 226 amino acids \u003cem\u003evs.\u003c/em\u003e median native protein length: 248 amino acids, P\u0026lt;0.001). However, breaking it down by origin, proteins from eukaryotic LGTs are slightly longer than native proteins (median length: 258, P\u0026lt;0.001), hence the observed shortness primarily pertains to prokaryotic LGTs. This difference likely reflects the intrinsic length disparity between eukaryotic and prokaryotic proteins\u003csup\u003e24\u003c/sup\u003e, challenging the assumption that the integration of larger DNA segments poses a major barrier to LGT (also see Supplementary Text: \u0026lsquo;Protein lengths across datasets\u0026rsquo;).\u003c/p\u003e\n\u003ch2\u003eTransferred proteins harbor donor-specific signatures\u003c/h2\u003e\n\u003cp\u003eTo explore the roles LGT genes play within their new hosts, we utilized DeepLoc\u003csup\u003e25\u003c/sup\u003e for localization predictions and EggNOG-mapper\u003csup\u003e26\u003c/sup\u003e for functional annotations. Similar to protein length, the functional signatures of LGTs differ between those donated by prokaryotes and eukaryotes. Overall, the predicted localizations of genes of prokaryotic origin depart more strongly from those of native genes, than those of eukaryotic origin. Prokaryotic LGTs are predicted to localize extracellularly at a rate 2.6 times higher than vertically-inherited genes (P\u0026lt;0.001; Figure 3B, 803 LGTs in total; Supplementary Figure 5E), suggesting a role in host-environment interactions. This is supported by significant enrichments in functional categories such as \u0026lsquo;Cell motility\u0026rsquo; and \u0026lsquo;Defense mechanisms\u0026rsquo; for prokaryotic LGTs (4.0 and 2.4-fold enrichment, respectively; P\u0026lt;0.001; Figure 3C), despite low absolute numbers of LGTs in these categories (19 and 36, respectively, Supplementary Figure 5G).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eConversely, eukaryotic LGTs are most enriched in the nucleus (1.4-fold higher average localization probability compared to vertically inherited genes, P\u0026lt;0.001; Figure 3B, 1515 LGTs in total; Supplementary Figure 5F), a surprising and significant finding without an immediate explanation in terms of adaptive function. Notably, eukaryotic LGTs also show strong overrepresentation in \u0026lsquo;Defense mechanisms\u0026rsquo; (3.1-fold enrichment, P\u0026lt;0.001, 53 LGTs in total; Supplementary Figure 4H), and more surprisingly, in informational processes like \u0026lsquo;Transcription\u0026rsquo; and \u0026lsquo;Cell cycle control\u0026rsquo; (2.2 and 2.1-fold enrichment, P\u0026lt;0.001; Supplementary Figure 5H), the latter two aligning with an enriched predicted nuclear localization. This suggests that LGTs can significantly influence key cellular and informational processes, challenging previous characterizations that emphasized extracellular localization and metabolic functions, which may have stemmed from analyzing predominantly prokaryote-to-eukaryote transfers. We highlight the remodeling of essential processes by foreign-origin genes, such as transcription factors and epigenetic regulators. For instance, several rhizarian lineages have acquired genes encoding Chromo domains (associated with histone modification and transcription regulation) from Opisthokonta. Similarly, Reticulofilosa acquired a translational regulator, Impact, from red algae, predicted to localize to plastids or mitochondria and play a role in stress response translation. These findings underscore the multifaceted contributions of LGTs to the evolution and functional diversification of recipient lineages.\u003c/p\u003e\n\u003ch2\u003eHighlighted cases of lateral gene transfers\u003c/h2\u003e\n\u003cp\u003eHere, we highlight several instances that exemplify the characteristics more commonly found in LGTs, as opposed to genes inherited vertically.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eIn Foraminifera, our analysis reveals that LGTs frequently fall into the COG category of \u0026apos;Nucleotide metabolism and transport,\u0026apos; and often originated from prokaryotes. A particularly interesting case is that of a Glutamine amidotransferase class-I enzyme (Figure 4A), which seems to have been independently transferred to other eukaryotic groups, such as alveolates and Discoba (Kinetoplastida), with the possibility of subsequent transfers among these groups. In Rhizaria, specifically within \u003cem\u003eR. filosa\u003c/em\u003e, this gene has undergone multiple duplications post-transfer, underscoring its biological significance. The availability of a genomic assembly for \u003cem\u003eR. filosa\u0026nbsp;\u003c/em\u003ebolsters our confidence that these are genuine paralogs, rather than assembly artifacts common in transcriptomic data. Although the enzymatic activity of this protein\u0026mdash;specifically, its role in catalyzing the removal of the ammonia group from glutamine\u0026mdash;is well documented, its specific cellular functions within foraminiferans remain elusive. Intriguingly, our investigation also reveals that some sequences closely match those found in giant viral genes (GVOGs), suggesting the possibility that this gene might have been introduced into the foraminiferan lineage through a viral intermediary.\u003c/p\u003e\n\u003cp\u003eOur other two cases involve the protoplast feeder \u003cem\u003eLeptophrys vorax\u003c/em\u003e, which carries many LGTs predicted to be secreted or membrane proteins facing the extracellular environment. These typically come from either plants (Figure 4B) or bacteria (Figure 4C). Most sequences that descended from these LGTs are predicted to be secreted or in the cell membrane (Figure 4B,C; hexagons and triangles). The case highlighted in Figure 4B exemplifies the frequent duplication events that were previously discussed, showcasing their impact on gene families following their transfer. This LGT involves a protein presumptively originating from plants, annotated as \u0026apos;Late embryogenesis abundant protein 2\u0026apos; (LEA-2, PF03168). Although this function appears to be specific to land plants, homologs can also be found in organisms such as the alga \u003cem\u003eKlebsormidium nitens\u003c/em\u003e and various prokaryotes\u003csup\u003e27\u003c/sup\u003e, where it has been hypothesized to have a role in stress tolerance. We can speculate that this transferred gene fulfills a similar role in \u003cem\u003eL. vorax\u003c/em\u003e.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eThe second \u003cem\u003eL. vorax\u003c/em\u003e signal-peptide containing LGT (Figure 4C) had no functional annotation by EggNOG-mapper. Its bacterial homologs are described as Hlyd-like secretion proteins or Cobalt-zinc-cadmium resistance proteins (CzcB), and they seem otherwise specific to bacteria (PF16576). This protein has been described as involved in resistance to metals like cadmium, zinc and cobalt\u003csup\u003e28\u003c/sup\u003e and might help \u003cem\u003eL. vorax\u003c/em\u003e handle varying metal concentrations that can be found in aquatic systems.\u0026nbsp;\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eThis study showcases the substantial contribution of lateral gene transfer to the gene content of eukaryotes, as exemplified by the large Rhizaria phylum. Our phylogenetic screen reveals that on average, LGTs account for approximately 30% of the proteins, which corresponds to individual species percentages ranging from 2.9% (\u003cem\u003eSorites\u0026nbsp;\u003c/em\u003esp.) to 64.5% (\u003cem\u003eMikrocytos mackini\u003c/em\u003e, see also Supplementary Text, \u0026lsquo;High percentage of LGT in \u003cem\u003eMikrocytos mackini\u003c/em\u003e\u0026rsquo;) and a median estimate of 26.1%. These numbers significantly surpass the previously estimated 1%\u003csup\u003e5\u003c/sup\u003e and are comparable to the LGT content found in prokaryotic genomes\u003csup\u003e29\u003c/sup\u003e. Such high percentages underscore the scope and depth of our phylogenetic analysis, combined with the inclusion of eukaryote-derived LGTs, a component often overlooked.\u003c/p\u003e\n\u003cp\u003eFirst and foremost, our findings emphasize the importance of considering LGT as a key evolutionary force shaping eukaryotic genomes. Although gene duplications are more common, LGTs represent a crucial source of genetic novelty, evidenced by the identification of over 13,000 LGT events across the Rhizarian tree of life. Ignoring LGT risks overestimating the number of species-specific genes and misunderstanding the evolutionary history of eukaryotic lineages. Moreover, LGT challenges conventional interpretations of patchy gene distributions in eukaryotes. Instead, it suggests a more dynamic evolutionary landscape than previously thought, marked by gene acquisition by LGT rather than mere vertical inheritance from a Last Eukaryotic Common Ancestor with an inconceivably large gene repertoire, followed by myriad losses.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eOne of the primary challenges in analyzing transcriptomic data from organisms that cannot be grown in pure culture is the difficulty in completely removing potential contamination by foreign sequences, which might be mistaken for lateral gene transfers (LGTs). However, our most rigorous approach to filter these out yielded a conservative estimate, which suggested that 20% of the genes analyzed by phylogeny have originated from LGT. This estimate represents a lower bound due to its exclusion of species-specific LGTs, which are most likely numerous (Supplementary Figure 6) given the considerable phylogenetic distances among many species in our dataset. To fully appreciate the impact of LGT in this group, a deeper and more extensive sampling of the Rhizaria clade is necessary\u003csup\u003e30,31\u003c/sup\u003e. The very same factors that make Rhizaria susceptible to sequence contamination \u0026ndash;complex feeding behaviors and (endo)symbiotic relationships\u0026ndash; may promote the integration of foreign genes, leading to many LGTs in their genomes. This process of symbiosis-mediated LGT could, in turn, enhance the symbiotic relationships themselves, a phenomenon observed in other evolutionary lineages\u003csup\u003e32\u003c/sup\u003e.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eOur findings highlight the numerical significance and unique characteristics of eukaryote-to-eukaryote LGTs, distinguishing them from the more commonly studied prokaryote-to-eukaryote ones. A pivotal question centers on if and how these numerous genes derived from eukaryotes integrate into the recipient\u0026apos;s cellular machinery. This question particularly applies to genes predicted to be involved in nuclear and cell cycle processes, given that these involve numerous components and interactions, with which such an LGT might need to interact. Eukaryote-derived LGT genes, however, may offer distinct advantages for integration. For instance, the presence of introns in these genes could potentially reduce their susceptibility to gene silencing mechanisms, like those mediated by the HUSH complex\u003csup\u003e33\u003c/sup\u003e. Furthermore, gene exchange via LGT may occur more readily between closely related species, as evidenced in grasses\u003csup\u003e34,35\u003c/sup\u003e, suggesting a propensity for such transfers within similar biological contexts. This potential bias for LGTs among more closely related lineages may imply that our current estimates of LGT events are conservative, since we did not consider LGTs from stramenopiles, alveolates, or other rhizarians, i.e. the closest relatives to the acceptor species.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eImportantly, our survey of LGTs within the \u0026ldquo;deep\u0026rdquo; parts of the Rhizaria phylogeny enabled us to map the trajectory of gene transfers over time, shedding light on how these genes evolved and potentially adapted within their new environments. Notably, we found that older LGT genes tend to have introns more frequently than their more recently acquired counterparts, aligning with expectations regarding gene evolution post-transfer, in particular from prokaryotic donors. Interestingly, in the case of \u003cem\u003eP. brassicae\u003c/em\u003e, we also observed that older LGT genes are situated in regions of the genome with higher gene density compared to newer LGTs, suggesting that these genes may migrate from areas of lower to higher gene density over time. This phenomenon of intragenomic migration, coupled with the acquisition of introns, could potentially facilitate increased gene expression. Consequently, these observations raise an intriguing question about the expression levels of older LGT genes compared to newer ones, suggesting that time might play a crucial role in the integration and functional significance of transferred genes within their new genomic contexts.\u003c/p\u003e\n\u003cp\u003eThrough our innovative machine learning approach, we ventured into uncharted territory to identify lateral gene transfer (LGT) candidates where traditional phylogenetic methods fall short. Our initial efforts focused on analyzing limited characteristics inherent to the protein products of genes, relying on predictive rather than experimental data. Incorporating additional information, such as transcriptional activity levels or specific genomic attributes of the genes in question, could significantly improve the accuracy of our model. For example, the surrounding gene density\u0026mdash;an aspect highlighted earlier\u0026mdash;may serve as an additional indicator for identifying LGT events. Ideally, adding such comprehensive data would enable our model not only to determine LGT likelihood, but also to predict a gene\u0026apos;s origin, be it prokaryotic or eukaryotic, thereby offering deeper insights into the evolutionary pathways of gene transfer.\u003c/p\u003e\n\u003cp\u003eIn conclusion, our investigation into lateral gene transfer (LGT) within Rhizaria illuminates the evolutionary dynamics of this enigmatic clade, providing valuable insights and methodologies for the functional characterization of LGT-derived genes. The ability to systematically assess LGT patterns, facilitated by our detection approaches, extends beyond Rhizaria to a broad spectrum of genomically characterized eukaryotic clades. This comprehensive analysis not only promises a more complete understanding of LGT\u0026apos;s quantitative and qualitative impacts across the eukaryotic tree but also highlights the intricate interplay between biological interactions and genetic exchange in these organisms. Our study underscores the nuanced nature of LGT, revealing the challenges in distinguishing LGT from vertical inheritance, particularly in closely related lineages where gene trees may not offer clear differentiation. This complexity necessitates refined analytical techniques to accurately identify genuine LGTs, which would advance our understanding of evolutionary processes in Rhizaria and across the diverse landscape of eukaryotic life.\u003c/p\u003e"},{"header":"Methods","content":"\u003ch2\u003eRhizaria species phylogeny reconstruction\u003c/h2\u003e\n\u003cp\u003eWe curated a dataset comprising predicted protein sequences from 29 Rhizaria lineages, sourced from both genomic (four) and transcriptomic data (25) (Supplementary Table 1). (Note here that the genomic assembly of \u003cem\u003eR. filosa\u003c/em\u003e exhibited considerable fragmentation, with, for instance, a median of one gene per scaffold. Consequently, we predominantly treated it as a \u0026apos;transcriptomic\u0026apos; dataset in most subsequent analyses, except when specifically examining introns.) To infer the species phylogeny of these 29 lineages, we employed PhyloFisher, an advanced phylogenetic pipeline designed for eukaryotic species tree inference\u003csup\u003e36\u003c/sup\u003e. Briefly, this pipeline facilitated the systematic search and manual curation of orthologs across 240 marker genes within our focal species, culminating in a concatenated alignment utilized for phylogenetic inference. We initiated the PhyloFisher process using its provided scripts, including config.py, fisher.py, informant.py, working_dataset_constructor.py, sgt_constructor.py, and forest.py, all with their default settings. Given that the PhyloFisher v.1.0 dataset encompasses diverse Rhizaria species, we leveraged the \u0026apos;phylogenetically-informed\u0026apos; approach within the \u0026apos;fisher.py\u0026apos; step by specifying the rhizarian species in that dataset as the \u0026apos;Blast Seed.\u0026apos; To classify the Rhizaria sequences from our dataset and, if necessary, those from the PhyloFisher dataset, as ortholog, paralog or contamination, we utilized ParaSorter, a graphical user interface tool integrated into PhyloFisher. Furthermore, we employed \u0026apos;apply_to_db.py\u0026apos; to incorporate our taxa into the PhyloFisher dataset and \u0026apos;select_orthologs.py\u0026apos; to exclude two marker genes (H2A and PYGB) due to the complexity of selecting orthologs using the single gene trees. Subsequently, we used \u0026apos;prep_final_dataset.py\u0026apos; and \u0026apos;matrix_constructor.py\u0026apos; to build the concatenated alignment. Details regarding the coverages of the 29 focal Rhizaria lineages in this alignment are provided in Supplementary Table 1 (\u0026apos;Supermatrix coverage\u0026apos;). The eukaryotic species tree encompassing these 29 Rhizaria lineages was inferred using IQ-TREE v.2.0.3\u003csup\u003e37\u003c/sup\u003e\u0026nbsp; under the LG+G4+C60+F model, with ultrafast bootstrapping (1000 replicates). We arbitrarily rooted the tree on the branch uniting Obazoa+Amoebozoa+CRuMs and Metamonada+Discoba. Within Rhizaria, all branches but one were well-supported (\u0026gt;80%). However, \u003cem\u003eM. mackini\u0026nbsp;\u003c/em\u003eclustered within Metamonada, probably due to a long-branch artifact\u003csup\u003e38\u003c/sup\u003e. Consequently, in the Rhizaria phylogeny utilized for mapping LGT events (as illustrated in Figure 1), we pruned the \u003cem\u003eM. mackini\u003c/em\u003e branch and grafted it as a sister to \u003cem\u003eGromia sphaerica\u003c/em\u003e, in accordance with its position reported previously\u003csup\u003e38\u003c/sup\u003e. The original phylogeny is available in Supplementary Figure 7.\u003c/p\u003e\n\u003ch2\u003e\u0026ldquo;Rhizaria-and-sister-clades\u0026rdquo; dataset assembly\u003c/h2\u003e\n\u003cp\u003eIn addition to Rhizaria, we extensively collected sequence data from members of Stramenopila and Alveolata (collectively referred to as Halvaria), given their status as the closest relatives of Rhizaria. As detailed below, any gene shared between Rhizaria and Halvaria was deemed ancestral to the SAR supergroup in our analyses. This comprehensive sampling of Rhizaria\u0026apos;s sister clades was undertaken to mitigate the risk of erroneous claims regarding lateral gene transfers (LGTs).\u0026nbsp;\u003c/p\u003e\n\u003ch2\u003eProtein families construction\u003c/h2\u003e\n\u003cp\u003eWe utilized OrthoFinder (version 2.3.8) to identify orthogroups within a dataset of 418 SAR proteomes\u003csup\u003e39\u003c/sup\u003e. From this comprehensive set of orthogroups, encompassing singleton \u0026apos;orthogroups\u0026apos; as well, we retained only those containing at least one sequence from our 29 Rhizaria representative proteomes (n=341,035). Subsequently, we compiled sequences from non-SAR eukaryotes and viruses available in the NR database (as of May 20th, 2020) and downsampled these sequences to 90% identity using CD-hit (version 4.8.1)\u003csup\u003e40\u003c/sup\u003e, in order to reduce the computational load associated with homology searches and subsequent analyses. For Bacteria and Archaea, we sourced data from GTDB release 89 (Parks et al., 2020), downsampling these sequences with CD-hit at a 70% identity threshold.\u003c/p\u003e\n\u003cp\u003eWe conducted homolog searches using the rhizarian sequences within the orthogroups as queries against all the data listed above (non-SAR eukaryotes, viruses, archaea and bacteria); we used DIAMOND v2.0.9\u003csup\u003e41\u003c/sup\u003e in \u0026apos;ultra-sensitive\u0026apos; mode, applying an e-value cutoff of 0.001. To mitigate biases stemming from uneven taxonomic representation in databases, we performed these searches separately against seven distinct databases: Archaea, Bacteria, Metazoa, Viridiplantae, Fungi, all other eukaryotes, and viruses. In each search, the maximum number of target sequences was restricted to 2000.\u003c/p\u003e\n\u003cp\u003eTo ensure that the rhizarian sequences serving as queries were not outliers within their respective orthogroups, those longer than twice the median length of the orthogroup were ignored. This precautionary step aimed to prevent potential artifacts arising from abnormally long sequences, such as additional protein domains, which could lead to spurious hits compared to other members within the orthogroup.\u003c/p\u003e\n\u003cp\u003eFinally, we incorporated outgroup hits from all seven outgroup datasets into the SAR orthogroups corresponding to the rhizarian queries.\u003c/p\u003e\n\u003ch2\u003eSingle protein tree inference\u003c/h2\u003e\n\u003cp\u003eFor the reconstruction of individual protein trees, we only selected the expanded orthogroups comprising a minimum of 15 and a maximum of 6000 members. Groups with fewer than 15 sequences were excluded due to the challenge of reliably inferring transfer directionality. Conversely, groups exceeding 6000 sequences were deemed computationally intensive with a corresponding phylogeny likely unresolved. Nevertheless, to accommodate these large orthogroups exceeding 6000 members, we \u0026lsquo;recreated\u0026rsquo; these by retrieving fewer homologues from each database (500 instead 2000). Any resulting orthogroups that still contained more than 6000 members at this step were not retained for further analysis.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eSubsequently, sequences within each orthogroup were aligned using MAFFT v7.407\u003csup\u003e42\u003c/sup\u003e in \u0026apos;auto\u0026apos; mode, followed by alignment trimming with BMGE v1.12\u003csup\u003e43\u003c/sup\u003e using specific parameters (-m BLOSUM30 -b 3 -g 0.7 -h 0.5). Alignments with fewer than 50 positions post-trimming were discarded to avoid spurious topologies arising from weak phylogenetic signal. Additionally, individual sequences displaying over 80% gaps in the trimmed alignment were removed. Alignments containing fewer than 15 sequences were also excluded from subsequent phylogenetic inference.\u003c/p\u003e\n\u003cp\u003eFinally, phylogenies for the remaining orthogroups were inferred using IQ-TREE v2.0.3, under the LG+F+R5 model. Branch support values were obtained via the SH-like approximate likelihood ratio test with 1000 replicates (-aLRT 1000) (Guindon et al., 2010). All 40,951 resulting trees are available as Newick files in Supplementary Dataset 1.\u003c/p\u003e\n\u003ch2\u003eLGT detection based on single gene trees\u003c/h2\u003e\n\u003cp\u003e\u003cstrong\u003eSummary of the workflow.\u0026nbsp;\u003c/strong\u003eWe comprehensively screened 40,951 single protein trees to identify potential lateral gene transfers (LGTs) using the ETE3 toolkit \u003csup\u003e44\u003c/sup\u003e. Specifically, we inferred the evolutionary origins of rhizarian sequences within these trees. These origins were categorized as either \u0026apos;vertical\u0026apos; (indicating parent-offspring inheritance), \u0026apos;lateral\u0026apos; (signifying an LGT event at the base of the rhizarian clade or taxon of interest), or \u0026apos;invention\u0026apos; (implying a novel gene, unique to Rhizaria). This categorization was based on the taxonomic affiliation of the clade within which the rhizarian ones were nested. To elaborate, a vertical origin was inferred if the closest sister clade, or second closest sister clade (sister to group formed by the rhizarian clade plus its closest outgroup) consisted of SAR sequences. Conversely, a lateral origin was suggested if these two sister clades were primarily composed of prokaryotes, or eukaryotes from non-SAR groups. Lastly, a clade was classified as a rhizarian invention if the entirety of the tree was predominantly composed of rhizarian sequences.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDetailed workflow.\u003c/strong\u003e The analysis process is outlined in Supplementary Figure 8. Initially, we annotated each leaf in the single gene trees with a taxonomic classification using NCBI Taxonomy\u003csup\u003e45\u003c/sup\u003e for eukaryotes and viruses (as of October 4th, 2021) and GTDB for prokaryotes (release 89). Leaves were categorized into groups: prokaryotes, non-SAR eukaryotes, halvarians (Stramenopila and Alveolata), rhizarians, viruses, or unknown (Supplementary Figure 8B). Trees without rhizarian sequences were excluded from the analysis. Recognizing viruses primarily as gene transfer agents rather than original gene sources, we removed viral sequences, while noting their proportion within each orthogroup. Sequences of \u0026lsquo;unknown\u0026rsquo; sources were also removed.\u003c/p\u003e\n\u003cp\u003eFor trees with 80% or more rhizarian sequences, we classified the tree as a rhizarian invention, attributing non-rhizarian sequences to either contamination or lateral gene transfers (LGTs) from Rhizaria. In other cases, to discern whether rhizarian sequences derived from vertical or lateral origins, we first identified all rhizarian clades in a tree after rooting the tree on a random non-rhizarian leaf with a non-rhizarian sister clade (Supplementary Figure 8C). Acknowledging potential inaccuracies due to single-gene tree limitations, we merged rhizarian clades that appeared artificially paraphyletic. Two criteria guided this merging: clades were considered monophyletic if only one non-rhizarian sequence was breaking this monophyly, presumed to be contamination or LGT (Supplementary Figure 8D); and clades were merged if fewer than two strongly supported branches (with SH-aLRT support values \u0026ge; 0.8) separated them, thereby correcting for the possible poor resolution of the tree. We merged such rhizarian clades by pruning one and grafting it at the base of the other, keeping the non-rhizarian sister clades intact (Supplementary Figure 8E). \u0026nbsp;After merging, the total number of rhizarian clades was tallied (Supplementary Figure 8F). Trees containing more than five rhizarian clades were deemed as likely too ambiguous for origin analysis due to insufficient phylogenetic signal.\u003c/p\u003e\n\u003cp\u003eWe aimed to then distinguish authentic \u0026quot;rhizarian\u0026quot; clades from those representing contaminant sequences. If a clade contained multiple rhizarian taxa, we considered it less likely to be the result of contamination. However, for clades consisting of sequences from a single species, we discarded the entire clade if at least one of its sequences displayed high identity to any stramenopile or alveolate sequence (\u0026gt;90%), or to any other eukaryotic or prokaryotic sequence (\u0026gt;80%), as determined by a DIAMOND BLASTP search with parameters ultra-sensitive -k 1 --max-hsps 1 --query-cover 70 (Supplementary Figures 8G, 9). To ensure a comprehensive search, we used our large SAR database for stramenopile and alveolate searches and NR database for other sequences.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eSubsequently, we determined the evolutionary origin of the remaining rhizarian clades by identifying their sister lineages. We first annotated internal nodes and leaves in the tree using NCBI Taxonomy and GTDB\u0026rsquo;s taxonomy (Supplementary Figure 8H). For a node to be annotated, we require its parent branch to be well-supported (\u0026gt;0.8 SH-like support). We annotated nodes at the lowest taxonomic level possible if it represented the last common ancestor of at least 80% of their descendant leaves. For example, we annotated a node that contained 20 Bilateria and three Euglenozoa leaves, as \u0026lsquo;Bilateria\u0026rsquo;. Allowing for such \u0026lsquo;interspersing\u0026rsquo; sequences takes into account the possibility that the node\u0026rsquo;s composition was affected by LGT, contamination or other artifacts, while likely reflecting the actual ancestral nature of the clade.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eAfter annotation, we examined the identities of sister clades of the rhizarian clade in an unrooted configuration. If both sister clades comprised Stramenopila and/or Alveolata, we rooted the tree on the ancestral rhizarian node (i.e. the last common ancestor of all rhizarian sequences in the clade of interest), in line with the generally accepted position of Rhizaria in the eukaryotic tree of life, as sister-group to both Stramenopila and Alveolata. Alternatively, we selected the sister clade with fewer SAR sequences and rooted on the leaf with the longest branch to the rhizarian clade. If neither sister clade contained SAR sequences, we chose the leaf with the longest branch to the rhizarian clade. By parsing the tree from leaves to root, we then identified the two first parent nodes with support values above 0.8, upward from the rhizarian clade (Supplementary Figure 8I). In some cases, only a single parent node was found.\u003c/p\u003e\n\u003cp\u003eEach parent was annotated broadly as \u0026apos;SAR\u0026apos;, \u0026apos;eukaryotes\u0026apos;, \u0026apos;prokaryotes\u0026apos;, or \u0026apos;mix\u0026apos; (if it contained both prokaryotic and eukaryotic sequences), and more specifically when possible (e.g., \u0026apos;Metazoa\u0026apos;). If either the first or second parent was identified as \u0026apos;SAR\u0026apos;, we inferred a vertical origin for the ancestral rhizarian sequence (Supplementary Figure 8J). Similarly, if the clade descending from the parent contained a diverse range of eukaryotes but lacked Stramenopile and Alveolata representation, we also inferred a vertical origin, suggesting potential gene loss from stramenopiles and alveolates. Conversely, if the clades descending from the parent nodes included (i) prokaryotes, (ii) specific lower-level eukaryotic groups (e.g., \u0026apos;Metazoa\u0026apos;), or (iii) a mix of prokaryotes and eukaryotes, we inferred lateral gene transfer (LGT) as the origin of the rhizarian homologs. Detailed configurations and associated inferred origins for the ancestral rhizarian sequences are provided in Supplementary Table 5. Subsequently, we inferred the taxonomic identity of the putative donor based on the identity of the first or second parent (see Supplementary Table 5 for details).\u003c/p\u003e\n\u003cp\u003eTo prevent erroneous origin inferences of highly divergent sequences (which are often erroneously placed in single gene trees), we excluded rhizarian clades from subsequent analyses if the median tip-to-tip distance between the rhizarian sequences and those from the first sister clade exceeded four expected substitutions per site (Supplementary Figure 8K).\u003c/p\u003e\n\u003cp\u003eFollowing the analysis of all single-gene trees, we validated the lateral origins identified in \u003cem\u003eB. natans\u003c/em\u003e and \u003cem\u003eP. brassicae\u003c/em\u003e using available genome data to ensure that the identified LGTs were not due to contamination (Supplementary Figure 8L). For this, we stipulated that a gene corresponding to an inferred LGT event must be located on a scaffold containing at least one gene inferred to have been inherited vertically. It is important to note that while the proteomes of \u003cem\u003eR. filosa\u003c/em\u003e and \u003cem\u003eGlobobulimina sp.\u003c/em\u003e were derived from genomic data, we could not use them for this validation step. The assembly of \u003cem\u003eR. filosa\u003c/em\u003e primarily consisted of very short scaffolds, typically containing only a single gene, and the annotated genomic sequence for \u003cem\u003eGlobobulimina sp.\u003c/em\u003e was not publicly available.\u003c/p\u003e\n\u003cp\u003eOverall, our analysis of the single-gene trees enabled us to determine the origins of 12% of the proteins within our 29 rhizarian lineages. Supplementary Figure 10 illustrates the percentage of sequences remaining after each step in our workflow.\u003c/p\u003e\n\u003ch2\u003ePredicting LGT using protein sequence features\u003c/h2\u003e\n\u003cp\u003eA substantial fraction (69%) of all initial rhizarian sequences were not assigned an evolutionary origin through our phylogenetic approach because they belonged to either small orthogroups (\u0026lt;15 sequences) or large orthogroups (\u0026gt;6000 sequences, even after downsampling of the outgroup).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eTo get a more complete picture of evolutionary dynamics in Rhizaria, we explored several approaches to predict the origins of these sequences by non-phylogenetic means. All approaches operate at the level of individual sequences. The individual results were then combined to infer the origin of the cluster they belong to. As a mean to benchmark these approaches, we first tested them on orthogroups that we could analyze through phylogenetics, and for which rhizarian sequences formed a single monophyletic clade in the tree; we regarded them as representing the most straightforward evolutionary history, and thus trusted that we correctly inferred their origins. This constituted a test set of 35,876 sequences which we inferred to belong to rhizarian clade of either vertical or lateral origin (21,224 and 14,652 sequences, respectively). Note that here, the assigned origin of a sequence corresponds to that of its ancestral rhizarian node (i.e., to the entire clade it belongs to; see \u0026lsquo;Detecting LGT with single gene trees\u0026rsquo;).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eWe first tested HGT index, which uses the ratio of the DIAMOND search bit score of the best \u0026lsquo;alien\u0026rsquo; hit (non-SAR eukaryotes and prokaryotes) divided by the bit score of the best \u0026lsquo;native\u0026rsquo; hit (stramenopile or alveolate). Considering the large overlap in HGT indices for sequences from vertical and lateral origins, we did not expect an excellent performance from this approach (Supplementary Figure 11A). Indeed, using the default threshold value of 1\u003csup\u003e46\u003c/sup\u003e as indicative of an HGT, this method showed a 70% accuracy rate (Supplementary Figure 11B), with a relatively high 31% rate of false positives (i.e., vertically inherited genes classified as HGT).\u003c/p\u003e\n\u003cp\u003eWe also tested a linear discriminant analysis, a machine learning technique, leveraging the scikit-learn software library\u003csup\u003e47\u003c/sup\u003e to distinguish between two types of genetic origins: LGT and vertical inheritance. This approach aimed to identify a boundary that best separates these two groups. Despite achieving an overall accuracy of only 59%, this method showed a significant improvement in reducing the misclassification of vertically inherited genes as LGT to just 6%, a notable decrease compared to the 31% misclassification rate using the default threshold in the HGT index method (Supplementary Figure 11C). However, the moderate success rate of this and similar HGT index-based approaches was anticipated, given the substantial overlap in HGT index values between sequences of vertical and lateral origins, indicating a limited ability of the HGT index alone to accurately predict genetic origin.\u003c/p\u003e\n\u003cp\u003eRecognizing the limitations of relying solely on the HGT index, we explored alternative strategies to enhance origin prediction by utilizing distinguishing features of LGTs (see \u0026lsquo;Transferred proteins harbor demarcating signatures\u0026rsquo;). Here, our approach involved analyzing individual protein sequences for diagnostic characteristics (Supplementary Table 3), rather than assessing gene families evolutionary histories. We meticulously compiled these features and prepared our data for machine learning analysis, employing one-hot encoding to manage categorical variables.\u003c/p\u003e\n\u003cp\u003eOur evaluation of various classifiers available in scikit-learn focused on those capable of handling our dataset\u0026apos;s mix of categorical and numeric features. Among the techniques we tested, Gradient Boosting emerged as the most effective. This \u0026lsquo;ensemble\u0026rsquo; method, which combines multiple decision trees, outperformed others in accuracy, the area under the curve (AUC), and the F1 score. We refined our Gradient Boosting model through hyperparameter tuning with a five-fold cross-validation on the training set. Based on this, we trained the classifier with 1,000 decision trees (n_estimator) and a maximum tree depth of 20 for each tree, to prevent overfitting. Additionally, we addressed potential biases from imbalanced data (disproportionate numbers of lateral vs. vertical genes) by balancing these two categories in our training sample.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eTraining on 80% of this balanced dataset and testing on the remaining 20%, our optimized Gradient Boosting classifier achieved a 78% accuracy rate. This performance significantly surpassed that of methods relying solely on the HGT index, also reducing the false positive rate to 26%, although this remains high (Supplementary Figure 11D). Further analysis of our classifier\u0026apos;s effectiveness included examining the significance of various features in the model, using the \u0026quot;feature_importances_\u0026quot; attribute to assess their roles based on Gini impurity (Supplementary Table 3).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eFurther building on our observations that eukaryotic lateral gene transfers (LGTs) differ significantly from prokaryotic LGTs, we explored a more complex model capable of distinguishing among three categories: \u0026apos;vertical\u0026apos; (inherited within the species), \u0026apos;LGT eukaryotic\u0026apos; (gene transfers from eukaryotes), and \u0026apos;LGT prokaryotic\u0026apos; (gene transfers from prokaryotes). Using a gradient boosting model adapted for this multi-class classification, we found that the accuracy dropped to 70% from the 78% achieved with the simpler two-class model. This decrease in accuracy primarily stemmed from difficulties in distinguishing between vertically inherited genes and those acquired from eukaryotic LGTs (Supplementary Figure 11E). Specifically, 27% of the genes we identified as vertically inherited were misclassified as eukaryotic LGTs, and 7% as prokaryotic LGTs. This supports our earlier findings that eukaryotic LGTs often closely resemble genes inherited vertically, more so than prokaryotic LGTs do.\u003c/p\u003e\n\u003cp\u003eGiven our goal was to estimate the overall rate of LGT rather than to precisely categorize eukaryotic and prokaryotic LGTs, we decided to continue using the more straightforward two-class model (Supplementary Figure 11D).\u003c/p\u003e\n\u003cp\u003eWe also sought to identify the origins of orthogroups that were not amenable to analysis with our previous phylogenetic methods. For OGs consisting of more than 80% rhizarian sequences (excluding viruses), we inferred an origin through a process known as \u0026apos;invention\u0026apos; in a manner similar to our phylogenetic analysis. For the remaining orthogroups, we applied our gradient boosting classifier. Recognizing that some orthogroups could contain genes from multiple origins within Rhizaria, we first divided these orthogroups into smaller clusters using MCL\u003csup\u003e48\u003c/sup\u003e with an inflation factor of 1.2. This approach resulted in most orthogroups forming a single cluster, with only a small percentage showing more complex patterns indicative of multiple origins.\u003c/p\u003e\n\u003cp\u003eFor each cluster resulting from MCL, we then used our classifier to predict the origin of each gene as \u0026apos;vertical\u0026apos; or \u0026apos;lateral\u0026apos;. We also calculated the average probability of belonging to either of these two categories for all genes within a cluster. A cluster\u0026apos;s origin was determined to be the category with the highest average probability, but only if this probability exceeded a 70% confidence threshold, discarding the most dubious cases.\u003c/p\u003e\n\u003cp\u003eAdditionally, to ensure the reliability of our predictions and avoid mislabeling due to potential contamination, we applied the same criteria used in our phylogenetic analysis. This involved checking for unusually high sequence similarities to non-rhizarian sequences and, in specific cases, confirming the presence of native genes on the same genetic scaffold when possible. The native nature of genes was deducted from the earlier phylogenetics-based detection strategy.\u003c/p\u003e\n\u003ch2\u003eComprehensive reconstruction of gene evolutionary histories\u003c/h2\u003e\n\u003cp\u003eAfter identifying the origins of rhizarian genes, we carried out a comprehensive reconstruction of their evolution across the Rhizaria tree of life. This reconstruction mapped out gene duplications, transfers, and losses, utilizing two sets of data: one derived solely from the phylogenetic LGT detection (Figure 1) and another that combined both phylogeny-based detection and machine learning or distribution-based predictions (Supplementary Figure 2).\u003c/p\u003e\n\u003cp\u003eWe started by compiling protein sequences from rhizarian clades or clusters. When an origin was determined based on species distribution, clusters corresponded directly to orthogroups. However, in some cases, orthogroups were subdivided during the MCL \u0026nbsp;stage. For rhizarian clades containing at least three sequences, we aligned them using the MAFFT \u0026lsquo;linsi\u0026rsquo; tool and refined the alignment with BMGE (settings: -m BLOSUM45 -b 3 -g 0.7 -h 0.5). We then employed GeneRax v2.0.1\u003csup\u003e49\u003c/sup\u003e to infer gene trees containing only rhizarian sequences and reconcile these with the overall Rhizaria species tree, using the LG+G substitution model and the UndatedDL reconciliation model. For nodes or clusters with fewer sequences, we used Notung v2.9.1.5\u003csup\u003e50\u003c/sup\u003e for pairs of sequences, or created a hypothetical phylogenetic branch for single sequences, ensuring each scenario was accommodated in our analysis. As input for Notung, we either used the pruned subtree corresponding to the rhizarian clade, or, for the results from the prediction approach, inferred a tree using the alignment and trimming procedure as described above (\u0026lsquo;Large-scale inference of single gene trees\u0026rsquo;) and subsequently FastTree v2.0\u003csup\u003e51\u003c/sup\u003e (settings: -lg -gamma). We annotated all resulting reconciled single gene trees with the information collected during the detection or prediction, most importantly the type of origin (vertical, lateral or invention) and, in case of a lateral origin, the donor clade.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eTo ensure a comprehensive reconstruction, we updated the reconciled gene trees to include inferred branches representing either missing ancestral nodes or gene losses, reflecting our analysis findings. These additions were crucial for accurately mapping gene evolution. For example, when our analysis (either detection or prediction) indicated that the origin of a rhizarian clade was vertical (i.e. inherited from an older eukaryotic ancestor), but that the root of the reconciled gene tree was mapped onto a more recent node than the last common ancestor of Rhizaria (e.g., \u0026apos;Foraminifera\u0026apos;), we addressed this discrepancy by adding the missing branches at the base of the tree. These additions, creating previously missing internal nodes such as \u0026apos;Retaria\u0026apos; and \u0026apos;Rhizaria\u0026apos;, effectively adjusted the ancestral rhizarian node back to the common ancestor of Rhizaria. These newly added sister branches were identified as representing \u0026apos;lost\u0026apos; branches. We also annotated them accordingly, including the clade we hypothesized they were lost from, like \u0026apos;Radiolaria\u0026apos; and \u0026apos;Cercozoa\u0026apos; in our example. By annotating these newly added branches and internal nodes, we captured a more detailed picture of rhizarian gene history, including duplications, transfers, and losses, and ancestral gene content.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eFor all the branches we added to the gene trees, we assigned a length of zero. After making these adjustments, we proceeded to gather or infer specific information for each node (both internal and external, except for those representing lost genes) within the updated reconciled gene trees. This information included: (1) the origin of the gene, categorized as vertical, lateral, invention, or duplication. This categorization was based on the outcomes of our detection or prediction analyses or whether the node was identified as a result of duplication, or represented a duplication itself; (2) the source clade, in cases where the gene\u0026apos;s origin was lateral; (3) the median distance from the node to the tree leaves; (4) the total count of gene duplications and losses among the node\u0026apos;s descendants, providing insight into the gene\u0026apos;s evolutionary dynamics.\u003c/p\u003e\n\u003cp\u003eWe introduced the concept of a \u0026apos;residence index\u0026apos; of a gene as a measure to capture its retention across the species tree. This metric is derived by first pruning the species tree for the species that retained this gene, which therefore reflects the branches in which this gene resided, and then summing the branch length of that pruned species tree. The \u0026apos;residence\u0026apos; metric is crucial for normalizing the estimates of gene duplication and loss. Indeed, if a gene is heavily retained, we have more chances of observing gene duplication and loss events than if it got lost early on in the tree (or is inferred as such due to missing data). Similarly, an older gene will have had more time to duplicate than a younger one. This metric thus allows us to more accurately compare duplication and loss frequencies between gene families regardless of their distribution across the species tree.Finally, we aggregated the data from all reconciled single gene trees to tally LGTs, gene duplications, and losses across all branches of the species tree. This comprehensive approach also enables us to infer which genes existed in the ancestral lineages leading to the current members of the Rhizaria included in our study.\u003c/p\u003e\n\u003ch2\u003eCharacterizing LGT evolution and function using gene evolutionary histories\u003c/h2\u003e\n\u003cp\u003eWe used reconciled single-gene trees to explore whether genes acquired through lateral gene transfer (LGT) exhibit distinctive evolutionary dynamics, functional, and structural characteristics compared to genes inherited vertically. For the most reliable insights, we exclusively employed data derived from a phylogeny-based detection approach. We used the set of genes deduced to have been present in a particular lineage within the rhizarian species tree based on the inferences described above. We compared genes with a presumed \u0026apos;vertical\u0026apos; origin in Rhizaria against those acquired \u003cem\u003evia\u003c/em\u003e LGT directly in the same lineage. Thus, we excluded genes that, despite being vertically inherited from the parent, were originally acquired in the Rhizaria through LGT in an older ancestor. This exclusion was based on the premise that such genes might blur the distinction between vertical and lateral transfer, as they may have already adapted to the rhizarian host and exhibit characteristics more akin to vertical inheritance.\u003c/p\u003e\n\u003cp\u003eWe investigated whether LGTs show distinct patterns of gene duplication and loss compared to vertically inherited genes. For each branch, this involved tallying the occurrences of gene duplications and losses among their descendants and normalizing these counts by the \u0026lsquo;residence index\u0026rsquo; described above. This aimed to account for missing data and to facilitate comparisons across taxonomic levels.\u003c/p\u003e\n\u003cp\u003eFor calculating normalized losses, we only considered losses detected in the initial reconciled single gene tree, excluding losses attributed to the reassignment of vertically inherited nodes within Rhizaria (see \u0026lsquo;Comprehensive reconstruction of gene evolutionary histories\u0026rsquo;). To determine if the origin of genes (vertical vs lateral) influenced the normalized rates of duplication and loss, we conducted a two-sided Mann-Whitney U test, assuming both categories had variable distributions (i.e., more than one unique normalized count). We compared the mean normalized counts for each gene origin type and analyzed the differences between these means. Similarly, for LGTs, we calculated the median branch length from root-to-tip and normalized these lengths using the root-to-tip median branch length in the species tree.\u003c/p\u003e\n\u003cp\u003eAdditionally, we examined the dynamics of protein domain acquisition and loss in genes that were either inherited vertically or acquired laterally. This analysis was based on Pfam domain annotations of genes at each node of the rhizarian tree, and the node right above it. For this, we used HMMER\u0026apos;s hmmscan against Pfam-A database version 3.1b2\u003csup\u003e52\u003c/sup\u003e to analyze sequences from the sister clade, results which we combined to the domain annotation already done for the rhizarian sequences. By tracking the domain changes (gains and losses) at each node within the rhizarian gene tree, relative to its ancestral state and throughout its descendants, we quantitatively assessed these changes. Like for gene duplications and losses, the counts of domain dynamics were normalized considering the age and retention of the gene in the rhizarian clade (using the \u0026lsquo;residence index\u0026rsquo; described above).\u003c/p\u003e\n\u003cp\u003eFor the structural and functional analysis of lateral gene transfers (LGTs), we predicted various attributes of rhizarian protein sequences (Supplementary Table 6). These included intrinsic protein disorder with MobiDB-lite v3.10.0\u003csup\u003e53\u003c/sup\u003e, coiled-coil regions and many protein family annotations from InterProScan v5.48-83.0\u003csup\u003e54\u003c/sup\u003e, protein subcellular localization predictions with SignalP v5.0b\u003csup\u003e55\u003c/sup\u003e, TargetP v2.0\u003csup\u003e56\u003c/sup\u003e and DeepLoc v1.0\u003csup\u003e25\u003c/sup\u003e, transmembrane domains with Phobius v1.01\u003csup\u003e57\u003c/sup\u003e and TMHMM v2.0c\u003csup\u003e58\u003c/sup\u003e, viral and NCLDV affinities using VOGDB and GVOG datasets from ViralRecall v2.0\u003csup\u003e59\u003c/sup\u003e, carbohydrate enzyme activities with dbCAN2\u003csup\u003e60\u003c/sup\u003e and COG functional categories with eggNOG-mapper v2.1.5\u003csup\u003e26\u003c/sup\u003e. For each ancestral node identified as having acquired genes via LGT, we annotated its descendants. This included collecting data on all measurable features from the leaf nodes. Depending on the feature\u0026apos;s nature (qualitative or quantitative), we applied appropriate methodologies for projection onto internal nodes. For instance, Pfam annotations were treated with Dollo parsimony, requiring presence in at least 10% of leaf nodes for consideration. Numerical attributes, like protein length, internal nodes were annotated by using median values across its descending leaves.\u003c/p\u003e\n\u003cp\u003eWe compared these annotated features between LGTs and vertically inherited genes, looking for statistically significant differences. This comparison was conducted either branch-specifically within the species tree or across all examined branches, employing chi-square tests for boolean data and Mann-Whitney U tests for numeric data (Figure 3). We reported either the presence or absence of certain features (for boolean data) or median values (for numeric data), ensuring comprehensive analysis of structural and functional gene properties. This extensive examination allows us to discern any significant distinctions between the two gene origins across various biological and molecular aspects.\u003c/p\u003e\n\u003ch2\u003eCharacterizing LGT intron content and genomic context using individual sequences\u003c/h2\u003e\n\u003cp\u003eTo analyze intron acquisition, we extracted gene information from the GFF files of the genome assemblies for \u003cem\u003eB. natans\u003c/em\u003e, \u003cem\u003eP. brassicae\u003c/em\u003e, and \u003cem\u003eR. filosa\u003c/em\u003e. This data included the count of introns per gene, the sum of introns lengths, the intronic proportion of each gene, and the average length of introns within a gene. We calculated the median values for these four metrics and determined the percentage of genes containing at least one intron. These calculations were performed separately for both vertically inherited and laterally acquired proteins within each species. Moreover, we categorized laterally acquired proteins based on their acquisition timeline, allowing for a comparative analysis of intron properties between \u0026apos;young\u0026apos; and \u0026apos;old\u0026apos; LGTs. We employed the Mann-Whitney U test or the chi-square test to examine if the intron characteristics differed between laterally acquired and vertically inherited genes. Additionally, we specifically analyzed LGTs believed to have been transferred from eukaryotes and those from prokaryotes, comparing their intron features with those of vertically inherited genes to discern any distinct patterns.\u003c/p\u003e\n\u003cp\u003eIn our examination of the genomic context surrounding LGTs, we were unable to include \u003cem\u003eR. filosa\u003c/em\u003e due to its highly fragmented genome. Consequently, we investigated gene densities, viral gene abundance, giant virus gene abundance, and transposable element (TE) abundance in \u003cem\u003eB. natans\u003c/em\u003e and \u003cem\u003eP. brassicae\u003c/em\u003e. To operationalize these genomic features, we calculated the distances to the nearest gene, virus-annotated gene (identified using VOG HMM profiles), giant virus-annotated gene (identified using GVOG HMM profiles), and TE for each species, examining both strands of DNA for a comprehensive assessment. TE annotations of the genomes were obtained with RepeatModeler v2.0.1 with the Run the LTR structural discovery pipeline\u003csup\u003e61\u003c/sup\u003e and RepeatMasker v4.1.1\u003csup\u003e62\u003c/sup\u003e with the Dfam TE Tools Container v1.2\u003csup\u003e63\u003c/sup\u003e. We determined the median distances of each genomic feature from both vertically inherited and laterally acquired proteins, further categorizing the latter by their time of acquisition to discern patterns between \u0026apos;new\u0026apos; and \u0026apos;old\u0026apos; LGTs. To assess whether the spatial genomic contexts of these proteins differed based on their origin or acquisition timing, we utilized the Mann-Whitney U test to compare the distribution of distances.\u003c/p\u003e\n\u003ch2\u003eSelection and gene tree inference of illustrative LGTs\u003c/h2\u003e\n\u003cp\u003eWe analyzed the enrichment of specific COG categories and subcellular localizations in LGTs compared to vertically inherited genes, and determined common donor groups for these LGTs. Focusing on lineages that exhibited both functional overrepresentation and a high frequency of specific donors, we meticulously examined the gene phylogenies of these LGTs for confirmation. We inferred a phylogeny using IQ-TREE under a more sophisticated evolutionary model (LG+C60+G) with 1000 ultrafast bootstraps. For large orthogroups, we streamlined the dataset by including only the 1000 non-rhizarian sequences closest to the rhizarian sequences, based on the shortest branch lengths in the original gene tree. Following thorough validation and careful review of the phylogenies, we selected three representative cases for detailed presentation in Figure 4. These selected examples showcase the variety of functions and evolutionary histories of LGTs within our study.\u003c/p\u003e\n\u003ch2\u003eStatistics and data visualization\u003c/h2\u003e\n\u003cp\u003eWe used the Python toolkit SciPy\u003csup\u003e64\u003c/sup\u003e for all statistical analyses performed in this study. If applicable, such as in the case of the Mann-Whitney U test, we used the \u0026lsquo;two-tailed\u0026rsquo; version of that test. Across all statistical tests, we performed a P-value correction for multiple testing errors, namely the Benjamini-Hochberg procedure, which decreases the false discovery rate (FDR). For this purpose, we used the statsmodels python package\u003csup\u003e65\u003c/sup\u003e. The thereafter applied significance threshold was P\u0026lt;0.05. All figures were generated with R\u0026apos;s ggplot2 package and with the ggtree package\u003csup\u003e66\u003c/sup\u003e, except Figure 4, which was made using iTOL\u003csup\u003e67\u003c/sup\u003e. For visualization of the phylogenies in Figure 4, and to increase their interpretability, we rerooted them in a way that emphasized the recipient lineage and its sister groups.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eCode availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eScripts for detecting and predicting the origins of Rhizaria genes are available at https://github.com/jolienvanhooff/lgtcallrhizaria\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis work was supported by the European Research Council under the European Union\u0026rsquo;s Horizon 2020 research and innovation programme (ERC Starting Grant to L.E. for Macro-EpiK, 803151). Bioinformatic analyses were run on a local cluster with the help of Philippe Deschamps, on the IFB Core Cluster, and the ABiMS Cluster. We thank the following people for sharing sequence data of various SAR species:\u0026nbsp;Christian Woehle and Alexandra-Sophie Roy (Kiel University), Rebecca Gast (Woods Hole Oceanographic Institution), Nick Irwin and Varsha Mathur (University of Oxford), Fabien Burki (Uppsala University), Denis Tikhonenkov (Papanin Institute for Biology of Inland Waters), J\u0026uuml;rgen Strassert (Leibniz Institute of Freshwater Ecology and Inland Fisheries), Tonje Marita Bjerkan Heggeset (SINTEF Materials and Chemistry), Kristina Terpis and Chris Lane (The University of Rhode Island), and Matthew Brown (Mississippi State University). We thank the PhyloFisher team (Mississippi State University) for providing support in inferring the eukaryotic species phylogeny, Edward Susko (Dalhousie University) for advice on implementing multiple testing correction, and Michael Seidl (Utrecht University) for advice on genomic density estimates. We thank Michelle Leger, David Moreira, Andrew Roger, Purificaci\u0026oacute;n L\u0026oacute;pez-Garc\u0026iacute;a, and Courtney Stairs for fruitful discussions.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor information\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors and affiliations\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEcologie Syst\u0026eacute;matique Evolution, CNRS, Universit\u0026eacute; Paris-Saclay, AgroParisTech, Gif-sur-Yvette, France\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eJolien J.E. van Hooff, Laura Eme\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eLaboratory of Microbiology, Wageningen University and Research, Wageningen, The Netherlands\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eJolien J.E. van Hooff (present address)\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eContributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eConceptualization: J.J.E.v.H. and L.E.; investigation and software development: J.J.E.v.H.; result analysis and interpretation: J.J.E.v.H. and L.E. supervision: L.E.; writing: J.J.E.v.H. and L.E. Funding acquisition: L.E.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCorresponding authors\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eCorrespondence to Laura Eme and Jolien van Hooff.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSupplementary Material\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eSupplementary Figures, Supplementary Tables and Text can be found as Supplementary Material. Supplementary Datasets 1-4 can be found at Figshare (https://figshare.com/projects/Lateral_gene_transfers_LGTs_in_Rhizaria/158240). Figure and table captions, dataset descriptions and Supplementary Text can be found in the file \u0026lsquo;Supplementary Information\u0026rsquo;.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eKu, C. \u003cem\u003eet al.\u003c/em\u003e Endosymbiotic origin and differential loss of eukaryotic genes. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e524\u003c/strong\u003e, 427\u0026ndash;432 (2015).\u003c/li\u003e\n\u003cli\u003eMartin, W. F. Too Much Eukaryote LGT. \u003cem\u003eBioessays\u003c/em\u003e \u003cstrong\u003e39\u003c/strong\u003e, 1700115 (2017).\u003c/li\u003e\n\u003cli\u003eLeger, M. M., Eme, L., Stairs, C. W. \u0026amp; Roger, A. J. Demystifying Eukaryote Lateral Gene Transfer (Response to Martin 2017 DOI: 10.1002/bies.201700115). \u003cem\u003eBioessays\u003c/em\u003e (2018) doi:10.1002/bies.201700242.\u003c/li\u003e\n\u003cli\u003eRoger, A. J. Reply to \u0026rsquo; Eukaryote lateral gene transfer is Lamarckian \u0026rsquo;. \u003cem\u003eNature Ecology \u0026amp; Evolution\u003c/em\u003e 2018 (2018).\u003c/li\u003e\n\u003cli\u003eVan Etten, J. \u0026amp; Bhattacharya, D. Horizontal Gene Transfer in Eukaryotes: Not if, but How Much? \u003cem\u003eTrends Genet.\u003c/em\u003e \u003cstrong\u003e36\u003c/strong\u003e, 915\u0026ndash;925 (2020).\u003c/li\u003e\n\u003cli\u003eSibbald, S. J., Eme, L., Archibald, J. M. \u0026amp; Roger, A. J. Lateral Gene Transfer Mechanisms and Pan-genomes in Eukaryotes. \u003cem\u003eTrends Parasitol.\u003c/em\u003e (2020) doi:10.1016/j.pt.2020.07.014.\u003c/li\u003e\n\u003cli\u003eTria, F. D. K. \u003cem\u003eet al.\u003c/em\u003e Gene Duplications Trace Mitochondria to the Onset of Eukaryote Complexity. \u003cem\u003eGenome Biol. Evol.\u003c/em\u003e \u003cstrong\u003e13\u003c/strong\u003e, (2021).\u003c/li\u003e\n\u003cli\u003eHusnik, F. \u003cem\u003eet al.\u003c/em\u003e Horizontal gene transfer from diverse bacteria to an insect genome enables a tripartite nested mealybug symbiosis. \u003cem\u003eCell\u003c/em\u003e \u003cstrong\u003e153\u003c/strong\u003e, 1567\u0026ndash;1578 (2013).\u003c/li\u003e\n\u003cli\u003eShi-Kunne, X., van Kooten, M., Depotter, J. R. L., Thomma, B. P. H. J. \u0026amp; Seidl, M. F. The Genome of the Fungal Pathogen Verticillium dahliae Reveals Extensive Bacterial to Fungal Gene Transfer. \u003cem\u003eGenome Biol. Evol.\u003c/em\u003e \u003cstrong\u003e11\u003c/strong\u003e, 855\u0026ndash;868 (2019).\u003c/li\u003e\n\u003cli\u003eEme, L., Gentekaki, E., Curtis, B., Archibald, J. M. \u0026amp; Roger, A. J. Lateral Gene Transfer in the Adaptation of the Anaerobic Parasite Blastocystis to the Gut. \u003cem\u003eCurr. Biol.\u003c/em\u003e \u003cstrong\u003e27\u003c/strong\u003e, 807\u0026ndash;820 (2017).\u003c/li\u003e\n\u003cli\u003eXu, F. \u003cem\u003eet al.\u003c/em\u003e On the reversibility of parasitism: adaptation to a free-living lifestyle via gene acquisitions in the diplomonad Trepomonas sp. PC1. \u003cem\u003eBMC Biol.\u003c/em\u003e \u003cstrong\u003e14\u003c/strong\u003e, 62 (2016).\u003c/li\u003e\n\u003cli\u003eHusnik, F. \u0026amp; McCutcheon, J. P. Functional horizontal gene transfer from bacteria to eukaryotes. \u003cem\u003eNat. Rev. Microbiol.\u003c/em\u003e \u003cstrong\u003e16\u003c/strong\u003e, 67\u0026ndash;79 (2018).\u003c/li\u003e\n\u003cli\u003eBurki, F. \u003cem\u003eet al.\u003c/em\u003e Phylogenomics Reshuffles the Eukaryotic Supergroups. \u003cem\u003ePLoS One\u003c/em\u003e \u003cstrong\u003e2\u003c/strong\u003e, e790 (2007).\u003c/li\u003e\n\u003cli\u003eHackett, J. D. \u003cem\u003eet al.\u003c/em\u003e Phylogenomic analysis supports the monophyly of cryptophytes and haptophytes and the association of rhizaria with chromalveolates. \u003cem\u003eMol. Biol. Evol.\u003c/em\u003e \u003cstrong\u003e24\u003c/strong\u003e, 1702\u0026ndash;1713 (2007).\u003c/li\u003e\n\u003cli\u003eBurki, F., Roger, A. J., Brown, M. W. \u0026amp; Simpson, A. G. B. The New Tree of Eukaryotes. \u003cem\u003eTrends Ecol. Evol.\u003c/em\u003e \u003cstrong\u003e35\u003c/strong\u003e, 43\u0026ndash;55 (2020).\u003c/li\u003e\n\u003cli\u003eStrassert, J. F. H., Irisarri, I., Williams, T. A. \u0026amp; Burki, F. A molecular timescale for eukaryote evolution with implications for the origin of red algal-derived plastids. \u003cem\u003eNat. Commun.\u003c/em\u003e \u003cstrong\u003e12\u003c/strong\u003e, 1\u0026ndash;13 (2021).\u003c/li\u003e\n\u003cli\u003eGoetz, E. J. \u003cem\u003eet al.\u003c/em\u003e Foraminifera as a model of the extensive variability in genome dynamics among eukaryotes. \u003cem\u003eBioessays\u003c/em\u003e \u003cstrong\u003e44\u003c/strong\u003e, e2100267 (2022).\u003c/li\u003e\n\u003cli\u003eNowack, E. C. M. \u003cem\u003eet al.\u003c/em\u003e Gene transfers from diverse bacteria compensate for reductive genome evolution in the chromatophore of Paulinella chromatophora. \u003cem\u003eProc. Natl. Acad. Sci. U. S. A.\u003c/em\u003e \u003cstrong\u003e113\u003c/strong\u003e, 12214\u0026ndash;12219 (2016).\u003c/li\u003e\n\u003cli\u003eGilbert, C. \u0026amp; Cordaux, R. Viruses as vectors of horizontal transfer of genetic material in eukaryotes. \u003cem\u003eCurr. Opin. Virol.\u003c/em\u003e \u003cstrong\u003e25\u003c/strong\u003e, 16\u0026ndash;22 (2017).\u003c/li\u003e\n\u003cli\u003eIrwin, N. A. T., Pittis, A. A., Richards, T. A. \u0026amp; Keeling, P. J. Systematic evaluation of horizontal gene transfer between eukaryotes and viruses. \u003cem\u003eNat Microbiol\u003c/em\u003e (2021) doi:10.1038/s41564-021-01026-3.\u003c/li\u003e\n\u003cli\u003eFan, X. \u003cem\u003eet al.\u003c/em\u003e Phytoplankton pangenome reveals extensive prokaryotic horizontal gene transfer of diverse functions. \u003cem\u003eScience Advances\u003c/em\u003e \u003cstrong\u003e6\u003c/strong\u003e, eaba0111 (2020).\u003c/li\u003e\n\u003cli\u003eCiach, M. A., Pawłowska, J. \u0026amp; Muszewska, A. Horizontal gene transfer in 44 early diverging fungi favors short, metabolic, extracellular proteins from associated bacteria. \u003cem\u003ebioRxiv\u003c/em\u003e 2021.12.02.471044 (2021) doi:10.1101/2021.12.02.471044.\u003c/li\u003e\n\u003cli\u003eVancaester, E., Depuydt, T., Osuna-Cruz, C. M. \u0026amp; Vandepoele, K. Comprehensive and Functional Analysis of Horizontal Gene Transfer Events in Diatoms. \u003cem\u003eMol. Biol. Evol.\u003c/em\u003e \u003cstrong\u003e37\u003c/strong\u003e, 3243\u0026ndash;3257 (2020).\u003c/li\u003e\n\u003cli\u003eNevers, Y., Glover, N. M., Dessimoz, C. \u0026amp; Lecompte, O. Protein length distribution is remarkably uniform across the tree of life. \u003cem\u003eGenome Biol.\u003c/em\u003e \u003cstrong\u003e24\u003c/strong\u003e, 135 (2023).\u003c/li\u003e\n\u003cli\u003eAlmagro Armenteros, J. J., S\u0026oslash;nderby, C. K., S\u0026oslash;nderby, S. K., Nielsen, H. \u0026amp; Winther, O. DeepLoc: prediction of protein subcellular localization using deep learning. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cstrong\u003e33\u003c/strong\u003e, 3387\u0026ndash;3395 (2017).\u003c/li\u003e\n\u003cli\u003eCantalapiedra, C. P., Hern\u0026aacute;ndez-Plaza, A., Letunic, I., Bork, P. \u0026amp; Huerta-Cepas, J. eggNOG-mapper v2: Functional Annotation, Orthology Assignments, and Domain Prediction at the Metagenomic Scale. \u003cem\u003eMol. Biol. Evol.\u003c/em\u003e (2021) doi:10.1093/molbev/msab293.\u003c/li\u003e\n\u003cli\u003eMertens, J., Aliyu, H. \u0026amp; Cowan, D. A. LEA Proteins and the Evolution of the WHy Domain. \u003cem\u003eAppl. Environ. Microbiol.\u003c/em\u003e \u003cstrong\u003e84\u003c/strong\u003e, (2018).\u003c/li\u003e\n\u003cli\u003eBouzat, J. L. \u0026amp; Hoostal, M. J. Evolutionary analysis and lateral gene transfer of two-component regulatory systems associated with heavy-metal tolerance in bacteria. \u003cem\u003eJ. Mol. Evol.\u003c/em\u003e \u003cstrong\u003e76\u003c/strong\u003e, 267\u0026ndash;279 (2013).\u003c/li\u003e\n\u003cli\u003eCorel, E. \u003cem\u003eet al.\u003c/em\u003e Bipartite Network Analysis of Gene Sharings in the Microbial World. \u003cem\u003eMol. Biol. Evol.\u003c/em\u003e \u003cstrong\u003e35\u003c/strong\u003e, 899\u0026ndash;913 (2018).\u003c/li\u003e\n\u003cli\u003eSibbald, S. J. \u0026amp; Archibald, J. M. More protist genomes needed. \u003cem\u003eNature Ecology \u0026amp;Amp; Evolution\u003c/em\u003e \u003cstrong\u003e1\u003c/strong\u003e, 0145 (2017).\u003c/li\u003e\n\u003cli\u003eBurki, F. \u0026amp; Keeling, P. J. Rhizaria. \u003cem\u003eCurr. Biol.\u003c/em\u003e \u003cstrong\u003e24\u003c/strong\u003e, R103\u0026ndash;7 (2014).\u003c/li\u003e\n\u003cli\u003eGilbert, C. \u0026amp; Maumus, F. Multiple horizontal acquisitions of plant genes in the whitefly Bemisia tabaci. \u003cem\u003ebioRxiv\u003c/em\u003e 2022.01.12.476015 (2022) doi:10.1101/2022.01.12.476015.\u003c/li\u003e\n\u003cli\u003eSeczynska, M., Bloor, S., Cuesta, S. M. \u0026amp; Lehner, P. J. Genome surveillance by HUSH-mediated silencing of intronless mobile elements. \u003cem\u003eNature\u003c/em\u003e (2021) doi:10.1038/s41586-021-04228-1.\u003c/li\u003e\n\u003cli\u003eHibdige, S. G. S., Raimondeau, P., Christin, P.-A. \u0026amp; Dunning, L. T. Widespread lateral gene transfer among grasses. \u003cem\u003eNew Phytol.\u003c/em\u003e (2021) doi:10.1111/nph.17328.\u003c/li\u003e\n\u003cli\u003eDunning, L. T. \u003cem\u003eet al.\u003c/em\u003e Lateral transfers of large DNA fragments spread functional genes among grasses. \u003cem\u003eProc. Natl. Acad. Sci. U. S. A.\u003c/em\u003e \u003cstrong\u003e116\u003c/strong\u003e, 4416\u0026ndash;4425 (2019).\u003c/li\u003e\n\u003cli\u003eTice, A. K. \u003cem\u003eet al.\u003c/em\u003e PhyloFisher: A phylogenomic package for resolving eukaryotic relationships. \u003cem\u003ePLoS Biol.\u003c/em\u003e \u003cstrong\u003e19\u003c/strong\u003e, e3001365 (2021).\u003c/li\u003e\n\u003cli\u003eMinh, B. Q. \u003cem\u003eet al.\u003c/em\u003e IQ-TREE 2: New Models and Efficient Methods for Phylogenetic Inference in the Genomic Era. \u003cem\u003eMol. Biol. Evol.\u003c/em\u003e \u003cstrong\u003e37\u003c/strong\u003e, 1530\u0026ndash;1534 (2020).\u003c/li\u003e\n\u003cli\u003eBurki, F. \u003cem\u003eet al.\u003c/em\u003e Phylogenomics of the Intracellular Parasite Mikrocytos mackini Reveals Evidence for a Mitosome in Rhizaria. \u003cem\u003eCurr. Biol.\u003c/em\u003e \u003cstrong\u003e23\u003c/strong\u003e, 1541\u0026ndash;1547 (2013).\u003c/li\u003e\n\u003cli\u003eEmms, D. M. \u0026amp; Kelly, S. OrthoFinder: phylogenetic orthology inference for comparative genomics. \u003cem\u003eGenome Biol.\u003c/em\u003e \u003cstrong\u003e20\u003c/strong\u003e, 1\u0026ndash;14 (2019).\u003c/li\u003e\n\u003cli\u003eFu, L., Niu, B., Zhu, Z., Wu, S. \u0026amp; Li, W. CD-HIT: accelerated for clustering the next-generation sequencing data. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cstrong\u003e28\u003c/strong\u003e, 3150\u0026ndash;3152 (2012).\u003c/li\u003e\n\u003cli\u003eBuchfink, B., Reuter, K. \u0026amp; Drost, H.-G. Sensitive protein alignments at tree-of-life scale using DIAMOND. \u003cem\u003eNat. Methods\u003c/em\u003e \u003cstrong\u003e18\u003c/strong\u003e, 366\u0026ndash;368 (2021).\u003c/li\u003e\n\u003cli\u003eKatoh, K. \u0026amp; Standley, D. M. MAFFT multiple sequence alignment software version 7: improvements in performance and usability. \u003cem\u003eMol. Biol. Evol.\u003c/em\u003e \u003cstrong\u003e30\u003c/strong\u003e, 772\u0026ndash;780 (2013).\u003c/li\u003e\n\u003cli\u003eCriscuolo, A. \u0026amp; Gribaldo, S. BMGE (Block Mapping and Gathering with Entropy): a new software for selection of phylogenetic informative regions from multiple sequence alignments. \u003cem\u003eBMC Evol. Biol.\u003c/em\u003e \u003cstrong\u003e10\u003c/strong\u003e, 210 (2010).\u003c/li\u003e\n\u003cli\u003eHuerta-Cepas, J., Serra, F. \u0026amp; Bork, P. ETE 3: Reconstruction, Analysis, and Visualization of Phylogenomic Data. \u003cem\u003eMol. Biol. Evol.\u003c/em\u003e \u003cstrong\u003e33\u003c/strong\u003e, 1635\u0026ndash;1638 (2016).\u003c/li\u003e\n\u003cli\u003eSchoch, C. L. \u003cem\u003eet al.\u003c/em\u003e NCBI Taxonomy: a comprehensive update on curation, resources and tools. \u003cem\u003eDatabase \u003c/em\u003e \u003cstrong\u003e2020\u003c/strong\u003e, (2020).\u003c/li\u003e\n\u003cli\u003eBoschetti, C. \u003cem\u003eet al.\u003c/em\u003e Biochemical diversification through foreign gene expression in bdelloid rotifers. \u003cem\u003ePLoS Genet.\u003c/em\u003e \u003cstrong\u003e8\u003c/strong\u003e, e1003035 (2012).\u003c/li\u003e\n\u003cli\u003ePedregosa, F. \u003cem\u003eet al.\u003c/em\u003e Scikit-learn: Machine learning in Python. \u003cem\u003ethe Journal of machine Learning research\u003c/em\u003e \u003cstrong\u003e12\u003c/strong\u003e, 2825\u0026ndash;2830 (2011).\u003c/li\u003e\n\u003cli\u003eEnright, A. J., Van Dongen, S. \u0026amp; Ouzounis, C. A. An efficient algorithm for large-scale detection of protein families. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cstrong\u003e30\u003c/strong\u003e, 1575\u0026ndash;1584 (2002).\u003c/li\u003e\n\u003cli\u003eMorel, B., Kozlov, A. M., Stamatakis, A. \u0026amp; Sz\u0026ouml;llősi, G. J. GeneRax: A Tool for Species-Tree-Aware Maximum Likelihood-Based Gene Family Tree Inference under Gene Duplication, Transfer, and Loss. \u003cem\u003eMol. Biol. Evol.\u003c/em\u003e \u003cstrong\u003e37\u003c/strong\u003e, 2763\u0026ndash;2774 (2020).\u003c/li\u003e\n\u003cli\u003eStolzer, M. \u003cem\u003eet al.\u003c/em\u003e Inferring duplications, losses, transfers and incomplete lineage sorting with nonbinary species trees. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cstrong\u003e28\u003c/strong\u003e, i409\u0026ndash;i415 (2012).\u003c/li\u003e\n\u003cli\u003ePrice, M. N., Dehal, P. S. \u0026amp; Arkin, A. P. FastTree 2--approximately maximum-likelihood trees for large alignments. \u003cem\u003ePLoS One\u003c/em\u003e \u003cstrong\u003e5\u003c/strong\u003e, e9490 (2010).\u003c/li\u003e\n\u003cli\u003eMistry, J. \u003cem\u003eet al.\u003c/em\u003e Pfam: The protein families database in 2021. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cstrong\u003e49\u003c/strong\u003e, D412\u0026ndash;D419 (2021).\u003c/li\u003e\n\u003cli\u003eNecci, M., Piovesan, D., Clementel, D., Doszt\u0026aacute;nyi, Z. \u0026amp; Tosatto, S. C. E. MobiDB-lite 3.0: fast consensus annotation of intrinsic disorder flavours in proteins. \u003cem\u003eBioinformatics\u003c/em\u003e (2020) doi:10.1093/bioinformatics/btaa1045.\u003c/li\u003e\n\u003cli\u003eJones, P. \u003cem\u003eet al.\u003c/em\u003e InterProScan 5: genome-scale protein function classification. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cstrong\u003e30\u003c/strong\u003e, 1236\u0026ndash;1240 (2014).\u003c/li\u003e\n\u003cli\u003eAlmagro Armenteros, J. J. \u003cem\u003eet al.\u003c/em\u003e SignalP 5.0 improves signal peptide predictions using deep neural networks. \u003cem\u003eNat. Biotechnol.\u003c/em\u003e \u003cstrong\u003e37\u003c/strong\u003e, 420\u0026ndash;423 (2019).\u003c/li\u003e\n\u003cli\u003eAlmagro Armenteros, J. J. \u003cem\u003eet al.\u003c/em\u003e Detecting sequence signals in targeting peptides using deep learning. \u003cem\u003eLife Science Alliance\u003c/em\u003e \u003cstrong\u003e2\u003c/strong\u003e, e201900429 (2019).\u003c/li\u003e\n\u003cli\u003eK\u0026auml;ll, L., Krogh, A. \u0026amp; Sonnhammer, E. L. L. A combined transmembrane topology and signal peptide prediction method. \u003cem\u003eJ. Mol. Biol.\u003c/em\u003e \u003cstrong\u003e338\u003c/strong\u003e, 1027\u0026ndash;1036 (2004).\u003c/li\u003e\n\u003cli\u003eKrogh, A., Larsson, B., von Heijne, G. \u0026amp; Sonnhammer, E. L. L. Predicting transmembrane protein topology with a hidden markov model: application to complete genomes11Edited by F. Cohen. \u003cem\u003eJ. Mol. Biol.\u003c/em\u003e \u003cstrong\u003e305\u003c/strong\u003e, 567\u0026ndash;580 (2001).\u003c/li\u003e\n\u003cli\u003eAylward, F. O. \u0026amp; Moniruzzaman, M. ViralRecall-A Flexible Command-Line Tool for the Detection of Giant Virus Signatures in \u0026rsquo;Omic Data. \u003cem\u003eViruses\u003c/em\u003e \u003cstrong\u003e13\u003c/strong\u003e, (2021).\u003c/li\u003e\n\u003cli\u003eZhang, H. \u003cem\u003eet al.\u003c/em\u003e dbCAN2: a meta server for automated carbohydrate-active enzyme annotation. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cstrong\u003e46\u003c/strong\u003e, W95\u0026ndash;W101 (2018).\u003c/li\u003e\n\u003cli\u003eFlynn, J. M. \u003cem\u003eet al.\u003c/em\u003e RepeatModeler2 for automated genomic discovery of transposable element families. \u003cem\u003eProc. Natl. Acad. Sci. U. S. A.\u003c/em\u003e \u003cstrong\u003e117\u003c/strong\u003e, 9451\u0026ndash;9457 (2020).\u003c/li\u003e\n\u003cli\u003eRepeatMasker Home Page. http://www.repeatmasker.org/.\u003c/li\u003e\n\u003cli\u003e\u003cem\u003eTETools: Dfam Transposable Element Tools Docker Container\u003c/em\u003e. (Github).\u003c/li\u003e\n\u003cli\u003eVirtanen, P. \u003cem\u003eet al.\u003c/em\u003e SciPy 1.0: fundamental algorithms for scientific computing in Python. \u003cem\u003eNat. Methods\u003c/em\u003e \u003cstrong\u003e17\u003c/strong\u003e, 261\u0026ndash;272 (2020).\u003c/li\u003e\n\u003cli\u003eSeabold, S. \u0026amp; Perktold, J. Statsmodels: Econometric and statistical modeling with python. in \u003cem\u003eProceedings of the 9th Python in Science Conference\u003c/em\u003e (SciPy, 2010). doi:10.25080/majora-92bf1922-011.\u003c/li\u003e\n\u003cli\u003eYu, G., Smith, D. K., Zhu, H., Guan, Y. \u0026amp; Lam, T. T.-Y. Ggtree : An r package for visualization and annotation of phylogenetic trees with their covariates and other associated data. \u003cem\u003eMethods Ecol. Evol.\u003c/em\u003e \u003cstrong\u003e8\u003c/strong\u003e, 28\u0026ndash;36 (2017).\u003c/li\u003e\n\u003cli\u003eLetunic, I. \u0026amp; Bork, P. Interactive Tree Of Life (iTOL) v5: an online tool for phylogenetic tree display and annotation. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cstrong\u003e49\u003c/strong\u003e, W293\u0026ndash;W296 (2021).\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"nature-portfolio","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"","title":"Nature Portfolio","twitterHandle":"","acdcEnabled":false,"dfaEnabled":false,"editorialSystem":"ejp","reportingPortfolio":"","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-4176859/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4176859/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"Lateral gene transfer (LGT) is a fundamental process that has contributed to the genetic makeup of various eukaryotic lineages. Yet our comprehension of its prevalence and significance across the eukaryotic domain remains incomplete, particularly when it comes to eukaryote-to-eukaryote transfers. The Rhizaria forms an expansive, ancient, and morphologically diverse clade of mostly free-living, single-celled phagotrophs, whose genetic diversity reflects their adaptability and evolutionary success across a range of habitats. Here, we undertake a comprehensive investigation of LGT within Rhizaria, tracing its role from the clade's emergence to the present day, employing advanced phylogenetic analyses and machine learning techniques. Contrary to previous assertions that LGT in eukaryotes is rare, our findings suggest that at least 8% of genes in contemporary rhizarian genomes were acquired through LGT at various points during their billion-year history. Our analysis reveals that although gene duplications exceed the number of LGT events, the duplicated genes that were originally acquired through LGT have a visible effect on the evolutionary trajectory of the recipient organisms. Our results also indicate that LGTs originating from other eukaryotes are more common than those from prokaryotes and exhibit unique patterns. By offering both quantitative and qualitative insights into the role of LGT in shaping the evolution of a major eukaryotic lineage, this work further demystifies lateral gene transfer in eukaryotes.","manuscriptTitle":"Lateral gene transfer leaves lasting traces in Rhizaria","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-03-29 05:30:05","doi":"10.21203/rs.3.rs-4176859/v1","editorialEvents":[],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"communications-biology","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"commsbio","sideBox":"Learn more about [Communications Biology](http://www.nature.com/commsbio/)","snPcode":"","submissionUrl":"","title":"Communications Biology","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"Communications Series","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"776e7845-846c-4e8a-966e-2d68ddbe2933","owner":[],"postedDate":"March 29th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[{"id":29996636,"name":"Biological sciences/Microbiology/Microbial genetics"},{"id":29996637,"name":"Biological sciences/Evolution/Phylogenetics"},{"id":29996638,"name":"Biological sciences/Computational biology and bioinformatics/High-throughput screening"},{"id":29996639,"name":"Biological sciences/Evolution/Molecular evolution"},{"id":29996640,"name":"Biological sciences/Evolution/Evolutionary genetics"}],"tags":[],"updatedAt":"2026-05-01T21:15:40+00:00","versionOfRecord":[],"versionCreatedAt":"2024-03-29 05:30:05","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-4176859","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4176859","identity":"rs-4176859","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.