Dominant contribution of Asgard archaea to eukaryogenesis | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Biological Sciences - Article Dominant contribution of Asgard archaea to eukaryogenesis Eugene Koonin, Victor Tobiasson, Jacob Luo, Yuri Wolf This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-5352492/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 14 Jan, 2026 Read the published version in Nature → Version 1 posted You are reading this latest preprint version Abstract The origin of eukaryotes is one of the key problems in evolutionary biology 1,2. The demonstration that the Last Eukaryotic Common Ancestor (LECA) already contained the mitochondrion, an endosymbiotic organelle derived from an alphaproteobacterium 3, and the discovery of Asgard archaea, the closest archaeal relatives of eukaryotes 4-8, inform and constrain evolutionary scenarios of eukaryogenesis 9. We undertook a comprehensive analysis of the origins of the core eukaryotic genes tracing to the LECA within a rigorous statistical framework centered around evolutionary hypotheses testing using constrained phylogenetic trees. The results reveal dominant contributions of Asgard archaea to the origin of most of the conserved eukaryotic functional systems and pathways. A limited contribution from Alphaproteobacteria was identified, primarily relating to the energy transformation systems and Fe-S cluster biogenesis, whereas ancestry from other bacterial phyla was scattered across the eukaryotic functional landscape, with almost no consistent trends. These findings suggest a model of eukaryogenesis in which key features of eukaryotic cell organization evolved in the Asgard ancestor, followed by the capture of the Alphaproteobacterial endosymbiont, and augmented by numerous but sporadic horizontal acquisition of genes from other bacteria both before and after endosymbiosis. Biological sciences/Evolution/Phylogenetics Biological sciences/Evolution/Molecular evolution Asgard archaea Origin of eukaryotes Asgard contribution to eukaryogenesis Endosymbiosis Horizontal gene transfer Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Introduction Eukaryotes drastically differ from archaea and bacteria (collectively, prokaryotes) by the highly complex cellular organization. The signature features of this organizational complexity include the endomembrane system, in particular, the nucleus, the eponymous eukaryotic organelle, the elaborate cytoskeleton, and the mitochondrion, the energy-converting organelle that evolved from an alphaproteobacterial endosymbiont 10 . Early models of eukaryogenesis featured a protoeukaryote lacking mitochondria that captured and domesticated an alphaproteobacterium at a relatively late stage of evolution 11,12 . However, subsequent research revealed no primary amitochondrial eukaryotes (although many have lost mitochondria through reductive evolution) 3 , suggesting that the Last Eukaryotic Common Ancestor (LECA) was already endowed with mitochondria along with the other signatures of the eukaryotic cellular organization 1,9 . These findings lend credence to scenarios of eukaryogenesis in which mitochondrial endosymbiosis triggered cellular reorganization, giving rise to the eukaryotic cellular complexity 9,13-15 . Phylogenomic analyses show that those conserved, core genes of eukaryotes that can be traced to the LECA are a mix originating from archaea and from various bacteria. The archaea-derived genes were found primarily in information processing systems (replication, core transcription and translation) whereas genes of apparent bacterial origin comprise the operational component of the eukaryotic gene complement, in particular, encoding metabolic enzymes 16-19 . The study of eukaryogenesis was transformed by the discovery and subsequent exploration of Asgard archaea (currently, phylum Promethearchaeota within kingdom Promethearchaeati 20 ; hereafter, Asgard, for brevity) that includes the closest known archaeal relatives of eukaryotes 4-8 . Phylogenetic analyses of highly conserved genes have been pushing the eukaryotic branch progressively deeper with the Asgard tree. The latest such analysis identified the order Hodarchaeales , within the Asgard class Heimdallarchaeia , as the likely sister group of eukaryotes 8 . Perhaps, even more notably, it has been shown that Asgard archaea encode homologs of a broad variety of eukaryote signature proteins beyond the core information processing componentry, in particular, many cytoskeletal proteins and those involved in membrane remodeling 5,6,8 . A variety of eukaryogenesis models have been proposed, differing with respect to the timing of the origin of the eukaryotic cellular organization and the exact nature and succession of the underlying events 1,2,9 . The endosymbiotic origin of the mitochondria from an alphaproteobacterium is indisputable, but the identity of the host of the proto-mitochondrial endosymbiont remains a matter of debate. The most straightforward models posit an Asgard archaeal host. However, a major conundrum for such scenarios of eukaryogenesis is the chemistry of cell membrane lipids and the enzymology of their biosynthesis, which are unrelated in archaea and bacteria, with eukaryotic being of the bacterial type 21 . Accordingly, any model of eukaryogenesis that includes an archaeal host would require a membrane replacement step. More complex, alternative models of eukaryogenesis, underpinned by metabolic symbiosis, postulate two endosymbiotic events whereby an Asgard archaeon was first engulfed by a bacterium, followed by the loss of the archaeal membrane, and then, a second endosymbiosis gave rise to the mitochondria 15,22,23 . Even if seemingly less parsimonious than single-endosymbiosis scenarios, such models of eukaryogenesis account for the continuity of bacterial membranes and also appear compatible with the syntrophic lifestyle of at least some Asgard archaea, which have been found in consortia with bacteria, Myxococcota (formerly Deltaproteobacteria), in particular 24,25 . Furthermore, a pivotal study by Pittis and Gabaldon in which relative timing of events in eukaryogenesis was inferred by comparing stem length from different ancestors in phylogenetic trees of conserved eukaryotic genes, suggested a late capture of the mitochondria, apparently preceded by acquisition of many genes from different bacteria, possibly, through an earlier endosymbiosis 26,27 . Here, we took advantage of the rapidly growing collection of archaeal, bacterial and eukaryotic genome sequences to infer the origins of core eukaryotic genes tracing to the LECA within a rigorous statistical framework based on evolutionary hypotheses testing using constrained phylogenetic trees. The results reveal the dominant contribution of Asgard to the origin of most functional classes of eukaryotic genes, which is compatible with a scenario of eukaryogenesis where many signature, complex features of eukaryotic cells evolved in the Asgard ancestor of eukaryotes prior to mitochondrial endosymbiosis. Main Phylogenomic analysis and global evolutionary associations of conserved eukaryotic proteins To identify associations between prokaryotic and eukaryotic protein families, separate hidden Markov model (HMM) databases for prokaryotes and eukaryotes were constructed using a custom, cascaded, sequence-to-profile clustering pipeline, implemented using mmseqs2 28 , followed by a multistep data-reduction and multiple sequence alignment (MSA) procedure to generate HMM profiles using hhsuite 29 32 (see Methods for details). Initially, a prokaryotic database of 75 million protein sequences was curated from 47,545 complete prokaryotic genomes obtained from the NCBI GenBank in November 2023 and supplemented with proteins extracted from 146 Asgard genome assemblies (Extended Data Figure 1) 6,8 . To avoid including genes present only within a narrow subset of species, possibly resulting from horizontal transfer from eukaryotes post-LECA, we reconstructed the “soft-core” pangenome for each of 26 curated prokaryotic taxonomic classes. These pangenomes include only those genes that are present in at least 67% of the families within each class of bacteria and archaea (see Methods) resulting in an initial database of 6.3 million nonredundant sequences. The initial eukaryotic database consisted of protein sequences from 993 species taken from EukProt v3 30 , cleaned using mmseqs2 to remove likely prokaryotic contaminants and clustered to 25 million nonredundant sequences. Both the eukaryotic and prokaryotic were clustered, clusters were realigned using muscle5 and turned into HMM profiles using HHsuite (see Methods). The resulting eukaryotic HMM dataset was queried against the prokaryotic dataset using hhblits 29 to identify sets of homologous protein sequences. Each eukaryotic cluster and all its significant prokaryotic hits constituted an individual sequence set, hereinafter referred to as a Eukaryotic/Prokaryotic Orthologous Cluster (EPOC). The EPOCs constitute groups of homologous proteins from eukaryotes and prokaryotes (each EPOC contains a unique set of eukaryotic proteins, but some clusters of prokaryotic proteins can be present in multiple EPOCs) that were used for phylogenetic tree construction, annotation, and evolutionary hypothesis testing. The final EPOCs include 5.7 million prokaryotic and 1.9 million eukaryotic sequences, mapping to 90% and 8% of the respective non-redundant datasets. To infer the most likely prokaryotic ancestry of the eukaryotic proteins in each EPOC, rather than relying on the tree topology directly, we employed a probabilistic approach for evolutionary hypothesis testing using constraint trees. Following the construction of an initial master tree, we exhaustively sampled all arrangements of likely sister clades relative to the eukaryotic outgroup and obtained Expected Likelihood Weights (ELW) for the set of possible sister clade models, (Extended Data Figure 2) 31 . Given that the ELW metric is analogous to model selection confidence, here we take it to be proportional to the probability of a sampled prokaryotic clade to be the true sister group of the given eukaryotic clade among a set of competing models. For each EPOC, our analysis dynamically accounts for long branch outliers and is robust to phylogenetically non-homogenous clades (see Methods). This analysis is further capable of resolving eukaryotic paraphyly, treating each eukaryotic clade within a EPOC as a single datapoint for downstream analysis. The resulting data included 14,300 EPOCs annotated using profiles generated from KEGG Orthology Groups (KOGs) 32 , each with an MSA generated using muscle5 33 , a maximum likelihood tree inferred using IQtree2 34 and associated ELW values for all candidate prokaryotic sister phyla. The analysis of prokaryotic ancestry was performed only for those eukaryotic clades that included more than 5 distinct taxonomic labels, with at least one coming from Amorphea and one from Diaphoretickes , the two expansive eukaryotic clades considered to emit from either the first or the second bifurcation in the evolution of eukaryotes 35,36 . Thus, these clades represent genes likely mapping back to the LECA. Considering the global average distribution of ELW values across all EPOCs covering 4330 unique KOGs, the single greatest average ELW (aELW), here referred to as an association, was with Asgard archaea. Further associations with Cyanobacteria , Actinomycetota , Betaproteobacteria and Alphaproteobacteria , as well as trace association with additional bacterial classes, were detected at lower levels. However, the number of included sequences and topological diversity of the trees varied substantially across the EPOCs, with many trees showing low maximum ELW values. Excluding EPOCs with a maximum ELW < 0.4 yielded a robust core set of 5590 EPOCs stemming from the LECA, covering 2540 KOGs and improving the interpretability of the results. This core set covers a wide range of ubiquitous metabolic functions, information processing pathways, transporters, and permeases, as well as regulatory and housekeeping proteins. By contrast, it does not include metabolic pathways limited to individual eukaryotic supergroups and thus can be considered a rough approximation of the core gene set of the LECA. Limiting the analysis to this core subset of well assigned eukaryotic families with wide taxonomic coverage notably increased the global Asgard association which now accounted for the 62 % of the ELW across more than 6000 unique data points across at least 2500 protein families. Our approach readily reproduced known evolutionary associations at the global functional level. Averaging ELW scores across EPOCs based on the KEGG ontology shows support for major Asgard association for, among other functional systems and pathways, the ribosome, RNA and DNA polymerases, Ras-like GTPases, Ubiquitin-mediated protein degradation, the proteasome, and large parts of the core metabolic network. In contrast, prominent alphaproteobacterial associations included proteins involved in oxidative phosphorylation, glutathione metabolism and Fe-S clusters biogenesis. Together, the Asgard and alphaproteobacterial associations amounted to an aELW of 0.55 + 0.06 = 0.61 across all of Metabolism and to 0.77 across Genetic Information Processing. Thus, we observed a far stronger overall association between Asgards and eukaryotes across diverse biological functions and pathways than previously described 5,6,8 although a consistent association between eukaryotic core metabolism and diverse bacterial phyla is still present. Broad, dominant Asgard contributions to eukaryogenesis In accord with the key contribution of Asgard archaea to eukaryogenesis, we observed associations of Asgard proteins with a wide array of cellular functions. In previous work, the strongest Asgards traces have been noted across the information processing systems, with unambiguous associations with DNA replication, core transcription, RNA processing as well as translation and protein trafficking 2,6,8 . Here, we consistently observed strong Asgard associations for genome replication and transcription and further detected pronounced Asgard traces for nucleotide excision repair, mismatch repair and homologous recombination. Additional well-known associations, such as ribosomal proteins, were extended to include translation factors, components of the co-translational membrane insertion machinery, protein targeting and aminoacyl-tRNA biosynthesis (Extended Data Figure 3, Extended Data File 1). Thus, all groups of core eukaryotic proteins involved in information processing appear to be almost exclusively of Asgard descent. We detected additional Asgard associations extending far beyond the information processing systems, including prominent contributions to the machinery involved in nucleocytoplasmic transport as well as downstream protein sorting, glycosylation and targeting. In particular, central components of the ER associated, N-linked glycan biosynthesis and transfer, including both cytoplasmic and lumenal monoglycosyltransferases, as well as the core of the oligosaccharyltransferase complex (OSTC) are strongly associated with Asgard (Extended Data Figure 3). Notably, enzymes associated with glycosylation maturation in the Golgi complex did not show strong Asgard or other prokaryotic associations in our analysis, possibly, due to extensive diversification of domain architecture in eukaryotes. The Asgard connections of the eukaryotic glycosylation machinery further included the synthesis of GPI-anchors, which post-translationally tether targeted proteins to the membrane 37 , here detected as unambiguously Asgard-derived. We also detected an Asgard origin of the 7-subunit (UDP-GlcNAc)-transferring (GPI-GnT)-monoglycosyltransferase complex responsible for initiating GPI-anchor synthesis, components required for the maturation of the GPI-anchor, as well as the transamidase complex and factors responsible for protein transfer onto the mature GPI-anchor (Extended Data Figure 3). Of major importance to eukaryogenesis is the provenance of the pathways for the biosynthesis of bacterial-type lipids, given that (at least) all binary archaea-bacteria symbiogenesis scenarios require a transition from archaeal to bacterial lipids in the membranes 9 . Although strong Asgard associations were observed for large parts of the overall metabolic network, we observed a high degree of mosaicism in the pathways for fatty acid synthesis and decay. The global aELW values favored Asgard origin for these pathways, but there were also notable associations with Actinomycetota . However, most KOGs within these pathways are represented by multiple EPOCs with conflicting assignments potentially obscuring any consistent signal (Extended Data Figure 3). By contrast, less perplexity was observed within the adjoining ER localized pathways for sphingolipid metabolism, a broad class of derived plasma membrane lipids in eukaryotes. Previous studies have highlighted possible convergent origins of this pathway in bacteria and eukaryotes 38,39 , but here we detected broad associations with Asgard. Of further note is the ER-associated isoprenoid biosynthesis pathway, here also found to be strongly associated with Asgard. In eukaryotes, isoprenoids form the precursor units for sterols, carotenoids and terpenoids, synthesized in the ER lumen via either the mevalonate pathway or the MEP/DOXP pathway. In Archaea, isoprenoids are the precursors for the ether-linked membrane lipids 40 . Here we found the mevalonate pathway, from Acetyl-COA to mevalonate and further to Farnesyl and Geranyl diphosphate, to be strongly Asgard-associated (Extended Data Figure 3), with the key enzymes hydroxymethylglutaryl-CoA synthase (HMGCS), mevalonate kinase (MVK), phosphomevalonate kinase (PMVK), mevalonate diphosphate decarboxylase (MVD) being clearly Asgard-derived. In conclusion, we detected Asgard associations across a wide range of cellular functions and metabolic pathways while noting a distinctly weaker Asgard signal for pathways involved in bacterial lipid biosynthesis, suggesting a complex evolutionary history. Limited and highly specific alphaproteobacterial contributions In line with the central role of mitochondria in eukaryotic energy metabolism, we primarily observed associations between Alphaproteobacteria and mitochondrially localized metabolic pathways. As expected, apart from the components of the mitochondrial translation system, the most prominent alphaproteobacterial associations were evident for complexes involved in oxidative phosphorylation and the associated ubiquinone synthesis (Extended Data Figure 3). Outside these central energy-transforming functions, we only detected sparse contributions from core alphaproteobacterial genes. One such prominent association was the pathway for iron sulfur cluster (ISC) biogenesis. As previously reported 41 , the ISC assembly machinery is of alphaproteobacterial origin, and in accord with these observations, the 4Fe-4S ISCA platforms as well as IBA57 and Fe-S cluster binding ferredoxin-1 and 2, were found to be strongly associated with Alphaproteobacteria (Figure 3, Extended Data Figure 3). However, the 2Fe-2S precursor scaffold ISCU showed mosaic associations, with a minor but clearly detectable Asgard contribution. Notably, the cysteine desulfurase NFS1 and the upstream pathways for the biosynthesis of sulfur-containing amino acids, cysteine, and methionine, were strongly Asgard- associated. The ISC biosynthesis is intimately linked to the general redox homeostasis and core sulfur metabolism via glutathione, directly coordinating Fe-S clusters during synthesis and transport 42 . In line with this role in Fe-S coordination, although some Asgard association persisted, we observed specific associations of glutathione metabolism with alphaproteobacteria, including glutathione hydrolase, dehydrogenase and reductase, as well as the family of glutathione transferases, GST, GSTP and GSTK1 (Extended Data Figure 3). Outside the mitochondria, ISC insertion depends on the cytosolic targeting complex CIA, which consists of CIAO1, CIA2B and MMS19 43 . For the CIA components CIA1 and CIA2B, we observed clear association with Asgard whereas MMS19 was not detected in our data. Taken together, these observations indicate that the contributions of alphaproteobacteria to the gene set of the LECA are functionally specific and apparently limited in scope, clearly centered around mitochondria-related functions. Paucity of functionally consistent contributions from other bacteria Although our analysis greatly expanded the Asgard contributions to eukaryogenesis, while also revealing a limited but prominent and functionally consistent alphaproteobacterial association, contributions from diverse other bacteria were consistently detected. For some biological functions, this diverse bacterial component accounted for the majority of the aELW, and roughly one third of the analyzed KOGs (680 of 2540), and EPOCs (1810 of 5590) were associated neither with known ancestors of endosymbionts, Alphaproteobacteria and Cyanobacteria , nor with Asgard. However, in a sharp contrast to Asgard associations including information processing, protein glycosylation and trafficking, and other functions as discussed above, or oxidative phosphorylation and sulfur metabolism for Alphaproteobacteria , EPOCs associated with diverse other bacteria showed few if any coherent functional trends. Considering the diverse set of bacteria, and all possible KEGG maps and modules, only Alphaproteobacteria were associated with pathway including more than 20 EPOCs and with a greater aELW than Asgard (Glutathione metabolism, Figure 4). For all other analyzed functional classes of eukaryotic genes, bacterial associations were weaker than the associations with Asgard (Figure 4). The second most individually prominent bacterial contribution was from Myxococcota , of the former deltaproteobacterial clade. Although globally weaker than Asgard associations, Myxococcota showed consistent associations with nicotinate and nucleotide synthesis, including both purine and pyrimidine synthesis, as well as nucleoside sugar metabolism. Myxococcota were unique in this regard as most bacteria showed diffuse associations across sugar and fatty acid metabolism, and/or diverse transporters. The nucleotide-related associations with Myxococcota were primarily limited to phosphatases and phosphoribosyltransferases acting on nucleotide sugars including 5 and 3’ nucleotidases, and the respective EPOCs showed little to no competing Asgard association. While noteworthy, these associations were limited to a few unique KOGs whereas all other associations with Myxococcota remained scattered across various pathways (Extended Data Figure 5). In addition to investigating metabolic pathways for global associations, we directly examined those individual EPOCs that were highly likely to be derived from diverse bacteria. For this analysis, we considered a stricter subset of the core EPOCs, requiring a eukaryotic outgroup containing at least 15 taxonomic clades and prokaryotic sister taxa with at least 20 sequences and ELW > 0.7. Only 16 of the 127 unique KOGs meeting these criteria were found to be associated with diverse bacterial lineages, mostly, Actinomycetota , FCB group and Betaproteobacteria , with minor contributions from Mycoplasmatota and Campylobacteriota (Extended Data Figure 6). The remaining 111 were Asgard-derived. The 16 prominent bacterial KOGs covered a wide range of cellular functions from MFS transporters to lipases to components of core sugar metabolism and cardiolipin synthesis once again with no noticeable trends. Taken together, although many associations were individually significant, no pathways appeared significantly enriched in associations with any single bacterial taxon, and we identified virtually no consistent trends among the bacterial associations. Instead, we interpret these diffuse associations as indications of highly specific contributions of limited functional scope from diverse bacteria other than the known endosymbiotic partners. Relative contributions of evolution pre- and post-LECA To compare the ancestral stem lengths in phylogenetic trees of core eukaryotic genes of different inferred origins we employed the methodology originally implemented by Pittis and Gabaldon 26 . Briefly, we define the raw stem length as the distance from the LECA node to the shared Last Common Ancestor (LCA) node of Eukarya and its most likely prokaryotic sister phyla. To account for differences in evolutionary rates, we divided this raw stem length by the median of eukaryotic branch lengths, measured from the LECA to each leaf (Figure 5C). Considering that the astronomic time post-LECA is the same for all genes, if the tempo and mode of pre- and post-LECA evolution were the same, the normalized stem length is proportional to the time elapsed since the divergence of the gene from its prokaryotic donor to LECA, that is, reflects the timing of acquisition. Using this approach, Pittis and Gabaldon found that proteins of alphaproteobacterial descent had significantly shorter normalized stem lengths than proteins of archaeal descent, consistent with a mitochondria-late scenario for eukaryogenesis 26 . Across the full set of eukaryotic stem lengths for our dataset (5,850 stems), we observed a wide distribution with a sharp maximum close to 0.05, highly reminiscent of the previous findings eukaryogenesis 26 (Figure 5B). However, in our analysis, the stems for genes of alphaproteobacterial origin were significantly (p ≈ 9.9x10 ‑6 ; Figure 5) longer than those of Asgard origin, and longer than those of other major bacterial contributors as well (Figure 5B). Comparison of the stem length distributions across functional classes of genes (Figure 5A) suggested an explanation for these observations. The shortest stems belong to the Genetic Information Processing functional category, belying our expectations of these genes being the longest-residing genes in the nascent eukaryotic lineage. We suggest that another major determinant of the relative stem length of a gene is the amount of adaptive evolution post-acquisition, which is necessary to adjust the newly acquired gene to the alien intracellular molecular environment. Thus, genes inherited from the Asgard ancestor, in particular, those involved in information processing, were pre-adapted to the cellular environment of the evolving protoeukaryotes, whereas genes acquired from radically different bacterial sources had to substantially adapt post-acquisition, increasing the apparent lengths of their pre-LECA stems. A case in point is the set of oxidative phosphorylation components, which were apparently acquired during the mitochondrial symbiogenesis, and thus, simultaneously, from the same donor. The distribution of their stem lengths (Figure 5A) was as broad as that for proteins of other functional classes (Figure 5A and Extended Data Figure 7), demonstrating that stem lengths are mostly determined by factors other than the acquisition time. Thus, our stem length analysis failed to provide unequivocal resolution of the temporal order of the prokaryotic contributions to eukaryogenesis. Nevertheless, the results appear to be compatible with the capture of the alphaproteobacterial endosymbiont by a host that was already on the path to eukaryotic-like complexification, along with capture of genes from various bacteria at different stages of eukaryogenesis. Discussion Reconstruction of eukaryogenesis is a moving target as strikingly demonstrated by the discovery and ongoing exploration of Asgard archaea which revolutionized the field 4-8 . Nevertheless, recent progress in prokaryote as well as eukaryote genome sequencing, along with the advances of metagenomics, created unprecedented opportunities for phylogenomics studies aimed at understanding the origins of eukaryotes. In this work, we took advantage of this expanded genome collection, coupling it with maximally sensitive HMM profile-profile searches, and performed a comprehensive phylogenetic analysis of core eukaryotic genes tentatively mapped to the LECA, within a statistical framework focused on testing evolutionary hypotheses using constraint trees. The principal outcome of this analysis is that a substantial majority of the LECA genes, for which the origin could be inferred with confidence, came from Asgard archaea. Particularly notable is the functional diversity of the apparent Asgard contribution to eukaryogenesis. Far from being limited to information processing and certain cellular processes, such as membrane remodeling, a strong Asgard trace was detected for most core functional systems and biochemical pathways of eukaryotes. In very few protein families did the detected bacterial contribution exceed that of Asgard archaea. The most conspicuous of the bacterial contributions, not unexpectedly, comes from Alphaproteobacteria , the ancestors of the mitochondria, and accounts for the core components of the electron transport chain complexes and iron-sulfur cluster metabolism, along with the mitochondrial translation system. However, although functionally relevant and consistent, the alphaproteobacterial contribution is relatively small and not comparable in scale to that of Asgard. The contributions of other bacterial phyla, although cumulatively substantial, fail to show any individually consistent trends, with the exception of one or two scattered metabolic pathways. Many previous studies, especially early ones, demonstrated the apparent chimeric origins of the core eukaryotic gene set, with the bacterial contribution quantitatively exceeding the archaeal one, and the latter being largely limited to information processing 17-19,44 . The discovery and exploration of Asgard archaea partially changed this notion by demonstrating that many genes involved in various cellular processes and systems, such as membrane remodeling and cytoskeleton, were of apparent Asgard origin 5,6,8 . The present work substantially extends the dominance of Asgard in the ancestry of the LECA genes, by revealing major Asgard contributions to nearly all functional systems and pathways of the eukaryotic cell. These findings imply a model of eukaryogenesis in which the ancestral Asgard archaeon already possessed many features characteristic of eukaryotic cells including cytoskeleton and the endomembrane system. These complex ancestral Asgard archaea might have led a predatory lifestyle conducive to engulfment of bacteria and serial capture of bacterial genes. Under this scenario, mitochondrial endosymbiosis, also perhaps facilitated by predation, while antedating the LECA, was a relatively late event that made a limited, even if functionally crucial, contribution to the gene composition of the emerging eukaryote (Figure 6). Other bacterial contributions appear to be piecemeal, coming from a limited capture of genes from diverse bacteria, likely, both before and after the mitochondrial endosymbiosis; there is no indication of another symbiotic event contributing to eukaryogenesis. This conclusion complements and extends a recent study that suggested emergence of some features of eukaryotic complexity prior to the mitochondrial endosymbiosis based on estimated timing of ancient gene duplications 45 . However, how eukaryotes acquired the bacterial-type membranes, remains an enigma. Whereas the eukaryotic pathways for isoprenoid synthesis are undoubtedly Asgard-derived, the ancestry of the enzymes involved in membrane lipid biosynthesis appears complex and remains unresolved in the present analysis, with multiple bacterial ancestors appearing likely. Our analysis of these pathways is likely confounded by the presence of two sets of lipid biosynthetic pathways, one mitochondrial and the other ER-associated. Nevertheless, the Asgard affinity of adjoining systems and pathways, such as the synthesis of GPI anchors, sphingolipids, or isoprenoids, shows that a substantial component of the lipid biosynthesis and modification capacity of eukaryotes persisted from the Asgard ancestor. If an Asgard archaeal cell was the principal scaffold on which the eukaryotic cell evolved, replacement of the archaeal membrane with a bacterial one must have occurred at an early stage of eukaryogenesis, likely, through a mixed membrane stage. Viable bacteria with a mixed bacterial-archaeal membrane have been experimentally engineered albeit resulting in a substantial fitness loss 46,47 . Membrane replacement might have even predated mitochondrial endosymbiosis, following early capture of the necessary bacterial enzymes by the Asgard archaeon (Figure 6). The conclusions of this work are subject to caveats stemming, above all, from the biased and still limited sampling of sequenced archaeal and bacterial genomes. In particular, almost all currently available genomes of Asgard archaea are incomplete metagenomic assemblies, with the exception of only three closed circular genomes 24,48,49 . Therefore, even though our current analysis points to a dominant contribution of Asgard to the eukaryotic gene core, it is nevertheless almost certainly an underestimate. The same pertains to Alphaproteobacteria both because of the sequencing bias whereby many of the sequenced alphaproteobacterial genomes come from symbionts and parasites with reduced gene complements, and because the ancestor of the mitochondria apparently belongs outside of the currently known diversity of Alphaproteobacteria 50 . Some of the other bacterial phyla are even more severely under-sampled, potentially, leading to underestimates of their contributions to the evolution of eukaryotes. For example, it is difficult to rule out that we miss a major signal from Myxococcota (formerly Deltaproteobacteria), that have been proposed as one of the partners in the syntrophy scenario of eukaryogenesis 15,22 , given the limited genomic information on these bacterial phyla. The expanding sampling of prokaryotic genome diversity might lead to a substantial modification or even complete revision of the model of eukaryogenesis proposed here. It nevertheless appears likely that the dominant contribution of Asgard archaea, the main finding of this work, is here to stay. Methods Database curation For prokaryotes 75 million sequences were curated from 47,545 completely sequenced prokaryotic genomes obtained from the NCBI GenBank (https://ftp.ncbi.nlm.nih.gov/genomes/) in November 2023 and supplemented with 441,150 proteins sequences from 146 assembled Asgard genomes 6,8 , selected to represent all clades in these two publications (where available, protein sequences were directly taken from GenBank annotations; otherwise, produced by Prodigal v2.6.3 51 , trained on the set of 12 complete or chromosome-level assembled Asgard genomes. To only include those sequences which are widely present within prokaryotic families, and to minimize the possibility of post- LECA horizontal transfer from eukaryotes, soft-core pangenomes were constructed for each of our 26 curated taxonomic groups, mostly conforming to the NCBI taxonomy rank “class” (see below). To construct the soft-core pangenomes, all sequences from each taxonomic group were clustered individually using mmseqs2 “mmseqs cluster -s 7 -c 0.9 --cov-mode 0” 28 . All clusters which did not contain sequences from at least 67% of all bacterial families within a prokaryotic class were rejected from the database. The eukaryotic database was constructed starting from a curated set of 72 genomes available on the NCBI GenBank. In order to provide a more accurate representation of eukaryotic diversity, this dataset was supplemented with EukProt v3 30 , a curated database aiming to provide a sparse representation of eukaryotic biology, to create Euk72Ep containing ~30.3 million sequences. As sequences included in EukProt v3 come from a diverse range of sources, some known to contain contaminant prokaryotic sequences, the database was screened for prokaryotic contamination. This screen was carried out by first constructing candidate prokaryotic and eukaryotic HMM databases as described below and querying both databases against the eukaryotic sequence database using hhblits 29 . All eukaryotic sequences with a top hit alignment score from a prokaryotic HMM were removed from the database prior to initial clustering. Curation of taxonomic labels In order to provide a compromise between taxonomic specificity and accuracy of clade assignment, we curated a set of taxonomic labels for both Euk72Ep and Prok2311As to act as representatives. To construct this set, we started with manually assigning all species in EukProt to their closest relative in the NCBI taxonomy. Then, each species from Prok2311As and Euk72Ep was assigned a taxonomic label corresponding to their “Class” rank (for example, Alphaproteobacteria , Thermococci , Mammalia ) based on the NCBI Taxonomy of November 2023 using tools from the ete3 3.1.3. toolkit 52 as well as custom scripts. For species without a corresponding “Class” rank, the closest relevant rank was manually assigned. This procedure yielded an initial list of 317 unique taxonomic class identifiers. As the rank “Class” does not evenly partition biological diversity, identifiers were further manually mapped to a set of 45 eukaryotic and 26 prokaryotic taxonomic superclasses, each present in the NCBI taxonomy (for example, Metazoa , Streptophyta , TACK archaea). Due to the partially unresolved taxonomic relationship within Asgard archaea, all candidate Asgard sequences from Prok2311 or from Refs. 19,20 were assigned the class label “Asgard”. Following the recent reclassification of Deltaproteobacteria into Myxococcota , Desulfobacterota , Bdellovibrionota and SAR324 species which could not be mapped into either of these clades using existing taxonomic annotation were classified as Deltaproteobacteria (0.025% of total data). Sequence clustering and profile database generation To transform Euk72Ep and Prok2311As into profile databases suitable for sensitive HMM-HMM searches, we implemented an unsupervised, cascaded sequence-profile clustering pipeline using the tools available in the mmseqs2 software suite 28 . In brief, each sequence database was initially clustered at 90% sequence identity and 80% pairwise coverage using “mmseqs linclust” and cluster representatives were chosen using “mmseqs result2repseq”. This procedure provides a non-redundant set of 6.3 million prokaryotic and 25.1 million eukaryotic sequences for further analysis. The non-redundant sequences are further collapsed into a set of initial profiles and consensus sequence pairs using “mmseqs cluster”, “mmseqs result2profile” and “mmseqs profile2consensus”. All consensus sequences are queried against their profiles using “mmseqs search” and results clustered using “mmseqs clust”. Clusters are mapped to their original non-redundant sequences using “mmseqs mergeclusters” and profiles and consensus sequences are constructed once again. In order to grow clusters based on their sequence-profile alignments, this procedure of “mmseqs profile2consensus - mmseqs search - mmseqs clust - mmseqs result2profile” is iterated until cluster sizes converge. An 80% pairwise coverage is maintained throughout all steps of the cascaded clustering protocol. The final clustered databases contain 14.1 million and 91.000 clusters for Euk72Ep and Prok2311As, respectively (Extended Data Figure 1). To transform the unsupervised clusters into a set of HMM profiles, we employed a mixed MSA construction strategy coupled with automatic data reduction to automatically create accurate alignments for a wide diversity of cluster sizes and sequence diversities. To construct HMMs, all non-singleton sequence clusters are first aligned using FAMSA 2.0.1. 53 to produce an initial alignment. Using the initial alignment, we then reduce the sequence space by first creating an approximate tree using FastTree 2.1.1. using “FastTree -gamma” 54 and iteratively removing the closest pairwise leaves using ete3 52 , keeping the leaf closest to the root, until the desired sequence set size is reached. For profile generation, we retain no more than 300 sequences for both Prok2311As and Euk72Ep. This reduced sequence set is then realigned with muscle5 5.1.linux64 33 using “muscle -diversified” with 5 replicates, and the maximum column confidence alignment is extracted. This alignment is then trimmed to keep columns with more than 0.2 bits of Shannon information. This process of data reduction/alignment/trimming is referred to as “prune-and-align” and employed later for downstream analysis. Trimmed alignments are turned into an HMM database using HHsuite with “hhconsensus -M 50”, hhmake and “cstranslate -f -x 0.3 -c 4 -I a3m”. The final HHsuite databases contained 26.000 and 1.6 million profiles for prokaryotes and eukaryotes, respectively, see (Extended Data Figure 1). Search for prokaryotic homologs of eukaryotic proteins Given that clade-specific protein families were not of direct relevance to this study, we reduced the search space by excluding HMM profiles from eukaryotic clusters with less than 10 sequences and a lowest common ancestor with a taxonomic rank below “Superkingdom” as calculated by “mmseqs lca”. This requires a sequence cluster to contain at least two sequences from different kingdoms, as defined by the NCBI taxonomy (ex. Amoebozoa and Haptista , or Viridiplanta and Opistokonta ), retaining 142.000 profiles of conserved eukaryotic proteins. To associate eukaryotic sequence clusters with homologous prokaryotic clusters, we queried the Euk72Ep HMM database against all Prok2311As HMM profiles using a single iteration of HHBlits 3.3.0 as “HHblits -n 1 -p 80” and retaining the hits with at least 80% HHblits probability and 80% pairwise profile length. These filters resulted in a total of 20,700 accepted eukaryotic query profiles targeting a total of 8,300 unique prokaryotic profiles. These clusters contain a total of 5.7 million and 1.9 million sequences. EPOC multiple sequence alignment construction Unless otherwise stated, all subsequent steps were carried out using custom python 3.9 scripts, with taxonomy and tree parsing carried out using ete3 52 . All sequences from a eukaryotic cluster with at least 10 sequences and a lowest common ancestor at the rank of “Superkingdom” (NCBI taxonomy), together with all members of the homologous prokaryotic clusters identified in the HMM-HMM search with at least 80% probability and pairwise sequence coverage, forms a EPOC. As EPOCs varied in size by more than 5 orders of magnitude (10 to 100.000 sequences), robust data subsampling is necessary for accurate downstream MSA generation and tree construction. We employed a variant of the “prune-and-align” strategy described above for profile generation taking into account the taxonomic distribution of the constituent sequences. Rather than iteratively pruning the closest leaves pairwise, we first label each leaf with its corresponding taxonomic label and only prune monophyletic groups, maintaining a count of all pruned leaves per clade. If the procedure fails to reach the target leaf number before converging, all leaves are ordered by the number of pruned relatives and leaves with the lowest number of relatives deleted (retaining the largest monophyletic groups) until the target size is reached. The result approximates a maximum diversity representation, preferentially pruning isolated singleton clades. Using this modified protocol, eukaryotic and prokaryotic sequences are cropped to a maximum of 30 and 70 sequences, respectively. Each EPOC therefore consists of a maximum of 100 representative sequences. All EPOCs are aligned using “muscle -diversified” with 5 replicates, and the maximum confidence alignment is extracted using the --maxcc argument. This alignment is then trimmed to keep columns with more than 0.15 bits of Shannon information content for tree reconstruction. EPOC tree construction and processing From the pruned and trimmed alignment, we created a maximum likelihood tree using IQtree2 2.3.5. “IQtree2 -B 1000 -bnni -m MFP -mset LG,Q.pfam --cmin 4 --cmax 12” with model parameters estimated by model finder plus 34 . Candidate models and rate category search ranges were selected based on a test set of 200 randomly chosen EPOCs from the filtered core set analyzed with full model parameter evaluation by Model finder 34 and 10 tree replicates. The LG and Q.pfam models were chosen as they consistently produced the highest ranking log-likelihoods. Rate categories were assigned based on the highest and lowest amount found among the top 10 average scoring models from the same set. As the master trees are built from sequences obtained through sensitive cascaded sequence-profile clustering we would in rare cases observe long-stemmed clade outliers, in addition to more common long stem leaf outliers. In order to assess and remove such erroneous data, we first estimated the log-normal distribution of all stem lengths using scipy.stats 1.11.1 and excluded any stems outside the 0-99.5 % Probability Point Function interval. This simultaneously removes long terminal leaf branches as well as highly diverged clades. In rare cases where this pruning would remove either the entire eukaryotic clade or more than 30 % of all leaves, the EPOC was discarded (387 EPOCs). The resulting trees were then re-rooted using a weighted midpoint approach so that the sum of all branch stem lengths on either side of the root was equal. To assign taxonomic clades, all leaf nodes were assigned taxonomic labels from our curated set and monophyletic clades were collapsed, taking as the representative the lowest branching leaf. We found that assigning clades purely based on a monophyletic definition of taxonomic purity severely hampered the analysis as leaves can be taxonomically incongruous with their neighbors. This effect most likely stems from cases of locally erroneous tree topology or local HGT, but severely complicates the downstream analysis. We therefore adopted the notion of a “soft LCA” representing the root of a close-to-monophyletic clade. To greedily rank all the best “soft LCAs”, for each taxonomic label, the tree was traversed from each monophyletic clade root to the tree root, calculating a score for each internal node: (number of leaves with label X in clade / clade size) * (number of leaves with label X in clade / total number of label X) This metric balances taxonomic “purity” and “scope” for each possible clade of label X. All such nodes are then ordered by score. This allows for paraphyly across all taxonomic labels. To avoid overinterpreting small clades, we only consider clades to be valid if they represent at least 3 sequences with clade purity > 0.8 for prokaryotes, and at least 5 sequences with purity > 0.8 for eukaryotes. Trees that fail to identify either any eukaryotic or any prokaryotic clades under these constraints, primarily derived from small EPOCs, are discarded (411 EPOCs). Trees with more than 3 valid eukaryotic clades were considered of high paraphyly and likewise discarded (781 EPOCs). Evolutionary hypothesis testing using constraint trees In order to assess the relative probabilities of each sampled prokaryotic “soft LCA'' to represent the eukaryotic sister clade, we employed evolutionary hypothesis testing using constraint trees as implemented in IQtree2 34 . For each EPOC and for each detected eukaryotic clade, as defined by its corresponding “soft LCA”, we generated a set of constraint trees, one for each possible prokaryotic sister. Each constraint tree had three defined clades, 1) the eukaryotic group as defined by all sequences below the eukaryotic LCA, 2) the prokaryotic sister group and 3) all other prokaryotic leaves. We then construct a set of local trees from slices of the original MSA using IQtree2, forcing the constraint tree topology using “iqtree2 -g”, using the same evolutionary model as the original, unconstrained “master tree”. The set of constrained trees is then ranked using “iqtree2 -z” calculating the relative model assignment confidence for each of the constrained topologies and resulting test metrics are saved for downstream analysis. As the number of trees to evaluate scales by the number of eukaryotic clades multiplied by the number of prokaryotic clades, we limit the sampling to a maximum of 3 clades per tree for eukaryotes, discarding rare cases of high eukaryotic paraphyly, and only consider the 12 closest prokaryotic clades to each individual eukaryotic clade, based on their topological distance (Extended Data Figure 2). The procedure results in an Expected Likelihood Weight (ELW) 31 for each of the evaluated model trees, here taken as being analogous to model selection confidence that an evaluated prokaryotic clade is the most likely sister to a eukaryotic clade. EPOC annotation To functionally annotate protein families within EPOCs, we generated profiles based on protein sequences in KEGG release 110 32 . For each Kegg Orthologous Group (KOG), we extracted Kegg metadata including KEGG pathways and BRITE labels via KEGG’s API. At the time of parsing (January 2024), KEGG contained 26,695 KOGs. Due to KEGG API constraints, protein sequences in KEGG were extracted from Uniprot, first by mapping KEGG proteins to Uniprot IDs via Uniprot’s ID mapping service (uniprot.org/id-mapping), then by extracting the sequences for these Uniprot IDs corresponding to each KOG 55 . For each KOG, we generated HHM profiles using the aforementioned prune-and-align pipeline. To increase specificity of labeling, we partitioned those KOGs which encompass both prokaryotic and eukaryotic sequences into separate taxonomic groups and sequence alignments were generated separately for prokaryotes and eukaryotes. Each eukaryotic HMM profile forming the basis for a EPOC was queried against our KEGG profile database using “HHblits -n 3 -p 80”. Results were subsequently filtered to a pairwise profile coverage of 0.5 using custom scripts. As a result, 13,100 of the 14,300 EPOCs could be annotated by at least one KOG profile at 80% probability and 0.5 pairwise coverage. For all downstream analyses, we considered only the top hit as relevant for annotation. Data filtering and removal of likely horizontal gene transfer cases Unless otherwise stated, all data is interpreted from the core set of EPOCs meeting the following criteria: 1) the eukaryotic constituent cluster profile identifies at least one target profile in our KEGG profile database at a probability of 80% and 50% coverage, 2) the eukaryotic clades include more than 5 distinct taxonomic labels, at least one from both Amorphea and Diaphoretickes, thus encompassing LECA, 3) only prokaryotic sister clades with 0.4 < ELW < 0.99 were considered reliable and were included in the analysis. Cases of ELW < 0.99 were not considered as these are found to be disproportionately eukaryotic-like, often branching from within the eukaryotic clade itself, and thus likely derived from late horizontal transfer post LECA. Due to the presence of such high ELW values within Alphaproteobacterial associations with Oxidative phosphorylation, in this case, as an exception, ELW values of 1 were included. Data visualization and plotting All plots were prepared using python v3.9 and pandas v.2.0.3 with Altair v.4.2.2. Layout, annotation, and vector editing was done using Inkscape v.1.1.1. All statistical tests were carried out using scipy.stats 1.11.1 Declarations Data availability All initial databases and taxonomy annotation as well as all final parsed and tabulated data is available either at Zenodo url:10.5281/zenodo.14002645 or and at https://ftp.ncbi.nih.gov/pub/wolf/ without any restrictions. All intermediate files, including MSAs and master trees for all EPOCs are available at Zenodo doi:10.5281/zenodo.14004407 Code availability All code and scripts used to generate the final parsed data is available on Github at: https://github.com/VictorTobiasson/eukgen Supplementary Data File 1 Excel file with four sheets: 1) List of revised taxonomy mapping species with without well curated taxonomy to closes existing taxonomic ranks. 2) List of revised taxonomy merging sparsely populated classes. 3) Manually curated list of genes involved in Fe-S assembly. 4) Tabular data of all tables included in figures. Author contributions E.V.K. initiated the study; V.T. and Y.I.W. developed the pipeline for data analysis; V.T. and J.L. collected the data; V.T., J.L., Y.I.W and E.V.K. performed data analysis; V.T. and E.V.K wrote the manuscript that was read, edited and approved by all authors. Acknowledgements The authors thank Toni Gabaldon and Koonin Group members for helpful discussions. The authors’ research is supported by the Intramural Research Program of the National Institutes of Health of the USA (National Library of Medicine). Competing interests The authors declare no competing interests. References Embley, T. M. & Martin, W. Eukaryotic evolution, changes and challenges. Nature 440 , 623-630 (2006). Vosseberg, J. et al. The emerging view on the origin and early evolution of eukaryotic cells. Nature 633 , 295-305 (2024). https://doi.org:10.1038/s41586-024-07677-6 van der Giezen, M. Hydrogenosomes and mitosomes: conservation and evolution of functions. J Eukaryot Microbiol 56 , 221-231 (2009). https://doi.org:JEU407 [pii] 10.1111/j.1550-7408.2009.00407.x Spang, A. et al. Complex archaea that bridge the gap between prokaryotes and eukaryotes. Nature 521 , 173-179 (2015). https://doi.org:10.1038/nature14447 Zaremba-Niedzwiedzka, K. et al. Asgard archaea illuminate the origin of eukaryotic cellular complexity. Nature 541 , 353-358 (2017). https://doi.org:10.1038/nature21031 Liu, Y. et al. Expanded diversity of Asgard archaea and their relationships with eukaryotes. Nature 593 , 553-557 (2021). https://doi.org:10.1038/s41586-021-03494-3 Eme, L., Spang, A., Lombard, J., Stairs, C. W. & Ettema, T. J. G. Archaea and the origin of eukaryotes. Nat Rev Microbiol 15 , 711-723 (2017). https://doi.org:10.1038/nrmicro.2017.133 Eme, L. et al. Inference and reconstruction of the heimdallarchaeial ancestry of eukaryotes. Nature 618 , 992-999 (2023). https://doi.org:10.1038/s41586-023-06186-2 Lopez-Garcia, P. & Moreira, D. The symbiotic origin of the eukaryotic cell. C R Biol 346 , 55-73 (2023). https://doi.org:10.5802/crbiol.118 Alberts, B. et al. Molecular Biology of the Cell . 7 edn, (W.W. Norton & Co, 2022). Cavalier-Smith, T. Eukaryotes with no mitochondria. Nature 326 , 332-333 (1987). https://doi.org:10.1038/326332a0 Sogin, M. L., Gunderson, J. H., Elwood, H. J., Alonso, R. A. & Peattie, D. A. Phylogenetic meaning of the kingdom concept: an unusual ribosomal RNA from Giardia lamblia. Science 243 , 75-77 (1989). https://doi.org:10.1126/science.2911720 Sagan, L. On the origin of mitosing cells. J Theor Biol 14 , 255-274 (1967). Martin, W., Hoffmeister, M., Rotte, C. & Henze, K. An overview of endosymbiotic models for the origins of eukaryotes, their ATP-producing organelles (mitochondria and hydrogenosomes), and their heterotrophic lifestyle. Biol Chem 382 , 1521-1539 (2001). https://doi.org:10.1515/BC.2001.187 Lopez-Garcia, P. & Moreira, D. The Syntrophy hypothesis for the origin of eukaryotes revisited. Nat Microbiol 5 , 655-667 (2020). https://doi.org:10.1038/s41564-020-0710-4 Brown, J. R. & Doolittle, W. F. Archaea and the prokaryote-to-eukaryote transition. Microbiol Mol Biol Rev 61 , 456-502 (1997). https://doi.org:10.1128/mmbr.61.4.456-502.1997 Rivera, M. C., Jain, R., Moore, J. E. & Lake, J. A. Genomic evidence for two functionally distinct gene classes. Proc Natl Acad Sci U S A 95 , 6239-6244 (1998). https://doi.org:10.1073/pnas.95.11.6239 Ribeiro, S. & Golding, G. B. The mosaic nature of the eukaryotic nucleus. Mol Biol Evol 15 , 779-788 (1998). https://doi.org:10.1093/oxfordjournals.molbev.a025983 Koonin, E. V. et al. A comprehensive evolutionary classification of proteins encoded in complete eukaryotic genomes. Genome Biol 5 , R7 (2004). Imachi, H. et al. Promethearchaeum syntrophicum gen. nov., sp. nov., an anaerobic, obligately syntrophic archaeon, the first isolate of the lineage 'Asgard' archaea, and proposal of the new archaeal phylum Promethearchaeota phyl. nov. and kingdom Promethearchaeati regn. nov. Int J Syst Evol Microbiol 74 (2024). https://doi.org:10.1099/ijsem.0.006435 Lombard, J., Lopez-Garcia, P. & Moreira, D. The early evolution of lipid membranes and the three domains of life. Nat Rev Microbiol 10 , 507-515 (2012). https://doi.org:10.1038/nrmicro2815 Moreira, D. & Lopez-Garcia, P. Symbiosis between methanogenic archaea and delta-proteobacteria as the origin of eukaryotes: the syntrophic hypothesis. J Mol Evol 47 , 517-530 (1998). https://doi.org:10.1007/pl00006408 Krupovic, M., Dolja, V. V. & Koonin, E. V. The virome of the last eukaryotic common ancestor and eukaryogenesis. Nat Microbiol 8 , 1008-1017 (2023). https://doi.org:10.1038/s41564-023-01378-y Imachi, H. et al. Isolation of an archaeon at the prokaryote-eukaryote interface. Nature 577 , 519-525 (2020). https://doi.org:10.1038/s41586-019-1916-6 Lopez-Garcia, P. & Moreira, D. Cultured Asgard Archaea Shed Light on Eukaryogenesis. Cell 181 , 232-235 (2020). https://doi.org:10.1016/j.cell.2020.03.058 Pittis, A. A. & Gabaldon, T. Late acquisition of mitochondria by a host with chimaeric prokaryotic ancestry. Nature 531 , 101-104 (2016). https://doi.org:10.1038/nature16941 Gabaldon, T. Origin and Early Evolution of the Eukaryotic Cell. Annu Rev Microbiol 75 , 631-647 (2021). https://doi.org:10.1146/annurev-micro-090817-062213 Steinegger, M. & Soding, J. MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nat Biotechnol 35 , 1026-1028 (2017). https://doi.org:10.1038/nbt.3988 Steinegger, M. et al. HH-suite3 for fast remote homology detection and deep protein annotation. BMC Bioinformatics 20 , 473 (2019). https://doi.org:10.1186/s12859-019-3019-7 Richter, D. J. et al. EukProt: A database of genome-scale predicted proteins across the diversity of eukaryotes Peer Community J 2 , e56 (2022). Strimmer, K. & Rambaut, A. Inferring confidence sets of possibly misspecified gene trees. Proc Biol Sci 269 , 137-142 (2002). https://doi.org:10.1098/rspb.2001.1862 Kanehisa, M., Furumichi, M., Tanabe, M., Sato, Y. & Morishima, K. KEGG: new perspectives on genomes, pathways, diseases and drugs. Nucleic Acids Res 45 , D353-D361 (2017). https://doi.org:10.1093/nar/gkw1092 Edgar, R. C. Muscle5: High-accuracy alignment ensembles enable unbiased assessments of sequence homology and phylogeny. Nat Commun 13 , 6968 (2022). https://doi.org:10.1038/s41467-022-34630-w Minh, B. Q. et al. IQ-TREE 2: New Models and Efficient Methods for Phylogenetic Inference in the Genomic Era. Mol Biol Evol 37 , 1530-1534 (2020). https://doi.org:10.1093/molbev/msaa015 Derelle, R. et al. Bacterial proteins pinpoint a single eukaryotic root. Proc Natl Acad Sci U S A 112 , E693-699 (2015). https://doi.org:10.1073/pnas.1420657112 Al Jewari, C. & Baldauf, S. L. Conflict over the Eukaryote Root Resides in Strong Outliers, Mosaics and Missing Data Sensitivity of Site-Specific (CAT) Mixture Models. Syst Biol 72 , 1-16 (2023). https://doi.org:10.1093/sysbio/syac029 Paulick, M. G. & Bertozzi, C. R. The glycosylphosphatidylinositol anchor: a complex membrane-anchoring structure for proteins. Biochemistry 47 , 6991-7000 (2008). https://doi.org:10.1021/bi8006324 Stankeviciute, G. et al. Convergent evolution of bacterial ceramide synthesis. Nat Chem Biol 18 , 305-312 (2022). https://doi.org:10.1038/s41589-021-00948-7 Biran, A., Santos, T. C. B., Dingjan, T. & Futerman, A. H. The Sphinx and the egg: Evolutionary enigmas of the (glyco)sphingolipid biosynthetic pathway. Biochim Biophys Acta Mol Cell Biol Lipids 1869 , 159462 (2024). https://doi.org:10.1016/j.bbalip.2024.159462 Hoshino, Y. & Gaucher, E. A. On the Origin of Isoprenoid Biosynthesis. Mol Biol Evol 35 , 2185-2197 (2018). https://doi.org:10.1093/molbev/msy120 Richards, T. A. & van der Giezen, M. Evolution of the Isd11-IscS complex reveals a single alpha-proteobacterial endosymbiosis for all eukaryotes. Mol Biol Evol 23 , 1341-1344 (2006). https://doi.org:10.1093/molbev/msl001 Cassier-Chauvat, C., Marceau, F., Farci, S., Ouchane, S. & Chauvat, F. The Glutathione System: A Journey from Cyanobacteria to Higher Eukaryotes. Antioxidants (Basel) 12 (2023). https://doi.org:10.3390/antiox12061199 Kassube, S. A. & Thoma, N. H. Structural insights into Fe-S protein biogenesis by the CIA targeting complex. Nat Struct Mol Biol 27 , 735-742 (2020). https://doi.org:10.1038/s41594-020-0454-0 Hedges, S. B. The origin and evolution of model organisms. Nat Rev Genet 3 , 838-849 (2002). https://doi.org:10.1038/nrg929 Vosseberg, J. et al. Timing the origin of eukaryotic cellular complexity with ancient duplications. Nat Ecol Evol 5 , 92-100 (2021). https://doi.org:10.1038/s41559-020-01320-z Caforio, A. et al. Converting Escherichia coli into an archaebacterium with a hybrid heterochiral membrane. Proc Natl Acad Sci U S A 115 , 3704-3709 (2018). https://doi.org:10.1073/pnas.1721604115 Hoekzema, M., Jiang, J. & Driessen, A. J. M. Optimizing Archaeal Lipid Biosynthesis in Escherichia coli. ACS Synth Biol 13 , 2470-2479 (2024). https://doi.org:10.1021/acssynbio.4c00235 Tamarit, D. et al. A closed Candidatus Odinarchaeum chromosome exposes Asgard archaeal viruses. Nat Microbiol 7 , 948-952 (2022). https://doi.org:10.1038/s41564-022-01122-y Rodrigues-Oliveira, T. et al. Actin cytoskeleton and complex cell architecture in an Asgard archaeon. Nature 613 , 332-339 (2023). https://doi.org:10.1038/s41586-022-05550-y Munoz-Gomez, S. A. et al. Site-and-branch-heterogeneous analyses of an expanded dataset favour mitochondria as sister to known Alphaproteobacteria. Nat Ecol Evol 6 , 253-262 (2022). https://doi.org:10.1038/s41559-021-01638-2 Hyatt, D. et al. Prodigal: prokaryotic gene recognition and translation initiation site identification. BMC Bioinformatics 11 , 119 (2010). https://doi.org:10.1186/1471-2105-11-119 Huerta-Cepas, J., Serra, F. & Bork, P. ETE 3: Reconstruction, Analysis, and Visualization of Phylogenomic Data. Mol Biol Evol 33 , 1635-1638 (2016). https://doi.org:10.1093/molbev/msw046 Deorowicz, S., Debudaj-Grabysz, A. & Gudys, A. FAMSA: Fast and accurate multiple sequence alignment of huge protein families. Sci Rep 6 , 33964 (2016). https://doi.org:10.1038/srep33964 Price, M. N., Dehal, P. S. & Arkin, A. P. FastTree 2--approximately maximum-likelihood trees for large alignments. PLoS One 5 , e9490 (2010). https://doi.org:10.1371/journal.pone.0009490 UniProt, C. UniProt: the universal protein knowledgebase in 2021. Nucleic Acids Res 49 , D480-D489 (2021). https://doi.org:10.1093/nar/gkaa1100 Additional Declarations There is NO Competing Interest. Supplementary Files SupplementarydataTobiasson27102024.xlsx Dataset 1 Tobiassonextendeddata.pdf Extended Data 1 Cite Share Download PDF Status: Published Journal Publication published 14 Jan, 2026 Read the published version in Nature → Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-5352492","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Biological Sciences - Article","associatedPublications":[],"authors":[{"id":372082455,"identity":"495d7598-1a58-475b-8b65-dd004317cb5a","order_by":0,"name":"Eugene Koonin","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA5UlEQVRIiWNgGAWjYBACCRCRwGADZDE2MDAUAHnMPERpSYNqMSBWCwPDYSgLpIWBgBbJ9uZjDx7uOJ/HP7u5gbnAYBuDOTvvAcYvFYdxapHmOZZukHjmdrHEnYMNzDMMbjNYNvMlMMucwa1FTiLHTCKx7XbiBonEBmYeoBaDwzwGzJJtaYS0nCNBizREywFULYwf22xwe7/nWBpQS3LiDKBfgIpv84C0HGY4g1uLxPHmY5I/2+wS+2e3P3zMU3FbzuD8GcOHPyokcGpBAQcYoJFymEDUYAGMP0jWMgpGwSgYBcMYAAABHk9fgmpqBgAAAABJRU5ErkJggg==","orcid":"https://orcid.org/0000-0003-3943-8299","institution":"National Institutes of Health","correspondingAuthor":true,"prefix":"","firstName":"Eugene","middleName":"","lastName":"Koonin","suffix":""},{"id":372082456,"identity":"6ddcafe2-0f23-4577-95c7-6e8d3b672c34","order_by":1,"name":"Victor Tobiasson","email":"","orcid":"https://orcid.org/0000-0001-8920-017X","institution":"National Center for Biotechnology Information","correspondingAuthor":false,"prefix":"","firstName":"Victor","middleName":"","lastName":"Tobiasson","suffix":""},{"id":372082457,"identity":"2a0ec2d9-c52b-41d8-8ce3-b1a9fd3c30e1","order_by":2,"name":"Jacob Luo","email":"","orcid":"https://orcid.org/0000-0003-0018-4850","institution":"","correspondingAuthor":false,"prefix":"","firstName":"Jacob","middleName":"","lastName":"Luo","suffix":""},{"id":372082458,"identity":"848ec99b-189e-43b1-8f2f-21b5cef4bf30","order_by":3,"name":"Yuri Wolf","email":"","orcid":"https://orcid.org/0000-0002-0247-8708","institution":"National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health","correspondingAuthor":false,"prefix":"","firstName":"Yuri","middleName":"","lastName":"Wolf","suffix":""}],"badges":[],"createdAt":"2024-10-29 08:35:15","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-5352492/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-5352492/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1038/s41586-025-09960-6","type":"published","date":"2026-01-14T05:00:00+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":67930793,"identity":"c7aebc5b-b785-46ae-ad17-4eb77133ec91","added_by":"auto","created_at":"2024-10-31 09:48:26","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":434478,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003ePhylogenetic associations of functional classes of eukaryotic proteins with major archaeal and bacterial taxa. A) \u003c/strong\u003eSimplified pipeline consisting of, from left to right, one eukaryotic and one prokaryotic sequence database transformed into HMM profiles. The query of eukaryotic against prokaryotic profiles generates 18.800 valid EPOCs, automatically annotated, and with detected prokaryotic clades assessed for their probability to be the closest eukaryotic sister phyla by permutation testing using IQTree2. \u003cstrong\u003eB) \u003c/strong\u003eGlobal average ELW for all EPOCs within the unfiltered and core datasets as well as top level breakdowns for metabolic enzymes and proteins involved in information processing as defined by the KEGG ontology “BRITE” at the top “A” level (KO:09100, KO:09120). \u003cstrong\u003eC) \u003c/strong\u003eFurther subdivision of average ELW values across the lower BRITE “B” level for each BRITE A category.\u003c/p\u003e","description":"","filename":"image1.png","url":"https://assets-eu.researchsquare.com/files/rs-5352492/v1/a07b6cc66930b69b42536e4c.png"},{"id":67930794,"identity":"c181ee4f-d23f-4470-bba4-3162dfbf2ea0","added_by":"auto","created_at":"2024-10-31 09:48:26","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":418802,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eAsgard origins of eukaryotic pathways. A)\u003c/strong\u003e Top 15 KEGG maps with strongest Asgard association, only concludes maps with more than 20 EPOCs. \u003cstrong\u003eB)\u003c/strong\u003eMevalonate pathway in eukaryotes with each protein colored by inferred ancestry. Split backgrounds correspond to mixed signals.\u003c/p\u003e","description":"","filename":"image2.png","url":"https://assets-eu.researchsquare.com/files/rs-5352492/v1/daa1d9cbc428ac5f312edbd8.png"},{"id":67930795,"identity":"a331648b-665d-4f4e-b8da-b5d6f5c06d4d","added_by":"auto","created_at":"2024-10-31 09:48:26","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":91041,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eAncestry of eukaryotic Fe-S cluster synthesis pathways.\u003c/strong\u003e Schematic for Fe-S cluster assembly including downstream cytosolic insertion. Specific ancestry of the mitochondrial 4Fe-4S synthesis and assembly pathways from alphaproteobacteria was observed whereas cytosolic or cytosolic-derived components were found to be associated with Asgard.\u003c/p\u003e","description":"","filename":"image3.png","url":"https://assets-eu.researchsquare.com/files/rs-5352492/v1/194063515bdb7bee474866f1.png"},{"id":67930797,"identity":"c959c584-1934-4e9d-ae34-b902ad275d55","added_by":"auto","created_at":"2024-10-31 09:48:26","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":531456,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eDiverse bacterial contributions to the origins of eukaryotic genes of different functional classes compared to the contributions of Asgard archaea. \u003c/strong\u003eMaximum ELW for the single highest enriched KEGG maps per prokaryotic taxa, relative to Asgard contributions for each respective map, is shown.\u003c/p\u003e","description":"","filename":"image4.png","url":"https://assets-eu.researchsquare.com/files/rs-5352492/v1/51bde57c40c820dc2cc2368d.png"},{"id":67930798,"identity":"a8ae827b-105a-4ae8-8acb-3aae77cf91fa","added_by":"auto","created_at":"2024-10-31 09:48:26","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":364274,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eAncestral stem length distributions for different groups of prokaryotes contributing to eukaryogenesis and different functional classes of genes. A) \u003c/strong\u003eDistribution of normalized eukaryotic stem lengths for different top-level eukaryotic functions as defined by BRITE A Metabolism (KO:09100), Genetic Information Processing (KO:09120), Oxidative phosphorylation (map00190 excluding KOGs for V-type ATPases). Includes all eukaryotic clade stem lengths between 0.01 and 1.0.\u003cstrong\u003e B) \u003c/strong\u003eDistribution of normalized eukaryotic stem lengths for prokaryotic taxa associated with eukaryotes. E-values calculated using Mann-Whitney U test to assess whether the Asgard distribution was significantly different from the distributions for the indicated taxa. \u003cstrong\u003eC)\u003c/strong\u003e Relative amounts of pre- and post-LECA evolution.\u003c/p\u003e","description":"","filename":"image5.png","url":"https://assets-eu.researchsquare.com/files/rs-5352492/v1/5ae04e802d28180fa762c95d.png"},{"id":67932312,"identity":"28793b1b-c54f-44ed-90c5-b7512f8737f0","added_by":"auto","created_at":"2024-10-31 09:56:26","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":568050,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eHypothetical scenario of eukaryogenesis\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"image6.png","url":"https://assets-eu.researchsquare.com/files/rs-5352492/v1/887fbe2f0e1eb88f75a7bd69.png"},{"id":100296088,"identity":"3490fa3d-b287-491e-b30c-6181682e757c","added_by":"auto","created_at":"2026-01-15 08:10:44","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":3223854,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5352492/v1/2337c34d-1ddc-42e2-b023-6094bc5dd8a1.pdf"},{"id":67930801,"identity":"93b7951f-89e6-4516-a678-d3f0ea70842b","added_by":"auto","created_at":"2024-10-31 09:48:26","extension":"xlsx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":239761,"visible":true,"origin":"","legend":"Dataset 1","description":"","filename":"SupplementarydataTobiasson27102024.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-5352492/v1/cde64bc428ff1fa661889203.xlsx"},{"id":67932311,"identity":"099db29b-baf5-4bc6-bb6d-671b4a3d73fb","added_by":"auto","created_at":"2024-10-31 09:56:26","extension":"pdf","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":1693874,"visible":true,"origin":"","legend":"Extended Data 1","description":"","filename":"Tobiassonextendeddata.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5352492/v1/4cc22413bdab36cd5ed2faae.pdf"}],"financialInterests":"There is \u003cb\u003eNO\u003c/b\u003e Competing Interest.","formattedTitle":"Dominant contribution of Asgard archaea to eukaryogenesis","fulltext":[{"header":"Introduction","content":"\u003cp\u003eEukaryotes drastically differ from archaea and bacteria (collectively, prokaryotes) by the highly complex cellular organization. The signature features of this organizational complexity include the endomembrane system, in particular, the nucleus, the eponymous eukaryotic organelle, the elaborate cytoskeleton, and the mitochondrion, the energy-converting organelle that evolved from an alphaproteobacterial endosymbiont \u003csup\u003e10\u003c/sup\u003e. Early models of eukaryogenesis featured a protoeukaryote lacking mitochondria that captured and domesticated an alphaproteobacterium at a relatively late stage of evolution \u003csup\u003e11,12\u003c/sup\u003e. However, subsequent research revealed no primary amitochondrial eukaryotes (although many have lost mitochondria through reductive evolution) \u003csup\u003e3\u003c/sup\u003e, suggesting that the Last Eukaryotic Common Ancestor (LECA) was already endowed with mitochondria along with the other signatures of the eukaryotic cellular organization \u003csup\u003e1,9\u003c/sup\u003e. These findings lend credence to scenarios of eukaryogenesis in which mitochondrial endosymbiosis triggered cellular reorganization, giving rise to the eukaryotic cellular complexity \u003csup\u003e9,13-15\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003ePhylogenomic analyses show that those conserved, core genes of eukaryotes that can be traced to the LECA are a mix originating from archaea and from various bacteria. The archaea-derived genes were found primarily in information processing systems (replication, core transcription and translation) whereas genes of apparent bacterial origin comprise the operational component of the eukaryotic gene complement, in particular, encoding metabolic enzymes \u003csup\u003e16-19\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003eThe study of eukaryogenesis was transformed by the discovery and subsequent exploration of Asgard archaea (currently, phylum \u003cem\u003ePromethearchaeota\u003c/em\u003e within kingdom \u003cem\u003ePromethearchaeati\u003c/em\u003e\u003csup\u003e20\u003c/sup\u003e; hereafter, Asgard, for brevity) that includes the closest known archaeal relatives of eukaryotes \u003csup\u003e4-8\u003c/sup\u003e. Phylogenetic analyses of highly conserved genes have been pushing the eukaryotic branch progressively deeper with the Asgard tree. The latest such analysis identified the order \u003cem\u003eHodarchaeales\u003c/em\u003e, within the Asgard class \u003cem\u003eHeimdallarchaeia\u003c/em\u003e, as the likely sister group of eukaryotes \u003csup\u003e8\u003c/sup\u003e. Perhaps, even more notably, it has been shown that Asgard archaea encode homologs of a broad variety of eukaryote signature proteins beyond the core information processing componentry, in particular, many cytoskeletal proteins and those involved in membrane remodeling \u003csup\u003e5,6,8\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003eA variety of eukaryogenesis models have been proposed, differing with respect to the timing of the origin of the eukaryotic cellular organization and the exact nature and succession of the underlying events \u003csup\u003e1,2,9\u003c/sup\u003e. The endosymbiotic origin of the mitochondria from an alphaproteobacterium is indisputable, but the identity of the host of the proto-mitochondrial endosymbiont remains a matter of debate. The most straightforward models posit an Asgard archaeal host. However, a major conundrum for such scenarios of eukaryogenesis is the chemistry of cell membrane lipids and the enzymology of their biosynthesis, which are unrelated in archaea and bacteria, with eukaryotic being of the bacterial type \u003csup\u003e21\u003c/sup\u003e. Accordingly, any model of eukaryogenesis that includes an archaeal host would require a membrane replacement step. More complex, alternative models of eukaryogenesis, underpinned by metabolic symbiosis, postulate two endosymbiotic events whereby an Asgard archaeon was first engulfed by a bacterium, followed by the loss of the archaeal membrane, and then, a second endosymbiosis gave rise to the mitochondria \u003csup\u003e15,22,23\u003c/sup\u003e. Even if seemingly less parsimonious than single-endosymbiosis scenarios, such models of eukaryogenesis account for the continuity of bacterial membranes and also appear compatible with the syntrophic lifestyle of at least some Asgard archaea, which have been found in consortia with bacteria, \u003cem\u003eMyxococcota \u003c/em\u003e(formerly Deltaproteobacteria), in particular \u003csup\u003e24,25\u003c/sup\u003e. Furthermore, a pivotal study by Pittis and Gabaldon in which relative timing of events in eukaryogenesis was inferred by comparing stem length from different ancestors in phylogenetic trees of conserved eukaryotic genes, suggested a late capture of the mitochondria, apparently preceded by acquisition of many genes from different bacteria, possibly, through an earlier endosymbiosis \u003csup\u003e26,27\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003eHere, we took advantage of the rapidly growing collection of archaeal, bacterial and eukaryotic genome sequences to infer the origins of core eukaryotic genes tracing to the LECA within a rigorous statistical framework based on evolutionary hypotheses testing using constrained phylogenetic trees. The results reveal the dominant contribution of Asgard to the origin of most functional classes of eukaryotic genes, which is compatible with a scenario of eukaryogenesis where many signature, complex features of eukaryotic cells evolved in the Asgard ancestor of eukaryotes prior to mitochondrial endosymbiosis.\u003c/p\u003e"},{"header":"Main","content":"\u003ch3\u003e\u003cstrong\u003ePhylogenomic analysis and global evolutionary associations of conserved eukaryotic proteins\u003c/strong\u003e\u003c/h3\u003e\n\u003cp\u003eTo identify associations between prokaryotic and eukaryotic protein families, separate hidden Markov model (HMM) databases for prokaryotes and eukaryotes were constructed using a custom, cascaded, sequence-to-profile clustering pipeline, implemented using mmseqs2 \u003csup\u003e28\u003c/sup\u003e, followed by a multistep data-reduction and multiple sequence alignment (MSA) procedure to generate HMM profiles using hhsuite \u003csup\u003e29\u003c/sup\u003e\u003csup\u003e32\u003c/sup\u003e (see Methods for details). Initially, a prokaryotic database of 75 million protein sequences was curated from 47,545 complete prokaryotic genomes obtained from the NCBI GenBank in November 2023 and supplemented with proteins extracted from 146 Asgard genome assemblies (Extended Data Figure 1) \u003csup\u003e6,8\u003c/sup\u003e. To avoid including genes present only within a narrow subset of species, possibly resulting from horizontal transfer from eukaryotes post-LECA, we reconstructed the \u0026ldquo;soft-core\u0026rdquo; pangenome for each of 26 curated prokaryotic taxonomic classes. These pangenomes include only those genes that are present in at least 67% of the families within each class of bacteria and archaea (see Methods) resulting in an initial database of 6.3 million nonredundant sequences. The initial eukaryotic database consisted of protein sequences from 993 species taken from EukProt v3 \u003csup\u003e30\u003c/sup\u003e, cleaned using mmseqs2 to remove likely prokaryotic contaminants and clustered to 25 million nonredundant sequences.\u003c/p\u003e\n\u003cp\u003eBoth the eukaryotic and prokaryotic were clustered, clusters were realigned using muscle5 and turned into HMM profiles using HHsuite (see Methods). The resulting eukaryotic HMM dataset was queried against the prokaryotic dataset using hhblits \u003csup\u003e29\u003c/sup\u003e to identify sets of homologous protein sequences. Each eukaryotic cluster and all its significant prokaryotic hits constituted an individual sequence set, hereinafter referred to as a Eukaryotic/Prokaryotic Orthologous Cluster (EPOC). The EPOCs constitute groups of homologous proteins from eukaryotes and prokaryotes (each EPOC contains a unique set of eukaryotic proteins, but some clusters of prokaryotic proteins can be present in multiple EPOCs) that were used for phylogenetic tree construction, annotation, and evolutionary hypothesis testing. The final EPOCs include 5.7 million prokaryotic and 1.9 million eukaryotic sequences, mapping to 90% and 8% of the respective non-redundant datasets.\u003c/p\u003e\n\u003cp\u003eTo infer the most likely prokaryotic ancestry of the eukaryotic proteins in each EPOC, rather than relying on the tree topology directly, we employed a probabilistic approach for evolutionary hypothesis testing using constraint trees. Following the construction of an initial master tree, we exhaustively sampled all arrangements of likely sister clades relative to the eukaryotic outgroup and obtained Expected Likelihood Weights (ELW) for the set of possible sister clade models, (Extended Data Figure 2) \u003csup\u003e31\u003c/sup\u003e. Given that the ELW metric is analogous to model selection confidence, here we take it to be proportional to the probability of a sampled prokaryotic clade to be the true sister group of the given eukaryotic clade among a set of competing models. For each EPOC, our analysis dynamically accounts for long branch outliers and is robust to phylogenetically non-homogenous clades (see Methods). This analysis is further capable of resolving eukaryotic paraphyly, treating each eukaryotic clade within a EPOC as a single datapoint for downstream analysis. The resulting data included 14,300 EPOCs annotated using profiles generated from KEGG Orthology Groups (KOGs) \u003csup\u003e32\u003c/sup\u003e, each with an MSA generated using muscle5 \u003csup\u003e33\u003c/sup\u003e, a maximum likelihood tree inferred using IQtree2 \u003csup\u003e34\u003c/sup\u003e and associated ELW values for all candidate prokaryotic sister phyla. The analysis of prokaryotic ancestry was performed only for those eukaryotic clades that included more than 5 distinct taxonomic labels, with at least one coming from\u003cem\u003e Amorphea\u003c/em\u003e and one from \u003cem\u003eDiaphoretickes\u003c/em\u003e, the two expansive eukaryotic clades considered to emit from either the first or the second bifurcation in the evolution of eukaryotes \u003csup\u003e35,36\u003c/sup\u003e. Thus, these clades represent genes likely mapping back to the LECA.\u003c/p\u003e\n\u003cp\u003eConsidering the global average distribution of ELW values across all EPOCs covering 4330 unique KOGs, the single greatest average ELW (aELW), here referred to as an association, was with Asgard archaea. Further associations with \u003cem\u003eCyanobacteria\u003c/em\u003e, \u003cem\u003eActinomycetota\u003c/em\u003e, \u003cem\u003eBetaproteobacteria\u003c/em\u003e and \u003cem\u003eAlphaproteobacteria\u003c/em\u003e, as well as trace association with additional bacterial classes, were detected at lower levels. However, the number of included sequences and topological diversity of the trees varied substantially across the EPOCs, with many trees showing low maximum ELW values. Excluding EPOCs with a maximum ELW \u0026lt; 0.4 yielded a robust core set of 5590 EPOCs stemming from the LECA, covering 2540 KOGs and improving the interpretability of the results. This core set covers a wide range of ubiquitous metabolic functions, information processing pathways, transporters, and permeases, as well as regulatory and housekeeping proteins. By contrast, it does not include metabolic pathways limited to individual eukaryotic supergroups and thus can be considered a rough approximation of the core gene set of the LECA. Limiting the analysis to this core subset of well assigned eukaryotic families with wide taxonomic coverage notably increased the global Asgard association which now accounted for the 62 % of the ELW across more than 6000 unique data points across at least 2500 protein families.\u003c/p\u003e\n\u003cp\u003eOur approach readily reproduced known evolutionary associations at the global functional level. Averaging ELW scores across EPOCs based on the KEGG ontology shows support for major Asgard association for, among other functional systems and pathways, the ribosome, RNA and DNA polymerases, Ras-like GTPases, Ubiquitin-mediated protein degradation, the proteasome, and large parts of the core metabolic network. In contrast, prominent alphaproteobacterial associations included proteins involved in oxidative phosphorylation, glutathione metabolism and Fe-S clusters biogenesis. Together, the Asgard and alphaproteobacterial associations amounted to an aELW of 0.55\u0026nbsp;+\u0026nbsp;0.06\u0026nbsp;=\u0026nbsp;0.61 across all of Metabolism and to 0.77 across Genetic Information Processing. Thus, we observed a far stronger overall association between Asgards and eukaryotes across diverse biological functions and pathways than previously described \u003csup\u003e5,6,8\u003c/sup\u003e although a consistent association between eukaryotic core metabolism and diverse bacterial phyla is still present.\u003c/p\u003e\n\u003ch3\u003e\u003cstrong\u003eBroad, dominant Asgard contributions to eukaryogenesis\u003c/strong\u003e\u003c/h3\u003e\n\u003cp\u003eIn accord with the key contribution of Asgard archaea to eukaryogenesis, we observed associations of Asgard proteins with a wide array of cellular functions. In previous work, the strongest Asgards traces have been noted across the information processing systems, with unambiguous associations with DNA replication, core transcription, RNA processing as well as translation and protein trafficking \u003csup\u003e2,6,8\u003c/sup\u003e. Here, we consistently observed strong Asgard associations for genome replication and transcription and further detected pronounced Asgard traces for nucleotide excision repair, mismatch repair and homologous recombination. Additional well-known associations, such as ribosomal proteins, were extended to include translation factors, components of the co-translational membrane insertion machinery, protein targeting and aminoacyl-tRNA biosynthesis (Extended Data Figure 3, Extended Data File 1). Thus, all groups of core eukaryotic proteins involved in information processing appear to be almost exclusively of Asgard descent.\u003c/p\u003e\n\u003cp\u003eWe detected additional Asgard associations extending far beyond the information processing systems, including prominent contributions to the machinery involved in nucleocytoplasmic transport as well as downstream protein sorting, glycosylation and targeting. In particular, central components of the ER associated, N-linked glycan biosynthesis and transfer, including both cytoplasmic and lumenal monoglycosyltransferases, as well as the core of the oligosaccharyltransferase complex (OSTC) are strongly associated with Asgard (Extended Data Figure 3). Notably, enzymes associated with glycosylation maturation in the Golgi complex did not show strong Asgard or other prokaryotic associations in our analysis, possibly, due to extensive diversification of domain architecture in eukaryotes. The Asgard connections of the eukaryotic glycosylation machinery further included the synthesis of GPI-anchors, which post-translationally tether targeted proteins to the membrane \u003csup\u003e37\u003c/sup\u003e, here detected as unambiguously Asgard-derived. We also detected an Asgard origin of the 7-subunit (UDP-GlcNAc)-transferring (GPI-GnT)-monoglycosyltransferase complex responsible for initiating GPI-anchor synthesis, components required for the maturation of the GPI-anchor, as well as the transamidase complex and factors responsible for protein transfer onto the mature GPI-anchor (Extended Data Figure 3).\u003c/p\u003e\n\u003cp\u003eOf major importance to eukaryogenesis is the provenance of the pathways for the biosynthesis of bacterial-type lipids, given that (at least) all binary archaea-bacteria symbiogenesis scenarios require a transition from archaeal to bacterial lipids in the membranes \u003csup\u003e9\u003c/sup\u003e. Although strong Asgard associations were observed for large parts of the overall metabolic network, we observed a high degree of mosaicism in the pathways for fatty acid synthesis and decay. The global aELW values favored Asgard origin for these pathways, but there were also notable associations with \u003cem\u003eActinomycetota\u003c/em\u003e. However, most KOGs within these pathways are represented by multiple EPOCs with conflicting assignments potentially obscuring any consistent signal (Extended Data Figure 3). By contrast, less perplexity was observed within the adjoining ER localized pathways for sphingolipid metabolism, a broad class of derived plasma membrane lipids in eukaryotes. Previous studies have highlighted possible convergent origins of this pathway in bacteria and eukaryotes \u003csup\u003e38,39\u003c/sup\u003e, but here we detected broad associations with Asgard. Of further note is the ER-associated isoprenoid biosynthesis pathway, here also found to be strongly associated with Asgard. In eukaryotes, isoprenoids form the precursor units for sterols, carotenoids and terpenoids, synthesized in the ER lumen via either the mevalonate pathway or the MEP/DOXP pathway. In Archaea, isoprenoids are the precursors for the ether-linked membrane lipids \u003csup\u003e40\u003c/sup\u003e. Here we found the mevalonate pathway, from Acetyl-COA to mevalonate and further to Farnesyl and Geranyl diphosphate, to be strongly Asgard-associated (Extended Data Figure 3), with the key enzymes hydroxymethylglutaryl-CoA synthase (HMGCS), mevalonate kinase (MVK), phosphomevalonate kinase (PMVK), mevalonate diphosphate decarboxylase (MVD) being clearly Asgard-derived.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eIn conclusion, we detected Asgard associations across a wide range of cellular functions and metabolic pathways while noting a distinctly weaker Asgard signal for pathways involved in bacterial lipid biosynthesis, suggesting a complex evolutionary history.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003ch3\u003e\u003cstrong\u003eLimited and highly specific alphaproteobacterial contributions\u003c/strong\u003e\u003c/h3\u003e\n\u003cp\u003eIn line with the central role of mitochondria in eukaryotic energy metabolism, we primarily observed associations between \u003cem\u003eAlphaproteobacteria\u003c/em\u003e and mitochondrially localized metabolic pathways. As expected, apart from the components of the mitochondrial translation system, the most prominent alphaproteobacterial associations were evident for complexes involved in oxidative phosphorylation and the associated ubiquinone synthesis (Extended Data Figure 3). Outside these central energy-transforming functions, we only detected sparse contributions from core alphaproteobacterial genes. One such prominent association was the pathway for iron sulfur cluster (ISC) biogenesis. As previously reported \u003csup\u003e41\u003c/sup\u003e, the ISC assembly machinery is of alphaproteobacterial origin, and in accord with these observations, the 4Fe-4S ISCA platforms as well as IBA57 and Fe-S cluster binding ferredoxin-1 and 2, were found to be strongly associated with \u003cem\u003eAlphaproteobacteria \u003c/em\u003e(Figure 3, Extended Data Figure 3). However, the 2Fe-2S precursor scaffold ISCU showed mosaic associations, with a minor but clearly detectable Asgard contribution. Notably, the cysteine desulfurase NFS1 and the upstream pathways for the biosynthesis of sulfur-containing amino acids, cysteine, and methionine, were strongly Asgard- associated. The ISC biosynthesis is intimately linked to the general redox homeostasis and core sulfur metabolism via glutathione, directly coordinating Fe-S clusters during synthesis and transport \u003csup\u003e42\u003c/sup\u003e. In line with this role in Fe-S coordination, although some Asgard association persisted, we observed specific associations of glutathione metabolism with alphaproteobacteria, including glutathione hydrolase, dehydrogenase and reductase, as well as the family of glutathione transferases, GST, GSTP and GSTK1 (Extended Data Figure 3). Outside the mitochondria, ISC insertion depends on the cytosolic targeting complex CIA, which consists of CIAO1, CIA2B and MMS19 \u003csup\u003e43\u003c/sup\u003e. For the CIA components CIA1 and CIA2B, we observed clear association with Asgard whereas MMS19 was not detected in our data. Taken together, these observations indicate that the contributions of alphaproteobacteria to the gene set of the LECA are functionally specific and apparently limited in scope, clearly centered around mitochondria-related functions.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ePaucity of functionally consistent contributions from other bacteria\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAlthough our analysis greatly expanded the Asgard contributions to eukaryogenesis, while also revealing a limited but prominent and functionally consistent alphaproteobacterial association, contributions from diverse other bacteria were consistently detected. For some biological functions, this diverse bacterial component accounted for the majority of the aELW, and roughly one third of the analyzed KOGs (680 of 2540), and EPOCs (1810 of 5590) were associated neither with known ancestors of endosymbionts, \u003cem\u003eAlphaproteobacteria\u003c/em\u003e and \u003cem\u003eCyanobacteria\u003c/em\u003e, nor with Asgard. However, in a sharp contrast to Asgard associations including information processing, protein glycosylation and trafficking, and other functions as discussed above, or oxidative phosphorylation and sulfur metabolism for \u003cem\u003eAlphaproteobacteria\u003c/em\u003e, EPOCs associated with diverse other bacteria showed few if any coherent functional trends.\u003c/p\u003e\n\u003cp\u003eConsidering the diverse set of bacteria, and all possible KEGG maps and modules, only \u003cem\u003eAlphaproteobacteria\u003c/em\u003e were associated with pathway including more than 20 EPOCs and with a greater aELW than Asgard (Glutathione metabolism, Figure 4). For all other analyzed functional classes of eukaryotic genes, bacterial associations were weaker than the associations with Asgard (Figure 4). The second most individually prominent bacterial contribution was from \u003cem\u003eMyxococcota\u003c/em\u003e, of the former deltaproteobacterial clade. Although globally weaker than Asgard associations, \u003cem\u003eMyxococcota\u003c/em\u003e showed consistent associations with nicotinate and nucleotide synthesis, including both purine and pyrimidine synthesis, as well as nucleoside sugar metabolism. \u003cem\u003eMyxococcota \u003c/em\u003ewere unique in this regard as most bacteria showed diffuse associations across sugar and fatty acid metabolism, and/or diverse transporters. The nucleotide-related associations with \u003cem\u003eMyxococcota\u003c/em\u003e were primarily limited to phosphatases and phosphoribosyltransferases acting on nucleotide sugars including 5 and 3\u0026rsquo; nucleotidases, and the respective EPOCs showed little to no competing Asgard association. While noteworthy, these associations were limited to a few unique KOGs whereas all other associations with \u003cem\u003eMyxococcota \u003c/em\u003eremained scattered across various pathways (Extended Data Figure 5).\u003c/p\u003e\n\u003cp\u003eIn addition to investigating metabolic pathways for global associations, we directly examined those individual EPOCs that were highly likely to be derived from diverse bacteria. For this analysis, we considered a stricter subset of the core EPOCs, requiring a eukaryotic outgroup containing at least 15 taxonomic clades and prokaryotic sister taxa with at least 20 sequences and ELW \u0026gt; 0.7. Only 16 of the 127 unique KOGs meeting these criteria were found to be associated with diverse bacterial lineages, mostly, \u003cem\u003eActinomycetota\u003c/em\u003e, FCB group and \u003cem\u003eBetaproteobacteria\u003c/em\u003e, with minor contributions from \u003cem\u003eMycoplasmatota\u003c/em\u003e and \u003cem\u003eCampylobacteriota \u003c/em\u003e(Extended Data Figure 6). The remaining 111 were Asgard-derived. The 16 prominent bacterial KOGs covered a wide range of cellular functions from MFS transporters to lipases to components of core sugar metabolism and cardiolipin synthesis once again with no noticeable trends. Taken together, although many associations were individually significant, no pathways appeared significantly enriched in associations with any single bacterial taxon, and we identified virtually no consistent trends among the bacterial associations. Instead, we interpret these diffuse associations as indications of highly specific contributions of limited functional scope from diverse bacteria other than the known endosymbiotic partners.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eRelative contributions of evolution pre- and post-LECA\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo compare the ancestral stem lengths in phylogenetic trees of core eukaryotic genes of different inferred origins we employed the methodology originally implemented by Pittis and Gabaldon \u003csup\u003e26\u003c/sup\u003e. Briefly, we define the raw stem length as the distance from the LECA node to the shared Last Common Ancestor (LCA) node of \u003cem\u003eEukarya\u003c/em\u003e and its most likely prokaryotic sister phyla. To account for differences in evolutionary rates, we divided this raw stem length by the median of eukaryotic branch lengths, measured from the LECA to each leaf (Figure 5C). Considering that the astronomic time post-LECA is the same for all genes, if the tempo and mode of pre- and post-LECA evolution were the same, the normalized stem length is proportional to the time elapsed since the divergence of the gene from its prokaryotic donor to LECA, that is, reflects the timing of acquisition. Using this approach, Pittis and Gabaldon found that proteins of alphaproteobacterial descent had significantly shorter normalized stem lengths than proteins of archaeal descent, consistent with a mitochondria-late scenario for eukaryogenesis \u003csup\u003e26\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003eAcross the full set of eukaryotic stem lengths for our dataset (5,850 stems), we observed a wide distribution with a sharp maximum close to 0.05, highly reminiscent of the previous findings eukaryogenesis \u003csup\u003e26\u003c/sup\u003e (Figure 5B). However, in our analysis, the stems for genes of alphaproteobacterial origin were significantly (p\u0026nbsp;\u0026asymp;\u0026nbsp;9.9x10\u003csup\u003e‑6\u003c/sup\u003e; Figure 5) longer than those of Asgard origin, and longer than those of other major bacterial contributors as well (Figure 5B). Comparison of the stem length distributions across functional classes of genes (Figure 5A) suggested an explanation for these observations. The shortest stems belong to the Genetic Information Processing functional category, belying our expectations of these genes being the longest-residing genes in the nascent eukaryotic lineage. We suggest that another major determinant of the relative stem length of a gene is the amount of adaptive evolution post-acquisition, which is necessary to adjust the newly acquired gene to the alien intracellular molecular environment. Thus, genes inherited from the Asgard ancestor, in particular, those involved in information processing, were pre-adapted to the cellular environment of the evolving protoeukaryotes, whereas genes acquired from radically different bacterial sources had to substantially adapt post-acquisition, increasing the apparent lengths of their pre-LECA stems. A case in point is the set of oxidative phosphorylation components, which were apparently acquired during the mitochondrial symbiogenesis, and thus, simultaneously, from the same donor. The distribution of their stem lengths (Figure 5A) was as broad as that for proteins of other functional classes (Figure 5A and Extended Data Figure 7), demonstrating that stem lengths are mostly determined by factors other than the acquisition time. Thus, our stem length analysis failed to provide unequivocal resolution of the temporal order of the prokaryotic contributions to eukaryogenesis. Nevertheless, the results appear to be compatible with the capture of the alphaproteobacterial endosymbiont by a host that was already on the path to eukaryotic-like complexification, along with capture of genes from various bacteria at different stages of eukaryogenesis.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eReconstruction of eukaryogenesis is a moving target as strikingly demonstrated by the discovery and ongoing exploration of Asgard archaea which revolutionized the field \u003csup\u003e4-8\u003c/sup\u003e. Nevertheless, recent progress in prokaryote as well as eukaryote genome sequencing, along with the advances of metagenomics, created unprecedented opportunities for phylogenomics studies aimed at understanding the origins of eukaryotes. In this work, we took advantage of this expanded genome collection, coupling it with maximally sensitive HMM profile-profile searches, and performed a comprehensive phylogenetic analysis of core eukaryotic genes tentatively mapped to the LECA, within a statistical framework focused on testing evolutionary hypotheses using constraint trees.\u003c/p\u003e\n\u003cp\u003eThe principal outcome of this analysis is that a substantial majority of the LECA genes, for which the origin could be inferred with confidence, came from Asgard archaea. Particularly notable is the functional diversity of the apparent Asgard contribution to eukaryogenesis. Far from being limited to information processing and certain cellular processes, such as membrane remodeling, a strong Asgard trace was detected for most core functional systems and biochemical pathways of eukaryotes. In very few protein families did the detected bacterial contribution exceed that of Asgard archaea. The most conspicuous of the bacterial contributions, not unexpectedly, comes from \u003cem\u003eAlphaproteobacteria\u003c/em\u003e, the ancestors of the mitochondria, and accounts for the core components of the electron transport chain complexes and iron-sulfur cluster metabolism, along with the mitochondrial translation system. However, although functionally relevant and consistent, the alphaproteobacterial contribution is relatively small and not comparable in scale to that of Asgard. The contributions of other bacterial phyla, although cumulatively substantial, fail to show any individually consistent trends, with the exception of one or two scattered metabolic pathways.\u003c/p\u003e\n\u003cp\u003eMany previous studies, especially early ones, demonstrated the apparent chimeric origins of the core eukaryotic gene set, with the bacterial contribution quantitatively exceeding the archaeal one, and the latter being largely limited to information processing \u003csup\u003e17-19,44\u003c/sup\u003e. The discovery and exploration of Asgard archaea partially changed this notion by demonstrating that many genes involved in various cellular processes and systems, such as membrane remodeling and cytoskeleton, were of apparent Asgard origin \u003csup\u003e5,6,8\u003c/sup\u003e. The present work substantially extends the dominance of Asgard in the ancestry of the LECA genes, by revealing major Asgard contributions to nearly all functional systems and pathways of the eukaryotic cell. These findings imply a model of eukaryogenesis in which the ancestral Asgard archaeon already possessed many features characteristic of eukaryotic cells including cytoskeleton and the endomembrane system. These complex ancestral Asgard archaea might have led a predatory lifestyle conducive to engulfment of bacteria and serial capture of bacterial genes. Under this scenario, mitochondrial endosymbiosis, also perhaps facilitated by predation, while antedating the LECA, was a relatively late event that made a limited, even if functionally crucial, contribution to the gene composition of the emerging eukaryote (Figure 6). Other bacterial contributions appear to be piecemeal, coming from a limited capture of genes from diverse bacteria, likely, both before and after the mitochondrial endosymbiosis; there is no indication of another symbiotic event contributing to eukaryogenesis. This conclusion complements and extends a recent study that suggested emergence of some features of eukaryotic complexity prior to the mitochondrial endosymbiosis based on estimated timing of ancient gene duplications \u003csup\u003e45\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003eHowever, how eukaryotes acquired the bacterial-type membranes, remains an enigma. Whereas the eukaryotic pathways for isoprenoid synthesis are undoubtedly Asgard-derived, the ancestry of the enzymes involved in membrane lipid biosynthesis appears complex and remains unresolved in the present analysis, with multiple bacterial ancestors appearing likely. Our analysis of these pathways is likely confounded by the presence of two sets of lipid biosynthetic pathways, one mitochondrial and the other ER-associated. Nevertheless, the Asgard affinity of adjoining systems and pathways, such as the synthesis of GPI anchors, sphingolipids, or isoprenoids, shows that a substantial component of the lipid biosynthesis and modification capacity of eukaryotes persisted from the Asgard ancestor.\u003c/p\u003e\n\u003cp\u003eIf an Asgard archaeal cell was the principal scaffold on which the eukaryotic cell evolved, replacement of the archaeal membrane with a bacterial one must have occurred at an early stage of eukaryogenesis, likely, through a mixed membrane stage. Viable bacteria with a mixed bacterial-archaeal membrane have been experimentally engineered albeit resulting in a substantial fitness loss \u003csup\u003e46,47\u003c/sup\u003e. Membrane replacement might have even predated mitochondrial endosymbiosis, following early capture of the necessary bacterial enzymes by the Asgard archaeon (Figure 6).\u003c/p\u003e\n\u003cp\u003eThe conclusions of this work are subject to caveats stemming, above all, from the biased and still limited sampling of sequenced archaeal and bacterial genomes. In particular, almost all currently available genomes of Asgard archaea are incomplete metagenomic assemblies, with the exception of only three closed circular genomes \u003csup\u003e24,48,49\u003c/sup\u003e. Therefore, even though our current analysis points to a dominant contribution of Asgard to the eukaryotic gene core, it is nevertheless almost certainly an underestimate. The same pertains to \u003cem\u003eAlphaproteobacteria \u003c/em\u003eboth because of the sequencing bias whereby many of the sequenced alphaproteobacterial genomes come from symbionts and parasites with reduced gene complements, and because the ancestor of the mitochondria apparently belongs outside of the currently known diversity of \u003cem\u003eAlphaproteobacteria\u003c/em\u003e\u003csup\u003e50\u003c/sup\u003e. Some of the other bacterial phyla are even more severely under-sampled, potentially, leading to underestimates of their contributions to the evolution of eukaryotes. For example, it is difficult to rule out that we miss a major signal from \u003cem\u003eMyxococcota \u003c/em\u003e(formerly Deltaproteobacteria), that have been proposed as one of the partners in the syntrophy scenario of eukaryogenesis \u003csup\u003e15,22\u003c/sup\u003e, given the limited genomic information on these bacterial phyla. The expanding sampling of prokaryotic genome diversity might lead to a substantial modification or even complete revision of the model of eukaryogenesis proposed here. It nevertheless appears likely that the dominant contribution of Asgard archaea, the main finding of this work, is here to stay.\u003c/p\u003e"},{"header":"Methods","content":"\u003ch3\u003eDatabase curation\u003c/h3\u003e\n\u003cp\u003eFor prokaryotes 75 million sequences were curated from 47,545 completely sequenced prokaryotic genomes obtained from the NCBI GenBank (https://ftp.ncbi.nlm.nih.gov/genomes/) in November 2023 and supplemented with 441,150 proteins sequences from 146 assembled Asgard genomes\u0026nbsp;\u003csup\u003e6,8\u003c/sup\u003e, selected to represent all clades in these two publications (where available, protein sequences were directly taken from GenBank annotations; otherwise, produced by Prodigal v2.6.3\u0026nbsp;\u003csup\u003e51\u003c/sup\u003e, trained on the set of 12 complete or chromosome-level assembled Asgard genomes. To only include those sequences which are widely present within prokaryotic families, and to minimize the possibility of post- LECA horizontal transfer from eukaryotes, soft-core pangenomes were constructed for each of our 26 curated taxonomic groups, mostly conforming to the NCBI taxonomy rank \u0026ldquo;class\u0026rdquo; (see below). To construct the soft-core pangenomes, all sequences from each taxonomic group were clustered individually using mmseqs2 \u0026ldquo;mmseqs cluster -s 7 -c 0.9 --cov-mode 0\u0026rdquo;\u0026nbsp;\u003csup\u003e28\u003c/sup\u003e. All clusters which did not contain sequences from at least 67% of all bacterial families within a prokaryotic class were rejected from the database.\u003c/p\u003e\n\u003cp\u003eThe eukaryotic database was constructed starting from a curated set of 72 genomes available on the NCBI GenBank. In order to provide a more accurate representation of eukaryotic diversity, this dataset was supplemented with EukProt v3\u0026nbsp;\u003csup\u003e30\u003c/sup\u003e, a curated database aiming to provide a sparse representation of eukaryotic biology, to create Euk72Ep containing ~30.3 million sequences. As sequences included in EukProt v3 come from a diverse range of sources, some known to contain contaminant prokaryotic sequences, the database was screened for prokaryotic contamination. This screen was carried out by first constructing candidate prokaryotic and eukaryotic HMM databases as described below and querying both databases against the eukaryotic sequence database using hhblits\u0026nbsp;\u003csup\u003e29\u003c/sup\u003e. All eukaryotic sequences with a top hit alignment score from a prokaryotic HMM were removed from the database prior to initial clustering.\u003c/p\u003e\n\u003ch3\u003e\u003cstrong\u003eCuration of taxonomic labels\u003c/strong\u003e\u003c/h3\u003e\n\u003cp\u003eIn order to provide a compromise between taxonomic specificity and accuracy of clade assignment, we curated a set of taxonomic labels for both Euk72Ep and Prok2311As to act as representatives. To construct this set, we started with manually assigning all species in EukProt to their closest relative in the NCBI taxonomy. Then, each species from Prok2311As and Euk72Ep was assigned a taxonomic label corresponding to their \u0026ldquo;Class\u0026rdquo; rank (for example, \u003cem\u003eAlphaproteobacteria\u003c/em\u003e, \u003cem\u003eThermococci\u003c/em\u003e, \u003cem\u003eMammalia\u003c/em\u003e) based on the NCBI Taxonomy of November 2023 using tools from the ete3 3.1.3. toolkit\u0026nbsp;\u003csup\u003e52\u003c/sup\u003e as well as custom scripts. For species without a corresponding \u0026ldquo;Class\u0026rdquo; rank, the closest relevant rank was manually assigned. This procedure yielded an initial list of 317 unique taxonomic class identifiers. As the rank \u0026ldquo;Class\u0026rdquo; does not evenly partition biological diversity, identifiers were further manually mapped to a set of 45 eukaryotic and 26 prokaryotic taxonomic superclasses, each present in the NCBI taxonomy (for example, \u003cem\u003eMetazoa\u003c/em\u003e, \u003cem\u003eStreptophyta\u003c/em\u003e, TACK archaea). Due to the partially unresolved taxonomic relationship within Asgard archaea, all candidate Asgard sequences from Prok2311 or from Refs. 19,20 were assigned the class label \u0026ldquo;Asgard\u0026rdquo;. Following the recent reclassification of Deltaproteobacteria into \u003cem\u003eMyxococcota\u003c/em\u003e, \u003cem\u003eDesulfobacterota\u003c/em\u003e, \u003cem\u003eBdellovibrionota\u003c/em\u003e and SAR324 species which could not be mapped into either of these clades using existing taxonomic annotation were classified as Deltaproteobacteria (0.025% of total data).\u003c/p\u003e\n\u003ch3\u003e\u003cstrong\u003eSequence clustering and profile database generation\u003c/strong\u003e\u003c/h3\u003e\n\u003cp\u003eTo transform Euk72Ep and Prok2311As into profile databases suitable for sensitive HMM-HMM searches, we implemented an unsupervised, cascaded sequence-profile clustering pipeline using the tools available in the mmseqs2 software suite\u0026nbsp;\u003csup\u003e28\u003c/sup\u003e. In brief, each sequence database was initially clustered at 90% sequence identity and 80% pairwise coverage using \u0026ldquo;mmseqs linclust\u0026rdquo; and cluster representatives were chosen using \u0026ldquo;mmseqs result2repseq\u0026rdquo;. This procedure provides a non-redundant set of 6.3 million prokaryotic and 25.1 million eukaryotic sequences for further analysis. The non-redundant sequences are further collapsed into a set of initial profiles and consensus sequence pairs using \u0026ldquo;mmseqs cluster\u0026rdquo;, \u0026ldquo;mmseqs result2profile\u0026rdquo; and \u0026ldquo;mmseqs profile2consensus\u0026rdquo;. All consensus sequences are queried against their profiles using \u0026ldquo;mmseqs search\u0026rdquo; and results clustered using \u0026ldquo;mmseqs clust\u0026rdquo;. Clusters are mapped to their original non-redundant sequences using \u0026ldquo;mmseqs mergeclusters\u0026rdquo; and profiles and consensus sequences are constructed once again. In order to grow clusters based on their sequence-profile alignments, this procedure of \u0026ldquo;mmseqs profile2consensus - mmseqs search - mmseqs clust - mmseqs result2profile\u0026rdquo; is iterated until cluster sizes converge. An 80% pairwise coverage is maintained throughout all steps of the cascaded clustering protocol. The final clustered databases contain 14.1 million and 91.000 clusters for Euk72Ep and Prok2311As, respectively (Extended Data Figure 1).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eTo transform the unsupervised clusters into a set of HMM profiles, we employed a mixed MSA construction strategy coupled with automatic data reduction to automatically create accurate alignments for a wide diversity of cluster sizes and sequence diversities. To construct HMMs, all non-singleton sequence clusters are first aligned using FAMSA 2.0.1.\u0026nbsp;\u003csup\u003e53\u003c/sup\u003e to produce an initial alignment. Using the initial alignment, we then reduce the sequence space by first creating an approximate tree using FastTree 2.1.1. using \u0026ldquo;FastTree -gamma\u0026rdquo;\u0026nbsp;\u003csup\u003e54\u003c/sup\u003e and iteratively removing the closest pairwise leaves using ete3\u0026nbsp;\u003csup\u003e52\u003c/sup\u003e, keeping the leaf closest to the root, until the desired sequence set size is reached. For profile generation, we retain no more than 300 sequences for both Prok2311As and Euk72Ep. This reduced sequence set is then realigned with muscle5 5.1.linux64\u0026nbsp;\u003csup\u003e33\u003c/sup\u003e using \u0026ldquo;muscle -diversified\u0026rdquo; with 5 replicates, and the maximum column confidence alignment is extracted. This alignment is then trimmed to keep columns with more than 0.2 bits of Shannon information. This process of data reduction/alignment/trimming is referred to as \u0026ldquo;prune-and-align\u0026rdquo; and employed later for downstream analysis. Trimmed alignments are turned into an HMM database using HHsuite with \u0026ldquo;hhconsensus -M 50\u0026rdquo;, hhmake and \u0026ldquo;cstranslate -f -x 0.3 -c 4 -I a3m\u0026rdquo;. The final HHsuite databases contained 26.000 and 1.6 million profiles for prokaryotes and eukaryotes, respectively, see (Extended Data Figure 1).\u0026nbsp;\u003c/p\u003e\n\u003ch3\u003e\u003cstrong\u003eSearch for prokaryotic homologs of eukaryotic proteins\u003c/strong\u003e\u003c/h3\u003e\n\u003cp\u003eGiven that clade-specific protein families were not of direct relevance to this study, we reduced the search space by excluding HMM profiles from eukaryotic clusters with less than 10 sequences and a lowest common ancestor with a taxonomic rank below \u0026ldquo;Superkingdom\u0026rdquo; as calculated by \u0026ldquo;mmseqs lca\u0026rdquo;. This requires a sequence cluster to contain at least two sequences from different kingdoms, as defined by the NCBI taxonomy (ex. \u003cem\u003eAmoebozoa\u003c/em\u003e and \u003cem\u003eHaptista\u003c/em\u003e, or \u003cem\u003eViridiplanta\u0026nbsp;\u003c/em\u003eand \u003cem\u003eOpistokonta\u003c/em\u003e), retaining 142.000 profiles of conserved eukaryotic proteins. To associate eukaryotic sequence clusters with homologous prokaryotic clusters, we queried the Euk72Ep HMM database against all Prok2311As HMM profiles using a single iteration of HHBlits 3.3.0 as \u0026ldquo;HHblits -n 1 -p 80\u0026rdquo; and retaining the hits with at least 80% HHblits probability and 80% pairwise profile length. These filters resulted in a total of 20,700 accepted eukaryotic query profiles targeting a total of 8,300 unique prokaryotic profiles. These clusters contain a total of 5.7 million and 1.9 million sequences.\u0026nbsp;\u003c/p\u003e\n\u003ch3\u003e\u003cstrong\u003eEPOC multiple sequence alignment construction\u003c/strong\u003e\u003c/h3\u003e\n\u003cp\u003eUnless otherwise stated, all subsequent steps were carried out using custom python 3.9 scripts, with taxonomy and tree parsing carried out using ete3\u0026nbsp;\u003csup\u003e52\u003c/sup\u003e. All sequences from a eukaryotic cluster with at least 10 sequences and a lowest common ancestor at the rank of \u0026ldquo;Superkingdom\u0026rdquo; (NCBI taxonomy), together with all members of the homologous prokaryotic clusters identified in the HMM-HMM search with at least 80% probability and pairwise sequence coverage, forms a EPOC. As EPOCs varied in size by more than 5 orders of magnitude (10 to 100.000 sequences), robust data subsampling is necessary for accurate downstream MSA generation and tree construction. We employed a variant of the \u0026ldquo;prune-and-align\u0026rdquo; strategy described above for profile generation taking into account the taxonomic distribution of the constituent sequences. Rather than iteratively pruning the closest leaves pairwise, we first label each leaf with its corresponding taxonomic label and only prune monophyletic groups, maintaining a count of all pruned leaves per clade. If the procedure fails to reach the target leaf number before converging, all leaves are ordered by the number of pruned relatives and leaves with the lowest number of relatives deleted (retaining the largest monophyletic groups) until the target size is reached. The result approximates a maximum diversity representation, preferentially pruning isolated singleton clades. Using this modified protocol, eukaryotic and prokaryotic sequences are cropped to a maximum of 30 and 70 sequences, respectively. Each EPOC therefore consists of a maximum of 100 representative sequences. All EPOCs are aligned using \u0026ldquo;muscle -diversified\u0026rdquo; with 5 replicates, and the maximum confidence alignment is extracted using the --maxcc argument. This alignment is then trimmed to keep columns with more than 0.15 bits of Shannon information content for tree reconstruction.\u0026nbsp;\u003c/p\u003e\n\u003ch3\u003e\u003cstrong\u003eEPOC tree construction and processing\u003c/strong\u003e\u003c/h3\u003e\n\u003cp\u003eFrom the pruned and trimmed alignment, we created a maximum likelihood tree using IQtree2 2.3.5. \u0026ldquo;IQtree2 -B 1000 -bnni -m MFP -mset LG,Q.pfam --cmin 4 --cmax 12\u0026rdquo; with model parameters estimated by model finder plus\u0026nbsp;\u003csup\u003e34\u003c/sup\u003e. Candidate models and rate category search ranges were selected based on a test set of 200 randomly chosen EPOCs from the filtered core set analyzed with full model parameter evaluation by Model finder\u0026nbsp;\u003csup\u003e34\u003c/sup\u003e and 10 tree replicates. The LG and Q.pfam models were chosen as they consistently produced the highest ranking log-likelihoods. Rate categories were assigned based on the highest and lowest amount found among the top 10 average scoring models from the same set.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eAs the master trees are built from sequences obtained through sensitive cascaded sequence-profile clustering we would in rare cases observe long-stemmed clade outliers, in addition to more common long stem leaf outliers. In order to assess and remove such erroneous data, we first estimated the log-normal distribution of all stem lengths using scipy.stats 1.11.1 and excluded any stems outside the 0-99.5 % Probability Point Function interval. This simultaneously removes long terminal leaf branches as well as highly diverged clades. In rare cases where this pruning would remove either the entire eukaryotic clade or more than 30 % of all leaves, the EPOC was discarded (387 EPOCs). The resulting trees were then re-rooted using a weighted midpoint approach so that the sum of all branch stem lengths on either side of the root was equal.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eTo assign taxonomic clades, all leaf nodes were assigned taxonomic labels from our curated set and monophyletic clades were collapsed, taking as the representative the lowest branching leaf. We found that assigning clades purely based on a monophyletic definition of taxonomic purity severely hampered the analysis as leaves can be taxonomically incongruous with their neighbors. This effect most likely stems from cases of locally erroneous tree topology or local HGT, but severely complicates the downstream analysis. We therefore adopted the notion of a \u0026ldquo;soft LCA\u0026rdquo; representing the root of a close-to-monophyletic clade. To greedily rank all the best \u0026ldquo;soft LCAs\u0026rdquo;, for each taxonomic label, the tree was traversed from each monophyletic clade root to the tree root, calculating a score for each internal node:\u003c/p\u003e\n\u003cp\u003e(number of leaves with label X in clade / clade size) * (number of leaves with label X in clade / total number of label X)\u003c/p\u003e\n\u003cp\u003eThis metric balances taxonomic \u0026ldquo;purity\u0026rdquo; and \u0026ldquo;scope\u0026rdquo; for each possible clade of label X. All such nodes are then ordered by score. This allows for paraphyly across all taxonomic labels. To avoid overinterpreting small clades, we only consider clades to be valid if they represent at least 3 sequences with clade purity \u0026gt;\u0026nbsp;0.8 for prokaryotes, and at least 5 sequences with purity \u0026gt; 0.8 for eukaryotes. Trees that fail to identify either any eukaryotic or any prokaryotic clades under these constraints, primarily derived from small EPOCs, are discarded (411 EPOCs). Trees with more than 3 valid eukaryotic clades were considered of high paraphyly and likewise discarded (781 EPOCs).\u003c/p\u003e\n\u003ch3\u003e\u003cstrong\u003eEvolutionary hypothesis testing using constraint trees\u003c/strong\u003e\u003c/h3\u003e\n\u003cp\u003eIn order to assess the relative probabilities of each sampled prokaryotic \u0026ldquo;soft LCA\u0026apos;\u0026apos; to represent the eukaryotic sister clade, we employed evolutionary hypothesis testing using constraint trees as implemented in IQtree2\u0026nbsp;\u003csup\u003e34\u003c/sup\u003e. For each EPOC and for each detected eukaryotic clade, as defined by its corresponding \u0026ldquo;soft LCA\u0026rdquo;, we generated a set of constraint trees, one for each possible prokaryotic sister. Each constraint tree had three defined clades, 1) the eukaryotic group as defined by all sequences below the eukaryotic LCA, 2) the prokaryotic sister group and 3) all other prokaryotic leaves. We then construct a set of local trees from slices of the original MSA using IQtree2, forcing the constraint tree topology using \u0026ldquo;iqtree2 -g\u0026rdquo;, using the same evolutionary model as the original, unconstrained \u0026ldquo;master tree\u0026rdquo;. The set of constrained trees is then ranked using \u0026ldquo;iqtree2 -z\u0026rdquo; calculating the relative model assignment confidence for each of the constrained topologies and resulting test metrics are saved for downstream analysis. As the number of trees to evaluate scales by the number of eukaryotic clades multiplied by the number of prokaryotic clades, we limit the sampling to a maximum of 3 clades per tree for eukaryotes, discarding rare cases of high eukaryotic paraphyly, and only consider the 12 closest prokaryotic clades to each individual eukaryotic clade, based on their topological distance (Extended Data Figure 2). The procedure results in an Expected Likelihood Weight (ELW)\u003csup\u003e31\u003c/sup\u003e for each of the evaluated model trees, here taken as being analogous to model selection confidence that an evaluated prokaryotic clade is the most likely sister to a eukaryotic clade.\u0026nbsp;\u003c/p\u003e\n\u003ch3\u003e\u003cstrong\u003eEPOC annotation\u003c/strong\u003e\u003c/h3\u003e\n\u003cp\u003eTo functionally annotate protein families within EPOCs, we generated profiles based on protein sequences in KEGG release 110\u0026nbsp;\u003csup\u003e32\u003c/sup\u003e. For each Kegg Orthologous Group (KOG), we extracted Kegg metadata including KEGG pathways and BRITE labels via KEGG\u0026rsquo;s API. At the time of parsing (January 2024), KEGG contained 26,695 KOGs. Due to KEGG API constraints, protein sequences in KEGG were extracted from Uniprot, first by mapping KEGG proteins to Uniprot IDs via Uniprot\u0026rsquo;s ID mapping service (uniprot.org/id-mapping), then by extracting the sequences for these Uniprot IDs corresponding to each KOG\u0026nbsp;\u003csup\u003e55\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003eFor each KOG, we generated HHM profiles using the aforementioned prune-and-align pipeline. To increase specificity of labeling, we partitioned those KOGs which encompass both prokaryotic and eukaryotic sequences into separate taxonomic groups and sequence alignments were generated separately for prokaryotes and eukaryotes. Each eukaryotic HMM profile forming the basis for a EPOC was queried against our KEGG profile database using \u0026ldquo;HHblits -n 3 -p 80\u0026rdquo;. Results were subsequently filtered to a pairwise profile coverage of 0.5 using custom scripts. As a result, 13,100 of the 14,300 EPOCs could be annotated by at least one KOG profile at 80% probability and 0.5 pairwise coverage. For all downstream analyses, we considered only the top hit as relevant for annotation.\u0026nbsp;\u003c/p\u003e\n\u003ch3\u003e\u003cstrong\u003eData filtering and removal of likely horizontal gene transfer cases\u003c/strong\u003e\u003c/h3\u003e\n\u003cp\u003eUnless otherwise stated, all data is interpreted from the core set of EPOCs meeting the following criteria: 1) the eukaryotic constituent cluster profile identifies at least one target profile in our KEGG profile database at a probability of 80% and 50% coverage, 2) the eukaryotic clades include more than 5 distinct taxonomic labels, at least one from both Amorphea and Diaphoretickes, thus encompassing LECA, 3) only prokaryotic sister clades with 0.4 \u0026lt; ELW \u0026lt; 0.99 were considered reliable and were included in the analysis. Cases of ELW \u0026lt; 0.99 were not considered as these are found to be disproportionately eukaryotic-like, often branching from within the eukaryotic clade itself, and thus likely derived from late horizontal transfer post LECA. Due to the presence of such high ELW values within Alphaproteobacterial associations with Oxidative phosphorylation, in this case, as an exception, ELW values of 1 were included.\u0026nbsp;\u003c/p\u003e\n\u003ch3\u003e\u003cstrong\u003eData visualization and plotting\u003c/strong\u003e\u003c/h3\u003e\n\u003cp\u003eAll plots were prepared using python v3.9 and pandas v.2.0.3 with Altair v.4.2.2. Layout, annotation, and vector editing was done using Inkscape v.1.1.1. All statistical tests were carried out using scipy.stats 1.11.1\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eData availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll initial databases and taxonomy annotation as well as all final parsed and tabulated data is available either at Zenodo url:10.5281/zenodo.14002645 or and at https://ftp.ncbi.nih.gov/pub/wolf/ without any restrictions. All intermediate files, including MSAs and master trees for all EPOCs are available at Zenodo doi:10.5281/zenodo.14004407\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCode availability\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll code and scripts used to generate the final parsed data is available on Github at: https://github.com/VictorTobiasson/eukgen\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSupplementary\u003c/strong\u003e\u003cstrong\u003e\u0026nbsp;Data File 1\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eExcel file with four sheets: 1) List of revised taxonomy mapping species with without well curated taxonomy to closes existing taxonomic ranks. 2) List of revised taxonomy merging sparsely populated classes. 3) Manually curated list of genes involved in Fe-S assembly. 4) Tabular data of all tables included in figures. \u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eE.V.K. initiated the study; V.T. and Y.I.W. developed the pipeline for data analysis; V.T. and J.L. collected the data; V.T., J.L., Y.I.W and E.V.K. performed data analysis; V.T. and E.V.K wrote the manuscript that was read, edited and approved by all authors.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors thank Toni Gabaldon and Koonin Group members for helpful discussions. The authors\u0026rsquo; research is supported by the Intramural Research Program of the National Institutes of Health of the USA (National Library of Medicine).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare no competing interests.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eEmbley, T. M. \u0026amp; Martin, W. Eukaryotic evolution, changes and challenges. \u003cem\u003eNature\u003c/em\u003e\u003cstrong\u003e440\u003c/strong\u003e, 623-630 (2006).\u003c/li\u003e\n\u003cli\u003eVosseberg, J.\u003cem\u003e et al.\u003c/em\u003e The emerging view on the origin and early evolution of eukaryotic cells. \u003cem\u003eNature\u003c/em\u003e\u003cstrong\u003e633\u003c/strong\u003e, 295-305 (2024). https://doi.org:10.1038/s41586-024-07677-6\u003c/li\u003e\n\u003cli\u003evan der Giezen, M. Hydrogenosomes and mitosomes: conservation and evolution of functions. \u003cem\u003eJ Eukaryot Microbiol\u003c/em\u003e\u003cstrong\u003e56\u003c/strong\u003e, 221-231 (2009). https://doi.org:JEU407 [pii] 10.1111/j.1550-7408.2009.00407.x\u003c/li\u003e\n\u003cli\u003eSpang, A.\u003cem\u003e et al.\u003c/em\u003e Complex archaea that bridge the gap between prokaryotes and eukaryotes. \u003cem\u003eNature\u003c/em\u003e\u003cstrong\u003e521\u003c/strong\u003e, 173-179 (2015). https://doi.org:10.1038/nature14447\u003c/li\u003e\n\u003cli\u003eZaremba-Niedzwiedzka, K.\u003cem\u003e et al.\u003c/em\u003e Asgard archaea illuminate the origin of eukaryotic cellular complexity. \u003cem\u003eNature\u003c/em\u003e\u003cstrong\u003e541\u003c/strong\u003e, 353-358 (2017). https://doi.org:10.1038/nature21031\u003c/li\u003e\n\u003cli\u003eLiu, Y.\u003cem\u003e et al.\u003c/em\u003e Expanded diversity of Asgard archaea and their relationships with eukaryotes. \u003cem\u003eNature\u003c/em\u003e\u003cstrong\u003e593\u003c/strong\u003e, 553-557 (2021). https://doi.org:10.1038/s41586-021-03494-3\u003c/li\u003e\n\u003cli\u003eEme, L., Spang, A., Lombard, J., Stairs, C. W. \u0026amp; Ettema, T. J. G. Archaea and the origin of eukaryotes. \u003cem\u003eNat Rev Microbiol\u003c/em\u003e\u003cstrong\u003e15\u003c/strong\u003e, 711-723 (2017). https://doi.org:10.1038/nrmicro.2017.133\u003c/li\u003e\n\u003cli\u003eEme, L.\u003cem\u003e et al.\u003c/em\u003e Inference and reconstruction of the heimdallarchaeial ancestry of eukaryotes. \u003cem\u003eNature\u003c/em\u003e\u003cstrong\u003e618\u003c/strong\u003e, 992-999 (2023). https://doi.org:10.1038/s41586-023-06186-2\u003c/li\u003e\n\u003cli\u003eLopez-Garcia, P. \u0026amp; Moreira, D. The symbiotic origin of the eukaryotic cell. \u003cem\u003eC R Biol\u003c/em\u003e\u003cstrong\u003e346\u003c/strong\u003e, 55-73 (2023). https://doi.org:10.5802/crbiol.118\u003c/li\u003e\n\u003cli\u003e Alberts, B.\u003cem\u003e et al.\u003c/em\u003e\u003cem\u003eMolecular Biology of the Cell\u003c/em\u003e. 7 edn, (W.W. Norton \u0026amp; Co, 2022).\u003c/li\u003e\n\u003cli\u003e Cavalier-Smith, T. Eukaryotes with no mitochondria. \u003cem\u003eNature\u003c/em\u003e\u003cstrong\u003e326\u003c/strong\u003e, 332-333 (1987). https://doi.org:10.1038/326332a0\u003c/li\u003e\n\u003cli\u003e Sogin, M. L., Gunderson, J. H., Elwood, H. J., Alonso, R. A. \u0026amp; Peattie, D. A. Phylogenetic meaning of the kingdom concept: an unusual ribosomal RNA from Giardia lamblia. \u003cem\u003eScience\u003c/em\u003e\u003cstrong\u003e243\u003c/strong\u003e, 75-77 (1989). https://doi.org:10.1126/science.2911720\u003c/li\u003e\n\u003cli\u003e Sagan, L. On the origin of mitosing cells. \u003cem\u003eJ Theor Biol\u003c/em\u003e\u003cstrong\u003e14\u003c/strong\u003e, 255-274 (1967).\u003c/li\u003e\n\u003cli\u003e Martin, W., Hoffmeister, M., Rotte, C. \u0026amp; Henze, K. An overview of endosymbiotic models for the origins of eukaryotes, their ATP-producing organelles (mitochondria and hydrogenosomes), and their heterotrophic lifestyle. \u003cem\u003eBiol Chem\u003c/em\u003e\u003cstrong\u003e382\u003c/strong\u003e, 1521-1539 (2001). https://doi.org:10.1515/BC.2001.187\u003c/li\u003e\n\u003cli\u003e Lopez-Garcia, P. \u0026amp; Moreira, D. The Syntrophy hypothesis for the origin of eukaryotes revisited. \u003cem\u003eNat Microbiol\u003c/em\u003e\u003cstrong\u003e5\u003c/strong\u003e, 655-667 (2020). https://doi.org:10.1038/s41564-020-0710-4\u003c/li\u003e\n\u003cli\u003e Brown, J. R. \u0026amp; Doolittle, W. F. Archaea and the prokaryote-to-eukaryote transition. \u003cem\u003eMicrobiol Mol Biol Rev\u003c/em\u003e\u003cstrong\u003e61\u003c/strong\u003e, 456-502 (1997). https://doi.org:10.1128/mmbr.61.4.456-502.1997\u003c/li\u003e\n\u003cli\u003e Rivera, M. C., Jain, R., Moore, J. E. \u0026amp; Lake, J. A. Genomic evidence for two functionally distinct gene classes. \u003cem\u003eProc Natl Acad Sci U S A\u003c/em\u003e\u003cstrong\u003e95\u003c/strong\u003e, 6239-6244 (1998). https://doi.org:10.1073/pnas.95.11.6239\u003c/li\u003e\n\u003cli\u003e Ribeiro, S. \u0026amp; Golding, G. B. The mosaic nature of the eukaryotic nucleus. \u003cem\u003eMol Biol Evol\u003c/em\u003e\u003cstrong\u003e15\u003c/strong\u003e, 779-788 (1998). https://doi.org:10.1093/oxfordjournals.molbev.a025983\u003c/li\u003e\n\u003cli\u003e Koonin, E. V.\u003cem\u003e et al.\u003c/em\u003e A comprehensive evolutionary classification of proteins encoded in complete eukaryotic genomes. \u003cem\u003eGenome Biol\u003c/em\u003e\u003cstrong\u003e5\u003c/strong\u003e, R7 (2004).\u003c/li\u003e\n\u003cli\u003e Imachi, H.\u003cem\u003e et al.\u003c/em\u003e Promethearchaeum syntrophicum gen. nov., sp. nov., an anaerobic, obligately syntrophic archaeon, the first isolate of the lineage 'Asgard' archaea, and proposal of the new archaeal phylum Promethearchaeota phyl. nov. and kingdom Promethearchaeati regn. nov. \u003cem\u003eInt J Syst Evol Microbiol\u003c/em\u003e\u003cstrong\u003e74\u003c/strong\u003e (2024). https://doi.org:10.1099/ijsem.0.006435\u003c/li\u003e\n\u003cli\u003e Lombard, J., Lopez-Garcia, P. \u0026amp; Moreira, D. The early evolution of lipid membranes and the three domains of life. \u003cem\u003eNat Rev Microbiol\u003c/em\u003e\u003cstrong\u003e10\u003c/strong\u003e, 507-515 (2012). https://doi.org:10.1038/nrmicro2815\u003c/li\u003e\n\u003cli\u003e Moreira, D. \u0026amp; Lopez-Garcia, P. Symbiosis between methanogenic archaea and delta-proteobacteria as the origin of eukaryotes: the syntrophic hypothesis. \u003cem\u003eJ Mol Evol\u003c/em\u003e\u003cstrong\u003e47\u003c/strong\u003e, 517-530 (1998). https://doi.org:10.1007/pl00006408\u003c/li\u003e\n\u003cli\u003e Krupovic, M., Dolja, V. V. \u0026amp; Koonin, E. V. The virome of the last eukaryotic common ancestor and eukaryogenesis. \u003cem\u003eNat Microbiol\u003c/em\u003e\u003cstrong\u003e8\u003c/strong\u003e, 1008-1017 (2023). https://doi.org:10.1038/s41564-023-01378-y\u003c/li\u003e\n\u003cli\u003e Imachi, H.\u003cem\u003e et al.\u003c/em\u003e Isolation of an archaeon at the prokaryote-eukaryote interface. \u003cem\u003eNature\u003c/em\u003e\u003cstrong\u003e577\u003c/strong\u003e, 519-525 (2020). https://doi.org:10.1038/s41586-019-1916-6\u003c/li\u003e\n\u003cli\u003e Lopez-Garcia, P. \u0026amp; Moreira, D. Cultured Asgard Archaea Shed Light on Eukaryogenesis. \u003cem\u003eCell\u003c/em\u003e\u003cstrong\u003e181\u003c/strong\u003e, 232-235 (2020). https://doi.org:10.1016/j.cell.2020.03.058\u003c/li\u003e\n\u003cli\u003e Pittis, A. A. \u0026amp; Gabaldon, T. Late acquisition of mitochondria by a host with chimaeric prokaryotic ancestry. \u003cem\u003eNature\u003c/em\u003e\u003cstrong\u003e531\u003c/strong\u003e, 101-104 (2016). https://doi.org:10.1038/nature16941\u003c/li\u003e\n\u003cli\u003e Gabaldon, T. Origin and Early Evolution of the Eukaryotic Cell. \u003cem\u003eAnnu Rev Microbiol\u003c/em\u003e\u003cstrong\u003e75\u003c/strong\u003e, 631-647 (2021). https://doi.org:10.1146/annurev-micro-090817-062213\u003c/li\u003e\n\u003cli\u003e Steinegger, M. \u0026amp; Soding, J. MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. \u003cem\u003eNat Biotechnol\u003c/em\u003e\u003cstrong\u003e35\u003c/strong\u003e, 1026-1028 (2017). https://doi.org:10.1038/nbt.3988\u003c/li\u003e\n\u003cli\u003e Steinegger, M.\u003cem\u003e et al.\u003c/em\u003e HH-suite3 for fast remote homology detection and deep protein annotation. \u003cem\u003eBMC Bioinformatics\u003c/em\u003e\u003cstrong\u003e20\u003c/strong\u003e, 473 (2019). https://doi.org:10.1186/s12859-019-3019-7\u003c/li\u003e\n\u003cli\u003e Richter, D. J.\u003cem\u003e et al.\u003c/em\u003e EukProt: A database of genome-scale predicted proteins across the diversity of eukaryotes \u003cem\u003ePeer Community J\u003c/em\u003e\u003cstrong\u003e2\u003c/strong\u003e, e56 (2022).\u003c/li\u003e\n\u003cli\u003e Strimmer, K. \u0026amp; Rambaut, A. Inferring confidence sets of possibly misspecified gene trees. \u003cem\u003eProc Biol Sci\u003c/em\u003e\u003cstrong\u003e269\u003c/strong\u003e, 137-142 (2002). https://doi.org:10.1098/rspb.2001.1862\u003c/li\u003e\n\u003cli\u003e Kanehisa, M., Furumichi, M., Tanabe, M., Sato, Y. \u0026amp; Morishima, K. KEGG: new perspectives on genomes, pathways, diseases and drugs. \u003cem\u003eNucleic Acids Res\u003c/em\u003e\u003cstrong\u003e45\u003c/strong\u003e, D353-D361 (2017). https://doi.org:10.1093/nar/gkw1092\u003c/li\u003e\n\u003cli\u003e Edgar, R. C. Muscle5: High-accuracy alignment ensembles enable unbiased assessments of sequence homology and phylogeny. \u003cem\u003eNat Commun\u003c/em\u003e\u003cstrong\u003e13\u003c/strong\u003e, 6968 (2022). https://doi.org:10.1038/s41467-022-34630-w\u003c/li\u003e\n\u003cli\u003e Minh, B. Q.\u003cem\u003e et al.\u003c/em\u003e IQ-TREE 2: New Models and Efficient Methods for Phylogenetic Inference in the Genomic Era. \u003cem\u003eMol Biol Evol\u003c/em\u003e\u003cstrong\u003e37\u003c/strong\u003e, 1530-1534 (2020). https://doi.org:10.1093/molbev/msaa015\u003c/li\u003e\n\u003cli\u003e Derelle, R.\u003cem\u003e et al.\u003c/em\u003e Bacterial proteins pinpoint a single eukaryotic root. \u003cem\u003eProc Natl Acad Sci U S A\u003c/em\u003e\u003cstrong\u003e112\u003c/strong\u003e, E693-699 (2015). https://doi.org:10.1073/pnas.1420657112\u003c/li\u003e\n\u003cli\u003e Al Jewari, C. \u0026amp; Baldauf, S. L. Conflict over the Eukaryote Root Resides in Strong Outliers, Mosaics and Missing Data Sensitivity of Site-Specific (CAT) Mixture Models. \u003cem\u003eSyst Biol\u003c/em\u003e\u003cstrong\u003e72\u003c/strong\u003e, 1-16 (2023). https://doi.org:10.1093/sysbio/syac029\u003c/li\u003e\n\u003cli\u003e Paulick, M. G. \u0026amp; Bertozzi, C. R. The glycosylphosphatidylinositol anchor: a complex membrane-anchoring structure for proteins. \u003cem\u003eBiochemistry\u003c/em\u003e\u003cstrong\u003e47\u003c/strong\u003e, 6991-7000 (2008). https://doi.org:10.1021/bi8006324\u003c/li\u003e\n\u003cli\u003e Stankeviciute, G.\u003cem\u003e et al.\u003c/em\u003e Convergent evolution of bacterial ceramide synthesis. \u003cem\u003eNat Chem Biol\u003c/em\u003e\u003cstrong\u003e18\u003c/strong\u003e, 305-312 (2022). https://doi.org:10.1038/s41589-021-00948-7\u003c/li\u003e\n\u003cli\u003e Biran, A., Santos, T. C. B., Dingjan, T. \u0026amp; Futerman, A. H. The Sphinx and the egg: Evolutionary enigmas of the (glyco)sphingolipid biosynthetic pathway. \u003cem\u003eBiochim Biophys Acta Mol Cell Biol Lipids\u003c/em\u003e\u003cstrong\u003e1869\u003c/strong\u003e, 159462 (2024). https://doi.org:10.1016/j.bbalip.2024.159462\u003c/li\u003e\n\u003cli\u003e Hoshino, Y. \u0026amp; Gaucher, E. A. On the Origin of Isoprenoid Biosynthesis. \u003cem\u003eMol Biol Evol\u003c/em\u003e\u003cstrong\u003e35\u003c/strong\u003e, 2185-2197 (2018). https://doi.org:10.1093/molbev/msy120\u003c/li\u003e\n\u003cli\u003e Richards, T. A. \u0026amp; van der Giezen, M. Evolution of the Isd11-IscS complex reveals a single alpha-proteobacterial endosymbiosis for all eukaryotes. \u003cem\u003eMol Biol Evol\u003c/em\u003e\u003cstrong\u003e23\u003c/strong\u003e, 1341-1344 (2006). https://doi.org:10.1093/molbev/msl001\u003c/li\u003e\n\u003cli\u003e Cassier-Chauvat, C., Marceau, F., Farci, S., Ouchane, S. \u0026amp; Chauvat, F. The Glutathione System: A Journey from Cyanobacteria to Higher Eukaryotes. \u003cem\u003eAntioxidants (Basel)\u003c/em\u003e\u003cstrong\u003e12\u003c/strong\u003e (2023). https://doi.org:10.3390/antiox12061199\u003c/li\u003e\n\u003cli\u003e Kassube, S. A. \u0026amp; Thoma, N. H. Structural insights into Fe-S protein biogenesis by the CIA targeting complex. \u003cem\u003eNat Struct Mol Biol\u003c/em\u003e\u003cstrong\u003e27\u003c/strong\u003e, 735-742 (2020). https://doi.org:10.1038/s41594-020-0454-0\u003c/li\u003e\n\u003cli\u003e Hedges, S. B. The origin and evolution of model organisms. \u003cem\u003eNat Rev Genet\u003c/em\u003e\u003cstrong\u003e3\u003c/strong\u003e, 838-849 (2002). https://doi.org:10.1038/nrg929\u003c/li\u003e\n\u003cli\u003e Vosseberg, J.\u003cem\u003e et al.\u003c/em\u003e Timing the origin of eukaryotic cellular complexity with ancient duplications. \u003cem\u003eNat Ecol Evol\u003c/em\u003e\u003cstrong\u003e5\u003c/strong\u003e, 92-100 (2021). https://doi.org:10.1038/s41559-020-01320-z\u003c/li\u003e\n\u003cli\u003e Caforio, A.\u003cem\u003e et al.\u003c/em\u003e Converting Escherichia coli into an archaebacterium with a hybrid heterochiral membrane. \u003cem\u003eProc Natl Acad Sci U S A\u003c/em\u003e\u003cstrong\u003e115\u003c/strong\u003e, 3704-3709 (2018). https://doi.org:10.1073/pnas.1721604115\u003c/li\u003e\n\u003cli\u003e Hoekzema, M., Jiang, J. \u0026amp; Driessen, A. J. M. Optimizing Archaeal Lipid Biosynthesis in Escherichia coli. \u003cem\u003eACS Synth Biol\u003c/em\u003e\u003cstrong\u003e13\u003c/strong\u003e, 2470-2479 (2024). https://doi.org:10.1021/acssynbio.4c00235\u003c/li\u003e\n\u003cli\u003e Tamarit, D.\u003cem\u003e et al.\u003c/em\u003e A closed Candidatus Odinarchaeum chromosome exposes Asgard archaeal viruses. \u003cem\u003eNat Microbiol\u003c/em\u003e\u003cstrong\u003e7\u003c/strong\u003e, 948-952 (2022). https://doi.org:10.1038/s41564-022-01122-y\u003c/li\u003e\n\u003cli\u003e Rodrigues-Oliveira, T.\u003cem\u003e et al.\u003c/em\u003e Actin cytoskeleton and complex cell architecture in an Asgard archaeon. \u003cem\u003eNature\u003c/em\u003e\u003cstrong\u003e613\u003c/strong\u003e, 332-339 (2023). https://doi.org:10.1038/s41586-022-05550-y\u003c/li\u003e\n\u003cli\u003e Munoz-Gomez, S. A.\u003cem\u003e et al.\u003c/em\u003e Site-and-branch-heterogeneous analyses of an expanded dataset favour mitochondria as sister to known Alphaproteobacteria. \u003cem\u003eNat Ecol Evol\u003c/em\u003e\u003cstrong\u003e6\u003c/strong\u003e, 253-262 (2022). https://doi.org:10.1038/s41559-021-01638-2\u003c/li\u003e\n\u003cli\u003e Hyatt, D.\u003cem\u003e et al.\u003c/em\u003e Prodigal: prokaryotic gene recognition and translation initiation site identification. \u003cem\u003eBMC Bioinformatics\u003c/em\u003e\u003cstrong\u003e11\u003c/strong\u003e, 119 (2010). https://doi.org:10.1186/1471-2105-11-119\u003c/li\u003e\n\u003cli\u003e Huerta-Cepas, J., Serra, F. \u0026amp; Bork, P. ETE 3: Reconstruction, Analysis, and Visualization of Phylogenomic Data. \u003cem\u003eMol Biol Evol\u003c/em\u003e\u003cstrong\u003e33\u003c/strong\u003e, 1635-1638 (2016). https://doi.org:10.1093/molbev/msw046\u003c/li\u003e\n\u003cli\u003e Deorowicz, S., Debudaj-Grabysz, A. \u0026amp; Gudys, A. FAMSA: Fast and accurate multiple sequence alignment of huge protein families. \u003cem\u003eSci Rep\u003c/em\u003e\u003cstrong\u003e6\u003c/strong\u003e, 33964 (2016). https://doi.org:10.1038/srep33964\u003c/li\u003e\n\u003cli\u003e Price, M. N., Dehal, P. S. \u0026amp; Arkin, A. P. FastTree 2--approximately maximum-likelihood trees for large alignments. \u003cem\u003ePLoS One\u003c/em\u003e\u003cstrong\u003e5\u003c/strong\u003e, e9490 (2010). https://doi.org:10.1371/journal.pone.0009490\u003c/li\u003e\n\u003cli\u003e UniProt, C. UniProt: the universal protein knowledgebase in 2021. \u003cem\u003eNucleic Acids Res\u003c/em\u003e\u003cstrong\u003e49\u003c/strong\u003e, D480-D489 (2021). https://doi.org:10.1093/nar/gkaa1100\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"nature-portfolio","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"","title":"Nature Portfolio","twitterHandle":"","acdcEnabled":false,"dfaEnabled":false,"editorialSystem":"ejp","reportingPortfolio":"","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Asgard archaea, Origin of eukaryotes, Asgard contribution to eukaryogenesis, Endosymbiosis, Horizontal gene transfer ","lastPublishedDoi":"10.21203/rs.3.rs-5352492/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-5352492/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"The origin of eukaryotes is one of the key problems in evolutionary biology 1,2. The demonstration that the Last Eukaryotic Common Ancestor (LECA) already contained the mitochondrion, an endosymbiotic organelle derived from an alphaproteobacterium 3, and the discovery of Asgard archaea, the closest archaeal relatives of eukaryotes 4-8, inform and constrain evolutionary scenarios of eukaryogenesis 9. We undertook a comprehensive analysis of the origins of the core eukaryotic genes tracing to the LECA within a rigorous statistical framework centered around evolutionary hypotheses testing using constrained phylogenetic trees. The results reveal dominant contributions of Asgard archaea to the origin of most of the conserved eukaryotic functional systems and pathways. A limited contribution from Alphaproteobacteria was identified, primarily relating to the energy transformation systems and Fe-S cluster biogenesis, whereas ancestry from other bacterial phyla was scattered across the eukaryotic functional landscape, with almost no consistent trends. These findings suggest a model of eukaryogenesis in which key features of eukaryotic cell organization evolved in the Asgard ancestor, followed by the capture of the Alphaproteobacterial endosymbiont, and augmented by numerous but sporadic horizontal acquisition of genes from other bacteria both before and after endosymbiosis.","manuscriptTitle":"Dominant contribution of Asgard archaea to eukaryogenesis","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-10-31 09:48:22","doi":"10.21203/rs.3.rs-5352492/v1","editorialEvents":[],"status":"published","journal":{"display":false,"email":"
[email protected]","identity":"nature","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"nature","sideBox":"Learn more about [Nature](http://www.nature.com/nature/)","snPcode":"","submissionUrl":"","title":"Nature","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"Nature","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"fd5517cb-a771-4851-bcf8-b4edb33b49a4","owner":[],"postedDate":"October 31st, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":39598353,"name":"Biological sciences/Evolution/Phylogenetics"},{"id":39598354,"name":"Biological sciences/Evolution/Molecular evolution"}],"tags":[],"updatedAt":"2026-01-15T08:10:38+00:00","versionOfRecord":{"articleIdentity":"rs-5352492","link":"https://doi.org/10.1038/s41586-025-09960-6","journal":{"identity":"nature","isVorOnly":false,"title":"Nature"},"publishedOn":"2026-01-14 05:00:00","publishedOnDateReadable":"January 14th, 2026"},"versionCreatedAt":"2024-10-31 09:48:22","video":"","vorDoi":"10.1038/s41586-025-09960-6","vorDoiUrl":"https://doi.org/10.1038/s41586-025-09960-6","workflowStages":[]},"version":"v1","identity":"rs-5352492","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-5352492","identity":"rs-5352492","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.