Analysis of NUDIX enzymes across fungi reveals previously unrecognized diversity

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-17

This study comprehensively analyzed fungal proteomes, identifying 25 NUDIX enzyme subfamilies, 13 of which are novel, suggesting diverse and physiologically relevant functions.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

Abstract

Abstract Background The NUDIX superfamily encompasses highly diverse enzymes involved in a plethora of biological functions such as mRNA metabolism, DNA repair, and lipid peroxidation. These hydrolases are found in all domains of life and show surprising versatility in terms of the substrates that they process. The knowledge about the diversity of fungal NUDIX proteins is fragmentary, being largely limited to a small number of characterized enzymes from yeasts. To address this knowledge gap systematically, we performed a detailed analysis of the NUDIX hydrolases across 183 fungal proteomes. Results Members of six of the known NUDIX families were present in fungi being particularly abundant in Glomeromycota. Phylogenetic analysis and sequence clustering grouped fungal NUDIX enzymes in 25 subfamilies, 13 of which did not cluster with previously known enzymes. These 13 newly identified subfamilies all belong to the canonical NUDIX family, and structural comparison revealed a typical NUDIX fold with α-β-α sandwich structure. Molecular docking suggested Ap3A and Ap4A as substrates with the highest binding affinity, but their possible cellular roles remain unclear. We also found evidence of expression of most of the genes that encode these enzymes, suggesting physiological relevance. Conclusions Our analysis offers a comprehensive perspective on the structural and sequence relationships of the NUDIX superfamily across fungi with potential to guide experimental characterization of their biological functions.
Full text 143,402 characters · extracted from preprint-html · click to expand
Analysis of NUDIX enzymes across fungi reveals previously unrecognized diversity | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Analysis of NUDIX enzymes across fungi reveals previously unrecognized diversity Zofia Pasterny, Drishtee Barua, Eugenio Mancera, Anna Muszewska This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6343747/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 01 Jul, 2025 Read the published version in BMC Genomics → Version 1 posted 10 You are reading this latest preprint version Abstract Background The NUDIX superfamily encompasses highly diverse enzymes involved in a plethora of biological functions such as mRNA metabolism, DNA repair, and lipid peroxidation. These hydrolases are found in all domains of life and show surprising versatility in terms of the substrates that they process. The knowledge about the diversity of fungal NUDIX proteins is fragmentary, being largely limited to a small number of characterized enzymes from yeasts. To address this knowledge gap systematically, we performed a detailed analysis of the NUDIX hydrolases across 183 fungal proteomes. Results Members of six of the known NUDIX families were present in fungi being particularly abundant in Glomeromycota. Phylogenetic analysis and sequence clustering grouped fungal NUDIX enzymes in 25 subfamilies, 13 of which did not cluster with previously known enzymes. These 13 newly identified subfamilies all belong to the canonical NUDIX family, and structural comparison revealed a typical NUDIX fold with α-β-α sandwich structure. Molecular docking suggested Ap3A and Ap4A as substrates with the highest binding affinity, but their possible cellular roles remain unclear. We also found evidence of expression of most of the genes that encode these enzymes, suggesting physiological relevance. Conclusions Our analysis offers a comprehensive perspective on the structural and sequence relationships of the NUDIX superfamily across fungi with potential to guide experimental characterization of their biological functions. NUDIX hydrolases fungi protein families substrate specificity Glomeromycota Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Background The NUDIX superfamily of enzymes encompasses various organic pyrophosphatases capable of cleaving nucleoside diphosphates linked to any moiety into nucleoside monophosphates and organic pyrophosphates, hence their name. Initially, the superfamily was called the MutT family after its first representative, the 8-oxoGTPase central for E. coli DNA repair [ 1 ]. However, apart from their role in DNA repair, nowadays these enzymes are known to be involved in the hydrolysis of detrimental metabolites, decapping, processing of mRNA, cell cycle regulation, and survival [ 2 ], and gating ion channels [ 3 ]. In agreement with the wide range of functions that they perform, the diversity of the substrates of these enzymes is high, including di- and triphosphates, nucleotide sugars, dinucleosides, diphosphoinositol polyphosphates, and RNA caps [ 4 ]. The biomedical and biotechnological potential of NUDIX proteins is starting to be recognized. Human Dcp2, NUDT2, NUDT12, and NUDT16 can be used to characterize 5′ capped RNA transcripts [ 5 ], while OsNUDX14 from Oryza sativa is thought to be a grain quality regulator and to determine plant development, being strongly expressed in mature leaves [ 6 ]. NUDIX enzymes with feasible druggable sites have also been identified, and therefore they could represent novel drug targets for human diseases, especially in cancer cells [ 2 ]. Table 1 Characteristics of the NUDIX Pfam families regarding the most common co-occuring domains, human and baker’s yeast representatives, function, and presence across kingdoms of life. Pfam Most common coexisting domains H. sapiens representatives S. cervisiae representatives Available PDB structures Function Presence (kingdoms) PF00293 in majority of the cases exists as a single domain NUDT1, NUDT2, NUDT3, NUDT4, NUDT4B, NUDT5, NUDT6, NUDT7, NUDT8, NUDT9, NUDT10, NUDT11, NUDT12, NUDT13, NUDT14, NUDT15, NUDT16, NUDT17, NUDT18, NUDT20, IDI1, IDI2 DCP2, DDP1, IDI1, NPY1, PCD1,YJR142W, YJR142W, YSA1 Among others: H. sapiens: NUDT1, NUDT2, NUDT7, NUDT12, NUDT15, NUDT16, IDI1 and IDI2, DDP1, DDP2, DPP3-alpha, ADP-ribose pyrophosphatase, Adenine DNA glycosylase, mitochondrial 39S ribosomal protein L46; E. coli: nudE, NADH pyrophosphatase, 7,8-dihydro-8-oxoguanine triphosphatase, nudF. various functions, but the most prominent one is sanitization of nucleotide pool all kingdoms PF03559 in majority of the cases exists as a single domain EvaA 2,3-dehydratase (A. orientalis) involved in antibiotic production pathways mainly bacteria PF09296 zf-NADH-PPase (PF09297), NUDIX domain (PF00293) NUDT12, NUDT13 Peroxisomal NADH pyrophosphatase NUDT12 (H. sapiens, M. musculus) deNADing NAD-capped RNA eukaryotes and bacteria PF11788 in majority of the cases exists as a single domain or along cannonical NUDIX domain (PF00293) MRLP46 Mitochondrial ribosomal protein L46 Mitochondrial ribosomal protein L46 (H. sapiens, S. scrofa, S. cerevisiae, T. brucei, L. major, N. crassa) role in mitochondrial protein synthesis eukaryotes PF13869 in majority of the cases exists as a single domain NUDT21 Cleavage and polyadenylation specificity factor subunit 5 (H. sapiens) 3' RNA cleavage and polyadenylation processing eukaryotes PF14815 HhH-GPD superfamily base excision DNA repair protein (PF14815) MUTYH Adenine DNA glycosylase (M. musculus, G. stearothermophilus, H. sapiens), MutT protein (B. bacteriovorus, B. fragilis, S. aureus, E. coli, B. henselae), 7,8-dihydro-8-oxoguanine-triphosphatase (K. pneumoniae) excises inappropriately matched adenine from DNA backbone to either 8-oxoguanine or guanine eukaryotes and bacteria PF15916 NUDIX domain (PF00293) MutT/nudix family protein (R. rubrum ATCC 11170) functionally uncharacterised, takes part in thiamine biosynthetic process eukaryotes and bacteria PF16262 in majority of the cases exists as a single domain Putative uncharacterized protein (J. denitrificans) unknown bacteria PF16705 FERM central domain (PF00373) Krev interaction trapped protein 1 Krev interaction trapped protein 1 (H. sapiens) activates β1 integrin by antagonization of ICAP1 (Integrin Cytoplasmic Associated Protein-1) eukaryotes (only animals) PF18290 NUDIX domain (PF00293) NUDT6 (H. sapiens) NUDT7 (A. thaliana), NUDT6 (H. sapiens) hydrolases catalyze the hydrolysis of nucleoside diphosphates which are often toxic metabolic intermediates and signalling molecules eukaryotes and bacteria Despite the functional differences between these enzymes, at the structural level, all NUDIX hydrolases are characterized by an α-β-α sandwich structure with a specific NUDIX motif, which contains the catalytic site and metal-binding residues. The NUDIX motif is composed of 23 amino acids: Gx 5 Ex 7 REUxEExGU where U is a bulky amino acid (typically isoleucine, leucine, or valine), and x represents any amino acid. In addition, the N-terminus of the helix fold has glutamates necessary for binding divalent cations that link the NUDIX hydrolases with the pyrophosphate of the substrate. In terms of the reaction, nucleophilic substitution by water occurs at particular phosphorus atoms within a diphosphate or polyphosphate chain [ 7 ]. NUDIX enzymes are ubiquitous in all kingdoms of life and viruses [ 3 , 8 , 9 ] although these proteins are evolutionarily related, they exhibit significant sequence divergence. For this reason the superfamily has been divided into ten Pfam families. Table 1 presents the characteristics of each family. In fungi there are 29 NUDIX enzymes reported at the UniProt database that are part of five of the ten families. All of these records belong to Dikarya taxons and mostly to the model yeasts Saccharomyces cerevisiae and Schizosaccharomyces pombe . Therefore, the phylogenetic distribution and functional roles of NUDIX enzymes in fungi remain poorly understood. In this work, we identified 25 subfamilies belonging to the NUDIX superfamily throughout the fungal kingdom, 13 of which are newly identified and several appear to be only present in fungi. NUDIX enzymes are particularly abundant in the phylum Glomeromycota. The newly identified subfamilies possess characteristics typical of NUDIX hydrolases such as α-β-α sandwich structure and all belong to the canonical NUDIX family. Molecular docking suggested particular substrate preferences and mining available transcriptomic data provided evidence of their physiological relevance. Overall, our work highlights the extensive diversity of NUDIX enzymes in fungi and positions these organisms as a potential source of newly identified NUDIX enzymes for biotechnological, antifungal or research applications. Methods Identification of NUDIX proteins across 183 fungal proteomes To identify NUDIX enzymes across the fungal tree, we first searched all Pfam families in a set of 183 fungal proteomes ( Supplementary file S1 , sheet S1) by amino acid sequence similarity using pfam_scan.pl with default settings against Pfam database v. 36 [ 10 ]. Then, from the full set of Pfam family assignments of the 183 fungal proteomes, we extracted all proteins that matched the NUDIX families grouped in the CL0261 clan. The redundancy across the identified NUDIX sequences was first reduced with cd-hit (70% sequence identity, 90% coverage, 5 letters word length) [ 11 ]. Full protein sequences were then aligned with MAFFT v7.40 using the iterative local alignment mode [ 12 , 13 ] and trimmed manually to the domain region using Jalview [ 14 ]. As a reference set of NUDIX enzymes, all human proteins with NUDIX domains were retrieved from the UniProt database (status reviewed) and all NUDIX domain-containing proteins with resolved structures were obtained from PDB [ 15 ]. The set of human and PDB sequences will be referred to as the r eference sequence set . The PDB sequences of NUDIX hydrolases include proteins from all kingdoms and viruses. Together, the reference sequence set contains representatives of all ten Pfam NUDIX families. The identified fungal NUDIX domain sequences together with the reference sequence set and Pfam consensus sequences were then used as queries in a blastp v. 2.13.0 + search [ 16 ] against the same set of 183 proteomes using an e-value cutoff of 0.001. The blastp resulted in 7,394 fungal sequences, out of which 2,801 turned out to belong to the NUDIX superfamily. The remaining blastp hits contain domains that often coexist with the NUDIX domain (see Supplementary file S1 , sheet S8 to see the list of these domains) and were identified with blastp by the full-length sequences of the reference set queries. The set of 2,801 fungal NUDIX hits together with the reference sequences was subjected to clustering based on amino acid sequence similarity using CLANS (with p-value of 1e − 10, attraction exponent value of 2 and remaining parameters set to default values). Phylogenetic analysis To trace the evolutionary relationships between NUDIX subfamilies, representative sequences from related clusters were aligned with mafft (--localpair –maxiterate 100, v7.407 (Rozewicki et al., 2019). The resulting alignment was then trimmed manually in Jalview and used for phylogenetic analyses. The phylogenetic tree was inferred for each of the alignments by maximum likelihood with IQTREE2 [ 17 ], considering 1000 bootstrap replicates (parameters: -B 1000) with automated model selection. The phylogenetic tree at the species level was generated with OrthoFinder [ 18 ] (parameters: mmseqs and dendroblast) for 41 species spanning the taxonomic diversity of the whole fungal kingdom. The Choanoflagellate Monosiga brevicollis and the fruit fly Drosophila melanogaster were used as outgroups. The species tree was used to represent the taxonomic distribution of individual protein families and subfamilies across the diverse fungal lineages. Trees were visualized and rendered in iToL [ 19 ]. Sequence and structural characterization identified subfamilies To characterize the sequence and structural properties of all fungal NUDIX subfamilies, all of the sequences were scanned against all Pfam family definitions with pfamscan.pl [ 20 ] to confirm the presence of the NUDIX domain (PF00293), and look for other coexisting domains. Where available, we used NCBI Conserved Domain Database (CDD) [ 21 ] annotations for the subfamilies, since CDD enables the identification of conserved domains across multiple protein sequences and defines a more detailed classification of subfamilies of the NUDIX proteins than the Pfam database. Complete protein sequences were then aligned with MAFFT v7.407 [ 13 ] and trimmed to the domain region using Jalview [ 14 ] to construct logos of the NUDIX signature motif with WebLogo 3 [ 22 ]. Low complexity regions were predicted with the SEG tool [ 23 ], coiled-coil regions with EMBOSS pepcoil [ 24 ], disordered protein regions with IUPred3 [ 25 ], subcellular localization with WoLF PSORT [ 26 ], and transmembrane regions were identified with TMHMM [ 27 ]. For each of the newly identified subfamilies, we also retrieved annotations of upstream and downstream genes with Entrez Direct [ 28 ] in order to verify their possible involvement in metabolic clusters. In order to verify the taxonomic representation of the new subfamilies, we extended the sequence search by using the newly identified fungal NUDIX sequences d as queries in a blastp v. 2.13.0 + search [ 16 ] against the NCBI non-redundant (NR) database with an e-value cutoff of 0.001 to collect all homologs. This set of sequences was mapped on the NCBI-Taxonomy database using the taxdump files (nodes and names) to extract their taxonomic assignment. We then calculated the percentage of fungal sequences in all of the subfamilies. Molecular docking for substrate identification In order to predict the potential specificities of each of the newly identified subfamilies, molecular docking was performed. Each of the subfamilies was represented by two selected sequences and their AlphaFold2 3D structure models [ 29 , 30 ]. These structures were then scanned against the FoldSeek database to find the most similar structures [ 31 ]. PyMOL was used to visualize and compare the structure of the representatives of the newly identified subfamilies with available PDB structures of NUDIX enzymes (The PyMOL Molecular Graphics System, Version 3.0 Schrödinger, LLC). Structures were trimmed to the NUDIX domain region in PyMOL and prepared for the molecular docking using pdb2pqr (parameters: --ff AMBER --ph-calc-method = propka --with-ph = xx) [ 32 ]. Molecular docking with AutoDock Vina v1.2.5 [ 33 , 34 ] was performed for the twelve most common NUDIX substrates: 8-oxo-dGDP, 8-oxo-dGTP, ADP-glucose, ADP-ribose, Ap3A, Ap4A, Ap6A, dATP, DHNTP, HSCoA, NAD + and NADPH. The substrates were obtained from the ChEBI database in SMILES format [ 29 ] and converted into three-dimensional PDBQT structures with Open Babel 3.1.1 [ 35 ] (parameters: --gen3D -p xx -opdbqt ). AutoDock Vina predicts the receptor-ligand binding conformations and ranks them according to the predicted change in free energy during binding [ 36 ]. For all the proteins and ligands tested, we ensured they were at the same pH and protonation state to ensure the biological relevance of the predictions. As the NUDIX hydrolases reach their maximum activity in a mildly basic environment we consider the protomers of the tested representatives at pH 7.0, 8.0 and 9.0. Analysis of gene expression Expression of the genes encoding the NUDIX proteins was verified using publicly available transcriptomic datasets of fungi. The data contained a total of 25 transcriptomes (six Mucoro- and Mortierellomycota species, four Glomero- and Basidiomycota, and five Ascomycota datasets). Description of transcriptome datasets, references and results of the gene expression analyses are available in Supplementary file S1 , sheets S5, S6, S7. The data files were obtained from the ENA server in fastq format, and their quality was checked with FASTQC (v0.11.8) [ 37 ]. Adapters were trimmed using fastp (v0.19.6) with default settings [ 38 ]. The adapter-trimmed reads were aligned with reference genomes of the respective organisms (downloaded from NCBI Datasets) using Hisat2 (v2.1.0) [ 39 ]. The SAM alignment files were compressed into binary format (BAM) using samtools (v.10) [ 40 ]. The aligned reads were mapped with GFF files using StringTie (v2.1.3b) [ 41 ] to generate read abundance in the form of Transcripts per Million (TPM) values. The average TPM was calculated for each dataset and values < 1 were regarded as statistically not significant and removed. The remaining TPM values were then subjected to Z-score normalization. Finally, NUDIX enzymes from the identified families and subfamilies were searched in the expression dataset of protein-coding genes. A schematic representation of the workflow is shown in Supplementary Figure S1 . Results NUDIX enzymes are widespread across the fungal kingdom To identify NUDIX enzymes across fungi, we scanned a set of 183 fungal proteomes from all the fungal phyla for the presence of representatives of the ten NUDIX families included in the Pfam database. Six out of the ten families had homologs in the analyzed proteomes (Fig. 1 ), and two of the four missing families (PF03559 and PF16262) have also not been found in other eukaryotes. The canonical NUDIX domain family (PF00293) was by far the most common family present in 182 out of the 183 proteomes. This family was the only one identified in all phyla with a median number of proteins per proteome ranging from 2 for parasitic Microsporidia to more than 14 in Basidiomycota and Mucoromycota. The presence of the remaining families was less uniform (Fig. 2 ): the DUF4743 family (PF15916) was present in 132 proteomes, the 39S mitochondrial ribosomal protein L46 family (PF11788) in 119, the NADH pyrophosphatase-like rudimentary family (PF09296) in 113, the NUDIX_4 domain family (PF14815) in 85, and the NUDIX_2 domain family (PF13869) in 84 proteomes. These five families had a median of close to one representative per phylum and some families were completely absent in several phyla. Overall, the phyla Mucoromycota, Glomeromycota, and Mortierellomycota contain the highest number of NUDIX enzymes, in particular from the canonical family (PF00293). The highest abundance was observed in species from Glomeromycota, including species such as Gigaspora rosea (101 NUDIX proteins), Gigaspora margarita (47 NUDIX proteins), and Diversispora epigaea (50 NUDIX proteins). The phyla Ascomycota, Basidiomycota, Chytridiomycota, Mortierellomycota, and Mucoromycota have representatives of all six families present in fungi. Glomeromycota and Zoopagomycota do not have members of the families PF14815 and PF09296, Kickxellomycota lacks PF14815 NUDIX hydrolases, Microsporidia has only representatives of PF00293 and PF13869 NUDIX proteins and Neocallimastigomycota are missing proteins of the PF11788 family. Our results show that NUDIX enzymes are widespread among fungi although not all known families are present in these organisms. Fungal NUDIX homologs can be clustered into 25 subfamilies To better understand the diversity of fungal NUDIX enzymes, we clustered them according to sequence similarity together with a well-curated reference set that includes representative enzymes from all kingdoms and from all the ten NUDIX Pfam families (see Methods for details). Clustering distinguished 25 groups of fungal proteins with the NUDIX domain. On the other hand, NUDIX reference sequences from four NUDIX families (PF03559, PF16262, PF16705, and PF18290) did not cluster with any fungal protein. This is in agreement with our initial finding that only six NUDIX families are present among fungi. The most abundant NUDIX family in fungi, the canonical NUDIX family (PF00293) formed a large centrally located cluster (Fig. 3 ). In this cluster, we found homologs of nine human NUDIX proteins (IDI1 and IDI2, NUDT1, NUDT3, NUDT5, NUDT7 and NUDT8, NUDT15 and NUDT20), while ten human NUDIX hydrolases did not cluster close to any fungal homolog (NUDT4, NUDT4B, NUDT6, NUDT9, NUDT10, NUDT11, NUDT14, NUDT16, NUDT17, NUDT18, KRIT1). The latter human proteins likely have no fungal orthologs. Fungal homologs of the other five NUDIX families formed small clusters with the corresponding representatives from the reference set. Supplementary Figure S2 presents the phylogenomic tree of representative fungal isolates (spanning the fungal tree of life) annotated with the distribution of NUDIX homologs of characterized NUDIX proteins from the reference sequence set. The clustering results were further assessed by generating maximum likelihood phylogenetic trees with representative sequences from selected clusters. Sequence diversity in the whole PF00293 family does not allow for reliable multiple sequence alignments so we built trees for sets of related clusters within the family. The phylogenetic trees also allowed drawing hypotheses about the evolution of specific NUDIX enzymes. For example, we found that four human NUDIX hydrolases—NUDT4, NUDT4B, NUDT10, and NUDT11—are paralogs of NUDT3. In general, these findings show that there is considerable diversity of NUDIX enzymes in fungi, with representatives of 25 different subfamilies. 13 newly identified subfamilies of NUDIX hydrolases Thirteen fungal NUDIX clusters formed by members of the canonical NUDIX domain family (PF00293) did not contain human or PDB representatives and, therefore, we consider them as newly identified subfamilies. Table 2 summarizes the characteristics of these subfamilies, Supplementary file S1 , sheet S2 provides a full list of proteins associated with each subfamily, and a detailed description of each of the newly identified subfamilies can be found in Supplementary file S2 , their distribution in fungal phyla is summarized in Supplementary Figure S3. The enzymes of the newly identified subfamilies showed considerable diversity in terms of their sequence and structure (Table 2 and Supplementary file S2 ). For example, 40% of the members from subfamily H are predicted to have regions embedded in the cell membrane, and genes encoding proteins from subfamily K are overrepresented in retrotransposons. In terms of cellular localization, most of the members of the new subfamilies are predicted to be localized in the nucleus. Table 2 Characteristics of newly identified fungal NUDIX subfamilies. From left to right, the columns correspond to the assigned subfamily name, the number of proteins in the subfamily, percentage of proteins with low complexity regions, coiled coil regions, percentage of the positions of proteins predicted as being disordered, fraction of proteins with predicted transmembrane helices, and the taxa in which the subfamily is present. All the values correspond to the NUDIX blastp search across 183 fungal proteomes. Subfamily # proteins % low complexity regions % coiled coil regions % positions disordered % transmembrane helices Distribution A 84 98% 43% 51% 0% Fungi B 20 85% 10% 8% 0% Fungi (83.77%), archea and bacteria C 32 69% 0% 5% 9% Fungi D 105 70% 18% 9% 2% Fungi (84%), Protozoa, plants E 36 89% 8% 23% 0% Fungi F 36 67% 6% 13% 0% Fungi G 140 47% 0% 13% 3% Bacteria (48.72%), Fungi (45.80%), Insects H 152 73% 0% 14,35% 39% Fungi (75.50%), Bacteria I 57 61% 25% 10,01% 0% Fungi J 40 38% 10% 11,77% 0% Fungi K 14 86% 57% 18% 0% Fungi (75%), Bacteria L 19 79% 37% 9% 5% Fungi M 22 50% 14% 14,35% 0% Fungi (97.03%), Bacteria All newly identified subfamilies have a well-preserved NUDIX signature motif with at least two glutamates acting as metal ligands and lysine and arginine residues that strengthen catalytic forces [ 7 ] (Fig. 5 A). However, we did observe several substitutions within the NUDIX motif of the newly identified subfamilies, explaining in part why they clustered apart from previously known NUDIX enzymes. High sequence diversity within the NUDIX box is not unprecedented. This is well known for members of the PF13869 family and even the loss of the NUDIX box has been documented for other NUDIX representatives such as members of the PF14815 family [ 3 ] ( Supplementary Figure S4 ). Figure 5 B shows the distribution of the newly identified subfamilies across a set of genomes representing all the phyla of the fungal kingdom. Beside Rozellomycota and Blastocladiomycota, all fungal phyla have representatives of several of the newly identified subfamilies. Glomeromycota are particularly enriched with these proteins, having representatives of 11 out of 13 subfamilies. Moreover, proteins of the newly identified subfamilies are rare outside fungi, with seven of the subfamilies (A, C, E, F, I, J, L) being specific to fungi. Furthermore, the remaining subfamilies that also have representatives in other kingdoms have less than 10% of the total number of proteins from taxa that are not fungi. The former with the exception of subfamily G, for which half of the sequences in the subfamily belong to Terrabacteria representatives. Several subfamilies have very narrow distribution across the fungal kingdom. Subfamilies C, F, I, J, K, and J have representatives only in the subphylum Diversiporales, and subfamilies D and E are found only in Mortierellomycota. Even subfamilies with such narrow taxonomic distribution show sequence diversity ( Supplementary Figure S5, Supplementary Figure S6 ). In addition, proteins from subfamilies A and M are most likely distantly related to human NUDT2 (Fig. 4 A, Supplementary Figure S7, Supplementary Figure S8 ), which is known to regulate ApNA levels associated with stress signaling pathways and protect against oxidative damage to nucleotides. Overall, the clustering with CLANS, the phylogenetic position, and differences in the NUDIX box motif, all support the separation of the newly identified subfamilies. Even the closest newly identified subfamilies J and K, which cluster together in CLANS, form sister clades in the phylogenetic tree and display differences in the NUDIX box motif despite their high overall sequence similarity (32.56%). Proteins of the newly identified fungal NUDIX subfamilies share binding affinity towards specific substrates NUDIX hydrolases are known to have a broad spectrum of substrates. To investigate the substrate preferences of the newly identified subfamilies, we performed molecular docking with representatives of each subfamily. The average affinity per subfamily for twelve different substrates is shown in Fig. 6 and Supplementary Figure S9 . Through all subfamilies, the highest binding affinity was observed towards NADPH, NAD+, Ap3A, and Ap4A. Despite that DHNTPase residues are conserved within the NUDIX box in the proteins from subfamily B ( Supplementary Figure S10 ), even these proteins show preference for Ap3A, and Ap4. These substrates are typical of human NUDT12 and NUDT13 which belong to the NUDIX family of NADH pyrophosphatase-like rudimentary NUDIX domain (PF09296). However, all sequences from the 13 newly identified subfamilies are members of the canonical NUDIX family (PF00293). Moreover, none of them possess the architecture nor the SQPWPFPXS motif that are typical of NADH pyrophosphatases and diphosphatases of the PF09296 family [ 42 ]. Therefore, it is possible that the affinity of the newly identified subfamilies has converged in parallel towards the same substrates. Subfamily C also draws attention since, unlike the rest of the subfamilies, its members show relatively low affinity to Ap4A and high affinity to ADP-ribose. [ 3 ]. The latter specificity is known in proteins involved in post-translational modification removal (NUDT16) and DNA repair (NUDT5). Even when empirical validation is needed to determine the exact function of these newly identified subfamilies, our docking results allow formulating predictions that could guide experimentation. Most of the newly identified NUDIX proteins are expressed To start understanding the functional role of the predicted NUDIX proteins including those from the newly identified subfamilies, we searched for evidence of expression of the genes that encode them in available transcriptomic data. Genome-wide gene expression data is widely available for Dikarya, but it is much more scarce for other fungi. Therefore, we only considered nine representative datasets for Dikarya taxa where the predicted NUDIX enzymes are present, while we included all available datasets in the NCBI SRA database for other fungi. In total, these represented transcriptomes of 25 species ( Supplementary file S1 , sheets S7 & S8). In these datasets, a gene was considered expressed if it had an average TPM value ≥ 1 (Methods). Overall, we found evidence of gene expression for 82% of all the predicted NUDIX enzymes in the 25 species with available RNASeq data. The fraction of expressed NUDIX enzymes varied between families, from 97% in the PF11788 family to 81% in the PF00293 family. At the species level, transcriptomes showed expression of a variable subset of known NUDIX members, from 6 (out of 14) in Cryptococcus neoformans and 6 (out of 10) Endogone sp. FLAS-F59701 to 104 (out of 130) in Gigaspora rosea . Apart from Rhizophagus irregularis which had only members of the canonical NUDIX family expressed, all other species had evidence of expression of at least four of the six families (Fig. 7 ). For proteins from the newly identified subfamilies, overall, we found evidence of expression for 71% of them. The former despite the fact that these subfamilies have a narrower taxonomic distribution and, in consequence, the available expression datasets of taxa that have them in their genomes is considerably limited. In total, we observed the expression of at least one member from each of the thirteen newly identified subfamilies, except for subfamily L. Transcriptomes of species of Glomeromycota showed evidence of expression of eight of the newly identified subfamilies (A, B, C, D, H, I, J, and K), followed by Mortierellomycota with six (A, B, D, E, G and H), Ascomycota with five (B, D, G, H and M), Mucoromycota with four (A, D, G and H), and Basidiomycota with three (G, H and M) (Fig. 7 ). For subfamilies I, J and K there was evidence of expression only in one single species. Moreover, the NUDIX genes analysed in the differential expression experiments on R. delemar and G. rosea were expressed across tested conditions pointing at a housekeeping function of many NUDIX hydrolases ( Supplementary file S1 , sheets S7 & S8). In summary, mining of available transcriptomes showed that the NUDIX enzymes that we predicted are widely expressed in fungi suggesting that they may be of biological importance for these organisms. Discussion Even when NUDIX enzymes were first described more than 70 years ago and despite the many biological functions that they perform, these proteins have been barely studied in fungi, especially beyond Dikarya. In this work, we showed the richness of the NUDIX superfamily across fungi and discovered new subfamilies unique to these organisms. There are only 29 fungal NUDIX enzymes with status reviewed at the UniProt database. These belong to five of the ten known NUDIX families (22 canonical NUDIX domain (PF00293), one NADH pyrophosphatase-like rudimentary NUDIX domain (PF09296), three 39S mitochondrial ribosomal protein L46 (PF11788), one NUDIX_4 domain (PF14815), and two DUF4743 domain (PF15916)). Our work significantly expands the currently known catalog of NUDIX hydrolases in fungi by identifying 2,801 candidate enzymes, including those listed in UniProt as well as enzymes from an additional family not reported for fungi in this database. We did not find representatives of four NUDIX families in the 183 fungal genomes that were analyzed, which was consistent with the fact that there are no fungal NUDIX proteins associated with these families in Uniprot or Pfam. Furthermore, all of the reviewed records of fungal NUDIX hydrolases in UniProt belong to Ascomycota species, while we showed that these enzymes are present across all fungal phyla, with Glomeromycota possessing the highest number and diversity of these enzymes in fungi. Performing detailed sequence similarity analyses we found that the fungal enzymes clustered in 25 subgroups, and that thirteen represent newly identified subfamilies as their members did not cluster with any previously known NUDIX enzyme. All of the newly identified subfamilies belong to the well-known canonical NUDIX family (PF00293). We found evidence that many of the genes that encode the newly identified subfamilies are expressed, suggesting physiological relevance. However, given the many functions that NUDIX enzymes perform, the specific cellular role of these enzymes is difficult to predict based on sequence properties. We observed common substrate preference for diadenosine triphosphate and tetraphosphate (Ap3A and Ap4A) in these enzymes. The level of both diadenosine metabolites is known to increase upon exposure to various stresses in bacteria and humans. Therefore they are considered to be alarmones, triggering a stress response that [ 43 ] influences a variety of cellular processes such as gene expression and DNA repair [ 43 , 44 ] Therefore, it is possible that the enzymes of the newly identified subfamilies may modulate the response to cellular stress in fungi. Alternatively, they could also be involved in metabolic pathways for nucleotide synthesis and degradation. Ap3A and Ap4A signaling in animals is needed for apoptosis and development. Programmed cell death types differ among fungi [ 45 ] including apoptosis-like process, heterokaryon incompatibility, and ferroptosis which is a consequence of iron release and lipid peroxidation. The diversity of fungal developmental strategies, death types, and lipids could also depend, at least partially, on the diversification of the repertoire of NUDIX enzymes. Some of the subfamilies showed preference for ADP-ribose which can also be related with DNA repair or signalling regulation. ADP-ribosylation is often found as a posttranslational modification of proteins [ 46 ]. On the other hand poly-ADP-ribose can accumulate when DNA is damaged [ 47 ]. There are known ADP-ribose processing NUDIX enzymes that can be remove it from proteins reversing glycosylation (NUDT16), and mono ADP-ribose can be converted to ribose-5-phosphate plus AMP/ATP (NUDT5). The newly identified subfamily C may extend this list of ADP-ribose processing NUDIX proteins. NUDIX enzymes from the newly identified subfamilies are particularly abundant in Glomeromycota. Fungi in this phylum are known to lack several DNA repair and replication genes [ 48 ] which could be associated with the abundance of NUDIX enzymes. Alternatively, NUDIX enzymes in Glomeromycota could be involved in lipid metabolism since the phylum is also known to be unique in this respect [ 49 ] and NUDIX enzymes involved in beta-oxidation of fatty acids are widespread in fungi [ 50 ]. In terms of the evolutionary origin of fungal NUDIX enzymes, members of all of the six NUDIX families that we identified in fungi are also present in animals. On the other hand, animals have representatives of two families that we did not find in fungi (PF16705 and PF18290). Since other eukaryotic taxa do have members of the Nudix_hydro PF18290 which is a N-terminal domain preceding the canonical PF00293 domain, its absence in fungi probably represents a lineage-specific loss. On the other hand, the PF16705 ‘Nudix or N-terminal NPxY motif-rich region of KRIT’ is limited to animals and likely represents an evolutionary novelty. Recently, during Pfam 37.1 release, a novel NUDIX domain (PF22327, Nudt16-like) was introduced for U8 snoRNA-decapping enzyme-like proteins from animals which is likely a second NUDIX family limited to this kingdom. Therefore, the NUDIX repertoire of the Opisthokonta ancestor most likely consisted of enzymes from at least the seven families present in animals and fungi (six common to both kingdoms, and Nudix_hydro missing from the analyzed fungi). At the level of subfamilies, only members of the canonical NUDIX PF00293 family formed more than one group when clustering by sequence similarity. This showed that diversification of NUDIX enzymes in fungi has mostly occurred within this family. Seven of these subfamilies included human representatives, suggesting that the divergence of these enzymes occurred before the animal-fungi split. The distribution of the thirteen newly identified subfamilies is more scattered throughout the fungal tree. Members of six of these subfamilies are found across multiple phyla, while six are exclusive to Glomeromycota and one is restricted to Mortierellomycota. Although it is difficult to draw conclusions regarding the specific set of events that gave rise to the observed subfamily distribution in the phylogeny, it is clear that a series of gene duplications followed by sequence diversification probably occurred throughout the evolution of Glomeromycota. Conclusions Our findings could have implications beyond the annotation of the undescribed enzyme universe. NUDIX hydrolases are known to be involved in the pathogenicity of a variety of bacteria [ 51 ], and there have been attempts to use them as antimicrobial targets [ 52 ]. Although much less is known about fungal pathogens, NUDIX hydrolases have been implicated in the oxidative stress response of Cryptococcus neoformans [ 53 ]. The restricted distribution of some of the subfamilies also opens up the possibility to employ them as taxonomic markers either in diagnostics or for biodiversity assessment. For instance, sequences belonging to the newly identified subfamily E and that were found only in Mortierellomycota seem ideal for this purpose. Recently, NUDIX hydrolases received attention as orchestrators of carotenoid biosynthesis in plants [ 54 ]. Therefore, they could also be manipulated in fungi for the production of metabolites of commercial interest. Overall, the diversity of fungal NUDIX hydrolases described here sets them as ideal targets for tailored manipulation in both medical and biotechnological settings. Declarations Data availability statement All metadata processed in this study are deposited in Zenodo: 10.5281/zenodo.15100536 All protein identifiers, genomic assemblies, transcriptomic datasets are listed in Supplementary file S1. The Supplementary Figures and description of novel subfamilies are available in Supplementary file S2. Funding This work was supported by National Science Centre grants (#2021/41/B/NZ2/02426 to A.M.). E.M. received a fellowship from Conahcyt for a sabbatical stay (I0200/111/2024). Authors' contributions Z.P., E.M. and A.M. designed the study, Z.P and D.B. performed analyses. Z.P., D.B., E. M. and A.M. drafted the manuscript, Z.P., and D.B. prepared the figures, A.M. conceptualized the project. Competing interests The author(s) declare that they have no competing interests. References Bessman MJ, Frick DN, O’Handley SF. The MutT Proteins or “Nudix” Hydrolases, a Family of Versatile, Widely Distributed, “Housecleaning” Enzymes *. J Biol Chem. 1996;271:25059–62. Carreras-Puigvert J, Zitnik M, Jemth A-S, Carter M, Unterlass JE, Hallström B, et al. A comprehensive structural, biochemical and biological profiling of the human NUDIX hydrolase family. Nature Communications. 2017;8:1–17. Srouji JR, Xu A, Park A, Kirsch JF, Brenner SE. The evolution of function within the Nudix homology clan. Proteins. 2017;85:775. McLennan AG. The Nudix hydrolase superfamily. Cell Mol Life Sci. 2006;63. Lukaszewicz M. Application of Mammalian Nudix Enzymes to Capped RNA Analysis. Pharmaceuticals (Basel, Switzerland). 2024;17. Liu Y, Zhang W, Wang Y, Xie L, Zhang Q, Zhang J, et al. Nudix hydrolase 14 influences plant development and grain chalkiness in rice. Frontiers in plant science. 2022;13. Mildvan AS, Xia Z, Azurmendi HF, Saraswat V, Legler PM, Massiah MA, et al. Structures and mechanisms of Nudix hydrolases. Arch Biochem Biophys. 2005;433:129–43. Sheikh S, O’Handley SF, Dunn CA, Bessman MJ. Identification and characterization of the Nudix hydrolase from the Archaeon, Methanococcus jannaschii, as a highly specific ADP-ribose pyrophosphatase. The Journal of biological chemistry. 1998;273. Xu W, Gauss P, Shen J, Dunn CA, Bessman MJ. The gene e.1 (nudE.1) of T4 bacteriophage designates a new member of the Nudix hydrolase superfamily active on flavin adenine dinucleotide, adenosine 5’-triphospho-5'-adenosine, and ADP-ribose. The Journal of biological chemistry. 2002;277. Bateman A, Coin L, Durbin R, Finn RD, Hollich V, Griffiths-Jones S, et al. The Pfam protein families database. Nucleic acids research. 2004;32 Database issue. Li W, Godzik A. Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences. Bioinformatics. 2006;22. Li W, Jaroszewski L, Godzik A. Clustering of highly homologous sequences to reduce the size of large protein databases. Bioinformatics. 2001;17. Rozewicki J, Li S, Amada KM, Standley DM, Katoh K. MAFFT-DASH: integrated protein sequence and structural alignment. Nucleic Acids Res. 2019;47:W5–10. Waterhouse AM, Procter JB, Martin DMA, Clamp M, Barton GJ. Jalview Version 2—a multiple sequence alignment editor and analysis workbench. Bioinformatics. 2009;25:1189–91. Berman HM, Westbrook J, Feng Z, Gilliland G, Bhat TN, Weissig H, et al. The Protein Data Bank. Nucleic acids research. 2000;28. Altschul SF, Gish W, Miller W, Myers EW, Lipman DJ. Basic local alignment search tool. J Mol Biol. 1990;215. Minh BQ, Schmidt HA, Chernomor O, Schrempf D, Woodhams MD, von Haeseler A, et al. IQ-TREE 2: New Models and Efficient Methods for Phylogenetic Inference in the Genomic Era. Mol Biol Evol. 2020;37:1530–4. Emms DM, Kelly S. OrthoFinder: phylogenetic orthology inference for comparative genomics. Genome Biol. 2019;20:238. Letunic I, Bork P. Interactive Tree Of Life (iTOL) v5: an online tool for phylogenetic tree display and annotation. Nucleic Acids Res. 2021;49. Mistry J, Chuguransky S, Williams L, Qureshi M, Salazar GA, Sonnhammer ELL, et al. Pfam: The protein families database in 2021. Nucleic Acids Res. 2021;49:D412–9. Wang J, Chitsaz F, Derbyshire MK, Gonzales NR, Gwadz M, Lu S, et al. The conserved domain database in 2023. Nucleic Acids Res. 2023;51. Crooks GE, Hon G, Chandonia JM, Brenner SE. WebLogo: a sequence logo generator. Genome Res. 2004;14. Wootton JC. Non-globular domains in protein sequences: automated segmentation using complexity measures. Comput Chem. 1994;18. Lupas A, Van Dyke M, Stock J. Predicting coiled coils from protein sequences. Science. 1991;252. Erdős G, Pajkos M, Dosztányi Z. IUPred3: prediction of protein disorder enhanced with unambiguous experimental annotation and visualization of evolutionary conservation. Nucleic Acids Res. 2021;49:W297–303. Horton P, Park K-J, Obayashi T, Fujita N, Harada H, Adams-Collier CJ, et al. WoLF PSORT: protein localization predictor. Nucleic Acids Res. 2007;35 Web Server issue:W585. Krogh A, Larsson B, von Heijne G, Sonnhammer EL. Predicting transmembrane protein topology with a hidden Markov model: application to complete genomes. J Mol Biol. 2001;305. Kans J. Entrez Direct: E-utilities on the Unix Command Line. In: Entrez Programming Utilities Help [Internet]. National Center for Biotechnology Information (US); 2024. Hastings J, Owen G, Dekker A, Ennis M, Kale N, Muthukrishnan V, et al. ChEBI in 2016: Improved services and an expanding collection of metabolites. Nucleic Acids Res. 2015;44:D1214–9. Varadi M, Bertoni D, Magana P, Paramval U, Pidruchna I, Radhakrishnan M, et al. AlphaFold Protein Structure Database in 2024: providing structure coverage for over 214 million protein sequences. Nucleic Acids Res. 2023;52:D368–75. van Kempen M, Kim SS, Tumescheit C, Mirdita M, Lee J, Gilchrist CLM, et al. Fast and accurate protein structure search with Foldseek. Nat Biotechnol. 2023;42:243–6. Dolinsky TJ, Czodrowski P, Li H, Nielsen JE, Jensen JH, Klebe G, et al. PDB2PQR: expanding and upgrading automated preparation of biomolecular structures for molecular simulations. Nucleic acids research. 2007;35 Web Server issue. Trott O, Olson AJ. AutoDock Vina: improving the speed and accuracy of docking with a new scoring function, efficient optimization and multithreading. J Comput Chem. 2010;31:455. Eberhardt J, Santos-Martins D, Tillack AF, Forli S. AutoDock Vina 1.2.0: New Docking Methods, Expanded Force Field, and Python Bindings. J Chem Inf Model. 2021;61. O’Boyle NM, Banck M, James CA, Morley C, Vandermeersch T, Hutchison GR. Open Babel: An open chemical toolbox. J Cheminform. 2011;3:1–14. Ivanova L, Karelson M. The Impact of Software Used and the Type of Target Protein on Molecular Docking Accuracy. Molecules (Basel, Switzerland). 2022;27. Bittencourt SA, S. A, Bittencourt a S. FastQC: a quality control tool for high throughput sequence data. https://www.scienceopen.com/document?vid=de674375-ab83-4595-afa9-4c8aa9e4e736. Accessed 18 Jun 2024. Chen S, Zhou Y, Chen Y, Gu J. fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics. 2018;34. Kim D, Langmead B, Salzberg SL. HISAT: a fast spliced aligner with low memory requirements. Nat Methods. 2015;12. Li H, Handsaker B, Wysoker A, Fennell T, Ruan J, Homer N, et al. The Sequence Alignment/Map format and SAMtools. Bioinformatics. 2009;25. Pertea M, Pertea GM, Antonescu CM, Chang TC, Mendell JT, Salzberg SL. StringTie enables improved reconstruction of a transcriptome from RNA-seq reads. Nat Biotechnol. 2015;33. García-Saura AG, Zapata-Pérez R, Martínez-Moñino AB, Hidalgo JF, Morte A, Pérez-Gilabert M, et al. The first comprehensive phylogenetic and biochemical analysis of NADH diphosphatases reveals that the enzyme from Tuber melanosporum is highly active towards NAD+. Sci Rep. 2019;9. Zegarra V, Mais C-N, Freitag J, Bange G. The mysterious diadenosine tetraphosphate (AP4A). microLife. 2023;4:uqad016. Oka K, Suzuki T, Onodera Y, Miki Y, Takagi K, Nagasaki S, et al. Nudix-type motif 2 in human breast carcinoma: a potent prognostic factor associated with cell proliferation. International journal of cancer. 2011;128. Gaspar ML, Pawlowska TE. Innate immunity in fungi: Is regulated cell death involved? PLoS Pathog. 2022;18:e1010460. Palazzo L, Thomas B, Jemth A-S, Colby T, Leidecker O, Feijs KLH, et al. Processing of protein ADP-ribosylation by Nudix hydrolases. Biochem J. 2015;468:293–301. Qi H, Rh GW, Beato M, Price BD. The ADP-ribose hydrolase NUDT5 is important for DNA repair. Cell Rep. 2022;41. Venice F, Desirò A, Silva G, Salvioli A, Bonfante P. The Mosaic Architecture of NRPS-PKS in the Arbuscular Mycorrhizal Fungus Gigaspora margarita Shows a Domain With Bacterial Signature. Front Microbiol. 2020;11. Diet of Arbuscular Mycorrhizal Fungi: Bread and Butter? Trends Plant Sci. 2017;22:652–60. Sokołowska B, Orłowska M, Okrasińska A, Piłsyk S, Pawłowska J, Muszewska A. What can be lost? Genomic perspective on the lipid metabolism of Mucoromycota. IMA Fungus. 2023;14:1–21. Kraszewska E, Drabinska J. Nudix proteins affecting microbial pathogenesis. Microbiology (Reading). 2020;166:1110–4. Sharma A, Tendulkar AV, Wangikar PP. Structure based prediction of functional sites with potential inhibitors to Nudix enzymes from disease causing microbes. Bioinformation. 2011;5. Lee KT, Kwon H, Lee D, Bahn YS. A Nudix Hydrolase Protein, Ysa1, Regulates Oxidative Stress Response and Antifungal Drug Susceptibility in Cryptococcus neoformans. Mycobiology. 2014;42. Rao S, Cao H, O’Hanna FJ, Zhou X, Lui A, Wrightstone E, et al. Nudix hydrolase 23 post-translationally regulates carotenoid biosynthesis in plants. The Plant cell. 2024;36. Additional Declarations No competing interests reported. Supplementary Files SupplementaryFileS1.xlsx SupplementaryFileS2.docx.pdf Cite Share Download PDF Status: Published Journal Publication published 01 Jul, 2025 Read the published version in BMC Genomics → Version 1 posted Editorial decision: Revision requested 25 Apr, 2025 Reviews received at journal 24 Apr, 2025 Reviewers agreed at journal 17 Apr, 2025 Reviews received at journal 13 Apr, 2025 Reviewers agreed at journal 04 Apr, 2025 Reviewers invited by journal 03 Apr, 2025 Editor invited by journal 02 Apr, 2025 Editor assigned by journal 02 Apr, 2025 Submission checks completed at journal 02 Apr, 2025 First submitted to journal 31 Mar, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6343747","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":444495266,"identity":"0fb4db0a-f169-41eb-b2d9-577d2bf1b6d1","order_by":0,"name":"Zofia Pasterny","email":"","orcid":"","institution":"Institute of Biochemistry and Biophysics, Polish Academy of Sciences","correspondingAuthor":false,"prefix":"","firstName":"Zofia","middleName":"","lastName":"Pasterny","suffix":""},{"id":444495267,"identity":"73253d22-dcea-4cf7-be66-6374414bbcdf","order_by":1,"name":"Drishtee Barua","email":"","orcid":"","institution":"Institute of Biochemistry and Biophysics, Polish Academy of Sciences","correspondingAuthor":false,"prefix":"","firstName":"Drishtee","middleName":"","lastName":"Barua","suffix":""},{"id":444495269,"identity":"86f27cd4-2d1f-4d47-b016-a0ccc9d05cb3","order_by":2,"name":"Eugenio Mancera","email":"","orcid":"","institution":"Center for Research and Advanced Studies of the National Polytechnic Institute","correspondingAuthor":false,"prefix":"","firstName":"Eugenio","middleName":"","lastName":"Mancera","suffix":""},{"id":444495270,"identity":"37392a6f-58fe-4bfb-8b6a-6ae4220ddf79","order_by":3,"name":"Anna Muszewska","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAArklEQVRIiWNgGAWjYNCCCokEAyhThkgtZxBaeIjTwdjGQIIW3Rnpjz98nGeRZ87AY/iZh+EOYS1mN3IMDGdukyi2bOAxluZheEaUFoZk3m0SiRsO8G6QzmE4TIyW9AeH/84Ba9n8m0gtCYbNjA1gLduItOXMG2PGnmNAvzTzf7P+Y0CMX44DQ+xHTV2eOXtb8s0ZFXfkCGpBAGYQYXCABB1QQIaWUTAKRsEoGPYAABsqO3PzgcS5AAAAAElFTkSuQmCC","orcid":"","institution":"Institute of Biochemistry and Biophysics, Polish Academy of Sciences","correspondingAuthor":true,"prefix":"","firstName":"Anna","middleName":"","lastName":"Muszewska","suffix":""}],"badges":[],"createdAt":"2025-03-31 10:08:10","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6343747/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6343747/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1186/s12864-025-11778-5","type":"published","date":"2025-07-01T15:57:15+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":81295048,"identity":"f9b0dacd-e784-4592-82ec-8afc805077ba","added_by":"auto","created_at":"2025-04-24 12:52:13","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":833791,"visible":true,"origin":"","legend":"\u003cp\u003eDistribution of the six NUDIX families present in fungi across phyla. The phylogenomic tree of representative fungal isolates spanning the diversity of fungi calculated with Orthofinder was annotated with the distribution of six NUDIX Pfam families. The tree is rooted with the fruit fly \u003cem\u003eD. melanogaster \u003c/em\u003eand the choanoflagellate \u003cem\u003eM. brevicollis\u003c/em\u003e. The figure was prepared in iToL.\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-6343747/v1/e9b1db421bbc23f43fc18b12.png"},{"id":81294544,"identity":"f58f3d20-017a-4722-8e2e-9d35adea2e0f","added_by":"auto","created_at":"2025-04-24 12:44:13","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":296075,"visible":true,"origin":"","legend":"\u003cp\u003eDistribution of the six NUDIX families present in fungi across phyla for 183 proteomes (listed in \u003cstrong\u003eSupplementary file S1\u003c/strong\u003e, sheet S11). Box plots presenting the abundance of NUDIX families across fungal phyla are shown for phyla with more than 2 representatives. In the case of proteins from the canonical NUDIX family (PF00293), the scale of the boxplot goes up to 50, leaving out the Glomeromycota representative \u003cem\u003eGigaspora rosea, \u003c/em\u003ewhich has 101 NUDIX enzymes. All remaining boxplots are scaled between 0 and 4 which fits all the occurrences.\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-6343747/v1/fd4521b58f1b69be083f6f23.png"},{"id":81294545,"identity":"6e4a8f31-a41d-4b61-955e-8ae5b3e4ebd8","added_by":"auto","created_at":"2025-04-24 12:44:13","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":813512,"visible":true,"origin":"","legend":"\u003cp\u003eClustering of fungal NUDIX sequences together with the reference set (human NUDIX sequences and PDB structures from all kingdoms representing the diversity of 10 Pfam NUDIX families). \u003cstrong\u003eA)\u003c/strong\u003e The set of 2,801 fungal sequences identified as NUDIX hydrolases clustered by sequence similarity using CLANS. The fungal NUDIX proteins are depicted as grey nodes, reference sequences (human NUDIX proteins and PDB representatives from the ten NUDIX Pfam families) are highlighted in red. Reference sequences are marked with respective Pfam names on the Figure beside the canonical NUDIX family (PF00293) which represents the remaining enzymes. \u003cstrong\u003eB) \u003c/strong\u003eCloseup of the clustering of proteins from the PF00293 where the newly identified subfamilies are marked by different colors as detailed in the figure legend.\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-6343747/v1/e688aae00880b7e2c05d042d.png"},{"id":81295050,"identity":"ee5542ec-04b6-4eb6-8d96-6ba466c547f1","added_by":"auto","created_at":"2025-04-24 12:52:13","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":808862,"visible":true,"origin":"","legend":"\u003cp\u003eMaximum likelihood phylogenetic trees of selected sequences representing the novel fungal NUDIX clusters along with the NUDIX hydrolases from \u003cem\u003eH. sapiens\u003c/em\u003e (tip labels with prefix ‘human’), and \u003cem\u003eM. musculus \u003c/em\u003e(tip labels with suffix mouse).\u003cem\u003e \u003c/em\u003e\u0026nbsp;Some of the fungal subfamilies are closely related and in such cases they are on the same tree. Fungal sequences are named with the capital letter of the NUDIX subfamily they belong to, followed by the NCBI protein accession and the species name. Bootstrap values for branches are showcased in rectangular boxes.\u003c/p\u003e","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-6343747/v1/a2ac8145444eb295b3cd3e21.png"},{"id":81294547,"identity":"68b0b233-a275-4884-8dd9-98d3e260bfd7","added_by":"auto","created_at":"2025-04-24 12:44:13","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":780961,"visible":true,"origin":"","legend":"\u003cp\u003eProteins of the newly identified subfamilies within the canonical NUDIX family PF00293 show sequence diversity in the NUDIX domain and are widely distributed in the fungal kingdom. \u003cstrong\u003eA)\u003c/strong\u003e Comparison of the NUDIX box motif between newly identified subfamilies. The cells in green indicate compatibility with the motif characteristic for the NUDIX superfamily, while the ones in orange show differences. The cells in white display the variable part of the motif. \u003cstrong\u003eB)\u003c/strong\u003e. Distribution of NUDIX sequences of newly identified subfamilies in the phylogenomic tree of fungi. The tree is the same as in \u003cstrong\u003eFigure 1\u003c/strong\u003e. Colored dots indicate the presence of proteins from a given subfamily in the genome of representative fungal species.\u003c/p\u003e","description":"","filename":"floatimage7.png","url":"https://assets-eu.researchsquare.com/files/rs-6343747/v1/5fbf94351bc56d9e1e47b971.png"},{"id":81294548,"identity":"fba3289b-1d10-4df6-a48e-5097b0bf9840","added_by":"auto","created_at":"2025-04-24 12:44:13","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":99327,"visible":true,"origin":"","legend":"\u003cp\u003eMolecular dockings show similarities in the binding affinities of the newly identified NUDIX fungal subfamilies. Molecular docking results for newly identified subfamilies against 12 substrates clustered by similarity between subfamilies. Results represent relative binding energy [kcal/mol] and are normalized by Z-scores with lower values indicating more favorable interactions as previously described (Xue et al. 2022).\u003c/p\u003e","description":"","filename":"floatimage8.png","url":"https://assets-eu.researchsquare.com/files/rs-6343747/v1/c25a0380b914671353bf02e9.png"},{"id":81294551,"identity":"84e85a05-8c6f-41bc-b50e-0584531bbfed","added_by":"auto","created_at":"2025-04-24 12:44:13","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":536787,"visible":true,"origin":"","legend":"\u003cp\u003ePredicted NUDIX enzymes, including those from newly identified subfamilies, are expressed across fungi. Evidence of gene expression for NUDIX families and the newly identified subfamilies in available transcriptomic datasets. Squares are only shown if the family/subfamily is present in the given taxa and they are filled if there is evidence of expression. Subfamily F is not present in the available transcriptomes.\u003c/p\u003e","description":"","filename":"floatimage9.png","url":"https://assets-eu.researchsquare.com/files/rs-6343747/v1/07a589abe79be141cb9e60e0.png"},{"id":86179875,"identity":"2e7848ba-0bff-4d2f-b494-f08b5effd91b","added_by":"auto","created_at":"2025-07-07 16:20:12","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":5416124,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6343747/v1/3b39b6eb-7f26-4adf-81b2-d8fd01c3de4d.pdf"},{"id":81296173,"identity":"eed22050-d0c3-4b6b-9353-a5c77286fd11","added_by":"auto","created_at":"2025-04-24 13:00:13","extension":"xlsx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":640251,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryFileS1.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-6343747/v1/15b80872e033fec17ab4e0cd.xlsx"},{"id":81294572,"identity":"6f03733c-a64c-4edb-8420-c2b855ab5755","added_by":"auto","created_at":"2025-04-24 12:44:14","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":5278668,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryFileS2.docx.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6343747/v1/47d957726fe740c0c83ddd32.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Analysis of NUDIX enzymes across fungi reveals previously unrecognized diversity","fulltext":[{"header":"Background","content":"\u003cp\u003eThe NUDIX superfamily of enzymes encompasses various organic pyrophosphatases capable of cleaving nucleoside diphosphates linked to any moiety into nucleoside monophosphates and organic pyrophosphates, hence their name. Initially, the superfamily was called the MutT family after its first representative, the 8-oxoGTPase central for \u003cem\u003eE. coli\u003c/em\u003e DNA repair [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. However, apart from their role in DNA repair, nowadays these enzymes are known to be involved in the hydrolysis of detrimental metabolites, decapping, processing of mRNA, cell cycle regulation, and survival [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e], and gating ion channels [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. In agreement with the wide range of functions that they perform, the diversity of the substrates of these enzymes is high, including di- and triphosphates, nucleotide sugars, dinucleosides, diphosphoinositol polyphosphates, and RNA caps [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. The biomedical and biotechnological potential of NUDIX proteins is starting to be recognized. Human Dcp2, NUDT2, NUDT12, and NUDT16 can be used to characterize 5\u0026prime; capped RNA transcripts [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e], while OsNUDX14 from \u003cem\u003eOryza sativa\u003c/em\u003e is thought to be a grain quality regulator and to determine plant development, being strongly expressed in mature leaves [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. NUDIX enzymes with feasible druggable sites have also been identified, and therefore they could represent novel drug targets for human diseases, especially in cancer cells [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e].\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eCharacteristics of the NUDIX Pfam families regarding the most common co-occuring domains, human and baker\u0026rsquo;s yeast representatives, function, and presence across kingdoms of life.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"7\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePfam\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMost common coexisting domains\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cem\u003eH. sapiens\u003c/em\u003e representatives\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cem\u003eS. cervisiae\u003c/em\u003e representatives\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eAvailable PDB structures\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eFunction\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003ePresence (kingdoms)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePF00293\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ein majority of the cases exists as a single domain\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNUDT1, NUDT2, NUDT3, NUDT4, NUDT4B, NUDT5, NUDT6, NUDT7, NUDT8, NUDT9, NUDT10, NUDT11, NUDT12, NUDT13, NUDT14, NUDT15, NUDT16, NUDT17, NUDT18, NUDT20, IDI1, IDI2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eDCP2, DDP1, IDI1, NPY1, PCD1,YJR142W, YJR142W, YSA1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eAmong others:\u003c/p\u003e \u003cp\u003eH. sapiens: NUDT1, NUDT2, NUDT7, NUDT12, NUDT15, NUDT16, IDI1 and IDI2, DDP1, DDP2, DPP3-alpha, ADP-ribose pyrophosphatase, Adenine DNA glycosylase, mitochondrial 39S ribosomal protein L46;\u003c/p\u003e \u003cp\u003eE. coli: nudE, NADH pyrophosphatase, 7,8-dihydro-8-oxoguanine triphosphatase, nudF.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003evarious functions, but the most prominent one is sanitization of nucleotide pool\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eall kingdoms\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePF03559\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ein majority of the cases exists as a single domain\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eEvaA 2,3-dehydratase (A. orientalis)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003einvolved in antibiotic production pathways\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003emainly bacteria\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePF09296\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ezf-NADH-PPase (PF09297), NUDIX domain (PF00293)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNUDT12, NUDT13\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003ePeroxisomal NADH pyrophosphatase NUDT12 (H. sapiens, M. musculus)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003edeNADing NAD-capped RNA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eeukaryotes and bacteria\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePF11788\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ein majority of the cases exists as a single domain or along cannonical NUDIX domain (PF00293)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMRLP46\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMitochondrial ribosomal protein L46\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eMitochondrial ribosomal protein L46 (H. sapiens, S. scrofa, S. cerevisiae, T. brucei, L. major, N. crassa)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003erole in mitochondrial protein synthesis\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eeukaryotes\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePF13869\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ein majority of the cases exists as a single domain\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNUDT21\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eCleavage and polyadenylation specificity factor subunit 5 (H. sapiens)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e3' RNA cleavage and polyadenylation processing\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eeukaryotes\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePF14815\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eHhH-GPD superfamily base excision DNA repair protein (PF14815)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMUTYH\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eAdenine DNA glycosylase (M. musculus, G. stearothermophilus, H. sapiens), MutT protein (B. bacteriovorus, B. fragilis, S. aureus, E. coli, B. henselae), 7,8-dihydro-8-oxoguanine-triphosphatase (K. pneumoniae)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eexcises inappropriately matched adenine from DNA backbone to either 8-oxoguanine or guanine\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eeukaryotes and bacteria\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePF15916\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNUDIX domain (PF00293)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eMutT/nudix family protein (R. rubrum ATCC 11170)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003efunctionally uncharacterised, takes part in thiamine biosynthetic process\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eeukaryotes and bacteria\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePF16262\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ein majority of the cases exists as a single domain\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003ePutative uncharacterized protein (J. denitrificans)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eunknown\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003ebacteria\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePF16705\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eFERM central domain (PF00373)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eKrev interaction trapped protein 1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eKrev interaction trapped protein 1 (H. sapiens)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eactivates β1 integrin by antagonization of ICAP1 (Integrin Cytoplasmic Associated Protein-1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eeukaryotes (only animals)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePF18290\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNUDIX domain (PF00293)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNUDT6 (H. sapiens)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eNUDT7 (A. thaliana), NUDT6 (H. sapiens)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003ehydrolases catalyze the hydrolysis of nucleoside diphosphates which are often toxic metabolic intermediates and signalling molecules\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eeukaryotes and bacteria\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eDespite the functional differences between these enzymes, at the structural level, all NUDIX hydrolases are characterized by an α-β-α sandwich structure with a specific NUDIX motif, which contains the catalytic site and metal-binding residues. The NUDIX motif is composed of 23 amino acids: Gx\u003csub\u003e5\u003c/sub\u003eEx\u003csub\u003e7\u003c/sub\u003eREUxEExGU where U is a bulky amino acid (typically isoleucine, leucine, or valine), and x represents any amino acid. In addition, the N-terminus of the helix fold has glutamates necessary for binding divalent cations that link the NUDIX hydrolases with the pyrophosphate of the substrate. In terms of the reaction, nucleophilic substitution by water occurs at particular phosphorus atoms within a diphosphate or polyphosphate chain [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eNUDIX enzymes are ubiquitous in all kingdoms of life and viruses [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e, \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e] although these proteins are evolutionarily related, they exhibit significant sequence divergence. For this reason the superfamily has been divided into ten Pfam families. Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e presents the characteristics of each family. In fungi there are 29 NUDIX enzymes reported at the UniProt database that are part of five of the ten families. All of these records belong to Dikarya taxons and mostly to the model yeasts \u003cem\u003eSaccharomyces cerevisiae\u003c/em\u003e and \u003cem\u003eSchizosaccharomyces pombe\u003c/em\u003e. Therefore, the phylogenetic distribution and functional roles of NUDIX enzymes in fungi remain poorly understood. In this work, we identified 25 subfamilies belonging to the NUDIX superfamily throughout the fungal kingdom, 13 of which are newly identified and several appear to be only present in fungi. NUDIX enzymes are particularly abundant in the phylum Glomeromycota. The newly identified subfamilies possess characteristics typical of NUDIX hydrolases such as α-β-α sandwich structure and all belong to the canonical NUDIX family. Molecular docking suggested particular substrate preferences and mining available transcriptomic data provided evidence of their physiological relevance. Overall, our work highlights the extensive diversity of NUDIX enzymes in fungi and positions these organisms as a potential source of newly identified NUDIX enzymes for biotechnological, antifungal or research applications.\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eIdentification of NUDIX proteins across 183 fungal proteomes\u003c/h2\u003e \u003cp\u003eTo identify NUDIX enzymes across the fungal tree, we first searched all Pfam families in a set of 183 fungal proteomes (\u003cb\u003eSupplementary file S1\u003c/b\u003e, sheet S1) by amino acid sequence similarity using pfam_scan.pl with default settings against Pfam database v. 36 [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. Then, from the full set of Pfam family assignments of the 183 fungal proteomes, we extracted all proteins that matched the NUDIX families grouped in the CL0261 clan. The redundancy across the identified NUDIX sequences was first reduced with cd-hit (70% sequence identity, 90% coverage, 5 letters word length) [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e]. Full protein sequences were then aligned with MAFFT v7.40 using the iterative local alignment mode [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e] and trimmed manually to the domain region using Jalview [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eAs a reference set of NUDIX enzymes, all human proteins with NUDIX domains were retrieved from the UniProt database (status reviewed) and all NUDIX domain-containing proteins with resolved structures were obtained from PDB [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]. The set of human and PDB sequences will be referred to as the r\u003cem\u003eeference sequence set\u003c/em\u003e. The PDB sequences of NUDIX hydrolases include proteins from all kingdoms and viruses. Together, the reference sequence set contains representatives of all ten Pfam NUDIX families.\u003c/p\u003e \u003cp\u003eThe identified fungal NUDIX domain sequences together with the reference sequence set and Pfam consensus sequences were then used as queries in a blastp v. 2.13.0\u0026thinsp;+\u0026thinsp;search [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e] against the same set of 183 proteomes using an e-value cutoff of 0.001. The blastp resulted in 7,394 fungal sequences, out of which 2,801 turned out to belong to the NUDIX superfamily. The remaining blastp hits contain domains that often coexist with the NUDIX domain (see \u003cb\u003eSupplementary file S1\u003c/b\u003e, sheet S8 to see the list of these domains) and were identified with blastp by the full-length sequences of the reference set queries. The set of 2,801 fungal NUDIX hits together with the reference sequences was subjected to clustering based on amino acid sequence similarity using CLANS (with p-value of 1e\u0026thinsp;\u0026minus;\u0026thinsp;10, attraction exponent value of 2 and remaining parameters set to default values).\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003ePhylogenetic analysis\u003c/h3\u003e\n\u003cp\u003eTo trace the evolutionary relationships between NUDIX subfamilies, representative sequences from related clusters were aligned with mafft (--localpair \u0026ndash;maxiterate 100, v7.407 (Rozewicki et al., 2019). The resulting alignment was then trimmed manually in Jalview and used for phylogenetic analyses. The phylogenetic tree was inferred for each of the alignments by maximum likelihood with IQTREE2 [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e], considering 1000 bootstrap replicates (parameters: -B 1000) with automated model selection. The phylogenetic tree at the species level was generated with OrthoFinder [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e] (parameters: mmseqs and dendroblast) for 41 species spanning the taxonomic diversity of the whole fungal kingdom. The Choanoflagellate \u003cem\u003eMonosiga brevicollis\u003c/em\u003e and the fruit fly \u003cem\u003eDrosophila melanogaster\u003c/em\u003e were used as outgroups. The species tree was used to represent the taxonomic distribution of individual protein families and subfamilies across the diverse fungal lineages. Trees were visualized and rendered in iToL [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e].\u003c/p\u003e\n\u003ch3\u003eSequence and structural characterization identified subfamilies\u003c/h3\u003e\n\u003cp\u003eTo characterize the sequence and structural properties of all fungal NUDIX subfamilies, all of the sequences were scanned against all Pfam family definitions with pfamscan.pl [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e] to confirm the presence of the NUDIX domain (PF00293), and look for other coexisting domains. Where available, we used NCBI Conserved Domain Database (CDD) [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e] annotations for the subfamilies, since CDD enables the identification of conserved domains across multiple protein sequences and defines a more detailed classification of subfamilies of the NUDIX proteins than the Pfam database. Complete protein sequences were then aligned with MAFFT v7.407 [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e] and trimmed to the domain region using Jalview [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e] to construct logos of the NUDIX signature motif with WebLogo 3 [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e]. Low complexity regions were predicted with the SEG tool [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e], coiled-coil regions with EMBOSS pepcoil [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e], disordered protein regions with IUPred3 [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e], subcellular localization with WoLF PSORT [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e], and transmembrane regions were identified with TMHMM [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]. For each of the newly identified subfamilies, we also retrieved annotations of upstream and downstream genes with Entrez Direct [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e] in order to verify their possible involvement in metabolic clusters.\u003c/p\u003e \u003cp\u003eIn order to verify the taxonomic representation of the new subfamilies, we extended the sequence search by using the newly identified fungal NUDIX sequences d as queries in a blastp v. 2.13.0\u0026thinsp;+\u0026thinsp;search [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e] against the NCBI non-redundant (NR) database with an e-value cutoff of 0.001 to collect all homologs. This set of sequences was mapped on the NCBI-Taxonomy database using the taxdump files (nodes and names) to extract their taxonomic assignment. We then calculated the percentage of fungal sequences in all of the subfamilies.\u003c/p\u003e\n\u003ch3\u003eMolecular docking for substrate identification\u003c/h3\u003e\n\u003cp\u003eIn order to predict the potential specificities of each of the newly identified subfamilies, molecular docking was performed. Each of the subfamilies was represented by two selected sequences and their AlphaFold2 3D structure models [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e, \u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e]. These structures were then scanned against the FoldSeek database to find the most similar structures [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]. PyMOL was used to visualize and compare the structure of the representatives of the newly identified subfamilies with available PDB structures of NUDIX enzymes (The PyMOL Molecular Graphics System, Version 3.0 Schr\u0026ouml;dinger, LLC). Structures were trimmed to the NUDIX domain region in PyMOL and prepared for the molecular docking using pdb2pqr (parameters: --ff AMBER --ph-calc-method\u0026thinsp;=\u0026thinsp;propka --with-ph\u0026thinsp;=\u0026thinsp;xx) [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e]. Molecular docking with AutoDock Vina v1.2.5 [\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e, \u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e] was performed for the twelve most common NUDIX substrates: 8-oxo-dGDP, 8-oxo-dGTP, ADP-glucose, ADP-ribose, Ap3A, Ap4A, Ap6A, dATP, DHNTP, HSCoA, NAD\u0026thinsp;+\u0026thinsp;and NADPH. The substrates were obtained from the ChEBI database in SMILES format [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e] and converted into three-dimensional PDBQT structures with Open Babel 3.1.1 [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e] (parameters: \u003cem\u003e--gen3D -p xx -opdbqt\u003c/em\u003e). AutoDock Vina predicts the receptor-ligand binding conformations and ranks them according to the predicted change in free energy during binding [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e]. For all the proteins and ligands tested, we ensured they were at the same pH and protonation state to ensure the biological relevance of the predictions. As the NUDIX hydrolases reach their maximum activity in a mildly basic environment we consider the protomers of the tested representatives at pH 7.0, 8.0 and 9.0.\u003c/p\u003e\n\u003ch3\u003eAnalysis of gene expression\u003c/h3\u003e\n\u003cp\u003eExpression of the genes encoding the NUDIX proteins was verified using publicly available transcriptomic datasets of fungi. The data contained a total of 25 transcriptomes (six Mucoro- and Mortierellomycota species, four Glomero- and Basidiomycota, and five Ascomycota datasets). Description of transcriptome datasets, references and results of the gene expression analyses are available in \u003cb\u003eSupplementary file S1\u003c/b\u003e, sheets S5, S6, S7. The data files were obtained from the ENA server in fastq format, and their quality was checked with FASTQC (v0.11.8) [\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e]. Adapters were trimmed using fastp (v0.19.6) with default settings [\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e]. The adapter-trimmed reads were aligned with reference genomes of the respective organisms (downloaded from NCBI Datasets) using Hisat2 (v2.1.0) [\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e]. The SAM alignment files were compressed into binary format (BAM) using samtools (v.10) [\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e]. The aligned reads were mapped with GFF files using StringTie (v2.1.3b) [\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e] to generate read abundance in the form of Transcripts per Million (TPM) values. The average TPM was calculated for each dataset and values\u0026thinsp;\u0026lt;\u0026thinsp;1 were regarded as statistically not significant and removed. The remaining TPM values were then subjected to Z-score normalization. Finally, NUDIX enzymes from the identified families and subfamilies were searched in the expression dataset of protein-coding genes. A schematic representation of the workflow is shown in \u003cb\u003eSupplementary Figure \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e\u003c/b\u003e.\u003c/p\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003eNUDIX enzymes are widespread across the fungal kingdom\u003c/h2\u003e \u003cp\u003eTo identify NUDIX enzymes across fungi, we scanned a set of 183 fungal proteomes from all the fungal phyla for the presence of representatives of the ten NUDIX families included in the Pfam database. Six out of the ten families had homologs in the analyzed proteomes (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e), and two of the four missing families (PF03559 and PF16262) have also not been found in other eukaryotes. The canonical NUDIX domain family (PF00293) was by far the most common family present in 182 out of the 183 proteomes. This family was the only one identified in all phyla with a median number of proteins per proteome ranging from 2 for parasitic Microsporidia to more than 14 in Basidiomycota and Mucoromycota. The presence of the remaining families was less uniform (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e): the DUF4743 family (PF15916) was present in 132 proteomes, the 39S mitochondrial ribosomal protein L46 family (PF11788) in 119, the NADH pyrophosphatase-like rudimentary family (PF09296) in 113, the NUDIX_4 domain family (PF14815) in 85, and the NUDIX_2 domain family (PF13869) in 84 proteomes. These five families had a median of close to one representative per phylum and some families were completely absent in several phyla.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eOverall, the phyla Mucoromycota, Glomeromycota, and Mortierellomycota contain the highest number of NUDIX enzymes, in particular from the canonical family (PF00293). The highest abundance was observed in species from Glomeromycota, including species such as \u003cem\u003eGigaspora rosea\u003c/em\u003e (101 NUDIX proteins), \u003cem\u003eGigaspora margarita\u003c/em\u003e (47 NUDIX proteins), and \u003cem\u003eDiversispora epigaea\u003c/em\u003e (50 NUDIX proteins). The phyla Ascomycota, Basidiomycota, Chytridiomycota, Mortierellomycota, and Mucoromycota have representatives of all six families present in fungi. Glomeromycota and Zoopagomycota do not have members of the families PF14815 and PF09296, Kickxellomycota lacks PF14815 NUDIX hydrolases, Microsporidia has only representatives of PF00293 and PF13869 NUDIX proteins and Neocallimastigomycota are missing proteins of the PF11788 family. Our results show that NUDIX enzymes are widespread among fungi although not all known families are present in these organisms.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eFungal NUDIX homologs can be clustered into 25 subfamilies\u003c/h3\u003e\n\u003cp\u003eTo better understand the diversity of fungal NUDIX enzymes, we clustered them according to sequence similarity together with a well-curated reference set that includes representative enzymes from all kingdoms and from all the ten NUDIX Pfam families (see Methods for details). Clustering distinguished 25 groups of fungal proteins with the NUDIX domain. On the other hand, NUDIX reference sequences from four NUDIX families (PF03559, PF16262, PF16705, and PF18290) did not cluster with any fungal protein. This is in agreement with our initial finding that only six NUDIX families are present among fungi.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe most abundant NUDIX family in fungi, the canonical NUDIX family (PF00293) formed a large centrally located cluster (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). In this cluster, we found homologs of nine human NUDIX proteins (IDI1 and IDI2, NUDT1, NUDT3, NUDT5, NUDT7 and NUDT8, NUDT15 and NUDT20), while ten human NUDIX hydrolases did not cluster close to any fungal homolog (NUDT4, NUDT4B, NUDT6, NUDT9, NUDT10, NUDT11, NUDT14, NUDT16, NUDT17, NUDT18, KRIT1). The latter human proteins likely have no fungal orthologs. Fungal homologs of the other five NUDIX families formed small clusters with the corresponding representatives from the reference set. \u003cb\u003eSupplementary Figure \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003e\u003c/b\u003e presents the phylogenomic tree of representative fungal isolates (spanning the fungal tree of life) annotated with the distribution of NUDIX homologs of characterized NUDIX proteins from the reference sequence set.\u003c/p\u003e \u003cp\u003eThe clustering results were further assessed by generating maximum likelihood phylogenetic trees with representative sequences from selected clusters. Sequence diversity in the whole PF00293 family does not allow for reliable multiple sequence alignments so we built trees for sets of related clusters within the family. The phylogenetic trees also allowed drawing hypotheses about the evolution of specific NUDIX enzymes. For example, we found that four human NUDIX hydrolases\u0026mdash;NUDT4, NUDT4B, NUDT10, and NUDT11\u0026mdash;are paralogs of NUDT3. In general, these findings show that there is considerable diversity of NUDIX enzymes in fungi, with representatives of 25 different subfamilies.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003e13 newly identified subfamilies of NUDIX hydrolases\u003c/h2\u003e \u003cp\u003eThirteen fungal NUDIX clusters formed by members of the canonical NUDIX domain family (PF00293) did not contain human or PDB representatives and, therefore, we consider them as newly identified subfamilies. Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e summarizes the characteristics of these subfamilies, \u003cb\u003eSupplementary file S1\u003c/b\u003e, sheet S2 provides a full list of proteins associated with each subfamily, and a detailed description of each of the newly identified subfamilies can be found in \u003cb\u003eSupplementary file S2\u003c/b\u003e, their distribution in fungal phyla is summarized in \u003cb\u003eSupplementary Figure S3.\u003c/b\u003e\u003c/p\u003e \u003cp\u003eThe enzymes of the newly identified subfamilies showed considerable diversity in terms of their sequence and structure (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e and \u003cb\u003eSupplementary file S2\u003c/b\u003e). For example, 40% of the members from subfamily H are predicted to have regions embedded in the cell membrane, and genes encoding proteins from subfamily K are overrepresented in retrotransposons. In terms of cellular localization, most of the members of the new subfamilies are predicted to be localized in the nucleus.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eCharacteristics of newly identified fungal NUDIX subfamilies. From left to right, the columns correspond to the assigned subfamily name, the number of proteins in the subfamily, percentage of proteins with low complexity regions, coiled coil regions, percentage of the positions of proteins predicted as being disordered, fraction of proteins with predicted transmembrane helices, and the taxa in which the subfamily is present. All the values correspond to the NUDIX blastp search across 183 fungal proteomes.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"7\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSubfamily\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003e# proteins\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003e% low complexity regions\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003e% coiled coil regions\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003e% positions disordered\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003e% transmembrane helices\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eDistribution\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eA\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e84\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e98%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e43%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e51%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eFungi\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eB\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e20\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e85%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e10%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e8%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eFungi (83.77%), archea and bacteria\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eC\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e32\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e69%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e5%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e9%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eFungi\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eD\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e105\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e70%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e18%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e9%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e2%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eFungi (84%), Protozoa, plants\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eE\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e36\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e89%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e8%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e23%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eFungi\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eF\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e36\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e67%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e6%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e13%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eFungi\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eG\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e140\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e47%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e13%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e3%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eBacteria (48.72%), Fungi (45.80%), Insects\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eH\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e152\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e73%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e14,35%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e39%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eFungi (75.50%), Bacteria\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eI\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e57\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e61%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e25%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e10,01%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eFungi\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eJ\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e40\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e38%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e10%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e11,77%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eFungi\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eK\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e14\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e86%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e57%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e18%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eFungi (75%), Bacteria\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eL\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e19\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e79%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e37%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e9%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e5%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eFungi\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eM\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e22\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e50%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e14%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e14,35%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eFungi (97.03%), Bacteria\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eAll newly identified subfamilies have a well-preserved NUDIX signature motif with at least two glutamates acting as metal ligands and lysine and arginine residues that strengthen catalytic forces [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e] (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eA). However, we did observe several substitutions within the NUDIX motif of the newly identified subfamilies, explaining in part why they clustered apart from previously known NUDIX enzymes. High sequence diversity within the NUDIX box is not unprecedented. This is well known for members of the PF13869 family and even the loss of the NUDIX box has been documented for other NUDIX representatives such as members of the PF14815 family [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e] (\u003cb\u003eSupplementary Figure S4\u003c/b\u003e).\u003c/p\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eB shows the distribution of the newly identified subfamilies across a set of genomes representing all the phyla of the fungal kingdom. Beside Rozellomycota and Blastocladiomycota, all fungal phyla have representatives of several of the newly identified subfamilies. Glomeromycota are particularly enriched with these proteins, having representatives of 11 out of 13 subfamilies. Moreover, proteins of the newly identified subfamilies are rare outside fungi, with seven of the subfamilies (A, C, E, F, I, J, L) being specific to fungi. Furthermore, the remaining subfamilies that also have representatives in other kingdoms have less than 10% of the total number of proteins from taxa that are not fungi. The former with the exception of subfamily G, for which half of the sequences in the subfamily belong to Terrabacteria representatives. Several subfamilies have very narrow distribution across the fungal kingdom. Subfamilies C, F, I, J, K, and J have representatives only in the subphylum Diversiporales, and subfamilies D and E are found only in Mortierellomycota. Even subfamilies with such narrow taxonomic distribution show sequence diversity (\u003cb\u003eSupplementary Figure S5, Supplementary Figure S6\u003c/b\u003e). In addition, proteins from subfamilies A and M are most likely distantly related to human NUDT2 (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eA, \u003cb\u003eSupplementary Figure S7, Supplementary Figure S8\u003c/b\u003e), which is known to regulate ApNA levels associated with stress signaling pathways and protect against oxidative damage to nucleotides.\u003c/p\u003e \u003cp\u003eOverall, the clustering with CLANS, the phylogenetic position, and differences in the NUDIX box motif, all support the separation of the newly identified subfamilies. Even the closest newly identified subfamilies J and K, which cluster together in CLANS, form sister clades in the phylogenetic tree and display differences in the NUDIX box motif despite their high overall sequence similarity (32.56%).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eProteins of the newly identified fungal NUDIX subfamilies share binding affinity towards specific substrates\u003c/h2\u003e \u003cp\u003eNUDIX hydrolases are known to have a broad spectrum of substrates. To investigate the substrate preferences of the newly identified subfamilies, we performed molecular docking with representatives of each subfamily. The average affinity per subfamily for twelve different substrates is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e and \u003cb\u003eSupplementary Figure S9\u003c/b\u003e. Through all subfamilies, the highest binding affinity was observed towards NADPH, NAD+, Ap3A, and Ap4A. Despite that DHNTPase residues are conserved within the NUDIX box in the proteins from subfamily B (\u003cb\u003eSupplementary Figure S10\u003c/b\u003e), even these proteins show preference for Ap3A, and Ap4. These substrates are typical of human NUDT12 and NUDT13 which belong to the NUDIX family of NADH pyrophosphatase-like rudimentary NUDIX domain (PF09296). However, all sequences from the 13 newly identified subfamilies are members of the canonical NUDIX family (PF00293). Moreover, none of them possess the architecture nor the SQPWPFPXS motif that are typical of NADH pyrophosphatases and diphosphatases of the PF09296 family [\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e]. Therefore, it is possible that the affinity of the newly identified subfamilies has converged in parallel towards the same substrates. Subfamily C also draws attention since, unlike the rest of the subfamilies, its members show relatively low affinity to Ap4A and high affinity to ADP-ribose. [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. The latter specificity is known in proteins involved in post-translational modification removal (NUDT16) and DNA repair (NUDT5). Even when empirical validation is needed to determine the exact function of these newly identified subfamilies, our docking results allow formulating predictions that could guide experimentation.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cb\u003eMost of the newly identified\u003c/b\u003e NUDIX \u003cb\u003eproteins are expressed\u003c/b\u003e\u003c/p\u003e \u003cp\u003eTo start understanding the functional role of the predicted NUDIX proteins including those from the newly identified subfamilies, we searched for evidence of expression of the genes that encode them in available transcriptomic data. Genome-wide gene expression data is widely available for Dikarya, but it is much more scarce for other fungi. Therefore, we only considered nine representative datasets for Dikarya taxa where the predicted NUDIX enzymes are present, while we included all available datasets in the NCBI SRA database for other fungi. In total, these represented transcriptomes of 25 species (\u003cb\u003eSupplementary file S1\u003c/b\u003e, sheets S7 \u0026amp; S8). In these datasets, a gene was considered expressed if it had an average TPM value\u0026thinsp;\u0026ge;\u0026thinsp;1 (Methods).\u003c/p\u003e \u003cp\u003eOverall, we found evidence of gene expression for 82% of all the predicted NUDIX enzymes in the 25 species with available RNASeq data. The fraction of expressed NUDIX enzymes varied between families, from 97% in the PF11788 family to 81% in the PF00293 family. At the species level, transcriptomes showed expression of a variable subset of known NUDIX members, from 6 (out of 14) in \u003cem\u003eCryptococcus neoformans\u003c/em\u003e and 6 (out of 10) \u003cem\u003eEndogone sp.\u003c/em\u003e FLAS-F59701 to 104 (out of 130) in \u003cem\u003eGigaspora rosea\u003c/em\u003e. Apart from \u003cem\u003eRhizophagus irregularis\u003c/em\u003e which had only members of the canonical NUDIX family expressed, all other species had evidence of expression of at least four of the six families (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eFor proteins from the newly identified subfamilies, overall, we found evidence of expression for 71% of them. The former despite the fact that these subfamilies have a narrower taxonomic distribution and, in consequence, the available expression datasets of taxa that have them in their genomes is considerably limited. In total, we observed the expression of at least one member from each of the thirteen newly identified subfamilies, except for subfamily L. Transcriptomes of species of Glomeromycota showed evidence of expression of eight of the newly identified subfamilies (A, B, C, D, H, I, J, and K), followed by Mortierellomycota with six (A, B, D, E, G and H), Ascomycota with five (B, D, G, H and M), Mucoromycota with four (A, D, G and H), and Basidiomycota with three (G, H and M) (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003e). For subfamilies I, J and K there was evidence of expression only in one single species. Moreover, the NUDIX genes analysed in the differential expression experiments on \u003cem\u003eR. delemar\u003c/em\u003e and \u003cem\u003eG. rosea\u003c/em\u003e were expressed across tested conditions pointing at a housekeeping function of many NUDIX hydrolases (\u003cb\u003eSupplementary file S1\u003c/b\u003e, sheets S7 \u0026amp; S8). In summary, mining of available transcriptomes showed that the NUDIX enzymes that we predicted are widely expressed in fungi suggesting that they may be of biological importance for these organisms.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eEven when NUDIX enzymes were first described more than 70 years ago and despite the many biological functions that they perform, these proteins have been barely studied in fungi, especially beyond Dikarya. In this work, we showed the richness of the NUDIX superfamily across fungi and discovered new subfamilies unique to these organisms. There are only 29 fungal NUDIX enzymes with status reviewed at the UniProt database. These belong to five of the ten known NUDIX families (22 canonical NUDIX domain (PF00293), one NADH pyrophosphatase-like rudimentary NUDIX domain (PF09296), three 39S mitochondrial ribosomal protein L46 (PF11788), one NUDIX_4 domain (PF14815), and two DUF4743 domain (PF15916)). Our work significantly expands the currently known catalog of NUDIX hydrolases in fungi by identifying 2,801 candidate enzymes, including those listed in UniProt as well as enzymes from an additional family not reported for fungi in this database. We did not find representatives of four NUDIX families in the 183 fungal genomes that were analyzed, which was consistent with the fact that there are no fungal NUDIX proteins associated with these families in Uniprot or Pfam. Furthermore, all of the reviewed records of fungal NUDIX hydrolases in UniProt belong to Ascomycota species, while we showed that these enzymes are present across all fungal phyla, with Glomeromycota possessing the highest number and diversity of these enzymes in fungi.\u003c/p\u003e \u003cp\u003ePerforming detailed sequence similarity analyses we found that the fungal enzymes clustered in 25 subgroups, and that thirteen represent newly identified subfamilies as their members did not cluster with any previously known NUDIX enzyme. All of the newly identified subfamilies belong to the well-known canonical NUDIX family (PF00293). We found evidence that many of the genes that encode the newly identified subfamilies are expressed, suggesting physiological relevance. However, given the many functions that NUDIX enzymes perform, the specific cellular role of these enzymes is difficult to predict based on sequence properties. We observed common substrate preference for diadenosine triphosphate and tetraphosphate (Ap3A and Ap4A) in these enzymes. The level of both diadenosine metabolites is known to increase upon exposure to various stresses in bacteria and humans. Therefore they are considered to be alarmones, triggering a stress response that [\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e] influences a variety of cellular processes such as gene expression and DNA repair [\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e, \u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e] Therefore, it is possible that the enzymes of the newly identified subfamilies may modulate the response to cellular stress in fungi. Alternatively, they could also be involved in metabolic pathways for nucleotide synthesis and degradation. Ap3A and Ap4A signaling in animals is needed for apoptosis and development. Programmed cell death types differ among fungi [\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e] including apoptosis-like process, heterokaryon incompatibility, and ferroptosis which is a consequence of iron release and lipid peroxidation. The diversity of fungal developmental strategies, death types, and lipids could also depend, at least partially, on the diversification of the repertoire of NUDIX enzymes.\u003c/p\u003e \u003cp\u003eSome of the subfamilies showed preference for ADP-ribose which can also be related with DNA repair or signalling regulation. ADP-ribosylation is often found as a posttranslational modification of proteins [\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e]. On the other hand poly-ADP-ribose can accumulate when DNA is damaged [\u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e]. There are known ADP-ribose processing NUDIX enzymes that can be remove it from proteins reversing glycosylation (NUDT16), and mono ADP-ribose can be converted to ribose-5-phosphate plus AMP/ATP (NUDT5). The newly identified subfamily C may extend this list of ADP-ribose processing NUDIX proteins.\u003c/p\u003e \u003cp\u003eNUDIX enzymes from the newly identified subfamilies are particularly abundant in Glomeromycota. Fungi in this phylum are known to lack several DNA repair and replication genes [\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e] which could be associated with the abundance of NUDIX enzymes. Alternatively, NUDIX enzymes in Glomeromycota could be involved in lipid metabolism since the phylum is also known to be unique in this respect [\u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e] and NUDIX enzymes involved in beta-oxidation of fatty acids are widespread in fungi [\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eIn terms of the evolutionary origin of fungal NUDIX enzymes, members of all of the six NUDIX families that we identified in fungi are also present in animals. On the other hand, animals have representatives of two families that we did not find in fungi (PF16705 and PF18290). Since other eukaryotic taxa do have members of the Nudix_hydro PF18290 which is a N-terminal domain preceding the canonical PF00293 domain, its absence in fungi probably represents a lineage-specific loss. On the other hand, the PF16705 \u0026lsquo;Nudix or N-terminal NPxY motif-rich region of KRIT\u0026rsquo; is limited to animals and likely represents an evolutionary novelty. Recently, during Pfam 37.1 release, a novel NUDIX domain (PF22327, Nudt16-like) was introduced for U8 snoRNA-decapping enzyme-like proteins from animals which is likely a second NUDIX family limited to this kingdom. Therefore, the NUDIX repertoire of the Opisthokonta ancestor most likely consisted of enzymes from at least the seven families present in animals and fungi (six common to both kingdoms, and Nudix_hydro missing from the analyzed fungi).\u003c/p\u003e \u003cp\u003eAt the level of subfamilies, only members of the canonical NUDIX PF00293 family formed more than one group when clustering by sequence similarity. This showed that diversification of NUDIX enzymes in fungi has mostly occurred within this family. Seven of these subfamilies included human representatives, suggesting that the divergence of these enzymes occurred before the animal-fungi split. The distribution of the thirteen newly identified subfamilies is more scattered throughout the fungal tree. Members of six of these subfamilies are found across multiple phyla, while six are exclusive to Glomeromycota and one is restricted to Mortierellomycota. Although it is difficult to draw conclusions regarding the specific set of events that gave rise to the observed subfamily distribution in the phylogeny, it is clear that a series of gene duplications followed by sequence diversification probably occurred throughout the evolution of Glomeromycota.\u003c/p\u003e"},{"header":"Conclusions","content":"\u003cp\u003eOur findings could have implications beyond the annotation of the undescribed enzyme universe. NUDIX hydrolases are known to be involved in the pathogenicity of a variety of bacteria [\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e], and there have been attempts to use them as antimicrobial targets [\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e]. Although much less is known about fungal pathogens, NUDIX hydrolases have been implicated in the oxidative stress response of \u003cem\u003eCryptococcus neoformans\u003c/em\u003e [\u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e]. The restricted distribution of some of the subfamilies also opens up the possibility to employ them as taxonomic markers either in diagnostics or for biodiversity assessment. For instance, sequences belonging to the newly identified subfamily E and that were found only in Mortierellomycota seem ideal for this purpose. Recently, NUDIX hydrolases received attention as orchestrators of carotenoid biosynthesis in plants [\u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e54\u003c/span\u003e]. Therefore, they could also be manipulated in fungi for the production of metabolites of commercial interest. Overall, the diversity of fungal NUDIX hydrolases described here sets them as ideal targets for tailored manipulation in both medical and biotechnological settings.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eData availability statement\u003c/h2\u003e\n\u003cp\u003eAll metadata processed in this study are deposited in Zenodo: 10.5281/zenodo.15100536\u003c/p\u003e\n\u003cp\u003eAll protein identifiers, genomic assemblies, transcriptomic datasets are listed in \u003cstrong\u003eSupplementary file S1.\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe Supplementary Figures and description of novel subfamilies are available in \u0026nbsp;\u003cstrong\u003eSupplementary file S2.\u003c/strong\u003e\u003c/p\u003e\n\u003ch2\u003eFunding\u003c/h2\u003e\n\u003cp\u003eThis work was supported by National Science Centre grants (#2021/41/B/NZ2/02426 to A.M.). E.M. received a fellowship from Conahcyt for a sabbatical stay (I0200/111/2024).\u003c/p\u003e\n\u003ch2\u003eAuthors\u0026apos; contributions\u003c/h2\u003e\n\u003cp\u003eZ.P., E.M. and A.M. designed the study, Z.P and D.B. performed analyses. Z.P., D.B., E. M. and A.M. drafted the manuscript, Z.P., and D.B. prepared the figures, A.M. conceptualized the project.\u0026nbsp;\u003c/p\u003e\n\u003ch2\u003eCompeting\u0026nbsp;interests\u003c/h2\u003e\n\u003cp\u003eThe author(s) declare that they have no competing interests.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eBessman MJ, Frick DN, O\u0026rsquo;Handley SF. The MutT Proteins or \u0026ldquo;Nudix\u0026rdquo; Hydrolases, a Family of Versatile, Widely Distributed, \u0026ldquo;Housecleaning\u0026rdquo; Enzymes *. J Biol Chem. 1996;271:25059\u0026ndash;62.\u003c/li\u003e\n\u003cli\u003eCarreras-Puigvert J, Zitnik M, Jemth A-S, Carter M, Unterlass JE, Hallstr\u0026ouml;m B, et al. A comprehensive structural, biochemical and biological profiling of the human NUDIX hydrolase family. Nature Communications. 2017;8:1\u0026ndash;17.\u003c/li\u003e\n\u003cli\u003eSrouji JR, Xu A, Park A, Kirsch JF, Brenner SE. The evolution of function within the Nudix homology clan. Proteins. 2017;85:775.\u003c/li\u003e\n\u003cli\u003eMcLennan AG. The Nudix hydrolase superfamily. Cell Mol Life Sci. 2006;63.\u003c/li\u003e\n\u003cli\u003eLukaszewicz M. Application of Mammalian Nudix Enzymes to Capped RNA Analysis. Pharmaceuticals (Basel, Switzerland). 2024;17.\u003c/li\u003e\n\u003cli\u003eLiu Y, Zhang W, Wang Y, Xie L, Zhang Q, Zhang J, et al. Nudix hydrolase 14 influences plant development and grain chalkiness in rice. Frontiers in plant science. 2022;13.\u003c/li\u003e\n\u003cli\u003eMildvan AS, Xia Z, Azurmendi HF, Saraswat V, Legler PM, Massiah MA, et al. Structures and mechanisms of Nudix hydrolases. Arch Biochem Biophys. 2005;433:129\u0026ndash;43.\u003c/li\u003e\n\u003cli\u003eSheikh S, O\u0026rsquo;Handley SF, Dunn CA, Bessman MJ. Identification and characterization of the Nudix hydrolase from the Archaeon, Methanococcus jannaschii, as a highly specific ADP-ribose pyrophosphatase. The Journal of biological chemistry. 1998;273.\u003c/li\u003e\n\u003cli\u003eXu W, Gauss P, Shen J, Dunn CA, Bessman MJ. The gene e.1 (nudE.1) of T4 bacteriophage designates a new member of the Nudix hydrolase superfamily active on flavin adenine dinucleotide, adenosine 5\u0026rsquo;-triphospho-5\u0026apos;-adenosine, and ADP-ribose. The Journal of biological chemistry. 2002;277.\u003c/li\u003e\n\u003cli\u003eBateman A, Coin L, Durbin R, Finn RD, Hollich V, Griffiths-Jones S, et al. The Pfam protein families database. Nucleic acids research. 2004;32 Database issue.\u003c/li\u003e\n\u003cli\u003eLi W, Godzik A. Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences. Bioinformatics. 2006;22.\u003c/li\u003e\n\u003cli\u003eLi W, Jaroszewski L, Godzik A. Clustering of highly homologous sequences to reduce the size of large protein databases. Bioinformatics. 2001;17.\u003c/li\u003e\n\u003cli\u003eRozewicki J, Li S, Amada KM, Standley DM, Katoh K. MAFFT-DASH: integrated protein sequence and structural alignment. Nucleic Acids Res. 2019;47:W5\u0026ndash;10.\u003c/li\u003e\n\u003cli\u003eWaterhouse AM, Procter JB, Martin DMA, Clamp M, Barton GJ. Jalview Version 2\u0026mdash;a multiple sequence alignment editor and analysis workbench. Bioinformatics. 2009;25:1189\u0026ndash;91.\u003c/li\u003e\n\u003cli\u003eBerman HM, Westbrook J, Feng Z, Gilliland G, Bhat TN, Weissig H, et al. The Protein Data Bank. Nucleic acids research. 2000;28.\u003c/li\u003e\n\u003cli\u003eAltschul SF, Gish W, Miller W, Myers EW, Lipman DJ. Basic local alignment search tool. J Mol Biol. 1990;215.\u003c/li\u003e\n\u003cli\u003eMinh BQ, Schmidt HA, Chernomor O, Schrempf D, Woodhams MD, von Haeseler A, et al. IQ-TREE 2: New Models and Efficient Methods for Phylogenetic Inference in the Genomic Era. Mol Biol Evol. 2020;37:1530\u0026ndash;4.\u003c/li\u003e\n\u003cli\u003eEmms DM, Kelly S. OrthoFinder: phylogenetic orthology inference for comparative genomics. Genome Biol. 2019;20:238.\u003c/li\u003e\n\u003cli\u003eLetunic I, Bork P. Interactive Tree Of Life (iTOL) v5: an online tool for phylogenetic tree display and annotation. Nucleic Acids Res. 2021;49.\u003c/li\u003e\n\u003cli\u003eMistry J, Chuguransky S, Williams L, Qureshi M, Salazar GA, Sonnhammer ELL, et al. Pfam: The protein families database in 2021. Nucleic Acids Res. 2021;49:D412\u0026ndash;9.\u003c/li\u003e\n\u003cli\u003eWang J, Chitsaz F, Derbyshire MK, Gonzales NR, Gwadz M, Lu S, et al. The conserved domain database in 2023. Nucleic Acids Res. 2023;51.\u003c/li\u003e\n\u003cli\u003eCrooks GE, Hon G, Chandonia JM, Brenner SE. WebLogo: a sequence logo generator. Genome Res. 2004;14.\u003c/li\u003e\n\u003cli\u003eWootton JC. Non-globular domains in protein sequences: automated segmentation using complexity measures. Comput Chem. 1994;18.\u003c/li\u003e\n\u003cli\u003eLupas A, Van Dyke M, Stock J. Predicting coiled coils from protein sequences. Science. 1991;252.\u003c/li\u003e\n\u003cli\u003eErdős G, Pajkos M, Doszt\u0026aacute;nyi Z. IUPred3: prediction of protein disorder enhanced with unambiguous experimental annotation and visualization of evolutionary conservation. Nucleic Acids Res. 2021;49:W297\u0026ndash;303.\u003c/li\u003e\n\u003cli\u003eHorton P, Park K-J, Obayashi T, Fujita N, Harada H, Adams-Collier CJ, et al. WoLF PSORT: protein localization predictor. Nucleic Acids Res. 2007;35 Web Server issue:W585.\u003c/li\u003e\n\u003cli\u003eKrogh A, Larsson B, von Heijne G, Sonnhammer EL. Predicting transmembrane protein topology with a hidden Markov model: application to complete genomes. J Mol Biol. 2001;305.\u003c/li\u003e\n\u003cli\u003eKans J. Entrez Direct: E-utilities on the Unix Command Line. In: Entrez Programming Utilities Help [Internet]. National Center for Biotechnology Information (US); 2024.\u003c/li\u003e\n\u003cli\u003eHastings J, Owen G, Dekker A, Ennis M, Kale N, Muthukrishnan V, et al. ChEBI in 2016: Improved services and an expanding collection of metabolites. Nucleic Acids Res. 2015;44:D1214\u0026ndash;9.\u003c/li\u003e\n\u003cli\u003eVaradi M, Bertoni D, Magana P, Paramval U, Pidruchna I, Radhakrishnan M, et al. AlphaFold Protein Structure Database in 2024: providing structure coverage for over 214 million protein sequences. Nucleic Acids Res. 2023;52:D368\u0026ndash;75.\u003c/li\u003e\n\u003cli\u003evan Kempen M, Kim SS, Tumescheit C, Mirdita M, Lee J, Gilchrist CLM, et al. Fast and accurate protein structure search with Foldseek. Nat Biotechnol. 2023;42:243\u0026ndash;6.\u003c/li\u003e\n\u003cli\u003eDolinsky TJ, Czodrowski P, Li H, Nielsen JE, Jensen JH, Klebe G, et al. PDB2PQR: expanding and upgrading automated preparation of biomolecular structures for molecular simulations. Nucleic acids research. 2007;35 Web Server issue.\u003c/li\u003e\n\u003cli\u003eTrott O, Olson AJ. AutoDock Vina: improving the speed and accuracy of docking with a new scoring function, efficient optimization and multithreading. J Comput Chem. 2010;31:455.\u003c/li\u003e\n\u003cli\u003eEberhardt J, Santos-Martins D, Tillack AF, Forli S. AutoDock Vina 1.2.0: New Docking Methods, Expanded Force Field, and Python Bindings. J Chem Inf Model. 2021;61.\u003c/li\u003e\n\u003cli\u003eO\u0026rsquo;Boyle NM, Banck M, James CA, Morley C, Vandermeersch T, Hutchison GR. Open Babel: An open chemical toolbox. J Cheminform. 2011;3:1\u0026ndash;14.\u003c/li\u003e\n\u003cli\u003eIvanova L, Karelson M. The Impact of Software Used and the Type of Target Protein on Molecular Docking Accuracy. Molecules (Basel, Switzerland). 2022;27.\u003c/li\u003e\n\u003cli\u003eBittencourt SA, S. A, Bittencourt a S. FastQC: a quality control tool for high throughput sequence data. https://www.scienceopen.com/document?vid=de674375-ab83-4595-afa9-4c8aa9e4e736. Accessed 18 Jun 2024.\u003c/li\u003e\n\u003cli\u003eChen S, Zhou Y, Chen Y, Gu J. fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics. 2018;34.\u003c/li\u003e\n\u003cli\u003eKim D, Langmead B, Salzberg SL. HISAT: a fast spliced aligner with low memory requirements. Nat Methods. 2015;12.\u003c/li\u003e\n\u003cli\u003eLi H, Handsaker B, Wysoker A, Fennell T, Ruan J, Homer N, et al. The Sequence Alignment/Map format and SAMtools. Bioinformatics. 2009;25.\u003c/li\u003e\n\u003cli\u003ePertea M, Pertea GM, Antonescu CM, Chang TC, Mendell JT, Salzberg SL. StringTie enables improved reconstruction of a transcriptome from RNA-seq reads. Nat Biotechnol. 2015;33.\u003c/li\u003e\n\u003cli\u003eGarc\u0026iacute;a-Saura AG, Zapata-P\u0026eacute;rez R, Mart\u0026iacute;nez-Mo\u0026ntilde;ino AB, Hidalgo JF, Morte A, P\u0026eacute;rez-Gilabert M, et al. The first comprehensive phylogenetic and biochemical analysis of NADH diphosphatases reveals that the enzyme from Tuber melanosporum is highly active towards NAD+. Sci Rep. 2019;9.\u003c/li\u003e\n\u003cli\u003eZegarra V, Mais C-N, Freitag J, Bange G. The mysterious diadenosine tetraphosphate (AP4A). microLife. 2023;4:uqad016.\u003c/li\u003e\n\u003cli\u003eOka K, Suzuki T, Onodera Y, Miki Y, Takagi K, Nagasaki S, et al. Nudix-type motif 2 in human breast carcinoma: a potent prognostic factor associated with cell proliferation. International journal of cancer. 2011;128.\u003c/li\u003e\n\u003cli\u003eGaspar ML, Pawlowska TE. Innate immunity in fungi: Is regulated cell death involved? PLoS Pathog. 2022;18:e1010460.\u003c/li\u003e\n\u003cli\u003ePalazzo L, Thomas B, Jemth A-S, Colby T, Leidecker O, Feijs KLH, et al. Processing of protein ADP-ribosylation by Nudix hydrolases. Biochem J. 2015;468:293\u0026ndash;301.\u003c/li\u003e\n\u003cli\u003eQi H, Rh GW, Beato M, Price BD. The ADP-ribose hydrolase NUDT5 is important for DNA repair. Cell Rep. 2022;41.\u003c/li\u003e\n\u003cli\u003eVenice F, Desir\u0026ograve; A, Silva G, Salvioli A, Bonfante P. The Mosaic Architecture of NRPS-PKS in the Arbuscular Mycorrhizal Fungus Gigaspora margarita Shows a Domain With Bacterial Signature. Front Microbiol. 2020;11.\u003c/li\u003e\n\u003cli\u003eDiet of Arbuscular Mycorrhizal Fungi: Bread and Butter? Trends Plant Sci. 2017;22:652\u0026ndash;60.\u003c/li\u003e\n\u003cli\u003eSokołowska B, Orłowska M, Okrasińska A, Piłsyk S, Pawłowska J, Muszewska A. What can be lost? Genomic perspective on the lipid metabolism of Mucoromycota. IMA Fungus. 2023;14:1\u0026ndash;21.\u003c/li\u003e\n\u003cli\u003eKraszewska E, Drabinska J. Nudix proteins affecting microbial pathogenesis. Microbiology (Reading). 2020;166:1110\u0026ndash;4.\u003c/li\u003e\n\u003cli\u003eSharma A, Tendulkar AV, Wangikar PP. Structure based prediction of functional sites with potential inhibitors to Nudix enzymes from disease causing microbes. Bioinformation. 2011;5.\u003c/li\u003e\n\u003cli\u003eLee KT, Kwon H, Lee D, Bahn YS. A Nudix Hydrolase Protein, Ysa1, Regulates Oxidative Stress Response and Antifungal Drug Susceptibility in Cryptococcus neoformans. Mycobiology. 2014;42.\u003c/li\u003e\n\u003cli\u003eRao S, Cao H, O\u0026rsquo;Hanna FJ, Zhou X, Lui A, Wrightstone E, et al. Nudix hydrolase 23 post-translationally regulates carotenoid biosynthesis in plants. The Plant cell. 2024;36.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"bmc-genomics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"gics","sideBox":"Learn more about [BMC Genomics](http://bmcgenomics.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/gics","title":"BMC Genomics","twitterHandle":"#BMCGenomics","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"NUDIX hydrolases, fungi, protein families, substrate specificity, Glomeromycota","lastPublishedDoi":"10.21203/rs.3.rs-6343747/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6343747/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e \u003cp\u003eThe NUDIX superfamily encompasses highly diverse enzymes involved in a plethora of biological functions such as mRNA metabolism, DNA repair, and lipid peroxidation. These hydrolases are found in all domains of life and show surprising versatility in terms of the substrates that they process. The knowledge about the diversity of fungal NUDIX proteins is fragmentary, being largely limited to a small number of characterized enzymes from yeasts. To address this knowledge gap systematically, we performed a detailed analysis of the NUDIX hydrolases across 183 fungal proteomes.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eMembers of six of the known NUDIX families were present in fungi being particularly abundant in Glomeromycota. Phylogenetic analysis and sequence clustering grouped fungal NUDIX enzymes in 25 subfamilies, 13 of which did not cluster with previously known enzymes. These 13 newly identified subfamilies all belong to the canonical NUDIX family, and structural comparison revealed a typical NUDIX fold with α-β-α sandwich structure. Molecular docking suggested Ap3A and Ap4A as substrates with the highest binding affinity, but their possible cellular roles remain unclear. We also found evidence of expression of most of the genes that encode these enzymes, suggesting physiological relevance.\u003c/p\u003e\u003ch2\u003eConclusions\u003c/h2\u003e \u003cp\u003eOur analysis offers a comprehensive perspective on the structural and sequence relationships of the NUDIX superfamily across fungi with potential to guide experimental characterization of their biological functions.\u003c/p\u003e","manuscriptTitle":"Analysis of NUDIX enzymes across fungi reveals previously unrecognized diversity","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-04-24 12:44:08","doi":"10.21203/rs.3.rs-6343747/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2025-04-25T20:01:35+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-04-24T21:00:28+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"261348224422865209571371744840917600963","date":"2025-04-17T15:37:41+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-04-14T00:43:42+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"181885761867820029774216838718331393233","date":"2025-04-04T14:38:18+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-04-03T15:59:08+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2025-04-02T17:04:08+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-04-02T08:25:18+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-04-02T08:25:06+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Genomics","date":"2025-03-31T09:52:59+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"bmc-genomics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"gics","sideBox":"Learn more about [BMC Genomics](http://bmcgenomics.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/gics","title":"BMC Genomics","twitterHandle":"#BMCGenomics","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"0e9540ae-9884-4531-95a6-2b762818bb4c","owner":[],"postedDate":"April 24th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2025-07-07T16:12:26+00:00","versionOfRecord":{"articleIdentity":"rs-6343747","link":"https://doi.org/10.1186/s12864-025-11778-5","journal":{"identity":"bmc-genomics","isVorOnly":false,"title":"BMC Genomics"},"publishedOn":"2025-07-01 15:57:15","publishedOnDateReadable":"July 1st, 2025"},"versionCreatedAt":"2025-04-24 12:44:08","video":"","vorDoi":"10.1186/s12864-025-11778-5","vorDoiUrl":"https://doi.org/10.1186/s12864-025-11778-5","workflowStages":[]},"version":"v1","identity":"rs-6343747","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6343747","identity":"rs-6343747","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-26T02:00:01.498150+00:00
License: CC-BY-4.0