Phylogenetic comparison and splice site conservation of the animal SMNDC1 gene family

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Alternative splicing is the process of multiple mRNAs from a single pre mRNA under the action of the spliceosome and other splicing factors. SMNDC1 (survival motor neuron domain containing 1) has been identified as a constituent of the spliceosome complex. Previous studies indicated that SMNDC1 is required for splicing catalysis in vitro and regulates intron retention in cancer. However, the phylogenetic relationships and expression profiles of SMNDC1 have not been systematically studied in the animal kingdom. To this end, in our work, the phylogenetic analysis of SMNDC1 genes was widely performed in the animal kingdom. Specifically, a total of 72 SMNDC1 genes were identified from 66 animal species. Bioinformatics analysis showed that the gene structure and function of SMNDC1 proteins are relatively conserved, and only a few members have two copies. In particular, the human SMNDC1 gene is highly expressed in multiple cancer types, including breast cancer, colon cancer and rectal cancer, indicating that SMNDC1 may play an essential role in cancer development and may be used as a valuable diagnostic or therapeutic protein target in clinical treatment. In summary, our findings facilitated a comprehensive overview of the animal SMNDC1 gene family, and provided a basic data and potential clues for the further study of molecular functions of SMNDC1.
Full text 119,743 characters · extracted from preprint-html · click to expand
Phylogenetic comparison and splice site conservation of the animal SMNDC1 gene family | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Help Center Sign In Submit a Preprint Cite Share Download PDF Article Phylogenetic comparison and splice site conservation of the animal SMNDC1 gene family Ouyang Guojun, Ya-Nan Leng, Mo-xian Chen, Bao-Xin Huang, Chao Sun, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-3896856/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Alternative splicing is the process of multiple mRNAs from a single pre mRNA under the action of the spliceosome and other splicing factors. SMNDC1 (survival motor neuron domain containing 1) has been identified as a constituent of the spliceosome complex. Previous studies indicated that SMNDC1 is required for splicing catalysis in vitro and regulates intron retention in cancer. However, the phylogenetic relationships and expression profiles of SMNDC1 have not been systematically studied in the animal kingdom. To this end, in our work, the phylogenetic analysis of SMNDC1 genes was widely performed in the animal kingdom. Specifically, a total of 72 SMNDC1 genes were identified from 66 animal species. Bioinformatics analysis showed that the gene structure and function of SMNDC1 proteins are relatively conserved, and only a few members have two copies. In particular, the human SMNDC1 gene is highly expressed in multiple cancer types, including breast cancer, colon cancer and rectal cancer, indicating that SMNDC1 may play an essential role in cancer development and may be used as a valuable diagnostic or therapeutic protein target in clinical treatment. In summary, our findings facilitated a comprehensive overview of the animal SMNDC1 gene family, and provided a basic data and potential clues for the further study of molecular functions of SMNDC1. Biological sciences/Genetics Biological sciences/Genetics/Rna splicing Alternative splicing phylogenetics splicing factor splice site selection SMNDC1 Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 INTRODUCTION Most genes in higher eukaryotes are composed of exon and intron intervals. Gene splicing is the process of removing introns and joining exons of genes to generate mature mRNA. Alternative splicing (AS) is the process of selecting different combinations of splice sites, which leads to the generation of multiple mRNAs from one pre-mRNA [ 1 ]. AS greatly enriches proteomic structural and functional diversity by producing multiple proteins from a single gene. Meanwhile, these protein isoforms may differ in properties such as enzymatic activity, subcellular localization, and ligand binding. Previous studies have shown that more than 90% of human genes experience AS events [ 2 – 5 ]. Increasing evidence has shown that AS is essential for normal biological processes, such as hematopoiesis [ 6 ], brain development [ 7 ], and muscle function [ 8 ]. Furthermore, it also plays an important role in the occurrence and development of various diseases, including Duchenne muscular dystrophy, spinal muscular atrophy, beta-thalassemia, myotonic dystrophy, isolated growth hormone deficiency type II and Frasier syndrome [ 9 – 15 ]. AS is performed by the spliceosome, which contains five kinds of snRNPs (small nuclear ribonucleoprotein particles, U1, U2, U4, U5, U6) and a variety of non-snRNPs splicing factors [ 16 , 17 ]. For instance, SMN (survival motor neuron) belongs to part of the small nuclear ribonucleoprotein (snRNP) complex in the cytoplasm, and is responsible for pre-mRNA splicing [ 18 ]. SMN not only has affinity for Sm ribonucleoproteins that form a ring involved in the splicing process [ 19 ], but is also implicated in binding methylated arginines [ 20 ]. Two highly homologous SMN genes were identified in the human genome, namely, SMN1 and SMN2. Meanwhile, mutation of SMN1 is the main cause of spinal muscular atrophy (SMA) [ 21 ]. SMNDC1 (survival motor neuron domain containing 1) is a paralogue of the SMN1 gene, which is implicated in pre-mRNA splicing, and has been identified to play a critical role in spliceosome assembly of the nucleus [ 20 , 22 – 24 ]. Importantly, SMNDC1 can also regulate the splicing efficiency. For instance, low SMNDC1 poison exon inclusion was associated with notably widespread reductions in intron retention [ 25 ]. In addition, the SMNDC1 poison exon controls SMNDC1 expression to modulate pancancer intron retention [ 25 ]. Meanwhile, the biological functions of SMNDC1 have been shown to play an important transcriptional role in skeletal muscle, the adult brain and the spinal cord [ 26 ]. Additionally, SMNDC1 mRNA has been identified as a potential target of fragile X mental retardation protein (FMRP), whose loss of expression leads to fragile X syndrome [ 27 ]. All of these studies demonstrate that SMNDC1 plays important roles in human disease. On the other hand, SMNDC1 is also known as survival of motor neuron-related splicing factor 30 (SPF30), which has been predicted by bioinformatics to form the evolution and alternative splicing profile analysis of the splicing factor 30 (SPF30) in plant species [ 28 ]. To this end, in our present work, we identified and analyzed the phylogenetic relationship of the SMNDC1 gene family in different animal species. Subsequently, the gene structure, protein domains, and conserved splicing patterns were elucidated, and their expression patterns in different tissues and different diseases were discussed. This study explored the potential functions of SMNDC1 s to provide theoretical support for further functional studies. RESULTS Identification of SMNDC1 genes in animal and construction of a phylogenetic tree To explore the functional differentiation of SMNDC1 family genes, in our work, 72 SMNDC1 protein sequences from 66 animal species were subjected to protein domain alignment analysis by the online software SMART. Specifically, 23 Primates, 19 Rodents and lagomorphs, 9 Fish, 2 birds and Reptiles (Anole lizard and Chinese softshell turtle), and one other animal (Lamprey) were identified. Specifically, 60 out of 66 species, including humans and all fish, have only one SMNDC1 gene, while 6 species have two copies of SMNDC1, including Angola colobus, Mouse Lemurs, Pig-tailed macaques, Rabbits, Golden Hamsters and Pigs. In addition, all animal species do not have more than 3 copies of the SMNDC1 gene, which is inconsistent with the results in plants. For example, Kalanchoe laxifora has four SPF30 (SMNDC1) genes and Triticum aestivum (wheat) possesses 3 copies [ 28 ].Furthermore, in order to understand the evolutionary history and phylogenetic relationships among the above identified SMNDC1 genes, a phylogenetic tree was constructed using the Bayesian method based on the amino acid sequences of 72 SMNDC1 members from 66 animal species (Fig. 1 ). From multiple transcript isoforms of one gene, the gene with the longest protein-coding sequence was selected as a representative. Bootstrap values are presented as a color gradient at the branches. Species from different taxonomies are marked with different colors. The tree grouped into four major clades including primates, rodents and lagomorphs (purple), birds and reptiles (pink), fish (light blue) and other mammals (green). Not surprisingly, genes from phylogenetically related animal species tend to cluster together in the tree. For example, the SMNDC1 gene from primate species including Homo sapiens and its close relatives belongs to a unique monophyletic group (Fig. 1 ). Taken together, the four main clades clustered reflect general animal phylogeny. Furthermore, the lengths of the branches indicate evolutionary distances between organisms, while the clear topology indicates the validity of the phylogenetic reconstruction of the SMNDC1 gene family in animals. The high-precision phylogenetic tree constructed by the current study can provide the basis for subsequent bioinformatics analysis. Analysis of protein domain/motif To further study the conservation of the animal SMNDC1 gene, a detailed analysis of its protein domains and conserved motifs was performed. The SMNDC1 proteins of 66 representative animal species were further aligned and used to construct a phylogenetic tree (Figs. 1 , 2 ). According to the results, the length of the identified SMNDC1 proteins from all animal species was characterized in a range of 178 to 288 amino acids. Most SMNDC1 proteins are approximately 238 amino acids in length (Table S2 ). Moreover, all SMNDC1 proteins have a characteristic central SMN domain. Specifically, the size of the conserved SMN domains was kept strictly at 60 amino acids. Moreover, the conserved motifs of animal SMNDC1 proteins were predicted by the MEME online tool. In detail, the top ten conserved motifs are illustrated in colored boxes, which cover most areas of the protein (Fig. 2 right panel). The vast majority of animal SMNDC1 sequences, including those of humans, contain 10 conserved motifs. Furthermore, the SMN domain was mainly concentrated in the middle 5 motifs (Fig. 2 right panel). Interestingly, animal species SMNDC1 with two copies have some differences in their motifs, one with 10 motifs, the other with less than 10 or with differences in motifs. For instance, ENSCANT00000039560.1 in Angola colobus has 10 motifs, while ENSCANT00000045349.1 has only 8 motifs, which implies potential functional diversification. Interaction Networks of SMNDC1 The crystal structure of human SMNDC1 is presented here (Fig. 3 ). The aromatic cage in the Tudor domain of SMNDC1 mediates dimethylarginine recognition through cation-π interactions with five important residues of the aromatic cage (Trp83, Tyr90, Phe108, Tyr111, and Asn113), as shown in Fig. 3 . In details, Tyr90 and Asn113 were highly conserved at ConSurf Grade 9. Trp83 was conserved at ConSurf Grade 8. Phe108(98.571%) and Tyr111 (97.143%) were conserved at ConSurf Grade 6 and ConSurf Grade 4. Since the SMNDC1 protein interaction network of SMNDC1 proteins may further reveal its involvement in various biological processes. In our study, to investigate the functional relationship between SMNDC1 and other proteins, the webtool STRING was used to construct the protein interaction networks of animal SMNDC1. Based on experiments and databases, three representative SMNDC1 protein sequences of human, mouse and Schizosaccharomyces pombe (yeast) were selected to generate an interaction network (Fig. 4 ). The resulting networks of human, mouse and yeast SMNDC1 networks grouped 10, 10 and 5 functional partners, respectively. In detail, the interacting proteins of human SMNDC1 can be divided into three categories: small nuclear ribonucleoprotein (SNRNP200 and SNRPB), splicing factor (SF3A3, SF3A2, SF3B2, SF3B4, SF3B5 and SF3B6) and pre-mRNA processing factor (PRPF6 and PRPF3). However, except the three interacting proteins described above, mouse SMNDC1 also interacts with U6 small nuclear RNA and mRNA degradation-associated protein (Lsm5 and Lsm6) and RNA-binding motif protein (Rbmx). Interestingly, the yeast SPF30 interacting protein is quite different from human and mouse, mainly including prp1 (U4/U6 x U5 tri-snRNP complex subunit Prp1), sap62 (zinc finger protein Sap62), itr2 (MFS myo-inositol transporter), dis3 (putative 3'-5' exoribonuclease subunit Dis3) and swi6 (chromodomain protein Swi6). In addition, we found that many interacting proteins of mammalian SMNDC1 have no apparent homologue in Schizosaccharomyces pombe, for example, splicing factor (SF3A3, SF3A2, SF3B2, SF3B4, SF3B5 and SF3B6) and pre-mRNA processing factor (PRPF6 and PRPF3). Taken together, the specific interaction studies and further functional verification of SMNDC1 may reveal its involvement in various biological processes. Analysis of gene structure and conserved motifs To further explore the conservation of gene structure and motif composition at the genome level, the longest SMNDC1 gene transcript of each coding sequence (CDS) was chosen for analysis (Fig. 5 ). According to the results, different genomic structures were observed, with the number of total exons ranging between two and seven. In most primates, the number of exons remains at five, moreover, the white partridge ENSMLET00000055620.1 has seven exons, while ENSMNET00000044298.1 and ENSMICT00000046893.2 have three exons, and ENSCANT00000045349.1 has only one exon. In fish, the number of exons remained stable at five and six. In summary, the SMNDC1 gene with five exons in the CDS accounts for approximately 89% of the total (Fig. 5 and Table S2 ), including SMNDC1 genes from representative species human and rodent and rabbit IDs. Among the 72 SMNDC1 family genes, 43 sequences had 5 exon-4 intron gene structure layouts, accounting for 59.7% of the total number of members. Twenty-two members had 6 exon-5 intron gene structure layouts, accounting for 30.5% of the total number of members. Additionally, ENSOCUT00000012273.2 and ENSMLET00000055620.1 possess the most exons with exon 7, while ENSCANT00000045349.1 and ENSSSCT00000039370.2 have the fewest exons with exon 2. Furthermore, ENSMNET00000044298.1, ENSMICT00000046893.2 and ENSMAUT00000006691.1 have 3 exons. Among all members with 6 exons and 5 introns, except one SMNDC1 from Armadillo (ENSDNOT00000016385.2), the other members all have an extra exon that was not a coding exon. Notably, there are two sequences of SMNDC1 genes from 6 species, which have different gene structures; for example, two sequences from Sus scrofa were found, one of which has 2 exons (ENSSSCT00000039370.2), and the other contains 5 exons (ENSSHAP00000001057.1). Collectively, the differences in the exon-intron distribution patterns of SMNDC1 among the above animal species, indicate that the structural changes of genes may be involved in the evolution of the gene family in the phylogeny of general animals. Furthermore, SMNDC1 in the same branch has obvious similarities in gene structure, indicating that they have a close evolutionary relationship. Based on the differences in gene structure between SMNDC1 genes, we further used MEME to determine whether there were differences in motif composition in their cDNA sequences. As shown in the results, the 10 most conserved motifs were identified from the cDNA sequence of SMNDC1 (supplementary Fig. 6, right panel). Overall, over half of the SMNDC1 sequences contained 10 conserved motifs. The motif position and number of the SMNDC1 gene in most animals showed little difference among primates, rodents and lagomorphs, other mammals (purple) and other vertebrates (pink). Interestingly, there were few differences between the observed motifs of SMNDC1 sequences with two different gene structures in one species. For example, two sequences from Oryctolagus cuniculus were found, one containing 9 motifs (ENSMICT000000041763.2) and the other containing 8 motifs (ENSMICT 000000046893.2). In conclusion, by comparing the conserved motifs at the RNA/cDNA and protein levels, it is found that the codon usage, number and similarity of these homologues are not different. The location of these motifs indicates the preservation of animal SMNDC1 between different proteins and cDNA. In addition, the comparison of cDNA showed that no conservative motif was found in the untranslated region, and the region was enriched with regulatory elements, which provided additional information for the conservative regulatory mechanism among these SMNDC1s. Transcript Isoforms and Conserved Splice Site Analysis To investigate the splicing patterns and conserved splicing sites of the animal SMNDC1 family genes, we performed an AS analysis of the animal SMNDC1 genes. According to the results, a total of 36 transcript isoforms from 15 animal SMNDC1 genes were summarized from the Ensembl database and linked to the phylogenetic relationships among selected species (Fig. 6 ). In particular, SMNDC1 in Rattus norvegicus and Mus musculus have the most numbers of isoforms, possess five transcript isoforms, while in the other 13 animals SMNDC1 contains two transcripts. In addition, conserved protein motifs were identified from potential protein products of the above transcript isoforms by using MEME (Fig. 6 right panel). From the results, the location of splicing is mainly located on the SMN protein domain. Meanwhile, the primary transcript has the longest peptide sequence and the most conserved motifs, while the spliced transcript has a shorter protein length and contains fewer motifs. In addition, alternative splicing types of SMNDC1 in 15 animal species are mainly alternative 3′ splice and alternative 5′ splice. Meanwhile exon skipping in Rattus norvegicus and Mus musculus was also detected. Furthermore, conserved splicing sites or conserved sequences were identified. Flanking sequences (31 bp in total) of animal SMNDC1 genes were analyzed to show their consensus in WebLogo and multiple alignment. According to the results, five representative splice sites were identified. (Fig. 7 A, B). Expression profile analysis of animal SMNDC1s To further investigate the potential functions of animal SMNDC1 in response to developmental cues or disease correlations, we analyzed the expression patterns of SMNDC1 genes from Homo sapiens and Mus musculus. In this work, we reconstructed the expression profiles of SMNDC1 in various biological aspects, such as developmental stages, different tissues and cell types, and disease conditions by using the BAR Heat Mapper Plus tool (Supplementary Figures S1 –S6).The data of Homo sapiens disease proteomics expression showed that SMNDC1 protein had high expression abundance in multiple cancer types, including breast cancer (breast tumor luminal, HER2 positive breast carcinoma and triple-negative breast cancer), colon cancer (colon adenocarcinoma and colon mucinous adenocarcinoma) and rectal cancer (rectal cell carcinoma and rectal mucinous adenocarcinoma) (Figure S3 ). Moreover, the SMNDC1 transcript of humans is widely expressed in whole body tissues, including skeletal muscle, adult brain, spinal cord, testis, liver, ovary and lung (Fig. 8 , Figure S2 ), while mouse SMNDC1 is highly expressed in brain tissue (Fig. 8 , Figure S2 ). Furthermore, cell type expression analysis showed that human SMNDC1 was highly expressed in granulocyte monocyte progenitor cells, hematopoietic multipotent progenitor cells and hematopoietic stem cells, while mouse SMNDC1 was expressed in naive thymus-derived CD4-positive, alpha-beta T cells, embryonic stem cells, accumulated in induced T-regulatory cells and T-helper 17 cells (Fig. 8 , Figure S4 ). On the other hand, human SMNDC1 was highly expressed in the fetal period and downregulated in the juvenile period (Figure S1 ), while mouse SMNDC1 was highly expressed in the embryonic period but did not abundantly accumulate in the fetus (Fig. 8 , Figure S4 ). In addition, we will pay more attention to the expression of the SMNDC1 gene in cancer and other diseases. Specifically, transcriptome data revealed that human SMNDC1 was expressed at higher gene expression levels in cancer tissues than in normal paracarcinoma tissue and normal tissues (Figure S3 ). Among them, human SMNDC1 had the highest expression level in ovarian adenocarcinoma, followed by esophageal adenocarcinoma (Figure S3 ), while the expression abundance of this protein was enriched in breast, colon and rectal cancer. In summary, we found that SMNDC1 is highly expressed in ovarian adenocarcinoma and digestive system diseases, and may be used as a valuable diagnostic or therapeutic protein target in clinical treatment. DISCUSSION AS is the main mechanism for maintaining protein diversity [ 29 ], and abnormalities in AS can lead to the occurrence of many diseases [ 30 , 31 ]. Specifically, abnormal AS promotes all stages of tumorigenesis, including cell proliferation [ 32 ], apoptosis [ 33 ], epithelial-mesenchymal transitions [ 34 ], tumor invasion and tumor metastasis [ 35 , 36 ]. Increasing studies have shown that alternative splicing may be used as a new biomarker in oncology and provide a large number of new targets for drug development, which is of great value for improving the prognosis of cancer patients [ 37 ].SMNDC1 is one of the key spliceosomes, and an in-depth comparison and phylogenetic analysis of the animal SMNDC1 family were conducted, which can provide a more comprehensive and in-depth understanding of the function of SMNDC1 in animals. Assessment of phylogeny relationships and putative functions in animal SMNDC1s SMNDC1 has been identified as an essential component of the spliceosome complex [ 20 ]. In the present work, we successfully identified 110 SMNDC1 genes from 61 animal species and reconstructed their phylogenetic relationships of these selected genes. SMNDC1 proteins can be broadly divided into four groups, including primates, rodents and lagomorphs, other mammals, and other vertebrates, which are closely related to the evolution of animal lineages. Moreover, only six species SMNDC1 genes contained 2 copies (Supplementary Table S1 ), and analysis of the protein structures and protein domains of these cDNAs revealed that this gene family maintains conserved functions (Figs. 2 , 5 ). In addition, the conservative splicing pattern of animal SMNDC1 s indicates that most transcriptional subtypes of animal SMNDC1 tend to form N-terminal truncated protein types (Fig. 6 ). Transcriptional isoforms of SMNDC1 in animals share similar gene structures, suggesting that they may have similar functions in regulating gene expression and protein interactions. On the other hand, a previous study showed that different spliceosomes have different biological functions. For instance, two splicing isoforms of ZNF148 have different effects on the proliferation, invasion and migration of human colorectal cancer cells, and exert mutual antagonistic effects [ 38 ]. In addition, a previous study reported that SMNDC1 is critical for regulating ovarian cancer tumor growth and metastasis [ 39 ]. We expect that different spliceosomes of SMNDC1 may be a potential target for the treatment of ovarian cancer, however, the isoform function of SMNDC1 still needs further study. Our work showed that the SMNDC1 proteins of animals have an SMN (Tudor) domain (Fig. 2 ), which has affinity for Sm ribonucleoproteins and is further involved in the splicing process [ 19 ]. Studies have reported that SMN plays a key role in the assembly of uridine-rich small nuclear ribonucleoprotein complexes [ 40 – 42 ] and in pre-mRNA splicing [ 22 , 24 ]. In our protein interaction work, we found that human SMNDC1 can interact with SNRNP200 (small nuclear ribonucleoprotein U5 subunit 200). SNRNP200 is closely related to the splicing of precursor mRNA, and can regulate the expression of related genes by affecting splicing, thereby affecting cell proliferation [ 43 ]. In addition, SNRNP200 also plays an important role in the pathogenesis of hereditary retinitis pigmentosa [ 44 ] and acute myeloid leukemia [ 45 ].The above results reveal that the interaction between SMNDC1 and SNRNP200 may play an important role in the function of SNRNP200 and its impact on diseases. Furthermore, SMNDC1 in mammals can interact with multiple splicing factors and pre-mRNA processing factors, suggesting that it plays an important role in AS regulation. In yeast, there is no SMN domain in the SMNDC1 protein and only one tudor-3 domain (Figure S7), which may prevent it from interacting with most alternative splicing factors. Functional diversity of animal SMNDC1 s based on their differential expression pattern SMNDC1 is a survival motor neuron protein that is required for spliceosome assembly [ 20 ]. Here, previous proteomic analysis revealed that SMNDC1 is critical for regulating ovarian cancer tumor growth and metastasis, consistent with its high expression level in ovarian adenocarcinoma (Figure S3 ), which will provide a new target and direction for anticancer drug development [ 39 ]. Splicing factors are frequently overexpressed in cancer [ 46 ]. Meanwhile, based on available proteomics datasets, our findings showed that high expression of SMNDC1 was observed in breast, colon and rectal cancers, implying its potential functional role in cancer development in these organs. In addition, based on the fact that SMNDC1 is required for splicing catalysis in vitro [ 24 ], researchers speculate that its poison exons may influence the extensive intron retention characteristic of most cancers [ 47 , 48 ]. Analysis of RNA-seq data from 512 lung adenocarcinoma samples showed that: low SMNDC1 poison exon inclusion was associated with notably widespread reductions in intron retention, further experimental validation showed that SMNDC1 poison exons control SMNDC1 expression to regulate intron retention [ 25 ]. This discovery may provide a new perspective for developing new treatments and defeating cancer. Spinal muscular atrophy (SMA) is a degenerative neuromuscular disease with muscle weakness and muscle atrophy, caused by deletion or mutation of the SMN1 gene. Its incidence in neonates is estimated to be 1:6 000 to 1:10 000 [ 49 , 50 ]. Moreover, SMNDC1 is a paralogue of the SMN1 gene, and may share a cellular function similar to that of the SMN1 gene. However, in spinal muscular atrophy being overshadowed by SMN, the biological function of SMNDC1 may not be fully elucidated [ 27 ]. Interestingly, SMNDC1 is highly expressed in the spinal cord and skeletal muscle tissue of Homo sapiens, which is consistent with the results of previous studies [ 20 ]. In addition, SMNDC1 is mainly expressed in fetal skeletal muscle tissue and not in adult tissue. These results indicated that the deletion or mutation of SMNDC1 may be another key factor in spinal muscular atrophy (SMA). Furthermore, in this work, we obtained comprehensive information on SMDNC1 AS in multiple animals, but biological experiments are still needed to validate these new predictions. The above findings may allow scientists to better identify biomarkers of disease substances and therapeutic targets. For example, SMNDC1 mRNA is a target of FMRP, and this result could complement the current understanding of the etiology of FXS [ 27 ].AS has become one of the hotspots in the era of functional genomics. The form of AS can be found by comparing transcripts and genomes, however, with the increasing amount of research samples and data analysis, the development of high-throughput experimental technology is particularly important [ 51 ]. The isoform level of SMNDC1 has not been thoroughly studied. Hence, it is necessary to further study the expression profile of animal smndc1 isoforms through SWATH-MS (sequential window acquisition of all theoretical mass spectra)-based quantitative approaches [ 52 ]. Our successful identification of the biological functions of the splicing-related protein SRP in plants, will provide a reference for us to further study the specific function of each SMNDC1 transcript isoform [ 53 ]. Comparison of SMNDC1 in animals, yeast and plants Although the splicing machinery is fairly conserved among eukaryotic species, the splicing mechanisms of humans, yeast and Arabidopsis are not identical. In particular, our work further analyzed and compared the genome structure and splice site patterns of SMNDC1 s from humans, yeast and Arabidopsis. Based on the results, we found that the SMN domain is retained between humans, Arabidopsis and rice, while a Tuor3 domain is present in yeast (Figure S7). Interestingly, the three exons encoding the domain of SMNDC1 were identical between the three species. EXPERIMENTAL METHODS Sequence identification and collection of the animal SMNDC1 proteins The SMNDC1 protein sequence (ENST00000369592.1) of Homo sapiens was used as a reference to perform the BLASTp search with an e-value cut of = 1e − 10 against all available animal genome sequences from the Ensembl database( http://asia.ensembl.org/index.html ) as described previously. The obtained protein regions were predicted by the online software HMMER ( https://www.ebi.ac.uk/Tools/hmmer/search/phmmer ). Consequently, the phylogenetic tree was constructed using the Bayesian method based on the amino acid sequences of 72 SMNDC1 members from 66 animal species. Phylogenetic analysis of the SMNDC1 gene family in animals The amino acid sequences of 72 SMNDC1 genes from 66 animal species were used for phylogenetic analysis by using the Bayesian method for genes with different transcript isoforms, the one with the longest protein coding sequence was used. Multiple sequence alignments of all selected SMNDC1 sequences were carried out using Mus-cle v3.8. Bayesian methods were used to construct a rooted phylogenetic tree of the SMNDC1 proteins using Mrbayes3.2. Maximum likelihood methods were also used to construct an additional tree by PhyML v3.0 for validating the result from the Bayesian tree [ 54 ]. The phylogenetic trees were edited using FigTree v1.4.3 [ 55 ]. Analysis of Gene Structures, Protein Domains and conserved motif Gene structure and cDNA conserved motifs were identified by the MEME online tool ( http://meme-suite.org/tools/meme ) (Bailey et al. 2009). Protein domains were predicted by HMMER website ( https://www.ebi.ac.uk/Tools/hmmer/ ) [ 56 ] and were drawn using TBtools [ 57 ]. and the exon–intron structures of all genes were downloaded and reconstructed from the Ensembl database. Analysis of Protein Interaction Networks The protein sequences of humans (ENSP00000363129.3), Mus musculus (ENSMUSP00000156644.1) and Saccharomyces cerevisiae (YLR298C_mRNA) were selected to obtain the interaction network on the STRING web server ( https://string-db.org/ ) [ 58 ]. Finally, the predicted functional partners of each SMNDC1 protein were presented in the form of an interaction network drawn by Cytoscape 3.8 software. AS Profile Analysis and Identification of Conserved Splice Sites All available alternative transcripts of animal SMNDC1 genes were downloaded from the Ensembl database. All available splicing isoforms of animal SMNDC1 genes were obtained again from Ensembl database. Selected splice junction sequences (15 bp on each side) were further examined using BLAST. Consensus sequences at representative splice sites were analyzed and visually represented by using WebLogo v3.0 ( https://weblogo.berkeley.edu/logo.cgi ) Expression Analysis of SMNDC1 from Online Microarray Datasets Expression data for animal SMNDC1 family members were downloaded from the Expression Atlas ( https://www.ebi.ac.uk/gxa/home ). The retrieved expression data were reorganized and presented as heatmaps by using online BAR HeatMapper Plus software ( http://bar.utoronto.ca/ntools/cgi-bin/ntools_heatmapper_plus.cgi ). CONCLUSION In this study, we identified a total of 110 SMNDC1 genes from 61 animal species and comprehensively analyzed their phylogenetic relationships, genomic organization, motif and protein domain enrichment and splicing pattern conservation, providing a foundation for molecular research on SMNDC1 proteins with respect to their roles in human diseases investigated in mammalian cell lines or animal models. In conclusion, the study of SMNDC1 is of great significance not only for the elucidation of related mechanisms, but also for the diagnosis and treatment of related diseases. Declarations COMPETING INTERESTS There are no competing interests to declare. AUTHOR CONTRIBUTIONS Conceptualization, CS, HMW, and BX-H; writing original draft preparation, OY-GJ, Y-NL, BX-H and CS; writing review and editing, OY-GJ, HMW, Y-NL, CS and M-XC; funding, M-XC. The final version of the manuscript was agreed by all authors. ACKNOWLEDGMENT This work was supported by the Program for Science Technology and Innovation Committee of Shenzhen (2021N062-JCYJ20210324115408023), the National Natural Science Foundation of China (NSFC32001932), and the Hong Kong Research Grant Council (AoE/M-05/12, AoE/M-403/16, GRF12100318, 12103219, 12103220). References Kornblihtt AR, Vibe-Pedersen K, Baralle FE: Human fibronectin: molecular cloning evidence for two mRNA species differing by an internal segment coding for a structural domain. EMBO J 1984, 3(1):221–226. Wang ET, Sandberg R, Luo S, Khrebtukova I, Zhang L, Mayr C, Kingsmore SF, Schroth GP, Burge CB: Alternative isoform regulation in human tissue transcriptomes. Nature 2008, 456(7221):470–476. Pal S, Gupta R, Davuluri RV: Alternative transcription and alternative splicing in cancer. Pharmacol Ther 2012, 136(3):283–294. Uhlen M, Fagerberg L, Hallstrom BM, Lindskog C, Oksvold P, Mardinoglu A, Sivertsson A, Kampf C, Sjostedt E, Asplund A et al : Proteomics. Tissue-based map of the human proteome. Science 2015, 347(6220):1260419. Hu Z, Scott HS, Qin G, Zheng G, Chu X, Xie L, Adelson DL, Oftedal BE, Venugopal P, Babic M et al : Revealing Missing Human Protein Isoforms Based on Ab Initio Prediction, RNA-seq and Proteomics. Sci Rep 2015, 5:10940. Wong ACH, Rasko JEJ, Wong JJ: We skip to work: alternative splicing in normal and malignant myelopoiesis. Leukemia 2018, 32(5):1081–1093. Matsuda T, Namura A, Oinuma I: Dynamic spatiotemporal patterns of alternative splicing of an F-actin scaffold protein, afadin, during murine development. Gene 2019, 689:56–68. Nakka K, Ghigna C, Gabellini D, Dilworth FJ: Diversification of the muscle proteome through alternative splicing. Skelet Muscle 2018, 8(1):8. Disset A, Bourgeois CF, Benmalek N, Claustres M, Stevenin J, Tuffery-Giraud S: An exon skipping-associated nonsense mutation in the dystrophin gene uncovers a complex interplay between multiple antagonistic splicing elements. Hum Mol Genet 2006, 15(6):999–1013. Cartegni L, Krainer AR: Disruption of an SF2/ASF-dependent exonic splicing enhancer in SMN2 causes spinal muscular atrophy in the absence of SMN1. Nat Genet 2002, 30(4):377–384. Svasti S, Suwanmanee T, Fucharoen S, Moulton HM, Nelson MH, Maeda N, Smithies O, Kole R: RNA repair restores hemoglobin expression in IVS2-654 thalassemic mice. Proc Natl Acad Sci U S A 2009, 106(4):1205–1210. Lin X, Miller JW, Mankodi A, Kanadia RN, Yuan Y, Moxley RT, Swanson MS, Thornton CA: Failure of MBNL1-dependent post-natal splicing transitions in myotonic dystrophy. Hum Mol Genet 2006, 15(13):2087–2097. Williams C, Hoppe HJ, Rezgui D, Strickland M, Forbes BE, Grutzner F, Frago S, Ellis RZ, Wattana-Amorn P, Prince SN et al : An exon splice enhancer primes IGF2:IGF2R binding site structure and function evolution. Science 2012, 338(6111):1209–1213. Wang GS, Cooper TA: Splicing in disease: disruption of the splicing code and the decoding machinery. Nat Rev Genet 2007, 8(10):749–761. Chabot B, Shkreta L: Defective control of pre-messenger RNA splicing in human disease. J Cell Biol 2016, 212(1):13–27. Zhou Z, Licklider LJ, Gygi SP, Reed R: Comprehensive proteomic analysis of the human spliceosome. Nature 2002, 419(6903):182–185. Will CL, Luhrmann R: Spliceosomal UsnRNP biogenesis, structure and function. Curr Opin Cell Biol 2001, 13(3):290–301. Kolb SJ, Battle DJ, Dreyfuss G: Molecular functions of the SMN complex. J Child Neurol 2007, 22(8):990–994. Cote J, Richard S: Tudor domains bind symmetrical dimethylated arginines. J Biol Chem 2005, 280(31):28476–28483. Talbot K, Miguel-Aliaga I, Mohaghegh P, Ponting CP, Davies KE: Characterization of a gene encoding survival motor neuron (SMN)-related protein, a constituent of the spliceosome complex. Hum Mol Genet 1998, 7(13):2149–2156. Lefebvre S, Burglen L, Reboullet S, Clermont O, Burlet P, Viollet L, Benichou B, Cruaud C, Millasseau P, Zeviani M et al : Identification and characterization of a spinal muscular atrophy-determining gene. Cell 1995, 80(1):155–165. Meister G, Hannus S, Plottner O, Baars T, Hartmann E, Fakan S, Laggerbauer B, Fischer U: SMNrp is an essential pre-mRNA splicing factor required for the formation of the mature spliceosome. EMBO J 2001, 20(9):2304–2314. Neubauer G, King A, Rappsilber J, Calvio C, Watson M, Ajuh P, Sleeman J, Lamond A, Mann M: Mass spectrometry and EST-database searching allows characterization of the multi-protein spliceosome complex. Nat Genet 1998, 20(1):46–50. Rappsilber J, Ajuh P, Lamond AI, Mann M: SPF30 is an essential human splicing factor required for assembly of the U4/U5/U6 tri-small nuclear ribonucleoprotein into the spliceosome. J Biol Chem 2001, 276(33):31142–31150. Thomas JD, Polaski JT, Feng Q, De Neef EJ, Hoppe ER, McSharry MV, Pangallo J, Gabel AM, Belleville AE, Watson J et al : RNA isoform screens uncover the essentiality and tumor-suppressor activity of ultraconserved poison exons. Nat Genet 2020, 52(1):84–94. Shahid R, Bugaut A, Balasubramanian S: The BCL-2 5' untranslated region contains an RNA G-quadruplex-forming motif that modulates protein expression. Biochemistry 2010, 49(38):8300–8306. McAninch DS, Heinaman AM, Lang CN, Moss KR, Bassell GJ, Rita Mihailescu M, Evans TL: Fragile X mental retardation protein recognizes a G quadruplex structure within the survival motor neuron domain containing 1 mRNA 5'-UTR. Mol Biosyst 2017, 13(8):1448–1457. Zhang D, Yang JF, Gao B, Liu TY, Hao GF, Yang GF, Fu LJ, Chen MX, Zhang J: Identification, evolution and alternative splicing profile analysis of the splicing factor 30 (SPF30) in plant species. Planta 2019, 249(6):1997–2014. Li Y, Sun N, Lu Z, Sun S, Huang J, Chen Z, He J: Prognostic alternative mRNA splicing signature in non-small cell lung cancer. Cancer Lett 2017, 393:40–51. Dvinge H, Kim E, Abdel-Wahab O, Bradley RK: RNA splicing factors as oncoproteins and tumour suppressors. Nat Rev Cancer 2016, 16(7):413–430. Scotti MM, Swanson MS: RNA mis-splicing in disease. Nat Rev Genet 2016, 17(1):19–32. Xie R, Chen X, Chen Z, Huang M, Dong W, Gu P, Zhang J, Zhou Q, Dong W, Han J et al : Polypyrimidine tract binding protein 1 promotes lymphatic metastasis and proliferation of bladder cancer via alternative splicing of MEIS2 and PKM. Cancer Lett 2019, 449:31–44. Tyson-Capper A, Gautrey H: Regulation of Mcl-1 alternative splicing by hnRNP F, H1 and K in breast cancer cells. RNA Biol 2018, 15(12):1448–1457. Pradella D, Naro C, Sette C, Ghigna C: EMT and stemness: flexible processes tuned by alternative splicing in development and cancer progression. Mol Cancer 2017, 16(1):8. Wang F, Fu X, Chen P, Wu P, Fan X, Li N, Zhu H, Jia TT, Ji H, Wang Z et al : SPSB1-mediated HnRNP A1 ubiquitylation regulates alternative splicing and cell migration in EGF signaling. Cell Res 2017, 27(4):540–558. Chen L, Yao Y, Sun L, Zhou J, Miao M, Luo S, Deng G, Li J, Wang J, Tang J: Snail Driving Alternative Splicing of CD44 by ESRP1 Enhances Invasion and Migration in Epithelial Ovarian Cancer. Cell Physiol Biochem 2017, 43(6):2489–2504. Blencowe BJ: Alternative splicing: new insights from global analyses. Cell 2006, 126(1):37–47. Liu Y, Huang W, Gao X, Kuang F: Regulation between two alternative splicing isoforms ZNF148(FL) and ZNF148(DeltaN), and their roles in the apoptosis and invasion of colorectal cancer. Pathol Res Pract 2019, 215(2):272–277. Giri K, Shameer K, Zimmermann MT, Saha S, Chakraborty PK, Sharma A, Arvizo RR, Madden BJ, McCormick DJ, Kocher JP et al : Understanding protein-nanoparticle interaction: a new gateway to disease therapeutics. Bioconjug Chem 2014, 25(6):1078–1090. Buhler D, Raker V, Luhrmann R, Fischer U: Essential role for the tudor domain of SMN in spliceosomal U snRNP assembly: implications for spinal muscular atrophy. Hum Mol Genet 1999, 8(13):2351–2357. Pellizzoni L, Kataoka N, Charroux B, Dreyfuss G: A novel function for SMN, the spinal muscular atrophy disease gene product, in pre-mRNA splicing. Cell 1998, 95(5):615–624. Pellizzoni L, Yong J, Dreyfuss G: Essential role for the SMN complex in the specificity of snRNP assembly. Science 2002, 298(5599):1775–1779. Ehsani A, Alluin JV, Rossi JJ: Cell cycle abnormalities associated with differential perturbations of the human U5 snRNP associated U5-200kD RNA helicase. PLoS One 2013, 8(4):e62125. Liu T, Jin X, Zhang X, Yuan H, Cheng J, Lee J, Zhang B, Zhang M, Wu J, Wang L et al : A novel missense SNRNP200 mutation associated with autosomal dominant retinitis pigmentosa in a Chinese family. PLoS One 2012, 7(9):e45464. Gillissen MA, Kedde M, Jong G, Moiset G, Yasuda E, Levie SE, Bakker AQ, Claassen YB, Wagner K, Bohne M et al : AML-specific cytotoxic antibodies in patients with durable graft-versus-leukemia responses. Blood 2018, 131(1):131–143. Urbanski LM, Leclair N, Anczukow O: Alternative-splicing defects in cancer: Splicing regulators and their downstream targets, guiding the way to novel cancer therapeutics. Wiley Interdiscip Rev RNA 2018, 9(4):e1476. Dvinge H, Bradley RK: Widespread intron retention diversifies most cancer transcriptomes. Genome Med 2015, 7(1):45. Jung H, Lee D, Lee J, Park D, Kim YJ, Park WY, Hong D, Park PJ, Lee E: Intron retention is a widespread mechanism of tumor-suppressor inactivation. Nat Genet 2015, 47(11):1242–1248. Lunn MR, Wang CH: Spinal muscular atrophy. Lancet 2008, 371(9630):2120–2133. Schorling DC, Pechmann A, Kirschner J: Advances in Treatment of Spinal Muscular Atrophy - New Phenotypes, New Challenges, New Implications for Care. J Neuromuscul Dis 2020, 7(1):1–13. Zhu FY, Chen MX, Ye NH, Shi L, Ma KL, Yang JF, Cao YY, Zhang Y, Yoshida T, Fernie AR et al : Proteogenomic analysis reveals alternative splicing and translation as part of the abscisic acid response in Arabidopsis seedlings. Plant J 2017, 91(3):518–533. Zhu FY, Chen MX, Chan WL, Yang F, Tian Y, Song T, Xie LJ, Zhou Y, Xiao S, Zhang J et al : SWATH-MS quantitative proteomic investigation of nitrogen starvation in Arabidopsis reveals new aspects of plant nitrogen stress responses. J Proteomics 2018, 187:161–170. Chen MX, Mei LC, Wang F, Boyagane Dewayalage IKW, Yang JF, Dai L, Yang GF, Gao B, Cheng CL, Liu YG et al : PlantSPEAD: a web resource towards comparatively analysing stress-responsive expression of splicing-related proteins in plant. Plant Biotechnol J 2021, 19(2):227–229. Guindon S, Dufayard JF, Lefort V, Anisimova M, Hordijk W, Gascuel O. New algorithms and methods to estimate maximum-likelihood phylogenies: assessing the performance of PhyML 3.0. Syst Biol. 2010, 59(3):307–21. Morariu V, Srinivasan B, Raykar V, Duraiswami R, Davis L. Automatic online tuning for fast Gaussian summation. Adv Neural Inf Process Syst. 2008. 1113–1120. Potter SC, Luciani A, Eddy SR, Park Y, Lopez R, Finn RD. HMMER web server: 2018 update. Nucleic Acids Res. 2018;46(W1): W200-W204. Chen MX, Zhang KL, Zhang M, Das D, Fang YM, Dai L, Zhang J, Zhu FY. Alternative splicing and its regulatory role in woody plants. Tree Physiol. 2020, 40(11):1475–1486. Szklarczyk, D., Franceschini, A., Wyder, S., Forslund, K., Heller, D., Huerta-Cepas, J., et al. STRING V10: Protein-Protein Interaction Networks, Integrated over the Tree of Life. Nucleic Acids Res. 2015, 43, D447–D452. Additional Declarations No competing interests reported. Supplementary Files SMNDC1TableS1.xlsx SMNDC1TableS2.xlsx SMNDC1TableS3.xlsx Supplementalfigures.docx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-3896856","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":274774129,"identity":"67602568-2562-455e-b1f5-6590f80675b6","order_by":0,"name":"Ouyang Guojun","email":"","orcid":"","institution":"Shenzhen Children's Hospital","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Ouyang","middleName":"","lastName":"Guojun","suffix":""},{"id":274774130,"identity":"d5171613-923e-4c3c-b2fc-6959b60c60bc","order_by":1,"name":"Ya-Nan Leng","email":"","orcid":"","institution":"Shenzhen Children's Hospital","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Ya-Nan","middleName":"","lastName":"Leng","suffix":""},{"id":274774131,"identity":"e8cde38a-202a-40a2-9057-e4831cde3733","order_by":2,"name":"Mo-xian Chen","email":"","orcid":"","institution":"Shenzhen Children's Hospital","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Mo-xian","middleName":"","lastName":"Chen","suffix":""},{"id":274774132,"identity":"60134c42-2d16-479a-a280-b0edf35698c9","order_by":3,"name":"Bao-Xin Huang","email":"","orcid":"","institution":"Shenzhen Children's Hospital","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Bao-Xin","middleName":"","lastName":"Huang","suffix":""},{"id":274774133,"identity":"7f7ee3b8-12a9-4309-8ddd-09748b1d00fe","order_by":4,"name":"Chao Sun","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA0UlEQVRIiWNgGAWjYDACCQST8UFCRQ1pWpgNHpw5RpoWNsmHLcyEdfDPbn748GfbHQaD42fMKhIb2Bj427sT8Fty55ixMW/bMwaDMzlmNxJ3yDBInDm7Aa8WA4kEM2nGtsMMBgdAWs6wAUVyCWlJ/yb5E6Tl/BuzgsQ2ZmK05JhJ8IK03MgxYyBKi8SNnGJjnnOHGSRvPCuWSDhzjIegX/hnpG98+KPsMAPf+eSNH39U1Mjxt/fi1wID9QsOQBg8RCkHA/kG4tWOglEwCkbBCAMAUJlJYtD3oxYAAAAASUVORK5CYII=","orcid":"","institution":"Shenzhen People's Hospital","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Chao","middleName":"","lastName":"Sun","suffix":""},{"id":274774134,"identity":"185935fe-f38e-4ecd-b259-61f94bdb4c20","order_by":5,"name":"Hong-Mei Wang","email":"","orcid":"","institution":"Shenzhen Children's Hospital","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Hong-Mei","middleName":"","lastName":"Wang","suffix":""}],"badges":[],"createdAt":"2024-01-25 10:44:14","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-3896856/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-3896856/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":51626532,"identity":"99d3fbed-a42f-4c41-8e16-1b1c87cbf216","added_by":"auto","created_at":"2024-02-26 07:48:25","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":436607,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003ePhylogenetic analysis of the SMNDC1 genes in animals. \u003c/strong\u003eThe phylogenetic tree was constructed using the Bayesian method based on the amino acid sequences of 72 SMNDC1 members from 66 animal species. The assigned gene identifiers or accession numbers for the genes are displayed. Bootstrap values are presented as a color gradient at the branches. Species from different taxonomies are marked with different colors.\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-3896856/v1/3806e99111ba575be1b47208.png"},{"id":51626533,"identity":"7de4b93c-148e-403e-b38c-f9510fd1bb94","added_by":"auto","created_at":"2024-02-26 07:48:25","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":1008289,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eProtein motif analysis of the animal SMNDC1 family. \u003c/strong\u003eThe phylogenetic relationship is listed in the left panel. Protein regions predicted by the online software HMMER are listed in the middle panel. Conserved motifs analysed by the MEME online tool are listed in the right panel. Different motifs are represented by different colored boxes, and fogged boxes indicate scanned sites. For conserved motifs, the height of a box indicates the significance of the site as taller boxes are more significant. The main regions (middle panel) correspond to the motifs (right panel), and the red frame represents the SMN region.\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-3896856/v1/0a6284e7f1f8a49d6f1effe5.png"},{"id":51626538,"identity":"c744cda4-4452-46a4-8407-96303718a7cc","added_by":"auto","created_at":"2024-02-26 07:48:25","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":179142,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eEvolutionary conservation analysis of amino acid positions in animal SMNDC1s.\u003c/strong\u003e The crystal structure of human SMNDC1 (PDB ID: 4A4H) is shown. The ribbon representation is colored according to ConSurf Grade (1-blue to 9-purple) by using all identified protein sequences of animal SMNDC1s.\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-3896856/v1/8fb54f31975a1aff5a4fef4e.png"},{"id":51626542,"identity":"7ede247e-8c09-4530-96cd-da77a1fc96db","added_by":"auto","created_at":"2024-02-26 07:48:26","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":491352,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eProtein-protein interaction networks of \u003c/strong\u003e\u003cem\u003e\u003cstrong\u003eHomo sapiens\u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003e, \u003c/strong\u003e\u003cem\u003e\u003cstrong\u003eMus musculus\u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003e and \u003c/strong\u003e\u003cem\u003e\u003cstrong\u003eSchizosaccharomyces pombe\u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003e. \u003c/strong\u003e\u0026nbsp;Known interactions were chosen for networks. Colored notes were query proteins and the first shell of interactions, and white notes were the second shell of interaction. Empty notes were proteins of unknown 3D structure, while filled notes were known or predicted 3D structure.\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-3896856/v1/82b5a3b6206d7b77bd44fe1d.png"},{"id":51626537,"identity":"5e4aa6b0-f14f-455b-895b-1cdabd42d243","added_by":"auto","created_at":"2024-02-26 07:48:25","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":1129154,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eComparisons of genomic organization and conserved DNA motif identification among animal SMNDC1 genes.\u003c/strong\u003e Gene structure (middle panel) and identified conserved cDNA motifs (right panel) by MEME analysis are shown against the vertical phylogenetic tree (left panel). The conserved sequences of the ten most identified DNA motifs are listed. Different motifs are represented by different colored boxes, and fogged boxes indicate scanned sites. For conserved motifs, the height of a box indicates the significance of the site as taller boxes are more significant.\u003c/p\u003e","description":"","filename":"5.png","url":"https://assets-eu.researchsquare.com/files/rs-3896856/v1/70998ba512d62a3282a9843b.png"},{"id":51626534,"identity":"039bffde-c764-418c-bbff-111f7db1da49","added_by":"auto","created_at":"2024-02-26 07:48:25","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":124164,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSummary of splicing isoforms for animal SMNDC1 genes. \u003c/strong\u003eTranscript isoforms from 15 animal SMNDC1 genes are summarized (left and middle panels). Conserved protein motifs of potential protein products from splicing isoforms are illustrated (right panel) with additional annotation to define exon-exon boundaries (lines between boxes). The colored arrowheads (except pink arrowheads) indicate the conserved splice site located in the region of the SMN domain. The pink arrows represent the same conserved splice sites where the new AS events were found.\u003c/p\u003e","description":"","filename":"6.png","url":"https://assets-eu.researchsquare.com/files/rs-3896856/v1/d1f4df53ddf5e27a9159ee25.png"},{"id":51626776,"identity":"81ebbd1a-106a-48cc-9fa7-2d21e0d62b62","added_by":"auto","created_at":"2024-02-26 07:56:26","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":419458,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSplice site dinucleotides (if any) are marked in blue and red boxes. \u003c/strong\u003eThe SMN domain and its corresponding conserved motifs predicted by MEME are listed in the right panel of Fig. 6. Five splicing sites \u003cstrong\u003e(A)\u003c/strong\u003e were found in different animal SMNDC1 genes, including four sites (1-4 arrow), without the detection of particular splicing events \u003cstrong\u003e(B)\u003c/strong\u003e, located in the region of SMN domain. One site (5 arrow) is the particular splicing event.\u003c/p\u003e","description":"","filename":"7.png","url":"https://assets-eu.researchsquare.com/files/rs-3896856/v1/a213a7da75739e958bc40081.png"},{"id":51626536,"identity":"17d4ef62-cdd1-4204-919f-af0e7b02d4c7","added_by":"auto","created_at":"2024-02-26 07:48:25","extension":"png","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":237064,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eLiver and disease expression in\u003c/strong\u003e\u003cem\u003e\u003cstrong\u003e Homo sapiens\u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003eand \u003c/strong\u003e\u003cem\u003e\u003cstrong\u003eMus musculu\u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003e. \u003c/strong\u003eHomo sapiens (A) and Mus_musculu (B) liver relational cell line, cell type, organism part, disease, and developmental stage expression. 1-3 of Cell line in (A) represent 675 Genentech - liver - hepatocellular carcinoma, 675 Genentech - liver - renal carcinoma, Cell lines - CCLE – hepatoblastoma. 1 and 2 of Individual in (A) indicated Pan-Cancer Analysis of Whole Genomes (by individual - liver – cholangiocarcinoma and Pan-Cancer Analysis of Whole Genomes (by individual) - liver - hepatocellular carcinoma.\u003c/p\u003e","description":"","filename":"8.png","url":"https://assets-eu.researchsquare.com/files/rs-3896856/v1/7e1e1eca553e9bc10e7da2a3.png"},{"id":54741699,"identity":"0984b62c-6ab9-4015-abef-f888f1f561a3","added_by":"auto","created_at":"2024-04-16 06:05:34","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":4742881,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3896856/v1/a49719cf-9403-4a19-a5ee-14b5ee9878a2.pdf"},{"id":51626531,"identity":"5ba69e60-6cc7-44a7-9b07-88f4d3b3c30a","added_by":"auto","created_at":"2024-02-26 07:48:25","extension":"xlsx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":16635,"visible":true,"origin":"","legend":"","description":"","filename":"SMNDC1TableS1.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-3896856/v1/eb14c6df6b173c5a576678ba.xlsx"},{"id":51626775,"identity":"4f9d096a-0c0b-4ba4-86fa-2058cf3a3afb","added_by":"auto","created_at":"2024-02-26 07:56:25","extension":"xlsx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":13178,"visible":true,"origin":"","legend":"","description":"","filename":"SMNDC1TableS2.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-3896856/v1/60fd69168797032c9ce67edc.xlsx"},{"id":51626540,"identity":"65051989-4940-4071-b966-cb4305634f9c","added_by":"auto","created_at":"2024-02-26 07:48:26","extension":"xlsx","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":5610305,"visible":true,"origin":"","legend":"","description":"","filename":"SMNDC1TableS3.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-3896856/v1/bcb602f3d14d056e8de7d537.xlsx"},{"id":51626543,"identity":"bd32e167-f2e6-423e-abdf-1b80f5ef80f0","added_by":"auto","created_at":"2024-02-26 07:48:26","extension":"docx","order_by":4,"title":"","display":"","copyAsset":false,"role":"supplement","size":637173,"visible":true,"origin":"","legend":"","description":"","filename":"Supplementalfigures.docx","url":"https://assets-eu.researchsquare.com/files/rs-3896856/v1/2441447f2589bf1bca0a95ba.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Phylogenetic comparison and splice site conservation of the animal SMNDC1 gene family","fulltext":[{"header":"INTRODUCTION","content":"\u003cp\u003eMost genes in higher eukaryotes are composed of exon and intron intervals. Gene splicing is the process of removing introns and joining exons of genes to generate mature mRNA. Alternative splicing (AS) is the process of selecting different combinations of splice sites, which leads to the generation of multiple mRNAs from one pre-mRNA [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. AS greatly enriches proteomic structural and functional diversity by producing multiple proteins from a single gene. Meanwhile, these protein isoforms may differ in properties such as enzymatic activity, subcellular localization, and ligand binding. Previous studies have shown that more than 90% of human genes experience AS events [\u003cspan additionalcitationids=\"CR3 CR4\" citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. Increasing evidence has shown that AS is essential for normal biological processes, such as hematopoiesis [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e], brain development [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e], and muscle function [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. Furthermore, it also plays an important role in the occurrence and development of various diseases, including Duchenne muscular dystrophy, spinal muscular atrophy, beta-thalassemia, myotonic dystrophy, isolated growth hormone deficiency type II and Frasier syndrome [\u003cspan additionalcitationids=\"CR10 CR11 CR12 CR13 CR14\" citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eAS is performed by the spliceosome, which contains five kinds of snRNPs (small nuclear ribonucleoprotein particles, U1, U2, U4, U5, U6) and a variety of non-snRNPs splicing factors [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e, \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]. For instance, SMN (survival motor neuron) belongs to part of the small nuclear ribonucleoprotein (snRNP) complex in the cytoplasm, and is responsible for pre-mRNA splicing [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. SMN not only has affinity for Sm ribonucleoproteins that form a ring involved in the splicing process [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e], but is also implicated in binding methylated arginines [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. Two highly homologous SMN genes were identified in the human genome, namely, SMN1 and SMN2. Meanwhile, mutation of SMN1 is the main cause of spinal muscular atrophy (SMA) [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eSMNDC1 (survival motor neuron domain containing 1) is a paralogue of the SMN1 gene, which is implicated in pre-mRNA splicing, and has been identified to play a critical role in spliceosome assembly of the nucleus [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e, \u003cspan additionalcitationids=\"CR23\" citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e]. Importantly, SMNDC1 can also regulate the splicing efficiency. For instance, low SMNDC1 poison exon inclusion was associated with notably widespread reductions in intron retention [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]. In addition, the SMNDC1 poison exon controls SMNDC1 expression to modulate pancancer intron retention [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]. Meanwhile, the biological functions of SMNDC1 have been shown to play an important transcriptional role in skeletal muscle, the adult brain and the spinal cord [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e]. Additionally, SMNDC1 mRNA has been identified as a potential target of fragile X mental retardation protein (FMRP), whose loss of expression leads to fragile X syndrome [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]. All of these studies demonstrate that SMNDC1 plays important roles in human disease. On the other hand, SMNDC1 is also known as survival of motor neuron-related splicing factor 30 (SPF30), which has been predicted by bioinformatics to form the evolution and alternative splicing profile analysis of the splicing factor 30 (SPF30) in plant species [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eTo this end, in our present work, we identified and analyzed the phylogenetic relationship of the SMNDC1 gene family in different animal species. Subsequently, the gene structure, protein domains, and conserved splicing patterns were elucidated, and their expression patterns in different tissues and different diseases were discussed. This study explored the potential functions of SMNDC1 s to provide theoretical support for further functional studies.\u003c/p\u003e"},{"header":"RESULTS","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eIdentification of SMNDC1 genes in animal and construction of a phylogenetic tree\u003c/h2\u003e \u003cp\u003eTo explore the functional differentiation of SMNDC1 family genes, in our work, 72 SMNDC1 protein sequences from 66 animal species were subjected to protein domain alignment analysis by the online software SMART. Specifically, 23 Primates, 19 Rodents and lagomorphs, 9 Fish, 2 birds and Reptiles (Anole lizard and Chinese softshell turtle), and one other animal (Lamprey) were identified. Specifically, 60 out of 66 species, including humans and all fish, have only one SMNDC1 gene, while 6 species have two copies of SMNDC1, including Angola colobus, Mouse Lemurs, Pig-tailed macaques, Rabbits, Golden Hamsters and Pigs. In addition, all animal species do not have more than 3 copies of the SMNDC1 gene, which is inconsistent with the results in plants. For example, Kalanchoe laxifora has four SPF30 (SMNDC1) genes and Triticum aestivum (wheat) possesses 3 copies [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e].Furthermore, in order to understand the evolutionary history and phylogenetic relationships among the above identified SMNDC1 genes, a phylogenetic tree was constructed using the Bayesian method based on the amino acid sequences of 72 SMNDC1 members from 66 animal species (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). From multiple transcript isoforms of one gene, the gene with the longest protein-coding sequence was selected as a representative. Bootstrap values are presented as a color gradient at the branches. Species from different taxonomies are marked with different colors. The tree grouped into four major clades including primates, rodents and lagomorphs (purple), birds and reptiles (pink), fish (light blue) and other mammals (green). Not surprisingly, genes from phylogenetically related animal species tend to cluster together in the tree. For example, the SMNDC1 gene from primate species including Homo sapiens and its close relatives belongs to a unique monophyletic group (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). Taken together, the four main clades clustered reflect general animal phylogeny. Furthermore, the lengths of the branches indicate evolutionary distances between organisms, while the clear topology indicates the validity of the phylogenetic reconstruction of the SMNDC1 gene family in animals. The high-precision phylogenetic tree constructed by the current study can provide the basis for subsequent bioinformatics analysis.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003eAnalysis of protein domain/motif\u003c/h2\u003e \u003cp\u003eTo further study the conservation of the animal SMNDC1 gene, a detailed analysis of its protein domains and conserved motifs was performed. The SMNDC1 proteins of 66 representative animal species were further aligned and used to construct a phylogenetic tree (Figs.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e,\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). According to the results, the length of the identified SMNDC1 proteins from all animal species was characterized in a range of 178 to 288 amino acids. Most SMNDC1 proteins are approximately 238 amino acids in length (Table \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003e). Moreover, all SMNDC1 proteins have a characteristic central SMN domain. Specifically, the size of the conserved SMN domains was kept strictly at 60 amino acids. Moreover, the conserved motifs of animal SMNDC1 proteins were predicted by the MEME online tool. In detail, the top ten conserved motifs are illustrated in colored boxes, which cover most areas of the protein (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e right panel). The vast majority of animal SMNDC1 sequences, including those of humans, contain 10 conserved motifs. Furthermore, the SMN domain was mainly concentrated in the middle 5 motifs (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e right panel). Interestingly, animal species SMNDC1 with two copies have some differences in their motifs, one with 10 motifs, the other with less than 10 or with differences in motifs. For instance, ENSCANT00000039560.1 in Angola colobus has 10 motifs, while ENSCANT00000045349.1 has only 8 motifs, which implies potential functional diversification.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003eInteraction Networks of SMNDC1\u003c/h2\u003e \u003cp\u003eThe crystal structure of human SMNDC1 is presented here (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). The aromatic cage in the Tudor domain of SMNDC1 mediates dimethylarginine recognition through cation-π interactions with five important residues of the aromatic cage (Trp83, Tyr90, Phe108, Tyr111, and Asn113), as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e. In details, Tyr90 and Asn113 were highly conserved at ConSurf Grade 9. Trp83 was conserved at ConSurf Grade 8. Phe108(98.571%) and Tyr111 (97.143%) were conserved at ConSurf Grade 6 and ConSurf Grade 4.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eSince the SMNDC1 protein interaction network of SMNDC1 proteins may further reveal its involvement in various biological processes. In our study, to investigate the functional relationship between SMNDC1 and other proteins, the webtool STRING was used to construct the protein interaction networks of animal SMNDC1. Based on experiments and databases, three representative SMNDC1 protein sequences of human, mouse and Schizosaccharomyces pombe (yeast) were selected to generate an interaction network (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e). The resulting networks of human, mouse and yeast SMNDC1 networks grouped 10, 10 and 5 functional partners, respectively. In detail, the interacting proteins of human SMNDC1 can be divided into three categories: small nuclear ribonucleoprotein (SNRNP200 and SNRPB), splicing factor (SF3A3, SF3A2, SF3B2, SF3B4, SF3B5 and SF3B6) and pre-mRNA processing factor (PRPF6 and PRPF3). However, except the three interacting proteins described above, mouse SMNDC1 also interacts with U6 small nuclear RNA and mRNA degradation-associated protein (Lsm5 and Lsm6) and RNA-binding motif protein (Rbmx). Interestingly, the yeast SPF30 interacting protein is quite different from human and mouse, mainly including prp1 (U4/U6 x U5 tri-snRNP complex subunit Prp1), sap62 (zinc finger protein Sap62), itr2 (MFS myo-inositol transporter), dis3 (putative 3'-5' exoribonuclease subunit Dis3) and swi6 (chromodomain protein Swi6). In addition, we found that many interacting proteins of mammalian SMNDC1 have no apparent homologue in Schizosaccharomyces pombe, for example, splicing factor (SF3A3, SF3A2, SF3B2, SF3B4, SF3B5 and SF3B6) and pre-mRNA processing factor (PRPF6 and PRPF3). Taken together, the specific interaction studies and further functional verification of SMNDC1 may reveal its involvement in various biological processes.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003eAnalysis of gene structure and conserved motifs\u003c/h2\u003e \u003cp\u003eTo further explore the conservation of gene structure and motif composition at the genome level, the longest SMNDC1 gene transcript of each coding sequence (CDS) was chosen for analysis (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e). According to the results, different genomic structures were observed, with the number of total exons ranging between two and seven. In most primates, the number of exons remains at five, moreover, the white partridge ENSMLET00000055620.1 has seven exons, while ENSMNET00000044298.1 and ENSMICT00000046893.2 have three exons, and ENSCANT00000045349.1 has only one exon. In fish, the number of exons remained stable at five and six. In summary, the SMNDC1 gene with five exons in the CDS accounts for approximately 89% of the total (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e and Table \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003e), including SMNDC1 genes from representative species human and rodent and rabbit IDs. Among the 72 SMNDC1 family genes, 43 sequences had 5 exon-4 intron gene structure layouts, accounting for 59.7% of the total number of members. Twenty-two members had 6 exon-5 intron gene structure layouts, accounting for 30.5% of the total number of members. Additionally, ENSOCUT00000012273.2 and ENSMLET00000055620.1 possess the most exons with exon 7, while ENSCANT00000045349.1 and ENSSSCT00000039370.2 have the fewest exons with exon 2. Furthermore, ENSMNET00000044298.1, ENSMICT00000046893.2 and ENSMAUT00000006691.1 have 3 exons. Among all members with 6 exons and 5 introns, except one SMNDC1 from Armadillo (ENSDNOT00000016385.2), the other members all have an extra exon that was not a coding exon. Notably, there are two sequences of SMNDC1 genes from 6 species, which have different gene structures; for example, two sequences from Sus scrofa were found, one of which has 2 exons (ENSSSCT00000039370.2), and the other contains 5 exons (ENSSHAP00000001057.1). Collectively, the differences in the exon-intron distribution patterns of SMNDC1 among the above animal species, indicate that the structural changes of genes may be involved in the evolution of the gene family in the phylogeny of general animals. Furthermore, SMNDC1 in the same branch has obvious similarities in gene structure, indicating that they have a close evolutionary relationship. Based on the differences in gene structure between SMNDC1 genes, we further used MEME to determine whether there were differences in motif composition in their cDNA sequences. As shown in the results, the 10 most conserved motifs were identified from the cDNA sequence of SMNDC1 (supplementary Fig.\u0026nbsp;6, right panel). Overall, over half of the SMNDC1 sequences contained 10 conserved motifs. The motif position and number of the SMNDC1 gene in most animals showed little difference among primates, rodents and lagomorphs, other mammals (purple) and other vertebrates (pink). Interestingly, there were few differences between the observed motifs of SMNDC1 sequences with two different gene structures in one species. For example, two sequences from Oryctolagus cuniculus were found, one containing 9 motifs (ENSMICT000000041763.2) and the other containing 8 motifs (ENSMICT 000000046893.2). In conclusion, by comparing the conserved motifs at the RNA/cDNA and protein levels, it is found that the codon usage, number and similarity of these homologues are not different. The location of these motifs indicates the preservation of animal SMNDC1 between different proteins and cDNA. In addition, the comparison of cDNA showed that no conservative motif was found in the untranslated region, and the region was enriched with regulatory elements, which provided additional information for the conservative regulatory mechanism among these SMNDC1s.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003eTranscript Isoforms and Conserved Splice Site Analysis\u003c/h2\u003e \u003cp\u003eTo investigate the splicing patterns and conserved splicing sites of the animal SMNDC1 family genes, we performed an AS analysis of the animal SMNDC1 genes. According to the results, a total of 36 transcript isoforms from 15 animal SMNDC1 genes were summarized from the Ensembl database and linked to the phylogenetic relationships among selected species (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e). In particular, SMNDC1 in Rattus norvegicus and Mus musculus have the most numbers of isoforms, possess five transcript isoforms, while in the other 13 animals SMNDC1 contains two transcripts. In addition, conserved protein motifs were identified from potential protein products of the above transcript isoforms by using MEME (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e right panel). From the results, the location of splicing is mainly located on the SMN protein domain. Meanwhile, the primary transcript has the longest peptide sequence and the most conserved motifs, while the spliced transcript has a shorter protein length and contains fewer motifs. In addition, alternative splicing types of SMNDC1 in 15 animal species are mainly alternative 3\u0026prime; splice and alternative 5\u0026prime; splice. Meanwhile exon skipping in Rattus norvegicus and Mus musculus was also detected. Furthermore, conserved splicing sites or conserved sequences were identified. Flanking sequences (31 bp in total) of animal SMNDC1 genes were analyzed to show their consensus in WebLogo and multiple alignment. According to the results, five representative splice sites were identified. (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eA, B).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eExpression profile analysis of animal SMNDC1s\u003c/h2\u003e \u003cp\u003eTo further investigate the potential functions of animal SMNDC1 in response to developmental cues or disease correlations, we analyzed the expression patterns of SMNDC1 genes from Homo sapiens and Mus musculus. In this work, we reconstructed the expression profiles of SMNDC1 in various biological aspects, such as developmental stages, different tissues and cell types, and disease conditions by using the BAR Heat Mapper Plus tool (Supplementary Figures \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e\u0026ndash;S6).The data of Homo sapiens disease proteomics expression showed that SMNDC1 protein had high expression abundance in multiple cancer types, including breast cancer (breast tumor luminal, HER2 positive breast carcinoma and triple-negative breast cancer), colon cancer (colon adenocarcinoma and colon mucinous adenocarcinoma) and rectal cancer (rectal cell carcinoma and rectal mucinous adenocarcinoma) (Figure \u003cspan refid=\"MOESM3\" class=\"InternalRef\"\u003eS3\u003c/span\u003e). Moreover, the SMNDC1 transcript of humans is widely expressed in whole body tissues, including skeletal muscle, adult brain, spinal cord, testis, liver, ovary and lung (Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003e, Figure \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003e), while mouse SMNDC1 is highly expressed in brain tissue (Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003e, Figure \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003e). Furthermore, cell type expression analysis showed that human SMNDC1 was highly expressed in granulocyte monocyte progenitor cells, hematopoietic multipotent progenitor cells and hematopoietic stem cells, while mouse SMNDC1 was expressed in naive thymus-derived CD4-positive, alpha-beta T cells, embryonic stem cells, accumulated in induced T-regulatory cells and T-helper 17 cells (Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003e, Figure \u003cspan refid=\"MOESM4\" class=\"InternalRef\"\u003eS4\u003c/span\u003e). On the other hand, human SMNDC1 was highly expressed in the fetal period and downregulated in the juvenile period (Figure \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e), while mouse SMNDC1 was highly expressed in the embryonic period but did not abundantly accumulate in the fetus (Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003e, Figure \u003cspan refid=\"MOESM4\" class=\"InternalRef\"\u003eS4\u003c/span\u003e). In addition, we will pay more attention to the expression of the SMNDC1 gene in cancer and other diseases. Specifically, transcriptome data revealed that human SMNDC1 was expressed at higher gene expression levels in cancer tissues than in normal paracarcinoma tissue and normal tissues (Figure \u003cspan refid=\"MOESM3\" class=\"InternalRef\"\u003eS3\u003c/span\u003e). Among them, human SMNDC1 had the highest expression level in ovarian adenocarcinoma, followed by esophageal adenocarcinoma (Figure \u003cspan refid=\"MOESM3\" class=\"InternalRef\"\u003eS3\u003c/span\u003e), while the expression abundance of this protein was enriched in breast, colon and rectal cancer. In summary, we found that SMNDC1 is highly expressed in ovarian adenocarcinoma and digestive system diseases, and may be used as a valuable diagnostic or therapeutic protein target in clinical treatment.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"DISCUSSION","content":"\u003cp\u003eAS is the main mechanism for maintaining protein diversity [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e], and abnormalities in AS can lead to the occurrence of many diseases [\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e, \u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]. Specifically, abnormal AS promotes all stages of tumorigenesis, including cell proliferation [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e], apoptosis [\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e], epithelial-mesenchymal transitions [\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e], tumor invasion and tumor metastasis [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e, \u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e]. Increasing studies have shown that alternative splicing may be used as a new biomarker in oncology and provide a large number of new targets for drug development, which is of great value for improving the prognosis of cancer patients [\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e].SMNDC1 is one of the key spliceosomes, and an in-depth comparison and phylogenetic analysis of the animal SMNDC1 family were conducted, which can provide a more comprehensive and in-depth understanding of the function of SMNDC1 in animals.\u003c/p\u003e \u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003eAssessment of phylogeny relationships and putative functions in animal SMNDC1s\u003c/h2\u003e \u003cp\u003eSMNDC1 has been identified as an essential component of the spliceosome complex [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. In the present work, we successfully identified 110 SMNDC1 genes from 61 animal species and reconstructed their phylogenetic relationships of these selected genes. SMNDC1 proteins can be broadly divided into four groups, including primates, rodents and lagomorphs, other mammals, and other vertebrates, which are closely related to the evolution of animal lineages. Moreover, only six species SMNDC1 genes contained 2 copies (Supplementary Table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e), and analysis of the protein structures and protein domains of these cDNAs revealed that this gene family maintains conserved functions (Figs.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e,\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e). In addition, the conservative splicing pattern of animal SMNDC1 s indicates that most transcriptional subtypes of animal SMNDC1 tend to form N-terminal truncated protein types (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e). Transcriptional isoforms of SMNDC1 in animals share similar gene structures, suggesting that they may have similar functions in regulating gene expression and protein interactions. On the other hand, a previous study showed that different spliceosomes have different biological functions. For instance, two splicing isoforms of ZNF148 have different effects on the proliferation, invasion and migration of human colorectal cancer cells, and exert mutual antagonistic effects [\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e]. In addition, a previous study reported that SMNDC1 is critical for regulating ovarian cancer tumor growth and metastasis [\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e]. We expect that different spliceosomes of SMNDC1 may be a potential target for the treatment of ovarian cancer, however, the isoform function of SMNDC1 still needs further study. Our work showed that the SMNDC1 proteins of animals have an SMN (Tudor) domain (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e), which has affinity for Sm ribonucleoproteins and is further involved in the splicing process [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. Studies have reported that SMN plays a key role in the assembly of uridine-rich small nuclear ribonucleoprotein complexes [\u003cspan additionalcitationids=\"CR41\" citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e] and in pre-mRNA splicing [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e, \u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e]. In our protein interaction work, we found that human SMNDC1 can interact with SNRNP200 (small nuclear ribonucleoprotein U5 subunit 200). SNRNP200 is closely related to the splicing of precursor mRNA, and can regulate the expression of related genes by affecting splicing, thereby affecting cell proliferation [\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e]. In addition, SNRNP200 also plays an important role in the pathogenesis of hereditary retinitis pigmentosa [\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e] and acute myeloid leukemia [\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e].The above results reveal that the interaction between SMNDC1 and SNRNP200 may play an important role in the function of SNRNP200 and its impact on diseases. Furthermore, SMNDC1 in mammals can interact with multiple splicing factors and pre-mRNA processing factors, suggesting that it plays an important role in AS regulation. In yeast, there is no SMN domain in the SMNDC1 protein and only one tudor-3 domain (Figure S7), which may prevent it from interacting with most alternative splicing factors.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eFunctional diversity of animal SMNDC1 s based on their differential expression pattern\u003c/h2\u003e \u003cp\u003eSMNDC1 is a survival motor neuron protein that is required for spliceosome assembly [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. Here, previous proteomic analysis revealed that SMNDC1 is critical for regulating ovarian cancer tumor growth and metastasis, consistent with its high expression level in ovarian adenocarcinoma (Figure \u003cspan refid=\"MOESM3\" class=\"InternalRef\"\u003eS3\u003c/span\u003e), which will provide a new target and direction for anticancer drug development [\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e]. Splicing factors are frequently overexpressed in cancer [\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e]. Meanwhile, based on available proteomics datasets, our findings showed that high expression of SMNDC1 was observed in breast, colon and rectal cancers, implying its potential functional role in cancer development in these organs. In addition, based on the fact that SMNDC1 is required for splicing catalysis in vitro [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e], researchers speculate that its poison exons may influence the extensive intron retention characteristic of most cancers [\u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e, \u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e]. Analysis of RNA-seq data from 512 lung adenocarcinoma samples showed that: low SMNDC1 poison exon inclusion was associated with notably widespread reductions in intron retention, further experimental validation showed that SMNDC1 poison exons control SMNDC1 expression to regulate intron retention [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]. This discovery may provide a new perspective for developing new treatments and defeating cancer. Spinal muscular atrophy (SMA) is a degenerative neuromuscular disease with muscle weakness and muscle atrophy, caused by deletion or mutation of the SMN1 gene. Its incidence in neonates is estimated to be 1:6 000 to 1:10 000 [\u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e, \u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e]. Moreover, SMNDC1 is a paralogue of the SMN1 gene, and may share a cellular function similar to that of the SMN1 gene. However, in spinal muscular atrophy being overshadowed by SMN, the biological function of SMNDC1 may not be fully elucidated [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]. Interestingly, SMNDC1 is highly expressed in the spinal cord and skeletal muscle tissue of Homo sapiens, which is consistent with the results of previous studies [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. In addition, SMNDC1 is mainly expressed in fetal skeletal muscle tissue and not in adult tissue. These results indicated that the deletion or mutation of SMNDC1 may be another key factor in spinal muscular atrophy (SMA). Furthermore, in this work, we obtained comprehensive information on SMDNC1 AS in multiple animals, but biological experiments are still needed to validate these new predictions. The above findings may allow scientists to better identify biomarkers of disease substances and therapeutic targets. For example, SMNDC1 mRNA is a target of FMRP, and this result could complement the current understanding of the etiology of FXS [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e].AS has become one of the hotspots in the era of functional genomics. The form of AS can be found by comparing transcripts and genomes, however, with the increasing amount of research samples and data analysis, the development of high-throughput experimental technology is particularly important [\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e]. The isoform level of SMNDC1 has not been thoroughly studied. Hence, it is necessary to further study the expression profile of animal smndc1 isoforms through SWATH-MS (sequential window acquisition of all theoretical mass spectra)-based quantitative approaches [\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e]. Our successful identification of the biological functions of the splicing-related protein SRP in plants, will provide a reference for us to further study the specific function of each SMNDC1 transcript isoform [\u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eComparison of SMNDC1 in animals, yeast and plants\u003c/h2\u003e \u003cp\u003eAlthough the splicing machinery is fairly conserved among eukaryotic species, the splicing mechanisms of humans, yeast and Arabidopsis are not identical. In particular, our work further analyzed and compared the genome structure and splice site patterns of SMNDC1 s from humans, yeast and Arabidopsis. Based on the results, we found that the SMN domain is retained between humans, Arabidopsis and rice, while a Tuor3 domain is present in yeast (Figure S7). Interestingly, the three exons encoding the domain of SMNDC1 were identical between the three species.\u003c/p\u003e \u003c/div\u003e "},{"header":"EXPERIMENTAL METHODS","content":"\u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003cdiv id=\"Sec14\" class=\"Section3\"\u003e \u003ch2\u003eSequence identification and collection of the animal SMNDC1 proteins\u003c/h2\u003e \u003cp\u003eThe SMNDC1 protein sequence (ENST00000369592.1) of Homo sapiens was used as a reference to perform the BLASTp search with an e-value cut of =\u0026thinsp;1e\u0026thinsp;\u0026minus;\u0026thinsp;10 against all available animal genome sequences from the Ensembl database(\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://asia.ensembl.org/index.html\u003c/span\u003e\u003cspan address=\"http://asia.ensembl.org/index.html\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) as described previously. The obtained protein regions were predicted by the online software HMMER (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ebi.ac.uk/Tools/hmmer/search/phmmer\u003c/span\u003e\u003cspan address=\"https://www.ebi.ac.uk/Tools/hmmer/search/phmmer\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). Consequently, the phylogenetic tree was constructed using the Bayesian method based on the amino acid sequences of 72 SMNDC1 members from 66 animal species.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003ePhylogenetic analysis of the SMNDC1 gene family in animals\u003c/h2\u003e \u003cp\u003eThe amino acid sequences of 72 SMNDC1 genes from 66 animal species were used for phylogenetic analysis by using the Bayesian method for genes with different transcript isoforms, the one with the longest protein coding sequence was used. Multiple sequence alignments of all selected SMNDC1 sequences were carried out using Mus-cle v3.8. Bayesian methods were used to construct a rooted phylogenetic tree of the SMNDC1 proteins using Mrbayes3.2. Maximum likelihood methods were also used to construct an additional tree by PhyML v3.0 for validating the result from the Bayesian tree [\u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e54\u003c/span\u003e]. The phylogenetic trees were edited using FigTree v1.4.3 [\u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e55\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003eAnalysis of Gene Structures, Protein Domains and conserved motif\u003c/h2\u003e \u003cp\u003eGene structure and cDNA conserved motifs were identified by the MEME online tool (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://meme-suite.org/tools/meme\u003c/span\u003e\u003cspan address=\"http://meme-suite.org/tools/meme\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) (Bailey et al. 2009). Protein domains were predicted by HMMER website (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ebi.ac.uk/Tools/hmmer/\u003c/span\u003e\u003cspan address=\"https://www.ebi.ac.uk/Tools/hmmer/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) [\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e] and were drawn using TBtools [\u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e57\u003c/span\u003e]. and the exon\u0026ndash;intron structures of all genes were downloaded and reconstructed from the Ensembl database.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003eAnalysis of Protein Interaction Networks\u003c/h2\u003e \u003cp\u003eThe protein sequences of humans (ENSP00000363129.3), Mus musculus (ENSMUSP00000156644.1) and Saccharomyces cerevisiae (YLR298C_mRNA) were selected to obtain the interaction network on the STRING web server (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://string-db.org/\u003c/span\u003e\u003cspan address=\"https://string-db.org/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) [\u003cspan citationid=\"CR58\" class=\"CitationRef\"\u003e58\u003c/span\u003e]. Finally, the predicted functional partners of each SMNDC1 protein were presented in the form of an interaction network drawn by Cytoscape 3.8 software.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec18\" class=\"Section2\"\u003e \u003ch2\u003eAS Profile Analysis and Identification of Conserved Splice Sites\u003c/h2\u003e \u003cp\u003eAll available alternative transcripts of animal SMNDC1 genes were downloaded from the Ensembl database. All available splicing isoforms of animal SMNDC1 genes were obtained again from Ensembl database. Selected splice junction sequences (15 bp on each side) were further examined using BLAST. Consensus sequences at representative splice sites were analyzed and visually represented by using WebLogo v3.0 (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://weblogo.berkeley.edu/logo.cgi\u003c/span\u003e\u003cspan address=\"https://weblogo.berkeley.edu/logo.cgi\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e)\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec19\" class=\"Section2\"\u003e \u003ch2\u003eExpression Analysis of SMNDC1 from Online Microarray Datasets\u003c/h2\u003e \u003cp\u003eExpression data for animal SMNDC1 family members were downloaded from the Expression Atlas (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ebi.ac.uk/gxa/home\u003c/span\u003e\u003cspan address=\"https://www.ebi.ac.uk/gxa/home\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). The retrieved expression data were reorganized and presented as heatmaps by using online BAR HeatMapper Plus software (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://bar.utoronto.ca/ntools/cgi-bin/ntools_heatmapper_plus.cgi\u003c/span\u003e\u003cspan address=\"http://bar.utoronto.ca/ntools/cgi-bin/ntools_heatmapper_plus.cgi\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e"},{"header":"CONCLUSION","content":"\u003cp\u003eIn this study, we identified a total of 110 SMNDC1 genes from 61 animal species and comprehensively analyzed their phylogenetic relationships, genomic organization, motif and protein domain enrichment and splicing pattern conservation, providing a foundation for molecular research on SMNDC1 proteins with respect to their roles in human diseases investigated in mammalian cell lines or animal models. In conclusion, the study of SMNDC1 is of great significance not only for the elucidation of related mechanisms, but also for the diagnosis and treatment of related diseases.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eCOMPETING INTERESTS\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThere are no competing interests to declare.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAUTHOR CONTRIBUTIONS\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eConceptualization, CS, HMW, and BX-H; writing original draft preparation, OY-GJ, Y-NL, BX-H and CS; writing review and editing, OY-GJ, HMW, Y-NL, CS and M-XC; funding, M-XC. The final version of the manuscript was agreed by all authors.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eACKNOWLEDGMENT\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis work was supported by the Program for Science Technology and Innovation Committee of Shenzhen (2021N062-JCYJ20210324115408023), the National Natural Science Foundation of China (NSFC32001932), and the Hong Kong Research Grant Council (AoE/M-05/12, AoE/M-403/16, GRF12100318, 12103219, 12103220).\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eKornblihtt AR, Vibe-Pedersen K, Baralle FE: Human fibronectin: molecular cloning evidence for two mRNA species differing by an internal segment coding for a structural domain. EMBO J 1984, 3(1):221\u0026ndash;226.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang ET, Sandberg R, Luo S, Khrebtukova I, Zhang L, Mayr C, Kingsmore SF, Schroth GP, Burge CB: Alternative isoform regulation in human tissue transcriptomes. Nature 2008, 456(7221):470\u0026ndash;476.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePal S, Gupta R, Davuluri RV: Alternative transcription and alternative splicing in cancer. Pharmacol Ther 2012, 136(3):283\u0026ndash;294.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eUhlen M, Fagerberg L, Hallstrom BM, Lindskog C, Oksvold P, Mardinoglu A, Sivertsson A, Kampf C, Sjostedt E, Asplund A \u003cem\u003eet al\u003c/em\u003e: Proteomics. Tissue-based map of the human proteome. \u003cem\u003eScience\u003c/em\u003e 2015, 347(6220):1260419.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHu Z, Scott HS, Qin G, Zheng G, Chu X, Xie L, Adelson DL, Oftedal BE, Venugopal P, Babic M \u003cem\u003eet al\u003c/em\u003e: Revealing Missing Human Protein Isoforms Based on Ab Initio Prediction, RNA-seq and Proteomics. Sci Rep 2015, 5:10940.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWong ACH, Rasko JEJ, Wong JJ: We skip to work: alternative splicing in normal and malignant myelopoiesis. Leukemia 2018, 32(5):1081\u0026ndash;1093.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMatsuda T, Namura A, Oinuma I: Dynamic spatiotemporal patterns of alternative splicing of an F-actin scaffold protein, afadin, during murine development. Gene 2019, 689:56\u0026ndash;68.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNakka K, Ghigna C, Gabellini D, Dilworth FJ: Diversification of the muscle proteome through alternative splicing. Skelet Muscle 2018, 8(1):8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDisset A, Bourgeois CF, Benmalek N, Claustres M, Stevenin J, Tuffery-Giraud S: An exon skipping-associated nonsense mutation in the dystrophin gene uncovers a complex interplay between multiple antagonistic splicing elements. Hum Mol Genet 2006, 15(6):999\u0026ndash;1013.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCartegni L, Krainer AR: Disruption of an SF2/ASF-dependent exonic splicing enhancer in SMN2 causes spinal muscular atrophy in the absence of SMN1. Nat Genet 2002, 30(4):377\u0026ndash;384.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSvasti S, Suwanmanee T, Fucharoen S, Moulton HM, Nelson MH, Maeda N, Smithies O, Kole R: RNA repair restores hemoglobin expression in IVS2-654 thalassemic mice. Proc Natl Acad Sci U S A 2009, 106(4):1205\u0026ndash;1210.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLin X, Miller JW, Mankodi A, Kanadia RN, Yuan Y, Moxley RT, Swanson MS, Thornton CA: Failure of MBNL1-dependent post-natal splicing transitions in myotonic dystrophy. Hum Mol Genet 2006, 15(13):2087\u0026ndash;2097.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWilliams C, Hoppe HJ, Rezgui D, Strickland M, Forbes BE, Grutzner F, Frago S, Ellis RZ, Wattana-Amorn P, Prince SN \u003cem\u003eet al\u003c/em\u003e: An exon splice enhancer primes IGF2:IGF2R binding site structure and function evolution. Science 2012, 338(6111):1209\u0026ndash;1213.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang GS, Cooper TA: Splicing in disease: disruption of the splicing code and the decoding machinery. Nat Rev Genet 2007, 8(10):749\u0026ndash;761.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChabot B, Shkreta L: Defective control of pre-messenger RNA splicing in human disease. J Cell Biol 2016, 212(1):13\u0026ndash;27.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhou Z, Licklider LJ, Gygi SP, Reed R: Comprehensive proteomic analysis of the human spliceosome. Nature 2002, 419(6903):182\u0026ndash;185.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWill CL, Luhrmann R: Spliceosomal UsnRNP biogenesis, structure and function. Curr Opin Cell Biol 2001, 13(3):290\u0026ndash;301.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKolb SJ, Battle DJ, Dreyfuss G: Molecular functions of the SMN complex. J Child Neurol 2007, 22(8):990\u0026ndash;994.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCote J, Richard S: Tudor domains bind symmetrical dimethylated arginines. J Biol Chem 2005, 280(31):28476\u0026ndash;28483.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTalbot K, Miguel-Aliaga I, Mohaghegh P, Ponting CP, Davies KE: Characterization of a gene encoding survival motor neuron (SMN)-related protein, a constituent of the spliceosome complex. Hum Mol Genet 1998, 7(13):2149\u0026ndash;2156.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLefebvre S, Burglen L, Reboullet S, Clermont O, Burlet P, Viollet L, Benichou B, Cruaud C, Millasseau P, Zeviani M \u003cem\u003eet al\u003c/em\u003e: Identification and characterization of a spinal muscular atrophy-determining gene. Cell 1995, 80(1):155\u0026ndash;165.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMeister G, Hannus S, Plottner O, Baars T, Hartmann E, Fakan S, Laggerbauer B, Fischer U: SMNrp is an essential pre-mRNA splicing factor required for the formation of the mature spliceosome. EMBO J 2001, 20(9):2304\u0026ndash;2314.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNeubauer G, King A, Rappsilber J, Calvio C, Watson M, Ajuh P, Sleeman J, Lamond A, Mann M: Mass spectrometry and EST-database searching allows characterization of the multi-protein spliceosome complex. Nat Genet 1998, 20(1):46\u0026ndash;50.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRappsilber J, Ajuh P, Lamond AI, Mann M: SPF30 is an essential human splicing factor required for assembly of the U4/U5/U6 tri-small nuclear ribonucleoprotein into the spliceosome. J Biol Chem 2001, 276(33):31142\u0026ndash;31150.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThomas JD, Polaski JT, Feng Q, De Neef EJ, Hoppe ER, McSharry MV, Pangallo J, Gabel AM, Belleville AE, Watson J \u003cem\u003eet al\u003c/em\u003e: RNA isoform screens uncover the essentiality and tumor-suppressor activity of ultraconserved poison exons. Nat Genet 2020, 52(1):84\u0026ndash;94.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShahid R, Bugaut A, Balasubramanian S: The BCL-2 5' untranslated region contains an RNA G-quadruplex-forming motif that modulates protein expression. Biochemistry 2010, 49(38):8300\u0026ndash;8306.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMcAninch DS, Heinaman AM, Lang CN, Moss KR, Bassell GJ, Rita Mihailescu M, Evans TL: Fragile X mental retardation protein recognizes a G quadruplex structure within the survival motor neuron domain containing 1 mRNA 5'-UTR. Mol Biosyst 2017, 13(8):1448\u0026ndash;1457.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang D, Yang JF, Gao B, Liu TY, Hao GF, Yang GF, Fu LJ, Chen MX, Zhang J: Identification, evolution and alternative splicing profile analysis of the splicing factor 30 (SPF30) in plant species. Planta 2019, 249(6):1997\u0026ndash;2014.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi Y, Sun N, Lu Z, Sun S, Huang J, Chen Z, He J: Prognostic alternative mRNA splicing signature in non-small cell lung cancer. Cancer Lett 2017, 393:40\u0026ndash;51.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDvinge H, Kim E, Abdel-Wahab O, Bradley RK: RNA splicing factors as oncoproteins and tumour suppressors. Nat Rev Cancer 2016, 16(7):413\u0026ndash;430.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eScotti MM, Swanson MS: RNA mis-splicing in disease. Nat Rev Genet 2016, 17(1):19\u0026ndash;32.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXie R, Chen X, Chen Z, Huang M, Dong W, Gu P, Zhang J, Zhou Q, Dong W, Han J \u003cem\u003eet al\u003c/em\u003e: Polypyrimidine tract binding protein 1 promotes lymphatic metastasis and proliferation of bladder cancer via alternative splicing of MEIS2 and PKM. Cancer Lett 2019, 449:31\u0026ndash;44.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTyson-Capper A, Gautrey H: Regulation of Mcl-1 alternative splicing by hnRNP F, H1 and K in breast cancer cells. RNA Biol 2018, 15(12):1448\u0026ndash;1457.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePradella D, Naro C, Sette C, Ghigna C: EMT and stemness: flexible processes tuned by alternative splicing in development and cancer progression. Mol Cancer 2017, 16(1):8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang F, Fu X, Chen P, Wu P, Fan X, Li N, Zhu H, Jia TT, Ji H, Wang Z \u003cem\u003eet al\u003c/em\u003e: SPSB1-mediated HnRNP A1 ubiquitylation regulates alternative splicing and cell migration in EGF signaling. Cell Res 2017, 27(4):540\u0026ndash;558.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen L, Yao Y, Sun L, Zhou J, Miao M, Luo S, Deng G, Li J, Wang J, Tang J: Snail Driving Alternative Splicing of CD44 by ESRP1 Enhances Invasion and Migration in Epithelial Ovarian Cancer. Cell Physiol Biochem 2017, 43(6):2489\u0026ndash;2504.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBlencowe BJ: Alternative splicing: new insights from global analyses. Cell 2006, 126(1):37\u0026ndash;47.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu Y, Huang W, Gao X, Kuang F: Regulation between two alternative splicing isoforms ZNF148(FL) and ZNF148(DeltaN), and their roles in the apoptosis and invasion of colorectal cancer. Pathol Res Pract 2019, 215(2):272\u0026ndash;277.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGiri K, Shameer K, Zimmermann MT, Saha S, Chakraborty PK, Sharma A, Arvizo RR, Madden BJ, McCormick DJ, Kocher JP \u003cem\u003eet al\u003c/em\u003e: Understanding protein-nanoparticle interaction: a new gateway to disease therapeutics. Bioconjug Chem 2014, 25(6):1078\u0026ndash;1090.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBuhler D, Raker V, Luhrmann R, Fischer U: Essential role for the tudor domain of SMN in spliceosomal U snRNP assembly: implications for spinal muscular atrophy. Hum Mol Genet 1999, 8(13):2351\u0026ndash;2357.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePellizzoni L, Kataoka N, Charroux B, Dreyfuss G: A novel function for SMN, the spinal muscular atrophy disease gene product, in pre-mRNA splicing. Cell 1998, 95(5):615\u0026ndash;624.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePellizzoni L, Yong J, Dreyfuss G: Essential role for the SMN complex in the specificity of snRNP assembly. Science 2002, 298(5599):1775\u0026ndash;1779.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEhsani A, Alluin JV, Rossi JJ: Cell cycle abnormalities associated with differential perturbations of the human U5 snRNP associated U5-200kD RNA helicase. PLoS One 2013, 8(4):e62125.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu T, Jin X, Zhang X, Yuan H, Cheng J, Lee J, Zhang B, Zhang M, Wu J, Wang L \u003cem\u003eet al\u003c/em\u003e: A novel missense SNRNP200 mutation associated with autosomal dominant retinitis pigmentosa in a Chinese family. PLoS One 2012, 7(9):e45464.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGillissen MA, Kedde M, Jong G, Moiset G, Yasuda E, Levie SE, Bakker AQ, Claassen YB, Wagner K, Bohne M \u003cem\u003eet al\u003c/em\u003e: AML-specific cytotoxic antibodies in patients with durable graft-versus-leukemia responses. Blood 2018, 131(1):131\u0026ndash;143.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eUrbanski LM, Leclair N, Anczukow O: Alternative-splicing defects in cancer: Splicing regulators and their downstream targets, guiding the way to novel cancer therapeutics. Wiley Interdiscip Rev RNA 2018, 9(4):e1476.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDvinge H, Bradley RK: Widespread intron retention diversifies most cancer transcriptomes. Genome Med 2015, 7(1):45.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJung H, Lee D, Lee J, Park D, Kim YJ, Park WY, Hong D, Park PJ, Lee E: Intron retention is a widespread mechanism of tumor-suppressor inactivation. Nat Genet 2015, 47(11):1242\u0026ndash;1248.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLunn MR, Wang CH: Spinal muscular atrophy. Lancet 2008, 371(9630):2120\u0026ndash;2133.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSchorling DC, Pechmann A, Kirschner J: Advances in Treatment of Spinal Muscular Atrophy - New Phenotypes, New Challenges, New Implications for Care. J Neuromuscul Dis 2020, 7(1):1\u0026ndash;13.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhu FY, Chen MX, Ye NH, Shi L, Ma KL, Yang JF, Cao YY, Zhang Y, Yoshida T, Fernie AR \u003cem\u003eet al\u003c/em\u003e: Proteogenomic analysis reveals alternative splicing and translation as part of the abscisic acid response in Arabidopsis seedlings. Plant J 2017, 91(3):518\u0026ndash;533.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhu FY, Chen MX, Chan WL, Yang F, Tian Y, Song T, Xie LJ, Zhou Y, Xiao S, Zhang J \u003cem\u003eet al\u003c/em\u003e: SWATH-MS quantitative proteomic investigation of nitrogen starvation in Arabidopsis reveals new aspects of plant nitrogen stress responses. J Proteomics 2018, 187:161\u0026ndash;170.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen MX, Mei LC, Wang F, Boyagane Dewayalage IKW, Yang JF, Dai L, Yang GF, Gao B, Cheng CL, Liu YG \u003cem\u003eet al\u003c/em\u003e: PlantSPEAD: a web resource towards comparatively analysing stress-responsive expression of splicing-related proteins in plant. Plant Biotechnol J 2021, 19(2):227\u0026ndash;229.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGuindon S, Dufayard JF, Lefort V, Anisimova M, Hordijk W, Gascuel O. New algorithms and methods to estimate maximum-likelihood phylogenies: assessing the performance of PhyML 3.0. Syst Biol. 2010, 59(3):307\u0026ndash;21.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMorariu V, Srinivasan B, Raykar V, Duraiswami R, Davis L. Automatic online tuning for fast Gaussian summation. Adv Neural Inf Process Syst. 2008. 1113\u0026ndash;1120.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePotter SC, Luciani A, Eddy SR, Park Y, Lopez R, Finn RD. HMMER web server: 2018 update. Nucleic Acids Res. 2018;46(W1): W200-W204.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen MX, Zhang KL, Zhang M, Das D, Fang YM, Dai L, Zhang J, Zhu FY. Alternative splicing and its regulatory role in woody plants. Tree Physiol. 2020, 40(11):1475\u0026ndash;1486.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSzklarczyk, D., Franceschini, A., Wyder, S., Forslund, K., Heller, D., Huerta-Cepas, J., et al. STRING V10: Protein-Protein Interaction Networks, Integrated over the Tree of Life. Nucleic Acids Res. 2015, 43, D447\u0026ndash;D452.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Alternative splicing, phylogenetics, splicing factor, splice site selection, SMNDC1","lastPublishedDoi":"10.21203/rs.3.rs-3896856/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-3896856/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eAlternative splicing is the process of multiple mRNAs from a single pre mRNA under the action of the spliceosome and other splicing factors. SMNDC1 (survival motor neuron domain containing 1) has been identified as a constituent of the spliceosome complex. Previous studies indicated that SMNDC1 is required for splicing catalysis in vitro and regulates intron retention in cancer. However, the phylogenetic relationships and expression profiles of SMNDC1 have not been systematically studied in the animal kingdom. To this end, in our work, the phylogenetic analysis of SMNDC1 genes was widely performed in the animal kingdom. Specifically, a total of 72 SMNDC1 genes were identified from 66 animal species. Bioinformatics analysis showed that the gene structure and function of SMNDC1 proteins are relatively conserved, and only a few members have two copies. In particular, the human SMNDC1 gene is highly expressed in multiple cancer types, including breast cancer, colon cancer and rectal cancer, indicating that SMNDC1 may play an essential role in cancer development and may be used as a valuable diagnostic or therapeutic protein target in clinical treatment. In summary, our findings facilitated a comprehensive overview of the animal SMNDC1 gene family, and provided a basic data and potential clues for the further study of molecular functions of SMNDC1.\u003c/p\u003e","manuscriptTitle":"Phylogenetic comparison and splice site conservation of the animal SMNDC1 gene family","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-02-26 07:48:20","doi":"10.21203/rs.3.rs-3896856/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"adbcda0b-18c6-4b40-929b-17dd5870018b","owner":[],"postedDate":"February 26th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":28956309,"name":"Biological sciences/Genetics"},{"id":28956310,"name":"Biological sciences/Genetics/Rna splicing"}],"tags":[],"updatedAt":"2024-04-16T05:57:26+00:00","versionOfRecord":[],"versionCreatedAt":"2024-02-26 07:48:20","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-3896856","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-3896856","identity":"rs-3896856","version":["v1"]},"buildId":"CiT4i_kKBbxQbnFL0ufpk","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-08-12T06:43:03.944938+00:00
License: CC-BY-4.0