Vector Semantics of Multidomain Protein Architectures

preprint OA: closed CC-BY-NC-ND-4.0
📄 Open PDF Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-16

This study utilized vector semantics to model protein domain content, finding that semantic similarity of multidomain architectures better predicts shared function than domain content similarity alone, and even identifies functionally similar proteins with no shared domains.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-16 · read from full text

The paper studies whether multidomain protein function is determined by domain content alone or whether a domain’s functional contribution depends on neighboring domains. Using NLP-inspired vector embedding (“vector semantics”) models to represent domain architectures with local contextual signals, the authors found that multidomain architectures that are semantically similar in the embedding space share more functional attributes than architectures chosen by domain-content similarity alone, indicating that context matters. They further observed that the models can identify pairs of multidomain architectures with high functional similarity even when the two proteins share no domains, which they interpret as evidence of “domain synonyms” captured by the embeddings. The paper’s main limitation is that it is an embedding-based exploratory framework rather than direct mechanistic validation of how specific domain-context interactions drive function. This paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Multidomain proteins are mosaics of domains , protein modules that are associated with a specific structure or function and are found in diverse combinations. This modular organization facilitates the evolution of novel protein functions, but the principles that govern the relationship between the domain content of a protein and its function is poorly understood. In particular, do domains always contribute the same function, or does the functional contribution of a domain depend on the neighboring domains in the protein? To answer this question, we used vector embeddings, which account for local contextual signals, to model the protein domain content of multidomain proteins. We observe that multidomain architectures that are semantically similar share more functional attributes than multidomain architectures selected based on domain content similarity, alone, suggesting that context is important for understanding the relationship between domain content and protein function. Surprisingly, vector semantics also identified multidomain architecture pairs with significantly high functional similarity, despite having no domains in common at all, suggesting that vector semantics may be discovering domain “synonyms”. Taken together, our results underscore the importance of contextual models for understanding the interplay between domain architecture evolution and functional innovation in multidomain proteins.
Full text 59,800 characters · extracted from preprint-html · click to expand
Vector Semantics of Multidomain Protein Architectures | bioRxiv /* */ /* */ <!-- <!-- /*! * yepnope1.5.4 * (c) WTFPL, GPLv2 */ (function(a,b,c){function d(a){return"[object Function]"==o.call(a)}function e(a){return"string"==typeof a}function f(){}function g(a){return!a||"loaded"==a||"complete"==a||"uninitialized"==a}function h(){var a=p.shift();q=1,a?a.t?m(function(){("c"==a.t?B.injectCss:B.injectJs)(a.s,0,a.a,a.x,a.e,1)},0):(a(),h()):q=0}function i(a,c,d,e,f,i,j){function k(b){if(!o&&g(l.readyState)&&(u.r=o=1,!q&&h(),l.onload=l.onreadystatechange=null,b)){"img"!=a&&m(function(){t.removeChild(l)},50);for(var d in y[c])y[c].hasOwnProperty(d)&&y[c][d].onload()}}var j=j||B.errorTimeout,l=b.createElement(a),o=0,r=0,u={t:d,s:c,e:f,a:i,x:j};1===y[c]&&(r=1,y[c]=[]),"object"==a?l.data=c:(l.src=c,l.type=a),l.width=l.height="0",l.onerror=l.onload=l.onreadystatechange=function(){k.call(this,r)},p.splice(e,0,u),"img"!=a&&(r||2===y[c]?(t.insertBefore(l,s?null:n),m(k,j)):y[c].push(l))}function j(a,b,c,d,f){return q=0,b=b||"j",e(a)?i("c"==b?v:u,a,b,this.i++,c,d,f):(p.splice(this.i++,0,a),1==p.length&&h()),this}function k(){var a=B;return a.loader={load:j,i:0},a}var l=b.documentElement,m=a.setTimeout,n=b.getElementsByTagName("script")[0],o={}.toString,p=[],q=0,r="MozAppearance"in l.style,s=r&&!!b.createRange().compareNode,t=s?l:n.parentNode,l=a.opera&&"[object Opera]"==o.call(a.opera),l=!!b.attachEvent&&!l,u=r?"object":l?"script":"img",v=l?"script":u,w=Array.isArray||function(a){return"[object Array]"==o.call(a)},x=[],y={},z={timeout:function(a,b){return b.length&&(a.timeout=b[0]),a}},A,B;B=function(a){function b(a){var a=a.split("!"),b=x.length,c=a.pop(),d=a.length,c={url:c,origUrl:c,prefixes:a},e,f,g;for(f=0;f<d;f++)g=a[f].split("="),(e=z[g.shift()])&&(c=e(c,g));for(f=0;f<b;f++)c=x[f](c);return c}function g(a,e,f,g,h){var i=b(a),j=i.autoCallback;i.url.split(".").pop().split("?").shift(),i.bypass||(e&&(e=d(e)?e:e[a]||e[g]||e[a.split("/").pop().split("?")[0]]),i.instead?i.instead(a,e,f,g,h):(y[i.url]?i.noexec=!0:y[i.url]=1,f.load(i.url,i.forceCSS||!i.forceJS&&"css"==i.url.split(".").pop().split("?").shift()?"c":c,i.noexec,i.attrs,i.timeout),(d(e)||d(j))&&f.load(function(){k(),e&&e(i.origUrl,h,g),j&&j(i.origUrl,h,g),y[i.url]=2})))}function h(a,b){function c(a,c){if(a){if(e(a))c||(j=function(){var a=[].slice.call(arguments);k.apply(this,a),l()}),g(a,j,b,0,h);else if(Object(a)===a)for(n in m=function(){var b=0,c;for(c in a)a.hasOwnProperty(c)&&b++;return b}(),a)a.hasOwnProperty(n)&&(!c&&!--m&&(d(j)?j=function(){var a=[].slice.call(arguments);k.apply(this,a),l()}:j[n]=function(a){return function(){var b=[].slice.call(arguments);a&&a.apply(this,b),l()}}(k[n])),g(a[n],j,b,n,h))}else!c&&l()}var h=!!a.test,i=a.load||a.both,j=a.callback||f,k=j,l=a.complete||f,m,n;c(h?a.yep:a.nope,!!i),i&&c(i)}var i,j,l=this.yepnope.loader;if(e(a))g(a,0,l,0);else if(w(a))for(i=0;i (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0];var j=d.createElement(s);var dl=l!='dataLayer'?'&l='+l:'';j.src='//www.googletagmanager.com/gtm.js?id='+i+dl;j.type='text/javascript';j.async=true;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-M677548'); Skip to main content Home About Submit ALERTS / RSS Search for this keyword Advanced Search New Results Vector Semantics of Multidomain Protein Architectures View ORCID Profile Xiaoyue Cui , View ORCID Profile Yuting Xiao , View ORCID Profile Maureen Stolzer , View ORCID Profile Dannie Durand doi: https://doi.org/10.1101/2025.07.07.663606 Xiaoyue Cui 1 Computational Biology, Carnegie Mellon University , Pittsburgh, PA, 15213, USA Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Xiaoyue Cui For correspondence: xiaoyuec{at}andrew.cmu.edu durand{at}cmu.edu Yuting Xiao 2 Biological Sciences, Carnegie Mellon University , Pittsburgh, PA, 15213, USA Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Yuting Xiao Maureen Stolzer 2 Biological Sciences, Carnegie Mellon University , Pittsburgh, PA, 15213, USA Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Maureen Stolzer Dannie Durand 1 Computational Biology, Carnegie Mellon University , Pittsburgh, PA, 15213, USA 2 Biological Sciences, Carnegie Mellon University , Pittsburgh, PA, 15213, USA Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Dannie Durand For correspondence: xiaoyuec{at}andrew.cmu.edu durand{at}cmu.edu Abstract Full Text Info/History Metrics Supplementary material Data/Code Preview PDF Abstract Multidomain proteins are mosaics of domains , protein modules that are associated with a specific structure or function and are found in diverse combinations. This modular organization facilitates the evolution of novel protein functions, but the principles that govern the relationship between the domain content of a protein and its function is poorly understood. In particular, do domains always contribute the same function, or does the functional contribution of a domain depend on the neighboring domains in the protein? To answer this question, we used vector embeddings, which account for local contextual signals, to model the protein domain content of multidomain proteins. We observe that multidomain architectures that are semantically similar share more functional attributes than multidomain architectures selected based on domain content similarity, alone, suggesting that context is important for understanding the relationship between domain content and protein function. Surprisingly, vector semantics also identified multidomain architecture pairs with significantly high functional similarity, despite having no domains in common at all, suggesting that vector semantics may be discovering domain “synonyms”. Taken together, our results underscore the importance of contextual models for understanding the interplay between domain architecture evolution and functional innovation in multidomain proteins. Introduction Multidomain proteins are mosaics of structural or functional units called domains. Domains act as protein modules, are found in otherwise unrelated proteins, and fold independently in many different contexts. The domain architecture (DA) of a protein is an abstract representation consisting of its domains in N- to C-terminal order. For example, the domain architecture of the Mpp6 protein shown in Fig. 1 can be represented as L27-L27-PDZ-SH3-GuK . The domain architecture abstraction is widely used to probe questions of protein evolution, including variation in the domain repertoire across taxonomic lineages [Karev et al., 2004, Tordai et al., 2005 , Ye and Godzik, 2004 , Cromar et al., 2016, Dohmen et al., 2020 ], plasticity in domain order [ Bashton and Chothia, 2002 , Kummerfeld and Teichmann, 2009 , Weiner 3rd et al., 2006], domain occurrence graphs [ Vogel et al., 2005 , Karev et al., 2002, Cromar et al., 2014, Przytycka et al., 2006 ], and domain promiscuity, i.e., the propensity of a domain to co-occur with many other domains [ Marcotte et al., 1999 , Basu et al., 2009 , 2008, Cohen-Gihon et al., 2011, Cromar et al., 2014, Weiner et al., 2008 ]. Download figure Open in new tab Fig. 1. Functional similarity between target domain architectures and the k = 1 nearest neighbor that lacks a common domain. Multidomain sequences evolve via the insertion, internal duplication, and deletion of domains, resulting in families with similar domain architectures, but with some variation in domain content, order and copy number (e.g., the Membrane Associated Guanylate Kinases (MAGuKs) [Te Velthuis et al., 2007 ] shown in Fig. 1 ). The modular organization of multidomain proteins supports rapid evolutionary exploration of new protein functions through the formation of different domain combinations that enable diverse biological roles. At the same time, this exploration is not unconstrained: the number of domain combinations observed in the multidomain repertoire is much smaller than expected by chance [ Yu et al., 2019 , Cui et al., 2022 , Vogel et al., 2005 ]. The principles that govern which domain combinations are realized in nature are poorly understood. Many forces could be contributing to those constraints, including structural compatibility, mutational processes, and cellular environment. Download figure Open in new tab Domain architectures of MAGuK subfamilies exhibit similar domain content in various combinations. (c) Structure of MPP6, with individual domains color coded as in (b). Here we focus on the relationship between the domain content of a protein and its function. In specific cases, the function of multidomain proteins can be understood in terms of the functions of individual domains. For example, the extracellular domains of a receptor kinase recognize a molecular signal; the intercellular kinase domain passes on that signal by phosphorylating an amino acid. More generally, proteins with identical domain architectures have more similar GO annotations, on average, than those that do not and proteins that share at least one domain have greater GO similarity than those that have no domains in common [Radivojac, 2022], suggesting that information is there to exploit. This is further supported by recent successes with protein function prediction methods that incorporate domain content information [e.g. You et al., 2018 , Ibtehaz et al., 2023 , Talo and Bozdag, 2025 , Wang et al., 2025 , Hallee et al., 2024 , Chen and Luo, 2024 , Hu and Zhao, 2025 ]. However, it is not clear whether a given domain confers the same functional properties in all contexts, or whether the functional contribution of a domain varies depending on neighboring domains in the protein. Vector Semantics Here we investigate vector semantic models, developed for natural language processing (NLP), as an analytical framework for probing the relationship between the domain content and the function of a multidomain protein. Just as documents are strings of words, domain architectures are strings of domains, suggesting that NLP models are a promising approach for exploratory analysis of the multidomain universe. Vector semantics are based on the assumption that the meanings of words are implicitly encoded in their usage in the language. The distribution of words across documents is a representation of word usage; hence, words that occur in similar contexts tend to have related meanings. Contextual information can be encoded as a word embedding , where each word corresponds to a vector that represents its distribution in a corpus of documents. The goal is to select a vectorization such that words that are proximal in the vector space will tend to have related meanings. This includes words with similar meanings (e.g. chilly, cool), but also words that have meanings that are dissimilar, but are topically related (e.g., bus, train ; dog, animal ; or cup, coffee ). Embeddings used in information retrieval are sparse, high dimensional encodings. In a term frequency - inverse document frequency (TF-IDF) encoding, each dimension corresponds to a document; the i th component of a word vector represents its frequency in document i , normalized by its frequency across the corpus [ Luhn, 1957 , Sparck Jones, 1972 ]. In a TF-IDF vectorization, words that occur in the same documents will be close in the vector space. In a Positive Pointwise Mutual Information (PPMI) [ Church and Hanks, 1990 ] embedding, each dimension corresponds to a word in the vocabulary; the i th component of a word vector is a function of its frequency of co-occurrence with the i th word in the vocabulary, relative to the independent overall frequencies of the two words. In a PPMI embedding, words will be proximal if they co-occur with the same words. Static language embeddings, e.g. Word2Vec [ Mikolov et al., 2013 ], learns compact vector representations from training data (e.g., document corpora) and more fine-grained contextual information is exploited. Given a word embedding, a document embedding can be constructed, in which each document is represented as a vector obtained by combining the vectors of the words that appear in the document. Here, the goal is a vectorization such that documents on similar subjects are proximal. The words used in proximal documents will tend to have meanings that are related. Document embeddings provide a framework for comparing documents and investigating the structure of a corpus. Multidomain proteins poses challenges for bioinformatic techniques that use sequence analysis to extract functional, structural, and evolutionary information. In families with variable domain architectures, such as the MAGuKs, some sequence segments are homologous and have discernible sequence similarity; other regions are not alignable. As a result, bioinformatic tools based on sequence comparison cannot be directly applied to multidomain protein families with mosaic sequences. At the same time, the variations in domain composition carry information that current tools are not designed to exploit. Examining word context using vector semantics provides information about the relationships between words, and hence documents, that is independent of whether the words contain similar letters or are derived from the same root in an ancient language, suggesting the promise of this approach for multidomain protein architectures Domain embeddings [ Buchan and Jones, 2020 , Melidis and Nejdl, 2021 ] have been used to investigate the potential of vector semantics for transferring functional information between neighboring domains. Moreover, domain embeddings have been shown to improve functional prediction of proteins [ Ibtehaz et al., 2023 , Wang et al., 2025 ]. For instance, Ibtehaz et al. [2023] learned domain embeddings from domain-GO co-occurrence data, achieving state-of-the-art results on GO annotation tasks. Further integrating complementary features such as sequence similarity and protein interaction information enhanced prediction accuracy. These results support the view that domain composition carries meaningful functional signals. Our contributions Here we investigate the use of domain architecture embeddings for studying functional relationships between multidomain proteins, using human domain architectures as a case study. The human genome is studied extensively and has high quality gene and domain annotations. In addition, there are many sources of detailed information about human multidomain protein families from which benchmarks can be constructed. Restricting the study to a single genome results in a relatively small corpus. However, data sets that span multiple genomes have biases that are hard to characterize and correct. It is difficult to determine whether the same domain architecture, found in multiple related genomes, arose through convergent formation of that architecture in independent lineages or via vertical descent from the same domain architecture in a common ancestor. In the latter case, the multiple instances of the domain architecture represent observations of the same event and should be discarded. Uneven taxon sampling in the underlying database further exacerbates this problem. In addition, if the functional roles of the domains change over the course of evolution, combining information from distant genomes would further confound the signal. While the analogy between documents as strings of words and domain architectures are strings of domains is compelling, the characteristic scales of domain architecture data and natural language corpora are very different. In natural language corpora, sentences are typically 15 to 20 words in length, taken from a vocabulary of hundreds of thousands of words. Depending on the application, the document data available is almost unlimited. In contrast, the average length of a human domain architecture is less than 5 (median = 3), drawn from ∼1100 domains. There are approximately 5000 unique domain architectures in the human genome. With this in mind, we experimented with a number of domain embedding strategies, with modifications tuned to the scale of our data. We used two sparse embeddings, TF-IDF and Pointwise Mutual Information (PMI). For natural language applications, pointwise mutual information is typically restricted to positive values, focusing on words that co-occur more often than expected. In domain architectures, underrepresentation of domain pairs can be informative. Further, problems with underflow that occur with low frequency word pairs do not arise because the small scale of DA data. We also experimented with Word2Vec using a skip-gram model, which accounts for local context; i.e., the domain order and content in a window of width 2 w + 1 centered on the current domain. In addition to the default window size of w = 5, we experimented with a smaller window ( w = 1) to allow for the large number of domain architectures of length 5 or shorter. To account for the reduced scale of domain architecture data, we considered smaller dimensionalities (5 and 10), as well as the default value (100). We did not consider contextual embeddings (e.g, BERT representations, Devlin et al. [2018]), which account for contextual differences in meaning (e.g., homonyms) because most domain architectures are too short to benefit from that level of representational learning. We apply this vector semantic framework to ask three questions about the relationship between domain content and protein function. First, do proteins with similar domain content also perform similar functions? To assess the association between domain architecture and protein function, for each embedding, we asked how accurately the function of a multidomain protein is predicted by the GO annotations of its neighbors in the embedding. Next, we asked whether this association is stronger when domain content similarity is assessed using vector embeddings, which carry implicit contextual information, compared with a direct comparison of protein domain content using Jaccard similarity, which provides a measure of the domains shared by two architectures, but does not use any of the additional information captured by a vector semantic embedding. Finally, we considered the case where neighboring domain architectures have no domains in common. In the natural language analogy, embeddings that capture words with similar or related meanings can address the “vocabulary problem”, where an information retrieval request fails because the word usage in the query does not match the word usage in the desired document [ Furnas et al., 1987 ]. We ask whether an analogous “domain vocabulary problem” exists for multidomain protein function and, if so, whether domain architecture embeddings can help to solve it. Materials and methods Dataset and pre-processing Domain annotations (start and end positions) for 58, 023 ENSEMBL protein sequences in the human genome (assembly GRCh38.p13) were downloaded from the SUPERFAMILY database, version 1.75 [Pandurangan et al., 2019]. SUPERFAMILY uses a hierarchical classification based on SCOP, in which domains are grouped into families and superfamilies. Domains within the same superfamily share a structural core, although they may have low sequence similarity. In this work, we used the superfamily-level assignments. The pre-processing module of DomArchov [ Cui et al., 2022 ] was used to extract the domain architectures of these sequences, resulting in a set of N = 5,031 unique DAs, with an average length of 4.2 domains per DA. These architectures comprise a set of 1,109 distinct domain superfamilies, denoted D . GO terms associated with the ENSEMBL protein sequences were obtained from ENSEMBL [Martin et al., 2023] and assigned to the corresponding domain architectures. Most DAs correspond to multiple ENSEMBL sequences. In these cases, the DA is annotated with those GO codes that are associated with at least 50% of all corresponding protein sequences. The list of GO associations is expanded to include to ancestors in the GO hierarchy connected by “is a” and “part of” relations. This resulted in 1,157 distinct Molecular Function (MF), 1,726 Biological Process (BP), and 378 Cellular Component (CC) GO codes are assigned to human DAs. Of 5,031 DAs, 4,244 have at least one GO code. On average, each DA is mapped to 21 GO terms. Within the individual ontologies, 3,910 DAs have at least one MF term; 2,631 DAs have at least one BP term; 1,295 DAs have at least one CC term. Any DA with only a protein-binding (GO:0005515) annotation is removed from analyses for the MF ontology following the practice discussed in Zhou et al. [2019]. Domain Embedding Domain embeddings were constructed from the 5,031 unique DAs using eight variants of three different vectorization strategies, where each domain D i ∈ D is represented as a vector, e ( D i ) = [ e 1 ( D i ), …, e m ( D i )]. The dimensionality, m , depends on how the embedding is constructed. Pointwise mutual information (PMI) The pointwise mutual information [ Church and Hanks, 1990 ] of domain D i followed by D j is defined to be where P ( D i ) is the empirical frequency of domain superfamily D i , and P ( D i , D j ) is the empirical frequency of the bigram D i D j . A pseudocount ψ is added to the count of all bigrams. Following Cui et al. [2022], we use ψ = 0.0009. In the PMI embedding, the j th element of e ( D i ) is e j ( D i ) = PMI( D i , D j ). The dimensionality of this embedding is the number of unique domain superfamilies in the training data, m = | D |. Term frequency-inverse document frequency (TF-IDF) The inverse document frequency [Sparck Jones, 1972 ] of D i is defined as idf( D i ) = N/ df( D i ), where df( D i ) is the number of DAs in which D i occurs and N is the total number of domain architectures in the dataset. For any domain D i and domain architecture A j , the term frequency tf( D i , A j ) is the number of copies of D i in architecture A j [ Luhn, 1957 ]. Here the j th element of e ( D i ) is e j ( D i ) = tf( D i , A j ) · idf( D i ) and m = N . Word2vec Word2Vec domain embeddings were constructed using the Python library gensim [ Řehůřek and Sojka, 2010 ] with the skip-gram model [ Mikolov et al., 2013 ] and a minimum word count of 1. Embeddings were calculated for vector sizes of m ∈ {5, 10, 100} and window sizes of w ∈ {1, 5}, resulting in six different Word2Vec embeddings. Each model is denoted by w2v( m, w ). The models were trained for 15 epochs. Domain architecture embedding A domain architecture embedding was constructed for each of the eight domain embeddings (i.e, TF-IDF, PMI, and 6 Word2Vec models) described above. Given domain architecture A = D 1 … D n , the embedding vector of A is obtained by averaging the embeddings of its constituent domains Assessment For assessments of functional similarity, only annotated multidomain architectures (i.e., DAs with 2 or more domains and at least one GO term) were considered. The numbers of annotated multidomain architectures in each of the three sub-ontologies are 3,193 (MF), 2,066 (BP), and 1,001 (CC). Let A O , O ∈ {MF, BP, CC} denote the set of annotated multidomain architectures in ontology O . Assessment centers around the functional attributes of the k nearest neighbors of a target domain architecture A ∈ A O , O ∈ {MF, BP, CC}. Pairwise domain content similarity Given a pair of domain architectures, A i and A j , proximity in the embedding space is quantified using pairwise cosine similarity, denoted by S C ( v ( A i ), v ( A j )). The similarity of two architectures can also be quantified based on shared domain content, independent of an embedding, using the Jaccard similarity: where n ij is the number of domain copies shared by A i and A j . If A i and A j share no domains, then S J ( A i , A j ) = 0. Functional similarity Comparison of two domain architectures For each ontology, let ℱ O ( A ) denote the set of GO terms annotated to domain architecture A . The functional similarity between domain architectures A i and A j , is defined to be the Jaccard similarity between ℱ O ( A i ) and ℱ O ( A j ): Comparison with sets of domain architectures The GO term annotation of a set domain architectures is defined to be the union of the GO terms associated with each architecture in the set, . The functional similarity between a single architecture A and set of architectures U is To avoid confusion due to use of the Jaccard similarity in two contexts, we use functional similarity to denote the similarity between two sets of GO codes and Jaccard similarity to refer to shared domain content. Annotated domain architecture neighborhoods Given an embedding E , let N O,k ( A ) be the annotated nearest neighborhood of size k for architecture A , separately for each of the three sub-ontologies O . We define this k-neighborhood to be the k DAs in A O that are closest to A in embedding E . Formally, is a k -neighborhood of A if If there are two or more DAs that are equidistant from A , the k-neighborhood may not be unique. In this case, we assign one of the sets of k nearest neighbors to N O,k ( A ) arbitrarily. We further consider the subsets of N O,k ( A ): the shared k-neighborhood , consisting of domain architectures that share at least one domain with A , and the non-sharing k-neighborhood , consisting of domain architectures that share no domains with A , The respective sizes of these neighborhoods are k + and k − , where k + + k − = k . When comparing domain content with Jaccard similarity, without reference to an embedding, the k -neighborhood of A is defined analogously. In this case, all DAs in the k -neighborhood of A have at least one domain in common with A , and thus and . Functional similarity in the k-neighborhood We ask whether multidomain architecture pairs that are close in the embedding tend to have similar functions. To assess whether domain architectures in the k-neighborhood of A have similar functions to A , we calculated , the functional similarity between A and its neighbors. The functional similarity is defined to be in shared neighborhoods and in non-sharing neighborhoods. For a given embedding and a given ontology, the mean functional similarity over all architectures in A O for each O ∈ {MF, BP, CC} provides a measure of how well that embedding places domain architectures with similar functions in close proximity. We quantify performance using precision, recall, and Matthew’s Correlation Coefficients (MCC). Functional consistency within shared k-neighborhoods. To assess how tightly focused sets of nearest neighbors are, we computed two measures for each shared k =neighborhood. First, we considered the mean number of distinct GO terms in the shared k -neighborhood, We also calculated , the mean functional similarity within the shared k -neighborhood obtained from all pairs of DAs in , For each embedding method (or Jaccard), we report the mean value of each measure across all target domain architectures. Statistical tests for non-sharing nearest neighborhoods We assessed the significance of observing neighboring domain architectures that are functionally similar, but have no common domains, using a randomization strategy with two test statistics. Given an embedding and an ontology, let be the set of architectures in A O that have at least one neighbor that with no common domain. The first test statistic is the mean functional similarity, averaged over all target architectures in , The second is the number of such pairs with functional similarity greater than 0.8: where I is the indicator function. For each test statistic, the expected value was estimated from 100,000 null replicates generated as follows: for each target architecture in , an non-sharing partner DA was selected at random, with replacement, from the set of all DAs that share no domain A . The value of the test statistic was calculated for set of null pairs generated in each replicate and averaged over all replicates to obtain the expected value. The empirical p -value is the proportion of replicates in which the null pairs achieve a greater value of the test statistic than the genuine pairs. Results Here, we use vector semantics to investigate the relationship between protein domain content and protein function. We construct domain architecture embeddings using two sparse, high-dimensional encodings (TF-IDF, PMI) and six learned low-dimensional embeddings trained with Word2Vec. For each embedding, we ask: given a target domain architecture, how well do the combined GO terms of its k nearest annotated neighbors retrieve the GO annotations of the target? We assess this by measuring the functional similarity between the target and the union of GO terms from its neighbors. As shown in Table S1, functional similarity is high across all embeddings. In addition, GO term inheritance metrics - precision, recall, and MCC, are also high (Table S2). These results indicate that proximity in embedding space captures aspects of shared biological function. To better understand the underlying functional signals captured by embeddings, we next consider two scenarios, when proximal domain architectures share at least one domain, and when they do not have any domain in common. These cases allow us to decompose the contribution of domain content and learned contextual patterns in domain organization. Nearest neighbors that share at least one domain First, we examine the scenario where nearest neighbors share at least one domain with the target. This allows us to assess whether domain architecture embeddings provide more information than a simple domain content comparison. As a control, we calculated the same statistics using the k DAs with the highest Jaccard similarity to the target, that is, neighbors with the most similar domain content. For multidomain architectures that share at least one domain, neighboring domain architectures have more similar GO annotations when the neighborhood is determined using a DA embedding than when domain content alone is considered (Fig. S1, Table S3). We compared how well GO terms in shared neighborhoods recapitulate the functional annotation of the target domain, for k = 1, 3, 5 and k = 10. For values of k greater than 1, The Word2Vec ( m = 5, w = 5) embedding consistently obtained the best precision. The best recall was obtained when neighboring DAs were identified based on shared domain content using Jaccard similarity or TF-IDF, which also reflects shared domain content. This is also true when k = 1 for the MF and BP ontologies ( Table 1 ). View this table: View inline View popup Download powerpoint Table 1. Accuracy of GO annotation transfer in the shared k -neighborhood for multidomain architectures as the mean over all A ∈ A O . This led us to hypothesize that shared domain content provides some functional information, but is not sufficiently specific. That is, some DAs that share domains have similar functions, but others do not. To investigate this, we compared the number of distinct GO terms in the shared-domain neighborhood ( k = 3 and k = 5) obtained with Jaccard and with the various embeddings. The number of terms retrieved using Jaccard similarity exceeds the number of terms retrieved using all embeddings in all three ontologies, and is larger than the best embedding by as much as 50% ( Table 2 ). Similarly, mean GO term similarity between domain architecture pairs within the shared-domain neighborhood is consistently smaller with Jaccard, compared with all other methods (Table S4). These results suggest that shared domain content alone is not sufficiently precise. Embeddings are deriving more information than simply the presence or absence of shared domains. View this table: View inline View popup Download powerpoint Table 2. functional consistency within shared k -neighborhoods measured as the mean over all A ∈ A O . Nearest neighbors that share no domains Embedding-based document similarity can identify texts on similar subjects even when they do not contain the same words. This is because they can contain words with similar or related meanings, captured by the embedding. By analogy, we ask whether embeddings are able to identify domain architectures with similar functions, even when they share no domains. To explore this, we first asked how often multidomain architectures that share no domain are nearest neighbors in an embedding. Indeed, as many as 30% of proximal multidomain architectures have no domains in common ( Table 3 ), depending on the embedding and the ontology. TF-IDF retrieves notably fewer such architectures than other embeddings, which is not surprising given that TF-IDF explicitly considers domains that are shared across architectures. View this table: View inline View popup Download powerpoint Table 3. Mean functional similarity averaged over all A ∈ A O ) in non-sharing neighborhoods ( k = 1). We next asked whether proximal domain architectures that share no domain with the target have similar GO terms ( Fig. 1 ). For BP and CC, the majority of embeddings (shown in bold in Table 3 ) achieve significantly higher mean functional similarity than expected by chance. For MF, two embeddings identify non-sharing neighbors with higher than expected similarity. The observed mean similarity exceeds that of the null model for all embeddings except TF-IDF. Larger neighborhoods ( k = 3, 5, 10) tell a similar story: non-sharing neighbors remain highly functionally relevant (Fig. S2, Table S5). To further assess whether embeddings identify non-sharing neighbors with meaningful similarities, we asked how often such pairs have high functional similarity. Indeed, embeddings consistently identify more pairs with functional similarity greater than 0.8 than expected by chance ( Table 4 ). The results are statistically significant for w2v(5,1) in all three ontologies,for w2v(5,5) in BP and CC, and additionally w2v(100,1) and w2v(10,5) for CC. View this table: View inline View popup Download powerpoint Table 4. Number of neighboring multidomain pairs ( k = 1) that lack a common domain with functional similarity > 0.8. The possibility that domain architecture embeddings can identify multidomain pairs that share functional annotations, even when there is no overlapping domain content, is intriguing. To further explore this result, we examined the 187 unique pairs of nearest neighbors ( k = 1) that lack a shared domain, but have functional similarity above 0.8, identified by any embedding. Interestingly, only a small number of these 187 pairs are found by more than one embedding. This low overlap suggests that each embedding captures different contextual signals, highlighting complementary aspects of functional relationships between domain architectures. The complete list of such pairs is provided in Supplementary Information (S1). The threshold of 0.8 was selected to be consistent with the mean functional similarity observed in nearest neighbors ( k = 1) that share at least one domain (Table S3). The best threshold for identifying meaningful similarity is a question for future work. DA pairs that lack shared domains, but nevertheless have high functional similarity, might be caused by errors in domain annotation. To screen for such cases, we compared the amino acid sequences that correspond to the 187 pairs with functional similarity greater than 0.8, in search of conserved regions that might correspond to a shared domain that was overlooked. Out of 14,486 pairwise sequence comparisons (recall that most DAs correspond to more than one sequence) only 16 sequence pairs had local alignments with an E value less than 1. Only two domain architecture pairs corresponded to sequence comparisons with E values below 1 (E = 0.77 and E= 0.074, respectively). At time of writing, the default significance threshold in the blastp interface is (E < 0.05). We discuss two of these pairs in greater detail. The Zinc finger transcription factor Trps1 and Homeobox Hox-B4 proteins ( Figure 2 , top panel) are both transcription factors involved in developmental regulation of the skeletal system [ Kaiser et al., 2007 , Dias et al., 2013 ]. Trps1 is associated with regulation of chondrocyte differentiation (GO:0032330), which is “part of” skeletal system development (GO:0001501). Mutations in the gene are the basis of tricho-rhino-phalangeal syndrome type I (TRPS I) a genetic disorder characterized by skeletal abnormalities [ Momeni et al., 2000 ]. HOXB4 is associated with embryonic skeletal system morphogenesis (GO:0048704) and bone marrow development (GO:0048539), also “part of” skeletal system development. HOXB4 plays crucial roles in vertebrate skeletal system development (reviewed in Morgan et al. [2004]) and regulates the balance between differentiation into osteogenic (bone formation) or hematopoietic lineages of human embryonic stem cells [ Kärner et al., 2009 ]. As transcription factors, TRPS1 and HOXB4 also share GO terms such as regulation of RNA biosynthetic process (GO:2001141). Both domain architectures contain DNA-binding domains (glucocorticoid receptor-like domains, zinc fingers, and homeodomains), but not the same DNA-binding domains. The blastp web interface identifies no significant similarity between these sequences at default settings. Download figure Open in new tab Fig. 2. Schematic representation of domain architectures of two pairs of multidomain proteins that share no domains, but have high functional term similarity (see text). The relative lengths of the proteins and domains are approximate. SWI/SNF-related matrix-associated actin-dependent regulator of chromatin subfamily A member 5 (Smarca5) and Poly [ADP-ribose] polymerase (PARP) ( Figure 2 , bottom panel) are DNA-binding proteins that participate in DNA repair [ Xie et al., 2015 , Toiber et al., 2013 ]. They share CC GO terms (nucleolus (GO:0005730) and site of double-strand break (GO:0035861), and MF GO terms (DNA binding (GO:0003677)). Both domain architectures encode DNA binding domains (HAND domains and homeodomains versus PARP zinc fingers and WGR domains), as well as domains with enzymatic activities (AAA-ATPases and poly(ADP-ribose)polymerases, respectively), but the specific domains in each of these categories differ in the two proteins. This pair has no significant sequence similarity at default parameter settings. In this pilot study, vector embeddings identified neighboring domain architecture pairs that share functional properties, but lack a common domain. There are more such pairs than expected by chance and the mean similarity of nearest neighbors that lack a shared domain is also greater than expected. However, many non-sharing nearest neighbors do not have strong functional similarity. Additional research is needed to determine how to identify the most promising pairs. Discussion Mounting evidence suggests that protein domains carry information about the functions of the proteins that encode them, but explicit models that relate domain content to protein function are lacking. In particular, we require a better understanding of how the functional relevance of a particular domain depends on its surrounding context. Here, we investigated the potential of vector semantic embeddings to capture contextual information. We observe that pairs of multidomain architectures identified using vector semantics share more functional attributes than pairs of multidomain architectures selected based on domain content similarity, alone. Our results suggest that these vector semantic models may capture combinatorial interactions of domains in the same protein. Surprisingly, in some cases, vector semantic embeddings identified multidomain architecture pairs that had high functional similarity, despite having no domains in common at all . Note that no information from sequence analysis or biological textual descriptions were used to construct the embeddings; these inferences are based on domain co-occurrence alone. Taken together, these results underscore the potential of protein domain vector semantics to elucidate the “design rules” of multidomain architectures. While intriguing, these results are preliminary. Further investigations that comprehend larger data sets and greater taxonomic breadth are required. The revolution in protein structure prediction driven by advances in deep learning has produced greatly expanded bioinformatic resources for protein domain analyses [ Paysan-Lafosse et al., 2025 , Lau et al., 2024 ]. Future work that incorporates this new knowledge is also imperative. The results presented here suggest that protein domain vector semantics may be able to discover the protein domain equivalent of words and phrases with similar or related meanings in natural languages. The science of lexical semantics deals with many types of related meanings [reviewed in Jurafsky and Martin, 2008 ]. Words may have similar meanings (e.g., sarcastic, ironic ) be used in the same context (e.g., coffee, cup ), be made of the same materials or have the similar constituent parts (e.g., car, bicycle ), or be associated with the same activity (e.g., suture, scalpel ). Characterizing different types of word associations and developing NLP models that are capable of distinguishing between them are active areas of NLP research [ Hill et al., 2015 ]. Our discovery of functionally similar multidomain architectures that harbor no common domains suggests that notions of semantic or topical relatedness may be relevant to protein domain function. What types of “topical relatedness” are meaningful in the protein domain context is an exciting open question. In one of the first works to tackle this question, Buchan and Jones [2020] experimented with vector arithmetic models of word analogies (e.g., king is to man as queen is to woman ) to explore whether such relations also exist for domains. Vector semantic models of multidomain proteins also hold promise for advancing bioinformatic applications. The problem of developing a Domain Ontology is closely related to questions in protein domain lexical semantics. The Gene Ontology (GO) is a powerful and general system for describing protein function. The hierarchy of well-defined functional terms using a controlled vocabulary supports a broad range of applications to query, manipulate, compare and reason with biological data. Currently, no equivalent domain function ontology exists; nor is there agreement on what a domain-specific ontology should look like. Should domain ontology terms describe inherent functional properties of the domains themselves? Or is the goal to annotate domains with terms that predict the functions of any protein that is found to encode the domain? Several lines of research are advancing functional annotation for domains. The InterPro2GO project is building a manually annotated domain resource, where annotations are based on experimental evidence of a domain’s function, and not simply the function of its associated proteins [Burge et al., 2012, Blum et al., 2021]. In an orthogonal approach, strategies are being developed to support automated mapping of Gene Ontology (GO) terms from proteins to domains [ Weiner et al., 2008 , Buchan and Jones, 2020 , Fang and Gough, 2013 , López and Pazos, 2013 , Ulusoy and Doğan, 2024 , Ibtehaz et al., 2023 ]. However, these annotations are couched in terms of the GO, which was developed to describe the functions of genes and proteins, not domains. Further examination of the appropriateness of the MF, BP, and CC trichotomy for describing domain function is warranted. It may be possible to partition the molecular functions of a protein by attributing specific molecular roles to domains, but it is less clear whether protein domains tend to be associated with specific biological processes or cellular compartments. This is an interesting research question that has practical consequences for the development of bioinformatic resources. Vector semantics of domains and domain architectures also holds promise for protein function prediction [Radivojac et al., 2013, Jiang et al., 2016, Zhou et al., 2019]. A better understanding of multidomain design rules will shed light on the predictive value of domain content and organization for multidomain protein functions. This is important: In the first Critical Assessment of protein Function Annotation (CAFA) challenge, all of the top predictors performed less well on multidomain proteins than on single domain proteins [Radivojac et al., 2013]. Protein function prediction methods are increasingly combining many different types of data [Radivojac et al., 2013, Jiang et al., 2016, Zhou et al., 2019, Boadu et al., 2025 ], including domain family information, but few methods have attempted to incorporate an explicit model relating domain content to the functional attributes of multidomain proteins. The success of several notable exceptions [ You et al., 2018 , Ibtehaz et al., 2023 , Talo and Bozdag, 2025 ] suggests the value of vector semantics in this context. Funding This work was supported in part by National Science Foundation Grants [DBI-1838344 and DBI-1759943]. Data availability The data and code used in this work are available on Zenodo (draft). Funder Information Declared National Science Foundation, https://ror.org/021nxhr62 , DBI-1838344 , DBI-1759943 Footnotes https://zenodo.org/records/15769961?preview=1&token=eyJhbGciOiJIUzUxMiJ9.eyJpZCI6IjYyOTg2MWI2LThmZmItNDMzZi05MDk0LTM2MmZlZjc4ZGFhMSIsImRhdGEiOnt9LCJyYW5kb20iOiI1ZDRmMmNjMzlkYjMwYjU5NjM5ZTBmMmUyNzYxNjc5MSJ9.8vqKqYa_J0aHWI8QUaT2k0BxfUN6jxeM2BzC-QTldG37zD96lm_xZi-fu5wKMK36V66_FrCH5k2G8OxxW-0LDg References ↵ M. Bashton and C. Chothia . The geometry of domain combination in proteins . J Mol Biol , 315 ( 4 ): 927 – 939 , Jan 2002 . OpenUrl CrossRef PubMed Web of Science M. Basu , L. Carmel , I. Rogozin , and E. Koonin . Evolution of protein domain promiscuity in eukaryotes . Genome Res , 18 ( 3 ): 449 – 461 , Jan 2008 . OpenUrl Abstract / FREE Full Text ↵ M. K. Basu , E. Poliakov , and I. B. Rogozin . Domain mobility in proteins: functional and evolutionary implications . Brief Bioinform , 10 ( 3 ): 205 – 216 , Jan 2009 . OpenUrl CrossRef PubMed M. Blum et al. The InterPro protein families and domains database: 20 years on . Nucleic Acids Res , 49 : D344 – D354 , Jan 2021 . OpenUrl CrossRef PubMed ↵ F. Boadu , A. Lee , and J. Cheng . Deep learning methods for protein function prediction . Proteomics , 25 : e2300471 , Jan 2025 . OpenUrl CrossRef PubMed ↵ D. W. Buchan and D. T. Jones . Learning a functional grammar of protein domains using natural language word embedding techniques . Proteins , 88 ( 4 ): 616 – 624 , 2020 . OpenUrl CrossRef PubMed S. Burge et al. Manual GO annotation of predictive protein signatures: the InterPro approach to GO curation . Database , 2012 : bar068 , 02 2012 . ISSN 1758-0463 . doi: 10.1093/database/bar068 . OpenUrl CrossRef PubMed ↵ Z. Chen and Q. Luo . DualNetGO: a dual network model for protein function prediction via effective feature selection . Bioinformatics , 40 ( 7 ): btae437 , 2024 . OpenUrl PubMed ↵ K. Church and P. Hanks . Word association norms, mutual information, and lexicography . Comput Linguist , 16 ( 1 ): 22 – 29 , 1990 . OpenUrl I. Cohen-Gihon et al. Evolution of domain promiscuity in eukaryotic genomes-a perspective from the inferred ancestral domain architectures . Mol Biosyst , 7 ( 3 ): 784 – 792 , Mar 2011 . OpenUrl CrossRef PubMed G. Cromar et al. New tricks for “old” domains: how novel architectures and promiscuous hubs contributed to the organization and evolution of the ECM . Genome Biol Evol , 6 : 2897 – 2917 , Oct 2014 . OpenUrl CrossRef PubMed G. L. Cromar et al. PhyloPro2.0: a database for the dynamic exploration of phylogenetically conserved proteins and their domain architectures across the Eukarya . Database (Oxford) , 2016 :pii:baw013, Mar 2016 . ↵ X. Cui , Y. Xue , C. McCormack , A. Garces , T. W. Rachman , Y. Yi , M. Stolzer , and D. Durand . Simulating domain architecture evolution . Bioinformatics , 38 ( Supplement 1 ): i134 – i142 , 2022 . OpenUrl CrossRef PubMed J. Devlin , M.-W. Chang , K. Lee , and K. Toutanova . Bert: Pre-training of deep bidirectional transformers for language understanding . arXiv preprint arxiv: 1810.04805 , 2018 . ↵ C. Dias , L. Isidoro , M. Santos , H. Santos , and J. S. Marques . Trichorhinophalangeal syndrome type I: a patient with two novel and different mutations in the TRPS1 gene . Case Rep Genet , 2013 ( 1 ): 748057 , 2013 . OpenUrl PubMed ↵ E. Dohmen , S. Klasberg , E. Bornberg-Bauer , S. Perrey , and C. Kemena . The modular nature of protein evolution: domain rearrangement rates across eukaryotic life . BMC Evol Biol , 20 ( 1 ): 30 , Feb 2020 . OpenUrl CrossRef PubMed ↵ H. Fang and J. Gough . DcGO: database of domain-centric ontologies on functions, phenotypes, diseases and more . Nucleic Acids Res , 41 : D536 – D544 , Jan 2013 . OpenUrl CrossRef PubMed Web of Science ↵ G. W. Furnas , T. K. Landauer , L. M. Gomez , and S. T. Dumais . The vocabulary problem in human-system communication . Commun ACM , 30 ( 11 ): 964 – 971 , 1987 . OpenUrl ↵ L. Hallee , N. Rafailidis , C. Horger , D. Hong , and J. P. Gleghorn . Annotation Vocabulary (Might Be) All You Need . bioRxiv , 2024 . ↵ F. Hill , R. Reichart , and A. Korhonen . Simlex-999: Evaluating semantic models with (genuine) similarity estimation . Comput Linguist , 41 ( 4 ): 665 – 695 , 2015 . OpenUrl ↵ S. Hu and B. Zhao . Protein function prediction using GO similarity-based heterogeneous network propagation . Sci Rep , 15 ( 1 ): 19131 , 2025 . OpenUrl PubMed ↵ N. Ibtehaz , Y. Kagaya , and D. Kihara . Domain-PFP allows protein function prediction using function-aware domain embedding representations . Commun Biol , 6 ( 1 ): 1103 , 2023 . OpenUrl PubMed Y. Jiang et al. An expanded evaluation of protein function prediction methods shows an improvement in accuracy . Genome Biol , 17 : 184 , Sep 2016 . OpenUrl CrossRef PubMed ↵ D. Jurafsky and J. H. Martin . Speech and Language Processing: Processing, Recognition. An Introduction to Natural Language Computational Linguistics, and Speech PEARSON INDIA, 2nd edition edition , January 2008 . ↵ F. J. Kaiser , H.-J. Lüdecke , and S. Weger . SUMOylation modulates transcriptional repression by TRPS1 . Biol Chem , 388 ( 4 ): 381 – 390 , 2007 . doi : doi: 10.1515/BC.2007.051 . URL https://doi.org/10.1515/BC.2007.051. OpenUrl CrossRef PubMed Web of Science P. Karev , Georgy et al. Birth and death of protein domains: a simple model of evolution explains power law behavior . BMC Evol Biol , 2 ( 1 ): 18 , 2002 . OpenUrl CrossRef PubMed P. Karev , Georgy et al. Gene family evolution: an in-depth theoretical and simulation analysis of non-linear birth-death-innovation models . BMC Evol Biol , 4 : 32 , Sep 2004 . OpenUrl CrossRef PubMed ↵ E. Kärner , C. Unger , R. Cerny , L. Ä hrlund-Richter, B. Ganss , M. S. Dilber , and M. Wendel . Differentiation of human embryonic stem cells into osteogenic or hematopoietic lineages: A dose-dependent effect of osterix over-expression . J Cell Physiol , 218 ( 2 ): 323 – 333 , 2009 . OpenUrl PubMed ↵ S. Kummerfeld and S. Teichmann . Protein domain organisation: adding order . BMC Bioinf , 10 : 39 , Jan 2009 . OpenUrl ↵ A. M. Lau , N. Bordin , S. M. Kandathil , I. Sillitoe , V. P. Waman , J. Wells , C. A. Orengo , and D. T. Jones . Exploring structural diversity across the protein universe with the Encyclopedia of Domains . Science , 386 ( 6721 ): eadq4946 , 2024 . doi: 10.1126/science.adq4946 . OpenUrl CrossRef ↵ D. López and F. Pazos . COPRED: prediction of fold, go molecular function and functional residues at the domain level . Bioinformatics , 29 : 1811 – 1812 , Jul 2013 . OpenUrl CrossRef PubMed ↵ H. P. Luhn . A statistical approach to mechanized encoding and searching of literary information . IBM J Res Dev , 1 ( 4 ): 309 – 317 , 1957 . OpenUrl ↵ E. M. Marcotte , M. Pellegrini , H.-L. Ng , D. W. Rice , T. O. Yeates , and D. Eisenberg . Detecting protein function and protein-protein interactions from genome sequences . Science , 285 ( 5428 ): 751 – 753 , 1999 . OpenUrl Abstract / FREE Full Text F. J. Martin et al. Ensembl 2023 . Nucleic Acids Res , 51 : D933 – D941 , Jan 2023 . OpenUrl CrossRef PubMed ↵ D. P. Melidis and W. Nejdl . Capturing protein domain structure and function using self-supervision on domain architectures . Algorithms , 14 ( 1 ): 28 , 2021 . OpenUrl ↵ T. Mikolov , K. Chen , G. Corrado , and J. Dean . Efficient estimation of word representations in vector space . arXiv preprint arxiv: 1301.3781 , 2013 . ↵ P. Momeni , G. Glöckner , O. Schmidt , D. Von Holtum , B. Albrecht , G. Gillessen-Kaesbach , R. Hennekam , P. Meinecke , B. Zabel , A. Rosenthal , et al. Mutations in a new gene, encoding a zinc-finger protein, cause trichorhino-phalangeal syndrome type I . Nat Genet , 24 ( 1 ): 71 – 74 , 2000 . OpenUrl CrossRef PubMed Web of Science R. Morgan , R. Pettengell , and J. Sohal . The double life of HOXB4 . FEBS Lett , 578 : 1 – 4 , Dec 2004 . OpenUrl CrossRef PubMed A. P. Pandurangan et al. The SUPERFAMILY 2.0 database: a significant proteome update and a new webserver . Nucleic Acids Res , 47 ( D1 ): D490 – D494 , 2019 . OpenUrl CrossRef PubMed ↵ T. Paysan-Lafosse , A. Andreeva , M. Blum , S. R. Chuguransky , T. Grego , B. L. Pinto , G. A. Salazar , M. L. Bileschi , F. Llinares-López , L. Meng-Papaxanthos , et al. The Pfam protein families database: embracing AI/ML . Nucleic Acids Res ., 53 ( D1 ): D523 – D534 , 2025 . doi: 10.1093/nar/gkae997 . OpenUrl CrossRef ↵ T. Przytycka , G. Davis , N. Song , and D. Durand . Graph theoretical insights into evolution of multidomain proteins . J Comput Biol , 13 ( 2 ): 351 – 363 , 2006 . OpenUrl CrossRef PubMed P. Radivojac et al. A large-scale evaluation of computational protein function prediction . Nat Methods , 10 : 221 – 227 , Mar 2013 . OpenUrl CrossRef PubMed Web of Science ↵ R. Řehůřek and P. Sojka . Software Framework for Topic Modelling with Large Corpora . In Proceedings of the LREC 2010 Workshop on New Challenges for NLP Frameworks, pages 45–50, Valletta, Malta , May 2010 . ELRA. ↵ K. Sparck Jones . A statistical interpretation of term specificity and its application in retrieval . J Doc , 28 ( 1 ): 11 – 21 , 1972 . OpenUrl CrossRef Web of Science ↵ M. Talo and S. Bozdag . ProtFun: A protein function prediction model using graph attention networks with a protein large language model . bioRxiv , 2025 . doi: 10.1101/2025.05.13.653854 . URL https://www.biorxiv.org/content/early/2025/05/17/2025.05.13.653854 . OpenUrl Abstract / FREE Full Text ↵ A. J. Te Velthuis , J. F. Admiraal , and C. P. Bagowski . Molecular evolution of the MAGUK family in metazoan genomes . BMC Evol Biol , 7 : 1 – 10 , 2007 . OpenUrl CrossRef PubMed ↵ D. Toiber , F. Erdel , K. Bouazoune , D. M. Silberman , L. Zhong , P. Mulligan , C. Sebastian , C. Cosentino , B. Martinez-Pastor , S. Giacosa , et al. SIRT6 recruits SNF2H to DNA break sites, preventing genomic instability through chromatin remodeling . Mol Cell , 51 ( 4 ): 454 – 468 , 2013 . OpenUrl CrossRef PubMed Web of Science ↵ H. Tordai , A. Nagy , K. Farkas , L. Banyai , and L. Patthy . Modules, multidomain proteins and organismic complexity . FEBS J , 272 ( 19 ): 5064 – 5078 , Oct 2005 . OpenUrl CrossRef PubMed ↵ E. Ulusoy and T. Doğan . Mutual annotation-based prediction of protein domain functions with Domain2GO . Protein Sci ., 33 ( 6 ): e4988 , 2024 . doi: 10.1002/pro.4988 . URL https://onlinelibrary.wiley.com/doi/abs/10.1002/pro.4988 . OpenUrl CrossRef PubMed ↵ C. Vogel , S. A. Teichmann , and J. Pereira-Leal . The relationship between domain duplication and recombination . J Mol Biol , 346 ( 1 ): 355 – 365 , Feb 2005 . OpenUrl CrossRef PubMed Web of Science ↵ W. Wang , Y. Shuai , M. Zeng , W. Fan , and M. Li . DPFunc: accurately predicting protein function via deep learning with domain-guided structure information . Nat Commun , 16 ( 1 ): 70 , Jan 2025 . OpenUrl CrossRef PubMed ↵ J. Weiner , A. D. Moore , and E. Bornberg-Bauer . Just how versatile are domains? BMC Evol Biol , 8 : 285 , Oct 2008 . OpenUrl CrossRef PubMed J. Weiner 3rd, F. Beaussart , and E. Bornberg-Bauer . Domain deletions and substitutions in the modular protein evolution . FEBS J , 273 ( 9 ): 2037 – 2047 , May 2006 . OpenUrl CrossRef PubMed ↵ S. Xie , O. Mortusewicz , H. T. Ma , P. Herr , R. Y. Poon , T. Helleday , and C. Qian . Timeless interacts with PARP-1 to promote homologous recombination repair . Mol Cell , 60 ( 1 ): 163 – 176 , 2015 . OpenUrl CrossRef PubMed ↵ Y. Ye and A. Godzik . Comparative analysis of protein domain organization . Genome Res , 14 ( 3 ): 343 – 353 , Mar 2004 . OpenUrl Abstract / FREE Full Text ↵ R. You , Z. Zhang , Y. Xiong , F. Sun , H. Mamitsuka , and S. Zhu . GOLabeler: improving sequence-based large-scale protein function prediction by learning to rank . Bioinformatics , 34 : 2465 – 2473 , Jul 2018 . OpenUrl CrossRef PubMed ↵ L. Yu , D. K. Tanwar , E. D. S. Penha , Y. I. Wolf , E. V. Koonin , and M. K. Basu . Grammar of protein domain architectures . Proc Natl Acad Sci U S A , 116 : 3636 – 3645 , Feb 2019 . OpenUrl Abstract / FREE Full Text N. Zhou et al. The CAFA challenge reports improved protein function prediction and new functional annotations for hundreds of genes through experimental screens . Genome Biol , 20 : 244 , Nov 2019 . OpenUrl CrossRef PubMed View the discussion thread. Back to top Previous Next Posted July 10, 2025. Download PDF Supplementary Material Data/Code Email Thank you for your interest in spreading the word about bioRxiv. NOTE: Your email address is requested solely to identify you as the sender of this article. Your Email * Your Name * Send To * Enter multiple addresses on separate lines or separate them with commas. You are going to email the following Vector Semantics of Multidomain Protein Architectures Message Subject (Your Name) has forwarded a page to you from bioRxiv Message Body (Your Name) thought you would like to see this page from the bioRxiv website. Your Personal Message CAPTCHA This question is for testing whether or not you are a human visitor and to prevent automated spam submissions. Share Vector Semantics of Multidomain Protein Architectures Xiaoyue Cui , Yuting Xiao , Maureen Stolzer , Dannie Durand bioRxiv 2025.07.07.663606; doi: https://doi.org/10.1101/2025.07.07.663606 Share This Article: Copy Citation Tools Vector Semantics of Multidomain Protein Architectures Xiaoyue Cui , Yuting Xiao , Maureen Stolzer , Dannie Durand bioRxiv 2025.07.07.663606; doi: https://doi.org/10.1101/2025.07.07.663606 Citation Manager Formats BibTeX Bookends EasyBib EndNote (tagged) EndNote 8 (xml) Medlars Mendeley Papers RefWorks Tagged Ref Manager RIS Zotero Tweet Widget Facebook Like Google Plus One Subject Area Evolutionary Biology Subject Areas All Articles Animal Behavior and Cognition (7642) Biochemistry (17715) Bioengineering (13907) Bioinformatics (42005) Biophysics (21472) Cancer Biology (18624) Cell Biology (25534) Clinical Trials (138) Developmental Biology (13391) Ecology (19935) Epidemiology (2067) Evolutionary Biology (24356) Genetics (15617) Genomics (22529) Immunology (17753) Microbiology (40437) Molecular Biology (17200) Neuroscience (88697) Paleontology (667) Pathology (2840) Pharmacology and Toxicology (4829) Physiology (7653) Plant Biology (15171) Scientific Communication and Education (2046) Synthetic Biology (4304) Systems Biology (9827) Zoology (2272)

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-22T02:00:06.705733+00:00
License: CC-BY-NC-ND-4.0