{"paper_id":"014f8eef-a962-427e-8ffd-088b2465ba7a","body_text":"Advances in biomolecular research, coupled with rapidly increasing availability of\ninformation from multiple genome sequencing initiatives, global gene expression\npatterns, large scale molecular interaction experiments and genome wide association\nstudies, have led to an exponential increase in biological data. The explosion of\ndata, accompanied by a plethora of theoretical tools for predicting gene function,\nhas created an information overload. The immense challenges in separating the\nbiological wheat from the chaff have necessitated the development of a variety of\nanalytical tools and databases to store and manage biological data and retrieve\nmeaningful information to facilitate further experimental characterisation.\nThe biological role of a gene or a protein is not only defined by its sequence and\nstructure but also by when and where it is expressed and its interactions with other\nbiomolecules (such as proteins, nucleic acids and metabolites). In the post-genomic\nera, attempts at function annotation increasingly employ data from different types\nof repositories. Biological data from a single type of data source, though useful,\nis often limited in extent to which it may help uncover functional associations;\neither because of a systematic bias towards specific genes, gene families and\npathways and/or inclusion of erroneous entries during data acquisition. With focus\nshifting from genes and proteins to biological systems, integrating information from\nmultiple data types is a more robust and accurate means of enhancing existing\ninterpretations and unravelling new functional associations as demonstrated in\nseveral studies  [1] ,\n [2] .\nHowever, biological data integration is a formidable task. Different computational\ntools and data sources may often employ different approaches and formats for input,\nstoring and retrieving relevant information that may often result in appreciable\ndifferences in data quality. This heterogeneity often restricts compatibility\nbetween different resources and limits the extent and efficiency of combined\nanalysis. Furthermore, investigation of diverse data types necessitates a flexible,\nuniform and simplified interface to query, retrieve and analyse data across diverse\nsources. Despite these hurdles, the immense potential benefits of a combined\ninvestigative approach have spawned several initiatives towards integrated data\nrepositories  [3] ,\n [4] ,  [5] ,  [6] ,  [7] . Among these, of\nparticular interest are data warehouses, which compile all the relevant information\nto a common platform  [6] ,  [8] ,  [9] ,  [10] ,  [11] ,  [12] ,  [13] ,  [14] . A data warehouse is particularly desirable, since it\npermits a wide range of queries based on diverse attributes (including genes,\nproteins, families, pathways, ontologies, diseases and expression profiles) and\npossesses the ability to produce unified output and the flexibility in selecting the\ntype and the order of the data sources. InterMine is a multi-purpose data warehouse\nframework ( http://www.intermine.org/ ), originally developed for FlyMine, an\nintegrated database for  Drosophila  and  Anopheles \ngenomics  [13] . It\nfeatures a sequence ontology-based data model and a user-friendly web interface\npermitting the end users to either design flexible and complex database queries, or\nchoose from a library of ‘templates’ consisting of predefined queries\nwith a simple form and description  [13] . In addition, InterMine provides default parsers for\nintegrating data from several resources with the framework for incorporating\ncustomised parsers and data sources. The flexibility in designing queries and\nintegrating diverse data types provides a powerful tool for the researchers. In\naddition to FlyMine, InterMine also powers modEncode ( http://intermine.modencode.org/ ), RatMine ( http://ratmine.mcw.edu/ratmine/begin.do ), YeastMine ( http://yeastmine.yeastgenome.org:8080/yeastmine/begin.do ) and\nMetabolicMine ( http://www.metabolicmine.org/ ).\nIdentification of suitable targets (such as genes, proteins, non-peptide gene\nproducts and pathways) for characterisation is one of the most critical steps in\nbiology, particularly in annotating gene function, drug discovery and understanding\nmolecular bases of diseases. An integrated approach that combines results from\nmultiple data types is best suited for optimal target discovery  [15] ,  [16] . The distinct\nmerits of the InterMine framework have inspired us to develop TargetMine, an\nintegrated resource for retrieval of target genes and proteins for experimental\ncharacterisation and drug discovery. In this paper, we describe the data sources\navailable in the present release of TargetMine and their access and query\ncapability. We also outline an objective protocol for target prioritisation with\nTargetMine that relies on the integration of diverse data types. Gene prioritisation\nrefers to the selection of most interesting or promising genes from a larger set of\ngenes for further analysis  [17] ,  [18] . Experimental evaluation of large gene lists to identify\nsuitable candidates is a formidable and often impossible task and therefore,\ncomputational tools for candidate gene prioritisation have emerged over the years.\nThese tools variously rely on functional associations, protein-protein interactions,\ngene expression data, sequence and structure properties or combinations thereof to\nselect candidate genes  [16] ,  [17] ,  [18] ,  [19] ,  [20] ,  [21] ,  [22] ,  [23] ,  [24] ,  [25] . TargetMine was designed specifically for target\nprioritisation within the framework of a data warehouse and our prioritisation\nprotocol, though less sophisticated than some standalone tools, is easier to use and\nprovides flexibility in the choice of data sources that may be employed for analysis\nof query gene sets. Finally, we discuss the possibilities of future implementations\nin the TargetMine data warehouse to provide maximum coverage of the biological\ntarget space.\n\nA detailed description of the InterMine system is available elsewhere  [13] . Here we\nrestrict ourselves to a brief overview of the InterMine data organisation.\nInterMine is an open source data warehouse framework. Each entry in the system\n(such as a gene or a protein) is considered an ‘object’. The\nInterMine object-based data model, consists of ‘classes’ and\nreflects the relationships between different data types. Each class contains\nobjects that share similar properties and a set of ‘attributes’ that\ncorrespond to various types of information (such as gene symbol and gene/protein\nidentifier) associated with each object of that class. The classes are linked\nwith each other by references that specify the associations between objects in\ndifferent classes. The InterMine data structure readily allows the navigation of\nthe stored biological data via the relationships between different data types,\nfacilitated by an inbuilt tool termed ‘query builder’. The query\nbuilder tool permits the users to select and constrain the data types for the\ndesired output. The list function enables the query process to be performed with\na user-supplied list of objects and export the lists as either comma separated\n(csv) or tab separated values (tsv). It also permits the user to convert\ngenes/proteins from one species to another based on KEGG orthology associations.\nThe InterMine Web Service allows the users to query TargetMine from their own\nweb pages and applications.\nIn addition to the existing InterMine classes, we have customised the InterMine\ndata model and created new classes to collate biological data types most likely\nto help facilitate target discovery ( Table S1 ). We will discuss some of these\nimplementations below. As of now, the biological data in TargetMine for most\npart is limited to human, rat, mouse and fruit fly, the best studied model\norganisms in biology. The data sources compiled in TargetMine are summarised in\n Table 1 .\n*H: human, R: rat, M: mouse, F: fruit fly\n( Drosophila ), E:  E. coli .\nOMIM data are presently not distributed with the TargetMine\ndemonstration version.\nStructural data for biological macromolecules, especially proteins, have been\nextremely important in explaining their molecular and biochemical functions,\nevolutionary relationships and understanding their explicit biological roles\n [26] .\nIt is well recognised that complementing protein sequence information with\nstructural data is a robust approach towards more accurate protein function\nannotation  [27]  and hence, more reliable target discovery. However,\nintegrating protein sequence and structural information from different sources\nremains a non-trivial task. In recognition of the obvious benefits of an\nintegrated protein sequence-structure repository, we customised and embellished\nthe default InterMine data model to combine protein sequence information from\nthe UniProt database  [28]  with protein structure information from the Protein\nData Bank (PDB)  [29]  and structural classification based on evolutionary\nrelationships in the Structural Classification of Proteins (SCOP) database  [30] . With our\ncustomised data model, the user can easily query for PDB structures\ncross-referenced (if available) with the protein of interest in the UniProt\nrepository and other databases such as DrugBank  [31]  (e.g., “Show all the\nprotein structures that contain the targets, as defined in DrugBank, of a given\nset of drugs” or “Given a list of proteins, show all the approved\ndrugs solved in complex with any structure of these proteins if present”).\nThe user can also retrieve disease associations, pathway associations and\npotential protein-drug associations, based on ligands associated with the\nprotein structures, for the protein of interest (e.g., “Show all the PDB\nentries that contain a given drug”).\nDifferent data sources use different numbering systems for specifying protein\nregions. To associate protein sequences (in the Protein class) with protein\nstructures (in the ProteinStructure class), we introduced two new classes\n(ProteinStructureRegion and PDBRegion;  Figure 1 ). We also introduced the\nProteinDomainRegion class to link the Protein class to the Protein domain class\nthat stores InterPro  [32]  domain annotations. The PDB-UniProt mapping was taken\nfrom SIFTS  [33]  and InterPro domain assignments from IPI  [34] . The\nintegration facilitated querying detailed domain and structural assignments; for\nexample, the user can query regions of a protein, for which structural\ninformation is available, and then retrieve domain annotations falling within\nthese regions.\nThe data model is depicted as a class diagram in the Unified Modeling\nLanguage ( http://www.uml.org ). Some details of the model are\nignored to reduce the complexity of the diagram.\nTranscription factors (TFs) are proteins that bind to specific DNA sequences,\nthereby regulating the expression (transcription) of their target genes  [35] . TFs are\nof immense significance in biomedical investigations and some TFs such as\nnuclear receptors are important drug targets  [36] ,  [37] . In view of the\nsignificance of these protein-DNA interactions to cellular physiology, we\nmodified the existing InterMine Interaction class, which describes gene-gene\ninteractions, to define a new class named ProteinDNAInteraction. The\nProteinDNAInteraction class contains specific attributes that reflect the unique\naspects of protein-DNA interactions, such as protein (TF) binding sites in the\nregulatory regions of the target genes. These data were retrieved from AMADEUS\n [38]  and\nOregAnno  [39]  resources and from assorted literature sources. Since\ndifferent resources adopt different approaches to compiling protein-DNA\ninteraction information, the combined source data were manually processed to\nuniformly assign Entrez gene identifiers to each participating gene and remove\nredundancies prior to the incorporation into TargetMine. The integration enabled\nus to make a complicated query such as: “Given a list of genes, retrieve\nall the TF-target relations observed within the list”.\nFor disease and phenotype association, we created new classes and data parsers to\nretrieve the data from OMIM database  [40]  and human genome\ndisease annotations  [41] . Enzymes play key roles in many biological processes\nand are attractive candidates for experimental investigation aimed at\nunderstanding cellular processes, diseases and identifying suitable drug\ntargets. We designed a new Enzyme class (linked to the Protein class) to gather\nall information on enzymes as curated in the Enzyme database  [42] . The\nEnzyme class was also directly linked to the Pathway class by parsing the KEGG\n [43]  mapping files, thereby providing links to their\npotential roles in cellular processes. Most genes and proteins function in\nassociation with other proteins and thus, the study of protein-protein\ninteractions (PPIs) is critical to understanding their roles in living systems.\nIn addition to the default InterMine Interaction class that was employed for\nstoring biomolecular interactions from the BioGRID database  [44] , we designed\na new ProteinInteraction class to collate all interactions curated in PPIview,\nan integrated repository of human PPIs  [45] . This integration\nfacilitated the querying of interacting partners of a gene/protein or a list of\ngenes/proteins of interest and infer overall interaction networks involving\nthese genes/proteins.\nIn addition, to expand the information space for sparsely annotated genes and\nproteins, we provided a framework for including  in silico \nannotations derived from selected protein prediction tools (FUGUE  [46] , Protein-DNA\nbinding propensity  [47]  and Protein-protein interaction sites  [48] ) and for\nincluding experimental data from in-house research.\nOur general protocol for target prioritisation using TargetMine is shown in  Figure 2 . First, we upload a\nlist of initial candidate genes or proteins (e.g., a set of differentially\nexpressed genes or a set of proteins that interact with a given protein) to\nTargetMine to create a TargetMine gene list. Enrichment of specific biological\nthemes (including but not limited to, KEGG pathways, Gene Ontology (GO) terms\n [49] \nand OMIM phenotypes) associated with the initial list is estimated by\nhypergeometric distribution and the inferred  p -values are\nfurther adjusted for multiple test corrections to control the false discovery\nrate using the Benajmini and Hochberg procedure  [50] . The significantly enriched\nbiological associations (that satisfied, in this instance, a condition of\n p ≤0.05 after a multiple test correction with the\nBenajmini and Hochberg procedure) can be visualised in the individual enrichment\nwidgets. We gather the genes mapped to the top  N  significant\nassociations (where  N  = 1,2,3…, an\nadjustable value reflecting incrementally relaxed thresholds) retrieved from\nKEGG (A), GO Biological Process (B) and OMIM (C) databases into separate lists\nand merge them (for example, by taking the union\nA B C of the retrieved\ngenes) to infer corresponding sets of prioritised genes, albeit no ranking is\nprovided at the moment. (We assume that an initial candidate list is from a\nsingle species and the enrichment calculation is performed using the data for\nthis species only.)\nTo evaluate the effectiveness of TargetMine in identifying suitable targets for\nfurther characterisation, we performed target gene prioritisation tests (as\ndescribed above) on 19 sets of known disease-associated genes compiled from the\nliterature  [51] \n( Table 2  and  Figures 3  and  4 ; see  Materials and Methods  for details). In all instances, our\nprioritisation approach was supported by high sensitivity and precision values,\nand enforcing a threshold of collecting only the genes mapped to top seven\nassociations (that satisfied a  p -value cutoff of\n p ≤0.05 after a multiple test correction with the\nBenajmini and Hochberg procedure) was by and large most suited to ensuring\nmaximum coverage and minimum over-prediction ( Table S2 ).\nThough for cirrhosis and cervical carcinoma, the number of false positives was\nslightly larger than those for the other diseases, the sensitivity and precision\nremained high.\nTP- True positive, FP- False positive (see text for details).\n(The full disease names and their abbreviations are listed in  Table 2 .) Each line\nrepresents the F-score for a particular disease data set as a function\nof the threshold (the top  N  significant associations\nconsidered). The error bars show the standard deviation across ten\nbenchmarking evaluations for each disease.\nWe have repeated the tests by changing the proportion of known curated genes in\nan input gene list (from one third to one tenth). Although both sensitivity and\nprecision decreased slightly, reasonable performance was maintained with a\ncutoff of six ( Table S3 ), suggesting that the method still works for situations\nwhere only one tenth of input genes are disease-associated. We have also\nevaluated the results from a method using only a single data source. By taking\nthe union of the collected genes from KEGG, GO and OMIM, the performance in most\ncases increased by about 0.1 points (measured by the F-score; see  Materials and Methods ), demonstrating the\nusefulness of the integration.\nThese results showed that the integration of diverse biological properties in\nTargetMine was a successful approach towards the identification of candidate\ngenes for further investigation. Besides, the operation in TargetMine is\nsemi-automatically accomplished by a few mouse clicks instead of preparing\nspecific data files and running external software. The TargetMine data model\npermits retrieval of stored data and its analysis in a single interface and thus\naids in efficient prioritisation. The ease of accomplishing such analysis via a\nsimple web interface further underscores the utility of TargetMine as an\neffective tool in investigation of genes and genomes. In our benchmark tests, we\nchose KEGG, GO Biological Process and OMIM as the best sources for highlighting\nthe functional associations of groups of genes but TargetMine also provides\nenrichment widgets for GO Molecular Function and Cellular Component, Drug and\nDisease Ontology (DO) associations, which may be used to assist in selecting\ncandidate genes. The user may also employ TF-target associations to identify\ncommon regulatory themes that may be associated with a set of co-expressed\nfunctionally similar genes.\nAs a data warehouse, TargetMine is not an alternative to large public databases\n(such as UniProt  [28] ) but rather, it is designed for use in individual\nlaboratories in academia and industry. In comparison to existing integrated\ndatabases, TargetMine provides an alternative usage that aims to rapidly and\nefficiently retrieve varied biological information for large gene sets in a\nsimplified manner. Most integrated databases are able to retrieve different\nbiological properties, but are largely designed for simple queries for a single\ngene. Though some may provide facilities for batch query, the users in many\ninstances need to employ external scripts for querying and post-processing the\nrelevant data. In contrast, TargetMine provides a simple interface for batch\nquery with numerous templates and the facility to construct complicated queries.\nThe output options permit user-defined displays on the type and the order of\ndifferent annotations. Besides, the enrichment widgets, as described above,\nprovide a quick preliminary analysis of the genes in the list and thus, greatly\nhelp in understanding the enriched themes associated with query sets and also\nhelp complement the analysis performed by specialised gene prioritisation tools.\nTherefore, TargetMine facilitates biological data gathering and data analysis in\na single user-friendly interface.\nAlthough some commercial resources such as Ingenuity® (Redwood City,\nCalifornia) and MetaCore™ (GeneGo, St. Joseph, MI) provide more\ninteraction and/or pathway data plus tools for statistical data analysis, they\nlargely emphasise on collating gene annotations and mostly lack protein level\nannotations such as domains and structures. Additionally, several data types\navailable in TargetMine such as Protein-DNA interactions, to the best of our\nknowledge, are not made available by other publicly available resources, some of\nwhich, including GeneDistiller  [52]  and PolySearch  [53] , can perform tasks similar\nto TargetMine's. However, the key difference is TargetMine's\nflexibility and its built-in prioritisation protocol; the data size and data\ntypes are readily customisable in TargetMine, providing a more flexible and\ncomprehensive framework for target discovery.\nTargetMine employs an “unsupervised” protocol for prioritisation, as\nopposed to most other comparable tools such as ToppGene  [21]  and Endeavour  [20] , which are\n“supervised” learning methods. Thus, while direct comparison with\nthese other tools is difficult (and our data warehouse will complement, not\nreplace, stand-alone tools), the preliminary results above suggest that\nTargetMine is well suited for target prioritisation. In our group, we have been\nusing TargetMine for analysing a diverse array of experimental data and we have\nverified experimentally that some of the prioritised genes have been associated\nwith the disease of interest  [54] .\nTargetMine is structured to accommodate increasingly available biological data\nfrom large-scale experiments. Inclusion of new data sources would enable\nenhanced repertoire of functional associations currently available in TargetMine\nand at the same time expand the coverage to newer systems relevant to candidate\ngene prioritisation and drug discovery. We plan to add new data including\nhost-pathogen interactions, specific gene and protein expression patterns,\nrelationships between potential targets and chemical compounds and/or moieties,\nprotein-compound interactions and single nucleotide polymorphisms (SNPs). We aim\nto supplement the newer data sources with further developments in the TargetMine\nweb interface, lists, templates and tools for data visualisation (such as novel\nwidgets) and analysis.\nTargetMine is an integrated data warehouse that enables complicated searches that\nare difficult to perform using existing comparable tools and therefore, assists\nin efficient target prioritisation. The benchmarking results for our proposed\nprotocol for target gene prioritisation suggested the effectiveness of\nTargetMine in target discovery. The flexibility in TargetMine structure ensures\nthat different types of biological data can be readily added and analysed to\ngenerate new hypotheses for further investigation. The inclusion of additional\ndata sources and analytical tools will greatly enhance the ability of TargetMine\nto investigate biological systems for better target discovery.\n\nInterMine was downloaded from  http://www.intermine.org . New\nparsers were written in Java and integrated into the InterMine code base. A list of\nURLs for the individual data sources can be found in  Table S4 . Part\nof OMIM data, not available in downloadable files, was retrieved from the online\nresource using custom PERL scripts and TF-target associations were manually\nprocessed prior to integration into TargetMine.\nTo benchmark our gene prioritisation protocol, we performed target gene\nprioritisation on 19 sets of known disease-associated genes (denoted by set\n x ) compiled from the literature  [51] . We first created test datasets\n(set  y ), where each curated gene set was merged with twice its\nnumber of unrelated randomly selected human genes (set  r ) to\nincorporate background “noise”. To avoid any bias incurred due to the\nselection of random genes, the process was repeated 10 times to infer 10 test gene\nsets for each curated gene list. The prioritisation tests ( Figures 2  and  3 ) were then performed for each test gene set. We\ngathered the genes mapped to up to the top 10 associations, retrieved from KEGG, GO\nand OMIM databases to infer prioritised genes (set  z ). These were\nthen compared with the curated gene sets ( x ∩ z )\nand the efficiency of the prioritisation procedure was estimated with sensitivity\nand precision measures ( Table S2 ). The True Positives\n( TP ) in  z  were defined as genes present in\n x , while those corresponding to  r  were defined\nas False Positives ( FP ). The False Negatives ( FN )\nwere those genes corresponding to  x  that were not included in\n z  at the specified threshold, while the True Negatives\n( TN ) were genes corresponding to  r  correctly\nleft out from the list of prioritised genes at a given threshold. Sensitivity,\nmeasuring the proportion of the known disease-associated genes that were correctly\nprioritised, was defined as\n TP /( TP + FN ) and\nprecision, measuring the proportion of the prioritised genes that were known\ndisease-associated genes, was defined as\n TP /( TP + FP ). The\nperformance of the prioritisation protocol was also assessed using the F-score\ndefined as 2(precision×sensitivity)/(precision+sensitivity)  [55] ,  [56] .\n\nA full list of newly defined classes in TargetMine.\n(XLS)\nClick here for additional data file.\nDetailed benchmarking results for candidate gene prioritisation with\nTargetMine using 19 sets of known disease-associated genes.\n(XLS)\nClick here for additional data file.\nDetailed benchmarking results for candidate gene prioritisation with\nTargetMine using 19 sets of known disease-associated genes with increased\nbackground noise.\n(XLS)\nClick here for additional data file.\nA list of URLs for the individual data sources in TargetMine.\n(XLS)\nClick here for additional data file.","source_license":"CC-BY-4.0","license_restricted":false}