Results
The PRISMA flow chart for the selection of studies and genes is shown in Fig. 1 . The PRISMA checklist is provided as supplementary information (File S 1 ). The PubMed search imported into PubTerm identified 2496 articles, which referred to 1172 genes. The genes were reduced to 1055 after filtering for human genes (Fig. 1 , Table S 1 ). Then, each gene was carefully curated, annotated, and categorized as described in methods. The curation revealed that 400 genes were somehow potentially related to CD while 655 genes were not related due to annotations errors, the gene mention was casual, no association was found, or no evidence of variation was shown (mainly in subsequent gene expression changes). From the 400 genes, 81 were finally categorized as associated with a related disease such as IBD in general, UC, familial diarrhea syndrome, colorectal cancer, or chronic lymphocytic leukemia. Ninety-three genes were classified to other genetic alterations because its gene was uncertain, which included genes identified through haplotypes or intergenic regions in GWAS. Thus 226 genes were confirmed as associated with CD from this curation (Fig. 1 ). Fig. 1 Summary of categorized genes for Crohn’s disease. The numbers at left in arrows at the bottom represent the genes from search [ 1 ], while the numbers at right correspond to search [ 2 ]. * SNX20 was added manually
Summary of categorized genes for Crohn’s disease. The numbers at left in arrows at the bottom represent the genes from search [ 1 ], while the numbers at right correspond to search [ 2 ]. * SNX20 was added manually
We also used GWAS Catalog [ 30 ] as a source of gene information. From 22 publications for CD risk, we obtained the 525 comprising variants. Variants were further filtered by removing those whose tagged gene was not reported, were not significant, or the gene or variant was duplicated, leaving only 133 genes (Table S 2 ). From these, 77 were already categorized in the PubMed curation described above. Thus 56 genes were added to our list of genes, 27 intergenic variants were assigned to other genetic alterations, and 29 to GWAS evidence within gene . A list of the considered PubMed abstracts is shown in Supplementary Information (files S 2 and S 3 ). In summary, we identified 256 genes associated with diverse aspects of CD (Fig. 1 ).
A total of 126 genes were found to have experimental evidence of variants in CD. The top 26 genes of this category mentioned in more than 15 abstracts are shown in Table 1 , while the genes with less than 15 abstracts are summarized in Table 2 and detailed in the Supplementary Information (Table S 4 ). The topmost frequent genes for this category are well-known for their association with CD [ 11 , 12 ], such as NOD2 , TNF , IL23R , ATG16L1, TLR4, IL10, SLC22A4, SLC22A5 , and IRGM (Table 1 ). Table 1 The subset of top genes with experimental variants associated with CD (abstracts > 15) Gene Abstracts Panels GWAS catalog studies ClinVar NOD2 751 16 8 Yes TNF 156 0 0 Yes IL23R 148 1 6 Yes ATG16L1 128 1 5 Yes TLR4 88 0 0 No IL10 86 2 1 Yes SLC22A4 73 0 0 Yes IRGM 60 1 3 Yes SLC22A5 50 0 0 Yes TNFSF15 45 0 4 No NOD1 41 0 0 No IL6 41 4 0 Yes IL1B 40 0 0 No STAT3 37 1 2 No NFKB1 36 0 0 No DLG5 35 0 0 No IL12B 34 0 0 No ABCB1 33 1 0 Yes KRAS 33 0 0 No PTPN22 31 0 0 No IL1RN 29 0 0 No IL23A 28 0 0 No CD14 28 0 0 No PTPN2 24 0 0 No IL10RA 20 2 0 Yes NLRP3 19 0 0 No MEFV 20 1 0 No IL4 18 0 0 No NKX2–3 17 0 0 No ICAM1 16 0 1 No IFNG 16 0 0 No TLR9 16 0 0 No Table 2 Genes with experimental variants associated with CD mentioned by less than 15 abstracts. Details are provided in Table S 1 . * denotes manual addition ACE CFTR EPX IL10RB MST1 PTGS2 TLR6 AGER CLEC2D ERAP2 IL16 MTRR REEP6 TNFAIP3 AGT CREM FAS IL18 MUC2 SFTPD TNFRSF1A APOE CSF1R FUT2 IL27 MUC3A SLC11A1 TNFRSF1B ATG16L2 CSF2RB GC IL4R MYO9B SLC15A1 TNFSF8 BDKRB1 CX3CR1 GSTT1 IRF1 NAT2 SLC22A1 TRAIP BPI CXCL16 HLA-DQA2 IRF5 NCF2 SLC39A8 UCP2 BTNL2 CXCR4 HLA-G JAK2 NCF4 SLCO3A1 ULK1 CALCOCO2 CYP2A6 HNF4A KCNN4 NFKBIA SMAD3 XIAP CARD9 DEFB1 HNRNPD KRT8 NOS2 SNX20* ZNF365 CCR2 DLG1 HSPA1L MAGI3 ORMDL3 TCN2 CCR5 DMBT1 HSPA4 MAP3K8 POU5F1 TIMP1 CD24 DNAH12 IFNA10 MIF PPARG TLR1 CD40LG DUOX2 IFNA4 MLN PTEN TLR5
The subset of top genes with experimental variants associated with CD (abstracts > 15)
Genes with experimental variants associated with CD mentioned by less than 15 abstracts. Details are provided in Table S 1 . * denotes manual addition
Besides the above 126 genes (Tables 1 and 2 ), we also found 71 genes associated with CD that were categorized as GWAS evidence within gene where an SNP is located within genomic coordinates, either an intron or exon (Table 3 ). Additionally, 18 genes were found to be specifically associated with treatment response in CD and 41 related to disease complications (Table 3 ). Table 3 Other genes associated with CD for diverse categories. * Retrieved from GWASCatalog. + Retrieved from a panel Category Feature Genes (Abstracts) GWAS evidence within the gene SNP within gene associated to CD from GWAS LRRK2 (19), STAT4 (14), IFNGR2 (3)*, CYLD (2), CLEC16A (2), LY75 (2), TLR8 (2), ZMIZ1 (2), CXCL12 (2), ZAP70 (2)+, TLR10 (2), PER3 (2), MTMR3 (2), MAGI1 (2), CD40 (2)*, TAB1 (1), CDYL2 (1), ELF1 (1)*, NLRP11 (1), IL2RB (1), MAP3K1 (1), TIMMDC1 (1), PDE2A (1), PRKCQ (1), SLC22A23 (1), TLE1 (1), TRPM2 (1), MORC4 (1), CYP4F2 (1), CLCA2 (1), SLC23A1 (1), GCKR (2)*, IL18RAP (4)*, BRD2 (2)*, INAVA (1)*, IL2RA (4)*+, SP140 (1)*, ITLN1 (1)*, BACH2 (4)*, GPR65 (2)*, IL1RL1 (4)*, PUS10 (1)*, ANKRD55*, OSMR*, CDH13*, DENND1B*, DNMT3A*, FOSL2*, JAZF1*, KSR1*, LPP*, TAB2*, NDFIP1*, NFATC1*, PLCL1*, RFT1*, RSPO3*, SLAMF8*, THADA*, UBE3D*, ADCY3*BANK1*,BSN*,CDC37*, DAP*,FCGR2A*,HORMAD2*,PLA2G4A*,SMURF1*,TRAF3IP2*,TSPAN14* Genetic evidence in treatment response Azathioprine TPMT (41), ITPA (4), GSTA1 (1) Anti-TNF TLR2 (26), IFNGR1 (2), TBX21 (2) Infliximab FCGR3A (11), IL17F
+ (7), CASP9 (7), ADAM17
+ (3), TRAF3IP2 (2), FCGR1A (1), TNFAIP6 (1), ZNF133 (1) Corticosteroid dependency and resistance NR3C1 (5) Thiopurine NUDT15 (7), NAT1 (2) Immunosuppressive therapy necessity IL1R1 (4) Genetic evidence in related complications Surgical intervention HSPA2 (6), CHRNA5 (1), CNTF (1), MMP9 (1), TSPAN14 (1), SMURF1 (1) Structuring behavior /Aggressive disease progression SERPINE1 (4), IDO1 (2) , SELL (2), HLA - DOA (1), CSF2RA (1), CYBA (1), FAAH (1), ZBTB44 (1), CACNA1E (1), XPO1 (1), KIAA1614 (1), SULF2 (1) Food intolerance FOXO3 (3), (mustard, ginger, tomatoes, wasabi) TNF production NLRP12 (3) Granuloma formation ATG4A (2), ATG2A (1), ATG4D (1), FNBP1L (1) Develop CD before 40 years of age CNR1 (2) Bone mineral density COL1A1 (2) Ileal CD FUT3 (2), MIR196A2 (2), MIR122 (1) Variation of GMSI level RCL1 (1) Pouch outcome DAGLB (1) Vitamin D levels SCUBE3 (1), PHF11 (1) Fibrostenotic CD PRPF31 (1) Favorable disease recurrence TIMP2 (1), CYP26B1 (1) Tuberculosis and CD IL22RA1 (1) Stenotic complications MMP3 (1), MMP1 (1) Linear growth affected DYM (1) Colon location in CD MIR124 – 1 (1)
Other genes associated with CD for diverse categories. * Retrieved from GWASCatalog. + Retrieved from a panel
Of the 126 genes corresponding to the categories of experimental evidence of variants plus 71 genes with GWAS evidence within gene, only 17 genes (< 12%) were found to be annotated in ClinVar [ 32 ]. In this context, to support our systematic categorization, a supplementary file is provided with the information referring to the location of the variants that were not found in ClinVar (Tables S 4 and S 5 ). This information was revised and obtained either by the original paper or by the information related to the SNP reported at dbSNP [ 43 ].
To provide an overview of the 126 genes with experimental evidence to CD, a functional bioinformatics analysis was performed [ 33 , 44 ]. For this, we assessed whether the genes prioritized in our study are indeed statistical and biologically relevant. We performed gene set enrichment analysis in multiple databases containing different biological terms, including pathways (KEGG), diseases (GAD), and gene ontology (GO) terms. Detailed results of all enriched gene sets are present in the Supplementary Information. Because the number of significant terms was high, repetitive, and difficult to interpret, we grouped the terms by biological and genetic similarity (see Methods ).
Regarding diseases, as expected, the most similar term to CD is IBD and UC, validating our strategy (Fig. 2 ). We observed a dense group of genes and diseases where CD is located close to other groups of autoimmune and inflammatory diseases, certain types of cancer, hypersensitivity disorders caused by allergies and intolerance, infections by virus and bacteria, pregnancy complications, and metabolic complications. Fig. 2 Functional and expression analysis of CD-associated Genes. A The y-axis comprises the GO, KEGG, and disease terms. The x-axis comprises the 126 genes used for the analysis. The heat map show values from 0 to 1, corresponding to the average presence of a gene within all terms merged in each group. Diseases include 10 groups, autoimmune ( Lupus Erythematosus Systemic (LES), Psoriasis (PS) , Diabetes type 1 (DT1), Vitiligo and Arthritis Rheumatoid (RA), Chronic/Inflammatory diseases ( Psoriasis, Endometriosis, Sarcoidosis, and Cystic fibrosis ), infections ( Leprosy, Tuberculosis, Sepsis, Dengue, Hepatitis, and HIV ), cancer ( Meningioma, cervical, lung, esophageal, liver, ovarian, stomach and prostate cancer), hypersensitivity ( Asthma, Atopy, Celiac disease, and dermatitis ), pregnancy complications ( Abortion, preeclampsia and premature birth ), vascular diseases ( Atherosclerosis, restenosis, and thromboembolism ), brain and mental diseases ( Depression, Migraine, Parkinson and Schizophrenia ) and metabolic complications ( Hypercholesterolemia, Obesity, Diabetes type 2 (DT2) and metabolic syndrome ). B Differential Gene Expression of Genes represented by fold changes of all the CD genes, which show to be significantly different in at least one comparison. * denotes significance at q < 0.1 (holm p -value adjusted). Figures were mainly rendered in R software ( https://cran.r-project.org/ )
Functional and expression analysis of CD-associated Genes. A The y-axis comprises the GO, KEGG, and disease terms. The x-axis comprises the 126 genes used for the analysis. The heat map show values from 0 to 1, corresponding to the average presence of a gene within all terms merged in each group. Diseases include 10 groups, autoimmune ( Lupus Erythematosus Systemic (LES), Psoriasis (PS) , Diabetes type 1 (DT1), Vitiligo and Arthritis Rheumatoid (RA), Chronic/Inflammatory diseases ( Psoriasis, Endometriosis, Sarcoidosis, and Cystic fibrosis ), infections ( Leprosy, Tuberculosis, Sepsis, Dengue, Hepatitis, and HIV ), cancer ( Meningioma, cervical, lung, esophageal, liver, ovarian, stomach and prostate cancer), hypersensitivity ( Asthma, Atopy, Celiac disease, and dermatitis ), pregnancy complications ( Abortion, preeclampsia and premature birth ), vascular diseases ( Atherosclerosis, restenosis, and thromboembolism ), brain and mental diseases ( Depression, Migraine, Parkinson and Schizophrenia ) and metabolic complications ( Hypercholesterolemia, Obesity, Diabetes type 2 (DT2) and metabolic syndrome ). B Differential Gene Expression of Genes represented by fold changes of all the CD genes, which show to be significantly different in at least one comparison. * denotes significance at q < 0.1 (holm p -value adjusted). Figures were mainly rendered in R software ( https://cran.r-project.org/ )
Gene sets from GO are divided into cellular components, biological processes, and molecular functions. Within biological processes terms, we observed significant enrichment in processes related to response to bacteria , positive regulation of nitric oxide (NO), ERK , and NFκβ , apoptosis, cell proliferation, response to lipopolysaccharide , inflammatory and immune response. For molecular function, cytokine activity , interleukine-1 receptor binding, receptor activity, and protein homodimerization activity were significantly enriched by the CD-risk genes analyzed. For cellular components, only membrane, plasma membrane, and extracellular region were significant . Overall the significantly enriched GO terms point to known CD terms such as the immune response, cytokine activity, and signaling receptors as the primary source of functional causes (in terms as Immune Response, Infection, Interleukin-1 receptor binding, Cytokine and defense response, Cytokine activity, Autoimmune Disease ).
Additionally, among the enriched pathways identified were JAK-STAT signaling pathway , cytokine receptor interaction , NOD-like receptor signaling pathway , NF-kappa , TNF signaling , Toll-like receptor (TLR) signaling, T cell receptor signaling , and Osteoclast differentiation . These signaling pathways converge on the activation of NF-κB, a protein complex that controls the transcription of DNA cytokine production and cell survival [ 45 ]. In addition, the KEGG comorbidities identified are infectious diseases caused by bacteria, protozoa and virus , and IBD, which are reliable associations due to the relationship between microbes and CD [ 46 ]. The mapping is, therefore, an excellent guide to connect genes and important biological aspects of CD (Fig. 2 ).
Among the genes present in the enriched biological terms, we noted two distinct groups depending on the frequency of their presence in the gene sets and their number of abstracts found by PubTerm, designated as common and sporadic (Fig. 2 ) . Briefly, the 34 common genes are highly related to diseases and biological terms and well-studied. In comparison, the 92 sporadic genes are associated with particular diseases or biological terms and not as studied in CD as the common genes. Among the common group, the most shared genes across concepts are NOD2, TNF, ICAM1 NFKBIA, NFKB1, TNFRSF1A, CD14, ACE TLR4, and TLR9, which are involved in both TNF and of NF-κβ signaling pathways [ 47 – 49 ] . Also, IFNG, IL1B, IL6, IL23R, IL10, IL4, IL12B, IL1RN, IL18, and IL4R, which are all well-known cytokines or related genes, and contribute to the inflammatory response and cytokine interaction process [ 50 – 52 ]. The sporadic group comprised 92 genes that were much less frequent among enriched terms, where ATG16L1, SLC22A4, IRGM, SLC22A5, TNFSF15, NOD1, PTPN2, PTPN22, and DLC5 being more frequently mentioned.
DNAH12, ERAP2, FUT2, ORMDL3 and, TRAIP were more specifically enriched in the CD disease term and less common for the remaining terms.
Once we noted two clear sets of common and sporadic genes that were associated with specific terms, we considered whether the genes might also be grouped by other phenotypes that could explain CD symptoms. Thus, we used the Gene Network tool, which clusters genes with similar Human Phenotype Ontology. Five sub-networks were identified from the 126 genes categorized as experimental evidence . We then merged three highly interconnected sub-networks and subsequently analyzed the three resultant modules (73, 26, and 26 genes respectively for Module 1 in green, Module 2 in purple, and Module 3 in blue, as shown in Fig. 3 ). Next, for each module, the genes were also functionally analyzed to identify their potential phenotypic consequences. For this, we also used the Gene Network tool. Only exclusive terms for each module that were significant after Bonferroni correction were analyzed (see Methods ). Fig. 3 Network analysis for the three main groups identified. Group 1 (green): 73 genes, Group 2 (purple): 26 genes and Group 3 (blue): 27 genes. Figure adapted from Gene Network tool
Network analysis for the three main groups identified. Group 1 (green): 73 genes, Group 2 (purple): 26 genes and Group 3 (blue): 27 genes. Figure adapted from Gene Network tool
For the first module (green), 60 phenotypes were retrieved; most of them related to severe symptoms such as Ocular complications, Altered immune system, Sepsis, Bowel incontinence, Heart complications, Endocrine abnormalities, Dysphagia and constipation, and muscle problems . For the second module (purple), 11 terms were identified mainly related to Neoplasm of the gastrointestinal tract, Hyperhidrosis, Respiratory complications, and Visual impairment . These symptoms are among common abnormalities detected in IBD patients [ 53 – 55 ]. Finally, for the third module (blue), only four consequences were identified, abnormality of glutamine metabolism , abnormality of the small intestine, and two remaining terms related to the facial skeleton . The information related to the module assigned to each gene is provided in Table S 3 .
A recent gene expression analysis of CD, UC, and controls in the colon and ileum showed that 1008 genes were differentially expressed [ 56 ]. We, therefore, explored whether the genes found genetically associated with CD in our systematic review were related to those differentially expressed (DE) between CD or UC relative to their normal ileum and colon gene expression. From the 126 genes, 67 genes were found to be DE for both UC and CD (Fig. 2 B). The overlap between the 67 genes with those 1008 is highly significant ( p = 1 − 290 , hypergeometric test), suggesting that our 126 genes are particularly enriched in DE genes. Except for SLC39A8 , FAS , IRF5 , HSPA1L , and PTEN , the vast majority were indeed more significantly associated with UC than with CD. Moreover, the majority of the genes were less expressed in CD relative to controls. We noted two solute carriers more expressed in CD than in normal colon or ileum; SLC22A4 (ergothioneine, carnitine, tetraethylammonium) probably for detoxification, and SLC22A5 (carnitine), whose variants are reported to affect the function of carnitine and organic cation transporters [ 57 ]. There are also two less expressed solute carriers, SLC11A1, with a role in the susceptibility of humans and animals to several infections [ 58 ] and SLC39A8, associated with gut microbiome composition [ 59 ].
The main therapeutic drugs for CD are azathioprine [ 60 ], infliximab, adalimumab, certolizumab pegol, ustekinumab, vedolizumab [ 61 – 63 ], prednisone, hydrocortisone, and hydrocortisone acetate [ 61 ]. To explore drug-gene interactions, an analysis was performed in DGIdb [ 41 ] using these drugs. There is a total of 78 genes that had a reported interaction with these CD therapeutic drugs, 10 of which were identified in our systematic review (Table 4 ). We reasoned that focusing on drugs that target the gene variants associated with IBD could be a strategy for CD drug repurposing [ 64 ]. Thus, to search for alternative drugs for CD, we used 13 other drugs commonly employed in the treatment of chronic autoimmune and inflammatory diseases [ 64 , 65 ]. We identified 13 CD genes (IL1B, IL1RN, IL6, ABCB1, XIAP, IFNG, ICAM1, NLRP3, JAK2, PPARG, PTGS2, APOE, and SMAD3) that show some interaction and therefore could be further explored as possible treatments for CD (Table 4 ). Table 4 Interactions between therapeutic drugs and the 126 genes Disease Drug Gene interaction Crohn’s disease Infliximab TNF , TLR4 Prednisone APOE , IFNG , ABCB1 Hydrocortisone NOS2 , ABCB1 , IL1B , AGT Adalimumab TNF Ustekinumab IL12B , IL23A Certolizumab pegol TNF Azathioprine – Hydrocortisone acetate – Vedolizumab – Chronic autoimmune and inflammatory diseases Canakinumab IL1B Rilonacept IL1B, IL1RN Metronidazole IL6 Rituximab ABCB1, XIAP Methylprednisolone ABCB1, IFNG Methotrexate IL1RN Natalizumab ICAM1 Anakinra NLRP3 Olsalazine IFNG Tofacitinib JAK2 Sulfasalazine PPARG, PTGS2 Triamcinolone APOE Dexamethasone SMAD3
Interactions between therapeutic drugs and the 126 genes
To compare the generated list of genes with genetic panels already in use, we benchmarked within those panels in the GTR [ 42 ]. We found 21 panels, of which 19 were specific to Crohn’s disease containing only 2 genes, NOD2 and IL6 . The two remaining tests were not specific for Crohn’s, IBD, and related diseases. These tests considered 70 genes, of which 22 were identified as functional variants for CD in our curation (including NOD2 and IL6) . Thus, from the 256 genes we found (Tables 3 and S 3 ), 225 genes were not included in any panels for CD or IBD-related disorders.
We also verified the identifications at the Open Targets (OT) platform, whose pipeline includes a fine mapping of variants [ 31 ]. Filtering 3093 genes for CD for genetic association score > 0.8, 178 genes were identified. Of these, 3 genes (SNN, SH2B3, and SKAP2) were not identified in the set of 1092 curated genes. From the 126 genes identified here having experimental evidence, 39 showed a low genetic association score for CD (0.5 to 0.79), 8 genes showed a good score (0.8 to 1.0) for IBD, and 52 genes do not have genetic data information in OT for CD nor IBD. Some of the 52 genes show variations that have not been reported in GWAS, explaining their absence in OT. Examples include NOD1 , ABCB1 , IL1RN , MEFV , and IL18 having variants that could not be easily identified in GWAS because there are triallelic changes [ 66 ], deletions [ 67 ], and VNTR [ 68 ]. This comparison shows that even GWAS information can leave aside some information of other variants detected through other technologies or methodologies.
Discussion
Through our methodology, we have identified 256 genes associated with some aspects of CD. Of them, 126 genes were associated with experimental evidence of variants in CD , 71 genes found in GWAS with a sequence variant within the gene , 41 genes for complications , and 18 genes for treatment response in CD.
There is an explosion of genetic data provided by the high throughput technologies such as genome-wide SNP arrays and next-generation sequencing. This growing list of associated genes has the potential to improve diagnosis and treatment, but progress has been slow. There is a need for better strategies for prioritization and curation. In this study, we found that from 126 genes with variants associated with CD, a total of 110 genes have not been included in any genetic panel for CD and related diseases. That is probably reflected by the lack of individual predictive value of most individual common SNPs. The small number of variants annotated in ClinVar [ 32 ] seems to be caused by some variants not found by GWAS. Also, the increasing tendency of acquiring genetic data suggests that more efforts and more accurate annotations, such as those provided here, are highly needed and valuable.
The functional bioinformatics analysis performed confirmed the relationship between CD and highlighted modules of genes in our systematic review. We identified autoimmune diseases that could have affected pathways similar to those of CD such as Type 1 Diabetes , Multiple sclerosis , Lupus , A rthritis Rheumatoid , and Psoriasis. These relationships among CD and other autoimmune diseases are already known and have been previously studied [ 69 , 70 ]. There are also relationships with hypersensitivity diseases such as asthma and celiac disease and metabolic complications such as T2D and hypercholesterolemia. Indeed, there are some Previous studies have shown an association of CD and IBD with asthma [ 71 ], type 1 diabetes [ 69 ], and T2D [ 72 ]. Those diseases identified in our analysis are likely to share a genetic background with CD due to their inflammation process and their condition as autoimmune diseases, as suggested by previous studies [ 69 , 71 , 72 ]. Our results highlight the genes which could be shared among conditions and allow focusing on future research efforts among these genes.
Additionally, we found functions related to immune response, cytokine activity, and receptors. It is clear that CD pathogenesis is caused by an immune imbalance [ 73 ], which was also reflected in our de novo analysis. Some hypotheses have attempted to explain its mechanisms, including delayed hypersensitivity, activation induced by food, and others [ 73 ]. These mechanisms converge into the immune response in an environment where self-tolerance has been lost and where cytokines have an active role in maintaining this pro-inflammatory state [ 73 , 74 ]. Additionally, other terms, such as apoptosis and response to Lipopolysaccharides (LPS), may provide interesting insights. LPS response is related to a monocyte/macrophage stimulation by enteric bacteria constituents [ 75 ], and resistance to apoptosis in patients with CD has also been reported [ 76 ].
We spotted signaling pathways specific for some important genes in CD converging on the activation of NF-κB, which is a protein complex that controls the transcription of DNA cytokine production and cell survival [ 45 ]. This is comparable with previous reports of abnormal activation of NF-κB, causing chronic inflammation in the bowel [ 45 ]. Similarly, pathways related to infections caused by protozoa, virus, and bacteria were identified consistent with the known relationship between microbes and CD [ 46 ]. Pathogen infections are one of the environmental factors which are likely to be a key component for CD; however, their roles or mechanisms of action remain speculative [ 77 ]. Additionally, the microbiota plays an important role [ 56 ]. Our results show that most of the genes related to pathogen infections are among the common genes and close to the pathways of NOD , TLR , and NFκβ, which could aid in the future understanding of the mechanisms of action specifically in CD. We also observed the pathway for Osteoclast differentiation , which has been recently studied, linking the function of IL-17 , and TNFa modulating bone resorption [ 78 ].
In our functional analysis, we spotted the following genes, NOD2, IL23R, IL6, IRGM, ATG16L1, and IL10, whose CD-predominant risk associations are known [ 48 , 79 – 82 ] . Among them, NOD2 has the highest contribution to CD risk alone, with 5% of penetrance and ~ 20 fold risk [ 82 ]. Other genes that are not currently present in any diagnostic test for CD or a related condition but which showed general importance for CD, related diseases, and biological process, molecular function, and pathways are TNF, ICAM1, NFKBIA, NFKB1, TNFRSF1A, CD14, ACE, TLR4, TLR9, IFNG, IL1B, IL4, IL12B, IL1RN, IL18, and IL4R, which are involved in TNF and NF-κB signaling pathways, in inflammatory response, and cytokine interaction processes [ 47 – 52 ]. These genes and others from our list could be used to design a more robust prediction panel for CD risk.
Our analysis highlighted poorly studied genes (10 or fewer abstracts). From these, FUT2, DNAH12, TRAIP, and ERAP, were identified in the functional analysis to be specific for CD (Fig. 2 ). Among the processes reported for these genes are ABH antigens expression [ 83 ], motile cilia function [ 84 ], regulation of innate immune signaling [ 85 ], and immune activation and inflammation [ 86 ].
Among these poorly studied genes, only FUT2 is currently present in a diagnostic panel related to IBD diseases. This fact remarks the importance of considering and further studying the biological implication of the less studied set of genes to increase our knowledge of this complex disease.
We identified three network modules of genes associated with specific symptoms. The first module comprising 73 genes was related to severe symptoms [ 87 ], such as Altered immune system, Sepsis, Bleeding, Muscle disorders , and Heart complications well-known in CD. The second module involved 26 genes related to hyperhidrosis , respiratory complications , and neoplasm of the gastrointestinal tract . These symptoms are among the common abnormalities detected in IBD patients [ 53 – 55 ]. The third module, including 26 genes, was associated with abnormality of glutamine metabolism and abnormality of the small intestine . Glutamine is an important supplementation in IBD patients [ 88 ], and its effects in IBD have been studied in animal models [ 89 ] and patients [ 90 , 91 ]. Thus, this module seems to map genes related to less severe consequences for CD. Thus, gene-symptom mapping may provide important insights into CD.
Differential expression analysis of the 126 genes identified in our systematic review revealed that a significant number of genes show dysregulated levels of expression in colon and ileum biopsies of both CD and UC when compared with not-IBD patients further supporting our gene prioritization approach. The greatest changes in expression are observed in the colon, with differential expression in pro-inflammatory genes ( NOD2, IL1B , and TNF ). Other observations are that changes are different among UC and CD and that a large proportion of the genes do not show evident gene expression. Thus, to further understand whether the functional implications of these changes in expression are causal for CD pathogenesis or whether CD patients carrying other specific variants show different gene expression profiles, further functional experiments are needed.
Among the 126 genes analyzed, 10 have a reported interaction with known CD therapeutic drugs, and 13 have a reported interaction with other autoimmune and inflammatory diseases. The treatment for CD is complex, and it is focused on controlling the symptoms and the remission of the disease [ 92 ]. Focusing on drugs that target the gene variants associated with IBD can be a strategy for CD drug development [ 64 ]. This highlights the necessity of considering more genes to study other possible interactions for CD beyond what is currently known and shows the importance of gene curation strategies, like the one proposed here. Further research on the genes highlighted here, and their mechanisms of interaction with CD diseases could improve the knowledge of the disease development and expand treatments.
Traditional drug development is costly and can take 10–15 years to develop an efficient drug [ 64 ]. Personalized medicine exhibits the clinical application of drug-gene interaction, where drugs are guided based on the individual’s genetics and disease progress. Targeting CD’s genetic risk regions that had been experimentally validated can improve the identification of possible drug candidates. This can be reflected in target-directed therapies, which is one of the main objectives of personalized medicine. The analysis of drug-gene interactions in a complex disease, such as MDD (major depressive disorder), allowed a better, prioritization of drug-genes sets and the identification of drugs indicating an effect on a disease, reflecting potential repurposing opportunities [ 93 ]. Nevertheless, validation studies are still required to ensure the drug-gene interaction and avoid side effects.
Our results support the consideration of several genes when studying CD. More importantly, the functional analysis provides a mapping between genes and key aspects of Crohn’s disease. The integration of other genes may also be important. For example, genes close by a non-coding GWAS SNP, i.e., intergenic variants, or those involved in related diseases, could play a role in CD etiology, but further validation or fine gene mapping is needed.