Results
Herein, we introduced a computational method to search and design drug candidate compounds for various diseases based on cell‐specific correlations between compounds and diseases. A schematic of the proposed method is presented in Figure 1 .
Overview of the proposed method. (A) Construction of an ideal therapeutic profile. All elements of the gene expression profile of the target disease were multiplied by −1. (B) Data completion for utilizing cell‐specific chemically induced profiles. Missing values in the chemically induced profiles on various cell lines were imputed using the tensor decomposition algorithm. (C) Similarity search with correlation coefficients. A cell line relevant to a specific disease was first selected. Then, the correlation coefficients between the ideal therapeutic and chemically induced profiles for the selected cells were calculated on the disease‐specific gene sets. Finally, the compounds were determined based on the correlation coefficients. (D) Generation of new chemical structures from seed compounds. In TRIOMPHE‐BOA, new compounds were generated by fusing multiple chemical structures of seed compounds on the latent space using Bayesian optimization.
The drug search comprised three steps: (i) construction of an optimal therapeutic gene expression profile (termed ideal therapeutic profile) that is inversely correlated with the disease‐specific gene expression profile (termed disease‐specific profile), (ii) data completion for utilizing cell‐specific chemically induced profiles, and (iii) similarity search with Pearson's product–moment correlation coefficients (hereafter referred to as correlation coefficients). In the first step, we defined the sets of differentially expressed genes (DEGs) in the disease‐specific profiles as disease‐specific genes and extracted the disease‐specific genes from all genes in the disease‐specific profiles. Then, all gene expression values of the disease‐specific genes were multiplied by −1, and the resulting values were set to the elements constructed in the ideal therapeutic profile (Figure 1A ). In the second step, missing values in the chemically induced profiles on various cell lines were imputed using the tensor decomposition algorithm [ 23 ] (Figure 1B ). In the third step, a cell line relevant to a given disease was first selected for the similarity search (Figure 1C ). Next, correlation coefficients of the ideal therapeutic profile with the chemically induced profiles on the selected cell were calculated. Subsequently, we selected the top 10 high‐scoring compounds as drug candidates.
The drug design procedure is an extension of the drug search to generate new chemical structures. It comprises the following two steps: selection of seed compounds and generation of new chemical structures. During the selection of seed compounds, we selected the top three high‐scoring compounds as seed compounds in the same steps as the drug search method. In the generation step, new chemical structures were generated from the fusion of the seed compounds using a deep learning‐based structural generator (Figure 1D ). Herein, transcriptome‐based inference and generation of molecules with desired phenotypes using the Bayesian optimization algorithm (TRIOMPHE‐BOA) was used as the structural generator [ 24 ]. This model generates new compounds by fusing multiple chemical structures of seed compounds on the latent space of a transformer‐based variational autoencoder called TransVAE [ 25 ], using Bayesian optimization. The performance of TRIOMPHE‐BOA depends on the selection of seed compounds.
We employed our proposed method to gastric cancer and atopic dermatitis because of high medical needs. Gastric cancer is the fifth most commonly occurring cancer with low survival rates [ 26 ]. Atopic dermatitis is a chronic inflammatory disease of the skin that has a significant effect on patients’ quality of life [ 27 ]. To evaluate the cell specificity, we designated cell lines derived from the tissues where the disease of interest developed as “tissue‐associated cells” and cell lines different from tissues where the disease of interest developed as “tissue‐unassociated cells.”
To evaluate the effects of considering cell specificity, we compared the results of drug search and design between using the tissue‐associated cells and using tissue‐unassociated cells. Thus, we selected AGS (a cell line derived from gastric tissue) and A375 (a cell line derived from skin tissue) for gastric cancer and atopic dermatitis as tissue‐associated cells, respectively. Then, the drug search and design for tissue‐associated cells were performed, which were referred to as “tissue‐associated search” and “tissue‐associated generation,” respectively (Figure 1C,D ). This procedure corresponds to the proposed method. In contrast, we selected MCF7 (a cell line derived from breast tissue) as a representative of tissue‐unassociated cells, because chemically induced profiles have been abundantly measured for MCF7, and the gene expression levels on MCF7 were usually used in most previous studies on Omics‐based drug discovery [ 5 , 6 , 7 , 8 , 9 – 10 ]. The drug search and design for the tissue‐unassociated cell were performed, which were referred to as “tissue‐unassociated search” and “tissue‐unassociated generation,” respectively. This procedure corresponds to the previous method that does not consider the cell specificity.
We performed the drug search to identify compounds that would counteract the disease‐specific gene expression patterns for gastric cancer and atopic dermatitis. We then examined the differences in the chemical structures of the compounds detected by the tissue‐associated and tissue‐unassociated searches. The detailed results are shown in Figure S1 in the Supporting Information.
We performed a hierarchical clustering analysis for 15,420 compounds in the chemically induced gene expression data. The Tanimoto distance based on the ECFP4 fingerprint was calculated between compounds, and a hierarchical clustering was performed using the Ward's algorithm (see the “Hierarchical clustering by chemical structure similarity” subsection in Materials and Methods for more details). As a result, the compounds were grouped into nine clusters, as shown in Figure 2A .
Structural features of the compounds detected by the tissue‐associated and tissue‐unassociated searches. (A) The dendrogram based on the structural dissimilarity. The gray dashed line indicates the threshold of clustering distance. The resulting clusters were highlighted in different colors. (B) The classifications of the detected compounds in gastric cancer. (C) The classifications of the detected compounds in atopic dermatitis. The horizontal axis represents the search method name, and the vertical axis represents the ranking of the detected compounds. The cluster number for each detected compound is written at the bottom of the corresponding compound.
Figure 2B shows the cluster numbers of the top 10 compounds detected by the tissue‐associated and tissue‐unassociated searches for gastric cancer. In the case of the tissue‐unassociated search, compounds were assigned clusters 3, 5, 7, and 8. On the other hand, in the case of the tissue‐associated search, compounds were assigned clusters 1, 2, 3, 5, 6, 7, and 8. The compounds in the tissue‐associated search were grouped into more diverse clusters, compared with those in the tissue‐unassociated search.
Figure 2C shows the cluster numbers of the top 10 compounds detected by the tissue‐associated and tissue‐unassociated searches for atopic dermatitis. In the case of the tissue‐unassociated search, compounds were assigned clusters 2, 3, 5, and 6, and many compounds were dominantly grouped into a specific cluster. On the other hand, in the case of the tissue‐associated search, compounds were assigned clusters 3, 5, 6, 7, and 9, and compounds tended to be grouped into diverse clusters. These results suggest that the proposed tissue‐associated search can detect compounds with diverse chemical structures.
Herein, the chemically induced counteraction of disease‐specific gene expression was evaluated. Thus, we standardized the disease‐specific and chemically induced profiles of the detected compounds on the disease‐specific genes. If the signs of values of a disease and a detected compound are opposite (i.e., plus and minus) in each element of the gene expression profiles, the compound is considered to counteract the disease‐specific gene expression pattern.
Figure 3 shows that the number of compounds that counteract each gene of the diseases in the tissue‐associated search is represented as bar graphs. In Figure 3A , more than half of the compounds counteracted 48 of 78 upregulated genes and 44 of 63 downregulated genes in the gene expression profiles of gastric cancer. Only nine upregulated and eight downregulated genes were not counteracted by any of the detected compounds. In Figure 3B , more than half of the compounds counteracted 29 of 41 upregulated and 39 of 54 downregulated genes in the gene expression profiles of atopic dermatitis. Furthermore, only two upregulated genes were not counteracted by any of the detected compounds. These results indicate that the compounds detected based on tissue‐associated cells tend to counteract with disease‐specific gene expression patterns.
Genes counteracted by the compounds detected based on the tissue‐associated cells. (A) The number of the compounds detected by the tissue‐associated search counteracting with each gene for gastric cancer. (B) The number of the compounds detected by the tissue‐associated search counteracting with each gene for atopic dermatitis. The red and green bars show the results for upregulated and downregulated genes in the disease of interest, respectively. Because the black dashed lines in the figures represent half of the detected compounds, the gene with a bar graph higher than the line was counteracted by many detected compounds.
We investigated the effects of the detected compounds on the strongly expressed disease‐specific genes that play critical roles in the disease; the detected compounds may strongly counteract the disease‐specific expression of such genes. This study regarded genes with absolute values > 1 in the gene expression profiles as strongly expressed. The visualization on the strongly expressed genes is shown in Figure 4 (see Figure S2 in the Supporting Information for more details).
Heatmaps of gene expression profiles in strongly expressed genes. (A) The heatmap of gene expression profiles of gastric cancer and the compounds detected by the tissue‐associated search in strongly expressed genes. (B) The heatmap of gene expression profiles of atopic dermatitis and the compounds detected by the tissue‐associated search in strongly expressed genes. The horizontal axis shows the strongly expressed genes, and the vertical axis represents the disease of interest and the ranking of the detected compounds. The low values of profiles are shown in green, whereas the high values are indicated in red.
Figure 4A shows the gene expression profiles of gastric cancer and the compounds detected by the tissue‐associated search in the strongly expressed genes. Among such genes, more than half of the compounds had absolute values > 1, and the plus and minus signs of their chemically induced profiles were opposite to those of the disease‐specific profiles for RPS5 , HSPD1 , HSPB1 , ITGB1BP1 , GFUS , P4HA2 , SYNE2 , DNAJC15 , TIMM17B , CCNA1 , OXCT1 , XBP1 , LIPA , and RGS2 . This finding indicates that many compounds detected by the tissue‐associated search strongly counteract with disease‐specific gene expression of gastric cancer on these genes. Figure 4B shows the gene expression profiles of atopic dermatitis and the compounds detected by the tissue‐associated search in the strongly expressed genes. Among such genes, more than half of the compounds had absolute values > 1, and the plus and minus signs of their chemically induced profiles were opposite to those of the disease‐specific profiles for SFN , PGAM1 , IGFBP3 , MIF , NFKBIA , ZFP36 , CREG1 , and FOS . This finding indicates that many compounds detected by the tissue‐associated search strongly counteracted with the disease‐specific gene expression of atopic dermatitis on these genes. These results suggest that the compounds detected by the tissue‐associated search have strong counteracting effects on the strongly expressed genes in gastric cancer and atopic dermatitis.
We performed pathway‐enrichment analysis using the DAVID tool [ 28 ] for the identified genes counteracted by the detected compounds. We then examined whether the detected compounds act on the pathways related to the disease of interest. The results of the pathway‐enrichment analysis are shown in Tables 1 and 2 .
Pathway‐enrichment analysis for the genes counteracted by the compounds in the tissue‐associated search for gastric cancer. The enriched pathways with the upregulated and downregulated genes in gastric cancer were shown. These genes were counteracted by the compounds detected by the tissue‐associated search. The P value was set at 0.05.
Pathway‐enrichment analysis for the genes counteracted by the compounds in the tissue‐associated search for atopic dermatitis. The enriched pathways with the upregulated and downregulated genes in atopic dermatitis are shown. These genes were counteracted by the compounds detected by the tissue‐associated search. The P value was set at 0.05.
In the case of the compounds detected by the tissue‐associated search for gastric cancer, the upregulated genes were enriched in cancer‐related pathways, such as pathway in cancer ( P = 1.74 × 10 −3 ) and microRNAs in cancer ( P = 4.51 × 10 −2 ), and in the stomach‐related pathways such as gastric acid secretion ( P = 4.33 × 10 −2 ) (Table 1A ). Furthermore, the downregulated genes were enriched in several pathways such as inflammatory bowel disease ( P = 2.34 × 10 −2 ) and the epithelial cell signaling in Helicobacter pylori infection ( P = 2.69 × 10 −2 ) (Table 1B ). The H. pylori infection is one of the risk factors of gastric cancer and promotes epithelial–mesenchymal transition and gastric cancer progression by worsening the tissue microenvironment [ 29 ].
In the case of the compounds detected by the tissue‐associated search for atopic dermatitis, the upregulated genes were enriched in the HIF‐1 signaling pathway ( P = 8.69 × 10 −5 ) and the mTOR signaling pathway ( P = 4.70 × 10 −2 ) (Table 2A ). It is reported that the mTOR signaling pathway plays a crucial role in the pathogenesis of atopic dermatitis by in vitro experiments, and it is also revealed that the pathway is activated by IL‐13, leading to increase IL‐13 activity and suppress epidermal barrier‐related proteins such as filaggrin, loricrin, and involucrin [ 30 ]. The downregulated genes were enriched in immune response‐related pathways, such as Th1 and Th2 cell differentiation ( P = 3.21 × 10 −2 ) and Th17 cell differentiation ( P = 4.30 × 10 −2 ) (Table 2B ). In summary, the results of the pathway‐enrichment analysis demonstrated that the detected compounds may act on the pathways related to the diseases of interest. Therefore, the detected compounds are expected to have therapeutic effects on these diseases, as determined by the pathway‐enrichment analysis.
The effect of considering cell specificity in drug design using structure generators was investigated. We then generated new compounds from seed compounds based on the tissue‐associated and tissue‐unassociated generations. The new compound is considered to have a drug‐like structure for the disease if the chemical structure of the new compound is similar to that of the registered drug. Thus, we calculated the structural similarity of the new compounds generated by the tissue‐associated and tissue‐unassociated generations to the registered drugs and compared the maximum values. The seed compounds by the tissue‐associated and tissue‐unassociated generations are shown in Figure S3 in the Supporting Information. The structural similarities between the new compounds and registered drugs are shown in Figure 5 .
New chemical structures generated by the tissue‐associated and tissue‐unassociated generations. The structures of newly generated compounds and their structural similarities with registered drugs are shown. The “T” indicates the Tanimoto coefficient in ECFP4 between the compound and the registered drug. (A) The structural similarity comparison between the new compounds generated by the tissue‐associated and tissue‐unassociated generations for gastric cancer. (B) The structural similarity comparison between the new compounds generated by the tissue‐associated and tissue‐unassociated generations for atopic dermatitis.
In the case of gastric cancer, the new compounds generated by the tissue‐associated generation showed the highest structural similarity to floxuridine (Tanimoto coefficient: 0.2162) (Figure 5A ). In contrast, the maximum structural similarity between the new compound generated by the tissue‐unassociated generation and floxuridine was 0.1429. This finding indicates that the structural similarity of the new compound by the tissue‐associated generation is higher than that of the new compound by the tissue‐unassociated generation. According to the chemical structures, 2‐hydroxy‐6‐methylpyridine in the new compound by the tissue‐associated generation was similar to 5‐fluorouracil structure in floxuridine, and both floxuridine and the new compounds are characterized by the ring structures connected to the six‐membered ring by a single bond.
In the case of atopic dermatitis, the new compound generated by the tissue‐associated generation showed the highest structural similarity to pyridoxal phosphate hydrate (Tanimoto coefficient: 0.2078) (Figure 5B ). In contrast, the maximum structural similarity between the new compound identified by the tissue‐unassociated generation and pyridoxal phosphate hydrate was 0.1446. This finding indicates that the structural similarity of the new compound by the tissue‐associated generation was higher than that of the new compound by the tissue‐unassociated generation. The registered drug and the new compound generated by the tissue‐associated generation are characterized by having 2‐methylpyridine. These results demonstrate that by considering cell specificity, our proposed method can generate new compounds with more drug‐like chemical structures than the conventional method.
Materials
The gene expression profiles of the patients were obtained from CRowd Extracted Expression of Differential Signatures (CREEDS) [ 37 ]. The gene expression profiles in CREEDS comprised scores calculated using the characteristic direction method [ 38 ]. This method compares gene expression levels in diseased tissue with those in control tissue. We then constructed disease‐specific profiles by averaging the gene expression profiles of patients with the same disease, which comprised 14,804 dimensions of the gene expression values. In total, we extracted DEGs that were common across the disease‐specific and chemically induced profiles from the 14,804 genes for each disease, which we termed disease‐specific genes (141 for gastric cancer and 95 for atopic dermatitis). Ideal therapeutic profiles were obtained by multiplying all elements of the disease‐specific profile by −1.
We obtained the level 5 gene expression data from LINCS [ 39 ]. The data consist of gene expression profiles of 15,630 compounds, 978 genes, and 94 cell lines, which have numerous missing values for a range of cell lines. Next, we used a tensor decomposition method [ 23 ] to complete the missing values. Then, we extracted chemically induced profiles in AGS for gastric cancer, A375 for atopic dermatitis, and MCF7 for the control from the completed tensor. Furthermore, the 15,630 compounds were converted into SMILES strings using a web‐application on the PubChem database [ 40 ], and 15,420 compounds remained. We obtained gene expression profiles of 15,420 compounds × 978 genes, where the genes correspond to the L1000 landmark genes, for each cell line. Furthermore, disease‐specific genes were extracted from L1000 landmark genes to search for drug‐like compounds.
Twenty registered drugs for gastric cancer and 46 registered drugs for atopic dermatitis were obtained from the KEGG DRUG database [ 41 ] (Table S1 in the Supporting Information). These drugs were used for evaluating the drug‐likeness of newly generated compounds in the tissue‐associated and tissue‐unassociated generations.
The newly generated compounds by the tissue‐associated and tissue‐unassociated generations were compared with the registered drugs against the target diseases. This is because the bioactivity of a compound is expected to be similar to that of a registered drug if the compound is structurally similar to the registered drug. The Tanimoto coefficient with extended‐connectivity fingerprint 4 (ECFP4) [ 42 ] was used to evaluate structural similarity. ECFP4 is a binary vector of chemical structures based on substructures within the distance of two atoms in a radius. In this study, the dimension of the binary vector of the ECFP4 was 2,048. The Tanimoto coefficient is defined as follows
(1)
T = N AB N A + N B − N AB
where A and B represent chemical structures; N
A , N
B , and N
AB are the total numbers of binary vectors in the ECFP4 of A and B and common to A and B, respectively. If the Tanimoto coefficient is 0, A is not structurally similar to B, and if it is close to 1, A is structurally similar to B.
To confirm the structural differences of compounds detected by the tissue‐associated and tissue‐unassociated searches, we performed a hierarchical clustering analysis of compounds based on their structural dissimilarities. First, the Tanimoto distances with the ECFP4 fingerprint were calculated for all possible pairs of compounds. The Tanimoto distance is defined as 1 minus the Tanimoto coefficient, and it was used as the degree of dissimilarity between the compounds. Then, a hierarchical clustering was performed using Ward's algorithm. In this study, the threshold of clustering distance was set to 13.
The compounds detected by the tissue‐associated search were evaluated based on how much the compounds counteracted with the disease‐specific gene expression. The gene expression profiles on the disease‐specific genes were standardized according to the following equation:
(2)
z = x − μ σ
where x represents the value of each gene in the gene expression profile and µ and σ represent the mean and standard deviation of the gene expression profile, respectively.
Conclusion
In this study, we proposed a novel computational method to obtain drug‐like compounds for the disease of interest. We then applied the proposed method to search for drug‐like compounds and generate new drug‐like chemical structures for gastric cancer and atopic dermatitis. In the drug search task, it was shown that the detected compounds with the tissue‐associated search had more diverse chemical structures compared with those with the tissue‐unassociated search. The compounds tended to counteract with many disease‐specific gene expressions, and some strongly expressed disease‐specific gene patterns. In the drug design task, it was shown that the newly generated compounds with the tissue‐associated generation possessed more drug‐like chemical structures than those with the tissue‐unassociated generation. The proposed method is expected to contribute exploration of compounds with desired phenotypes for omics‐based drug discovery.
Discussion
In this study, we make a hypothesis that the use of chemically induced profiles in the same cell type as target diseases will enable us to propose drug candidate compounds with more desirable phenotypes than conventional methods. One of the limitations of previous studies is that they do not consider cell specificities of chemically induced profiles. An advantage of this study over previous studies is the consideration of cell specificity. A previous study for COVID‐19 [ 5 ] used the gene expression profiles of 1,309 compounds and five cell lines in Connectivity Map (CMap) [ 7 ]. Despite targeting COVID‐19, the previous study used chemically induced profiles obtained from cell lines such as MCF7 (derived from breast tissue) and SKMEL5 (derived from skin tissue). Another previous study for endometriosis [ 6 ] also used CMap and analyzed chemically induced gene expression profiles obtained from cell types different from the disease. In these previous studies, cell types of the disease and compounds were not matched. On the other hand, we aim to consider cell specificities of chemically induced profiles for disease‐associated tissues. Thus, we used AGS (derived from stomach tissue) for gastric cancer, and A375 (derived from skin tissue) for atopic dermatitis.
A unique feature of this study is the enhancement of the searchable number of drug candidate compounds. In a previous study for sorafenib‐resistant hepatocellular carcinoma [ 17 ], chemically induced profiles in HepG2 cells (derived from liver tissue) were used. The previous study considered cell specificity and was consistent with the concept in our approach. However, the number of searchable compounds for which chemically induced profiles were observed in HepG2 cells is limited, and the previous study searched for drug candidate compounds from 3,740 compounds. An advantage of this study is the enhancement of the number of drug candidate compounds using the data completion via a tensor imputation. Therefore, we were able to use 15,420 compounds as drug candidates for each cell line, enabling to explore a broader range of compounds than previous studies.
We chose gastric cancer and atopic dermatitis because these diseases have extremely high medical needs. Gastric cancer is the fifth most commonly occurring cancer and the third leading cause of cancer death. For example, over 1,000,000 people were newly diagnosed with gastric cancer, and 783,000 people died from the disease in 2018. Furthermore, it is reported that the incidence rates have increased greatly in the eastern Asia including Japan [ 26 ]. Therefore, there is a growing need for drug development. Atopic dermatitis is a chronic inflammatory skin disease, which significantly affects patients’ quality of life. Low satisfaction with medical treatment for atopic dermatitis in 10 countries was reported [ 27 ]. While topical corticosteroids have been used for the treatment for atopic dermatitis, alternative medicines are desired because of the side effects by long‐term use of the drugs [ 31 ].
The real‐world applications in practice would be to use single‐cell RNA sequencing (scRNA‐seq) data in addition to bulk RNA data. scRNA‐seq data in disease states are often used to investigate the heterogeneity of cells in diseased tissues, and previous studies have identified the associations between diseases and cell subpopulations [ 32 , 33 ]. Furthermore, the analysis of scRNA‐seq data for drug responses has been receiving much attention [ 34 , 35 ]. By using scRNA‐seq data, our proposed method is expected to explore therapeutic compounds considering not only cell specificity but also cell heterogeneity. For example, in neurodegenerative diseases, the understanding of cell heterogeneity including cell subpopulation‐specific gene expression changes is growing, which reveals the mechanisms in the disease onset and progression [ 36 ]. Cell heterogeneity‐based approaches may open new avenues for finding new therapeutic drugs for such diseases. Therefore, the combination of our proposed method and scRNA‐seq data would be a promising research direction in future.
In this study, we demonstrated the potential of a computational method using cell‐specific gene expression data. However, the proposed method has several limitations. First, there is room for improvement of several hyperparameters in each prediction process, such as the number of seed compounds. Automatic parameter optimization would improve the usability of the method. Second, we evaluated the usefulness of the methods from computational viewpoints, such as reproducibility with structural and gene expression similarity, and visualized the gene expression patterns. Experimental validation would be an important future work. Finally, the prediction results may rely on the data completion method of chemically induced profiles because the original gene expression data have many missing values. To our knowledge, the state‐of‐the‐art completion method is used in this study [ 23 ]. However, more accurate completion methods could be used in the future. Therefore, selecting an appropriate method for data completion in practice is imperative.
Introduction
The development of new drugs is extremely difficult and expensive. Typically, it takes 10–15 years and billions of dollars to bring a new drug to the market [ 1 , 2 ], which is one of the reasons for the low success rate in drug discovery. Among the US Food and Drug Administration drugs registered from 2018 to 2022, several small compounds declined, accounting for a large proportion of the drugs [ 3 ]. The enormous cost and prolonged duration required for drug discovery provide a strong incentive to seek efficient drug development and computational methods. However, most computational methods have a chemistry‐oriented approach, and the corresponding algorithms are mainly based on the chemical structures of the drugs.
Recently, the use of omics data, including gene expression profiles, has received significant attention in drug discovery. Gene expression profiles represent the ratio of mRNA expression levels obtained under two experimental conditions and allow us to understand comprehensive cellular responses, which cannot be obtained from chemical structures alone. For example, gene expression profiles that compared a state in which a compound is exposed to a cell with that in which only a solvent is added to the same cell (termed chemically induced profiles) were used to elucidate the mechanism of action of the compound [ 4 ].
Omics‐based drug searches and designs are based on the correlations between chemically induced profiles and disease‐specific gene expression profiles. The inverse correlation method identifies compounds that counteract disease‐specific gene expression patterns; it was established based on the assumption that the ideal drug restores the gene expression pattern from a diseased state to a normal state. Many compounds have been identified as new drug candidates for various diseases, such as COVID‐19 [ 5 ], endometriosis [ 6 ], neurodegenerative diseases [ 7 , 8 ], inflammatory bowel diseases [ 9 , 10 ], chronic inflammatory skin diseases [ 11 , 12 – 13 ], and diverse cancer types [ 14 , 15 , 16 , 17 , 18 , 19 – 20 ]. Furthermore, a deep learning‐based structural generator using the inverse correlation method was proposed [ 21 ]. However, most previous studies have a critical limitation that the cell specificity of chemically induced profiles for cells relevant to given diseases has not been considered. Although gene expression responses after drug exposure have been reported to be highly cell‐specific [ 22 ], most studies have used chemically induced profiles obtained from specific cell lines, such as MCF7 or those without distinguishing between cell types or tissues. This is mainly because most elements in chemically induced gene expression data are missing or unobserved for a range of cells. It is a problem that cell types of target diseases are not matched with those of chemical exposures, because gene expression values in chemically induced profiles are highly dependent on cell types. Therefore, it is essential to match cell types of the disease and compounds in omics‐based drug discovery.
In this study, we proposed a novel computational method for drug search and design for various diseases using cell‐specific correlations between drugs and diseases. The missing value issue in chemically induced profiles was resolved by data completion using a tensor decomposition algorithm, which allowed us to characterize the cell‐specific gene expression patterns for cells relevant to given diseases. We then applied the proposed method to search for drug candidates and generate new chemical structures of compounds for gastric cancer and atopic dermatitis. Thus, the usefulness of our drug search was demonstrated by evaluating the chemical structure diversity of detected compounds and the involvement in diseases at the pathway levels. Additionally, the usefulness of our drug design was demonstrated by evaluating the newly generated compounds in terms of the reproducibility of for the approved drugs. Taken together, the proposed method is expected to be useful in omics‐based drug discovery.
Coi Statement
The authors declare no conflicts of interest.
Supplementary Material
Supplementary Material
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.