Predicting substrates for orphan Solute Carrier Proteins using multi- omics datasets

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-15

This study developed a predictive algorithm using multi-omics data to identify substrates and drug interactions for orphan Solute Carrier Proteins.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-15 · read from full text

The paper develops a bioinformatic algorithm to de-orphan solute carrier (SLC) proteins by correlating SLC gene expression across large cancer cell line panels (CCLE2019, NCI-60, and CCL180) with intracellular metabolite concentrations measured by LC-MS. Using updated substrate annotations, the authors show that their correlation-based approach recovers known SLC–substrate pairs with higher mean sensitivity/specificity than simulated random SLC–metabolite pairs, and that the predictive correlations are concordant across metabolomics and transcriptomics datasets generated with different methods. The study further reports that incorporating CRISPR loss-of-function screen data and metabolic pathway adjacency data improves performance, and that combining SLC expression with drug sensitivity profiles enables predictions of new SLC–drug interactions, though the approach depends on metabolite coverage and cell-line context. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Abstract Solute carriers (SLC) are integral membrane proteins responsible for transporting a wide variety of metabolites, signaling molecules and drugs across cellular membranes. Despite key roles in metabolism, signaling and pharmacology, around one third of SLC proteins are ‘orphans’ whose substrates are unknown. Experimental determination of SLC substrates is technically challenging given the wide range of possible physiological candidates. Here, we develop a predictive algorithm to identify correlations between SLC expression levels and intracellular metabolite concentrations by leveraging existing cancer multi-omics datasets. Our predictions recovered known SLC-substrate pairs with high sensitivity and specificity compared to simulated random pairs. CRISPR loss-of-function screen data and metabolic pathway adjacency data further improved the performance of our algorithm. In parallel, we combined drug sensitivity data with SLC expression profiles to predict new SLC-drug interactions. Together, we provide a novel bioinformatic pipeline to predict new substrate predictions for SLCs, offering new opportunities to de-orphanise SLCs with important implications for understanding their roles in health and disease.
Full text 139,919 characters · extracted from preprint-html · click to expand
Predicting substrates for orphan Solute Carrier Proteins using multi- omics datasets | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Predicting substrates for orphan Solute Carrier Proteins using multi- omics datasets Y. Zhang, S. Newstead, P. Sarkies This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4713269/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 11 Feb, 2025 Read the published version in BMC Genomics → Version 1 posted 13 You are reading this latest preprint version Abstract Solute carriers (SLC) are integral membrane proteins responsible for transporting a wide variety of metabolites, signaling molecules and drugs across cellular membranes. Despite key roles in metabolism, signaling and pharmacology, around one third of SLC proteins are ‘orphans’ whose substrates are unknown. Experimental determination of SLC substrates is technically challenging given the wide range of possible physiological candidates. Here, we develop a predictive algorithm to identify correlations between SLC expression levels and intracellular metabolite concentrations by leveraging existing cancer multi-omics datasets. Our predictions recovered known SLC-substrate pairs with high sensitivity and specificity compared to simulated random pairs. CRISPR loss-of-function screen data and metabolic pathway adjacency data further improved the performance of our algorithm. In parallel, we combined drug sensitivity data with SLC expression profiles to predict new SLC-drug interactions. Together, we provide a novel bioinformatic pipeline to predict new substrate predictions for SLCs, offering new opportunities to de-orphanise SLCs with important implications for understanding their roles in health and disease. Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Introduction Solute Carriers (SLCs) represent the second largest family of membrane proteins in the human genome after G-protein Coupled Receptors (GPCRs). According to the latest classification database a total of 456 protein-coding genes are classified as a SLC (Ferrada & Superti-Furga, 2022 ). SLC proteins are localised throughout the cell and regulate the flux of many different classes of small molecules including ions, sugars, amino acids, peptides, vitamins and nucleotides across the cell and organelle membranes (Hediger et al., 2013 ; Meixner et al., 2020 ). SLCs are alpha-helical integral membrane proteins and operate under the influence of ion or metabolite gradients to transport molecules across membranes using an alternating access cycle. Several versions of the alternating access cycle have evolved, however, they all share the same fundamental process: an outward open state, wherein a central binding site is available to the non-cytoplasmic side of the membrane; an occluded state, where ions and/or metabolites are trapped inside the transporter; and an inward-facing state, where the transporter has opened its binding site to the cytoplasm. The breadth of substrate specificities and subcellular localization of SLCs gives them critical roles in regulating cellular metabolism (Song et al., 2020 ), energy production (Aulakh et al., 2022 ; Kunji et al., 2016 ; Majd et al., 2018 ), signal transduction (Nimmanon et al., 2017 ), and the maintenance of physical characteristics such as cell volume (Okada, 2004 ). Inherited mutations in SLCs have been linked to at least 100 monogenic disorders (Lin et al., 2015 ). Beyond their physiological substrates, SLCs also transport drug molecules and are thus of crucial importance in pharmacokinetics, drug sensitivity and therapy outcomes (Koepsell & Endou, 2004 ; Mikkaichi et al., 2004 ; Winter et al., 2014 ). The key roles played by SLCs in regulating cellular metabolism and their potential to control the efficacy of drug treatments mean that characterisation of both physiological and non-physiological substrates of SLCs is of great importance. However, many SLCs have no substrate annotated and are described as ‘Orphan Transporters’. A recent survey reported that around 28% of SLCs had no experimentally determined substrate (Meixner et al., 2020 ). The key challenge to de-orphanising SLCs is the inability to predict substrates based on sequence or structure alone, as many transporters within the same family recognise chemically diverse molecules (Ferrada & Superti-Furga, 2022 ). For example, members of the SLC13 family share 40% − 50% amino acid identity, but separate into two different groups based on function: whilst SLC13A1 and SLC13A4 transport sulfates, SLC13A2, SLC13A3 and SLC13A5 recognise carboxylates (Bergeron et al., 2013 ). Therefore, it is currently not possible to accurately infer SLC substrates from amino acid sequence, and experimental determination of substrates is required. However, screening enough substrates to cover all the possibilities is not practical. Thus, methods that could produce predictions for the likely substrate of any given SLC would be extremely valuable in narrowing down the list of substrates to test in functional assays as well as suggesting insights into cellular functions of uncharacterised SLCs. Here, we describe a new method to use existing multi-omics high throughput datasets to predict SLC substrates. We demonstrate that our method carries predictive power in recovering known SLC-substrate pairs when applied to multiple major cell line panels. We use our method to generate predictions for orphan SLCs. In parallel, we develop our method to produce new predictions for the effects of SLC expression on sensitivity to drugs. We hope that our predictions act to generate new hypotheses for the substrates, cellular roles, and therapeutic implications of uncharacterised SLC proteins. Results Correlation analysis between metabolomics and expression datasets successfully predicts SLC substrates We set out to de-orphanise SLC proteins by investigating the potential effects their expression might have on cellular metabolite concentrations. We reasoned that the transporter activity of SLCs might result in correlations between SLC expression level and the intracellular concentrations of their corresponding substrates (Figure 1A). We used a major cancer cell panel profiling 225 metabolites with liquid chromatography-mass spectrometry (LC-MS) across almost a thousand cancer cell lines from the Cancer Cell Line Encyclopedia 2019 (CCLE2019) (Li et al., 2019) . We selected a list of SLC and SLC-like genes (S1 Table) from previous curations (Gyimesi & Hediger, 2022; Meixner et al., 2020). For each SLC or SLC-like gene in the list, we related their normalised transcript levels to metabolite concentration in 913 cell lines, and calculated the Z-score of the absolute values of the correlation coefficients (Spearman’s ρ) for each metabolite to account for varying degrees of correlation strength across different metabolites (Methods). Upon inspection of this set we observed many cases where expression of an SLC correlated most strongly with its known substrate. For example, SLC6A6, a Na + /Cl - -dependent β-alanine and taurine transporter (Ramamoorthy et al., 1994), correlated most strongly to β-alanine and taurine, whilst SLC6A8, a Na + /Cl - -dependent creatine transporter (Skelton et al., 2011), correlated most strongly to creatine (Figure 1B). Notably, expression of SLC6A8 also strongly correlated with two other metabolites, phosphocreatine and creatinine, which are direct derivatives of creatine (Taegtmeyer & Ingwall, 2013; Wyss & Kaddurah-Daouk, 2000). These examples indicated that the expression and metabolite level variation across cancer cell lines might be more generally predictive of the functions of SLCs. To test this, we first expanded the correlation analysis to two other major cell line panels, NCI-60 (Shoemaker, 2006) and CCL180 (Cherkaoui et al., 2022) (Table S2-S4). These three datasets were generated using different methodologies to measure metabolite levels and gene expression; nevertheless, our analysis demonstrated significant concordance in SLC-metabolite correlations between them, suggesting that our method generated robust predictions (Figure 1C). We next investigated how well our correlation method was able to capture the known substrates of SLCs. We updated and expanded the database of SLC annotations based on a previous report (Meixner et al., 2020), selected a list of known substrates that can be found in the metabolomics we used (Table S5; see also Table S1 for references). Since the annotated names for the same metabolites was often different across the databases, we manually examined the metabolomic annotations in the three metabolomics we used to ensure consistent nomenclature across them (Table S6). We tested what fraction of the known substrates were recovered by our correlation analysis and compared this to a random expectation generated by shuffling the SLCs and substrates (“simulated random pairs”). All three datasets had a mean Z-score normalised Spearman’s ρ of known SLC-substrate pairs (Table S5; see also Table S1 for references) significantly higher than that of simulated random pairs (Figure 1D), indicating that correlation analysis was able to successfully predict known interactions. In order to use our method to generate prediction for novel SLC-substrate pairs, we sought to define a cutoff for the correlation strength indicative of a strong prediction (Methods). To do this we systematically varied the normalized Spearman’s ρ and attempted to maximise the fractional difference between true positive (fraction of the set of known pairs above the cutoff) and false negatives (fraction of the set of simulated random pairs above the cutoff). We defined this threshold separately for each of the three datasets (Figure S1A-C). To unify the correlations from the three datasets we assigned each metabolite/SLC pair with a score such that a higher score corresponded to a stronger prediction. To assign a score for a particular metabolite/SLC pair we first normalized the absolute value of the correlation coefficient to give a Z-score. We then compared the Z-score to the Z-scores of the correlation coefficients for all the experimentally determined metabolite/SLC pairs. To give an example of how the score is calculated, the correlation between SLC6A8 and creatine has a Z-score of 3.08 in the NCI60 panel, 8.89 in the CCLE2019 panel and is not represented in the CCL180 panel because creatine was not measured. 3.08 is within the top 10% of known metabolite/SLC pairs in the NCI60, and is therefore assigned a score of 10; 8.89 is in fact the best correlation within the known set for the CCLE and so is assigned a score of 11. These scores are multiplied by 3 to give a total score of 63 for the creatine/SLC6A8 pair. When compared against the mean confidence score of the known SLC-substrate pairs our score had good predictive power compared with the simulated random sets (Figure 1E) and worked across a range of confidence score cutoffs (Figure 1F). Taken together, these results indicate that correlation analysis is able to provide an accurate indication of SLC-substrate pairs and thus could be a potential method to predict new substrates for orphan SLCs. Data from gene dependency screens improves SLC substrate predictions We next considered whether data from genome-wide gene dependency screens (Bock et al., 2022) could be incorporated into our predictions of SLC substrates. We reasoned that cell growth may be dependent on a specific SLC if the cells grow more slowly after the expression of that SLC is depleted, due to loss of that particular metabolite(s) (Figure 2A). Metabolites whose concentrations are significantly different between dependent and non-dependent cell lines might therefore be candidate substrates for the SLC in question. We used the CRISPR-Cas9 dependency screen (Tsherniak et al., 2017) recording cell growth in over a thousand cell lines upon CRISPR knockdown, of which 625 cell lines overlap with CCLE2019 metabolomics profiling. To infer the dependency of cell lines on specific genes we used the gene effect score as defined previously (Meyers et al., 2017). Genes annotated with negative gene effect scores indicated that cells exhibit reduced proliferation upon their deletion compared to normal cells. For each SLC gene in the curated list, we ranked the cell lines based on their corresponding gene effect scores, excluding positive scores where growth was improved by loss of the SLC. We computed a p -value following multiple test corrections for the difference in each metabolite between cell lines with the top 20% of negative scores (more dependent) and bottom 20% of negative scores (less dependent) (Table S7). Across a range of p -value cutoffs the fraction of significant pairs recovered from the known set is higher than simulated random pairs. The fractional difference is maximised when the p -value cutoff for significance is set to 0.077 (Figure S1D). For p -values smaller than the cutoff, a confidence score is assigned based on its position within a similarly calculated decile-based quantile of known pairs, scaled to 0.1 (Figure S1E). Thus, incorporating the CRISPR-Cas9 dependency screen as an additional data source improved the prediction performance, as the confidence score difference between known pairs and simulated random pairs increased by 7 % (Figure 2C). Inclusion of adjacent metabolites improves substrate prediction for SLCs Previous results consolidated that the correlation analysis and CRISPR-Cas9 dependency carry predictive power towards recovering known SLC-substrate pairs (references?). However, substrates might easily dissipate into downstream derivatives, leading to poor correlation and reduced predictive power. Furthermore, the prediction of substrates will be reinforced if the SLC correlates with the derivatives of the substrates as well. To address these ideas, we created a metabolite adjacency matrix (Table S8) from annotated KEGG metabolic pathways. This was done by extracting the number of conversion steps required for one metabolite node to reach another, with each unit of adjacency representing a conversion edge between two metabolite nodes. We reasoned that the expression of the SLC that transports the substrate molecule may correlate with its proximal derivatives, while for its distant derivatives, the distribution of correlation strength will be more random and thus more similar to metabolites that are not related to the substrate of the SLC (Figure 3A). To validate this hypothesis, we generated adjacency tables containing the derivatives that represent different steps of conversion away from the original SLC-substrate table, from proximal to distant. For each SLC-derivative pair in the table, we measured the similarity between the correlation of SLC-derivative pairs and the original SLC-substrate pairs by calculating their Spearman’s ρ difference. Subsequently, these differences were compared against the control tables containing randomly sampled non-adjacent metabolites. Non-adjacent metabolites cannot be linked to the substrate node via any continuous path. We performed one-tailed Wilcoxon tests to compare the Spearman’s ρ differences between the adjacent tables and non-adjacent controls, testing the null hypothesis that the differences in adjacent tables are not smaller than non-adjacent controls. Additionally, for each non-adjacent control in the previous comparison, 100 non-adjacent controls were compared with in one-to-one manner to ensure robustness. Our results demonstrate that proximal derivatives (those requiring fewer conversion steps) showed greater similarity to the original substrates compared to distant derivatives (Figure 3B) and that the difference between derivatives and non-adjacent molecules decreased with increasing distance from the substrate (Figure S2). We next determined the optimal correlation coefficient and the increase in confidence score to be added to maximise the difference between known SLC-metabolite pairs and simulated random pairs(Figure S1F-I). This improved the fractional recovery of known SLC-metabolite pairs by 90% (Figure 3D), indicating that metabolite adjacency information bolstered the accuracy of our prediction algorithm. Predicted substrates for Orphan SLCs Our method thus confirms that true SLC-substrate pairs tend to appear in a higher position compared to simulated random pairs when pairs for each SLC were ranked according to their confidence scores, measured with median rank (Figure 4A). We further demonstrated that the fractional difference peaked if we only considered predictions ranked ≤ 178, with the fraction of true positives reaching 50% (Figure 4B). However, in order to generate a number of predictions for orphan SLCs that could be reasonably tested experimentally, we sought to reduce the number of predictions further. We reasoned that we could improve predictive power for a smaller number of possible substrates by simultaneously identifying over-represented metabolite pathways within the set. We curated a list of 623 metabolites across the three metabolomics datasets that could be linked to 57 metabolic pathways (Table S9). Using the known SLC-metabolite pairs we showed that 20 metabolite predictions were optimal to successfully predict enriched metabolic pathways containing the known substrate (Figure 4C). On this basis we used our prediction algorithm to create a list of substrate predictions with high confidence scores for 128 orphan SLCs (Table S10). We identified many predictions that are in line with experimental data. For example, we found strong associations between the orphan SLC CLN3 and several glycerol phosphate related metabolites (e.g. phosphatidylcholine, alpha-glycerophosphocholine, alpha-glycerophosphate, glycerylphosphorylethanolamine), agreeing with recent research indicating that CLN3 mutant in zebrafish leads to glycerophosphodiesters (GPDs) accumulation in early development (Heins-Marroquin et al., 2024) . We predicted that MTCH1 could be associated with metabolites involved in glutathione synthesis (glycine, glutathione, glutamate, pyroglutamate, NADPH), which aligns with the recent observation that MTCH1-deficiency correlates with NAD+ depletion in mitochondria (Wang et al., 2023). Moreover, our results converge with a previous attempt to predict SLC substrate predictions that used sequence information (Meixner et al., 2020). In this publication, SLC25A45, SLC22A25 and SLC35E2B were all predicted to have nucleobase-containing substrates, and our algorithm also predicted a variety of nucleobases as substrates for these transporters (Table S10). Together, our predictions could be used to generate plausible hypotheses for novel SLC substrates, which can be used to narrow down sebsets of metabolites for downstream experimental verification and leading to faster de-orphanisation. Leveraging drug repurposing panels provides new predicted SLC-drug interactions Solute carriers are known to play an important role in determining drug pharmacokinetics, safety and efficacy profiles (Alam et al., 2023) . A key goal of the recently established International Transporter Consortium is to identify key transporters involved in drug transport and highlight potential issues around adverse drug-drug interactions involving transporters during clinical trials. Therefore, in parallel to the prediction of physiologically relevant substrates, we investigated whether interrogation of omics datasets could be used to identify drug molecules that are substrates for specific SLC proteins. We reasoned that expression of SLCs might affect drug efficacy, thus alter the shape of the dose-response curve reporting the relationship between viability and drug concentration. For example, when considering anticancer drugs, if cell death is improved or attenuated with higher SLC expression levels, one possible indication is that the drug is a substrate for transport by the SLC in question (Figure 5A). We investigated our hypothesis using the cancer repurposing screen profiling 1448 active drugs against 578 cancer cell lines across 8 doses (Corsello et al., 2020) . The cancer cell lines screened were ranked based on each SLC’s transcript levels in CCLE2019, with the highest and lowest 20% marked with “high expression” and “low expression”, respectively. We fitted non-linear regression models into the annotated screen data, and compared the curves between high expression and low expression cell lines. Our analysis captured the difference with accuracy, as it revealed consistency with previously validated results. For example, SLC35F2 expression sensitized cells to the drug YM-155, a known substrate imported by this SLC (Winter et al., 2014). SLC19A2 encodes a plasma membrane thiamine transporter (Dutta et al., 1999), but thiamine uptake is not a dose-dependent factor impacting cell viability (Figure 5B). To test if the analysis shows systematic predictive power in the drug repurposing screen, we selected a list of known transport activity of drug molecule by SLCs (Table S11), and used this to benchmark our predictions. Known pairs showed a higher difference between cell lines marked with high and low expression levels in a dosage dependent manner, compared to simulated random pairs (Figure 5C), with optimal p -value cutoff maximising the fractional difference at 0.17 (Figure S1H). To remove the general impact of drug properties, we calculated a drug-specific significance threshold. For every drug, we randomly picked 20% of cell lines and separated these into high and low expression for 100 times and compared their model predictions. To filter out insignificant pairs we used a drug-specific significant threshold two standard deviations away from the mean of log-transformed p -value (Table S12). Subsequently, drug predictions were listed as we ranked pairs with absolute mean difference and dose-dependent strength, and removed any drug targeting for specific mutations. We then selected the top 50 predictions (Table S13). The SLC-drug pair with the best prediction statistics was an experimentally validated transport activity of YM-155 by SLC35F2 (Winter et al., 2014). Our algorithm also predicted previously unknown links; for example, we predicted an interaction between the orphan SLC Patched Domain Containing 4 (PTCHD4) and the drug molecule idasanutlin, which acts as a small molecule antagonist of p53 activity suppressor Mouse double minute 2 homolog (MDM2) (Figure 5B; Ding et al. , 2013). We also noticed a group of SLCs (SLC3A1, SLC7A7, SLC16A4, SLC23A1, SLC37A1, SLC37A2, SLC41A2, NPC1L1, CLN3) that interact with the small molecule inhibitor RITA, which leads to induction of cell apoptosis by (re)activating wild-type or mutant p53 (Wiegering et al., 2017). SLC3A1 associates with attenuation while the other associate with sensitisation of drug killing effect (Figure S3A). Importantly, our predictions worked across a range of cell lines (Figure S3B), demonstrating…. In summary, our work provided a possible route to predict SLC-drug interactions in parallel to physiological substrate determination, aiding the process of exploring SLC as a therapeutic target reservoir or alerting drug discovery teams to potential downstream issues with cell toxicity or adverse impacts on drug pharmacokinetics. Discussion Deorphanisation of SLCs has been a large collective effort in the community (César-Razquin et al., 2015 ; Superti-Furga et al., 2020 ). However, experimental substrate determination is hindered by the technical difficulties in expressing and purifying functional membrane proteins and the huge range of potential compounds that could be tested even if the SLC is isolated for functional study. Therefore, accurate prediction of SLC substrates would be an important development for the field. Here we developed a new method to predict SLC substrates, and demonstrated we can recover known interactions, indicating its potential for deorphanisation of SLCs. Below, we discuss how the algorithm compares to previous methods, its strengths and limitations, and prospects for future improvement. An obvious approach to attempt to predict SLC substrates is to utilise structures of transporters with known substrates to generate a set of rules that could be subsequently used to predict substrates for orphan SLCs. This approach was attempted by Meixner et al. ( 2020 ), who trained a machine learning algorithm with systematic SLC substrate annotations and structural features, such as sequence and topological domains. Probabilities were produced for 115 orphan SLCs against 18 selected substrate terms. Over the subsequent 3 years, substrates have been experimentally defined for 28 of these orphan SLCs, offering the opportunity to evaluate this method (Table S14). Prediction of the substrates of 4 SLCs aligned with experimental results (SLC39A11, TMEM165, SLC16A6, SLC16A17). For example, TMEM165, now characterised as a lysosomal Ca 2+ importer (Zajac et al., 2024 ), was predicted to have high probability in transporting divalent metal cations. SLC6A17, demonstrated to transport glutamine in mice synaptic vesicles (Jia et al., 2023 ), was predicted to transport L-amino acids. However, for the remainder of SLCs where substrates were determined, the algorithm either did not provide any predictions (16/28), or the predictions produced do not match experimental determination completely (8/28). For example, ANKH was predicted to have high probability of transporting metal ions, but turned out to mediate ATP and citrate export (Szeri et al., 2020 ). This evaluation showed that the prediction based on structural features has limitations. Our method disregarded SLC’s structural features entirely, and focused on the how varying levels of SLC expression affects metabolite levels. Our approach is completely agnostic about known biochemical properties, which means that it has the potential to predict surprising or unexpected interactions, even if an experimentally defined substrate has been annotated. This could be important as many substrates are defined experimentally using a limited range of compounds and using in vitro assays, therefore may not capture the full spectrum of activities for an SLC in vivo. For example, inositol was found to have great correlation with recently characterised facilitative taurine transporter SLC16A6 (Higuchi et al., 2022 ) in all three datasets, but not with other SLC16 family members, potentially indicating of that this may be an additional substrate of SLC16A6. The biochemical naivety of our model also presents limitations. As we rely on metabolomics data sets, our method is limited to metabolites measured and prevalent in these experiments. Most notably, ion concentrations are not measured, therefore we cannot predict ion transporter activity, an obvious drawback since many SLCs transport small ions either exclusively or as counterions to drive transport of another metabolite (Pizzagalli et al., 2021 ). Moreover, many of the correlations with metabolites may reflect secondary effects downstream of the primary substrate which is either not measured or itself rapidly metabolised to other substrates. Our inclusion of adjacency information (Figure X) may go some way towards addressing this; nevertheless it means that predicted substrates should be taken as an indication of the group of possible substrates rather than a clearly defined single substrate. Thus, at present we would suggest that our method offers an opportunity to generate predictions which would still have to be validated with experiments in vitro with purified proteins. Our predictions thus present a resource to the community to expedite experimental substrate identification. In addition to identifying physiological substrates for SLCs, we also used omics data to identify potential drugs that might be substrates. A previous, experimental, approach to this, screened the impact of knocking out specific SLCs on 60 representative cytotoxic drugs, highlighting the broad role of SLCs in drug efficacy (Girardi et al., 2020 ). Of the 201 prominent associations (47 drugs, 101 SLCs) they reported, 39 drugs and 97 SLCs are also included in our analysis. Two-thirds (26/39) of these drugs were found to have associated SLCs predicted in our study, with five associated SLCs ranked highly in our predictions: SLC1A4 & Triptolide (3rd out of 92), SLC19A1 & Methotrexate (6th out of 72), SLC2A1 & Idarubicin (6th out of 58), SLC15A1 & 6-Mercaptopurine (7th out of 61), and SLC12A4 & 5-Azacitidine (11th out of 86). The remaining 13 drug associations were eliminated as they fell outside the range of drug-specific thresholds; however, half of these would have associated SLCs found in the top 10% of predictions for each drug, indicating concordance and predictive power. Our analysis could be a good complement to this previous screen by including a much larger dataset of cell lines and types (469 cell lines compared to 1 in the previous study). Additionally, we examined the effect of individual SLC expression, which might be masked by severe knockout phenotypes. By disregarding structural considerations in our analysis, we allow for the emergence of unexpected drug associations. However, this approach might lead to predicted drug associations not directly related to SLC transport of the drug itself, such as the metabolic environment of the cell and downstream events of SLC activity. Nevertheless, our analysis could aid in characterizing unknown SLC-drug interactions. The potential for polymorphisms in SLCs within human populations has been demonstrated to be a promising angle for personalized medicine(Giacomini et al., 2022 ); our drug interaction results potentially extend this to include differences in expression as a method to predict the sensitivity of specific tumors to particular drugs in personalized medicine. Methods Data acquisition for CCLE2019 dataset CCLE2019 RNA-Seq and metabolomics data were downloaded from the DepMap Portal (depmap.org, CCLE 2019 omics) as read counts (file “CCLE_RNAseq_genes_counts_20180929.gct.gz”) and mean concentration levels (file “CCLE_metabolomics_20190502.csv”). Data acquisition for NCI60 dataset NCI60 RNA-Seq data was derived from alignment and normalisation as performed previously (Perez & Sarkies, 2023 ). NCI-60 metabolomics data were downloaded from the NCI DTP Data Portal (wiki.nci.nih.gov) as mean concentration levels (file “WEB_DATA_METABOLON.ZIP”). Data acquisition for CCL180 dataset CCL180 metabolomics data were downloaded from the ETH Research Data Collection ( https://doi.org/10.3929/ethz-b-000511784 ) as concentration levels and annotations (file “primary analysis (metabolomics)”). Data acquisition for CRISPR-Cas9 dependency screen CRISPR-Cas9 dependency screen dataset was downloaded from DepMap portal (depmap.org, DepMap Public 23Q2 omics) as gene effect scores (file “CRISPRGeneEffect.csv”). Data acquisition for drug repurposing screen Drug repurposing screen dataset was downloaded from Cancer Dependency Map Portal ( https://depmap.org/repurposing ), including cell line annotation (file “secondary-screen-cell-line-info.csv”), treatment metadata (“secondary-screen-replicate-collapsed-treatment-info.csv”) and viability log-fold (“secondary-screen-replicate-collapsed-logfold-change.csv”). Correlating SLC expression levels to metabolites Prior to all analysis, RNA-Seq read counts were normalised with Median Ratio Normalisation (MRN) by ‘DESeq2’ package in R to account for gene expression difference across different tissue types and cancer cell lines. Normalisations were applied to both CCLE 2019 and NCI60 raw counts across all cell lines. The data was first converted to a DESeqDataSet (dds) object using ‘DESeqDataSetFromMatrix()’ function, and the sum of gene reads in each cell line was calculated and filtered if lower than 10. The resulting dds object was normalised by applying ‘estimateSizeFactors()’ function, and the normalised pseudocounts were extracted by ‘counts()’ function with argument ‘normalized = TRUE’. All subsequent analyses used the resulting normalised pseudocounts. Correlation analysis was applied between CCLE 2019 pseudocounts and CCLE 2019 metabolomics (“CCLE2019”), CCLE 2019 pseudocounts and CCL180 metabolomics (“CCL180”), NCI-60 pseudocounts and NCI-60 metabolomics (“NCI60”). Spearman’s correlations were computed across mutually overlapping cell lines between pseudocounts and metabolite levels using the ‘cor.test()’ function in R with argument ‘method = “spearman”’. Resultingcorrelation p -values were adjusted for each gene using Benjamini-Hochberg Procedure using the ‘p.adjust()’ function with ‘method = “BH”’. Correlation coefficients (ρ) might not be able to present correlation strength accurately across dataset due to the change of correlation distributions. Therefore, for metabolite \(\:a\) and SLC \(\:z\) , the normalised ρ coefficient \(\:\stackrel{\sim}{{\rho\:}_{(a,z)}}\) is computed from the following formula: $$\:\stackrel{\sim}{{\rho\:}_{(a,z)}}=\:\frac{\left|{\rho\:}_{(a,z)}\right|-\:\stackrel{-}{\left|{\rho\:}_{\left(a\right)}\right|}}{{\sigma\:}_{\left|{\rho\:}_{\left(a\right)}\right|}}$$ where each \(\:{\rho\:}_{(a,z)}\) will be transformed to represent the number of absolute standard deviation ( \(\:{\sigma\:}_{\left|{\rho\:}_{\left(a\right)}\right|}\) ) away from the absolute mean ( \(\:\stackrel{-}{\left|{\rho\:}_{\left(a\right)}\right|}\) ) based on the correlation distribution of metabolite a to every SLC, and thus represent only correlation strength of the pair. Concordance assessments Between datasets, only mutually overlapping SLC and metabolite terms were assessed. The resulting raw ρ values for each overlapping SLC and metabolite were taken and correlated using the ‘cor.test()’ function in R with argument ‘method = “spearman”’. Benchmarking Known pair tables (Table S5, S11) were manually extracted from the SLC ontology annotation (Table S1 ) for overlapping SLC and metabolite terms. Metabolites or drug molecules appearing in the known pair tables were shuffled and randomly assigned to SLCsthat are not known to transport it, while keeping the SLC column unchanged, generating 100 simulated random pair tables. Mean statistics (e.g. normalised ρ, adjusted p , confidence score) were calculated per table to measure predictive power. Across threshold of discovery for corresponding statistics, fractional difference was calculated as the difference between true positive fraction (fraction left in known pair tables) and false positive fraction (fraction left in simulated random pair tables) to measure the validity of threshold chosen. To benchmark the drug repurposing screen predictions, the difference between “high expression” and “low expression” cell lines for each SLC were compared to the difference between 100 sets of the same numbers of randomly selected cell lines. A value that two standard deviation away from the mean -log 10 ( p -value) and absolute mean difference was taken per drug as the filtering threshold. Metabolite adjacency Metabolite adjacency table (Table S8) was generated from human KEGG pathway by ‘MetaboSignal’ package in R (Rodriguez-Martinez et al., 2016 ). Specifically, all human metabolic pathways were extracted and subsetted using ‘MS_getPathIds()’ function with argument ‘organism_code = “hsa”’. Reaction network was built based on metabolic pathways using ‘MS_reactionNetwork()’ function. Subsequently, node distance was calculated using ‘MS_nodeBW()’ function with argument ‘node = “out”’ and ‘normalized = TRUE’. Metabolite prediction algorithm For every SLC-metabolite pair \(\:i\) , the prediction algorithm computes a confidence score by evaluating its correlation statistics across the three datasets ( \(\:{A}_{i}\) ), resulted adjusted p -value of annotated CRISPR-Cas9 screen ( \(\:{B}_{i}\) ), and adjacent metabolites ( \(\:{C}_{i}\) ). Threshold of discovery and score reward per discovery is specified in Figure S1 , validified with fractional difference between true positive and false positive. For every SLC-metabolite pair, if value implicated in \(\:{A}_{i}\) , \(\:{B}_{i}\) , or \(\:{C}_{i}\) is smaller than its respective threshold of discovery or the metabolite is not measured in the dataset not exist, a confidence score of 0 was assigned. Otherwise, the confidence score will be measured with respect to the decile-based percentile of known SLC-substrate tables, such that a decile of 10 corresponds to the top 10% and a decile of 1 the lowest 10%; correlations greater than or equal to the top correlation (i.e. greater than decile 10) were assigned a score of 11. The adjusted p -value of \(\:{B}_{i}\) was measured in log-transformed format. The decile number was multiplied by 3 for \(\:{A}_{i}\) , 1.1 for \(\:{B}_{i}\) , and 0.9 for \(\:{C}_{i}\) . Drug prediction algorithm For each pair between drug molecule \(\:a\) and SLC \(\:z\) , the dose responses were only calculated for cell lines annotated based on the highest and lowest 20% of SLC \(\:z\) expression. Local polynomial regression models were fitted to the two expression types using ‘loess()’ function, viability against log transformed dose (-3.21 to 1). 421 data points, ranging from − 3.21 to 1 and separated by 0.01, predicted by model using ‘predict()’ function was generated to capture the shape information of fitted models. Curve shapes were compared with a paired t-test. The resulted p -value and mean differences were recorded. Declarations Code availability All the code for processing data, as well as generating the figures and tables in the manuscripts, is available as supplemental material and will be uploaded to GitHub upon publication. Ethics approval and consent to participate No specific ethics approval was required for this study Consent for publication No specific consent for publication is required for this study. Funding No specific funding was obtained for this project Authors' contributions Conceptualization: PS, SN Formal analysis: YZ, PS Data analysis and interpretation: YZ, PS Manuscript first draft: YZ, PS Manuscript review and editing: YZ, SN, PS All authors read and approved the final manuscript. Acknowledgements We thank the members of the EpiEvo group and other colleagues at the Department of Biochemistry, University of Oxford for helpful discussion and comments on the project. We thank Dr Marcos Francisco Perez for aligning and normalising raw NCI-60 transcriptomics data. We also thank Dr Louise Fets (London Institute of Medical Sciences) for helpful discussions. References Alam, S., Doherty, E., Ortega-Prieto, P., Arizanova, J., & Fets, L. (2023). Membrane transporters in cell physiology, cancer metabolism and drug response. Disease Models & Mechanisms , 16 (11). https://doi.org/10.1242/dmm.050404 Aulakh, S. K., Varma, S. J., & Ralser, M. (2022). Metal ion availability and homeostasis as drivers of metabolic evolution and enzyme function. Current Opinion in Genetics & Development , 77 , 101987. https://doi.org/10.1016/j.gde.2022.101987 Bergeron, M. J., Clémençon, B., Hediger, M. A., & Markovich, D. (2013). SLC13 family of Na+-coupled di- and tri-carboxylate/sulfate transporters. Molecular Aspects of Medicine , 34 (2–3), 299–312. https://doi.org/10.1016/j.mam.2012.12.001 Bock, C., Datlinger, P., Chardon, F., Coelho, M. A., Dong, M. B., Lawson, K. A., Lu, T., Maroc, L., Norman, T. M., Song, B., Stanley, G., Chen, S., Garnett, M., Li, W., Moffat, J., Qi, L. S., Shapiro, R. S., Shendure, J., Weissman, J. S., & Zhuang, X. (2022). High-content CRISPR screening. In Nature Reviews Methods Primers (Vol. 2, Issue 1). Springer Nature. https://doi.org/10.1038/s43586-021-00093-4 César-Razquin, A., Snijder, B., Frappier-Brinton, T., Isserlin, R., Gyimesi, G., Bai, X., Reithmeier, R. A., Hepworth, D., Hediger, M. A., Edwards, A. M., & Superti-Furga, G. (2015). A Call for Systematic Research on Solute Carriers. Cell , 162 (3), 478–487. https://doi.org/10.1016/j.cell.2015.07.022 Cherkaoui, S., Durot, S., Bradley, J., Critchlow, S. E., Dubuis, S., Masiero, M., Wegmann, R., Snijder, B., Othman, A., Bendtsen, C., & Zamboni, N. (2022). A functional analysis of 180 cancer cell lines reveals conserved intrinsic metabolic programs. Molecular Systems Biology , 18 (11). https://doi.org/10.15252/msb.202211033 Corsello, S. M., Nagari, R. T., Spangler, R. D., Rossen, J., Kocak, M., Bryan, J. G., Humeidi, R., Peck, D., Wu, X., Tang, A. A., Wang, V. M., Bender, S. A., Lemire, E., Narayan, R., Montgomery, P., Ben-David, U., Garvie, C. W., Chen, Y., Rees, M. G., … Golub, T. R. (2020). Discovering the anticancer potential of non-oncology drugs by systematic viability profiling. Nature Cancer , 1 , 235–248. https://doi.org/10.1038/s43018-019-0018-6 Dutta, B., Huang, W., Molero, M., Kekuda, R., Leibach, F. H., Devoe, L. D., Ganapathy, V., & Prasad, P. D. (1999). Cloning of the Human Thiamine Transporter, a Member of the Folate Transporter Family *. Journal of Biological Chemistry , 274 (45), 31925–31929. https://doi.org/10.1074/jbc.274.45.31925 Ferrada, E., & Superti-Furga, G. (2022). A structure and evolutionary-based classification of solute carriers. IScience , 25 (10), 105096. https://doi.org/10.1016/j.isci.2022.105096 Giacomini, K. M., Yee, S. W., Koleske, M. L., Zou, L., Matsson, P., Chen, E. C., Kroetz, D. L., Miller, M. A., Gozalpour, E., & Chu, X. (2022). New and Emerging Research on Solute Carrier and ATP Binding Cassette Transporters in Drug Discovery and Development: Outlook From the International Transporter Consortium. Clinical Pharmacology & Therapeutics , 112 (3), 540–561. https://doi.org/10.1002/CPT.2627 Girardi, E., César-Razquin, A., Lindinger, S., Papakostas, K., Konecka, J., Hemmerich, J., Kickinger, S., Kartnig, F., Gürtl, B., Klavins, K., Sedlyarov, V., Ingles-Prieto, A., Fiume, G., Koren, A., Lardeau, C.-H., Kumaran Kandasamy, R., Kubicek, S., Ecker, G. F., & Superti-Furga, G. (2020). A widespread role for SLC transmembrane transporters in resistance to cytotoxic drugs. Nature Chemical Biology , 16 (4), 469–478. https://doi.org/10.1038/s41589-020-0483-3 Gyimesi, G., & Hediger, M. A. (2022). Systematic in silico discovery of novel solute carrier-like proteins from proteomes. PLOS ONE , 17 (7), e0271062–e0271062. https://doi.org/10.1371/journal.pone.0271062 Hediger, M. A., Clémençon, B., Burrier, R. E., & Bruford, E. A. (2013). The ABCs of membrane transporters in health and disease (SLC series): Introduction. Molecular Aspects of Medicine , 34 (2–3), 95–107. https://doi.org/10.1016/J.MAM.2012.12.009 Heins-Marroquin, U., Singh, R. R., Perathoner, S., Gavotto, F., Ruiz, C. M., Patraskaki, M., Gomez-Giro, G., Borgmann, F. K., Meyer, M., Carpentier, A., Warmoes, M. O., Jäger, C., Mittelbronn, M., Schwamborn, J. C., Cordero-Maldonado, M. L., Crawford, A. D., Schymanski, E. L., & Linster, C. L. (2024). CLN3 deficiency leads to neurological and metabolic perturbations during early development. Life Science Alliance , 7 (3), e202302057–e202302057. https://doi.org/10.26508/lsa.202302057 Higuchi, K., Sugiyama, K., Tomabechi, R., Kishimoto, H., & Inoue, K. (2022). Mammalian monocarboxylate transporter 7 (MCT7/Slc16a6) is a novel facilitative taurine transporter. The Journal of Biological Chemistry , 298 (4), 101800. https://doi.org/10.1016/j.jbc.2022.101800 Jia, X., Zhu, J., Bian, X., Liu, S., Yu, S., Liang, W., Jiang, L., Mao, R., Zhang, W., & Rao, Y. (2023). Importance of glutamine in synaptic vesicles revealed by functional studies of SLC6A17 and its mutations pathogenic for intellectual disability. ELife , 12 , RP86972. https://doi.org/10.7554/eLife.86972 Koepsell, H., & Endou, H. (2004). The SLC22 drug transporter family. Pflügers Archiv: European Journal of Physiology , 447 (5), 666–676. https://doi.org/10.1007/s00424-003-1089-9 Kunji, E. R. S., Aleksandrova, A., King, M. S., Majd, H., Ashton, V. L., Cerson, E., Springett, R., Kibalchenko, M., Tavoulari, S., Crichton, P. G., & Ruprecht, J. J. (2016). The transport mechanism of the mitochondrial ADP/ATP carrier. Biochimica et Biophysica Acta (BBA) - Molecular Cell Research , 1863 (10), 2379–2393. https://doi.org/10.1016/j.bbamcr.2016.03.015 Li, H., Ning, S., Ghandi, M., Kryukov, G. V, Gopal, S., Deik, A., Souza, A., Pierce, K., Keskula, P., Hernandez, D., Ann, J., Shkoza, D., Apfel, V., Zou, Y., Vazquez, F., Barretina, J., Pagliarini, R. A., Galli, G. G., Root, D. E., … Sellers, W. R. (2019). The landscape of cancer cell line metabolism. Nature Medicine , 25 (5), 850–860. https://doi.org/10.1038/s41591-019-0404-8 Lin, L., Yee, S. W., Kim, R. B., & Giacomini, K. M. (2015). SLC transporters as therapeutic targets: Emerging opportunities. In Nature Reviews Drug Discovery (Vol. 14, Issue 8, pp. 543–560). Nature Publishing Group. https://doi.org/10.1038/nrd4626 Majd, H., King, M. S., Smith, A. C., & Kunji, E. R. S. (2018). Pathogenic mutations of the human mitochondrial citrate carrier SLC25A1 lead to impaired citrate export required for lipid, dolichol, ubiquinone and sterol synthesis. Biochimica et Biophysica Acta (BBA) - Bioenergetics , 1859 (1), 1–7. https://doi.org/10.1016/j.bbabio.2017.10.002 Meixner, E., Goldmann, U., Sedlyarov, V., Scorzoni, S., Rebsamen, M., Girardi, E., & Superti‐Furga, G. (2020). A substrate‐based ontology for human solute carriers. Molecular Systems Biology , 16 (7). https://doi.org/10.15252/msb.20209652 Meyers, R. M., Bryan, J. G., McFarland, J. M., Weir, B. A., Sizemore, A. E., Xu, H., Dharia, N. V, Montgomery, P. G., Cowley, G. S., Pantel, S., Goodale, A., Lee, Y., Ali, L. D., Jiang, G., Lubonja, R., Harrington, W. F., Strickland, M., Wu, T., Hawes, D. C., … Tsherniak, A. (2017). Computational correction of copy number effect improves specificity of CRISPR–Cas9 essentiality screens in cancer cells. Nature Genetics , 49 (12), 1779–1784. https://doi.org/10.1038/ng.3984 Mikkaichi, T., Suzuki, T., Onogawa, T., Tanemoto, M., Mizutamari, H., Okada, M., Chaki, T., Masuda, S., Tokui, T., Eto, N., Abe, M., Satoh, F., Unno, M., Hishinuma, T., Inui, K. I., Ito, S., Goto, J., & Abe, T. (2004). Isolation and characterization of a digoxin transporter and its rat homologue expressed in the kidney. Proc. Natl. Acad. Sci. U. S. A. , 101 (10), 3569–3574. https://doi.org/10.1073/pnas.0304987101 Nimmanon, T., Ziliotto, S., Morris, S., Flanagan, L., & Taylor, K. M. (2017). Phosphorylation of zinc channel ZIP7 drives MAPK, PI3K and mTOR growth and proliferation signalling. Metallomics , 9 (5), 471–481. https://doi.org/10.1039/c6mt00286b Okada, Y. (2004). Ion Channels and Transporters Involved in Cell Volume Regulation and Sensor Mechanisms. Cell Biochemistry and Biophysics , 41 (2), 233–258. https://doi.org/10.1385/cbb:41:2:233 Perez, M. F., & Sarkies, P. (2023). Histone methyltransferase activity affects metabolism in human cells independently of transcriptional regulation. PLoS Biology , 21 (10 October). https://doi.org/10.1371/journal.pbio.3002354 Pizzagalli, M. D., Bensimon, A., & Superti-Furga, G. (2021). A guide to plasma membrane solute carrier proteins. In FEBS Journal (Vol. 288, Issue 9, pp. 2784–2835). John Wiley and Sons Inc. https://doi.org/10.1111/febs.15531 Ramamoorthy, S., Leibach, F. H., Mahesh, V. B., Han, H., Yang-Feng, T., Blakely, R. D., & Ganapathy, V. (1994). Functional characterization and chromosomal localization of a cloned taurine transporter from human placenta. Biochemical Journal , 300 (3), 893–900. https://doi.org/10.1042/bj3000893 Rodriguez-Martinez, A., Ayala, R., Posma, J. M., Neves, A. L., Gauguier, D., Nicholson, J. K., & Dumas, M.-E. (2016). MetaboSignal: a network-based approach for topological analysis of metabotype regulationviametabolic and signaling pathways. Bioinformatics , 33 (5), btw697. https://doi.org/10.1093/bioinformatics/btw697 Shoemaker, R. H. (2006). The NCI60 human tumour cell line anticancer drug screen. Nature Reviews Cancer , 6 (10), 813–823. https://doi.org/10.1038/nrc1951 Skelton, M. R., Schaefer, T. L., Graham, D. L., deGrauw, T. J., Clark, J. F., Williams, M. T., & Vorhees, C. V. (2011). Creatine Transporter (CrT; Slc6a8) Knockout Mice as a Model of Human CrT Deficiency. PLoS ONE , 6 (1), e16187. https://doi.org/10.1371/journal.pone.0016187 Song, W., Li, D., Tao, L., Luo, Q., & Chen, L. (2020). Solute carrier transporters: the metabolic gatekeepers of immune cells. Acta Pharmaceutica Sinica B , 10 (1), 61–78. https://doi.org/10.1016/j.apsb.2019.12.006 Superti-Furga, G., Lackner, D., Wiedmer, T., Ingles-Prieto, A., Barbosa, B., Girardi, E., Goldmann, U., Gürtl, B., Klavins, K., Klimek, C., Lindinger, S., Liñeiro-Retes, E., Müller, A. C., Onstein, S., Redinger, G., Reil, D., Sedlyarov, V., Wolf, G., Crawford, M., … Steppan, C. M. (2020). The RESOLUTE consortium: unlocking SLC transporters for drug discovery. Nature Reviews Drug Discovery , 19 (7), 429–430. https://doi.org/10.1038/d41573-020-00056-6 Szeri, F., Lundkvist, S., Donnelly, S., Engelke, U., Rhee, K., Williams, C. J., Sundberg, J. P., Wevers, R. A., Tomlinson, R. E., Jansen, R. S., & Wetering, K. (2020). The membrane protein ANKH is crucial for bone mechanical performance by mediating cellular export of citrate and ATP. PLOS Genetics , 16 (7), e1008884–e1008884. https://doi.org/10.1371/journal.pgen.1008884 Taegtmeyer, H., & Ingwall, J. S. (2013). Creatine—A Dispensable Metabolite? Circulation Research , 112 (6), 878–880. https://doi.org/10.1161/circresaha.113.300974 Tsherniak, A., Vazquez, F., Montgomery, P. G., Weir, B. A., Kryukov, G., Cowley, G. S., Gill, S., Harrington, W. F., Pantel, S., Krill-Burger, J. M., Meyers, R. M., Ali, L., Goodale, A., Lee, Y., Jiang, G., Hsiao, J., Gerath, W. F. J., Howell, S., Merkel, E., … Hahn, W. C. (2017). Defining a Cancer Dependency Map. Cell , 170 (3), 564-576.e16. https://doi.org/10.1016/j.cell.2017.06.010 Wang, X., Ji, Y., Qi, J., Zhou, S., Wan, S., Fan, C., Gu, Z., An, P., Luo, Y., & Luo, J. (2023). Mitochondrial carrier 1 (MTCH1) governs ferroptosis by triggering the FoxO1-GPX4 axis-mediated retrograde signaling in cervical cancer cells. Cell Death and Disease , 14 (8). https://doi.org/10.1038/s41419-023-06033-2 Wiegering, A., Matthes, N., Mühling, B., Koospal, M., Quenzer, A., Peter, S., Germer, C.-T., Linnebacher, M., & Otto, C. (2017). Reactivating p53 and Inducing Tumor Apoptosis (RITA) Enhances the Response of RITA-Sensitive Colorectal Cancer Cells to Chemotherapeutic Agents 5-Fluorouracil and Oxaliplatin. Neoplasia , 19 (4), 301–309. https://doi.org/10.1016/j.neo.2017.01.007 Winter, G. E., Radic, B., Mayor-Ruiz, C., Blomen, V. A., Trefzer, C., Kandasamy, R. K., Huber, K. V. M., Gridling, M., Chen, D., Klampfl, T., Kralovics, R., Kubicek, S., Fernandez-Capetillo, O., Brummelkamp, T. R., & Superti-Furga, G. (2014). The solute carrier SLC35F2 enables YM155-mediated DNA damage toxicity. Nature Chemical Biology , 10 (9), 768–773. https://doi.org/10.1038/nchembio.1590 Wyss, M., & Kaddurah-Daouk, R. (2000). Creatine and Creatinine Metabolism. Physiological Reviews , 80 (3), 1107–1213. https://doi.org/10.1152/physrev.2000.80.3.1107 Zajac, M., Mukherjee, S., Anees, P., Oettinger, D., Henn, K., Srikumar, J., Zou, J., Saminathan, A., & Krishnan, Y. (2024). A mechanism of lysosomal calcium entry. Science Advances , 10 (7). https://doi.org/10.1126/sciadv.adk2317 Additional Declarations No competing interests reported. Supplementary Files CodeForplottingtheresults.zip CodeForGeneratingResults.zip Supplementarytable.zip Cite Share Download PDF Status: Published Journal Publication published 11 Feb, 2025 Read the published version in BMC Genomics → Version 1 posted Editorial decision: Revision requested 23 Oct, 2024 Reviews received at journal 22 Oct, 2024 Reviews received at journal 24 Sep, 2024 Reviews received at journal 12 Sep, 2024 Reviewers agreed at journal 05 Sep, 2024 Reviewers agreed at journal 05 Sep, 2024 Reviewers agreed at journal 31 Aug, 2024 Reviewers agreed at journal 29 Aug, 2024 Reviewers invited by journal 22 Aug, 2024 Editor invited by journal 12 Jul, 2024 Editor assigned by journal 11 Jul, 2024 Submission checks completed at journal 11 Jul, 2024 First submitted to journal 09 Jul, 2024 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4713269","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":333653509,"identity":"520a0ae8-6e36-4333-bc94-4cd7f13df341","order_by":0,"name":"Y. Zhang","email":"","orcid":"","institution":"University of Oxford","correspondingAuthor":false,"prefix":"","firstName":"Y.","middleName":"","lastName":"Zhang","suffix":""},{"id":333653510,"identity":"bc927063-9979-447d-9983-ef1815d419d6","order_by":1,"name":"S. Newstead","email":"","orcid":"","institution":"University of Oxford","correspondingAuthor":false,"prefix":"","firstName":"S.","middleName":"","lastName":"Newstead","suffix":""},{"id":333653511,"identity":"13c3bee3-fee0-42c2-974c-326e7fca0f46","order_by":2,"name":"P. Sarkies","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABG0lEQVRIie2RsWrDMBCGTxjs5YxXmxTyCjIeTKDBr5LgoUvp0sVDMTIGZzFk1RD6DIVCZhlBJ2f3EGhNoLu3eqsTSNOAQujWQR9IiEMfd78EoNH8a5AwgETgcPqpGeqrx/JBqf+mDIsUAq4qobWsdgmk47CUWds/b28cayOgK2DsMAyoQpmU0vBrkP5qk+WBvf5Er3yYEV6DzwUGM4VCm9j0GAjCHVKMyFoibZAamAB5AQyESnnfWT2DNBqURd+vjgqF6KLSGObwVMac21kBNjt1me8V1WCTMg48RmXMscpH+LbPck8rXrsxl+ajKn5oVW3HknTK8a7tvp62kWPV/kdX3E6Xi/zVVQ32az8xRHAvfqSqs0aj0WjO+QbmrFvUFxiuhgAAAABJRU5ErkJggg==","orcid":"","institution":"University of Oxford","correspondingAuthor":true,"prefix":"","firstName":"P.","middleName":"","lastName":"Sarkies","suffix":""}],"badges":[],"createdAt":"2024-07-09 15:44:31","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4713269/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4713269/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1186/s12864-025-11330-5","type":"published","date":"2025-02-11T15:57:55+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":61750842,"identity":"e7cb3919-547a-4ef9-896e-d8167585cb61","added_by":"auto","created_at":"2024-08-05 07:41:36","extension":"jpeg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":541142,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eCorrelating SLC transcript levels to substrate concentrations reveals known substrates of SLC.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eA)\u003c/strong\u003e Schematic representation shows the principle of correlation analysis. SLC expression levels are hypothesised to correlate with the intracellular concentration of their corresponding substrates. Blue, SLC exporting substrate; red, SLC importing substrate. Figure created with elements from BioRender.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eB)\u003c/strong\u003e Scatter plot shows normalised Spearman’s ρ and adjusted \u003cem\u003ep\u003c/em\u003e-value for 225 metabolites correlating with the expression level of SLC6A6 and SLC6A8 across 913 cell lines from CCLE. Key substrates are labelled.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eC)\u003c/strong\u003e Mutual concordance of SLC-metabolite correlation outcomes from different datasets, each processed with the pipeline specified in \u003cstrong\u003eMethods\u003c/strong\u003e. Nodes, datasets, edge, concordance parameter measured in Spearman’s ρ. CCLE2019 and NCI60, \u003cem\u003ep\u003c/em\u003e = 4.30 x 10\u003csup\u003e-293\u003c/sup\u003e; CCLE2019 and CCL180, \u003cem\u003ep\u003c/em\u003e = 1.50 x 10\u003csup\u003e-221\u003c/sup\u003e; NCI60 and CCL180, \u003cem\u003ep\u003c/em\u003e = 2.38 x 10\u003csup\u003e-24\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eD)\u003c/strong\u003e Violin plot shows the mean normalised Spearman’s ρ in known SLC-metabolite set compared to 100 simulated random sets in each dataset. Colored violin and dots, mean normalised Spearman’s ρ of 100 simulated random sets; Red dashed line, mean normalised Spearman’s ρ of known sets; \u003cem\u003ep\u003c/em\u003e-value was derived from one-tailed t-test against the null hypothesis that average normalised Spearman’s ρ of known set is not higher than its corresponding 100 simulated random sets.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eE)\u003c/strong\u003e Boxplot shows the mean confidence score of known SLC-metabolite sets compared to that of 100 simulated random sets. \u003cem\u003ep\u003c/em\u003e-value was derived from one-tailed t-test against the null hypothesis that mean normalised Spearman’s ρ of known set is not higher than its corresponding 100 simulated random sets.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eF)\u003c/strong\u003e Across confidence cutoff selected, fraction of interactions recovered (having confidence score better than cutoff) in known set compared to simulated random set. Red dashed line, recovered fraction value in known set; grey violin distribution, recovered fraction value in simulated random set; \u003cem\u003ep\u003c/em\u003e-value was derived from one-tailed Wilcoxon test against the null hypothesis that fractions recovered in known set is not higher than its corresponding 100 simulated random sets.\u003c/p\u003e","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-4713269/v1/65c8044ae31f128b28e58af4.jpeg"},{"id":61750840,"identity":"2c7726b4-cd5b-42bc-bfe1-4abafd3e64d5","added_by":"auto","created_at":"2024-08-05 07:41:35","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":396801,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eCRISPR loss-of-function screen provides alternative source for SLC prediction.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eA)\u003c/strong\u003e Schematic representation shows the principle of CRISPR loss-of-function data analysis to predict SLC substrates. Figure created with elements from BioRender.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eB)\u003c/strong\u003e Violin plot shows the recovered fractions in known sets compared to simulated random sets across different cutoff selected for resulting adjusted \u003cem\u003ep\u003c/em\u003e-values. The \u003cem\u003ep\u003c/em\u003e-value comparing the fraction distribution between known and simulated random was derived from one-tailed Wilcoxon test against the null hypothesis that fractions recovered in known set is not higher than its corresponding 100 simulated random sets.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eC)\u003c/strong\u003e Boxplot showing the effect of incorporation of CRISPR loss-of-function screen into prediction algorithm on the confidence score difference between true positive and false positives. The \u003cem\u003ep\u003c/em\u003e-value comparing the confidence score difference was derived from one-tailed Wilcoxon test against the null hypothesis that confidence score difference calculated using both correlation analysis and CRISPR loss-of-function screen (“Four”) is not better than only using correlation analysis (“Three”).\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-4713269/v1/22b6a80f2c04ced4f422c707.png"},{"id":61751565,"identity":"bfd02f66-9d14-4463-b1e3-5e9061e43183","added_by":"auto","created_at":"2024-08-05 07:49:35","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":312648,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eInclusion of metabolite adjacency to the prediction substantially differentiates known interaction from simulated random interaction.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eA)\u003c/strong\u003e Schematic representation shows the relationship between adjacent metabolites (\u003cem\u003ei.e.\u003c/em\u003e, derivatives of the substrate) and non-adjacent metabolites (\u003cem\u003ei.e.\u003c/em\u003e, not derivatives of the substrate). Pipeline details specified in \u003cstrong\u003eMethod\u003c/strong\u003e. Figure created with elements from BioRender.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eB)\u003c/strong\u003e Boxplot shows the general similarity of Spearman’s ρ between derivative across unit of adjacency. Red dashed line, Spearman’s ρ differences between adjacent derivatives and substrates are compared to non-adjacent controls across unit of adjacency. Grey boxes, Spearman’s ρ differences between non-adjacent controls are compared to another 100 non-adjacent controls across unit of adjacency.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eC)\u003c/strong\u003e Boxplot shows the effect of the incorporation of metabolite adjacency into the prediction algorithm (“Five” vs “Four”). The \u003cem\u003ep\u003c/em\u003e-value comparing the confidence score difference was derived from one-tailed Wilcoxon test against the null hypothesis that confidence score difference calculated adding metabolite adjacency (“Five”) is not better than only using correlation analysis (“Four”).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eD)\u003c/strong\u003e Violin plot shows the fractional difference between true positive and false positives calculated with the inclusion of metabolite adjacency compared to the algorithm without including metabolite adjacency over a range of confidence score cutoffs. The \u003cem\u003ep\u003c/em\u003e-value comparing the confidence score difference was derived from one-tailed Wilcoxon test against the null hypothesis that confidence score difference calculated adding metabolite adjacency (“Five”) is not better than only using correlation analysis (“Four”).\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-4713269/v1/05bf1b623e39e11a6abdb5e1.png"},{"id":61750843,"identity":"dae98a5d-59a7-4205-beb7-e98d01ea94a3","added_by":"auto","created_at":"2024-08-05 07:41:36","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":147198,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eTop-ranking predictions of known transport activity .\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eA)\u003c/strong\u003e Boxplot shows the median rank of known SLC-substrate pairs compared to simulated random pairs. The \u003cem\u003ep\u003c/em\u003e-value comparing the median rank was derived from one-tailed Wilcoxon test against the null hypothesis that the median rank in the known set is not closer to top than in simulated random set.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eB)\u003c/strong\u003e Fractional difference between true positive (TP%) and false positive (FP%) when only rank above the given value is considered as predicted. Black line, median fractional difference; grey, range of fractional difference.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eC)\u003c/strong\u003e Point plot shows the number of SLC with substrate converged with the pathway enriched from metabolites ranked above the given rank value in the prediction list.\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-4713269/v1/682114ee6fcd3ca1363c52e0.png"},{"id":61751566,"identity":"69bcc7ec-945b-4696-9884-28c462c535d3","added_by":"auto","created_at":"2024-08-05 07:49:36","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":409827,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eCombining drug sensitivities and SLC expression profile reveals valuable associations between SLC and drug efficacy.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eA)\u003c/strong\u003e Schematic representation shows the principle of leveraging the drug dose response curve to predict SLC-drug associations. If a cytotoxic drug compound is transported by a SLC, expression of the SLC may either enhance or attenuate its killing effect. Figure created with elements from BioRender.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eB)\u003c/strong\u003e Non-linear regression shows an example of how SLC35F2 expression affects the killing efficacy of the known substrate drug YM-155 (Left); an example of how SLC19A2 expression does not affect sensitivity to thiamine (Middle); a prediction example of idasanutlin efficacy might associate with PTCHD4, which is not an identified link. Statistics were computed based on paired t-test of model prediction capturing the shape of fitted regression.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eC)\u003c/strong\u003e Boxplots show comparison between simulated random interactions and known interactions either through fitting non-linear regression models (Model, left) or performing correlation analysis between dosage and cellular viability (Dose effect, right); \u003cem\u003ep\u003c/em\u003e-values derive from one-tailed Wilcoxon test (left) and one-tailed t-test (right) against a null distribution that statistics of the known set is not better than those of the simulated random.\u003c/p\u003e","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-4713269/v1/543647f08b2bb2234d46483a.png"},{"id":76488251,"identity":"232ddb76-7850-476a-9189-0fd0abd42095","added_by":"auto","created_at":"2025-02-17 16:13:38","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2606624,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4713269/v1/7a686221-861a-4143-9db8-cb4b54305287.pdf"},{"id":61750846,"identity":"42e9fd16-3c35-491a-856a-d257c809fa89","added_by":"auto","created_at":"2024-08-05 07:41:36","extension":"zip","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":24699493,"visible":true,"origin":"","legend":"","description":"","filename":"CodeForplottingtheresults.zip","url":"https://assets-eu.researchsquare.com/files/rs-4713269/v1/4ff0ff768800bfbbce44e815.zip"},{"id":61751564,"identity":"598792cc-5e20-40d8-8832-bac84de377c0","added_by":"auto","created_at":"2024-08-05 07:49:35","extension":"zip","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":125719,"visible":true,"origin":"","legend":"","description":"","filename":"CodeForGeneratingResults.zip","url":"https://assets-eu.researchsquare.com/files/rs-4713269/v1/245d6bcb0606eabbd555b76b.zip"},{"id":61750848,"identity":"bc022841-ee22-4c68-b848-7059cf84bc5b","added_by":"auto","created_at":"2024-08-05 07:41:36","extension":"zip","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":61245108,"visible":true,"origin":"","legend":"","description":"","filename":"Supplementarytable.zip","url":"https://assets-eu.researchsquare.com/files/rs-4713269/v1/21803be0e991f041baba9185.zip"}],"financialInterests":"No competing interests reported.","formattedTitle":"Predicting substrates for orphan Solute Carrier Proteins using multi- omics datasets","fulltext":[{"header":"Introduction","content":"\u003cp\u003eSolute Carriers (SLCs) represent the second largest family of membrane proteins in the human genome after G-protein Coupled Receptors (GPCRs). According to the latest classification database a total of 456 protein-coding genes are classified as a SLC (Ferrada \u0026amp; Superti-Furga, \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e2022\u003c/span\u003e). SLC proteins are localised throughout the cell and regulate the flux of many different classes of small molecules including ions, sugars, amino acids, peptides, vitamins and nucleotides across the cell and organelle membranes (Hediger et al., \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e2013\u003c/span\u003e; Meixner et al., \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e2020\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eSLCs are alpha-helical integral membrane proteins and operate under the influence of ion or metabolite gradients to transport molecules across membranes using an alternating access cycle. Several versions of the alternating access cycle have evolved, however, they all share the same fundamental process: an outward open state, wherein a central binding site is available to the non-cytoplasmic side of the membrane; an occluded state, where ions and/or metabolites are trapped inside the transporter; and an inward-facing state, where the transporter has opened its binding site to the cytoplasm.\u003c/p\u003e \u003cp\u003eThe breadth of substrate specificities and subcellular localization of SLCs gives them critical roles in regulating cellular metabolism (Song et al., \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e2020\u003c/span\u003e), energy production (Aulakh et al., \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2022\u003c/span\u003e; Kunji et al., \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e2016\u003c/span\u003e; Majd et al., \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e2018\u003c/span\u003e), signal transduction (Nimmanon et al., \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e2017\u003c/span\u003e), and the maintenance of physical characteristics such as cell volume (Okada, \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e2004\u003c/span\u003e). Inherited mutations in SLCs have been linked to at least 100 monogenic disorders (Lin et al., \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e2015\u003c/span\u003e). Beyond their physiological substrates, SLCs also transport drug molecules and are thus of crucial importance in pharmacokinetics, drug sensitivity and therapy outcomes (Koepsell \u0026amp; Endou, \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e2004\u003c/span\u003e; Mikkaichi et al., \u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e2004\u003c/span\u003e; Winter et al., \u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e2014\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eThe key roles played by SLCs in regulating cellular metabolism and their potential to control the efficacy of drug treatments mean that characterisation of both physiological and non-physiological substrates of SLCs is of great importance. However, many SLCs have no substrate annotated and are described as \u0026lsquo;Orphan Transporters\u0026rsquo;. A recent survey reported that around 28% of SLCs had no experimentally determined substrate (Meixner et al., \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). The key challenge to de-orphanising SLCs is the inability to predict substrates based on sequence or structure alone, as many transporters within the same family recognise chemically diverse molecules (Ferrada \u0026amp; Superti-Furga, \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e2022\u003c/span\u003e). For example, members of the SLC13 family share 40% \u0026minus;\u0026thinsp;50% amino acid identity, but separate into two different groups based on function: whilst SLC13A1 and SLC13A4 transport sulfates, SLC13A2, SLC13A3 and SLC13A5 recognise carboxylates (Bergeron et al., \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e2013\u003c/span\u003e). Therefore, it is currently not possible to accurately infer SLC substrates from amino acid sequence, and experimental determination of substrates is required. However, screening enough substrates to cover all the possibilities is not practical. Thus, methods that could produce predictions for the likely substrate of any given SLC would be extremely valuable in narrowing down the list of substrates to test in functional assays as well as suggesting insights into cellular functions of uncharacterised SLCs.\u003c/p\u003e \u003cp\u003eHere, we describe a new method to use existing multi-omics high throughput datasets to predict SLC substrates. We demonstrate that our method carries predictive power in recovering known SLC-substrate pairs when applied to multiple major cell line panels. We use our method to generate predictions for orphan SLCs. In parallel, we develop our method to produce new predictions for the effects of SLC expression on sensitivity to drugs. We hope that our predictions act to generate new hypotheses for the substrates, cellular roles, and therapeutic implications of uncharacterised SLC proteins.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003e\u003cstrong\u003eCorrelation analysis between metabolomics and expression datasets successfully predicts SLC substrates\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe set out to de-orphanise SLC proteins by investigating the potential effects their expression might have on cellular metabolite concentrations. We reasoned that the transporter activity of SLCs might result in correlations between SLC expression level and the intracellular concentrations of their corresponding substrates (Figure 1A). We used a major cancer cell panel profiling 225 metabolites with liquid chromatography-mass spectrometry (LC-MS) across almost a thousand cancer cell lines from the Cancer Cell Line Encyclopedia 2019 (CCLE2019) \u003cspan lang=\"EN-GB\"\u003e(Li et al., 2019)\u003c/span\u003e. We selected a list of SLC and SLC-like genes (S1 Table) from previous curations (Gyimesi \u0026amp; Hediger, 2022; Meixner et al., 2020).\u0026nbsp;For each SLC or SLC-like gene in the list, we related their normalised transcript levels to metabolite concentration in 913 cell lines, and calculated the\u0026nbsp;Z-score of the absolute values of the correlation coefficients\u0026nbsp;(Spearman\u0026rsquo;s\u0026nbsp;\u0026rho;)\u0026nbsp;for each metabolite to account for varying degrees of correlation strength across different metabolites (Methods). Upon inspection of this set we observed many cases where expression of an SLC\u0026nbsp;correlated most strongly with its known substrate. For example, SLC6A6, a Na\u003csup\u003e+\u003c/sup\u003e/Cl\u003csup\u003e-\u003c/sup\u003e-dependent \u0026beta;-alanine and taurine transporter (Ramamoorthy et al., 1994), correlated most strongly to \u0026beta;-alanine and taurine, whilst SLC6A8, a Na\u003csup\u003e+\u003c/sup\u003e/Cl\u003csup\u003e-\u003c/sup\u003e-dependent creatine transporter (Skelton et al., 2011), correlated most strongly to creatine (Figure 1B). Notably, expression of SLC6A8 also strongly correlated with two other metabolites, phosphocreatine and creatinine, which are direct derivatives of creatine (Taegtmeyer \u0026amp; Ingwall, 2013; Wyss \u0026amp; Kaddurah-Daouk, 2000).\u003c/p\u003e\n\u003cp\u003eThese examples indicated that the expression and metabolite level variation across cancer cell lines might be more generally predictive of the functions of SLCs. To test this, we first expanded the correlation analysis to two other major cell line panels, NCI-60 (Shoemaker, 2006) and CCL180 (Cherkaoui et al., 2022) (Table S2-S4). These three datasets were generated using different methodologies to measure metabolite levels and gene expression; nevertheless, our analysis demonstrated significant concordance in SLC-metabolite correlations between them, suggesting that our method generated robust predictions (Figure 1C). We next investigated how well our correlation method was able to capture the known substrates of SLCs. We updated and expanded the database of SLC annotations based on a previous report (Meixner et al., 2020), selected a list of known substrates that can be found in the metabolomics we used (Table S5; see also Table S1 for references). \u0026nbsp;Since the annotated names for the same metabolites was often different across the databases, we manually examined the metabolomic annotations in the three metabolomics we used to ensure consistent nomenclature across them (Table S6). We tested what fraction of the known substrates were recovered by our correlation analysis and compared this to a random expectation generated by shuffling the SLCs and substrates (\u0026ldquo;simulated random pairs\u0026rdquo;). All three datasets had a mean Z-score normalised Spearman\u0026rsquo;s \u0026rho; of known SLC-substrate pairs (Table S5; see also Table S1 for references) \u0026nbsp;significantly higher than that of simulated random pairs (Figure 1D), indicating that correlation analysis was able to successfully predict known interactions.\u003c/p\u003e\n\u003cp\u003eIn order to use our method to generate prediction for novel SLC-substrate pairs, we sought to define a cutoff for the correlation strength indicative of a strong prediction (Methods). To do this we systematically varied the normalized Spearman\u0026rsquo;s \u0026rho; and attempted to maximise the fractional difference between true positive (fraction of the set of known pairs above the cutoff) and false negatives (fraction of the set of simulated random pairs above the cutoff). We defined this threshold separately for each of the three datasets (Figure S1A-C). To unify the correlations from the three datasets we assigned each metabolite/SLC pair with a score such that a higher score corresponded to a stronger prediction. To assign a score for a particular metabolite/SLC pair we first normalized the absolute value of the correlation coefficient to give a Z-score. \u0026nbsp;We then compared the Z-score to the Z-scores of the correlation coefficients for all the experimentally determined metabolite/SLC pairs. To give an example of how the score is calculated, the correlation between SLC6A8 and creatine has a Z-score of 3.08 in the NCI60 panel, 8.89 in the CCLE2019 panel and is not represented in the CCL180 panel because creatine was not measured. 3.08 is within the top 10% of known metabolite/SLC pairs in the NCI60, and is therefore assigned a score of 10; 8.89 is in fact the best correlation within the known set for the CCLE and so is assigned a score of 11. These scores are multiplied by 3 to give a total score of 63 for the creatine/SLC6A8 pair. When compared against the mean confidence score of the known SLC-substrate pairs our score had good predictive power compared with the simulated random sets (Figure 1E) and worked across a range of confidence score cutoffs (Figure 1F). Taken together, these results indicate that correlation analysis is able to provide an accurate indication of SLC-substrate pairs and thus could be a potential method to predict new substrates for orphan SLCs.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData from gene dependency screens improves SLC substrate predictions\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe next considered whether data from genome-wide gene dependency screens (Bock et al., 2022) could be incorporated into our predictions of SLC substrates. We reasoned that cell growth may be dependent on a specific SLC if the cells grow more slowly after the expression of that SLC is depleted, due to loss of that particular metabolite(s) (Figure 2A). Metabolites whose concentrations are significantly different between dependent and non-dependent cell lines might therefore be candidate substrates for the SLC in question. We used the CRISPR-Cas9 dependency screen (Tsherniak et al., 2017) recording cell growth in over a thousand cell lines upon CRISPR knockdown, of which 625 cell lines overlap with CCLE2019 metabolomics profiling. To infer the dependency of cell lines on specific genes we used the gene effect score as defined previously \u0026nbsp;(Meyers et al., 2017). Genes annotated with negative gene effect scores indicated that cells exhibit reduced proliferation upon their deletion compared to normal cells. For each SLC gene in the curated list, we ranked the cell lines based on their corresponding gene effect scores, excluding positive scores where growth was improved by loss of the SLC. We computed a \u003cem\u003ep\u003c/em\u003e-value following multiple test corrections for the difference in each metabolite between cell lines with the top 20% of negative scores (more dependent) and bottom 20% of negative scores (less dependent) (Table S7). Across a range of \u003cem\u003ep\u003c/em\u003e-value cutoffs the fraction of significant pairs recovered from the known set is higher than simulated random pairs. The fractional difference is maximised when the \u003cem\u003ep\u003c/em\u003e-value cutoff for significance is set to 0.077 (Figure S1D). For \u003cem\u003ep\u003c/em\u003e-values smaller than the cutoff, a confidence score is assigned based on its position within a similarly calculated decile-based quantile of known pairs, scaled to 0.1 (Figure S1E). Thus, incorporating the CRISPR-Cas9 dependency screen as an additional data source improved the prediction performance, as the confidence score difference between known pairs and simulated random pairs increased by 7 % (Figure 2C).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eInclusion of adjacent metabolites improves substrate prediction for SLCs\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003ePrevious results consolidated that the correlation analysis and CRISPR-Cas9 dependency carry predictive power towards recovering known SLC-substrate pairs (references?). However, substrates might easily dissipate into downstream derivatives, leading to poor correlation and reduced predictive power. Furthermore, the prediction of substrates will be reinforced if the SLC correlates with the derivatives of the substrates as well. To address these ideas, we created a metabolite adjacency matrix (Table S8) from annotated KEGG metabolic pathways. This was done by extracting the number of conversion steps required for one metabolite node to reach another, with each unit of adjacency representing a conversion edge between two metabolite nodes. We reasoned that the expression of the SLC that transports the substrate molecule may correlate with its proximal derivatives, while for its distant derivatives, the distribution of correlation strength will be more random and thus more similar to metabolites that are not related to the substrate of the SLC (Figure 3A). To validate this hypothesis, we generated adjacency tables containing the derivatives that represent different steps of conversion away from the original SLC-substrate table, from proximal to distant. For each SLC-derivative pair in the table, we measured the similarity between the correlation of SLC-derivative pairs and the original SLC-substrate pairs by calculating their Spearman\u0026rsquo;s \u0026rho; difference. Subsequently, these differences were compared against the control tables containing randomly sampled non-adjacent metabolites. Non-adjacent metabolites cannot be linked to the substrate node via any continuous path. We performed one-tailed Wilcoxon tests to compare the Spearman\u0026rsquo;s \u0026rho; differences between the adjacent tables and non-adjacent controls, testing the null hypothesis that the differences in adjacent tables are not smaller than non-adjacent controls. Additionally, for each non-adjacent control in the previous comparison, 100 non-adjacent controls were compared with in one-to-one manner to ensure robustness. Our results demonstrate that proximal derivatives (those requiring fewer conversion steps) showed greater similarity to the original substrates compared to distant derivatives (Figure 3B) and that the difference between derivatives and non-adjacent molecules decreased with increasing distance from the substrate (Figure S2). \u0026nbsp;We next determined the optimal correlation coefficient and the increase in confidence score to be added to maximise the difference between known SLC-metabolite pairs and simulated random pairs(Figure S1F-I). This improved the fractional recovery of known SLC-metabolite pairs by 90% (Figure 3D), indicating that metabolite adjacency information bolstered the accuracy of our prediction algorithm. \u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ePredicted substrates for Orphan SLCs\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eOur method thus confirms that true SLC-substrate pairs tend to appear in a higher position compared to simulated random pairs when pairs for each SLC were ranked according to their confidence scores, measured with median rank (Figure 4A). We further demonstrated that the fractional difference peaked if we only considered predictions ranked \u0026le; 178, with the fraction of true positives reaching 50% (Figure 4B). However, in order to generate a number of predictions for orphan SLCs that could be reasonably tested experimentally, we sought to reduce the number of predictions further. We reasoned that we could improve predictive power for a smaller number of possible substrates by simultaneously identifying over-represented metabolite pathways within the set. We curated a list of 623 metabolites across the three metabolomics datasets that could be linked to 57 metabolic pathways (Table S9). Using the known SLC-metabolite pairs we showed that 20 metabolite predictions were optimal to successfully predict enriched metabolic pathways containing the known substrate (Figure 4C).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eOn this basis we used our prediction algorithm to create a list of substrate predictions with high confidence scores for 128 orphan SLCs (Table S10). We identified many predictions that are in line with experimental data. For example, we found strong associations between the orphan SLC CLN3 and several glycerol phosphate related metabolites (e.g. phosphatidylcholine, alpha-glycerophosphocholine, alpha-glycerophosphate, glycerylphosphorylethanolamine), agreeing with recent research indicating that CLN3 mutant in zebrafish leads to glycerophosphodiesters (GPDs) accumulation in early development \u003cspan lang=\"EN-GB\"\u003e(Heins-Marroquin et al., 2024)\u003c/span\u003e. We predicted that MTCH1 could be associated with metabolites involved in glutathione synthesis (glycine, glutathione, glutamate, pyroglutamate, NADPH),\u0026nbsp;which aligns with the recent observation\u0026nbsp;that MTCH1-deficiency correlates with NAD+ depletion in mitochondria (Wang et al., 2023).\u0026nbsp;Moreover, our results converge with a previous attempt to predict SLC substrate predictions that used sequence information\u0026nbsp;(Meixner et al., 2020). In this publication, SLC25A45, SLC22A25 and SLC35E2B were all predicted to have nucleobase-containing substrates, and our algorithm also predicted a variety of nucleobases as substrates for these transporters (Table S10). Together, our predictions could be used to generate plausible hypotheses for novel SLC substrates, which can be used to narrow down sebsets of metabolites for downstream experimental verification and leading to faster de-orphanisation.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eLeveraging drug repurposing panels provides new predicted SLC-drug interactions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eSolute carriers \u0026nbsp;are known to play an important role in determining drug pharmacokinetics, safety and efficacy profiles \u003cspan lang=\"EN-GB\"\u003e(Alam et al., 2023)\u003c/span\u003e. A key goal of the recently established International Transporter Consortium is to identify key transporters involved in drug transport and highlight potential issues around adverse drug-drug interactions involving transporters during clinical trials. Therefore, in parallel to the prediction of physiologically relevant substrates, we investigated whether interrogation of omics datasets could be used to identify drug molecules that are substrates for specific SLC proteins. We reasoned that expression of SLCs might affect drug efficacy, thus alter the shape of the dose-response curve reporting the relationship between viability and drug concentration. For example, when considering anticancer drugs, if cell death is improved or attenuated with higher SLC expression levels, one possible indication is that the drug is a substrate for transport by the SLC in question (Figure 5A). We investigated our hypothesis using the cancer repurposing screen profiling 1448 active drugs against 578 cancer cell lines across 8 doses \u003cspan lang=\"EN-GB\"\u003e(Corsello et al., 2020)\u003c/span\u003e. The cancer cell lines screened were ranked based on each SLC\u0026rsquo;s transcript levels in CCLE2019, with the highest and lowest 20% marked with \u0026ldquo;high expression\u0026rdquo; and \u0026ldquo;low expression\u0026rdquo;, respectively. We fitted non-linear regression models into the annotated screen data, and compared the curves between high expression and low expression cell lines. Our analysis captured the difference with accuracy, as it revealed consistency with previously validated results. For example, SLC35F2 expression sensitized cells to the drug YM-155, a known substrate imported by this SLC (Winter et al., 2014). SLC19A2 encodes a plasma membrane thiamine transporter (Dutta et al., 1999), but thiamine uptake is not a dose-dependent factor impacting cell viability (Figure 5B).\u003c/p\u003e\n\u003cp\u003eTo test if the analysis shows systematic predictive power in the drug repurposing screen, we selected a list of known transport activity of drug molecule by SLCs (Table S11), and used this to benchmark our predictions. Known pairs showed a higher difference between cell lines marked with high and low expression levels in a dosage dependent manner, compared to simulated random pairs (Figure 5C), with optimal \u003cem\u003ep\u003c/em\u003e-value cutoff maximising the fractional difference at 0.17 (Figure S1H).\u003c/p\u003e\n\u003cp\u003eTo remove the general impact of drug properties, we calculated a drug-specific significance threshold. For every drug, we randomly picked 20% of cell lines and separated these into high and low expression for 100 times and compared their model predictions. To filter out insignificant pairs we used a drug-specific significant threshold two standard deviations away from the mean of log-transformed \u003cem\u003ep\u003c/em\u003e-value (Table S12). Subsequently, drug predictions were listed as we ranked pairs with absolute mean difference and dose-dependent strength, and removed any drug targeting for specific mutations. We then selected the top 50 predictions (Table S13). The SLC-drug pair with the best prediction statistics was an experimentally validated transport activity of YM-155 by SLC35F2 (Winter et al., 2014). Our algorithm also predicted previously unknown links; for example, we predicted an interaction between the orphan SLC Patched Domain Containing 4 (PTCHD4) and the drug molecule idasanutlin, which acts as a small molecule antagonist of p53 activity suppressor Mouse double minute 2 homolog (MDM2) (Figure 5B; Ding \u003cem\u003eet al.\u003c/em\u003e, 2013). We also noticed a group of SLCs (SLC3A1, SLC7A7, SLC16A4, SLC23A1, SLC37A1, SLC37A2, SLC41A2, NPC1L1, CLN3) that interact with the small molecule inhibitor RITA, which leads to induction of cell apoptosis by (re)activating wild-type or mutant p53 (Wiegering et al., 2017). SLC3A1 associates with attenuation while the other associate with sensitisation of drug killing effect (Figure S3A). Importantly, our predictions worked across a range of cell lines (Figure S3B), demonstrating\u0026hellip;. In summary, our work provided a possible route to predict SLC-drug interactions in parallel to physiological substrate determination, aiding the process of exploring SLC as a therapeutic target reservoir or alerting drug discovery teams to potential downstream issues with cell toxicity or adverse impacts on drug pharmacokinetics.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eDeorphanisation of SLCs has been a large collective effort in the community (C\u0026eacute;sar-Razquin et al., \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e2015\u003c/span\u003e; Superti-Furga et al., \u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). However, experimental substrate determination is hindered by the technical difficulties in expressing and purifying functional membrane proteins and the huge range of potential compounds that could be tested even if the SLC is isolated for functional study. Therefore, accurate prediction of SLC substrates would be an important development for the field. Here we developed a new method to predict SLC substrates, and demonstrated we can recover known interactions, indicating its potential for deorphanisation of SLCs. Below, we discuss how the algorithm compares to previous methods, its strengths and limitations, and prospects for future improvement.\u003c/p\u003e \u003cp\u003eAn obvious approach to attempt to predict SLC substrates is to utilise structures of transporters with known substrates to generate a set of rules that could be subsequently used to predict substrates for orphan SLCs. This approach was attempted by Meixner et al. (\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e2020\u003c/span\u003e), who trained a machine learning algorithm with systematic SLC substrate annotations and structural features, such as sequence and topological domains. Probabilities were produced for 115 orphan SLCs against 18 selected substrate terms. Over the subsequent 3 years, substrates have been experimentally defined for 28 of these orphan SLCs, offering the opportunity to evaluate this method (Table S14). Prediction of the substrates of 4 SLCs aligned with experimental results (SLC39A11, TMEM165, SLC16A6, SLC16A17). For example, TMEM165, now characterised as a lysosomal Ca\u003csup\u003e2+\u003c/sup\u003e importer (Zajac et al., \u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e2024\u003c/span\u003e), was predicted to have high probability in transporting divalent metal cations. SLC6A17, demonstrated to transport glutamine in mice synaptic vesicles (Jia et al., \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e2023\u003c/span\u003e), was predicted to transport L-amino acids. However, for the remainder of SLCs where substrates were determined, the algorithm either did not provide any predictions (16/28), or the predictions produced do not match experimental determination completely (8/28). For example, ANKH was predicted to have high probability of transporting metal ions, but turned out to mediate ATP and citrate export (Szeri et al., \u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). This evaluation showed that the prediction based on structural features has limitations.\u003c/p\u003e \u003cp\u003eOur method disregarded SLC\u0026rsquo;s structural features entirely, and focused on the how varying levels of SLC expression affects metabolite levels. Our approach is completely agnostic about known biochemical properties, which means that it has the potential to predict surprising or unexpected interactions, even if an experimentally defined substrate has been annotated. This could be important as many substrates are defined experimentally using a limited range of compounds and using in vitro assays, therefore may not capture the full spectrum of activities for an SLC in vivo. For example, inositol was found to have great correlation with recently characterised facilitative taurine transporter SLC16A6 (Higuchi et al., \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e2022\u003c/span\u003e) in all three datasets, but not with other SLC16 family members, potentially indicating of that this may be an additional substrate of SLC16A6.\u003c/p\u003e \u003cp\u003eThe biochemical naivety of our model also presents limitations. As we rely on metabolomics data sets, our method is limited to metabolites measured and prevalent in these experiments. Most notably, ion concentrations are not measured, therefore we cannot predict ion transporter activity, an obvious drawback since many SLCs transport small ions either exclusively or as counterions to drive transport of another metabolite (Pizzagalli et al., \u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e2021\u003c/span\u003e). Moreover, many of the correlations with metabolites may reflect secondary effects downstream of the primary substrate which is either not measured or itself rapidly metabolised to other substrates. Our inclusion of adjacency information (Figure X) may go some way towards addressing this; nevertheless it means that predicted substrates should be taken as an indication of the group of possible substrates rather than a clearly defined single substrate. Thus, at present we would suggest that our method offers an opportunity to generate predictions which would still have to be validated with experiments in vitro with purified proteins. Our predictions thus present a resource to the community to expedite experimental substrate identification.\u003c/p\u003e \u003cp\u003eIn addition to identifying physiological substrates for SLCs, we also used omics data to identify potential drugs that might be substrates. A previous, experimental, approach to this, screened the impact of knocking out specific SLCs on 60 representative cytotoxic drugs, highlighting the broad role of SLCs in drug efficacy (Girardi et al., \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). Of the 201 prominent associations (47 drugs, 101 SLCs) they reported, 39 drugs and 97 SLCs are also included in our analysis. Two-thirds (26/39) of these drugs were found to have associated SLCs predicted in our study, with five associated SLCs ranked highly in our predictions: SLC1A4 \u0026amp; Triptolide (3rd out of 92), SLC19A1 \u0026amp; Methotrexate (6th out of 72), SLC2A1 \u0026amp; Idarubicin (6th out of 58), SLC15A1 \u0026amp; 6-Mercaptopurine (7th out of 61), and SLC12A4 \u0026amp; 5-Azacitidine (11th out of 86). The remaining 13 drug associations were eliminated as they fell outside the range of drug-specific thresholds; however, half of these would have associated SLCs found in the top 10% of predictions for each drug, indicating concordance and predictive power.\u003c/p\u003e \u003cp\u003eOur analysis could be a good complement to this previous screen by including a much larger dataset of cell lines and types (469 cell lines compared to 1 in the previous study). Additionally, we examined the effect of individual SLC expression, which might be masked by severe knockout phenotypes. By disregarding structural considerations in our analysis, we allow for the emergence of unexpected drug associations. However, this approach might lead to predicted drug associations not directly related to SLC transport of the drug itself, such as the metabolic environment of the cell and downstream events of SLC activity. Nevertheless, our analysis could aid in characterizing unknown SLC-drug interactions. The potential for polymorphisms in SLCs within human populations has been demonstrated to be a promising angle for personalized medicine(Giacomini et al., \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e2022\u003c/span\u003e); our drug interaction results potentially extend this to include differences in expression as a method to predict the sensitivity of specific tumors to particular drugs in personalized medicine.\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003eData acquisition for CCLE2019 dataset\u003c/h2\u003e \u003cp\u003eCCLE2019 RNA-Seq and metabolomics data were downloaded from the DepMap Portal (depmap.org, CCLE 2019 omics) as read counts (file \u0026ldquo;CCLE_RNAseq_genes_counts_20180929.gct.gz\u0026rdquo;) and mean concentration levels (file \u0026ldquo;CCLE_metabolomics_20190502.csv\u0026rdquo;).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eData acquisition for NCI60 dataset\u003c/h2\u003e \u003cp\u003eNCI60 RNA-Seq data was derived from alignment and normalisation as performed previously (Perez \u0026amp; Sarkies, \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). NCI-60 metabolomics data were downloaded from the NCI DTP Data Portal (wiki.nci.nih.gov) as mean concentration levels (file \u0026ldquo;WEB_DATA_METABOLON.ZIP\u0026rdquo;).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eData acquisition for CCL180 dataset\u003c/h2\u003e \u003cp\u003eCCL180 metabolomics data were downloaded from the ETH Research Data Collection (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3929/ethz-b-000511784\u003c/span\u003e\u003cspan address=\"10.3929/ethz-b-000511784\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) as concentration levels and annotations (file \u0026ldquo;primary analysis (metabolomics)\u0026rdquo;).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eData acquisition for CRISPR-Cas9 dependency screen\u003c/h2\u003e \u003cp\u003eCRISPR-Cas9 dependency screen dataset was downloaded from DepMap portal (depmap.org, DepMap Public 23Q2 omics) as gene effect scores (file \u0026ldquo;CRISPRGeneEffect.csv\u0026rdquo;).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003eData acquisition for drug repurposing screen\u003c/h2\u003e \u003cp\u003eDrug repurposing screen dataset was downloaded from Cancer Dependency Map Portal (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://depmap.org/repurposing\u003c/span\u003e\u003cspan address=\"https://depmap.org/repurposing\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e), including cell line annotation (file \u0026ldquo;secondary-screen-cell-line-info.csv\u0026rdquo;), treatment metadata (\u0026ldquo;secondary-screen-replicate-collapsed-treatment-info.csv\u0026rdquo;) and viability log-fold (\u0026ldquo;secondary-screen-replicate-collapsed-logfold-change.csv\u0026rdquo;).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003eCorrelating SLC expression levels to metabolites\u003c/h2\u003e \u003cp\u003ePrior to all analysis, RNA-Seq read counts were normalised with Median Ratio Normalisation (MRN) by \u0026lsquo;DESeq2\u0026rsquo; package in R to account for gene expression difference across different tissue types and cancer cell lines. Normalisations were applied to both CCLE 2019 and NCI60 raw counts across all cell lines. The data was first converted to a DESeqDataSet (dds) object using \u0026lsquo;DESeqDataSetFromMatrix()\u0026rsquo; function, and the sum of gene reads in each cell line was calculated and filtered if lower than 10. The resulting \u003cem\u003edds\u003c/em\u003e object was normalised by applying \u0026lsquo;estimateSizeFactors()\u0026rsquo; function, and the normalised pseudocounts were extracted by \u0026lsquo;counts()\u0026rsquo; function with argument \u0026lsquo;normalized\u0026thinsp;=\u0026thinsp;TRUE\u0026rsquo;. All subsequent analyses used the resulting normalised pseudocounts.\u003c/p\u003e \u003cp\u003eCorrelation analysis was applied between CCLE 2019 pseudocounts and CCLE 2019 metabolomics (\u0026ldquo;CCLE2019\u0026rdquo;), CCLE 2019 pseudocounts and CCL180 metabolomics (\u0026ldquo;CCL180\u0026rdquo;), NCI-60 pseudocounts and NCI-60 metabolomics (\u0026ldquo;NCI60\u0026rdquo;). Spearman\u0026rsquo;s correlations were computed across mutually overlapping cell lines between pseudocounts and metabolite levels using the \u0026lsquo;cor.test()\u0026rsquo; function in R with argument \u0026lsquo;method = \u0026ldquo;spearman\u0026rdquo;\u0026rsquo;.\u003c/p\u003e \u003cp\u003eResultingcorrelation \u003cem\u003ep\u003c/em\u003e-values were adjusted for each gene using Benjamini-Hochberg Procedure using the \u0026lsquo;p.adjust()\u0026rsquo; function with \u0026lsquo;method = \u0026ldquo;BH\u0026rdquo;\u0026rsquo;. Correlation coefficients (ρ) might not be able to present correlation strength accurately across dataset due to the change of correlation distributions. Therefore, for metabolite \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:a\\)\u003c/span\u003e\u003c/span\u003e and SLC \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:z\\)\u003c/span\u003e\u003c/span\u003e, the normalised ρ coefficient \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\stackrel{\\sim}{{\\rho\\:}_{(a,z)}}\\)\u003c/span\u003e\u003c/span\u003eis computed from the following formula:\u003cdiv id=\"Equa\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equa\" name=\"EquationSource\"\u003e\n$$\\:\\stackrel{\\sim}{{\\rho\\:}_{(a,z)}}=\\:\\frac{\\left|{\\rho\\:}_{(a,z)}\\right|-\\:\\stackrel{-}{\\left|{\\rho\\:}_{\\left(a\\right)}\\right|}}{{\\sigma\\:}_{\\left|{\\rho\\:}_{\\left(a\\right)}\\right|}}$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003ewhere each \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\rho\\:}_{(a,z)}\\)\u003c/span\u003e\u003c/span\u003e will be transformed to represent the number of absolute standard deviation (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\sigma\\:}_{\\left|{\\rho\\:}_{\\left(a\\right)}\\right|}\\)\u003c/span\u003e\u003c/span\u003e) away from the absolute mean (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\stackrel{-}{\\left|{\\rho\\:}_{\\left(a\\right)}\\right|}\\)\u003c/span\u003e\u003c/span\u003e) based on the correlation distribution of metabolite \u003cem\u003ea\u003c/em\u003e to every SLC, and thus represent only correlation strength of the pair.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003eConcordance assessments\u003c/h2\u003e \u003cp\u003eBetween datasets, only mutually overlapping SLC and metabolite terms were assessed. The resulting raw ρ values for each overlapping SLC and metabolite were taken and correlated using the \u0026lsquo;cor.test()\u0026rsquo; function in R with argument \u0026lsquo;method = \u0026ldquo;spearman\u0026rdquo;\u0026rsquo;.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003eBenchmarking\u003c/h2\u003e \u003cp\u003eKnown pair tables (Table S5, S11) were manually extracted from the SLC ontology annotation (Table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e) for overlapping SLC and metabolite terms. Metabolites or drug molecules appearing in the known pair tables were shuffled and randomly assigned to SLCsthat are not known to transport it, while keeping the SLC column unchanged, generating 100 simulated random pair tables. Mean statistics (e.g. normalised ρ, adjusted \u003cem\u003ep\u003c/em\u003e, confidence score) were calculated per table to measure predictive power. Across threshold of discovery for corresponding statistics, fractional difference was calculated as the difference between true positive fraction (fraction left in known pair tables) and false positive fraction (fraction left in simulated random pair tables) to measure the validity of threshold chosen.\u003c/p\u003e \u003cp\u003eTo benchmark the drug repurposing screen predictions, the difference between \u0026ldquo;high expression\u0026rdquo; and \u0026ldquo;low expression\u0026rdquo; cell lines for each SLC were compared to the difference between 100 sets of the same numbers of randomly selected cell lines. A value that two standard deviation away from the mean -log\u003csub\u003e10\u003c/sub\u003e(\u003cem\u003ep\u003c/em\u003e-value) and absolute mean difference was taken per drug as the filtering threshold.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec18\" class=\"Section2\"\u003e \u003ch2\u003eMetabolite adjacency\u003c/h2\u003e \u003cp\u003eMetabolite adjacency table (Table S8) was generated from human KEGG pathway by \u0026lsquo;MetaboSignal\u0026rsquo; package in R (Rodriguez-Martinez et al., \u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e2016\u003c/span\u003e). Specifically, all human metabolic pathways were extracted and subsetted using \u0026lsquo;MS_getPathIds()\u0026rsquo; function with argument \u0026lsquo;organism_code = \u0026ldquo;hsa\u0026rdquo;\u0026rsquo;. Reaction network was built based on metabolic pathways using \u0026lsquo;MS_reactionNetwork()\u0026rsquo; function. Subsequently, node distance was calculated using \u0026lsquo;MS_nodeBW()\u0026rsquo; function with argument \u0026lsquo;node = \u0026ldquo;out\u0026rdquo;\u0026rsquo; and \u0026lsquo;normalized\u0026thinsp;=\u0026thinsp;TRUE\u0026rsquo;.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec19\" class=\"Section2\"\u003e \u003ch2\u003eMetabolite prediction algorithm\u003c/h2\u003e \u003cp\u003eFor every SLC-metabolite pair \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:i\\)\u003c/span\u003e\u003c/span\u003e, the prediction algorithm computes a confidence score by evaluating its correlation statistics across the three datasets (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{A}_{i}\\)\u003c/span\u003e\u003c/span\u003e), resulted adjusted \u003cem\u003ep\u003c/em\u003e-value of annotated CRISPR-Cas9 screen (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{B}_{i}\\)\u003c/span\u003e\u003c/span\u003e), and adjacent metabolites (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{C}_{i}\\)\u003c/span\u003e\u003c/span\u003e). Threshold of discovery and score reward per discovery is specified in Figure \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e, validified with fractional difference between true positive and false positive. For every SLC-metabolite pair, if value implicated in \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{A}_{i}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{B}_{i}\\)\u003c/span\u003e\u003c/span\u003e, or \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{C}_{i}\\)\u003c/span\u003e\u003c/span\u003e is smaller than its respective threshold of discovery or the metabolite is not measured in the dataset not exist, a confidence score of 0 was assigned. Otherwise, the confidence score will be measured with respect to the decile-based percentile of known SLC-substrate tables, such that a decile of 10 corresponds to the top 10% and a decile of 1 the lowest 10%; correlations greater than or equal to the top correlation (i.e. greater than decile 10) were assigned a score of 11. The adjusted \u003cem\u003ep\u003c/em\u003e-value of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{B}_{i}\\)\u003c/span\u003e\u003c/span\u003e was measured in log-transformed format. The decile number was multiplied by 3 for \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{A}_{i}\\)\u003c/span\u003e\u003c/span\u003e, 1.1 for \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{B}_{i}\\)\u003c/span\u003e\u003c/span\u003e, and 0.9 for \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{C}_{i}\\)\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec20\" class=\"Section2\"\u003e \u003ch2\u003eDrug prediction algorithm\u003c/h2\u003e \u003cp\u003eFor each pair between drug molecule \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:a\\)\u003c/span\u003e\u003c/span\u003e and SLC \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:z\\)\u003c/span\u003e\u003c/span\u003e, the dose responses were only calculated for cell lines annotated based on the highest and lowest 20% of SLC \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:z\\)\u003c/span\u003e\u003c/span\u003e expression. Local polynomial regression models were fitted to the two expression types using \u0026lsquo;loess()\u0026rsquo; function, viability against log transformed dose (-3.21 to 1). 421 data points, ranging from \u0026minus;\u0026thinsp;3.21 to 1 and separated by 0.01, predicted by model using \u0026lsquo;predict()\u0026rsquo; function was generated to capture the shape information of fitted models. Curve shapes were compared with a paired t-test. The resulted \u003cem\u003ep\u003c/em\u003e-value and mean differences were recorded.\u003c/p\u003e \u003c/div\u003e "},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eCode availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll the code for processing data, as well as generating the figures and tables in the manuscripts, is available as supplemental material and will be uploaded to GitHub upon publication.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNo specific ethics approval was required for this study\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNo specific consent for publication is required for this study.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNo specific funding was obtained for this project\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026apos; contributions\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eConceptualization: PS, SN\u003c/p\u003e\n\u003cp\u003eFormal analysis: YZ, PS\u003c/p\u003e\n\u003cp\u003eData analysis and interpretation: YZ, PS\u003c/p\u003e\n\u003cp\u003eManuscript first draft: YZ, PS\u003c/p\u003e\n\u003cp\u003eManuscript review and editing: YZ, SN, PS\u003c/p\u003e\n\u003cp\u003eAll authors read and approved the final manuscript.\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe thank the members of the EpiEvo group and other colleagues at the Department of Biochemistry, University of Oxford for helpful discussion and comments on the project. We thank Dr Marcos Francisco Perez for aligning and normalising raw NCI-60 transcriptomics data. We also thank Dr Louise Fets (London Institute of Medical Sciences) for helpful discussions. \u0026nbsp;\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eAlam, S., Doherty, E., Ortega-Prieto, P., Arizanova, J., \u0026amp; Fets, L. (2023). Membrane transporters in cell physiology, cancer metabolism and drug response. \u003cem\u003eDisease Models \u0026amp; Mechanisms\u003c/em\u003e, \u003cem\u003e16\u003c/em\u003e(11). https://doi.org/10.1242/dmm.050404\u003c/li\u003e\n\u003cli\u003eAulakh, S. K., Varma, S. J., \u0026amp; Ralser, M. (2022). Metal ion availability and homeostasis as drivers of metabolic evolution and enzyme function. \u003cem\u003eCurrent Opinion in Genetics \u0026amp; Development\u003c/em\u003e, \u003cem\u003e77\u003c/em\u003e, 101987. https://doi.org/10.1016/j.gde.2022.101987\u003c/li\u003e\n\u003cli\u003eBergeron, M. J., Cl\u0026eacute;men\u0026ccedil;on, B., Hediger, M. A., \u0026amp; Markovich, D. (2013). SLC13 family of Na+-coupled di- and tri-carboxylate/sulfate transporters. \u003cem\u003eMolecular Aspects of Medicine\u003c/em\u003e, \u003cem\u003e34\u003c/em\u003e(2\u0026ndash;3), 299\u0026ndash;312. https://doi.org/10.1016/j.mam.2012.12.001\u003c/li\u003e\n\u003cli\u003eBock, C., Datlinger, P., Chardon, F., Coelho, M. A., Dong, M. B., Lawson, K. A., Lu, T., Maroc, L., Norman, T. M., Song, B., Stanley, G., Chen, S., Garnett, M., Li, W., Moffat, J., Qi, L. S., Shapiro, R. S., Shendure, J., Weissman, J. S., \u0026amp; Zhuang, X. (2022). High-content CRISPR screening. In \u003cem\u003eNature Reviews Methods Primers\u003c/em\u003e (Vol. 2, Issue 1). Springer Nature. https://doi.org/10.1038/s43586-021-00093-4\u003c/li\u003e\n\u003cli\u003eC\u0026eacute;sar-Razquin, A., Snijder, B., Frappier-Brinton, T., Isserlin, R., Gyimesi, G., Bai, X., Reithmeier, R. A., Hepworth, D., Hediger, M. A., Edwards, A. M., \u0026amp; Superti-Furga, G. (2015). A Call for Systematic Research on Solute Carriers. \u003cem\u003eCell\u003c/em\u003e, \u003cem\u003e162\u003c/em\u003e(3), 478\u0026ndash;487. https://doi.org/10.1016/j.cell.2015.07.022\u003c/li\u003e\n\u003cli\u003eCherkaoui, S., Durot, S., Bradley, J., Critchlow, S. E., Dubuis, S., Masiero, M., Wegmann, R., Snijder, B., Othman, A., Bendtsen, C., \u0026amp; Zamboni, N. (2022). A functional analysis of 180 cancer cell lines reveals conserved intrinsic metabolic programs. \u003cem\u003eMolecular Systems Biology\u003c/em\u003e, \u003cem\u003e18\u003c/em\u003e(11). https://doi.org/10.15252/msb.202211033\u003c/li\u003e\n\u003cli\u003eCorsello, S. M., Nagari, R. T., Spangler, R. D., Rossen, J., Kocak, M., Bryan, J. G., Humeidi, R., Peck, D., Wu, X., Tang, A. A., Wang, V. M., Bender, S. A., Lemire, E., Narayan, R., Montgomery, P., Ben-David, U., Garvie, C. W., Chen, Y., Rees, M. G., \u0026hellip; Golub, T. R. (2020). Discovering the anticancer potential of non-oncology drugs by systematic viability profiling. \u003cem\u003eNature Cancer\u003c/em\u003e, \u003cem\u003e1\u003c/em\u003e, 235\u0026ndash;248. https://doi.org/10.1038/s43018-019-0018-6\u003c/li\u003e\n\u003cli\u003eDutta, B., Huang, W., Molero, M., Kekuda, R., Leibach, F. H., Devoe, L. D., Ganapathy, V., \u0026amp; Prasad, P. D. (1999). Cloning of the Human Thiamine Transporter, a Member of the Folate Transporter Family *. \u003cem\u003eJournal of Biological Chemistry\u003c/em\u003e, \u003cem\u003e274\u003c/em\u003e(45), 31925\u0026ndash;31929. https://doi.org/10.1074/jbc.274.45.31925\u003c/li\u003e\n\u003cli\u003eFerrada, E., \u0026amp; Superti-Furga, G. (2022). A structure and evolutionary-based classification of solute carriers. \u003cem\u003eIScience\u003c/em\u003e, \u003cem\u003e25\u003c/em\u003e(10), 105096. https://doi.org/10.1016/j.isci.2022.105096\u003c/li\u003e\n\u003cli\u003eGiacomini, K. M., Yee, S. W., Koleske, M. L., Zou, L., Matsson, P., Chen, E. C., Kroetz, D. L., Miller, M. A., Gozalpour, E., \u0026amp; Chu, X. (2022). New and Emerging Research on Solute Carrier and ATP Binding Cassette Transporters in Drug Discovery and Development: Outlook From the International Transporter Consortium. \u003cem\u003eClinical Pharmacology \u0026amp; Therapeutics\u003c/em\u003e, \u003cem\u003e112\u003c/em\u003e(3), 540\u0026ndash;561. https://doi.org/10.1002/CPT.2627\u003c/li\u003e\n\u003cli\u003eGirardi, E., C\u0026eacute;sar-Razquin, A., Lindinger, S., Papakostas, K., Konecka, J., Hemmerich, J., Kickinger, S., Kartnig, F., G\u0026uuml;rtl, B., Klavins, K., Sedlyarov, V., Ingles-Prieto, A., Fiume, G., Koren, A., Lardeau, C.-H., Kumaran Kandasamy, R., Kubicek, S., Ecker, G. F., \u0026amp; Superti-Furga, G. (2020). A widespread role for SLC transmembrane transporters in resistance to cytotoxic drugs. \u003cem\u003eNature Chemical Biology\u003c/em\u003e, \u003cem\u003e16\u003c/em\u003e(4), 469\u0026ndash;478. https://doi.org/10.1038/s41589-020-0483-3\u003c/li\u003e\n\u003cli\u003eGyimesi, G., \u0026amp; Hediger, M. A. (2022). Systematic in silico discovery of novel solute carrier-like proteins from proteomes. \u003cem\u003ePLOS ONE\u003c/em\u003e, \u003cem\u003e17\u003c/em\u003e(7), e0271062\u0026ndash;e0271062. https://doi.org/10.1371/journal.pone.0271062\u003c/li\u003e\n\u003cli\u003eHediger, M. A., Cl\u0026eacute;men\u0026ccedil;on, B., Burrier, R. E., \u0026amp; Bruford, E. A. (2013). The ABCs of membrane transporters in health and disease (SLC series): Introduction. \u003cem\u003eMolecular Aspects of Medicine\u003c/em\u003e, \u003cem\u003e34\u003c/em\u003e(2\u0026ndash;3), 95\u0026ndash;107. https://doi.org/10.1016/J.MAM.2012.12.009\u003c/li\u003e\n\u003cli\u003eHeins-Marroquin, U., Singh, R. R., Perathoner, S., Gavotto, F., Ruiz, C. M., Patraskaki, M., Gomez-Giro, G., Borgmann, F. K., Meyer, M., Carpentier, A., Warmoes, M. O., J\u0026auml;ger, C., Mittelbronn, M., Schwamborn, J. C., Cordero-Maldonado, M. L., Crawford, A. D., Schymanski, E. L., \u0026amp; Linster, C. L. (2024). CLN3 deficiency leads to neurological and metabolic perturbations during early development. \u003cem\u003eLife Science Alliance\u003c/em\u003e, \u003cem\u003e7\u003c/em\u003e(3), e202302057\u0026ndash;e202302057. https://doi.org/10.26508/lsa.202302057\u003c/li\u003e\n\u003cli\u003eHiguchi, K., Sugiyama, K., Tomabechi, R., Kishimoto, H., \u0026amp; Inoue, K. (2022). Mammalian monocarboxylate transporter 7 (MCT7/Slc16a6) is a novel facilitative taurine transporter. \u003cem\u003eThe Journal of Biological Chemistry\u003c/em\u003e, \u003cem\u003e298\u003c/em\u003e(4), 101800. https://doi.org/10.1016/j.jbc.2022.101800\u003c/li\u003e\n\u003cli\u003eJia, X., Zhu, J., Bian, X., Liu, S., Yu, S., Liang, W., Jiang, L., Mao, R., Zhang, W., \u0026amp; Rao, Y. (2023). Importance of glutamine in synaptic vesicles revealed by functional studies of SLC6A17 and its mutations pathogenic for intellectual disability. \u003cem\u003eELife\u003c/em\u003e, \u003cem\u003e12\u003c/em\u003e, RP86972. https://doi.org/10.7554/eLife.86972\u003c/li\u003e\n\u003cli\u003eKoepsell, H., \u0026amp; Endou, H. (2004). The SLC22 drug transporter family. \u003cem\u003ePfl\u0026uuml;gers Archiv: European Journal of Physiology\u003c/em\u003e, \u003cem\u003e447\u003c/em\u003e(5), 666\u0026ndash;676. https://doi.org/10.1007/s00424-003-1089-9\u003c/li\u003e\n\u003cli\u003eKunji, E. R. S., Aleksandrova, A., King, M. S., Majd, H., Ashton, V. L., Cerson, E., Springett, R., Kibalchenko, M., Tavoulari, S., Crichton, P. G., \u0026amp; Ruprecht, J. J. (2016). The transport mechanism of the mitochondrial ADP/ATP carrier. \u003cem\u003eBiochimica et Biophysica Acta (BBA) - Molecular Cell Research\u003c/em\u003e, \u003cem\u003e1863\u003c/em\u003e(10), 2379\u0026ndash;2393. https://doi.org/10.1016/j.bbamcr.2016.03.015\u003c/li\u003e\n\u003cli\u003eLi, H., Ning, S., Ghandi, M., Kryukov, G. V, Gopal, S., Deik, A., Souza, A., Pierce, K., Keskula, P., Hernandez, D., Ann, J., Shkoza, D., Apfel, V., Zou, Y., Vazquez, F., Barretina, J., Pagliarini, R. A., Galli, G. G., Root, D. E., \u0026hellip; Sellers, W. R. (2019). The landscape of cancer cell line metabolism. \u003cem\u003eNature Medicine\u003c/em\u003e, \u003cem\u003e25\u003c/em\u003e(5), 850\u0026ndash;860. https://doi.org/10.1038/s41591-019-0404-8\u003c/li\u003e\n\u003cli\u003eLin, L., Yee, S. W., Kim, R. B., \u0026amp; Giacomini, K. M. (2015). SLC transporters as therapeutic targets: Emerging opportunities. In \u003cem\u003eNature Reviews Drug Discovery\u003c/em\u003e (Vol. 14, Issue 8, pp. 543\u0026ndash;560). Nature Publishing Group. https://doi.org/10.1038/nrd4626\u003c/li\u003e\n\u003cli\u003eMajd, H., King, M. S., Smith, A. C., \u0026amp; Kunji, E. R. S. (2018). Pathogenic mutations of the human mitochondrial citrate carrier SLC25A1 lead to impaired citrate export required for lipid, dolichol, ubiquinone and sterol synthesis. \u003cem\u003eBiochimica et Biophysica Acta (BBA) - Bioenergetics\u003c/em\u003e, \u003cem\u003e1859\u003c/em\u003e(1), 1\u0026ndash;7. https://doi.org/10.1016/j.bbabio.2017.10.002\u003c/li\u003e\n\u003cli\u003eMeixner, E., Goldmann, U., Sedlyarov, V., Scorzoni, S., Rebsamen, M., Girardi, E., \u0026amp; Superti‐Furga, G. (2020). A substrate‐based ontology for human solute carriers. \u003cem\u003eMolecular Systems Biology\u003c/em\u003e, \u003cem\u003e16\u003c/em\u003e(7). https://doi.org/10.15252/msb.20209652\u003c/li\u003e\n\u003cli\u003eMeyers, R. M., Bryan, J. G., McFarland, J. M., Weir, B. A., Sizemore, A. E., Xu, H., Dharia, N. V, Montgomery, P. G., Cowley, G. S., Pantel, S., Goodale, A., Lee, Y., Ali, L. D., Jiang, G., Lubonja, R., Harrington, W. F., Strickland, M., Wu, T., Hawes, D. C., \u0026hellip; Tsherniak, A. (2017). Computational correction of copy number effect improves specificity of CRISPR\u0026ndash;Cas9 essentiality screens in cancer cells. \u003cem\u003eNature Genetics\u003c/em\u003e, \u003cem\u003e49\u003c/em\u003e(12), 1779\u0026ndash;1784. https://doi.org/10.1038/ng.3984\u003c/li\u003e\n\u003cli\u003eMikkaichi, T., Suzuki, T., Onogawa, T., Tanemoto, M., Mizutamari, H., Okada, M., Chaki, T., Masuda, S., Tokui, T., Eto, N., Abe, M., Satoh, F., Unno, M., Hishinuma, T., Inui, K. I., Ito, S., Goto, J., \u0026amp; Abe, T. (2004). Isolation and characterization of a digoxin transporter and its rat homologue expressed in the kidney. \u003cem\u003eProc. Natl. Acad. Sci. U. S. A.\u003c/em\u003e, \u003cem\u003e101\u003c/em\u003e(10), 3569\u0026ndash;3574. https://doi.org/10.1073/pnas.0304987101\u003c/li\u003e\n\u003cli\u003eNimmanon, T., Ziliotto, S., Morris, S., Flanagan, L., \u0026amp; Taylor, K. M. (2017). Phosphorylation of zinc channel ZIP7 drives MAPK, PI3K and mTOR growth and proliferation signalling. \u003cem\u003eMetallomics\u003c/em\u003e, \u003cem\u003e9\u003c/em\u003e(5), 471\u0026ndash;481. https://doi.org/10.1039/c6mt00286b\u003c/li\u003e\n\u003cli\u003eOkada, Y. (2004). Ion Channels and Transporters Involved in Cell Volume Regulation and Sensor Mechanisms. \u003cem\u003eCell Biochemistry and Biophysics\u003c/em\u003e, \u003cem\u003e41\u003c/em\u003e(2), 233\u0026ndash;258. https://doi.org/10.1385/cbb:41:2:233\u003c/li\u003e\n\u003cli\u003ePerez, M. F., \u0026amp; Sarkies, P. (2023). Histone methyltransferase activity affects metabolism in human cells independently of transcriptional regulation. \u003cem\u003ePLoS Biology\u003c/em\u003e, \u003cem\u003e21\u003c/em\u003e(10 October). https://doi.org/10.1371/journal.pbio.3002354\u003c/li\u003e\n\u003cli\u003ePizzagalli, M. D., Bensimon, A., \u0026amp; Superti-Furga, G. (2021). A guide to plasma membrane solute carrier proteins. In \u003cem\u003eFEBS Journal\u003c/em\u003e (Vol. 288, Issue 9, pp. 2784\u0026ndash;2835). John Wiley and Sons Inc. https://doi.org/10.1111/febs.15531\u003c/li\u003e\n\u003cli\u003eRamamoorthy, S., Leibach, F. H., Mahesh, V. B., Han, H., Yang-Feng, T., Blakely, R. D., \u0026amp; Ganapathy, V. (1994). Functional characterization and chromosomal localization of a cloned taurine transporter from human placenta. \u003cem\u003eBiochemical Journal\u003c/em\u003e, \u003cem\u003e300\u003c/em\u003e(3), 893\u0026ndash;900. https://doi.org/10.1042/bj3000893\u003c/li\u003e\n\u003cli\u003eRodriguez-Martinez, A., Ayala, R., Posma, J. M., Neves, A. L., Gauguier, D., Nicholson, J. K., \u0026amp; Dumas, M.-E. (2016). MetaboSignal: a network-based approach for topological analysis of metabotype regulationviametabolic and signaling pathways. \u003cem\u003eBioinformatics\u003c/em\u003e, \u003cem\u003e33\u003c/em\u003e(5), btw697. https://doi.org/10.1093/bioinformatics/btw697\u003c/li\u003e\n\u003cli\u003eShoemaker, R. H. (2006). The NCI60 human tumour cell line anticancer drug screen. \u003cem\u003eNature Reviews Cancer\u003c/em\u003e, \u003cem\u003e6\u003c/em\u003e(10), 813\u0026ndash;823. https://doi.org/10.1038/nrc1951\u003c/li\u003e\n\u003cli\u003eSkelton, M. R., Schaefer, T. L., Graham, D. L., deGrauw, T. J., Clark, J. F., Williams, M. T., \u0026amp; Vorhees, C. V. (2011). Creatine Transporter (CrT; Slc6a8) Knockout Mice as a Model of Human CrT Deficiency. \u003cem\u003ePLoS ONE\u003c/em\u003e, \u003cem\u003e6\u003c/em\u003e(1), e16187. https://doi.org/10.1371/journal.pone.0016187\u003c/li\u003e\n\u003cli\u003eSong, W., Li, D., Tao, L., Luo, Q., \u0026amp; Chen, L. (2020). Solute carrier transporters: the metabolic gatekeepers of immune cells. \u003cem\u003eActa Pharmaceutica Sinica B\u003c/em\u003e, \u003cem\u003e10\u003c/em\u003e(1), 61\u0026ndash;78. https://doi.org/10.1016/j.apsb.2019.12.006\u003c/li\u003e\n\u003cli\u003eSuperti-Furga, G., Lackner, D., Wiedmer, T., Ingles-Prieto, A., Barbosa, B., Girardi, E., Goldmann, U., G\u0026uuml;rtl, B., Klavins, K., Klimek, C., Lindinger, S., Li\u0026ntilde;eiro-Retes, E., M\u0026uuml;ller, A. C., Onstein, S., Redinger, G., Reil, D., Sedlyarov, V., Wolf, G., Crawford, M., \u0026hellip; Steppan, C. M. (2020). The RESOLUTE consortium: unlocking SLC transporters for drug discovery. \u003cem\u003eNature Reviews Drug Discovery\u003c/em\u003e, \u003cem\u003e19\u003c/em\u003e(7), 429\u0026ndash;430. https://doi.org/10.1038/d41573-020-00056-6\u003c/li\u003e\n\u003cli\u003eSzeri, F., Lundkvist, S., Donnelly, S., Engelke, U., Rhee, K., Williams, C. J., Sundberg, J. P., Wevers, R. A., Tomlinson, R. E., Jansen, R. S., \u0026amp; Wetering, K. (2020). The membrane protein ANKH is crucial for bone mechanical performance by mediating cellular export of citrate and ATP. \u003cem\u003ePLOS Genetics\u003c/em\u003e, \u003cem\u003e16\u003c/em\u003e(7), e1008884\u0026ndash;e1008884. https://doi.org/10.1371/journal.pgen.1008884\u003c/li\u003e\n\u003cli\u003eTaegtmeyer, H., \u0026amp; Ingwall, J. S. (2013). Creatine\u0026mdash;A Dispensable Metabolite? \u003cem\u003eCirculation Research\u003c/em\u003e, \u003cem\u003e112\u003c/em\u003e(6), 878\u0026ndash;880. https://doi.org/10.1161/circresaha.113.300974\u003c/li\u003e\n\u003cli\u003eTsherniak, A., Vazquez, F., Montgomery, P. G., Weir, B. A., Kryukov, G., Cowley, G. S., Gill, S., Harrington, W. F., Pantel, S., Krill-Burger, J. M., Meyers, R. M., Ali, L., Goodale, A., Lee, Y., Jiang, G., Hsiao, J., Gerath, W. F. J., Howell, S., Merkel, E., \u0026hellip; Hahn, W. C. (2017). Defining a Cancer Dependency Map. \u003cem\u003eCell\u003c/em\u003e, \u003cem\u003e170\u003c/em\u003e(3), 564-576.e16. https://doi.org/10.1016/j.cell.2017.06.010\u003c/li\u003e\n\u003cli\u003eWang, X., Ji, Y., Qi, J., Zhou, S., Wan, S., Fan, C., Gu, Z., An, P., Luo, Y., \u0026amp; Luo, J. (2023). Mitochondrial carrier 1 (MTCH1) governs ferroptosis by triggering the FoxO1-GPX4 axis-mediated retrograde signaling in cervical cancer cells. \u003cem\u003eCell Death and Disease\u003c/em\u003e, \u003cem\u003e14\u003c/em\u003e(8). https://doi.org/10.1038/s41419-023-06033-2\u003c/li\u003e\n\u003cli\u003eWiegering, A., Matthes, N., M\u0026uuml;hling, B., Koospal, M., Quenzer, A., Peter, S., Germer, C.-T., Linnebacher, M., \u0026amp; Otto, C. (2017). Reactivating p53 and Inducing Tumor Apoptosis (RITA) Enhances the Response of RITA-Sensitive Colorectal Cancer Cells to Chemotherapeutic Agents 5-Fluorouracil and Oxaliplatin. \u003cem\u003eNeoplasia\u003c/em\u003e, \u003cem\u003e19\u003c/em\u003e(4), 301\u0026ndash;309. https://doi.org/10.1016/j.neo.2017.01.007\u003c/li\u003e\n\u003cli\u003eWinter, G. E., Radic, B., Mayor-Ruiz, C., Blomen, V. A., Trefzer, C., Kandasamy, R. K., Huber, K. V. M., Gridling, M., Chen, D., Klampfl, T., Kralovics, R., Kubicek, S., Fernandez-Capetillo, O., Brummelkamp, T. R., \u0026amp; Superti-Furga, G. (2014). The solute carrier SLC35F2 enables YM155-mediated DNA damage toxicity. \u003cem\u003eNature Chemical Biology\u003c/em\u003e, \u003cem\u003e10\u003c/em\u003e(9), 768\u0026ndash;773. https://doi.org/10.1038/nchembio.1590\u003c/li\u003e\n\u003cli\u003eWyss, M., \u0026amp; Kaddurah-Daouk, R. (2000). Creatine and Creatinine Metabolism. \u003cem\u003ePhysiological Reviews\u003c/em\u003e, \u003cem\u003e80\u003c/em\u003e(3), 1107\u0026ndash;1213. https://doi.org/10.1152/physrev.2000.80.3.1107\u003c/li\u003e\n\u003cli\u003eZajac, M., Mukherjee, S., Anees, P., Oettinger, D., Henn, K., Srikumar, J., Zou, J., Saminathan, A., \u0026amp; Krishnan, Y. (2024). A mechanism of lysosomal calcium entry. \u003cem\u003eScience Advances\u003c/em\u003e, \u003cem\u003e10\u003c/em\u003e(7). https://doi.org/10.1126/sciadv.adk2317\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"bmc-genomics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"gics","sideBox":"Learn more about [BMC Genomics](http://bmcgenomics.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/gics","title":"BMC Genomics","twitterHandle":"#BMCGenomics","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-4713269/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4713269/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eSolute carriers (SLC) are integral membrane proteins responsible for transporting a wide variety of metabolites, signaling molecules and drugs across cellular membranes. Despite key roles in metabolism, signaling and pharmacology, around one third of SLC proteins are \u0026lsquo;orphans\u0026rsquo; whose substrates are unknown. Experimental determination of SLC substrates is technically challenging given the wide range of possible physiological candidates. Here, we develop a predictive algorithm to identify correlations between SLC expression levels and intracellular metabolite concentrations by leveraging existing cancer multi-omics datasets. Our predictions recovered known SLC-substrate pairs with high sensitivity and specificity compared to simulated random pairs. CRISPR loss-of-function screen data and metabolic pathway adjacency data further improved the performance of our algorithm. In parallel, we combined drug sensitivity data with SLC expression profiles to predict new SLC-drug interactions. Together, we provide a novel bioinformatic pipeline to predict new substrate predictions for SLCs, offering new opportunities to de-orphanise SLCs with important implications for understanding their roles in health and disease.\u003c/p\u003e","manuscriptTitle":"Predicting substrates for orphan Solute Carrier Proteins using multi- omics datasets","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-08-05 07:41:31","doi":"10.21203/rs.3.rs-4713269/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2024-10-23T08:22:17+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2024-10-22T22:10:24+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2024-09-24T20:20:38+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2024-09-12T08:14:10+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"214751873261883764274890871293051173423","date":"2024-09-05T17:45:01+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"64780935902484902355157183876896508182","date":"2024-09-05T07:08:24+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"150566266131480549315852734567980633092","date":"2024-08-31T14:01:33+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"133499210862548203902627649102404247807","date":"2024-08-29T08:51:55+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2024-08-22T06:31:56+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2024-07-12T11:50:11+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2024-07-11T04:34:32+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2024-07-11T04:33:37+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Genomics","date":"2024-07-09T15:41:54+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"bmc-genomics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"gics","sideBox":"Learn more about [BMC Genomics](http://bmcgenomics.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/gics","title":"BMC Genomics","twitterHandle":"#BMCGenomics","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"1ae2b606-6ae3-46f3-a4d2-a2960cab059c","owner":[],"postedDate":"August 5th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2025-02-17T16:09:56+00:00","versionOfRecord":{"articleIdentity":"rs-4713269","link":"https://doi.org/10.1186/s12864-025-11330-5","journal":{"identity":"bmc-genomics","isVorOnly":false,"title":"BMC Genomics"},"publishedOn":"2025-02-11 15:57:55","publishedOnDateReadable":"February 11th, 2025"},"versionCreatedAt":"2024-08-05 07:41:31","video":"","vorDoi":"10.1186/s12864-025-11330-5","vorDoiUrl":"https://doi.org/10.1186/s12864-025-11330-5","workflowStages":[]},"version":"v1","identity":"rs-4713269","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4713269","identity":"rs-4713269","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-22T02:00:06.705733+00:00
License: CC-BY-4.0