Integrating single-cell and bulk transcriptomic perturbation resources reveals complementary therapeutic spaces for drug repurposing | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Integrating single-cell and bulk transcriptomic perturbation resources reveals complementary therapeutic spaces for drug repurposing Enock Niyonkuru, Umair Khan, Xinyu Tang, Laura Almonte, Eden Chun, and 9 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-10606217/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 5 You are reading this latest preprint version Abstract Background Transcriptome-based drug repurposing can accelerate therapeutic discovery, but is limited by fragmented resources, inconsistent quality control, and reliance on single perturbation databases. Methods We developed CDRPipe ( C omputational D rug R epurposing Pipe line), a unified framework that interrogates disease signatures against drug perturbation signatures generated by distinct experimental technologies. Specifically, CDRPipe harmonizes microarray perturbation profiles from the Connectivity Map (CMap; 1,968 quality-filtered experiments) with pseudo-bulk profiles derived from large-scale single-cell RNA sequencing experiments in the Tahoe-100M database (56,827 experiments). CDRPipe standardizes preprocessing, computes rank-based connectivity scores and evaluates significance using empirical null models. We applied CDRPipe to 233 curated disease signatures from GEO and CREEDS and evaluated performance using known drug-disease associations from Open Targets. Results Single-cell-derived pseudo-bulk profiles recovered more annotated therapeutics than microarray profiles (mean recall 47.3% vs 18.5%; Wilcoxon signed-rank test p < 10 –11 ), though these differences partly reflect differences in drug library composition and clinical annotation coverage. Importantly, the two resources were highly complementary, with only 3.5% overlap in recovered drugs, indicating that integrating predictions across independent perturbation resources expands therapeutic coverage and enables identification of high-confidence consensus candidates. Case studies in autoimmune disease and endometriosis further demonstrate that CDRPipe recovers clinically relevant therapies while revealing technology-dependent patterns of discovery. Conclusions Integrating heterogeneous transcriptomic perturbation resources improves the robustness and interpretability of transcriptional drug repurposing. Consensus predictions supported by independent resources represent higher-confidence candidates for downstream experimental and clinical validation, and CDRPipe provides an openly available framework to support this integrated approach. Drug repurposing Transcriptomics Signature reversal Connectivity Map Single-cell RNA sequencing Tahoe-100M Disease Signature Drug Signature Computational pharmacology Figures Figure 1 Figure 2 Figure 3 Figure 4 Background The contemporary landscape of pharmaceutical research and development (R&D) is characterized by a paradoxical trend often referred to as “Eroom’s Law”, the observation that drug discovery is becoming slower and more expensive over time, despite improvements in technology [ 1 ]. Developing a therapeutic agent de novo is a capital-intensive endeavor, estimated to cost between 2 billion and 3 billion and requiring 10 to 17 years to move from target identification to market approval [ 2 ]. Compounding this financial burden is a high attrition rate; approximately 90% of drug candidates entering clinical trials fail to achieve regulatory approval, often due to lack of efficacy or unforeseen toxicity in late-stage development [ 3 ]. This systemic inefficiency leaves a vast array of diseases, particularly rare conditions and complex chronic disorders, without effective targeted treatments. In response to these challenges, drug repurposing, the identification of novel therapeutic indications for existing drugs with established safety profiles, has emerged as a critical strategy to accelerate therapeutic discovery and reduce translational risk [ 4 , 5 ]. Historically, repurposing successes were largely unanticipated, driven by clinical observation of off-target effects, as exemplified by the redeployment of sildenafil for erectile dysfunction or thalidomide for multiple myeloma [ 1 , 6 ]. However, the explosion of high-throughput “omics” data has catalyzed a paradigm shift toward systematic, computational drug repurposing [ 5 , 7 , 8 ]. Central to this modern approach is the hypothesis of transcriptional reversal (or signature reversion). This principle posits that if a disease state is characterized by a specific gene expression signature (a set of up- and downregulated genes), a small molecule capable of inducing the opposite transcriptional profile may neutralize the disease phenotype and restore a healthy physiological state ([ 9 ]). This concept was pioneered by the Connectivity Map (CMap) project, which provided the first large-scale reference database of drug-induced transcriptional perturbations, allowing researchers, including our own team, to systematically connect small molecules, genes, and diseases [ 10 – 12 ]. Despite the conceptual elegance and early successes of transcriptomic drug repurposing, the field faces persistent limitations related to data heterogeneity, platform bias, and biological resolution. In bulk transcriptomic experiments, gene expression is measured across all cells present in a tissue or sample and represented as a single composite profile. As a consequence, transcriptional programs originating from rare but disease-driving cell populations can be masked by signals from more abundant but less relevant cell types, reducing the sensitivity of bulk-derived signatures to capture true disease mechanisms. This dilution of cell-type-specific signals can cause therapeutically relevant compounds, those targeting the actual disease-driving pathways, to fail to show significant transcriptional reversal against the composite bulk signature, leading to missed therapeutic associations. At the same time, compounds that induce broad transcriptional suppression or cytostatic responses may be artifactually favored over those with more targeted mechanisms of action [ 13 ]. Furthermore, reliance on a single perturbation technology can introduce platform-specific biases that limit generalizability [ 14 ]. The emergence of large-scale single-cell RNA sequencing (scRNA-seq) perturbation databases presents an opportunity to diversify the transcriptional resources available for drug repurposing [ 14 – 16 ]. While single-cell perturbation profiling does not directly resolve cell-type heterogeneity in bulk disease signatures, it offers complementary advantages: broader genomic coverage, a wider diversity of cellular contexts for measuring drug responses, and the ability to capture drug-induced transcriptional effects that bulk microarray platforms may fail to detect. For example, Tahoe-100M is a massive single-cell perturbation atlas comprising approximately 100 million scRNA-seq profiles that capture transcriptional responses to small-molecule treatments across dozens of cancer cell lines. Yet, the integration of these high-dimensional, disparate data types remains a significant computational bottleneck [ 17 ], often forcing researchers to choose between the extensive drug coverage of legacy bulk databases and the granular resolution of modern single-cell platforms. To address these challenges, we developed CDRPipe, a unified framework for transcriptional drug repurposing that enables direct comparison and integration of transcriptional perturbation signatures derived from heterogeneous measurement technologies. CDRPipe standardizes preprocessing, applies consistent nonparametric connectivity scoring, and evaluates statistical significance within a shared empirical framework. We hypothesized that therapeutically relevant compounds would exhibit consistent transcriptional reversal across independent perturbation resources, and that pseudo-bulk signatures derived from single-cell experiments would improve the recovery of known drug-disease associations compared to microarray-based approaches alone. By applying CDRPipe to 233 curated disease signatures from CREEDS (Crowd-Extracted Expression of Differential Signatures) [ 18 ] and validating predictions against curated drug-disease knowledge bases, such as Open Targets [ 19 ], we demonstrate that multi-database integration substantially enhances the robustness, interpretability, and translational relevance of transcriptional drug repurposing. We further illustrate the utility of this framework through case studies in autoimmune diseases, where we recover established immunomodulators and identify novel targeted therapies, and in endometriosis, where we replicate prior findings while uncovering new enzyme-centric therapeutic candidates. Methods Disease and Drug Signatures This study was designed to systematically compare and integrate two connectivity mapping approaches for transcriptional drug repurposing. We conducted a retrospective computational analysis using publicly available disease signatures and drug perturbation profiles, with validation against known drug-disease associations. The primary outcomes were recall (recovery rate) of known therapeutics and identification of consensus candidates across platforms. Sample sizes were determined by the availability of curated disease signatures (n = 233) in CREEDS and their alignment with Open Targets (n = 203 evaluable diseases). Disease signatures were obtained from CREEDS, a manually curated database of disease-associated differential expression signatures derived from public studies in the Gene Expression Omnibus (GEO) [ 18 , 20 ]. Each signature contrasts diseased tissues or cells with matched normal or control samples within the same study and is defined by sets of up-regulated and down-regulated genes representing transcriptomic changes associated with the disease state relative to healthy controls. Signatures were generated through a large-scale crowdsourcing effort with manual sample selection, standardized disease annotation, batch-effect correction using surrogate variable analysis, and differential expression prioritization via the Characteristic Direction method, and were all derived from microarray-based experiments [ 18 ]. They were quality-controlled during original curation; we applied no additional experiment-level filtering prior to analysis. Across the 233 signatures, the mean size was approximately 975 genes (range 38 − 4,867) prior to filtering and 419 genes (range 2–2,525) after applying significance and fold-change thresholds (Fig. 1 ). We used the original Connectivity Map (CMap) resource, which provides direct microarray expression profiles of drug-treated cells [ 12 ]. Following quality-control filtering based on inter-replicate consistency (Pearson correlation r ≥ 0.15 between replicate signatures for the same drug-dose-cell line combination), we retained 1,968 of 6,100 original experiments (32.3%), spanning 13,071 unique gene features across five cell lines and 1,309 unique drugs. We used pseudo-bulk perturbation profiles derived from large-scale single-cell RNA sequencing experiments in the Tahoe-100M atlas [ 16 ]. From this resource, we used 56,827 drug-response experiments profiling responses across diverse cell types and conditions. Gene-level significance filtering (p ≤ 0.05) was applied during signature generation to retain only significantly perturbed genes; for connectivity scoring, the retained log2 fold-change values were converted to genome-wide ranks, which served as input for the nonparametric KS-based scoring procedure (see below). No additional experiment-level filtering was required given the resource’s inherently high data quality. After mapping to a shared gene universe, 22,168 genes were available for analysis, and the library covered 50 cell lines and 379 unique drugs. The two libraries provide largely non-overlapping chemical space: CMap and Tahoe-100M share 85 drugs and 12,544 genes but no cell lines (Fig. 1 F, Table 1 ). Table 1 Comparison of drug signature characteristics between CMap and Tahoe-100M Category CMap Tahoe-100M Shared Unique drugs 1,309 379 85 Unique genes 13,071 25,084 12,544 Cell lines 5 50 0 Gene matrix dimensions ~ 6,100 × 13,071 56,827 × 62,710 — Technology Microarray HG U133A Single-cell RNA seq with aggregated pseudo bulk matrices Total profiles 6,100 experiment profiles > 100,000,000 single cell profiles and 56,827 aggregated experiment profiles Expression values per experiment ranked probe set signatures basemean, log2FC, p value, p adjusted CDRPipe Computational Pipeline Disease signatures were standardized by mapping gene symbols to Entrez IDs using g:Profiler (gconvert function, organism = “hsapiens”, target = “ENTREZGENE_ACC”). Genes were filtered by adjusted p-value ( 0.25). Only genes present in the respective drug perturbation database gene universe were retained. For each disease signature and drug perturbation profile, we computed a modified Kolmogorov-Smirnov (KS) connectivity score quantifying transcriptional reversal. Briefly, disease up-regulated genes were expected to be ranked low in drug signatures (indicating downregulation by drug), while disease down-regulated genes were expected to be ranked high (Additional file 1: Algorithm S1). The connectivity score combines KS statistics for up and down gene sets: Score = KS up − KS down Where Score = 0 if sgn(KS up ) = sgn(KS down ) and both exist. Negative scores indicate transcriptional reversal (drug reverses disease signature); positive scores indicate mimicry. Statistical significance was assessed using an empirical null distribution generated from 100,000 random gene sets matched in size to the query disease signature. Two-sided p-values were computed as the proportion of random scores with absolute value exceeding the observed score. To prevent zero p-values, we applied the permutation-based minimum: p = 1/(N + 1) when no random scores exceeded the observed value [ 21 ]. False discovery rate correction was performed using the qvalue package; when qvalue failed due to distribution assumptions, Benjamini-Hochberg adjustment was applied. Drug candidates were identified using a significance threshold of q < 0.05 combined with negative connectivity score (reversal direction). For each disease, the most significant instance per drug (minimum score across cell lines and concentrations) was retained to avoid duplicate counting. Validation Against Known Drug-Disease Associations To validate our pipeline against established clinical knowledge, we obtained known drug-disease associations from the Open Targets Platform (2024 release) [ 19 ]. Disease entities from CREEDS were systematically mapped to Open Targets using a hierarchical approach, identifying 151 diseases (64.8%) through exact name matching and 52 diseases (22.3%) via ontology synonyms, while 30 diseases (12.9%) remained unmatched (Table 2 ). This yielded 203 diseases with at least one known therapy for validation. We included disease-drug pairs with evidence of any recorded development activity, including drugs in clinical trials across all phases; in total, Open Targets provided over 118,000 associations for these conditions, of which 2,668 disease-drug pairs involved drugs present in the CMap or Tahoe-100M libraries and were evaluated in our analysis. Table 2 Classification of disease signatures by therapeutic area Therapeutic Area N % Examples Cancer/Tumor 27 15.0% Lung adenocarcinoma, Colorectal cancer, Melanoma, Glioblastoma multiforme, Ovarian cancer Genetic/Congenital 20 11.1% Cystic fibrosis, Duchenne muscular dystrophy, Down syndrome, Huntington's disease, Marfan syndrome Nervous System 18 10.0% Alzheimer's disease, Multiple sclerosis, Epilepsy syndrome, Amyotrophic lateral sclerosis,.. Immune System 13 7.2% Rheumatoid arthritis, Psoriasis, Sjögren's syndrome, Ankylosing spondylitis, Psoriatic arthritis Gastrointestinal 11 6.1% Ulcerative colitis, Crohn's disease, Barrett's esophagus, NASH, Gastroesophageal reflux disease Musculoskeletal 11 6.1% Osteoarthritis, Systemic lupus erythematosus, Scleroderma, Osteoporosis, Arthritis Cardiovascular 9 5.0% Myocardial infarction, Hypertension, Atherosclerosis, Dilated cardiomyopathy, Pulmonary hypertension Infectious Disease 9 5.0% Tuberculosis, Influenza, Hepatitis C, HIV infectious disease, Dengue disease Respiratory 8 4.4% Asthma, Chronic obstructive pulmonary disease, Allergic asthma, Acute lung injury, Bronchopulmonary dysplasia Hematologic 7 3.9% Sickle cell anemia, Aplastic anemia, Myelodysplastic syndrome, Chronic myeloid leukemia, Acute T cell leukemia Psychiatric 7 3.9% Schizophrenia, Autism spectrum disorder, Alcohol abuse, Nicotine dependence, Bipolar disorder Reproductive/Breast 7 3.9% Breast cancer, Endometriosis, Prostate cancer, Endometrial cancer, Polycystic ovary syndrome Endocrine System 6 3.3% Polycystic ovary syndrome, Papillary thyroid carcinoma, Adrenoleukodystrophy, Androgen insensitivity syndrome Phenotype 6 3.3% Ischemic stroke, Septic shock, Hypoxia, Dystonia, Intellectual disability Urinary System 5 2.8% Renal cell carcinoma, Cystitis, Nephroblastoma, Interstitial cystitis, Diabetic nephropathy Skin/Integumentary 5 2.8% Atopic dermatitis, Eczema, Psoriasis vulgaris, Actinic keratosis, Urticaria Metabolic 5 2.8% Diabetes mellitus, Obesity, Morbid obesity, MELAS syndrome, Familial hypercholesterolemia Pancreas 3 1.7% Type 1 diabetes mellitus, Type 2 diabetes mellitus, Pancreatic ductal adenocarcinoma Visual System 2 1.1% Glaucoma, Primary open angle glaucoma Pregnancy/Perinatal 1 0.6% Pre-eclampsia For each disease we defined I as the number of unique drugs predicted, S as the number of predicted drugs validated in Open Targets, and P as the recoverable ceiling (known drugs present in the platform’s library). Precision was calculated as S/I and recall (recovery rate) as S/P. Case Study Analyses For the autoimmune disease case study, 18 autoimmune diseases with available Open Targets matches were analyzed in depth. For each disease, CDRPipe was run independently against both CMap and Tahoe-100M perturbation libraries using the standard pipeline parameters (q < 0.05, negative connectivity score). Drug-level analysis classified each recovered drug by its platform source (CMap only, Tahoe-100M only, or both). Clinical validation was performed by cross-referencing recovered drugs against Phase 4 clinical trial data from Open Targets for each specific disease indication. Recovery rates were compared between platforms using the Wilcoxon signed-rank test, with effect size quantified by Cohen’s d. For the endometriosis case study, six disease signatures from Oskotsky et al. (Unstratified, Stage I-II, Stage III-IV, Proliferative Phase, Early Secretory Phase, Mid-Secretory Phase) were re-analyzed using both CMap and Tahoe-100M. To ensure exact replication of the original findings, the CMap analysis used the parameters reported in Oskotsky et al.: p-value 1.1, random seed = 2009, and 1,000 permutations. The same parameters were then applied to Tahoe-100M for direct comparison. Top candidates were ranked by connectivity score across signatures, and concordance between platforms was assessed by comparing the top 50 candidates from each. Diseases were categorized into nine therapeutic areas using keyword matching: Oncology, Metabolic, Neurodegenerative, Cardiovascular, Infectious, Autoimmune, Allergic/Respiratory, Rare/Genetic, and Other. Category-level performance was summarized as mean known drug hits per disease. Statistical Analysis All statistical analyses were performed in R (version 4.2+). Paired comparisons of Tahoe-100M versus CMap recovery rates were performed using the Wilcoxon signed-rank test. Effect sizes were quantified using Cohen’s d. P-values < 0.05 were considered statistically significant. For the overall recovery rate comparison, 153 diseases with evaluable data for both platforms were included in the paired analysis. Complementarity was quantified as the percentage of recovered drugs identified by both platforms, Category-level recommendations were based on mean known drug hits with a threshold of > 2 drugs difference for primary method recommendation. Results Study Overview To systematically evaluate the impact of cellular resolution on drug repurposing, we integrated heterogeneous transcriptional resources into a unified analytical framework (Fig. 1 ). We applied CDRPipe to 233 curated CREEDS disease signatures, interrogating each against two complementary drug perturbation libraries with distinct measurement technologies: the microarray-based Connectivity Map (CMap) and the single-cell-derived Tahoe-100M atlas (Table 1 ). CDRPipe applies uniform preprocessing and a nonparametric connectivity score to quantify how strongly each drug reverses a given disease signature, with significance assessed by empirical permutation-based null models (see Methods). Predictions for all 233 signatures were benchmarked against known drug-disease associations from Open Targets (203 evaluable diseases; 2,668 assessable drug-disease pairs). The following sections report recovery performance, cross-platform complementarity, and disease-specific case studies. Prediction Recovery and Validation We assessed performance using precision and recall against Open Targets (defined in the Methods). Because Open Targets captures drugs with any recorded development activity, including early-phase clinical trial candidates, and is biased toward compounds with established clinical annotation, these metrics reflect recovery of documented associations within the benchmark rather than absolute predictive accuracy. To account for overlapping annotations, diseases sharing identical therapeutic area combinations in Open Targets were grouped. Tahoe-100M recovered more annotated drug-disease pairs than CMap (2,198 vs. 948), a 2.3-fold difference. Tahoe-100M also covered more diseases (171 vs. 155) and achieved a higher overall recovery rate (22.1% versus 18.1%). These differences partly reflect the composition of each drug library: Tahoe-100M contains many newer, clinically oriented small-molecule compounds that may have inherently higher overlap with Open Targets annotations, whereas CMap’s library includes older and more exploratory chemical entities. Applying these per-disease metrics, lung adenocarcinoma illustrates the approach: Tahoe-100M predicted 61 unique drugs (I = 61), 10 of which were validated (S = 10), yielding a precision of 16.4%; of the 12 recoverable known drugs present in the library (P = 12), the pipeline recovered 10, achieving 83.3% recall. At the disease level across all 233 evaluated diseases, Tahoe-100M recovered a higher proportion of Open Targets-annotated drug-disease pairs than CMap (Fig. 2 ). Mean precision was 1.8% (SD 3.3%) for Tahoe-100M versus 1.3% (SD 2.8%) for CMap (measured across n = 223 and n = 208 diseases with at least one prediction, respectively) (Fig. 2 A). Among diseases with recoverable drugs (P > 0), Tahoe-100M achieved a mean recall of 47.3% (SD 38.5%) compared to 18.5% (SD 25.1%) for CMap (n = 156 and n = 165 diseases, respectively; Fig. 2 B). These differences should be interpreted cautiously: they partly reflect the greater overlap between Tahoe-100M’s drug library and Open Targets annotations rather than representing a direct measure of predictive superiority. Nonetheless, many diseases achieved > 20% recall with Tahoe-100M compared to substantially fewer with CMap, and multiple diseases achieved > 50% recall with Tahoe-100M, a threshold not reached by any disease in CMap analyses (Fig. 2 C). Variable sample sizes reflect platform-specific prediction coverage and the filtering requirements for each metric (I > 0 for precision calculations and P > 0 for recall calculations). Complementarity of Platforms and Mechanistic Profiles Analysis of the overlap between drug-disease pairs predicted by each platform revealed low concordance (Fig. 2 D). Among drug-disease pairs supported by Open Targets, only 144 were identified by both CMap and Tahoe-100M (Jaccard index 4.8%). When considering all predicted pairs, including those not supported by Open Targets, the overlap remained limited (173 pairs; Jaccard index 1.2%). Despite both platforms being applied to the same set of 233 disease signatures, the majority of predicted pairs were unique to a single platform, with each exhibiting different recovery volumes and enrichment patterns across therapeutic areas and drug target classes (Fig. 2 D, 2 E). This divergence is consistent with the limited chemical space overlap between the two drug libraries (Fig. 1 F) and is further reflected in the distinct disease-to-target combinations favored by each platform (Fig. 2 F). Together, these results indicate that CMap and Tahoe-100M sample largely non-overlapping therapeutic spaces rather than redundantly identifying the same candidates, and that the primary value of integrating both platforms lies in this complementarity. While both platforms validated predictions across diverse therapeutic areas, with cancer/tumor being the largest category for both, Tahoe-100M’s recovered predictions were significantly enriched for oncology, with 28.7% of Tahoe-100M’s recovered drug-disease pairs falling in oncology compared to 14.3% for CMap (Fig. 2 D– 2 F), consistent with its enzyme-centric drug library bias (Fig. 1 C). Biological Concordance of Novel Predictions To determine if novel predictions adhered to the same therapeutic logic as clinically validated drugs, we evaluated the biological concordance between “recovered” (validated) predictions and “all” discoveries by comparing their joint distributions across disease categories and drug target classes. Heatmaps of these distributions revealed that Tahoe-100M’s differential discovery patterns were highly similar between recovered and all predictions (Fig. 2 G, 2 H), yielding a cosine similarity of 0.987 and a Jensen-Shannon divergence of 0.150 (Fig. 2 I, 2 L). This consistency indicates that Tahoe-100M’s novel candidates closely mirror the mechanistic profile of its validated successes. CMap showed lower concordance between its recovered and all discovery distributions (cosine similarity = 0.850; Jensen-Shannon divergence = 0.221; Fig. 2 I, 2 L). The butterfly chart of target class distributions (Fig. 2 J) shows that CMap exhibits notable shifts between the “all discoveries” and “recovered” profiles, particularly for membrane receptors, whereas Tahoe-100M distributions remain consistent. This is quantified in the lollipop chart (Fig. 2 K), where CMap exhibits large percentage-point deviations between total discovery and validated recovery (e.g., − 18.9% for membrane receptors), while Tahoe-100M clusters tightly within a ± 2% shift zone. This higher divergence for CMap implies that it explores a broader chemical space, potentially identifying mechanistically novel relationships that differ from established clinical precedents. We note that these platform-specific mechanistic profiles partly reflect the inherent composition of each drug library (Fig. 1 B, 1 C): Tahoe-100M’s enzyme-enriched library naturally yields enzyme-centric predictions, while CMap’s receptor-heavy library favors receptor-targeted discoveries. Case Study 1: Autoimmune Diseases To evaluate pipeline performance, we analyzed 18 autoimmune diseases mapped to Open Targets via exact name or synonym matching, details about the diseases available in Additional file 1: Table S1 . Note that the “known drugs” column in Additional file 1: Table S1 reflects all drugs with any recorded development activity for that indication in Open Targets (including clinical trial candidates at all phases), rather than only approved therapies; this accounts for the high counts observed for some diseases. Autoimmune diseases constitute a rigorous, biologically grounded benchmark for transcriptional drug repurposing, as their pathogenesis is driven by aberrant immune activation that manifests as reproducible gene expression signatures [ 22 ]. Despite clinical heterogeneity, these disorders converge on shared inflammatory axes, including interferon signaling, cytokine-mediated activation, and T-cell dysregulation, allowing for a systematic evaluation of the pipeline’s ability to recover established immunomodulators [ 23 ]. Furthermore, the persistence of treatment resistance in autoimmune conditions underscores the clinical urgency for scalable approaches that nominate candidates based on molecular reversal rather than phenotype alone. As shown in Fig. 3 B, Tahoe-100M significantly outperformed CMap in recovering known therapeutics in the Open Targets database (mean recovery rate: 76.5% vs. 20.9%; Wilcoxon signed-rank test p < 0.001; Cohen’s d = 2.22), representing a 3.7-fold improvement (Fig. 3 ). Disease-specific analysis revealed Tahoe-100M’s particular strength in inflammatory conditions: Crohn’s disease (75.0% vs 7.4%), psoriasis (100.0% vs 75.0%), and rheumatoid arthritis (64.0% vs 17.8%) (Fig. 3 A). However, CMap demonstrated better performance for type 1 diabetes mellitus, which may reflect the fact that its treatment paradigm centers on glucose regulation rather than immunosuppression, potentially favoring CMap’s drug library composition. We further validated these findings by examining disease-specific Phase 4 clinical trial data (Fig. 3 C). Of the recovered drug-disease pairs, 26.3% (46 of 175) involved drugs validated in Phase 4 specifically for that indication. The most frequently recovered agents were cornerstone immunosuppressants, including dexamethasone (identified in 12 diseases), methotrexate (11 diseases), and hydrocortisone (7 diseases). Several robust disease-specific patterns emerged: rheumatoid arthritis demonstrated the highest concordance with Phase 4 data, recovering 15 distinct Phase 4-approved drugs. This is consistent with rheumatoid arthritis having one of the broadest therapeutic arsenals among autoimmune diseases, providing a larger validation set. Both platforms consistently predicted methotrexate, celecoxib, naproxen, dexamethasone, and meloxicam, representing greater therapeutic validation than other autoimmune conditions evaluated. Similarly, the psoriasis spectrum (including psoriatic arthritis) showed consistent recovery of dimethyl fumarate, tazarotene, and clobetasol propionate. Gastrointestinal autoimmune diseases (Crohn’s disease and ulcerative colitis) preferentially recovered locally-acting corticosteroids such as budesonide. Sjögren’s syndrome recovered pilocarpine and leflunomide, reflecting its distinct pathophysiology centered on exocrine gland dysfunction. Despite Tahoe-100M’s overall superiority, the two methods showed high complementarity with only 4.0% overlap in recovered drug-disease pairs. Across 175 recovered drug-disease pairs, Tahoe-100M uniquely contributed 104 (59.4%), while CMap uniquely identified 64 (36.6%) that Tahoe-100M missed. Beyond validated associations, both platforms generated substantial novel predictions: CMap produced 4,900 novel (unvalidated) drug-disease pairs involving 198 drugs that had no Open Targets confirmation, while Tahoe-100M produced 8,766 novel pairs involving 32 exclusively novel drugs, a difference reflecting Tahoe-100M’s smaller but more clinically oriented library, where most drugs already have some Open Targets annotation (Fig. 3 F, 3 G). Drugs predicted by both methods showed 2.6-fold higher precision (5.2%) compared to single-method predictions (1.8–2.0%), suggesting consensus predictions represent higher-confidence repurposing candidates. Across the full panel of 18 autoimmune diseases, we identified 94 unique recovered drugs representing the combined outputs of both databases. Only 7 of 175 drug-disease pairs (4.0%) were identified by both platforms, reinforcing that distinct signals are captured by CMap and Tahoe-100M (Fig. 3 C, 3 D). Qualitative analysis of the recovered candidates revealed distinct platform-specific therapeutic spaces. CMap excelled at recovering established medications, including NSAIDs, corticosteroids, and traditional immunosuppressants. In contrast, Tahoe-100M uniquely identified emerging targeted therapies, including JAK inhibitors (tofacitinib, filgotinib) across four diseases, BTK inhibitors (tirabrutinib), and SGLT2 inhibitors for diabetes. This difference partly reflects the composition of each drug library: Tahoe-100M includes newer, clinically oriented compounds not present in CMap, while CMap’s library is enriched for older, well-established pharmacological agents. These findings underscore the value of a multi-platform approach: each platform captures a distinct pharmacological space, and their integration maximizes overall therapeutic coverage. While the volume of novel candidates is substantial, all predictions are rank-ordered by their connectivity score, providing a principled prioritization for downstream evaluation. The high candidate counts partly reflect the use of uniform statistical thresholds across all 18 diseases; disease-specific parameter tuning could narrow the output but risks introducing subjective bias. Importantly, the ranking ensures that the most promising candidates, those with the strongest expression-reversal signatures, appear at the top, enabling systematic experimental follow-up in order of predicted efficacy rather than requiring exhaustive evaluation of the full list. Case Study 2: Endometriosis To further validate our pipeline, we applied Tahoe-100M to endometriosis, benchmarking our results against the prior CMap-based analysis by Oskotsky et al [ 6 ]. Endometriosis presents a unique opportunity for computational discovery due to the critical unmet need for non-hormonal therapies and the historical stagnation of traditional research and development [ 24 ]. Biologically, however, it is an ideal candidate for this approach as the disease shares significant molecular hallmarks with malignancy, including invasion, angiogenesis, and metastasis [ 25 ]. This transcriptional overlap is particularly advantageous given that the reference data in both Tahoe-100M and CMap are derived largely from cancer cell lines. We began by replicating the findings of Oskotsky et al [ 6 ] using the CDRPipe framework. We achieved 100% replication of their reported results using the original study parameters: a p-value less than 0.05, an absolute log2 fold change greater than 1.1, a random seed of 2009, and 1,000 permutations (Fig. 4 A). Following this validation, we applied the same rigorous parameters to the Tahoe-100M pipeline (Fig. 4 B). While Oskotsky et al. stratified signatures by disease stage and menstrual phase, we focused our comparative ranking on the “Unstratified” signature to identify robust, broad-spectrum candidates. We subsequently evaluated the concordance between the two platforms by analyzing the top 50 drug candidates identified by each. The overlap was modest, highlighting the distinct yet complementary nature of the libraries. Among the top 50 hits in CMap, only six were present in Tahoe-100M, and conversely, only seven of the top 50 Tahoe hits were present in CMap. Despite these differences, we identified two drugs in common within the top rankings: terfenadine and irinotecan. This divergence suggests that the platforms capture different aspects of the disease biology. Tahoe-100M is heavily enriched for oncology and signaling inhibitors, surfacing numerous kinase inhibitors and epigenetic regulators such as FGFR inhibitors (pemigatinib, erdafitinib, futibatinib), CDK/AKT pathway drugs (dinaciclib, ipatasertib), and HDAC inhibitors (panobinostat, tucidinostat). This profile aligns well with modern characterizations of endometriosis as a proliferative, invasive, and inflammatory disease with cancer-like transcriptomic properties[ 26 , 27 ]. In contrast, CMap leans toward symptom modulation, prioritizing hormonal agents (levonorgestrel, medrysone), anti-inflammatory drugs (mesalazine, fenoprofen, resveratrol), and neuropsychiatric compounds (clomipramine). The clinical relevance of these findings is supported by strong literature evidence for top hits from both pipelines. Both platforms successfully recovered established therapies: CMap identified levonorgestrel, a hormonal (progestin) therapy commonly used to treat endometriosis related pain, administered orally, as implants, and used in IUDs [ 28 – 30 ], while Tahoe-100M identified medroxyprogesterone acetate, which is widely used for hormonal suppression of lesions [ 31 ]. Beyond standard of care, Tahoe-100M identified pentoxifylline as the top candidate common across all six disease signatures; its potential to prevent endometriosis recurrence after conservative surgery has been supported by controlled trials [ 32 , 33 ]. Similarly, CMap identified fenoprofen and resveratrol, both of which have documented efficacy in reducing lesion size and angiogenesis in animal models [ 6 , 34 ]. Several novel hits also show mechanistic promise; simvastatin (CMap) has been reported to inhibit stromal cell proliferation in vitro [ 35 ], while valproic acid (CMap) has shown potential in preclinical models to inhibit lesion growth [ 26 , 27 , 36 ]. Furthermore, Tahoe-100M identified drospirenone, a progestin with anti-androgenic and anti-inflammatory properties relevant to pain control [ 37 ]. While toxicity limits the clinical application of some hits, such as the topoisomerase inhibitor irinotecan (found in both pipelines), their identification underscores the pipeline’s ability to correctly identify antiproliferative mechanisms relevant to the disease pathology. Discussion The principal finding of this study is the high complementarity between CMap and Tahoe-100M: only 3.5% of recovered drugs were identified by both platforms, yet each captured a distinct and clinically relevant segment of the therapeutic landscape. While Tahoe-100M uniquely recovered 64.7% of candidates, capturing emerging targeted therapies like JAK and BTK inhibitors, CMap contributed 31.8% of hits, excelling at established therapeutics such as NSAIDs and corticosteroids. When benchmarked against Open Targets annotations, Tahoe-100M recovered a greater proportion of documented drug-disease associations than CMap (47.3% vs. 18.5%; 2.6-fold difference), with pronounced differences in autoimmune (4.3-fold) and oncology (3.0-fold) categories. These differences in annotated recovery rates should be interpreted cautiously, however: they partly reflect the greater overlap between Tahoe-100M’s drug library and Open Targets annotations, and the larger experimental breadth of Tahoe-100M (56,827 vs. 1,968 experiments capturing greater pharmacological and compound diversity), rather than representing a direct measure of superior predictive accuracy. This divergence is driven by distinct mechanistic biases. Tahoe-100M displays a strongly “enzyme-centric” profile (44–46% of discoveries), reflecting its sensitivity to kinase cascades that create coherent transcriptomic signatures in oncology and inflammatory conditions. In contrast, CMap exhibits a “receptor-oriented” profile (35.8% membrane receptors), effectively capturing the rapid, membrane-proximal effects of agonists and ion channel modulators. These findings demonstrate that integrating heterogeneous perturbation resources is essential to maximize therapeutic coverage Our disease category analysis provides actionable guidance for method selection. Tahoe-100M is recommended as the primary approach for oncology, autoimmune, and cardiovascular drug repurposing, where it demonstrated consistent 1.5–3.0-fold advantages. For metabolic and neurodegenerative diseases, CMap showed modest advantages, suggesting these categories may benefit from CMap-first or complementary approaches. The endometriosis case study illustrates an important caveat: for hormone-responsive conditions with established CMap-derived findings, CMap remains optimal for replication (62.5% recovery of Oskotsky et al.’s drugs), while Tahoe-100M provides value as a complementary discovery tool contributing unique candidates. This disease-specific performance pattern highlights that no single drug signature dataset is universally superior; method selection should be tailored to the disease biology and study objectives. Our findings have several implications for computational drug repurposing practice; the first is that consensus predictions merit prioritization. Drugs identified by both CMap and Tahoe-100M showed 2.6-fold higher precision than single-method predictions, suggesting consensus hits represent higher-confidence candidates for experimental validation. Secondly, multi-method approaches maximize coverage. The high complementarity between methods (only 3.5–7.6% overlap in recovered drugs) indicates that single-method studies substantially underestimate the drug repurposing landscape. Integrating both CMap and Tahoe-100M captures distinct therapeutic signals. Lastly, platform-specific strengths inform therapeutic focus. CMap’s strength in established drug classes and Tahoe-100M’s identification of emerging targeted therapies suggest that platform selection can be tuned to study objectives, validation of known mechanisms versus discovery of novel interventions. Our comparative analysis informs evidence-based platform selection for drug repurposing campaigns. Tahoe-100M is recommended for enzyme-driven pathologies (oncology, inflammation, metabolic disorders) and programs requiring high precision and biological concordance. CMap is optimal for receptor-mediated pathways (neurology, psychiatry, cardiovascular) and exploratory programs targeting ion channels or transporters. Given the minimal overlap (1.2%) yet distinct mechanistic strengths, we recommend an integrated strategy. Dual-platform hits represent high-confidence candidates suitable for prioritized validation, while platform-specific predictions should be pursued based on alignment between the disease mechanism and the platform’s established target-class bias. Several limitations should be considered. First, validation relied on Open Targets, which creates a bias toward approved and late-stage candidates while missing off-label or early-stage therapeutics. Our definition of a true association, any drug evaluated for a given disease, favors sensitivity but may obscure specificity regarding efficacy. Second, despite standardized processing, disease signatures from CREEDS retain heterogeneity due to diverse experimental designs and tissue sources in the underlying GEO data. Third, our approach evaluates transcriptional reversal, which does not guarantee a specific therapeutic mechanism; observed reversals could stem from generalized stress responses rather than disease-specific correction. Fourth, many identified compounds, particularly from CMap, could not be mapped to Open Targets. This reflects the presence of proprietary or research-grade chemicals in CMap that lack public annotation, potentially leading to an underestimation of their relevance. Finally, the perturbation datasets differ technically: CMap uses bulk microarray profiling on limited cell lines, while Tahoe-100M utilizes single-cell RNA sequencing across a wider cellular panel. Although both rely heavily on neoplastic lines, they capture conserved biological processes relevant to non-cancer contexts. The superior resolution of the Tahoe-100M likely drives its improved precision. Ultimately, our findings suggest that CMap and Tahoe-100M are complementary, and platform selection should be driven by the specific translational context. While this study identifies thousands of drug repurposing candidates validated against known therapeutics, bridging the gap to clinical application requires a rigorous translational framework [ 38 ]. Top-ranked candidates, particularly those achieving “triple-validation” across CMap, Tahoe-100M, and prior literature, must first undergo experimental verification in high-fidelity models, such as patient-derived organoids or microphysiological systems, to confirm efficacy beyond standard cell lines [ 39 – 41 ]. These efforts should be complemented by deep mechanistic profiling to ensure that signature reversal is driven by therapeutically relevant target engagement rather than off-target toxicity, a step particularly critical for drugs with polypharmacology [ 42 ]. For candidates already approved for other indications, this translation can be accelerated by systematically reviewing real-world evidence and electronic health records to de-risk safety profiles in new patient populations [ 43 ]. Furthermore, our observation that Tahoe-100M achieves variable recovery rates across autoimmune conditions underscores the critical need for transcriptomic patient stratification to identify responsive subtypes [ 44 ]. Ultimately, randomized clinical trials remain the definitive gold standard; future studies must transition validated candidates into Phase II proof-of-concept or innovative basket trials to establish efficacy and safety in human subjects prior to broad clinical implementation [ 45 ]. Conclusions This comprehensive evaluation demonstrates that Tahoe-100M and CMap provide complementary perspectives on the drug repurposing landscape, with Tahoe-100M showing overall superior recovery of known therapeutics and particular strength in oncology and autoimmune conditions. The high complementarity between methods, with only 3.5–7.6% overlap in recovered drugs, argues for multi-method approaches as standard practice in computational drug repurposing. Consensus predictions from both platforms represent high-confidence candidates meriting experimental prioritization. Our disease category-level recommendations and case study analyses provide actionable guidance for method selection across diverse therapeutic areas. Abbreviations AKT protein kinase B BTK Bruton tyrosine kinase CDK cyclin-dependent kinase CDRPipe Computational Drug Repurposing Pipeline CMap Connectivity Map CREEDS Crowd-Extracted Expression of Differential Signatures FDR false discovery rate FGFR fibroblast growth factor receptor GEO Gene Expression Omnibus HDAC histone deacetylase JAK Janus kinase KS Kolmogorov-Smirnov NSAID nonsteroidal anti-inflammatory drug QC quality control R&D research and development scRNA-seq single-cell RNA sequencing SD standard deviation SGLT2 sodium-glucose cotransporter 2. Declarations Ethics approval and consent to participate Not applicable. This study analysed only publicly available, previously published, de-identified datasets and did not involve new experiments on human participants, human tissue, or animals. Consent for publication Not applicable. Competing interests The authors declare that they have no competing interests. Funding This work was supported by the National Institutes of Health (grants 1R21HD114953 to MS and LCG; 1R01AI180118 to MS; P30 AR070155 to MS; P01HD106414 to LCG; and 1R01AG100879-0 to MS) and by the UCSF March of Dimes Prematurity Research Center. The funders had no role in the design of the study; the collection, analysis, and interpretation of the data; or the writing of the manuscript. Author Contribution EN, UK, and MS conceptualised the study. EN, UK, MS, JN, LCG, and TO developed the methodology. EN performed the investigation. EN and BO developed the software. EN, UK, XT, LA, EC, BA, and CPS performed data curation and formal analysis. EN prepared the visualisations. MS supervised the study. MS and LCG acquired funding. EN and MS wrote the original draft. UK, XT, LA, EC, BA, CPS, BO, BG, DKS, JN, LCG, and TO reviewed and edited the manuscript. All authors read and approved the final manuscript. Acknowledgement The authors thank the CREEDS, Connectivity Map, Tahoe-100M, and Open Targets teams for making their data resources publicly available. Data Availability All data analysed in the current study are publicly available. The CDRPipe R package and R Shiny application are available at https://github.com/enockniyonkuru/drug_repurposing. The comparative analysis code supporting this manuscript is available at https://github.com/enockniyonkuru/cdrpipe-comparative-analysis. The R Shiny application is freely accessible at https://cdrpipe.org/. Disease signatures from CREEDS are available at https://maayanlab.cloud/CREEDS/. Known drug-disease associations from the Open Targets Platform (2024 release) are available at https://platform.opentargets.org/. Connectivity Map (CMap) perturbation profiles are available from the Broad Institute at https://www.broadinstitute.org/connectivity-map-cmap. Tahoe-100M perturbation profiles are available from Hugging Face at https://huggingface.co/datasets/tahoebio/Tahoe-100M. References Gamal H, Shoeib EM, Hajjaj A, Abdullah HEA, Elramy EH, Ellah DAA, et al. Incorporating AI, in silico, and CRISPR technologies to uncover the potential of repurposed drugs in cancer therapy. RSC Pharm. 2025;2:1019–33. Camps I, Künzel SR, Schubert M, Singh RK. Editorial: opportunities and challenges in drug repurposing. Front Pharmacol. 2025;16:1709217. Sun D, Gao W, Hu H, Zhou S. Why 90% of clinical drug development fails and how to improve it? Acta Pharm Sin B. 2022;12:3049–62. Golla U, Patel S, Shah N, Talamo S, Bhalodia R, Claxton D, et al. From deworming to cancer therapy: benzimidazoles in hematological malignancies. Cancers. 2024;16:3454. Dudley JT, Sirota M, Shenoy M, Pai RK, Roedder S, Chiang AP, et al. Computational repositioning of the anticonvulsant topiramate for inflammatory bowel disease. Sci Transl Med. 2011;3:96ra76. Oskotsky TT, Bhoja A, Bunis D, Le BL, Tang AS, Kosti I, et al. Identifying therapeutic candidates for endometriosis through a transcriptomics-based drug repositioning approach. iScience. 2024;27:109388. Sirota M, Schaub MA, Batzoglou S, Robinson WH, Butte AJ. Autoimmune disease classification by inverse association with SNP alleles. PLoS Genet. 2009;5:e1000792. Chen K, Wang H. Discovery of disease relationships via transcriptomic signature analysis powered by agentic AI. arXiv. 2025. https://doi.org/10.48550/arXiv.2508.04742 Iorio F, Rittman T, Ge H, Menden M, Saez-Rodriguez J. Transcriptional data: a new gateway to drug repositioning? Drug Discov Today. 2013;18:350–7. Sirota M, Dudley JT, Kim J, Chiang AP, Morgan AA, Sweet-Cordero A, et al. Discovery and preclinical validation of drug indications using compendia of public gene expression data. Sci Transl Med. 2011;3:96ra77. Musa A, Ghoraie LS, Zhang S-D, Glazko G, Yli-Harja O, Dehmer M, et al. A review of connectivity map and computational approaches in pharmacogenomics. Brief Bioinform. 2017;19:506–23. Lamb J, Crawford ED, Peck D, Modell JW, Blat IC, Wrobel MJ, et al. The Connectivity Map: using gene-expression signatures to connect small molecules, genes, and disease. Science. 2006;313:1929–35. Smith I, Scott K, Haibe-Kains B. Similarity bias from consensus perturbational signatures from the L1000 Connectivity Map. bioRxiv. 2024. https://doi.org/10.1101/2022.01.24.477615 Aissa AF, Islam ABMMK, Ariss MM, Go CC, Rader AE, Conrardy RD, et al. Single-cell transcriptional changes associated with drug tolerance and response to combination therapies in cancer. Nat Commun. 2021;12:1628. Peidli S, Green TD, Shen C, Gross T, Min J, Garda S, et al. scPerturb: harmonized single-cell perturbation data. Nat Methods. 2024;21:531–40. Zhang J, Ubas AA, de Borja R, Svensson V, Thomas N, Thakar N et al. Tahoe-100M: a giga-scale single-cell perturbation atlas for context-dependent gene function and cellular modeling. bioRxiv. 2025. https://doi.org/10.1101/2025.02.20.639398 Lei W, Yuan M, Long M, Zhang T, Huang Y-E, Liu H, et al. scDR: predicting drug response at single-cell resolution. Genes. 2023;14:268. Wang Z, Monteiro CD, Jagodnik KM, Fernandez NF, Gundersen GW, Rouillard AD, et al. Extraction and analysis of signatures from the Gene Expression Omnibus by the crowd. Nat Commun. 2016;7:12846. Carvalho-Silva D, Pierleoni A, Pignatelli M, Ong C, Fumis L, Karamanis N, et al. Open Targets Platform: new developments and updates two years on. Nucleic Acids Res. 2019;47:D1056–65. Barrett T, Wilhite SE, Ledoux P, Evangelista C, Kim IF, Tomashevsky M, et al. NCBI GEO: archive for functional genomics data sets - update. Nucleic Acids Res. 2013;41:D991–5. Phipson B, Smyth GK. Permutation P-values should never be zero: calculating exact P-values when permutations are randomly drawn. Stat Appl Genet Mol Biol. 2010;9. https://doi.org/10.2202/1544-6115.1585 . :Article 39. Petitdemange A, Blaess J, Sibilia J, Felten R, Arnaud L. Shared development of targeted therapies among autoimmune and inflammatory diseases: a systematic repurposing analysis. Ther Adv Musculoskelet Dis. 2020;12:1759720X20969261. Sumida TS, Lincoln MR, He L, Park Y, Ota M, Oguchi A, et al. An autoimmune transcriptional circuit drives FOXP3 + regulatory T cell dysfunction. Sci Transl Med. 2024;16:eadp1720. Zondervan KT, Becker CM, Koga K, Missmer SA, Taylor RN. Viganò P. Endometriosis. Nat Rev Dis Primers. 2018;4:9. Anglesio MS, Papadopoulos N, Ayhan A, Nazeran TM, Noë M, Horlings HM, et al. Cancer-associated mutations in endometriosis without cancer. N Engl J Med. 2017;376:1835–48. Lin Y-H, Chen Y-H, Chang H-Y, Au H-K, Tzeng C-R, Huang Y-H. Chronic niche inflammation in endometriosis-associated infertility: current understanding and future therapeutic strategies. Int J Mol Sci. 2018;19:2385. Perrone U, Barra F, Anatrà M, Paudice M, Vellone VG, Gullo G, et al. Targeting inflammation in endometriosis: emerging therapeutic options. Expert Opin Investig Drugs. 2025;34:995–1009. Gibbons T, Georgiou EX, Cheong YC, Wise MR. Levonorgestrel-releasing intrauterine device (LNG-IUD) for symptomatic endometriosis following surgery. Cochrane Database Syst Rev. 2021;2021:CD005072. Giudice LC, Liu B, Irwin JC. Endometriosis and adenomyosis unveiled through single-cell glasses. Am J Obstet Gynecol. 2025;232:S105–23. Becker CM, Bokor A, Heikinheimo O, Horne A, Jansen F, Kiesel L, et al. ESHRE guideline: endometriosis. Hum Reprod Open. 2022;2022:hoac009. Vercellini P et al. [Reference to be confirmed by the authors: cited in text as Vercellini 2003 regarding medroxyprogesterone acetate for hormonal suppression of endometriotic lesions.]. Kamencic H, Thiel JA. Pentoxifylline after conservative surgery for endometriosis: a randomized, controlled trial. J Minim Invasive Gynecol. 2008;15:62–6. Grammatis AL, Georgiou EX, Becker CM. Pentoxifylline for the treatment of endometriosis-associated pain and infertility. Cochrane Database Syst Rev. 2021;8:CD007677. Bruner-Tran KL, Osteen KG, Taylor HS, Sokalska A, Haines K, Duleba AJ. Resveratrol inhibits development of experimental endometriosis in vivo and reduces endometrial stromal cell invasiveness in vitro. Biol Reprod. 2011;84:106–12. Bruner-Tran KL, Osteen KG, Duleba AJ. Simvastatin protects against the development of endometriosis in a nude mouse model. J Clin Endocrinol Metab. 2009;94:2489–94. Liu M, Liu X, Zhang Y, Guo S-W. Valproic acid and progestin inhibit lesion growth and reduce hyperalgesia in experimentally induced endometriosis in rats. Reprod Sci. 2012;19:360–73. Shim JY, Garbo G, Grimstad FW, Scatoni A, Barrera EP, Boskey ER. Use of the drospirenone-only contraceptive pill in adolescents with endometriosis. J Pediatr Adolesc Gynecol. 2024;37:402–6. Pushpakom S, Iorio F, Eyers PA, Escott KJ, Hopper S, Wells A, et al. Drug repurposing: progress, challenges and recommendations. Nat Rev Drug Discov. 2019;18:41–58. Low LA, Mummery C, Berridge BR, Austin CP, Tagle DA. Organs-on-chips: into the next decade. Nat Rev Drug Discov. 2021;20:345–61. Vlachogiannis G, Hedayat S, Vatsiou A, Jamin Y, Fernández-Mateos J, Khan K, et al. Patient-derived organoids model treatment response of metastatic gastrointestinal cancers. Science. 2018;359:920–6. Oskotsky TT, Tang X, Arthurs E, Govil A, Abbasi F, Bhoja A et al. A transcriptomics-based computational drug repositioning pipeline identifies simvastatin and primaquine as novel therapeutics for endometriosis pain. bioRxiv. 2025. https://doi.org/10.1101/2025.05.28.656743 Moffat JG, Vincent F, Lee JA, Eder J, Prunotto M. Opportunities and challenges in phenotypic drug discovery: an industry perspective. Nat Rev Drug Discov. 2017;16:531–43. Corrigan-Curay J, Sacks L, Woodcock J. Real-world evidence and real-world data for evaluating drug safety and effectiveness. JAMA. 2018;320:867–8. Barturen G, Beretta L, Cervera R, Van Vollenhoven R, Alarcón-Riquelme ME. Moving towards a molecular taxonomy of autoimmune rheumatic diseases. Nat Rev Rheumatol. 2018;14:75–93. Saranraj K, Kiran PU. Drug repurposing: clinical practices and regulatory pathways. Perspect Clin Res. 2025;16:61–8. Additional Declarations No competing interests reported. Supplementary Files CDRPipeAdditionalFile1Supplement.docx Cite Share Download PDF Status: Under Review Version 1 posted Reviewers agreed at journal 23 Aug, 2026 Reviewers invited by journal 23 Aug, 2026 Editor assigned by journal 13 Aug, 2026 Submission checks completed at journal 06 Aug, 2026 First submitted to journal 05 Aug, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-10606217","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":706741300,"identity":"c7376bfb-e90b-4023-b862-116c465d2963","order_by":0,"name":"Enock Niyonkuru","email":"","orcid":"","institution":"University of California, San Francisco","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Enock","middleName":"","lastName":"Niyonkuru","suffix":""},{"id":706741302,"identity":"c3639603-f8ca-45ba-a063-ddcee001a338","order_by":1,"name":"Umair Khan","email":"","orcid":"","institution":"University of California, San Francisco","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Umair","middleName":"","lastName":"Khan","suffix":""},{"id":706741304,"identity":"0302b23f-3f06-4010-8b8e-cc0103f0733a","order_by":2,"name":"Xinyu Tang","email":"","orcid":"","institution":"University of California, San Francisco","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Xinyu","middleName":"","lastName":"Tang","suffix":""},{"id":706741308,"identity":"7e1546d4-d8fe-4f56-b271-e9ccd707e7dc","order_by":3,"name":"Laura Almonte","email":"","orcid":"","institution":"University of California, San Francisco","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Laura","middleName":"","lastName":"Almonte","suffix":""},{"id":706741315,"identity":"9fd18c68-12ed-4f04-8b72-e2378dde25c2","order_by":4,"name":"Eden Chun","email":"","orcid":"","institution":"University of California, San Francisco","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Eden","middleName":"","lastName":"Chun","suffix":""},{"id":706741322,"identity":"418a4222-2488-47d2-83d8-0f03efd7109c","order_by":5,"name":"Brenda Ametepe","email":"","orcid":"","institution":"University of California, San Francisco","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Brenda","middleName":"","lastName":"Ametepe","suffix":""},{"id":706741324,"identity":"73d95094-f431-494f-b9fc-d7eb2322ceb6","order_by":6,"name":"Carlota Pereda Serras","email":"","orcid":"","institution":"University of California, San Francisco","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Carlota","middleName":"Pereda","lastName":"Serras","suffix":""},{"id":706741327,"identity":"eec0672f-e009-409c-bbf0-62d172554226","order_by":7,"name":"Boris Oskotsky","email":"","orcid":"","institution":"University of California, San Francisco","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Boris","middleName":"","lastName":"Oskotsky","suffix":""},{"id":706741332,"identity":"72f71338-d72b-4d01-a6d0-66bd6edfc10a","order_by":8,"name":"Brice Gaudillière","email":"","orcid":"","institution":"Stanford School of Medicine","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Brice","middleName":"","lastName":"Gaudillière","suffix":""},{"id":706741333,"identity":"e3261a1d-1b58-478e-8a61-ac973bb475cb","order_by":9,"name":"David K. Stevenson","email":"","orcid":"","institution":"Stanford University School of Medicine","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"David","middleName":"K.","lastName":"Stevenson","suffix":""},{"id":706741336,"identity":"1698dcf8-474d-4111-b0b9-b4f929cdea9e","order_by":10,"name":"Jessica Neely","email":"","orcid":"","institution":"University of California, San Francisco","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Jessica","middleName":"","lastName":"Neely","suffix":""},{"id":706741338,"identity":"e5a6eef6-a824-4211-9c17-67df9702d830","order_by":11,"name":"Linda C. Giudice","email":"","orcid":"","institution":"University of California, San Francisco","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Linda","middleName":"C.","lastName":"Giudice","suffix":""},{"id":706741341,"identity":"21982c93-818d-41c7-a273-4ce24c181ece","order_by":12,"name":"Tomiko Oskotsky","email":"","orcid":"","institution":"University of California, San Francisco","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Tomiko","middleName":"","lastName":"Oskotsky","suffix":""},{"id":706741344,"identity":"51b7d581-6d5e-4667-80fb-c3ee6a84bc95","order_by":13,"name":"Marina Sirota","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA7klEQVRIiWNgGAWjYJCCAyCCDcysAGJmIJ+HSC2MDQxnQKwEwlpggLGBsY0ILbrtZx8eLqhhsOeTPnz8wc95h+Xl2xgYH7xtw63F7Ey6weEZxxgS2/jSEht7tx023HCMgdlwLj4tB9IYDvOwMSSw8fAYNvBuO5xgIN/AJs2LT8v5Z0At/xjsQVoa/845nAB0GPtvvFpuAG0BKmBsA2pp5m04nMBwjIGNGb8WoC28fRKJbTxsibNljqUD/cLYLDnnHD6HpTF/5vlmYy/fw3zg45saa2CIMR/88KYMtxYokEDmAKN0FIyCUTAKRgFlAADXVku7iu1cAgAAAABJRU5ErkJggg==","orcid":"","institution":"University of California, San Francisco","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Marina","middleName":"","lastName":"Sirota","suffix":""}],"badges":[],"createdAt":"2026-08-05 18:09:19","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-10606217/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-10606217/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":118787477,"identity":"46214fbe-5cfd-4e7f-a579-336dcf3eb6ec","added_by":"auto","created_at":"2026-08-31 10:56:04","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":116942,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eStudy Workflow, Dataset Characteristics, and Quality Control Metrics.\u003c/strong\u003e \u003cstrong\u003e(A)\u003c/strong\u003e The CDRpipe framework integrates transcriptomic drug perturbation libraries from \u003cstrong\u003eCMap\u003c/strong\u003e(microarray-based) and \u003cstrong\u003eTahoe-100M\u003c/strong\u003e (single-cell RNA-seq) with disease signatures from the \u003cstrong\u003eCREEDS\u003c/strong\u003e manual cohort. Following platform-specific preprocessing and quality control (QC), the Comparative Drug Repurposing (CDR) engine predicts therapeutic associations, which are subsequently validated against known indications from \u003cstrong\u003eOpen Targets\u003c/strong\u003e. \u003cstrong\u003e(B)\u003c/strong\u003e Target class distribution for 457 CMap drugs, showing a receptor-oriented profile dominated by membrane receptors (33.4%). \u003cstrong\u003e(C)\u003c/strong\u003e Target class distribution for 170 Tahoe-100M drugs, revealing an enzyme-centric bias (48.8%) reflecting the library’s focus on intracellular signaling. \u003cstrong\u003e(D)\u003c/strong\u003e Comparative signature strength distribution; Tahoe-100M exhibits a shift toward higher fold changes compared to CMap. \u003cstrong\u003e(E)\u003c/strong\u003e Cross-cell line consistency, showing comparable correlation coefficients for compounds profiled across varying lineages. \u003cstrong\u003e(F)\u003c/strong\u003eVenn diagram of the chemical space overlap. Only 36 drugs are shared across all three datasets, underscoring the complementary pharmacological coverage provided by the multi-platform approach. \u003cstrong\u003eG)\u003c/strong\u003e Distribution of the 180 diseases analyzed, categorized by primary therapeutic area. Oncology (n=27) and Genetic/Congenital disorders (n=20) represent the largest cohorts\u003cstrong\u003e. (H)\u003c/strong\u003eCorrelation between up-regulated and down-regulated gene strength. Most diseases cluster along the diagonal, indicating balanced regulatory signatures. \u003cstrong\u003e(I)\u003c/strong\u003e Distribution of gene counts per signature, confirming a balanced representation of up-regulated and down-regulated genes (mean = 209 each).\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-10606217/v1/741e36c89cf37c6d306f043e.png"},{"id":118787545,"identity":"f4d14679-6ffd-4b4e-a27c-4b7b7627cc11","added_by":"auto","created_at":"2026-08-31 10:56:18","extension":"jpeg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":184718,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eComparative Performance Benchmarking and Pharmacological Discovery Patterns.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e(A–C) Precision and Recall Performance Across Diseases\u003c/strong\u003e. Benchmarking of the CDRPipe engine using Open Targets as a gold standard.\u003cstrong\u003e(A) Precision distribution density\u003c/strong\u003e displaying right-skewed profiles for both platforms. Low mean precision (CMap: 1.3%; Tahoe-100M: 1.8%) reflects the inherent challenge of predicting validated relationships from transcriptomic signatures alone.\u003cstrong\u003e(B) Recall distribution density\u003c/strong\u003e for CMap (n=165 diseases) and Tahoe-100M (n=156). Dashed vertical lines indicate mean recall (CMap: 18.5%, Tahoe-100M: 47.3%), illustrating Tahoe-100M's significantly superior recovery of known drug-disease relationships. (\u003cstrong\u003eC) Precision versus Recall scatter plot\u003c/strong\u003e by disease. Tahoe-100M (blue) clusters at higher recall values than CMap (orange) while maintaining comparable precision, demonstrating superior overall recovery of therapeutically validated drug-disease relationships.\u003cstrong\u003e(D–F) Platform-Specific Prediction Profiles\u003c/strong\u003e.\u003cstrong\u003e(D) Recovery by therapeutic area.\u003c/strong\u003eTahoe-100M demonstrates superior recovery volume across most categories, with particularly strong enrichment in Oncology, Gastrointestinal, and Immune System disorders.\u003cstrong\u003e(E) Recovery by drug target class.\u003c/strong\u003e Tahoe-100M displays a distinct mechanistic bias toward Enzymes, reflecting its sensitivity to signaling pathway inhibitors, whereas CMap shows a more balanced distribution across Membrane Receptors and Ion Channels. \u003cstrong\u003e(F) Top Disease-to-Target combinations.\u003c/strong\u003e The \"Cancer/Tumor → Enzyme\" axis dominates Tahoe-100M's performance, driven by the successful recovery of kinase inhibitors. In contrast, CMap recoveries are more evenly distributed across diverse mechanisms. \u003cstrong\u003e(G -L) Biological Concordance Between Validated and Novel Predictions\u003c/strong\u003e. \u003cstrong\u003e(G - H) Differential discovery patterns.\u003c/strong\u003e Heatmaps display the discovery frequency difference between Tahoe-100M (blue) and CMap (red) for \u003cstrong\u003e(G)\u003c/strong\u003e recovered (validated) and \u003cstrong\u003e(H)\u003c/strong\u003e all (total) computational predictions. The high similarity between panels indicates that Tahoe-100M’s novel predictions faithfully reflect the biological mechanisms of validated drugs. \u003cstrong\u003e(I) Cosine similarity\u003c/strong\u003e between recovered and all prediction target-class distributions for each platform. Tahoe-100M (0.987) approaches the theoretical maximum of 1.0, indicating near-identical mechanistic profiles, whereas CMap (0.850) shows greater divergence. \u003cstrong\u003e(J) Target class distributions\u003c/strong\u003e (Butterfly Chart). Comparison of \"All Discoveries\" (light) versus \"Recovered\" (dark) predictions. CMap shows notable shifts between discovery and validation phases, while Tahoe-100M distributions remain consistent. \u003cstrong\u003e(K) Distribution shift\u003c/strong\u003e (Lollipop Chart). Quantifies the percentage point shift between total discovery and validated recovery. CMap exhibits large deviations (e.g., -18.9% for membrane receptors), whereas Tahoe-100M clusters tightly within a ±2% \"Minimal Shift\" zone, demonstrating superior biological stability and predictive reliability. \u003cstrong\u003e(L) Jensen-Shannon divergence\u003c/strong\u003e between the same distributions. Lower values indicate greater similarity; Tahoe-100M (0.150) exhibits less distributional shift than CMap (0.221), confirming that Tahoe-100M's novel predictions more faithfully recapitulate the target-class composition of its validated hits.\u003c/p\u003e","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-10606217/v1/235efade9f4e60fbdb524502.jpeg"},{"id":118787478,"identity":"24edbcdb-6a78-40c4-9113-401132fc0ba9","added_by":"auto","created_at":"2026-08-31 10:56:04","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":128459,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eCase Study: Autoimmune Diseases.\u003c/strong\u003e \u003cstrong\u003e(A) Drug Hits vs. Recovery Rate.\u003c/strong\u003eScatter plot comparing the total number of predicted drug hits against the recovery rate of known therapeutics for 18 autoimmune diseases. Tahoe-100M (blue) demonstrates a strong positive correlation where increased predictions translate to higher recovery, whereas CMap (orange) shows higher variability and lower ceilings for recovery. \u003cstrong\u003e(B) Statistical Comparison.\u003c/strong\u003e Box plots quantifying the difference in known drug recovery rates. Tahoe-100M significantly outperforms CMap (Mean: 76.5% vs 20.9%; Wilcoxon signed-rank test p \u0026lt; 0.001, Cohen's d = 2.22), representing a 3.7-fold improvement in identifying validated treatments. \u003cstrong\u003e(C) Drug Recovery Matrix and Clinical Validation.\u003c/strong\u003e Heatmap displaying the 94 unique drugs recovered across the autoimmune cohort. Cells are colored by platform source: \u003cstrong\u003eOrange\u003c/strong\u003e (CMap Only), \u003cstrong\u003eBlue\u003c/strong\u003e (Tahoe Only), and \u003cstrong\u003ePurple\u003c/strong\u003e (Both). \u003cstrong\u003eRed borders\u003c/strong\u003e denote drugs validated in Phase 4 clinical trials specifically for that disease indication. The matrix highlights the low overlap between platforms (sparse purple cells, ~4.0%) and the high clinical relevance of the predictions, particularly for Tahoe-100M in diseases like Rheumatoid Arthritis and Psoriasis. \u003cstrong\u003e(D) Platform Complementarity (Recovered).\u003c/strong\u003e Donut chart showing the distribution of 175 recovered drug-disease pairs by platform source: Tahoe-only (104, 59.4%), CMap-only (64, 36.6%), and Both (7, 4.0%). \u003cstrong\u003e(E) Per-Disease Recovered Drug Sources.\u003c/strong\u003e Horizontal stacked bar chart displaying the per-disease breakdown of recovered drugs by platform source, sorted by total count. Psoriasis and rheumatoid arthritis dominate both in volume and in consensus predictions (purple), while several diseases (e.g., ITP, Crohn's, ankylosing spondylitis) show exclusively single-platform recovery, underscoring the orthogonal therapeutic spaces sampled by each platform. \u003cstrong\u003e(F) Per-Disease Novel Drug Predictions.\u003c/strong\u003e Horizontal stacked bar chart showing the per-disease breakdown of novel (unvalidated) drug predictions by platform source. Sjögren's syndrome (1,008) and relapsing-remitting multiple sclerosis (913) generate the most novel predictions, reflecting the large candidate pools available for these conditions. \u003cstrong\u003e(G) Platform Overlap in Novel Predictions.\u003c/strong\u003e Donut chart showing the distribution of 8,183 novel drug-disease pairs: CMap-only (3,245, 39.7%), Tahoe-only (4,923, 60.2%), and Both (15, 0.2%). The near-zero overlap (0.2%) in novel predictions is even more pronounced than in recovered drugs (4.0%), reinforcing that the two platforms sample largely non-overlapping therapeutic spaces.\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-10606217/v1/63cd43dbaf7de4022b9f2967.png"},{"id":118787481,"identity":"cd2732b3-1677-495f-b589-806839e56a67","added_by":"auto","created_at":"2026-08-31 10:56:04","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":179237,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eComparison of Top Drug Repurposing Candidates for Endometriosis via CMap and Tahoe-100M Pipelines.\u003c/strong\u003e \u003cstrong\u003e(A) CMap Top 50 Candidates.\u003c/strong\u003e Heatmap displaying the top-ranked drugs identified using the CMap microarray database. Rows represent individual drugs, and columns correspond to the six endometriosis signatures (Unstratified, Stage I/II, Stage III/IV, Proliferative Phase, Early Secretory Phase, Mid-Secretory Phase). Darker purple indicates a stronger predicted reversal of the disease signature. The list is enriched for symptom-modulating agents, including anti-inflammatories (fenoprofen) and hormonal therapies. \u003cstrong\u003e(B) Tahoe-100M Top Candidates.\u003c/strong\u003e Heatmap displaying the top 44 candidates identified using the single-cell Tahoe-100M database. In contrast to CMap, this list is heavily enriched for oncology-related kinase inhibitors and signaling modulators (e.g., pentoxifylline, futibatinib), reflecting the invasive, proliferative nature of endometriosis. \u003cstrong\u003e(A-B)\u003c/strong\u003e Black arrows \u003cstrong\u003e(\u0026gt;\u0026gt;\u0026gt; drug \u0026lt;\u0026lt;\u0026lt;)\u003c/strong\u003e highlight \u003cstrong\u003eirinotecan\u003c/strong\u003e and \u003cstrong\u003eterfenadine\u003c/strong\u003e, the only two drugs shared between the top rankings of both platforms. This minimal overlap illustrates the complementary nature of the resources: CMap captures established management strategies, while Tahoe-100M uncovers novel disease-modifying mechanisms akin to cancer therapy.\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-10606217/v1/900cb1f76c6a1709ea2d0d4f.png"},{"id":118787676,"identity":"d7dfe856-1be6-4767-a5bc-aefec40f8ba7","added_by":"auto","created_at":"2026-08-31 10:56:29","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":871715,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-10606217/v1/1a1ef49a-863a-4ad4-85ae-8eda46f93a3b.pdf"},{"id":118787476,"identity":"8cafd239-d812-4a58-8153-165106ca5fb0","added_by":"auto","created_at":"2026-08-31 10:56:04","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":2004015,"visible":true,"origin":"","legend":"","description":"","filename":"CDRPipeAdditionalFile1Supplement.docx","url":"https://assets-eu.researchsquare.com/files/rs-10606217/v1/0f2d8bfe4644e0b941cba7f4.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Integrating single-cell and bulk transcriptomic perturbation resources reveals complementary therapeutic spaces for drug repurposing","fulltext":[{"header":"Background","content":"\u003cp\u003eThe contemporary landscape of pharmaceutical research and development (R\u0026amp;D) is characterized by a paradoxical trend often referred to as \u0026ldquo;Eroom\u0026rsquo;s Law\u0026rdquo;, the observation that drug discovery is becoming slower and more expensive over time, despite improvements in technology [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. Developing a therapeutic agent de novo is a capital-intensive endeavor, estimated to cost between 2\u0026nbsp;billion and 3\u0026nbsp;billion and requiring 10 to 17 years to move from target identification to market approval [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. Compounding this financial burden is a high attrition rate; approximately 90% of drug candidates entering clinical trials fail to achieve regulatory approval, often due to lack of efficacy or unforeseen toxicity in late-stage development [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. This systemic inefficiency leaves a vast array of diseases, particularly rare conditions and complex chronic disorders, without effective targeted treatments.\u003c/p\u003e \u003cp\u003eIn response to these challenges, drug repurposing, the identification of novel therapeutic indications for existing drugs with established safety profiles, has emerged as a critical strategy to accelerate therapeutic discovery and reduce translational risk [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. Historically, repurposing successes were largely unanticipated, driven by clinical observation of off-target effects, as exemplified by the redeployment of sildenafil for erectile dysfunction or thalidomide for multiple myeloma [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. However, the explosion of high-throughput \u0026ldquo;omics\u0026rdquo; data has catalyzed a paradigm shift toward systematic, computational drug repurposing [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eCentral to this modern approach is the hypothesis of transcriptional reversal (or signature reversion). This principle posits that if a disease state is characterized by a specific gene expression signature (a set of up- and downregulated genes), a small molecule capable of inducing the opposite transcriptional profile may neutralize the disease phenotype and restore a healthy physiological state ([\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]). This concept was pioneered by the Connectivity Map (CMap) project, which provided the first large-scale reference database of drug-induced transcriptional perturbations, allowing researchers, including our own team, to systematically connect small molecules, genes, and diseases [\u003cspan additionalcitationids=\"CR11\" citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eDespite the conceptual elegance and early successes of transcriptomic drug repurposing, the field faces persistent limitations related to data heterogeneity, platform bias, and biological resolution. In bulk transcriptomic experiments, gene expression is measured across all cells present in a tissue or sample and represented as a single composite profile. As a consequence, transcriptional programs originating from rare but disease-driving cell populations can be masked by signals from more abundant but less relevant cell types, reducing the sensitivity of bulk-derived signatures to capture true disease mechanisms. This dilution of cell-type-specific signals can cause therapeutically relevant compounds, those targeting the actual disease-driving pathways, to fail to show significant transcriptional reversal against the composite bulk signature, leading to missed therapeutic associations. At the same time, compounds that induce broad transcriptional suppression or cytostatic responses may be artifactually favored over those with more targeted mechanisms of action [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. Furthermore, reliance on a single perturbation technology can introduce platform-specific biases that limit generalizability [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. The emergence of large-scale single-cell RNA sequencing (scRNA-seq) perturbation databases presents an opportunity to diversify the transcriptional resources available for drug repurposing [\u003cspan additionalcitationids=\"CR15\" citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]. While single-cell perturbation profiling does not directly resolve cell-type heterogeneity in bulk disease signatures, it offers complementary advantages: broader genomic coverage, a wider diversity of cellular contexts for measuring drug responses, and the ability to capture drug-induced transcriptional effects that bulk microarray platforms may fail to detect. For example, Tahoe-100M is a massive single-cell perturbation atlas comprising approximately 100\u0026nbsp;million scRNA-seq profiles that capture transcriptional responses to small-molecule treatments across dozens of cancer cell lines. Yet, the integration of these high-dimensional, disparate data types remains a significant computational bottleneck [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e], often forcing researchers to choose between the extensive drug coverage of legacy bulk databases and the granular resolution of modern single-cell platforms.\u003c/p\u003e \u003cp\u003eTo address these challenges, we developed CDRPipe, a unified framework for transcriptional drug repurposing that enables direct comparison and integration of transcriptional perturbation signatures derived from heterogeneous measurement technologies. CDRPipe standardizes preprocessing, applies consistent nonparametric connectivity scoring, and evaluates statistical significance within a shared empirical framework. We hypothesized that therapeutically relevant compounds would exhibit consistent transcriptional reversal across independent perturbation resources, and that pseudo-bulk signatures derived from single-cell experiments would improve the recovery of known drug-disease associations compared to microarray-based approaches alone.\u003c/p\u003e \u003cp\u003eBy applying CDRPipe to 233 curated disease signatures from CREEDS (Crowd-Extracted Expression of Differential Signatures) [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e] and validating predictions against curated drug-disease knowledge bases, such as Open Targets [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e], we demonstrate that multi-database integration substantially enhances the robustness, interpretability, and translational relevance of transcriptional drug repurposing. We further illustrate the utility of this framework through case studies in autoimmune diseases, where we recover established immunomodulators and identify novel targeted therapies, and in endometriosis, where we replicate prior findings while uncovering new enzyme-centric therapeutic candidates.\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eDisease and Drug Signatures\u003c/h2\u003e \u003cp\u003eThis study was designed to systematically compare and integrate two connectivity mapping approaches for transcriptional drug repurposing. We conducted a retrospective computational analysis using publicly available disease signatures and drug perturbation profiles, with validation against known drug-disease associations. The primary outcomes were recall (recovery rate) of known therapeutics and identification of consensus candidates across platforms. Sample sizes were determined by the availability of curated disease signatures (n\u0026thinsp;=\u0026thinsp;233) in CREEDS and their alignment with Open Targets (n\u0026thinsp;=\u0026thinsp;203 evaluable diseases).\u003c/p\u003e \u003cp\u003eDisease signatures were obtained from CREEDS, a manually curated database of disease-associated differential expression signatures derived from public studies in the Gene Expression Omnibus (GEO) [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e, \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. Each signature contrasts diseased tissues or cells with matched normal or control samples within the same study and is defined by sets of up-regulated and down-regulated genes representing transcriptomic changes associated with the disease state relative to healthy controls. Signatures were generated through a large-scale crowdsourcing effort with manual sample selection, standardized disease annotation, batch-effect correction using surrogate variable analysis, and differential expression prioritization via the Characteristic Direction method, and were all derived from microarray-based experiments [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. They were quality-controlled during original curation; we applied no additional experiment-level filtering prior to analysis. Across the 233 signatures, the mean size was approximately 975 genes (range 38\u0026thinsp;\u0026minus;\u0026thinsp;4,867) prior to filtering and 419 genes (range 2\u0026ndash;2,525) after applying significance and fold-change thresholds (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eWe used the original Connectivity Map (CMap) resource, which provides direct microarray expression profiles of drug-treated cells [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. Following quality-control filtering based on inter-replicate consistency (Pearson correlation r\u0026thinsp;\u0026ge;\u0026thinsp;0.15 between replicate signatures for the same drug-dose-cell line combination), we retained 1,968 of 6,100 original experiments (32.3%), spanning 13,071 unique gene features across five cell lines and 1,309 unique drugs.\u003c/p\u003e \u003cp\u003eWe used pseudo-bulk perturbation profiles derived from large-scale single-cell RNA sequencing experiments in the Tahoe-100M atlas [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]. From this resource, we used 56,827 drug-response experiments profiling responses across diverse cell types and conditions. Gene-level significance filtering (p\u0026thinsp;\u0026le;\u0026thinsp;0.05) was applied during signature generation to retain only significantly perturbed genes; for connectivity scoring, the retained log2 fold-change values were converted to genome-wide ranks, which served as input for the nonparametric KS-based scoring procedure (see below). No additional experiment-level filtering was required given the resource\u0026rsquo;s inherently high data quality. After mapping to a shared gene universe, 22,168 genes were available for analysis, and the library covered 50 cell lines and 379 unique drugs.\u003c/p\u003e \u003cp\u003eThe two libraries provide largely non-overlapping chemical space: CMap and Tahoe-100M share 85 drugs and 12,544 genes but no cell lines (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eF, Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eComparison of drug signature characteristics between CMap and Tahoe-100M\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCategory\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCMap\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eTahoe-100M\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eShared\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eUnique drugs\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1,309\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e379\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e85\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eUnique genes\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e13,071\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e25,084\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e12,544\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eCell lines\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e50\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eGene matrix dimensions\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e~\u0026thinsp;6,100 \u0026times; 13,071\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e56,827 \u0026times; 62,710\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u0026mdash;\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eTechnology\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMicroarray HG U133A\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eSingle-cell RNA seq with aggregated pseudo bulk matrices\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eTotal profiles\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e6,100 experiment profiles\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u0026gt;\u0026thinsp;100,000,000 single cell profiles and 56,827 aggregated experiment profiles\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eExpression values per experiment\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eranked probe set signatures\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ebasemean, log2FC, p value, p adjusted\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eCDRPipe Computational Pipeline\u003c/h3\u003e\n\u003cp\u003eDisease signatures were standardized by mapping gene symbols to Entrez IDs using g:Profiler (gconvert function, organism = \u0026ldquo;hsapiens\u0026rdquo;, target = \u0026ldquo;ENTREZGENE_ACC\u0026rdquo;). Genes were filtered by adjusted p-value (\u0026lt;\u0026thinsp;0.05, when available) and absolute log fold-change (\u0026gt;\u0026thinsp;0.25). Only genes present in the respective drug perturbation database gene universe were retained.\u003c/p\u003e \u003cp\u003eFor each disease signature and drug perturbation profile, we computed a modified Kolmogorov-Smirnov (KS) connectivity score quantifying transcriptional reversal. Briefly, disease up-regulated genes were expected to be ranked low in drug signatures (indicating downregulation by drug), while disease down-regulated genes were expected to be ranked high (Additional file 1: Algorithm S1). The connectivity score combines KS statistics for up and down gene sets:\u003c/p\u003e \u003cp\u003eScore\u0026thinsp;=\u0026thinsp;KS\u003csub\u003eup\u003c/sub\u003e \u0026minus; KS\u003csub\u003edown\u003c/sub\u003e\u003c/p\u003e \u003cp\u003eWhere Score\u0026thinsp;=\u0026thinsp;0 if sgn(KS\u003csub\u003eup\u003c/sub\u003e)\u0026thinsp;=\u0026thinsp;sgn(KS\u003csub\u003edown\u003c/sub\u003e) and both exist.\u003c/p\u003e \u003cp\u003eNegative scores indicate transcriptional reversal (drug reverses disease signature); positive scores indicate mimicry.\u003c/p\u003e \u003cp\u003eStatistical significance was assessed using an empirical null distribution generated from 100,000 random gene sets matched in size to the query disease signature. Two-sided p-values were computed as the proportion of random scores with absolute value exceeding the observed score. To prevent zero p-values, we applied the permutation-based minimum: p\u0026thinsp;=\u0026thinsp;1/(N\u0026thinsp;+\u0026thinsp;1) when no random scores exceeded the observed value [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]. False discovery rate correction was performed using the qvalue package; when qvalue failed due to distribution assumptions, Benjamini-Hochberg adjustment was applied.\u003c/p\u003e \u003cp\u003eDrug candidates were identified using a significance threshold of q\u0026thinsp;\u0026lt;\u0026thinsp;0.05 combined with negative connectivity score (reversal direction). For each disease, the most significant instance per drug (minimum score across cell lines and concentrations) was retained to avoid duplicate counting.\u003c/p\u003e\n\u003ch3\u003eValidation Against Known Drug-Disease Associations\u003c/h3\u003e\n\u003cp\u003eTo validate our pipeline against established clinical knowledge, we obtained known drug-disease associations from the Open Targets Platform (2024 release) [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. Disease entities from CREEDS were systematically mapped to Open Targets using a hierarchical approach, identifying 151 diseases (64.8%) through exact name matching and 52 diseases (22.3%) via ontology synonyms, while 30 diseases (12.9%) remained unmatched (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). This yielded 203 diseases with at least one known therapy for validation. We included disease-drug pairs with evidence of any recorded development activity, including drugs in clinical trials across all phases; in total, Open Targets provided over 118,000 associations for these conditions, of which 2,668 disease-drug pairs involved drugs present in the CMap or Tahoe-100M libraries and were evaluated in our analysis.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eClassification of disease signatures by therapeutic area\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTherapeutic Area\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eN\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003e%\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eExamples\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCancer/Tumor\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e27\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e15.0%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eLung adenocarcinoma, Colorectal cancer, Melanoma, Glioblastoma multiforme, Ovarian cancer\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGenetic/Congenital\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e20\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e11.1%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eCystic fibrosis, Duchenne muscular dystrophy, Down syndrome, Huntington's disease, Marfan syndrome\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNervous System\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e18\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e10.0%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eAlzheimer's disease, Multiple sclerosis, Epilepsy syndrome, Amyotrophic lateral sclerosis,..\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eImmune System\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e13\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e7.2%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eRheumatoid arthritis, Psoriasis, Sj\u0026ouml;gren's syndrome, Ankylosing spondylitis, Psoriatic arthritis\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGastrointestinal\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e11\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e6.1%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eUlcerative colitis, Crohn's disease, Barrett's esophagus, NASH, Gastroesophageal reflux disease\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMusculoskeletal\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e11\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e6.1%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eOsteoarthritis, Systemic lupus erythematosus, Scleroderma, Osteoporosis, Arthritis\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCardiovascular\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e5.0%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMyocardial infarction, Hypertension, Atherosclerosis, Dilated cardiomyopathy, Pulmonary hypertension\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eInfectious Disease\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e5.0%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eTuberculosis, Influenza, Hepatitis C, HIV infectious disease, Dengue disease\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRespiratory\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e4.4%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eAsthma, Chronic obstructive pulmonary disease, Allergic asthma, Acute lung injury, Bronchopulmonary dysplasia\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHematologic\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3.9%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSickle cell anemia, Aplastic anemia, Myelodysplastic syndrome, Chronic myeloid leukemia, Acute T cell leukemia\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePsychiatric\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3.9%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSchizophrenia, Autism spectrum disorder, Alcohol abuse, Nicotine dependence, Bipolar disorder\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eReproductive/Breast\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3.9%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eBreast cancer, Endometriosis, Prostate cancer, Endometrial cancer, Polycystic ovary syndrome\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEndocrine System\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3.3%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003ePolycystic ovary syndrome, Papillary thyroid carcinoma, Adrenoleukodystrophy, Androgen insensitivity syndrome\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePhenotype\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3.3%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eIschemic stroke, Septic shock, Hypoxia, Dystonia, Intellectual disability\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eUrinary System\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2.8%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eRenal cell carcinoma, Cystitis, Nephroblastoma, Interstitial cystitis, Diabetic nephropathy\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSkin/Integumentary\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2.8%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eAtopic dermatitis, Eczema, Psoriasis vulgaris, Actinic keratosis, Urticaria\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMetabolic\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2.8%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eDiabetes mellitus, Obesity, Morbid obesity, MELAS syndrome, Familial hypercholesterolemia\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePancreas\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.7%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eType 1 diabetes mellitus, Type 2 diabetes mellitus, Pancreatic ductal adenocarcinoma\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eVisual System\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.1%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eGlaucoma, Primary open angle glaucoma\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePregnancy/Perinatal\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.6%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003ePre-eclampsia\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eFor each disease we defined I as the number of unique drugs predicted, S as the number of predicted drugs validated in Open Targets, and P as the recoverable ceiling (known drugs present in the platform\u0026rsquo;s library). Precision was calculated as S/I and recall (recovery rate) as S/P.\u003c/p\u003e\n\u003ch3\u003eCase Study Analyses\u003c/h3\u003e\n\u003cp\u003eFor the autoimmune disease case study, 18 autoimmune diseases with available Open Targets matches were analyzed in depth. For each disease, CDRPipe was run independently against both CMap and Tahoe-100M perturbation libraries using the standard pipeline parameters (q\u0026thinsp;\u0026lt;\u0026thinsp;0.05, negative connectivity score). Drug-level analysis classified each recovered drug by its platform source (CMap only, Tahoe-100M only, or both). Clinical validation was performed by cross-referencing recovered drugs against Phase 4 clinical trial data from Open Targets for each specific disease indication. Recovery rates were compared between platforms using the Wilcoxon signed-rank test, with effect size quantified by Cohen\u0026rsquo;s d.\u003c/p\u003e \u003cp\u003eFor the endometriosis case study, six disease signatures from Oskotsky et al. (Unstratified, Stage I-II, Stage III-IV, Proliferative Phase, Early Secretory Phase, Mid-Secretory Phase) were re-analyzed using both CMap and Tahoe-100M. To ensure exact replication of the original findings, the CMap analysis used the parameters reported in Oskotsky et al.: p-value\u0026thinsp;\u0026lt;\u0026thinsp;0.05, absolute log2 fold change\u0026thinsp;\u0026gt;\u0026thinsp;1.1, random seed\u0026thinsp;=\u0026thinsp;2009, and 1,000 permutations. The same parameters were then applied to Tahoe-100M for direct comparison. Top candidates were ranked by connectivity score across signatures, and concordance between platforms was assessed by comparing the top 50 candidates from each.\u003c/p\u003e \u003cp\u003eDiseases were categorized into nine therapeutic areas using keyword matching: Oncology, Metabolic, Neurodegenerative, Cardiovascular, Infectious, Autoimmune, Allergic/Respiratory, Rare/Genetic, and Other. Category-level performance was summarized as mean known drug hits per disease.\u003c/p\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003eStatistical Analysis\u003c/h2\u003e \u003cp\u003eAll statistical analyses were performed in R (version 4.2+). Paired comparisons of Tahoe-100M versus CMap recovery rates were performed using the Wilcoxon signed-rank test. Effect sizes were quantified using Cohen\u0026rsquo;s d. P-values\u0026thinsp;\u0026lt;\u0026thinsp;0.05 were considered statistically significant. For the overall recovery rate comparison, 153 diseases with evaluable data for both platforms were included in the paired analysis. Complementarity was quantified as the percentage of recovered drugs identified by both platforms, Category-level recommendations were based on mean known drug hits with a threshold of \u0026gt;\u0026thinsp;2 drugs difference for primary method recommendation.\u003c/p\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003eStudy Overview\u003c/h2\u003e \u003cp\u003eTo systematically evaluate the impact of cellular resolution on drug repurposing, we integrated heterogeneous transcriptional resources into a unified analytical framework (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). We applied CDRPipe to 233 curated CREEDS disease signatures, interrogating each against two complementary drug perturbation libraries with distinct measurement technologies: the microarray-based Connectivity Map (CMap) and the single-cell-derived Tahoe-100M atlas (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). CDRPipe applies uniform preprocessing and a nonparametric connectivity score to quantify how strongly each drug reverses a given disease signature, with significance assessed by empirical permutation-based null models (see Methods). Predictions for all 233 signatures were benchmarked against known drug-disease associations from Open Targets (203 evaluable diseases; 2,668 assessable drug-disease pairs). The following sections report recovery performance, cross-platform complementarity, and disease-specific case studies.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003ePrediction Recovery and Validation\u003c/h3\u003e\n\u003cp\u003eWe assessed performance using precision and recall against Open Targets (defined in the Methods). Because Open Targets captures drugs with any recorded development activity, including early-phase clinical trial candidates, and is biased toward compounds with established clinical annotation, these metrics reflect recovery of documented associations within the benchmark rather than absolute predictive accuracy. To account for overlapping annotations, diseases sharing identical therapeutic area combinations in Open Targets were grouped. Tahoe-100M recovered more annotated drug-disease pairs than CMap (2,198 vs. 948), a 2.3-fold difference. Tahoe-100M also covered more diseases (171 vs. 155) and achieved a higher overall recovery rate (22.1% versus 18.1%). These differences partly reflect the composition of each drug library: Tahoe-100M contains many newer, clinically oriented small-molecule compounds that may have inherently higher overlap with Open Targets annotations, whereas CMap\u0026rsquo;s library includes older and more exploratory chemical entities.\u003c/p\u003e \u003cp\u003eApplying these per-disease metrics, lung adenocarcinoma illustrates the approach: Tahoe-100M predicted 61 unique drugs (I\u0026thinsp;=\u0026thinsp;61), 10 of which were validated (S\u0026thinsp;=\u0026thinsp;10), yielding a precision of 16.4%; of the 12 recoverable known drugs present in the library (P\u0026thinsp;=\u0026thinsp;12), the pipeline recovered 10, achieving 83.3% recall.\u003c/p\u003e \u003cp\u003eAt the disease level across all 233 evaluated diseases, Tahoe-100M recovered a higher proportion of Open Targets-annotated drug-disease pairs than CMap (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). Mean precision was 1.8% (SD 3.3%) for Tahoe-100M versus 1.3% (SD 2.8%) for CMap (measured across n\u0026thinsp;=\u0026thinsp;223 and n\u0026thinsp;=\u0026thinsp;208 diseases with at least one prediction, respectively) (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA). Among diseases with recoverable drugs (P\u0026thinsp;\u0026gt;\u0026thinsp;0), Tahoe-100M achieved a mean recall of 47.3% (SD 38.5%) compared to 18.5% (SD 25.1%) for CMap (n\u0026thinsp;=\u0026thinsp;156 and n\u0026thinsp;=\u0026thinsp;165 diseases, respectively; Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eB). These differences should be interpreted cautiously: they partly reflect the greater overlap between Tahoe-100M\u0026rsquo;s drug library and Open Targets annotations rather than representing a direct measure of predictive superiority. Nonetheless, many diseases achieved\u0026thinsp;\u0026gt;\u0026thinsp;20% recall with Tahoe-100M compared to substantially fewer with CMap, and multiple diseases achieved\u0026thinsp;\u0026gt;\u0026thinsp;50% recall with Tahoe-100M, a threshold not reached by any disease in CMap analyses (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eC). Variable sample sizes reflect platform-specific prediction coverage and the filtering requirements for each metric (I\u0026thinsp;\u0026gt;\u0026thinsp;0 for precision calculations and P\u0026thinsp;\u0026gt;\u0026thinsp;0 for recall calculations).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eComplementarity of Platforms and Mechanistic Profiles\u003c/h2\u003e \u003cp\u003eAnalysis of the overlap between drug-disease pairs predicted by each platform revealed low concordance (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eD). Among drug-disease pairs supported by Open Targets, only 144 were identified by both CMap and Tahoe-100M (Jaccard index 4.8%). When considering all predicted pairs, including those not supported by Open Targets, the overlap remained limited (173 pairs; Jaccard index 1.2%). Despite both platforms being applied to the same set of 233 disease signatures, the majority of predicted pairs were unique to a single platform, with each exhibiting different recovery volumes and enrichment patterns across therapeutic areas and drug target classes (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eD, \u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eE). This divergence is consistent with the limited chemical space overlap between the two drug libraries (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eF) and is further reflected in the distinct disease-to-target combinations favored by each platform (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eF). Together, these results indicate that CMap and Tahoe-100M sample largely non-overlapping therapeutic spaces rather than redundantly identifying the same candidates, and that the primary value of integrating both platforms lies in this complementarity.\u003c/p\u003e \u003cp\u003eWhile both platforms validated predictions across diverse therapeutic areas, with cancer/tumor being the largest category for both, Tahoe-100M\u0026rsquo;s recovered predictions were significantly enriched for oncology, with 28.7% of Tahoe-100M\u0026rsquo;s recovered drug-disease pairs falling in oncology compared to 14.3% for CMap (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eD\u0026ndash;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eF), consistent with its enzyme-centric drug library bias (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eC).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eBiological Concordance of Novel Predictions\u003c/h2\u003e \u003cp\u003eTo determine if novel predictions adhered to the same therapeutic logic as clinically validated drugs, we evaluated the biological concordance between \u0026ldquo;recovered\u0026rdquo; (validated) predictions and \u0026ldquo;all\u0026rdquo; discoveries by comparing their joint distributions across disease categories and drug target classes. Heatmaps of these distributions revealed that Tahoe-100M\u0026rsquo;s differential discovery patterns were highly similar between recovered and all predictions (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eG, \u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eH), yielding a cosine similarity of 0.987 and a Jensen-Shannon divergence of 0.150 (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eI, \u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eL). This consistency indicates that Tahoe-100M\u0026rsquo;s novel candidates closely mirror the mechanistic profile of its validated successes.\u003c/p\u003e \u003cp\u003eCMap showed lower concordance between its recovered and all discovery distributions (cosine similarity\u0026thinsp;=\u0026thinsp;0.850; Jensen-Shannon divergence\u0026thinsp;=\u0026thinsp;0.221; Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eI, \u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eL). The butterfly chart of target class distributions (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eJ) shows that CMap exhibits notable shifts between the \u0026ldquo;all discoveries\u0026rdquo; and \u0026ldquo;recovered\u0026rdquo; profiles, particularly for membrane receptors, whereas Tahoe-100M distributions remain consistent. This is quantified in the lollipop chart (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eK), where CMap exhibits large percentage-point deviations between total discovery and validated recovery (e.g., \u0026minus;\u0026thinsp;18.9% for membrane receptors), while Tahoe-100M clusters tightly within a\u0026thinsp;\u0026plusmn;\u0026thinsp;2% shift zone. This higher divergence for CMap implies that it explores a broader chemical space, potentially identifying mechanistically novel relationships that differ from established clinical precedents. We note that these platform-specific mechanistic profiles partly reflect the inherent composition of each drug library (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eB, \u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eC): Tahoe-100M\u0026rsquo;s enzyme-enriched library naturally yields enzyme-centric predictions, while CMap\u0026rsquo;s receptor-heavy library favors receptor-targeted discoveries.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eCase Study 1: Autoimmune Diseases\u003c/h2\u003e \u003cp\u003eTo evaluate pipeline performance, we analyzed 18 autoimmune diseases mapped to Open Targets via exact name or synonym matching, details about the diseases available in Additional file 1: Table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e. Note that the \u0026ldquo;known drugs\u0026rdquo; column in Additional file 1: Table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e reflects all drugs with any recorded development activity for that indication in Open Targets (including clinical trial candidates at all phases), rather than only approved therapies; this accounts for the high counts observed for some diseases. Autoimmune diseases constitute a rigorous, biologically grounded benchmark for transcriptional drug repurposing, as their pathogenesis is driven by aberrant immune activation that manifests as reproducible gene expression signatures [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e]. Despite clinical heterogeneity, these disorders converge on shared inflammatory axes, including interferon signaling, cytokine-mediated activation, and T-cell dysregulation, allowing for a systematic evaluation of the pipeline\u0026rsquo;s ability to recover established immunomodulators [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e]. Furthermore, the persistence of treatment resistance in autoimmune conditions underscores the clinical urgency for scalable approaches that nominate candidates based on molecular reversal rather than phenotype alone. As shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eB, Tahoe-100M significantly outperformed CMap in recovering known therapeutics in the Open Targets database (mean recovery rate: 76.5% vs. 20.9%; Wilcoxon signed-rank test p\u0026thinsp;\u0026lt;\u0026thinsp;0.001; Cohen\u0026rsquo;s d\u0026thinsp;=\u0026thinsp;2.22), representing a 3.7-fold improvement (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). Disease-specific analysis revealed Tahoe-100M\u0026rsquo;s particular strength in inflammatory conditions: Crohn\u0026rsquo;s disease (75.0% vs 7.4%), psoriasis (100.0% vs 75.0%), and rheumatoid arthritis (64.0% vs 17.8%) (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eA). However, CMap demonstrated better performance for type 1 diabetes mellitus, which may reflect the fact that its treatment paradigm centers on glucose regulation rather than immunosuppression, potentially favoring CMap\u0026rsquo;s drug library composition.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eWe further validated these findings by examining disease-specific Phase 4 clinical trial data (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eC). Of the recovered drug-disease pairs, 26.3% (46 of 175) involved drugs validated in Phase 4 specifically for that indication. The most frequently recovered agents were cornerstone immunosuppressants, including dexamethasone (identified in 12 diseases), methotrexate (11 diseases), and hydrocortisone (7 diseases). Several robust disease-specific patterns emerged: rheumatoid arthritis demonstrated the highest concordance with Phase 4 data, recovering 15 distinct Phase 4-approved drugs. This is consistent with rheumatoid arthritis having one of the broadest therapeutic arsenals among autoimmune diseases, providing a larger validation set. Both platforms consistently predicted methotrexate, celecoxib, naproxen, dexamethasone, and meloxicam, representing greater therapeutic validation than other autoimmune conditions evaluated. Similarly, the psoriasis spectrum (including psoriatic arthritis) showed consistent recovery of dimethyl fumarate, tazarotene, and clobetasol propionate. Gastrointestinal autoimmune diseases (Crohn\u0026rsquo;s disease and ulcerative colitis) preferentially recovered locally-acting corticosteroids such as budesonide. Sj\u0026ouml;gren\u0026rsquo;s syndrome recovered pilocarpine and leflunomide, reflecting its distinct pathophysiology centered on exocrine gland dysfunction.\u003c/p\u003e \u003cp\u003eDespite Tahoe-100M\u0026rsquo;s overall superiority, the two methods showed high complementarity with only 4.0% overlap in recovered drug-disease pairs. Across 175 recovered drug-disease pairs, Tahoe-100M uniquely contributed 104 (59.4%), while CMap uniquely identified 64 (36.6%) that Tahoe-100M missed. Beyond validated associations, both platforms generated substantial novel predictions: CMap produced 4,900 novel (unvalidated) drug-disease pairs involving 198 drugs that had no Open Targets confirmation, while Tahoe-100M produced 8,766 novel pairs involving 32 exclusively novel drugs, a difference reflecting Tahoe-100M\u0026rsquo;s smaller but more clinically oriented library, where most drugs already have some Open Targets annotation (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eF, \u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eG). Drugs predicted by both methods showed 2.6-fold higher precision (5.2%) compared to single-method predictions (1.8\u0026ndash;2.0%), suggesting consensus predictions represent higher-confidence repurposing candidates. Across the full panel of 18 autoimmune diseases, we identified 94 unique recovered drugs representing the combined outputs of both databases. Only 7 of 175 drug-disease pairs (4.0%) were identified by both platforms, reinforcing that distinct signals are captured by CMap and Tahoe-100M (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eC, \u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eD).\u003c/p\u003e \u003cp\u003eQualitative analysis of the recovered candidates revealed distinct platform-specific therapeutic spaces. CMap excelled at recovering established medications, including NSAIDs, corticosteroids, and traditional immunosuppressants. In contrast, Tahoe-100M uniquely identified emerging targeted therapies, including JAK inhibitors (tofacitinib, filgotinib) across four diseases, BTK inhibitors (tirabrutinib), and SGLT2 inhibitors for diabetes. This difference partly reflects the composition of each drug library: Tahoe-100M includes newer, clinically oriented compounds not present in CMap, while CMap\u0026rsquo;s library is enriched for older, well-established pharmacological agents. These findings underscore the value of a multi-platform approach: each platform captures a distinct pharmacological space, and their integration maximizes overall therapeutic coverage. While the volume of novel candidates is substantial, all predictions are rank-ordered by their connectivity score, providing a principled prioritization for downstream evaluation. The high candidate counts partly reflect the use of uniform statistical thresholds across all 18 diseases; disease-specific parameter tuning could narrow the output but risks introducing subjective bias. Importantly, the ranking ensures that the most promising candidates, those with the strongest expression-reversal signatures, appear at the top, enabling systematic experimental follow-up in order of predicted efficacy rather than requiring exhaustive evaluation of the full list.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003eCase Study 2: Endometriosis\u003c/h2\u003e \u003cp\u003eTo further validate our pipeline, we applied Tahoe-100M to endometriosis, benchmarking our results against the prior CMap-based analysis by Oskotsky et al [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. Endometriosis presents a unique opportunity for computational discovery due to the critical unmet need for non-hormonal therapies and the historical stagnation of traditional research and development [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e]. Biologically, however, it is an ideal candidate for this approach as the disease shares significant molecular hallmarks with malignancy, including invasion, angiogenesis, and metastasis [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]. This transcriptional overlap is particularly advantageous given that the reference data in both Tahoe-100M and CMap are derived largely from cancer cell lines. We began by replicating the findings of Oskotsky et al [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e] using the CDRPipe framework. We achieved 100% replication of their reported results using the original study parameters: a p-value less than 0.05, an absolute log2 fold change greater than 1.1, a random seed of 2009, and 1,000 permutations (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eA). Following this validation, we applied the same rigorous parameters to the Tahoe-100M pipeline (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eB). While Oskotsky et al. stratified signatures by disease stage and menstrual phase, we focused our comparative ranking on the \u0026ldquo;Unstratified\u0026rdquo; signature to identify robust, broad-spectrum candidates.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eWe subsequently evaluated the concordance between the two platforms by analyzing the top 50 drug candidates identified by each. The overlap was modest, highlighting the distinct yet complementary nature of the libraries. Among the top 50 hits in CMap, only six were present in Tahoe-100M, and conversely, only seven of the top 50 Tahoe hits were present in CMap. Despite these differences, we identified two drugs in common within the top rankings: terfenadine and irinotecan. This divergence suggests that the platforms capture different aspects of the disease biology. Tahoe-100M is heavily enriched for oncology and signaling inhibitors, surfacing numerous kinase inhibitors and epigenetic regulators such as FGFR inhibitors (pemigatinib, erdafitinib, futibatinib), CDK/AKT pathway drugs (dinaciclib, ipatasertib), and HDAC inhibitors (panobinostat, tucidinostat). This profile aligns well with modern characterizations of endometriosis as a proliferative, invasive, and inflammatory disease with cancer-like transcriptomic properties[\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e, \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]. In contrast, CMap leans toward symptom modulation, prioritizing hormonal agents (levonorgestrel, medrysone), anti-inflammatory drugs (mesalazine, fenoprofen, resveratrol), and neuropsychiatric compounds (clomipramine).\u003c/p\u003e \u003cp\u003eThe clinical relevance of these findings is supported by strong literature evidence for top hits from both pipelines. Both platforms successfully recovered established therapies: CMap identified levonorgestrel, a hormonal (progestin) therapy commonly used to treat endometriosis related pain, administered orally, as implants, and used in IUDs [\u003cspan additionalcitationids=\"CR29\" citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e], while Tahoe-100M identified medroxyprogesterone acetate, which is widely used for hormonal suppression of lesions [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]. Beyond standard of care, Tahoe-100M identified pentoxifylline as the top candidate common across all six disease signatures; its potential to prevent endometriosis recurrence after conservative surgery has been supported by controlled trials [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e, \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e]. Similarly, CMap identified fenoprofen and resveratrol, both of which have documented efficacy in reducing lesion size and angiogenesis in animal models [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e]. Several novel hits also show mechanistic promise; simvastatin (CMap) has been reported to inhibit stromal cell proliferation in vitro [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e], while valproic acid (CMap) has shown potential in preclinical models to inhibit lesion growth [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e, \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e, \u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e]. Furthermore, Tahoe-100M identified drospirenone, a progestin with anti-androgenic and anti-inflammatory properties relevant to pain control [\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e]. While toxicity limits the clinical application of some hits, such as the topoisomerase inhibitor irinotecan (found in both pipelines), their identification underscores the pipeline\u0026rsquo;s ability to correctly identify antiproliferative mechanisms relevant to the disease pathology.\u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eThe principal finding of this study is the high complementarity between CMap and Tahoe-100M: only 3.5% of recovered drugs were identified by both platforms, yet each captured a distinct and clinically relevant segment of the therapeutic landscape. While Tahoe-100M uniquely recovered 64.7% of candidates, capturing emerging targeted therapies like JAK and BTK inhibitors, CMap contributed 31.8% of hits, excelling at established therapeutics such as NSAIDs and corticosteroids. When benchmarked against Open Targets annotations, Tahoe-100M recovered a greater proportion of documented drug-disease associations than CMap (47.3% vs. 18.5%; 2.6-fold difference), with pronounced differences in autoimmune (4.3-fold) and oncology (3.0-fold) categories. These differences in annotated recovery rates should be interpreted cautiously, however: they partly reflect the greater overlap between Tahoe-100M\u0026rsquo;s drug library and Open Targets annotations, and the larger experimental breadth of Tahoe-100M (56,827 vs. 1,968 experiments capturing greater pharmacological and compound diversity), rather than representing a direct measure of superior predictive accuracy.\u003c/p\u003e \u003cp\u003eThis divergence is driven by distinct mechanistic biases. Tahoe-100M displays a strongly \u0026ldquo;enzyme-centric\u0026rdquo; profile (44\u0026ndash;46% of discoveries), reflecting its sensitivity to kinase cascades that create coherent transcriptomic signatures in oncology and inflammatory conditions. In contrast, CMap exhibits a \u0026ldquo;receptor-oriented\u0026rdquo; profile (35.8% membrane receptors), effectively capturing the rapid, membrane-proximal effects of agonists and ion channel modulators. These findings demonstrate that integrating heterogeneous perturbation resources is essential to maximize therapeutic coverage\u003c/p\u003e \u003cp\u003eOur disease category analysis provides actionable guidance for method selection. Tahoe-100M is recommended as the primary approach for oncology, autoimmune, and cardiovascular drug repurposing, where it demonstrated consistent 1.5\u0026ndash;3.0-fold advantages. For metabolic and neurodegenerative diseases, CMap showed modest advantages, suggesting these categories may benefit from CMap-first or complementary approaches. The endometriosis case study illustrates an important caveat: for hormone-responsive conditions with established CMap-derived findings, CMap remains optimal for replication (62.5% recovery of Oskotsky et al.\u0026rsquo;s drugs), while Tahoe-100M provides value as a complementary discovery tool contributing unique candidates. This disease-specific performance pattern highlights that no single drug signature dataset is universally superior; method selection should be tailored to the disease biology and study objectives.\u003c/p\u003e \u003cp\u003eOur findings have several implications for computational drug repurposing practice; the first is that consensus predictions merit prioritization. Drugs identified by both CMap and Tahoe-100M showed 2.6-fold higher precision than single-method predictions, suggesting consensus hits represent higher-confidence candidates for experimental validation. Secondly, multi-method approaches maximize coverage. The high complementarity between methods (only 3.5\u0026ndash;7.6% overlap in recovered drugs) indicates that single-method studies substantially underestimate the drug repurposing landscape. Integrating both CMap and Tahoe-100M captures distinct therapeutic signals. Lastly, platform-specific strengths inform therapeutic focus. CMap\u0026rsquo;s strength in established drug classes and Tahoe-100M\u0026rsquo;s identification of emerging targeted therapies suggest that platform selection can be tuned to study objectives, validation of known mechanisms versus discovery of novel interventions.\u003c/p\u003e \u003cp\u003eOur comparative analysis informs evidence-based platform selection for drug repurposing campaigns. Tahoe-100M is recommended for enzyme-driven pathologies (oncology, inflammation, metabolic disorders) and programs requiring high precision and biological concordance. CMap is optimal for receptor-mediated pathways (neurology, psychiatry, cardiovascular) and exploratory programs targeting ion channels or transporters. Given the minimal overlap (1.2%) yet distinct mechanistic strengths, we recommend an integrated strategy. Dual-platform hits represent high-confidence candidates suitable for prioritized validation, while platform-specific predictions should be pursued based on alignment between the disease mechanism and the platform\u0026rsquo;s established target-class bias.\u003c/p\u003e \u003cp\u003eSeveral limitations should be considered. First, validation relied on Open Targets, which creates a bias toward approved and late-stage candidates while missing off-label or early-stage therapeutics. Our definition of a true association, any drug evaluated for a given disease, favors sensitivity but may obscure specificity regarding efficacy. Second, despite standardized processing, disease signatures from CREEDS retain heterogeneity due to diverse experimental designs and tissue sources in the underlying GEO data. Third, our approach evaluates transcriptional reversal, which does not guarantee a specific therapeutic mechanism; observed reversals could stem from generalized stress responses rather than disease-specific correction. Fourth, many identified compounds, particularly from CMap, could not be mapped to Open Targets. This reflects the presence of proprietary or research-grade chemicals in CMap that lack public annotation, potentially leading to an underestimation of their relevance. Finally, the perturbation datasets differ technically: CMap uses bulk microarray profiling on limited cell lines, while Tahoe-100M utilizes single-cell RNA sequencing across a wider cellular panel. Although both rely heavily on neoplastic lines, they capture conserved biological processes relevant to non-cancer contexts. The superior resolution of the Tahoe-100M likely drives its improved precision. Ultimately, our findings suggest that CMap and Tahoe-100M are complementary, and platform selection should be driven by the specific translational context.\u003c/p\u003e \u003cp\u003eWhile this study identifies thousands of drug repurposing candidates validated against known therapeutics, bridging the gap to clinical application requires a rigorous translational framework [\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e]. Top-ranked candidates, particularly those achieving \u0026ldquo;triple-validation\u0026rdquo; across CMap, Tahoe-100M, and prior literature, must first undergo experimental verification in high-fidelity models, such as patient-derived organoids or microphysiological systems, to confirm efficacy beyond standard cell lines [\u003cspan additionalcitationids=\"CR40\" citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e]. These efforts should be complemented by deep mechanistic profiling to ensure that signature reversal is driven by therapeutically relevant target engagement rather than off-target toxicity, a step particularly critical for drugs with polypharmacology [\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e]. For candidates already approved for other indications, this translation can be accelerated by systematically reviewing real-world evidence and electronic health records to de-risk safety profiles in new patient populations [\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e]. Furthermore, our observation that Tahoe-100M achieves variable recovery rates across autoimmune conditions underscores the critical need for transcriptomic patient stratification to identify responsive subtypes [\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e]. Ultimately, randomized clinical trials remain the definitive gold standard; future studies must transition validated candidates into Phase II proof-of-concept or innovative basket trials to establish efficacy and safety in human subjects prior to broad clinical implementation [\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e].\u003c/p\u003e"},{"header":"Conclusions","content":"\u003cp\u003eThis comprehensive evaluation demonstrates that Tahoe-100M and CMap provide complementary perspectives on the drug repurposing landscape, with Tahoe-100M showing overall superior recovery of known therapeutics and particular strength in oncology and autoimmune conditions. The high complementarity between methods, with only 3.5\u0026ndash;7.6% overlap in recovered drugs, argues for multi-method approaches as standard practice in computational drug repurposing. Consensus predictions from both platforms represent high-confidence candidates meriting experimental prioritization. Our disease category-level recommendations and case study analyses provide actionable guidance for method selection across diverse therapeutic areas.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cdiv class=\"DefinitionList\"\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eAKT\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eprotein kinase B\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eBTK\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eBruton tyrosine kinase\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eCDK\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003ecyclin-dependent kinase\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eCDRPipe\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eComputational Drug Repurposing Pipeline\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eCMap\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eConnectivity Map\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eCREEDS\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eCrowd-Extracted Expression of Differential Signatures\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eFDR\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003efalse discovery rate\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eFGFR\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003efibroblast growth factor receptor\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eGEO\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eGene Expression Omnibus\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eHDAC\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003ehistone deacetylase\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eJAK\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eJanus kinase\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eKS\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eKolmogorov-Smirnov\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eNSAID\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003enonsteroidal anti-inflammatory drug\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eQC\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003equality control\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eR\u0026amp;D\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eresearch and development\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003escRNA-seq\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003esingle-cell RNA sequencing\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eSD\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003estandard deviation\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eSGLT2\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003esodium-glucose cotransporter 2.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"Declarations","content":"\u003ch2\u003eEthics approval and consent to participate\u003c/h2\u003e\n\u003cp\u003eNot applicable. This study analysed only publicly available, previously published, de-identified datasets and did not involve new experiments on human participants, human tissue, or animals.\u003c/p\u003e\n\u003ch2\u003eConsent for publication\u003c/h2\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003ch2\u003eCompeting interests\u003c/h2\u003e\n\u003cp\u003eThe authors declare that they have no competing interests.\u003c/p\u003e\n\u003ch2\u003eFunding\u003c/h2\u003e\n\u003cp\u003eThis work was supported by the National Institutes of Health (grants 1R21HD114953 to MS and LCG; 1R01AI180118 to MS; P30 AR070155 to MS; P01HD106414 to LCG; and 1R01AG100879-0 to MS) and by the UCSF March of Dimes Prematurity Research Center. The funders had no role in the design of the study; the collection, analysis, and interpretation of the data; or the writing of the manuscript.\u003c/p\u003e\n\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\n\u003cp\u003eEN, UK, and MS conceptualised the study. EN, UK, MS, JN, LCG, and TO developed the methodology. EN performed the investigation. EN and BO developed the software. EN, UK, XT, LA, EC, BA, and CPS performed data curation and formal analysis. EN prepared the visualisations. MS supervised the study. MS and LCG acquired funding. EN and MS wrote the original draft. UK, XT, LA, EC, BA, CPS, BO, BG, DKS, JN, LCG, and TO reviewed and edited the manuscript. All authors read and approved the final manuscript.\u003c/p\u003e\n\u003ch2\u003eAcknowledgement\u003c/h2\u003e\n\u003cp\u003eThe authors thank the CREEDS, Connectivity Map, Tahoe-100M, and Open Targets teams for making their data resources publicly available.\u003c/p\u003e\n\u003ch2\u003eData Availability\u003c/h2\u003e\n\u003cp\u003eAll data analysed in the current study are publicly available. The CDRPipe R package and R Shiny application are available at https://github.com/enockniyonkuru/drug_repurposing. The comparative analysis code supporting this manuscript is available at https://github.com/enockniyonkuru/cdrpipe-comparative-analysis. The R Shiny application is freely accessible at https://cdrpipe.org/. Disease signatures from CREEDS are available at https://maayanlab.cloud/CREEDS/. Known drug-disease associations from the Open Targets Platform (2024 release) are available at https://platform.opentargets.org/. Connectivity Map (CMap) perturbation profiles are available from the Broad Institute at https://www.broadinstitute.org/connectivity-map-cmap. Tahoe-100M perturbation profiles are available from Hugging Face at https://huggingface.co/datasets/tahoebio/Tahoe-100M.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eGamal H, Shoeib EM, Hajjaj A, Abdullah HEA, Elramy EH, Ellah DAA, et al. Incorporating AI, in silico, and CRISPR technologies to uncover the potential of repurposed drugs in cancer therapy. RSC Pharm. 2025;2:1019\u0026ndash;33.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCamps I, K\u0026uuml;nzel SR, Schubert M, Singh RK. Editorial: opportunities and challenges in drug repurposing. Front Pharmacol. 2025;16:1709217.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSun D, Gao W, Hu H, Zhou S. Why 90% of clinical drug development fails and how to improve it? Acta Pharm Sin B. 2022;12:3049\u0026ndash;62.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGolla U, Patel S, Shah N, Talamo S, Bhalodia R, Claxton D, et al. From deworming to cancer therapy: benzimidazoles in hematological malignancies. Cancers. 2024;16:3454.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDudley JT, Sirota M, Shenoy M, Pai RK, Roedder S, Chiang AP, et al. Computational repositioning of the anticonvulsant topiramate for inflammatory bowel disease. Sci Transl Med. 2011;3:96ra76.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOskotsky TT, Bhoja A, Bunis D, Le BL, Tang AS, Kosti I, et al. Identifying therapeutic candidates for endometriosis through a transcriptomics-based drug repositioning approach. iScience. 2024;27:109388.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSirota M, Schaub MA, Batzoglou S, Robinson WH, Butte AJ. Autoimmune disease classification by inverse association with SNP alleles. PLoS Genet. 2009;5:e1000792.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen K, Wang H. Discovery of disease relationships via transcriptomic signature analysis powered by agentic AI. arXiv. 2025. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.48550/arXiv.2508.04742\u003c/span\u003e\u003cspan address=\"10.48550/arXiv.2508.04742\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eIorio F, Rittman T, Ge H, Menden M, Saez-Rodriguez J. Transcriptional data: a new gateway to drug repositioning? Drug Discov Today. 2013;18:350\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSirota M, Dudley JT, Kim J, Chiang AP, Morgan AA, Sweet-Cordero A, et al. Discovery and preclinical validation of drug indications using compendia of public gene expression data. Sci Transl Med. 2011;3:96ra77.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMusa A, Ghoraie LS, Zhang S-D, Glazko G, Yli-Harja O, Dehmer M, et al. A review of connectivity map and computational approaches in pharmacogenomics. Brief Bioinform. 2017;19:506\u0026ndash;23.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLamb J, Crawford ED, Peck D, Modell JW, Blat IC, Wrobel MJ, et al. The Connectivity Map: using gene-expression signatures to connect small molecules, genes, and disease. Science. 2006;313:1929\u0026ndash;35.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSmith I, Scott K, Haibe-Kains B. Similarity bias from consensus perturbational signatures from the L1000 Connectivity Map. bioRxiv. 2024. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1101/2022.01.24.477615\u003c/span\u003e\u003cspan address=\"10.1101/2022.01.24.477615\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAissa AF, Islam ABMMK, Ariss MM, Go CC, Rader AE, Conrardy RD, et al. Single-cell transcriptional changes associated with drug tolerance and response to combination therapies in cancer. Nat Commun. 2021;12:1628.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePeidli S, Green TD, Shen C, Gross T, Min J, Garda S, et al. scPerturb: harmonized single-cell perturbation data. Nat Methods. 2024;21:531\u0026ndash;40.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang J, Ubas AA, de Borja R, Svensson V, Thomas N, Thakar N et al. Tahoe-100M: a giga-scale single-cell perturbation atlas for context-dependent gene function and cellular modeling. bioRxiv. 2025. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1101/2025.02.20.639398\u003c/span\u003e\u003cspan address=\"10.1101/2025.02.20.639398\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLei W, Yuan M, Long M, Zhang T, Huang Y-E, Liu H, et al. scDR: predicting drug response at single-cell resolution. Genes. 2023;14:268.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang Z, Monteiro CD, Jagodnik KM, Fernandez NF, Gundersen GW, Rouillard AD, et al. Extraction and analysis of signatures from the Gene Expression Omnibus by the crowd. Nat Commun. 2016;7:12846.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCarvalho-Silva D, Pierleoni A, Pignatelli M, Ong C, Fumis L, Karamanis N, et al. Open Targets Platform: new developments and updates two years on. Nucleic Acids Res. 2019;47:D1056\u0026ndash;65.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBarrett T, Wilhite SE, Ledoux P, Evangelista C, Kim IF, Tomashevsky M, et al. NCBI GEO: archive for functional genomics data sets - update. Nucleic Acids Res. 2013;41:D991\u0026ndash;5.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePhipson B, Smyth GK. Permutation P-values should never be zero: calculating exact P-values when permutations are randomly drawn. Stat Appl Genet Mol Biol. 2010;9. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.2202/1544-6115.1585\u003c/span\u003e\u003cspan address=\"10.2202/1544-6115.1585\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. :Article 39.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePetitdemange A, Blaess J, Sibilia J, Felten R, Arnaud L. Shared development of targeted therapies among autoimmune and inflammatory diseases: a systematic repurposing analysis. Ther Adv Musculoskelet Dis. 2020;12:1759720X20969261.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSumida TS, Lincoln MR, He L, Park Y, Ota M, Oguchi A, et al. An autoimmune transcriptional circuit drives FOXP3\u0026thinsp;+\u0026thinsp;regulatory T cell dysfunction. Sci Transl Med. 2024;16:eadp1720.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZondervan KT, Becker CM, Koga K, Missmer SA, Taylor RN. Vigan\u0026ograve; P. Endometriosis. Nat Rev Dis Primers. 2018;4:9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAnglesio MS, Papadopoulos N, Ayhan A, Nazeran TM, No\u0026euml; M, Horlings HM, et al. Cancer-associated mutations in endometriosis without cancer. N Engl J Med. 2017;376:1835\u0026ndash;48.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLin Y-H, Chen Y-H, Chang H-Y, Au H-K, Tzeng C-R, Huang Y-H. Chronic niche inflammation in endometriosis-associated infertility: current understanding and future therapeutic strategies. Int J Mol Sci. 2018;19:2385.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePerrone U, Barra F, Anatr\u0026agrave; M, Paudice M, Vellone VG, Gullo G, et al. Targeting inflammation in endometriosis: emerging therapeutic options. Expert Opin Investig Drugs. 2025;34:995\u0026ndash;1009.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGibbons T, Georgiou EX, Cheong YC, Wise MR. Levonorgestrel-releasing intrauterine device (LNG-IUD) for symptomatic endometriosis following surgery. Cochrane Database Syst Rev. 2021;2021:CD005072.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGiudice LC, Liu B, Irwin JC. Endometriosis and adenomyosis unveiled through single-cell glasses. Am J Obstet Gynecol. 2025;232:S105\u0026ndash;23.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBecker CM, Bokor A, Heikinheimo O, Horne A, Jansen F, Kiesel L, et al. ESHRE guideline: endometriosis. Hum Reprod Open. 2022;2022:hoac009.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVercellini P et al. [Reference to be confirmed by the authors: cited in text as Vercellini 2003 regarding medroxyprogesterone acetate for hormonal suppression of endometriotic lesions.].\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKamencic H, Thiel JA. Pentoxifylline after conservative surgery for endometriosis: a randomized, controlled trial. J Minim Invasive Gynecol. 2008;15:62\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGrammatis AL, Georgiou EX, Becker CM. Pentoxifylline for the treatment of endometriosis-associated pain and infertility. Cochrane Database Syst Rev. 2021;8:CD007677.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBruner-Tran KL, Osteen KG, Taylor HS, Sokalska A, Haines K, Duleba AJ. Resveratrol inhibits development of experimental endometriosis in vivo and reduces endometrial stromal cell invasiveness in vitro. Biol Reprod. 2011;84:106\u0026ndash;12.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBruner-Tran KL, Osteen KG, Duleba AJ. Simvastatin protects against the development of endometriosis in a nude mouse model. J Clin Endocrinol Metab. 2009;94:2489\u0026ndash;94.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu M, Liu X, Zhang Y, Guo S-W. Valproic acid and progestin inhibit lesion growth and reduce hyperalgesia in experimentally induced endometriosis in rats. Reprod Sci. 2012;19:360\u0026ndash;73.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShim JY, Garbo G, Grimstad FW, Scatoni A, Barrera EP, Boskey ER. Use of the drospirenone-only contraceptive pill in adolescents with endometriosis. J Pediatr Adolesc Gynecol. 2024;37:402\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePushpakom S, Iorio F, Eyers PA, Escott KJ, Hopper S, Wells A, et al. Drug repurposing: progress, challenges and recommendations. Nat Rev Drug Discov. 2019;18:41\u0026ndash;58.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLow LA, Mummery C, Berridge BR, Austin CP, Tagle DA. Organs-on-chips: into the next decade. Nat Rev Drug Discov. 2021;20:345\u0026ndash;61.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVlachogiannis G, Hedayat S, Vatsiou A, Jamin Y, Fern\u0026aacute;ndez-Mateos J, Khan K, et al. Patient-derived organoids model treatment response of metastatic gastrointestinal cancers. Science. 2018;359:920\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOskotsky TT, Tang X, Arthurs E, Govil A, Abbasi F, Bhoja A et al. A transcriptomics-based computational drug repositioning pipeline identifies simvastatin and primaquine as novel therapeutics for endometriosis pain. bioRxiv. 2025. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1101/2025.05.28.656743\u003c/span\u003e\u003cspan address=\"10.1101/2025.05.28.656743\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMoffat JG, Vincent F, Lee JA, Eder J, Prunotto M. Opportunities and challenges in phenotypic drug discovery: an industry perspective. Nat Rev Drug Discov. 2017;16:531\u0026ndash;43.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCorrigan-Curay J, Sacks L, Woodcock J. Real-world evidence and real-world data for evaluating drug safety and effectiveness. JAMA. 2018;320:867\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBarturen G, Beretta L, Cervera R, Van Vollenhoven R, Alarc\u0026oacute;n-Riquelme ME. Moving towards a molecular taxonomy of autoimmune rheumatic diseases. Nat Rev Rheumatol. 2018;14:75\u0026ndash;93.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSaranraj K, Kiran PU. Drug repurposing: clinical practices and regulatory pathways. Perspect Clin Res. 2025;16:61\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"genome-medicine","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"Learn more about [Genome Medicine](https://genomemedicine.biomedcentral.com/)","snPcode":"13073","submissionUrl":"https://submission.springernature.com/new-submission/13073/3","title":"Genome Medicine","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"BMC/SO AJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Drug repurposing, Transcriptomics, Signature reversal, Connectivity Map, Single-cell RNA sequencing, Tahoe-100M, Disease Signature, Drug Signature, Computational pharmacology","lastPublishedDoi":"10.21203/rs.3.rs-10606217/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-10606217/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e \u003cp\u003eTranscriptome-based drug repurposing can accelerate therapeutic discovery, but is limited by fragmented resources, inconsistent quality control, and reliance on single perturbation databases.\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e \u003cp\u003eWe developed CDRPipe (\u003cb\u003eC\u003c/b\u003eomputational \u003cb\u003eD\u003c/b\u003erug \u003cb\u003eR\u003c/b\u003eepurposing \u003cb\u003ePipe\u003c/b\u003eline), a unified framework that interrogates disease signatures against drug perturbation signatures generated by distinct experimental technologies. Specifically, CDRPipe harmonizes microarray perturbation profiles from the Connectivity Map (CMap; 1,968 quality-filtered experiments) with pseudo-bulk profiles derived from large-scale single-cell RNA sequencing experiments in the Tahoe-100M database (56,827 experiments). CDRPipe standardizes preprocessing, computes rank-based connectivity scores and evaluates significance using empirical null models. We applied CDRPipe to 233 curated disease signatures from GEO and CREEDS and evaluated performance using known drug-disease associations from Open Targets.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eSingle-cell-derived pseudo-bulk profiles recovered more annotated therapeutics than microarray profiles (mean recall 47.3% vs 18.5%; Wilcoxon signed-rank test p\u0026thinsp;\u0026lt;\u0026thinsp;10\u003csup\u003e\u0026ndash;11\u003c/sup\u003e), though these differences partly reflect differences in drug library composition and clinical annotation coverage. Importantly, the two resources were highly complementary, with only 3.5% overlap in recovered drugs, indicating that integrating predictions across independent perturbation resources expands therapeutic coverage and enables identification of high-confidence consensus candidates. Case studies in autoimmune disease and endometriosis further demonstrate that CDRPipe recovers clinically relevant therapies while revealing technology-dependent patterns of discovery.\u003c/p\u003e\u003ch2\u003eConclusions\u003c/h2\u003e \u003cp\u003eIntegrating heterogeneous transcriptomic perturbation resources improves the robustness and interpretability of transcriptional drug repurposing. Consensus predictions supported by independent resources represent higher-confidence candidates for downstream experimental and clinical validation, and CDRPipe provides an openly available framework to support this integrated approach.\u003c/p\u003e","manuscriptTitle":"Integrating single-cell and bulk transcriptomic perturbation resources reveals complementary therapeutic spaces for drug repurposing","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-08-31 10:54:36","doi":"10.21203/rs.3.rs-10606217/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"reviewerAgreed","content":"85526313338205503596942378378205742956","date":"2026-08-24T02:40:55+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-08-23T22:00:56+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-08-13T05:01:29+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-08-07T00:01:13+00:00","index":"","fulltext":""},{"type":"submitted","content":"Genome Medicine","date":"2026-08-05T18:04:21+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"genome-medicine","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"Learn more about [Genome Medicine](https://genomemedicine.biomedcentral.com/)","snPcode":"13073","submissionUrl":"https://submission.springernature.com/new-submission/13073/3","title":"Genome Medicine","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"BMC/SO AJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"a42f0bbf-f995-44bc-8ff9-6f3c02ba7376","owner":[],"postedDate":"August 31st, 2026","published":true,"recentEditorialEvents":[{"type":"reviewerAgreed","content":"85526313338205503596942378378205742956","date":"2026-08-24T02:40:55+00:00","index":12,"fulltext":""},{"type":"reviewersInvited","content":"11","date":"2026-08-23T22:00:56+00:00","index":"","fulltext":""}],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-08-31T10:54:39+00:00","versionOfRecord":[],"versionCreatedAt":"2026-08-31 10:54:36","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-10606217","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-10606217","identity":"rs-10606217","version":["v1"]},"buildId":"0shC4O-rRljfh4Nq9OyMh","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.