{"paper_id":"5b5c8ee6-e067-47c2-884c-c8b5bed9eb7c","body_text":"TRACE: A FINE-TUNED BIOMEDICAL LANGUAGE MODEL FOR DIRECTIONALLY \nINFORMED DRUG REPURPOSING FROM TRANSCRIPTOME-WIDE ASSOCIATION \nSTUDIES \nChristopher O. Otieno \nDivision of Epidemiology, Department of Medicine, Vanderbilt Health \nNashville, TN 37232, USA \nEmail: christopher.otieno@vanderbilt.edu \nHannah M. Seagle \nDivision of Epidemiology, Department of Medicine, Vanderbilt Health \nNashville, TN 37232, USA \nEmail: hannah.m.seagle@vanderbilt.edu \nAlexis T. Akerele  \nDivision of Quantitative and Clinical Science, Department of Obstetrics and Gynecology, Department of \nBiomedical Informatics, Vanderbilt Health \nNashville, TN 37232, USA \nEmail: alexis.pigg@vumc.org \nJames Jaworski \nDivision of Epidemiology, Department of Medicine, Vanderbilt Health \nNashville, TN 37232, USA \nEmail: james.jaworski@vumc.org \nLindsay Guare \nDepartment of Pathology and Laboratory Medicine, University of Pennsylvania \nPhiladelphia, PA 19104, USA \nEmail: Lindsay.Guare@Pennmedicine.upenn.edu \nShefali Setia-Verma  \nDepartment of Pathology and Laboratory Medicine, Department of Biostatistics, Epidemiology and \nInformatics, University of Pennsylvania \nPhiladelphia, PA 19104, USA \nEmail: Shefali.SetiaVerma@pennmedicine.upenn.edu \nDigna R. Velez Edwards \nDivision of Quantitative and Clinical Science, Department of Obstetrics and Gynecology, Department of \nBiomedical Informatics, Vanderbilt Health \nNashville, TN 37232, USA \nEmail: digna.r.velez.edwards@vumc.org \nTodd L. Edwards \nDivision of Epidemiology, Department of Medicine, Vanderbilt Health \nNashville, TN 37232, USA \nEmail: todd.l.edwards@vumc.org \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 28, 2026. ; https://doi.org/10.64898/2026.08.25.26361263doi: medRxiv preprint \nNOTE: This preprint reports new research that has not been certified by peer review and should not be used to guide clinical practice.\n\nAbstract/Summary \nTranscriptome-wide association studies (TWAS) can identify genes where genetically predicted gene \nexpression is associated with disease risk, but translating those signals into therapeutic opportunities \nremains time-consuming, manual, and difficult to reproduce. We developed TRACE (TWAS -driven \nRepurposing through AI-assisted Curation of Evidence), a gene- and phenotype-agnostic computational \npipeline that accepts a TWAS gene and effect -size direction, normalizes the gene symbol, retrieves \nFDA-approved drug -gene candidates from four online resources, collects related peer-reviewed \nliterature from PubMed, and uses a fine -tuned biomedical language model to classify whether the \nliterature supports a direct drug -gene relationship, the mechanism of action, and the direction of effect. \nThe pipeline then compares the drug -derived direction with the direction implied by the TWAS effect \nestimate to rank candidate therapeutic pairs and flag potential drug safety concerns. The local classifier, \nbuilt on BiomedBERT, was trained using pipeline -derived labels, BioCreative VI ChemProt  gold-\nstandard chemical -protein relation examples, and author -reviewed active -learning cases, reaching a \nheld-out macro F1 of 0.809 across three simultaneous classification tasks . We validated the pipeline \nagainst a manually curated endometriosis gold standard of 43 drug -gene pairs spanning six TWAS-\nidentified genes, developed through S -PrediXcan analysis of endometriosis GWAS summary statistics, \nmanual querying of four drug -gene interaction databases for each gene, literature review of drug -gene \nmechanistic evidence, and Mendelian randomization validation of candidate pairs. External validation \nused two independently published genetically informed drug -repurposing studies in metabolic \ndysfunction-associated steatotic  liver disease (MASLD) and type 2 diabetes (T2D). The pipeline \nrecovered 90.7% of endometriosis pairs, 88.2% of MASLD pairs, and 92.9% of T2D pairs that were \npresent in at least one queried database. Applied to 99 endometriosis -associated TWAS genes, the \npipeline identified 1,089 FDA -approved drug -gene pairs, 32 candidate therapeutic pairs, and 77 \npotential safety concerns, including independent recovery of leuprolide acetate, an established \nendometriosis therapy. This framework provides a scalable, literature -grounded bridge from TWAS \ndiscovery to prioritized therapeutic hypotheses, while preserving uncertainty through manual -review \nflags and requiring downstream Mendelian randomization, electronic health record -based validation, \nand experimental follow-up before clinical interpretation. \nKeywords: drug repurposing; transcriptome-wide association study; biomedical language model; natural \nlanguage processing; precision medicine; endometriosis \n1. Introduction \nDrug development remains expensive, time-consuming, and failure -prone: recent empirical \nestimates place the average cost of developing a new drug in the range of hundreds of millions to \nseveral billion dollars, and reviews of drug repurposing emphasize the appeal of starting from \ncompounds with established safety, clinical, preclinical, and formulation knowledge that may \nreduce risk, cost, and development timelines 1, 2. \nRepurposing approved or investigational drugs can therefore shorten the path from biological \ninsight to testable intervention, but prioritization still requires a credible link between disease \nbiology, drug target, mechanism, and expected direction of effect 2. Human genetic support \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 28, 2026. ; https://doi.org/10.64898/2026.08.25.26361263doi: medRxiv preprint \n\nimproves that link: drug mechanisms with genetic support have been estimated to be 2.6 times \nmore likely to succeed clinically than mechanisms without such support3. \nTranscriptome-wide association studies ( TWAS) methods provide one way to connect genetic \nvariation to disease biology by testing whether genetically predicted gene expression is associated \nwith a phenotype. PrediXcan tests genetically predicted expression against traits using reference \ntranscriptome models, and S-PrediXcan extends this framework to genome-wide association study \n(GWAS) summary statistics 4, 5. \nFor drug repurposing, the direction of a TWAS or S -PrediXcan association can help prioritize \ntherapeutic mechanisms. If increased genetically predicted expression of a gene is associated with \nincreased disease risk, then an inhibitor or downregulator is directionally concordant with \ntherapeutic benefit. If decreased genetically predicted expression is associated with increased \ndisease risk, then an activator or upregulator is directionally concordant. Prior genetically informed \ndrug-repurposing studies have used this logic to connect genetically predicted expression, drug \ntargets, Mendelian randomization (MR), and electronic health record (EHR) - based validation and \nnominate candidate therapeutic inputs for multiple phenotypes 6, 7. \nThe bottleneck is that the step between a TWAS gene list and a directionally appropriate drug \ncandidate is still commonly performed by manual curation. Investigators search drug -gene \ndatabases, read mechanistic literature, decide whether evidence suggests a drug inhibits, activates, \nupregulates, or downregulates the implicated gene, and then compares that direction with the \nTWAS effect estimate 7, 8. This approach is biologically interpretable but does not scale easily to \nhundreds of genes and is difficult to reproduce across teams. \nA related class of drug -repurposing approaches based on systems biology principles compares \ndisease-level transcriptomic signatures with drug -induced gene network expression profiles, as \nintroduced by the Connectivity Map and extended by later connectivity -mapping resources9. These \napproaches scale well and are valuable for whole -signature reversal, but they do not necessarily \nprovide a gene -specific, literature-grounded explanation for a particular TWAS -nominated target. \nThis can be an important element of subsequent efforts in animal models to expand indications and \nmove to clinical trials.  \nWe developed TRACE, a computational pipeline that automates the gene-by-gene curation step \nand implements the directionality logic that makes TWAS results useful for therapeutic hypothesis \ngeneration. The pipeline combines TWAS -derived effect direction, multi -database drug -gene \nretrieval, FDA approval filtering, PubMed evidence retrieval, and a fine -tuned biomedical \nlanguage model that classifies drug -gene relationship support, mechanism, and direction from \nabstracts. We validate the system against three previously completed drug repurposing projects, \ntwo of which have been published,  and apply it to endometriosis -associated TWAS genes as a \nproof-of-concept for precision -medicine drug repurposing  and nominate several candidates with \nrepurposing potential. \n2. Methods \n2.1. TRACE Pipeline overview \nThe pipeline accepts  a gene symbol and a risk direction derived from a TWAS or PrediXcan \neffect estimate as input. It returns ranked FDA -approved drug-gene pairs classified as candidate \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 28, 2026. ; https://doi.org/10.64898/2026.08.25.26361263doi: medRxiv preprint \n\ntherapeutic pairs, potential safety concerns, unclear -direction pairs requiring manual review, or \ndrug-gene pairs without a specified disease -risk direction. The same code can be applied to any \nphenotype with a gene list and effect -size directions. Risk directions must be supplied by the user \nbased on the sign of the effect estimate from their TWAS or S -PrediXcan output; a positive effect \nestimate corresponds to increased_gene_increases_risk and a negative effect estimate corresponds \nto decreased_gene_increases_risk. Once genes have been grouped by direction, the pipeline can be \nrun on all genes in a group using the provided batch shell script, which loops over a user -supplied \ngene list and submits each gene with the appropriate direction flag. Users with large gene lists, for \nexample several hundred genes from a GWAS  or TWAS, do not need to run genes individually. \nGenes sharing the same direction can be batched together in a single command, and the two \ndirection groups can be submitted sequentially or in parallel on a computing cluster. The pipeline \nsource code and batch script examples are publicly available at \nhttps://github.com/otienoco/TRACE. \n \n \nFig. 1. Overview of the directionally informed TWAS-to-therapeutics pipeline. A gene and TWAS effect \ndirection are converted into FDA-approved drug-gene candidates, PubMed evidence, local BiomedBERT \nclassifications, and prioritized therapeutic or safety hypotheses. \n2.2. Gene normalization and drug candidate retrieval \nInput gene symbols are resolved to official HGNC nomenclature through the HGNC REST \nApplication Programming Interface ( API), which supports queries against approved symbols, \naliases, and previous symbols, allowing submitted symbols to be mapped to canonical HGNC \nidentifiers before downstream database queries10. \nCandidate drugs are retrieved from four drug-gene interaction resources. Drug Gene Interaction \nDatabase ( DGIdb) 4.0 aggregates drug -gene interactions and druggable -gene information from \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 28, 2026. ; https://doi.org/10.64898/2026.08.25.26361263doi: medRxiv preprint \n\npublications, databases, and web resources11. Open Targets is an open -source platform designed to \nsupport systematic drug-target identification and prioritization 12. CTD manually curates chemical -\ngene, chemical-disease, and related toxicogenomic relationships 13. Pharos is the web interface for \nthe NIH Illuminating the Druggable Genome Knowledge Management Center and provides \nintegrated protein, target, and druggability information 14. \nDrug candidates from the four resources are pooled and deduplicated by normalized drug \nname. Candidates are then filtered for confirmed United States FDA approval using the openFDA \ndrug-labeling API, with RxNorm used to normalize brand names, generic names, and related \nclinical drug identifiers before matching across sources 15, 16 . Only confirmed FDA -approved \ncompounds are advanced to literature retrieval and language-model classification steps. \n2.3. PubMed evidence retrieval \nFor each FDA -approved drug -gene pair, PubMed abstracts are retrieved using the NCBI E -\nutilities API. Queries combine the official gene symbol and normalized drug name; PMIDs \nprovided by source drug -gene databases are added as supplementary evidence when available. Up \nto seven abstracts per pair are retained by default, with each record storing the title, abstract text, \nPMID, and publication year. This value is configurable in config.py for users who wish to increase \nor decrease literature coverage per pair. \n2.4. Evidence classification with a fine-tuned biomedical language model \nEach abstract is classified on three axes: whether it supports a direct gene -level drug -gene \nrelationship, the mechanism of the drug action on the gene, and the direction of the drug effect on \ngene expression or activity. Abstracts that document a relationship but do not state a clear direction \nare assigned an unclear-direction label rather than being forced into a mechanistic category. \nThe local classifier is built on BiomedBERT  (microsoft/BiomedNLP-BiomedBERT-base-\nuncased-abstract-fulltext), a BERT -based biomedical language model pretrained on PubMed \nabstracts and PubMed Central full-text articles17. We added three classification heads to the [CLS] \ntoken representation for relationship detection, mechanism classification, and direction \nclassification. The model was trained with a weighted cross -entropy loss summed across the three \ntasks. Training used the AdamW optimizer with a linear warmup schedule, a learning rate of \n2×10⁻⁵, a batch size of eight, and five epochs on CPU. \nDuring development, Claude API outputs were used to generate structured candidate labels and \nto identify disagreement cases for active learning. However, all validation and application results \nreported here were generated with the locally fine -tuned model only, with no external API calls \nduring result generation. This design preserves transparency about model development while \nensuring that the reported pipeline can run locally and without per-query API cost. \n2.5. Training data and held-out model performance \nTraining data were built iteratively across four versions. Version 1 included 3,866 pipeline -\nderived examples. Version 2 added independent -disease validation examples. Version 3 \nincorporated the BioCreative VI ChemProt corpus, a manually annotated gold -standard corpus for \nchemical-protein relation extraction. ChemProt  contains PubMed abstracts annotated for chemical \nand gene/protein entities and compound -protein relation types, including mechanistic relationships \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 28, 2026. ; https://doi.org/10.64898/2026.08.25.26361263doi: medRxiv preprint \n\nrelevant to drug -target interpretation. This addition increased the number of expert-annotated \ntraining examples and exposed the classifier to curated chemical -protein relation labels rather than \nonly pipeline -derived labels 18. Version 4 added high -confidence active -learning disagreement \ncases reviewed during development. \n \n \nVersion Examples Gold-standard examples Key addition Macro F1 \nV1 3,866 188 Initial pipeline-labeled \nexamples \n0.620 \nV2 5,806 188 Independent-disease \nvalidation queries \n0.648 \nV3 18,040 12,422 BioCreative VI ChemProt \ncorpus \n0.808 \nV4 18,880 12,422 Active-learning \ndisagreement labels \n0.809 \nTable 1. Training dataset versions and held-out macro F1. \nThe final model was evaluated on a held -out set of 2,833 examples. Relationship detection \nachieved 91% accuracy and 0.88 macro F1; mechanism classification achieved 84% accuracy and \n0.78 macro F1; direction classification achieved 83% accuracy and 0.73 macro F1. Mean macro F1 \nacross the three tasks was 0.809. \nTask Accuracy Macro F1 Weighted F1 \nRelationship detection 91% 0.88 0.91 \nMechanism classification 84% 0.78 0.84 \nDirection classification 83% 0.73 0.83 \nOverall mean n/a 0.809 n/a \nTable 2. Held-out performance of the final local classifier by task. \n2.6. Directionality classification and prioritization \nEach gene is assigned  a therapeutic direction from its TWAS effect estimate. A positive effect \nestimate, meaning increased genetically predicted expression is associated with increased disease \nrisk, implies that inhibition or downregulation is the directionally concordant therapeutic \nhypothesis. A negative effect estimate implies that activation or upregulation is the directionally \nconcordant therapeutic hypothesis. This follows the genetically informed drug -repurposing logic \nused in prior S-PrediXcan and MR-based studies6, 7. \nEach drug-gene pair is then classified by comparing the literature -derived drug-effect direction \nwith the gene-risk direction. A concordant pair is labeled candidate_therapeutic_pair; a discordant \npair is labeled potential_safety_concern; a pair with documented relationship evidence but \ninsufficient directionality is labeled unclear_direction_manual_review; and a pair with no disease -\nrisk direction is labeled drug_gene_pair_only . Candidate pairs are ranked using a composite score \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 28, 2026. ; https://doi.org/10.64898/2026.08.25.26361263doi: medRxiv preprint \n\nincorporating evidence strength, number of supporting databases, classifier confidence, FDA \napproval, and supporting PubMed abstract count. \n2.7. Validation strategy \nWe evaluated  TRACE against three disease settings. Internal validation used 43 manually \ncurated endometriosis drug-gene pairs across six TWAS -identified genes. External validation used \npublished drug -gene pairs from a metabolic dysfunction-associated steatotic liver disease \n(MASLD) genetically informed repurposing study and from a type 2 diabetes EHR -validated \ngenetically supported drug -repurposing pipeline 7, 8 . Drug aliases and salt forms were matched \nusing a manually curated alias table, and all validation runs used the local classifier only. \n3. Results \n3.1. Internal validation of TRACE in endometriosis \nAgainst the 43 -pair manually curated endometriosis gold standard, the pipeline recovered \n90.7% of pairs overall, including 88.9% of candidate therapeutic pairs and 92.0% of potential \nsafety concern pairs. Exact-classification accuracy across the full gold standard was 46.5%, largely \nbecause many retrieved pairs were conservatively assigned to unclear_direction_manual_review \nrather than forced into candidate or safety categories when abstracts did not state a clear \nmechanism or direction. \nMetric Value \nOverall recall 90.7% (39/43) \nTherapeutic pair recall 88.9% (16/18) \nSafety pair recall 92.0% (23/25) \nExact classification across full gold standard  46.5% (20/43) \nTable 3. Internal validation results in the endometriosis gold standard. \nThe four unrecovered endometriosis pairs involved broad drug -class terms such as progestin \nthat are not represented as specific FDA -approved compounds in the queried drug -gene databases. \nThis failure mode reflects a database representation issue rather than a literature -classification \nfailure. \n3.2. External validation in MASLD and type 2 diabetes \nIn the MASLD validation set, the pipeline recovered 88.2% of published drug -gene pairs. In \nthe type 2 diabetes validation set, raw recall was 65.0%; however, six of seven unrecovered pairs \nwere explained by database coverage or mapping limitations, including withdrawn drugs, \nliterature-supported pairs absent from the four queried databases, and calcium -channel subunit \nmapping inconsistencies. Restricting the denominator to pairs present in at least one queried \ndatabase yielded 92.9% adjusted recall. \nDisease Source Pairs Recall Adjusted recall \nMASLD Seagle et al., 2025 34 88.2% (30/34) n/a \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 28, 2026. ; https://doi.org/10.64898/2026.08.25.26361263doi: medRxiv preprint \n\n(PMID: 40780050) \nType 2 diabetes Shuey et al., 2023 (PMID: \n37399599) \n20 65.0% (13/20) 92.9% (13/14) \nTable 4. External validation across two independent disease applications. \nAcross the three disease settings, the pipeline recovered 88 -93% of curated or database -\nfindable pairs, suggesting that the main remaining limitation is drug -gene database coverage rather \nthan PubMed retrieval or local language-model classification. \n3.3. Application to endometriosis-associated TWAS genes \nWe applied TRACE to 99 endometriosis -associated TWAS genes identified through S -\nPrediXcan analysis across 49 human tissues from the Genotype -Tissue Expression (GTEx) project \nversion 8. Genes were retained if they passed a false discovery rate threshold of 0.05 and a \ncolocalization posterior probability threshold of 0.60, yielding 49 genes requiring an inhibitory or \ndownregulating therapeutic effect, 45 genes requiring an activating or upregulating effect, and five \ngenes with ambiguous naming.  Of the 99 genes, 75 had at least one FDA -approved drug \ninteraction across the four queried databases. Across those druggable genes, the pipeline identified \n1,089 FDA -approved drug -gene pairs, including 32 candidate therapeutic pairs and 77 potential \nsafety concerns.  The full 99 -gene analysis was completed in approximately two hours on a \nstandard high-performance computing cluster node using the local classifier, at no per -query API \ncost, with individual gene queries requiring approximately 10 to 20 minutes depending on the \nnumber of FDA-approved candidates retrieved. \nThe pipeline recovered 39 of 43 manually curated endometriosis validation pairs when run \nwith the local model, matching the recall obtained during development with API -assisted \nclassification. This indicates that the local classifier reproduced the retrieval and prioritization \nperformance needed for the application analysis while avoiding external API use during final result \ngeneration. \nBeyond the manually curated genes, the pipeline surfaced 27 additional candidate therapeutic \npairs across 21 genes. For GNRH1, the pipeline independently identified leuprolide acetate, a \nGnRH agonist that has randomized -trial evidence supporting its efficacy in endometriosis -\nassociated pain and disease suppression 19-21. We treat this result as an internal positive control \nshowing that the pipeline can rediscover a clinically established endometriosis therapy from \nautomated evidence retrieval and classification alone. \nGene Drug Pipeline mechanism Biological note \nGNRH1 Leuprolide acetate Upregulator/agonist Established endometriosis therapy; \nrandomized-trial support for \nleuprolide in endometriosis (PMID: \n2118858; PMID: 9916956; PMID: \n9464714). \nCALCRL Erenumab Antagonist CGRP receptor antagonist used for \nmigraine prevention; CGRP has been \nstudied in endometriosis and \nendometriosis-migraine comorbidity \n(PMID: 33934575; PMID: 38658933). \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 28, 2026. ; https://doi.org/10.64898/2026.08.25.26361263doi: medRxiv preprint \n\nRIT1 Selumetinib Inhibitor MEK inhibitor; MAPK signaling has \nbeen implicated in endometriosis \nbiology (PMID: 33590617; PMID: \n21303778). \nPGR Diethylstilbestrol Upregulator Identified through CTD; listed as a \nhypothesis requiring careful manual \nreview rather than a therapeutic \nrecommendation. \nTable 5. Selected candidate therapeutic pairs identified by the pipeline in the endometriosis application. \n4. Discussion \nWe developed and validated TRACE, a scalable computational pipeline that automates a key \nbottleneck in genetically informed drug repurposing:  converting TWAS -nominated genes and \neffect directions into literature -supported, directionally classified drug -gene hypotheses. Th is \nmethod differs from prior manual genetically informed repurposing workflows by extracting drug -\ngene direction from PubMed evidence at the pair level, applying the same logic uniformly across \nlarge gene lists, and replacing API -assisted development labels with a local fine -tuned biomedical \nlanguage model for final validation and application , creating a scalable, reproducible drug \nrepurposing pipeline. \nThe main strength of the framework is that it preserves the interpretability of gene -specific \ncuration while improving scalability and reproducibility. Rather than assigning a drug direction \nfrom a broad therapeutic class or clinical indication, the model evaluates abstracts for evidence \nabout the specific drug-gene pair. This is important because the same compound can have different \nmechanisms at different targets or in different biological contexts. \nThe validation results support the use of the pipeline as a high -recall prioritization and triage \ntool. In endometriosis, MASLD, and type 2 diabetes, the pipeline recovered most curated or \ndatabase-findable pairs. When it could not assign a confident direction, it frequently used the \nunclear_direction_manual_review category rather than forcing a potentially incorrect candidate or \nsafety label. This conservative behavior reduces overclaiming and makes the output more \nappropriate for translational hypothesis generation. \nThe endometriosis application illustrates how the system can expand manual review . A human \nteam might reasonably focus on a few high -priority genes, but the pipeline applies the same \nprocess to all genes with assigned risk direction, producing candidate pairs, safety concerns, and \nmanual-review cases across the full TWAS gene set. The independent recovery of leuprolide \nacetate provides  a positive control 19, 21, 22 , while candidates such as erenumab and selumetinib \nillustrate how the method can propose biologically plausible hypotheses that require downstream \nvalidation. \nThe pipeline does not establish therapeutic efficacy. It produces prioritized hypotheses that \nshould be followed by the same validation steps used in prior genetically informed repurposing \nwork, including M R with gene -expression or protein -level proxies, EHR -based validation when \nappropriate, and experimental assays or structural modeling when biologically informative6-8. \nThis study also has limitations. First, the pipeline depends on the coverage and quality of \nsource drug-gene databases; pairs absent from DGIdb, Open Targets, CTD, and Pharos cannot be \nrecovered unless future versions add additional resources. Second, restricting the search to FDA -\n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 28, 2026. ; https://doi.org/10.64898/2026.08.25.26361263doi: medRxiv preprint \n\napproved drugs increases immediate translational relevance but excludes investigational drugs and \ncompounds approved only outside the United States.  Future versions of this package will include \nan option to label by FDA status rather than filter.  Third, some mechanism classes remain \nunderrepresented in the training data, limiting classifier performance for rare classes. Fourth, the \nunclear-direction category was common, reflecting the fact that abstracts often document drug -\ngene relationships without stating a direct expression or activity direction. \nFuture work will focus on expanding active -learning labels across additional diseases and \nincorporating additional drug-target and pharmacogenomic resources. Because the current input is \nsimply a gene symbol and effect -size direction, TRACE is readily applicable to any complex trait \nor disease with TWAS, PrediXcan , or mapped gene results. As GWAS and TWAS resources \ncontinue to grow across diverse ancestries and tissues, tools that can rapidly and reproducibly \ntranslate genetic discovery into directionally informed therapeutic hypotheses will become \nincreasingly valuable. TRACE provides that bridge, offering a scalable, cost -free, and literature -\ngrounded approach to drug repurposing that can accelerate the path from genetic association to \nprioritized candidates for experimental and clinical follow-up. \n \nAcknowledgments \nResearch reported in this publication was supported by the Eunice Kennedy Shriver National \nInstitute of Child Health and Human Development of the National Institutes of Health under award \nnumber R01HD110567. H.M.S. was supported by American Heart Association grant \n26PRE1550935 and National Institutes of Health grant TL1TR002244. C.O.O. was supported by \nNational Institutes of Health grant TL1TR002244. A.T.A. was supported by National Institutes of \nHealth grant T32 CA160056.  Preprint of an article submitted for consideration in Pacific \nSymposium on Biocomputing © 2027 World Scientific Publishing Co., Singapore, \nhttp://psb.stanford.edu/. The authors wish to dedicate this work to the memory of Dr. Ephraim \nOtieno, MD, whose thoughtful curiosity and encouragement during the development of this work \nare deeply appreciated. \n \nData Availability \nThe TRACE pipeline source code is publicly available at https://github.com/otienoco/TRACE. The \nfine-tuned BiomedBERT classifier weights are available at \nhttps://huggingface.co/otienoco/TRACE-classifier and are downloaded automatically on first run \nwhen using the local classifier. \nLLM Disclosure \nThe Claude Application Programming Interface (API) (Anthropic) was used during pipeline \ndevelopment as an evidence classifier and as one source of labels for active -learning disagreement \ncases, as described in Methods. All validation and endometriosis application results reported in this \nmanuscript were generated using the locally fine -tuned model only, with no Claude API \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 28, 2026. ; https://doi.org/10.64898/2026.08.25.26361263doi: medRxiv preprint \n\ninvolvement during result generation. Claude was also used as a coding assistant during pipeline \ndevelopment.  \n \n \n \nReferences \n \n1. Wouters OJ, McKee M, Luyten J. Estimated research and development investment needed \nto bring a new medicine to market, 2009-2018. Jama. 2020;323(9):844–53. \n2. Pushpakom S, Iorio F, Eyers PA, Escott KJ, Hopper S, Wells A, et al. Drug repurposing: \nprogress, challenges and recommendations. Nature reviews Drug discovery. 2019;18(1):41–58. \n3. Minikel EV, Painter JL, Dong CC, Nelson MR. Refining the impact of genetic evidence on \nclinical success. Nature. 2024;629(8012):624–9. \n4. Gamazon ER, Wheeler HE, Shah KP, Mozaffari SV, Aquino -Michaels K, Carroll RJ, et al. \nA gene -based association method for mapping traits using reference transcriptome data. Nature \ngenetics. 2015;47(9):1091–8. \n5. Barbeira AN, Dickinson SP, Bonazzola R, Zheng J, Wheeler HE, Torres JM, et al. \nExploring the phenotypic consequences of tissue specific gene expression variation inferred from \nGWAS summary statistics. Nature communications. 2018;9(1):1825. \n6. Khankari NK, Keaton JM, Walker VM, Lee KM, Shuey MM, Clarke SL, et al. Using \nMendelian randomisation to identify opportunities for type 2 diabetes prevention by repurposing \nmedications used for lipid management. EBioMedicine. 2022;80. \n7. Shuey MM, Lee KM, Keaton J, Khankari NK, Breeyear JH, Walker VM, et al. A \ngenetically supported drug repurposing pipeline for diabetes treatment using electronic health \nrecords. EBioMedicine. 2023;94. \n8. Seagle HM, Akerele AT, DeCorte JA, Hellwege JN, Breeyear JH, Kim J, et al. Genomics -\ninformed drug-repurposing strategy identifies two therapeutic targets for preventing liver disease \nassociated with metabolic dysfunction. The American Journal of Human Genetics. \n2025;112(8):1778–91. \n9. Lamb J, Crawford ED, Peck D, Modell JW, Blat IC, Wrobel MJ, et al. The Connectivity \nMap: using gene -expression signatures to connect small molecules, genes, and disease. science. \n2006;313(5795):1929–35. \n10. Seal RL, Braschi B, Gray K, Jones TE, Tweedie S, Haim -Vilmovsky L, et al. Genenames. \norg: the HGNC resources in 2023. Nucleic acids research. 2023;51(D1):D1003–D9. \n11. Freshour SL, Kiwala S, Cotto KC, Coffman AC, McMichael JF, Song JJ, et al. Integration \nof the Drug–Gene Interaction Database (DGIdb 4.0) with open crowdsource efforts. Nucleic acids \nresearch. 2021;49(D1):D1144–D51. \n12. Ochoa D, Hercules A, Carmona M, Suveges D, Baker J, Malangone C, et al. The next -\ngeneration Open Targets Platform: reimagined, redesigned, rebuilt. Nucleic acids research. \n2023;51(D1):D1353–D9. \n13. Davis AP, Wiegers TC, Johnson RJ, Sciaky D, Wiegers J, Mattingly CJ. Comparative \ntoxicogenomics database (CTD): update 2023. Nucleic acids research. 2023;51(D1):D1257–D62. \n14. Nguyen D -T, Mathias S, Bologa C, Brunak S, Fernandez N, Gaulton A, et al. Pharos: \ncollating protein information to shed light on the druggable genome. Nucleic acids research. \n2017;45(D1):D995–D1002. \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 28, 2026. ; https://doi.org/10.64898/2026.08.25.26361263doi: medRxiv preprint \n\n15. Kass-Hout TA, Xu Z, Mohebbi M, Nelsen H, Baker A, Levine J, et al. OpenFDA: an \ninnovative platform providing access to a wealth of FDA’s publicly available data. Journal of the \nAmerican Medical Informatics Association. 2016;23(3):596–600. \n16. Nelson SJ, Zeng K, Kilbourne J, Powell T, Moore R. Normalized names for clinical drugs: \nRxNorm at 6 years. Journal of the American Medical Informatics Association. 2011;18(4):441–8. \n17. Gu Y, Tinn R, Cheng H, Lucas M, Usuyama N, Liu X, et al. Domain -specific language \nmodel pretraining for biomedical natural language processing. ACM Transactions on Computing \nfor Healthcare (HEALTH). 2021;3(1):1–23. \n18. Krallinger M, Rabal O, Akhondi SA, Pérez MP, Santamaría J, Rodríguez GP, et al., editors. \nOverview of the BioCreative VI chemical -protein interaction Track. Proceedings of the sixth \nBioCreative challenge evaluation workshop; 2017. \n19. Dlugi AM, Miller JD, Knittle J, Group LS. Lupron depot (leuprolide acetate for depot \nsuspension) in the treatment of endometriosis: a randomized, placebo -controlled, double -blind \nstudy. Fertility and sterility. 1990;54(3):419–27. \n20. Ling FW. Randomized controlled trial of depot leuprolide in patients with chronic pelvic \npain and clinically suspected endometriosis. Obstetrics & Gynecology. 1999;93(1):51–8. \n21. Hornstein MD, Surrey ES, Weisberg GW, Casino LA, Group LA -BS. Leuprolide acetate \ndepot and hormonal add -back in endometriosis: a 12 -month study. Obstetrics & Gynecology. \n1998;91(1):16–24. \n22. Wheeler JM, Knittle JD, Miller JD. Depot leuprolide versus danazol in treatment of women \nwith symptomatic endometriosis: I. Efficacy results. American journal of obstetrics and \ngynecology. 1992;167(5):1367–71. \n \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 28, 2026. ; https://doi.org/10.64898/2026.08.25.26361263doi: medRxiv preprint","source_license":"CC-BY-4.0","license_restricted":false}