Identification of novel causally related genes in adenomyosis: An integrated summary data-based Mendelian randomization study and bioinformatics analysis

article OA: gold CC0 ⤵ 2 in-corpus citations
⚙ AI-generated summary by gemini-2.5-flash-lite, 2026-06-12 ⓘ

This study identified 24 novel protein-coding genes, including ARHGEF35 and INTS1, causally linked to adenomyosis, and found DNA2 and INTS1 were highly expressed in affected patients.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

⚙ AI-generated deep summary by claude@2026-06, 2026-06-12 · read from full text ⓘ

This paper used an integrated summary-data Mendelian randomization framework to test whether genetically regulated gene expression causally influences adenomyosis risk, performing a transcriptome-wide SMR analysis using FinnGen adenomyosis GWAS summary statistics (4,267 cases and 107,564 controls) with cis-eQTL instruments from GTEx whole blood and uterine tissue. The key outputs were prioritized candidate “causal” genes based on SMR and a multi-SNP SMR sensitivity analysis (qFDR<0.05), followed by GO/KEGG functional enrichment and validation of candidate genes for differential expression using GEO datasets from adenomyosis patients. A major caveat discussed by the authors is that MR/SMR relies on instrument validity assumptions (eQTL association with expression, independence from confounding, and conditional independence), and they note the potential for bias if pleiotropy remains despite multi-SNP sensitivity approaches. This paper is centrally about endometriosis/adenomyosis — specifically adenomyosis, aiming to identify novel causally related genes and pathways for adenomyosis using summary-based MR and bioinformatics.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Adenomyosis (AM) is recognized as a complex gynecological and endocrine disorder that contributes to infertility and elevates the risk of pregnancy complications; however, its underlying genetic basis remains unidentified. This study aimed to identify potentially causative genes that may relate to AM. We conducted a summary data-based Mendelian randomization analysis using single nucleotide polymorphisms as an instrumental variable, along with expression quantitative trait loci data from whole blood and uterus as exposures and AM as the outcome. Summary data-based Mendelian randomization incorporating multiple single nucleotide polymorphisms was employed as a sensitivity analysis to reduce the false-positive rate. The false discovery rate was used to adjust for multiple tests. Furthermore, bioinformatics analysis was performed to elucidate the biological functions in which the novel target risk genes may be involved and to evaluate the diagnostic performance of risk genes based on data from the Gene Expression Omnibus database. We have identified 24 novel protein-coding genes potentially causally linked to AM, none of which have been previously reported in the context of this disease. The most relevant candidate genes are ARHGEF35, AMT, RCVRN, GMPPB, and INTS1. Bioinformatics analysis indicates that these genes play critical roles in essential biological functions, including base-excision repair, negative regulation of various cell cycle processes, and metabolism-related pathways in AM. Differential gene expression analysis was conducted using the Gene Expression Omnibus database, comparing data from AM patients and healthy controls. Specifically, DNA2 and INTS1 displayed high expression levels, whereas EFCAB2, HLA-DQA2, and RPS26 exhibited low expression levels. The receiver operating characteristic curve analysis for the Predictive Diagnostic Index revealed an area under the curve of 0.8 for the combined analysis of the 5 risk genes. Our study identifies novel potentially causal genes associated with AM, suggesting that these genes hold promise as therapeutic targets and biomarkers for early diagnosis. These findings significantly enhance our understanding of the underlying mechanisms and provide new insights into potential curative targets for AM.
Full text 39,569 characters · extracted from pmc-nxml · 7 sections · click to expand

Intro

Adenomyosis (AM) is a prevalent gynecological endocrine disorder characterized by the invasion and proliferation of endometrial glands and stroma within the myometrium. [ 1 ] Despite its frequent occurrence, the etiology and pathogenesis of AM remain poorly understood, often rendering it a clinically enigmatic condition. [ 2 ] AM presents with endocrine abnormalities, including abnormal uterine bleeding, menorrhagia, reduced fertility, dysmenorrhea, dyspareunia, and an increased risk of adverse pregnancy outcomes. [ 3 ] These symptoms significantly impair the quality of life for affected women. [ 4 – 6 ] It is also a significant factor associated with adverse reproductive outcomes and elevated obstetric risks globally. [ 5 ] Unfortunately, effective conservative treatment options remain limited for the majority of AM patients, often necessitating hysterectomy as a definitive intervention. [ 7 , 8 ] This surgical approach, while providing symptom relief, results in irreversible loss of fertility and may lead to significant complications, including pelvic organ prolapse and urinary system dysfunction. [ 9 ] Consequently, managing AM presents a considerable clinical challenge. Given these problems, studying AM’s etiology is particularly important but still lags behind other gynecological diseases. [ 4 ] These challenges stem from previous studies’ limitations in fully elucidating the precise molecular mechanisms underlying AM. [ 6 ] Therefore, there is an urgent need to deepen our understanding of the causal factors associated with the onset and progression of AM to elucidate its complex mechanisms and develop effective therapeutic targets. Previous studies have indicated that hormones, genetic, epigenetic, immune, inflammatory, and environmental factors contribute to the development of AM [ 6 , 10 , 11 ] and have explored risk loci associated with these conditions. [ 12 ] However, the genomic characteristics of AM remain incompletely understood. [ 6 ] Consequently, there is a scarcity of studies capable of inferring the causality of risk genes in AM. It is important to note that the causal relationships underlying disease cannot be inferred solely from genome-wide association studies (GWAS). [ 13 ] Inspiringly, Mendelian randomization (MR) studies offer a promising avenue by using genetic variation as a natural experiment to draw causal inferences from observational data, potentially aiding in identifying risk genes. [ 14 ] In contrast, a MR framework, which integrates summary statistics from multiple GWAS, can better elucidate causal associations of genetic variants with disease. [ 15 ] The MR approach adheres to Mendel law of inheritance, which postulates that “alleles of parents are randomly assigned to offspring.” It utilizes genetic variation as instrumental variables (IVs) to derive causal inferences between exposure factors and outcomes in observational studies. [ 16 ] These IVs are typically single nucleotide polymorphisms (SNPs), which are randomly assigned at conception and remain unaffected by confounders and reverse causation, thus enhancing the reliability of MR in causal inference studies. [ 17 ] Furthermore, expression quantitative trait loci (eQTL) can reveal associations between genetic variants and gene expression traits. [ 18 ] Thus, summary data-based MR (SMR) analysis, which incorporates summary data from GWAS and eQTL within the MR framework, can effectively predict risk genes that may causally relate to diseases or traits. [ 19 ] Additionally, SMR analysis utilizing summary statistics from multiple SNPs has been employed as a sensitivity analysis to avoid potentially biased findings. [ 20 ] Regrettably, SMR analysis, which could aid in identifying risk genes associated with genetic variants, has yet to be applied in AM. To elucidate the causal relationship between host gene expression and AM, our study represents the first transcriptome-wide SMR analysis utilizing the top SNP in available eQTL as an instrumental variable. This approach allowed us to identify the causal links between potential risk genes as exposure factors and AM as the outcome. Additionally, we conducted functional enrichment analyses based on the Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway and gene ontology (GO) annotations to better understand the biological functions associated with the identified risk genes. Using the Gene Expression Omnibus (GEO) dataset, we further validated and analyzed candidate risk genes from AM patients. The resulting list of identified risk genes is anticipated to serve as biomarkers for early diagnosis. Consequently, our study offers a novel perspective on the genetic mechanisms underlying AM, contributing to research aimed at identifying effective therapeutic targets, thereby laying a foundation for improved clinical management of AM.

Author

Data curation: Jianhui Fu. Formal analysis: Qiaomei Yang, Jingxuan Hong, Hao Lin, Jianhui Fu. Funding acquisition: Qiaomei Yang, Junying Jiang. Investigation: Qiaomei Yang, Fuchun Zhong, Li Chen, Hao Lin. Methodology: Xianhua Liu, Li Chen, Xinye Zheng. Project administration: Jingxuan Hong. Resources: Xinye Zheng. Software: Xianhua Liu. Supervision: Jingxuan Hong, Junying Jiang. Validation: Jingxuan Hong, Hao Lin. Writing – original draft: Qiaomei Yang, Fuchun Zhong. Writing – review & editing: Junying Jiang.

Methods

Figure 1 illustrates the overall design of this study. Initially, we extracted the comprehensive summary data for AM from publicly accessible GWAS data. Then, we used the SMR analysis to evaluate the causality of transcriptome-wide gene expression as the exposure factor and AM as the outcome of this study. We used available eQTL studies to provide information on the effect of genetic variants on gene expression levels. The cis-eQTLs were obtained from whole blood and uterus from the Genotype-Tissue Expression (GTEx) project. We performed SMR analysis using the summary statistics from multiple SNPs as the sensitivity analysis to reduce the false-positive rate (q FDR<0.05). Subsequently, to enhance our understanding of the biological functions of potential causal genes, we conducted functional enrichment analyses utilizing GO annotations and KEGG pathways. Finally, we further validated and analyzed the candidate risk genes using datasets from the GEO specific to AM. The flowchart of the study. Our study adopted a multi-stage analysis methodology. The integrative summary-based Mendelian randomization analysis of GWAS summary data of adenomyosis with whole blood and uterus cis-eQTLs data was performed to identify potential causal genes of adenomyosis. Functional enrichment analysis of risk genes was conducted to identify significant gene ontologies and molecular pathways. Differential expression analysis was validated on transcriptomic gene expression data of adenomyosis versus healthy individuals. Identified risk genes by SMR analysis were differentially expressed in adenomyosis patients and promising to be biomarkers. eQTL = expression quantitative trait loci, GWAS = genome-wide association studies, SMR = summary data-based Mendelian randomization. Since gene expression levels vary across tissue types, eQTL studies are typically performed for a specific tissue type. [ 21 ] In this study, probes are utilized to measure a gene’s expression, and the cis-eQTLs for that gene are defined as any SNPs located within 1 Mb of the gene probe that show a significant association with the gene’s expression, as indicated by a P eQTL value of <5E-8. [ 20 ] The cis-eQTLs were obtained from the whole blood and uterine from v8 of the Genotype-Tissue Expression (GTEx) project with a sample size of 670 and 129, respectively. [ 22 ] The mapping genes for SNPs used in the MR analysis were obtained from the GWAS database ( https://www.ebi.ac.uk/gwas/ ). We collected the GWAS summary data of AM from FinnGen ( https://storage.googleapis.com/finngen-public-data-r9/summary_stats/finngen_R9_N14_ENDOMETRIOSIS_UTERUS.gz ), a public–private partnership that combines genotype and health data from Finnish biobanks and registries; a total of 4267 AM cases and 107,564 standard general population cases were analyzed in this study. This data used deidentified data from publicly available studies, so no ethical approval or informed consent was required. We selected IV for SMR analysis following the standard approach described previously. [ 23 ] The top cis-eQTL within the cis-region of a probe having the most potent effect on the gene’s expression was captured as the IV for the SMR analysis to evaluate the association between a gene’s expression level in the specific tissue and the outcome. Only genes with at least 1 cis-eQTL( P eQTL  < 5E-8) were included in the analysis. [ 24 ] SNPs with an allele frequency of 0.2 among the 3 datasets: GWAS summary statistics, eQTL summary statistics, and linkage disequilibrium (LD) reference panel data. [ 25 ] The mathematical model underlying MR relies on 3 main assumptions for accurately inferring causality: the IV has to be (i) associated with the exposure, (ii) independent of any confounder of the exposure–outcome association, and (iii) conditionally independent of the outcome given the exposure and confounders. [ 15 , 17 ] The schematic diagram of the MR study is presented in Figure 2 A. Criteria for the selection of IV in both uterine tissues and whole blood were as described above. Illustration of the MR and SMR method. (A) Flow chart of the MR. The red cross sign means heritable variables are not associated with confounders and outcomes. (B) In the illustration of the primary analysis of SMR, the most associated cis-eQTL for each gene (SNPtop) was used as the genetic instrument to estimate the effect of the expression level of the gene on the outcome. (C) Illustration of the sensitivity analysis of multi-SNPs SMR, all independent cis-eQTLs for a gene ( r 2  < 0.1 with the top cis-eQTL) which were significant ( P eQTL < 5E − 8) were used as the genetic instruments to estimate the effect of the expression level of the gene on the outcome. Both analyses were repeated for whole blood and uterus and adenomyosis outcomes. eQTL = expression quantitative trait loci, GWAS = genome-wide association studies, MR = Mendelian randomization, SMR = summary data-based Mendelian randomization, SNPs = single nucleotide polymorphism. The initial threshold set for identifying independent cis-eQTLs was r 2  < 0.9. Nonetheless, SNPs with strong LD are considered to have r 2 values of 0.1 or greater, which may fall within this threshold. Therefore, we defined independent SNPs using an LD threshold of r 2  < 0.1. [ 26 ] To calculate LD estimates for the analysis, we utilized the reference panel from phase 3 of the 1000 Genomes Project ( https://www.internationalgenome.org/data-portal/sample ) [ 20 ] . The LD reference panel was generated using VCFtools 0.1.15, [ 27 ] BCFtools 1.11, [ 28 ] and PLINK 1.90 ( https://www.cog-genomics.org/plink/1.9/ ). [ 29 ] In addition, we employed the multi-SNP SMR method as a sensitivity analysis to mitigate potential bias that could arise from using a single variant as the genetic instrument in SMR, which may confound causal associations with horizontal pleiotropy. By incorporating multiple instrumental variables, particularly within the cis-region of the probe, the likelihood of horizontal pleiotropy was reduced, while the statistical power of the analysis was enhanced. [ 30 ] For the multi-SNP SMR analysis, we used all independent cis-eQTLs, defined as SNPs with an r 2 value of <0.1 relative to the top cis-eQTL, and with P -values <5E − 8 within the cis-region, as instrumental variables. Detailed descriptions of the SMR and multi-SNP SMR methods have been provided in previous studies. [ 23 , 25 ] The SMR approach integrates summary statistics from GWAS and eQTL studies within the MR framework to prioritize genes with expression levels potentially causally related to an outcome trait. [ 31 ] MR study is an IV analysis that uses genetic variants, such as SNPs, as exposure proxies. In this study, we performed an SMR analysis based on established methodologies from prior research, which describe the SMR and multi-SNP SMR techniques, illustrated in Figure 2 B and C. [ 20 ] All data analyses were completed using SMR v1.03 ( https://cnsgenomics.com/software/smr/ ) and R version 4.2.1 ( https://www.r-project.org/ ). We conducted a primary analysis, utilizing the cis-eQTL with the strongest association signal for each gene as the single genetic instrument. [ 32 , 33 ] While using a single SNP instrument in MR analysis can be effective, [ 34 ] employing a single eQTL as the instrument may lead to biased results and inflated false-positive rates. [ 35 ] Recognizing that much of the genetic variation accounts for only a tiny fraction of trait variation, primary analyses using single SNP instruments require large sample sizes, often in the 10s of 1000s, to achieve sufficient statistical power. [ 36 ] Consequently, using a single variant as the genetic instrument in primary analysis may be unable to differentiate between association due to causation and horizontal pleiotropy, whereby the association with the exposure might be independent of the association with the outcome. Palmer et al proposed using multiple SNP instruments in MR analysis to address the limitations of single SNP instruments. [ 37 ] Furthermore, a causal estimate derived from multiple causal SNPs is generally more precise than 1 derived from a single SNP. [ 36 , 38 ] Therefore, we conducted a multi-SNP SMR as a sensitivity analysis in our study, [ 20 ] incorporating all independent cis-eQTLs with P  < 5E-8 within the cis-region as instrumental variables. [ 23 ] Only gene–trait associations that showed significant results in primary and sensitivity analyses were considered potential pathogenic genes. We corrected for multiple testing using the FDR q value, with q values <0.05 considered statistically significant. The analysis was conducted using available eQTL data for AM outcomes in both whole blood and uterine tissues. After conducting the SMR analysis, we performed the HEIDI test to differentiate between pleiotropy and linkage in the associations identified by the SMR analysis. [ 39 , 40 ] Associations with a P HEIDI  < .05 were considered to exhibit significant heterogeneity and were therefore excluded from further consideration. [ 23 ] The HEIDI test evaluates the null hypothesis that the association identified by the SMR test is attributable to pleiotropy, explicitly determining whether it originates from a single causal variant rather than 2 variants in high LD. During the HEIDI test, cis-eQTLs exhibiting strong LD ( r 2  > 0.9) or weak LD ( r 2  < 0.05) were excluded. The HEIDI test was also performed only when the number of cis-eQTLs included in the analysis was ≥ 3. [ 25 ] We performed GO and KEGG pathway enrichment analysis with R package clusterProfiler to gain insight into the biological function of the potential causal genes associated with AM. The GO analysis annotates genes in a hierarchically structured manner for biological processes (BP), molecular functions, and cellular components. The enrichment analysis results were visualized with an enriched plot of the R package. The high-throughput transcriptomic datasets relevant to AM were sourced from the GEO database ( https://www.ncbi.nlm.nih.gov/geo ). Due to the limited availability of transcriptomic data on AM in the GEO database and the small sample sizes of individual datasets, datasets GSE190580 and GSE193928 were integrated for a more comprehensive analysis. [ 41 ] The GSE190580 dataset provided preprocessed expression data as a gene expression matrix. These data were already normalized and were directly used for downstream analyses. The GSE193928 dataset consisted of raw sequencing data, which required comprehensive preprocessing. Specifically, the raw sequencing data from dataset GSE193928 were subjected to quality control using FastQC (version 0.12.1) to assess the quality of the raw reads. Subsequently, trimming and adapter removal were performed using Trim Galore (version 0.6.7). The cleaned sequencing reads were aligned to the human genome (GRCh38) utilizing the STAR aligner (version 2.7.10a), a widely adopted tool for efficient alignment of RNA-Seq data. [ 42 ] Post-alignment gene expression quantification was conducted using featureCounts, [ 43 ] allowing for precisely counting aligned reads to annotated genes. To mitigate systematic discrepancies and ensure data consistency across different datasets, the remove BatchEffect function from the limma package (version 3.56.2) was employed to address and correct batch effects. Differential expression analysis was performed using DESeq2 (version 1.40.2). This analysis focused on identifying risk genes associated with AM, which was informed by SMR analysis. DESeq2 offers robust modeling of count data for deriving statistically significant differential expression results.nrichment plots were generated using R to visualize the analysis outcomes, emphasizing biologically relevant patterns in gene expression. Following this, a logistic regression model was constructed to develop a diagnostic tool based on characteristic genes identified as key discriminators of AM. To evaluate the performance of the diagnostic model, a receiver operating characteristic (ROC) curve was generated using the pROC package (version 1.18.4). All statistical analyses were conducted using R Studio (version 4.2.1). In the SMR analysis, associations with P -values below .05 (including raw P , P FDR , and P HEIDI ) were considered strong indicators of causal relationships. Due to the study’s exploratory nature and the limited sample size, multiple testing correction was not applied. Instead, a threshold of P   1.5 was used to identify differentially expressed genes in the bioinformatics analysis, consistent with previous studies. [ 44 ]

Results

In the transcriptome-wide analysis of whole blood, we used univariable SMR analysis integrated with multi-SNP SMR sensitivity analysis to elucidate potential causal relationships between gene expression profiles and AM susceptibility. The identification of significant associations was performed under the HEIDI test framework ( P HEIDI  > .05) with multiple testing corrections (FDR q value < 0.05). We identified 356 genes in whole blood with significant AM-related associations (Table S1, Supplemental Digital Content, https://links.lww.com/MD/P130 ). The top 5 most associated protein-coding genes were C6orf211 , also known as Acidic Residue Methyltransferase 1 ( ARMT1 ), Carboxypeptidase Vitellogenic Like ( CPVL ), Killer Cell Immunoglobulin Like Receptor-Two Ig Domains And Short Cytoplasmic Tail 4 ( KIR2DS4 ), Rho Guanine Nucleotide Exchange Factor 35 ( ARHGEF35 ), and RNA Polymerase III Subunit G ( POLR3G ). Notably, rs2758600 SNP tagging genes ARHGEF35 and OR2A20P showed significant association with AM in SMR analyses, but OR2A20P is a Pseudogene. The top 20 risk genes associated with AM in whole blood are displayed in Table 1 . The top 20 risk genes show an association with AM in whole blood. eQTL = expression quantitative trait loci, GWAS = genome-wide association studies, HEIDI = heterogeneity in the dependent instrument, ProbeChr = probe chromosome, SMR = summary data-based Mendelian randomization, SNP = single nucleotide polymorphism. We identified a total of 80 gene-AM associations in the uterus that were significant in both SMR and multi-SNPs SMR analysis and that were not being rejected by the HEIDI test (FDR q value .05) (Table S2, Supplemental Digital Content, https://links.lww.com/MD/P131 ). The top 5 protein-coding genes were ARHGEF35 , GDP-Mannose Pyrophosphorylase B ( GMPPB ), Zinc Finger Protein 880 ( ZNF880 ), Recoverin ( RCVRN ), and Hyaluronidase 3 ( HYAL3 ). The top 20 genes associated with AM in the uterus are displayed in Table 2 . The top 20 risk genes show an association with AM in the uterus. eQTL = expression quantitative trait loci, GWAS = genome-wide association studies, HEIDI = heterogeneity in the dependent instrument, ProbeChr = probe chromosome, SMR = summary data-based Mendelian randomization, SNP = single nucleotide polymorphism. We identified the same 39 risk genes in both whole blood and uterine tissue with gene–trait associations of AM by univariable SMR and multi-SNPs SMR analysis with satisfying for FDR q value .05 displayed in Table 3 . Among these candidate risk genes, 15 of them are nonprotein coding genes, including pseudogene ( HLA-DRB6 , RP11-958N24.2 , ROCK1P , CCT6P3 , AC125232.1 , RP11-460N20.5 , and P SMD5-AS1 ), lincRNA ( AC093627.10 , RP11-384K6.6 , CKMT2-AS1 , and RP11-196G11.2 ), sense intronic ( CTA-390C10.10 and AF131215.2 ), antisense ( AP006621.6 and RP11-597D13.9 ), and the rest 24 are all protein-coding genes. Notably, the top 10 most significantly associated genes are all protein-coding genes, including ARHGEF35 , Aminomethyltransferase ( AMT ), RCVRN , GDP-Mannose Pyrophosphorylase B( GMPPB ), Integrator Complex Subunit 1 ( INTS1 ), Hyaluronidase 3 ( HYAL3 ), Ribosomal Protein S26 ( RPS26 ), R.B. Transcriptional Corepressor Like 2 ( RBL2 ), Zinc Finger Protein 880 ( ZNF880 ) and Nei Like DNA Glycosylase 2 ( NEIL2 ). The significantly correlated same-risk genes with AM were present in the whole blood and uterus. AM = adenomyosis, SMR = summary data-based Mendelian randomization. SMR analysis assessed the causal correlation of target gene expression in both whole blood and uterus with AM risk. ARHGEF35 was selected as a representative gene to illustrate the methodology and findings, with the results presented in Figure 3 and Table 4 . This analytical approach is consistent with established methodologies in similar studies, such as MR analyses examining the association of target gene expression in blood with breast cancer risk, [ 45 ] where representative genes are frequently utilized to demonstrate key insights and validate analytical frameworks. Our study inferred the effect of ARHGEF35 on the trait risk of AM was negatively correlated. We also discovered ARHGEF35 acting on gene trait that was the most significantly associated with AM both in whole blood and uterus (PSMR = 1.23E − 3, β=−0.1626 versus PSMR = 2.23E − 3, β=−0.0912, respectively). Risk target genes were identified in this study that the exact negative correlation with ARHGEF35 and AM in gene expression-trait included AMT , RPS26 , EFCAB2 , POLR2J2 , PLA2G16 , NSUN2 , HYAL3 , MRPL39 , and FAM118A . The positively associated risk genes of gene expression-trait risk in AM were GMPPB , RBL2 , ZNF880 , LY6G5B , INTS1 , DNA2 , RCVRN , DUS3L , MX1 , GPX7 , NEIL2 , HLA-DQA2 , ERV3-1 , WRAP73 , and DUS3L . The relationship of risk genes acting on both expression levels and traits in AM. eQTL = expression quantitative trait loci, GWAS = genome-wide association studies, HEIDI = heterogeneity in the dependent instrument, SMR = summary data-based Mendelian randomization. SMR analysis association of risk loci for gene expression in both blood and the uterus of adenomyosis. (A) In the top panel, gray dots represent the P -values for SNPs from the adenomyosis GWAS in whole blood. The bottom plot shows the eQTL P -values of the ENSG00000213214 probe SNP tagging ARHGEF35 . Highlighted in red are ARHGEF35 genes that passed the SMR and HEIDI tests. (B) The correlation of ARHGEF35 between eQTL effect sizes and GWAS effect sizes in blood shows a negative correlation. The orange dashed lines represent the estimate of the effect size of the MR association at the top cis-eQTL.(C) In the top panel, gray dots represent the P -values of SNP from adenomyosis GWAS in the uterus. The lower panel shows the eQTL P -values of SNP from ENSG00000213214 probe marker ARHGEF35 in the uterine of adenomyosis. Highlighted in red is the ARHGEF35 gene tested by SMR and HEIDI. (D) The correlation of ARHGEF35 between eQTL effect sizes and GWAS effect sizes in the uterus shows a negative correlation. The orange dashed lines represent the estimate of the effect size of the MR association at the top cis-eQTL. Error bars are the standard errors for SNP effects. eQTL = expression quantitative trait loci, GWAS = genome-wide association studies, HEIDI = heterogeneity in the dependent instrument, MR = Mendelian randomization, SMR = summary data-based Mendelian randomization, SNPs = single nucleotide polymorphism. To better understand the potential biological functions of the identified pathogenic genes in this study, functional enrichment analysis was conducted for the genes identified via SMR analysis. The adjusted P -value (rather than the raw P -value) was used for both GO and KEGG analyses to control for false discovery rates. Enrichment analysis of the risk genes identified in both the uterus and whole blood was presented in Figure 4 and Table 5 . The GO enrichment results indicated that the causal relationship of genes was primarily enriched in BP such as base-excision repair, negative regulation of cell cycle phase transition, negative regulation of cell cycle process, tRNA modification, and negative regulation of cell cycle (Fig. 4 A). The cellular component enrichment suggested the risk genes were associated with the spindle, sperm flagellum, chromatoid body,9 + 2 motile cilium, and mitochondrial matrix (Fig. 4 B). The molecular function enrichment revealed that the causal genes were involved in catalytic activity, acting on a tRNA, hydrolase activity, acting on glycosyl bonds, catalytic activity, acting on RNA, hydroxymethyl, formyl- and related transferase activity, and endodeoxyribonuclease activity, producing 5’-phosphomonoester (Fig. 4 C). Additionally, the KEGG pathway analysis showed that the pathways with the lowest adjusted P -values included glycosaminoglycan degradation, lipoic acid metabolism, alpha-linolenic acid metabolism, linoleic acid metabolism, and glyoxylate and dicarboxylate metabolism (Fig. 4 D). Top 10 enrichment analysis items of the risk genes for protein-coding with AM. BP = biological processes, CC = cellular components, GO = gene ontology, MF = molecular functions, KEGG = Kyoto Encyclopedia of Genes and Genomes. Enrichment analysis of identification risk genes. (A) Bubble diagram depicting GO biological processes enrichment analysis of risk genes. (B) Bubble diagram illustrating GO cellular components enrichment analysis of risk genes. (C) Bubble diagram displaying the GO molecular functions enrichment analysis of risk genes. (D) Bubble diagram showing the KEGG pathway enrichment analysis of risk genes. GO = gene ontology, KEGG = Kyoto Encyclopedia of Genes and Genomes. To investigate the expression profiles of causal genes in AM patients, we analyzed 2 datasets obtained from the GEO database, namely GSE190580 and GSE193928 . Both datasets comprised uterine tissue samples, encompassing 47 samples in total: 18 from AM patients and 29 from non-AM (control) individuals. This analysis identified 4014 differentially expressed genes (Fig. 5 A). Subsequently, we intersected these 4014 differentially expressed genes with 24 causative genes previously identified via SMR analysis, identifying 5 key risk genes: DNA2 , HLA-DQA2 , INTS1 , EFCAB2 , and RPS26 . Further quantitative analysis revealed that DNA2 , INTS1 , and RPS26 were markedly overexpressed in AM tissues, whereas HLA-DQA2 and EFCAB2 exhibited significantly reduced expression levels (Fig. 5 B and C). This differential expression pattern underscores their potential as biomarkers for AM. Validation of key risk genes and Identification of biomarkers in adenomyosis. (A) Distribution statistics of causal risk genes and differential expression genes. (B) Venn diagram of the intersection of causal risk genes and differential expression genes. (C) Box plots of 5 differentially expressed risk genes.* P  < .05, ** P  < .01,*** P  < .001. (D) ROC curves of DNA2 , HLA-DQA2 , INTS1 , EFCAB2 , and RPS26 expression levels and of including conjoint analysis. ROC = receiver operating characteristic. We performed logistic regression modeling to evaluate the differentially expressed risk genes’ diagnostic potential. The combined analysis of the 5 risk genes yielded a Predictive Diagnostic Index, with the ROC curve demonstrating an area under the curve (AUC) of 0.8, indicating potential predictive performance to the current cohort. Individual risk genes also showed notable diagnostic potential, with AUC values as follows: DNA2 (AUC = 0.69), HLA-DQA2 (AUC = 0.68), INTS1 (AUC = 0.68), EFCAB2 (AUC = 0.67), and RPS26 (AUC = 0.6) (Fig. 5 D). These findings suggest that each gene contributes uniquely to the model’s diagnostic capability, highlighting their potential utility in clinical applications.

Discussion

In this study, we conducted the SMR analysis between both whole blood and uterus transcriptome-wide eQTL datasets as exposure factors and AM as the outcome. We identified 39 potentially risk genes associated with AM, and 24 of these genes are protein-coding, with the remaining 15 for nonprotein coding. Sensitivity analyses confirmed the reliability of the results. GO and KEGG pathway enrichment analyses indicated that the identified risk genes are associated with specific BP and pathways, including base-excision repair, regulation of cell cycle progression, tRNA modification, and lipid metabolism. In addition, we identified 5 differentially expressed genes that may be associated with AM. These findings provide novel insights into the potential molecular mechanisms underlying AM and highlight promising therapeutic targets for future intervention. While these genes were found to be differentially expressed in our study, their specificity to AM has not been evaluated, as they may also be involved in other conditions. Further studies are needed to determine whether these genes are uniquely linked to AM or are shared across multiple diseases. Our study lays a foundation for further research and offers a scientific direction for exploring the pathogenesis and treatment of AM. To our knowledge, this is the first study to use SMR to investigate the genetic basis of AM. This study enhances our understanding of the genetic factors involved in AM by identifying 24 novel protein-coding genes. These include genes with a positive correlation, such as RCVRN , GMPPB, INTS1 , RBL2 , ZNF880 , and others, as well as those with a negative correlation, including ARHGEF35 , AMT , HYAL3 , RPS26 , NSUN2 , among others. While the functions of these identified genes are known, none have been previously linked to AM. Further research is required to determine their association with the AM conclusively. Differential gene expression analysis was conducted using the GEO database, comparing data from AM patients and healthy controls. The analysis revealed high DNA2 and INTS1 expression and low expression of EFCAB2 , HLA-DQA2 , and RPS26 . Logistic regression modeling was employed further to investigate these differentially expressed genes’ diagnostic potential. The combined analysis of these 5 risk genes yielded a preliminary predictive model, as evidenced by ROC curve analysis with an AUC of 0.8. These findings suggest that these genes may serve as promising biomarkers or therapeutic targets, providing a scientific foundation for developing early diagnostic strategies for AM. Our study identified that the 10 most relevant risk genes associated with AM are all protein-coding. The genes GMPPB , RCVRN , INTS1 , RBL2 , ZNF880 , and NEIL2 show a positive correlation with the condition, while ARHGEF35 , AMT , HYAL3 , and RPS26 exhibit a negative correlation. Many of these genes have established roles in other diseases. For instance, GMPPB overexpression promotes glioblastoma cell proliferation, migration, and invasion. [ 46 ] RBL2 provides antiapoptotic effects in cardiomyocytes. [ 47 ] NEIL2 inhibits viral replication and prevents excessive inflammatory responses in CoV-2 pathogenesis. [ 48 ] HYAL3 is a potential marker for bladder cancer prognosis, [ 49 ] and RPS26 mutations result in fewer physical malformations than Diamond–Blackfan anemia cases. [ 50 ] The function of these genes is consistent with the occurrence and development of AM to some extent. Thus, these genes may also play an important decision-making role in the pathogenesis of AM and are worthy of further investigation. To elucidate the potential biological mechanisms underlying AM, we conducted comprehensive enrichment analyses to identify novel GO BP and KEGG pathways associated with the disease. Our findings revealed significant enrichment in BP related to base-excision repair, regulation of cell cycle phase transitions, and tRNA modification and processing. Notably, deficiencies in base-excision repair genes, such as NEIL2 and DNA2 identified in our study, have been previously implicated in inflammation, aging, and neurodegenerative disorders, suggesting their potential as therapeutic targets for AM. [ 51 ] Furthermore, we identified key risk genes involved in cell cycle regulation, including RBL2, NSUN2, and DNA2, which are particularly interesting given their established roles in cancer biology. [ 52 ] KEGG pathway analysis further highlighted the involvement of lipid-related metabolism pathways, providing novel insights into the metabolic dysregulation associated with AM. These findings collectively suggest a potential causal relationship between the identified risk genes and AM pathogenesis, offering valuable directions for future mechanistic investigations. However, it is essential to note that, as indicated in Table 5 , most enriched terms failed to achieve statistical significance following rigorous multiple testing correction using adjusted P -values to control for false discovery rates. While these results provide intriguing preliminary insights into the molecular mechanisms of AM, the current evidence remains insufficient to establish significant functional enrichment conclusively. Therefore, these findings should be interpreted cautiously, and their biological relevance to AM pathogenesis requires further substantiation. Future studies incorporating more enormous, more diverse cohorts, advanced functional genomic approaches (e.g., CRISPR-based screening, single-cell RNA sequencing), [ 53 , 54 ] and targeted experimental validation will be crucial to confirm these associations and delineate the precise molecular pathways driving AM development and progression. Such efforts may ultimately contribute to identifying novel therapeutic targets and developing more effective treatment strategies for this complex gynecological disorder. Moreover, our combined analysis of 5 risk genes showed a substantial area under the receiver operating characteristic curve (AUC = 0.8), suggesting potential predictive performance within the current cohort. This offers preliminary evidence for identifying risk genes associated with AM, highlighting their diagnostic potential. However, these findings may not apply to broader populations due to the specific sample set. Thus, the results should be seen as exploratory, and their clinical relevance requires validation in more extensive, diverse cohorts. Future research should increase sample sizes and use rigorous validation methods, like cross-validation or independent testing, to confirm the diagnostic efficacy and causal roles of these risk genes, such as sensitivity, specificity, and positive predictive value (PPV), which are critical for evaluating the clinical utility of the predictive model. [ 55 , 56 ] Functional and mechanistic studies will also be essential to understand their roles in AM pathogenesis. Although this study provides novel insights into AM’s molecular mechanisms and potential therapeutic targets, clinical application requires further validation and optimization through larger-scale investigations. This study has some notable limitations that should be acknowledged. Firstly, the SMR analysis relied exclusively on SNPs derived from individuals of European ancestry, which may restrict the generalizability of the findings to populations of non-European descent or from other continents. Future SMR studies should include diverse populations, such as those from Asia and Africa. Secondly, the SMR analysis primarily focuses on innate genetic risk factors for AM, potentially overlooking the influence of acquired characteristics, such as social and environmental factors. Additionally, the lack of multiple testing corrections in our bioinformatics analysis may increase the risk of false positives. Due to the limited availability of AM-specific public datasets, we did not partition the data into training and validation sets for external validation or calculate key diagnostic metrics such as sensitivity, specificity, or positive predictive value. Furthermore, while the SMR analysis identified potential causal genes associated with AM, the underlying biological mechanisms remain unexplored. Despite these limitations, our study has several strengths. This is the first study to apply SMR analysis to identify putative functional causal genes linked to AM. The SMR approach is powerful for inferring causal relationships between risk factors and clinically relevant outcomes. [ 57 ] This method utilizes genetic variants as instrumental variables to assess the effects of putative risk factors on gene expression, phenotypic traits, or disease states, thereby providing robust evidence for causality in complex biological systems. The focus on European ancestry helped minimize population heterogeneity, thereby enhancing the validity of the results. Furthermore, our study provides a comprehensive analysis of the genetic contribution to AM, including expression profiling and rigorous bioinformatics association analyses, which may offer valuable preliminary insights and lay a foundation for future research. Finally, we validated risk genes with differential expression in AM patients, underscoring their potential as precise diagnostic markers and therapeutic targets.

Conclusions

In conclusion, our study identified 39 risk genes potentially causally associated with AM through SMR analysis based on genome-wide transcriptomic data of blood and uterine tissues. Notably, 24 of these genes are novel protein-coding genes. Enrichment analysis has highlighted new potential biological pathways implicated in AM. Using GEO database datasets, we further validated the differential expression of key risk genes, including DNA2, INTS1, EFCAB2, HLA-DQA2, and RPS26. These findings emphasize the significance of genetic variation in the pathogenesis of AM, advancing our understanding of its molecular underpinnings. This research lays the groundwork for discovering targeted therapeutic interventions and developing novel biomarkers for early diagnosis. Future studies are essential to explore the relationship between these predominantly identified genes and AM, potentially guiding new approaches to treatment and diagnosis.

Acknowledgments

We thank citexs ( www.citexs.com ) for editing the English language.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

⚙ Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml ⓘ

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Condition tags

adenomyosisinfertility

MeSH descriptors

Adenomyosis Adenomyosis Adenomyosis Adenomyosis Adenomyosis Adenomyosis Adenomyosis Adenomyosis Adenomyosis Adenomyosis Adenomyosis Adenomyosis Adenomyosis Adenomyosis Adenomyosis Adenomyosis Adenomyosis Adenomyosis Adenomyosis Adenomyosis

Citation neighborhood

Papers in the corpus that this work cites (lower rings, blue) and that cite this one (upper rings, green). Dot size scales with the paper's in-corpus citation count — bigger dot = more influential within the endo/adeno field. Click a dot to open that paper. [ expand to 2 hops ] — adds papers reached through this work's immediate citers/citees. Heavier; up to 60 extra dots.

References (60)

Cited by (4)

Source provenance

europepmc
last seen: 2026-09-27T09:11:36.575535+00:00
openalex
last seen: 2026-06-10T17:14:06.276822+00:00
pmc
last seen: 2026-05-13T20:22:03.195721+00:00
pubmed
last seen: 2026-09-28T06:09:00.134067+00:00
License: CC0 · commercial use OK