Abstract
The menstrual cycle is one of the most fundamental biological rhythms in human physiology, yet its systemic molecular changes remain poorly understood. Here we show that the menstrual cycle is accompanied by widespread changes in the circulating proteome. By profiling nearly 3,000 plasma proteins in 2,760 women from the UK Biobank, we identified 198 proteins that vary across the cycle, forming distinct temporal patterns aligned with menstrual phases. These proteins include reproductive hormones, cytokines and growth factors, many of which are enriched in endometrial tissue and expressed in epithelial and stromal cell types, highlighting their biological specificity. Several proteins were linked to common reproductive disorders, including endometriosis, leiomyoma and abnormal bleeding. Finally, we developed a proteomic score on the basis of 75 proteins that accurately predicts menstrual cycle phase. Together, these findings provide a systems-level atlas of menstrual cycle biology and inform biomarker discovery in women’s health.
Similar content being viewed by others
Main
The human menstrual cycle is a complex, hormonally regulated process that not only regulates cyclical changes within the reproductive tract but also exerts systemic effects across multiple organ systems. While the roles of estrogen, progesterone, luteinizing hormone (LH) and follicle-stimulating hormone (FSH) in driving endometrial proliferation, ovulation and luteal transformation are well established, accumulating evidence indicates that menstrual cycle-related hormonal fluctuations also modulate immune function, metabolism and cardiovascular physiology1,2,3,4,5. However, the molecular pathways underlying these broader systemic effects remain incompletely understood.
Circulating proteins carry signals from multiple tissues and biological pathways, providing a dynamic readout of both local and systemic physiology. Plasma proteomics therefore offers a unique opportunity to map menstrual cycle-dependent molecular changes at scale. So far, most studies of menstrual biology have relied on targeted hormone assays, small panels of immune mediators6 or analyses of menstrual blood and endometrial tissue, typically with modest sample sizes and limited temporal resolution7,8. Large-scale, untargeted proteomic profiling across the menstrual cycle could reveal previously unrecognized protein dynamics, uncover phase-specific biomarkers and improve the interpretation of proteomic studies in women of reproductive age.
Here, we performed comprehensive plasma proteomic profiling of 2,760 healthy women from the UK Biobank (UKB), quantifying over 2,900 proteins across the menstrual cycle. By integrating menstrual cycle timing with protein, tissue and cell-type-specific analyses, we created a high-resolution atlas of cyclical protein variation. We further applied Mendelian randomization (MR) to explore potential causal links between cycle-associated proteins and female reproductive health, and developed a proteomic score capable of predicting menstrual cycle phase from a single blood sample. These data provide insights into the molecular physiology of the menstrual cycle and may enhance health diagnostics in female individuals of reproductive age.
Results
Plasma proteomic signatures of menstrual cycle day
A total of 2,760 women with 2,917 proteins measured were included in this study. Protein identifiers and assay quality control metrics are provided in Supplementary Table 1. The median age was 45 years, the median menstrual cycle length was 28 days and the median time since last menstruation was 12 days (Table 1). We identified 198 proteins (Fig. 1a and Supplementary Table 2) that were significantly associated with cycle day at a false discovery rate (FDR) <0.05. The five strongest associations were observed for PROK1 (P = 2.6 × 10−92), MMP10 (P = 1.1 × 10−61), RLN1 (P = 8.3 × 10−42), CXCL13 (P = 1.5 × 10−41) and CGA (P = 3.8 × 10−36). CGA is the common α-subunit shared by FSH, LH and human chorionic gonadotropin (hCG), and FSHB, the β-subunit of FSH (P = 4.7 × 10−17); both represent core reproductive endocrine markers and served as positive controls. To investigate the stability of the results across age groups, we divided the population at the median age. In younger women (≤45 years; N = 1,421), 100 proteins significantly associated with cycle day at FDR 45 years; N = 1,339). We observed a strong concordance in effect sizes (r = 0.936, P = 9.2 × 10−301) for all 198 cycle-dependent proteins across the two age strata. To determine how much of the proteome varies with menstrual cycle timing, the variance explained (R2) was calculated. Cycle day explained >5% of variance in only nine proteins (0.31%), whereas the corresponding numbers were 391 (13.4%) for age and 332 (11.3%) for body mass index (BMI) (Fig. 1b,c and Supplementary Table 3). This suggests that cycle timing has large effects on only a subset of proteins, while age and BMI explain a larger proportion of variance across a broader set of proteins.
Temporal proteomic signatures
K-means clustering was applied to explore dynamic changes in protein expression across the menstrual cycle. Four distinct clusters were identified, each with a characteristic temporal profile (Fig. 2). Cluster 1 (n = 53 proteins) included proteins with peak expression on day 1 of menstruation, cluster 2 (n = 69) showed peak expression during the early follicular phase, cluster 3 (n = 20) peaked around the late follicular and periovulatory phase and cluster 4 (n = 56) showed the highest expression during the mid- and late luteal phase. Representative proteins from each cluster included MMP10, PAEP and LEFTY2 in cluster 1; CGA, FSHB and TNFSF10 in cluster 2; STC2, SFRP4 and CRLF1 in cluster 3; and PROK1, RLN1 and CXCL13 in cluster 4. The full list of proteins assigned to each cluster is provided in Supplementary Table 4.
Tissue and cell type enrichment
To investigate the tissue specificity of proteins associated with the cycle day, we ran tissue enrichment analyses using data on 40 tissues from the Human Protein Atlas. Significant enrichment was detected for proteins specifically expressed in the endometrium and placenta (Fig. 3a and Supplementary Table 5). Given the observed tissue enrichment in the endometrium and its central role in the menstrual cycle, single-cell RNA sequencing (scRNA-seq) data from ten individuals were incorporated to identify potentially relevant cell types for the cycle-associated proteins (Fig. 3b). We analyzed the transcriptomes of 71,032 single cells using original cell cluster annotations to explore protein expression at the cellular level. In total, 184 significant associations (FDR <0.05) were identified across 112 proteins, with the majority of proteins expressed in stromal fibroblasts (n = 37), unciliated (n = 36) and ciliated epithelia (n = 23; Fig. 3c). Among proteins with the strongest and most cell-type-specific expression were PAEP in unciliated epithelial cells, a progesterone-regulated secretory protein essential for endometrial receptivity, and CXCL8 (interleukin-8) in epithelial and immune compartments. CXCL8 is a key chemokine that recruits neutrophils, promotes matrix remodeling and angiogenic signaling, and is known to rise around menstruation9 (Supplementary Table 6).
Pathway enrichment analyses
Pathway enrichment analyses were performed to investigate the biological relevance of proteins associated with the menstrual cycle and its temporal patterns. For the significantly associated proteins, some of the strongest enrichments were observed for ‘ovulation cycle’ and ‘female pregnancy’, reflecting the physiological specificity of the proteomic signals (Supplementary Table 7). Among the most significantly overrepresented pathways were ‘hormone activity’ and ‘cytokine activity’, reflecting the cyclic secretion of key hormones (for example, FSHB, CGA, PRL and OXT) and inflammatory mediators (for example, IL17, CXCL13 and CXCL8) that coordinate ovulation, endometrial regeneration and immune modulation (Supplementary Table 7). Cluster-specific analyses revealed distinct temporal patterns. Cluster 1 proteins, which peaked at the onset of menstruation, were enriched in ‘protease inhibitor activity’, suggesting roles in the controlled degradation and remodeling of endometrial tissue during menses (Supplementary Table 7). Proteins in cluster 2, elevated during the follicular phase, were enriched in ‘cytokine activity’, ‘hormone activity’ and ‘insulin-like growth factor binding’, consistent with immune-regulated tissue regeneration and endocrine signals that prepare the endometrium and ovary for ovulation (Supplementary Table 7). Cluster 3, containing proteins peaking in the late follicular and periovulatory window, was enriched in ‘hormone activity’ and reproductive growth pathways, including GH1, PRL, OXT and NPPC, reflecting the endocrine surge that coordinates ovulation and reproductive tract readiness (Supplementary Table 7). Finally, cluster 4 proteins, elevated during the luteal phase, showed enrichment in ‘cytokine signaling’, ‘pattern recognition receptor activity’ and MAPK pathway regulation, consistent with immune surveillance, stromal remodeling and progesterone-driven endometrial maturation following ovulation (Supplementary Table 7).
Associations with female reproductive traits
The broader relevance of the identified cycle-associated proteins was assessed in the full UKB cohort with available Olink proteomic data (23,674 women) by examining their associations with reproductive diseases (Fig. 4a and Supplementary Table 8). We identified 60 significant protein-prevalent disease associations (FDR < 0.05), involving 44 unique proteins across 12 conditions (Fig. 4b and Supplementary Table 9). The strongest association was observed between CGA and a history of excessive, frequent and irregular bleeding (odds ratio (OR) 1.35 per standard deviation (s.d.) increase in protein levels; 95% confidence interval (CI) 1.27–1.45; P = 1.4 × 10−16). The largest effect estimates were seen for ITIH4 (OR 1.50 per s.d. increase; 95% CI 1.30–2.65; P = 6.2 × 10−4) and for IGDCC4 (OR 0.51 per s.d. increase; 95% CI 0.36–0.73; P = 2.6 × 10−4) in relation to ovarian cancer. For incident disease outcomes, we identified 89 significant protein–disease associations involving cycle day proteins (FDR < 0.05), comprising 50 unique proteins across eight conditions (Fig. 4c and Supplementary Table 10). The strongest association was observed between CHRDL2 and incident leiomyoma (OR 2.05; 95% CI 1.87–2.25; P = 4.5 × 10−52). The largest effect estimate was seen for PAEP and excessive, frequent and irregular bleeding (OR 2.44; 95% CI 2.09–2.86; P = 7.6 × 10−29).
MR
We next investigated whether proteins associated with prevalent and/or incident female reproductive diseases play a causal role in disease development or are merely consequences of the disease. To address this, MR analyses were performed on significant protein–disease associations using protein quantitative trait locus (pQTL) data from the UKB10 and genome-wide association study (GWAS) summary statistics for relevant diseases from a recent genome-wide meta-analysis of 42 female reproductive diseases11. In cis-MR analyses, where cis-pQTLs served as exposures and disease GWAS data as outcomes, potential causal relationships were identified for 14 protein–disease pairs (Supplementary Table 11). Of these, 13 associations were robust across all sensitivity analyses (directionally consistent across MR models and without evidence of horizontal pleiotropy), and 12 reached at least nominal significance (P < 0.05) when applying an alternative linkage disequilibrium (LD) threshold (r2 = 0.01; Supplementary Table 12). Examples included associations between higher SEZ6L2 and risk of endometriosis, higher FSHB and risk of endometriosis, excessive menstrual bleeding, leiomyoma, female genital polyps, and PAEP and risk of excessive menstrual bleeding. Colocalization analyses provided further support to associations between FSHB and endometriosis (PP4 = 0.999), excessive menstrual bleeding (PP4 = 0.993), and leiomyoma (PP4 = 0.982), indicating a shared genetic signal underlying both protein levels and disease risk. To evaluate potential reverse causality, we conducted reverse MR analyses in which disease traits were used as exposures and protein levels as outcomes. None of the tested outcomes showed significant associations in the reverse direction (Supplementary Table 13).
Proteomic prediction of the menstrual cycle
The UKB cohort was randomly split into a 70% training set (n = 1,931) and a 30% validation set (n = 829) to develop a proteomic feature capturing the menstrual cycle (cycleDayProtS). Using this approach, the model identified 75 proteins with non-zero coefficients (Fig. 5a and Supplementary Table 14). The resulting cycleDayProtS was strongly correlated with cycle day in the holdout validation set (r = 0.56, P = 2.3 × 10−85; Fig. 5b), and its trajectory increased steadily across the follicular phase, reaching a peak in the mid-late luteal phase (Fig. 5c). To benchmark how well cycleDayProtS explained menstrual cycle timing, we compared its performance with serum estradiol. As expected, estradiol levels showed a characteristic biphasic pattern, with peaks in the late follicular phase and mid-luteal phase (Fig. 5d). Estradiol correlated weakly with cycleDayProtS (r = 0.07, P = 0.11) and explained negligible variance in cycle day (R2 = 0.002). By contrast, cycleDayProtS explained substantially more of the variability in cycle day (R2 = 0.418). Including both estradiol and cycleDayProtS in the same model did not improve prediction (combined R2 = 0.415), indicating that a random measurement of estradiol contributes little additional information beyond the proteomic signal.
Discussion
By profiling over 2,900 circulating proteins in 2,760 women, we mapped the circulating proteome across the menstrual cycle. Overall, 198, including classical reproductive hormones, cytokines, growth factors, extracellular-matrix enzymes and tissue-specific secreted factors, showed significant variation with cycle day. Using unsupervised clustering, four distinct temporal trajectories that aligned with key physiological phases of the menstrual cycle were identified. Most cycle-linked proteins were expressed in endometrial tissue and cell types, highlighting the tissue-specific nature of these molecular signals. Beyond physiological relevance, many of these proteins associated with common reproductive disorders, and cis-MR analyses supported potential causal roles for a subset.
Previous proteomic studies of the menstrual cycle have focused mainly on specific tissues or secretions, including endometrial tissue, uterine fluid, menstrual blood and cervical mucus, and have typically addressed specific aspects such as endometrial receptivity, folliculogenesis or endometriosis in relatively small cohorts12,13,14,15,16. Our study provides a large-scale, untargeted plasma proteomic map of the menstrual cycle. The use of plasma enables systematic evaluation of proteins that integrate signals from multiple tissues, thereby extending relevance beyond local uterine biology.
Our findings confirm and expand prior knowledge of menstrual biology. Established regulators such as FSHB, the β-subunit of FSH, and CGA, the common α-subunit shared by FSH, hCG and LH, showed expected cyclical variation. We also identified several other less-characterized proteins, which also demonstrated tightly timed expression: MMP-10, LEFTY2 and PAEP peaked around day 1 of menstruation, while PROK1, CXCL13, RLN1, RLN2 and AREG peaked in the mid-to-late luteal phase, suggesting novel roles in late-cycle endometrial remodeling and immune modulation. Beyond its established roles in sexual arousal and lactation17, oxytocin also showed marked cyclical variation, peaking in the late follicular and periovulatory phase, possibly to enhance sexual receptivity and promote reproductive tract motility around ovulation. Another interesting finding was that renin, the rate-limiting enzyme in the renin–angiotensin–aldosterone system, was among the top proteins associated with cycle day, showing a peak in the late luteal phase. Previous studies have consistently demonstrated higher levels of renin, renin activity and aldosterone during the luteal compared with the follicular phase, probably reflecting the influence of estradiol and progesterone18,19. Beyond its systemic role in fluid and blood pressure regulation, renin and other renin–angiotensin–aldosterone system components are also produced locally in the endometrium, where their cyclic variation may contribute to regulation of endometrial blood flow and initiation of menstruation20. Many of the dynamically regulated proteins were significantly enriched in endometrial tissue and cell types, particularly within epithelial and stromal compartments, reinforcing the biological specificity and relevance of the observed protein patterns.
Understanding the downstream relevance of cycle-associated proteins is important for understanding their role in women’s health. We found that many proteins were associated with multiple different conditions. Although these disorders appear clinically distinct, they share underlying biological processes, including inflammatory, angiogenic and steroid-responsive pathways, consistent with the widespread pleiotropy observed in female reproductive traits11. To explore the potential causal effects of these relationships, we employed MR and identified evidence of causality in 12 unique protein–disease pairs. For example, higher levels of FSHB were causally linked to increased risk of endometriosis, abnormal uterine bleeding and uterine leiomyoma. Under normal physiology, FSH rises during the early follicular phase, stimulating follicle recruitment and estradiol production. Although primarily viewed as an ovarian hormone, FSH receptors have also been detected in endometrial epithelial and stromal cells21, suggesting that the endometrium may respond directly to FSH signaling. This raises the possibility that cycle-dependent fluctuations in FSHB could influence endometrial remodeling, proliferation or receptivity, particularly during the proliferative phase when FSH levels peak. Endometriosis is strongly estrogen dependent, and higher FSHB may promote greater estrogenic drive, thereby supporting the survival and proliferation of ectopic endometrial tissue. In addition, aberrant or amplified FSH signaling could modify stromal–epithelial interactions, inflammation or decidualization directly, processes known to be dysregulated in endometriosis22. Together, these observations suggest that FSHB may not simply reflect altered ovarian physiology but may also exert direct endometrial effects relevant to disease mechanisms, reinforcing the broader concept that menstrual cycle biology can contribute to gynecological pathogenesis rather than merely mirroring it.
We developed a menstrual cycle proteomic score that captures the coordinated molecular changes across the menstrual cycle. The protein score showed a strong association with cycle day and a characteristic rise through the follicular phase, peaking in the mid-late luteal phase. Compared with serum estradiol, the protein score explained a larger proportion of variance in menstrual cycle timing, indicating that the multiprotein signature provides a more robust and informative molecular readout of cycle progression from a single blood sample.
Importantly, although many proteins tracked with cycle day, its overall contribution to proteomic variation was modest. Cycle effects were concentrated to a small subset with clear phase-specific dynamics, whereas age and BMI affected a broader set of proteins across the proteome, as has been reported previously23. From a biomarker-discovery perspective, this suggests that phase-matching may be most informative for proteins with known cycle sensitivity, particularly those related to ovarian, pituitary or endometrial biology.
Limitations
merit discussion. First, despite that several previous publications have documented the validity of self-reported menstrual cycle characteristics24,25, some degree of misclassification is expected, such as recall error and cycle-to-cycle variability26,27. However, such nondifferential misclassification would bias our associations towards the null. Second, the median age of participants was 45 years, raising the possibility of perimenopausal changes that could influence proteomic profiles. In addition, recent work has demonstrated nonlinear shifts in circulating proteomic signatures in midlife28. Although we excluded women reporting menopause and those with irregular and abnormally long cycles, and observed high concordance of cycle-associated effect estimates across age strata, some age-related proteomic variation within this age range could partially overlap with cycle-associated signatures and should be considered in interpretation. Third, because only a single sample was available per participant, we were unable to characterize within-person fluctuations in protein levels. Longitudinal studies are needed to map true intraindividual trajectories. Fourth, the Olink panel, while broad, is biased towards secreted and signaling proteins. Fifth, participants were predominantly of European ancestry and generally healthy. Findings may not generalize to more diverse or disease-enriched cohorts. Sixth, the UKB biochemical panel did not include key reproductive hormones such as FSH, LH or progesterone. These markers show characteristic phase-dependent patterns and would probably provide stronger benchmarks for menstrual cycle timing than estradiol alone. Sixth, although cis-MR suggested potential causal relationships for several protein–disease pairs, colocalization analyses supported a shared causal variant for only a subset of these associations. This discrepancy may reflect limited statistical power of disease GWASs at individual loci, the presence of multiple causal variants within regions, or LD between distinct causal signals29. Seventh, although the disease GWASs used as outcomes in the MR analyses represent the largest datasets currently available, reverse MR analyses may still have been underpowered to detect evidence to support reverse causation. Finally, in the absence of an independent replication dataset, the protein–disease associations identified here should be viewed as hypothesis-generating and warrant confirmation in external cohorts.
The menstrual cycle shapes the female proteome through coordinated, phase-specific changes across hundreds of circulating proteins. These rhythms reflect systemic endocrine and immune activity and may contribute to female reproductive disease risk. Our publicly available atlas provides important insights into menstrual timing in biomarker studies and precision health research.
Methods
Study population
The UKB is a large prospective cohort study comprising over 500,000 participants aged 40–69 years at recruitment (2006–2010) from across the UK. Extensive phenotyping was performed, including questionnaires, physical assessments, imaging, whole-genome sequencing and linkage to longitudinal electronic health records30. As part of the UK Biobank Plasma Proteomics Project (UKB-PPP)10, approximately 54,000 participants underwent proteomic profiling in plasma samples, an initiative that was funded by a consortium of 13 biopharmaceutical companies. Proteomic profiling procedures, including sample selection and handling, have been described in detail elsewhere10. In the present study, we analyzed data from female participants who had completed a questionnaire on reproductive factors. We excluded individuals with missing data on key covariates (age, sex, BMI, smoking status and self-reported ancestry) or those who failed proteomic quality control, as detailed below. The UKB received ethical approval from the North West Multi-Centre Research Ethics Committee as a research tissue bank (reference 11/NW/0382), and all participants provided written informed consent. All analyses were conducted under UKB application no. 43247.
Phenotype derivation of menstrual cycle day
We used questionnaire data to derive menstrual cycle day. To calculate menstrual cycle day, we used responses to two questions: “How many days is your usual menstrual cycle?” (data field 3710), referring to the typical number of days between menstrual periods, and “How many days since your last menstrual period?” (data field 3700), defined as the number of days since the first day of the most recent period. These questions were only answered by women who indicated they were still menstruating and had not yet reached menopause, based on their response to the question “Have you had your menopause (periods stopped)?” (data field 2724). Cycle day was calculated as the time since the last menstrual period minus the reported cycle length, plus one, such that day 1 corresponds to the first day of menstruation. We restricted this analysis to women under the age of 55 years who reported a regular menstrual cycle length between 21 days and 35 days and were not currently using oral contraceptives (data field 2804).
Proteomic profiling
Untargeted proteomic profiling was performed using the antibody-based Olink Explore 3072 proximity extension assay, which measures 2,941 protein analytes and 2,923 unique proteins10. Detailed assay procedures have been described previously10. Participants with more than 20% missing protein values were excluded, as were proteins missing in more than 10% of participants. Remaining missing values were imputed using the k-nearest neighbor method (k = 10) implemented in the impute R package (version 1.72.3). Protein levels were rank-based inverse normalized and scaled to a mean of 0 and s.d. of 1. We further excluded outlier individuals with a first or second principal component score more than 5 s.d. from the mean, or with a median normalized protein expression (NPX) value more than 5 s.d. from the overall cohort median, as has been done previously31. We assessed the association between protein levels and menstrual cycle day using linear regression, adjusted for age, BMI, smoking status, self-reported ancestry, time between blood draw and proteomic measurement, and batch. An FDR <0.05 was used to denote statistical significance. To compare the proportion of variance explained by menstrual cycle, age and BMI, we fit three separate models per protein, each using a natural cubic spline for one predictor (cycle day, age or BMI) to capture potential nonlinear relationships. For each predictor, spline knots were fixed globally at the 10th, 50th and 90th percentiles, with boundary knots at the observed minimum and maximum, as biologically defined inflection points are not known for most analytes32. For every protein–predictor pair, we reported the unadjusted R2 as the fraction of variance explained by that predictor alone. We then summarized the distributions of R2 across proteins (medians and the proportion with R2 > 0.05) to benchmark the proteome-wide influence of cycle timing against age and BMI.
Clustering analysis
To characterize temporal patterns of circulating proteins across the menstrual cycle, we focused on proteins significantly associated with cycle day (FDR <0.05). Each participant’s derived cycle day was rescaled linearly to a 28-day cycle. For each protein, mean expression was calculated across normalized cycle days, z-score normalized and clustered on the basis of temporal expression profiles using k-means clustering, with the optimal number of clusters determined using the gap statistic. Distinct expression trajectories were visualized with curves smoothed using locally estimated scatterplot smoothing (LOESS), and representative proteins with peak expression in specific phases were identified for each cluster. For interpretative purposes, menstrual phases were overlaid according to established endometrial physiology: menses/regenerative (days 1–6), follicular/proliferative (days 7–13), periovulatory (days 14–15), early luteal/secretory (days 16–19), mid-luteal/secretory (days 20–23), and late luteal/secretory (days 24–28), to illustrate how the data-driven clusters align with known physiological transitions.
Pathway and tissue enrichment analysis
We performed Gene Ontology (GO) enrichment analysis separately for each of the four clusters, as well as for the full set of cycle-associated proteins and for clusters, using the clusterProfiler R package (version 4.17.0)33. Gene symbols were converted to Entrez Gene IDs, and enrichment was conducted across GO ontologies. Significance was determined using an FDR < 0.05. Tissue enrichment of associated proteins was conducted using the TissueEnrich R package (version 1.6.0)34. Genes encoding proteins included on the Olink panel were used as the background set. For tissue-specific enrichment, we used RNA expression data from the Human Protein Atlas35, considering all genes reported as expressed in each tissue. Statistical significance was defined as an FDR < 0.05.
Cell type enrichment analysis
For the proteins significantly associated with the menstrual cycle, we prioritized potentially causal cell types using previously published scRNA-seq data on 71,032 single cells from ten healthy human endometria and their original cell clusters, generated with the 10× Chromium platform by Wang et al.36. In that study, both a lower-throughput Fluidigm C1 dataset (19 donors) and a higher-throughput 10× Chromium dataset (ten donors) were produced. We specifically used the 10× dataset because its substantially larger number of cells provides greater power to detect enrichment in less abundant endometrial cell populations. Quality control was performed in multiple steps with Seurat (version 4.4.0)37. First, genes expressed in fewer than three cells and cells with fewer than 200 transcripts were removed. We then excluded cells with ≤200 or ≥8,000 detected genes and ≥20% mitochondrial reads. The gene expression matrix was normalized and scaled with the NormalizeData() and ScaleData() functions. The top 1,000 variable genes were selected by FindVariableFeatures() for PCA, and batch effects were corrected with Harmony (version 1.0) using the first 20 principal components. Cell identities were taken from the original publication. The two unciliated epithelial subclusters were merged into a single ‘Unciliated epithelia’ category for all downstream analyses. Within each cell type, we investigated differential expression of significantly associated proteins using the FindMarkers() option with the following parameters: only.pos = TRUE, min.pct = 0.1, logfc.threshold = 0.25, test.use = “Wilcoxon” and assay = “RNA”.
Phenome-wide association study
To investigate whether proteins that vary physiologically across the menstrual cycle are also related to reproductive disorders, we conducted a phenome-wide association study (PheWAS) in the full UKB cohort with available Olink proteomic measurements (N = 23,674 women). This allowed us to first identify proteins that vary under normal physiological conditions and subsequently assess whether deviations from these physiological patterns are linked to reproductive pathology. We examined 42 female reproductive health diagnoses11, defined using the International Classification of Diseases, Tenth Revision codes (detailed in Supplementary Table 7), and included only outcomes with at least 20 cases. Associations between protein levels and both prevalent (that is, diagnoses before enrollment) and incident (that is, diagnoses occurring after enrollment) diseases were tested using logistic regression models, adjusted for age, BMI, smoking status, self-reported ancestry, time between blood draw and proteomic measurement, and batch. Statistical significance was defined as an FDR <0.05.
MR
We performed two-sample MR analyses using circulating protein levels as exposures and female reproductive health diagnoses as outcomes. cis-pQTLs from the UKB-PPP10 were used as instrumental variables, defined as genetic variants located within 1 Mb of the transcription start site of the corresponding protein-coding gene. For each protein, we selected genome-wide significant variants (P < 5 × 10−8) and performed LD clumping at r2 < 0.1 to maximize variance explained while retaining largely independent cis signals, consistent with prior proteome-wide MR studies. Variants were clumped using PLINK (version 1.9)38 and an LD reference panel derived from 10,000 randomly selected European-ancestry participants from the UKB. GWAS summary statistics for outcomes were obtained from a recent genome-wide meta-analysis of female reproductive traits11. Instrumental variants absent from the outcome GWAS were replaced with proxy variants in high LD (r2 ≥ 0.8), identified using PLINK and 1000 Genomes Project EUR panel39. For each missing variant, the proxy with the highest r2 within a ±1-Mb window was selected. To minimize weak instrument bias, we restricted analyses to instruments with F-statistics >10 (ref. 40). Causal effects were estimated using the random-effects inverse-variance weighted (IVW) method or the Wald ratio when only one single-nucleotide polymorphism was available, as implemented in the TwoSampleMR R package (version 0.6.17)41. Statistical significance was defined as an FDR <0.05 for IVW or Wald ratio estimates. We conducted multiple sensitivity analyses to assess robustness. First, we performed weighted median and MR-Egger analyses and retained only associations directionally concordant with the primary analysis. Second, we assessed heterogeneity using Cochran’s Q statistic and excluded associations showing evidence of heterogeneity (P < 0.05). Third, directional horizontal pleiotropy was evaluated using the MR-Egger intercept test for all IVW-based analyses, and associations with evidence of pleiotropy (intercept P < 0.05) were excluded. Fourth, associations passing these filters were reanalyzed using a more stringent LD threshold (r2 < 0.01). Fifth, to further support shared causal variants and mitigate confounding due to LD, we performed colocalization analyses using the coloc R package. We tested whether cis-pQTLs and female reproductive trait associations shared the same causal variant within a ±500-kb region using default priors. A posterior probability for colocalization (PP4) >0.8 was considered strong evidence of a shared causal signal. Finally, reverse MR analyses were conducted to evaluate potential reverse causality by testing whether female reproductive health diagnoses influenced circulating protein levels, using the same instruments, filtering criteria and analytical framework.
Proteomic score
To derive a menstrual cycle protein score, we randomly split the main cohort (2,760 women) into a training set (70%) and an independent test set (30%). For both splits, missing protein values were separately imputed using the k-nearest neighbor method (k = 10) and then rank-based inverse normalized, to avoid information leakage. In the training set, we fit a linear regression model, including 2,917 proteins that passed quality control, and penalization by the least absolute shrinkage and selection operator (LASSO) implemented in the glmnet R package (verision 4.1.8). Tenfold cross-validation was used to select the optimal regularization parameter (λ) on the basis of the mean squared error. Covariates included age, BMI, smoking status, ethnicity and batch. The resulting protein score was calculated in the test set as a weighted sum of protein values using non-zero coefficients from the LASSO model. To benchmark the predictive ability of the proteomic score against established hormonal biomarkers, we extracted serum estradiol concentrations (data field 30800). We then fit linear models predicting cycle day using (1) estradiol alone, (2) the proteomic score alone and (3) both predictors jointly. We modeled both exposures using natural cubic splines with knots placed at the 10th, 50th and 90th percentiles of its distribution. Model performance was compared using R2 in the held-out test set.
Statistics and reproducibility
No statistical method was used to predetermine sample size. Sample size was determined by the number of UKB participants with available plasma proteomic measurements and complete information on menstrual cycle characteristics and covariates. Data were excluded only according to predefined quality control criteria described in the Methods. In brief, participants with missing key covariates or failing proteomic quality control were removed, as were individuals with excessive missing protein measurements. Proteins with high missingness were excluded, and remaining missing values were imputed using the k-nearest neighbor method. The experiments were not randomized. The investigators were not blinded to allocation during experiments and outcome assessment. Proteomic measurements were generated as part of the UKB-PPP using standardized protocols and internal quality control procedures independent of the present analyses. All statistical analyses were performed using R. Associations between circulating proteins and menstrual cycle day were assessed using linear regression models with adjustment for relevant covariates. Multiple testing was controlled using the Benjamini–Hochberg FDR procedure, with FDR <0.05 considered statistically significant. Additional analyses, including clustering, enrichment analyses, phenome-wide association analyses, MR and derivation of the proteomic score using LASSO regression with cross-validation, are described in detail above.
Reporting summary
Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.
Data availability
Individual-level proteomic, phenotypic and electronic health record data are available from UKB upon application (https://www.ukbiobank.ac.uk/). Summary-level data are available in Supplementary Tables 1–14. The female reproductive disease GWAS summary statistics are publicly accessible in the GWAS Catalog. Accession numbers were provided by Pujol Gualdo et al.11 and can also be found in Supplementary Table 8. Data from the Human Protein Atlas are publicly available (https://www.proteinatlas.org/) and embedded in the tissueEnrich R package35. Endometrial scRNA-seq data can be found at NCBI’s Gene Expression Omnibus (series accession code GSE111976).
Code availability
The code used to conduct these analyses is available via GitHub at https://github.com/ghousen/menstrualCycle. Settings for individual software/packages are described in the Methods. The following software and packages were used for data analysis: impute version 1.72.3 (https://www.bioconductor.org/packages/release/bioc/html/impute.html), stats R package for k-means clustering, PLINK 2.0 (https://www.cog-genomics.org/plink/2.0/), TwoSampleMR version 0.5.6 (https://mrcieu.github.io/TwoSampleMR/), clusterProfiler for GO enrichment (https://bioconductor.org/packages/release/bioc/html/clusterProfiler.html), TissueEnrich R package for tissue-specific enrichment, Seurat version 4.4.0 (https://github.com/satijalab/seurat), glmnet version 4.1.8 for LASSO regression and R version 4.1.2 (https://www.r-project.org/).
References
Rosen Vollmar, A. K., Mahalingaiah, S. & Jukic, A. M. The menstrual cycle as a vital sign: a comprehensive review. F. S. Rev. 6, 100081 (2025).
Notbohm, H. L. et al. The effects of menstrual cycle phases on immune function and inflammation at rest and after acute exercise: a systematic review and meta-analysis. Acta Physiol. 238, e14013 (2023).
Draper, C. F. et al. Menstrual cycle rhythmicity: metabolic patterns in healthy women. Sci. Rep. 8, 14568 (2018).
Benton, M. J., Hutchins, A. M. & Dawes, J. J. Effect of menstrual cycle on resting metabolism: a systematic review and meta-analysis. PLoS ONE 15, e0236025 (2020).
Wang, Y.-X. et al. Menstrual cycle regularity and length across the reproductive lifespan and risk of cardiovascular disease. JAMA Netw. Open 5, e2238513 (2022).
Crona Guterstam, Y. et al. The cytokine profile of menstrual blood. Acta Obstet. Gynecol. Scand. 100, 339–346 (2021).
Hood, B. L. et al. Proteomics of the human endometrial glandular epithelium and stroma from the proliferative and secretory phases of the menstrual cycle. Biol. Reprod. 92, 106 (2015).
Yang, H., Zhou, B., Prinz, M. & Siegel, D. Proteomic analysis of menstrual blood. Mol. Cell. Proteom. 11, 1024–1035 (2012).
Osteen, K. G., Keller, N. R., Feltus, F. A. & Melner, M. H. Paracrine regulation of matrix metalloproteinase expression in the normal human endometrium. Gynecol. Obstet. Invest. 48, 2–13 (1999).
Sun, B. B. et al. Plasma proteomic associations with genetics and health in the UK Biobank. Nature 622, 329–338 (2023).
Pujol Gualdo, N. et al. Atlas of genetic and phenotypic associations across 42 female reproductive health diagnoses. Nat. Med. 31, 1626–1634 (2025).
Apostolov, A. et al. Multi-omics analysis of uterine fluid extracellular vesicles reveals a resemblance with endometrial tissue across the menstrual cycle: biological and translational insights. Hum. Reprod. Open 2025, hoaf010 (2025).
Ji, S. et al. DIA-based analysis of the menstrual blood proteome identifies association between CXCL5 and IL1RN and endometriosis. J. Proteom. 289, 104995 (2023).
Ye, L. & Dimitriadis, E. Endometrial receptivity—lessons from ‘omics’. Biomolecules 15, 106 (2025).
Penariol, L. B. C. et al. What do the transcriptome and proteome of menstrual blood-derived mesenchymal stem cells tell us about endometriosis? Int. J. Mol. Sci. 23, 11515 (2022).
Alikhani, M. et al. Proteome analysis of endometrial tissue from patients with PCOS reveals proteins predicted to impact the disease. Mol. Biol. Rep. 47, 8763–8774 (2020).
Engel, S., Klusmann, H., Ditzen, B., Knaevelsrud, C. & Schumacher, S. Menstrual cycle-related fluctuations in oxytocin concentrations: a systematic review and meta-analysis. Front. Neuroendocrinol. 52, 144–155 (2019).
Chidambaram, M. et al. Variation in the renin angiotensin system throughout the normal menstrual cycle. J. Am. Soc. Nephrol. 13, 446–452 (2002).
Li, X. F. & Ahmed, A. Compartmentalization and cyclic variation of immunoreactivity of renin and angiotensin converting enzyme in human endometrium throughout the menstrual cycle. Hum. Reprod. 12, 2804–2809 (1997).
Johnson, I. R. Renin substrate, active and acid-activatable renin concentrations in human plasma and endometrium during the normal menstrual cycle. BJOG 87, 875–882 (1980).
La Marca, A., Carducci Artenisio, A., Stabile, G., Rivasi, F. & Volpe, A. Evidence for cycle-dependent expression of follicle-stimulating hormone receptor in human endometrium. Gynecol. Endocrinol. 21, 303–306 (2005).
Zondervan, K. T. et al. Endometriosis. Nat. Rev. Dis. Primers 4, 9 (2018).
Carrasco-Zanini, J. et al. Mapping biological influences on the human plasma proteome beyond the genome. Nat. Metab. 6, 2010–2023 (2024).
Jukic, A. M. Z. et al. Accuracy of reporting of menstrual cycle length. Am. J. Epidemiol. 167, 25–33 (2008).
Must, A. et al. Recall of early menstrual history and menarcheal body size: after 30 years, how well do women remember? Am. J. Epidemiol. 155, 672–679 (2002).
Soumpasis, I., Grace, B. & Johnson, S. Real-life insights on menstrual cycles and ovulation using big data. Hum. Reprod. Open 2020, hoaa011 (2020).
Bull, J. R. et al. Real-world menstrual cycle characteristics of more than 600,000 menstrual cycles. npj Digit. Med. 2, 83 (2019).
Shen, X. et al. Nonlinear dynamics of multi-omics profiles during human aging. Nat. Aging 4, 1619–1634 (2024).
Zuber, V. et al. Combining evidence from Mendelian randomization and colocalization: review and comparison of approaches. Am. J. Hum. Genet. 109, 767–782 (2022).
Bycroft, C. et al. The UK Biobank resource with deep phenotyping and genomic data. Nature 562, 203–209 (2018).
Carrasco-Zanini, J. et al. Proteomic signatures improve risk prediction for common and rare diseases. Nat. Med. 30, 2489–2498 (2024).
Austin, P. C., Fang, J. & Lee, D. S. Using fractional polynomials and restricted cubic splines to model non-proportional hazards or time-varying covariate effects in the Cox regression model. Stat. Med. 41, 612–624 (2022).
Yu, G., Wang, L.-G., Han, Y. & He, Q.-Y. clusterProfiler: an R package for comparing biological themes among gene clusters. OMICS 16, 284–287 (2012).
Jain, A. & Tuteja, G. TissueEnrich: tissue-specific gene enrichment analysis. Bioinformatics 35, 1966–1967 (2019).
Uhlén, M. et al. Tissue-based map of the human proteome. Science 347, 1260419 (2015).
Wang, W. et al. Single-cell transcriptomic atlas of the human endometrium during the menstrual cycle. Nat. Med. 26, 1644–1653 (2020).
Satija, R., Farrell, J. A., Gennert, D., Schier, A. F. & Regev, A. Spatial reconstruction of single-cell gene expression data. Nat. Biotechnol. 33, 495–502 (2015).
Chang, C. C. et al. Second-generation PLINK: rising to the challenge of larger and richer datasets. Gigascience 4, 7 (2015).
Auton, A. et al. A global reference for human genetic variation. Nature 526, 68–74 (2015).
Sanderson, E., Spiller, W. & Bowden, J. Testing and correcting for weak and pleiotropic instruments in two-sample multivariable Mendelian randomization. Stat. Med. 40, 5434–5452 (2021).
Hemani, G. et al. The MR-Base platform supports systematic causal inference across the human phenome. eLife 7, e34408 (2018).
Acknowledgments
This work was supported by AUFF Recruitment grant (grant no. AUFF-E-2024-7-10 to J.G.). The funders had no role in study design, data collection and analysis, decision to publish or preparation of the manuscript.
Author information
Authors and Affiliations
Contributions
I.R. and J.G. conceived these analyses. I.R., S.A.R. and J.G. performed formal analyses. I.R., L.R., P.R.L., L.D.A., O.H.L., C.K.E., A.P., S.S. and J.G. provided resources. I.R., S.A.R., S.S. and J.G. performed data curation. I.R. and J.G. drafted the manuscript. I.R., S.A.R., S.S. and J.G. performed data visualization. L.R., P.R.L., L.D., O.H.L., C.K.E., A.P., S.S. and J.G. supervised the study. All authors contributed to the critical review and revision of the paper.
Corresponding authors
Ethics declarations
Competing interests
J.G. has received lecture fees from Illumina and is a former employee of Novo Nordisk A/S. A.P. has received independent research grants and lecture fees from Abbott, IBSA, Gedeon Richter, Ferring and Merck A/S. A.P. is part of research advisory boards for Gedeon Richter and Ferring. All other authors declare no competing interests.
Peer review
Peer review information
Nature Medicine thanks Lauren Houghton, Peter Rogers, Sara Stinson and the other, anonymous, reviewer(s) for their contribution to the peer review of this work. Primary Handling Editor: Ashley Castellanos-Jankiewicz, in collaboration with the Nature Medicine team.
Additional information
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Supplementary information
Supplementary Tables 1–14 (download XLSX )
Supplementary Tables 1–14.
Rights and permissions
Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
About this article
Cite this article
Riishede, I., Rode, L., Lundegaard, P.R. et al. Plasma proteomic signature of the human menstrual cycle. Nat Med (2026). https://doi.org/10.1038/s41591-026-04326-5
Received:
Accepted:
Published:
Version of record:
DOI: https://doi.org/10.1038/s41591-026-04326-5
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.