Intro
Primary dysmenorrhea (PD) refers to recurrent, cramping lower abdominal pain occurring around menstruation in the absence of identifiable pelvic pathology, and typically develops within the first year after menarche once ovulatory cycles are established. 1 In addition to pelvic pain, many affected individuals experience systemic symptoms such as nausea, vomiting, diarrhea, headache, fatigue, and, in some cases, syncope, indicating that PD extends beyond a localized gynecological disorder. 1 PD is highly prevalent worldwide and represents one of the most common pain conditions among women of reproductive age. A recent large-scale systematic review including data from more than 70 countries reported a pooled prevalence of 71.3% (95% CI 68.7–73.8%), with the highest burden observed in adolescents and young adults. 2 Despite its frequency, the broader impact of PD is often underestimated; the condition is associated with substantial reductions in quality of life, frequent school and work absenteeism, and impaired productivity, resulting in a socioeconomic burden comparable to that of other chronic pain conditions. 3–5
The mechanisms underlying primary dysmenorrhea are complex and involve hormonal, inflammatory, vascular, and neural processes. The most widely accepted explanation is the prostaglandin hypothesis, which links menstrual pain to excessive production of prostaglandin F 2 α (PGF 2 α) in the endometrium following progesterone withdrawal. Elevated PGF 2 α promotes strong uterine contractions and reduces uterine blood flow, resulting in ischemic pain. 6–9 Other mediators, such as vasopressin and oxytocin, further intensify vasoconstriction and uncoordinated uterine contractions, contributing to symptom severity. 10 , 11 Emerging evidence also points to an inflammatory and neuroimmune contribution, with increased pro-inflammatory cytokines, immune imbalance, and reduced endogenous pain inhibition, which may lower pain thresholds and increase vulnerability to long-term pain sensitization. 4 , 12 Together, these overlapping pathways help explain why treatments targeting a single mechanism often provide incomplete relief.
Current international guidelines recommend nonsteroidal anti-inflammatory drugs (NSAIDs) as first-line therapy and hormonal contraceptives as second-line treatment for PD. While NSAIDs effectively reduce pain by inhibiting cyclooxygenase-mediated prostaglandin synthesis, up to 15–25% of patients exhibit inadequate analgesic response, and long-term use is constrained by gastrointestinal, renal, and cardiovascular adverse effects. 13–15 Hormonal contraceptives, although effective, are unsuitable for many women due to contraindications, fertility considerations, thromboembolic risk, and poor tolerability. 7 , 14 Consequently, a substantial therapeutic gap persists for patients who are refractory to NSAIDs, unwilling or unable to use hormonal therapy, or seeking treatments with better long-term safety profiles.
Guizhi Fuling Capsules (GZFL) are a standardized pharmaceutical formulation derived from the classical prescription Guizhi Fuling Wan. The active formula comprises five herbal materials: Cinnamomi Ramulus (Guizhi), Poria (Fuling), Moutan Cortex (Mudanpi), Persicae Semen (Taoren), and Paeoniae Radix Alba (Baishao), each at 240 g per 1000 capsules in the 2020 Chinese Pharmacopoeia monograph. The finished product is not exclusively herbal: the pharmacopoeial manufacturing process adds dextrin, and manufacturer technical documents for the Jiangsu Kanion product also describe beta-cyclodextrin in the production pathway. Batch-specific complete excipient lists were not consistently reported in the included trial publications, so these standard/manufacturer-reported excipients should not be interpreted as verified for every trial batch. The pharmacopoeial quality standard includes microscopic and thin-layer chromatographic identification, gas-chromatographic identification of cinnamaldehyde, an HPLC fingerprint with similarity of at least 0.85, and per-capsule content limits for paeonol (at least 1.8 mg), paeoniflorin (at least 3.0 mg), and amygdalin (at least 0.90 mg). Published quality-control work additionally describes quantitative fingerprinting and monitoring of representative compounds. 10 , 16–18
Current management includes pharmacological and non-pharmacological approaches; a recent randomized trial of myofascial release and Kinesio Taping further illustrates continuing interest in short-term pain-relief strategies for primary dysmenorrhoea. 19 A GZFL trial protocol and a later randomized report likewise reflect contemporary research activity. 20 , 21 Earlier reviews of Chinese herbal medicine or Guizhi Fuling formulations included mixed dysmenorrhoea populations, heterogeneous formulations and comparators, or predominantly older trials. 22–24 The present review therefore focuses on randomized trials evaluating GZFL-based regimens in primary dysmenorrhoea and incorporates contemporary trials. Its objectives were to estimate effects on pain-related outcomes and overall response, evaluate adverse events and inflammatory biomarkers, explore heterogeneity by treatment strategy, comparator, and dose, and assess certainty using GRADE.
Methods
This meta-analysis was conducted in strict accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines 25 ( Appendix 1 ). The study protocol was prospectively registered in the International Prospective Register of Systematic Reviews (PROSPERO) under registration number CRD420251271957.
A comprehensive and systematic literature search was conducted in eight electronic databases, including four English-language databases (PubMed, Embase, Web of Science, and the Cochrane Central Register of Controlled Trials [CENTRAL]) and four Chinese-language databases (China National Knowledge Infrastructure [CNKI], Wanfang Data, VIP Database, and SinoMed). All databases were searched from their inception to 1 December 2025, without restrictions on publication status. The search strategy combined controlled vocabulary (such as Medical Subject Headings [MeSH] in PubMed and Emtree terms in Embase) with free-text terms to maximize sensitivity. Boolean operators (“AND”, “OR”) were used to combine search terms related to the population (primary dysmenorrhea), intervention (Guizhi Fuling Capsules and its variants), and study design (randomized controlled trials). Search terms were adapted to the syntax and indexing system of each database. The detailed search strategies for all databases are provided in the Appendix 2 .
P: Eligible studies were randomized controlled trials (parallel-group RCTs) enrolling women of reproductive age with a clinical diagnosis of primary dysmenorrhea, defined as menstrual pain in the absence of identifiable pelvic pathology (eg, endometriosis, adenomyosis, uterine fibroids, pelvic inflammatory disease, congenital uterine anomalies).
I: The intervention was an oral GZFL-based regimen, used either as monotherapy or as part of an add-on or multimodal treatment strategy. Eligible controls were placebo or conventional medicines used for primary dysmenorrhoea, including analgesic/NSAID and hormonal therapies. When co-interventions differed between groups, the comparison was retained as an estimate of the complete GZFL-based strategy rather than the isolated pharmacological effect of GZFL; treatment strategy was recorded and explored in subgroup analysis.
O: The primary outcome was pain relief rate. Secondary outcomes were VAS pain intensity, total effective rate, inflammatory biomarkers (TNF-alpha, IL-6, and IL-10), and adverse events. Pain relief rate was defined as the proportion meeting a clearly reported higher-threshold categorical pain-response criterion in each trial (for example, clinical cure, marked response, or at least 50% pain reduction). Because labels varied, categories were mapped at study level and the same threshold was applied to both groups within each trial. VAS was treated exclusively as a continuous pain-intensity measure. Total effective rate was defined as the proportion classified by trialists within any non-failure response category, usually cured, markedly effective, or effective, based on pain and/or composite symptom criteria. These outcomes were not treated as interchangeable, and the trial-specific instruments, thresholds, and categories are reported in Appendix 8 . The primary and secondary outcome hierarchy was prespecified in PROSPERO (CRD420251271957) and is reported here consistently with the registration.
We excluded duplicate reports (retaining the most complete or recent), non-randomized designs (eg, reviews, observational studies, case reports), trials in secondary dysmenorrhea or mixed populations without extractable PD data, and studies involving pregnancy or lactation.
A standardized data-extraction form was developed and pilot-tested. Two reviewers independently extracted study characteristics, participant data, interventions, comparators, follow-up, outcome definitions, measurement time points, and numerical results. We also extracted whether each report explicitly stated research ethics approval and participant informed consent. Absence of reporting was recorded as “not reported” and was not interpreted as evidence that approval or consent had not been obtained. Disagreements were resolved by discussion or third-reviewer adjudication. When multiple time points were available, end-of-treatment data were prioritized. For the three-arm Liu 2013 trial, the two GZFL dose groups were combined for the primary meta-analyses to avoid counting the shared placebo group twice; only the dose subgroup analysis retained separate active arms and split the shared placebo denominator and events.
Risk of bias was assessed independently by two reviewers using the Cochrane Risk of Bias tool (RoB 1.0) across the following domains: (1) random sequence generation (selection bias), (2) allocation concealment (selection bias), (3) blinding of participants and personnel (performance bias), (4) blinding of outcome assessment (detection bias), (5) incomplete outcome data (attrition bias), (6) selective reporting (reporting bias), and (7) other sources of bias (eg, baseline imbalance, inappropriate co-interventions, early stopping, or deviations from the intended protocol). Each domain was judged as low, high, or unclear risk of bias, with supporting justification recorded. Any disagreements were resolved by discussion, with arbitration by a third reviewer when necessary.
The certainty of evidence for each outcome was appraised using GRADE. Randomized evidence started at high certainty and could be downgraded for risk of bias, inconsistency, indirectness, imprecision, or publication bias. Outcome-specific reasons for every downgrade are reported in the Results and in footnotes to the evidence profile ( Appendices 6A and B ). Two reviewers made judgments independently, resolving disagreements by discussion or third-reviewer adjudication.
All analyses used aggregate data from the published reports and were recalculated in R 4.4.2 using the meta package (version 8.2–0). Dichotomous pain relief and total effective rate were expressed as risk ratios (RRs), adverse events as odds ratios (ORs), and continuous outcomes as mean differences (MDs), each with 95% confidence intervals (CIs). Statistical heterogeneity was assessed using the Chi-square test and I 2 statistic. A common-effect model was used when P > 0.10 and I 2 < 50%; otherwise a random-effects model was used. Leave-one-out sensitivity analyses assessed robustness and influential trials. Funnel plots were inspected when at least 10 studies were available. Prespecified subgroup analyses examined treatment strategy (monotherapy versus add-on/multimodal), comparator (placebo, analgesic/NSAID, or hormonal therapy), and daily dose (standard versus higher dose), with formal interaction tests. The two Liu 2013 active arms were combined in primary analyses; for the dose subgroup only, the shared placebo group was split approximately equally.
Results
A total of 294 records were identified through database searches. We removed 176 duplicate database records before screening, leaving 118 unique records. After title and abstract screening, 49 full-text reports were assessed and 35 were excluded. Of these, four were duplicate or overlapping full-text reports whose relationship could only be established after detailed full-text assessment; these differed from database record duplicates removed before screening. Fourteen randomized controlled trials were included ( Figure 1 ).
Figure 1 PRISMA flow diagram. Duplicate database records were removed before screening; duplicate or overlapping full-text reports were identified only after full-text assessment. The PRISMA flow diagram outlines the study selection process. It begins with the identification of 294 records through database searches, including PubMed, Embase, CENTRAL, Web of Science, CNKI, VIP, Wanfang and SinoMed/CBM. No additional records were identified through other sources. Next, 176 duplicate records were removed before screening, leaving 118 unique records. During the screening phase, these records were screened by title and abstract, resulting in 69 exclusions. The eligibility phase involved assessing 49 full-text reports, with 35 excluded for reasons such as not being randomized trials, non-primary dysmenorrhea, inappropriate comparator, insufficient outcome data, duplicate reports, or unavailable full text. Finally, 14 randomized controlled trials were included in the qualitative review and quantitative synthesis. A PRISMA flow diagram of study selection process with identification, screening, eligibility and inclusion steps.
PRISMA flow diagram. Duplicate database records were removed before screening; duplicate or overlapping full-text reports were identified only after full-text assessment.
Fourteen RCTs involving 1280 participants were included, and all were conducted in China. Eight trials evaluated GZFL as monotherapy and six evaluated add-on or multimodal GZFL-based strategies. Comparators included placebo, analgesic/NSAID therapy, and hormonal therapy. Treatment was short-term and was generally administered over three menstrual cycles. In trials with non-identical co-interventions, the effect estimate represents the combined treatment strategy and cannot isolate the contribution of GZFL. These differences in strategy and comparator are clinically important sources of heterogeneity ( Table 1 ).
Table 1 Baseline Characteristics of Included Studies Study Baseline Sverity (Pain) Age (T/C) Disease Duration, Year (T/C) Participants Intervention Control Treatment Duration (Cycles) Sun 2004 26 Symptom-score grading (14) 19.13±0.97/18.82±1.07 4.86±0.50/4.78±0.55 30/30 GZFL (0.93 g, tid) Placebo (tid) 3 Wang 2016 27 VAS ≥4 18–30 ≥0.5 15/15 GZFL (1.395 g, tid) Placebo (tid) 3 Sun 2010 28 Symptom-score grading (14) 30.5 (16–49) 0.5–15 40/40 GZFL + Con (0.93 g, tid) Tramadol (bid) 3 Wang 2021 29 VRS grading (II–III) 18.9±4.8/19.3±4.6 0.5–6/0.42–5 41/41 GZFL (6 g, bid) Desogestrel/Ethinylestradiol (qd) 3 Lu 2017 30 Pain severity (moderate/severe) 25.5±1.3/26.5±0.9 6.5±1.1/6.5±1.2 30/30 GZFL + Con (1.35 g, bid) Desogestrel/Ethinylestradiol (qd) 3 Zhong 2019 31 VAS reported 26.31±4.17/27.05±4.69 2.14±0.55/2.18±0.59 36/36 GZFL + Con (0.93 g, tid) Ibuprofen (0.25 g, bid) 3 Zhang 2020 32 Symptom-score severity strata 23.63±5.90/24.00±5.78 2.56±1.09/2.56±1.09 30/30 GZFL + Con (0.93 g, tid) Ibuprofen (0.25 g, bid) 3 Li 2022 33 VAS 5.67±1.50 / 5.60±1.36 28.72±8.45/27.61±7.47 6 ± 4.63/6 ± 3.15 54/54 GZFL (0.93 g, tid) Dydrogesterone (10 mg, bid) 3 Li 2014 34 NR 24.2±3.4/23.7±3.8 NR 75/75 GZFL + Con (0.93 g, tid) Medroxyprogesterone (4 mg, qd) 3 Liu 2013 35 VAS ≥4 24.29±2.15/24.49±2.20/23.89±2.73 5±7/5.21±7/5±6 74/76/71 GZFL (0.93 g, tid)/ (1.395 g, tid) Placebo (0.93 g, tid) 3 Mao 2023 36 VAS ≥4 21 (20,22)/20 (19,22) 5± 3.52/6 ± 1.48 64/64 GZFL (3 g, bid) Placebo (3 g, bid) 3 Cheng 2011 37 VAS ≥4 24.56±1.96/24.24±1.66 9.95±1.05/9.02±0.9 30/30 GZFL (0.93 g, tid) Placebo (0.93 g, tid) 3 Zhang 2012 38 VAS ≥4 21.30/22.70 4.64/4.75 39/20 GZFL (0.93 g, tid) Placebo (0.93 g, tid) 3 Huang 2022 39 CMSS severity score 12.48±1.26 / 12.46±1.24 32.53/32.49 1.17/1.16 55/55 GZFL (0.93 g, tid) Dydrogesterone (10 mg, bid) 3 Note : Treatment duration is expressed in menstrual cycles. Study references are shown as superscript Arabic numerals. Abbreviations : T/C, treatment/control; GZFL, Guizhi Fuling Capsules; Con, concomitant conventional medication; qd, once daily; bid, twice daily; tid, three times daily; NR, not reported.
Baseline Characteristics of Included Studies
Note : Treatment duration is expressed in menstrual cycles. Study references are shown as superscript Arabic numerals.
Abbreviations : T/C, treatment/control; GZFL, Guizhi Fuling Capsules; Con, concomitant conventional medication; qd, once daily; bid, twice daily; tid, three times daily; NR, not reported.
Research ethics approval was explicitly reported in 4 of 14 trial reports and informed consent in 8 of 14; the remaining reports did not clearly report these items. Ethics reporting was not a prespecified eligibility criterion, and studies were not excluded solely because approval information was absent from the publication ( Appendix 7 ).
Risk of bias was assessed using the Cochrane RoB 1 tool ( Figures 2 and 3 ). All 14 trials clearly described how the random sequence was generated and were judged to have low risk of bias for this domain. However, only a few studies reported adequate methods to conceal group assignment; most did not provide enough detail and were rated as unclear risk for allocation concealment. Blinding was reported inconsistently. Seven trials were double-blind and were judged at low risk of bias related to blinding, whereas one trial was open-label and was judged at high risk for bias in this area. The remaining trials did not report sufficient information about blinding and were therefore rated as unclear risk. Across all studies, risk of bias was low for incomplete outcome data, selective reporting, and other potential sources of bias, with minimal or balanced attrition and no clear signs of baseline imbalance or major deviations from the planned methods. Overall, risk of bias was generally low, with the main limitations arising from incomplete reporting of allocation concealment and blinding.
Figure 2 Risk-of-bias graph. Green indicates low risk, yellow indicates unclear risk, and red indicates high risk. Included studies: Sun 2004; 26 Wang 2016; 27 Sun 2010; 28 Wang 2021; 29 Lu 2017; 30 Zhong 2019; 31 Zhang 2020; 32 Li 2022; 33 Li 2014; 34 Liu 2013; 35 Mao 2023; 36 Cheng 2011; 37 Zhang 2012; 38 Huang 2022 39 . A stacked horizontal bar graph displays risk of bias across seven domains. The X-axis shows percentages from 0 to 100. The Y-axis lists: Random sequence generation (selection bias), Allocation concealment (selection bias), Blinding of participants and personnel (performance bias), Blinding of outcome assessment (detection bias), Incomplete outcome data (attrition bias), Selective reporting (reporting bias), Other bias. Bar values: Random sequence generation is 100% low risk. Allocation concealment is 50% low risk and 50% unclear risk. Blinding of participants and personnel is 45% low risk, 48% unclear risk, 7% high risk. Blinding of outcome assessment is 45% low risk, 48% unclear risk, 7% high risk. Incomplete outcome data, Selective reporting and Other bias are all 100% low risk. The legend indicates Low, Unclear and High risk of bias. A stacked horizontal bar graph showing risk of bias across seven assessment domains.
Figure 3 Risk-of-bias summary. A green plus sign indicates low risk, a yellow question mark indicates unclear risk, and a red minus sign indicates high risk. Included studies: Sun 2004; 26 Wang 2016; 27 Sun 2010; 28 Wang 2021; 29 Lu 2017; 30 Zhong 2019; 31 Zhang 2020; 32 Li 2022; 33 Li 2014; 34 Liu 2013; 35 Mao 2023; 36 Cheng 2011; 37 Zhang 2012; 38 Huang 2022 39 . Risk-of-bias summary for 14 studies using Cochrane RoB 1 tool, showing low, unclear and high risks.
Risk-of-bias graph. Green indicates low risk, yellow indicates unclear risk, and red indicates high risk. Included studies: Sun 2004; 26 Wang 2016; 27 Sun 2010; 28 Wang 2021; 29 Lu 2017; 30 Zhong 2019; 31 Zhang 2020; 32 Li 2022; 33 Li 2014; 34 Liu 2013; 35 Mao 2023; 36 Cheng 2011; 37 Zhang 2012; 38 Huang 2022 39 .
Risk-of-bias summary. A green plus sign indicates low risk, a yellow question mark indicates unclear risk, and a red minus sign indicates high risk. Included studies: Sun 2004; 26 Wang 2016; 27 Sun 2010; 28 Wang 2021; 29 Lu 2017; 30 Zhong 2019; 31 Zhang 2020; 32 Li 2022; 33 Li 2014; 34 Liu 2013; 35 Mao 2023; 36 Cheng 2011; 37 Zhang 2012; 38 Huang 2022 39 .
Eleven independent trials including 1014 unique participants reported pain relief rate. The two GZFL dose arms in Liu 2013 were combined for this primary analysis to avoid double-counting the shared placebo group. Heterogeneity was substantial (I 2 = 62%, P = 0.003). The random-effects estimate was associated with a higher trial-defined pain-relief rate for GZFL-based regimens than for controls (RR = 1.62, 95% CI 1.23–2.14; P = 0.0006; Figure 4 ).
Figure 4 Forest plot of pain relief rate. The two GZFL dose groups in Liu 2013 were combined for the primary analysis to avoid double-counting the shared placebo group. Squares show individual-study estimates, horizontal lines show 95% CIs, and the diamond shows the random-effects pooled estimate. Studies: Sun 2004; 26 Sun 2010; 28 Wang 2021; 29 Lu 2017; 30 Zhong 2019; 31 Zhang 2020; 32 Li 2014; 34 Liu 2013; 35 Cheng 2011; 37 Zhang 2012; 38 Huang 2022 39 . A forest plot displays study data with risk ratio (RR) estimates. Columns include Study; Experimental, Events and Total; Control, Events and Total; RR; 95% CI; Weight. X-axis labeled Risk ratio, with ticks at 0.01, 0.1, 1, 10, 100. Studies listed: Sun 2004 (RR 7.00, CI 0.38-129.84, weight 0.9%), Sun 2010 (RR 1.33, CI 0.87-2.04, 13.4%), Wang 2021 (RR 2.00, CI 0.54-7.46, 3.6%), Lu 2017 (RR 2.12, CI 1.09-4.16, 9.1%), Zhong 2019 (RR 2.00, CI 0.98-4.08, 8.6%), Zhang 2020 (RR 1.47, CI 1.03-2.09, 14.8%), Li 2014 (RR 2.23, CI 1.60-3.12, 15.2%), Liu 2013 (RR 1.42, CI 1.09-1.85, 16.5%), Cheng 2011 (RR 13.00, CI 0.77-220.80, 0.9%), Zhang 2012 (RR 8.72, CI 1.25-60.87, 1.9%), Huang 2022 (RR 0.85, CI 0.61-1.19, 15.1%). Random effects model: RR 1.62, CI 1.23-2.14, weight 100%. Heterogeneity: I superscript 2 = 62.4%, tau superscript 2 = 0.1037, p = 0.0030. A forest plot of pain relief rate risk ratio across studies, with pooled estimate above 1.
Forest plot of pain relief rate. The two GZFL dose groups in Liu 2013 were combined for the primary analysis to avoid double-counting the shared placebo group. Squares show individual-study estimates, horizontal lines show 95% CIs, and the diamond shows the random-effects pooled estimate. Studies: Sun 2004; 26 Sun 2010; 28 Wang 2021; 29 Lu 2017; 30 Zhong 2019; 31 Zhang 2020; 32 Li 2014; 34 Liu 2013; 35 Cheng 2011; 37 Zhang 2012; 38 Huang 2022 39 .
Prespecified subgroup interaction tests for pain relief were not statistically significant for treatment strategy (P = 0.86), comparator type (P = 0.40), or daily dose (P = 0.70). These analyses were exploratory because each subgroup contained few trials ( Appendices 3A – C ).
Three trials including 295 participants reported extractable end-of-treatment VAS means, standard deviations, and sample sizes and were included in this synthesis. Huang 2022 reported the Cox Menstrual Symptom Scale rather than VAS and was therefore not included. The other eligible trials reported categorical pain responses, composite symptom outcomes, other pain measures, or insufficient VAS data. The random-effects pooled MD was −1.26 (95% CI −4.21 to 1.70; P = 0.40), with considerable heterogeneity (I 2 = 98%); the confidence interval crossed no difference and did not establish a clear VAS benefit ( Figure 5 ).
Figure 5 Forest plot of VAS pain score. Squares show individual-study MDs, horizontal lines show 95% CIs, and the diamond shows the random-effects pooled estimate. Studies: Li 2022; 33 Mao 2023; 36 Zhang 2012 38 . Forest plot with study table and effect estimates. Columns: Study; Experimental (N, Mean, SD); Control (N, Mean, SD). Studies: Li 2022 (Exp N 54, Mean 3.04, SD 1.59; Ctrl N 54, Mean 1.52, SD 1.84), Mao 2023 (Exp N 64, Mean 3.75, SD 2.08; Ctrl N 64, Mean 5.52, SD 2.08), Zhang 2012 (Exp N 39, Mean 1.65, SD 1.38; Ctrl N 20, Mean 5.19, SD 1.56). Random effects totals: Exp 157, Ctrl 138. X-axis: Mean difference (lower scores favor GZFL), ticks: -4, -2, 0, 2, 4. Y-axis: studies and Random effects model. Each study: square and line. Right columns: MD, 95% CI, Weight. Li 2022: MD 1.52, 95% CI 0.87-2.17, Weight 33.5%. Mao 2023: MD -1.77, 95% CI -2.49 to -1.05, Weight 33.3%. Zhang 2012: MD -3.54, 95% CI -4.35 to -2.73, Weight 33.2%. Pooled diamond: MD -1.26, 95% CI -4.21 to 1.70, Weight 100%. Heterogeneity: I superscript 2 = 98%, tau superscript 2 = 6.6678, p < 0.0001. A forest plot of VAS pain score mean difference across three studies, with a pooled estimate near minus 1.
Forest plot of VAS pain score. Squares show individual-study MDs, horizontal lines show 95% CIs, and the diamond shows the random-effects pooled estimate. Studies: Li 2022; 33 Mao 2023; 36 Zhang 2012 38 .
Eleven independent trials including 1014 unique participants reported total effective rate, a trial-defined composite response outcome. Liu 2013 dose arms were combined and end-of-treatment data were used. Heterogeneity was considerable (I 2 = 89%, P < 0.0001). The random-effects estimate was associated with a higher total effective rate for GZFL-based regimens than controls (RR = 1.42, 95% CI 1.17–1.73; P = 0.0005), although variable definitions and heterogeneity limit confidence in the magnitude ( Figure 6 ).
Figure 6 Forest plot of total effective rate. The two GZFL dose groups in Liu 2013 were combined for the primary analysis. Squares show individual-study RRs, horizontal lines show 95% CIs, and the diamond shows the random-effects pooled estimate. Studies: Sun 2004; 26 Sun 2010; 28 Wang 2021; 29 Lu 2017; 30 Zhong 2019; 31 Zhang 2020; 32 Li 2014; 34 Liu 2013; 35 Cheng 2011; 37 Zhang 2012; 38 Huang 2022 39 . Forest plot with a study table at left and risk ratio plot at right. Table columns are Study; Experimental, Events and Total; Control, Events and Total. Studies and counts: Sun 2004, 27 of 30 vs 12 of 30; Sun 2010, 38 of 40 vs 32 of 40; Wang 2021, 37 of 41 vs 28 of 41; Lu 2017, 30 of 30 vs 21 of 30; Zhong 2019, 35 of 36 vs 30 of 36; Zhang 2020, 29 of 30 vs 25 of 30; Li 2014, 73 of 75 vs 55 of 75; Liu 2013 combined doses, 130 of 150 vs 41 of 71; Cheng 2011, 27 of 30 vs 5 of 30; Zhang 2012, 38 of 39 vs 4 of 20; Huang 2022, 45 of 55 vs 53 of 55. Right side columns list risk ratio, 95 percent confidence interval and weight: Sun 2004, 2.25, 1.43 to 3.54, 7.2 percent; Sun 2010, 1.19, 1.00 to 1.41, 10.8 percent; Wang 2021, 1.32, 1.05 to 1.67, 10.1 percent; Lu 2017, 1.42, 1.13 to 1.78, 10.1 percent; Zhong 2019, 1.17, 1.00 to 1.36, 10.9 percent; Zhang 2020, 1.16, 0.98 to 1.38, 10.8 percent; Li 2014, 1.33, 1.15 to 1.53, 11.1 percent; Liu 2013 combined doses, 1.50, 1.22 to 1.85, 10.4 percent; Cheng 2011, 5.40, 2.40 to 12.13, 3.9 percent; Zhang 2012, 4.87, 2.02 to 11.72, 3.5 percent; Huang 2022, 0.85, 0.74 to 0.97, 11.1 percent. The graph x axis label is Risk ratio, unitless, with ticks at 0.1, 0.5, 1, 2 and 10. A vertical reference line is at 1. The random effects model diamond is centered at 1.42 with 95 percent confidence interval 1.17 to 1.73; totals shown are 556 experimental and 458 control; overall weight 100.0 percent. Heterogeneity line reads: i superscript 2 equals 88.8 percent, T superscript 2 equals 0.0851, p less than 0.0001. A forest plot of total effective rate showing most studies favoring experimental over control.
Forest plot of total effective rate. The two GZFL dose groups in Liu 2013 were combined for the primary analysis. Squares show individual-study RRs, horizontal lines show 95% CIs, and the diamond shows the random-effects pooled estimate. Studies: Sun 2004; 26 Sun 2010; 28 Wang 2021; 29 Lu 2017; 30 Zhong 2019; 31 Zhang 2020; 32 Li 2014; 34 Liu 2013; 35 Cheng 2011; 37 Zhang 2012; 38 Huang 2022 39 .
Four independent trials including 512 unique participants contributed adverse-event data. The two Liu 2013 GZFL dose arms were combined to avoid double-counting the shared placebo group. Heterogeneity was low (I 2 = 17%, P = 0.31), so a common-effect model was used. GZFL-based regimens were associated with fewer reported adverse events than controls (OR = 0.42, 95% CI 0.23–0.76; P = 0.004; Figure 7 ).
Figure 7 Forest plot of adverse events. The two GZFL dose groups in Liu 2013 were combined to avoid double-counting the shared placebo group. Squares show individual-study ORs, horizontal lines show 95% CIs, and the diamond shows the common-effect pooled estimate. Studies: Lu 2017; 30 Li 2022; 33 Liu 2013; 35 Mao 2023 36 . Forest plot with table and odds ratio chart for adverse events. Left table: Study; Experimental Events/Total; Control Events/Total. Rows: Lu 2017, Exp 10/30; Ctrl 19/30. Li 2022, Exp 5/51; Ctrl 15/52. Liu 2013, Exp 6/150; Ctrl 4/71. Mao 2023, Exp 3/64; Ctrl 2/64. Right columns: OR, 95% CI, Weight: Lu 2017 OR 0.29, CI 0.10-0.84, Weight 38.2%. Li 2022 OR 0.27, CI 0.09-0.81, Weight 40.4%. Liu 2013 OR 0.70, CI 0.19-2.56, Weight 15.7%. Mao 2023 OR 1.52, CI 0.25-9.45, Weight 5.7%. Chart x-axis: Odds ratio, ticks 0.1, 0.5, 1, 2, 10. Each study shows a square at its OR with a line for 95% CI. Vertical line at 1. Common effect model: Exp 295, Ctrl 217, diamond at OR 0.42, CI 0.23-0.76, Weight 100%. Text: Heterogeneity, I superscript 2 = 17.1%, p = 0.3057. A forest plot of adverse events showing odds ratios mostly below 1 and a pooled estimate below 1.
Forest plot of adverse events. The two GZFL dose groups in Liu 2013 were combined to avoid double-counting the shared placebo group. Squares show individual-study ORs, horizontal lines show 95% CIs, and the diamond shows the common-effect pooled estimate. Studies: Lu 2017; 30 Li 2022; 33 Liu 2013; 35 Mao 2023 36 .
Across the four independent trials, adverse events were generally mild and no serious adverse events were reported. Lu 2017 reported 10/30 events with the GZFL-based regimen and 19/30 with the hormonal comparator. Li 2022 (the study corresponding to the 5/51 versus 15/52 data noted by the reviewer) reported fewer events with GZFL than with dydrogesterone. Liu 2013 reported infrequent events across two GZFL dose arms and placebo, and Mao 2023 reported 3/64 versus 2/64 events. Higher control-group rates may reflect comparator-specific adverse effects and study design; these findings do not establish a protective effect of GZFL.
Inflammatory biomarkers were reported in few trials. IL-6 was reported in three trials (n = 195; random-effects MD −1.48, 95% CI −3.31 to 0.35; I 2 = 89%), IL-10 in two trials (n = 132; common-effect MD −1.94, 95% CI −2.84 to −1.05; I 2 = 0%), and TNF-alpha in two trials (n = 132; common-effect MD −4.28, 95% CI −5.60 to −2.96; I 2 = 0%) ( Figure 8 ). Wang 2016 reported percentage change in TNF-alpha for a high-dose group and placebo rather than endpoint concentrations in pg/mL, so it was described narratively and not combined with the two endpoint studies. Although IL-10 was lower with GZFL-based regimens, IL-10 is an anti-inflammatory cytokine; therefore, this result does not support a simple anti-inflammatory mechanism. The biomarker findings remain exploratory and mechanistically inconclusive.
Figure 8 Forest plots of inflammatory biomarkers: ( A ) IL-6—Zhong 2019, Zhang 2020, and Li 2022; ( B ) IL-10—Zhong 2019 and Zhang 2020; and ( C ) TNF-alpha—Zhong 2019 and Zhang 2020. Squares show individual-study MDs, horizontal lines show 95% CIs, and diamonds show the selected pooled estimates (random effects for IL-6; common effect for IL-10 and TNF-alpha). Studies: Zhong 2019; 31 Zhang 2020; 32 Li 2022 33 . A) IL-6 forest plot: Zhong 2019 showed MD -2.28 (95% CI -3.44 to -1.12), weight 32.7%. Zhang 2020 had MD -2.44 (95% CI -3.72 to -1.16), weight 31.8%. Li 2022 reported MD 0.11 (95% CI -0.61 to 0.83), weight 35.6%. Pooled results: MD -1.48 (95% CI -3.31 to 0.35), weight 100%, heterogeneity I superscript 2 = 89.3%, p < 0.0001. B) IL-10 forest plot: Zhong 2019 MD -1.91 (95% CI -3.15 to -0.67), weight 52.7%. Zhang 2020 MD -1.98 (95% CI -3.29 to -0.67), weight 47.3%. Pooled results: MD -1.94 (95% CI -2.84 to -1.05), weight 100%, heterogeneity I superscript 2 = 0%, p = 0.9392. C) TNF-alpha forest plot: Zhong 2019 MD -4.29 (95% CI -6.09 to -2.49), weight 54.2%. Zhang 2020 MD -4.27 (95% CI -6.23 to -2.31), weight 45.8%. Pooled results: MD -4.28 (95% CI -5.60 to -2.96), weight 100%, heterogeneity I superscript 2 = 0%, p = 0.9882. Three forest plots of inflammatory biomarkers showing mostly negative mean differences across studies.
Forest plots of inflammatory biomarkers: ( A ) IL-6—Zhong 2019, Zhang 2020, and Li 2022; ( B ) IL-10—Zhong 2019 and Zhang 2020; and ( C ) TNF-alpha—Zhong 2019 and Zhang 2020. Squares show individual-study MDs, horizontal lines show 95% CIs, and diamonds show the selected pooled estimates (random effects for IL-6; common effect for IL-10 and TNF-alpha). Studies: Zhong 2019; 31 Zhang 2020; 32 Li 2022 33 .
Subgroup analyses for pain relief rate were conducted by treatment strategy (GZFL monotherapy vs add-on therapy), comparator type (placebo, NSAIDs, hormonal therapy), and daily dose (standard ~2.7–2.8 g/day vs higher ≥4 g/day). Across these analyses, effect estimates consistently favoured GZFL, and tests for subgroup differences were not significant, providing no clear evidence that the magnitude of benefit differed materially by regimen, comparator, or dose. Heterogeneity was generally low to moderate within most subgroups, with the main exception being comparisons against hormonal therapy, where heterogeneity was substantial and the pooled estimate was imprecise. Overall, the subgroup findings support a broadly consistent direction of effect for GZFL on pain relief, while suggesting that any apparent differences across subgroups are more likely driven by between-trial variability and limited subgroup power than by true effect modification ( Appendices 3A – C ).
Leave-one-out analysis of pain relief produced similar pooled estimates, and every omission retained a statistically significant result. For VAS, omitting Li 2022 changed the pooled MD to −2.65 (95% CI −4.38 to −0.91) and reduced I 2 to 90%; omitting either of the other trials left a confidence interval crossing no difference, and no single omission resolved the heterogeneity. For IL-6, omitting Li 2022 reduced I 2 from 89% to 0% and yielded MD −2.35 (95% CI −3.21 to −1.49); omitting either of the other trials left I 2 above 91% with confidence intervals crossing no difference. The two-study IL-10 and TNF-alpha syntheses were too sparse for stable influence attribution. These analyses support caution when interpreting continuous outcomes and biomarkers ( Appendices 4A and B ). The pain-relief funnel plot is shown in Appendix 5 .
Certainty was moderate for pain relief and was downgraded for risk-of-bias concerns. Although heterogeneity was substantial, effects were generally in the same direction and prespecified subgroup interactions were not statistically significant, so no additional inconsistency downgrade was applied. Total effective-rate evidence was low certainty because of risk of bias and considerable inconsistency. VAS evidence was very low certainty because of risk of bias, very serious inconsistency, and imprecision, as the confidence interval crossed no difference. Adverse-event evidence was low certainty because of risk-of-bias concerns and imprecision from the small number of events and trials; it was not downgraded for inconsistency because I 2 was 17%. Certainty was very low for IL-6 because of risk-of-bias concerns, serious inconsistency, and imprecision. IL-10 and TNF-alpha were low-certainty because of risk-of-bias concerns and imprecision from only two small trials, despite I 2 = 0%. Outcome-specific reasons are provided in Appendices 6A and B .
Discussion
In these short-term randomized trials, GZFL-based regimens were associated with higher trial-defined pain relief and total effective rates. In contrast, the pooled VAS estimate was imprecise, crossed no difference, and was highly heterogeneous. The direction of the main dichotomous outcome was stable in sensitivity analysis, but substantial heterogeneity affected pain relief, VAS, total effective rate, and biomarkers. The findings should therefore be interpreted as short-term associations rather than definitive causal effects or proof of a uniform benefit across regimens.
Previous reviews of Chinese herbal medicines and Guizhi Fuling formulations have reported potential benefits for dysmenorrhoea and other gynaecological conditions, but often included heterogeneous formulations, mixed clinical indications, or older trials. 22 , 23 The present review narrows the question to randomized evidence relevant to primary dysmenorrhoea and adds more recent trials. Its findings are broadly compatible with earlier reports, while the current GRADE assessment and heterogeneity analyses indicate greater uncertainty than efficacy rates alone might suggest.
A multicentre, randomized, double-blind, placebo-controlled trial reported higher pain-response rates with two GZFL doses than with placebo and suggested persistence during follow-up. 35 This trial supports the direction of the pooled pain-relief result, but one trial cannot establish long-term comparative effectiveness or generalizability beyond similar Chinese clinical settings.
Comparisons with conventional medicines should be interpreted according to the actual treatment strategy. Eight trials evaluated monotherapy and six evaluated add-on or multimodal strategies, with placebo, analgesic/NSAID, or hormonal comparators. For comparisons with unequal co-interventions, the estimate reflects the complete combined strategy rather than an isolated GZFL effect. These clinically different questions should not be treated as interchangeable, and their combination may contribute to heterogeneity.
Subgroup analyses did not identify statistically significant differences by monotherapy versus add-on use, comparator type, or dose. However, these analyses contained few studies and limited power, so absence of a subgroup interaction should not be interpreted as proof that effects are identical across clinical settings.
Substantial or considerable heterogeneity in pain relief, VAS, total effective rate, and biomarkers limits confidence in pooled magnitudes. Potential contributors include different outcome definitions, scales, measurement times, participant characteristics, GZFL regimens, and comparator therapies. Influence analyses reduced heterogeneity for some outcomes after omitting individual studies but did not remove the broader uncertainty.
Animal, in vitro, and network-pharmacology studies propose that GZFL constituents may influence cyclooxygenase activity, prostaglandin production, inflammatory signalling, uterine contraction, and microcirculation. 8 , 10 , 14 , 40–45 These findings provide biological hypotheses but are not direct evidence that the same mechanisms mediate clinical effects in women with primary dysmenorrhoea.
The clinical biomarker findings do not establish a coherent anti-inflammatory mechanism. TNF-alpha was lower in two small endpoint studies, IL-6 was inconclusive and heterogeneous, and the lower IL-10 estimate is not readily compatible with a simple anti-inflammatory interpretation. Mechanistic claims must therefore remain tentative.
Pharmacokinetic and experimental findings on individual constituents may inform future research, but they cannot be used to infer clinical efficacy, optimal dose, or treatment duration from this meta-analysis. Human mechanistic studies with prespecified biomarkers are needed. 45–47
Adverse events were generally mild, but only four independent trials contributed to the pooled safety analysis. The higher event rates in the control groups of Lu 2017 and Li 2022 may reflect adverse-effect profiles of hormonal comparators, concomitant treatment, or reporting differences. These small, short-term data do not demonstrate superior long-term safety or protection against adverse events.
This review focuses on randomized evidence for GZFL-based regimens in primary dysmenorrhoea and evaluates pain-related outcomes, overall response, biomarkers, adverse events, subgroup effects, sensitivity, and GRADE certainty. These features help distinguish it from broader earlier reviews, although conclusions remain dependent on the limitations of the included reports.
Several limitations warrant consideration. Allocation concealment, blinding, and outcome definitions were incompletely reported in some trials. Ethics approval was explicitly reported in only 4 of 14 reports and informed consent in 8, although non-reporting cannot be interpreted as non-approval. All trials were conducted in China and follow-up was short, limiting generalizability and long-term safety inference. Substantial heterogeneity affected continuous outcomes and biomarkers, and mechanistic evidence was largely animal or in vitro. Future RCTs should use longer follow-up, standardized pain and response outcomes, credible blinding/placebo or double-dummy designs, prespecified safety monitoring, and transparent ethics reporting.