Data
Data used for the current review are available from the corresponding author upon reasonable request.
Methods
The study protocol was approved by the Peking Union Medical College Hospital Ethics Committee (JS-2238, 12/24/2019). The study was registered with the Clinical Trials.gov Protocol Registration and Results System (PRS) ( https://register.clinicaltrials.gov : NCT04245137 ). All participants provided informed consent prior to participating in the study.
An overview of the study design is presented in Fig. 1 (I). This large, multicenter, population-based study was conducted between 2020 and 2024 in 21 diverse centres. The centres covered almost all 7 regions in China with different health care levels (from county to provincial) and were willing to participate in this study. We recruited volunteers who responded to community-based recruitment advertisements. All candidates underwent an eligibility assessment. The inclusion criteria were nulliparous women or at least 2 years since their last delivery. Healthy women were defined as: 19 (i) no symptoms of PFDs: leaking urine in the past month when coughing/sneezing/lifting heavy objects; leaking urine in the past month when having urgency/intention to urinate before going to the bathroom; unable to control bowel movements for the past month, with a tendency to defecate and later overflowing before going to the bathroom; often see or feel lumps coming out of the vagina; pain in the abdomen, lumbosacral region, and buttocks below the navel, lasting for more than 6 months. (ii) PFM strength M3–M5 by modified Oxford grading Scale (MOS) (M0: no contraction; M1: flicker; M2: weak; M3: moderate; M4: good; M5: strong). 20 (iii) any point less than 0 according to the pelvic organ prolapses quantification system (POP-Q). 21 (iv) No tenderness and/or tender points during pelvic floor palpation. According to International Continence Society (ICS) terminology, PFM disorder/dysfunction refers to any deviation from normal function that causes distress to patients and has related signs and investigations. 20 In our study, normal PFM indicated a certain level of muscle strength (≥M3), without muscle spasms or clinical manifestations of PFD. Not meeting any of the above criteria was considered abnormal, with subdivision including urinary incontinence, pelvic organ prolapse, pelvic pain, and low strength (≤M2). Fig. 1 Study process flowchart . (I) The research flowchart (II) The architecture of AI-Diagnostician-PFD. ML models, Machine Learning Models, DL models, Deep Learning Models.
Study process flowchart . (I) The research flowchart (II) The architecture of AI-Diagnostician-PFD. ML models, Machine Learning Models, DL models, Deep Learning Models.
The exclusion criteria were: (i) no history of sexual intercourse or intolerance to vaginal examination; (ii) pregnancy or breastfeeding; (iii) acute inflammatory conditions affecting the lower reproductive tract or lower urinary system; (iv) radical surgery or radiotherapy for malignant tumour in pelvic cavity; (v) history of hysterectomy and subtotal hysterectomy; (vi) pelvic floor surgery; (vii) neurological dysfunction that significantly affect muscle function; and (viii) cannot cooperate with PFM contraction.
From April 2020 to May 2024, 1668 participants were recruited from 21 centres across China. The participants initially signed the informed consent form and completed a baseline questionnaire with demographic and medical information. The participants were required to undertake a series of assessments, including digital PFM strength measurement quantified by the MOS, prolapse assessment according to the POP-Q system during the Valsalva manoeuvre, and PFM palpation. The participants performed an MVC when instructed to move their PFM ‘inwards and upwards as hard as you can’, relax for 10s, and repeat the contraction. The correct contraction was identified by vaginal palpation directed at 4 and 8 o'clock and observation of an inward lift at the perineum during contraction, with discouragement of breath-holding and synergistic contractions of abdominal, hip adductors and glutaeal muscles (slight abdominal contractions are allowed). POP-Q and pelvic floor pain assessment continued if digital strength was normal. All assessments were performed by a single experienced female physical therapist at each centre. Physical therapists specialized in pelvic floor rehabilitation for at least 3 years. At the project kickoff meeting, the PI conducted a face-to-face assessment of each centre's physical therapists, including research processes and specific operational techniques. Only those who passed the assessment were eligible to enrol as participants. In addition, the research team also conducted real-time and irregular on-site quality control in each center.
The vaginal sEMG was assessed according to the Glazer protocol by the same assessor of physical examination. The intravaginal disposable sensor (VET-S, Nanjing Vishee Medical Technology Co., Ltd., China) was spindle-shaped with a half-curvature design at the front end. The electrode set consisted of a plastic substrate with metal sheets on both sides, measuring 46 mm±2 mm in length. The electrode was placed with the two metal sheets in 3 and 9 o'clock position, respectively, in contact with lateral vaginal walls, primarily measuring pubococcygeus muscle. The vaginal electrode collected EMG signals from both sides of PFMs. Thus, an abdominal electrode patch was pasted 2 cm lateral to the umbilicus, and the other was pasted 2 cm downwards, aligning with the above one. The reference electrode was placed on the electrically-neutral anterior superior iliac spine.
The evaluation was conducted with participants in a semi-recumbent posture, maintaining an approximate 120-degree trunk-thigh angle during measurements. The legs were naturally extended and rotated outward with the heels slightly separated, maintaining complete relaxation during the procedure. The Glazer protocol was installed into the sEMG-based device, which recorded and analysed the sEMG signals from PFMs performing a series of contractions and relaxations. The protocol consisted of five steps: (1) average mean amplitude (μV) and variability for 60s pre-baseline rest; (2) average peak amplitude and relaxation time for 5 rapid maximal contractions (10s rests in between): participants were instructed to contract the PFM as quickly as possible and to relax the PFM immediately after each contraction; (3) average mean amplitude, variability, and relaxation time for five 10s tonic contractions (10-s rests in between); (4) average mean amplitude and variability for 60-s endurance contraction; (5) average mean amplitude and variability to evaluate the recovery of pelvic floor muscles for 60-s post-baseline rest. In total, 11 sEMG parameter data were collected in the steps above. Furthermore, the sEMG signals were analysed considering the mean of 20-Hz high-pass and 500-Hz low-pass filters.
A comprehensive multidimensional sEMG database was established by integrating pelvic floor sEMG data, demographic information, medical history, menstrual and delivery history, POP-Q, visual analogue scale (VAS) pain score, and digital palpation by MOS. The database encompasses 63 personal metrics, providing a robust foundation for subsequent analysis ( Appendix p2–p5 ).
Preprocessing was performed to ensure completeness based on the multidimensional sEMG database. To maintain completeness, samples with missing values were excluded for the 11 pelvic floor sEMG parameters. Similarly, samples with incomplete data for the POP-Q scores or modified Oxford grading scales were removed, as both metrics were essential for determining the diagnostic labels.
Outlier detection was performed using the Interquartile Range (IQR) method for each pelvic floor sEMG parameter, with the lower and upper limits calculated using the 5th and 95th percentiles.
Following this, we labelled the refined data, segregating the samples into two categories. One group of healthy women with normal pelvic floor muscle function was labelled “normal”. The second group, including individuals with urinary incontinence, pelvic organ prolapse, chronic pelvic pain, and muscle strength scores of M0-M2, was labelled as “abnormal.”
Considering the factors of geographical coverage, we divided the dataset based on the proportional sample sizes of the centres to ensure that the training and evaluation were representative and to minimize bias caused by variations in data distribution. Specifically, data from 15 centres were split into 60% for the training dataset and 40% for the test dataset. In contrast, an independent validation dataset was constructed from 6 additional centres. Both groups contained similar proportions of normal and abnormal samples and comprehensively covered the eastern, central, and western regions.
We designed an AI-based model named AI-Diagnostician-PFD, which derives the reference ranges for the sEMG parameters and provides an intelligent diagnosis of PFDs. The model's architecture was illustrated as shown in Fig. 1 (II).
The primary objective of the AI-Diagnostician-PFD was to derive accurate reference ranges for the 11 sEMG parameters. To achieve this, we utilized an ML model based on a genetic algorithm (GA) for multi-objective optimization. 22 Inspired by natural selection and evolution principles, GA iteratively refines a population of candidate solutions through selection, crossover, and mutation operations. This study explicitly employed GA to optimize the thresholds for the sEMG parameters, ensuring that the derived reference ranges accurately distinguish between normal and abnormal PFM function. The optimization process utilized true labels (0 for normal, 1 for abnormal) to guide the algorithm in aligning the derived reference ranges with clinical definitions. Balancing sensitivity and specificity, the algorithm systematically identified optimal parameter boundaries aligning with clinical definitions. Detailed information about the implementation of the GA and its integration into reference range calculations is provided in Appendix p11–p12 .
The second task of AI-Diagnostician-PFD was to achieve a high-precision diagnosis of PFDs. To accomplish this, we designed and trained an integrated ML model. The model adopted the idea of ensemble learning, 23 integrating four basic classifiers: Stochastic Gradient Descent Classifier (SGD Classifier), 24 Support Vector Classifier (SVC), 25 Random Forest (RF), 26 and multi-layer perceptron (MLP), 27 to optimize diagnostic performance.
These classifiers were selected for their established performance in ML and DL tasks. Each base classifier was trained individually, and their predictions were aggregated through a weighted voting mechanism. The weights were determined based on the performance of each classifier on the test dataset. The final prediction was derived by iteratively training a meta-classifier that utilized the outputs of the base classifiers as inputs, enhancing the diagnostic precision of the model.
The evaluation process was conducted in two parts. The first focused on validating the AI-Reference ranges, while the second evaluated the performance of AI-Diagnostician-PFD. A. Evaluation of AI-Reference Ranges
Evaluation of AI-Reference Ranges
In the first part of the evaluation, we evaluated the diagnostic performance of AI-Reference ranges. For comparison, we also implemented the standard methodologies outlined in the NCCLS-C28-A2 publication, titled “ How to Define and Determine Reference ranges in the Clinical Laboratory ,” issued by the National Committee for Clinical Laboratory Standards (NCCLS). Specifically, we followed the “ Approved Guidelines Second Edition ” to calculate the reference ranges for the normal population, using three different methods: the normal distribution method, the non-parametric method, and the multivariate non-parametric method. The calculation was performed using normal population samples to ensure the reference ranges accurately reflect the distribution of healthy individuals. The computational process was outlined in the Appendix p6–10 .
We transformed the original sEMG data into binary values for each reference range; the data points falling within the range were mapped to 0, while those outside were mapped to 1. These binary features were then used in the AI-Diagnostician-PFD for further evaluation.
During the validation process, we compared the diagnostic performance of five reference ranges: the standard Glazer ranges, the normal distribution method ranges, the non-parametric method ranges, the multivariate non-parametric method ranges, and our AI-Reference ranges. In particular, we validated the diagnostic performance of the five reference ranges at the external validation dataset and each of the six centres in the validation dataset. B. Evaluation of AI-Diagnostician-PFD
Evaluation of AI-Diagnostician-PFD
In the second part of the evaluation, we assessed the reliability and diagnostic performance of AI-Diagnostician-PFD. To achieve this, we compared its performance with six baseline models: three DL models, including an MLP, a Convolutional Neural Network (CNN) 28 and a Transformer 29 ; as well as three classical machine learning models, including SGD Classifier, SVC, and Random Forest. These baseline models represent a diverse range of learning paradigms, providing a comprehensive evaluation framework for assessing the diagnostic capabilities of AI-Diagnostician-PFD.
The DL models (CNN and Transformer) were specifically designed to leverage the original sEMG data, capturing complex patterns and dependencies in the data. The MLP consisted of a single layer with ReLU activation function, optimized using Adam. The CNN included two convolutional layers, followed by one fully connected layer with dropout, and utilized ReLU for non-linearity. Furthermore, the Transformer model used one encoder layer with four multi-head attention mechanisms and was trained with an SGD optimizer.
The classical ML models were trained on binary-transformed data based on the AI-Reference ranges, where values within the AI-Reference ranges were mapped to 0. Those outside were mapped to 1. C. Evaluation process
Evaluation process
The evaluation process began with an initial validation on the internal test dataset, followed by validation on the independent external validation dataset. Evaluation metrics, including AUC, sensitivity, and specificity, were used to quantify the model's performance.
The study's funders had no role in the study design, data collection, analysis, interpretation, or report writing.
Results
In this study, 1668 participants were recruited from 21 centres in different regions of China. Based on enrolment criteria and handling outliers, 1092 normal and 513 abnormal women were included for analysis. Table 1 shows the general characteristics. The names of the 21 centres and the sample info are in Appendix p13 . Table 1 Sociodemographic and clinical characteristics of the multidimensional sEMG database. Numbers (n = 1605) Normal (n = 1092) Abnormal (n = 513) Age ≤20 6 3 (0.27%) 3 (0.58%) 21–30 554 328 (30.04%) 226 (44.05%) 31–40 655 481 (44.05%) 174 (33.92%) 41–50 245 199 (18.22%) 46 (8.97%) ≥ 51 145 81 (7.42%) 64 (12.48%) BMI ≤18.5 131 87 (7.97%) 44 (8.58%) 18.5–25 1223 865 (79.21%) 358 (69.79%) ≥25 251 140 (12.82%) 111 (21.64%) Number of births 0 615 339 (31.04%) 276 (53.80%) 1 670 531 (48.63%) 139 (27.10%) 2 281 206 (18.86%) 75 (14.62%) 3 26 9 (0.82%) 17 (3.31%) 4 6 3 (0.27%) 3 (0.58%) 5 5 2 (0.18%) 3 (0.58%) 6 2 2 (0.18%) 0 (0.00%) Delivery history Vaginal delivery 591 422 (38.64%) 169 (32.94%) Cesarean section 321 280 (25.64%) 41 (7.99%) Vaginal delivery and cesarean section 58 52 (4.76%) 6 (1.17%) No 635 338 (30.95%) 297 (57.89%) POP-Q I 1048 864 (79.12%) 184 (35.87%) II-mild (−1 cm ≤ any point 0 cm) 116 0 (0.00%) 116 (22.61%) III 2 0 (0.00%) 2 (0.39%) Family history of pelvic floor dysfunction UI 158 110 (10.07%) 48 (9.36%) POP 46 33 (3.02%) 13 (2.53%) UI and POP 44 25 (2.29%) 19 (3.70%) No 1189 794 (72.71%) 395 (77.00%) N/A 168 130 (11.90%) 38 (7.41%) History of pelvic surgery Yes 114 68 (6.23%) 46 (8.97%) No 1491 1024 (93.77%) 467 (91.03%) Medical history Diabetes 11 7 (0.64%) 4 (0.78%) Chronic constipation 73 33 (3.02%) 40 (7.80%) Chronic cough 7 3 (0.27%) 4 (0.78%) Chronic bronchitis 3 1 (0.09%) 2 (0.39%) Hypertension 13 11 (1.01%) 2 (0.39%) Diabetes and hypertension 1 1 (0.09%) 0 (0.00%) Diabetes and chronic constipation 2 1 (0.09%) 1 (0.19%) No 1495 1035 (94.78%) 460 (89.67%) Smoking history Yes 25 19 (1.74%) 6 (1.17%) No 1580 1073 (98.26%) 507 (98.83%) Gynaecological disease No obvious pelvic mass detected in the past year 1400 960 (87.91%) 440 (85.77%) Endometriosis 18 14 (1.28%) 4 (0.78%) Fibroid 98 71 (6.50%) 27 (5.26%) Other 121 70 (6.41%) 51 (9.94%) Pelvic floor rehabilitation history No 1368 890 (81.50%) 478 (93.18%) Occasionally, irregular,<30 contractions per day 226 193 (17.67%) 33 (6.43%) ≥30 contractions per day on average 11 9 (0.82%) 2 (0.39%) Menstrual history Premenopausal 1462 1015 (92.95%) 447 (87.13%) Perimenopausal period 15 9 (0.82%) 6 (1.17%) Postmenopausal 128 68 (6.23%) 60 (11.70%) Modified Oxford grading scale <3 451 0 (0.00%) 451 (87.91%) ≥3 1154 1092 (100.00%) 62 (12.09%) Data split Training dataset 782 528 (48.35%) 254 (49.51%) Test dataset 522 352 (32.24%) 170 (33.14%) Validation dataset 301 212 (19.41%) 89 (17.35%)
Sociodemographic and clinical characteristics of the multidimensional sEMG database.
The AI-Reference ranges for the sEMG parameters derived by AI-Diagnostician-PFD, along with three reference ranges calculated using the NCCLS standard methods and the existing Glazer standard ranges, were summarized in Table 2 , presenting a total of five types of reference ranges. Table 2 sEMG parameter reference ranges by different methods. Five types of ranges Descriptive statistics Normality& non-parametric Non-parametric Multivariate non-parametric Glazer AI-reference Mean SD Median Quartile (25%, 75%) Pre resting Mean μV [0.89, 17.95] [1.14, 18.27] [2.63, 13.69] [2.0, 4.0] [1.06, 8.85] 7.54 4.23 Variability <0.48 <0.48 <0.20 <0.2 <0.32 0.09 (0.08, 0.13) Rapid contraction Maximum μV [16.73, 99.78] [16.75, 104.57] [28.78, 85.80] [35.0, 45.0] [23.68, 88.38] 44.59 21.49 Relaxation time s <2.96 <2.96 <0.40 <0.5 <0.32 0.28 (0.22, 0.48) Continuous contraction Mean μV [11.38, 76.05] [11.38, 76.05] [19.05, 56.27] [30.0, 40.0] [23.99, 37.81] 27.64 (19.38, 35.59) Variability <0.52 <0.54 <0.44 <0.2 <0.49 0.34 0.10 Relaxation time s <6.30 <6.30 <0.86 <1.0 <0.15 0.30 (0.12, 0.94) Endurance contraction Mean μV [10.47, 71.71] [10.47, 71.71] [17.12, 53.44] [25.0, 35.0] [13.12, 39.07] 25.13 (17.74, 33.06) Variability <0.49 <0.49 <0.33 <0.2 <0.60 0.21 (0.17, 0.27) Post resting Mean μV [0.89, 17.52] [0.95, 18.17] [2.64, 13.28] [2.0, 4.0] [2.76, 16.51] 7.19 4.06 Variability <0.41 <0.41 <0.19 <0.2 <0.38 0.10 (0.08, 0.15)
sEMG parameter reference ranges by different methods.
Descriptive statistics were utilized to summarize the sEMG data: mean and standard deviation (SD) were calculated for normal distribution parameters. At the same time, median and quartile were used for non-normal distribution parameters. The validation for normally distributed parameters was provided in the Appendix p6 .
We evaluated the diagnostic performance of five reference ranges on both the internal test dataset and external validation dataset. The AI-Reference ranges outperformed the other methods in terms of diagnostic performance. The ROC curves and performance metrics for each reference range are shown in Fig. 2 A. Fig. 2 ROC Curves and Diagnostic Performance of Different Reference Ranges . (I) The results of the internal validation on the test dataset. (II) The results of the validation on the external validation dataset (A) ROC curves and corresponding AUCs for the diagnostic performance of five methods, with 95% confidence intervals calculated using 1000 bootstrap iterations. (B) A boxplot displaying the AUCs and their Interquartile Range (IQR); the improvement rates were annotated in pink font. (C) Comparisons of the mean sensitivity of different sEMG parameter reference ranges. (D) Comparisons of the mean specificity of different sEMG parameter reference ranges. AUC, area under the receiver operating characteristic curve. ROC, receiver operating characteristic.
ROC Curves and Diagnostic Performance of Different Reference Ranges . (I) The results of the internal validation on the test dataset. (II) The results of the validation on the external validation dataset (A) ROC curves and corresponding AUCs for the diagnostic performance of five methods, with 95% confidence intervals calculated using 1000 bootstrap iterations. (B) A boxplot displaying the AUCs and their Interquartile Range (IQR); the improvement rates were annotated in pink font. (C) Comparisons of the mean sensitivity of different sEMG parameter reference ranges. (D) Comparisons of the mean specificity of different sEMG parameter reference ranges. AUC, area under the receiver operating characteristic curve. ROC, receiver operating characteristic.
In comparison to the Glazer reference ranges, the AI-Reference ranges exhibited superior performance, with AUCs of 0.81 (95% CI: 8.13 × 10 −1 , 8.16 × 10 −1 ) and 0.79 (95% CI: 7.90 × 10 −1 , 7.94 × 10 −1 ) on the test and validation datasets, respectively. Both these values exceeded the Glazer's AUCs of 0.76 (95% CI: 7.56 × 10 −1 , 7.59 × 10 −1 ) and 0.68 (95% CI: 6.74 × 10 −1 , 6.78 × 10 −1 ), indicating 11% improvement. To formally assess the statistical significance of these differences, we performed DeLong's test for AUC comparison. The results showed that the AI method significantly outperformed the Glazer protocol, with p-values of 0.18 × 10 −3 for the test dataset and 0.35 × 10 −5 for the validation dataset (both p < 0.05). These findings confirmed that the AI method significantly improved over the Glazer standard.
Furthermore, the sensitivity of the AI-Reference ranges surpassed that of the Glazer reference ranges by 0.05 and 0.14 points on the test and validation datasets, respectively. Similarly, the specificity also demonstrated an advantage, being 0.06 and 0.14 points higher on the same datasets. All indicators of AI-Reference ranges were higher than those of NCCLS standard methods.
We evaluated the reference ranges separately for each participating center for the external validation dataset. The model demonstrated consistently good performance across all centres, with detailed results in the Appendix on p. 14–15 . It is worth noting that for Foshan Maternity and Child Healthcare Hospital, the sample size was relatively small (25 samples, with only 1 labelled as abnormal), which might limit the representativeness of the evaluation metrics for this specific center. However, the aggregated results across all centres supported the robustness and generalizability of the model.
We conducted a comprehensive evaluation of the diagnostic prowess of AI-Diagnostician-PFD against six baseline models, with the results presented in Fig. 3 . On both the internal test dataset and the external validation dataset, AI-Diagnostician-PFD achieved an AUC that was 1% higher than the best-performing baseline model. Additionally, its sensitivity and specificity surpassed the best-performing baseline model by 2% in the external validation dataset. Fig. 3 Diagnostic Perfo rmance of AI-Diagnostician-PFD and Baseline Models . (I) The results of the internal validation on the test dataset. (II) The results of the validation on the external validation dataset. (A) Comparison of AUC across models. (B) Comparison of sensitivity across models. (C) Comparison of specificity across models. AUC = area under the receiver operating characteristic curve. The x-axis denotes the classification models: SGD Classifier (blue), Support Vector Machine Classifier (light blue), Random Forest (grey), Convolutional Neural Network (light green), Multilayer Perceptron (green), Transformer (dark green), and AI-Diagnostician-PFD (red). The y-axis represents the metric values for AUC, sensitivity, and specificity. Among all models, AI-Diagnostician-PFD consistently outperforms others across all three metrics in both datasets.
Diagnostic Perfo rmance of AI-Diagnostician-PFD and Baseline Models . (I) The results of the internal validation on the test dataset. (II) The results of the validation on the external validation dataset. (A) Comparison of AUC across models. (B) Comparison of sensitivity across models. (C) Comparison of specificity across models. AUC = area under the receiver operating characteristic curve. The x-axis denotes the classification models: SGD Classifier (blue), Support Vector Machine Classifier (light blue), Random Forest (grey), Convolutional Neural Network (light green), Multilayer Perceptron (green), Transformer (dark green), and AI-Diagnostician-PFD (red). The y-axis represents the metric values for AUC, sensitivity, and specificity. Among all models, AI-Diagnostician-PFD consistently outperforms others across all three metrics in both datasets.
Discussion
Using AI methodologies, we constructed a diagnostic model named AI-Diagnostician-PFD. In comparison to the reference ranges of Glazer, AI-Reference ranges demonstrate a significantly higher AUC, indicating an 11% increase in the external validation dataset. The result highlights the advantage of using AI-based methods to establish reference ranges, as they can account for more complex patterns and variations in the data, leading to more precise and accurate results. Establishing the reference range of quantitative sEMG parameters would help standardize pelvic muscle diagnosis and functional evaluation.
In diagnosing PFDs, AI-Diagnostician-PFD emerged as the top performer, surpassing all ML and DL models. These findings underscore the superiority of AI-Diagnostician-PFD over both traditional ML and cutting-edge DL models, as evidenced by its performance across all crucial diagnostic metrics. This achievement underscores the advanced diagnostic capabilities of AI-Diagnostician-PFD. These findings confirm the potential of the AI-Diagnostician-PFD model to differentiate healthy and unhealthy people and help to enhance clinical decision-making. The improvement in diagnostic accuracy is due to the broader range of the new reference interval compared to the current Glazer standard, with AUC results indicating an 11% increase in diagnostic accuracy. Next, we will verify its diagnostic accuracy in more populations, especially PFD patients with hypotonic or hypertonic PFD. The framework developed in this study demonstrated significant translational potential, enabling precise identification of PFDs and providing clinical value at both individual and population levels.
According to the ICS report of PFM assessment published in 2011, there is no universally recognized standard for distinguishing between normal or abnormal PFM function. 20 In our study, according to our healthy definition, some young participants (21–30 yrs old) who had not given birth were regarded as abnormal because of their low muscle strength. In the literature review, we find a similar phenomenon. Ferreira's observational study documented that 35% of collegiate populations exhibited suboptimal contractile performance of PFM graded as weak or moderate under dual-rater assessment. 30 Dietz et al. also revealed that almost half of nulligravid females of reproductive age demonstrated inadequate PFM contractions when unassisted by instructions. 31 so the value of physical examination has limitations, and quantitative investigation such as sEMG may play a crucial complementary role in diagnosing PFD.
The Glazer protocol assessment is currently recognized as one pelvic floor surface electromyography evaluation method, which has a great auxiliary decision-making function for diagnosing and treating PFDs. Compared with currently recognized Glazer standards, the reference range of contraction by AI was wider, which reflected individual variability in muscle function. PFM function is influenced by various factors, such as age, hormone levels, childbirth et al. From the mechanism of disease occurrence, neuromuscular dysfunction is only one aspect; ligaments and fascia also play essential roles. From the medical perspective, the prediction of a disease requires the integration of multidimensional factors. Therefore, the AI model incorporated 11 pelvic floor sEMG parameters to assist in predicting PFDs, resulting in higher accuracy in diagnosis.
Our research had several strengths. First, the study design was a multicenter, population-based cross-sectional study, and experienced operators obtained measurements. We recruited women with different ages, medical histories, and reproductive histories for sEMG data collection. Second, our multidimensional database had 1605 cases with 63 attributes, higher in data volume and dimension than existing studies. 8 , 13 , 32 Third, our AI-Diagnostician-PFD was based on a multi-objective optimization strategy. The concept of multi-objective optimization is that when multiple objectives need to be achieved in a particular scenario, there may be internal conflicts between objectives. The optimization of one objective is at the cost of the deterioration of other objectives, making it difficult to find a unique optimal solution. The multi-objective optimization method used in this article makes coordination and compromise among multiple electromyographic parameters, making the overall goal (identifying abnormal pelvic floor populations) as optimal as possible. Compared to a single ML model, the advantage of ensemble learning models is that they can organically combine multiple single learning models to obtain a unified ensemble learning model, thereby obtaining more accurate, stable, and robust results. AI-Diagnostician-PFD was designed with ensemble learning, combining multiple weak learners into strong learners to achieve better diagnostic performance.
The weakness of our study concerns the uneven distribution of age among the participants. Naturally, external validation across diverse demographic cohorts and expanded sample sizes would confirm the generalizability of our conclusions. Next, we did not differentiate different types of PFDs among unhealthy women, such as urinary incontinence, pelvic pain, etc. Future research will further describe the characteristic manifestations of sEMG parameters in different types of PFDs. Exploring the model's capabilities to classify subtypes of PFDs is an important direction for future research. Third, we assessed the participants using only one assessor without considering the inter-observer variability.
This research developed normative reference values for pelvic floor sEMG parameters by employing AI algorithms to identify PFDs intelligently. These findings demonstrated a viable framework for implementing AI-based solutions in PFDs management. The model is a potential tool to assist physical therapists and could be embedded in clinical workstations to provide free real-time guidance.
Contributors
Juan Chen and Lan Zhu conceived and designed the study, contributed to data collection and wrote the article. Jiahui Yao and Tengjiao Wang contributed to AI analyses and wrote the article. Wei Chen, Heyuan Wang and Feng Zhang contributed to model implementation and result visualization. Xiaoying Xu, Huan Ge, Hongmei Zhou, Jin Cen, Dan Li, Bengui Jiang, Li He, Tingting Fu, Zhengxian Xu, Lei Chu, Shuxia Zhang, Dongmei Yao, Linyi Wei, Liu Huang, Anjing Ge, Cuiping Jin, Zimu Fu, Qin Liu, Xuefeng Yu, Chengmao Zhao participated in the execution of the study and data collection. Lan Zhu contributed to all aspects of study design, obtained funding, revised the article and assumed overall responsibility for the study. All authors were responsible for draughting the article. All authors have read and approved the final version for publication.
Introduction
Pelvic floor dysfunctions (PFDs) refer to a collection of medical conditions involving impaired pelvic floor function, such as urinary incontinence, fecal incontinence, pelvic organ prolapses, and pelvic pain. Notably, the prevalence of symptomatic PFDs varies from 11.9% to 67.9% in both low- and middle-income countries and developed countries. Additionally, one-quarter of all adult women suffer from at least one PFD. 1 , 2 , 3 , 4 , 5 These conditions not only profoundly impact patients' quality of life by physical discomfort and emotional distress but also create substantial societal burdens on public health systems globally.
Optimal treatment planning originates from rigorous diagnosis achieved through sequential clinical processes: initial history-taking and physical examination, followed by specialized pelvic floor muscle assessment and imaging investigation. Evaluating pelvic floor muscle (PFM) function is the first step of physical therapy for patients with PFDs. Surface Electromyography (sEMG) is a well-established method to quantify neuromuscular activation and detect early muscle abnormalities. In clinical practice, sEMG using a vaginal probe is one of the most commonly used clinical tools for assessing and treating individuals with PFM dysfunctions. Importantly, the correlation between sEMG and PFDs has been validated. 6 , 7 , 8
The standardized Glazer protocol system, employing sEMG signal acquisition and processing technologies, has been widely adopted in clinical practice. As such, approximately 11% of published studies utilize sEMG using the Glazer protocol. 9 For sEMG assessment, several test-retest studies showed moderate to excellent reliability of rest, maximal voluntary contraction (MVC), and endurance in healthy and PFD patients. 10 , 11 Moreover, investigations were conducted to address the reliability of the Glazer assessment. 12 The reference ranges for the sEMG parameters in the Glazer protocol were initially formulated with limited participant pools and restricted dimensions during the 1990s, 13 making it challenging to diagnose PFDs in various populations and conditions precisely. Therefore, a more accurate and reasonable reference range of pelvic floor sEMG parameters is needed for standardized assessment.
New technological breakthroughs have emerged in healthcare of the urogynecological field with the rapid advancements in artificial intelligence (AI), including machine learning (ML) and deep learning (DL). 14 , 15 , 16 These algorithmic systems are reshaping clinical decision-making paradigms by diagnosis refinement, patient-specific therapeutic strategy formulation, real-time treatment response evaluation, and prognostic prediction. Currently, there are models based on ML and DL for distinguishing electromyographic signals between normal individuals and patients with neurological disorders. 17 , 18 According to previous research, no research has been found using AI methods to derive reference ranges for pelvic floor sEMG parameters and intelligent diagnosis of PFDs.
The large-scale collection of pelvic floor sEMG data is a solid foundation for data analysis in diagnosing PFDs. This study aimed to construct a multidimensional sEMG database to derive more reasonable reference ranges for Glazer protocol parameters and a high-precision model for accurate diagnosis of PFDs by AI.
Coi Statement
None of the authors have conflicts of interest to disclose.
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.