{"paper_id":"416769e3-308e-44b4-b0ea-154c0270629c","body_text":"Received June 23, 2020, accepted July 7, 2020, date of publication July 10, 2020, date of current version July 29, 2020.\nDigital Object Identifier 10.1 109/ACCESS.2020.3008473\nAI-Enabled Diagnosis of Spontaneous Rupture\nof Ovarian Endometriomas: A PSO\nEnhanced Random Forest Approach\nMINGYAN ZHOU\n1,*, FENG LIN 2,*, QIAN HU\n 1,\nZHENZHOU TANG\n 1, (Senior Member, IEEE), AND CHU JIN 3\n1College of Computer Science and Artiﬁcial Intelligence, Wenzhou University, Wenzhou 325035, China\n2Department of Gynecology, The First Afﬁliated Hospital of Wenzhou Medical University, Wenzhou 325015, China\n3Renji College, Wenzhou Medical University, Wenzhou 325015, China\nCorresponding authors: Zhenzhou Tang (mr.tangzz@gmail.com) and Chu Jin (634137089@qq.com)\n∗M. Zhou and F. Lin contributed equally to this work.\nThis work was supported in part by the Zhejiang Provincial Natural Science Foundation of China under Grant LZ20F010008 and Grant\nLQ18H280007, in part by the Natural Science Foundation of China under Grant 81803781, and in part by the Fundamental Scientiﬁc\nResearch Project of Wenzhou City under Grant G20180008 and Grant G20190021\n.\nABSTRACT Spontaneous rupture of ovarian endometriomas (OEs) may cause serious injury to patients.\nHowever, in traditional clinical diagnosis, it is vulnerable to be ignored or misdiagnosed for the symptoms of\nthe acute abdomen caused by it are quite similar to those of some common gynecological emergencies, which\nleads to serious complications to patients. In view of this, this study investigates the AI-enabled, early and\naccurate diagnosis of spontaneous rupture of OEs. Although artiﬁcial intelligence (AI) has been proved to be\na powerful tool to help to make accurate clinical diagnosis, however, as far as we know, there is so far no report\non AI-enabled diagnosis of spontaneous rupture of OEs yet. Speciﬁcally, this study proposes a particle swarm\noptimization (PSO) enhanced random forest (RF) classiﬁcation model, called PSO-RF, to make diagnosis,\nwhere RF is used to rank feature importance and make diagnosis considering that an OE is ruptured or not\nis a typical 0-1 classiﬁcation problem, and PSO is leveraged to ﬁne-tune the essential parameters of RF. The\nperformance of the proposed PSO-RF model is evaluated with practical data collected from a local hospital\nand fully benchmarked by comparing with eight other machine learning models whose key parameters are\nsufﬁciently optimized as well by grid search or PSO, for the sake of fairness. The experiment results show that\nthe proposed PSO-RF model outperforms all the other models, with the accuracy of 97.47%, the area under\nthe ROC curve (AUC) of 0.996, the sensitivity of 94.12% and the speciﬁcity of 98.39%. It can be concluded\nthat the PSO-RF model is a highly effective AI-enabled tool for preoperatively diagnosing spontaneous\nrupture of OEs.\nINDEX TERMS Ovarian endometrioma, spontaneous rupture, random forest, particle swarm optimization,\nmachine learning.\nI. INTRODUCTION\nOvarian endometrioma is a kind of endometriosis, patients\nof which have ectopic endometrial tissues in ovary [1] and\nsuffer from various clinical symptoms including pelvic pain,\nirregular menstrual bleeding, dysmenorrheal, dyspareunia\nand infertility [2]. It is one of the most common diseases in\nreproductive women aging from 25 to 45 years old, with a\nhigh incidence of 10%-15% [3]. Although ovarian endometri-\noma is a common benign disease, the size of the cyst typically\nincreases gradually, and with a small probability, spontaneous\nThe associate editor coordinating the review of this manuscript and\napproving it for publication was Derek Abbott\n.\nrupture of cyst may occur during or after menstruation [4],\n[5]. If the rupture can not be diagnosed and treated correctly\nin time, the severe cases may cause massive hemorrhage\nin the abdominal cavity, leading to hemorrhagic shock [6].\nHowever, the symptoms of the acute abdomen caused by\nspontaneous rupture of ovarian endometrioma are similar\nto those of ectopic pregnancy, rupture and hemorrhage of\ncorpus luteum, torsion of ovarian cyst pedicle and appendici-\ntis [7], [8]. Consequently, it is vulnerable to be ignored or\nmisdiagnosed clinically which causes serious complications\nto patients. Early and accurate diagnosis the spontaneous\nrupture of ovarian endometrioma could signiﬁcantly help to\navoid serious injury.\nVOLUME 8, 2020 This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ 132253\n\nM. Zhou et al.: AI-Enabled Diagnosis of Spontaneous Rupture of Ovarian Endometriomas: A PSO Enhanced Random Forest Approach\nConsidering the high misdiagnosis of spontaneous rupture\nof ovarian endometrimas, researches on it has been widely\ncarried out. Some of them have tried to ﬁnd suitable biomark-\ners to assist diagnosis of rupture of ovarian endometrimas\npreoperatively and have reported that spontaneous rupture\nof ovarian endometrimas are typically accompanied by high\nlevels of serum CA125 and CA19-9 [1], [7], [9]–[12]. In\nparticular, Xinyu Dai et al. have evaluated the clinical sig-\nniﬁcance of serum CA125 and CA19-9 in the patients with\nspontaneous rupture of ovarian endometromas by statisti-\ncal methods [2]. The results have shown that the levels of\nCA125 and CA19-9 in patients with spontaneous rupture\nof ovarian endometromas increased signiﬁcantly. Currently,\nit is believed that the combined biomarkers of CA125 and\nCA19-9 is helpful for the early diagnosis of spontaneous\nrupture of ovarian endometriomas.\nIn the other hand, artiﬁcial intelligence (AI), especially\nmachine learning (ML), has been widely used in medi-\ncal diagnosis [13]. The innovative AI-enabled methods are\nimportant assistances in precision medicine, which may help\nto make accurate clinical diagnosis at low cost and early.\nIt has been frequently reported recently that the innova-\ntive AI-enabled approaches are important tools in precision\nmedicine, which may help to make accurate clinical diagnosis\nat low cost and early.\nBy leveraging various machine learning methods,\nKawakami et al. have developed an ovarian cancer-speciﬁc\npredictive framework for clinical stage, histotype, residual\ntumor burden, and prognosis based on multiple biomarkers\n[14]. The results have revealed that AI-enabled models can\nprovide accurate diagnoses and prognostic predictions for\npatients with epithelial ovarian cancer prior to initial inter-\nvention. In [15], an optimized random forest (RF) classiﬁer\nhas been used to detect nodules inside the lungs. In [16],\nY . Shi et. al have proposed a hybrid QPSO-RF (Quantum\nparticle swarm optimization - RF) model to predict the\ndisease progression of Idiopathic Pulmonary Fibrosis using\nComputed Tomography (CT). Speciﬁcally, QPSO algorithm\nhas been used to search the optimal feature subsets and RF has\nbeen used for classiﬁcation. In [17], four machine learning\nalgorithms, namely support vector machine (SVM), partial\nleast squares discriminant analysis (PLS-DA), RF, and logis-\ntic regression (LR), have been applied to develop the model\nfor identifying metabolomic biomarkers in cervicovaginal\nﬂuid for EC detection. Pergialiotis et al. have investigated\nthe diagnostic accuracy of three different machine learning\nalgorithms, i.e. LR, artiﬁcial neural networks and classiﬁ-\ncation and regression trees (CARTs) for the prediction of\nendometrial cancer in postmenopausal women with vagi-\nnal bleeding or endometrial thickness ≥ 5 mm, as deter-\nmined by ultrasound examination [18]. E. Pashaei et al.\nhave proposed a binary version of black hole algorithm for\nsolving feature selection problem in biological data [19].\nA novel one-dimensional deep densely connected neural\nnetwork (DDNN) has been constructed in [20] to detect\natrial ﬁbrillation in 12-lead electrocardiogram waveforms.\nThe proposed model has been proved to be sufﬁciently\neffective to be applied to clinical diagnosis of atrial ﬁbril-\nlation. In the study of Obrzut et al., the usefulness of six\nAI-enabled methods have been evaluated for 5-year overall\nsurvival prediction in patients with cervical cancer treated\nby radical hysterectomy [21]. It has demonstrated that the\nbest model, namely the probabilistic neural network model,\nachieves a high prediction accuracy of 0.892 and sensitivity\nof 0.975.\nMachine learning models have also been used in the\ndiagnosis of endometrial tumor related diseases. In [22], deci-\nsion trees have been used to analyze the effectiveness of treat-\nment of patients with recurrent pelvic cyst who underwent\nsurgical intervention. The study of Günakan has constructed\na naive Bayes model to make predictions of lymph node\ninvolvement in endometrial cancer [23]. In [24], automated\nimage analysis and RF models have been used to classify\nnormal, premalignant, and malignant endometrial tissues.\nHowever, as far as we known, no study on ML-enabled\ndiagnostics of spontaneous rupture of ovarian endometriomas\nhas been reported so far.\nInspired by the above observations, this paper proposes\na machine learning model to make accurate diagnosis of\nspontaneous rupture of ovarian endometriomas preopera-\ntively. Speciﬁcally, we ﬁrst model the problem of cyst rup-\nture as a 0-1 classiﬁcation problem. Practical physiological\ndata involving plenty of features are collected from women\nwith ovarian endometriomas who have treated in the First\nAfﬁliated Hospital of Wenzhou Medical University between\n2006 and 2017. Considering the effectiveness of random for-\nest algorithm [25] in feature selection and classiﬁcation and\ntherefore it has been widely used to make intelligent medical\ndiagnoses [15]–[17], [26], [27], we design a random forest\nmodel to rank the importance of different features and further\nsolve the 0-1 classiﬁcation problem. In order to improve the\nperformance of random forest model, we further leverage\nparticle swarm optimization (PSO) algorithm [28] to optimize\nthe parameters of the random forest model. To the best of our\nknowledge, this is the ﬁrst time to adopt the machine learning\ntechnique in the diagnosis of spontaneous rupture of ovarian\nendometriomas.\nThe proposed PSO enhanced RF model, PSO-RF for\nshort, is comprehensively benchmarked by comparing\nwith other ﬁne-tuned machine learning models, including\nthe gridsearch optimized random forest model (GO-RF),\nthe naive Bayesian classiﬁcation model [29] (NBC), the grid-\nsearch optimized k-nearest neighbor [30] model (GO-KNN),\nthe PSO enhanced KNN (PSO-KNN), the gridsearch opti-\nmized lightGBM [31] model (GO-lightGBM), the PSO\nenhanced lightGBM model (PSO-lightGBM), the gridsearch\noptimized logical regression [32] model (GO-LR), and the\nPSO enhanced logical regression model (PSO-LR). The\nresults show that the PSO-RF model outperforms all the other\nmodels, with the accuracy of 95.57%, the area under the\nROC curve (AUC) of 0.996, the sensitivity of 94.12% and\nthe speciﬁcity of 98.39%.\n132254 VOLUME 8, 2020\n\nM. Zhou et al.: AI-Enabled Diagnosis of Spontaneous Rupture of Ovarian Endometriomas: A PSO Enhanced Random Forest Approach\nTABLE 1. Baseline characteristics of endometriosis patients.\nThe rest of this paper is organized as follows. The\nmethodology of the proposed PSO-RF is given in detail in\nSection II. Section III presents the experimental veriﬁcation\nand analysis. Finally, Section IV concludes the paper and\nproposes the future work.\nII. METHODS\nA. DATA ACQUISITION AND PREPROCESSING\nPhysiological data were collected by complete blood counts,\nlaparoscopic surgery and laparotomy from premenopausal\nfemale patients with ovarian endometriomas. The dataset\nincludes 193 records where 53 patients have been diag-\nnosed with spontaneous ruptured ovarian endometriomas\nand 140 patients with unruptured ovarian endometriomas.\nEach sample in the dataset includes the age, the leukocyte\ncount and the physiological features obtained by laparo-\nscopic surgery or laparotomy, such as the carcinoembryonic\nantigen (CEA) level, the CA125 level, the alpha fetopro-\ntein (AFP) level, the CA19-9 level, the position of the ovarian\nendometrioma, the endometrioma stage 1 and the size of the\novarian endometrioma, as shown in Table1. All the data have\nbeen normalized by x∗ = x−min\nmax − min .\nThe dataset has been divided into the training set and the\ntest set by repeated random sampling until the P values with\nrespect to all the features between the two sets are greater\nthan 0.2. This results in a division of 114 patients in the\ntraining set (32 patients with ruptured ovarian endometriomas\nand 82 patients with unruptured ovarian endometriomas in\nthe training set) and 79 patients in the test set (21 patients\nwith ruptured ovarian endometriomas and 58 patients with\nunruptured ovarian endometriomas), as listed in Table 1.\n1These patients were classiﬁed into four stages according to the revised\nclassiﬁcation made by the American Society for Reproductive Medicine\n(ASRM).\nFIGURE 1. Structure of a RF model.\nB. RANDOM FOREST MODEL\nRandom forest (RF) is an ensemble learning model for\nclassiﬁcation and regression. RF runs efﬁciently on large\ndata bases, seldom overﬁts, and has the ability of giving\nestimates of which features are important in the classiﬁcation.\nIn particular, RF is able to achieve a good accuracy rate for\nthe classiﬁcation of discrete high-dimensional datasets. With\nthis in mind, in this work, RF is leveraged to diagnose the\nspontaneous rupture of ovarian endometriomas.\nThe RF model consists of multiple classiﬁcation and\nregression trees (CARTs) [33]. Each CART is a binary deci-\nsion tree that is grown by recursively partitioning the data in\na parent node into child nodes. The structure of the RF model\nis shown in Fig. 1.\n1) SAMPLING AND SPLITTING\nThe overall samples in the training set are further divided\ninto two parts. Speciﬁcally, 60% of the samples are used\nto train the CARTs, the set of which is denoted as D =\n{s1, s2, . . . ,s|D|} where si, i = 1, 2, . . . ,|D| is the ith\nVOLUME 8, 2020 132255\n\nM. Zhou et al.: AI-Enabled Diagnosis of Spontaneous Rupture of Ovarian Endometriomas: A PSO Enhanced Random Forest Approach\nmulti-dimensional sample in D. The rest are the pretest sam-\nples used to measure the classiﬁcation accuracy rate of this\nCARTs, the set of which is denoted as ˆD = { ˆs1, ˆs2, . . . ,ˆs| ˆD|}\nwhere ˆsi, i = 1, 2, . . . ,| ˆD| is the ith multi-dimensional\nsample in ˆD.\nEach CART is grown with a bootstrap sample set which\nis taken by randomly sampling with replacement from D.\nConcretely, one sample is randomly drawn from D as a train-\ning sample at a time, and then, this sample is put back. This\nprocess of sampling with replacement is repeated m times so\nthat to generate a new sample set with the same size as D.\nThe new created sample set is then used as the root, e.g. the\ntraining set of a CART, denoted as Ri, i = 1, . . . ,NT , where\nNT is the number of CA TRs in RF. Apparently, in this way,\nsome samples may appear multiple times in the new training\nset, while others, namely about 36.8% 2 of the samples in the\noriginal dataset, may never appear. Therefore, the samples\nthat do not appear in the new dataset are taken as the test set.\nThe training set of each CART will be split as a binary\ntree. Let NF be the number of features of each sample. At\neach node of the decision tree, randomly select Nf , Nf ≪ NF\nfeatures for splitting. In the other words, at each node,\nthe splitting feature is selected among a random subset of the\noriginal set of features. This randomness further enhances the\ngeneralization of RF model.\nThe criteria of splitting in this work is the Gini index\nof impurity. Formally, denote the sample set at node t as\nSt = {s t,1, . . . ,st,|St |} where st,i = {f t,i,1, . . . ,ft,i,NF } is the\nith sample in node t and ft,i,j is the value of the jth feature in\nst,i. The class assignment of st,i is denoted as Ct,i ∈ {−1, 1}\nwhere -1 indicates an unruptured ovarian endometrioma and\n1 indicates a ruptured ovarian endometrioma. The Gini index\nat node t is given by\nGt =\n∑\nk=−1,1\npk|t (1 − pk|t ), (1)\nwhere pk|t is the probability of class k estimated from the\nsamples in node t, e.g.\npk|t = 1\n|St |\n∑\nst,i∈St\nCt,i = k. (2)\nAt each step in the algorithm, the node t is split at the\nfeature j and value ft,j into a pair of child nodes, denoted as\ntL = {x ∈ t|xj ≤ ft,j} and tR = {x ∈ t|xj > ft,j}, respectively,\nwhere j and ft,j are determined by\narg max\nj,ft,j\n1Gt |j,ft,j, (3)\nwhere\n1Gt |j,ft,j =\n(\nGt − pL|t GL|t − pR|t GR|t\n)⏐⏐j,ft,j , (4)\n2Assuming there are m samples in the original data set, the probability\nof each sample being selected is 1/m, and the probability of a sample\nnot being selected in all of the m sampling is (1 − 1/m)m. Moreover,\nlimm→∞ (1 − 1/m)m = e−1 ≈ 0.368\npL|t and pR|t are the probabilities of assigning a sample to the\nleft child and right child, respectively.\nThe process of splitting repeats until the terminal condition\nis satisﬁed, that is, 1) the depth of the CART reaches the\npredeﬁned maximum value dmax, or 2) |St | is less than the\npredeﬁned minimum value Ns,min. In this way, a CART\nis formed. The set of the CARTs in RF is denoted as\nT = {T 1, . . . ,TNT }.\n2) VOTING\nFinally, all the CARTs involved in a RF cast votes for their\ninputs, after which RF aggregates their results and determines\nthe output by majority voting of the CARTs. Taking in con-\nsideration the different classiﬁcation abilities of CARTs in RF\nmodel, in this work, the voting weight of each CART is set\nseparately according to its classiﬁcation accuracy. In essence,\nafter the training, the weight of each CART is estimated by\nthe pretest samples in ˆD as follows:\nwTi = N correct\ni\n⏐⏐⏐ ˆD\n⏐⏐⏐\n, i = 1, 2, . . . ,NT , (5)\nwhere N correct\ni is the number of samples classiﬁed correctly\nby Ti.\nThe voting strategy is to summarize the classiﬁcation\nresults of all the CARTs, and then take the weighted mode of\nthe classiﬁcations as the ﬁnal classiﬁcation results as follows:\nCsi = arg max\nk=−1,1\n∑\nTi∈T\nwTi ·\n(\nCsi | Ti = k\n)\n. (6)\nwhere Csi|Ti is the classiﬁcation result of feature si made by\nCART Ti.\nIn summary, the procedure of random forest model is\nshown in Algorithm 1.\n3) FEATURE IMPORTANCE\nOne of the most attractive advantages of RF is that it gives\nestimates of what features are important in the classiﬁcation.\nThere are many biomarkers and factors associated with the\nrupture of ovarian endometriomas among which some may\nhave very weak correlation with it. So, it will be of great\nsigniﬁcance to rank the importance of all the features to elim-\ninate irrelevant or weak correlated features, so as to reduce\nthe complexity of RF model and improve the accuracy of the\nmodel.\nRF uses Gini index to evaluate the importance of a feature.\nSpeciﬁcally, the importance of feature j to node t of Ti,\ndenoted as V im\nj,t,i, is evaluated as follows.\nV im\nj,t,i = max\nft,j\n1Gj,t,i|ft,j . (7)\nLet Mj,i be the set of nodes where feature j is included in Ti.\nThe importance of feature j to Ti can be calculated as\nV im\nj,i =\n∑\nt∈Mj,i\nV im\nj,t,i. (8)\nFurther, the importance of feature j to RF model is\nV im\nj =\n∑\nTi∈T\nV im\nj,i . (9)\n132256 VOLUME 8, 2020\n\nM. Zhou et al.: AI-Enabled Diagnosis of Spontaneous Rupture of Ovarian Endometriomas: A PSO Enhanced Random Forest Approach\nAlgorithm 1 Procedure of Random Forest Model\n1: Initialize NT , NF , Nf , dmax, Ns,min\n2: Preprocess the data set and obtain D and ˆD\n3: for i in [1, NT ] do\n4: for j in [1, |D|] do\n5: Randomly draw sj from D with replacement.\n6: Insert sj to Ri\n7: end for\n8: t ← Ri\n9: SPLIT(St , Nf )\n10: end for\n11: Aggregates the results by Eq. (6).\n12:\n13: function split(St , Nf )\n14: if Termination condition is satisﬁed then return\n15: else\n16: Randomly select Nf out of NF features at t.\n17: Determine the feature among Nf features and the\nvalue for splitting by Eq. (3) and Eq. (4).\n18: Split the tagged node into tL and tR.\n19: SPLIT(StL , Nf )\n20: SPLIT(StR, Nf )\n21: end if\n22: end function\nFinally, V im\nj is normalized as follows:\n˜V im\nj =\nV im\nj\n∑\nfi∈F V im\ni\n, (10)\nwhere F is the set of features.\nC. PARAMETER OPTIMIZATION BASED ON PSO\nGenerally, in RF, some parameters such as the maximum\ndepth of a CART (dmax), the minimum number of samples in a\nnode (Ns,min) and the number of CARTs (NT ), etc., need to be\ntune for achieving the optimal performance of RF. A feasible\nsolution is to ﬁne-tune these parameters by an appropriate\noptimization algorithm, such as various heuristic algorithms.\nIn this study, we leverage particle swarm optimization [34],\nwhich is a metaheuristic algorithm simulating the behavior of\nbirds catching food, to tune the essential parameters of the RF\nmodel to further improve the accuracy. The essence of PSO\nis to use the current position, global extremum and individual\nextremum information to guide the next iteration position of\nparticles, which enables it to approach the optimal solution\nwith a fast convergence speed, and hence effectively optimize\nthe parameters of the model [16], [35]. Another signiﬁcant\nadvantage of PSO is that it can adjust the maximum step size\nat each iteration, making it possible to ﬁnd an approximate\noptimal solution in a wide range of possible parameters\n[36]. In addition, PSO has some attractive features, such as\neasy to implement, less parameters, computationally cheap\nand so on.\nPSO searches the optimal solution through agents, or called\nparticles. A large number of particles form a swarm.\nA particle i is characterized by its location vector Li and\nvelocity vector vi which are updated iteratively by\nLi(t + 1) = Li(t) + vi(t + 1), (11)\nand\nvi(t + 1) = wivi(t) + c1r1\n(\nLi,best − Li(t)\n)\n+c2r2\n(\nLgbest − Li(t)\n)\n, (12)\nwhere wi is the selected weighting factor for particle i,\nLi,best is the location at which particle i previously had the\nbest ﬁtness measure, Lgbest is the global optimal location\nof the whole swarm, c1 and c2 are the cognitive acceler-\nation constant and the social acceleration constant, respec-\ntively, generally with the value of 2, and r1 and r2 are two\nrandom parameters within the range [0, 1]. The iteration\nterminates if 1) reach the maximum number of iterations,\nor 2) convergence is reached based on the ﬁtness measure.\nD. PSO-RF\nIn summary, the procedure of the proposed PSO-RF is as\nfollows:\nS1 Determine the parameters of PSO, including the\nmaximum number of iterations, the maximum speed of\nparticles and the search space, i.e., the feasible region of\nsolutions.\nS2 Randomly generate a swarm where each particle is a\nthree dimensional vector {dmax, Ns,min, NT }. The size of\nthe swarm is denoted as NPSO.\nS3 Start PSO iteration procedure to obtain the optimal dmax,\nNs,min and NT .\nS3.1 For each particle in the swarm, perform the RF\nalgorithm deﬁned in Algorithm 1 to evaluate the\nﬁtness, i.e., the diagnostic accuracy.\nS3.2 For each particle i, i ∈ {1, 2, . . . ,NPSO} in the\nswarm, compare the particle’s current ﬁtness with\nthe ﬁtness of Li,best . If the current location ﬁts\nbetter, then replace Li,best with the current location.\nS3.3 Update Lgbest if the new global optimal location of\nthe current swarm ﬁts better.\nS3.4 For each particle i, i ∈ {1, 2, . . . ,NPSO} in the\nswarm, update vi and Li by Eq.(12) and Eq.(11).\nS3.5 Repeat S3.1 - S3.4 until convergence is reached.\nS4 Perform the RF algorithm deﬁned in Algorithm 1 with\nthe optimal dmax, Ns,min and NT obtained in Step S3 to\nmake diagnosis.\nThe Pseudo code and the ﬂow diagram of the PSO-RF\nprocedure are shown in Algorithm 2 and Fig. 2, respectively.\nIII. RESULTS\nAll the involved algorithms and models have been imple-\nmented using Python 3.8. The computational analyses have\nbeen conducted on a Server running Windows Server\n2016 Standard 64 bits operating system with Intel (R) Xeon\n(R) CPU E5-2620 v4, 32GB of RAM and Nvidia GeForce\nGTX 1080 Ti graphics card. Unless otherwise stated,\nVOLUME 8, 2020 132257\n\nM. Zhou et al.: AI-Enabled Diagnosis of Spontaneous Rupture of Ovarian Endometriomas: A PSO Enhanced Random Forest Approach\nFIGURE 2. PSO-RF flowchart.\nAlgorithm 2 Procedure of PSO-RF\n1: Initialize the parameters for PSO.\n2: for each particle i, i ∈ {1, 2, . . . ,NPSO} do\n3: Randomly initiate vi and Li.\n4: end for\n5: repeat\n6: for each particle i, i ∈ {1, 2, . . . ,NPSO} do\n7: Perform Algorithm 1 to evaluate the ﬁtness.\n8: if ﬁtness(particle i) > ﬁtness(Li,best ) then\n9: Li,best ← Li\n10: end if\n11: if ﬁtness(particle i) > ﬁtness(Lgbest ) then\n12: Lgbest ← Li\n13: end if\n14: end for\n15: until Termination conditions are satisﬁed.\n16: Perform Algorithm 1 with Lgbest .\nthe parameter settings of all the involved algorithms and\nmodels are listed in Table 2.\nA. PERFORMANCE INDICATORS\n1) ACCURACY\nDiagnosis of spontaneous rupture of ovarian endometrioma\nis a typical binary classiﬁcation problem, in which the\ndiagnosis results are labeled either as positive (P), indi-\ncating a ruptured ovarian endometrioma, or negative (N),\nindicating an unruptured ovarian endometrioma. There are\nfour possible combinations of the diagnosis results and the\nactuality, namely TP, FP, TN and FN which are explained as\nfollows:\n• TP: both the diagnosis result and the actuality are P .\n• FP: the diagnosis result is p and the actuality is N.\n• TN: both the diagnosis result and the actuality are N.\n• FN: the diagnosis result is n and the actuality is P .\nTABLE 2. Parameters and settings of algorithms.\nThe classiﬁcation accuracy rate, or called the diagnostic\naccuracy rate, of the proposed PSO-RF model is\nevaluated by\nACC = TP + TN\nTP + TN + FP + FN (13)\n132258 VOLUME 8, 2020\n\nM. Zhou et al.: AI-Enabled Diagnosis of Spontaneous Rupture of Ovarian Endometriomas: A PSO Enhanced Random Forest Approach\n2) SENSITIVITY & SPECIFICITY\nBesides accuracy rate, sensitivity and speciﬁcity are also used\nto measure the effectiveness of PSO-RF model. Sensitivity,\nalso known as true-positive rate (TPR), is used to measure\nthe ability of correctly identifying the cases with ruptured\novarian endometriomas, whereas speciﬁcity, also known as\ntrue-negative rate (TNR), is used to measure the ability of the\nidentifying those without the disease. The higher the sensi-\ntivity or the speciﬁcity is, the more effective the model is.\nSensitivity (TPR) and speciﬁcity (TNR) can be calculated by\nTPR = TP\nTP + FN = TP\nP , (14)\nand\nTNR = TN\nFP + TN = TN\nN , (15)\nrespectively.\n3) ROC AND AUC\nMoreover, the receiver operating characteristic curve (ROC)\nis used to illustrate the diagnostic ability of the PSO-RF.\nROC is a comprehensive indicator that is capable of evaluat-\ning TPR and false-positive rate (FPR) at various thresholds.\nIt is a two-dimensional graph in which TPR is plotted on\nthe Y axis and FPR is plotted on the X axis. FPR can be\ncalculated by\nFPR = FP\nFP + TN = FP\nN = 1 − TNR. (16)\nThe AUC (area under the curve) is a concomitant of ROC\nobtained by calculating the area under the ROC. AUC tells\nhow much the model is capable of distinguishing between\nruptured and unruptured ovarian endometriomas. Higher the\nAUC, better the model is at predicting positives as positives\nand negatives as negatives.\nB. RESUL TS\n1) FEATURE SELECTION\nThe ranking of the feature importance for diagnosing the\nrupture of ovarian endometriomas derived by PSO-RF is\nshown in Fig. 3. The P values and the Pearson correlation\ncoefﬁcients (PCC) for the features related to the rupture of\nOE are listed in Table 3.\nIt can be observed the ranking obtained by PSO-RF\nis on the whole consistent with those obtained by sta-\ntistical analyses. The level of CA19-9 and CA125 are\nthe top two important biomarkers, with the normalized\nimportance of 0.316 and 0.273, respectively, the P values\nof 4.08 × 10−11 and 2.47 × 10−14, respectively, and the PPC\nvalues of 0.4522 and 0.5128, respectively. This conclusion is\nconsistent with the clinical experience. Serum CA19-9 and\nCA125, tumor associated antigens, have been used clini-\ncally to identify the recurrence and severity of endometriosis.\nThe endometrioma cyst ﬂuid was suspected to be rich of\nbiomarkers such as CA19-9 and CA125. However, in nor-\nmal conditions, the large CA19-9 and CA125 glycoprotein\nFIGURE 3. Ranking of the importance of features of ovarian\nendometriomas.\nTABLE 3. Association of ruptured/unruptured OEs with clinicopathologic\nfeatures.\nmolecules were preventing by the thick wall of endometrioma\ncyst from entering the peripheral circulation. The sponta-\nneous rupture of the OE may lead the biomarkers into the\nblood circulation, resulting the high elevation of the CA19-9\nand CA125 in the serum [11]. The level of CEA, with the\nnormalized importance of 0.231, P value of 7.35×10−11, and\nPCC value of -0.4468, is the third important biomarker. The\nleukocytes count and AFP level are also signiﬁcant biomark-\ners, however, are less important than CA19-9, CA125 and\nCEA.\nIn order to investigate the optimal combinations of features\nto be input into the PSO-RF model, the effectiveness of any\nsingle feature and various feature combinations on diagnosis\nof ruptured ovarian endometriomas is evaluated. Speciﬁcally,\nFig. 4 plots the ROC curves and the corresponding AUC\nvalues of the classiﬁcation results derived from PSO-RF\nusing a single feature. Not surprisingly, CA19-9, CA125 and\nCEA output the most effective results, with the values of\nAUC of 0.948, 0.902 and 0.819, respectively. In addition,\nalthough the importance of the Leukocyte level is a bit\nhigher than that of the AFP level, the AUC value of AFP\nlevel (0.806) is much greater than that of Leukocyte level\n(0.725). The ROC Curves, AUC values and accuracy rates\nof classiﬁcations using various features combinations derived\nfrom PSO-RF are depicted in Fig. 4(b). It can be seen that\nthe best performance is achieved in the case of inputting\nVOLUME 8, 2020 132259\n\nM. Zhou et al.: AI-Enabled Diagnosis of Spontaneous Rupture of Ovarian Endometriomas: A PSO Enhanced Random Forest Approach\nFIGURE 4. (a) ROCs Curves and values of AUC of all the features derived\nfrom PSO-RF for segregating ruptured ovarian endometriomas from\novarian endometriomas, (b) ROCs Curves, values of AUC and accuracy of\nvarious features combinations derived from PSO-RF for segregating\nruptured ovarian endometriomas from ovarian endometriomas.\nthe following four features: CA19-9 level, CA125 level,\nleukocyte count and AFP level, with the AUC of 0.996 and\nthe accuracy of 97.47%. The feature combination of CA19-9\nlevel, CA125 level, leukocyte count, AFP level and CEA level\nalso results in the same high accuracy of 97.47%, however,\nthe AUC value is a bit lower, with the value of 0.99.\nThe box and jitter plots representing the distributions of\nthe selected four features for PSO-RF-based classiﬁcation\nof ruptured and unruptured OEs are shown in Fig. 5. It can\nbe observed that the distributions are signiﬁcantly different\nbetween the ruptured and unruptured OEs. The levels of\nCA19-9, CA125, leukocyte count and AFP of the patients\nwith ruptured OEs are all signiﬁcantly higher than those of\npatients with unruptured OEs.\n2) COMPARISONS WITH OTHER ALGORITHMS\nThe proposed PSO-RF model is comprehensively\nbenchmarked by comparing with eight ﬁne-tuned machine\nlearning models, including four gridsearch optimized (GO)\nmodels, namely the GO-RF, GO-lightGBM, GO-LR\nand GO-KNN, three PSO enhanced models, namely\nFIGURE 5. Box and jitter plots representing distribution of the selected\nbiomarkers between the ruptured and unruptured ovarian\nendometriomas.\nPSO-lightGBM, PSO-LR and PSO-KNN, and the NBC\nmodel. The parameters obtained by grid search are shown\nin Table 2. NBC model has no parameter tuning for there is\nno parameter to be ﬁne-tuned in it.\nTable 4 summarizes the performance of all the nine\nmodels, including the classiﬁcation accuracy rates, the sen-\nsitivity and the speciﬁcity. Besides the performance results\nobtained from the dataset fragmentation shown in Table 1,\nTable. 4 also lists the accuracy rates obtained by hold-\nout cross-validations (40% holdout samples, 100 times) and\n10-fold cross-validations, respectively, to evaluate the per-\nformance more comprehensively. The distributions of the\ncross-validation results are presented as box plots, as shown\nin Fig. 6. For each test in the 40% holdout cross-validations,\nthe dataset is divided into a training set and a test set by\nrepeated random sampling until the P values with respect\nto all the four features between the two sets are greater\nthan 0.2 to ensure that there is no signiﬁcant correlation\nbetween the samples of the two sets. As to the 10-fold\ncross-validations, the P values are not guaranteed to be greater\nthan 0.2.\n132260 VOLUME 8, 2020\n\nM. Zhou et al.: AI-Enabled Diagnosis of Spontaneous Rupture of Ovarian Endometriomas: A PSO Enhanced Random Forest Approach\nTABLE 4. Performance comparisons among PSO-RF and benchmark models.\nFIGURE 6. Box plots of cross-validations.\nIt can be observed that the performances of PSO-RF in\nboth cross-validations are the highest. Speciﬁcally, for 40%\nholdout cross-validations, PSO-RF achieves the highest aver-\nage accuracy rate of 95.47%, with the minimum standard\ndeviation (S.D.) of 1.36%. The red box of PSO-RF in Fig. 6\nis the highest one, and also the narrowest one, with an\ninterquartile range (IQR) of 1.27% and a min-max range\nof 7.59%. PSO-RF also performs the best in the 10-fold\ncross-validations, with the highest average accuracy rate\nof 92.84%. The GO-RF model and the PSO-lightGBM model\nalso achieve high average accuracy rates, i.e., 93.58% and\n93.68% in 40% holdout cross-validations, and 91.26% and\n92.54% in 10-fold cross-validations, respectively.\nTable 4 also shows that PSO does improve the performance\nof these models. It slightly improves the accuracy rates of\nthe RF model, the lightGBM model, the LR model and the\nKNN model by 1.89%, 2.13%, 1.52% and 1.01%, respec-\ntively, in the 40% holdout cross-validations, and 1.58%,\n0.97%, 2.53% and 0.47%, respectively, in the 10-fold cross-\nvalidations. The reason why the performance improvements\nare limited is that the parameters of the original models\nhave also been optimized by grid search. However, the time\ncomplexity of PSO is O(NiterNPSO log NPSO) where Niter is\nthe number of iterations and NPSO is the size of the swarm.\nWhereas that of the gridsearch approach is O\n(∏\ni ni\n)\nwhere\nni, i = 1, 2, . . . ,j is number of potential values of feature i\nand j is the number of features to be ﬁne-tuned, which is much\nhigher than that of PSO.\nThe superiority of the PSO-RF model is also presented\nin Fig. 7 where the ROCs of these models are plotted. It can\nbe observed intuitively that the curve of PSO-RF is the closest\none to the upper left corner (AUC = 0.996), implying that it\noutperforms all the others for its highest the overall accuracy.\nThe ROCs of the PSO-lightGBM model and the GO-RF\nmodel are also quite closer to the upper left corner, with\nslightly lower AUC values, i.e., 0.996 for the PSO-lightGBM\nmodel and 0.995 for GO-RF model, respectively.\nThe confusion matrices of the PSO-RF model and the\nbenchmark models are shown in Fig. 8. PSO-RF performs\npretty well. There is only one false positive case and one\nfalse negative case, respectively. By contrast, one false nega-\ntive and two false positives are made by the GO-RF model\nand the PSO-lightGBM model. The PSO-LR model, the\nGO-lightGBM model and the NBC model recognize all the\npatients with unruptured ovarian endometriomas, however,\nmake three, four and ﬁve false positives, respectively. In\naddition, the PSO-KNN model makes two false positives and\ntwo false negatives, the GO-LR makes one false negative\nand four false positives, whereas the GO-KNN model makes\nfour false negatives and two false positives, all of which are\ninferior to the proposed PSO-RF model.\nIn addition to the classiﬁcation accuracy, the efﬁciency\nof these classiﬁers is also evaluated. For PSO enhanced\nmodels, the convergence performances are evaluated since\nthey train the models iteratively to approach the sub-optimal\nsolutions. Fig. 9 shows the comparisons of the convergence\nVOLUME 8, 2020 132261\n\nM. Zhou et al.: AI-Enabled Diagnosis of Spontaneous Rupture of Ovarian Endometriomas: A PSO Enhanced Random Forest Approach\nFIGURE 7. ROCs of PSO-RF and benchmark models.\nFIGURE 8. Confusion matrices of PSO-RF and benchmark models.\nperformance among PSO enhanced models. The ﬁtness\nof all the algorithm is the accuracy rate. It can be\nobserved that the accuracy rate of the PSO-RF model grad-\nually rises in the beginning (within 80 iterations). Then\nit converges to the stable value of 97.47%. The other\nthree PSO enhanced models, namely the PSO-LR model,\nthe PSO-KNN model and the PSO-lightGBM model con-\nverge after the 120th iteration, moreover, with lower accuracy\nrates.\nFig. 9 only presents the convergence performance along\nwith iterations. However, the time complexities per iteration\nof difference models vary greatly. In view of this, next,\nwe evaluate the per iteration time complexities of these mod-\nels. Considering that RF, lightGBM, KNN, NBC and LR are\nnot iterative algorithms, the CPU time required to train a\nmodel and then to make a classiﬁcation are measured instead\nof the convergence rate. The results averaged from 100 tests\nare shown in Fig. 9(b), where it can be observed that the CPU\ntime needed by RF is the longest. The reason is that the time\ncomplexity of RF is O(NT dmax · mn), that is, it is determined\nnot only by the number of samples (m) and the number of\nfeatures (n), but also by the number of CARTs (N T ) and the\ndepth of a CART (d max). As a contrast, the time complexity\nof LR is only O(mn); that of NBC is O(cmn) where c is the\nnumber of classes; that of KNN is O(kmn) where k is the\nnumber of neighbors; and the time complexity of lightGBM\nis co-determined by the procedures of gradient-based one-\nside sampling and exclusive feature bundling, with the time\ncomplexities of O(lmn) and O(lmnb), respectively, where l\nis the number of leaves and nb ≫ n is the number of\nbundles. Speciﬁcally, in this work, NT = 94, dmax = 2,\nc = 2, k = 7 and l = 9, so that NT · dmax ≫ k, l, 2, 1,\nresults in a much longer CPU time of RF than those of other\nmodels. Nevertheless, about 0.2 second CPU time training\nand classiﬁcation can fully satisfy the requirements of the\nclassiﬁcation system.\n132262 VOLUME 8, 2020\n\nM. Zhou et al.: AI-Enabled Diagnosis of Spontaneous Rupture of Ovarian Endometriomas: A PSO Enhanced Random Forest Approach\nFIGURE 9. Evaluations on convergence rates and CPU times.\nIV. CONCLUSIONS\nSpontaneous rupture of OE is vulnerable to be ignored or\nmisdiagnosed clinically, which may probably lead to serious\ncomplications to patients. In this work, we proposed a PSO\nenhanced RF model, namely PSO-RF, to assist in the preoper-\native diagnosis of spontaneous rupture of ovarian endometri-\nomas. As far as we know, it is the ﬁrst work on ML-enabled\ndiagnostics of spontaneous rupture of OEs. The data used in\nthis work were collected by complete blood counts, laparo-\nscopic surgery and laparotomy from premenopausal female\npatients with ovarian endometriomas who have treated in\nthe First Afﬁliated Hospital of Wenzhou Medical University.\nWe ﬁrst leveraged the RF model to rank the importance\nof the features and select the proper feature combination.\nThen, we employed PSO to ﬁne-tune the essential param-\neters of the RF model to further improve the diagnosis\naccuracy.\nThe PSO-RF model was comprehensively benchmarked\nby comparing with other machine learning models, including\nthree PSO enhanced models and ﬁve gridsearch-optimized\nmodels. The results showed that the PSO-RF model outper-\nforms all the other models in accuracy, with the accuracy\nof 97.47%, the AUC value of 0.996, the sensitivity of 94.12%\nand the speciﬁcity of 98.39%, due to the excellent classi-\nﬁcation ability of RF and parameter tuning based on PSO.\nAlthough the PSO-RF model costs the highest time complex-\nity, the overall complexity is totally acceptable practically.\nOur work concludes that the PSO-RF model is a highly\neffective AI-enabled tool for diagnose spontaneous rupture of\novarian endometriomas. It is fully qualiﬁed to assist surgeons\nin diagnosing whether a patient’s ovarian endometrioma is\nruptured preoperatively.\nV. ACKNOWLEDGMENT\n(Mingyan Zhou and Feng Lin contributed equally to\nthis work.)\nREFERENCES\n[1] A. Simonelli, R. Guadagni, P . De Franciscis, N. Colacurci, M. Pieri,\nP . Basilicata, P . Pedata, M. Lamberti, N. Sannolo, and N. Miraglia, ‘‘Envi-\nronmental and occupational exposure to bisphenol a and endometriosis:\nUrinary and peritoneal ﬂuid concentration levels,’’Int. Arch. Occupational\nEnviron. Health, vol. 90, no. 1, pp. 49–61, Jan. 2017.\n[2] X. Dai, C. Jin, Y . Hu, Q. Zhang, X. Y an, F. Zhu, and F. Lin, ‘‘High CA-\n125 and CA19-9 levels in spontaneous ruptured ovarian endometriomas,’’\nClinica Chim. Acta, vol. 450, pp. 362–365, Oct. 2015.\n[3] T. A. Gelbaya and L. G. Nardo, ‘‘Evidence-based management of\nendometrioma,’’ Reproductive Biomed. Online, vol. 23, no. 1, pp. 15–24,\nJul. 2011.\n[4] J. H. Pratt and W. R. Shamblin, ‘‘Spontaneous rupture of endome-\ntrial cysts of the ovary presenting as an acute abdominal emer-\ngency,’’ Amer . J. Obstetrics Gynecol., vol. 108, no. 1, pp. 56–62,\nSep. 1970.\n[5] K. Tanaka, Y . Kobayashi, K. Dozono, H. Shibuya, Y . Nishigaya,\nM. Momomura, H. Matsumoto, and M. Iwashita, ‘‘Elevation of\nplasma D-dimer levels associated with rupture of ovarian endometriotic\ncysts,’’ Taiwanese J. Obstetrics Gynecol., vol. 54, no. 3, pp. 294–296,\nJun. 2015.\n[6] M. Králíčková, A. S. Laganà, F. Ghezzi, and V . V etvicka, ‘‘Endometriosis\nand risk of ovarian cancer: What do we know?’’Arch. Gynecol. Obstetrics,\nvol. 301, pp. 1–10, Nov. 2019.\n[7] V . F. D. Amaral, R. A. Ferriani, M. F. S. D. Sá, A. A. Nogueira,\nJ. C. R. E. Silva, A. C. J. D. S. R. E. Silva, and M. D. D. Moura,\n‘‘Positive correlation between serum and peritoneal ﬂuid CA-125 levels\nin women with pelvic endometriosis,’’ Sao Paulo Med. J., vol. 124, no. 4,\npp. 223–227, 2006.\n[8] F. Zullo, E. Spagnolo, G. Saccone, M. Acunzo, S. Xodo, M. Ceccaroni, and\nV . Berghella, ‘‘Endometriosis and obstetrics complications: A systematic\nreview and meta-analysis,’’Fertility Sterility, vol. 108, no. 4, pp. 667–672,\nOct. 2017.\n[9] D. B. Manders, S. C. Purinton, J. S. Lea, D. L. Richardson, D. S. Miller, and\nS. M. Kehoe, ‘‘Predictive value of carcinoembryonic antigen and cancer\nantigen 125 in identifying mucinous ovarian tumors,’’ Gynecol. Oncol.,\nvol. 123, no. 2, pp. 445–446, Nov. 2011.\n[10] T. Ren, S. Wang, J. Sun, J.-M. Qu, Y . Xiang, K. Shen, and J. H. Lang,\n‘‘Endometriosis is the independent prognostic factor for survival in Chi-\nnese patients with epithelial ovarian carcinoma,’’ J. Ovarian Res., vol. 10,\nno. 1, p. 67, Dec. 2017.\n[11] B.-J. Park, T.-E. Kim, and Y .-W. Kim, ‘‘Massive peritoneal ﬂuid and\nmarkedly elevated serum CA125 and CA19-9 levels associated with an\novarian endometrioma,’’ J. Obstetrics Gynaecol. Res., vol. 35, no. 5,\npp. 935–939, Oct. 2009.\n[12] P .-H. Wang, Y .-T. Huang, K.-K. Ng, H.-H. Chou, H.-Y . Lu, S.-H. Ng,\nand G. Lin, ‘‘Detecting recurrent ovarian cancer: Revisit the values of\nwhole-body CT and serum CA 125 levels,’’ Acta Radiol., vol. 60, no. 10,\npp. 1360–1366, Oct. 2019.\n[13] K. R. Foster, R. Koprowski, and J. D. Skufca, ‘‘Machine learning, medical\ndiagnosis, and biomedical engineering research-commentary,’’ Biomed.\nEng. OnLine, vol. 13, no. 1, p. 94, 2014.\n[14] E. Kawakami, J. Tabata, N. Y anaihara, T. Ishikawa, K. Koseki, Y . Iida,\nM. Saito, H. Komazaki, J. S. Shapiro, C. Goto, Y . Akiyama, R. Saito,\nM. Saito, H. Takano, K. Y amada, and A. Okamoto, ‘‘Application of arti-\nﬁcial intelligence for preoperative diagnostic and prognostic prediction in\nepithelial ovarian cancer based on blood biomarkers,’’ Clin. Cancer Res.,\nvol. 25, no. 10, pp. 3006–3015, May 2019.\n[15] M. P . Paing, K. Hamamoto, S. Tungjitkusolmun, S. Visitsattapongse,\nand C. Pintavirooj, ‘‘Automatic detection of pulmonary nodules using\nthree-dimensional chain coding and optimized random forest,’’ Appl. Sci.,\nvol. 10, no. 7, p. 2346, Mar. 2020.\nVOLUME 8, 2020 132263\n\nM. Zhou et al.: AI-Enabled Diagnosis of Spontaneous Rupture of Ovarian Endometriomas: A PSO Enhanced Random Forest Approach\n[16] Y . Shi, W. K. Wong, J. G. Goldin, M. S. Brown, and G. H. J. Kim,\n‘‘Prediction of progression in idiopathic pulmonary ﬁbrosis using CT\nscans at baseline: A quantum particle swarm optimization–random forest\napproach,’’ Artif. Intell. Med., vol. 100, Sep. 2019, Art. no. 101709.\n[17] S.-C. Cheng, K. Chen, C.-Y . Chiu, K.-Y . Lu, H.-Y . Lu, M.-H. Chiang,\nC.-K. Tsai, C.-J. Lo, M.-L. Cheng, T.-C. Chang, and G. Lin, ‘‘Metabolomic\nbiomarkers in cervicovaginal ﬂuid for detecting endometrial cancer\nthrough nuclear magnetic resonance spectroscopy,’’Metabolomics, vol. 15,\nno. 11, p. 146, Nov. 2019.\n[18] V . Pergialiotis, A. Pouliakis, C. Parthenis, V . Damaskou, C. Chrelias,\nN. Papantoniou, and I. Panayiotides, ‘‘The utility of artiﬁcial neural net-\nworks and classiﬁcation and regression trees for the prediction of endome-\ntrial cancer in postmenopausal women,’’ Public Health, vol. 164, pp. 1–6,\nNov. 2018.\n[19] E. Pashaei and N. Aydin, ‘‘Binary black hole algorithm for feature selec-\ntion and classiﬁcation on biological data,’’ Appl. Soft Comput., vol. 56,\npp. 94–106, Jul. 2017.\n[20] W. Cai, Y . Chen, J. Guo, B. Han, Y . Shi, L. Ji, J. Wang, G. Zhang, and\nJ. Luo, ‘‘Accurate detection of atrial ﬁbrillation from 12-lead ECG\nusing deep neural network,’’ Comput. Biol. Med., vol. 116, Jan. 2020,\nArt. no. 103378.\n[21] B. Obrzut, M. Kusy, A. Semczuk, M. Obrzut, and J. Kluska, ‘‘Prediction\nof 5–year overall survival in cervical cancer patients treated with radical\nhysterectomy using computational intelligence methods,’’ BMC Cancer,\nvol. 17, no. 1, p. 840, 2017.\n[22] Y .-F. Wang, M.-Y . Chang, R.-D. Chiang, L.-J. Hwang, C.-M. Lee, and\nY .-H. Wang, ‘‘Mining medical data: A case study of endometriosis,’’\nJ. Med. Syst., vol. 37, no. 2, p. 9899, Apr. 2013.\n[23] E. Günakan, S. Atan, A. N. Haberal, İ. A. Küçükyildiz, E. Gökçe, and\nA. Ayhan, ‘‘A novel prediction method for lymph node involvement in\nendometrial cancer: Machine learning,’’ Int. J. Gynecol. Cancer, vol. 29,\nno. 2, pp. 320–324, Feb. 2019.\n[24] M. J. Downing, D. J. Papke, S. Tyekucheva, and G. L. Mutter, ‘‘A new clas-\nsiﬁcation of benign, premalignant, and malignant endometrial tissues using\nmachine learning applied to 1413 candidate variables,’’ Int. J. Gynecol.\nPathol., vol. 39, no. 4, pp. 333–343, 2020.\n[25] L. Breiman, ‘‘Random forests,’’ Mach. Learn., vol. 45, no. 1, pp. 5–32,\n2001.\n[26] A. Subasi, E. Alickovic, and J. Kevric, ‘‘Diagnosis of chronic kidney\ndisease by using random forest,’’ in Proc. CMBEBIH, Singapore, 2017,\npp. 589–594.\n[27] E. Alickovic and A. Subasi, ‘‘Medical decision support system for diagno-\nsis of heart arrhythmia using DWT and random forests classiﬁer,’’ J. Med.\nSyst., vol. 40, no. 4, p. 108, Apr. 2016.\n[28] S. M. Raafat, A. M. Hasan, T. W. A. Khairi, and K. G. Ali, ‘‘Opti-\nmized performance of consensus algorithm in multi agent system\nusing PSO,’’ Al-Nahrain J. Eng. Sci., vol. 21, no. 2, pp. 292–299,\nApr. 2018.\n[29] M. A. P . Andrade and M. A. M. Ferreira, ‘‘Bayesian networks use in\nsimple maternity problems,’’ Appl. Math. Sci., vol. 8, pp. 6963–6967,\n2014.\n[30] P . Sala, A. Colatutto, D. Fabbro, L. Mariuzzi, S. Marzinotto,\nB. Toffoletto, A. R. Perosa, and G. Damante, ‘‘Immunoglobulin k\nlight chain deﬁciency: A rare, but probably underestimated, humoral\nimmune defect,’’ Eur . J. Med. Genet., vol. 59, no. 4, pp. 219–222,\nApr. 2016.\n[31] G. Ke, Q. Meng, T. Finley, T. Wang, W. Chen, W. Ma, Q. Y e, and T.-Y . Liu,\n‘‘LightGBM: A highly efﬁcient gradient boosting decision tree,’’ in Proc.\nAdv. Neural Inf. Process. Syst., 2017, pp. 3146–3154.\n[32] S. Dreiseitl and L. Ohno-Machado, ‘‘Logistic regression and artiﬁcial\nneural network classiﬁcation models: A methodology review,’’ J. Biomed.\nInformat., vol. 35, nos. 5–6, pp. 352–359, Oct. 2002.\n[33] L. Breiman, J. H. Friedman, R. A. Olshen, and C. J. Stone, Classiﬁcation\nRegression Trees. Boca Raton, FL, USA: CRC Press, 1984.\n[34] J. Kennedy and R. Eberhart, ‘‘Particle swarm optimization,’’ in Proc. Int.\nConf. Neural Netw., vol. 4, Nov./Dec. 1995, pp. 1942–1948.\n[35] Y . Zhang, S. Y u, R. Xie, J. Li, A. Leier, T. T. Marquez-Lago, T. Akutsu,\nA. I. Smith, Z. Ge, J. Wang, T. Lithgow, and J. Song, ‘‘PeNGaRoo, a\ncombined gradient boosting and ensemble learning framework for pre-\ndicting non-classical secreted proteins,’’ Bioinformatics, vol. 36, no. 3,\npp. 704–712, Aug. 2019.\n[36] B. Xue, M. Zhang, W. N. Browne, and X. Y ao, ‘‘A survey on evolutionary\ncomputation approaches to feature selection,’’ IEEE Trans. Evol. Comput.,\nvol. 20, no. 4, pp. 606–626, Aug. 2016.\nMINGYAN ZHOU is currently pursuing the\nbachelor’s degree with the College of Computer\nScience and Artiﬁcial Intelligence, Wenzhou Uni-\nversity, Wenzhou, China. She is currently work-\ning on machine learning applications in intelligent\nmedicine.\nFENG LIN received the B.S. degree in obstetrics\nand gynecology and the M.S. degree in gynecol-\nogy from Wenzhou Medical University, Zhejiang,\nChina, in 2003 and 2011, respectively. She is\ncurrently an Associated Physician and a Profes-\nsor with the First Afﬁliated Hospital of Wenzhou\nMedical University. She mainly specializes in\nmalignant tumors of gynecology and women\ninfertility, and is good at surgical techniques in\nlaparoscopy and hysteroscopy.\nQIAN HU received the B.S. degree in communi-\ncation engineering from the Nanjing Institute of\nPosts and Telecommunications, Nanjing, China,\nin 2001, and the M.S. degree in circuits and sys-\ntems from Zhejiang University, Hangzhou, China,\nin 2005. She is currently an Associated Professor\nwith the College of Computer Science and Artiﬁ-\ncial Intelligence, Wenzhou University, Wenzhou,\nChina. Her research interests are in the areas of\nmachine learning, intelligent systems, and wireless\nnetworks.\nZHENZHOU TANG (Senior Member, IEEE)\nreceived the B.S. degree in electronic engineering\nand the M.S. degree in communications and\ninformation system from Zhejiang University,\nin 2001 and 2004, respectively, and the Ph.D.\ndegree in communications and information sys-\ntem from the Dalian University of Technology,\nin 2015. In March 2004, he joined Wenzhou\nUniversity. He is currently a Professor with the\nCollege of Computer Science and Artiﬁcial Intel-\nligence, Wenzhou University, and the Director of the Intelligent Network\nInnovation Research Center, Wenzhou University. From February 2015 to\nAugust 2015, he was a Visiting Scholar with the Department of Electrical\nand Computer Engineering, Worcester Polytechnic Institute, Worcester, MA,\nUSA. From November 2015 to June 2018, he was a Postdoctoral Research\nFellow with Zhejiang University. Since September 2018, he has been a\njoint Ph.D. Supervisor with Chonnam National University, South Korea.\nHe is currently serving as an Associate Editor of IEEE ACCESS. His current\nresearch interests include machine learning, intelligent systems, and wireless\nnetworks.\nCHU JIN received the M.Sc. degree in information\nsecurity from the Royal Holloway and Bed-\nford New College, University of London, U.K.,\nin 2002. He is currently a Lecturer with the Renji\nCollege, Wenzhou Medical University, Wenzhou,\nChina. His research interests are in the areas of\ncomputer security and network security.\n132264 VOLUME 8, 2020","source_license":"CC0","license_restricted":false}