Preterm birth and maternal heart disease: A machine learning analysis using the Korean national health insurance database.

OA: gold CC-BY-4.0

Abstract

BackgroundMaternal heart disease is suspected to affect preterm birth (PTB); however, validated studies on the association between maternal heart disease and PTB are still limited. This study aimed to build a prediction model for PTB using machine learning analysis and nationwide population data, and to investigate the association between various maternal heart diseases and PTB.MethodsA population-based, retrospective cohort study was conducted using data obtained from the Korea National Health Insurance claims database, that included 174,926 primiparous women aged 25-40 years who delivered in 2017. The random forest variable importance was used to identify the major determinants of PTB and test its associations with maternal heart diseases, i.e., arrhythmia, ischemic heart disease (IHD), cardiomyopathy, congestive heart failure, and congenital heart disease first diagnosed before or during pregnancy.ResultsAmong the study population, 12,701 women had PTB, and 12,234 women had at least one heart disease. The areas under the receiver-operating-characteristic curves of the random forest with oversampling data were within 88.53 to 95.31. The accuracy range was 89.59 to 95.22. The most critical variables for PTB were socioeconomic status and age. The random forest variable importance indicated the strong associations of PTB with arrhythmia and IHD among the maternal heart diseases. Within the arrhythmia group, atrial fibrillation/flutter was the most significant risk factor for PTB based on the Shapley additive explanation value.ConclusionsCareful evaluation and management of maternal heart disease during pregnancy would help reduce PTB. Machine learning is an effective prediction model for PTB and the major predictors of PTB included maternal heart disease such as arrhythmia and IHD.
Full text 31,097 characters · extracted from pmc-nxml · 5 sections · click to expand

Intro

Approximately 15 million neonates are born prematurely (defined as live birth at < 37 0/7 weeks of gestation) worldwide, accounting for about 11% of global births [ 1 , 2 ]. The reported rate of preterm birth (PTB) has been increasing in many countries [ 1 , 2 ]. PTB is the most important cause of death in infants and children, accounting for approximately 18% of deaths in children under the age of five years [ 1 – 3 ]. Cost-effective interventions, particularly focused on controlling maternal risk factors, have been estimated to prevent as much as three quarters of mortality due to PTB [ 2 ]. Additionally, identifying maternal PTB risk factors could help us better understand the etiology of PTB. The number of pregnant women with underlying diseases such as hypertension, diabetes, and obesity increase with maternal aging [ 4 , 5 ]. This leads to an increased number of pregnant women with heart disease (i.e., ischemic heart disease, cardiomyopathy, or arrhythmia) [ 4 – 6 ]. Furthermore, an increasing number of women with congenital heart disease (CHD) are reaching the reproductive age [ 4 ]. Although most women with CHD can carry a pregnancy and deliver safely, there are still concerns [ 4 , 7 ]. Pregnancy complicated by maternal heart disease is associated with maternal and fetal morbidity and mortality [ 4 , 7 ]. In addition, both CHD and acquired heart disease are known to affect PTB [ 4 , 7 , 8 ]. In a study of 5,739 pregnant women with acquired heart disease and CHD enrolled in the Registry Of Pregnancy And Cardiac disease (ROPAC) from 2007 to 2018, the prevalence of PTB in mothers with heart disease has been reported to be about 16% [ 8 ]. Another German study reported a prevalence of PTB of 11.7% in 2,114 pregnant women with CHD [ 7 ]. Overall, it has been consistently reported that the prevalence of PTB is higher in pregnant women with heart disease than in the general population, but there are differences in the prevalence of PTB reported in each country [ 7 – 9 ]. Moreover, most of the reported studies are the results of developed countries in the West, and there are no studies targeting Asian populations yet. Hence, this study aimed to build a prediction model for PTB using machine learning analysis and nationwide population data, and to investigate the association between various maternal heart diseases and PTB.

Results

A total of 174,926 women who delivered in 2017 were included in the analysis and 12,701 (7.83%) had preterm birth (PTB 4) ( Table 1 ). Among the total study population, 12,234 women had at least one heart disease. Arrhythmia was the most common maternal heart disease, followed by IHD and congestive heart failure (total population incidence: 4.18%, 2.86%, and 0.48% respectively). Hypertension, the major underlying disease for heart disease, was found in 12.36% of study population. The incidence of hypertension, arrhythmia, IHD, cardiomyopathy, and congestive heart failure was significantly higher in women who had PTB than in those who gave birth at term ( Table 1 ). The prevalence of PTB in pregnant woman with heart disease is presented in S3 Table . The prevalence of PTB in pregnant women with cardiomyopathy was the highest at 16.0%, and the prevalence of PTB among all pregnant women with heart disease was higher than that among pregnant women without heart disease. Values are median (interquartile range) or n (%). CHD = congenital heart disease. Table 2 presents the areas under the receiver-operating characteristic curves (AUC) of the random forest. The AUC with oversampling data was 88.53–95.31. Its logistic regression counterparts were within the range 50.10–53.54. The performance measures of the random forest with oversampling data were far beyond those of a logistic regression. Oversampling is an approach that matches the sizes of two groups (participants with and without PTB) to train the machines to balance the two groups. Logistic regression requires an unrealistic assumption of ceteris paribus , i.e., “all the other variables staying constant,” which is not required in a random forest. Hence, the findings of the logistic regression are best considered supplementary. PTB 1—PTB with preterm premature rupture of membranes (PPROM) only; PTB 2—PTB with spontaneous preterm labor without PPROM; PTB 3—PTB 1 or PTB 2; PTB 4—PTB 3 or other indicated PTB due to maternal or fetal indications. AUC = area under the receiver operating characteristic curve; PTB = preterm birth The random forest variable importance for PTB is shown in Fig 1 . These values were the averages for PTB 1–4. Table 3 presents the variable importance of the prediction model for PTB 4. Among the 36 variables, major determinants of PTB were socioeconomic status (0.3377), age (0.2881), gestational diabetes (0.0391), anemia (0.0329), sepsis (0.0311), abnormal menstruation (0.0285), benzodiazepine use (0.0249), TCAs use (0.0221), progesterone use (0.0214), hypertension (0.0213), vaginitis (0.0211), hyperlipidemia (0.0186), pelvic inflammatory disease (0.0184), recurrent miscarriage or infertility (0.0162), arrhythmia (0.0146), hypnotic/sedative drugs (0.0124), and IHD (0.0107). The variable importance of the prediction model for PTB 1–3 is presented in S3 Table . It should be noted that the variable importance measures of the random forest for the oversampling data were very similar to those for the original data ( Table 3 and S4 Table ). Notably, the SHAP value in Fig 2 shows the sign and magnitude of the effect of major determinants on PTB. For instance, the presence of recurrent miscarriages/infertility was consistently associated with an increased risk of PTB. In contrast, though anemia had a significant effect on PTB ( Table 3 ), the direction of the effect was inconsistent ( Fig 2 ). PTB 1—PTB with preterm premature rupture of membranes (PPROM) only; PTB 2—PTB with spontaneous preterm labor without PPROM; PTB 3—PTB 1 or PTB 2; PTB 4—PTB 3 or other indicated PTB due to maternal or fetal indications. PTB = preterm birth; CHD = congenital heart disease. PTB 4 indicated PTB with preterm premature rupture of membranes or spontaneous preterm labor or other indicated PTB due to maternal or fetal indications. PTB = preterm birth; CHD = congenital heart disease. PTB 4 indicated PTB with preterm premature rupture of membranes or spontaneous preterm labor or other indicated PTB due to maternal or fetal indications. PTB = preterm birth; CHD = congenital heart disease. Among the maternal heart diseases, arrhythmia (ranked 15 th on variable importance) was the most significant determinant of PTB, followed by IHD (17 th ), congestive heart failure (21 st ), acyanotic CHD (26 th ), and cardiomyopathy (27 th ), in that order. Based on SHAP values, the presence of IHD, congestive heart failure, and cardiomyopathy was associated with an increased PTB risk ( Fig 2 and S5 Table ). Although the variable importance of IHD was lower than that of hypertension, the presence of IHD more consistently increased the risk of PTB than hypertension. On the other hand, the presence of arrhythmia affected both the increasing and decreasing risk of PTB according to the SHAP value. To further delineate the effect of arrhythmia on PTB, we analyzed the arrhythmia subgroups. The subgroups included in the analysis were SVT, AF/AFL, conduction disorder, WPW syndrome, VA, and SSS. The incidence of maternal conduction disorders and AF/AFL was higher in the PTB group than in the term birth group ( Table 4 ). Based on the SHAP values, AF/AFL and conduction disorders particularly increased the risk of PTB among arrhythmia subgroups ( Fig 3 and S6 Table ). AF = atrial fibrillation; AFL = atrial flutter; SVT = supraventricular tachycardia; WPW = Wolff-Parkinson-White syndrome; VA = ventricular arrhythmia; SSS = sick sinus syndrome. Values are n (%). AF = atrial fibrillation; AFL = atrial flutter; SVT = supraventricular tachycardia; WPW = Wolff-Parkinson-White syndrome; VA = ventricular arrhythmia; SSS = sick sinus syndrome.

Conclusions

Machine learning is an effective prediction model for PTB and the major predictors of PTB included maternal heart disease such as arrhythmia and IHD. We used the random forest and considered a large collection of 36 demographic, socioeconomic, obstetric and medical variables to record the highest AUC of 0.95 for the prediction of PTB. Careful evaluation and management of maternal heart disease during pregnancy would help reduce PTB. Further research is needed on this strategy.

Materials|Methods

This nationwide population-based cohort study included singleton primiparous women who had delivered in 2017. We restricted the inclusion criteria to primiparous women to adjust prior PTB. Women aged 25–40 years who delivered before 37 0/7 weeks of gestation were included in the study. Data were extracted from the Korea National Health Insurance Service claims database. The Korean National Health Insurance Service (NHIS) claims data covers almost all citizens of Korea (approximately 50 million) [ 10 ]. The Korean NHIS data includes diagnosis codes based on International Classification of Disease, Tenth Revision (ICD-10), demographic information on age, sex, income decile, residential area, etc., and information on medication prescriptions, tests, and procedures performed during outpatient visits or hospitalizations since 2002. For primiparous women who gave birth in 2017, all medical history from 2002, when the Korean NHIS data began to be established, to 2016, the year immediately before delivery, was investigated. A total of 174,926 women were included in the analysis. The study was approved by the Institutional Review Board (IRB) of the Korea University Anam Hospital on November 5, 2018 (no. 2018AN0365). The requirement for informed consent was waived due to the retrospective nature of the study. An explanation of each variable according to the International Classification of Disease, Tenth Revision (ICD-10) code is presented in S1 Table . The dependent variable was PTB (birth before 37 0/7 weeks of gestation) in 2017. Four categories of PTB were introduced according to the ICD-10 code: (1) PTB 1—PTB with preterm premature rupture of membranes (PPROM) only; (2) PTB 2—PTB with spontaneous preterm labor without PPROM; (3) PTB 3—PTB 1 or PTB 2; (4) PTB 4—PTB 3 or other indicated PTB due to maternal or fetal indications. Thirty-six independent variables covered the following information: (1) demographic/socioeconomic determinants in 2017 including age and socioeconomic status measured by an insurance fee with a range of 0 (the lowest group) to 20 (the highest group); (2) obstetric and gynecologic diseases in 2002–2016, namely, gestational diabetes, hypertensive disorders during pregnancy (HDP; including, gestational hypertension, preeclampsia and eclampsia), pelvic inflammatory disease, vaginitis, endometriosis, pelvic organ prolapse, abnormal menstruation, recurrent miscarriage or infertility; (3) heart diseases in 2002–2016, including, CHD (acyanotic CHD, cyanotic CHD, severe lesion, shunt lesion, left or right side lesion, other lesion), arrhythmias (including conduction disorder, Wolff-Parkinson-White [WPW] syndrome, supraventricular tachycardia [SVT], atrial fibrillation/flutter [AF/AFL], ventricular arrhythmia [VA], and sick sinus syndrome [SSS]), cardiomyopathy, congestive heart failure, and ischemic heart disease (IHD); (4) other significant medical histories, including hypertension, diabetes, anemia, hyperlipidemia, pulmonary embolism, endocarditis, sepsis, stroke and cardiac arrest; (5) medication history in 2002–2016, particularly, benzodiazepine, calcium channel blocker (CCB), nitrate, progesterone, hypnotic/sedative drug (antihistamine, zolpidem, eszopiclone, pentobarbital sodium, and benzodiazepine derivates), and tricyclic antidepressant (TCA). These variables were selected based on previous studies and available data [ 11 – 13 ]. These data on disease and medication history were screened using ICD-10 and Anatomical Therapeutic Chemical (ATC) codes, respectively ( S1 and S2 Tables ). Logistic regression and random forest analyses were used to predict PTB [ 11 – 13 ]. A random forest is a group of decision trees that makes decisions on the dependent variable with a majority vote. A random forest with 100 decision trees was employed in this study: 100 training sets were sampled with replacements, 100 decision trees were trained with the training sets, 100 decision trees made 100 predictions, and the random forest took a majority vote on the dependent variable. The data of all the included observations were split into training and validation sets in an 80:20 ratio (139,940 vs. 34,986 cases). The validation criterion of the trained models was accuracy, which is the ratio of correct predictions among the 34,986 cases. A random forest variable importance was introduced to identify the major determinants of PTB and to test its association with 36 variables. The random forest variable importance of a certain variable (e.g., arrhythmia) can be defined as “the decrease of node impurity (GINI) in case a new branch is created based on the predictor in an average decision tree in the random forest”. Let’s assume that the random forest variable importance of arrhythmia for PTB is 0.0146. This indicates that node impurity (GINI) decreases by 0.0146 in case a new branch is created based on arrhythmia in an average decision tree in the random forest. The performance of the random forest increases as node impurity (GINI) decreases. In this context, the random forest variable importance of arrhythmia measures the contribution of arrhythmia for the performance of the random forest. A variable with the ranking of 18th or higher can be considered to be a major determinant in this study, given that it is a top 50% among 36 variables here. Furthermore, we calculated the Shapley additive explanation (SHAP) values to identify the direction of association between maternal heart disease and PTB in the prediction model. Here, the SHAP value of maternal heart disease measured the difference between the model’s predicted probability of PTB for each participant with and without maternal heart disease. Let’s assume that the SHAP value of atrial fibrillation for PTB is 0.1576. This indicates that the probability of PTB (predicted by the random forest) increases by 0.1576 in case the variable atrial fibrillation is added to the random forest. The SHAP value of atrial fibrillation can be considered to be an equivalence of machine learning to the odds ratio of logistic regression. For the arrhythmia group, which showed an even distribution for the increase or decrease in the risk of PTB in the overall SHAP value analysis, it was assumed that each disease within the category of arrhythmia would have a significantly different effect or mechanism on pregnant women chronically, and a subgroup analysis of arrhythmias was performed. Python (CreateSpace: Scotts Valley, 2009) was employed for the analysis from December 15, 2021 to April 15, 2022. It needs to be noted that in practice experts in artificial intelligence use random forest variable importance to derive the rankings and values of all predictors for the prediction of the dependent variable. Then, they employ the SHAP plots to evaluate the directions of associations between the predictors and the dependent variable. Linear or logistic regression used to play this role before the SHAP approach took it over. This is because the SHAP approach has a notable strength compared to linear or logistic regression: the former considers all realistic scenarios, un-like the latter. Let us assume that there are three predictors of PTB, i.e., socioeconomic status, age and maternal heart disease. As defined above, the SHAP value of maternal heart disease for PTB for a particular participant is the difference between what machine learning predicts for the prob-ability of PTB with and without maternal heart disease for the participant. Here, the SHAP value for the participant is the average of the following four scenarios for the participant: (1) socioeconomic status excluded, age excluded; (2) socioeconomic status excluded, age included; (3) socio-economic status included, age excluded; and (4) socioeconomic status included, age included. In other words, the SHAP value combines the results of all possible sub-group analyses, which are ignored in linear or logistic regression with an unrealistic assumption of ceteris paribus, i.e., “all the other variables staying constant”.

Supplementary Material

(DOCX) Click here for additional data file. (DOCX) Click here for additional data file. (DOCX) Click here for additional data file. (DOCX) Click here for additional data file. (DOCX) Click here for additional data file. (DOCX) Click here for additional data file.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-08-16T09:21:09.727480+00:00
unpaywall
last seen: 2026-05-21T05:10:58.409756+00:00
License: CC-BY-4.0