{"paper_id":"82beb4df-789f-4570-957c-11e62a9f2ac3","body_text":"Gestational diabetes mellitus (GDM) is a common pregnancy complication, defined as glucose intolerance with onset or first recognition during pregnancy, in women without prior diabetes history prior to pregnancy.[ 1 ,  2 ] During the last 20 years the prevalence of GDM has increased worldwide and it is expected to continue to rise along with the increase in pre-conception obesity and pregnant women affected by obesity.[ 3 ] GDM affects approximately 15% of all pregnancies, depending on population characteristics, and this prevalence may in fact be higher under the new diagnostic criteria.[ 4 ,  5 ] GDM is associated with an increased risk of maternal and infant morbidity, including macrosomia, large for gestational age (LGA), cesarean section delivery and preterm birth, but it is also considered to be a risk factor for long-term complications, such as type 2 diabetes mellitus and cardiovascular disease in the mother and the offspring.[ 6 – 9 ] The etiology of GDM is multifactorial and has not completely been established yet, while several risk factors may contribute to its onset. Age, overweight or obesity, ethnicity, family history of diabetes, and history of GDM are some of the proposed risk factors for GDM.[ 10 – 13 ]\nMeta-analyses of randomized clinical trials for GDM prevention that evaluated a range of dietary and lifestyle interventions during pregnancy, including diet and exercise, lifestyle advice, nutritional manipulation, and behavior modification, showed inconsistent findings, with some meta-analyses reporting significant deceased incidence of GDM [ 14 – 18 ], while others were null. [ 19 – 24 ]\nUnder the prism of the abundance of observational significant associations, we conducted an umbrella review of meta-analyses on risk factors for GDM. Using a standardized approach, we aimed to assess the credibility of those findings to identify which associations are with robust epidemiological evidence.\n\nThis study was performed according to the guidelines for systematic reviews under the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA).[ 25 ]\nWe conducted an umbrella review, which is a systematic collection and evaluation of multiple systematic reviews and meta-analyses performed on a specific research topic.[ 26 ] An umbrella review examines comparisons of a large number of existing systematic reviews and meta-analyses on risk factors into one accessible and usable document.[ 26 ,  27 ] The methods of performing an umbrella review are standardized and, in this work, we followed the same principles used in previously published umbrella reviews across various fields of research.[ 28 – 31 ] We used a ranking system to grade the evidence from meta-analyses of observational studies in terms of the significance of the summary effect, 95% prediction interval, presence of large heterogeneity, small study effects, and excess significance bias.\nTwo researchers (KG and SP) independently searched PubMed and ISI Web of Science from inception to 23 of December 2018 to identify meta-analyses of observational studies examining associations regarding risk factors for GDM. The search strategy used the keywords (“gestational diabetes” OR “pregnancy diabetes” OR “pregnancy hyperglycemia” OR “3 h abnormal gtt test” OR “insulin during pregnancy” OR “antidiabetics during pregnancy” OR “metformin in pregnancy”) AND (“systematic review” OR “meta-analysis”). All identified publications went through a three-step parallel review of title, abstract, and full text, performed by KG and SP, based on predefined inclusion and exclusion criteria. We also screened the references of the retrieved articles for possible eligible papers. Any disagreement was resolved with discussion.\nWe included meta-analyses of observational studies (i.e., cross-sectional, case-control and cohort studies), which investigated risk factors for GDM. Meta-analyses were retained if they included at least three studies in which information was provided per included study on a measure of association, its standard error, the number of cases and the total population. We did not apply any language restrictions in the selection of eligible studies. We included only meta-analyses of epidemiological studies in humans. We excluded studies in which risk factors were used for screening, diagnostic, or prognostic purposes, or meta-analyses that examined GDM as a risk factor for other medical conditions. We also excluded studies on women with pre-existing type II diabetes. We excluded systematic reviews and meta-analyses of genetic risk factors, narrative reviews, letters to the editor, meta-analyses of Randomised Control Trials (RCTs), and systematic reviews without a quantitative synthesis of data. If an article presented meta-analyses on other pregnancy outcomes including GDM, we only extracted information on the latter. When more than one meta-analysis on the same research question was eligible, the meta-analysis with the largest number of component studies with data on individual studies’ effect sizes was retained for the main analysis to avoid duplication of the study populations.\nData extraction was performed independently by two investigators (KG, SP), and in case of discrepancies, the final decision was reached by consensus, involving a third investigator, when necessary (EE). From each eligible meta-analysis, we extracted information on the first author, year of publication, the examined risk factors, the number of studies included, the study-specific relative risk estimates (risk ratio, odds ratio, or standardized mean differences) along with the corresponding confidence intervals (CI). Also, we recorded the reported summary meta-analytic estimates using both fixed and random effect methods along with the corresponding confidence intervals, the total population, and number of cases for each study. We also recorded whether the selected meta-analyses applied any criteria to evaluate the quality of the included studies.\nFor each meta-analysis, we re-calculated the summary effect and its 95% CI by using both fixed and random effect models.[ 32 ,  33 ] We also calculated the 95% prediction intervals (PI) for the summary random effects estimates, which further accounts for between-study heterogeneity and indicates the uncertainty for the effect that would be expected in a new study addressing the same association.[ 34 ,  35 ] We considered the largest study as the most precise with a difference between the point estimate and the upper or lower 95% confidence interval less than 0.20 (characterized as small effect size for a continuous outcome according to Cohen’s d definition.[ 36 ] We also recorded whether the largest study presented a statistically significant effect as part of the grading criteria.\nWe assessed heterogeneity among studies, and we reported the P value of the χ 2 -based Cochran Q test and the I 2  metric for inconsistency, which could reflect either diversity or bias. I 2  metric ranges between 0% and 100% and quantifies the variability in effect estimates that is due to heterogeneity rather than sampling error.[ 37 ] Values exceeding 50% or 75% are usually considered to represent large or very large heterogeneity, respectively. Confidence intervals were calculated as per Ioannidis et al.[ 38 ]\nMoreover, we assessed whether there is evidence for small study effect meaning whether smaller studies tend to give substantially larger estimates of effect size compared with larger studies. Small study effects can indicate publication and other selective reporting biases, but they can also reflect genuine heterogeneity, chance, or other reasons for differences between small and large studies.[ 39 ] We used the regression asymmetry test proposed by Egger et al for this assessment.[ 40 ] A P value <0.10 with more conservative effect in larger studies was considered evidence of small-study effects.\nWe further applied the excess significant test to evaluate whether there is a relative excess of significant findings in published literature due to any reason (e.g. publication bias, selective reporting of outcomes or analyses). This is a chi-squared-based test, in which the number of expected positive studies is estimated and compared against the number of observed number of studies with statistically significant results (P<0.05).[ 41 ] A binomial test was then used to evaluate whether the number of positive studies in a meta-analysis is too large according to the power that these studies have to detect plausible effects at α = 0.05. Briefly, a comparison between observed vs. expected is performed separately for each meta-analysis and it is also extended to research areas of many meta-analyses after summing the observed and expected from each meta-analysis. The expected number of significant studies for each meta-analysis is calculated by the sum of the statistical power estimates for each component study.[ 41 ] The power of each component study was estimated using the fixed or random effects summary, or the effect size of the largest study (smallest SE) as the plausible effect size.[ 42 ] The power of each study was calculated with an algorithm using a non-central t distribution.[ 43 ] Excess statistical significance for single meta-analyses was claimed at P<0.10 (one-sided P<0.05, with observed > expected as previously proposed).[ 41 ] We classified risk factors into categories based on biological pathways or types of exposures involved: biomarkers, nutrition and lifestyle factors, diseases and disorders, infections, and other factors. We examined excess of statistical significance separately in each of these categories as selective reporting bias may arise in different categories of research.\nWe characterized as convincing the associations fulfilling the following criteria: a significant effect under the random-effects model at P<10 −6  [ 44 ,  45 ], more than 1000 cases, between-study heterogeneity was not large (I 2 <50%), the 95% PI excluding the null value, and no evidence of small-study effects or excess of significance bias. Additionally, associations with more than 1000 cases, a significant effect at P<10 −6 , and a nominally statistically significant effect present at the largest study were characterized as highly suggestive. We considered as suggestive the associations with significant effect at P<10 −3  and more than 1000 cases. The remaining statistically significant associations at P<0.05 under random-effects model were graded as weak associations.\nTwo independent investigators (KG, SP) assessed the methodological quality of all included systematic reviews and meta-analyses of observational studies using the Assessment of Multiple Systematic Reviews (AMSTAR) tool.[ 46 ] The AMSTAR is an 11-item instrument with scores ranging from 0 to 11 related to vital features of the methodological rigor across systematic reviews and meta-analyses with higher scores indicating greater quality. AMSTAR scores are graded as high (8–11), medium (4–7), and low quality (0–3).[ 46 ,  47 ]\nAll authors had full access to all the data in the study. Statistical analyses were performed in STATA version 14 (STATA Corp, College Station, TX).\n\nOverall, the literature search identified 699 publications of which 616 were excluded after the title and abstract review. Of the 83 articles screened in full text, 22 articles did not report the appropriate information for the calculation of excess of statistical significance (either because the total sample size was missing or the study-specific relative risk estimates were missing), 10 articles were excluded because the outcome of interest was not gestational diabetes, 8 because were editorials or narrative reviews, 5 because were meta-analyses of RCTs, 6 articles excluded because a larger systematic review or meta-analysis including all previous studies investigating the same risk factor was available, and 2 articles were excluded because included only 2 component studies ( Fig 1 ). The 30 eligible papers [ 17 ,  48 – 76 ] included data on 61 different meta-analyses (comparisons) in five broad areas (biomarkers [n = 23 comparisons], nutrition and lifestyle [n = 20 comparisons], diseases and disorders [n = 8 comparisons], infections [n = 2 comparisons], and other factors [n = 8 comparisons]). There were 3 to 40 studies per meta-analysis, with a median of 9 studies. The publication date of the eligible articles ranged between 2009 and 2018. The median number of case and control participants in each study was 84 and 325, respectively. The median number of case and control subjects in each meta-analysis was 1747 and 13850, respectively. The number of cases was greater than 1000 in 38 (62%) meta-analyses ( Table 1 ).\nAbbreviations: Random effects, summary odds ratio (95% CI) using random effects model; Largest effect, odds ratio (95% CI) of the largest study in the meta-analysis; Egger, p-value from Egger's regression asymmetry test for evaluation of publication bias; P, p-value; NP, not pertinent, because the estimated is larger than the observed, and there is no evidence of excess of statistical significance based on the assumption made for the plausible effect size; BMI, Body Mass Index; GDM, gestational diabetes mellitus; PA, physical activity\n* Summary random effects odds ratio (95% CI) of each meta-analysis, except for three meta-analyses (Fu S 2016, Aune D 2016, Pandey S 2012 and Xiao Y 2018) where the RR was used.\n‡ Odds ratio (95% CI) of the largest study in each meta-analysis, except for three meta-analyses (Fu S 2016, Aune D 2016, Pandey S 2012 and Xiao Y 2018) where the RR was used.\n§ P-value from the Egger regression asymmetry test for evaluation of publication bias\n|| I 2  metric of inconsistency and P-value of the Cochran Q test for evaluation of heterogeneity\n≠  95% Prediction Interval\nFourteen papers (47%) used the Newcastle Ottawa Scale (NOS) to qualitatively assess the included primary studies. Three papers (10%) used the Cochrane Collaboration’s risk of bias tool, three (10%) papers used the STrengthening the Reporting of OBservational studies in Epidemiology (STROBE) Statement as a quality assessment tool, and three (10%) papers used other assessment tools. Six papers (20%) did not perform any quality assessment.  S1 Table  summarizes these 30 papers providing data on 61 meta-analyses (comparisons), which included 697 individual study estimates.\nS1 Table  demonstrates the quality assessment of the included meta-analyses using the AMSTAR tool. The median AMSTAR quality score was 7.5 (IQR: 6.25–8.75). All of the meta-analyses included a comprehensive literature search and provided a comprehensive list of the characteristics of the included studies. Most of the meta-analyses did not include a list of the excluded studies while most of the meta-analyses used appropriate methods for data analysis, addressed and incorporated publication bias considerations and the authors reported the conflicts of interest.\nOf the 61 meta-analyses (comparisons), 51 (82%) had nominally statistically significant findings at P<0.05 using the random effects model, while only 15 (25%) remained significant after the application of the more stringent p-value threshold of P<10 −6  ( Table 1 ). The fifteen risk factors that presented a significant effect for an association with GDM at P<10 −6  were the following: pre-pregnancy BMI (as a continuous variable), dietary total iron intake, low vs. normal BMI (cohort studies), overweight vs. normal BMI (cohort studies), BMI >30 vs. normal weight, BMI ~30–35 vs. normal weight, BMI >35 vs. normal weight, overweight vs. non-overweight (cohort studies), overweight vs. non-overweight (case-control), obese vs. non-obese (cohort studies), snoring, sleep-disordered breathing, hypothyroidism, polycystic ovary syndrome, and family history of diabetes. Additional information on all 61 meta-analyses is available online ( S2 Table ).\nAcross the five areas of risk factors there were differences in the proportion of associations that had nominally statistically significant summary effects. Based on the random effects calculations at P<0.05, the proportion of studies with nominally statistically significant summary effects was: 91% for biomarkers, 88% for diseases and disorders and 85% for nutrition and lifestyle. On the contrary, this was seen only in 50% of the meta-analyses on other risk factors and infections, respectively.\nSixteen (26%) meta-analyses had large heterogeneity estimates (I 2  ≥ 50% and I 2  ≤ 75%) and 15 (25%) meta-analyses had very large heterogeneity estimates (I 2  > 75%) ( Table 1 ). When we calculated the 95% prediction intervals, in 18 (30%) meta-analyses the null value was excluded. This included seven biomarkers [maternal iron deficiency, ferritin levels, DQ6, thyroid antibodies (case-control studies), thyroid antibodies (all studies), 25(OH)D5 <50 nmol/l, 25(OH)D <75 nmol/l], eight nutrition and lifestyle factors [prenatal exercise (cohort studies), low vs. normal BMI (cohort studies), overweight vs. normal BMI (cohort studies), BMI >30 vs. normal weight, BMI ~30–35 vs. normal weight, BMI >35 vs. normal weight, overweight vs. non-overweight (cohort studies), obese vs. non-obese (cohort studies)], two diseases and disorders (subclinical hypothyroidism and hypothyroidism), and one other risk factor (family history of diabetes) ( Table 1 ).\nEvidence for statistically significant small-study effects (Egger test P<0.10 and random effects summary estimate larger compared to the point estimate of the largest study in the meta-analysis) was identified in 5 out of 61 (8%) meta-analyses ( S2 Table , available online). These included four meta-analyses on biomarkers (hemoglobin concentration, mean platelet volume, serum retinol-binding protein-4, DQ2), and one on other factors (Extreme sleep duration). Eight (13%) associations had hints of excess statistical significance bias with statistically significant (P<0.05) excess of positive studies under any of the three assumptions for the plausible effect size—the fixed effects summary, the random effects summary or the results of the largest study ( S2 Table ). Four (50%) of them pertained to biomarkers, three (38%) pertained to nutrition and lifestyle, and one (12%) pertained to other risk factors.  Table 2  shows the results of excess of statistical significance bias according to category of risk factor.\n* NP, not pertinent, because the estimated is larger than the observed, and there is no evidence of excess of statistical significance based on the assumption made for the plausible effect size.\n† Expected number of statistically significant studies using the summary fixed effects estimate of each meta-analysis as the plausible effect size.\n‡ P value of the excess of statistically significant test. All statistical tests were two-sided.\n§ Expected number of statistically significant studies using the summary random effects estimate of each meta-analysis as the plausible effect size.\n‖ Expected number of statistically significant studies using the effect of the largest study of each meta-analysis as the plausible effect size.\n¶ Expected number of statistically significant studies using the most conservative of the three estimates (fixed effects summary, random effects summary, largest study) of each meta-analysis as the plausible effect size.\nAfter applying our credibility criteria, four risk factors, low vs. normal BMI (cohort studies), BMI ~30–35 vs. normal weight, BMI >35 vs. normal weight, and hypothyroidism (all types) presented convincing evidence for an association with GDM, supported by more than 1000 cases, P<10 −6  under the random effect model, no hints for small-study effects and for excess statistical significance, not large heterogeneity (I 2 <50%), and a 95% PI excluding the null value. Ten risk factors [pre-pregnancy BMI (as a continuous variable), overweight vs. normal BMI (cohort), BMI >30 vs. normal weight, overweight vs. non-overweight (cohort), overweight vs. non-overweight (case-control), obese vs. non-obese (cohort), snoring, sleep-disordered breathing, polycystic ovary syndrome, family history of diabetes] presented highly suggestive evidence for GDM.\nNine risk factors were supported by suggestive evidence and twenty-seven associations presented weak evidence (P<0.05). An overall assessment of statistically significant associations for GDM is presented in  Table 3 .\nAbbreviations: BMI, Body Mass Index; GDM, gestational diabetes mellitus.\na  P indicates the P-values of the meta-analysis random effects model.\nb  Small study effect is based on the P-value from the Egger’s regression asymmetry test (P<0.10).\nc  Based on the P-value (P<0.05) of the excess significance test using the largest study (smallest standard error) in a meta-analysis as the plausible effect size.\n\nIn this umbrella review we evaluated the current evidence, derived from meta-analyses of observational studies on the association between various risk factors and GDM. Overall, from the 61 associations that have been examined, only a minority had strongly significant results with no suggestion of bias, as can be inferred by substantial heterogeneity between studies, small study effects, and excess significance bias. Four risk factors were supported by convincing evidence, including low vs. normal BMI (cohort studies), BMI ~30–35 vs. normal weight, BMI >35 vs. normal weight, and hypothyroidism. Another ten risk factors from various fields [pre-pregnancy BMI (as a continuous variable), overweight vs. normal BMI (cohort), BMI >30 vs. normal weight, overweight vs. non-overweight (cohort), overweight vs. non-overweight (case-control), obese vs. non-obese (cohort), snoring, sleep-disordered breathing, polycystic ovary syndrome, family history of diabetes], achieved highly suggestive evidence for an association with GDM.\nIt is well-known that maternal weight, as determined from pre-conception BMI, is critical on the development of insulin resistance and type II diabetes as well as GDM. This summary of observational studies shows that the more robust associations were related to overweight and obesity, as three out of four associations that met the criteria for convincing evidence and six out of ten highly suggestive associations were concentrated on maternal pre-pregnancy BMI and the risk of GDM. The association of low BMI vs. normal BMI was the only protective factor, which it was supported by convincing evidence for protection against GDM.\nOur findings further support the current guidelines regarding pregnancy weight, nutrition and activity, issued from the National Institute for Health and Clinical Excellence (NICE), the Institute of Medicine (IOM) and the American College of Obstetricians and Gynecologists (ACOG), which they accepted lifestyle change as an essential component of prevention and management of GDM.[ 77 – 79 ] NICE recommendations include specific guidelines for healthy eating, low-fat diet and moderate physical activity before, during, and after pregnancy.[ 78 ] Preventive measures against gestational diabetes may include diet and exercise as described on the most recent Cochrane review of interventions from moderate quality evidence. Nevertheless, the variability of the diet and exercise components tested in the included studies, make the evidence insufficient to inform practice.[ 80 ] Large, well-designed, RCTs are needed to confirm the effectiveness of pre-conception weight and gestational weight gain reduction and the effects of dietary interventions in pregnancy for preventing GDM in different categories of pre-pregnancy BMI with special focus on overweight and obese women.\nThe observed association between obesity and GDM is biologically plausible. Normal pregnancy is characterized by a state of insulin resistance defined as an impaired response to insulin. This physiological insulin resistance also occurs in women with GDM on a background of chronic insulin resistance due to obesity to which the insulin resistance of pregnancy is partially additive. Obesity can cause major changes in maternal intermediary metabolism, where co-existing conditions associated with increased insulin resistance, higher serum lipids, and lower plasma levels of adiponectin, appear to play a central role to the development of GDM.[ 81 – 83 ]\nThe association between hypothyroidism, which includes both subclinical and overt hypothyroidism, and risk of GDM, was supported by convincing evidence. Increased levels of human chorionic gonadotropin (hCG) in the first trimester of pregnancy directly stimulate the thyroid gland to increase production of thyroid hormone, which leads in decreased secretion of thyroid stimulating hormone (TSH).[ 84 ] Proposed mechanisms that describe the relationship between hypothyroidism and gestational diabetes are supported from studies that show that both overt and subclinical hypothyroidism can lead to significantly increased insulin resistance.[ 85 – 88 ] Although, these findings would suggest that routine screening of thyroid hormones during pregnancy could be essential, universal thyroid screening in pregnancy is controversial.[ 89 ] The most recent ACOG recommendations suggest testing only women at high risk of thyroid disease before they become pregnant or when they are early in pregnancy.[ 90 ] On the contrary, the American Thyroid Association [ 91 ] and the Endocrine Society [ 92 ] call for universal thyroid-function screening early in pregnancy. On the side of this controversy, women with known thyroid disease could be offered GDM screening earlier in pregnancy.\nIn the current umbrella review, we applied a transparent and replicable set of criteria and statistical tests to evaluate and categorize the level of existing observational evidence within in five broad areas with the goal to detect biases that work on a field-wide level. Although, 82% of the included meta-analyses report a nominally (P<0.05) statistically significant random-effects summary estimate, when stringent P value was considered (P<10 −6 ), the proportion of significant associations decreased to 25%. Thirty-one (51%) associations had large or very large heterogeneity, while when we calculated the 95% prediction intervals, which further account for heterogeneity, we found that the null value was excluded in more than half of the associations. Only four of the assessed risk factors found to provide convincing evidence, indicating that several published meta-analyses of observational studies in the field could be susceptible to biases and the reported associations in the existing studies are often exaggerated.\nThe ability to modify those factors, mainly those related to overweight and obesity, through clinical interventions or public health policy measures remains to be established. Furthermore, there is no guarantee that even a convincing observational association for a modifiable risk factor would necessarily translate into large preventive benefits for GDM if these risk factors were to be modified.[ 93 ] With obesity becoming a global epidemic, the assessment of the strength of the evidence supporting the impact of overweight and obesity in GDM could allow the identification of women at high risk for adverse outcomes and allow better prevention. Obesity is generating an unfavorable metabolic environment from early gestation; therefore, initiation of interventions for weight loss during pregnancy might be belated to prevent or reverse adverse effects, which highlights the need of weight management strategies before conception.[ 94 ] GDM does not only increase the risk for maternal and fetal complication in pregnancy, but also significantly increases a woman’s risk of type 2 diabetes, metabolic syndrome (characterized by glucose intolerance, central obesity, dyslipidemia, and insulin resistance), and cardiovascular disease (CVD) after pregnancy.[ 95 – 98 ]\nUmbrella reviews focus on existing systematic reviews and meta-analyses and therefore some studies may have not been included either because the original systematic reviews did not identify them, or they were too recent to be included. In the current assessment we used all available data from observational studies, therefore the meta-analysis estimates may partly reflect the biases from which the original studies suffer from. Statistical tests of bias in the body of evidence (small study effect and excess significance tests) offer hints of bias, not definitive proof thereof, while the Egger test is difficult to interpret when the between-study heterogeneity is large. These tests have low power if the meta-analyses include less than 10 studies and they may not identify the exact source of bias.[ 39 ,  41 ,  99 ] Furthermore, we did not appraise the quality of the individual studies on our own, since this should be included in the original meta-analysis and it was beyond the scope of the current umbrella review. However, we recorded whether and how they performed a quality assessment of the synthesized studies. Lastly, we cannot exclude the possibility of selective reporting for some associations in several studies. For example, perhaps some risk factors were more likely to be reported, if they had statistically significant results.\n\nThe present umbrella review of meta-analyses identified 61 unique risk factors for GDM. Our analysis identified four risk factors with convincing evidence and strong epidemiological credibility pertaining to hypothyroidism and BMI (specifically, low vs. normal BMI (cohort studies), BMI ~30–35 vs. normal weight, BMI >35 vs. normal weight). Diet and lifestyle modifications in pregnancy should be tested in large randomized trials. Our findings suggest that women with known thyroid disease could be offered screening for GDM earlier in pregnancy. As previously suggested, the use of standardized definitions and protocols for exposures, outcomes, and statistical analyses may diminish the threat of biases, allow for the computation of more precise estimates and will promote the development and training of prediction models that could promote public health.\n\nNote: Y: Yes, N: No, CA: Cannot Answer. Item 1: Was an ‘‘a priori” design provided? Item 2: Was there duplicate study selection and data extraction? Item 3: Was a comprehensive literature search performed? Item 4: Was the status of publication (i.e., grey literature) used as an inclusion criterion? Item 5: Was a list of studies (included and excluded) provided? Item 6: Were the characteristics of the included studies provided? Item 7: Was the scientific quality of the included studies assessed and documented? Item 8: Was the scientific quality of the included studies used appropriately in formulating conclusions? Item 9: Were the methods used to combine the findings of studies appropriate? Item 10: Was the likelihood of publication bias assessed? Item 11: Was the conflict of interest included?\n(DOCX)\nClick here for additional data file.\nAbbreviations: Random effects, summary odds ratio (95% CI) using random effects model; Largest effect, odds ratio (95% CI) of the largest study in the meta-analysis; Egger, p-value from Egger's regression asymmetry test for evaluation of publication bias; P, p-value; NP, not pertinent, because the estimated is larger than the observed, and there is no evidence of excess of statistical significance based on the assumption made for the plausible effect size; BMI, Body Mass Index; GDM, gestational diabetes mellitus; PA, physical activity.* Summary random effects odds ratio (95% CI) of each meta-analysis, except for three meta-analyses (Fu S 2016, Aune D 2016, Pandey S 2012 and Xiao Y 2018) where the RR was used. † Summary fixed effects odds ratio (95% CI) of each meta-analysis, except for three meta-analyses (Fu S 2016, Aune D 2016, Pandey S 2012 and Xiao Y 2018) where the RR was used.‡ Odds ratio (95% CI) of the largest study in each meta-analysis, except for three meta-analyses (Fu S 2016, Aune D 2016, Pandey S 2012 and Xiao Y 2018) where the RR was used.§ P-value from the Egger regression asymmetry test for evaluation of publication bias|| I2 metric of inconsistency (95% confidence intervals of I2) and P-value of the Cochran Q test for evaluation of heterogeneity.\n≠ 95% Prediction Interval ¶ Observed number of statistically significant studies # Expected number of statistically significant studies using the summary fixed effects estimate of each meta-analysis as the plausible effect size** P-value of the excess statistical significance test.\n¥ Expected number of statistically significant studies using the summary random effects estimate of each meta-analysis as the plausible effect size ȣ Expected number of statistically significant studies using the effect of the largest study of each meta-analysis as the plausible effect size.\n(DOCX)\nClick here for additional data file.","source_license":"CC-BY-4.0","license_restricted":false}