Methods
The Cochrane Handbook for Systematic Reviews of Interventions ( Higgins and Green, 2011 ) was followed to conduct this review and meta-analysis and the findings were presented according to the PRISMA guideline. Registration number: PROSPERO 2018 CRD42018104879 (for poor responders of IVF), PROSPERO CRD42019150069 (for normal/high responders of IVF).
There was no restriction on language. We included studies from January 1990 (since the introduction of the concept of poor or high ovarian response in IVF) to April 2020. Abstracts or conference proceedings were also reviewed and included, avoiding duplication, only if all required information was available. Studies were excluded if complete information was not obtained despite personal request.
Participants. Couples underwent IVF/ICSI due to any cause, where the female partners were known or anticipated to have normal, high or poor response to ovarian stimulation. We went by the inclusion criteria as described by the authors to define the population as poor, normal (unselected) or hyper-responders and grouped the trials accordingly.
Poor responders: where women were predicted to have low ovarian reserve based on elevated basal follicle stimulation hormone (FSH) and/or low anti-Mullerian hormone (AMH) and/or low antral follicle count (AFC) and/or low ovarian response in the previous cycle and those who met the Bologna criteria ( Ferraretti et al. , 2011 ).
Normal responders: where the age of the women or ovarian reserve or previous ovarian response, as described by the authors, predicted to result in a not too low, or too high ovarian response. The definition of normal responders is based on predicted response only; some women might have had an unexpected exaggerated response while some others an unpredicted poor response. This limitation has been accepted, in absence of any better marker to denote ‘normal responders’.
Hyper-responders: where women were predicted to yield high ovarian response based on high AMH and/or high AFC and/or exaggerated follicular response in the previous cycle, except where a diagnosis of typical polycystic ovary syndrome (PCOS) was made.
If there is no mention of age or ovarian reserve in the primary study, we have classified them as ‘unselected population’ and included the data under the normal responders for meta-analysis.
Intervention. MD-IVF: Treatment protocol using ‘mild’ or low-dose (≤150 IU daily) gonadotrophin (FSH or hMG) alone, or in combination with oral compounds (e.g. CC/AIs) or oral compounds alone irrespective of agonist or antagonist protocol.
Comparison. CD-IVF: Protocols with gonadotrophin exposure higher than that of mild or low-dose arm in terms of daily dose and or duration.
The search conformed to the standard descriptions of ‘mild’ and ‘conventional’ stimulation IVF protocols ( Nargund et al. , 2007 ; Zegers-Hochschild et al. , 2009 ); but because of the varying description of these terms in the literature, we were obliged to define them on the basis of gonadotrophin dosage. This permitted the comparison of the outcomes of mild and conventional stimulation dosages of gonadotrophins (FSH and hMG) on the same population, whether daily or de facto ‘cumulative’.
Exclusion criteria . Studies comparing oocyte or embryo yield only with no data on any of the primary outcomes measured in this review were excluded. Studies comparing a ‘standard’ 150 IU daily dose in one arm with a wide range of ‘individualised’ stimulation dosage in the other arm based on ovarian reserve were excluded.
Primary outcomes . Live birth rate (LBR) per woman randomised; OHSS and cycle cancellation rates (CCRs) per cycle started.
Secondary outcomes . Cumulative LBR, ongoing pregnancy rate (OPR), clinical pregnancy rates (CPRs) (with separate note on biochemical pregnancies) as defined in the ICMART glossary ( Zegers-Hochschild et al. , 2009 ), total dose of gonadotrophin used, number of oocytes, number of embryos, number of high-grade embryos per started cycle and cost comparison. The number of embryos transferred may not be a true reflection of total number of embryos created, therefore was not considered.
All outcomes were derived from the first or only treatment cycle with fresh embryo transfer conducted in the individual trials, except while reporting the cumulative outcomes. Cumulative live birth, whether adding data from all subsequent frozen embryo transfer cycle(s) or subsequent fresh cycles as well as frozen cycles in a given study period, were expressed as per-patient randomised. Cumulative secondary outcomes, e.g. incidence of OHSS, cycle cancellations or mean number of oocytes or embryos, were therefore reported on a per started cycle basis, counting the outcomes from all fresh cycles together.
An electronic search was conducted in Medline, Embase, PreMedline and Cochrane Central from January 1990 (inception of the concept of low or high responder) to April 2020. Databases were searched using relevant medical subject headings, free-text terms and study type filters where appropriate, without language restrictions. Advance access articles of key journals were checked for related papers. The reference list of all reviews or individual RCTs was also hand-searched to find any additional RCT. Duplications arising from a conference abstract and subsequent full-text paper were excluded.
((IVF, ICSI, ovarian stimulation) AND ((mild IVF stimulation, oral agents, aromatase inhibitors, clomiphene, letrozole, anastrozole) OR ((gonadotropin, FSH, follitropin, hMG, menotrophin) AND (dose, low dose))) AND randomised controlled trials]. Because of the diversity in protocols, the terms related to CD-IVF were not included in the electronic search; however, individual abstracts were reviewed to confirm eligibility of the CD-IVF protocols and to identify trials on poor, normal or high responders in IVF. The electronic search was performed by National Guideline Alliance (NGA) of Royal College of Obstetricians and Gynaecologists.
First an electronic search was made using the search terms and databases described above. Full text of all shortlisted studies (RCTs) was reviewed by two reviewers (A.K.D. and N.F.) independently; conflict if any was resolved by any of the other reviewers (S.C. or G.N.). References of all included and excluded full-text papers and other related systematic reviews were hand-searched to look for additional RCTs. Cochrane Handbook for Systematic Reviews of Interventions ( Higgins and Green, 2011 ) was consulted to prepare the data-extraction form, obtain the features of included studies, assess RoB and outcome data. Review Manager 5 (version 5.3) software was used to construct the RoB graph, Funnel plots and Forest plots in this review (Review Manager (RevMan) (Computer program) Version 5.3. Copenhagen: The Nordic Cochrane Centre, The Cochrane Collaboration, 2014).
The following information and data were extracted.
Year and location of the trial (single or multi-centre), type of trials (2-arm/3-arm etc.), study population with sample size calculation, method of randomisation, method of allocation concealment, exclusion of participants after randomisation, proportion of and reasons for losses at follow up, reports of ethical approval and consent.
Age, ovarian reserve of the women, e.g. FSH, AMH, AFC, ovarian response in the previous IVF/ICSI cycles (if mentioned) to categorise women in poor, normal and high ovarian response groups. In addition, whether in accordance with Bologna criteria ( Ferraretti et al. , 2011 ) for poor responders, exclusion criteria of individual trials were also noted.
Treatment protocols in the intervention and comparator group(s) with regards to the type of medications (oral and injectable), dose, time of commencement, method of suppression of premature ovulation, dose adjustment or pre-treatment or co-intervention, if any, ovulation trigger type and dose, cancellation criteria and luteal phase regimen were noted.
What outcomes were reported, how the outcomes were defined and the timing of outcome measurement (e.g. per woman randomised/started cycle or per embryo transfer) were recorded. Cumulative live birth data were extracted as an aggregation of both the first fresh and all subsequent frozen transfer cycle(s) or further fresh cycle(s); data from each subsequent fresh or frozen cycle(s) were not analysed separately. In the cost analysis, whether total cost per cycle or per woman or cumulative cost of fresh and frozen cycles were noted.
RoB was assessed under the headings of Sequence generation, Allocation concealment, Blinding of participants and assessors, Selective outcome reporting and Other sources of bias as outlined in Cochrane Handbook for Systematic Reviews of Interventions ( Higgins and Green, 2011 ) . Blinding of patients and clinicians was neither possible nor applicable for this particular type of intervention and outcomes (e.g. pregnancy rates). We considered studies with absence of blinding as low RoB, as it was unlikely to influence outcomes. The RoB was considered ‘unclear’ if the information was insufficient in any type of bias.
For dichotomous data, relative risk (RR) and for continuous data, mean differences (MD) between treatment groups were calculated with 95% CI. In case of anticipated heterogeneity, a random effect model was used. In presence of heterogeneous data, the standardised mean difference (SMD) was used.
Authors were contacted for missing data by email at least twice.
The clinical and methodological characteristics of all included studies were examined ( Table I ); sub-group meta-analysis was performed as described below. Statistical heterogeneity was assessed by the Chi 2 test. The I 2 statistic assessed the impact of the heterogeneity on the meta-analysis; an I 2 of >50% indicated significant heterogeneity, in which case a ‘random effect model’ was applied, otherwise, a ‘fixed-effect model’ was used as a default.
Characteristics of included randomised controlled trials.
RCT, randomised controlled trials; #, number; yrs, years; BMI body mass index, *kg/m 2 ; Prev., previous; d, day; D, (cycle) day; GnRH-ant, GnRH antagonist; GnRH-a, GnRH agonist; DR, downregulation; adj, adjustment; LPS, luteal phase support; ?, not stated; ET, embryo transfer; OPR, ongoing pregnancy rate; NS, not significant; OHSS, ovarian hyperstimulation syndrome; CCR, cycle cancellation rate; P, progesterone; PR, pregnancy rate; Gn, gonadotrophin; AMH, anti-Mullerian hormone; AFC, antral follicle count; LBR, live birth rate; Cum LBR, cumulative live birth rate; CPR, clinical pregnancy rate; CC, clomiphene citrate; E2, oestradiol; Let, letrozole rhCG, recombinant hCG; rFSH, recombinant FSH; SET, single-embryo transfer; Micro, micronized; POR, poor ovarian reserve; COC, combined oral contraceptive.
A funnel plot was generated with all included studies on CCR outcome. We did not limit our search by any language or time.
For each outcome, meta-analyses were performed separately for poor responders, normal and high responders. Subgroup analysis was performed with different types of mild stimulation protocols: low-dose versus high-dose gonadotrophin only protocols; CC+ mild-dose gonadotrophin versus CD-IVF protocols and Letrozole+ mild-dose gonadotrophin versus CD-IVF protocols.
The methodology described by the Cochrane Hand book for Systematic Review of Intervention was followed in the meta-analysis of multi-arm studies ( Higgins and Green, 2011 ). If MD-IVF was compared with two different CD-IVF protocols, both the events and populations (denominators) in MD-IVF were equally divided and incorporated under respective sub-groups. If MD-IVF or CD-IVF consisted of two different doses or types of gonadotrophin, they were combined into one taking the average of both events and populations. For continuous data in the above situations, the mean and SD of the common groups were kept the same, only the population was equally split into two subgroups.
We performed sensitivity analysis by repeating meta-analyses of all outcomes in the following ways: excluding and including small studies with RoB; excluding studies with permitted dose adjustment; gonadotrophin with and without oral compounds; applying a fixed as well as a random effect model; and applying RR and peto odd-ratio (OR) as the method of determining effect size.
Results
The study selection process is demonstrated in the flow chart ( Fig. 1 ). Three publications were found by hand searching ( Out et al. , 2004 ; Tan et al. , 2005 ; Mukherjee et al. , 2012 ), the rest by electronic search. Forty-five shortlisted publications underwent full-test review for further assessment of eligibility criteria. Table II narrated the list of excluded studies with reasons. A large RCT applied single-embryo transfer policy in a ‘minimal’ group, with double-embryo transfer in a ‘conventional’ IVF group, but had both fresh and frozen-thawed transfer in both groups ( Heijnen et al. , 2005 )—this study was excluded for pregnancy outcomes per randomisation as these outcomes could have been affected by the differential embryo-transfer policy. However, cumulative pregnancy outcome, CCR and laboratory parameters would not have been affected hence this study was included in the meta-analyses for these outcomes. Finally, 31 RCTs were included: 15 RCTs in the poor, 14 RCTs in the normal and 2 RCTs in the hyper-responder group.
Flow-chart of the study selection process. RCT, randomised controlled trial.
The list of excluded studies. *
The table explains on what basis some of the studies that were included in other related systematic reviews were considered not eligible for this review.
Table I summarised the studies included in this review and meta-analysis. All included papers were written in English except one ( Martinez et al. , 2003 ), which was written in Spanish and the translation was by Google Translator . Two were conference abstracts with sufficient data for meta-analysis ( Huang et al. , 2015 ; Elnashar et al. , 2016 ).
Seven included studies were multi-centre trials ( Out et al. , 2004 ; Tan et al. , 2005 ; Heijnen et al. , 2007 ; Ragni et al. , 2012 ; Oudshoorn et al. , 2017 ; van Tilborg et al. , 2017 ; Youssef et al. , 2018 ), the rest were from a single centre. Four trials conducted three-arm comparison ( Harrison et al. , 1994 ; Ashrafi et al. , 2005 ; Bastu et al. , 2016 ; Yu et al. , 2018 ), one was a four-arm trial ( Martinez et al. , 2003 ), the rest were two-arm studies. Sample size calculation was done in six trials among poor responders: two for oocyte number ( Revelli et al. , 2014 ; Bastu et al. , 2016 ), one for CPR ( Yu et al. , 2018 ), one for OPRs ( Youssef et al. , 2017 ), one for LBR ( Ragni et al. , 2012 ) and two for cumulative live birth ( van Tilborg et al. , 2017 ; Liu et al. , 2020 ). Of the studies on unselected patients, four were powered for oocyte numbers ( Out et al. , 2004 ; Tan et al. , 2005 ; Baart et al. , 2007 ; Blockeel et al. , 2011 ), two for pregnancy rate (PR) ( Tummon et al. , 1992 ; Dhont et al. , 1995 ) and one for LBR ( Heijnen et al. , 2007 ). Both the RCTs on hyper-responders were large: one had adequate power for number of oocytes (n = 412) ( Casano et al. , 2012 ), the other for cumulative LBRs ( Oudshoorn et al. , 2017 ).
Recruitment in five RCTs was as per the Bologna consensus on poor ovarian response (POR) ( Ragni et al. , 2012 ; Huang et al. , 2015 ; Bastu et al. , 2016 ; Pilehvari et al. , 2016 ; Liu et al. , 2020 ); others were based on different combinations of age, FSH, AMH, AFC and previous poor response ( Table I ). Selection of patients in the non-PCOS hyper-responder group was on the sole criterion of AFC in both the RCTs ( Casano et al. , 2012 ; Oudshoorn et al. , 2017 ). Unselected patients/normal responders were recruited in absence of high or low ovarian reserve, mostly on the first cycle of IVF (detailed in Table I ).
Interventions in each individual trial were detailed in Table I . Comparison between low- and high-dose gonadotrophins only stimulation (without oral medication) was reported in six RCTs on poor responders ( Ashrafi et al. , 2005 ; Klinkert et al. , 2005 ; Kim et al. , 2009 ; van Tilborg et al. , 2017 ; Youssef et al. , 2017 ; Yu et al. , 2018 ); seven RCTs on normal responders ( Hohmann et al. , 2003 ; Out et al. , 2004 ; Tan et al. , 2005 ; Baart et al. , 2007 ; Heijnen et al. , 2007 ; Lou and Huang, 2010 ; Blockeel et al. , 2011 ) and both the RCTs on hyper-responders. Ten trials in the patient with POR used oral compounds in the MD-IVF arm either alone (CC) ( Ragni et al. , 2012 ) or in combination with low-dose gonadotrophins: CC was used in five and Letrozole in three trials ( Table I ). Among the normal responder group, six RCTs used CC+ gonadotrophin and two with Letrozole combination. Consistently, CC was used at 100 mg daily dose for 5 days, commencing on cycle Day 2–4, except the RCT by Ragini et al. where 150 mg daily dose was used. The dose for Letrozole was 5 mg daily, starting from Day 2–5, except in two trials: one used 2.5 mg daily ( Goswami et al. , 2004 ) and the other 10 mg daily dose ( Elnashar et al. , 2016 ). In all trials, the starting dose of gonadotrophin for MD-IVF was 150 IU daily, except in two RCTs on poor responder where a 75 IU dose was used ( Goswami et al. , 2004 ; Yu et al. , 2018 ); in one trial for the normal ( Tan et al. , 2005 ) and one for the high responders ( Oudshoorn et al. , 2017 )- 100 IU daily dose was used in both studies. However, the timing of commencement of gonadotrophin varied ( Table I ). Dose adjustment was allowed in 12 RCTs, fixed dose in 13 and not mentioned in remaining six trials ( Table I ). Pre-treatment was given in three RCTs ( Dhont et al. , 1995 ; Mohsen and El Din, 2013 ; Youssef et al. , 2017 ). Cycle cancellation criteria varied between the studies ( Table I ).
The definition of cumulative LBR differed among the studies: the RCT by Casano et al. (2012) and Liu et al. (2020) aggregated the outcome of fresh and all subsequent frozen-thawed transfer; while other three trials included all fresh and frozen cycles within a specified time-period of 12 months ( Heijnen et al. , 2007 ) or 18 months ( Oudshoorn et al. , 2017 ; van Tilborg et al. , 2017 ). Three studies reported pregnancy rates as positive beta-hCG ( Dhont et al. , 1995 ; Hohmann et al. , 2003 ; Blockeel et al. , 2011 ) and two trials did not specify whether it was clinical pregnancy ( Tummon et al. , 1992 ; Elnashar et al. , 2016 ) and therefore excluded from the meta-analysis on CPR. The criterion for cycle cancellation was not uniform ( Table I ). The clinical criteria for reporting of OHSS varied between the trials and were not clear in some studies. Three RCTs estimated total and mean per-patient cost of all fresh and frozen cycles together ( Heijnen et al. , 2007 ; Oudshoorn et al. , 2017 ; van Tilborg et al. , 2017 ); one trial reported total and per-patient cost of only fresh cycle ( Ragni et al. , 2012 ) and the remaining two reported the medication cost of stimulated cycles ( Lou and Huang, 2010 ; Mukherjee et al. , 2012 ).
A summary of RoB was graphically presented in Fig. 2 .
Risk of bias graph from the included studies .
All RCTs were found to be ‘low-risk’ for random sequence generation, except five trials where the risk was unclear ( Dhont et al. , 1995 ; Ashrafi et al. , 2005 ; Mukherjee et al. , 2012 ; Elnashar et al. , 2016 ; Pilehvari et al. , 2016 ). Allocation concealment was deemed to have low risk in all but seven RCTs where the risk was unclear ( Tummon et al. , 1992 ; Dhont et al. , 1995 ; Martinez et al. , 2003 ; Lou and Huang, 2010 ; Pilehvari et al. , 2016 ; Yu et al. , 2018 ; Liu et al. , 2020 ). Performance and detection bias: All RCTs were of ‘low-risk’ for performance bias, as the blinding of both patients and assessors was neither possible nor required for these objective outcome measures. Attrition bias: The outcome data were not complete in one trial (high risk) ( Huang et al. , 2015 ), and not clear in the three other studies ( Ashrafi et al. , 2005 ; Mohsen and El Din, 2013 ; Elnashar et al. , 2016 ), the rest were of ‘low risk’. Reporting bias: All RCTs had ‘low risk’ for reporting bias. Other bias: Baseline characteristics of both sides were not clear in eight RCTs ( Ashrafi et al. , 2005 ; Mukherjee et al. , 2012 ; Ragni et al. , 2012 ; Huang et al. , 2015 ; Bastu et al. , 2016 ; Elnashar et al. , 2016 ).
Poor responders. Five RCTs compared LBRs (n = 1248), two of them compared mild and conventional-dose gonadotrophin only stimulation ( Kim et al. , 2009 ; van Tilborg et al. , 2017 ), one CC and high-dose antagonist protocol ( Ragni et al. , 2012 ) and two with letrozole combination, of which the study by Yu et al. , also had a 3rd arm with low-dose gonadotrophin only protocol ( Yu et al. , 2018 ; Liu et al. , 2020 ). There was no evidence of a difference in LBRs: RR 0.91 (CI 0.68, 1.22) ( Fig. 3A ). There was no statistical heterogeneity ( I 2 0%) and four RCTs were of low RoB ( Kim et al. , 2009 ; Ragni et al. , 2012 ; van Tilborg et al. , 2017 ; Liu et al. , 2020 ). The finding remained unchanged in sensitivity analysis, when the smaller RCTs with possible RoB were excluded or whether trials with dose adjustments were included or excluded. The inference was the same, whether gonadotrophin only protocol or CC/Letrozole protocols were used. Due to the presence of significant clinical heterogeneity, the quality of evidence (QoE) was moderate ( Table III ).
Forest plot of mild versus conventional-dose IVF: live birth rate per randomisation . A for poor responders, B for normal responders, C for hyper-responder. MD-IVF, mild-dose IVF; CD-IVF, conventional-dose IVF.
Summary of evidence.
No difference ⊕⊕⊕⊖
RR 0.91 [0.68, 1.22]
RCT= 5, n= 1248
✓ I 2 0%
✓ Narrow CI
✓ 2 large RCTs low RoB
✓ No RCT contradicted
× 1 study with unclear RoB
× Clinical heterogeneity
No difference ⊕⊕⊕⊖
RR 0.88 [CI 0.69, 1.12]
RCT= 3, n= 573,
✓ I 2 0%
✓ Narrow CI
✓ ↓ Clinical heterogeneity
✓ No RCT contradicted
× Studies with unclear RoB
No difference ⊕⊕⊕⊖
RR 0.98 [CI 0.79, 1.22]
RCT= 2, n= 931
✓ I 2 0%
✓ Narrow CI
✓ Only 1 unclear RoB
✓ No RCT contradicted
× ↑ Clinical heterogeneity
↓ with MD-IVF ⊕⊕⊕⊖
RR 0.26 [CI 0.14, 0.49]
RCT= 9, n= 1925
✓ I 2 0%
✓ Narrow CI
✓ Large effect size
× Unclear RoB
× Clinical heterogeneity
↓ with MD-IVF ⊕⊕⊕⊖
RR 0.47 [CI 0.31, 0.72]
RCT=2, n=931
✓ I 2 0%
✓ Narrow CI
✓ Large effect size
✓ Low RoB (1 unclear)
× Clinical heterogeneity
No difference ⊕⊖⊖⊖
RR 1.33 [CI 0.96, 1.85]
RCT= 15, n= 3459
× I 2 64%
× Wide CI
× Most RCTs with RoB
× Clinical heterogeneity
↑ with MD-IVF ⊕⊖⊖⊖
RR 2.08 [CI 1.38, 3.14]*
RCT= 12, n= 2654
× I 2 48%
× Wide CI
× Small RCTs, unclear RoB
× Clinical heterogeneity
No difference ⊕⊕⊕⊖
RR 1.31 [CI 0.98, 1.77]
RCT= 2, n= 1348
✓ I 2 0%
✓ 2 large RCTs low RoB
× Moderately wide CI
Clinical heterogeneity
No difference ⊕⊕⊕⊖
RR 1.02 [CI 0.81, 1.25]
RCT= 7, n= 2006
✓ I 2 0%
✓ Narrow CI
✓ 3 large RCTs low RoB
✓ No RCT contradicted
× Clinical heterogeneity
No difference ⊕⊕⊕⊖
RR 1.10 [CI 0.88, 1.38]
RCT= 7, n= 1026
✓ I 2 0%
✓ ↓ Clinical heterogeneity
✓ No RCT contradicted
× Small studies with unclear RoB
No difference ⊕⊕⊖⊖
RR 0.86 [CI 0.61, 1.23]
RCT= 1, n= 521
✓ Large RCT
✓ Low RoB
× Based on just 1 RCT with
two different protocols
↓ with MD-IVF ⊕⊕⊖⊖
SMD -0.43 [CI -0.58, -0.28]
RCT= 14, n= 2773,
✓ Large effect size, narrow CI
× I 2 67%
× RCTs with RoB
× Clinical heterogeneity
↓ with MD-IVF ⊕⊖⊖⊖
SMD -1.34 [CI -1.94, -0.75]
RCT= 13, n= 3499,
× I 2 98%
× Wide CI
× RCTs with unclear RoB,
× Clinical heterogeneity
No difference ⊕⊕⊖⊖
SMD -0.31 [CI -0.74, 0.13]
RCT=2, n=931
✓ Low RoB (1 unclear)
× I 2 91%
× Wide CI
× Clinical heterogeneity
↓ with MD-IVF ⊕⊕⊖⊖
SMD -0.39 [CI -0.59, -0.20]
RCT= 9, n= 1559,
✓ Narrow CI
✓ 2 large RCTs with low RoB
× I 2 59%
× 1 RCT with high RoB
× Clinical heterogeneity
No difference ⊕⊖⊖⊖
SMD -0.30 [-0.58, 0.08]
RCT= 7, n= 1884,
× I 2 79%
× Small studies with wide CI
× Multiple unclear RoB
× Clinical heterogeneity
No difference ⊕⊕⊖⊖
MD -0.12 [-0.30, 0.05]
RCT= 4, n= 723
✓ I 2 0%
✓ 2 large RCTs with low RoB
× 1 small RCT with high RoB
× Clinical heterogeneity
No difference ⊕⊕⊖⊖
MD -0.18 [-0.49, 0.13]
RCT= 6, n= 551,
✓ I 2 0%
× Only 3 small RCTs (wide CI) with unclear RoB
× Clinical heterogeneity
No difference ⊕⊕⊖⊖
Meta-analysis not possible
✓ All 3 RCTs including 1 large one with low RoB reported no difference
No difference ⊕⊕⊖⊖
RR 1.07 [0.93, 1.23]
RCT= 3, n= 656
✓ I 2 0%
✓ No RCT contradicted
× 3 small RCTs, unclear RoB
No difference ⊕⊕⊖⊖
46.7% vs 42.1% [p>0.05]
RCT= 1, n= 412
✓ Only 1 RCT but large with low RoB
↓ with MS-IVF ⊕⊕⊖⊖
SMD -3.17 [-3.80 -2.54]
RCT= 13, n= 2314
✓ Large effect size
✓ No RCT contradicted
✓ 3 large RCTs low RoB
× I 2 96%
× RCTs with unclear/ high RoB
× Clinical heterogeneity
↓ with MS-IVF ⊕⊕⊖⊖
SMD of -5.86 [CI -7.06, -4.66]
RCT= 11, n= 2583
✓ Large effect size
✓ No RCT contradicted
× I 2 99%
× RCTs with unclear/ high RoB
× Clinical heterogeneity
↓ with MS-IVF ⊕⊕⊖⊖
SMD
-394.00 [-481.20 -306.80]
RCT= 1, n= 412
✓ Only 1 RCT but large with low RoB
⊕⊕⊕⊖, moderate quality of evidence; ⊕⊕⊖⊖, low quality of evidence; ⊕⊖⊖⊖, very low quality of evidence; RR, relative risk; MD, mean difference; SMD, standardised mean difference; RoB, risk of bias.
Normal responders. Three included studies reported LBRs (n = 573), all compared CC+ gonadotrophin ( Harrison et al. , 1994 ; Dhont et al. , 1995 ; Lin et al. , 2006 ) and long downregulation protocol. There was no difference in LBRs: RR 0.88 (CI 0.69, 1.12) ( Fig. 3B ). There was no statistical heterogeneity ( I 2 0%) and very little clinical heterogeneity between the trials. The finding did not alter in the sensitivity analysis. However, the large RCT had multiple areas of unclear RoB ( Dhont et al. , 1995 ); the other two were small trials ( Harrison et al. , 1994 ; Lin et al. , 2006 ); hence the QoE was moderate ( Table III ). The evidence with gonadotrophin only protocols for this outcome among normal responders was lacking.
Hyper-responders. Two large RCTs looked for livebirth, both were powered for their primary outcomes; both studies applied gonadotrophin only stimulation protocols ( Casano et al. , 2012 ; Oudshoorn et al. , 2017 ). The meta-analysis found LBRs did not differ between the groups: RR 0.98 (CI 0.79, 1.22). There was no statistical heterogeneity ( I 2 0%) and the QoE was moderate on GRADE analysis owing to clinical heterogeneity ( Table III ).
One RCT on poor responders reported OHSS rates ( van Tilborg et al. , 2017 ). The incidence was not significantly different between doses (1.8% with 150 IU dose vs. 1.2% with 225–450 IU dose, P = 0.45).
Normal responders. Nine RCTs (n = 1925) estimated OHSS rates: four small trials ( Tan et al. , 2005 ; Baart et al. , 2007 ; Lou and Huang, 2010 ; Blockeel et al. , 2011 ) and a large one ( Heijnen et al. , 2007 ) with gonadotrophin-only regimens showed lower incidence of OHSS (RR 0.43 (CI 0.21, 0.90)) in the MD-IVF group. Meta-analysis of four RCTs with oral compounds ( Dhont et al. , 1995 ; Lin et al. , 2006 ; Karimzadeh et al. , 2010 ; Mukherjee et al. , 2012 ), as well as all eight studies together, also found the risk of OHSS to be significantly lower with MD-IVF (RR 0.26 (CI 0.14, 0.49)) ( Fig. 4A ). Overall, the effect-size was large; there was no statistical heterogeneity ( I 2 0%) and the CI was narrow. Multiple studies had one or more areas of unclear RoB; in addition, clinical heterogeneity, including varied criteria for reporting OHSS, made this evidence of a moderate quality ( Table III ).
Forest plot of mild versus conventional-dose IVF: incidence of ovarian hyperstimulation syndrome per started cycle .
Hyper-responders. Both RCTs on hyper-responders were with a gonadotrophin-only regimen ( Casano et al. , 2012 ; Oudshoorn et al. , 2017 ). Meta-analysis of the pooled data found a significantly lower incidence of any grade of OHSS with MD-IVF, with a RR of 0.47 (0.31, 0.72) ( Fig. 4B ). There was no statistical heterogeneity ( I 2 0%) and no RoB. The QoE was moderate due to methodological diversity ( Table III ).
Poor responders. All 15 RCTs investigated CCRs (n = 3459). There was no difference in the risk of cycle cancellation between both arms, with an RR of 1.33 (CI 0.96, 1.85). The CCR was found to be higher when the trials with dose adjustments were excluded (RR 1.73 (CI 1.02, 2.93)) ( Fig. 5A ). The presence of significant statistical ( I 2 63%) as well as clinical heterogeneity, wide CI and unclear RoB in most trials led to a very low QoE for this outcome ( Table III ).
Forest plot of mild versus conventional-dose IVF: cycle cancellation rate per started cycle .
Normal responders. Seven RCTs ( Hohmann et al. , 2003 ; Out et al. , 2004 ; Tan et al. , 2005 ; Baart et al. , 2007 ; Heijnen et al. , 2007 ; Lou and Huang, 2010 ; Blockeel et al. , 2011 ) with gonadotrophin only regimen (n = 1430) found no difference in CCRs in the meta-analysis (RR 1.60 (CI 0.96, 2.67)), while pooled data from five RCTs comprising a CC+ gonadotrophin regimen ( Tummon et al. , 1992 ; Harrison et al. , 1994 ; Dhont et al. , 1995 ; Lin et al. , 2006 ; Karimzadeh et al. , 2010 ) (n = 1224) showed a higher risk of cycle cancellation with MD-IVF (RR 2.87 (CI 1.46, 5.64)) ( Fig. 5B ). However, when three trials that did not use GnRH-agonist or antagonist for LH suppression ( Tummon et al. , 1992 ; Harrison et al. , 1994 ; Dhont et al. , 1995 ) were taken out of the meta-analysis, the CCR became comparable with CD-IVF. Overall, the CCR was higher with MD-IVF: RR 2.08 (CI 1.38, 3.14). The I 2 indicating statistical heterogeneity was 48%. The QoE was very low, due to clinical heterogeneity, a wide CI and multiple small studies with unclear RoB ( Table III ).
Hyper-responders. CCRs were no different between MD-IVF and CD-IVF in both the RCTs with a gonadotrophin only agonist/antagonist protocol; RR 1.31 (CI 0.98, 1.77). Heterogeneity was absent ( I 2 0%), and studies were large with low RoB. This evidence, which was based on only two RCTs with diverse protocols, was of moderate quality.
This outcome was investigated in five RCTs in total (n = 2037): two with hyper-responders ( Casano et al. , 2012 ; Oudshoorn et al. , 2017 ), one with normal ( Heijnen et al. , 2007 ) and two with poor responders ( van Tilborg et al. , 2017 ; Liu et al. , 2020 ). All studies were large; all except one ( Liu et al. , 2020 ) were based on gonadotrophin only protocols. None of the individual RCTs found a difference between MD-IVF and CD-IVF protocols, and so the meta-analysis was of the pooled data (RR 0.96 (CI 0.86, 1.07)) ( Fig. 6 ). All studies had a low RoB, one had an unclear RoB, I 2 was 0% and the conclusion remained unchanged when the only study that allowed dose adjustment ( Casano et al. , 2012 ) was excluded from the meta-analysis. However, the trials were conducted in different population types and the description of ‘cumulative’ livebirth was different (as described above).
Forest plot of mild versus conventional-dose IVF: cumulative live birth rate per randomisation .
Poor responders. Seven RCTs (n = 2006) reported OPRs: three compared low-dose with high-dose gonadotrophin protocols ( Klinkert et al. , 2005 ; van Tilborg et al. , 2017 ; Youssef et al. , 2017 ), two trials with CC ( Martinez et al. , 2003 ; Revelli et al. , 2014 ) and the other two with Letrozole incorporated protocols ( Bastu et al. , 2016 ; Liu et al. , 2020 ). Meta-analysis of pooled data found no difference in OPRs: RR of 1.01 (CI 0.81, 1.25). There was no statistical heterogeneity between the studies ( I 2 0%), and all four large trials were of low RoB ( Revelli et al. , 2014 ; van Tilborg et al. , 2017 ; Youssef et al. , 2018 ; Liu et al. , 2020 ). However, due to two small RCTs having an area of ‘unclear RoB’ and clinical heterogeneity among the study protocols, the overall QoE was moderate ( Table III ). If smaller studies with ‘unclear RoB’ or studies with dose adjustments were excluded, the inference of the meta-analysis remained the same. The inclusion of large RCTs having low RoB strengthened the QoE in the subgroup comparing low-dose with high-dose gonadotrophin only.
Normal responders. Seven RCTs (n = 1026) in this category estimated OPR, six with gonadotrophin only stimulation ( Hohmann et al. , 2003 ; Out et al. , 2004 ; Tan et al. , 2005 ; Baart et al. , 2007 ; Lou and Huang, 2010 ; Blockeel et al. , 2011 ) and one was with CC combination ( Karimzadeh et al. , 2010 ). The RR of pooled data 1.10 (CI 0.88, 1.38) was not significant. Both statistical and clinical heterogeneity were low, and CI was narrow; however, this finding is based on predominantly gonadotrophin only protocols and presence of unclear RoB in multiple studies led to moderate QoE ( Table III ).
Hyper-responders. Only one RCT in this population reported OPR ( Oudshoorn et al. , 2017 ). This large RCT with low RoB found no difference on OPR with a RR 0.86 (CI 0.61, 1.23).
Poor responders. Twelve RCTs (n = 2211) on poor responders reported CPR. There was no significant difference in CPRs with an RR of 0.96 (CI 0.79, 1.16). The CI was narrow, statistical heterogeneity was absent ( I 2 0%) and the finding of the meta-analysis remained the same in sensitivity analysis. However, diversity in the clinical protocol resulted in a moderate QoE.
Normal responders . Three RCTs on gonadotrophin only ( Out et al. , 2004 ; Tan et al. , 2005 ; Lou and Huang, 2010 ) and four on oral compound+ gonadotrophin ( Harrison et al. , 1994 ; Lin et al. , 2006 ; Karimzadeh et al. , 2010 ; Mukherjee et al. , 2012 ) analysed the CPR. Meta-analysis showed no difference in the CPR between mild and conventional-dose arms (RR 1.10 (CI 0.92, 1.31)). Three RCTs reported PR, defined as positive pregnancy test (urine or serum β-hCG): of them one study reported a significantly lower PR per cycle with MD-IVF ( Dhont et al. , 1995 ) while the other two trials found no difference ( Hohmann et al. , 2003 ; Blockeel et al. , 2011 ). Two studies did not specify whether it was positive test or clinical pregnancy ( Tummon et al. , 1992 ; Elnashar et al. , 2016 ), and these five studies were excluded from the meta-analysis on CPR. There was no statistical heterogeneity ( I 2 0%) and CI was narrow; however, the studies were of small sample size with multiple ‘unclear RoB’ and diverse treatment protocols. Consequently, the QoE was low.
Hyper-responders
. Meta-analysis of two large RCTs with low RoB in this group found no difference in CPR (RR 0.91 (CI 0.82, 1.01)). Differences in the methodology made the QoE moderate.
Poor responders. All RCTs compared the number of oocytes retrieved; meta-analysis of 14 trials (n = 2773) that reported the mean number of oocytes found a significantly lower number of oocytes recovered in MD-IVF group, with an SMD of −0.43 (CI −0.58, −0.28). The other study ( Klinkert et al. , 2005 ) that expressed the figures in the median found no difference in the oocyte number. The effect size was large; but most of the studies had area(s) of ‘unclear RoB’ and one had an area of ‘high RoB’; there was significant statistical ( I 2 67%) and clinical heterogeneity, therefore, the QoE was low ( Table III ).
Normal responders. Meta-analysis from 13 RCTs (n = 3499) revealed fewer oocytes with low-dose stimulation (SMD −1.34 (CI −1.94, −0.75)), whether it was with a gonadotrophin only protocol (seven trials) or with CC (five trials). The only trial with Letrozole in the MD-IVF arm did not find any difference in the mean oocyte numbers ( Mukherjee et al. , 2012 ). The QoE, however, was very low in the presence of high statistical ( I 2 98%) and clinical heterogeneity, multiple RoB and wide CI ( Table III ).
Hyper-responders. Pooled data from two large RCTs with low RoB found no difference in the mean oocyte number (SMD −0.31 (−0.74, 0.13)). The QoE was low due to significant statistical ( I 2 91%) and clinical heterogeneity ( Table III ).
No data on this outcome for the hyper-responders were available.
Poor responders. Ten RCTs on poor responders compared total number of embryos created; a meta-analysis of nine of them (n = 1559) ( Goswami et al. , 2004 ; Kim et al. , 2009 ; Mohsen and El Din, 2013 ; Huang et al. , 2015 ; Bastu et al. , 2016 ; van Tilborg et al. , 2017 ; Youssef et al. , 2018 ; Yu et al. , 2018 ; Liu et al. , 2020 ) found a lower mean of total embryos with MD-IVF than CD-IVF (SMD −0.39 (CI −0.59, −0.20)), while the other trial expressed in the median (range) found no difference ( Heijnen et al. , 2005 ). Two large RCTs with low RoB used a gonadotrophin only regimen and found fewer embryos from MD-IVF (SMD −0.25 (CI −0.38, −0.12); the finding was the same with CC/Letrozole regimens (SMD −0.36 (CI −0.62, −0.10)), which were used in smaller studies with multiple areas of ‘unclear RoB’ and an area of high RoB. Overall, significant statistical ( I 2 59%) and clinical heterogeneity resulted in low QoE ( Table III ). Exclusion of small studies with RoB did not change the inference.
Normal responders. Seven RCTs (n = 1884) compared the mean of total embryos. Three trials were on gonadotrophin only protocols ( Tan et al. , 2005 ; Baart et al. , 2007 ; Heijnen et al. , 2007 ), the rest were a CC+ gonadotrophin regimen ( Tummon et al. , 1992 ; Harrison et al. , 1994 ; Lin et al. , 2006 ; Karimzadeh et al. , 2010 ). The mean of total embryos created was lower in the MD-IVF group (SMD −0.30 (−0.58, −0.08)). The difference was not significant in trials with CC+ gonadotrophin. Although the nature of the studies was more homogeneous, a high level of statistical heterogeneity ( I 2 79%), predominantly small studies with wide CI and trials with multiple unclear RoB made this evidence of very low quality ( Table III ).
Poor responders. Seven RCTs compared ‘top/high-grade’ embryos between MD-IVF and CD-IVF: four of them compared the mean number ( Kim et al. , 2009 ; Huang et al. , 2015 ; Youssef et al. , 2018 ; Liu et al. , 2020 )—meta-analysis of these trials (n = 723) showed no difference (SMD −0.12 (CI −0.30, 0.05)); three studies compared the proportion (%) of good-quality embryos ( Revelli et al. , 2014 ; Pilehvari et al. , 2016 ; Yu et al. , 2018 ). A meta-analysis was not possible due to unavailability of denominators. However, all three studies found the proportion of good-quality embryos to be no different between the two approaches. A large RCT (n = 640) with low RoB reported the proportion of embryos scoring >8 points to be 57.6% with MD-IVF and 54.8% with CD-IVF, the difference was not statistically significant ( Revelli et al. , 2014 ). Overall, clinical heterogeneity was significant, plus wide CI, and three studies had multiple areas of unclear bias (one had a high RoB ( Huang et al. , 2015 )); hence the QoE was low ( Table III ).
Normal responders. High-grade embryos were compared in six RCTs, three of them (n = 551) reported as mean ( Harrison et al. , 1994 ; Out et al. , 2004 ; Mukherjee et al. , 2012 ), and three (total population 656) as a proportion ( Baart et al. , 2007 ; Karimzadeh et al. , 2010 ; Elnashar et al. , 2016 ). All studies found no difference in mean or percentage of high-grade embryos. Meta-analysis of the mean number showed an MD of −0.18 (−0.49, 0.13) and the proportion of high-grade embryos showed an RR of 1.07 (0.93, 1.23). Although statistical heterogeneity was absent ( I 2 0%), this evidence is based on mostly clinically heterogenous small trials with multiple unclear RoB and was therefore of low quality ( Table III ).
Hyper-responders. Only one large RCT reported the proportion of high-grade embryos to be 46.7% versus 42.1% ( P > 0.05) in MD-IVF and HD-IVF groups, respectively ( Casano et al. , 2012 ).
Poor responders. All included RCTs, except the one that did not use any gonadotrophin in the MD-IVF protocol ( Ragni et al. , 2012 ), compared total amount of gonadotrophin used between the groups. There was a high level of clinical and statistical heterogeneity ( I 2 96%) among the studies and many RCTs had area(s) of ‘unclear bias’. However, all individual trials found less gonadotrophin requirement in the MD-IVF programme, with a large effect-size (SMD −3.17, CI −3.80, −2.54) in the meta-analysis of 13 trials that measured the mean of stimulation dose ( Table III ). The two RCTs reporting the median dose also found the same result ( Klinkert et al. , 2005 ; Liu et al. , 2020 ).
Normal responders. Eleven trials (n = 2583) compared gonadotrophin use among normal responders, four of them used gonadotrophin only protocols ( Out et al. , 2004 ; Tan et al. , 2005 ; Heijnen et al. , 2007 ; Lou and Huang, 2010 ); five studies used CC+ gonadotrophin as MD-IVF ( Tummon et al. , 1992 ; Harrison et al. , 1994 ; Dhont et al. , 1995 ; Lin et al. , 2006 ; Karimzadeh et al. , 2010 ); the remaining two trials were based on Letrozole ( Mukherjee et al. , 2012 ; Elnashar et al. , 2016 ). All individual studies reported lower gonadotrophin use with MD-IVF. Meta-analysis found an SMD of −5.86 (CI −7.06, −4.66) with a large effect size. However, significant statistical ( I 2 99%) and clinical heterogeneity, along with studies with unclear RoB, led to a low QoE ( Table III ).
Hyper-responders. One large RCT with low RoB that compared gonadotrophin dose found a lower total dose used in the MD-IVF group: MD −394.00 (CI −481.20, −306.80) ( Casano et al. , 2012 ).
Two included RCTs on poor responders ( Ragni et al. , 2012 ; van Tilborg et al. , 2017 ), three among normal ( Heijnen et al. , 2007 ; Lou and Huang, 2010 ; Mukherjee et al. , 2012 ) and one in the hyper-responder patient-category ( Oudshoorn et al. , 2017 ) performed a cost-analysis. Both trials on the poor responder were large with low RoB. One study found MD-IVF was associated with a per-cycle cost-saving of € 2620 with no use of gonadotrophin in the MD-IVF arm and a 450 IU daily dose in the CD-IVF arm ( Ragni et al. , 2012 ). The other RCT reported a reduced cumulative treatment cost with the lower gonadotrophin dose regimen by €1099 ( van Tilborg et al. , 2017 ). Cost with standard deviation was not reported in both the studies; therefore, a meta-analysis could not be performed. A larger RCT among the normal responders found a cumulative cost difference of €-2412.00 ( Heijnen et al. , 2007 ). The other study by Lou reported the treatment cost of 1056 ± 111 and 16 776 ± 3921 yuan (€136 ± 14.3 versus €2160.04 ± 505) ( P < 0.001) for mild and conventional treatment, respectively ( Lou and Huang, 2010 ). Converting yuan to euro at the current conversion rate, a meta-analysis of these two trials on normal responders also found MD-IVF to be less expensive (MD −2028.21 (CI −2208.00, −1848.41)). The RCT by Mukherjee et al. reported 34% less average cost with letrozole-based protocol. The only RCT on the hyper-responders, however, did not find any approach cheaper than the other ( Oudshoorn et al. , 2017 ). Overall, the study protocols including health-economic models were different between the trials, hence the QoE was low.