Case
When the outcomes of research can affect policy and clinical decisions, researchers must be especially considerate of bias affecting their results ( Lash et al. , 2014 ). One example of where this research/policy interaction occurs is in the study of the relationship between MAR and later health outcomes. For example, observational studies have linked MAR to cardiovascular disorders, diabetes, thyroid disease, obesity, and cancer ( Dayan et al. , 2017 ; Vassard et al. , 2018 ; Murugappan et al. , 2021 ; Cullinane et al. , 2022 ; Yiallourou et al. , 2022 ; Magnus et al. , 2023 ; Sachdev et al. , 2023 ; Wei et al. , 2023 ). Given that more than 1.9 million IVF cycles are performed each year globally, with a similar number of ovulation induction cycles estimated ( Chambers et al. , 2021 ), even a slight increase in the risk of poor health outcomes following such treatments may represent a significant burden of disease and concern for patients and clinicians.
To illustrate how bias may impact conclusions about MAR and later health outcomes, we provide an example using the relationship between MAR and cancer. There has been long-standing concern that the use of MAR may increase the risk of certain cancers, particularly hormone-sensitive ovarian, endometrial, and breast cancer ( Fishel and Jackson, 1989 ).
Over the past 30 years, numerous studies in different countries have attempted to elucidate the relationship between MAR and cancer risk by analysing large-scale linked administrative health datasets, with meta-analyses published almost yearly summarizing the association ( Gennari et al. , 2015 ; Momenimovahed et al. , 2019 ; Barcroft et al. , 2021 ; Beebeejaun et al. , 2021 ; Cullinane et al. , 2022 ). Though recent research has generally suggested limited to no association between hormone-related cancers and MAR ( Barcroft et al. , 2021 ), there is still substantial debate on the topic owing to the gravity of the concern for women seeking MAR treatment, the biological plausibility of associations ( Fathalla, 1971 ; Auranen et al. , 2005 ; Huhtaniemi, 2010 ; Liang and Shang, 2013 ; Brown and Hankinson, 2015 ), and the potential for random error and systematic bias in studies reliant on secondary use of administrative datasets ( Momenimovahed et al. , 2019 ).
The reported effect sizes from studies showing a positive association between MAR and cancer are generally small ( Barcroft et al. , 2021 ), indicating that even a small amount of bias (such as confounding or selection bias) may affect the outcome of a given study. Furthermore, there are many potential confounders that can affect both exposure to MAR and the development of cancer in women. These can generally be split into four groups: health history (e.g. infertility, anovulation, endometriosis, cancer screening, smoking, excessive alcohol consumption, oral contraceptive use); reproductive history (e.g. parity, age at first/last birth); social factors (e.g. socioeconomic advantage, rurality); and familial history (e.g. genetic predisposition to cancer). Owing to the challenges associated with collecting, accessing, linking, and analysing population-based data, to our knowledge no existing studies have accounted for all known confounders in their analyses ( Momenimovahed et al. , 2019 ; Bates, 2021 ). To illustrate how confounding can cause bias, consider the simple directed acyclic graph (DAG) in Fig. 1 showing the relationship between MAR, endometriosis, and cancer. If we conducted a cohort study, we might find an increased risk of cancer associated with MAR treatment. However, if we did not account for the impact of endometriosis, we could not be sure if this effect is driven by a causal effect of MAR on cancer, or by the fact that those with endometriosis are both more likely to receive MAR treatments and develop certain types of cancer, particularly ovarian cancer. An expanded DAG showing proposed relationships between all mentioned confounders is provided in Supplementary Fig. S1 and justifications for each causal link are provided in Supplementary Table S1 .
Example of the confounding relationship between endometriosis, MAR, and cancer. Causal direction is indicated by the arrows. MAR: medically assisted reproduction. Figures first made with the help of Dagitty ( https://dagitty.net ) ( Textor et al. , 2016 ).
As well as unmeasured confounding, accurately measuring factors associated with receiving MAR can be difficult. For example, it is likely that infertility is underreported in administrative medical records, as it is estimated that over half of all couples do not seek treatment for infertility ( Boivin et al. , 2007 ). Similarly, there may be substantial underreporting of endometriosis in administrative medical records, with one study showing that over 40% of suspected endometriosis cases were not captured in such records ( Rowlands et al. , 2021 ). Given the potential for bias, it is not surprising that many meta-analyses on MAR and cancer report a high proportion of total variability caused by between-study heterogeneity ( Gennari et al. , 2015 ; Del Pup et al. , 2018 ; Cullinane et al. , 2022 ).
Causal
A central challenge for epidemiology is causal inference, that is how to determine if two events are causally related.( Pearl, 2009 ) Generally, when trying to answer a causal question, researchers must deal with two sources of error: random error and bias ( Hernan and Robins, 2020 ). Random error describes how, when sampling data from an underlying distribution, the samples will exhibit some random deviation from the true mean value of that distribution. If this deviation is too high, and the number of samples is too low, then true relationships in the data will not be discernible. Bias describes how the observed relationship between two variables differs systematically from the true value owing to factors other than the causal link between the two ( Hernan and Robins, 2020 ). Bias can be split into three components ( Lash et al. , 2014 ):
Selection bias : bias introduced through the tendency for some types of participants to be sampled more or less often than others relative to the exposure and outcome.
Confounding : bias introduced through factors that influence both the exposure and the outcome (‘confounders’), such that it produces a non-causal association, or masks a true causal association, between the exposure and outcome. Measured confounders are those known and measured, and their impact on the true causal association between exposure and outcome is controlled for in the research design and/or analysis. Unmeasured confounders are those either unknown to the researcher, or known but unmeasured (e.g. when data on a confounder is too difficult to collect).
Measurement error : bias introduced through the incorrect measurement of the outcome, exposure, or any confounding or mediating variables.
Randomized controlled trials (RCTs), though not immune to selection bias, address confounding through randomization of participants to treatment/non-treatment conditions, and can attempt to minimize issues with measurement through methods such as blinding. However, practical and ethical concerns can often restrict the use of RCTs in research. A prime example of this restriction is the study of medically assisted reproduction (MAR), such as ovulation induction and IVF, and health outcomes. That is, if one is interested in a health outcome that generally takes a long time after treatment to manifest (e.g. dementia), it is not only impractical, but potentially unethical, to ask women not to seek a treatment that increases the likelihood of conception in their childbearing years. When conducting an RCT is not possible, researchers must rely on observational data. As there is no inherent randomization, studies using these data can be subject to substantial bias ( Lu, 2009 ).
Though both random error and bias are important to consider in the design of observational health studies, random error is generally measured through established statistical methods (such as the generation of confidence intervals) ( Rothman et al. , 2008 ; Barraza et al. , 2019 ; Fox et al. , 2022 ), while bias is rarely quantified ( Petersen et al. , 2021 ). Furthermore, the consequences of random and systematic error are not equivalent. Though both random error and bias can decrease researchers’ ability to observe a true causal relationship between an exposure and an outcome, bias can also often lead to erroneously concluding an exposure/outcome relationship exists when it does not, and in extreme cases may reverse the direction of the true causal relationship ( Barraza et al. , 2019 ). Therefore, the assessment and, ideally, quantification of bias is an important step when assessing the internal validity of studies and the strength of evidence they produce.
Despite both the importance of understanding sources of bias in observational designs and the availability of various methods to assess such error, estimation of bias remains rare in the epidemiological literature ( Petersen et al. , 2021 ). This may be linked to a lack of awareness about when estimating bias is useful, or knowledge of how bias may affect research outcomes. However, even when researchers are aware of potential sources of bias, they may be unaware of the available methods to estimate it. To address this concern, in this article, we provide an overview of two common approaches to estimating the potential impact of bias: quantitative bias analysis; and E -values. We then show how these techniques can help assess the strength of evidence for health outcomes following MAR using cancer risk as an example. In doing so, we hope to convince the reader that such methods are an important step to consider when conducting observational research, particularly in the study of health outcomes following MAR.
Section
The E -value is simple and quick to compute using the formulae provided by VanderWeele and Ding (2017) , even without person-level data. We calculate some E-values here for demonstration. Table 1 shows the E -values for selected results for four recent studies on women’s cancer and MAR, based on the hazard ratios reported in the papers, and potential uncontrolled confounders for each study ( Reigstad et al. , 2017 ; Lundberg et al. , 2019 ; Lindquist et al. , 2022 ; Sandvei et al. , 2023 ).
Selected E -values calculated from recent studies on assisted reproductive therapy and cancer.
E -value CI derived from values of hazard ratio CI.
E -values calculated using the E -value calculator at: https://www.evalue-calculator.com/ ( VanderWeele and Ding, 2017 ; Mathur et al. , 2018 ). All calculations assume the cancer outcome occurs in fewer than 15% of cases (to assume a reasonable approximation of risk ratios by hazard ratios).
From Table 1 , two points are evident. First, even in recent studies, there is variability in the potential confounders that have been controlled. Though most studies adjust for age and parity, they variably adjust for other key confounders like smoking, weight, and familial cancer history. Second, in most studies, it would only take one moderately strong uncontrolled confounder to nullify any observed association, particularly if the E -value is anchored to the lower bound of the confidence interval, as this represents the minimum level of confounding required to nullify the observed result. Though many studies list potential uncontrolled confounders in their limitations, E -values can be directly compared against existing literature. For example, endometriosis is not controlled for in any study in Table 1 , despite being variably linked to both the development of cancers in Table 1 ( Kvaskoff et al. , 2020 ; Ye et al. , 2022 ), and decreased fertility ( Prescott et al. , 2016 ; Thong et al. , 2020 ; Mattsson et al. , 2021 ).
If we consider the study on thyroid cancer by Lindquist et al. (2022) , the lower-bound E -value of the relationship between MAR and thyroid cancer is 1.34. This E -value indicates that a relationship of 1.34 on the risk ratio scale between both endometriosis and MAR and endometriosis and thyroid cancer, would be sufficient to drive the observed association between MAR and thyroid cancer. We first note that the estimated risk ratio (compared to the general population) for infertility for women with endometriosis is 2.12 ( Prescott et al. , 2016 ), and the estimated risk ratio for thyroid cancer for women with endometriosis is 1.39 ( Kvaskoff et al. , 2020 ). Given that both risk ratios are above the lower-bound E -value for the relationship between endometriosis and thyroid cancer, and endometriosis and infertility, it is plausible that endometriosis may be driving the observed relationship between MAR and thyroid cancer. By contrast, consider the relationship between MAR and ovarian cancer, where the reported lower-bound E -values are between 2.34 and 2.85. As the risk ratio for ovarian cancer for women with endometriosis is around 1.93 ( Kvaskoff et al. , 2020 ), it is unlikely that endometriosis is solely responsible for driving an observed association between MAR and ovarian cancer.
After estimating the potential for bias to affect the results through an E -value, the question remains of how our interpretation of the findings should change. It is important to state that efforts to measure the potential for bias should not serve to increase confidence in findings, given a major benefit of measuring bias is to counteract overconfidence in research findings ( Lash et al. , 2014 ). Furthermore, bias estimates are only interpretable in the face of domain-specific knowledge. That is, a high E -value for studies on ovarian cancer does not in itself mean the results are more likely to be true, particularly if there are plausible strong unmeasured confounders that may be affecting the results. Similarly, a low E -value should not necessarily decrease confidence in the results, especially if there are no plausible unmeasured confounders that could drive the effect. However, in cases where there are theoretically plausible unmeasured confounders, we should be particularly wary of low E -values. Based on our previous assessment of the association between endometriosis, MAR, and thyroid cancer, we may be more sceptical that a true causal relationship exists between MAR and thyroid cancer, or at the very least recommend that future studies attempt to control for endometriosis. Importantly, we may be more cautious about recommending against undertaking MAR treatments because of an increased risk of thyroid cancer.
Approaches
Substantial work has gone into developing ways to estimate the impact of bias. One approach, known as ‘quantitative bias analysis’, describes a set of approaches in which the researcher assesses how the relationship and strength of the causal relationship between an exposure variable and an outcome variable may change when one or more biases are explicitly assumed and modelled ( Lash et al. , 2014 ; Fox et al. , 2022 ). These methods range from mathematically and computationally straightforward (e.g. simple bias-sensitivity analysis) to highly complex (e.g. multiple bias modelling) ( Lash et al. , 2014 ; Fox et al. , 2022 ), and depending on the method chosen can be used to assess the principal types of bias (selection bias, unmeasured confounding, and measurement error) ( Lash et al. , 2014 ).
Another method of estimating the impact of bias caused by unmeasured confounding that has been proposed in causal questions is the ‘ E -value’ ( VanderWeele and Ding, 2017 ). Rather than assessing how the strength of a causal association may change by adjusting for different assumed biases, the E -value provides an absolute measure of the minimum association (in the risk ratio scale) an unmeasured confounder would need to have on both the exposure and the outcome to potentially reduce any observed association to the null, thus ‘explaining away’ any observed effect. In comparison to quantitative bias analysis, the E -value is easily computed and can be reported for a range of study designs. However, it does not account for either selection bias or measurement error nor does it explicitly consider that there may be multiple interacting unmeasured confounders ( Greenland, 2020 ).
The general equation for calculating an E -value (on the risk ratio scale) is given by
E = RR + R R RR - 1 ,
where E is the E -value and RR is the risk ratio of the exposure and outcome. To give an example of how the E -value is used, consider a scenario where the risk ratio of an association between an exposure and outcome is 1.50. Using the equation, we calculate our E-value to be 2.37. This E -value tells us that, for this observed association between the exposure and outcome of 1.50 to be reduced to the null, an unmeasured confounder would need to have a causal association on the risk ratio scale of strength at least 2.37 to the exposure and 2.37 to the outcome.
Discussion
Estimating and reporting the potential impact of bias is an important step to consider when conducting causal research on observational data, particularly when the conclusions may have implications for health policy and practice recommendations. Though overall the reporting of bias has been gaining traction, it has not been widely taken up in the epidemiological literature. Due to the risk of bias influencing the results of observational studies on MAR and fertility, it is imperative that researchers are aware of the potential pitfalls of bias, and subsequently the ways to estimate its impact on study findings. To promote the explicit consideration of bias in epidemiological studies, we outline two methods (quantitative bias analysis and E -values) researchers can use to estimate the potential impact of bias in their studies, and show how these might be used to enhance our understanding of the true relationship between MAR and later health outcomes (using cancer as an example).
Importantly, the point of estimating bias is not just to conduct better research but also to assist in the development of clinical practice guidelines. For example, reported links between MAR and certain health outcomes may impact both the health messaging doctors provide to prospective MAR patients and their decision to provide the treatment based on their patients’ risk profile for those outcomes. This decision may have life-changing consequences for the patient, who may decide not to undergo treatment based on the advice they receive. Therefore, such advice should be given with careful consideration to the weight of evidence provided by research findings. Estimating bias is therefore an important step in assessing this weight of evidence for clinicians and policy-makers, who can then adjust their recommendations accordingly.
We must emphasize that estimating the impact of bias is not a silver bullet to improving study conclusions. Any methodological limitations that existed prior to considering bias continue to exist once that error is estimated. Furthermore, estimating the impact of bias should only ever serve to decrease our confidence in observed results ( Lash et al. , 2014 ), and estimates of bias should not be used to ‘rule out’ potential sources of bias. Though one unmeasured confounder may not appear strong enough to drive an effect alone, it may be part of a larger set of uncontrolled confounders that are consequential. However, estimating bias does provide a mechanism to better assess the strength of evidence provided, beyond looking simply at whether studies have a large population capture or high precision in coefficient estimates.
It is also important to stress that this piece is not a criticism of existing work on cancer and MAR. Rather it serves as an example of where quantifying bias may be particularly important for stakeholders. Especially in the case of unmeasured confounding, it is often impossible to accurately determine values of various confounders from administrative datasets (e.g. age of menarche or family history of cancer) that are relevant to questions of cancer and MAR. We stress that an inability to control for unmeasured confounding does not mean such studies are not useful. However, the ongoing question of MAR’s relationship to cancer does highlight that measuring bias may be particularly informative when effect sizes are generally small and implications for health outcomes large.
Though we have presented both quantitative bias analysis and E-values as methods to estimate bias, there are also other approaches. For example, the use of negative controls, where researchers model an outcome expected to have the same potential sources of bias as the outcome of interest, but no relationship to the exposure, is another method of estimating the impact of bias ( Arnold et al. , 2016 ; Shi et al. , 2020 ). In general, simpler estimation approaches can be informative, but tend to be more limited in the scope of error they can account for. For example, E -values may give us information on unmeasured confounding, but they do not address complex causal structures, measurement errors, or selection bias. By contrast, more complicated approaches can be more informative, but also prohibitively difficult to run, requiring both specialist expertise and sometimes vast computational resources. It is worth noting proponents of quantitative bias analysis acknowledge the difficulties in implementation inherent in such methods and are making efforts to increase the accessibility of such analyses (such as through the distribution of readily modifiable implementation code) ( Fox et al. , 2023 ).
Researchers who choose to report E -values should be aware of their criticisms. E -values are fundamentally theory agnostic, incorporating only information about the observed risk ratio association between the exposure and outcome ( Greenland, 2020 ). This fact can lead to both researchers being overly optimistic and pessimistic about their research findings ( Sjölander and Greenland, 2022 ), as they may incorrectly assume a low E -value automatically indicates a problematic result (even when strong confounders have been controlled for) ( Greenland, 2020 ), or a high E -value indicates a trustworthy result (even when there may be multiple strong uncontrolled confounders). It is therefore critical researchers do not use simpler methods like E -values as a proxy for theory development. Rather, they should be used in conjunction with careful consideration of background knowledge and research, as would be required if one were to perform quantitative bias analysis.
A major strength of quantitative bias analysis over E -values is that any uncertainty in the true relationships between confounders, exposure, and treatment can be explicitly modelled. Though this strength means that quantitative bias analysis can provide a more comprehensive test of bias than E -values, there is no guarantee the researchers’ model of these relationships is correct or useful, and there is no way to test this empirically ( Lash et al. , 2014 ). As a poorly specified bias model may lead to incorrect conclusions about the level of bias, researchers must be careful when designing and implementing these bias analyses. For example, in this commentary, we have used infertility owing to endometriosis as a proxy for receiving MAR treatment, but not everyone who experiences infertility will receive MAR treatment. As such, it may be important to include this pathway from infertility to MAR treatment in a bias model. If you were comparing only infertile women who received and did not receive MAR (an approach taken for the Dutch OMEGA cohort, for example) ( van Leeuwen et al. , 2011 ), this extra step may not be necessary. Despite the risk of poorly specified bias models, we believe the benefits outweigh the risk assuming researchers think carefully about the specific problem they are studying.
In conclusion, quantitative bias analysis and E -values are two methods that can be used to assess how bias may affect research conclusions in observational studies. We argue this assessment is important in questions where there are potentially many sources of bias (such as studies on health outcomes associated with MAR), particularly where the findings may have important implications for health policy and clinical practice. They are not a remedy for all methodological problems faced in observational research but can help to encourage researchers and readers to more carefully consider the strength of evidence provided to arguments around the presence or absence of causal associations.
Quantitative
Quantitative bias analysis could be used to assess the impact of this bias by showing the precise change in the effect estimate assuming different strengths of confounder associations between MAR and cancer risk. For example, consider how the inability to control for endometriosis may bias an association between MAR and ovarian cancer in an Australian cohort. To test this association, we could take the following steps:
Using either expert knowledge or prior literature, estimate the population prevalence of endometriosis (the confounder): we know around 10% of women of childbearing age have endometriosis in Australia ( Australian Institute of Health and Welfare, 2023 ).
Using prior literature, estimate the risk ratio of endometriosis to ovarian cancer: according to a systematic review-informed meta-analysis, women with endometriosis have a risk ratio of around 1.93 for ovarian cancer compared to women without endometriosis ( Kvaskoff et al. , 2020 ).
Using prior literature, estimate the risk ratio of endometriosis to receiving MAR: one prospective study reported the risk ratio of infertility for women with endometriosis to be around 2.12 ( Prescott et al. , 2016 ).
Examine how much the estimated effect of MAR on cancer would change if we had accounted for confounding between endometriosis, MAR, and ovarian cancer.
A simplified working example of this process with numerical estimates can be seen in Supplementary Data File S1 . Taking these steps, we might find that any effect observed between MAR and ovarian cancer is attenuated when accounting for the higher rates of endometriosis in women who received MAR. Such a finding would lessen our confidence that we have observed a true causal effect between MAR and ovarian cancer, and not an effect observed caused by the confounding effect of endometriosis.
Quantitative bias analysis allows us to check many sources of bias in an analysis. For example, given not everyone who is infertile receives MAR, we might choose to test multiple plausible risk ratio values (say between 1.50 and 2.50) for the relationship between MAR and endometriosis and observe how the overall effect of MAR on the risk of ovarian cancer changes. If we had endometriosis in our dataset, but at a lower prevalence than the population level (perhaps because of underreporting), we could test for measurement error by simulating an increased proportion of endometriosis.
It is also possible to use quantitative bias analysis to test for more complicated types of confounding. For example, we could model the joint impact of two confounders (endometriosis and smoking), or model two confounders where one is assumed to affect both the exposure and the outcome, and the other confounder (e.g. age at menarche and endometriosis). Both these bias models can be seen in Fig. 2 . In either case, the process of doing quantitative bias analysis is the same—use prior knowledge to assume the strength of the relationship between the confounders, the exposure, and the outcome, and then see how your estimate of the effect of exposure on the outcome changes when accounting for this confounding.
Example of confounding bias models that could be tested using quantitative bias analysis. ( A ) Two confounders affecting exposure and outcome. ( B ) Two confounders where both affect the exposure and the outcome and one confounder also affects the other confounder. Figures first made with the help of Dagitty ( https://dagitty.net ) ( Textor et al. , 2016 ).
The main limitation of quantitative bias analysis is it tends to require a substantial amount of time and statistical expertise. Furthermore, though simple analyses can be conducted with summary data, more complex analyses (such as those involving person-time calculations) require access to person-level data ( Fox et al. , 2022 ). An alternative to quantitative bias analysis is to use E -values.
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.