Methods
This review follows the guidelines set by the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) ( Appendix 1 ), the protocol for the review was not registered in advance. 34
Searches were conducted in PubMed and Web of Science for articles without date restrictions. The last search was conducted on June 1, 2021. A combination of MeSH and general terms were used to develop a search in PubMed ( Appendix 2 ) which was then amended for Web of Science. No date or geographic restrictions were used. We also searched bibliographies and used citation tracking in Scopus to identify other potentially relevant studies missed by the database searches.
Study inclusion criteria were determined a priori. To be eligible, peer-reviewed health-related computational modeling studies published in English needed to include one or more probabilistic model inputs generated through a de novo expert opinion elicitation process. Health-related modeling studies were defined as those in which the model reports a health outcome in humans. We consider only probabilistic inputs, such as transitions between different behaviors or health risks. As such, studies that elicited only cost estimates, resource use estimates, or utility values (such as quality-adjusted life years or disability-adjusted life years) were excluded. Additionally, studies were excluded if the elicitation process was previously reported in an earlier publication that meets the inclusion criteria. There were no restrictions on the methods used for elicitation. Studies that state the use of expert opinion but reported no information on how it was conducted or used were excluded.
Following the removal of duplicates, all study titles and abstracts were screened to determine eligibility by one reviewer (CC). Abstract and title screening was purposefully inclusive and any abstract that was deemed potentially relevant, or where relevance was not clear was included. The full texts of selected studies were then screened to determine final eligibility by two reviewers and reasons for exclusion were documented (CC & MK). Any uncertainty over exclusion was discussed and resolved by consensus.
A data abstraction template was developed in Excel prior to data abstraction. Two reviewers independently abstracted relevant information from selected studies (CC, MK). Selected studies were classified into two groups, those that conducted a formal EE, and those that used indeterminate methods to elicit expert opinion. The determination was based solely on the information available in the study report and appendices. The former group, classified as “formal EE methods”, consists of studies that present a formal method to elicit and synthesize estimates from experts. The latter group, classified as “indeterminate EE methods”, includes studies that reported using experts to provide model inputs, but did not present a formal process, or provided insufficient details to determine the process by which expert judgments were elicited and synthesized. These studies are characterized by limited information regarding how expert judgements were elicited, collected, and collated.
For those classified as formal EE studies, four types of information were abstracted: 1) study-specific details (authors, date, location, study design, modeling method, and disease or condition of interest); 2) expert selection methodology (the number of, definition of, and identification and selection process for experts); 3) elicitation methods (the elicited parameters, elicitation process, aggregation methods, uncertainty estimation); and 4) incorporation of elicitation results (type of model input, and how uncertainty was modeled). Given the limited amount of information provided in studies classified as indeterminate EE methods, it was not possible to consistently extract study characteristics for comparison. We abstracted 1) study-specific details (authors, date, location, modeling method, and disease or condition of interest); 2) expert opinion details (elicited parameters, details on experts and methods, other sources used to complement expert opinion).
There is no widely accepted standard for assessing the quality of an EE used for generating input parameters of a health-related modeling study. Recent work by Bojke et al . 18 has presented a framework, but consensus has not yet been reached on how to assess EEs. Therefore, we assessed the reporting quality of EEs based on an amended version of the reporting guidelines for expert opinion from Iglesias et al. 9 Studies received one point for each of the criteria they satisfied for a total of 16 points ( Appendix 3 ). Briefly, these criteria included a description of the research rationale, the definition of an expert and expert selection criteria, the development of a structured elicitation protocol, the piloting of the elicitation process, as well as clearly reporting the methodology for administering the EE, collecting data, aggregating data, and the use of performance measures. Finally, studies needed to present and interpret their EE results in the context of their model. Given the limited information provided by studies classified as indeterminate, it was not feasible to conduct a similar quality assessment.
Following full-text review and data abstraction, direct comparisons between studies were made to highlight the differences between formal EE studies and those that used indeterminate methods. Comparisons also illustrate the range of formal EE methods currently used to recruit experts, elicit data, aggregate responses, and account for uncertainty. Methods were compared qualitatively across studies to highlight the various ways in which expert opinion was utilized. For those classified as indeterminate EE methods, we present a summary of the types of experts and number of experts consulted.
Results
Searches identified 1,520 unique studies. The screening process and reasons for exclusion are outlined in the PRISMA diagram ( Figure 1 ). Following title and abstract review, 355 articles were deemed as potentially eligible. The full-text review excluded 203 studies. The most common reason for exclusion was that the study did not use computational modeling (n=84). An additional 36 studies did not report health outcomes, while 34 used expert elicitation to generate model inputs that were not probabilities (cost estimates, resource use estimates, or utility values). Notably, 12 studies were excluded because they cited expert opinion as a source of model input parameters in the abstract of the paper but provided no further information in the full text.
In total, 40 studies met the inclusion criteria for using formal EE methods for model inputs. An additional 112 studies used indeterminate methods to elicit expert opinion.
Of the 40 EEs studies included ( Table 1 ), the most common study types were cost-effectiveness, cost-utility, or cost-benefit analysis (n=29). Other study types included the evaluation of hypothetical policy interventions (n=2), EE methods papers that included modeling (n=3), and a series of one-off study types, such as a risk management analysis or value of information analysis. The type of computational model used in the majority of studies was decision-analytic (n=24). 35 A wide range of diseases/conditions were examined with no single one being most prevalent.
Experts were generally defined by their topical experience. For instance, clinical experts on the health condition of interest were selected in 35 studies ( Table 2 ). In seven studies, non-clinical experts were recruited. These included epidemiologists, health economists, and academic researchers who were consulted in combination with clinicians in five studies. 12 , 36 – 39 Studies recruited ‘academic researchers’ in two instances. 13 , 14 One study defined experts as medical, pharmacy and public health students who were recruited along with ‘professional researchers.’ 40
Expert selection methods were unreported in 20 of the 40 studies. Eight studies reported a selection method that used convenience sampling in which experts were selected due to their proximity to the study team. 41 – 48 Of the remaining studies, four used publication records; 13 , 14 , 49 , 50 two used prescribing records to identify clinicians who have used the medications being evaluated; 51 , 52 two consulted external advisors; 53 , 54 and the rest used snowball sampling; 50 , 55 trial involvement; 12 all certified obstetrician/gynecologist in the country of study with years of experience; 56 or international reputation. 57 The average number of experts recruited per EE was 11 with a range of 1 to 30. No studies gave reasons for their choice regarding the number of experts recruited. Additionally, it was not possible to verify the suitability of experts or to assess the potential for bias introduced in the selection process, because most studies did not provide sufficient details on the selection process of experts or provide a list of the experts.
Elicitation methods for each included study are reported in Table 3 . No methods describing the process for implementing the EE were reported in three studies. 37 , 58 , 59 Among the remaining studies that did report elicitation methods, considerable variation existed regarding the level of detail included in the methods descriptions. The most frequent methods were Delphi or modified Delphi process (n=10) 13 , 36 , 49 , 54 , 56 , 57 , 60 – 63 while two used the SHELF method, 38 , 50 both methods rely on multiple rounds of discussion to build consensus among experts. Elicitations were frequently done virtually or over the phone (n=13) 14 , 36 , 43 , 49 , 51 , 52 , 55 , 56 , 60 , 64 – 67 while only five specified that the elicitation was conducted in-person. 39 , 41 , 45 , 64 , 68 Web- or spreadsheet-based elicitation tools were used to facilitate data collection in 9 elicitations. 12 , 40 , 45 , 47 , 50 , 64 , 67 , 69 , 70
Eleven studies elicited point estimates from experts. Of those, the most frequent estimates were the median (n=5), mean (n=3), and ‘most likely value’ (n=3). Expert uncertainty was elicited by asking experts to provide ranges or different quantiles in seven studies. While seven other studies used the histogram method or a modified histogram method (where experts mark a frequency chart) to elicit expert uncertainty. 41 , 45 , 47 , 50 , 67 , 69 , 70
Aggregation methods for each included study are reported in Table 3 . No aggregation method was reported in ten studies, 36 , 39 , 54 , 58 – 61 , 65 , 68 , 71 one study did not aggregate responses as only one expert was consulted, 67 and one study modeled individual expert responses separately in lieu of aggregation. 64 Among those that reported aggregation methods, four used behavioural aggregation (where experts were asked to reach consensus) 37 , 49 , 63 , 72 and the remainder used mathematical aggregation. The most frequently used form of mathematical aggregation was linear pooling (n=18). 12 , 40 , 41 , 43 – 47 , 50 – 53 , 56 , 57 , 62 , 66 , 69 , 70 Performance measures, 73 , 74 where expert’s estimates are weighted based on their responses to seed variables (also known as calibration questions) with answers known to the facilitator but unknown or not readily available to the experts, were only used in two studies. 64 , 66 No other weighting, such as by expert’s reported level of confidence, was used. Those that did not use linear pooling reported a diverse set of aggregation methods. These included aggregation based on a simulated distribution of expert estimates such as through the use of Latin Hypercube sampling 14 , random-effects meta-analysis, 50 or Monte Carlo simulation. 42 One study compared three alternative pooling approaches and concluded that differences in results based on pooling method existed, but that it remains unclear which method was preferable. 66
Methods that modelers used to account for the uncertainty in EE-based inputs are reported in Table 3 . In most cases (n=26), one-way or probabilistic sensitivity analysis was used to evaluate the uncertainty of outcomes due to elicited parameters by varying the input value across expert estimates or using a distribution generated from the estimates. Another approach was to simulate the responses of each expert individually (n=3). In seven studies, sensitivity analyses were not conducted.
We found substantial variation in the reporting quality of EEs ( Appendix 4 ). The average score based on our 16-point quality scale was 9.4. The most underreported characteristics were lack of an explicit reference to an elicitation protocol (n=3), reporting whether or not performance measures were used (n=6), and lack of piloting the elicitation (n=7). Less than half of studies reported whether training materials were used or made available, described the number or phrasing of questions, or presented the individual as well as aggregated expert responses.
Of the 112 studies that presented indeterminate EE methods ( Appendix 5 ), the most common study design was cost-effectiveness or cost-utility analysis (n=89). Experts were asked to provide estimates for a wide range of parameters, from the sensitivity and specificity of testing instruments, to state transition probabilities or the occurrence of adverse events. Eight studies cited the use of experts but did not specify the parameters that were elicited.
Among the studies where the type of expert was specified, clinical experts were most common (n=36). In eight instances, combinations of experts (clinicians, researchers, and/or patients, among others) were consulted and in five studies experts were members of the study team. In 63 studies, no specifications on the type of expert were provided.
Eighteen studies, elicited information from 2-5 experts, five utilized 6-10 experts, three elicited information from only one expert, and two included more than 10 experts. A total of 77 studies provided no information on the number of experts used in the elicitations, while seven studies consulted an unspecified number of experts.
A total of 47 studies included no details on the size or composition of the panels. Across the board, almost no information was provided regarding how expert opinion was elicited or aggregated. Expert opinion was supplemented by additional data mainly from the published literature or literature reviews in 65 studies.
Discussion
Our systematic review focused on the use of expert opinion to inform computational modeling health studies. Expert elicitation is a common approach used to generate parameter estimates employed in computational modeling health outcome studies. We found 152 studies utilizing some form of expert opinion for the generation of probabilistic model input parameters. However, detailed descriptions of the methods used to elicit expert opinions were often inconsistent and incomplete. Twelve studies could not be included because, while the use of expert opinion was highlighted in the abstract, they failed to mention its use in the text of the study. We classified 40 studies as using formal EE methods and 112 as using indeterminate EE methods.
Consistent with previous reviews of formal EE methods, we found poor levels of EE methods reporting. 10 , 11 Unlike previous reviews, our results extend beyond studies exclusively in pharmacoeconomics and cost-effectiveness research. While most of the included studies were still cost-effectiveness analyses, we identified 11 non-cost-effectiveness studies that used formal EE methods. Additionally, our study is the first to evaluate indeterminate EE methods in addition to those with formal EE methods. Indeterminate EE methods were more common than formal EE studies and reported minimal information about how experts were defined and selected or how inputs were generated, raising substantial concerns regarding the reliability and generalizability of the elicitation process, the transparency of derived modeling inputs and the validity of modeling results.
Among formal EE method studies, lack protocols and piloting of elicitations were most absent. However, a substantial number of studies either failed to report key methodological information, such as the expert recruitment/sampling process and how expert estimates were collected and collated, or reported incomplete results from elicitation without including all experts’ responses prior to aggregation. Studies with indeterminate methods did not report any of this information.
Our review identified two important avenues of future work on the use of EE. The first, is the development of EE best practices. The second is determining when EE is most valuable to researchers conducting health-related modeling studies.
With expert opinion playing a key role in studies that are directed at health policy and medical reimbursement decisions, ensuring the highest methodological quality is essential. Yet, modeling and cost-effectiveness best practices, such as those set forth by the 2012 ISPOR-SMDM Modeling Good Research Practices Task Force and the Second Panel on Cost-Effectiveness in Health and Medicine, among others, make limited mention of when and how expert opinion should be relied upon for input parameter estimation. 75 – 78 While the more recent ISPOR guidelines discuss the value of EE, limited guidance is offered on the proper methodologies. 79 The lack of consensus and guidance on methods have likely contributed to the limited reporting found in this and previous reviews. Earlier this year, Bojke et al. sought to develop reference case methods for EEs in the healthcare setting offering nine principles to guide the use of EE. 18 However, modelers are still left with considerable discretion to determine when and how expert opinion should be elicited, aggregated, modeled, and communicated, and readers with almost no ability to evaluate the quality of the resulting parameters and results if editors and reviewers do not know of these guiding principles.
While a lack of reporting does not necessarily equate to poor quality or indicate bias, 80 persistent underreporting of essential information can reduce the credibility and replicability of modeling studies that use EE. Recent reporting guidelines provide an extensive outline of what health-related EEs should report, but do not go so far as to highlight best methods. 9 At a minimum, studies should report the composition of the expert panel and selection process, the preparation materials provided and the elicitation and aggregation methods. These pieces of information are essential to ensure the replicability of EEs and allow reviewers and readers to assess potential bias. A range of contextual biases can influence EE results including anchoring, availability bias, the base rate fallacy, and overconfidence which can be amplified by inappropriate EE methods. 19 , 22 , 81 , 82 Better reporting is needed to facilitate the evaluation of potential bias and highlight any value of the EE. In addition, consistent and extensive poor reporting may call into question the general value of EEs, thereby making it difficult to use this methodology when there is a clear need for a well-conducted EE.
In the absence of methodological details, at best, it is not clear whether EE provides modelers information that improves upon a review of the literature. However, at worst, it could enable the use of EE results which highly bias the model findings. Research societies whose members make extensive use of EE and computational modeling should reach a consensus on how to conduct and report on their use. The development of reporting guidelines and best practices can promote transparency and allow for a fuller assessment of any sources of bias. 80 , 83 , 84
Beyond research society guidelines, journals and peer-reviewers also have a responsibility to present high-quality EE research. Since expert knowledge has been considered by some to be a lower form of evidence than observational or clinical trial data, 2 editors and reviewers may disregard stand-alone EE manuscripts or modeling studies that rely heavily on EEs regardless of their methodological rigor. Thereby, journals may limit the discussion, review and development of EE methods that could advance the field. Journals, particularly those that publish high-quality model-based research, should open venues to facilitate the publication and peer review of stand-alone EEs with emphasis on adhering to principles and reporting guidelines for EEs. 9 , 18 Separate peer-reviewed publication of EEs should be encouraged, as this would enable independent review of the EE and the subsequent modeling study.
Greater reporting of methods could allow for a more robust discussion about the application of EEs, which are generally considered to be of value when data is scarce 10 , 11 or a lack of consensus exists. 5 However, limited research exists on when the use of EE provides value over other methods to generate model input parameters. In some instances, EE is the only way to generate necessary model inputs given the lack of existing data. In others, it can be used to synthesize disparate sources or build consensus when evidence is conflictual. As seen in 65 of the studies with indeterminate methods, EE can also be combined with other sources of data. Rarely, justifications were provided as to why expert opinion was better than the information available in the published literature. The value of EE may also vary by the type of intervention being studied, where some interventions may be better suited to other types of data collection methods. EE may be useful when primary data collection is unethical, infeasible, or time sensitive. Even in studies that conducted formal EEs, few provided a robust discussion of the merits of conducting an EE relative to other data collection methods.
Despite a large number of studies that represent the spectrum of expert opinion, we note that our review is limited to only those studies where expert opinion was conducted to generate probability inputs for modeling health outcomes. EE may be used for the development of other types of model input parameters (such as costs, or utilities) or to inform the model development process, but these may require alternative methods to those described here. Additional insights on the use of EE may be gleaned from studies that focus on different types of outcomes beyond health. Regardless of the methods used, EEs for any model input parameters should be subject to the same critical review as EEs that elicit probabilistic inputs. Additionally, we only included peer-reviewed journal articles and did not search the grey literature which is likely to be much larger than what is represented in this review. Expert opinion generally, and formal EE in particular, is used by a diverse array of government agencies. 5 , 10 , 11 , 14 Reports in the grey literature may present other EEs that have been used directly for policy modeling. We also limited our search to elicitations that were used to inform computational modeling. Studies of EEs conducted outside of modeling settings may provide valuable information to modelers on methods for conducting and aggregating data from EEs.
Despite these challenges and limitations, our review illustrates the extensive use of expert opinion, both formally and indeterminately in health-related modeling studies to generate input probabilities. We highlight the pervasive lack of reporting of important methodological details about expert elicitation. Among formal EEs, we found considerable variation in the methods used for expert selection, data collection and aggregation, and in accounting for uncertainty. The comparison of these methods may be useful to modelers looking to implement their own elicitations. Further work is needed to outline when EE provides additional benefits beyond that of other methods such as literature reviews and primary data analysis. Going forward, the lack of reporting among studies with indeterminate EE methods should raise concerns. As EE continues to proliferate, reporting guidelines and recommended methods need to be established to guide future researchers looking to incorporate EEs into their model-based analyses of health outcomes.
Introduction
Computational models are widely used to evaluate the effects of policies and interventions in silico for health research. 1 , 2 Modeling can incorporate disparate sources of data in a coherent framework, compare different interventions that may not be ethical or feasible to evaluate through trials or observational studies, and account for uncertainty in the variables or effect sizes of the interventions of interest. 3 , 4 Crucial to any modeling project is the use of the best available data inputs for the question at hand. However, in many instances, modeling is applied to a problem or area with limited or even non-existent data. For example, the measures of the impact of new policies on behaviors or health outcomes, especially those that have not been previously implemented, are often unavailable. While the likelihood of some events can be extrapolated from similar policies that have been implemented, the potential impacts of new policies may be substantially different from past policies. In instances where policy effects are unknown, elicitation of expert judgement, also referred to as expert elicitation (EE), has been used as a method to obtain input parameters for computational models. 5
Decision-makers have used EEs to quantify the unknown potential effects and levels of uncertainty of a given policy or intervention. Alternatively, existing evidence on the effects of an intervention may be inconsistent and elicitation can be used to build consensus. For example, EE has been used by the US Environmental Protection Agency, Food and Drug Administration, and other federal agencies, as well as international bodies such as the Intergovernmental Panel on Climate Change, 2 , 5 for model development 6 and to gauge uncertainty in climate modeling. 7 , 8 In the field of health, the use of EE is most prevalent in pharmacoeconomic modeling to inform reimbursement decisions, 9 – 11 but has been used in models to inform cancer screening recommendations, 12 childhood obesity interventions 13 and tobacco control policy. 14 – 16 The process has been used extensively by the National Institute for Health and Care Excellence in the UK for treatment coverage decision-making. 17
EE refers to a variety of methods by which opinions of authorities in the field are collected and collated. Despite the widespread use of expert opinion as a source of data for health analyses, no consensus nor gold standard exist on the proper methods of planning, implementing, presenting, and using data from an EE. 9 , 18 There exists a range of formal to informal EE methods. 10 , 11 , 18 – 24 The levels of complexity and rigor of the methods vary. For instance, elicitations can be conducted in-person, over the phone, online via video conferences, or with paper or electronic surveys. The measures elicited can include point estimates (e.g., medians or means), ranges of uncertainty (e.g., quartiles or confidence intervals), or probability distributions. Probability distributions are often developed by asking each expert placing a series of markers on a frequency chart to represent points on a probability density function. This method is commonly referred to as the histogram method, 22 or alternatively ‘chips and bins’ 25 or the Trial Roulette method. 26
Differences also appear in the methods by which expert responses are aggregated. Behavioral approaches to data aggregation include the nominal group technique or decision conferencing, where experts exchange views to draw forth a single expression of their judgment. 22 Mathematical aggregation may involve linear pooling (i.e., simple averaging), logarithmic opinion pooling, weighted averaging, or Bayesian methods. In addition, a series of hybrid methods exist that combine both behavioral and mathematical approaches, such as the commonly used Delphi method. 9 , 27 In this approach, the experts provide initial estimates and are then asked to revise their estimates based on those of the other experts. This can take place over multiple rounds in an attempt to build consensus among all members of the elicitation. 27 All interaction in a Delphi study is facilitated by an investigator who understands the objectives and can effectively engage experts.
Finally, packages, such as the SHeffield ELicitation Framework (SHELF) and EXPLICIT, have been developed to aid researchers in the conduct elicitations. These packages provide templates and software that can assist with implementation and aggregation. SHELF in particular supports facilitator-guided group interactions and information sharing to arrive at a consensus. 28 EXPLICIT is an Excel-based package that provides real-time visualizations based on the expert’s elicited values. 29
Previous reviews on EE methods have highlighted these different approaches, but at the same time emphasize the need for increased methodological rigor and standardization. 10 , 11 To that end, Iglesias and colleagues, developed guidelines for reporting results from EEs. 9 More recently still, Bojke et al. developed a set of guiding principles and criticisms to foster the development of a reference case for health-related EEs. 18
Given the widespread use of EE in modeling health behaviors and outcomes, it is important to understand how EEs have been conducted in practice in order to better assess their utility for computational modeling studies. The purpose of this review is to systematically identify and compare methods used to elicit expert opinion to generate input parameters in computational modeling studies with particular emphasis on the way in which decisions regarding EE methods are reported. Two earlier reviews evaluated the use of EE limited to health technology assessments, 10 , 11 while a third looked at the use of EE when examining enteric illness. 20 We expand upon this work by including all health-related modeling studies as well as those that do not use formal EE methods. Consistent with previous reviews, 10 , 11 we focused exclusively on studies that elicited probability estimates, given the challenges in eliciting probabilities. 5 , 19 While EE can also be used to estimate costs or other outcomes, these inputs are more directly observable and their elicitation involves a separate set of issues. 30 – 33
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.