Methods
The pairwise comparison tasks were developed and administered to a representative
sample of the UK general population. Responders were asked to express their
preference between pairs of health profiles. The next section describes some of the
key aspects of the experiment.
A selection had to be made of possible bolt-ons from the candidate list
identified by Finch et al., 5 , 19 , 20 which included life
satisfaction, speech, vision, hearing, sleep, cognition, energy, and
relationships, among others. The bolt-ons examined in this study are hearing,
sleep, cognition, energy, and relationships. Relationships, energy, and hearing
were selected as they had large, moderate, and small coefficients, respectively,
when regressed over a proxy of HRQoL, the Health Visual Analogue Scale (VAS). 20 Cognition and sleep were selected as the former had a large impact on
preferences for health states in one study, 13 while the latter did not have a significant impact on preferences for
health states in another. 14
Descriptors and labels of bolt-on dimensions were developed to closely resemble
the format of the EQ-5D-5L. These were assessed by the research team in terms of
their coherence with the EQ-5D-5L wording, their suitability for the lay public,
and their consistency across dimensions and with the construct measured. Where
there were inconsistencies, descriptors and labels were reworded and initial
wordings replaced. If it was not possible to establish the best wording, then
multiple wordings were examined in the face validity testing phase.
The face validity of bolt-on variants was tested in 2 focus groups. The first
focus group recruited 5 members of the general public, and the second focus
group had 6 patients affected by chronic health conditions (i.e., chronic
obstructive pulmonary disease, type 1 diabetes, chronic fatigue syndrome, and
endometriosis). Participants were asked to comment on the bolt-ons relevance,
clarity (ease of understanding and responding to), and acceptability in
different populations. A topic guide was used to aid the discussion. Focus group
recordings were analyzed using content analysis, which involved a systematic
identification of sections of the transcripts related to the aspects of interest
such as clarity. Results were used to modify descriptors and labels if face
validity problems were identified and also to select the best descriptors where
multiple wording was presented. Final descriptors and labels for the 5 bolt-ons
are presented in Table
1 . Alternative descriptors and labels, as well as bolt-ons
descriptors and labels for life satisfaction, speech, and vision, are presented
in Supplemental Table S1 .
Descriptors and Labels of Bolt-on Dimensions Tested in the Pairwise Experiment a
Reproduced by permission of EuroQol Research Foundation. Reproduction
of this version is not allowed. For reproduction, use or
modification of the EQ-5D (any version), please register your study
by using the online EQ registration page: www.euroqol.org .
Given the large number of bolt-on dimensions, a decision had to be made between
selecting numerous health states with fewer respondents per health state and
selecting a smaller number of health states but eliciting preferences from a
larger sample of respondents. As this was the first study using pairwise choices
to support selection of bolt-ons, the latter was chosen to increase the
confidence in the results obtained, which in turn can better inform future
research.
To develop the health states, EQ-5D-5L pairs were needed on to which the selected
bolt-on dimensions could be added. The ideal pairs of EQ-5D-5L health states
would be those in which there was equal preference (50:50) across each pairwise
choice to maximize the ability to assess the impact on preferences when bolt-ons
were added. Three pairs of health states were selected from the valuation study
used to develop the EQ-5D-5L tariff for England, 22 , 23 which employed both TTO
and discrete-choice experiments based on pairwise choices. The selected pairs of
health states are health state pair 1 (state A 11122 v. state B 23111), health
state pair 2 (state A 52211 v. state B 11325), and health state pair 3 (state A
33142 v. state B 34333). (EQ-5D-5L has 5 dimensions and the numbers represent
the severity levels: 1, no problem; 2, slight problems; 3, moderate problems; 4,
severe problems; and 5, extreme problems/unable.)
There are 25 possible combinations of bolt-on levels for each bolt-on for each
pairwise choice (i.e., 1 v. 1, 1 v. 2, 1 v. 3, etc.). Due to resource
limitations, it was not feasible to test all these possible combinations. Hence,
3 levels per bolt-on were chosen for this study: levels 1, 3, and 5. The first
level was included since it allows an assessment of whether the simple presence
of a bolt-on dimension changes preferences for the pairs of health states
presented. The third and fifth levels were selected to allow investigation of
the impact of severity on the relative importance of bolt-ons. The selected
health states and bolt-on levels used in this study are presented in Table 2 .
Choice Questions by Bolt-ons
There were therefore 3 EQ-5D health state pairs and 5 bolt-ons with 3 severity
levels each investigated in this study. Bolt-ons at severity 3 and 5 were always
added to health state A in each pair (i.e., 11122, 52211, and 33142). In total,
48 pairwise questions were included in the survey, 3 of which did not include
bolt-ons.
The pairwise choices were administered in an online survey. The survey had 4
components presented in the following order: 1) background and sociodemographic
questions, 2) self-reported health assessed through the EQ-5D-5L + bolt-ons, 3)
familiarization session, and 4) 8 pairwise choice questions. Each pairwise
comparison asked respondents to select the profile they preferred (an example of
the pairwise question for a bolt-on at level 3 is presented in Figure 1 ). No indifference
option was provided, in line with previous research, 22 which implied that respondents had to choose option A or option B.
Example of pairwise comparison as presented to responders.
A design was employed for the survey, in which each pairwise question was
assigned to a block. Each block included 8 pairwise questions. This was
considered a feasible number of tasks per participant based on previous research. 24 Participants allocated to blocks 1, 2, and 3 completed 1 task comparing
pairs of EQ-5D-5L states without bolt-ons in each block and 7 tasks comparing
pairs of EQ-5D-5L states with bolt-ons. Participants allocated to blocks 4, 5,
and 6 completed 8 tasks comparing pairs of EQ-5D-5L states all with bolt-ons. To
avoid focusing effects, each block included all 3 EQ-5D-5L health state pairs at
least once and different combinations of bolt-ons.
The survey presented 2 levels of randomization. First, participants were
randomized to 1 of the 6 blocks (although the order within each block was not
randomized). Subsequently, a randomization of the side in terms of which options
participants saw as option A and option B was performed to avoid any position
bias.
It was estimated that to detect a 10% difference between responses with and
without a bolt-on dimension, using a 2-sided test with a power of 0.8 and
significance level of 1%, 170 responders per health state pair were required. 25 As the 48 health state pairs were presented in 6 blocks, the target
sample for the study was of 1020 participants (i.e., 170 * 6).
Participants were recruited using an existing UK online panel administered by
Research Now, a market research company, using quotas for sex, age, education,
whether they had children, religion, and marital status to achieve a
representative sample. The panel is made up of individuals who have previously
signed up to answer surveys in return for points that can be exchanged for
goods. Each responder used a weblink to access the survey and was for this
reason able to self-complete it at his or her own convenience after providing
informed consent. The survey was administered in May 2017. The University of
Sheffield provided ethical approval.
The background characteristics of the participants allocated to the different
blocks were compared. Fisher exact test and χ 2 tests were used to
identify the presence of statistically significant differences in age, sex, and
social and economic status across the 6 blocks.
To test whether bolt-ons had a statistically significant impact on preferences,
we used logistic regressions with clustered sandwich estimators. Use of cluster
sandwich estimators accounted for the possible intraindividual correlation
generated by the panel structure of the data. First, a logistic model was
estimated by regressing respondents’ choice over dummy variables signaling the
presence or absence of each bolt-on (i.e., hearing, sleep, cognition, energy,
relationships, and a dummy identifying the health state pairs). To ensure the
generalizability of our findings, we tested confounding by adjusting for
background sociodemographic characteristics. Marginal effects (i.e., log odds
ratios for the bolt-ons β coefficients) are presented.
Second, we tested the null hypothesis that additions of bolt-on level dummies to
the main effect bolt-on model resulted in a significant improvement in model
fit. The Wald test (i.e., pseudo-score test) was used for this purpose.
Following the Wald test results, a main effect model using dummies for each
bolt-on level and for the health state pairs was estimated. Marginal effects
(i.e., log odds ratios for the β coefficients of the bolt-on level dummies) are
presented, once again adjusting for background sociodemographics.
Third, to further assess bolt-ons’ impact on different health state pairs, models
were estimated separately for each of the bolt-on options at level 1, level 3,
or level 5, for each of the 3 pairs investigated. Marginal effects (i.e., log
odds ratios for the β coefficients for each bolt-on for each level and health
state) are once again presented.
For all models, marginal effects were used to compare the impact of different
bolt-ons. For example, the marginal effect for health state 11122 with hearing
at level 5 was compared with the marginal effect for the health state 11122 with
relationships at level 5 and so on. Analyses were conducted using STATA/MP 14.1
(StataCorp, Cary, NC).
Results
In total, 1581 individuals entered the survey, but 342 were excluded as they did not
select all options in the consent form. A further 169 were excluded as they did not
complete the survey, 5 as they “speeded” through the survey (threshold for speeding
is calculated as the median completion time divided by 3), and 25 as their quota was
already full. The final analysis set comprised 1040 participants. Each pairwise
choice was completed by a minimum number of 167 respondents and a maximum of 180,
depending on the block. The mean time taken to completion was 9.19 minutes (range,
2.12–245.33 minutes), and the median time was 7.05 minutes. Participants in block 2
took the shortest mean time (7.23 minutes), while participants in block 4 took the
longest mean time (10.52 minutes). Background characteristics and health of the
sample are presented in Table
3 . No statistically significant differences were found between
participants allocated to the 6 blocks in terms of sex, marital status, profession,
and highest education achieved. Differences in age were seen between blocks 1 and 4
( P < 0.001), with block 1 appearing more normally
distributed with the majority of respondents being 45 to 54 years old, and block 4
presenting a relatively uniform number of responders in each age category.
Self-reported health status was generally similar across blocks except for
responders in block 2 who reported more problems in self-care than responders in
block 6, and this difference was statistically significant ( P <
0.05).
Background Characteristics and Health of the Sample
GCSE, General Certificate of Secondary Education.
In the main effect bolt-on model, hearing had the largest impact (–0.16), followed by
cognition (–0.15), relationships (–0.12), sleep (–0.09), and energy (–0.08). None of
the sociodemographic confounding variables had a statistically significant impact on
the results, as well as health state pairs.
When testing extensions of the bolt-on main effect model using the Wald test, the
null hypothesis that dummies for bolt-ons at levels 3 and 5 were simultaneously
equal to zero was rejected, showing that their inclusion resulted in a statistically
significant improvement in model fit. Additions of bolt-ons at level 1 did not
result in a significant improvement in model fit.
When estimating main effects for the bolt-on levels, cognition had the largest
marginal effect at level 3 (–0.16). For the rest of the domains, at level 3,
relationships had a small marginal effect (–0.04) and energy the second smallest
(–0.05). Sleep and hearing reported the same marginal effects (–0.12). At level 5,
hearing had the largest impact (–0.43), followed by cognition (–0.31), relationships
(–0.29), energy (–0.27), and sleep (–0.19). Table 4 reports the marginal effects and
clustered sandwich estimator standard errors for the bolt-on main effects, as well
as the bolt-on level main effects for all health state pairs together, and provides
a ranking of bolt-on marginal effects for severity levels 3 and 5. Once again, none
of the sociodemographic confounding variables had a statistically significant impact
on the results, as well as health state pairs.
Marginal Effects and Clustered Sandwich Estimator Standard Errors for the
Bolt-on Main Effect and Bolt-on Level Main Effects for All Health State
Pairs and Ranking of Bolt-ons for Severity Levels 3 and 5 in Terms of
Marginal Effect Size, Adjusting for Sociodemographics
−, Not applicable.
Statistically significant at P < 0.01.
Statistically significant at P < 0.05.
Table 5 presents the
marginal effects of the β coefficients from the logistic models with clustered
sandwich estimators’ standard errors and ranking of bolt-ons for severity levels 3
and 5 in terms of marginal effect size. The marginal effects for all bolt-ons at
level 1 were generally small and not statistically significant. For example, in
health state pair 2, energy at level 1 had a nonstatistically significant marginal
effect of −0.01, while hearing at level 1 had a nonstatistically significant
marginal effect of −0.00 in health state pair 3. Of all bolt-ons tested at level 1,
only relationships reported a statistically significant marginal effect and only for
health state pair 3 (i.e., −0.18).
Marginal Effects of the β Coefficients from the Logistic Models with
Clustered Sandwich Estimators’ Standard Errors and Ranking of Bolt-ons for
Severity Levels 3 and 5 in Terms of Marginal Effect Size, Adjusting for Sociodemographics a
ME, marginal effect; SE, robust standard error.
Health state pair 1: state A 11122 v. state B 23111; health state pair 2:
state A 52211 v. state B 11325; health state pair 3: state A 33142 v.
state B 34333. Bolt-ons at severity 3 and 5 were always added on health
states A.
Statistically significant at P < 0.01.
Statistically significant at P < 0.05.
Bolt-ons at level 3 generally resulted in statistically significant negative marginal
effects in health state pairs 1 and 3. In health state pair 2, marginal effects for
hearing and energy were small and not statistically significant. Marginal effects
ranged between −0.27 (cognition) and −0.16 (energy) in health state pair 1 and
between −0.31 (cognition) and −0.19 (relationships) in health state pair 3, while
they were smaller when a bolt-on at level 3 was added to health state pair 2,
ranging between −0.11 (relationships) and 0.01 (energy). Cognition had the largest
marginal effects in all 3 pairs, showing that this bolt-on substantially reduced the
probability of responders choosing the health state to which it was added. By
contrast, energy at level 3 had the smallest marginal effect for health states 1 and
2 and the second smallest marginal effect for health state pair 3. There was no
consistent ordering for sleep, hearing, and relationships.
As expected, bolt-ons at level 5 always resulted in statistically significant
negative marginal effects, which were smaller than marginal effects for bolt-ons at
level 3. Marginal effects ranged between −0.43 (hearing) and −0.29 (sleep) in health
state pair 1, between −0.29 (hearing) and −0.18 (energy and cognition) in health
state pair 2, and between −0.43 (hearing) and −0.24 (sleep) in health state pair 3.
Hearing consistently reported the largest marginal effects for all 3 health state
pairs. This was followed by cognition, which reported the second largest marginal
effect in health state pairs 1 and 3. Sleep had the lowest impact in 2 of the 3
health state pairs (1 and 3).
Discussion
This study investigated the potential of using pairwise comparisons to determine
whether bolt-on dimensions previously identified through factor analysis change
preferences for EQ-5D-5L health states. The aim was to test the use of a simple
low-cost pairwise comparison as a method for informing the selection of bolt-on
dimensions.
The study showed that each of the individual bolt-ons had a significant impact on
preferences for the EQ-5D-5L. The extent of this impact varied according to the
bolt-ons and their severity level, as well as the health states to which they were
added. Additions of bolt-ons at level 1 generally resulted in small and not
statistically significant marginal effects. Additions of bolt-ons at level 3
generally produced negative and statistically significant marginal effects, showing
a reduction in the log odds ratio of individuals choosing the health state to which
the moderate level was added. Additions of bolt-ons at level 5 generally resulted in
even larger negative marginal effects compared to bolt-ons at level 3. The
dimensions that had the largest impact were hearing and cognition, while sleep and
energy had less impact.
These findings agree with those of previous research in that they show that hearing
and cognition make a significant impact on the judgments people place on the EQ-5D
health states. 13 , 15 Our findings also show that sleep has less impact on
preferences for EQ-5D health states, although in contrast to a previous study, its
impact was significant. 14
This study found that at severity level 5, hearing had the largest marginal effect,
followed by cognition, relationships, and energy with relatively similar marginal
effects and sleep with the smallest marginal effect. By contrast, at severity level
3, cognition reported the largest marginal effect, followed by sleep, hearing, and
relationships, with energy registering the smallest marginal effect. This suggests
that the relative weight responders place on different health problems is not
constant across levels of severity between bolt-ons. This finding is relevant for
selecting bolt-on dimensions, as it highlights the need for a judgment on what
decision rule needs to be followed. One possibility might be choosing the bolt-ons
that have the greatest impact on preferences, in terms of log odds, compared to the
same health state without bolt-ons, based on the worst severity level.
Alternatively, bolt-ons might be selected based on the overall log odds (i.e., main
effect bolt-on model). Either way, other considerations remain fundamental for the
final selection, such as what other dimensions are already present in the
descriptive system of the examined measure and the psychometric evidence for the
impact of the bolt-on for the performance of the measure in terms of validity and
responsiveness in the population of interest.
Another important issue is that the inclusion of bolt-ons in the parent measure
requires a complete revaluation of the extended measure’s tariff. This study
provides some evidence showing that the addition of a bolt-on at level 1 might have
little impact on the value of the remaining core items/dimensions of the GPBM,
although it does not clarify what impact the bolt-ons at levels 3 and 5 would have
on the core dimensions of the parent measure. It still needs to be clarified whether
interactions are present to inform which modeling approach should be preferred in
bolt-on studies.
This study has a number of limitations. Only 3 pairs of EQ-5D-5L health states were
selected. This design appropriately responded to the methodological questions
investigated in this study, but it has the disadvantage of not being powered for
advanced econometric investigations (e.g., interactions). Moreover, previous
research has shown that preferences for bolt-ons might vary depending on the
severity of the health states to which they are added, 15 and for this reason, using other pairs from the 3 selected might have
generated different results. The results reported in this study show that the
relative weight responders place on different health problems is not constant but
rather depends on the severity level of the additional dimension, and further
testing with levels 2 and 4 for the same bolt-ons is required to confirm these
findings. In addition, although each block included different bolt-ons and health
state pairs to minimize focusing effects, the order in which pairwise questions were
presented within each block was not randomized, which may have an impact on
preferences. Finally, the study used pairwise choices while other methods could have
been used, including a ranking exercise or another choice-based task that would
allow comparability with more widely used valuation methods such as TTO.
Nevertheless, this study provides important evidence in that it proposes a flexible
and easy to use method for selecting bolt-on dimensions. Further research is
recommended on testing other bolt-ons, levels, and potentially including duration in
the comparisons to see how this compares to the previous studies that used TTO.
Supplementary Material
Click here for additional data file.
Supplemental material, sj-docx-1-mdm-10.1177_0272989X20969686 for Selecting
Bolt-on Dimensions for the EQ-5D: Testing the Impact of Hearing, Sleep,
Cognition, Energy, and Relationships on Preferences Using Pairwise Choices by
Aureliano Paolo Finch, John Brazier and Clara Mukuria in Medical Decision
Making
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.