Methods
This review adhered to the PRISMA 2020 checklist ( Page et al., 2021 ), provided as Supplementary File S1 .
We included randomized controlled trial (RCT) studies published in English in peer-reviewed journals with n ≥ 10 per arm. Sample size was evaluated at posttreatment to determine eligibility. Included studies aimed to evaluate psychological interventions targeting symptoms of depression and/or depression diagnosis among young people with long-term physical health conditions.
We included youths (≤18 years) and young adults (19–29 years) diagnosed with a long-term or chronic physical health condition. Eligible participants were diagnosed with ≥1 of the following long-term physical health conditions with an expected duration of ≥3 months: human immunodeficiency virus/acquired immunodeficiency syndrome, asthma, cancer, chronic pain (e.g., fibromyalgia, migraine), cleft palate, cystic fibrosis, deafness/hearing impairment, diabetes, endometriosis, epilepsy, heart diseases, inflammatory bowel disease (IBD), kidney diseases, liver diseases, sickle cell anemia, skin diseases (e.g., eczema), spina bifida, traumatic brain injury, and visual impairment. We additionally included OW/OB ( Hampl et al., 2023 ). These included physical health conditions are consistent with the literature and previous reviews using similar populations ( Catanzano et al., 2020 ; Law et al., 2019 ).
We included studies that investigated any evidence-based psychological intervention for their efficacy in treating depression symptoms/diagnoses delivered to children, adolescents, and/or young adults. An “evidence-based psychological intervention” was defined as treatment by a behavioral healthcare provider (e.g., therapist, psychologist, or provider trained by a licensed mental health professional) with an established basis of scientific evidence supporting its use for young people with elevated depression symptoms and/or a depressive disorder. Consistent with past reviews (e.g., Thabrew et al., 2018a ), intervention approaches could include behavior, cognitive-behavioral, third-wave, psychodynamic, humanistic, integrative, systemic, and other psychologically oriented therapies. Interventions delivered face-to-face, via e-health/telehealth, or by phone were included. Interventions that combined psychological and psychopharmacological treatments were eligible. Solely self-guided interventions (i.e., no involvement of a trained practitioner) or exclusively pharmacological interventions were excluded from the current review. Comparator conditions included any control condition that did not involve a psychological/behavioral intervention, such as nonpsychological treatment (e.g., health education, psychopharmacological treatment alone), treatment-as-usual (TAU; e.g., standard medical care without psychological intervention), or waitlist.
Outcome measures focused on youth/young adult symptoms and functioning. We evaluated the difference between the experimental group and the control group separately for all outcomes. Additionally, we extracted data regarding publication details (data source, DOI), study methods (study design, randomization procedure), participant characteristics (baseline demographic information, type of chronic illness), study methods (randomization procedures), and intervention characteristics (description of treatment and control conditions, delivery modality and provider, dosage, manualization, parent/caregiver involvement, and weekly assignments or homework).
Depression symptoms : Change from baseline to posttreatment in severity of depression symptoms was measured using validated self- and parent-report scales (e.g., Children’s Depression Inventory [CDI]; Kovacs, 1992 ).
Diagnosis of a depressive disorder : Change from baseline to posttreatment in depressive disorder diagnosis needed to be measured via semistructured clinical interview in accordance with Diagnostic and Statistical Manual of Mental Disorders or International Classification of Diseases criteria ( American Psychiatric Association, 2013 ; World Health Organization, 2022 ).
Anxiety symptoms : Change from baseline to posttreatment in severity of anxiety symptoms was measured using validated self-report scales (e.g., Beck Youth Inventory [BYI-II]; Beck et al., 2001 ).
Perceived stress : Change from baseline to posttreatment in perceived stress was measured via self-report scales (e.g., Perceived Stress Scale [PSS]; Cohen et al., 1983 ).
Functional disability : Change from baseline to posttreatment in functional disability was measured using validated self-report scales (e.g., Functional Disability Inventory [FDI]; Walker & Greene, 1991 ).
Quality of life : Change from baseline to posttreatment in quality of life was measured using validated self-report scales (e.g., Pediatric Quality of Life Inventory; Varni et al., 2002 ).
We searched the following electronic databases without date restrictions through July 1, 2023: CENTRAL (Cochrane Library), MEDLINE (OVID) 1946 to present, MEDLINE (Ebsco) 1974 to present, PsycINFO (Ebsco) 1806 to present, and PubMed. The same key words were used across databases, and specified type of study, participant age, participant health condition, type of intervention, and depression/depressive symptoms. Search was limited to abstracts. For the full line-by-line search strategy for each database, please see Supplementary Table S1 .
We searched clinicaltrials.gov and WHO ICTRIP (apps.who/int/trialsearch) for any ongoing trials or trials yet to be published in peer-reviewed journals. We also searched reference lists of review papers identified through the electronic searches specified above to ensure identification of any further RCTs that met the current systematic review/meta-analysis eligibility criteria.
Titles and abstracts of articles identified by the initial search were examined by two independent review authors (S.E.D.P. and M.B.) to determine studies requiring full-text review. The full text for all identified studies was independently investigated by both authors, and discrepancies were resolved through discussion and consultation with another author with expertise in meta-analytic methods (J.H.). Reasons for study exclusion included: Depression symptoms/diagnosis was not a presenting psychological concern (e.g., Viola et al., 2022 ), participants were not 0–18/18–29 (e.g., Myers et al., 2022 ), paper was a systematic review/meta-analysis (e.g., Thabrew et al., 2018b ), participants were not diagnosed with a long-term physical health condition (e.g., Sahler et al., 2002 ), control conditions involved psychological/behavioral intervention (e.g., Lofrano-Prado et al., 2022 ), study did not include a continuous measure of depression symptom severity (e.g., van Dijk-Lokkart et al., 2016 ), unable to reach authors for data needed to be included (e.g., Martinovic et al., 2006 ), intervention was not primarily psychological (e.g., Rich et al., 2016 ), intervention did not target the reduction of depression symptoms and/or remission of a depressive disorder as a primary outcome (e.g., Barrera et al., 2022 ), study was not published in English and/or in a peer-reviewed journal (e.g., NCT06247527 ), study was not an RCT (e.g., Heerman et al., 2019 ), study did not measure change in depressive symptoms (e.g., Brown et al., 2016 ) and intervention was not delivered by a trained professional (e.g., NCT03655067 ). All review authors agreed to the final list of included studies.
Two authors (S.E.D.P. and M.B.) independently extracted data on article details (authorship, title, year, country, DOI), participant characteristics and demographics (sample size, age, sex assigned at birth and/or gender, race/ethnicity, socioeconomic status, physical health condition(s), baseline differences between groups), intervention characteristics (theoretical orientation of intervention, intervention name, intervention description, dosage referring to frequency and length of sessions), delivery mode (in-person vs. telehealth), manualization, clinician type (trained professional leading group), parental involvement, at-home assignments, type and content of comparison condition(s), methodological characteristics (study design, grouping, setting, randomization methods), and study outcomes (measures, mean effect size, SD , confidence interval [CI], CI level). S.E.D.P. and M.B. reviewed all studies deemed eligible to ensure that only one study from each independent sample was included in the final review. No duplicate samples were identified. S.E.D.P., M.B., and J.H. extracted data on study results for the analyses. Identifying information for all included studies (participant demographics, intervention characteristics, methodological characteristics, and study outcomes) is presented in detail in Supplementary Table S2 .
Six studies were identified in the literature search as potentially eligible for inclusion in this review and ultimately were deemed ineligible because analysis on change from baseline to posttreatment on outcomes of interest was unavailable ( Ataie Moghanloo et al., 2015 ; Grey et al., 2009 ; Li et al., 2016 ; Martinovic et al., 2006 ; Zhang et al., 2019a ).
Risk of bias (RoB) was assessed using the Cochrane RoB version 2 tool for RCTs ( Higgins et al., 2019b ). Data for assessing RoB were obtained from published articles. Risk was categorized for each domain as “low,” “some concerns”, or “high.” S.E.D.P. and M.B. independently reviewed each included article for all RoB domains. J.H. resolved disparate ratings by assigning the most conservative rating. No studies were determined to have high RoB; thus, no studies were excluded from analyses due to high RoB concerns. The following domains were rated:
Randomization process (whether the allocation sequence was random and concealed until participants were assigned to intervention conditions, whether intervention groups were found to have baseline differences that may suggest problems with random assignment)
Deviations from intended interventions (any noted deviation from intervention content and the analytic approach to estimating the effect of assignment to intervention)
Missing outcome data (whether data for each outcome variable were available for all, or nearly all, participants who were randomized to intervention conditions)
Measurement of the outcome (appropriateness of the measures administered to assess outcomes and whether outcome assessment may have been influenced by knowledge of the intervention received)
Selection of the reported result (whether data reporting was consistent with a prespecified analysis plan; if results reported were ascertained from multiple eligible outcome measurements and/or multiple analyses of the data)
Overall bias (the highest risk rating across the five bias domains was used as the overall rating)
Reporting biases were assessed along with the RoB assessment in this review. Funnel plots were used to assess reporting biases, and Egger’s tests were conducted to validate these conclusions ( Egger et al., 1997 ).
We anticipated that most studies would report continuous data on all outcomes of interest immediately posttreatment. For studies that reported repeated follow-up assessments, we extracted data from the first available follow-up, consistent with the Cochrane Handbook and previous reviews (e.g., Deeks et al., 2024 ; Law et al., 2019 ). Treatment effects were evaluated on the change from baseline to the first available posttreatment follow-up for all continuous outcomes.
We extracted and analyzed continuous outcome data when reported. We employed standardized mean differences (SMDs) with 95% CIs to evaluate treatment effects for continuous data. We interpreted effect sizes as indicated by Cohen (2013) : small (.2), moderate (.5), and large (.8).
It was expected that studies would randomize at the individual (i.e., patient) level. We adhered to the Cochrane Handbook if cluster-randomization occurred ( Higgins et al., 2024 ). For studies that included multiple intervention or control groups, arms were collapsed and the control group was split equally across intervention arms to allow for comparison. We planned to include the first-step comparison of treatment and control groups for any cross-over trials and did not include data from the second step, where arms are crossed over, to avoid carryover effects.
We contacted authors regarding any data that were missing from published manuscripts or preregistration materials. Adhering to the Cochrane Handbook guidelines, we used data available in published materials whenever possible (e.g., M , SD ; Higgins et al., 2019a ). We did not impute missing data that we were unable to obtain from the study authors. If a study reported both intent-to-treat analyses and protocol analyses, we planned to preferentially extract intent-to-treat data.
Heterogeneity was visually inspected using forest plots and calculated using χ 2 and Ι 2 . In accordance with Cochrane Handbook guidelines, 0%–40% might not be important, 30%–60% may indicate moderate heterogeneity, 50%–90% may indicate substantial heterogeneity, and 75%–100% indicates considerable heterogeneity ( Deeks et al., 2024 ). In cases when heterogeneity was substantial or considerable, we planned to conduct sensitivity analyses if appropriate.
Data were analyzed using R ( R Core Team, 2021 ) and Review Manager Version 5.4 ( Review Manager [RevMan], 2020 ). Pooled SMD for change from baseline to posttreatment between experimental and control conditions, along with 95% CIs, were calculated using an inverse variance approach with a linear model. Outcome data were analyzed using fixed-effects models. Findings across studies were described when it was not possible to combine data.
The quality of the evidence identified in this review was assessed using the GRADE system and guidelines provided in the Cochrane Handbook ( GRADEpro GDT: GRADEpro Guideline Development Tool, 2024 ; Schüneman et al., 2013 ). RCT study designs are strong and generally produce high certainty of evidence, but certain limitations may reduce this certainty and result in lower grades. Studies were assessed on five factors that may downgrade the quality of a body of evidence: RoB (limitations in design and execution; outcomes may be downgraded if they included studies that published unclear or high RoB ratings as assessed using the Cochrane RoB version 2 tool for RCTs), inconsistency (studies that report heterogeneity >30% are downgraded; Deeks et al., 2024 ), indirectness (indicates applicability of the population, intervention, comparator condition, or outcome; outcomes are downgraded if, for example, important differences in populations are identified or differences in intervention or comparator conditions are sufficient enough to influence outcomes), imprecision (preciseness of estimates determined by the study sample size and associated CIs; studies with sample size that is less than the Optimal Information Size [OIS] as determined by prespecified clinically meaningful effect size are downgraded), and publication bias (downgraded if studies fail to publish null findings or outcomes). These considerations collectively inform assessment of the quality of available evidence and result in a grade of high, moderate, low, or very low.
Results
There were 2,191 records identified from searches of electronic databases and other resources, of which 208 abstracts were identified as potentially relevant. Following review of the full-text reports of these studies, 202 studies were excluded. Six trials were included in the qualitative analysis and contributed data to the meta-analysis. See Figure 1 for PRISMA flow diagram.
PRISMA flow diagram.
Six trials were included in this review ( Freedenberg et al., 2017 ; Kashikar-Zuck et al., 2005 , 2012 ; Shomaker et al., 2016 ; Szigethy et al., 2007 ). Two of the six studies listed an inclusion criterion related to elevated depression symptoms, and all six studies excluded participants who endorsed symptoms consistent with clinical depression disorders. Supplementary Table S2 provides a detailed description of each study’s characteristics.
By design, all studies were RCTs. Studies were published between 2005 ( Kashikar-Zuck et al., 2005 ) and 2017 ( Freedenberg et al., 2017 ). One cross-over trial was included ( Kashikar-Zuck et al., 2005 ), for which, as planned a priori, only data from the first phase were retained in analyses. One trial ( Szigethy et al., 2007 ) had long-term follow-up data published in a later paper ( Thompson et al., 2012 ); only data from the initial follow-up were included. No cluster RCTs were identified in the current review.
Participants were 11–18 years of age, predominantly White (62%) and female (80%). Freedenberg et al. (2017) did not report participant race/ethnicity. Trials involved participants with various long-term physical health conditions: cardiac diagnoses ( Freedenberg et al., 2017 ), chronic daily headaches ( Hickman et al., 2015 ), juvenile primary fibromyalgia syndrome ( Kashikar-Zuck et al., 2005 ), juvenile fibromyalgia syndrome ( Kashikar-Zuck et al., 2012 ), OW/OB ( Shomaker et al., 2016 ), and IBD ( Szigethy et al., 2007 ).
Inclusion and exclusion criteria varied across studies. All trials specified an age range for inclusion and required that participants were English-speaking and demonstrated symptoms or diagnosis of the physical health condition of interest. Other commonly cited inclusion criteria included: mild functional disability ( Kashikar-Zuck et al., 2005 , 2012 ) and mild and/or moderate depression symptoms ( Hickman et al., 2015 ; Shomaker et al., 2016 ). All studies listed mental health diagnosis or psychiatric symptoms that necessitated treatment as exclusionary.
Five trials evaluated cognitive-behavioral interventions. One trial ( Freedenberg et al., 2017 ) evaluated a modified mindfulness-based stress reduction program ( Sibinga et al., 2011 ). Three trials involved youth only and had no parent/caregiver intervention involvement ( Freedenberg et al., 2017 ; Hickman et al., 2015 ; Shomaker et al., 2016 ); three trials involved youth and parents in some, but not all, intervention sessions ( Kashikar-Zuck et al., 2005 , 2012 ; Szigethy et al., 2007 ). Dosage varied across studies. Kashikar-Zuck et al. (2005) did not report the length of sessions for the experimental condition. Experimental groups in all six trials were facilitated by trained mental health providers.
Control conditions varied from TAU ( Szigethy et al., 2007 ) to video online semistructured discussion groups ( Freedenberg et al., 2017 ). Dosage for the control conditions was lower as compared to the experimental condition for four trials; of these, one indicated that there was no time requirement for the control condition ( Kashikar-Zuck et al., 2005 ) and three trials reported that dosage for the control condition was variable and/or lower than the experimental condition ( Freedenberg et al., 2017 ; Hickman et al., 2015 ; Szigethy et al., 2007 ). Two trials matched dosage across experimental and control conditions ( Kashikar-Zuck et al., 2012 ; Shomaker et al., 2016 ).
Changes in severity of depression symptoms were measured using the Hospital Anxiety and Depression Scale (HADS; Zigmond & Snaith, 1983 ) in one trial ( Freedenberg et al., 2017 ), the BYI-II ( Beck et al., 2001 ) in one trial ( Hickman et al., 2015 ), the CDI ( Kovacs, 1992 ) in three trials ( Kashikar-Zuck et al., 2005 , 2012 ; Szigethy et al., 2007 ), and the Center for Epidemiological Studies-Depression scale (CES-D; Radloff, 1977 ) in one trial ( Shomaker et al., 2016 ). In one study ( Szigethy et al., 2007 ), depression symptoms were measured using the CDI, and self-report and parent-report responses were summed into a total CDI severity score. All other studies used self-report exclusively. Of the six studies included in the present review, two found statistically significant effects of evidence-based psychological intervention on change in depression symptoms among young people with long-term physical health conditions ( Kashikar-Zuck et al., 2012 ; Szigethy et al., 2007 ), suggesting that some cognitive-behavioral interventions targeting symptoms of depression were more effective than control conditions for reducing symptoms of depression in this population, while four studies found no statistically significant differences on change in depression symptoms between experimental and control conditions ( Freedenberg et al., 2017 ; Hickman et al., 2015 ; Kashikar-Zuck et al., 2005 ; Shomaker et al., 2016 ).
We were unable to evaluate differences between the experimental group and control group for changes in diagnosis of a depressive disorder, as this outcome was only reported by Szigethy et al. (2007) . This study concluded that there were no significant differences in the number of depression symptoms (indicative of a depression diagnosis) between intervention conditions.
Change in severity of anxiety symptoms was measured in two studies. One study used the HADS ( Zigmond & Snaith, 1983 ) to measure this outcome ( Freedenberg et al., 2017 ), and the other study ( Hickman et al., 2015 ) used the BYI-II ( Beck et al., 2001 ). Hickman et al. (2015) found that the experimental group showed greater reductions in anxiety symptoms compared to the control group, but Freedenberg et al. (2017) found no significant differences between conditions for change in anxiety symptoms.
Changes in perceived stress were measured in two studies. Freedenberg et al. (2017) used the Responses to Stress Questionnaire ( Connor-Smith et al., 2000 ), and Hickman et al. (2015) used the PSS ( Cohen et al., 1983 ). Freedenberg et al. (2017) found that the experimental group showed greater reductions in perceived stress compared to the control group, but Hickman et al. (2015) found no significant differences between conditions for change in perceived stress.
Changes in functional disability were measured in three studies. One study ( Hickman et al., 2015 ) used the Pediatric Migraine Disability Assessment ( Hershey et al., 2001 ), and Kashikar-Zuck et al. (2005 , 2012 ) used the FDI ( Walker & Greene, 1991 ). Kashikar-Zuck et al. (2012) found that the experimental group showed significantly greater reductions in functional disability as compared to the control group, whereas the other two studies found that both intervention and control conditions contributed to reductions in functional disability ( Hickman et al., 2015 ; Kashikar-Zuck et al., 2005 ).
We were unable to evaluate differences between the experimental group and control group in change in quality of life as this outcome was only reported in one trial ( Kashikar-Zuck et al., 2012 ). No significant differences between intervention conditions were noted for quality of life by this study’s authors.
RoB for all study outcomes is presented in Tables 1–4 . Bias was rated “low” for all studies in each of the following domains: deviation from intended intervention, missing outcome data, and selection of the reported result. Bias arising from the randomization process was rated “some concerns” for Freedenberg et al. (2017) . In this trial, the experimental group reported significantly higher anxiety symptoms at baseline as compared to the control condition, and there were no details provided on sequence concealment. Measurement of the outcome was rated as “some concerns” across all included studies. All studies relied on self-reported measurement of outcomes, and although all studies described randomization procedures intended to keep participants blinded to condition, it is still possible that participants may have been able to discern if they received the experimental intervention condition or a control condition. Thus, though unlikely, participants’ self-report could be biased by knowledge of the intervention received. Overall bias was rated “some concerns” for all six included studies. Inter-rater reliability (IRR) of the coding of these studies was perfect for three of the five domains, thus, an IRR statistic is not available for these domains. IRR ranged from κ = 0–.22 in other domains. Due to the wide range of IRR scores, the most conservative rating was considered by selecting the lowest score to maintain stringent bias evaluation.
Depressive symptoms analysis result.
Anxiety symptoms analysis result.
Perceived stress analysis result.
Functional disability analysis result.
Of the six trials included in the meta-analysis, five studies published data that were needed to calculate SMDs and 95% CIs in change in all outcomes between the experimental and control conditions. For one study ( Kashikar-Zuck et al., 2012 ), additional data needed for inclusion were obtained via email correspondence.
The pooled SMD in change from baseline to posttreatment from six studies (376 individuals) was −.30 with 95% CI (−.51, −.10). The overall treatment effect was significant in the change in depression symptoms from baseline to posttreatment ( Z = 2.92, p = .004; Table 1 ).
Insufficient data were available to assess this outcome.
The pooled SMD in anxiety symptoms change from baseline to posttreatment from two articles (78 individuals) was −.10 with 95% CI (−.54, .35). The overall treatment effect was not significant in the change from baseline to posttreatment ( Z = .42, p = .67; Table 2 ).
The pooled SMD in perceived stress change from baseline to posttreatment from two articles (78 individuals) was .16 with 95% CI (−.29, .60). The overall treatment effect was not significant ( Z = .69, p = .49; Table 3 ).
The pooled SMD in functional disability change from baseline to posttreatment from three articles (171 individuals) was −.35 with 95% CI (−.66, −.05). The overall treatment effect was significant in the change in functional disability from baseline to posttreatment ( Z = 2.28, p = .02; Table 4 ).
Insufficient data were available to assess this outcome.
Funnel plots showed symmetry around the pooled outcomes of interest ( Supplementary Figures S1–S4 ), and Egger’s tests were conducted for the primary outcome, change in depressive symptoms, and the secondary outcome, change in functional disability, from baseline to validate this conclusion. There is no feasible statistical test for publication bias for the secondary outcomes of perceived stress and anxiety symptoms as there were only two studies included for these outcomes; however, the funnel plots suggested that there was no asymmetry in these two outcomes. The Egger’s test results suggested no significant asymmetry for either change in depressive symptoms from baseline ( p = .937) nor change in functional disability from baseline ( p = .193). Though the Egger’s test results suggested no significant asymmetry, we acknowledge that this test may lack the statistical power to detect bias when the number of studies is small, especially for change in functional disability ( k = 3). Thus, we conducted the precision-effect test (PET) and precision-effect estimate with SE (PEESE; Stanley, 2008 ; Stanley & Doucouliagos, 2014 ) on change in functional disability from baseline. Neither PET ( p = .148) nor PEESE ( p = .141) corrections were significant, which is consistent with our observation from funnel plots and Egger’s test results suggesting that publication bias was minor.
Study heterogeneity was characterized as nonsignificant in two of four outcomes of interest; the I 2 s were 0% for anxiety symptoms and perceived stress. Study heterogeneity was characterized as moderate for depression symptoms functional disability because the I 2 s for these outcomes were 34% and 41%, respectively. The χ 2 test results were χ 2 =7.6 ( p = .18), χ 2 =.00 ( p = .95), χ 2 =.03 ( p = .87), χ 2 =3.42 ( p = .18), for depression symptoms, anxiety symptoms, stress, and functional disability, respectively.
GRADE was used to assess the quality of the evidence. The primary study outcome, depression symptoms, was rated as moderate quality on the assessment. Perceived stress and functional disability were rated as moderate quality, and anxiety symptoms were rated as low quality ( Table 5 ). Functional disability and depression were downgraded due to evidence of inconsistency, as evidenced in study heterogeneity I 2 values in the range of 30%–60%, which is considered evidence of moderate heterogeneity ( Deeks et al., 2024 ). Anxiety symptoms and perceived stress were both downgraded due to evidence of imprecision, determined based on OIS ( Supplementary Table S3 ). OIS was calculated based on the prespecified clinically meaningful effect size with significance level of .05 and 80% power. If the total sample size available for a specific outcome in the meta-analysis was less than the OIS, imprecision was rated down one level. If imprecision was rated down one level due to OIS and there were very few events and CIs around both relative and absolute estimates of effect that included both appreciable benefit and appreciable harm, imprecision was rated down two levels. Consequently, perceived stress was rated down one level and anxiety was rated down two levels. There was no evidence of RoB or indirectness for any outcome measures included in the meta-analysis. No studies reported change in diagnosis of depression or quality of life and these outcomes are not included in the quality of evidence table.
Summary of findings for psychological interventions for depression symptoms in physical health conditions.
Notes
Study heterogeneity was characterized as moderate for outcomes with I 2 values falling in the range of 30%–60%; thus, depression symptoms and functional disability were downgraded 1 level for inconsistency.
OIS were calculated based on the prespecified meaningful effect size with significance level of .05 and for 80% power. If total sample size available for a specific outcome in the meta-analysis was less than the OIS, we downgraded one level; if there were very few events and CIs around both relative and absolute estimates of effect, we downgraded two levels.
RoB=risk of bias; CI=confidence interval; SMD=standard mean difference.
GRADE Working Group grades of evidence:
High quality ⨁⨁⨁⨁: Further research is very unlikely to change confidence in the estimate of effect.
Moderate quality ⨁⨁⨁: Further research is likely to have an important impact on confidence in the estimate of effect.
Low quality ⨁⨁: Further research is very likely to have an important impact on confidence in the estimate of effect and is likely to change the estimate.
Discussion
In the current meta-analysis, we sought to describe and evaluate the efficacy of empirically supported psychological interventions to address symptoms of depression in young people with long-term physical health conditions. Results suggest a significant overall treatment effect for depression symptoms, suggesting that psychological interventions show promise for reducing depression symptoms in adolescents with long-term physical health conditions. Interventions targeting depression symptoms also showed potential for reducing functional disability among adolescents with long-term physical health conditions. While previous reviews have suggested that psychological interventions may be associated with fewer depression symptoms posttreatment ( Thabrew et al., 2018a ), this is the first review to evaluate change in symptoms of depression from baseline to posttreatment. This difference is a primary reason many published studies were ineligible for the present review but included in others and is a vital criterion for work aiming to evaluate the effectiveness of psychological interventions.
A recent review of behavioral interventions for nonspecific psychological concerns in pediatric populations with health conditions determined low-quality evidence ( Thabrew et al., 2018a ). In contrast, the current RCTs focused on interventions specific to depression symptoms that additionally evaluated change in symptoms from baseline to posttreatment and showed that, overall, adolescents with health conditions benefited. These findings converge with the literature that psychological interventions are superior to control for reducing depression in the general population of young people ( Cuijpers et al., 2023 ; Eckshtain et al., 2020 ; Weisz et al., 2017 ; Wuthrich et al., 2023 ). Greater empirical attention to evaluating interventions targeting depression symptoms in young people with long-term physical health conditions is warranted, because youths with physical health conditions and comorbid, untreated psychopathology report more impaired physical functioning ( Ding et al., 2008 ), lower quality of life ( Johnson et al., 2004 ), and have poorer disease management ( Sildorf et al., 2018 ) and earlier mortality ( Olusunmade et al., 2019 ) than those without psychopathology. Depression also relates to poorer medical treatment adherence ( Plevinsky et al., 2020 ) and lower engagement in positive health behaviors (e.g., Yourell et al., 2023 ).
The current findings from a review of studies through July 1, 2023 highlight the potential of depression interventions for young people with long-term physical health conditions but also draw attention to several notable gaps. First, of the six studies included in the present review, only two individual reports found psychological intervention to significantly reduce depression symptoms among young people with long-term physical health conditions ( Kashikar-Zuck et al., 2012 ; Szigethy et al., 2007 ), suggesting that some behavioral interventions targeting depression may reduce symptoms in this population, while other interventions may not be of benefit to young people at risk for elevated symptoms of depression ( Freedenberg et al., 2017 ; Hickman et al., 2015 ; Kashikar-Zuck et al., 2005 ; Shomaker et al., 2016 ). Variations in results of individual reports could also be attributed to other factors, such as limited power or sample size, which highlights a benefit of meta-analytic approaches that pool study samples (e.g., Cheung & Vijayakumar, 2016 ).
Also, only two of the included studies listed an inclusion criterion based on elevated baseline depression symptoms ( Hickman et al., 2015 ; Shomaker et al., 2016 ), and all six excluded young people with clinical depression disorders. The extant literature identified by the present review thus cannot elucidate the efficaciousness of interventions for reducing elevated depression symptoms and/or clinical depression. Given the rise in depression in the general population ( Centers for Disease Control and Prevention (CDC), 2024 ), tests of empirically supported interventions for youth with long-term physical health conditions who experience depression concerns is merited. Based on the extant literature on depression interventions in youth without long-term physical health conditions and the positive results of the current meta-analysis in samples of youth with long-term physical health conditions and largely heterogeneous depression, we would anticipate benefits of evidence-based depression interventions for a more selected or indicated group, but this hypothesis requires testing.
Two studies included in the present review additionally evaluated intervention efficacy for reducing anxiety symptoms and perceived stress ( Freedenberg et al., 2017 ; Hickman et al., 2015 ), and the overall treatment effect for these secondary outcomes was also nonsignificant, suggesting that more research is needed to determine whether young people with long-term health conditions and elevated anxiety or stress may benefit from psychological interventions aiming to reduce distress.
Further, there was insufficient data to evaluate subgroup (i.e., differences based on participant characteristics such as age and health condition) or sensitivity (i.e., differences based on trial characteristics such as delivery modality, intervention length, control group type) analyses. Socioeconomic status and household income were not reported by all studies. Future research should include these constructs and perform subgroup analyses to investigate the potential role they play in moderating change in depression symptoms. All trials tested face-to-face interventions, and thus, efficacy of telehealth depression interventions in young people with long-term physical health conditions remains inconclusive ( Thabrew et al., 2018b ). Likewise, extant trials could not elucidate variations in other intervention characteristics (e.g., treatment length, tailoring/CBPR elements) that may modulate efficacy. All studies involved youth aged 11–18 years, leaving unknown the overall effect of interventions for symptoms of depression in children and young adults with health conditions. Similarly, participants were predominantly female, and in most samples, predominantly White. All trials were conducted in the United States. Likewise, it is important to note that all included studies listed mental health diagnosis or symptoms that necessitated treatment as exclusionary, limiting generalizability of findings to subthreshold and uncomplicated depression presentations. While it is valuable to understand whether interventions that are designed to reduce or treat depression can benefit subclinical samples at risk for depression, greater attention should be paid to the treatment of clinically significant depression concerns in this population in future research.
The present review considered OW and OB simultaneously as indicative of poor chronic health; while this is consistent with the current practice guidelines for child and adolescent obesity and with research on OW/OB generally ( Hampl et al., 2023 ), there is significant heterogeneity in what diagnoses may constitute a “long-term health condition” and closer attention and specificity is needed toward this consideration. The present review included health conditions that have been included in past reviews of similar topics ( Catanzano et al., 2020 ; Law et al., 2019 ), but there are many health conditions that were not listed in search terms for the present review (e.g., chronic fatigue syndrome, allergies); future research is needed to expand upon present findings. Research is needed to elaborate depression intervention efficaciousness across more diverse, globally representative samples. As the literature grows for RCTs evaluating empirically supported interventions for depression symptoms in young people with physical health conditions, subgroup and sensitivity analyses based upon participant and trial characteristics will be beneficial for optimizing interventions for distinct groups/contexts.
A very small number of studies were identified as eligible following the approach and method that were established, peer reviewed and approved, and preregistered prior to conducting the current meta-analysis. With recognition that post hoc power analyses are limited and insufficient (e.g., Hoenig & Heisey, 2001 ; Zhang et al., 2019b ), results of post hoc power analyses using a range of standardized effect sizes do indicate that the present review was sufficiently powered to detect medium-to-large effects for the primary outcome (depression symptoms) and secondary outcome functional disability, and had insufficient power for other secondary outcomes. Still, considering the small number of articles included, the results should be applied only to the studies included in this meta-analysis. Authors of six potentially eligible studies did not respond to requests for additional data and the limited number of eligible studies required the use of fixed models that cannot be generalized beyond the scope of this review. Four of the six eligible studies included in this review were identified through other sources, primarily studies included in similar meta-analyses and references, and additional databases that were not searched may produce additional studies. This limits the replicability of the current review. Additionally, the focus of this review, conservatively, was on baseline to posttreatment change effects; future work is needed to understand whether the effects of psychological interventions for depression are sustained over time. Finally, the included studies were all published in English, limiting generalizability of the current results.
Notwithstanding, this review is strengthened by inclusion of five distinct pediatric populations represented in six clinical trials. Evidence-based interventions specifically designed to treat depression hold promise for reducing depression symptoms and functional disability among youth with various health conditions. More research is needed on other intervention approaches and modalities and with youth diagnosed with various health conditions to better understand the generalizability of these effects. More trials focused on reducing depression symptoms as a primary outcome are needed. Many trials evaluated depression symptoms as a secondary outcome or did not publish data needed to evaluate efficacy ( Bhana et al., 2014 ; Schache et al., 2020 ). Future RCTs should prioritize evaluating interventions with the primary aim of reducing depression symptoms in young people with physical health conditions. Further, larger trials with this population are needed that evaluate participant characteristics (e.g., age, sex/gender, health condition type, race/ethnicity, SES, country of origin) to gain deeper insight into who is most likely to benefit and identify ways to enhance interventions to maximize the positive impact for the greatest possible number of young people.
Importance
While there is an expansive literature regarding the treatment of psychological disorders among young people, literature on the treatment of psychological disorders comorbid with physical health conditions is limited ( Koban et al., 2021 ; Levine et al., 2021 ; Suls & Green, 2019 ). Depression is one of the most prevalent, impairing psychological concerns affecting young people, and among pediatric populations, symptoms of depression adversely relate to disease self-management, health outcomes, and global wellbeing ( Piao et al., 2022 ). Effective interventions for psychological conditions share many common factors, but the most effective interventions involve distinct components targeting depression symptomatology (e.g., behavioral activation; McCauley et al., 2016 ).
To our knowledge, no prior reviews have investigated the efficacy of psychological interventions that specifically target depression symptoms and/or diagnoses among youths and/or young adults with long-term physical health conditions. A small number of reviews have evaluated the efficacy of psychological interventions for mental health concerns more broadly in this population ( Thabrew et al., 2018a , b ), and other reviews of interventions for young people with chronic health conditions have included depression diagnosis/symptoms as a secondary outcome (e.g., Levy et al., 2010 ; Scott et al., 2021 ). Similar extant reviews conducted by Thabrew et al. (2018a , b) have included studies that evaluated symptoms at posttreatment without consideration of change in symptoms from baseline to posttreatment; findings support low-quality evidence that psychological therapies may be associated with fewer depressive symptoms posttreatment ( Thabrew et al., 2018a ). The present review is timely, and necessary to understand the extent to which psychological interventions are efficacious for reducing depression symptoms/diagnoses among young people with long-term physical health conditions.
Objectives
Primary objectives of the current review and meta-analysis were to: (a) describe empirically supported psychological intervention modalities that have been used to address depression symptoms and/or disorders in youths and young adults with existing physical health conditions, (b) assess the efficacy of psychological interventions for decreasing depression symptoms and/or remission of disorders, and (c) evaluate participant/trial characteristics (e.g., age, sex, health condition type) as moderators. Secondary objectives were to evaluate change in health/wellbeing-related outcomes, including anxiety symptoms, perceived stress, quality of life, and functional disability.
Description
Psychological interventions aiming to reduce depression symptoms in general populations of children, adolescents, and young adults demonstrate small-to-moderate effects that are superior to control conditions ( Cuijpers et al., 2023 ; Eckshtain et al., 2020 ; Weisz et al., 2017 ; Wuthrich et al., 2023 ). Psychological interventions appear as efficacious as medication for youth depression treatment and have shown efficacy across individual and group deliveries ( Cuijpers et al., 2023 ; Robberegt et al., 2023 ). Cognitive-behavioral therapy has the largest evidence base; yet, other approaches such as interpersonal psychotherapy ( Bian et al., 2023 ), mindfulness-based intervention ( Reangsing et al., 2023 ), and caregiver-inclusive interventions informed by family systems theories ( Dippel et al., 2022 ) are similarly effective ( Cuijpers et al., 2023 ).
There is limited knowledge about factors that affect the magnitude of youths’ response to interventions targeting depression symptoms, in general, and even less information about how baseline characteristics (e.g., existing physical health conditions, age, sex/gender) interact with, or moderate, depression treatment efficacy ( Van Der Lee et al., 2007 ). Adolescents and young adults possess more developed cognitive/executive functioning than children aged <12 years ( Prencipe et al., 2011 ); however, data on whether depression intervention efficacy is affected by age are inconsistent ( Cuijpers et al., 2020 , 2021 ; Nilsen et al., 2013 ). Likewise, it remains unclear if sex assigned at birth and/or gender moderate depression intervention efficacy ( Courtney et al., 2022 ).
Intervention characteristics also may play a role. In particular, community-engaged tailoring of depression interventions for young people with physical health conditions may be more efficacious than a generalized approach, as the former ensures that learned skills are most relevant to patients’ distinct needs and contexts ( Czajkowski et al., 2015 ). Delivery modality (e.g., eHealth/telehealth compared to face-to-face) differences are not well understood ( Barney et al., 2020 ). Moreover, empirical data about depression intervention dosage are needed to evaluate whether lengthier evidence-based interventions (i.e., ∼16-session average in randomized trials; Catanzano et al., 2020 ; Weisz et al., 2017 ) are equivocal to shorter interventions.
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.