Measuring algorithmic bias to analyze the reliability of AI tools that predict depression risk using smartphone sensed-behavioral data

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract AI tools intend to transform mental healthcare by providing remote estimates of depression risk using behavioral data collected by sensors embedded in smartphones. While these tools accurately predict elevated symptoms in small, homogenous populations, recent studies show that these tools are less accurate in larger, more diverse populations. In this work, we show that accuracy is reduced because sensed-behaviors are unreliable predictors of depression across individuals; specifically the sensed-behaviors that predict depression risk are inconsistent across demographic and socioeconomic subgroups. We first identified subgroups where a developed AI tool underperformed by measuring algorithmic bias, where subgroups with depression were incorrectly predicted to be at lower risk than healthier subgroups. We then found inconsistencies between sensed-behaviors predictive of depression across these subgroups. Our findings suggest that researchers developing AI tools predicting mental health from behavior should think critically about the generalizability of these tools, and consider tailored solutions for targeted populations.
Full text 155,581 characters · extracted from preprint-html · click to expand
Measuring algorithmic bias to analyze the reliability of AI tools that predict depression risk using smartphone sensed-behavioral data | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Measuring algorithmic bias to analyze the reliability of AI tools that predict depression risk using smartphone sensed-behavioral data Daniel A. Adler, Caitlin A. Stamatis, Jonah Meyerhoff, David C. Mohr, and 4 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-3044613/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 21 Apr, 2024 Read the published version in npj Mental Health Research → Version 1 posted 8 You are reading this latest preprint version Abstract AI tools intend to transform mental healthcare by providing remote estimates of depression risk using behavioral data collected by sensors embedded in smartphones. While these tools accurately predict elevated symptoms in small, homogenous populations, recent studies show that these tools are less accurate in larger, more diverse populations. In this work, we show that accuracy is reduced because sensed-behaviors are unreliable predictors of depression across individuals; specifically the sensed-behaviors that predict depression risk are inconsistent across demographic and socioeconomic subgroups. We first identified subgroups where a developed AI tool underperformed by measuring algorithmic bias, where subgroups with depression were incorrectly predicted to be at lower risk than healthier subgroups. We then found inconsistencies between sensed-behaviors predictive of depression across these subgroups. Our findings suggest that researchers developing AI tools predicting mental health from behavior should think critically about the generalizability of these tools, and consider tailored solutions for targeted populations. Health sciences/Medical research/Biomarkers/Predictive markers Biological sciences/Computational biology and bioinformatics/Machine learning Health sciences/Diseases/Psychiatric disorders/Depression Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Introduction Mental healthcare systems are simultaneously facing a shortage of mental health specialty care providers and a large number of patients whose treatment needs remain unmet 1,2 . This service gap is driving research into AI-driven mental health monitoring tools, where sensed-behavioral data, defined as inferred behavioral data gathered by sensors and software embedded in everyday devices (e.g. smartphones, wearables), are repurposed to remotely monitor depression symptoms 3–7 . Sensed-behavioral data has also been referred to as personal, behavioral, or passive sensing data in other work 7 . AI tools that leverage sensed-behavioral data intend to near-continuously identify individuals experiencing elevated symptoms in-between clinical encounters and consequently deliver preventive care 8 . These tools can also be integrated into digital therapeutics to automate precision interventions 9 . Initial work showed that depression risk could be predicted from sensed-behavioral data at a similar accuracy to general practitioners 10 in small populations 5,11 . More recent work shows that these AI tools predict depression risk at an accuracy only slightly better than a coin flip in larger, more diverse samples 4,6,12,13 . This prior work has not specifically explored why accuracy is reduced in larger samples, and it is not clear how to improve AI tools for clinical use. In this work, we hypothesized that accuracy is reduced in larger, more diverse populations because sensed-behaviors are unreliable predictors of depression risk: sensed-behaviors that predict depression are inconsistent across demographic and socioeconomic (SES) subgroups 14 . We intentionally use the term reliability due to its importance in both a psychometric and AI context. In a psychometric context, reliability refers to the consistency of a tool, typically a symptom assessment, across different contexts (e.g. raters, time) 14,15 . In AI, reliability is related to generalizability, if an AI tool is consistently accurate in different contexts (e.g. different populations, over time, etc.) 12 . Given these definitions, researchers in AI fairness have argued that aspects of psychometric reliability are important in an AI context: similar inputs (e.g. sensed-behaviors) to an AI model should yield similar outputs (e.g. estimated depression risk) 16 . In this paper, we adapt these ideas to study a specific aspect of reliability important for mental health AI tools deployed in large populations, i.e. if similar sensed-behaviors are consistently related to depression risk across different groups of individuals. We hypothesize that if the sensed-behaviors predictive of depression risk are inconsistent across groups, AI models that use sensed-behaviors to predict depression risk will be inaccurate because similar sensed-behavioral patterns will indicate different levels of depression risk for different subgroups. For example, imagine that mobility positively correlates with depression risk in subgroup A, and negatively correlates with depression risk in subgroup B. An AI model trained across subgroups using exclusively mobility data, blind to subgroup information as is typically the case in this literature 3–5,17 , will receive unreliable information – high mobility can simultaneously indicate both low and high depression risk – and will make incorrect predictions for one of the subgroups. We note upfront that in this manuscript we do not consider temporal aspects of reliability, though we acknowledge that this is important in discussions of psychometric reliability, specifically if the AI tool is consistently accurate for the same individual, with predictions made under similar conditions 18 . We tested this hypothesis by identifying population subgroups where a depression risk prediction tool underperformed, and then analyzed sensed-behavioral differences across these subgroups. We identified subgroups where the tool underperformed by measuring algorithmic ranking bias (hereafter referred to as “bias”), where individuals experiencing depression from one subgroup (e.g. older individuals) were incorrectly ranked by the tool to be at lower risk than healthier individuals from other subgroups (e.g. younger individuals) 19–22 . Reliability was analyzed by measuring ranking bias because if individuals in large populations have inconsistent relationships between sensed-behavior and mental health, behaviors that represent high depression risk for one subgroup may represent lower risk for another subgroup. For example, imagine an AI tool predicting that higher phone use increases depression risk. Studies 23,24 show that younger individuals have higher phone use than older individuals. Thus, the AI tool may incorrectly rank older individuals with depression to be at lower risk than healthier younger individuals, decreasing model accuracy (Fig. 1 a). Against this backdrop, we developed an AI tool that estimated depression symptom risk using behavioral data collected from individuals’ smartphones, using similar sensed-behaviors and outcome measures from recent work 3–5,13,25 (Fig. 1 b). The data used to develop and analyze the AI tool was collected during a U.S.-based National Institute of Mental Health (NIMH)-funded study 3,25–29 , one of the largest, most geographically diverse studies of its kind. We then measured bias across attributes including age, sex at birth, race, household income, health insurance and employment to identify subgroups where the tool underperformed. We studied these specific attributes because of known behavioral differences across demographic and SES subgroups 23,24,30–32 that could impact the reliability of the developed AI tool. Finally, we interpreted why the tool underperformed by identifying inconsistencies between the AI tool and sensed-behaviors predicting depression across subgroups. A summary of this analysis can be found in Fig. 1 . Results Data Collection We analyzed data from a U.S.-based, NIMH-funded study conducted from 2019–2021 to identify associations between behavioral data collected from smartphones and depression symptoms 3,25–29 . Smartphone sensed-behavioral data on GPS location, phone usage (timestamp of screen unlock), and sleep were near-continuously collected from participants across the United States for 16 weeks and the PHQ-8, a self-reported measure of two week depression symptoms 33,34 , frequently used in mental health research 3,5,25,27 , was collected every three weeks beginning on the first week of the study (e.g. on weeks 1, 4, 7, …, known as weekly reporting periods). Sensed-behaviors were summarized over two weeks to align with collected PHQ-8 depression symptoms for prediction (see Table 1 ). For example, sensed-behaviors collected during weeks 3 and 4 were summarized to predict PHQ-8 values collected during week 4. Table 1 Sensed-behaviors. An overview of the sensed-behavioral data used in this analysis. The same set of sensed-behaviors were collected from all participants, and were summarized over two week periods to align with self-reported PHQ-8 symptoms. Please see the methods for more details. Category Derived Sensed-Behaviors Location Variance (variability in GPS location), number of unique locations, entropy (variability in unique locations), normalized entropy (entropy normalized by number of unique locations), duration of time spent at home, percentage of collected samples in-transition (participant moving at > 1 km/hour), and circadian movement (24-hour regularity in movement). Location sensed-behaviors were directly calculated over two week periods. For example, we calculated the number of unique locations over two weeks. Phone usage Duration of phone usage and number of screen unlocks each day and within four 6-hour periods (12-6AM, 6-12PM, 12-6PM, 6-12AM). The average and standard deviation of each phone usage sensed-behavior was calculated over two weeks, and the number of days with daily phone use and use within each 6-hour period was summed. Sleep Average sleep onset (beginning of sleep), average duration, and variability in duration over two weeks. Table 2 summarizes the data used for analysis. 3,900 samples were analyzed from 650 individuals, a large cohort and sample size compared to most studies to date analyzing associations between sensed-behaviors and mental health 4,5,25,35,36 . A sample was a set of sensed-behaviors, summarized over two weeks, with a corresponding PHQ-8 self-report. 46% of self-reported PHQ-8 values were ≥ 10, indicating clinically-significant depression (CSD) 33 . The majority of participants were relatively young to middle aged (75% 25 to 54 years old), female (74%), white (82%), middle to high income (61% annual household income ≥ $ 40,000), insured (93%) and employed (62%). We focused our results on groups with at least 15 participants 37 . The sensed-behavior distributions across the population for each subgroup can be found in the supplementary materials. Table 2 Study cohort . Data was collected within an NIMH-funded study to understand the relationships between digitally-collected behavioral data and depression symptoms 3,25–29 . Participants contributed six total samples (summarized behavioral data and depression outcome measures) throughout the course of the study. A sample was a set of sensed-behaviors, summarized over two weeks, with a corresponding PHQ-8 self-report. Entire Study Number of participants 650 Samples per participant 6 Total number of samples 3,900 % Clinically-significant depression 46 Attribute Group Number of participants (%) Attribute Group Number of participants (%) Age 18 to 25 60 (9) Household Income < 20,000 98 (15) 25 to 34 181 (28) 20,000 to 39,999 144 (22) 35 to 44 168 (26) 40,000 to 59,999 124 (19) 45 to 54 135 (21) 60,000 to 99,999 161 (25) 55 to 64 81 (12) 100,000+ 110 (17) 65 to 74 22 (3) Don't know 10 (2) 75 to 84 3 (0) Prefer not to answer 3 (0) Sex at Birth Female 482 (74) Health Insurance Status Insured 603 (93) Male 168 (26) Uninsured 43 (7) Race White 534 (82) Don't know 3 (0) Black/African American 61 (9) Prefer not to answer 1 (0) Asian/Asian American 22 (3) Employment Status Employed 401 (62) More than one race 24 (4) Unemployed 90 (14) Other 6 (1) Disability 72 (11) Prefer not to answer 3 (0) Retired 34 (5) Other 52 (8) Prefer not to answer 1 (0) Measuring Bias to Identify Subgroups where AI Models Underperform The PHQ-8 asked participants to self-report depression symptoms experienced over 14 days, and PHQ-8’s were delivered multiple times throughout each weekly reporting period. We trained AI models using 14 days of smartphone sensed-behavioral data to predict if the average PHQ-8 value across each weekly reporting period (days 7 through 14, see Fig. 1 c) indicated clinically-significant depression (CSD, PHQ-8 score ≥ 10 33 ) symptoms. Surveys were delivered multiple times each reporting week, and individual surveys may only reflect “briefly” elevated symptoms (e.g. work stress on the day the survey was administered). For this reason, PHQ-8 values were averaged over each reporting week to predict a more stable estimate of self-reported symptoms. In addition, by predicting average PHQ-8 values instead of individual self-reports, we could make use of all available data while ensuring that samples did not overlap temporally. Model performance was assessed by performing 5-fold cross-validation, partitioning on subjects, and predictions across folds were concatenated to calculate model performance. To analyze performance variability due to specific cross-validation splits, we performed 100 cross-validation trials, shuffling participants into different folds during each trial. AI models output a predicted risk score from 0–1 of experiencing CSD. We used the predicted risk to calculate common ranking bias metrics 20–22 (Fig. 2 ) across the subgroups in Table 2 . These metrics were based upon the area under the receiver operating curve (AUC), which measured the probability models correctly predicted that CSD samples were ranked higher (in the predicted risk) than RH samples. We first calculated the AUC within each subgroup (the “Subgroup AUC”). Note that equal Subgroup AUCs do not guarantee high AUC across an entire sample. For example, Fig. 2 a shows simulated data where an algorithm correctly predicted CSD risk within subgroups, but younger individuals, compared to older individuals, had higher predicted risk overall. Thus, healthy younger individuals were incorrectly predicted to be at higher risk than older individuals experiencing CSD. Two additional performance metrics assessed such errors. Specifically, the background-negative-subgroup-positive, or BNSP AUC (Fig. 2 b) measured the probability that individuals experiencing CSD (the “positive” label) from a subgroup were correctly predicted to have higher risk than RH (the “negative label”) individuals from other subgroups (“the background”), and the background-positive-subgroup-negative, or BPSN AUC (Fig. 2 c), measured the probability RH individuals from a subgroup were correctly predicted to have lower risk than background individuals experiencing CSD. The highest performing AI model (a random forest, 100 trees, max depth of 10, balanced class weights, see methods) achieved a median (95% confidence interval, CI) AUC of 0.55 (0.54 to 0.57) across trials. Note that this low AUC was expected: it is comparable to the cross-validation performance of similar depression symptom prediction tools developed in larger, more diverse populations 4,6,13 , and motivates the objective of this work to study the reliability of these tools in larger populations. Figure 3 shows the model results by each metric across subgroups. The Subgroup AUC was lower for males (median, 95% CI 0.52, 0.49 to 0.55), Black/African Americans (0.50, 0.46 to 0.54), individuals from low income households (< $ 20,000, 0.46, 0.43 to 0.50), uninsured (0.45, 0.41 to 0.51), and unemployed (0.46, 0.42 to 0.50) individuals, compared to the median subgroup AUC for each attribute (e.g. “Sex at Birth”) across trials. The BNSP AUC increased with age (from 0.50, 0.46 to 0.52 for 18 to 25 year olds, to 0.67, 0.62 to 0.73 for 65 to 74 year olds), but decreased with household income (from 0.60, 0.58 to 0.63 for individuals from < $ 20,000 income households, to 0.45, 0.42 to 0.48 for individuals from $ 100,000 + income households). Individuals who were White (0.49, 0.46 to 0.52), male (0.52, 0.49 to 0.55), insured (0.47, 0.43 to 0.50), employed (0.43, 0.41 to 0.45), or identified with an “Other” type of employment (0.55, 0.52 to 0.59) also had lower BNSP AUC, compared to the median BNSP AUC for each attribute. The BPSN AUC findings showed complementary trends: RH older individuals (e.g. 65 to 74, 0.46, 0.40 to 0.50), unemployed (0.38, 0.36 to 0.41), uninsured (0.47, 0.43 to 0.50), Black/African American (0.48, 0.45 to 0.50), females (0.52, 0.49 to 0.55), and individuals coming from lower income households (e.g. < $ 20,000 0.42, 0.39 to 0.44) had a lower BPSN AUC. Results were reasonably consistent across different types of models, within subgroup base rates (% samples with PHQ-8 ≥ 10) were sometimes, but not always, associated with the BNSP/BPSN AUC, and subgroup sample size did not appear to be associated with the Subgroup AUC (see supplementary materials). Isolating the Effects of Subgroup Membership on Model Underperformance We wished to account for intersectional identities (e.g. female and employed) and isolate the effect of subgroup membership on model underperformance. For an ideal classifier, the predicted risk would be low for RH subgroups, and high for CSD subgroups. In addition, we would expect groups with higher base rates (% of samples with PHQ-8 ≥ 10) to have a higher average predicted risk. We thus modeled expected differences from groups with either the lowest (for RH) or highest (for CSD) average risk across trials. Generalized estimating equations (GEE, exchangeable correlation structure) 38 , a type of linear regression, was used to estimate the average effect of subgroup membership on the predicted risk after controlling across all other attributes. GEE was used instead of linear regression to correct for the non-independence of repeated samples across trials 38 . The regression results can be found in Fig. 4 . The individuals with the lowest average predicted risk who were not experiencing depression were 18 to 25 years old, male, White, had a household income of $ 100,000+, were insured, and employed. The predicted risk was expected to be higher (95% CI lower-bound > 0) for RH individuals who were older than 34 (e.g. for 65 to 74 year olds, mean, 95% confidence interval 0.02, 0.01 to 0.04), identified as Asian/Asian American (0.02, 0.01 to 0.03), Black/African American (0.01, 0.00 to 0.01), came from < $ 60,000 income households (e.g. for < $ 20,000, 0.02, 0.01 to 0.03), were unemployed (0.03, 0.03 to 0.04), and/or on disability (0.01, 0.00 to 0.02). For individuals who were experiencing CSD, models predicted the highest average risk for 65 to 74 year olds, Females, Asian/Asian Americans, individuals who came from households with incomes of $ 20,000 to $ 39,999, were insured, and/or retired. The predicted risk for individuals experiencing CSD was expected to be lower (95% CI upper-bound < 0) if individuals were 18 to 25 (–0.02, − 0.04 to − 0.01), male (–0.01, − 0.02 to − 0.00), Black/African American (–0.02, − 0.03 to − 0.00), more than one race (–0.02, − 0.03 to − 0.00), White (–0.02, − 0.03 to − 0.01), came from any household with an annual income < $ 20,000 or ≥ $ 40,000 (e.g. $ 100,000+ − 0.03, − 0.03 to − 0.02), and/or were employed (–0.02, − 0.03 to − 0.01). Predicted risk distributions often overlapped across subgroups with higher or lower risk, though there were general trends across subgroups (e.g. the predicted risk increased with age and unemployment in RH individuals, and risk decreased with income level for both CSD and RH individuals, see Fig. 4 for more details). Interpreting Sensed-Behaviors across Subgroups where Models Underperformed We hypothesized that models underperformed because sensed-behaviors predictive of CSD were inconsistent across subgroups. We thus conducted an analysis to understand differences between how AI tools predicted CSD risk and the different relationships between sensed-behaviors and CSD across subgroups. First, we retrained the AI model on the entire data, and used Shapley additive explanations (SHAP) 39 to interpret how the AI tool predicted CSD risk from sensed-behaviors. We then compared SHAP values with coefficients from explanatory logistic regression models estimating how subgroup membership affected the relationship between each sensed-behavior and depression. We found different relationships between the SHAP values (Fig. 5 a) and sensed-behaviors associated with CSD across subgroups (Fig. 5 b, comparisons across each attribute and feature can be found in the supplementary materials). For example, the AI tool predicted that higher morning phone usage (6AM − 12PM) was generally associated with lower predicted depression risk. Higher morning phone usage decreased depression risk for 18 to 25 year olds (mean, 95% CI effect on depression, standardized units: − 0.77, − 1.07 to − 0.47), but increased risk for 65 to 74 year olds (0.60, 0.07 to 1.12). Younger individuals, overall, also had higher morning phone use (standardized median, 95% CI 18 to 25 year olds: 0.32, − 2.27 to 1.60) compared to older individuals (65 to 74 year olds: − 0.62, − 1.96 to 0.76). Figure 5 a also shows that specific mobility features, including the circadian movement (regularity in 24-hour movement), location entropy (regularity in travel to unique locations), and the percentage of collected GPS samples in transition (approximated speed > 1 km/hour) were often associated with lower predicted CSD risk. Circadian movement decreased CSD risk for employed individuals (–0.16, − 0.24 to − 0.07), but increased CSD risk for individuals who were on disability (0.44, 0.21 to 0.66). Circadian movement and location entropy also decreased depression risk for individuals from middle income ( $ 60,000 to $ 99,999) households (circadian movement: − 0.21, − 0.35 and − 0.07; location entropy: − 0.34, − 0.48 to − 0.20), but increased risk for individuals from low income (< $ 20,000) households (circadian movement: 0.30, 0.09 to 0.51; location entropy: 0.35, 0.14 to 0.57). Finally, a higher percentage of GPS samples in transition decreased depression risk for insured individuals (–0.15, − 0.22 to − 0.08), but increased risk for uninsured individuals (0.32, 0.11 to 0.52). Discussion In this study, we hypothesized that sensed-behaviors are unreliable measures of depression in larger populations, reducing the accuracy of AI tools that use sensed-behaviors to predict depression risk. To test this hypothesis, we developed an AI tool that predicted clinically-significant depression (CSD) from sensed-behaviors and measured algorithmic bias to identify specific age, race, sex at birth, and socioeconomic subgroups where the tool underperformed. We then found differences between SHAP values estimating how the AI tool predicted CSD from sensed-behaviors, and explanatory logistic regression models estimating the associations between sensed-behaviors and CSD across subgroups. In this discussion, we show how differences in sensed-behaviors across subgroups may explain the identified bias and AI underperformance in larger, more diverse populations. Measuring bias showed that models predicted older, female, Black/African American, low income, unemployed, and individuals on disability were at higher risk of experiencing CSD (high BNSP, low BPSN AUC), and younger, male, White, high income, insured, and employed individuals were at lower risk of experiencing CSD (high BPSN, low BNSP AUC), independent of outcomes. Comparing SHAP values to explanatory logistic regression coefficients suggests why AI models incorrectly predicted depression risk. For example, our findings show that younger individuals had higher daytime phone usage than older individuals. Models predicted that higher daytime phone usage was associated with lower CSD risk (Fig. 5 a), potentially explaining why younger individuals, overall, had lower predicted risk, and older adults had higher predicted risk (Fig. 3 ). Differences could be attributed to younger individuals using phones for entertainment and social activities that support well-being, while older individuals may prefer to use their phones for necessary communication or information gathering 23 . In another example, the model predicted that mobility, measured through circadian movement, location entropy, and GPS samples in transition, was associated with lower CSD risk (see Fig. 5 ). Prior work has identified a negative association between these same mobility features and CSD 5,25 , suggesting that mobility decreases depression risk. While we found the expected negative associations across majority, higher SES ( $ 60,000 to 99,999 household income, insured, and employed) subgroups, we found the opposite, positive association across less-represented lower SES (< $ 20,000 household income, on disability, uninsured) subgroups, potentially explaining the reduced model performance (lower Subgroup AUC) in these groups. There are many possible explanations for the identified differences in behavior. First, underlying reasons to be mobile (e.g. navigating bureaucracy to receive government payments) may increase stress for individuals who are lower income and/or on disability 31 , increasing depression risk. Second, the analyzed data was collected during the early-to-mid stages of the COVID-19 pandemic, when mobility for low SES essential workers may indicate work travel and increased COVID-19 risk, contributing to stress 32 and depression. These findings suggest that sensed-behaviors approximating phone use and mobility used to predict depression in prior work 3–6,25 do not reliably predict depression in larger populations because of subgroup differences. While existing work developing similar AI tools has strived to achieve generalizability 4,40 , our findings question this goal. Instead, it may be more practical to improve reliability by developing models for specific, targeted populations 41,42 . In addition, it may be helpful to train AI models using both sensed-behaviors and demographic information. In prior work and this study, AI models were trained using exclusively sensed-behavioral data 3–5,17 . However, prior work suggests that models may not be more predictive even with added demographic information 43 . This shows that additional methods are needed to clearly define subgroups, beyond demographics, with more homogenous relationships between sensed-behaviors and depression symptoms. Another method to improve reliability is to develop personalized models, trained on participants’ data over time 6,44 . While personalization seems appealing, researchers should ensure that personalized predictions are meaningful. For example, we experimented with personalized models using a procedure suggested from prior work 44 . The model AUC improved (0.68) compared to the presented results (0.55), but we achieved a better AUC (> 0.80) by developing a naive model re-predicting participants’ first self-reported PHQ-8 value for all future outcomes. Given at least one participant self-report is often needed for personalization, models should show greater accuracy than these naive benchmarks. Even if accuracy improves, models can still be biased 19,37 , and it is important to consider the clinical and public health implications of using biased risk scores for depression screening. For example, more frequent exposure to stress 45 contributes to higher rates of depression in lower SES populations 46 , but overestimating depression risk for healthy low SES individuals allocates mental health resources away from other individuals who need care. Similar issues persist for underestimating depression risk. For example, models predicted lower risk for males experiencing depression compared to healthier females (see Fig. 3 ). Males are less likely to seek treatment for their mental health than females 47 , and AI tools underestimating male depression risk may further reduce the likelihood that males seek care. Uncovering these biases are important before algorithmic tools are used in clinical settings. To reduce these harms, researchers can use methods described in this and other work 37 to identify subgroups where AI tools underperform by measuring bias. Resources could then be directed to develop new or retrain existing models for these subgroups. Simultaneously, clinical personnel using these tools can be trained to identify algorithmic bias and mitigate its effects 48 . In addition, depositing de-identified sensed-behavior and mental health outcomes data in research repositories could increase available data to analyze the reliability of AI tools 12 . Finally, our findings show the importance of developing AI tools using data from populations that have similar behavioral patterns to the populations where these tools will be deployed. More thorough reporting of model training data 49 , and monitoring AI tools in “silent mode”, in which predictions are made but not used for decision making 50 , could prevent AI tools developed in dissimilar populations from causing harm. Finally, it is important to consider how the choice to classify depression symptom severity influenced our results, specifically choosing to predict binarized PHQ-8 values instead of raw PHQ-8 scores. Predicting binarized symptom scores is a fairly common practice in both the depression prediction literature 3–5,17 , as well as in the clinical AI literature, broadly 51,52 . This practice is motivated by an interest to use AI tools for near-continuous symptom monitoring, in which an action (e.g. follow-up by a care provider) is triggered at a specific elevated symptom threshold. This motivation may be difficult to realize if the field continues to use depression symptom scales as outcomes; as recent work shows, these symptom scales do not produce categorical response distributions, with a clear decision boundary distinguishing individuals experiencing versus not experiencing symptoms. Instead, responses tend to exist along a continuum 14 . It is also important to consider if subgroup differences affect the interpretation and self-reporting of depression symptom scales. Despite this consideration, prior work provides evidence that the PHQ-8 exhibits measurement invariance across demographic and socioeconomic groups 53,54 . Thus, it may be unlikely that the bias identified in this work was due to group differences in self-reporting symptoms, but our findings could be partially attributed to the mistreatment of depression symptom scales as categorical in nature. This work had limitations. First, we analyzed data from a single study, though the studied cohort was larger in size, geographic representation, and timespan compared to prior work. In addition, the study cohort was majority White, employed, and female, though we did not find that sample size was associated with classification accuracy. Only inter-individual variability was considered, not intra-individual variability, and thus these findings do not extend to longitudinal monitoring contexts, where changes in sensed-behaviors may indicate changes in depression risk. In addition, data was only analyzed from participants who provided complete outcomes data (participants who reported at least one PHQ-8 value during each of the 6 weekly reporting periods). Data was exclusively collected from individuals who owned Android devices, only specific data types (GPS and phone usage) were analyzed. Only smartphone sensed-behaviors were analyzed, and data collected from other types of devices (e.g. wearables) were not analyzed. Finally, data collection took place from 2019 to 2021, when COVID-19 restrictions varied across the United States, which may influence our findings. Future work can examine if these results replicate over larger, more diverse cohorts, in both demographic and socioeconomic attributes, as well as the devices used for data collection. In addition, future work can explore if sensed-behaviors are reliable predictors of depression in longitudinal monitoring contexts, though recent work suggests that sensed-behaviors have low predictive power, even when used for longitudinal monitoring 25 . In conclusion, we present one method to assess the reliability of AI tools that use sensed-behaviors to predict depression risk. Specifically we measured ranking bias in a developed AI tool to identify subgroups where the tool underperformed, and then we interpreted why models underperformed by comparing the AI tool to sensed-behaviors predictive of depression across subgroups. Researchers and practitioners developing AI-driven mental health monitoring tools using behavioral data should think critically about whether these tools are likely to generalize, and consider developing tailored solutions that are well-validated in specific, targeted populations. Methods Cohort In this work, we performed a secondary analysis of data collected during a U.S.-based National Institute of Mental Health (NIMH) funded study. The motivation for this study was to identify smartphone sensed-behavioral patterns predictive of major depressive disorder (MDD) 3,25–29 . Participants were recruited from across the United States using digital registries and online advertisements, intentionally oversampling for individuals experiencing depression. Eligible participants lived in the United States, could read/write English, and owned an Android smartphone and data plan. In addition, eligible participants with at least moderate depression symptom severity based upon the Patient Health Questionnaire-8 (PHQ-8) ≥ 10 were oversampled to create a sample with elevated symptoms. Individuals were excluded from the study if they self-reported a diagnosis of bipolar disorder, any psychotic disorder, shared a smartphone with another individual, or were unwilling to share data. Eligible participants were asked to provide electronic informed consent after receiving a complete description of the study. Eligible participants had the option to not provide consent, and could withdraw from the study at any point. Consented participants downloaded a study smartphone application 55 and completed a baseline assessment to self-report demographic and lifestyle information. The study application passively collected GPS location, sampled every 5 minutes, and smartphone interactions (screen unlock and time of unlock) for 16 weeks. Individuals completed depression symptom assessments every 3 weeks within the smartphone application (the PHQ-8 33,34 ). Data collection took place from 2019–2021, and all study procedures were approved by the Northwestern University Institutional Review Board (study #STU00205316). Sensed-Behavioral Features We calculated sensed-behavioral features from the collected smartphone data to predict depression risk. Following established methods from prior work 3,5,25 , we calculated GPS mobility features including the location variance (variability in GPS), number of unique locations, location entropy (variability in unique locations), normalized location entropy (entropy normalized by number of unique locations), duration of time spent at home, percentage of collected samples in-transition (participant moving at > 1 km/hour), and circadian movement (24-hour regularity in movement) 5 . We also calculated phone usage features from the screen unlock data 40 , including the duration of phone use and the number of screen unlocks each day and within four 6-hour periods (12-6AM, 6-12PM, 12-6PM, 6-12AM). Finally, we used a standard algorithm 40,56 to approximate daily sleep onset and duration from screen unlock data. Classifying Depression Symptoms The PHQ-8 asks participants to self-report depression symptoms that occurred during the past two weeks. Symptoms are reported from 0 (not experiencing the symptom) to 3 (frequently experiencing the symptom). Scores are summed and thresholded to classify severity, where summed scores of 10 or greater indicate a higher likelihood of experiencing a clinically-significant depression 33 . We thus followed prior work 5,25 to calculate sensed-behavioral features in the two week period up to and including each weekly PHQ-8 reporting period. Behavioral features were input into machine learning models to predict clinically-significant symptoms (PHQ-8 ≥ 10). Data Preprocessing Screen unlock and sleep features were summarized to align with the PHQ-8 40 . The average and standard deviation of each daily and 6-hour epoch feature were calculated across the two week prediction period, and the number of days with daily phone use and use within each 6-hour epoch were summed. GPS features were directly calculated over the two weeks. As recommended by Saeb et al. 5 , skewed features were log-transformed. Missing data was filled using multivariate imputation 57 and then standardized (mean = 0, standard deviation = 1) based upon each training dataset prior to being input into predictive models. AI Model Training and Validation We trained machine learning models commonly used to predict mental health status from smartphone behavioral data including regularized (L2-norm) logistic regression (LR) 3,5 , support vector machines (SVM) 4,58 , and tree-based ensemble models including random forest (RF) and gradient boosting trees (GBT) 3,40 . We varied the strength of the LR and SVM regularization parameter (0.01, 0.1, 1.0), used a radial basis function SVM kernel, varied class balancing weights in the RF and SVM (unbalanced/balanced), varied the number of ensemble tree estimators (10, 100), depth (3, 10, or until pure), and the GBT learning rate (0.01, 0.1, 1.0) and loss (deviance and exponential). Non-logistic prediction models were calibrated using Platt scaling to approximately match the predicted risk to the proportion of individuals experiencing clinically-significant symptoms at each risk level 59 . Logistic regression models, as shown in prior work 59 , output calibrated probabilities. Models were implemented using the scikit-learn Python library 60 . Multiple PHQ-8 surveys were administered each weekly reporting period (e.g. week 1, 4, 7, etc.). Survey scores in each reporting week were averaged to remove overlap between sensor and outcomes data. Data was analyzed from study participants who self-reported at least one PHQ-8 during each reporting week, resulting in 6 predictions per participant. Data from all other participants were removed (288 participants removed, 31% of total) to focus this analysis towards algorithmic bias due to group differences rather than bias due to missing outcomes data. Declarations Competing Interests D.A. and T.C. have submitted patent applications related to this work. T.C. is a co-founder and equity holder of HealthRhythms, Inc. and has received grants from Click Therapeutics related to digital therapeutics. D.C.M has accepted honoraria and consulting fees from Boehringer-Ingelheim, Otsuka Pharmaceuticals, Optum Behavioral Health, Centerstone Research Institute, and the One Mind Foundation, royalties from Oxford Press, and has an ownership interest in Adaptive Health, Inc. J.M. has accepted consulting fees from Boehringer Ingelheim. G.J.A. holds equity in HealthRhythms, Inc. and Lyra Health, Inc., and has accepted consulting fees and honoraria from BetterUp and Quantum Health. Author Contributions D.A. conducted the analysis and wrote the draft manuscript. T.C., F.W., J.M., C.A.S., and D.C.M. provided supervisory support throughout the analysis. D.C.M. and J.M. were involved in data collection. All authors participated in drafting and revising the manuscript. Acknowledgements D.A. is supported by the National Science Foundation Graduate Research Fellowship under Grant No. DGE-2139899, and a Digital Life Initiative Doctoral Fellowship. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the authors and do not necessarily reflect the views of the funders. Data collection was supported by NIMH Grant No. R01MH111610 to D.C.M. J.M. is supported by K08MH128640. C.A.S. is supported by T32MH115882. Computing costs were funded by a Microsoft Azure Cloud Computing Grant through the Cornell Center for Data Science for Enterprise and Society, awarded to T.C. T.C. and F.W. were also supported by a multi-investigator seed grant awarded from the Cornell Academic Integration program. Data Availability Sensed-behavioral data cannot be made publicly available due to potentially identifying information (e.g. GPS location) that may compromise participant privacy. De-identified self-reported data (the PHQ-8) will be made available through the NIMH Data Archive. Code Availability A repository for all code used for analysis can be found at the following link: https://github.com/dadler6/reliability_depression_ml . References Cai, A. et al. Trends In Mental Health Care Delivery By Psychiatrists And Nurse Practitioners In Medicare, 2011–19. Health Aff. (Millwood) 41, 1222–1230 (2022). Mohr, D. C. et al. Banbury Forum Consensus Statement on the Path Forward for Digital Mental Health Treatment. Psychiatr. Serv. appi.ps.202000561 (2021) doi: 10.1176/appi.ps.202000561 . Liu, T. et al. The relationship between text message sentiment and self-reported depression. J. Affect. Disord. 302, 7–14 (2022). Xu, X. et al. GLOBEM: Cross-Dataset Generalization of Longitudinal Human Behavior Modeling. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 6, 190:1-190:34 (2023). Saeb, S. et al. Mobile Phone Sensor Correlates of Depressive Symptom Severity in Daily-Life Behavior: An Exploratory Study. J. Med. Internet Res. 17, (2015). Meegahapola, L. et al. Generalization and Personalization of Mobile Sensing-Based Mood Inference Models: An Analysis of College Students in Eight Countries. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 6, 176:1-176:32 (2023). Mohr, D. C., Shilton, K. & Hotopf, M. Digital phenotyping, behavioral sensing, or personal sensing: names and transparency in the digital age. Npj Digit. Med. 3, 1–2 (2020). Lee, E. E. et al. Artificial Intelligence for Mental Health Care: Clinical Applications, Barriers, Facilitators, and Artificial Wisdom. Biol. Psychiatry Cogn. Neurosci. Neuroimaging S245190222100046X (2021) doi: 10.1016/j.bpsc.2021.02.001 . Frank, E. et al. Personalized digital intervention for depression based on social rhythm principles adds significantly to outpatient treatment. Front. Digit. Health 4, (2022). Mitchell, A. J., Vaze, A. & Rao, S. Clinical diagnosis of depression in primary care: a meta-analysis. The Lancet 374, 609–619 (2009). Wang, R. et al. Tracking Depression Dynamics in College Students Using Mobile Phone and Wearable Sensing. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 2, 43:1–43:26 (2018). Adler, D. A. et al. A call for open data to develop mental health digital biomarkers. BJPsych Open 8, (2022). Müller, S. R., Chen, X. (Leslie), Peters, H., Chaintreau, A. & Matz, S. C. Depression predictions from GPS-based mobility do not generalize well to large demographically heterogeneous samples. Sci. Rep. 11, 14007 (2021). Fried, E. I., Flake, J. K. & Robinaugh, D. J. Revisiting the theoretical and methodological foundations of depression measurement. Nat. Rev. Psychol. 1–11 (2022) doi: 10.1038/s44159-022-00050-2 . Beck, A. T. Reliability of psychiatric diagnoses: 1. a critique of systematic studies. Am. J. Psychiatry 119, 210–216 (1962). Jacobs, A. Z. & Wallach, H. Measurement and Fairness. Proc. 2021 ACM Conf. Fairness Account. Transpar. 375–385 (2021) doi: 10.1145/3442188.3445901 . Jacobson, N. C., Weingarden, H. & Wilhelm, S. Digital biomarkers of mood disorders and symptom change. Npj Digit. Med. 2, 1–3 (2019). Boateng, G. O., Neilands, T. B., Frongillo, E. A., Melgar-Quiñonez, H. R. & Young, S. L. Best Practices for Developing and Validating Scales for Health, Social, and Behavioral Research: A Primer. Front. Public Health 6, (2018). Obermeyer, Z., Powers, B., Vogeli, C. & Mullainathan, S. Dissecting racial bias in an algorithm used to manage the health of populations. Science 366, 447–453 (2019). Borkan, D., Dixon, L., Sorensen, J., Thain, N. & Vasserman, L. Nuanced Metrics for Measuring Unintended Bias with Real Data for Text Classification. Preprint at http://arxiv.org/abs/1903.04561 (2019). Kallus, N. & Zhou, A. The Fairness of Risk Scores Beyond Classification: Bipartite Ranking and the XAUC Metric. in Advances in Neural Information Processing Systems vol. 32 (Curran Associates, Inc., 2019). Vogel, R., Bellet, A. & Clémençon, S. Learning Fair Scoring Functions: Bipartite Ranking under ROC-based Fairness Constraints. in Proceedings of The 24th International Conference on Artificial Intelligence and Statistics 784–792 (PMLR, 2021). Andone, I. et al. How age and gender affect smartphone usage. in Proceedings of the 2016 ACM International Joint Conference on Pervasive and Ubiquitous Computing: Adjunct 9–12 (ACM, 2016). doi: 10.1145/2968219.2971451 . Horwood, S., Anglim, J. & Mallawaarachchi, S. R. Problematic smartphone use in a large nationally representative sample: Age, reporting biases, and technology concerns. Comput. Hum. Behav. 122, 106848 (2021). Meyerhoff, J. et al. Evaluation of Changes in Depression, Anxiety, and Social Anxiety Using Smartphone Sensor Features: Longitudinal Cohort Study. J. Med. Internet Res. 23, e22844 (2021). Mohr, D. C. LifeSense: Transforming Behavioral Assessment of Depression Using Personal Sensing Technology. https://reporter.nih.gov/search/N6YCr94ZvkOVUNu1i5HNaQ/project-details/9982127 (2017). Stamatis, C. A. et al. Prospective associations of text-message-based sentiment with symptoms of depression, generalized anxiety, and social anxiety. Depress. Anxiety 39, 794–804 (2022). Meyerhoff, J. et al. Analyzing text message linguistic features: Do people with depression communicate differently with their close and non-close contacts? Behav. Res. Ther. 166, 104342 (2023). Stamatis, C. A. et al. The association of language style matching in text messages with mood and anxiety symptoms. Procedia Comput. Sci. 206, 151–161 (2022). Greissl, S. et al. Is unemployment associated with inefficient sleep habits? A cohort study using objective sleep measurements. J. Sleep Res. 31, e13516 (2022). Iezzoni, L. I., McCarthy, E. P., Davis, R. B. & Siebens, H. Mobility Difficulties Are Not Only a Problem of Old Age. J. Gen. Intern. Med. 16, 235–243 (2001). Levy, B. L., Vachuska, K., Subramanian, S. V. & Sampson, R. J. Neighborhood socioeconomic inequality based on everyday mobility predicts COVID-19 infection in San Francisco, Seattle, and Wisconsin. Sci. Adv. 8, eabl3825 (2022). Kroenke, K. et al. The PHQ-8 as a measure of current depression in the general population. J. Affect. Disord. 114, 163–173 (2009). Wu, Y. et al. Equivalency of the diagnostic accuracy of the PHQ-8 and PHQ-9: a systematic review and individual participant data meta-analysis. Psychol. Med. 50, 1368–1380 (2020). Opoku Asare, K. et al. Predicting Depression From Smartphone Behavioral Markers Using Machine Learning Methods, Hyperparameter Optimization, and Feature Importance Analysis: Exploratory Study. JMIR MHealth UHealth 9, e26540 (2021). Corponi, F. et al. Automated mood disorder symptoms monitoring from multivariate time-series sensory data: Getting the full picture beyond a single number. 2023.03.25.23287744 Preprint at https://doi.org/10.1101/2023.03.25.23287744 (2023). Seyyed-Kalantari, L., Zhang, H., McDermott, M. B. A., Chen, I. Y. & Ghassemi, M. Underdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under-served patient populations. Nat. Med. 27, 2176–2182 (2021). Ballinger, G. A. Using Generalized Estimating Equations for Longitudinal Data Analysis. Organ. Res. Methods 7, 127–150 (2004). Lundberg, S. M. & Lee, S.-I. A Unified Approach to Interpreting Model Predictions. in Advances in Neural Information Processing Systems vol. 30 (Curran Associates, Inc., 2017). Adler, D. A., Wang, F., Mohr, D. C. & Choudhury, T. Machine learning for passive mental health symptom prediction: Generalization across different longitudinal mobile sensing studies. PLOS ONE 17, e0266516 (2022). Sperrin, M., Riley, R. D., Collins, G. S. & Martin, G. P. Targeted validation: validating clinical prediction models in their intended population and setting. Diagn. Progn. Res. 6, 24 (2022). Mitchell, M. et al. Model Cards for Model Reporting. ArXiv181003993 Cs (2019) doi: 10.1145/3287560.3287596 . Pratap, A. et al. The accuracy of passive phone sensors in predicting daily mood. Depress. Anxiety 36, 72–81 (2019). Wang, R. et al. CrossCheck: toward passive sensing and detection of mental health changes in people with schizophrenia. in Proceedings of the 2016 ACM International Joint Conference on Pervasive and Ubiquitous Computing - UbiComp ’16 886–897 (ACM Press, 2016). doi: 10.1145/2971648.2971740 . Williams, D. R., Mohammed, S. A., Leavell, J. & Collins, C. Race, Socioeconomic Status and Health: Complexities, Ongoing Challenges and Research Opportunities. Ann. N. Y. Acad. Sci. 1186, 69–101 (2010). Everson, S. A., Maty, S. C., Lynch, J. W. & Kaplan, G. A. Epidemiologic evidence for the relation between socioeconomic status and depression, obesity, and diabetes. J. Psychosom. Res. 53, 891–895 (2002). Chatmon, B. N. Males and Mental Health Stigma. Am. J. Mens Health 14, 1557988320949322 (2020). Rajkomar, A., Hardt, M., Howell, M. D., Corrado, G. & Chin, M. H. Ensuring Fairness in Machine Learning to Advance Health Equity. Ann. Intern. Med. 169, 866–872 (2018). Gebru, T. et al. Datasheets for Datasets. ArXiv180309010 Cs (2020). Wiens, J. et al. Do no harm: a roadmap for responsible machine learning for health care. Nat. Med. 25, 1337–1340 (2019). Wong, A. et al. External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients. JAMA Intern. Med. (2021) doi: 10.1001/jamainternmed.2021.2626 . Adams, R. et al. Prospective, multi-site study of patient outcomes after implementation of the TREWS machine learning-based early warning system for sepsis. Nat. Med. 1–6 (2022) doi: 10.1038/s41591-022-01894-0 . Galenkamp, H., Stronks, K., Snijder, M. B. & Derks, E. M. Measurement invariance testing of the PHQ-9 in a multi-ethnic population in Europe: the HELIUS study. BMC Psychiatry 17, 349 (2017). Villarreal-Zegarra, D., Copez-Lonzoy, A., Bernabé-Ortiz, A., Melendez-Torres, G. J. & Bazo-Alvarez, J. C. Valid group comparisons can be made with the Patient Health Questionnaire (PHQ-9): A measurement invariance study across groups by demographic characteristics. PLOS ONE 14, e0221717 (2019). Audacious Software. Passive Data Kit. https://passivedatakit.org/ . Abdullah, S., Matthews, M., Murnane, E. L., Gay, G. & Choudhury, T. Towards circadian computing: ‘early to bed and early to rise’ makes some of us unhealthy and sleep deprived. in Proceedings of the 2014 ACM International Joint Conference on Pervasive and Ubiquitous Computing 673–684 (ACM, 2014). doi: 10.1145/2632048.2632100 . Buuren, S. van & Groothuis-Oudshoorn, K. mice: Multivariate Imputation by Chained Equations in R. J. Stat. Softw. 45, 1–67 (2011). Tseng, V. W.-S. et al. Using behavioral rhythms and multi-task learning to predict fine-grained symptoms of schizophrenia. Sci. Rep. 10, 15100 (2020). Niculescu-Mizil, A. & Caruana, R. Predicting good probabilities with supervised learning. in Proceedings of the 22nd international conference on Machine learning - ICML ’05 625–632 (ACM Press, 2005). doi: 10.1145/1102351.1102430 . Pedregosa, F. et al. Scikit-learn: Machine Learning in Python. ArXiv12010490 Cs (2018). Additional Declarations Competing interest reported. D.A. and T.C. have submitted patent applications related to this work. T.C. is a co-founder and equity holder of HealthRhythms, Inc. and has received grants from Click Therapeutics related to digital therapeutics. D.C.M has accepted honoraria and consulting fees from Boehringer-Ingelheim, Otsuka Pharmaceuticals, Optum Behavioral Health, Centerstone Research Institute, and the One Mind Foundation, royalties from Oxford Press, and has an ownership interest in Adaptive Health, Inc. J.M. has accepted consulting fees from Boehringer Ingelheim. G.J.A. holds equity in HealthRhythms, Inc. and Lyra Health, Inc., and has accepted consulting fees and honoraria from BetterUp and Quantum Health. Supplementary Files SupplementaryWorkbookonSensedBehaviors.xlsx SupplementaryInformation.docx nreditorialpolicychecklist20230823.pdf nrreportingsummary20230823.pdf Cite Share Download PDF Status: Published Journal Publication published 21 Apr, 2024 Read the published version in npj Mental Health Research → Version 1 posted Editorial decision: Revision requested 23 Dec, 2023 Reviews received at journal 21 Dec, 2023 Reviews received at journal 30 Oct, 2023 Reviewers agreed at journal 25 Sep, 2023 Reviewers agreed at journal 16 Sep, 2023 Reviewers invited by journal 16 Sep, 2023 Submission checks completed at journal 14 Sep, 2023 First submitted to journal 16 Aug, 2023 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-3044613","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":233373681,"identity":"24b6d0c6-fffa-4ec5-bb58-a319cdc70c37","order_by":0,"name":"Daniel A. Adler","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA0ElEQVRIiWNgGAWjYDCCAwyMDyR+2EA4DwrAIgS1MBtY9qRBOAkGxGlhk6hgO0yCFr7jhx9I3OA5L88vffbgA6AWOb4bCfi1SJ5JMzCcYXHbcGZfXrIBUIuxJCEtBgcSDJIleG4nGJzhMZMAakncQFDL+ecfDv9hOwfSYv4DqKWesJYbOYYNEmwHwLaAvJ9gQNAvN94UM0j2JBvO7OExBjpMwnDmmQf4tfCdT9/+Q+KHnTw/D4/hhw8VNvJ8xwnYgg4kSFM+CkbBKBgFowA7AAA/yEbJttyGbgAAAABJRU5ErkJggg==","orcid":"","institution":"Cornell Tech","correspondingAuthor":true,"prefix":"","firstName":"Daniel","middleName":"A.","lastName":"Adler","suffix":""},{"id":233373682,"identity":"5c8a4f3f-5f4e-46c2-acaa-5fbeb3aa5706","order_by":1,"name":"Caitlin A. Stamatis","email":"","orcid":"","institution":"Northwestern University Feinberg School of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Caitlin","middleName":"A.","lastName":"Stamatis","suffix":""},{"id":233373683,"identity":"e40aa2d2-cb31-41b7-a35e-8e7f9f695d25","order_by":2,"name":"Jonah Meyerhoff","email":"","orcid":"","institution":"Northwestern University Feinberg School of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Jonah","middleName":"","lastName":"Meyerhoff","suffix":""},{"id":233373684,"identity":"d90030d3-e9a8-47df-abc1-c83cd4259b83","order_by":3,"name":"David C. Mohr","email":"","orcid":"","institution":"Northwestern University Feinberg School of Medicine","correspondingAuthor":false,"prefix":"","firstName":"David","middleName":"C.","lastName":"Mohr","suffix":""},{"id":233373685,"identity":"07bd631a-abbb-41dc-a233-c01ae42b474b","order_by":4,"name":"Fei Wang","email":"","orcid":"","institution":"Weill Cornell Medicine","correspondingAuthor":false,"prefix":"","firstName":"Fei","middleName":"","lastName":"Wang","suffix":""},{"id":233373686,"identity":"986a5cc9-7ca9-46c0-9291-c94b235b3723","order_by":5,"name":"Gabriel J. Aranovich","email":"","orcid":"","institution":"Cornell Tech","correspondingAuthor":false,"prefix":"","firstName":"Gabriel","middleName":"J.","lastName":"Aranovich","suffix":""},{"id":233373687,"identity":"1e46a1b7-2eb6-4db8-9160-1c24e9841687","order_by":6,"name":"Srijan Sen","email":"","orcid":"","institution":"Michigan Medicine","correspondingAuthor":false,"prefix":"","firstName":"Srijan","middleName":"","lastName":"Sen","suffix":""},{"id":233373688,"identity":"29597188-786c-44a9-9f05-202147e5c581","order_by":7,"name":"Tanzeem Choudhury","email":"","orcid":"","institution":"Cornell Tech","correspondingAuthor":false,"prefix":"","firstName":"Tanzeem","middleName":"","lastName":"Choudhury","suffix":""}],"badges":[],"createdAt":"2023-06-09 19:29:18","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-3044613/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-3044613/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1038/s44184-024-00057-y","type":"published","date":"2024-04-22T00:00:00+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":55063771,"identity":"89831f75-b276-4726-8a69-33aa77fb33bb","added_by":"auto","created_at":"2024-04-22 03:15:37","extension":"jpeg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":480561,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eAnalyzing reliability in AI tools that predict depression symptom risk. a. \u003c/strong\u003eWe hypothesized that sensed-behaviors, like phone use, unreliably predict depression in larger populations because behaviors representing high depression risk for some subgroups (e.g. older individuals) may represent lower risk for other subgroups (e.g. younger individuals). RH is relatively healthy and CSD is clinically-significant depression. Histograms show simulated data describing the count of individuals (y-axis) with specific daytime phone usage (x-axis). Colors indicate individuals experiencing CSD (orange) versus RH (light-blue). Plots are split by age subgroups. Black boxes show that increased phone usage is not a reliable predictor of depression because RH younger individuals have higher phone use than CSD older individuals. \u003cstrong\u003eb. \u003c/strong\u003eThe analysis pipeline. Behavioral data from smartphones and mental health outcomes collected during a U.S.-based NIMH-funded study\u003csup\u003e3,25–29\u003c/sup\u003e were used to train and validate AI models that predicted depression symptom risk from the behavioral data. We then measured algorithmic ranking bias in the developed tool to identify subgroups where the predicted CSD risk was incorrectly ranked lower than RH subgroups, and compared sensed-behaviors across subgroups where algorithms underperformed. \u003cstrong\u003ec. \u003c/strong\u003eSimilar to prior work\u003csup\u003e3,25\u003c/sup\u003e, 14 days of sensed-behavioral data were used to predict whether the PHQ-8 value across each weekly reported period indicated clinically-significant depression symptoms (PHQ-8 ≥10\u003csup\u003e33\u003c/sup\u003e).\u003c/p\u003e","description":"","filename":"Figure1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-3044613/v1/09ce58fa79ad4875277495a7.jpeg"},{"id":55063414,"identity":"4bc8ef81-c725-4189-9c52-cacdcfc9bb5d","added_by":"auto","created_at":"2024-04-22 03:07:37","extension":"jpeg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":389963,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eMeasuring algorithmic ranking bias. \u003c/strong\u003eWe considered three metrics from prior work to assess algorithmic ranking bias\u003csup\u003e20–22\u003c/sup\u003e. The predicted risk is the probability, output by the AI tool, that individuals were experiencing clinically-significant depression (CSD). Histograms show simulated example predictions from an AI tool, describing the count of individuals (y-axis) who fell into a predicted risk bin (x-axis). Colors indicate individuals experiencing CSD (orange) versus RH (light-blue). Plots are split by age subgroups (younger/older). The AUC is the area under the receiver operating curve. The red and dark-blue boxes, and corresponding text color below each plot, highlight the subgroups compared for each metric.\u003cstrong\u003e a. \u003c/strong\u003eThe high Subgroup AUCs show that the predicted risk for individuals experiencing CSD was greater than the predicted risk for relatively healthy (RH) individuals within both age subgroups. But, this AI tool was biased to predict higher risk for younger individuals, overall, than older individuals. This bias is quantified using the \u003cstrong\u003eb. \u003c/strong\u003eBackground-Negative-Subgroup-Positive (BNSP) AUC and \u003cstrong\u003ec. \u003c/strong\u003eBackground-Positive-Subgroup-Negative (BPSN) AUC, which respectively show that younger individuals with CSD (“positive samples”) were correctly ranked higher (high BNSP) than RH (“negative samples) samples from all other groups (older individuals, the “background”), but RH younger individuals were incorrectly ranked higher (low BPSN) than background samples with CSD. Older individuals show the complementary result (low BNSP, high BPSN). This bias reduces the model AUC when measured across the entire sample (assuming equal number of older and younger individuals, AUC=0.75), compared to the AUC in each subgroup (1.00).\u003c/p\u003e","description":"","filename":"Figure2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-3044613/v1/0b26b91ad4d01fc93d66bf74.jpeg"},{"id":55063419,"identity":"7e3f9397-6e15-423a-b886-86c1e5961b60","added_by":"auto","created_at":"2024-04-22 03:07:38","extension":"jpeg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":1211844,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eMeasuring bias in predicted depression risk. \u003c/strong\u003eBias was assessed by measuring the area under the receiver operating curve comparing positive (clinically-significant depression, CSD) and negative (relatively healthy, RH) samples within subgroups (Subgroup AUC, left column), subgroup positive samples to negative samples from all other groups, called “the background” (background-negative-subgroup-positive, or BNSP AUC, middle column), and subgroup negative samples to background positive samples (background-positive-subgroup-negative, or BPSN AUC, right column)\u003csup\u003e20,22\u003c/sup\u003e. Point values indicate the median value across trials. Error bars show 95% confidence intervals (2.5 and 97.5 percentiles). Dotted lines and shaded areas show the distribution (median and 95% confidence intervals) of either the median (if \u0026gt;2 groups) or highest performing groups across trials.\u003c/p\u003e","description":"","filename":"Figure3.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-3044613/v1/d7cdc090fb3a26f83d87d4b9.jpeg"},{"id":55063415,"identity":"04143aa3-a8f1-4341-850b-c83761df6bb4","added_by":"auto","created_at":"2024-04-22 03:07:38","extension":"jpeg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":546268,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eIsolating subgroups where models underperformed. \u003c/strong\u003eFor an ideal classifier, the predicted risk would be low for relatively healthy (RH) individuals, and high for individuals with clinically-significant depression (CSD). We thus modeled expected differences from the subgroups with either the lowest (for RH, left) or highest (for CSD, right) average predicted risk across trials. Group effects were calculated using generalized estimating equations (GEE)\u003csup\u003e38\u003c/sup\u003e, a type of linear model, to analyze the average effect of group membership on the predicted risk, controlling across all attributes. GEE accounted for the non-independence of repeated samples across trials\u003csup\u003e38\u003c/sup\u003e. Separate regression models were created for each outcome (RH, CSD) to remove the effects of the group base rate. Points represent the GEE coefficient (expected effect), and error bars are 95% confidence intervals around the estimated effect. Dotted vertical lines highlight an expected group effect of 0.\u003c/p\u003e","description":"","filename":"Figure4.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-3044613/v1/9f02377cda1a8f1e78775942.jpeg"},{"id":55063417,"identity":"a2731410-ef2c-487f-a690-2112da7210c7","added_by":"auto","created_at":"2024-04-22 03:07:38","extension":"jpeg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":753875,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eInterpreting the relationships between sensed-behaviors and depression. a. \u003c/strong\u003eShapley additive explanations (SHAP)\u003csup\u003e39\u003c/sup\u003e were used to interpret how the AI tool predicted depression risk using sensed-behaviors. Sensed-behaviors are ordered, descending, on the y-axis by their average impact on the predicted risk (the “SHAP value”, x-axis). Only the top 10 sensed-behaviors with the highest average impact are listed, for space. Colors dictate whether a higher sensed-behavior “feature” value (red) is associated with higher or lower predicted risk. For example, higher average (“Avg”) phone unlocks from 6-12PM was generally associated with lower predicted risk. Averages and deviations summarize sensed-behaviors over 14 days (see Figure 1c). \u003cstrong\u003eb. \u003c/strong\u003eExample coefficients (β, 95% CI, standardized units) from explanatory logistic regression models estimating the associations between sensed-behaviors and depression across subgroups, as well as the median and 95% CI of the sensed-behavior distribution. Full coefficients and statistics can be found in the supplementary materials.\u003c/p\u003e","description":"","filename":"Figure5.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-3044613/v1/9528771d9df036ac203c3a18.jpeg"},{"id":55246711,"identity":"7ca4e3d2-d000-4cb1-bdfb-5bd6ef524a4b","added_by":"auto","created_at":"2024-04-24 16:32:02","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1232231,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3044613/v1/b5ee5310-a2d0-45cf-85f6-ab899b4d853e.pdf"},{"id":55063409,"identity":"c7bca5f5-82ef-4c8c-98b4-03adcbfbd25d","added_by":"auto","created_at":"2024-04-22 03:07:37","extension":"xlsx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":26262,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryWorkbookonSensedBehaviors.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-3044613/v1/a07804adab58efed69590173.xlsx"},{"id":55063772,"identity":"30708fe9-345d-4c17-a640-375759bba12e","added_by":"auto","created_at":"2024-04-22 03:15:37","extension":"docx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":500719,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryInformation.docx","url":"https://assets-eu.researchsquare.com/files/rs-3044613/v1/12b9b17fdd4029b8fbeffb8e.docx"},{"id":55063412,"identity":"d0662e45-aedb-4f22-82bd-1e5b73361996","added_by":"auto","created_at":"2024-04-22 03:07:37","extension":"pdf","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":882620,"visible":true,"origin":"","legend":"","description":"","filename":"nreditorialpolicychecklist20230823.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3044613/v1/e861e5c8003f73a80778e2b5.pdf"},{"id":55063773,"identity":"3399a5e1-e208-4496-a395-eebd6aa2d149","added_by":"auto","created_at":"2024-04-22 03:15:38","extension":"pdf","order_by":4,"title":"","display":"","copyAsset":false,"role":"supplement","size":686994,"visible":true,"origin":"","legend":"","description":"","filename":"nrreportingsummary20230823.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3044613/v1/5b85a55ffbf82a3faf672562.pdf"}],"financialInterests":"Competing interest reported. D.A. and T.C. have submitted patent applications related to this work. T.C. is a co-founder and equity holder of HealthRhythms, Inc. and has received grants from Click Therapeutics related to digital therapeutics. D.C.M has accepted honoraria and consulting fees from Boehringer-Ingelheim, Otsuka Pharmaceuticals, Optum Behavioral Health, Centerstone Research Institute, and the One Mind Foundation, royalties from Oxford Press, and has an ownership interest in Adaptive Health, Inc. J.M. has accepted consulting fees from Boehringer Ingelheim. G.J.A. holds equity in HealthRhythms, Inc. and Lyra Health, Inc., and has accepted consulting fees and honoraria from BetterUp and Quantum Health.","formattedTitle":"Measuring algorithmic bias to analyze the reliability of AI tools that predict depression risk using smartphone sensed-behavioral data","fulltext":[{"header":"Introduction","content":"\u003cp\u003eMental healthcare systems are simultaneously facing a shortage of mental health specialty care providers and a large number of patients whose treatment needs remain unmet\u003csup\u003e1,2\u003c/sup\u003e. This service gap is driving research into AI-driven mental health monitoring tools, where sensed-behavioral data, defined as inferred behavioral data gathered by sensors and software embedded in everyday devices (e.g. smartphones, wearables), are repurposed to remotely monitor depression symptoms\u003csup\u003e3\u0026ndash;7\u003c/sup\u003e. Sensed-behavioral data has also been referred to as personal, behavioral, or passive sensing data in other work\u003csup\u003e7\u003c/sup\u003e. AI tools that leverage sensed-behavioral data intend to near-continuously identify individuals experiencing elevated symptoms in-between clinical encounters and consequently deliver preventive care\u003csup\u003e8\u003c/sup\u003e. These tools can also be integrated into digital therapeutics to automate precision interventions\u003csup\u003e9\u003c/sup\u003e. Initial work showed that depression risk could be predicted from sensed-behavioral data at a similar accuracy to general practitioners\u003csup\u003e10\u003c/sup\u003e in small populations\u003csup\u003e5,11\u003c/sup\u003e. More recent work shows that these AI tools predict depression risk at an accuracy only slightly better than a coin flip in larger, more diverse samples\u003csup\u003e4,6,12,13\u003c/sup\u003e. This prior work has not specifically explored why accuracy is reduced in larger samples, and it is not clear how to improve AI tools for clinical use.\u003c/p\u003e \u003cp\u003eIn this work, we hypothesized that accuracy is reduced in larger, more diverse populations because sensed-behaviors are unreliable predictors of depression risk: sensed-behaviors that predict depression are inconsistent across demographic and socioeconomic (SES) subgroups\u003csup\u003e14\u003c/sup\u003e. We intentionally use the term reliability due to its importance in both a psychometric and AI context. In a psychometric context, reliability refers to the consistency of a tool, typically a symptom assessment, across different contexts (e.g. raters, time)\u003csup\u003e14,15\u003c/sup\u003e. In AI, reliability is related to generalizability, if an AI tool is consistently accurate in different contexts (e.g. different populations, over time, etc.)\u003csup\u003e12\u003c/sup\u003e. Given these definitions, researchers in AI fairness have argued that aspects of psychometric reliability are important in an AI context: similar inputs (e.g. sensed-behaviors) to an AI model should yield similar outputs (e.g. estimated depression risk)\u003csup\u003e16\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eIn this paper, we adapt these ideas to study a specific aspect of reliability important for mental health AI tools deployed in large populations, i.e. if similar sensed-behaviors are consistently related to depression risk across different groups of individuals. We hypothesize that if the sensed-behaviors predictive of depression risk are inconsistent across groups, AI models that use sensed-behaviors to predict depression risk will be inaccurate because similar sensed-behavioral patterns will indicate different levels of depression risk for different subgroups. For example, imagine that mobility positively correlates with depression risk in subgroup A, and negatively correlates with depression risk in subgroup B. An AI model trained across subgroups using exclusively mobility data, blind to subgroup information as is typically the case in this literature\u003csup\u003e3\u0026ndash;5,17\u003c/sup\u003e, will receive unreliable information \u0026ndash; high mobility can simultaneously indicate both low and high depression risk \u0026ndash; and will make incorrect predictions for one of the subgroups. We note upfront that in this manuscript we do not consider temporal aspects of reliability, though we acknowledge that this is important in discussions of psychometric reliability, specifically if the AI tool is consistently accurate for the same individual, with predictions made under similar conditions\u003csup\u003e18\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eWe tested this hypothesis by identifying population subgroups where a depression risk prediction tool underperformed, and then analyzed sensed-behavioral differences across these subgroups. We identified subgroups where the tool underperformed by measuring \u003cem\u003ealgorithmic ranking bias\u003c/em\u003e (hereafter referred to as \u0026ldquo;bias\u0026rdquo;), where individuals experiencing depression from one subgroup (e.g. older individuals) were incorrectly ranked by the tool to be at lower risk than healthier individuals from other subgroups (e.g. younger individuals)\u003csup\u003e19\u0026ndash;22\u003c/sup\u003e. Reliability was analyzed by measuring ranking bias because if individuals in large populations have inconsistent relationships between sensed-behavior and mental health, behaviors that represent high depression risk for one subgroup may represent lower risk for another subgroup. For example, imagine an AI tool predicting that higher phone use increases depression risk. Studies\u003csup\u003e23,24\u003c/sup\u003e show that younger individuals have higher phone use than older individuals. Thus, the AI tool may incorrectly rank older individuals with depression to be at lower risk than healthier younger individuals, decreasing model accuracy (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ea).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eAgainst this backdrop, we developed an AI tool that estimated depression symptom risk using behavioral data collected from individuals\u0026rsquo; smartphones, using similar sensed-behaviors and outcome measures from recent work\u003csup\u003e3\u0026ndash;5,13,25\u003c/sup\u003e (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eb). The data used to develop and analyze the AI tool was collected during a U.S.-based National Institute of Mental Health (NIMH)-funded study\u003csup\u003e3,25\u0026ndash;29\u003c/sup\u003e, one of the largest, most geographically diverse studies of its kind. We then measured bias across attributes including age, sex at birth, race, household income, health insurance and employment to identify subgroups where the tool underperformed. We studied these specific attributes because of known behavioral differences across demographic and SES subgroups\u003csup\u003e23,24,30\u0026ndash;32\u003c/sup\u003e that could impact the reliability of the developed AI tool. Finally, we interpreted why the tool underperformed by identifying inconsistencies between the AI tool and sensed-behaviors predicting depression across subgroups. A summary of this analysis can be found in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eData Collection\u003c/h2\u003e \u003cp\u003eWe analyzed data from a U.S.-based, NIMH-funded study conducted from 2019\u0026ndash;2021 to identify associations between behavioral data collected from smartphones and depression symptoms\u003csup\u003e3,25\u0026ndash;29\u003c/sup\u003e. Smartphone sensed-behavioral data on GPS location, phone usage (timestamp of screen unlock), and sleep were near-continuously collected from participants across the United States for 16 weeks and the PHQ-8, a self-reported measure of two week depression symptoms\u003csup\u003e33,34\u003c/sup\u003e, frequently used in mental health research\u003csup\u003e3,5,25,27\u003c/sup\u003e, was collected every three weeks beginning on the first week of the study (e.g. on weeks 1, 4, 7, \u0026hellip;, known as weekly reporting periods). Sensed-behaviors were summarized over two weeks to align with collected PHQ-8 depression symptoms for prediction (see Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). For example, sensed-behaviors collected during weeks 3 and 4 were summarized to predict PHQ-8 values collected during week 4.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003e\u003cb\u003eSensed-behaviors.\u003c/b\u003e An overview of the sensed-behavioral data used in this analysis. The same set of sensed-behaviors were collected from all participants, and were summarized over two week periods to align with self-reported PHQ-8 symptoms. Please see the methods for more details.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCategory\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDerived Sensed-Behaviors\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLocation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eVariance (variability in GPS location), number of unique locations, entropy (variability in unique locations), normalized entropy (entropy normalized by number of unique locations), duration of time spent at home, percentage of collected samples in-transition (participant moving at \u0026gt;\u0026thinsp;1 km/hour), and circadian movement (24-hour regularity in movement). Location sensed-behaviors were directly calculated over two week periods. For example, we calculated the number of unique locations over two weeks.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePhone usage\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDuration of phone usage and number of screen unlocks each day and within four 6-hour periods (12-6AM, 6-12PM, 12-6PM, 6-12AM). The average and standard deviation of each phone usage sensed-behavior was calculated over two weeks, and the number of days with daily phone use and use within each 6-hour period was summed.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSleep\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAverage sleep onset (beginning of sleep), average duration, and variability in duration over two weeks.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e summarizes the data used for analysis. 3,900 samples were analyzed from 650 individuals, a large cohort and sample size compared to most studies to date analyzing associations between sensed-behaviors and mental health\u003csup\u003e4,5,25,35,36\u003c/sup\u003e. A sample was a set of sensed-behaviors, summarized over two weeks, with a corresponding PHQ-8 self-report. 46% of self-reported PHQ-8 values were \u0026ge;\u0026thinsp;10, indicating clinically-significant depression (CSD)\u003csup\u003e33\u003c/sup\u003e. The majority of participants were relatively young to middle aged (75% 25 to 54 years old), female (74%), white (82%), middle to high income (61% annual household income \u0026ge;\u003cspan\u003e$\u003c/span\u003e40,000), insured (93%) and employed (62%). We focused our results on groups with at least 15 participants\u003csup\u003e37\u003c/sup\u003e. The sensed-behavior distributions across the population for each subgroup can be found in the supplementary materials.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003e\u003cb\u003eStudy cohort\u003c/b\u003e. Data was collected within an NIMH-funded study to understand the relationships between digitally-collected behavioral data and depression symptoms\u003csup\u003e3,25\u0026ndash;29\u003c/sup\u003e. Participants contributed six total samples (summarized behavioral data and depression outcome measures) throughout the course of the study. A sample was a set of sensed-behaviors, summarized over two weeks, with a corresponding PHQ-8 self-report.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colspan=\"3\" morerows=\"3\" nameend=\"c3\" namest=\"c1\" rowspan=\"4\"\u003e \u003cp\u003eEntire Study\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e \u003cp\u003eNumber of participants\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003e650\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e \u003cp\u003e\u003cb\u003eSamples per participant\u003c/b\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e \u003cp\u003e\u003cb\u003eTotal number of samples\u003c/b\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003e3,900\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e \u003cp\u003e\u003cb\u003e% Clinically-significant depression\u003c/b\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003e46\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eAttribute\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eGroup\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003eNumber of participants (%)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003eAttribute\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003eGroup\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003eNumber of participants (%)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"6\" rowspan=\"7\"\u003e \u003cp\u003e\u003cb\u003eAge\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e18 to 25\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e60 (9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\" morerows=\"6\" rowspan=\"7\"\u003e \u003cp\u003e\u003cb\u003eHousehold Income\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;20,000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e98 (15)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e25 to 34\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e181 (28)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e20,000 to 39,999\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e144 (22)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e35 to 44\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e168 (26)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e40,000 to 59,999\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e124 (19)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e45 to 54\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e135 (21)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e60,000 to 99,999\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e161 (25)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e55 to 64\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e81 (12)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e100,000+\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e110 (17)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e65 to 74\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e22 (3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eDon't know\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e10 (2)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e75 to 84\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e3 (0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003ePrefer not to answer\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e3 (0)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u003cb\u003eSex at Birth\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eFemale\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e482 (74)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\" morerows=\"3\" rowspan=\"4\"\u003e \u003cp\u003e\u003cb\u003eHealth Insurance Status\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eInsured\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e603 (93)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMale\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e168 (26)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eUninsured\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e43 (7)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"5\" rowspan=\"6\"\u003e \u003cp\u003e\u003cb\u003eRace\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eWhite\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e534 (82)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eDon't know\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e3 (0)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eBlack/African American\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e61 (9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003ePrefer not to answer\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e1 (0)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAsian/Asian American\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e22 (3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\" morerows=\"5\" rowspan=\"6\"\u003e \u003cp\u003e\u003cb\u003eEmployment Status\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eEmployed\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e401 (62)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMore than one race\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e24 (4)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eUnemployed\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e90 (14)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eOther\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e6 (1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eDisability\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e72 (11)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrefer not to answer\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e3 (0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eRetired\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e34 (5)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eOther\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e52 (8)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003ePrefer not to answer\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e1 (0)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003eMeasuring Bias to Identify Subgroups where AI Models Underperform\u003c/h2\u003e \u003cp\u003eThe PHQ-8 asked participants to self-report depression symptoms experienced over 14 days, and PHQ-8\u0026rsquo;s were delivered multiple times throughout each weekly reporting period. We trained AI models using 14 days of smartphone sensed-behavioral data to predict if the average PHQ-8 value across each weekly reporting period (days 7 through 14, see Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ec) indicated clinically-significant depression (CSD, PHQ-8 score\u0026thinsp;\u0026ge;\u0026thinsp;10\u003csup\u003e33\u003c/sup\u003e) symptoms. Surveys were delivered multiple times each reporting week, and individual surveys may only reflect \u0026ldquo;briefly\u0026rdquo; elevated symptoms (e.g. work stress on the day the survey was administered). For this reason, PHQ-8 values were averaged over each reporting week to predict a more stable estimate of self-reported symptoms. In addition, by predicting average PHQ-8 values instead of individual self-reports, we could make use of all available data while ensuring that samples did not overlap temporally. Model performance was assessed by performing 5-fold cross-validation, partitioning on subjects, and predictions across folds were concatenated to calculate model performance. To analyze performance variability due to specific cross-validation splits, we performed 100 cross-validation trials, shuffling participants into different folds during each trial.\u003c/p\u003e \u003cp\u003eAI models output a predicted risk score from 0\u0026ndash;1 of experiencing CSD. We used the predicted risk to calculate common ranking bias metrics\u003csup\u003e20\u0026ndash;22\u003c/sup\u003e (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e) across the subgroups in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e. These metrics were based upon the area under the receiver operating curve (AUC), which measured the probability models correctly predicted that CSD samples were ranked higher (in the predicted risk) than RH samples. We first calculated the AUC within each subgroup (the \u0026ldquo;Subgroup AUC\u0026rdquo;). Note that equal Subgroup AUCs do not guarantee high AUC across an entire sample. For example, Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ea shows simulated data where an algorithm correctly predicted CSD risk within subgroups, but younger individuals, compared to older individuals, had higher predicted risk overall. Thus, healthy younger individuals were incorrectly predicted to be at higher risk than older individuals experiencing CSD. Two additional performance metrics assessed such errors. Specifically, the background-negative-subgroup-positive, or BNSP AUC (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eb) measured the probability that individuals experiencing CSD (the \u0026ldquo;positive\u0026rdquo; label) from a subgroup were correctly predicted to have higher risk than RH (the \u0026ldquo;negative label\u0026rdquo;) individuals from other subgroups (\u0026ldquo;the background\u0026rdquo;), and the background-positive-subgroup-negative, or BPSN AUC (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ec), measured the probability RH individuals from a subgroup were correctly predicted to have lower risk than background individuals experiencing CSD.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe highest performing AI model (a random forest, 100 trees, max depth of 10, balanced class weights, see methods) achieved a median (95% confidence interval, CI) AUC of 0.55 (0.54 to 0.57) across trials. Note that this low AUC was expected: it is comparable to the cross-validation performance of similar depression symptom prediction tools developed in larger, more diverse populations\u003csup\u003e4,6,13\u003c/sup\u003e, and motivates the objective of this work to study the reliability of these tools in larger populations.\u003c/p\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e shows the model results by each metric across subgroups. The Subgroup AUC was lower for males (median, 95% CI 0.52, 0.49 to 0.55), Black/African Americans (0.50, 0.46 to 0.54), individuals from low income households (\u0026lt;\u003cspan\u003e$\u003c/span\u003e20,000, 0.46, 0.43 to 0.50), uninsured (0.45, 0.41 to 0.51), and unemployed (0.46, 0.42 to 0.50) individuals, compared to the median subgroup AUC for each attribute (e.g. \u0026ldquo;Sex at Birth\u0026rdquo;) across trials. The BNSP AUC increased with age (from 0.50, 0.46 to 0.52 for 18 to 25 year olds, to 0.67, 0.62 to 0.73 for 65 to 74 year olds), but decreased with household income (from 0.60, 0.58 to 0.63 for individuals from \u0026lt;\u003cspan\u003e$\u003c/span\u003e20,000 income households, to 0.45, 0.42 to 0.48 for individuals from \u003cspan\u003e$\u003c/span\u003e100,000\u0026thinsp;+\u0026thinsp;income households). Individuals who were White (0.49, 0.46 to 0.52), male (0.52, 0.49 to 0.55), insured (0.47, 0.43 to 0.50), employed (0.43, 0.41 to 0.45), or identified with an \u0026ldquo;Other\u0026rdquo; type of employment (0.55, 0.52 to 0.59) also had lower BNSP AUC, compared to the median BNSP AUC for each attribute.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe BPSN AUC findings showed complementary trends: RH older individuals (e.g. 65 to 74, 0.46, 0.40 to 0.50), unemployed (0.38, 0.36 to 0.41), uninsured (0.47, 0.43 to 0.50), Black/African American (0.48, 0.45 to 0.50), females (0.52, 0.49 to 0.55), and individuals coming from lower income households (e.g. \u0026lt;\u003cspan\u003e$\u003c/span\u003e20,000 0.42, 0.39 to 0.44) had a lower BPSN AUC. Results were reasonably consistent across different types of models, within subgroup base rates (% samples with PHQ-8\u0026thinsp;\u0026ge;\u0026thinsp;10) were sometimes, but not always, associated with the BNSP/BPSN AUC, and subgroup sample size did not appear to be associated with the Subgroup AUC (see supplementary materials).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003eIsolating the Effects of Subgroup Membership on Model Underperformance\u003c/h2\u003e \u003cp\u003eWe wished to account for intersectional identities (e.g. female and employed) and isolate the effect of subgroup membership on model underperformance. For an ideal classifier, the predicted risk would be low for RH subgroups, and high for CSD subgroups. In addition, we would expect groups with higher base rates (% of samples with PHQ-8\u0026thinsp;\u0026ge;\u0026thinsp;10) to have a higher average predicted risk. We thus modeled expected differences from groups with either the lowest (for RH) or highest (for CSD) average risk across trials. Generalized estimating equations (GEE, exchangeable correlation structure)\u003csup\u003e38\u003c/sup\u003e, a type of linear regression, was used to estimate the average effect of subgroup membership on the predicted risk after controlling across all other attributes. GEE was used instead of linear regression to correct for the non-independence of repeated samples across trials\u003csup\u003e38\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eThe regression results can be found in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e. The individuals with the lowest average predicted risk who were not experiencing depression were 18 to 25 years old, male, White, had a household income of \u003cspan\u003e$\u003c/span\u003e100,000+, were insured, and employed. The predicted risk was expected to be higher (95% CI lower-bound\u0026thinsp;\u0026gt;\u0026thinsp;0) for RH individuals who were older than 34 (e.g. for 65 to 74 year olds, mean, 95% confidence interval 0.02, 0.01 to 0.04), identified as Asian/Asian American (0.02, 0.01 to 0.03), Black/African American (0.01, 0.00 to 0.01), came from \u0026lt;\u003cspan\u003e$\u003c/span\u003e60,000 income households (e.g. for \u0026lt;\u003cspan\u003e$\u003c/span\u003e20,000, 0.02, 0.01 to 0.03), were unemployed (0.03, 0.03 to 0.04), and/or on disability (0.01, 0.00 to 0.02). For individuals who were experiencing CSD, models predicted the highest average risk for 65 to 74 year olds, Females, Asian/Asian Americans, individuals who came from households with incomes of \u003cspan\u003e$\u003c/span\u003e20,000 to \u003cspan\u003e$\u003c/span\u003e39,999, were insured, and/or retired. The predicted risk for individuals experiencing CSD was expected to be lower (95% CI upper-bound\u0026thinsp;\u0026lt;\u0026thinsp;0) if individuals were 18 to 25 (\u0026ndash;0.02, \u0026minus;\u0026thinsp;0.04 to \u0026minus;\u0026thinsp;0.01), male (\u0026ndash;0.01, \u0026minus;\u0026thinsp;0.02 to \u0026minus;\u0026thinsp;0.00), Black/African American (\u0026ndash;0.02, \u0026minus;\u0026thinsp;0.03 to \u0026minus;\u0026thinsp;0.00), more than one race (\u0026ndash;0.02, \u0026minus;\u0026thinsp;0.03 to \u0026minus;\u0026thinsp;0.00), White (\u0026ndash;0.02, \u0026minus;\u0026thinsp;0.03 to \u0026minus;\u0026thinsp;0.01), came from any household with an annual income \u0026lt;\u003cspan\u003e$\u003c/span\u003e20,000 or \u0026ge;\u003cspan\u003e$\u003c/span\u003e40,000 (e.g. \u003cspan\u003e$\u003c/span\u003e100,000+ \u0026minus;\u0026thinsp;0.03, \u0026minus;\u0026thinsp;0.03 to \u0026minus;\u0026thinsp;0.02), and/or were employed (\u0026ndash;0.02, \u0026minus;\u0026thinsp;0.03 to \u0026minus;\u0026thinsp;0.01). Predicted risk distributions often overlapped across subgroups with higher or lower risk, though there were general trends across subgroups (e.g. the predicted risk increased with age and unemployment in RH individuals, and risk decreased with income level for both CSD and RH individuals, see Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e for more details).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003eInterpreting Sensed-Behaviors across Subgroups where Models Underperformed\u003c/h2\u003e \u003cp\u003eWe hypothesized that models underperformed because sensed-behaviors predictive of CSD were inconsistent across subgroups. We thus conducted an analysis to understand differences between how AI tools predicted CSD risk and the different relationships between sensed-behaviors and CSD across subgroups. First, we retrained the AI model on the entire data, and used Shapley additive explanations (SHAP)\u003csup\u003e39\u003c/sup\u003e to interpret how the AI tool predicted CSD risk from sensed-behaviors. We then compared SHAP values with coefficients from explanatory logistic regression models estimating how subgroup membership affected the relationship between each sensed-behavior and depression.\u003c/p\u003e \u003cp\u003eWe found different relationships between the SHAP values (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003ea) and sensed-behaviors associated with CSD across subgroups (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eb, comparisons across each attribute and feature can be found in the supplementary materials). For example, the AI tool predicted that higher morning phone usage (6AM \u0026minus;\u0026thinsp;12PM) was generally associated with lower predicted depression risk. Higher morning phone usage decreased depression risk for 18 to 25 year olds (mean, 95% CI effect on depression, standardized units: \u0026minus;\u0026thinsp;0.77, \u0026minus;\u0026thinsp;1.07 to \u0026minus;\u0026thinsp;0.47), but increased risk for 65 to 74 year olds (0.60, 0.07 to 1.12). Younger individuals, overall, also had higher morning phone use (standardized median, 95% CI 18 to 25 year olds: 0.32, \u0026minus;\u0026thinsp;2.27 to 1.60) compared to older individuals (65 to 74 year olds: \u0026minus;\u0026thinsp;0.62, \u0026minus;\u0026thinsp;1.96 to 0.76).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003ea also shows that specific mobility features, including the circadian movement (regularity in 24-hour movement), location entropy (regularity in travel to unique locations), and the percentage of collected GPS samples in transition (approximated speed\u0026thinsp;\u0026gt;\u0026thinsp;1 km/hour) were often associated with lower predicted CSD risk. Circadian movement decreased CSD risk for employed individuals (\u0026ndash;0.16, \u0026minus;\u0026thinsp;0.24 to \u0026minus;\u0026thinsp;0.07), but increased CSD risk for individuals who were on disability (0.44, 0.21 to 0.66). Circadian movement and location entropy also decreased depression risk for individuals from middle income (\u003cspan\u003e$\u003c/span\u003e60,000 to \u003cspan\u003e$\u003c/span\u003e99,999) households (circadian movement: \u0026minus;\u0026thinsp;0.21, \u0026minus;\u0026thinsp;0.35 and \u0026minus;\u0026thinsp;0.07; location entropy: \u0026minus;\u0026thinsp;0.34, \u0026minus;\u0026thinsp;0.48 to \u0026minus;\u0026thinsp;0.20), but increased risk for individuals from low income (\u0026lt;\u003cspan\u003e$\u003c/span\u003e20,000) households (circadian movement: 0.30, 0.09 to 0.51; location entropy: 0.35, 0.14 to 0.57). Finally, a higher percentage of GPS samples in transition decreased depression risk for insured individuals (\u0026ndash;0.15, \u0026minus;\u0026thinsp;0.22 to \u0026minus;\u0026thinsp;0.08), but increased risk for uninsured individuals (0.32, 0.11 to 0.52).\u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eIn this study, we hypothesized that sensed-behaviors are unreliable measures of depression in larger populations, reducing the accuracy of AI tools that use sensed-behaviors to predict depression risk. To test this hypothesis, we developed an AI tool that predicted clinically-significant depression (CSD) from sensed-behaviors and measured algorithmic bias to identify specific age, race, sex at birth, and socioeconomic subgroups where the tool underperformed. We then found differences between SHAP values estimating how the AI tool predicted CSD from sensed-behaviors, and explanatory logistic regression models estimating the associations between sensed-behaviors and CSD across subgroups. In this discussion, we show how differences in sensed-behaviors across subgroups may explain the identified bias and AI underperformance in larger, more diverse populations.\u003c/p\u003e \u003cp\u003eMeasuring bias showed that models predicted older, female, Black/African American, low income, unemployed, and individuals on disability were at higher risk of experiencing CSD (high BNSP, low BPSN AUC), and younger, male, White, high income, insured, and employed individuals were at lower risk of experiencing CSD (high BPSN, low BNSP AUC), independent of outcomes. Comparing SHAP values to explanatory logistic regression coefficients suggests why AI models incorrectly predicted depression risk. For example, our findings show that younger individuals had higher daytime phone usage than older individuals. Models predicted that higher daytime phone usage was associated with lower CSD risk (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003ea), potentially explaining why younger individuals, overall, had lower predicted risk, and older adults had higher predicted risk (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). Differences could be attributed to younger individuals using phones for entertainment and social activities that support well-being, while older individuals may prefer to use their phones for necessary communication or information gathering\u003csup\u003e23\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eIn another example, the model predicted that mobility, measured through circadian movement, location entropy, and GPS samples in transition, was associated with lower CSD risk (see Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e). Prior work has identified a negative association between these same mobility features and CSD\u003csup\u003e5,25\u003c/sup\u003e, suggesting that mobility decreases depression risk. While we found the expected negative associations across majority, higher SES (\u003cspan\u003e$\u003c/span\u003e60,000 to 99,999 household income, insured, and employed) subgroups, we found the opposite, positive association across less-represented lower SES (\u0026lt;\u003cspan\u003e$\u003c/span\u003e20,000 household income, on disability, uninsured) subgroups, potentially explaining the reduced model performance (lower Subgroup AUC) in these groups. There are many possible explanations for the identified differences in behavior. First, underlying reasons to be mobile (e.g. navigating bureaucracy to receive government payments) may increase stress for individuals who are lower income and/or on disability\u003csup\u003e31\u003c/sup\u003e, increasing depression risk. Second, the analyzed data was collected during the early-to-mid stages of the COVID-19 pandemic, when mobility for low SES essential workers may indicate work travel and increased COVID-19 risk, contributing to stress\u003csup\u003e32\u003c/sup\u003e and depression. These findings suggest that sensed-behaviors approximating phone use and mobility used to predict depression in prior work\u003csup\u003e3\u0026ndash;6,25\u003c/sup\u003e do not reliably predict depression in larger populations because of subgroup differences.\u003c/p\u003e \u003cp\u003eWhile existing work developing similar AI tools has strived to achieve generalizability\u003csup\u003e4,40\u003c/sup\u003e, our findings question this goal. Instead, it may be more practical to improve reliability by developing models for specific, targeted populations\u003csup\u003e41,42\u003c/sup\u003e. In addition, it may be helpful to train AI models using both sensed-behaviors and demographic information. In prior work and this study, AI models were trained using exclusively sensed-behavioral data\u003csup\u003e3\u0026ndash;5,17\u003c/sup\u003e. However, prior work suggests that models may not be more predictive even with added demographic information\u003csup\u003e43\u003c/sup\u003e. This shows that additional methods are needed to clearly define subgroups, beyond demographics, with more homogenous relationships between sensed-behaviors and depression symptoms.\u003c/p\u003e \u003cp\u003eAnother method to improve reliability is to develop personalized models, trained on participants\u0026rsquo; data over time\u003csup\u003e6,44\u003c/sup\u003e. While personalization seems appealing, researchers should ensure that personalized predictions are meaningful. For example, we experimented with personalized models using a procedure suggested from prior work\u003csup\u003e44\u003c/sup\u003e. The model AUC improved (0.68) compared to the presented results (0.55), but we achieved a better AUC (\u0026gt;\u0026thinsp;0.80) by developing a naive model re-predicting participants\u0026rsquo; first self-reported PHQ-8 value for all future outcomes. Given at least one participant self-report is often needed for personalization, models should show greater accuracy than these naive benchmarks.\u003c/p\u003e \u003cp\u003eEven if accuracy improves, models can still be biased\u003csup\u003e19,37\u003c/sup\u003e, and it is important to consider the clinical and public health implications of using biased risk scores for depression screening. For example, more frequent exposure to stress\u003csup\u003e45\u003c/sup\u003e contributes to higher rates of depression in lower SES populations\u003csup\u003e46\u003c/sup\u003e, but overestimating depression risk for healthy low SES individuals allocates mental health resources away from other individuals who need care. Similar issues persist for underestimating depression risk. For example, models predicted lower risk for males experiencing depression compared to healthier females (see Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). Males are less likely to seek treatment for their mental health than females\u003csup\u003e47\u003c/sup\u003e, and AI tools underestimating male depression risk may further reduce the likelihood that males seek care. Uncovering these biases are important before algorithmic tools are used in clinical settings.\u003c/p\u003e \u003cp\u003eTo reduce these harms, researchers can use methods described in this and other work\u003csup\u003e37\u003c/sup\u003e to identify subgroups where AI tools underperform by measuring bias. Resources could then be directed to develop new or retrain existing models for these subgroups. Simultaneously, clinical personnel using these tools can be trained to identify algorithmic bias and mitigate its effects\u003csup\u003e48\u003c/sup\u003e. In addition, depositing de-identified sensed-behavior and mental health outcomes data in research repositories could increase available data to analyze the reliability of AI tools\u003csup\u003e12\u003c/sup\u003e. Finally, our findings show the importance of developing AI tools using data from populations that have similar behavioral patterns to the populations where these tools will be deployed. More thorough reporting of model training data\u003csup\u003e49\u003c/sup\u003e, and monitoring AI tools in \u0026ldquo;silent mode\u0026rdquo;, in which predictions are made but not used for decision making\u003csup\u003e50\u003c/sup\u003e, could prevent AI tools developed in dissimilar populations from causing harm.\u003c/p\u003e \u003cp\u003eFinally, it is important to consider how the choice to classify depression symptom severity influenced our results, specifically choosing to predict binarized PHQ-8 values instead of raw PHQ-8 scores. Predicting binarized symptom scores is a fairly common practice in both the depression prediction literature\u003csup\u003e3\u0026ndash;5,17\u003c/sup\u003e, as well as in the clinical AI literature, broadly\u003csup\u003e51,52\u003c/sup\u003e. This practice is motivated by an interest to use AI tools for near-continuous symptom monitoring, in which an action (e.g. follow-up by a care provider) is triggered at a specific elevated symptom threshold. This motivation may be difficult to realize if the field continues to use depression symptom scales as outcomes; as recent work shows, these symptom scales do not produce categorical response distributions, with a clear decision boundary distinguishing individuals experiencing versus not experiencing symptoms. Instead, responses tend to exist along a continuum\u003csup\u003e14\u003c/sup\u003e. It is also important to consider if subgroup differences affect the interpretation and self-reporting of depression symptom scales. Despite this consideration, prior work provides evidence that the PHQ-8 exhibits measurement invariance across demographic and socioeconomic groups\u003csup\u003e53,54\u003c/sup\u003e. Thus, it may be unlikely that the bias identified in this work was due to group differences in self-reporting symptoms, but our findings could be partially attributed to the mistreatment of depression symptom scales as categorical in nature.\u003c/p\u003e \u003cp\u003eThis work had limitations. First, we analyzed data from a single study, though the studied cohort was larger in size, geographic representation, and timespan compared to prior work. In addition, the study cohort was majority White, employed, and female, though we did not find that sample size was associated with classification accuracy. Only inter-individual variability was considered, not intra-individual variability, and thus these findings do not extend to longitudinal monitoring contexts, where changes in sensed-behaviors may indicate changes in depression risk. In addition, data was only analyzed from participants who provided complete outcomes data (participants who reported at least one PHQ-8 value during each of the 6 weekly reporting periods). Data was exclusively collected from individuals who owned Android devices, only specific data types (GPS and phone usage) were analyzed. Only smartphone sensed-behaviors were analyzed, and data collected from other types of devices (e.g. wearables) were not analyzed. Finally, data collection took place from 2019 to 2021, when COVID-19 restrictions varied across the United States, which may influence our findings. Future work can examine if these results replicate over larger, more diverse cohorts, in both demographic and socioeconomic attributes, as well as the devices used for data collection. In addition, future work can explore if sensed-behaviors are reliable predictors of depression in longitudinal monitoring contexts, though recent work suggests that sensed-behaviors have low predictive power, even when used for longitudinal monitoring\u003csup\u003e25\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eIn conclusion, we present one method to assess the reliability of AI tools that use sensed-behaviors to predict depression risk. Specifically we measured ranking bias in a developed AI tool to identify subgroups where the tool underperformed, and then we interpreted why models underperformed by comparing the AI tool to sensed-behaviors predictive of depression across subgroups. Researchers and practitioners developing AI-driven mental health monitoring tools using behavioral data should think critically about whether these tools are likely to generalize, and consider developing tailored solutions that are well-validated in specific, targeted populations.\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003eCohort\u003c/h2\u003e \u003cp\u003eIn this work, we performed a secondary analysis of data collected during a U.S.-based National Institute of Mental Health (NIMH) funded study. The motivation for this study was to identify smartphone sensed-behavioral patterns predictive of major depressive disorder (MDD)\u003csup\u003e3,25\u0026ndash;29\u003c/sup\u003e. Participants were recruited from across the United States using digital registries and online advertisements, intentionally oversampling for individuals experiencing depression. Eligible participants lived in the United States, could read/write English, and owned an Android smartphone and data plan. In addition, eligible participants with at least moderate depression symptom severity based upon the Patient Health Questionnaire-8 (PHQ-8)\u0026thinsp;\u0026ge;\u0026thinsp;10 were oversampled to create a sample with elevated symptoms. Individuals were excluded from the study if they self-reported a diagnosis of bipolar disorder, any psychotic disorder, shared a smartphone with another individual, or were unwilling to share data. Eligible participants were asked to provide electronic informed consent after receiving a complete description of the study. Eligible participants had the option to not provide consent, and could withdraw from the study at any point.\u003c/p\u003e \u003cp\u003eConsented participants downloaded a study smartphone application\u003csup\u003e55\u003c/sup\u003e and completed a baseline assessment to self-report demographic and lifestyle information. The study application passively collected GPS location, sampled every 5 minutes, and smartphone interactions (screen unlock and time of unlock) for 16 weeks. Individuals completed depression symptom assessments every 3 weeks within the smartphone application (the PHQ-8\u003csup\u003e33,34\u003c/sup\u003e). Data collection took place from 2019\u0026ndash;2021, and all study procedures were approved by the Northwestern University Institutional Review Board (study #STU00205316).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003eSensed-Behavioral Features\u003c/h2\u003e \u003cp\u003eWe calculated sensed-behavioral features from the collected smartphone data to predict depression risk. Following established methods from prior work\u003csup\u003e3,5,25\u003c/sup\u003e, we calculated GPS mobility features including the location variance (variability in GPS), number of unique locations, location entropy (variability in unique locations), normalized location entropy (entropy normalized by number of unique locations), duration of time spent at home, percentage of collected samples in-transition (participant moving at \u0026gt;\u0026thinsp;1 km/hour), and circadian movement (24-hour regularity in movement)\u003csup\u003e5\u003c/sup\u003e. We also calculated phone usage features from the screen unlock data\u003csup\u003e40\u003c/sup\u003e, including the duration of phone use and the number of screen unlocks each day and within four 6-hour periods (12-6AM, 6-12PM, 12-6PM, 6-12AM). Finally, we used a standard algorithm\u003csup\u003e40,56\u003c/sup\u003e to approximate daily sleep onset and duration from screen unlock data.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eClassifying Depression Symptoms\u003c/h2\u003e \u003cp\u003eThe PHQ-8 asks participants to self-report depression symptoms that occurred during the past two weeks. Symptoms are reported from 0 (not experiencing the symptom) to 3 (frequently experiencing the symptom). Scores are summed and thresholded to classify severity, where summed scores of 10 or greater indicate a higher likelihood of experiencing a clinically-significant depression\u003csup\u003e33\u003c/sup\u003e. We thus followed prior work\u003csup\u003e5,25\u003c/sup\u003e to calculate sensed-behavioral features in the two week period up to and including each weekly PHQ-8 reporting period. Behavioral features were input into machine learning models to predict clinically-significant symptoms (PHQ-8\u0026thinsp;\u0026ge;\u0026thinsp;10).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eData Preprocessing\u003c/h2\u003e \u003cp\u003eScreen unlock and sleep features were summarized to align with the PHQ-8\u003csup\u003e40\u003c/sup\u003e. The average and standard deviation of each daily and 6-hour epoch feature were calculated across the two week prediction period, and the number of days with daily phone use and use within each 6-hour epoch were summed. GPS features were directly calculated over the two weeks. As recommended by Saeb et al.\u003csup\u003e5\u003c/sup\u003e, skewed features were log-transformed. Missing data was filled using multivariate imputation\u003csup\u003e57\u003c/sup\u003e and then standardized (mean\u0026thinsp;=\u0026thinsp;0, standard deviation\u0026thinsp;=\u0026thinsp;1) based upon each training dataset prior to being input into predictive models.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eAI Model Training and Validation\u003c/h2\u003e \u003cp\u003eWe trained machine learning models commonly used to predict mental health status from smartphone behavioral data including regularized (L2-norm) logistic regression (LR)\u003csup\u003e3,5\u003c/sup\u003e, support vector machines (SVM)\u003csup\u003e4,58\u003c/sup\u003e, and tree-based ensemble models including random forest (RF) and gradient boosting trees (GBT)\u003csup\u003e3,40\u003c/sup\u003e. We varied the strength of the LR and SVM regularization parameter (0.01, 0.1, 1.0), used a radial basis function SVM kernel, varied class balancing weights in the RF and SVM (unbalanced/balanced), varied the number of ensemble tree estimators (10, 100), depth (3, 10, or until pure), and the GBT learning rate (0.01, 0.1, 1.0) and loss (deviance and exponential). Non-logistic prediction models were calibrated using Platt scaling to approximately match the predicted risk to the proportion of individuals experiencing clinically-significant symptoms at each risk level\u003csup\u003e59\u003c/sup\u003e. Logistic regression models, as shown in prior work\u003csup\u003e59\u003c/sup\u003e, output calibrated probabilities. Models were implemented using the scikit-learn Python library\u003csup\u003e60\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eMultiple PHQ-8 surveys were administered each weekly reporting period (e.g. week 1, 4, 7, etc.). Survey scores in each reporting week were averaged to remove overlap between sensor and outcomes data. Data was analyzed from study participants who self-reported at least one PHQ-8 during each reporting week, resulting in 6 predictions per participant. Data from all other participants were removed (288 participants removed, 31% of total) to focus this analysis towards algorithmic bias due to group differences rather than bias due to missing outcomes data.\u003c/p\u003e \u003c/div\u003e"},{"header":"Declarations","content":"\u003ch2\u003eCompeting Interests\u003c/h2\u003e \u003cp\u003eD.A. and T.C. have submitted patent applications related to this work. T.C. is a co-founder and equity holder of HealthRhythms, Inc. and has received grants from Click Therapeutics related to digital therapeutics. D.C.M has accepted honoraria and consulting fees from Boehringer-Ingelheim, Otsuka Pharmaceuticals, Optum Behavioral Health, Centerstone Research Institute, and the One Mind Foundation, royalties from Oxford Press, and has an ownership interest in Adaptive Health, Inc. J.M. has accepted consulting fees from Boehringer Ingelheim. G.J.A. holds equity in HealthRhythms, Inc. and Lyra Health, Inc., and has accepted consulting fees and honoraria from BetterUp and Quantum Health.\u003c/p\u003e\u003ch2\u003eAuthor Contributions\u003c/h2\u003e \u003cp\u003eD.A. conducted the analysis and wrote the draft manuscript. T.C., F.W., J.M., C.A.S., and D.C.M. provided supervisory support throughout the analysis. D.C.M. and J.M. were involved in data collection. All authors participated in drafting and revising the manuscript.\u003c/p\u003e\u003ch2\u003eAcknowledgements\u003c/h2\u003e \u003cp\u003eD.A. is supported by the National Science Foundation Graduate Research Fellowship under Grant No. DGE-2139899, and a Digital Life Initiative Doctoral Fellowship. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the authors and do not necessarily reflect the views of the funders. Data collection was supported by NIMH Grant No. R01MH111610 to D.C.M. J.M. is supported by K08MH128640. C.A.S. is supported by T32MH115882. Computing costs were funded by a Microsoft Azure Cloud Computing Grant through the Cornell Center for Data Science for Enterprise and Society, awarded to T.C. T.C. and F.W. were also supported by a multi-investigator seed grant awarded from the Cornell Academic Integration program.\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e \u003cp\u003eSensed-behavioral data cannot be made publicly available due to potentially identifying information (e.g. GPS location) that may compromise participant privacy. De-identified self-reported data (the PHQ-8) will be made available through the NIMH Data Archive.\u003c/p\u003e\u003ch2\u003eCode Availability\u003c/h2\u003e \u003cp\u003eA repository for all code used for analysis can be found at the following link: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/dadler6/reliability_depression_ml\u003c/span\u003e\u003cspan address=\"https://github.com/dadler6/reliability_depression_ml\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eCai, A. \u003cem\u003eet al.\u003c/em\u003e Trends In Mental Health Care Delivery By Psychiatrists And Nurse Practitioners In Medicare, 2011\u0026ndash;19. Health Aff. (Millwood) 41, 1222\u0026ndash;1230 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMohr, D. C. et al. Banbury Forum Consensus Statement on the Path Forward for Digital Mental Health Treatment. Psychiatr. Serv. appi.ps.202000561 (2021) doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1176/appi.ps.202000561\u003c/span\u003e\u003cspan address=\"10.1176/appi.ps.202000561\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu, T. \u003cem\u003eet al.\u003c/em\u003e The relationship between text message sentiment and self-reported depression. J. Affect. Disord. 302, 7\u0026ndash;14 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXu, X. \u003cem\u003eet al.\u003c/em\u003e GLOBEM: Cross-Dataset Generalization of Longitudinal Human Behavior Modeling. \u003cem\u003eProc. ACM Interact. Mob. Wearable Ubiquitous Technol.\u003c/em\u003e 6, 190:1-190:34 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSaeb, S. \u003cem\u003eet al.\u003c/em\u003e Mobile Phone Sensor Correlates of Depressive Symptom Severity in Daily-Life Behavior: An Exploratory Study. J. Med. Internet Res. 17, (2015).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMeegahapola, L. \u003cem\u003eet al.\u003c/em\u003e Generalization and Personalization of Mobile Sensing-Based Mood Inference Models: An Analysis of College Students in Eight Countries. \u003cem\u003eProc. ACM Interact. Mob. Wearable Ubiquitous Technol.\u003c/em\u003e 6, 176:1-176:32 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMohr, D. C., Shilton, K. \u0026amp; Hotopf, M. Digital phenotyping, behavioral sensing, or personal sensing: names and transparency in the digital age. Npj Digit. Med. 3, 1\u0026ndash;2 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLee, E. E. et al. Artificial Intelligence for Mental Health Care: Clinical Applications, Barriers, Facilitators, and Artificial Wisdom. Biol. Psychiatry Cogn. Neurosci. Neuroimaging S245190222100046X (2021) doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.bpsc.2021.02.001\u003c/span\u003e\u003cspan address=\"10.1016/j.bpsc.2021.02.001\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFrank, E. \u003cem\u003eet al.\u003c/em\u003e Personalized digital intervention for depression based on social rhythm principles adds significantly to outpatient treatment. Front. Digit. Health 4, (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMitchell, A. J., Vaze, A. \u0026amp; Rao, S. Clinical diagnosis of depression in primary care: a meta-analysis. The Lancet 374, 609\u0026ndash;619 (2009).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang, R. \u003cem\u003eet al.\u003c/em\u003e Tracking Depression Dynamics in College Students Using Mobile Phone and Wearable Sensing. \u003cem\u003eProc. ACM Interact. Mob. Wearable Ubiquitous Technol.\u003c/em\u003e 2, 43:1\u0026ndash;43:26 (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAdler, D. A. \u003cem\u003eet al.\u003c/em\u003e A call for open data to develop mental health digital biomarkers. BJPsych Open 8, (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eM\u0026uuml;ller, S. R., Chen, X. (Leslie), Peters, H., Chaintreau, A. \u0026amp; Matz, S. C. Depression predictions from GPS-based mobility do not generalize well to large demographically heterogeneous samples. Sci. Rep. 11, 14007 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFried, E. I., Flake, J. K. \u0026amp; Robinaugh, D. J. Revisiting the theoretical and methodological foundations of depression measurement. Nat. Rev. Psychol. 1\u0026ndash;11 (2022) doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s44159-022-00050-2\u003c/span\u003e\u003cspan address=\"10.1038/s44159-022-00050-2\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBeck, A. T. Reliability of psychiatric diagnoses: 1. a critique of systematic studies. Am. J. Psychiatry 119, 210\u0026ndash;216 (1962).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJacobs, A. Z. \u0026amp; Wallach, H. Measurement and Fairness. Proc. 2021 ACM Conf. Fairness Account. Transpar. 375\u0026ndash;385 (2021) doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1145/3442188.3445901\u003c/span\u003e\u003cspan address=\"10.1145/3442188.3445901\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJacobson, N. C., Weingarden, H. \u0026amp; Wilhelm, S. Digital biomarkers of mood disorders and symptom change. Npj Digit. Med. 2, 1\u0026ndash;3 (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBoateng, G. O., Neilands, T. B., Frongillo, E. A., Melgar-Qui\u0026ntilde;onez, H. R. \u0026amp; Young, S. L. Best Practices for Developing and Validating Scales for Health, Social, and Behavioral Research: A Primer. Front. Public Health 6, (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eObermeyer, Z., Powers, B., Vogeli, C. \u0026amp; Mullainathan, S. Dissecting racial bias in an algorithm used to manage the health of populations. Science 366, 447\u0026ndash;453 (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBorkan, D., Dixon, L., Sorensen, J., Thain, N. \u0026amp; Vasserman, L. Nuanced Metrics for Measuring Unintended Bias with Real Data for Text Classification. Preprint at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://arxiv.org/abs/1903.04561\u003c/span\u003e\u003cspan address=\"http://arxiv.org/abs/1903.04561\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKallus, N. \u0026amp; Zhou, A. The Fairness of Risk Scores Beyond Classification: Bipartite Ranking and the XAUC Metric. in Advances in Neural Information Processing Systems vol.\u0026nbsp;32 (Curran Associates, Inc., 2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVogel, R., Bellet, A. \u0026amp; Cl\u0026eacute;men\u0026ccedil;on, S. Learning Fair Scoring Functions: Bipartite Ranking under ROC-based Fairness Constraints. in Proceedings of The 24th International Conference on Artificial Intelligence and Statistics 784\u0026ndash;792 (PMLR, 2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAndone, I. et al. How age and gender affect smartphone usage. in Proceedings of the 2016 ACM International Joint Conference on Pervasive and Ubiquitous Computing: Adjunct 9\u0026ndash;12 (ACM, 2016). doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1145/2968219.2971451\u003c/span\u003e\u003cspan address=\"10.1145/2968219.2971451\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHorwood, S., Anglim, J. \u0026amp; Mallawaarachchi, S. R. Problematic smartphone use in a large nationally representative sample: Age, reporting biases, and technology concerns. Comput. Hum. Behav. 122, 106848 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMeyerhoff, J. \u003cem\u003eet al.\u003c/em\u003e Evaluation of Changes in Depression, Anxiety, and Social Anxiety Using Smartphone Sensor Features: Longitudinal Cohort Study. J. Med. Internet Res. 23, e22844 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMohr, D. C. LifeSense: Transforming Behavioral Assessment of Depression Using Personal Sensing Technology. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://reporter.nih.gov/search/N6YCr94ZvkOVUNu1i5HNaQ/project-details/9982127\u003c/span\u003e\u003cspan address=\"https://reporter.nih.gov/search/N6YCr94ZvkOVUNu1i5HNaQ/project-details/9982127\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2017).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eStamatis, C. A. \u003cem\u003eet al.\u003c/em\u003e Prospective associations of text-message-based sentiment with symptoms of depression, generalized anxiety, and social anxiety. Depress. Anxiety 39, 794\u0026ndash;804 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMeyerhoff, J. \u003cem\u003eet al.\u003c/em\u003e Analyzing text message linguistic features: Do people with depression communicate differently with their close and non-close contacts? Behav. Res. Ther. 166, 104342 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eStamatis, C. A. \u003cem\u003eet al.\u003c/em\u003e The association of language style matching in text messages with mood and anxiety symptoms. Procedia Comput. Sci. 206, 151\u0026ndash;161 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGreissl, S. \u003cem\u003eet al.\u003c/em\u003e Is unemployment associated with inefficient sleep habits? A cohort study using objective sleep measurements. J. Sleep Res. 31, e13516 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eIezzoni, L. I., McCarthy, E. P., Davis, R. B. \u0026amp; Siebens, H. Mobility Difficulties Are Not Only a Problem of Old Age. J. Gen. Intern. Med. 16, 235\u0026ndash;243 (2001).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLevy, B. L., Vachuska, K., Subramanian, S. V. \u0026amp; Sampson, R. J. Neighborhood socioeconomic inequality based on everyday mobility predicts COVID-19 infection in San Francisco, Seattle, and Wisconsin. Sci. Adv. 8, eabl3825 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKroenke, K. \u003cem\u003eet al.\u003c/em\u003e The PHQ-8 as a measure of current depression in the general population. J. Affect. Disord. 114, 163\u0026ndash;173 (2009).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWu, Y. \u003cem\u003eet al.\u003c/em\u003e Equivalency of the diagnostic accuracy of the PHQ-8 and PHQ-9: a systematic review and individual participant data meta-analysis. Psychol. Med. 50, 1368\u0026ndash;1380 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOpoku Asare, K. \u003cem\u003eet al.\u003c/em\u003e Predicting Depression From Smartphone Behavioral Markers Using Machine Learning Methods, Hyperparameter Optimization, and Feature Importance Analysis: Exploratory Study. JMIR MHealth UHealth 9, e26540 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCorponi, F. et al. Automated mood disorder symptoms monitoring from multivariate time-series sensory data: Getting the full picture beyond a single number. 2023.03.25.23287744 Preprint at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1101/2023.03.25.23287744\u003c/span\u003e\u003cspan address=\"10.1101/2023.03.25.23287744\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSeyyed-Kalantari, L., Zhang, H., McDermott, M. B. A., Chen, I. Y. \u0026amp; Ghassemi, M. Underdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under-served patient populations. Nat. Med. 27, 2176\u0026ndash;2182 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBallinger, G. A. Using Generalized Estimating Equations for Longitudinal Data Analysis. Organ. Res. Methods 7, 127\u0026ndash;150 (2004).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLundberg, S. M. \u0026amp; Lee, S.-I. A Unified Approach to Interpreting Model Predictions. in Advances in Neural Information Processing Systems vol. 30 (Curran Associates, Inc., 2017).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAdler, D. A., Wang, F., Mohr, D. C. \u0026amp; Choudhury, T. Machine learning for passive mental health symptom prediction: Generalization across different longitudinal mobile sensing studies. PLOS ONE 17, e0266516 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSperrin, M., Riley, R. D., Collins, G. S. \u0026amp; Martin, G. P. Targeted validation: validating clinical prediction models in their intended population and setting. Diagn. Progn. Res. 6, 24 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMitchell, M. et al. Model Cards for Model Reporting. ArXiv181003993 Cs (2019) doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1145/3287560.3287596\u003c/span\u003e\u003cspan address=\"10.1145/3287560.3287596\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePratap, A. \u003cem\u003eet al.\u003c/em\u003e The accuracy of passive phone sensors in predicting daily mood. Depress. Anxiety 36, 72\u0026ndash;81 (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang, R. et al. CrossCheck: toward passive sensing and detection of mental health changes in people with schizophrenia. in Proceedings of the 2016 ACM International Joint Conference on Pervasive and Ubiquitous Computing - UbiComp \u0026rsquo;16 886\u0026ndash;897 (ACM Press, 2016). doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1145/2971648.2971740\u003c/span\u003e\u003cspan address=\"10.1145/2971648.2971740\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWilliams, D. R., Mohammed, S. A., Leavell, J. \u0026amp; Collins, C. Race, Socioeconomic Status and Health: Complexities, Ongoing Challenges and Research Opportunities. Ann. N. Y. Acad. Sci. 1186, 69\u0026ndash;101 (2010).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEverson, S. A., Maty, S. C., Lynch, J. W. \u0026amp; Kaplan, G. A. Epidemiologic evidence for the relation between socioeconomic status and depression, obesity, and diabetes. J. Psychosom. Res. 53, 891\u0026ndash;895 (2002).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChatmon, B. N. Males and Mental Health Stigma. Am. J. Mens Health 14, 1557988320949322 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRajkomar, A., Hardt, M., Howell, M. D., Corrado, G. \u0026amp; Chin, M. H. Ensuring Fairness in Machine Learning to Advance Health Equity. Ann. Intern. Med. 169, 866\u0026ndash;872 (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGebru, T. et al. Datasheets for Datasets. ArXiv180309010 Cs (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWiens, J. \u003cem\u003eet al.\u003c/em\u003e Do no harm: a roadmap for responsible machine learning for health care. Nat. Med. 25, 1337\u0026ndash;1340 (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWong, A. et al. External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients. JAMA Intern. Med. (2021) doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1001/jamainternmed.2021.2626\u003c/span\u003e\u003cspan address=\"10.1001/jamainternmed.2021.2626\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAdams, R. et al. Prospective, multi-site study of patient outcomes after implementation of the TREWS machine learning-based early warning system for sepsis. Nat. Med. 1\u0026ndash;6 (2022) doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41591-022-01894-0\u003c/span\u003e\u003cspan address=\"10.1038/s41591-022-01894-0\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGalenkamp, H., Stronks, K., Snijder, M. B. \u0026amp; Derks, E. M. Measurement invariance testing of the PHQ-9 in a multi-ethnic population in Europe: the HELIUS study. BMC Psychiatry 17, 349 (2017).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVillarreal-Zegarra, D., Copez-Lonzoy, A., Bernab\u0026eacute;-Ortiz, A., Melendez-Torres, G. J. \u0026amp; Bazo-Alvarez, J. C. Valid group comparisons can be made with the Patient Health Questionnaire (PHQ-9): A measurement invariance study across groups by demographic characteristics. PLOS ONE 14, e0221717 (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAudacious Software. Passive Data Kit. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://passivedatakit.org/\u003c/span\u003e\u003cspan address=\"https://passivedatakit.org/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAbdullah, S., Matthews, M., Murnane, E. L., Gay, G. \u0026amp; Choudhury, T. Towards circadian computing: \u0026lsquo;early to bed and early to rise\u0026rsquo; makes some of us unhealthy and sleep deprived. in Proceedings of the 2014 ACM International Joint Conference on Pervasive and Ubiquitous Computing 673\u0026ndash;684 (ACM, 2014). doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1145/2632048.2632100\u003c/span\u003e\u003cspan address=\"10.1145/2632048.2632100\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBuuren, S. van \u0026amp; Groothuis-Oudshoorn, K. mice: Multivariate Imputation by Chained Equations in R. J. Stat. Softw. 45, 1\u0026ndash;67 (2011).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTseng, V. W.-S. \u003cem\u003eet al.\u003c/em\u003e Using behavioral rhythms and multi-task learning to predict fine-grained symptoms of schizophrenia. Sci. Rep. 10, 15100 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNiculescu-Mizil, A. \u0026amp; Caruana, R. Predicting good probabilities with supervised learning. in Proceedings of the 22nd international conference on Machine learning - ICML \u0026rsquo;05 625\u0026ndash;632 (ACM Press, 2005). doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1145/1102351.1102430\u003c/span\u003e\u003cspan address=\"10.1145/1102351.1102430\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePedregosa, F. et al. Scikit-learn: Machine Learning in Python. ArXiv12010490 Cs (2018).\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"npj-mental-health-research","isNatureJournal":false,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"npjmentalhealth","sideBox":"Learn more about [npj Mental Health Research](https://www.nature.com/npjmentalhealth/)","snPcode":"44184","submissionUrl":"https://mts-npjmentalhealth.nature.com/cgi-bin/main.p...","title":"npj Mental Health Research","twitterHandle":"@npjmentalhealth\n","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"npj","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-3044613/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-3044613/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eAI tools intend to transform mental healthcare by providing remote estimates of depression risk using behavioral data collected by sensors embedded in smartphones. While these tools accurately predict elevated symptoms in small, homogenous populations, recent studies show that these tools are less accurate in larger, more diverse populations. In this work, we show that accuracy is reduced because sensed-behaviors are unreliable predictors of depression across individuals; specifically the sensed-behaviors that predict depression risk are inconsistent across demographic and socioeconomic subgroups. We first identified subgroups where a developed AI tool underperformed by measuring algorithmic bias, where subgroups with depression were incorrectly predicted to be at lower risk than healthier subgroups. We then found inconsistencies between sensed-behaviors predictive of depression across these subgroups. Our findings suggest that researchers developing AI tools predicting mental health from behavior should think critically about the generalizability of these tools, and consider tailored solutions for targeted populations.\u003c/p\u003e","manuscriptTitle":"Measuring algorithmic bias to analyze the reliability of AI tools that predict depression risk using smartphone sensed-behavioral data","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-04-22 03:07:32","doi":"10.21203/rs.3.rs-3044613/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2023-12-23T16:32:13+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2023-12-21T21:52:48+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2023-10-30T14:33:48+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"9512b805-c599-472a-9b51-9115255bb5b3","date":"2023-09-25T17:22:26+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"174d67d3-9b75-429f-8e18-c9ce400727df","date":"2023-09-16T14:59:15+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2023-09-16T14:38:34+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2023-09-14T15:28:43+00:00","index":"","fulltext":""},{"type":"submitted","content":"npj Mental Health Research","date":"2023-08-16T23:12:18+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"npj-mental-health-research","isNatureJournal":false,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"npjmentalhealth","sideBox":"Learn more about [npj Mental Health Research](https://www.nature.com/npjmentalhealth/)","snPcode":"44184","submissionUrl":"https://mts-npjmentalhealth.nature.com/cgi-bin/main.p...","title":"npj Mental Health Research","twitterHandle":"@npjmentalhealth\n","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"npj","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"38d8b043-fbce-4250-8c26-5ca1c4ad165a","owner":[],"postedDate":"April 22nd, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":24685628,"name":"Health sciences/Medical research/Biomarkers/Predictive markers"},{"id":24685629,"name":"Biological sciences/Computational biology and bioinformatics/Machine learning"},{"id":24685630,"name":"Health sciences/Diseases/Psychiatric disorders/Depression"}],"tags":[],"updatedAt":"2024-04-24T16:31:57+00:00","versionOfRecord":{"articleIdentity":"rs-3044613","link":"https://doi.org/10.1038/s44184-024-00057-y","journal":{"identity":"npj-mental-health-research","isVorOnly":false,"title":"npj Mental Health Research"},"publishedOn":"2024-04-22 00:00:00","publishedOnDateReadable":"April 22nd, 2024"},"versionCreatedAt":"2024-04-22 03:07:32","video":"","vorDoi":"10.1038/s44184-024-00057-y","vorDoiUrl":"https://doi.org/10.1038/s44184-024-00057-y","workflowStages":[]},"version":"v1","identity":"rs-3044613","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-3044613","identity":"rs-3044613","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00