Learning from the Margins: Using Machine Learning Models to Predict Epilepsy Outcomes from Socioeconomic Factors

preprint OA: closed CC-BY-4.0
AI-generated summary by claude@2026-07, 2026-07-16

This study utilized machine learning models trained on demographic and socioeconomic factors to predict seizure-free status in adults with epilepsy, finding that social determinants independently predicted outcomes and highlighting social gradients in epilepsy care.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-16 · read from full text

This retrospective observational preprint studied 1,609 adults with epilepsy attending a specialist epilepsy service in England (2020–2024), using routinely collected demographic and area-level socioeconomic data from the Index of Multiple Deprivation alongside limited high-level clinical descriptors (without medication-specific information). The authors trained logistic regression, random forest, and gradient boosting models to predict whether patients were seizure-free at the most recent clinical review, applying SMOTE within training folds to address class imbalance and evaluating performance via stratified cross-validation and multiple classification metrics. They found that demographic and socioeconomic variables had consistent associations with seizure outcomes and substantially improved model performance, with random forest and gradient boosting achieving accuracy, precision, and recall above 83%, and that individuals in more socioeconomically deprived areas were overrepresented among those not seizure-free; they report comparable performance when models used predominantly non-clinical features, indicating independent prognostic value of social determinants. The main limitation explicitly noted is that the work is a preprint and not peer reviewed, and clinical characterization is constrained by the absence of medication-specific data. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Abstract Background Epilepsy is one of the most common serious neurological conditions globally and is associated with substantial morbidity, reduced quality of life, and increased socioeconomic disadvantage. Although clinical predictors of seizure outcomes have been widely studied, demographic and socioeconomic determinants of treatment response remain underrepresented in predictive modelling. Greater understanding of how these factors influence outcomes may inform earlier identification of patients at risk of poor seizure control and contribute to strategies aimed at reducing health inequalities. Methods A retrospective observational study was undertaken using routinely collected clinical, demographic, and socioeconomic data from 1,609 adults with epilepsy attending a specialist epilepsy service in England between 2020 and 2024. Demographic variables included age, sex, and ethnicity. Socioeconomic position was captured using area-level measures from the Index of Multiple Deprivation, including income, employment, education, health, housing, and living environment domains. Limited high-level clinical descriptors were included, without medication-specific information. Logistic Regression, Random Forest, and Gradient Boosting models were used to predict seizure-free status at the most recent clinical review. To address class imbalance, the Synthetic Minority Oversampling Technique (SMOTE) was applied within training folds. Model performance was assessed using stratified cross-validation and evaluated using accuracy, precision, recall, and precision–recall area under the curve. Results Demographic and socioeconomic variables showed consistent associations with epilepsy treatment outcomes and contributed substantially to model performance. After accounting for class imbalance, Random Forest and Gradient Boosting models achieved accuracy, precision, and recall exceeding 83%. Individuals residing in more socioeconomically deprived areas were overrepresented among those who were not seizure-free, indicating persistent social gradients in epilepsy outcomes. Comparable predictive performance was observed when models were trained using predominantly non-clinical features, suggesting that social determinants carry independent prognostic value. Conclusions Demographic and socioeconomic factors are important predictors of epilepsy treatment outcomes and can be leveraged within machine learning frameworks to support population-level risk stratification. Incorporating social determinants of health into predictive approaches may help identify patients at increased risk of poor outcomes and inform more equitable, context-aware epilepsy care pathways.
Full text 141,954 characters · extracted from preprint-html · click to expand
Learning from the Margins: Using Machine Learning Models to Predict Epilepsy Outcomes from Socioeconomic Factors | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Learning from the Margins: Using Machine Learning Models to Predict Epilepsy Outcomes from Socioeconomic Factors Md Shadab Mashuk, Lana Lai, Yang Lu, Shumit Saha, Natasha Carmichael, and 4 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8917468/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 5 You are reading this latest preprint version Abstract Background Epilepsy is one of the most common serious neurological conditions globally and is associated with substantial morbidity, reduced quality of life, and increased socioeconomic disadvantage. Although clinical predictors of seizure outcomes have been widely studied, demographic and socioeconomic determinants of treatment response remain underrepresented in predictive modelling. Greater understanding of how these factors influence outcomes may inform earlier identification of patients at risk of poor seizure control and contribute to strategies aimed at reducing health inequalities. Methods A retrospective observational study was undertaken using routinely collected clinical, demographic, and socioeconomic data from 1,609 adults with epilepsy attending a specialist epilepsy service in England between 2020 and 2024. Demographic variables included age, sex, and ethnicity. Socioeconomic position was captured using area-level measures from the Index of Multiple Deprivation, including income, employment, education, health, housing, and living environment domains. Limited high-level clinical descriptors were included, without medication-specific information. Logistic Regression, Random Forest, and Gradient Boosting models were used to predict seizure-free status at the most recent clinical review. To address class imbalance, the Synthetic Minority Oversampling Technique (SMOTE) was applied within training folds. Model performance was assessed using stratified cross-validation and evaluated using accuracy, precision, recall, and precision–recall area under the curve. Results Demographic and socioeconomic variables showed consistent associations with epilepsy treatment outcomes and contributed substantially to model performance. After accounting for class imbalance, Random Forest and Gradient Boosting models achieved accuracy, precision, and recall exceeding 83%. Individuals residing in more socioeconomically deprived areas were overrepresented among those who were not seizure-free, indicating persistent social gradients in epilepsy outcomes. Comparable predictive performance was observed when models were trained using predominantly non-clinical features, suggesting that social determinants carry independent prognostic value. Conclusions Demographic and socioeconomic factors are important predictors of epilepsy treatment outcomes and can be leveraged within machine learning frameworks to support population-level risk stratification. Incorporating social determinants of health into predictive approaches may help identify patients at increased risk of poor outcomes and inform more equitable, context-aware epilepsy care pathways. Epilepsy demographics health equity machine learning prediction neurology IMD socioeconomic factors Logistic Regression Random Forest Gradient Boosting SMOTE Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 1. Introduction Epilepsy is a chronic neurological disorder defined by the occurrence of unprovoked, recurrent seizures, and it remains one of the most common severe brain disorders globally [1]. It affects over 50 million people worldwide and approximately 600,000 individuals in the United Kingdom [2, 3]. While advances in pharmacological and surgical treatment have improved seizure control for many, around 30% of people with epilepsy continue to experience drug-resistant seizures [4, 5]. For these individuals, the consequences can extend beyond the neurological domain, affecting employment, education, mental health, and social integration [6–8]. The socioeconomic burden of epilepsy is significant and multifaceted. In the UK, the National Health Service (NHS) incurs substantial direct costs for epilepsy care, estimated at over £2 billion annually [9]. However, this figure does not capture the full economic impact, which includes indirect costs such as loss of productivity, reduced workforce participation, and informal caregiving [10–13]. Studies have shown that people with epilepsy are more likely to experience unemployment, social isolation, poverty, and reduced educational attainment compared to the general population [11–13]. These factors often create a cycle in which epilepsy both contributes to and is exacerbated by socioeconomic disadvantage. Importantly, socioeconomic inequalities can influence access to healthcare, treatment quality, medication adherence, and overall clinical outcomes [14, 15]. In the UK, individuals from ethnically minoritised communities and those living in areas of high deprivation have been shown to receive delayed diagnoses, experience greater difficulty accessing specialist care, and report poorer seizure control [16, 17]. The Index of Multiple Deprivation (IMD), a measure of relative deprivation in England produced by the Ministry of Housing, Communities and Local Government (MHCLG), captures area-level indicators such as income, employment, education, housing, and health, and can provide a robust tool for exploring these disparities [18]. Yet, its integration into predictive models for epilepsy treatment outcomes has been limited. Current approaches to epilepsy management primarily emphasise clinical variables, such as seizure type, aetiology, and medication regimens [19–22]. While these are crucial, growing evidence suggests that non-clinical factors such as age, gender, ethnicity, socioeconomic status, and geographic location play a vital role in determining treatment outcomes [7, 23, 24]. Neglecting these variables risks overlooking key determinants of therapeutic success, especially in diverse and socioeconomically heterogeneous populations. Indeed, evidence indicates that UK patients from more deprived backgrounds often face systemic barriers to timely diagnosis, equitable access to care, and sustained follow-up [25]. More specifically, evidence indicates that people from socioeconomically disadvantaged communities are more likely to experience delays in receiving an epilepsy diagnosis and referrals to epilepsy specialists, potentially resulting in misdiagnosis or underdiagnosis, particularly for those with functional seizures or atypical presentations [17, 26, 27]. Inadequate access to healthcare resources in these communities, such as fewer neurology clinics, longer waiting times, and limited transportation options, may further compound diagnostic delays [28], which can worsen epilepsy management and outcomes [29, 30]. Socioeconomic factors can also influence medication adherence and follow-up attendance. Individuals with lower incomes or those facing housing insecurity may have difficulty maintaining a consistent medication regimen, affording prescriptions, or attending regular clinical reviews, contributing to poorer seizure control [15]. In this regard, a population-based study in Denmark found that lower socioeconomic status was significantly associated with higher mortality and morbidity among individuals with epilepsy [31]. Similarly, UK-based studies have reported worse seizure outcomes and quality of life among those from more deprived areas [14]. These disparities suggest that non-clinical factors may be as important as clinical variables in predicting treatment response and long-term outcomes. Additionally, epilepsy is frequently comorbid with mental health disorders such as anxiety, depression, and post-traumatic stress disorder [32, 33]. These comorbidities are more prevalent and poorly managed in populations experiencing socioeconomic hardship [34, 35]. Stress-related unemployment, poverty, and social exclusion may also increase seizure frequency and severity, creating a vicious cycle in which epilepsy contributes to deprivation and vice versa [36, 37]. Recent advances in artificial intelligence (AI) and machine learning (ML) have shown significant promise in transforming the diagnosis, monitoring, and treatment of epilepsy. ML algorithms are being widely applied to analyse complex electroencephalography (EEG) signals for automatic seizure detection and prediction, with reported accuracies approaching or exceeding 98% in some models ([38, 39]). Techniques such as artificial neural networks [40], support vector machines [38], and ensemble methods like XGBoost [41] are used to classify seizures and identify atypical brain activity. Deep learning has further enabled the interpretation of EEG, neuroimaging, and wearable data to support tasks such as seizure detection, medication response prediction, and surgical outcome forecasting [42, 43]. Despite these technical advancements, there remain critical gaps in translating AI into equitable clinical practice. Indeed, to date, most machine learning (ML) studies predicting antiseizure medication (ASM) response have treated demographic variables such as age and sex as simple covariates, without examining how these factors interact with one another or with broader socioeconomic influences. For example, [19] a transformer-based model for first-ASM response included routine demographic variables but did not consider socioeconomic context or test interaction effects. Subasi et al.’s [38] brivaracetam prediction model drew primarily on clinical and genetic features with similarly limited demographic treatment. Reviews of the field (e.g., Xie et al. [44]) confirm that while datasets increasingly include EEG, MRI, genetic, and demographic information, most models remain “shallow” in their use of demographic data, focusing on main effects rather than systematically exploring intersectionality. This is a missed opportunity, given robust evidence that deprivation and other socioeconomic factors shape epilepsy incidence, access to care, and treatment outcomes. Models that ignore these dimensions risk producing biased or incomplete predictions. Recent work with large language models (LLMs) has begun to engage with these concerns. Abdaltawab et al. [45], for instance, trained an LLM on unstructured clinical notes to investigate disparities in treatment outcomes across demographic groups, illustrating the potential of such methods to reveal hidden inequities. Broader reviews of AI in epilepsy similarly highlight the importance of generalisability and ethical deployment, while acknowledging that socioeconomic dimensions remain underexplored. Together, these findings underscore a critical gap: current AI approaches in epilepsy rarely address the intersection of demographic and socioeconomic factors. Addressing this gap is essential for developing predictive models that are not only accurate but also equitable and patient-centred, moving beyond surface-level demographic adjustment to genuinely context-aware modelling of treatment outcomes. Given these findings, it is increasingly clear that a comprehensive, patient-centred approach to epilepsy care must include socioeconomic and demographic assessments. Predictive models that incorporate these features, as explored in this study, have the potential to support earlier interventions, optimise resource allocation, and reduce health inequalities. Furthermore, such approaches align with the NHS Long Term Plan’s goal of delivering personalised, equitable care and addressing the wider determinants of health [46]. Considering this, the present study investigated the potential of demographic and socioeconomic variables such as ethnicity, age and deprivation index to predict epilepsy treatment outcomes using machine learning approaches. In doing so, we seek to inform a more holistic approach to epilepsy care that recognises the interplay between medical and social determinants of health and supports data-driven strategies to improve patient outcomes across diverse communities. 2. Methodology This study involved a comprehensive analysis of demographic and clinical data from patients with epilepsy managed by the [masked] over a five-year period (2020–2024). All data were anonymised in strict accordance with institutional data protection protocols. The study population included individuals receiving care at the Salford Royal Hospital NHS Foundation Trust Specialist Epilepsy Clinic (n = 1,609). Clinical information was extracted from the electronic patient record system, while demographic data were sourced from the patient administration system. The demographic data included were: age, gender and ethnicity. Ethnicity was classified into five main groups (Black, Asian, White, Mixed and Other) as per the Office for National Statistics (ONS) census. The overall Index of Multiple Deprivation (IMD) and six other socioeconomic indicators were assessed: income, employment, education and skills, health and disability, barriers to housing and services, and living environment. Each indicator was reported in quintiles, of which lower values indicated higher levels of deprivation. We also included routinely collected clinical data such as seizure classification, medication history and non-pharmacological treatments. Data were extracted by a single member of the clinical team (NC) who determined treatment outcomes based on clinical records, and any ambiguities were resolved in discussion with a second clinician (RM). The study was approved by the Research & Development (R&D) department of the Salford Royal Hospital NHS Foundation Trust as a Health Improvement Project (22HIP46). 2.1 Data cleaning and preprocessing The original dataset contained 36 features, comprising both categorical and numerical variables. The first step involved cleaning and refining the data to address missing or ambiguous values. The existing labels, and their clinical relevance, were determined by the research team in consultation with RM, a clinical expert in epilepsy. Some categorical variables—such as ethnicity, epilepsy type, and aetiology—were reclassified into more meaningful and manageable categories (see Appendix 1 for details). All categorical variables were then converted into numerical form using label encoding. Features that were unclear, potentially misleading, or inconsistent due to data entry issues were removed. After this cleaning and preprocessing stage, 24 features remained and were used to analyse epilepsy treatment outcome (the target variable) and to develop the predictive models. 2.2 Tools and techniques To investigate the potential of demographic and socioeconomic factors in predicting epilepsy treatment outcomes, we applied a combination of classical and ensemble machine learning techniques alongside statistical analysis for feature evaluation and model development. All modelling tasks were conducted using Python-based data science libraries. First, exploratory data analysis (EDA) was conducted to examine the distribution of demographic and socioeconomic variables, including age, gender, ethnicity, and indices of deprivation. A correlation heatmap was produced to provide an initial overview of potential associations between the 24 feature variables and the binary treatment outcome (seizure-free vs. not-seizure-free). To assess feature importance systematically, two complementary methods were used: the univariate Chi-square (Chi²) test, which evaluates statistical independence between categorical variables and the target variable; and the Random Forest feature importance ranking, which identifies predictive variables based on their contribution to ensemble decision-tree splits. These analyses informed a reduced feature set comprising primarily demographic and socioeconomic variables for subsequent modelling. Particularly, three supervised learning algorithms were employed to develop classification models: Logistic Regression (a linear baseline model for binary classification) Random Forest (an ensemble method using multiple decision trees with bootstrap aggregation), and Gradient Boosting Classifier (an advanced boosting method that incrementally corrects classification errors). Model performance was evaluated in two experimental settings: Using the original dataset, stratified 5-fold cross-validation was performed to ensure balanced representation of the target classes across folds. Models were trained and evaluated using both the full feature set and the reduced demographic/socioeconomic subset. Hyperparameter tuning was conducted using RandomizedSearchCV to identify optimal configurations for each model. Addressing class imbalance, the Synthetic Minority Oversampling Technique (SMOTE) was applied to the training data to synthetically generate additional samples of the minority class (seizure-free outcomes). This resulted in a balanced dataset, upon which the same cross-validation and model evaluation procedures were repeated. To prevent data leakage, SMOTE was applied only to the training folds within each cross-validation split, and never to the validation folds. Model evaluation was therefore always conducted on data that were not synthetically generated. This ensured that reported performance metrics reflect generalisation to real patient data rather than artefacts of oversampling. For all approaches, data preprocessing included label mapping, imputation of missing values, and feature normalisation using a Min-Max scaler. Model performance was assessed using standard classification metrics—accuracy, precision and recall, with particular attention to precision and recall due to the clinical importance of accurately identifying seizure-free outcomes. In addition, Precision-Recall (PR) curves and Area Under the Curve (AUC) metrics were used to provide a more informative evaluation under imbalanced conditions. This multi-model, multi-phase approach enabled us to assess both the feasibility and reliability of using non-clinical factors in predicting epilepsy outcomes, as well as to evaluate model performance under realistic data constraints. 3. Results 3.1 Study Population In this cohort of 1,609 patients attending an epilepsy clinic, approximately 25% (n = 392) were classified as seizure-free at the most recent documented clinical assessment. Seizure freedom was defined as the absence of reported epileptic seizures at the most recent clinical review, as documented in the electronic patient record, consistent with routine clinical practice in specialist epilepsy services. As Table 1 (below) illustrates, the mean age was 39.9 years for the not-seizure-free group and 42.7 years for the seizure-free group, respectively. Approximately half the cohort were females in both groups, with the most common seizure type being focal seizures (64.8% for the non-seizure group and 63.0% for the seizure-free group). Most of the cohort was white (82.6%), and nearly half of the patients were from the most socioeconomically deprived quintile. The most commonly identified aetiology was genetic (9.0% in the not-seizure-free group and 11.0% in the seizure-free group). 3.2 Predictive analysis 3.2.1. Feature importance analysis As seen in Section 2 above, the final processed epilepsy data set consisted of 24 feature variables that are related to the epilepsy treatment outcome, which ranges from clinical conditions as well as demographic and socioeconomic factors. Thus, it is important to understand which features contribute most towards the prediction of a “seizure-free” epilepsy treatment outcome, with a particular focus on demographic and social indicators. First, we assessed whether socioeconomic and demographic indicators were independently predictive of clinical indicators. A cross-correlation heatmap was analysed with all 24 features as seen in Fig. 1 . The heatmap shows there is negligible correlation between the clinical indicators and socioeconomic indicators, whereas the socioeconomic and demographic indicators were comparatively more strongly correlated with each other. This demonstrates that the social and demographic indicators demonstrated a strong predictive signal even in the absence of clinical features. Additionally, the ranking of the correlation value of demographic and socioeconomic features against the target “Treatment Outcome” in Fig. 2 shows that they are positively correlated with the target. Finally, to evaluate the feature importance, two different techniques are used: the univariate Chi² test and the Random Forest model, as seen in Figures 3 and 4. Based on this observation, the following reduced feature set, comprising primarily demographic and socioeconomic variables alongside high-level clinical descriptors, was identified as contributing reasonably well towards epilepsy treatment outcome prediction. The feature set encompassed: 'Age', 'Gender', 'Ethnicity', 'IMD', 'Income', 'Employment', 'Education Skills', 'Health Disability', 'Barriers to Housing Services', 'Living Environment', 'Income Deprivation Affecting Children Index (IDACI)', 'Income Deprivation Affecting Older People Index (IDAOPI)', 'Epilepsy Type' and 'Aetiology' compared to the clinical features. Epilepsy type and aetiology were retained as coarse-grained clinical descriptors, as they are routinely recorded at diagnosis and do not encode treatment-specific or medication-level information. 3.2.2. Model training Once the feature importance is determined, three state-of-the-art machine learning models, namely, Logistic Regression, Random Forest and Gradient Boosting, were used for the performance evaluation and the best performing one was identified for predicting the epilepsy treatment outcome. Initially, the models were trained on all feature sets, and subsequently on the selected demographic and socioeconomic variables, to determine whether these factors can predict treatment outcomes, even in the presence of minimal clinical information. Due to the small dataset size and class imbalance in the target variable “current treatment outcome,” two distinct approaches to model training and prediction were employed. In the first approach, the original cleaned and preprocessed dataset was used to develop the models, applying the stratified K-fold cross-validation technique with hyperparameter tuning. In this case, model performance was evaluated initially using the full set of feature list and then with the demographic feature list as discussed above in Section 3.2.1 . In the second approach, due to target class imbalances, a technique known as Synthetic Minority Oversampling Technique (SMOTE) was used to oversample and synthetically increase and create a balanced dataset with equal number of target class labels. Then the first approach is repeated with demographic features to observe any improvement of model performance. In both approaches above, data were cleaned, missing values were imputed, label mappings were performed as discussed in Section 2 and shown in Appendix 1, and the data were finally normalised using a Min-Max scaler for model training. 3.2.3. Model performance 3.2.3.1 Approach one - Analysis without applying SMOTE The preprocessed and cleaned original dataset, comprising 1609 rows of data, was used for model training. Stratified 5-fold cross-validation and RandomizedSearchCV were used to select the best model parameters for the three models. The following table shows the best training parameters for each model. The table below shows the model performance predicting the epilepsy treatment outcome using the optimal hyperparameter combination. As seen in Table 2 , when all the features, including medication and clinical variables, were used, the model performance gave good results across all three models, with accuracy and recall in the range of 75%. Logistic Regressions’ performance was comparatively slightly better, with both Random Forest and Gradient Boosting performing almost equally well. Table 3 shows the model prediction performance when the demographic and socioeconomic features were used only. Interestingly, the performance was similar to the previous one, with Random Forest demonstrating the best performance overall, slightly better than Gradient Boosting. The precision fell slightly for Gradient Boosting and Logistic Regression, showing significantly poor precision around 57.2%. Table 2 Model performance for approach one with all predictor variables Model Mean Accuracy Mean Precision Mean Recall Hyperparameter Logistic regression 75.6% 71.6% 75.6% {'C': 0.216, 'multi_class': 'ovr', 'penalty': 'l2', 'solver': 'newton-cg'} Random Forest 75.6% 70.6% 75.6% {'n_estimators': 1000, 'min_samples_split': 5, 'min_samples_leaf': 4, 'max_features': 'auto', 'max_depth': 40, 'criterion': 'gini', 'bootstrap': False} Gradient Boosting 75.0% 70.0% 75.0% {'n_estimators': 50, 'max_depth': 3, 'loss': 'log_loss', 'learning_rate': 0.178, 'criterion': 'squared_error'} Table 3 Model performance for approach one with demographic predictor variables only Model Mean Accuracy Mean Precision Mean Recall Hyperparameter Logistic regression 75.6% 57.2% 75.6% {'C': 1.57, 'multi_class': 'ovr', 'penalty': 'l2', 'solver': 'newton-cg'} Random Forest 75.9% 72.3% 75.9% {'n_estimators': 833, 'min_samples_split': 10, 'min_samples_leaf': 4, 'max_features': 'sqrt', 'max_depth': 80, 'criterion': 'gini', 'bootstrap': False} Gradient Boosting 74.8% 67.4% 74.8% {'n_estimators': 50, 'max_depth': 3, 'loss': 'log_loss', 'learning_rate': 0.178, 'criterion': 'squared_error'} In general, the results show that even without medication and clinical features, the model can predict with similar accuracy, precision and recall, thus demonstrating that demographic and socioeconomic features represent a strong correlation towards predicting epilepsy treatment outcome. 3.2.3.2 Approach Two - Analysis applying SMOTE In the original dataset, the count of the target class labels was as follows: Not-seizure-free – 1217 and seizure-free – 392. Since there is a clear class imbalance, SMOTE technique was used to oversample the data and create a balanced dataset of 1217 instances of each class label. Hyperparameter tuning was done on the new dataset again, and then the 5-fold stratified cross-validation was applied. The results when SMOTE was applied can be seen in Table 4 . The performance of Logistic Regression fell to around 66%, whereas Random Forest and Gradient Boosting showed marked improvement in the overall performance, with both accuracies, recall and precision hovering above 80%. Table 4 Model performance for approach two with smote and demographic predictor variables only Model Mean Accuracy Mean Precision Mean Recall Hyperparameter Logistic regression 66.0% 67.0% 66.0% {'C': 0.065, 'multi_class': 'ovr', 'solver': 'sag'} Random Forest 84.0% 84.0% 84.0% {'n_estimators': 777, 'min_samples_split': 2, 'min_samples_leaf': 1, 'max_features': 'sqrt', 'max_depth': 60, 'criterion': 'gini', 'bootstrap': True} Gradient Boosting 83.0% 83.0% 82.0% {'n_estimators': 100, 'max_depth': 5, 'loss': 'log_loss', 'learning_rate': 0.094, 'criterion': 'squared_error'} After applying SMOTE to address class imbalance, both the Random Forest and Gradient Boosting classifiers achieved over 83% accuracy, precision, and recall, substantially outperforming Logistic Regression. These findings highlight the strong predictive value of non-clinical features, including age, ethnicity, gender, and multiple domains of the Index of Multiple Deprivation (IMD), which showed high feature importance. As the dataset was imbalanced, we evaluated model robustness using precision–recall (PR) curves, as shown in Fig. 5 and Fig. 6 . The PR curves showed that before applying SMOTE, the average area under the curve (AUC) was only around 30%, indicating that the models could not simultaneously achieve high precision and recall. After applying SMOTE, the PR curves showed marked improvement, with Gradient Boosting and Random Forest achieving average AUCs of about 90% and Logistic Regression improving to around 55%. This demonstrates that class imbalance was a major constraint and that balancing the data substantially improved predictive performance. 4. Discussion This study demonstrates that demographic and socioeconomic factors, even with minimal clinical information, can meaningfully predict epilepsy treatment outcomes, with machine learning (ML) models achieving performance comparable to those incorporating clinical variables. After addressing class imbalance with SMOTE, both Random Forest and Gradient Boosting achieved accuracies, precision, and recall above 83%, substantially outperforming Logistic Regression. These results highlight the strong predictive value of non-clinical features—particularly age, gender, ethnicity, and multiple domains of the Index of Multiple Deprivation (IMD)—as determinants of treatment response. They also point out that the likelihood of achieving seizure freedom can be substantially influenced by these indicators alone. Our findings align with growing evidence that socioeconomic disadvantage is closely linked to poorer epilepsy outcomes. Prior studies have shown that patients from deprived areas experience delays in diagnosis, lower adherence to medication, and worse seizure control [ 14 , 17 , 37 ]. However, such factors are rarely incorporated into predictive modelling frameworks, which typically emphasise clinical or neurophysiological data such as seizure type, aetiology, or EEG features [ 19 , 43 ]. By demonstrating that demographic and socioeconomic data alone can achieve comparable predictive performance, this study provides novel evidence that these social determinants are not merely contextual but represent key risk indicators in their own right. There are several plausible mechanisms underpinning these results. Socioeconomic disadvantage can affect epilepsy outcomes through multiple, interacting pathways that are structural, behavioural, and biological. Structurally, people living in deprived areas are less likely to access timely specialist review and more likely to experience diagnostic delay or misdiagnosis—patterns that are amplified in some ethnically minoritised groups [ 17 , 27 , 28 ]. Behaviourally, financial strain, unstable housing, caregiving burden, and transport constraints reduce continuity of follow-up and undermine adherence to antiseizure medication [ 15 , 30 ]. Biologically, chronic psychosocial stress associated with poverty and insecurity can increase seizure susceptibility via neuroendocrine and inflammatory pathways, compounding clinical risk [ 34 , 36 ]. In our cohort, younger adults from the most deprived quintiles were disproportionately represented among those who were not seizure-free, consistent with a cumulative disadvantage model in which early-life and young-adult exposures erode treatment responsiveness over time. These observations accord with epidemiological evidence linking lower socioeconomic position with higher morbidity and mortality in epilepsy [ 31 ] and reinforce the need to address wider determinants of health in routine care. Taken together, these findings support integrating demographic and socioeconomic profiling into early risk-stratification to identify patients who may benefit from targeted adherence support, social prescribing, proactive mental-health input, or prioritised referral to specialist services. Deprivation-aware decision support could also guide resource allocation, enabling teams to anticipate complexity and tailor follow-up intensity where need is greatest. This is squarely aligned with the NHS Long Term Plan’s commitment to personalised, inequality-reducing care [ 46 ] and with recent equity frameworks in epilepsy that call for the routine consideration of social context in clinical pathways [ 14 , 26 ]. Importantly, incorporating demographic and socioeconomic predictors into AI-driven models is not only a route to better predictive accuracy but also a mechanism for advancing equity. The broader health-AI literature shows that algorithms trained without attention to social and demographic context can inadvertently amplify disparities [ 39 , 47 ]. In epilepsy, equity-aware modelling implies (i) including robust measures of deprivation and other social determinants; (ii) testing interactions (intersectionality) between demographic and socioeconomic factors; (iii) reporting subgroup performance, calibration, and error patterns; and (iv) externally validating models across diverse settings [ 15 , 22 ]. Embedding these practices helps ensure that decision-support tools mitigate, rather than entrench, structural inequities—an essential step for person-centred epilepsy care that recognises the interplay of medical, social, and demographic determinants. Framed within a person-centred model, our results argue for care pathways that combine clinical profiling with explicit assessment of social needs, and for predictive tools that surface actionable risks (e.g., likely adherence challenges or access barriers) to inform shared decision-making. By “learning from the margins”—i.e., allowing the experiences of the most disadvantaged to shape model design and service responses—clinics can deliver more precise, compassionate, and fair epilepsy care. 5. Limitations This was a single-centre study using retrospective data, which may limit generalisability to other populations. The dataset was relatively small and initially imbalanced, and although SMOTE improved performance, oversampling may inflate accuracy by generating synthetic data points. Additionally, although SMOTE improved model performance, oversampling techniques may still inflate apparent accuracy by reducing class heterogeneity, and real-world performance may be lower when applied to unseen clinical populations. The analyses performed were also cross-sectional, preventing assessment of temporal dynamics or causal relationships. Furthermore, some demographic and clinical fields were incomplete or prone to data entry errors, which could introduce bias. External validation on larger, multicentre, and ethnically diverse datasets will be crucial for confirming the robustness of these models. We also acknowledge that seizure freedom is a clinically nuanced outcome and may vary with follow-up duration and reporting accuracy; however, the definition used reflects real-world decision-making in specialist epilepsy clinics. 6. Conclusion and Future Directions This study provides evidence that demographic and socioeconomic determinants, when analysed through machine learning frameworks, can serve as meaningful predictors of epilepsy treatment outcomes. Although performance was constrained by class imbalance in the original dataset, the application of SMOTE markedly enhanced model robustness, with Gradient Boosting and Random Forest achieving precision–recall performance comparable to models incorporating clinical features. These findings support the integration of social determinants into routine risk stratification for epilepsy care, contributing to more equitable treatment planning. Further research with larger and more diverse cohorts is warranted to validate generalisability and to explore the integration of social determinants into clinical decision-support systems. Abbreviations AI Artificial Intelligence ASM Antiseizure Medication AUC Area Under the Curve Chi² Chi–square test EDA Exploratory Data Analysis EEG Electroencephalography IMD Index of Multiple Deprivation IDACI Income Deprivation Affecting Children Index IDAOPI Income Deprivation Affecting Older People Index ILAE International League Against Epilepsy LLM Large Language Model ML Machine Learning MRI Magnetic Resonance Imaging NHS National Health Service NICE National Institute for Health and Care Excellence ONS Office for National Statistics PR Precision–Recall PTSD Post–Traumatic Stress Disorder R&D Research and Development SMOTE Synthetic Minority Oversampling Technique UK United Kingdom WHO World Health Organisation Declarations Ethics Approval and Consent to Participate This study was approved by the Research and Development Department of Salford Royal Hospital NHS Foundation Trust as a registered Health Improvement Project (reference: 22HIP46). All study procedures were conducted in accordance with applicable institutional policies and UK data protection regulations, and in compliance with the ethical principles of the Declaration of Helsinki. The study involved retrospective analysis of routinely collected clinical and demographic data. All data were anonymised prior to access and analysis, and no identifiable patient information was available to the research team. In accordance with NHS Health Research Authority guidance, individual informed consent was not required, as the study used anonymised routinely collected data and involved no direct patient contact or additional risk to participants. Consent for Publication Not applicable. This manuscript does not contain any individual person’s identifiable data in any form (including individual details, images, or videos). All data analysed in this study were fully anonymised prior to access and analysis, and no identifiable information is presented in the manuscript. Availability of Data and Materials The datasets generated and/or analysed during the current study are not publicly available due to institutional governance restrictions and UK data protection regulations, as they contain sensitive clinical information derived from NHS patient records. Fully anonymised data may be made available from the corresponding author on reasonable request, subject to approval by Salford Royal Hospital NHS Foundation Trust and relevant data governance procedures. Requests will be considered in line with applicable ethical approvals, data sharing agreements, and legal requirements . Competing Interests The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest. Funding This work was supported by the Alan Turing Institute – University of Manchester Sandpit Project Grant, 2022. Authors’ Contributions MSM, LYHL, YL and DDB conceived and designed the study. NC curated and collected the data. MSM, LYHL, YL, and NC performed the data analysis. SS, MS, and AW contributed to the analyses and ensured the correct application of methods. RM and DDB provided clinical insights and guided the selection and interpretation of features for analysis. All authors reviewed and approved the final manuscript. Acknowledgments The authors would like to thank the clinical and administrative staff at Salford Royal Hospital NHS Foundation Trust for their support in facilitating data access and governance approvals. We are particularly grateful to the members of the specialist epilepsy service for their assistance during data extraction and clarification of clinical records. We also acknowledge the patients whose routinely collected clinical data made this research possible. This work was supported by the Turing–Manchester Sandpit Project Grant (2022), and the authors thank the organisers and participants of the sandpit for valuable interdisciplinary discussions that informed the development of this study. References Feigin VL, Vos T, Nair BS, Hay SI, Abate YH, Abd Al Magied AHA, et al. Global, regional, and national burden of epilepsy, 1990–2021: a systematic analysis for the Global Burden of Disease Study 2021. Lancet Public Health. 2025;10:e203–27. https://doi.org/10.1016/S2468-2667(24)00302-5 . Epilepsy Society. What is epilepsy? 2024. Available from: https://epilepsysociety.org.uk/about-epilepsy/what-epilepsy World Health Organization. Epilepsy. Key facts. 2024. Available from: https://www.who.int/news-room/fact-sheets/detail/epilepsy Dalic L, Cook M. Managing drug-resistant epilepsy: challenges and solutions. Neuropsychiatr Dis Treat. 2016;12:2605–16. https://doi.org/10.2147/ndt.s84852 . Mesraoua B, Deleu D, Kullmann DM, Shetty AK, Boon P, Perucca E, et al. Novel therapies for epilepsy in the pipeline. Epilepsy Behav. 2019;97:282–90. https://doi.org/10.1016/j.yebeh.2019.04.042 . Karakis I, Janocko NJ, Morton ML, Groover O, Teagarden DL, Villarreal HK, et al. Stigma in psychogenic nonepileptic seizures. Epilepsy Behav. 2020;111. https://doi.org/10.1016/j.yebeh.2020.107269 . Strzelczyk A, Aledo-Serrano A, Coppola A, Didelot A, Bates E, Sainz-Fuertes R, et al. The impact of epilepsy on quality of life: Findings from a European survey. Epilepsy Behav. 2023;142:109179. https://doi.org/10.1016/j.yebeh.2023.109179 . Yeni K. Stigma and psychosocial problems in patients with epilepsy. Explor Neurosci. 2023;2:251–63. https://doi.org/10.37349/en.2023.00026 . National Institute for Health and Care Excellence. Epilepsies in children, young people and adults. 2025. Hussain SA, Ortendahl JD, Bentley TGK, Harmon AL, Gupta S, Begley CE, et al. The economic burden of caregiving in epilepsy: An estimate based on a survey of US caregivers. Epilepsia. 2020;61:319–29. https://doi.org/10.1111/epi.16429 . Maas L, Kellenaers J, van Mastrigt G, van Kuijk SMJ, Vlooswijk MCG, Hiligsmann M, et al. Societal costs and quality of life analysis in patients undergoing resective epilepsy surgery: A one-year follow-up. Epilepsy Behav Rep. 2023;24:100635. https://doi.org/10.1016/j.ebr.2023.100635 . Marquina C, Foster E, Chen Z, Vaughan DN, Abbott DF, Tailby C, et al. Work productivity, quality of life, and care needs: An unfolding epilepsy burden revealed in the Australian Epilepsy Project pilot study. Epilepsia Open. 2024;9:739–49. https://doi.org/10.1002/epi4.12919 . Sarac E, Yildiz E. The effect of epilepsy self-management on productivity at work. Epilepsy Behav. 2024;157:109839. https://doi.org/10.1016/j.yebeh.2024.109839 . Bush KJ, Cullen E, Mills S, Chin RFM, Thomas RH, Kingston A, et al. Assessing the extent and determinants of socioeconomic inequalities in epilepsy in the UK: a systematic review and meta-analysis of evidence. Lancet Public Health. 2024;9:e614–28. https://doi.org/10.1016/s2468-2667(24)00132-4 . Correa DJ, Gutierrez CA. Health Disparities and Inequities in Epilepsy. Achieving Equity in Neurological Practice: Principles and Pathways. Springer; 2024. pp. 91–123. Hayanga B, Stafford M, Bécares L. Ethnic inequalities in multiple long-term health conditions in the United Kingdom: a systematic review and narrative synthesis. BMC Public Health. 2023;23. https://doi.org/10.1186/s12889-022-14940-w . Maloney EM, Corcoran P, Costello DJ, O’Reilly ÉJ. Association between social deprivation and incidence of first seizures and epilepsy: A prospective population-based cohort. Epilepsia. 2022;63:2108–19. https://doi.org/10.1111/epi.17313 . Health Equity Evidence Centre. Understanding the Index of Multiple Deprivation (IMD) in public health research. 2024. Hakeem H, Feng W, Chen Z, Choong J, Brodie MJ, Fong S-L, et al. Development and validation of a deep learning model for predicting treatment response in patients with newly diagnosed epilepsy. JAMA Neurol. 2022;79:986–96. Perucca E. The pharmacological treatment of epilepsy: recent advances and future perspectives. Acta Epileptol. 2021;3:22. Piccenna L, O’Dwyer R, Leppik I, Beghi E, Giussani G, Costa C, et al. Management of epilepsy in older adults: A critical review by the ILAE Task Force on Epilepsy in the elderly. Epilepsia. 2023;64:567–85. https://doi.org/10.1111/epi.17426 . Ratcliffe C, Pradeep V, Marson A, Keller SS, Bonnett LJ. Clinical prediction models for treatment outcomes in newly diagnosed epilepsy: A systematic review. Epilepsia. 2024;65:1811–46. Fiest KM, Sauro KM, Wiebe S, Patten SB, Kwon C-S, Dykeman J, et al. Prevalence and incidence of epilepsy. Neurology. 2017;88:296–303. https://doi.org/10.1212/wnl.0000000000003509 . Kharkar S, Pillai J, Rochestie D, Haneef Z. Socio-Demographic Influences on Epilepsy Outcomes in an Inner-City Population. Seizure. 2014;23. https://doi.org/10.1016/j.seizure.2014.01.002 . Department of Health and Social Care. Independent Investigation of the National Health Service in England (accessible version). 2024. Bensken WP, Alberti PM, Khan OI, Williams SM, Stange KC, Vaca GF-B, et al. A framework for health equity in people living with epilepsy. Epilepsy Res. 2022;188:107038. https://doi.org/10.1016/j.eplepsyres.2022.107038 . Kiriakopoulos ET, García Sosa R, Blank L, Johnson EL, Gutierrez C. Shining a Light: Advancing Health Equity in Overlooked Epilepsy Communities. Epilepsy Curr. 2024. https://doi.org/10.1177/15357597241258081 . Alessi N, Perucca P, McIntosh AM. Missed, mistaken, stalled: identifying components of delay to diagnosis in epilepsy. Epilepsia. 2021;62:1494–504. Pellinen J. Treatment gaps in epilepsy. Front Epidemiol. 2022. https://doi.org/10.3389/fepid.2022.976039 . 2. Perzynski AT, Ramsey RK, Colón-Zimmermann K, Cage J, Welter E, Sajatovic M. Barriers and facilitators to epilepsy self-management for patients with physical and psychological co-morbidity. Chronic Illn. 2017;13:188–203. https://doi.org/10.1177/1742395316674540 . Jennum P, Gyllenborg J, Kjellberg J. The social and economic consequences of epilepsy: A controlled national study. Epilepsia. 2011;52:949–56. https://doi.org/10.1111/j.1528-1167.2010.02946.x . Mula M, Kanner AM, Jetté N, Sander JW. Psychiatric comorbidities in people with epilepsy. Neurol Clin Pract. 2021;11:e112–20. Tsigebrhan R, Derese A, Kariuki SM, Fekadu A, Medhin G, Newton CR, et al. Co-morbid mental health conditions in people with epilepsy and association with quality of life in low-and middle-income countries: a systematic review and meta-analysis. Health Qual Life Outcomes. 2023;21:5. Marbin D, Gutwinski S, Schreiter S, Heinz A. Perspectives in poverty and mental health. Front Public Health. 2022;10:975482. Smith MV, Mazure CM. Mental health and wealth: depression, gender, poverty, and parenting. Annu Rev Clin Psychol. 2021;17:181–205. Lolk K, Werenberg Dreier J, Christensen J. Individual and neighborhood-level socioeconomic deprivation and risk of epilepsy after traumatic brain Injury: A register-based cohort study. Epilepsy Behav. 2024;156:109807. https://doi.org/10.1016/j.yebeh.2024.109807 . Pickrell WO, Lacey AS, Bodger OG, Demmler JC, Thomas RH, Lyons RA, et al. Epilepsy and deprivation, a data linkage study. Epilepsia. 2015;56:585–91. https://doi.org/10.1111/epi.12942 . Subasi A, Kevric J, Abdullah Canbaz M. Epileptic seizure detection using hybrid machine learning methods. Neural Comput Appl. 2019;31:317–25. https://doi.org/10.1007/s00521-017-3003-y . Tran LV, Tran HM, Le TM, Huynh TTM, Tran HT, Dao SVT. Application of Machine Learning in Epileptic Seizure Detection. Diagnostics. 2022;12:2879. https://doi.org/10.3390/diagnostics12112879 . Chakrabarti S, Swetapadma A, Ranjan A, Pattnaik PK. Time domain implementation of pediatric epileptic seizure detection system for enhancing the performance of detection and easy monitoring of pediatric patients. Biomed Signal Process Control. 2020;59:101930. https://doi.org/10.1016/j.bspc.2020.101930 . Torlay L, Perrone-Bertolotti M, Thomas E, Baciu M. Machine learning–XGBoost analysis of language networks to classify patients with epilepsy. Brain Inf. 2017;4:159–69. https://doi.org/10.1007/s40708-017-0065-7 . Han K, Liu C, Friedman D. Artificial intelligence/machine learning for epilepsy and seizure diagnosis. Epilepsy Behav. 2024;155:109736. https://doi.org/10.1016/j.yebeh.2024.109736 . Smolyansky ED, Hakeem H, Ge Z, Chen Z, Kwan P. Machine learning models for decision support in epilepsy management: A critical review. Epilepsy Behav. 2021;123:108273. https://doi.org/10.1016/j.yebeh.2021.108273 . Xie K, Ojemann WKS, Gallagher RS, Shinohara RT, Lucas A, Hill CE, et al. Disparities in seizure outcomes revealed by large language models. J Am Med Inf Assoc. 2024;31:1348–55. https://doi.org/10.1093/jamia/ocae047 . Abdaltawab A, Chang L-C, Mansour M, Koubeissi M. How accurate are machine learning models in predicting antiseizure medication responses: A systematic review. Epilepsy Behav. 2025;163. https://doi.org/10.1016/j.yebeh.2024.110212 . NHS England. NHS Long Term Plan. 2019. Obermeyer Z, Powers B, Vogeli C, Mullainathan S. Dissecting racial bias in an algorithm used to manage the health of populations. Science. 2019;366:447–53. https://doi.org/10.1126/science.aax2342 . Table 1 Table 1 is available in the Supplementary Files section. Additional Declarations No competing interests reported. Supplementary Files SupplementaryMaterialMLtoPredictEpilepsyOutcomesfromSocioeconomicFactorsv1.303026.docx Table1.docx Cite Share Download PDF Status: Under Review Version 1 posted Reviewers invited by journal 24 Mar, 2026 Editor invited by journal 27 Feb, 2026 Editor assigned by journal 25 Feb, 2026 Submission checks completed at journal 25 Feb, 2026 First submitted to journal 19 Feb, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8917468","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":612208515,"identity":"0081a5de-8720-4146-998b-55fd22bb2863","order_by":0,"name":"Md Shadab Mashuk","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABB0lEQVRIie3RMUvDQBjG8bccpMubPSUln+GRQKQo+FVydHAJIrh0kNCpXdLOFfwQnTofHNQlxdWxJatDiyABgzSJm3C2o8P9pyN3P+5eQmSz/cOgulvBRDGxUET3ATXrdsdEBoohnlvi1EcRniY3iklULWGcR6KGJI/pHfL8Y18CweU07uxL0qGZ1LMka/2AzWzVy4Cwn29FLyMdmQlBJI6Sy1d3RYxKLryYfCJ9bSJoCH+nNeHiUAENEV8niTsRcrnJyOcf4jS3GB8GzdDuXMunfB35/XoWj3eTQYZb4/h4me4K/kzlPB8Wh/cRAq871G/l6OpibDKCSP361Bn/8VdsNpvNdk5HSgNYBt/Z2A4AAAAASUVORK5CYII=","orcid":"","institution":"University of Salford","correspondingAuthor":true,"prefix":"","firstName":"Md","middleName":"Shadab","lastName":"Mashuk","suffix":""},{"id":612208523,"identity":"09beaeb0-57b0-48f3-ac20-cfd76fbad366","order_by":1,"name":"Lana Lai","email":"","orcid":"","institution":"University of Manchester","correspondingAuthor":false,"prefix":"","firstName":"Lana","middleName":"","lastName":"Lai","suffix":""},{"id":612208530,"identity":"5a63734f-63de-4609-b40a-99627c918b0a","order_by":2,"name":"Yang Lu","email":"","orcid":"","institution":"Loughborough University","correspondingAuthor":false,"prefix":"","firstName":"Yang","middleName":"","lastName":"Lu","suffix":""},{"id":612208536,"identity":"544f3787-b623-4185-a555-2ba8685247cd","order_by":3,"name":"Shumit Saha","email":"","orcid":"","institution":"Meharry Medical College","correspondingAuthor":false,"prefix":"","firstName":"Shumit","middleName":"","lastName":"Saha","suffix":""},{"id":612208541,"identity":"26ca9bed-1464-40a5-b8b4-44514e617732","order_by":4,"name":"Natasha Carmichael","email":"","orcid":"","institution":"University of Manchester","correspondingAuthor":false,"prefix":"","firstName":"Natasha","middleName":"","lastName":"Carmichael","suffix":""},{"id":612208544,"identity":"91cb08cf-4e1b-4bd0-b558-c25c939195de","order_by":5,"name":"Matthew Shardlow","email":"","orcid":"","institution":"Manchester Metropolitan University","correspondingAuthor":false,"prefix":"","firstName":"Matthew","middleName":"","lastName":"Shardlow","suffix":""},{"id":612208546,"identity":"85eaadda-c510-470e-862f-ddfe86dcc3a1","order_by":6,"name":"Ashley Williams","email":"","orcid":"","institution":"Manchester Metropolitan University","correspondingAuthor":false,"prefix":"","firstName":"Ashley","middleName":"","lastName":"Williams","suffix":""},{"id":612208548,"identity":"10091035-3806-448d-951c-7bfccf956efd","order_by":7,"name":"Rajiv Mohanraj","email":"","orcid":"","institution":"Northern Care Alliance NHS Foundation Trust","correspondingAuthor":false,"prefix":"","firstName":"Rajiv","middleName":"","lastName":"Mohanraj","suffix":""},{"id":612208550,"identity":"92ba9814-132a-48c6-a5c6-7af22c824ee3","order_by":8,"name":"Daniela Di Basilio","email":"","orcid":"","institution":"Lancaster University","correspondingAuthor":false,"prefix":"","firstName":"Daniela","middleName":"Di","lastName":"Basilio","suffix":""}],"badges":[],"createdAt":"2026-02-19 12:09:44","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8917468/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8917468/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":105499772,"identity":"7e3aef1f-0cd3-43c7-b515-e3f4fd04d6ea","added_by":"auto","created_at":"2026-03-26 17:15:47","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":460268,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eCross-correlation heatmap of all 24 features\u003c/em\u003e\u003c/p\u003e","description":"","filename":"image1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8917468/v1/ca078d029559c6561db0af0c.jpg"},{"id":105499773,"identity":"39ab2fd6-f511-4eef-bbe9-0c26e1757f6a","added_by":"auto","created_at":"2026-03-26 17:15:47","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":92202,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eCorrelation value against target variable (epilepsy treatment outcome)\u003c/em\u003e\u003c/p\u003e","description":"","filename":"image2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8917468/v1/70ab26220f9e00461531fb92.jpg"},{"id":105566503,"identity":"864dc2d7-9945-4c61-96f3-801146f739b3","added_by":"auto","created_at":"2026-03-27 12:56:34","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":67322,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eFeature importance ranking using Chi² univariate testing\u003c/em\u003e\u003c/p\u003e","description":"","filename":"image3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8917468/v1/72b1168b1754c164d2c48678.jpg"},{"id":105566368,"identity":"f76d2904-2c56-4cd8-ac50-9d2c3a8802fc","added_by":"auto","created_at":"2026-03-27 12:56:17","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":66539,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eFeature contribution towards model prediction using random forest model\u003c/em\u003e\u003c/p\u003e","description":"","filename":"image4.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8917468/v1/8ccfbd24eedf0b55685a9e32.jpg"},{"id":105499775,"identity":"7214ee28-27db-4188-9af3-325c6036744d","added_by":"auto","created_at":"2026-03-26 17:15:47","extension":"jpg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":103502,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003ePrecision-recall curve of the three models for each fold when the original dataset is used\u003c/em\u003e\u003c/p\u003e","description":"","filename":"image5.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8917468/v1/18909e64df96a2113758b31e.jpg"},{"id":105566644,"identity":"69ce76d9-33df-4869-b660-68e97d053f12","added_by":"auto","created_at":"2026-03-27 12:56:53","extension":"jpg","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":101567,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003ePrecision-recall curve of the three models for each fold when smote is used to balance the dataset\u003c/em\u003e\u003c/p\u003e","description":"","filename":"image6.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8917468/v1/f19a3a9d9208667c6a6bf555.jpg"},{"id":105904233,"identity":"66118966-bab9-45f3-a19d-141d245b3bb2","added_by":"auto","created_at":"2026-04-01 10:06:33","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1715158,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8917468/v1/0fe8c536-cca6-4c10-9b70-0eca39d6e1bb.pdf"},{"id":105752029,"identity":"1f3aee71-884a-4117-a28b-1008b6b59839","added_by":"auto","created_at":"2026-03-30 15:53:20","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":923962,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryMaterialMLtoPredictEpilepsyOutcomesfromSocioeconomicFactorsv1.303026.docx","url":"https://assets-eu.researchsquare.com/files/rs-8917468/v1/441feee5ddc7f1081db81202.docx"},{"id":105566958,"identity":"3a1fa3ad-8437-4808-9e74-055ebf48a81f","added_by":"auto","created_at":"2026-03-27 12:57:48","extension":"docx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":15746,"visible":true,"origin":"","legend":"","description":"","filename":"Table1.docx","url":"https://assets-eu.researchsquare.com/files/rs-8917468/v1/9d44d02606b197393266d6a7.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Learning from the Margins: Using Machine Learning Models to Predict Epilepsy Outcomes from Socioeconomic Factors","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eEpilepsy is a chronic neurological disorder defined by the occurrence of unprovoked, recurrent seizures, and it remains one of the most common severe brain disorders globally [1]. It affects over 50 million people worldwide and approximately 600,000 individuals in the United Kingdom [2, 3]. While advances in pharmacological and surgical treatment have improved seizure control for many, around 30% of people with epilepsy continue to experience drug-resistant seizures [4, 5]. For these individuals, the consequences can extend beyond the neurological domain, affecting employment, education, mental health, and social integration [6\u0026ndash;8]. The socioeconomic burden of epilepsy is significant and multifaceted. In the UK, the National Health Service (NHS) incurs substantial direct costs for epilepsy care, estimated at over \u0026pound;2 billion annually [9]. However, this figure does not capture the full economic impact, which includes indirect costs such as loss of productivity, reduced workforce participation, and informal caregiving [10\u0026ndash;13]. Studies have shown that people with epilepsy are more likely to experience unemployment, social isolation, poverty, and reduced educational attainment compared to the general population [11\u0026ndash;13].\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eThese factors often create a cycle in which epilepsy both contributes to and is exacerbated by socioeconomic disadvantage. Importantly, socioeconomic inequalities can influence access to healthcare, treatment quality, medication adherence, and overall clinical outcomes [14, 15]. In the UK, individuals from ethnically minoritised communities and those living in areas of high deprivation have been shown to receive delayed diagnoses, experience greater difficulty accessing specialist care, and report poorer seizure control [16, 17]. The Index of Multiple Deprivation (IMD), a measure of relative deprivation in England produced by the Ministry of Housing, Communities and Local Government (MHCLG), captures area-level indicators such as income, employment, education, housing, and health, and can provide a robust tool for exploring these disparities [18]. Yet, its integration into predictive models for epilepsy treatment outcomes has been limited.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eCurrent approaches to epilepsy management primarily emphasise clinical variables, such as seizure type, aetiology, and medication regimens [19\u0026ndash;22]. While these are crucial, growing evidence suggests that non-clinical factors such as age, gender, ethnicity, socioeconomic status, and geographic location play a vital role in determining treatment outcomes [7, 23, 24]. Neglecting these variables risks overlooking key determinants of therapeutic success, especially in diverse and socioeconomically heterogeneous populations. Indeed, evidence indicates that UK patients from more deprived backgrounds often face systemic barriers to timely diagnosis, equitable access to care, and sustained follow-up [25]. More specifically, evidence indicates that people from socioeconomically disadvantaged communities are more likely to experience delays in receiving an epilepsy diagnosis and referrals to epilepsy specialists, potentially resulting in misdiagnosis or underdiagnosis, particularly for those with functional seizures or atypical presentations [17, 26, 27].\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eInadequate access to healthcare resources in these communities, such as fewer neurology clinics, longer waiting times, and limited transportation options, may further compound diagnostic delays [28], which can worsen epilepsy management and outcomes \u0026nbsp; [29, 30]. Socioeconomic factors can also influence medication adherence and follow-up attendance. Individuals with lower incomes or those facing housing insecurity may have difficulty maintaining a consistent medication regimen, affording prescriptions, or attending regular clinical reviews, contributing to poorer seizure control [15].\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eIn this regard, a population-based study in Denmark found that lower socioeconomic status was significantly associated with higher mortality and morbidity among individuals with epilepsy [31]. Similarly, UK-based studies have reported worse seizure outcomes and quality of life among those from more deprived areas [14]. These disparities suggest that non-clinical factors may be as important as clinical variables in predicting treatment response and long-term outcomes. Additionally, epilepsy is frequently comorbid with mental health disorders such as anxiety, depression, and post-traumatic stress disorder [32, 33]. These comorbidities are more prevalent and poorly managed in populations experiencing socioeconomic hardship [34, 35]. Stress-related unemployment, poverty, and social exclusion may also increase seizure frequency and severity, creating a vicious cycle in which epilepsy contributes to deprivation and vice versa [36, 37].\u003c/p\u003e\n\u003cp\u003eRecent advances in artificial intelligence (AI) and machine learning (ML) have shown significant promise in transforming the diagnosis, monitoring, and treatment of epilepsy. ML algorithms are being widely applied to analyse complex electroencephalography (EEG) signals for automatic seizure detection and prediction, with reported accuracies approaching or exceeding 98% in some models ([38, 39]). Techniques such as artificial neural networks [40], support vector machines [38], and ensemble methods like XGBoost [41] are used to classify seizures and identify atypical brain activity. Deep learning has further enabled the interpretation of EEG, neuroimaging, and wearable data to support tasks such as seizure detection, medication response prediction, and surgical outcome forecasting [42, 43]. \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u003c/p\u003e\n\u003cp\u003eDespite these technical advancements, there remain critical gaps in translating AI into equitable clinical practice. Indeed, to date, most machine learning (ML) studies predicting antiseizure medication (ASM) response have treated demographic variables such as age and sex as simple covariates, without examining how these factors interact with one another or with broader socioeconomic influences. For example, [19] a transformer-based model for first-ASM response included routine demographic variables but did not consider socioeconomic context or test interaction effects. Subasi et al.\u0026rsquo;s [38] brivaracetam prediction model drew primarily on clinical and genetic features with similarly limited demographic treatment. Reviews of the field (e.g., Xie et al. \u0026nbsp;[44]) confirm that while datasets increasingly include EEG, MRI, genetic, and demographic information, most models remain \u0026ldquo;shallow\u0026rdquo; in their use of demographic data, focusing on main effects rather than systematically exploring intersectionality. This is a missed opportunity, given robust evidence that deprivation and other socioeconomic factors shape epilepsy incidence, access to care, and treatment outcomes. Models that ignore these dimensions risk producing biased or incomplete predictions.\u003c/p\u003e\n\u003cp\u003eRecent work with large language models (LLMs) has begun to engage with these concerns. Abdaltawab et al. [45], for instance, trained an LLM on unstructured clinical notes to investigate disparities in treatment outcomes across demographic groups, illustrating the potential of such methods to reveal hidden inequities. Broader reviews of AI in epilepsy similarly highlight the importance of generalisability and ethical deployment, while acknowledging that socioeconomic dimensions remain underexplored. Together, these findings underscore a critical gap: current AI approaches in epilepsy rarely address the intersection of demographic and socioeconomic factors. Addressing this gap is essential for developing predictive models that are not only accurate but also equitable and patient-centred, moving beyond surface-level demographic adjustment to genuinely context-aware modelling of treatment outcomes. \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u003c/p\u003e\n\u003cp\u003eGiven these findings, it is increasingly clear that a comprehensive, patient-centred approach to epilepsy care must include socioeconomic and demographic assessments. Predictive models that incorporate these features, as explored in this study, have the potential to support earlier interventions, optimise resource allocation, and reduce health inequalities. Furthermore, such approaches align with the NHS Long Term Plan\u0026rsquo;s goal of delivering personalised, equitable care and addressing the wider determinants of health [46]. Considering this, the present study investigated the potential of demographic and socioeconomic variables such as ethnicity, age and deprivation index to predict epilepsy treatment outcomes using machine learning approaches. In doing so, we seek to inform a more holistic approach to epilepsy care that recognises the interplay between medical and social determinants of health and supports data-driven strategies to improve patient outcomes across diverse communities.\u003c/p\u003e"},{"header":"2. Methodology","content":"\u003cp\u003eThis study involved a comprehensive analysis of demographic and clinical data from patients with epilepsy managed by the [masked] over a five-year period (2020\u0026ndash;2024). All data were anonymised in strict accordance with institutional data protection protocols. The study population included individuals receiving care at the Salford Royal Hospital NHS Foundation Trust Specialist Epilepsy Clinic (n\u0026thinsp;=\u0026thinsp;1,609). Clinical information was extracted from the electronic patient record system, while demographic data were sourced from the patient administration system. The demographic data included were: age, gender and ethnicity. Ethnicity was classified into five main groups (Black, Asian, White, Mixed and Other) as per the Office for National Statistics (ONS) census. The overall Index of Multiple Deprivation (IMD) and six other socioeconomic indicators were assessed: income, employment, education and skills, health and disability, barriers to housing and services, and living environment. Each indicator was reported in quintiles, of which lower values indicated higher levels of deprivation. We also included routinely collected clinical data such as seizure classification, medication history and non-pharmacological treatments. Data were extracted by a single member of the clinical team (NC) who determined treatment outcomes based on clinical records, and any ambiguities were resolved in discussion with a second clinician (RM).\u003c/p\u003e \u003cp\u003eThe study was approved by the Research \u0026amp; Development (R\u0026amp;D) department of the Salford Royal Hospital NHS Foundation Trust as a Health Improvement Project (22HIP46).\u003c/p\u003e \u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e2.1 Data cleaning and preprocessing\u003c/h2\u003e \u003cp\u003eThe original dataset contained 36 features, comprising both categorical and numerical variables. The first step involved cleaning and refining the data to address missing or ambiguous values. The existing labels, and their clinical relevance, were determined by the research team in consultation with RM, a clinical expert in epilepsy. Some categorical variables\u0026mdash;such as ethnicity, epilepsy type, and aetiology\u0026mdash;were reclassified into more meaningful and manageable categories (see Appendix 1 for details). All categorical variables were then converted into numerical form using label encoding. Features that were unclear, potentially misleading, or inconsistent due to data entry issues were removed. After this cleaning and preprocessing stage, 24 features remained and were used to analyse epilepsy treatment outcome (the target variable) and to develop the predictive models.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e2.2 Tools and techniques\u003c/h2\u003e \u003cp\u003eTo investigate the potential of demographic and socioeconomic factors in predicting epilepsy treatment outcomes, we applied a combination of classical and ensemble machine learning techniques alongside statistical analysis for feature evaluation and model development. All modelling tasks were conducted using Python-based data science libraries. First, exploratory data analysis (EDA) was conducted to examine the distribution of demographic and socioeconomic variables, including age, gender, ethnicity, and indices of deprivation. A correlation heatmap was produced to provide an initial overview of potential associations between the 24 feature variables and the binary treatment outcome (seizure-free vs. not-seizure-free). To assess feature importance systematically, two complementary methods were used: the univariate Chi-square (Chi\u0026sup2;) test, which evaluates statistical independence between categorical variables and the target variable; and the Random Forest feature importance ranking, which identifies predictive variables based on their contribution to ensemble decision-tree splits. These analyses informed a reduced feature set comprising primarily demographic and socioeconomic variables for subsequent modelling. Particularly, three supervised learning algorithms were employed to develop classification models:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eLogistic Regression (a linear baseline model for binary classification)\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eRandom Forest (an ensemble method using multiple decision trees with bootstrap aggregation), and\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eGradient Boosting Classifier (an advanced boosting method that incrementally corrects classification errors).\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eModel performance was evaluated in two experimental settings:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eUsing the original dataset, stratified 5-fold cross-validation was performed to ensure balanced representation of the target classes across folds. Models were trained and evaluated using both the full feature set and the reduced demographic/socioeconomic subset. Hyperparameter tuning was conducted using RandomizedSearchCV to identify optimal configurations for each model.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eAddressing class imbalance, the Synthetic Minority Oversampling Technique (SMOTE) was applied to the training data to synthetically generate additional samples of the minority class (seizure-free outcomes). This resulted in a balanced dataset, upon which the same cross-validation and model evaluation procedures were repeated. To prevent data leakage, SMOTE was applied only to the training folds within each cross-validation split, and never to the validation folds. Model evaluation was therefore always conducted on data that were not synthetically generated. This ensured that reported performance metrics reflect generalisation to real patient data rather than artefacts of oversampling.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eFor all approaches, data preprocessing included label mapping, imputation of missing values, and feature normalisation using a Min-Max scaler. Model performance was assessed using standard classification metrics\u0026mdash;accuracy, precision and recall, with particular attention to precision and recall due to the clinical importance of accurately identifying seizure-free outcomes. In addition, Precision-Recall (PR) curves and Area Under the Curve (AUC) metrics were used to provide a more informative evaluation under imbalanced conditions. This multi-model, multi-phase approach enabled us to assess both the feasibility and reliability of using non-clinical factors in predicting epilepsy outcomes, as well as to evaluate model performance under realistic data constraints.\u003c/p\u003e \u003c/div\u003e"},{"header":"3. Results","content":"\u003cdiv id=\"Sec6\" class=\"Section2\"\u003e\n \u003ch2\u003e3.1 Study Population\u003c/h2\u003e\n \u003cp\u003eIn this cohort of 1,609 patients attending an epilepsy clinic, approximately 25% (n\u0026thinsp;=\u0026thinsp;392) were classified as seizure-free at the most recent documented clinical assessment. Seizure freedom was defined as the absence of reported epileptic seizures at the most recent clinical review, as documented in the electronic patient record, consistent with routine clinical practice in specialist epilepsy services.\u003c/p\u003e\n \u003cp\u003eAs Table \u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e (below) illustrates, the mean age was 39.9 years for the not-seizure-free group and 42.7 years for the seizure-free group, respectively. Approximately half the cohort were females in both groups, with the most common seizure type being focal seizures (64.8% for the non-seizure group and 63.0% for the seizure-free group). Most of the cohort was white (82.6%), and nearly half of the patients were from the most socioeconomically deprived quintile. The most commonly identified aetiology was genetic (9.0% in the not-seizure-free group and 11.0% in the seizure-free group).\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec7\" class=\"Section2\"\u003e\n \u003ch2\u003e3.2 Predictive analysis\u003c/h2\u003e\n \u003cdiv id=\"Sec8\" class=\"Section3\"\u003e\n \u003ch2\u003e3.2.1. Feature importance analysis\u003c/h2\u003e\n \u003cp\u003eAs seen in Section \u003cspan refid=\"Sec2\" class=\"InternalRef\"\u003e2\u003c/span\u003e above, the final processed epilepsy data set consisted of 24 feature variables that are related to the epilepsy treatment outcome, which ranges from clinical conditions as well as demographic and socioeconomic factors. Thus, it is important to understand which features contribute most towards the prediction of a \u0026ldquo;seizure-free\u0026rdquo; epilepsy treatment outcome, with a particular focus on demographic and social indicators. First, we assessed whether socioeconomic and demographic indicators were independently predictive of clinical indicators. A cross-correlation heatmap was analysed with all 24 features as seen in Fig. \u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e. The heatmap shows there is negligible correlation between the clinical indicators and socioeconomic indicators, whereas the socioeconomic and demographic indicators were comparatively more strongly correlated with each other. This demonstrates that the social and demographic indicators demonstrated a strong predictive signal even in the absence of clinical features. Additionally, the ranking of the correlation value of demographic and socioeconomic features against the target \u0026ldquo;Treatment Outcome\u0026rdquo; in Fig. \u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e shows that they are positively correlated with the target.\u003c/p\u003e\n \u003c/div\u003e\u003cp\u003e\u0026nbsp; Finally, to evaluate the feature importance, two different techniques are used: the univariate Chi\u0026sup2; test and the Random Forest model, as seen in Figures 3 and 4. Based on this observation, the following reduced feature set, comprising primarily demographic and socioeconomic variables alongside high-level clinical descriptors, was identified as contributing reasonably well towards epilepsy treatment outcome prediction. The feature set encompassed: \u0026apos;Age\u0026apos;, \u0026apos;Gender\u0026apos;, \u0026apos;Ethnicity\u0026apos;, \u0026apos;IMD\u0026apos;, \u0026apos;Income\u0026apos;, \u0026apos;Employment\u0026apos;, \u0026apos;Education Skills\u0026apos;, \u0026apos;Health Disability\u0026apos;, \u0026apos;Barriers to Housing Services\u0026apos;, \u0026apos;Living Environment\u0026apos;, \u0026apos;Income Deprivation Affecting Children Index (IDACI)\u0026apos;, \u0026apos;Income Deprivation Affecting Older People Index (IDAOPI)\u0026apos;, \u0026apos;Epilepsy Type\u0026apos; and \u0026apos;Aetiology\u0026apos; compared to the clinical features. Epilepsy type and aetiology were retained as coarse-grained clinical descriptors, as they are routinely recorded at diagnosis and do not encode treatment-specific or medication-level information.\u003c/p\u003e\n\n\u003c/div\u003e\u003cdiv id=\"Sec9\" class=\"Section3\"\u003e \u003ch2\u003e3.2.2. Model training\u003c/h2\u003e \u003cp\u003eOnce the feature importance is determined, three state-of-the-art machine learning models, namely, Logistic Regression, Random Forest and Gradient Boosting, were used for the performance evaluation and the best performing one was identified for predicting the epilepsy treatment outcome. Initially, the models were trained on all feature sets, and subsequently on the selected demographic and socioeconomic variables, to determine whether these factors can predict treatment outcomes, even in the presence of minimal clinical information.\u003c/p\u003e \u003cp\u003eDue to the small dataset size and class imbalance in the target variable \u0026ldquo;current treatment outcome,\u0026rdquo; two distinct approaches to model training and prediction were employed. In the first approach, the original cleaned and preprocessed dataset was used to develop the models, applying the stratified K-fold cross-validation technique with hyperparameter tuning. In this case, model performance was evaluated initially using the full set of feature list and then with the demographic feature list as discussed above in Section \u003cspan refid=\"Sec8\" class=\"InternalRef\"\u003e3.2.1\u003c/span\u003e. In the second approach, due to target class imbalances, a technique known as Synthetic Minority Oversampling Technique (SMOTE) was used to oversample and synthetically increase and create a balanced dataset with equal number of target class labels. Then the first approach is repeated with demographic features to observe any improvement of model performance.\u003c/p\u003e \u003cp\u003eIn both approaches above, data were cleaned, missing values were imputed, label mappings were performed as discussed in Section \u003cspan refid=\"Sec2\" class=\"InternalRef\"\u003e2\u003c/span\u003e and shown in Appendix 1, and the data were finally normalised using a Min-Max scaler for model training.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec10\" class=\"Section3\"\u003e \u003ch2\u003e3.2.3. Model performance\u003c/h2\u003e \u003cdiv id=\"Sec11\" class=\"Section4\"\u003e \u003ch2\u003e3.2.3.1 Approach one - Analysis without applying SMOTE\u003c/h2\u003e \u003cp\u003eThe preprocessed and cleaned original dataset, comprising 1609 rows of data, was used for model training. Stratified 5-fold cross-validation and RandomizedSearchCV were used to select the best model parameters for the three models. The following table shows the best training parameters for each model. The table below shows the model performance predicting the epilepsy treatment outcome using the optimal hyperparameter combination. As seen in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e, when all the features, including medication and clinical variables, were used, the model performance gave good results across all three models, with accuracy and recall in the range of 75%. Logistic Regressions\u0026rsquo; performance was comparatively slightly better, with both Random Forest and Gradient Boosting performing almost equally well. Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e shows the model prediction performance when the demographic and socioeconomic features were used only. Interestingly, the performance was similar to the previous one, with Random Forest demonstrating the best performance overall, slightly better than Gradient Boosting. The precision fell slightly for Gradient Boosting and Logistic Regression, showing significantly poor precision around 57.2%.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003e\u003cem\u003eModel performance for approach one with all predictor variables\u003c/em\u003e\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMean Accuracy\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMean Precision\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMean Recall\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eHyperparameter\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLogistic regression\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e75.6%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e71.6%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e75.6%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e{'C': 0.216, 'multi_class': 'ovr', 'penalty': 'l2', 'solver': 'newton-cg'}\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRandom Forest\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e75.6%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e70.6%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e75.6%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e{'n_estimators': 1000, 'min_samples_split': 5, 'min_samples_leaf': 4, 'max_features': 'auto', 'max_depth': 40, 'criterion': 'gini', 'bootstrap': False}\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGradient Boosting\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e75.0%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e70.0%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e75.0%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e{'n_estimators': 50, 'max_depth': 3, 'loss': 'log_loss', 'learning_rate': 0.178, 'criterion': 'squared_error'}\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003e\u003cem\u003eModel performance for approach one with demographic predictor variables only\u003c/em\u003e\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMean Accuracy\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMean Precision\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMean Recall\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eHyperparameter\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLogistic regression\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e75.6%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e57.2%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e75.6%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e{'C': 1.57, 'multi_class': 'ovr', 'penalty': 'l2', 'solver': 'newton-cg'}\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRandom Forest\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e75.9%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e72.3%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e75.9%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e{'n_estimators': 833, 'min_samples_split': 10, 'min_samples_leaf': 4, 'max_features': 'sqrt', 'max_depth': 80, 'criterion': 'gini', 'bootstrap': False}\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGradient Boosting\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e74.8%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e67.4%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e74.8%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e{'n_estimators': 50, 'max_depth': 3, 'loss': 'log_loss', 'learning_rate': 0.178, 'criterion': 'squared_error'}\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eIn general, the results show that even without medication and clinical features, the model can predict with similar accuracy, precision and recall, thus demonstrating that demographic and socioeconomic features represent a strong correlation towards predicting epilepsy treatment outcome.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section4\"\u003e \u003ch2\u003e3.2.3.2 Approach Two - Analysis applying SMOTE\u003c/h2\u003e \u003cp\u003eIn the original dataset, the count of the target class labels was as follows: Not-seizure-free \u0026ndash; 1217 and seizure-free \u0026ndash; 392. Since there is a clear class imbalance, SMOTE technique was used to oversample the data and create a balanced dataset of 1217 instances of each class label. Hyperparameter tuning was done on the new dataset again, and then the 5-fold stratified cross-validation was applied.\u003c/p\u003e \u003cp\u003eThe results when SMOTE was applied can be seen in Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e. The performance of Logistic Regression fell to around 66%, whereas Random Forest and Gradient Boosting showed marked improvement in the overall performance, with both accuracies, recall and precision hovering above 80%.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003e\u003cem\u003eModel performance for approach two with smote and demographic predictor variables only\u003c/em\u003e\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMean Accuracy\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMean Precision\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMean Recall\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eHyperparameter\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLogistic regression\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e66.0%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e67.0%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e66.0%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e{'C': 0.065, 'multi_class': 'ovr', 'solver': 'sag'}\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRandom Forest\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e84.0%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e84.0%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e84.0%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e{'n_estimators': 777, 'min_samples_split': 2, 'min_samples_leaf': 1, 'max_features': 'sqrt', 'max_depth': 60, 'criterion': 'gini', 'bootstrap': True}\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGradient Boosting\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e83.0%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e83.0%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e82.0%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e{'n_estimators': 100, 'max_depth': 5, 'loss': 'log_loss', 'learning_rate': 0.094, 'criterion': 'squared_error'}\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eAfter applying SMOTE to address class imbalance, both the Random Forest and Gradient Boosting classifiers achieved over 83% accuracy, precision, and recall, substantially outperforming Logistic Regression. These findings highlight the strong predictive value of non-clinical features, including age, ethnicity, gender, and multiple domains of the Index of Multiple Deprivation (IMD), which showed high feature importance. As the dataset was imbalanced, we evaluated model robustness using precision\u0026ndash;recall (PR) curves, as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e and Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e. The PR curves showed that before applying SMOTE, the average area under the curve (AUC) was only around 30%, indicating that the models could not simultaneously achieve high precision and recall. After applying SMOTE, the PR curves showed marked improvement, with Gradient Boosting and Random Forest achieving average AUCs of about 90% and Logistic Regression improving to around 55%. This demonstrates that class imbalance was a major constraint and that balancing the data substantially improved predictive performance.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"4. Discussion","content":"\u003cp\u003eThis study demonstrates that demographic and socioeconomic factors, even with minimal clinical information, can meaningfully predict epilepsy treatment outcomes, with machine learning (ML) models achieving performance comparable to those incorporating clinical variables. After addressing class imbalance with SMOTE, both Random Forest and Gradient Boosting achieved accuracies, precision, and recall above 83%, substantially outperforming Logistic Regression. These results highlight the strong predictive value of non-clinical features\u0026mdash;particularly age, gender, ethnicity, and multiple domains of the Index of Multiple Deprivation (IMD)\u0026mdash;as determinants of treatment response. They also point out that the likelihood of achieving seizure freedom can be substantially influenced by these indicators alone. Our findings align with growing evidence that socioeconomic disadvantage is closely linked to poorer epilepsy outcomes. Prior studies have shown that patients from deprived areas experience delays in diagnosis, lower adherence to medication, and worse seizure control [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e, \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e, \u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e]. However, such factors are rarely incorporated into predictive modelling frameworks, which typically emphasise clinical or neurophysiological data such as seizure type, aetiology, or EEG features [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e, \u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e]. By demonstrating that demographic and socioeconomic data alone can achieve comparable predictive performance, this study provides novel evidence that these social determinants are not merely contextual but represent key risk indicators in their own right. There are several plausible mechanisms underpinning these results. Socioeconomic disadvantage can affect epilepsy outcomes through multiple, interacting pathways that are structural, behavioural, and biological. Structurally, people living in deprived areas are less likely to access timely specialist review and more likely to experience diagnostic delay or misdiagnosis\u0026mdash;patterns that are amplified in some ethnically minoritised groups [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e, \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e, \u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e]. Behaviourally, financial strain, unstable housing, caregiving burden, and transport constraints reduce continuity of follow-up and undermine adherence to antiseizure medication [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e, \u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e]. Biologically, chronic psychosocial stress associated with poverty and insecurity can increase seizure susceptibility via neuroendocrine and inflammatory pathways, compounding clinical risk [\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e, \u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e]. In our cohort, younger adults from the most deprived quintiles were disproportionately represented among those who were not seizure-free, consistent with a cumulative disadvantage model in which early-life and young-adult exposures erode treatment responsiveness over time. These observations accord with epidemiological evidence linking lower socioeconomic position with higher morbidity and mortality in epilepsy [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e] and reinforce the need to address wider determinants of health in routine care. Taken together, these findings support integrating demographic and socioeconomic profiling into early risk-stratification to identify patients who may benefit from targeted adherence support, social prescribing, proactive mental-health input, or prioritised referral to specialist services. Deprivation-aware decision support could also guide resource allocation, enabling teams to anticipate complexity and tailor follow-up intensity where need is greatest. This is squarely aligned with the NHS Long Term Plan\u0026rsquo;s commitment to personalised, inequality-reducing care [\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e] and with recent equity frameworks in epilepsy that call for the routine consideration of social context in clinical pathways [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e, \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e]. Importantly, incorporating demographic and socioeconomic predictors into AI-driven models is not only a route to better predictive accuracy but also a mechanism for advancing equity. The broader health-AI literature shows that algorithms trained without attention to social and demographic context can inadvertently amplify disparities [\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e, \u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e]. In epilepsy, equity-aware modelling implies (i) including robust measures of deprivation and other social determinants; (ii) testing interactions (intersectionality) between demographic and socioeconomic factors; (iii) reporting subgroup performance, calibration, and error patterns; and (iv) externally validating models across diverse settings [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e, \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e]. Embedding these practices helps ensure that decision-support tools mitigate, rather than entrench, structural inequities\u0026mdash;an essential step for person-centred epilepsy care that recognises the interplay of medical, social, and demographic determinants. Framed within a person-centred model, our results argue for care pathways that combine clinical profiling with explicit assessment of social needs, and for predictive tools that surface actionable risks (e.g., likely adherence challenges or access barriers) to inform shared decision-making. By \u0026ldquo;learning from the margins\u0026rdquo;\u0026mdash;i.e., allowing the experiences of the most disadvantaged to shape model design and service responses\u0026mdash;clinics can deliver more precise, compassionate, and fair epilepsy care.\u003c/p\u003e"},{"header":"5. Limitations","content":"\u003cp\u003eThis was a single-centre study using retrospective data, which may limit generalisability to other populations. The dataset was relatively small and initially imbalanced, and although SMOTE improved performance, oversampling may inflate accuracy by generating synthetic data points. Additionally, although SMOTE improved model performance, oversampling techniques may still inflate apparent accuracy by reducing class heterogeneity, and real-world performance may be lower when applied to unseen clinical populations. The analyses performed were also cross-sectional, preventing assessment of temporal dynamics or causal relationships. Furthermore, some demographic and clinical fields were incomplete or prone to data entry errors, which could introduce bias. External validation on larger, multicentre, and ethnically diverse datasets will be crucial for confirming the robustness of these models. We also acknowledge that seizure freedom is a clinically nuanced outcome and may vary with follow-up duration and reporting accuracy; however, the definition used reflects real-world decision-making in specialist epilepsy clinics.\u003c/p\u003e"},{"header":"6. Conclusion and Future Directions","content":"\u003cp\u003eThis study provides evidence that demographic and socioeconomic determinants, when analysed through machine learning frameworks, can serve as meaningful predictors of epilepsy treatment outcomes. Although performance was constrained by class imbalance in the original dataset, the application of SMOTE markedly enhanced model robustness, with Gradient Boosting and Random Forest achieving precision\u0026ndash;recall performance comparable to models incorporating clinical features. These findings support the integration of social determinants into routine risk stratification for epilepsy care, contributing to more equitable treatment planning. Further research with larger and more diverse cohorts is warranted to validate generalisability and to explore the integration of social determinants into clinical decision-support systems.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cdiv class=\"DefinitionList\"\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eAI\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eArtificial Intelligence\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eASM\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eAntiseizure Medication\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eAUC\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eArea Under the Curve\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eChi\u0026sup2;\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eChi\u0026ndash;square test\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eEDA\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eExploratory Data Analysis\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eEEG\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eElectroencephalography\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eIMD\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eIndex of Multiple Deprivation\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eIDACI\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eIncome Deprivation Affecting Children Index\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eIDAOPI\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eIncome Deprivation Affecting Older People Index\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eILAE\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eInternational League Against Epilepsy\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eLLM\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eLarge Language Model\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eML\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eMachine Learning\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eMRI\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eMagnetic Resonance Imaging\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eNHS\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eNational Health Service\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eNICE\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eNational Institute for Health and Care Excellence\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eONS\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eOffice for National Statistics\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003ePR\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003ePrecision\u0026ndash;Recall\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003ePTSD\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003ePost\u0026ndash;Traumatic Stress Disorder\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eR\u0026amp;D\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eResearch and Development\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eSMOTE\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eSynthetic Minority Oversampling Technique\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eUK\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eUnited Kingdom\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eWHO\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eWorld Health Organisation\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics Approval and Consent to Participate\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study was approved by the Research and Development Department of Salford Royal Hospital NHS Foundation Trust as a registered Health Improvement Project (reference: 22HIP46). All study procedures were conducted in accordance with applicable institutional policies and UK data protection regulations, and in compliance with the ethical principles of the Declaration of Helsinki.\u003c/p\u003e\n\u003cp\u003eThe study involved retrospective analysis of routinely collected clinical and demographic data. All data were anonymised prior to access and analysis, and no identifiable patient information was available to the research team. In accordance with NHS Health Research Authority guidance, individual informed consent was not required, as the study used anonymised routinely collected data and involved no direct patient contact or additional risk to participants.\u003c/p\u003e\n\n\u003cp\u003e\u003cstrong\u003eConsent for Publication\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003eThis manuscript does not contain any individual person\u0026rsquo;s identifiable data in any form (including individual details, images, or videos). All data analysed in this study were fully anonymised prior to access and analysis, and no identifiable information is presented in the manuscript.\u003c/p\u003e\n\n\u003cp\u003e\u003cstrong\u003eAvailability of Data and Materials\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe datasets generated and/or analysed during the current study are not publicly available due to institutional governance restrictions and UK data protection regulations, as they contain sensitive clinical information derived from NHS patient records.\u003c/p\u003e\n\u003cp\u003eFully anonymised data may be made available from the corresponding author on reasonable request, subject to approval by Salford Royal Hospital NHS Foundation Trust and relevant data governance procedures. Requests will be considered in line with applicable ethical approvals, data sharing agreements, and legal requirements\u003cstrong\u003e.\u003c/strong\u003e\u003c/p\u003e\n\n\u003cp\u003e\u003cstrong\u003eCompeting Interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.\u003c/p\u003e\n\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis work was supported by the Alan Turing Institute \u0026ndash; University of Manchester Sandpit Project Grant, 2022.\u003c/p\u003e\n\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026rsquo; Contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eMSM, LYHL, YL and DDB conceived and designed the study. NC curated and collected the data. MSM, LYHL, YL, and NC performed the data analysis. SS, MS, and AW contributed to the analyses and ensured the correct application of methods. RM and DDB provided clinical insights and guided the selection and interpretation of features for analysis. All authors reviewed and approved the final manuscript.\u003c/p\u003e\n\n\u003cp\u003e\u003cstrong\u003eAcknowledgments\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors would like to thank the clinical and administrative staff at Salford Royal Hospital NHS Foundation Trust for their support in facilitating data access and governance approvals. We are particularly grateful to the members of the specialist epilepsy service for their assistance during data extraction and clarification of clinical records. We also acknowledge the patients whose routinely collected clinical data made this research possible. This work was supported by the Turing\u0026ndash;Manchester Sandpit Project Grant (2022), and the authors thank the organisers and participants of the sandpit for valuable interdisciplinary discussions that informed the development of this study.\u003c/p\u003e\n\n\n"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eFeigin VL, Vos T, Nair BS, Hay SI, Abate YH, Abd Al Magied AHA, et al. Global, regional, and national burden of epilepsy, 1990\u0026ndash;2021: a systematic analysis for the Global Burden of Disease Study 2021. Lancet Public Health. 2025;10:e203\u0026ndash;27. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/S2468-2667(24)00302-5\u003c/span\u003e\u003cspan address=\"10.1016/S2468-2667(24)00302-5\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEpilepsy Society. What is epilepsy? 2024. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://epilepsysociety.org.uk/about-epilepsy/what-epilepsy\u003c/span\u003e\u003cspan address=\"https://epilepsysociety.org.uk/about-epilepsy/what-epilepsy\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWorld Health Organization. Epilepsy. Key facts. 2024. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.who.int/news-room/fact-sheets/detail/epilepsy\u003c/span\u003e\u003cspan address=\"https://www.who.int/news-room/fact-sheets/detail/epilepsy\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDalic L, Cook M. Managing drug-resistant epilepsy: challenges and solutions. Neuropsychiatr Dis Treat. 2016;12:2605\u0026ndash;16. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.2147/ndt.s84852\u003c/span\u003e\u003cspan address=\"10.2147/ndt.s84852\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMesraoua B, Deleu D, Kullmann DM, Shetty AK, Boon P, Perucca E, et al. Novel therapies for epilepsy in the pipeline. Epilepsy Behav. 2019;97:282\u0026ndash;90. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.yebeh.2019.04.042\u003c/span\u003e\u003cspan address=\"10.1016/j.yebeh.2019.04.042\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKarakis I, Janocko NJ, Morton ML, Groover O, Teagarden DL, Villarreal HK, et al. Stigma in psychogenic nonepileptic seizures. Epilepsy Behav. 2020;111. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.yebeh.2020.107269\u003c/span\u003e\u003cspan address=\"10.1016/j.yebeh.2020.107269\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eStrzelczyk A, Aledo-Serrano A, Coppola A, Didelot A, Bates E, Sainz-Fuertes R, et al. The impact of epilepsy on quality of life: Findings from a European survey. Epilepsy Behav. 2023;142:109179. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.yebeh.2023.109179\u003c/span\u003e\u003cspan address=\"10.1016/j.yebeh.2023.109179\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYeni K. Stigma and psychosocial problems in patients with epilepsy. Explor Neurosci. 2023;2:251\u0026ndash;63. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.37349/en.2023.00026\u003c/span\u003e\u003cspan address=\"10.37349/en.2023.00026\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNational Institute for Health and Care Excellence. Epilepsies in children, young people and adults. 2025.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHussain SA, Ortendahl JD, Bentley TGK, Harmon AL, Gupta S, Begley CE, et al. The economic burden of caregiving in epilepsy: An estimate based on a survey of US caregivers. Epilepsia. 2020;61:319\u0026ndash;29. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1111/epi.16429\u003c/span\u003e\u003cspan address=\"10.1111/epi.16429\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMaas L, Kellenaers J, van Mastrigt G, van Kuijk SMJ, Vlooswijk MCG, Hiligsmann M, et al. Societal costs and quality of life analysis in patients undergoing resective epilepsy surgery: A one-year follow-up. Epilepsy Behav Rep. 2023;24:100635. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.ebr.2023.100635\u003c/span\u003e\u003cspan address=\"10.1016/j.ebr.2023.100635\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMarquina C, Foster E, Chen Z, Vaughan DN, Abbott DF, Tailby C, et al. Work productivity, quality of life, and care needs: An unfolding epilepsy burden revealed in the Australian Epilepsy Project pilot study. Epilepsia Open. 2024;9:739\u0026ndash;49. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1002/epi4.12919\u003c/span\u003e\u003cspan address=\"10.1002/epi4.12919\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSarac E, Yildiz E. The effect of epilepsy self-management on productivity at work. Epilepsy Behav. 2024;157:109839. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.yebeh.2024.109839\u003c/span\u003e\u003cspan address=\"10.1016/j.yebeh.2024.109839\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBush KJ, Cullen E, Mills S, Chin RFM, Thomas RH, Kingston A, et al. Assessing the extent and determinants of socioeconomic inequalities in epilepsy in the UK: a systematic review and meta-analysis of evidence. Lancet Public Health. 2024;9:e614\u0026ndash;28. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/s2468-2667(24)00132-4\u003c/span\u003e\u003cspan address=\"10.1016/s2468-2667(24)00132-4\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCorrea DJ, Gutierrez CA. Health Disparities and Inequities in Epilepsy. Achieving Equity in Neurological Practice: Principles and Pathways. Springer; 2024. pp. 91\u0026ndash;123.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHayanga B, Stafford M, B\u0026eacute;cares L. Ethnic inequalities in multiple long-term health conditions in the United Kingdom: a systematic review and narrative synthesis. BMC Public Health. 2023;23. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/s12889-022-14940-w\u003c/span\u003e\u003cspan address=\"10.1186/s12889-022-14940-w\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMaloney EM, Corcoran P, Costello DJ, O\u0026rsquo;Reilly \u0026Eacute;J. Association between social deprivation and incidence of first seizures and epilepsy: A prospective population-based cohort. Epilepsia. 2022;63:2108\u0026ndash;19. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1111/epi.17313\u003c/span\u003e\u003cspan address=\"10.1111/epi.17313\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHealth Equity Evidence Centre. Understanding the Index of Multiple Deprivation (IMD) in public health research. 2024.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHakeem H, Feng W, Chen Z, Choong J, Brodie MJ, Fong S-L, et al. Development and validation of a deep learning model for predicting treatment response in patients with newly diagnosed epilepsy. JAMA Neurol. 2022;79:986\u0026ndash;96.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePerucca E. The pharmacological treatment of epilepsy: recent advances and future perspectives. Acta Epileptol. 2021;3:22.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePiccenna L, O\u0026rsquo;Dwyer R, Leppik I, Beghi E, Giussani G, Costa C, et al. Management of epilepsy in older adults: A critical review by the ILAE Task Force on Epilepsy in the elderly. Epilepsia. 2023;64:567\u0026ndash;85. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1111/epi.17426\u003c/span\u003e\u003cspan address=\"10.1111/epi.17426\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRatcliffe C, Pradeep V, Marson A, Keller SS, Bonnett LJ. Clinical prediction models for treatment outcomes in newly diagnosed epilepsy: A systematic review. Epilepsia. 2024;65:1811\u0026ndash;46.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFiest KM, Sauro KM, Wiebe S, Patten SB, Kwon C-S, Dykeman J, et al. Prevalence and incidence of epilepsy. Neurology. 2017;88:296\u0026ndash;303. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1212/wnl.0000000000003509\u003c/span\u003e\u003cspan address=\"10.1212/wnl.0000000000003509\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKharkar S, Pillai J, Rochestie D, Haneef Z. Socio-Demographic Influences on Epilepsy Outcomes in an Inner-City Population. Seizure. 2014;23. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.seizure.2014.01.002\u003c/span\u003e\u003cspan address=\"10.1016/j.seizure.2014.01.002\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDepartment of Health and Social Care. Independent Investigation of the National Health Service in England (accessible version). 2024.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBensken WP, Alberti PM, Khan OI, Williams SM, Stange KC, Vaca GF-B, et al. A framework for health equity in people living with epilepsy. Epilepsy Res. 2022;188:107038. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.eplepsyres.2022.107038\u003c/span\u003e\u003cspan address=\"10.1016/j.eplepsyres.2022.107038\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKiriakopoulos ET, Garc\u0026iacute;a Sosa R, Blank L, Johnson EL, Gutierrez C. Shining a Light: Advancing Health Equity in Overlooked Epilepsy Communities. Epilepsy Curr. 2024. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1177/15357597241258081\u003c/span\u003e\u003cspan address=\"10.1177/15357597241258081\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAlessi N, Perucca P, McIntosh AM. Missed, mistaken, stalled: identifying components of delay to diagnosis in epilepsy. Epilepsia. 2021;62:1494\u0026ndash;504.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePellinen J. Treatment gaps in epilepsy. Front Epidemiol. 2022. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3389/fepid.2022.976039\u003c/span\u003e\u003cspan address=\"10.3389/fepid.2022.976039\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. 2.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePerzynski AT, Ramsey RK, Col\u0026oacute;n-Zimmermann K, Cage J, Welter E, Sajatovic M. Barriers and facilitators to epilepsy self-management for patients with physical and psychological co-morbidity. Chronic Illn. 2017;13:188\u0026ndash;203. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1177/1742395316674540\u003c/span\u003e\u003cspan address=\"10.1177/1742395316674540\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJennum P, Gyllenborg J, Kjellberg J. The social and economic consequences of epilepsy: A controlled national study. Epilepsia. 2011;52:949\u0026ndash;56. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1111/j.1528-1167.2010.02946.x\u003c/span\u003e\u003cspan address=\"10.1111/j.1528-1167.2010.02946.x\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMula M, Kanner AM, Jett\u0026eacute; N, Sander JW. Psychiatric comorbidities in people with epilepsy. Neurol Clin Pract. 2021;11:e112\u0026ndash;20.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTsigebrhan R, Derese A, Kariuki SM, Fekadu A, Medhin G, Newton CR, et al. Co-morbid mental health conditions in people with epilepsy and association with quality of life in low-and middle-income countries: a systematic review and meta-analysis. Health Qual Life Outcomes. 2023;21:5.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMarbin D, Gutwinski S, Schreiter S, Heinz A. Perspectives in poverty and mental health. Front Public Health. 2022;10:975482.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSmith MV, Mazure CM. Mental health and wealth: depression, gender, poverty, and parenting. Annu Rev Clin Psychol. 2021;17:181\u0026ndash;205.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLolk K, Werenberg Dreier J, Christensen J. Individual and neighborhood-level socioeconomic deprivation and risk of epilepsy after traumatic brain Injury: A register-based cohort study. Epilepsy Behav. 2024;156:109807. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.yebeh.2024.109807\u003c/span\u003e\u003cspan address=\"10.1016/j.yebeh.2024.109807\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePickrell WO, Lacey AS, Bodger OG, Demmler JC, Thomas RH, Lyons RA, et al. Epilepsy and deprivation, a data linkage study. Epilepsia. 2015;56:585\u0026ndash;91. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1111/epi.12942\u003c/span\u003e\u003cspan address=\"10.1111/epi.12942\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSubasi A, Kevric J, Abdullah Canbaz M. Epileptic seizure detection using hybrid machine learning methods. Neural Comput Appl. 2019;31:317\u0026ndash;25. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s00521-017-3003-y\u003c/span\u003e\u003cspan address=\"10.1007/s00521-017-3003-y\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTran LV, Tran HM, Le TM, Huynh TTM, Tran HT, Dao SVT. Application of Machine Learning in Epileptic Seizure Detection. Diagnostics. 2022;12:2879. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3390/diagnostics12112879\u003c/span\u003e\u003cspan address=\"10.3390/diagnostics12112879\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChakrabarti S, Swetapadma A, Ranjan A, Pattnaik PK. Time domain implementation of pediatric epileptic seizure detection system for enhancing the performance of detection and easy monitoring of pediatric patients. Biomed Signal Process Control. 2020;59:101930. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.bspc.2020.101930\u003c/span\u003e\u003cspan address=\"10.1016/j.bspc.2020.101930\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTorlay L, Perrone-Bertolotti M, Thomas E, Baciu M. Machine learning\u0026ndash;XGBoost analysis of language networks to classify patients with epilepsy. Brain Inf. 2017;4:159\u0026ndash;69. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s40708-017-0065-7\u003c/span\u003e\u003cspan address=\"10.1007/s40708-017-0065-7\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHan K, Liu C, Friedman D. Artificial intelligence/machine learning for epilepsy and seizure diagnosis. Epilepsy Behav. 2024;155:109736. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.yebeh.2024.109736\u003c/span\u003e\u003cspan address=\"10.1016/j.yebeh.2024.109736\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSmolyansky ED, Hakeem H, Ge Z, Chen Z, Kwan P. Machine learning models for decision support in epilepsy management: A critical review. Epilepsy Behav. 2021;123:108273. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.yebeh.2021.108273\u003c/span\u003e\u003cspan address=\"10.1016/j.yebeh.2021.108273\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXie K, Ojemann WKS, Gallagher RS, Shinohara RT, Lucas A, Hill CE, et al. Disparities in seizure outcomes revealed by large language models. J Am Med Inf Assoc. 2024;31:1348\u0026ndash;55. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1093/jamia/ocae047\u003c/span\u003e\u003cspan address=\"10.1093/jamia/ocae047\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAbdaltawab A, Chang L-C, Mansour M, Koubeissi M. How accurate are machine learning models in predicting antiseizure medication responses: A systematic review. Epilepsy Behav. 2025;163. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.yebeh.2024.110212\u003c/span\u003e\u003cspan address=\"10.1016/j.yebeh.2024.110212\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNHS England. NHS Long Term Plan. 2019.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eObermeyer Z, Powers B, Vogeli C, Mullainathan S. Dissecting racial bias in an algorithm used to manage the health of populations. Science. 2019;366:447\u0026ndash;53. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1126/science.aax2342\u003c/span\u003e\u003cspan address=\"10.1126/science.aax2342\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"},{"header":"Table 1","content":"\u003cp\u003eTable 1 is available in the Supplementary Files section.\u003c/p\u003e\n"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"bmc-neurology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"nurl","sideBox":"Learn more about [BMC Neurology](http://bmcneurol.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/nurl","title":"BMC Neurology","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Epilepsy, demographics, health equity, machine learning, prediction, neurology, IMD, socioeconomic factors, Logistic Regression, Random Forest, Gradient Boosting, SMOTE","lastPublishedDoi":"10.21203/rs.3.rs-8917468/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8917468/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e \u003cp\u003eEpilepsy is one of the most common serious neurological conditions globally and is associated with substantial morbidity, reduced quality of life, and increased socioeconomic disadvantage. Although clinical predictors of seizure outcomes have been widely studied, demographic and socioeconomic determinants of treatment response remain underrepresented in predictive modelling. Greater understanding of how these factors influence outcomes may inform earlier identification of patients at risk of poor seizure control and contribute to strategies aimed at reducing health inequalities.\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e \u003cp\u003eA retrospective observational study was undertaken using routinely collected clinical, demographic, and socioeconomic data from 1,609 adults with epilepsy attending a specialist epilepsy service in England between 2020 and 2024. Demographic variables included age, sex, and ethnicity. Socioeconomic position was captured using area-level measures from the Index of Multiple Deprivation, including income, employment, education, health, housing, and living environment domains. Limited high-level clinical descriptors were included, without medication-specific information. Logistic Regression, Random Forest, and Gradient Boosting models were used to predict seizure-free status at the most recent clinical review. To address class imbalance, the Synthetic Minority Oversampling Technique (SMOTE) was applied within training folds. Model performance was assessed using stratified cross-validation and evaluated using accuracy, precision, recall, and precision\u0026ndash;recall area under the curve.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eDemographic and socioeconomic variables showed consistent associations with epilepsy treatment outcomes and contributed substantially to model performance. After accounting for class imbalance, Random Forest and Gradient Boosting models achieved accuracy, precision, and recall exceeding 83%. Individuals residing in more socioeconomically deprived areas were overrepresented among those who were not seizure-free, indicating persistent social gradients in epilepsy outcomes. Comparable predictive performance was observed when models were trained using predominantly non-clinical features, suggesting that social determinants carry independent prognostic value.\u003c/p\u003e\u003ch2\u003eConclusions\u003c/h2\u003e \u003cp\u003eDemographic and socioeconomic factors are important predictors of epilepsy treatment outcomes and can be leveraged within machine learning frameworks to support population-level risk stratification. Incorporating social determinants of health into predictive approaches may help identify patients at increased risk of poor outcomes and inform more equitable, context-aware epilepsy care pathways.\u003c/p\u003e","manuscriptTitle":"Learning from the Margins: Using Machine Learning Models to Predict Epilepsy Outcomes from Socioeconomic Factors","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-03-26 17:15:41","doi":"10.21203/rs.3.rs-8917468/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"reviewersInvited","content":"","date":"2026-03-24T22:58:39+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2026-02-27T14:26:18+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-02-26T01:29:20+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-02-26T01:28:38+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Neurology","date":"2026-02-19T11:56:55+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"bmc-neurology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"nurl","sideBox":"Learn more about [BMC Neurology](http://bmcneurol.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/nurl","title":"BMC Neurology","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"87a7bf9a-8799-431f-95d8-4fc5ea4f3f18","owner":[],"postedDate":"March 26th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-03-26T17:15:41+00:00","versionOfRecord":[],"versionCreatedAt":"2026-03-26 17:15:41","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8917468","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8917468","identity":"rs-8917468","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-29T02:00:03.542394+00:00
License: CC-BY-4.0