Predicting Multiple Sclerosis Outcomes: A Machine Learning Approach Integrating Patient-Reported Outcomes and Clinical Data | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Predicting Multiple Sclerosis Outcomes: A Machine Learning Approach Integrating Patient-Reported Outcomes and Clinical Data Minerva VIGUERA MORENO, María Eugenia MARZO SOLA, Fernando MARTIN-SANCHEZ, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6634220/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 12 You are reading this latest preprint version Abstract Multiple sclerosis (MS) is a chronic, immune-mediated disease with variable progression that complicates clinical management( 1 ). Traditional assessments, such as the Expanded Disability Status Scale (EDSS), have limitations, prompting the integration of patient-reported outcome measures (PROMs) alongside clinician-reported outcomes (CROs) to capture a comprehensive view of disease impact( 2 ). This study investigates the use of machine learning (ML) techniques to predict three key clinical outcomes in MS: EDSS level, spasticity severity, and the development of new lesions detected through magnetic resonance. Data were collected from 240 MS patients over 18 months, including baseline demographics, CROs, and PROMs. The ML pipeline involved feature encoding, data splitting (80:20), oversampling for underrepresented classes, and hyperparameter tuning via grid search and cross-validation. Regression models (evaluated with MSE and R²) and classification models were trained depending on predicted variable characteristics. Results indicate that models incorporating both PROMs and CROs achieved superior performance in predicting EDSS and spasticity, while ML models reliably identified new lesions with 100% true positive rate. PROMs collected before clinical assessment combined with basal demographics also lead to predictive models with acceptable performance. These findings support the potential of integrating PROMs into clinical decision-making and telemedicine for improved MS management. Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Figure 9 Introduction Multiple sclerosis (MS) is a chronic, immune-mediated disease of the central nervous system with highly variable progression and clinical presentation. This variability complicates clinical management and long-term treatment planning. MS patients may experience motor and sensory deficits, cognitive impairments, and exacerbations, leading to increased disability over time( 3 ). The Expanded Disability Status Scale (EDSS) is the most used tool to assess neurological impairment, but it has limitations such as inter-rater variability and an overemphasis on ambulation( 4 ). Other important monitoring aspects include spasticity levels, new lesions detected by RMI, and exacerbations. Spasticity can greatly affect mobility and quality of life (QoL), while new lesions, typically detected via Magnetic Resonance Imaging (MRI) indicate active disease, and exacerbations mark acute neurological deterioration( 5 , 6 ). Clinical tests like the 25-foot walk, 9-hole peg test, and symbol digit modalities (SDM) test provide objective measures of physical and cognitive function but may not fully capture the patient’s experience( 5 ). In contrast, PROMs such as the MSIS-29, Fatigue Severity Scale (FSS), Health Assessment Questionnaire (HAQ), and Neuro-QoL instruments offer valuable insights into health and QoL from the patient's perspective. Integrating PROMs with CROs enhances patient-centered care, telemedicine, and individualized treatment decisions( 2 ). This study explores the use of machine learning (ML) techniques to predict key clinical outcomes in MS – EDSS level, spasticity severity, and newly developed RMI lesions by analyzing data from 240 patients monitored over 18 months. Previous studies have explored the possibility of predicting MS progression using ML or neural networks( 7 ); however, the utility of patient-reported outcome measures (PROMs) as predictive variables remains unclear( 8 – 10 ). We compare the predictive power of ML models trained on various combinations of baseline information, CROs, and PROMs, aiming to assess the utility of PROMs in clinical decision-making and to identify optimal features and algorithms for short-term disease progression prediction. It is important to emphasize that the PROMs in this study were collected prior to the determination of the study outcomes, ensuring that the predictive analysis relies solely on preliminary information. Methodology 3.1 Study Design This study follows a prospective cohort design involving 240 patients with MS treated at a second-level hospital in Spain. Data collection was conducted over 18 months, capturing baseline clinical characteristics, routine clinical assessments, and patient-reported outcomes at regular intervals. The dataset comprises 206 complete cases that include baseline data, consultation information, and PROMs collected up to 180 days prior to each consultation (median: 73 days)( 11 ). 3.2 Data Collection REDCap projects were designed to collect research information( 12 ). Baseline data were collected at the study's initiation, capturing both demographic (age, gender) and clinical information (time since diagnosis, RMI lesions at diagnosis, previous exacerbations and evolution form). Data from outpatient consultations were collected biannually and included EDSS level, spasticity severity, new MRI lesions and new exacerbations in the last semester. Routine clinical tests, such as the 25-foot walk test, 9-hole peg test, and SDM, were also performed. PROMs were completed by patients every three months and included assessments of physical and psychological impact (MSIS-29), fatigue (FSS), functional ability (HAQ), and cognitive and depressive symptoms (Neuro-QoL instruments)( 11 ). 3.3. Dataset Preparation The collected data underwent preprocessing to ensure quality and consistency. Raw data were validated and cleaned to address missing or inconsistent entries punctual missing data was imputed using KNN. PROMs, CROs, and baseline information were integrated into a unified dataset, without any identifiable information to comply with ethical and legal standards. To assess the utility of different data types, this dataset was used to prepare three subsets for algorithms training proposals: PROMs dataset included PROMs and baseline information. CROs dataset was composed by clinical tests, regular consultation data, and baseline information. All-data combined both. 3.4 Machine Learning Workflow The objective was to train and compare ML models and data mining techniques to predict three key outcomes: EDSS level, spasticity severity, and new RMI lesions. The ML pipeline included codification, splitting the data for training and testing, oversampling to balance underrepresented categories, hyperparameter tuning via grid search and cross-validation, and model evaluation. Regression models were assessed using mean squared error (MSE) and R-squared, while classification models were evaluated using accuracy, precision, recall, F1-score, and ROC-AUC, focusing on the clinical significance of these parameters. 3.4.1 EDSS Inference Inferring EDSS levels posed a methodological challenge. Previous studies successfully classified MS patients by functional level but typically used stratified categories (e.g., three levels), limiting clinical utility due to significant functional differences within each stratum( 9 ). This study approaches EDSS inference as a regression problem, considering its 0–10 scale with 0.5 increments. A preliminary analysis revealed an uneven distribution of EDSS levels, with an underrepresentation of higher scores (above 8). This imbalance negatively impacted model performance, biasing results toward the most frequent values. To address this, underrepresented categories were manually oversampled by generating synthetic samples while introducing random noise to reduce overfitting( 13 ). Normalization was applied before training Support Vector Machine (SVM) and Ridge Regression. Four regression algorithms were trained and optimized through grid search hyperparameter selection, including Ridge Regression, SVM, Random Forest (RF) Regressor, and Gradient Boosting (GB). 3.4.2 Spasticity level estimation Spasticity severity was measured using the Ashworth Scale, which quantifies muscle stiffness and requires in-person evaluation. Predicting spasticity severity using ML could enhance telemedicine capabilities and support clinical decision-making before patient consultations. The Ashworth Scale classifies spasticity into five levels (0, 1, 1.5, 2, 3), leading to a classification approach using a RF Classifier. However, higher severity classes were underrepresented, affecting model performance. Three approaches were tested: no oversampling (RF adjusted for disbalanced classes), manual oversampling, and SMOTE( 13 ). Despite these adjustments, precision for classes 2 and 3 remained suboptimal, leading to their combination into a single category. 3.4.3 New RMI lesion prediction New lesions concurrence is a marker of MS progression, typically assessed via MRI performed annually. The high cost and limited availability of MRI necessitate alternative decision-support methods. This study framed lesion prediction as a classification task to determine patients unlikely to develop new lesions. The focus was minimizing false negatives (FN), as failing to detect a new lesion could lead to inappropriate clinical decisions. Due to a strong class imbalance (11 positive cases), ADASYN oversampling was applied before training classification algorithms( 13 ). RF (n_estimators = 10, max_depth = 10), SVM (Kernel = Sigmoid) y XGBClassifier (objective='binary:logistic') were tested. Results 4.1 EDSS Inference Hyperparameters selection using grid search allowed for the optimization of each algorithm. Table 1 summarizes the MSE and R² values for each model across the datasets. Table 1 MSE and R2 for each model and dataset PROMs CROs All_data MSE R 2 MSE R 2 MSE R 2 RR 0.87 0.78 1.23 0.68 0.81 0.80 RF 0.40 0.90 0.32 0.92 0.27 0.93 SMV 0.43 0.89 0.46 0.88 0.40 0.90 GB 0.36 0.91 0.30 0.92 0.39 0.90 Complementing algorithm performance, we analyzed features importance for each optimized model to better understand the algorithm’s decision-making. The output values were adjusted to the nearest value on the scale, and a residual plot was generated. 4.1.1 PROMs and GB ( learning_rate = 0.2, n_estimators = 100) Using PROMs in model training (see Fig. 1 ), revealed the features with the highest weight were the disease course and duration, previous relapses, HAQ and FSS scores (physical function and fatigue). Variables related to cognitive function and depression have intermediate importance, while age, gender, and initial neural lesions are the least important. The residual analysis, shown in Fig. 2 , revealed 5 cases in the test set were classified with an absolute error greater than 1 point (the acceptability threshold set in this study), corresponding to 94% of subjects being correctly categorized. Additionally, there is a tendency to underestimate the EDSS level (negative residuals), which is consistent with the overrepresentation of lower levels. Furthermore, a better fit is observed in the higher levels, suggesting that the generation of synthetic samples in these underrepresented categories may have led to overfitting at high values. 4.1.2 CROs and GBRegressor ( learning_rate = 0.2, n_estimators = 100). The analysis of features importance (Fig. 3 ) indicates that disease course and duration, along with parameters measuring the physical impact of the disease (both gross and fine motor skills), carry the most weight. Cognitive impairment has a moderate importance in this model. Other variables, such as age, gender, or initial neural lesions, appear to have less influence. Residuals analysis (See Fig. 4 ) shows that only 4 cases had an absolute error greater than 1 point, meaning that 95% of the predictions fall within the acceptable range. Additionally, the system tends to underestimate the EDSS level, particularly in the mid-range values, while fitting relatively well for higher values as described when using PROMs. 4.1.3All-data and RF. ( max_depth = 10, min_samples_leaf = 1, min_samples_split = 3, n_estimators = 100). The features importance, shown in Fig. 5 , is consistent with what was expected regarding their relationship with disease progression. Disease course along with 25-foots test, HAQ score (also reflecting physical function), time since diagnosis, 9-Holes, and fatigue (FSS), were the variables with the greatest weight. In contrast, age, gender, spinal lesions, and variables related to cognitive function appear to have less importance. Regarding the residual distribution (Fig. 6 ), it is observed that only 3 out of the 87 cases in the test set deviated more than 1 point from the actual value, which is the clinically acceptable threshold. Therefore, it can be stated that 96.5% of the predictions fall within the acceptable variability range, indicating clinical utility. 4.2 Spasticity level Random Forest Classifier was trained with previously determined hyperparameters (n_estimators = 100, max_depth = 10, min_samples_split = 2, min_samples_leaf = 1). Table 2 presents the overall accuracy achieved with different data combinations and imbalance-handling methods, both for original and grouped categories. The highest accuracies were obtained using all available data: 0.86 for the five original categories and 0.9 for grouped categories. Using only PROMs dataset resulted in slightly lower but comparable accuracies (0.81 and 0.90, respectively), while employing CROs yielded performance equivalent to using all variables. Table 2 Overall accuracy for RF prediction of Ashworth Scale Categories Oversampling Dataset PROs CROs All-Data Original None 0,79 0,76 0,81 Manual 0,81 0,79 0,83 SMOTE 0,81 0,83 0,86 Grouped ( 2 – 3 ) None 0,86 0,88 0,86 Manual 0,90 0,88 0,86 SMOTE 0,86 0,9 0,9 Since overall accuracy does not fully reflect the clinical meaning of classification errors, confusion matrices and other metrics were extracted for the best-performing models. The confusion matrices reveal poorer performance for higher spasticity values, regardless of the training data (Table 3 ); however, grouping categories 2 and 3 improved the classification of these higher values, as demonstrated by precision and recall shown in Table 4 . Table 3 Confusion Matrices for RF optimized classification Categories PROs (manual) CROs (SMOTE) All-Data (SMOTE) Original \(\:\left(\begin{array}{c}\begin{array}{ccccc}27&\:1&\:0&\:0&\:0\\\:2&\:3&\:0&\:0&\:0\\\:0&\:0&\:3&\:1&\:0\\\:0&\:1&\:0&\:1&\:1\\\:0&\:0&\:0&\:2&\:0\end{array}\end{array}\right)\) \(\:\left(\begin{array}{c}\begin{array}{ccccc}28&\:0&\:0&\:0&\:0\\\:2&\:3&\:0&\:0&\:0\\\:1&\:0&\:2&\:1&\:0\\\:0&\:0&\:0&\:2&\:1\\\:0&\:0&\:0&\:2&\:0\end{array}\end{array}\right)\) \(\:\left(\begin{array}{c}\begin{array}{ccccc}27&\:0&\:1&\:0&\:0\\\:1&\:4&\:0&\:0&\:0\\\:0&\:0&\:3&\:1&\:0\\\:0&\:0&\:0&\:2&\:1\\\:0&\:0&\:0&\:2&\:0\end{array}\end{array}\right)\) Grouped ( 2 – 3 ) \(\:\left(\begin{array}{c}\begin{array}{cccc}27&\:1&\:0&\:0\\\:1&\:4&\:0&\:0\\\:0&\:0&\:3&\:1\\\:0&\:1&\:0&\:4\end{array}\end{array}\right)\) \(\:\left(\begin{array}{c}\begin{array}{cccc}28&\:0&\:0&\:0\\\:2&\:3&\:0&\:0\\\:1&\:0&\:2&\:1\\\:0&\:0&\:0&\:5\end{array}\end{array}\right)\) \(\:\left(\begin{array}{c}\begin{array}{cccc}26&\:1&\:1&\:0\\\:1&\:4&\:0&\:0\\\:0&\:0&\:3&\:1\\\:0&\:0&\:0&\:5\end{array}\end{array}\right)\) Table 4 Performance metrics for each dataset with grouped categories PROs CROs ALL-data precision recall precision recall precision recall Cat 0 0,96 0,96 0,9 1 0,96 0,93 Cat 1 0,67 0,8 1 0,6 0.8 0,8 Cat 1.5 1 0,75 1 0,5 0,75 0,75 Cat 2–3 0.80 0,80 0,83 1 0,83 1 Accuracy 0,9 0,9 0,90 Macro_avg 0,86 0,83 0,93 0,78 0,84 0,87 Weighted_avg 0,91 0,90 0,92 0,90 0,91 0,90 Figures 7 , 8 , and 9 graphically illustrate the importance of the variables used in training across the different optimized models. Greater importance was found for physical-related features across all models, consistently with the impact of spasticity on musculature ( 6 ). 4.3 New RMI lesion prediction RF, SVM and XGB produced underwhelming results when trained using all available variables across different data combinations. Aligned with the dataset’s negative bias, all cases were classified as negative. However, when the variable set was narrowed down to those that had higher importance in initial models, satisfactory results were achieved by minimizing FN and correctly classifying true positives (TP). Consequently, all possible combinations of few variables were tested for each algorithm with hyperparameters previously selected. Overall, the results indicate that it is possible to avoid FN when classifying patients with new lesions—considered an unacceptable error. In all cases, the TP rate reached 100%, meaning that the models correctly identified all patients with new lesions. Moreover, using few input variables appears to be associated with a lower number of false positives (FP), potentially optimizing the classification system and reducing unnecessary MRIs in patients without new lesions. The variable combinations that led to this optimized classification and obtained performance are shown in Table 5. Table 5 Optimized model features and performance by training dataset PROs CROs All_data Algorithm XGB SVM RF features FSS Course RMI_brain_lesions Prev_exacerbations RMI_brain_lesions RMI_spinal_lesions Gender FSS Course Prev_exacerbations 25-foot test Conf_matrix \(\:\left(\begin{array}{cc}42&\:18\\\:0&\:2\end{array}\right)\) \(\:\left(\begin{array}{cc}36&\:24\\\:0&\:2\end{array}\right)\) \(\:\left(\begin{array}{cc}55&\:5\\\:0&\:2\end{array}\right)\) AUC_ROC 0.80 0.84 0.92 Specifically, for PROMs data, the best combination included FSS, disease course, and basal brain MRI lesions. For the CROs dataset, the optimal variables were a new relapse in the last 6 months, spinal and brain lesions at onset, and gender. When both data types were combined, the rate of FP decrease to 5, leading to the most optimized model in terms of false positive minimization while maintaining 0% of FN rate. Overall, these results demonstrate the potential of ML models in predicting EDSS levels, spasticity severity, and new lesions with clinically meaningful accuracy. Including PROMs—either alone or combined with CROs—improved predictive performance, reinforcing their clinical value. Discussion This study examined the relationship between various data types collected from MS patients and their usefulness for predicting EDSS levels, spasticity severity, and the potential occurrence of new lesions without relying on MRI images. For EDSS inference, the R² and MSE values across the three datasets (PROMs, CROs, and All-data) indicate that the models effectively capture data variability and can determine the EDSS score with an absolute error of less than 1 point in 90–95% of cases. Although the optimized model using only PROMs and baseline characteristics performed slightly worse than those including clinical measurements, the results were comparable. It is important to note that these PROMs were collected up to six months before the patient’s specialized consultation, which might partly explain some errors if deterioration occurred later. However, this finding supports the utility of PROMs in telemedicine contexts, where they can serve as predictive or monitoring variables with similar performance to clinical data. The best performance was achieved when combining CROs with PROMs, reinforcing the importance of including both in clinical practice to minimize deviations in EDSS categorization. The relative importance of variables was consistent across models, with disease course and time leading, followed by variables related to the physical impact of MS, aligned with the higher weigh of physical impact in EDSS( 4 ). When both CROs and PROMs were included, their variable importance was consistent with the type of impact they measure, further supporting the inclusion of PROMs in predictive models. Our approach has allowed for the inference of the EDSS level without the need for stratification and by minimizing the MSE, improving the accuracy of previous models and expanding their applicability This approach allows for a more precise estimation of the EDSS level than previous studies, eliminating the need for stratification and thereby enhancing its applicability in supporting clinical decisions( 14 ). Regarding spasticity level estimation, global accuracy metrics (approximately 0.90 for CROs and All-data, and 0.88 for PROMs) show that the models correctly classify about 90% of cases. Although using only PROMs slightly reduced precision, the overall performance remained similar. The stable weighted average (~ 0.90) across configurations suggests that while frequent classes are well classified, improvements for less frequent, higher spasticity categories may be possible with larger datasets and alternative balancing strategies. Notably, the model demonstrated excellent performance in identifying patients without spasticity (Category 0) with high recall and precision, though this might be partly due to the dataset bias toward lower spasticity values. For mild spasticity (Category 1) All-data set appears to be most useful in detecting without generating FP. Intermediate spasticity (Category 1.5) was suboptimal detected overall, though training with PROMs yielded promising accuracy. Finally, for high spasticity (grouped Categories 2–3), recall was excellent (100% in CROs and All-data) with variable precision. In terms of variable importance, clinical variables such as time since diagnosis and disease course consistently held high weight, while neurological lesions, prior relapses, and gender were less influential. The importance of PROMs and their clinical counterparts aligns with expectations regarding their impact on patients’ physical disability and QoL. Estimating the spasticity level can serve as support for telemedicine applications, potentially reducing the need for regular outpatient consultations and, consequently, the burden of care on patients. Finally, for predicting new MRI lesions, PROMs collected before consultation correctly classified 73% of cases, achieving 100% precision for TP and zero FN. Models trained with CROs and All-data also achieved 100% true positive rates, albeit with notably lower overall precision for CROs. The fact that 95% of cases were correctly classified makes the model—trained with FSS, the 25-foot test, and baseline features—a strong candidate to support clinicians in managing MS. These results suggest that such models could be valuable in guiding clinical decisions regarding whether to perform an MRI, potentially reducing tests as well as associated costs and burdens. Previous studies relayed on MRI to determine new lesions( 15 ). Limitations include the single-center nature, the small-size of our cohort and the need for synthetic oversampling of underrepresented classes, which might introduce overfitting. Future work should include more patients from different hospitals and regions to validate these findings. Conclusions Overall, the results demonstrate a strong performance of various classification and regression algorithms combined with data mining techniques in predicting key MS progression variables in a small, single-center cohort. The use of both clinical data and PROMs proved useful for predicting EDSS and spasticity levels and MRI lesions. These findings can support clinical decisions and telemedicine systems. The particularly promising results in predicting new lesion occurrence (100% recall with no FN) could help avoid unnecessary MRIs, reducing both costs and patient burden. The consistent variable importance across models supports the reliability of the findings. However, limitations—small cohort size and single-center design—highlight the need for larger, multi-center studies to improve model performance and generalizability. Declarations Ethics Approval and Consent to Participate The study was approved by the Ethics Committee of Hospital San Pedro, La Rioja (Spain). All participants provided written informed consent prior to inclusion. The research was conducted in accordance with the ethical principles of the Declaration of Helsinki (2013 revision) and all relevant national regulations for biomedical research involving human subjects. Data Availability Statement The datasets generated and analyzed during the current study are not publicly available due to institutional restrictions and patient confidentiality agreements. However, anonymized data may be made available from the corresponding author upon reasonable request and pending ethical approval. Funding The authors received no specific funding for this work. Conflict of Interest Statement The authors declare that there is no conflict of interest regarding the publication of this article. Declaration of generative AI and AI-assisted technologies in the writing process. During the preparation of this work the author(s) used ChatGPT in order to improve translatation quality, consistency and readability. After using this tool/service, the author(s) reviewed and edited the content as needed and take(s) full responsibility for the content of the publication. References Weinshenker B. Natural history of multiple sclerosis. Ann Neurol. D’Amico E, Haase R, Ziemssen T, Review. Patient-reported outcomes in multiple sclerosis care. Mult Scler Relat Disord. 2019;33:61–6. Filippi M, Preziosa P, Langdon D, Lassmann H, Paul F, Rovira À, et al. Identifying Progression in Multiple Sclerosis: New Perspectives. Annals of Neurology. Volume 88. John Wiley and Sons Inc.; 2020. pp. 438–52. van Munster CEP, Uitdehaag BMJ. Outcome Measures in Clinical Trials for Multiple Sclerosis. CNS Drugs. Volume 31. Springer International Publishing; 2017. pp. 217–36. Inojosa H, Schriefer D, Ziemssen T. Clinical outcome measures in multiple sclerosis: A review. Autoimmun Rev [Internet]. 2020 May 1 [cited 2024 Aug 21];19(5). Available from: https://pubmed.ncbi.nlm.nih.gov/32173519/ Izquierdo G. Multiple sclerosis symptoms and spasticity management: new data. Neurodegener Dis Manag [Internet]. 2017 Nov 1 [cited 2025 Mar 16];7(6s):7–11. Available from: https://pubmed.ncbi.nlm.nih.gov/29143581/ Reeve K, On BI, Havla J, Burns J, Gosteli-Peter MA, Alabsawi A et al. Prognostic models for predicting clinical disease progression, worsening and activity in people with multiple sclerosis. Cochrane Database Syst Rev [Internet]. 2023 Sep 8 [cited 2024 Aug 21];9(9). Available from: https://pubmed.ncbi.nlm.nih.gov/37681561/ Seccia R, Romano S, Salvetti M, Crisanti A, Palagi L, Grassi F. Machine learning use for prognostic purposes in multiple sclerosis. Life. 2021;11(2):1–18. Branco D, Martino BD, Esposito A, Tedeschi G, Bonavita S, Lavorgna L. Machine learning techniques for prediction of multiple sclerosis progr ession. Soft Computing - A Fusion of Foundations, Methodologies and Applicatio ns. Haouam K, Benmalek M. Machine Learning Algorithms for Early Prediction of Multiple Sclerosis Progression: A Comparative Study. Advances in Artificial Intelligence and Machine Learning. Viguera M, Martin-Sanchez F, Marzo ME. Using Redcap to Support the Development of a Learning Healthcare System for Patients with Multiple Sclerosis. Studies in Health Technology and Informatics. IOS Press BV; 2023. pp. 377–80. Harris PA, Taylor R, Thielke R, Payne J, Gonzalez N, Conde JG. Research electronic data capture (REDCap)-A metadata-driven methodology and workflow process for providing translational research informatics support. J Biomed Inf. 2009;42(2):377–81. Alkhawaldeh IM, Albalkhi I, Naswhan AJ. Challenges and limitations of synthetic minority oversampling techniques in machine learning. World J Methodol. 2023;13(5):373–8. Dahshan A. Unlocking the Secrets of Multiple Sclerosis Progression: How Machine L earning is Changing the Game. Mult Scler Relat Disord. Bonacchi R, Meani A, Pagani E, Marchesi O, Filippi M, Rocca MA. The role of cerebellar damage in explaining disability and cognition in multiple sclerosis phenotypes: a multiparametric MRI study. J Neurol. 2022;269(7):3841–57. Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Editorial decision: Revision requested 13 Nov, 2025 Reviews received at journal 24 Jul, 2025 Reviews received at journal 07 Jul, 2025 Reviews received at journal 06 Jul, 2025 Reviewers agreed at journal 03 Jul, 2025 Reviewers agreed at journal 26 Jun, 2025 Reviewers agreed at journal 24 Jun, 2025 Reviewers invited by journal 19 Jun, 2025 Editor invited by journal 20 May, 2025 Editor assigned by journal 16 May, 2025 Submission checks completed at journal 16 May, 2025 First submitted to journal 10 May, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6634220","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":474709892,"identity":"4c376ace-013b-49c7-bc8d-68f2ee8c3b7e","order_by":0,"name":"Minerva VIGUERA MORENO","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA/ElEQVRIiWNgGAWjYFACxgYwZX+8gfEAXJCHGC0MZw4wEKsFBm4kEKlFd0Zy82ueinvyjDMfPzjMm3NHjkEi9+CDNww29ri0mN1IbLPmOVNs2CydZnCYd9szYwaJvGTDOQxpiQ24tJw52GbM25bA2CadANJyOLFBIsdMmofhcAJOW8Ba/iXY90ge/wDTYv6bh+E/bocdb2x+zNuQkDhDggdhCzMPwwFYSGLT0sY451hC8gaenIKDc7cdNmbjeWMsOccgGbdfDrM//vCmJsF2A/vxjQ/ebjssx8+eY/jhTYUdTocBAZsEKhdMGuDRwMDA/AGv9CgYBaNgFIwCADB2Wo6qYCxBAAAAAElFTkSuQmCC","orcid":"","institution":"Universidad Nacional de Educación a Distancia (UNED)","correspondingAuthor":true,"prefix":"","firstName":"Minerva","middleName":"VIGUERA","lastName":"MORENO","suffix":""},{"id":474709893,"identity":"20dcbe6d-0d3b-43b7-9151-e5ba59ac8f8e","order_by":1,"name":"María Eugenia MARZO SOLA","email":"","orcid":"","institution":"Hospital San Pedro","correspondingAuthor":false,"prefix":"","firstName":"María","middleName":"Eugenia MARZO","lastName":"SOLA","suffix":""},{"id":474709894,"identity":"4d240018-6b72-4163-a5d9-1241f799faa5","order_by":2,"name":"Fernando MARTIN-SANCHEZ","email":"","orcid":"","institution":"Hospital Universitario La Paz","correspondingAuthor":false,"prefix":"","firstName":"Fernando","middleName":"","lastName":"MARTIN-SANCHEZ","suffix":""},{"id":474709895,"identity":"6e8099c2-f78f-45ef-b2ad-26d28fbdfefc","order_by":3,"name":"Ricardo SANCHEZ DE MADARIAGA","email":"","orcid":"","institution":"Instituto de Salud Carlos III","correspondingAuthor":false,"prefix":"","firstName":"Ricardo","middleName":"SANCHEZ","lastName":"DE MADARIAGA","suffix":""}],"badges":[],"createdAt":"2025-05-10 10:23:18","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6634220/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6634220/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":85272449,"identity":"7986c4d6-de40-447b-8980-0c6aae18aff6","added_by":"auto","created_at":"2025-06-24 06:40:27","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":33323,"visible":true,"origin":"","legend":"\u003cp\u003eFeatures importance for GB trained with PROMs dataset.\u003c/p\u003e","description":"","filename":"image1.png","url":"https://assets-eu.researchsquare.com/files/rs-6634220/v1/30d2e94aa29db72c5e4f367b.png"},{"id":85273225,"identity":"7d0c4a9c-1e5c-42fc-a47a-bc1c90791ffc","added_by":"auto","created_at":"2025-06-24 06:48:27","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":28554,"visible":true,"origin":"","legend":"\u003cp\u003eResidual plot for GB trained with PROMs dataset.\u003c/p\u003e","description":"","filename":"image2.png","url":"https://assets-eu.researchsquare.com/files/rs-6634220/v1/8e3b9be5a63a3404c70a7f41.png"},{"id":85272450,"identity":"dde4e7b5-e961-429b-8b4b-a555ffe08abe","added_by":"auto","created_at":"2025-06-24 06:40:27","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":29425,"visible":true,"origin":"","legend":"\u003cp\u003eFeatures importance for GB trained with CROs dataset\u003c/p\u003e","description":"","filename":"image3.png","url":"https://assets-eu.researchsquare.com/files/rs-6634220/v1/573bc5d2c7303142763f6cf7.png"},{"id":85272451,"identity":"1be08e07-840c-4ced-88ba-8c90ce702ed5","added_by":"auto","created_at":"2025-06-24 06:40:27","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":27486,"visible":true,"origin":"","legend":"\u003cp\u003eResidual plot for GB trained with CROs dataset\u003c/p\u003e","description":"","filename":"image4.png","url":"https://assets-eu.researchsquare.com/files/rs-6634220/v1/ce64784a8869b5ede0803d70.png"},{"id":85272474,"identity":"e953c207-ae06-4d70-9365-d2fc3a31d14b","added_by":"auto","created_at":"2025-06-24 06:40:27","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":38309,"visible":true,"origin":"","legend":"\u003cp\u003eFeatures importance for RF trained with All_data\u003c/p\u003e","description":"","filename":"image5.png","url":"https://assets-eu.researchsquare.com/files/rs-6634220/v1/a4433632a7414636bfa1db1a.png"},{"id":85272476,"identity":"d794f667-9e3b-415f-a04d-00a3c1ae800e","added_by":"auto","created_at":"2025-06-24 06:40:27","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":28075,"visible":true,"origin":"","legend":"\u003cp\u003eResidual plot for RF trained with All_data\u003c/p\u003e","description":"","filename":"image6.png","url":"https://assets-eu.researchsquare.com/files/rs-6634220/v1/89b545df972a6949a993a0a4.png"},{"id":85272478,"identity":"bcd207ad-22be-4928-9ad1-7fbc3c8821e6","added_by":"auto","created_at":"2025-06-24 06:40:27","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":38352,"visible":true,"origin":"","legend":"\u003cp\u003eFeature importance for PROs dataset\u003c/p\u003e","description":"","filename":"image7.png","url":"https://assets-eu.researchsquare.com/files/rs-6634220/v1/8773bc0c2cb2c0455d242fb3.png"},{"id":85272489,"identity":"857c2bb5-0140-4d15-88ac-a17e29cd8d24","added_by":"auto","created_at":"2025-06-24 06:40:28","extension":"png","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":37311,"visible":true,"origin":"","legend":"\u003cp\u003eFeature importance for CROs dataset\u003c/p\u003e","description":"","filename":"image8.png","url":"https://assets-eu.researchsquare.com/files/rs-6634220/v1/14f74038669a30f3301a3918.png"},{"id":85272481,"identity":"8ab706d6-ea36-407d-bcac-a3598daf46ce","added_by":"auto","created_at":"2025-06-24 06:40:27","extension":"png","order_by":9,"title":"Figure 9","display":"","copyAsset":false,"role":"figure","size":33258,"visible":true,"origin":"","legend":"\u003cp\u003eFeature importance for ALL-data dataset\u003c/p\u003e","description":"","filename":"image9.png","url":"https://assets-eu.researchsquare.com/files/rs-6634220/v1/7c895d0430fb2138a513e28b.png"},{"id":85274105,"identity":"8d480907-38af-4536-ba04-7b639028d58e","added_by":"auto","created_at":"2025-06-24 06:56:28","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1075667,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6634220/v1/b7016bf2-eadb-43e9-a326-050bdeb4ef68.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Predicting Multiple Sclerosis Outcomes: A Machine Learning Approach Integrating Patient-Reported Outcomes and Clinical Data","fulltext":[{"header":"Introduction","content":"\u003cp\u003eMultiple sclerosis (MS) is a chronic, immune-mediated disease of the central nervous system with highly variable progression and clinical presentation. This variability complicates clinical management and long-term treatment planning. MS patients may experience motor and sensory deficits, cognitive impairments, and exacerbations, leading to increased disability over time(\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e). The Expanded Disability Status Scale (EDSS) is the most used tool to assess neurological impairment, but it has limitations such as inter-rater variability and an overemphasis on ambulation(\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eOther important monitoring aspects include spasticity levels, new lesions detected by RMI, and exacerbations. Spasticity can greatly affect mobility and quality of life (QoL), while new lesions, typically detected via Magnetic Resonance Imaging (MRI) indicate active disease, and exacerbations mark acute neurological deterioration(\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e, \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eClinical tests like the 25-foot walk, 9-hole peg test, and symbol digit modalities (SDM) test provide objective measures of physical and cognitive function but may not fully capture the patient\u0026rsquo;s experience(\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e). In contrast, PROMs such as the MSIS-29, Fatigue Severity Scale (FSS), Health Assessment Questionnaire (HAQ), and Neuro-QoL instruments offer valuable insights into health and QoL from the patient's perspective. Integrating PROMs with CROs enhances patient-centered care, telemedicine, and individualized treatment decisions(\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eThis study explores the use of machine learning (ML) techniques to predict key clinical outcomes in MS \u0026ndash; EDSS level, spasticity severity, and newly developed RMI lesions by analyzing data from 240 patients monitored over 18 months. Previous studies have explored the possibility of predicting MS progression using ML or neural networks(\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e); however, the utility of patient-reported outcome measures (PROMs) as predictive variables remains unclear(\u003cspan additionalcitationids=\"CR9\" citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e). We compare the predictive power of ML models trained on various combinations of baseline information, CROs, and PROMs, aiming to assess the utility of PROMs in clinical decision-making and to identify optimal features and algorithms for short-term disease progression prediction. It is important to emphasize that the PROMs in this study were collected prior to the determination of the study outcomes, ensuring that the predictive analysis relies solely on preliminary information.\u003c/p\u003e"},{"header":"Methodology","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e3.1 Study Design\u003c/h2\u003e \u003cp\u003eThis study follows a prospective cohort design involving 240 patients with MS treated at a second-level hospital in Spain. Data collection was conducted over 18 months, capturing baseline clinical characteristics, routine clinical assessments, and patient-reported outcomes at regular intervals. The dataset comprises 206 complete cases that include baseline data, consultation information, and PROMs collected up to 180 days prior to each consultation (median: 73 days)(\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e3.2 Data Collection\u003c/h2\u003e \u003cp\u003eREDCap projects were designed to collect research information(\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e). Baseline data were collected at the study's initiation, capturing both demographic (age, gender) and clinical information (time since diagnosis, RMI lesions at diagnosis, previous exacerbations and evolution form).\u003c/p\u003e \u003cp\u003eData from outpatient consultations were collected biannually and included EDSS level, spasticity severity, new MRI lesions and new exacerbations in the last semester. Routine clinical tests, such as the 25-foot walk test, 9-hole peg test, and SDM, were also performed. PROMs were completed by patients every three months and included assessments of physical and psychological impact (MSIS-29), fatigue (FSS), functional ability (HAQ), and cognitive and depressive symptoms (Neuro-QoL instruments)(\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e3.3. Dataset Preparation\u003c/h2\u003e \u003cp\u003eThe collected data underwent preprocessing to ensure quality and consistency. Raw data were validated and cleaned to address missing or inconsistent entries punctual missing data was imputed using KNN. PROMs, CROs, and baseline information were integrated into a unified dataset, without any identifiable information to comply with ethical and legal standards.\u003c/p\u003e \u003cp\u003eTo assess the utility of different data types, this dataset was used to prepare three subsets for algorithms training proposals: \u003cb\u003ePROMs dataset\u003c/b\u003e included PROMs and baseline information. \u003cb\u003eCROs dataset\u003c/b\u003e was composed by clinical tests, regular consultation data, and baseline information. \u003cb\u003eAll-data\u003c/b\u003e combined both.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e3.4 Machine Learning Workflow\u003c/h2\u003e \u003cp\u003eThe objective was to train and compare ML models and data mining techniques to predict three key outcomes: EDSS level, spasticity severity, and new RMI lesions.\u003c/p\u003e \u003cp\u003eThe ML pipeline included codification, splitting the data for training and testing, oversampling to balance underrepresented categories, hyperparameter tuning via grid search and cross-validation, and model evaluation. Regression models were assessed using mean squared error (MSE) and R-squared, while classification models were evaluated using accuracy, precision, recall, F1-score, and ROC-AUC, focusing on the clinical significance of these parameters.\u003c/p\u003e \u003cdiv id=\"Sec7\" class=\"Section3\"\u003e \u003ch2\u003e3.4.1 EDSS Inference\u003c/h2\u003e \u003cp\u003eInferring EDSS levels posed a methodological challenge. Previous studies successfully classified MS patients by functional level but typically used stratified categories (e.g., three levels), limiting clinical utility due to significant functional differences within each stratum(\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e). This study approaches EDSS inference as a regression problem, considering its 0\u0026ndash;10 scale with 0.5 increments.\u003c/p\u003e \u003cp\u003eA preliminary analysis revealed an uneven distribution of EDSS levels, with an underrepresentation of higher scores (above 8). This imbalance negatively impacted model performance, biasing results toward the most frequent values. To address this, underrepresented categories were manually oversampled by generating synthetic samples while introducing random noise to reduce overfitting(\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e). Normalization was applied before training Support Vector Machine (SVM) and Ridge Regression.\u003c/p\u003e \u003cp\u003eFour regression algorithms were trained and optimized through grid search hyperparameter selection, including Ridge Regression, SVM, Random Forest (RF) Regressor, and Gradient Boosting (GB).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section3\"\u003e \u003ch2\u003e3.4.2 Spasticity level estimation\u003c/h2\u003e \u003cp\u003eSpasticity severity was measured using the Ashworth Scale, which quantifies muscle stiffness and requires in-person evaluation. Predicting spasticity severity using ML could enhance telemedicine capabilities and support clinical decision-making before patient consultations.\u003c/p\u003e \u003cp\u003eThe Ashworth Scale classifies spasticity into five levels (0, 1, 1.5, 2, 3), leading to a classification approach using a RF Classifier. However, higher severity classes were underrepresented, affecting model performance. Three approaches were tested: no oversampling (RF adjusted for disbalanced classes), manual oversampling, and SMOTE(\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e). Despite these adjustments, precision for classes 2 and 3 remained suboptimal, leading to their combination into a single category.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section3\"\u003e \u003ch2\u003e3.4.3 New RMI lesion prediction\u003c/h2\u003e \u003cp\u003eNew lesions concurrence is a marker of MS progression, typically assessed via MRI performed annually. The high cost and limited availability of MRI necessitate alternative decision-support methods. This study framed lesion prediction as a classification task to determine patients unlikely to develop new lesions. The focus was minimizing false negatives (FN), as failing to detect a new lesion could lead to inappropriate clinical decisions.\u003c/p\u003e \u003cp\u003eDue to a strong class imbalance (11 positive cases), ADASYN oversampling was applied before training classification algorithms(\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e). RF (n_estimators\u0026thinsp;=\u0026thinsp;10, max_depth\u0026thinsp;=\u0026thinsp;10), SVM (Kernel\u0026thinsp;=\u0026thinsp;Sigmoid) y XGBClassifier (objective='binary:logistic') were tested.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003e4.1 EDSS Inference\u003c/h2\u003e \u003cp\u003eHyperparameters selection using grid search allowed for the optimization of each algorithm. Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e summarizes the MSE and R\u0026sup2; values for each model across the datasets.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eMSE and R2 for each model and dataset\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"7\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e\u0026nbsp;\u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003ePROMs\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e \u003cp\u003eCROs\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c7\" namest=\"c6\"\u003e \u003cp\u003eAll_data\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMSE\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eR\u003csup\u003e2\u003c/sup\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMSE\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eR\u003csup\u003e2\u003c/sup\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eMSE\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eR\u003csup\u003e2\u003c/sup\u003e\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.87\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.78\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e1.23\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.68\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.81\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.80\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRF\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.40\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.90\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.32\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.92\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.27\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.93\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSMV\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.43\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.89\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.46\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.88\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.40\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.90\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGB\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.36\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.91\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.30\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.92\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.39\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.90\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eComplementing algorithm performance, we analyzed features importance for each optimized model to better understand the algorithm\u0026rsquo;s decision-making. The output values were adjusted to the nearest value on the scale, and a residual plot was generated.\u003c/p\u003e \u003cdiv id=\"Sec12\" class=\"Section3\"\u003e \u003ch2\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e4.1.1 PROMs and GB\u003c/span\u003e (\u003cem\u003elearning_rate\u0026thinsp;=\u0026thinsp;0.2, n_estimators\u0026thinsp;=\u0026thinsp;100)\u003c/em\u003e\u003c/h2\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eUsing PROMs in model training (see Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e), revealed the features with the highest weight were the disease course and duration, previous relapses, HAQ and FSS scores (physical function and fatigue). Variables related to cognitive function and depression have intermediate importance, while age, gender, and initial neural lesions are the least important.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe residual analysis, shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e, revealed 5 cases in the test set were classified with an absolute error greater than 1 point (the acceptability threshold set in this study), corresponding to 94% of subjects being correctly categorized. Additionally, there is a tendency to underestimate the EDSS level (negative residuals), which is consistent with the overrepresentation of lower levels. Furthermore, a better fit is observed in the higher levels, suggesting that the generation of synthetic samples in these underrepresented categories may have led to overfitting at high values.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section3\"\u003e \u003ch2\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e4.1.2 CROs and GBRegressor (\u003c/span\u003elearning_rate\u0026thinsp;=\u0026thinsp;0.2, n_estimators\u0026thinsp;=\u0026thinsp;100).\u003c/h2\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe analysis of features importance (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e) indicates that disease course and duration, along with parameters measuring the physical impact of the disease (both gross and fine motor skills), carry the most weight. Cognitive impairment has a moderate importance in this model. Other variables, such as age, gender, or initial neural lesions, appear to have less influence.\u003c/p\u003e \u003cp\u003eResiduals analysis (See Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e) shows that only 4 cases had an absolute error greater than 1 point, meaning that 95% of the predictions fall within the acceptable range. Additionally, the system tends to underestimate the EDSS level, particularly in the mid-range values, while fitting relatively well for higher values as described when using PROMs.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section3\"\u003e \u003ch2\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e4.1.3All-data and RF. (\u003c/span\u003emax_depth\u0026thinsp;=\u0026thinsp;10, min_samples_leaf\u0026thinsp;=\u0026thinsp;1, min_samples_split\u0026thinsp;=\u0026thinsp;3, n_estimators\u0026thinsp;=\u0026thinsp;100).\u003c/h2\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe features importance, shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e, is consistent with what was expected regarding their relationship with disease progression. Disease course along with 25-foots test, HAQ score (also reflecting physical function), time since diagnosis, 9-Holes, and fatigue (FSS), were the variables with the greatest weight. In contrast, age, gender, spinal lesions, and variables related to cognitive function appear to have less importance.\u003c/p\u003e \u003cp\u003eRegarding the residual distribution (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e), it is observed that only 3 out of the 87 cases in the test set deviated more than 1 point from the actual value, which is the clinically acceptable threshold. Therefore, it can be stated that 96.5% of the predictions fall within the acceptable variability range, indicating clinical utility.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003e4.2 Spasticity level\u003c/h2\u003e \u003cp\u003eRandom Forest Classifier was trained with previously determined hyperparameters (n_estimators\u0026thinsp;=\u0026thinsp;100, max_depth\u0026thinsp;=\u0026thinsp;10, min_samples_split\u0026thinsp;=\u0026thinsp;2, min_samples_leaf\u0026thinsp;=\u0026thinsp;1). Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e presents the overall accuracy achieved with different data combinations and imbalance-handling methods, both for original and grouped categories. The highest accuracies were obtained using all available data: 0.86 for the five original categories and 0.9 for grouped categories. Using only PROMs dataset resulted in slightly lower but comparable accuracies (0.81 and 0.90, respectively), while employing CROs yielded performance equivalent to using all variables.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eOverall accuracy for RF prediction of Ashworth Scale\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eCategories\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eOversampling\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"3\" nameend=\"c5\" namest=\"c3\"\u003e \u003cp\u003eDataset\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ePROs\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eCROs\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eAll-Data\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eOriginal\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNone\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0,79\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0,76\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0,81\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eManual\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0,81\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0,79\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0,83\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSMOTE\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0,81\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0,83\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0,86\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eGrouped (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNone\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0,86\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0,88\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0,86\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eManual\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0,90\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0,88\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0,86\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSMOTE\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0,86\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0,9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0,9\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eSince overall accuracy does not fully reflect the clinical meaning of classification errors, confusion matrices and other metrics were extracted for the best-performing models. The confusion matrices reveal poorer performance for higher spasticity values, regardless of the training data (Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e); however, grouping categories 2 and 3 improved the classification of these higher values, as demonstrated by precision and recall shown in Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eConfusion Matrices for RF optimized classification\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCategories\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePROs (manual)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eCROs (SMOTE)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eAll-Data (SMOTE)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eOriginal\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\left(\\begin{array}{c}\\begin{array}{ccccc}27\u0026amp;\\:1\u0026amp;\\:0\u0026amp;\\:0\u0026amp;\\:0\\\\\\:2\u0026amp;\\:3\u0026amp;\\:0\u0026amp;\\:0\u0026amp;\\:0\\\\\\:0\u0026amp;\\:0\u0026amp;\\:3\u0026amp;\\:1\u0026amp;\\:0\\\\\\:0\u0026amp;\\:1\u0026amp;\\:0\u0026amp;\\:1\u0026amp;\\:1\\\\\\:0\u0026amp;\\:0\u0026amp;\\:0\u0026amp;\\:2\u0026amp;\\:0\\end{array}\\end{array}\\right)\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\left(\\begin{array}{c}\\begin{array}{ccccc}28\u0026amp;\\:0\u0026amp;\\:0\u0026amp;\\:0\u0026amp;\\:0\\\\\\:2\u0026amp;\\:3\u0026amp;\\:0\u0026amp;\\:0\u0026amp;\\:0\\\\\\:1\u0026amp;\\:0\u0026amp;\\:2\u0026amp;\\:1\u0026amp;\\:0\\\\\\:0\u0026amp;\\:0\u0026amp;\\:0\u0026amp;\\:2\u0026amp;\\:1\\\\\\:0\u0026amp;\\:0\u0026amp;\\:0\u0026amp;\\:2\u0026amp;\\:0\\end{array}\\end{array}\\right)\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\left(\\begin{array}{c}\\begin{array}{ccccc}27\u0026amp;\\:0\u0026amp;\\:1\u0026amp;\\:0\u0026amp;\\:0\\\\\\:1\u0026amp;\\:4\u0026amp;\\:0\u0026amp;\\:0\u0026amp;\\:0\\\\\\:0\u0026amp;\\:0\u0026amp;\\:3\u0026amp;\\:1\u0026amp;\\:0\\\\\\:0\u0026amp;\\:0\u0026amp;\\:0\u0026amp;\\:2\u0026amp;\\:1\\\\\\:0\u0026amp;\\:0\u0026amp;\\:0\u0026amp;\\:2\u0026amp;\\:0\\end{array}\\end{array}\\right)\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGrouped \u003c/p\u003e \u003cp\u003e(\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\left(\\begin{array}{c}\\begin{array}{cccc}27\u0026amp;\\:1\u0026amp;\\:0\u0026amp;\\:0\\\\\\:1\u0026amp;\\:4\u0026amp;\\:0\u0026amp;\\:0\\\\\\:0\u0026amp;\\:0\u0026amp;\\:3\u0026amp;\\:1\\\\\\:0\u0026amp;\\:1\u0026amp;\\:0\u0026amp;\\:4\\end{array}\\end{array}\\right)\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\left(\\begin{array}{c}\\begin{array}{cccc}28\u0026amp;\\:0\u0026amp;\\:0\u0026amp;\\:0\\\\\\:2\u0026amp;\\:3\u0026amp;\\:0\u0026amp;\\:0\\\\\\:1\u0026amp;\\:0\u0026amp;\\:2\u0026amp;\\:1\\\\\\:0\u0026amp;\\:0\u0026amp;\\:0\u0026amp;\\:5\\end{array}\\end{array}\\right)\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\left(\\begin{array}{c}\\begin{array}{cccc}26\u0026amp;\\:1\u0026amp;\\:1\u0026amp;\\:0\\\\\\:1\u0026amp;\\:4\u0026amp;\\:0\u0026amp;\\:0\\\\\\:0\u0026amp;\\:0\u0026amp;\\:3\u0026amp;\\:1\\\\\\:0\u0026amp;\\:0\u0026amp;\\:0\u0026amp;\\:5\\end{array}\\end{array}\\right)\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003ePerformance metrics for each dataset with grouped categories\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"7\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e\u0026nbsp;\u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003ePROs\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e \u003cp\u003eCROs\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c7\" namest=\"c6\"\u003e \u003cp\u003eALL-data\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eprecision\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003erecall\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eprecision\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003erecall\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eprecision\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003erecall\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCat 0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0,96\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0,96\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0,9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0,96\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0,93\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCat 1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0,67\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0,8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0,6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0,8\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCat 1.5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0,75\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0,5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0,75\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0,75\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCat 2\u0026ndash;3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.80\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0,80\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0,83\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0,83\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003e0,9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e \u003cp\u003e0,9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c7\" namest=\"c6\"\u003e \u003cp\u003e0,90\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMacro_avg\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0,86\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0,83\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0,93\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0,78\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0,84\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0,87\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWeighted_avg\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0,91\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0,90\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0,92\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0,90\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0,91\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0,90\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eFigures \u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003e, \u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003e, and \u003cspan refid=\"Fig9\" class=\"InternalRef\"\u003e9\u003c/span\u003e graphically illustrate the importance of the variables used in training across the different optimized models. Greater importance was found for physical-related features across all models, consistently with the impact of spasticity on musculature (\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003e4.3 New RMI lesion prediction\u003c/h2\u003e \u003cp\u003eRF, SVM and XGB produced underwhelming results when trained using all available variables across different data combinations. Aligned with the dataset\u0026rsquo;s negative bias, all cases were classified as negative.\u003c/p\u003e \u003cp\u003eHowever, when the variable set was narrowed down to those that had higher importance in initial models, satisfactory results were achieved by minimizing FN and correctly classifying true positives (TP). Consequently, all possible combinations of few variables were tested for each algorithm with hyperparameters previously selected.\u003c/p\u003e \u003cp\u003eOverall, the results indicate that it is possible to avoid FN when classifying patients with new lesions\u0026mdash;considered an unacceptable error. In all cases, the TP rate reached 100%, meaning that the models correctly identified all patients with new lesions. Moreover, using few input variables appears to be associated with a lower number of false positives (FP), potentially optimizing the classification system and reducing unnecessary MRIs in patients without new lesions. The variable combinations that led to this optimized classification and obtained performance are shown in Table 5.\u003c/p\u003e\n\u003cp\u003eTable 5 Optimized model features and performance by training dataset\u003c/p\u003e\u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"No\" id=\"Taba\" border=\"1\"\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePROs\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eCROs\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eAll_data\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAlgorithm\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eXGB\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eSVM\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eRF\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003efeatures\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eFSS\u003c/p\u003e \u003cp\u003eCourse\u003c/p\u003e \u003cp\u003eRMI_brain_lesions\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ePrev_exacerbations\u003c/p\u003e \u003cp\u003eRMI_brain_lesions\u003c/p\u003e \u003cp\u003eRMI_spinal_lesions\u003c/p\u003e \u003cp\u003eGender\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFSS\u003c/p\u003e \u003cp\u003eCourse\u003c/p\u003e \u003cp\u003ePrev_exacerbations\u003c/p\u003e \u003cp\u003e25-foot test\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eConf_matrix\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\left(\\begin{array}{cc}42\u0026amp;\\:18\\\\\\:0\u0026amp;\\:2\\end{array}\\right)\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\left(\\begin{array}{cc}36\u0026amp;\\:24\\\\\\:0\u0026amp;\\:2\\end{array}\\right)\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\left(\\begin{array}{cc}55\u0026amp;\\:5\\\\\\:0\u0026amp;\\:2\\end{array}\\right)\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAUC_ROC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.80\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.84\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.92\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eSpecifically, for PROMs data, the best combination included FSS, disease course, and basal brain MRI lesions. For the CROs dataset, the optimal variables were a new relapse in the last 6 months, spinal and brain lesions at onset, and gender. When both data types were combined, the rate of FP decrease to 5, leading to the most optimized model in terms of false positive minimization while maintaining 0% of FN rate.\u003c/p\u003e \u003cp\u003eOverall, these results demonstrate the potential of ML models in predicting EDSS levels, spasticity severity, and new lesions with clinically meaningful accuracy. Including PROMs\u0026mdash;either alone or combined with CROs\u0026mdash;improved predictive performance, reinforcing their clinical value.\u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eThis study examined the relationship between various data types collected from MS patients and their usefulness for predicting EDSS levels, spasticity severity, and the potential occurrence of new lesions without relying on MRI images.\u003c/p\u003e \u003cp\u003eFor EDSS inference, the R\u0026sup2; and MSE values across the three datasets (PROMs, CROs, and All-data) indicate that the models effectively capture data variability and can determine the EDSS score with an absolute error of less than 1 point in 90\u0026ndash;95% of cases. Although the optimized model using only PROMs and baseline characteristics performed slightly worse than those including clinical measurements, the results were comparable. It is important to note that these PROMs were collected up to six months before the patient\u0026rsquo;s specialized consultation, which might partly explain some errors if deterioration occurred later. However, this finding supports the utility of PROMs in telemedicine contexts, where they can serve as predictive or monitoring variables with similar performance to clinical data. The best performance was achieved when combining CROs with PROMs, reinforcing the importance of including both in clinical practice to minimize deviations in EDSS categorization. The relative importance of variables was consistent across models, with disease course and time leading, followed by variables related to the physical impact of MS, aligned with the higher weigh of physical impact in EDSS(\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e). When both CROs and PROMs were included, their variable importance was consistent with the type of impact they measure, further supporting the inclusion of PROMs in predictive models. Our approach has allowed for the inference of the EDSS level without the need for stratification and by minimizing the MSE, improving the accuracy of previous models and expanding their applicability This approach allows for a more precise estimation of the EDSS level than previous studies, eliminating the need for stratification and thereby enhancing its applicability in supporting clinical decisions(\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eRegarding spasticity level estimation, global accuracy metrics (approximately 0.90 for CROs and All-data, and 0.88 for PROMs) show that the models correctly classify about 90% of cases. Although using only PROMs slightly reduced precision, the overall performance remained similar. The stable weighted average (~\u0026thinsp;0.90) across configurations suggests that while frequent classes are well classified, improvements for less frequent, higher spasticity categories may be possible with larger datasets and alternative balancing strategies. Notably, the model demonstrated excellent performance in identifying patients without spasticity (Category 0) with high recall and precision, though this might be partly due to the dataset bias toward lower spasticity values. For mild spasticity (Category 1) All-data set appears to be most useful in detecting without generating FP. Intermediate spasticity (Category 1.5) was suboptimal detected overall, though training with PROMs yielded promising accuracy. Finally, for high spasticity (grouped Categories 2\u0026ndash;3), recall was excellent (100% in CROs and All-data) with variable precision. In terms of variable importance, clinical variables such as time since diagnosis and disease course consistently held high weight, while neurological lesions, prior relapses, and gender were less influential. The importance of PROMs and their clinical counterparts aligns with expectations regarding their impact on patients\u0026rsquo; physical disability and QoL.\u003c/p\u003e \u003cp\u003eEstimating the spasticity level can serve as support for telemedicine applications, potentially reducing the need for regular outpatient consultations and, consequently, the burden of care on patients.\u003c/p\u003e \u003cp\u003eFinally, for predicting new MRI lesions, PROMs collected before consultation correctly classified 73% of cases, achieving 100% precision for TP and zero FN. Models trained with CROs and All-data also achieved 100% true positive rates, albeit with notably lower overall precision for CROs. The fact that 95% of cases were correctly classified makes the model\u0026mdash;trained with FSS, the 25-foot test, and baseline features\u0026mdash;a strong candidate to support clinicians in managing MS. These results suggest that such models could be valuable in guiding clinical decisions regarding whether to perform an MRI, potentially reducing tests as well as associated costs and burdens. Previous studies relayed on MRI to determine new lesions(\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eLimitations include the single-center nature, the small-size of our cohort and the need for synthetic oversampling of underrepresented classes, which might introduce overfitting. Future work should include more patients from different hospitals and regions to validate these findings.\u003c/p\u003e"},{"header":"Conclusions","content":"\u003cp\u003eOverall, the results demonstrate a strong performance of various classification and regression algorithms combined with data mining techniques in predicting key MS progression variables in a small, single-center cohort. The use of both clinical data and PROMs proved useful for predicting EDSS and spasticity levels and MRI lesions. These findings can support clinical decisions and telemedicine systems. The particularly promising results in predicting new lesion occurrence (100% recall with no FN) could help avoid unnecessary MRIs, reducing both costs and patient burden.\u003c/p\u003e \u003cp\u003eThe consistent variable importance across models supports the reliability of the findings. However, limitations\u0026mdash;small cohort size and single-center design\u0026mdash;highlight the need for larger, multi-center studies to improve model performance and generalizability.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics Approval and Consent to Participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe study was approved by the Ethics Committee of Hospital San Pedro, La Rioja (Spain). All participants provided written informed consent prior to inclusion. The research was conducted in accordance with the ethical principles of the Declaration of Helsinki (2013 revision) and all relevant national regulations for biomedical research involving human subjects.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData Availability Statement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe datasets generated and analyzed during the current study are not publicly available due to institutional restrictions and patient confidentiality agreements. However, anonymized data may be made available from the corresponding author upon reasonable request and pending ethical approval.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors received no specific funding for this work.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConflict of Interest Statement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare that there is no conflict of interest regarding the publication of this article.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDeclaration of generative AI and AI-assisted technologies in the writing process.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eDuring the preparation of this work the author(s) used ChatGPT in order to improve translatation quality, consistency and readability. After using this tool/service, the author(s) reviewed and edited the content as needed and take(s) full responsibility for the content of the publication.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eWeinshenker B. Natural history of multiple sclerosis. Ann Neurol.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eD\u0026rsquo;Amico E, Haase R, Ziemssen T, Review. Patient-reported outcomes in multiple sclerosis care. Mult Scler Relat Disord. 2019;33:61\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFilippi M, Preziosa P, Langdon D, Lassmann H, Paul F, Rovira \u0026Agrave;, et al. Identifying Progression in Multiple Sclerosis: New Perspectives. Annals of Neurology. Volume 88. John Wiley and Sons Inc.; 2020. pp. 438\u0026ndash;52.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003evan Munster CEP, Uitdehaag BMJ. Outcome Measures in Clinical Trials for Multiple Sclerosis. CNS Drugs. Volume 31. Springer International Publishing; 2017. pp. 217\u0026ndash;36.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eInojosa H, Schriefer D, Ziemssen T. Clinical outcome measures in multiple sclerosis: A review. Autoimmun Rev [Internet]. 2020 May 1 [cited 2024 Aug 21];19(5). Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://pubmed.ncbi.nlm.nih.gov/32173519/\u003c/span\u003e\u003cspan address=\"https://pubmed.ncbi.nlm.nih.gov/32173519/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eIzquierdo G. Multiple sclerosis symptoms and spasticity management: new data. Neurodegener Dis Manag [Internet]. 2017 Nov 1 [cited 2025 Mar 16];7(6s):7\u0026ndash;11. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://pubmed.ncbi.nlm.nih.gov/29143581/\u003c/span\u003e\u003cspan address=\"https://pubmed.ncbi.nlm.nih.gov/29143581/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eReeve K, On BI, Havla J, Burns J, Gosteli-Peter MA, Alabsawi A et al. Prognostic models for predicting clinical disease progression, worsening and activity in people with multiple sclerosis. Cochrane Database Syst Rev [Internet]. 2023 Sep 8 [cited 2024 Aug 21];9(9). Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://pubmed.ncbi.nlm.nih.gov/37681561/\u003c/span\u003e\u003cspan address=\"https://pubmed.ncbi.nlm.nih.gov/37681561/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSeccia R, Romano S, Salvetti M, Crisanti A, Palagi L, Grassi F. Machine learning use for prognostic purposes in multiple sclerosis. Life. 2021;11(2):1\u0026ndash;18.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBranco D, Martino BD, Esposito A, Tedeschi G, Bonavita S, Lavorgna L. Machine learning techniques for prediction of multiple sclerosis progr ession. Soft Computing - A Fusion of Foundations, Methodologies and Applicatio ns.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHaouam K, Benmalek M. Machine Learning Algorithms for Early Prediction of Multiple Sclerosis Progression: A Comparative Study. Advances in Artificial Intelligence and Machine Learning.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eViguera M, Martin-Sanchez F, Marzo ME. Using Redcap to Support the Development of a Learning Healthcare System for Patients with Multiple Sclerosis. Studies in Health Technology and Informatics. IOS Press BV; 2023. pp. 377\u0026ndash;80.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHarris PA, Taylor R, Thielke R, Payne J, Gonzalez N, Conde JG. Research electronic data capture (REDCap)-A metadata-driven methodology and workflow process for providing translational research informatics support. J Biomed Inf. 2009;42(2):377\u0026ndash;81.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAlkhawaldeh IM, Albalkhi I, Naswhan AJ. Challenges and limitations of synthetic minority oversampling techniques in machine learning. World J Methodol. 2023;13(5):373\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDahshan A. Unlocking the Secrets of Multiple Sclerosis Progression: How Machine L earning is Changing the Game. Mult Scler Relat Disord.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBonacchi R, Meani A, Pagani E, Marchesi O, Filippi M, Rocca MA. The role of cerebellar damage in explaining disability and cognition in multiple sclerosis phenotypes: a multiparametric MRI study. J Neurol. 2022;269(7):3841\u0026ndash;57.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"bmc-medical-informatics-and-decision-making","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"midm","sideBox":"Learn more about [BMC Medical Informatics and Decision Making](http://bmcmedinformdecismak.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/midm/default.aspx","title":"BMC Medical Informatics and Decision Making","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-6634220/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6634220/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eMultiple sclerosis (MS) is a chronic, immune-mediated disease with variable progression that complicates clinical management(\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e). Traditional assessments, such as the Expanded Disability Status Scale (EDSS), have limitations, prompting the integration of patient-reported outcome measures (PROMs) alongside clinician-reported outcomes (CROs) to capture a comprehensive view of disease impact(\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e). This study investigates the use of machine learning (ML) techniques to predict three key clinical outcomes in MS: EDSS level, spasticity severity, and the development of new lesions detected through magnetic resonance. Data were collected from 240 MS patients over 18 months, including baseline demographics, CROs, and PROMs. The ML pipeline involved feature encoding, data splitting (80:20), oversampling for underrepresented classes, and hyperparameter tuning via grid search and cross-validation. Regression models (evaluated with MSE and R\u0026sup2;) and classification models were trained depending on predicted variable characteristics. Results indicate that models incorporating both PROMs and CROs achieved superior performance in predicting EDSS and spasticity, while ML models reliably identified new lesions with 100% true positive rate. PROMs collected before clinical assessment combined with basal demographics also lead to predictive models with acceptable performance. These findings support the potential of integrating PROMs into clinical decision-making and telemedicine for improved MS management.\u003c/p\u003e","manuscriptTitle":"Predicting Multiple Sclerosis Outcomes: A Machine Learning Approach Integrating Patient-Reported Outcomes and Clinical Data","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-06-24 06:40:22","doi":"10.21203/rs.3.rs-6634220/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2025-11-13T15:15:45+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-07-24T09:53:51+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-07-07T06:18:49+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-07-07T02:32:00+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"247594278773201177413201118447762897610","date":"2025-07-03T12:08:45+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"119077971094540622576984044726943635301","date":"2025-06-26T13:06:24+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"38337138169824049406494889284730600520","date":"2025-06-24T06:48:14+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-06-19T06:33:00+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2025-05-20T17:26:48+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-05-16T17:01:18+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-05-16T16:58:28+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Medical Informatics and Decision Making","date":"2025-05-10T10:21:04+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"bmc-medical-informatics-and-decision-making","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"midm","sideBox":"Learn more about [BMC Medical Informatics and Decision Making](http://bmcmedinformdecismak.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/midm/default.aspx","title":"BMC Medical Informatics and Decision Making","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"b1fb03cc-b1b8-4624-b4c7-aee8e2eca33e","owner":[],"postedDate":"June 24th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2025-12-30T18:08:11+00:00","versionOfRecord":[],"versionCreatedAt":"2025-06-24 06:40:22","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6634220","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6634220","identity":"rs-6634220","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.