Results
Out of 3077 records screened, 94 were selected for analysis in this review. PRISMA diagram [ 20 ] (Fig. 1 ) illustrates the process of paper selection. The reasons for paper exclusions were: no full-text available (38.1%), no PROMs used as input variables (35%), or no AI models used (12%), or methods did not aim to predict patient outcomes (8.75%). Fig. 1 PRISMA flow diagram
PRISMA flow diagram
Among the identified studies, 33 (35%) were conducted in USA, 31 (33%) in Europe, 6 (6%) in Canada, 15 (16%) in Asia, 1 (1%) in South America, 1 (1%) in New Zealand, and 1 (1%) in Turkey. Six (6%) studies were conducted internationally (USA and Canada ( n = 2, 2%); UK and USA ( n = 2, 2%); Canada and Sweden ( n = 1, 1%); Europe, US, Australia and Israel (n = 1, 1%)). Identified papers focused on orthopedics ( n = 38, 40%), oncology ( n = 22, 23%), mental health ( n = 17, 18%), respiratory ( n = 8, 9%), neurology ( n = 4, 4%) and other domains ( n = 5, 5%), which appeared only once: hearing, endometriosis, palliative care, sub-health state, and cardiovascular. The studies were published between 2010 and 2023 (Fig. 2 ). The data were obtained either from existing registry/database ( n = 44, 47%), or pre-existing or current research studies ( n = 47, 50%, not reported: n = 3, 3%).The self-reported input variables were combined with clinical and demographic data ( n = 63, 67%), only demographic data ( n = 14, 15%), only clinical data ( n = 3, 3%), or other types of data ( n = 4, 4%), such as wearable, electroencephalographic, bio-mechanical, or family data. Ten studies used only self-reported data for predictions. Most papers ( n = 63, 67%) were predicting self-reported outcomes, 14 (15%) of which were Minimally Clinically Important Differences (MCID) between pre- and post-clinical event data collection. Other papers used either only objectively measured outcome ( n = 28, 30%), or a combination of self-reported and objective outcomes ( n = 3, 3%). Sample sizes of the papers varied from 20 to 1,434,868 (mean = 25,888, median = 1022, 1 st quartile = 429.75, 3rd quartile = 2879.75). The quartiles do not indicate clear boundaries between the data, as there are small differences in the sample sizes. Hence, a boundary-based approach was followed instead of quartile-based: very small ( 20,000), as presented in Fig. 3 . The vast majority of studies ( n = 69, 73%) used condition-specific PROMs, such as orthopedic-specific Knee Injury and Osteoarthritis Outcome Score (KOOS) [ 22 ] or cancer-specific EORTC Core Quality of Life Questionnaire (QLQ-C30) [ 23 ]. In 31 papers (33%) condition-specific measures with generic questionnaires, for example EuroQol- 5D (EQ- 5D) [ 24 ]. Twelve papers (13%) used generic measures only. Out of 81 (86%) papers that reported the types of questionnaires used, 18 (22%) focused on physical health, 11 (14%) on mental health, and 52 (64%) on both.
Fig. 2 Year of publication of all 94 studies (top figure) and studies based on health domain (bottom figure) Fig. 3 Number of papers categorised based on sample sizes in each healthcare domain
Year of publication of all 94 studies (top figure) and studies based on health domain (bottom figure)
Number of papers categorised based on sample sizes in each healthcare domain
Thirty papers did not report the missingness in the dataset. Therefore, it is uncertain if the datasets in these studies did not have any missing data, or if missing data were not disclosed. Out of 64 papers (68%) which reported missing data, only 1 (2%) stated that there was no missingness in the dataset. Ten papers (16%) which reported having missing data did not report how the missingness was handled or addressed. Out of all papers, only 53 (56%) reported the technique for data imputation (Fig. 4 c). The 2 most common techniques were complete case analysis ( n = 16, 30%), and mean/median/mode imputation ( n = 15, 28%). Most papers ( N = 89, 95%) used classification as a prediction method. Fourteen (16%) of these did not provide any information about the class distribution. All papers which reported class distribution ( n = 75, 80%) performed binary classification. Out of these papers, only 11 (15%) had balanced classes (maximum imbalance ratio of 60:40 between the minority and majority class [ 25 ]). Sixty-four papers (68%) used dataset with imbalanced classes, 29 (45%) of which did not mention the class imbalance problem. Thirty-five papers (55%) acknowledged the issue but 13 (37%) of them left the data imbalanced. In total, 22 papers (23%) reported the need for balancing the classes, but there was inconsistency in the methods across papers (Fig. 4 ).
Fig. 4 Reporting of pre-processing and model development methods in the studies. Sub-figure a ) Frequency of hyperparameter tuning and values reporting. Sub-figure b ) Proportion of hyperparameter tuning techniques. Sub-figure c ) Missingness reporting and imputation in papers. Sub-figure d ) Handling class imbalance in studies
Reporting of pre-processing and model development methods in the studies. Sub-figure a ) Frequency of hyperparameter tuning and values reporting. Sub-figure b ) Proportion of hyperparameter tuning techniques. Sub-figure c ) Missingness reporting and imputation in papers. Sub-figure d ) Handling class imbalance in studies
Most papers ( n = 84, 89%) used multiple AI models for outcomes prediction. Forty papers (43%) used only traditional machine learning (ML) models, 5 (5%) only deep learning (DL) models, and 49 (52%) both ML and DL. The most frequently used models were regression models ( n = 61, 65%), including linear, logistic, ridge and LASSO regressions; boosting methods ( n = 53, 56%), including adaptive boosting, extreme gradient boosting and gradient boosting machine; random forest ( n = 50, 53%); artificial neural network ( n = 43, 46%), including single-or multi-layer perceptrons; and support vector machine ( n = 39, 41%) (Fig. 5 ). Most studies ( n = 74, 81%) applied AI models on data recorded in one time-point. The remaining studies trained their models on data collected in multiple time-points (Table 2 ). Out of these, 3 studies (3%) reported using models that process the temporal dependencies in the data, such as long-short term memory (LSTM) model [ 26 , 27 ], and recurrent neural network with gated recurrent units (GRU) [ 28 ]. Nine studies (12%) considered temporality through coding it in the feature sets, and 5 papers (7%) did not address temporality at all (Table 3 ). Half of the papers in this review ( n = 47, 50%) reported performing hyperparameter tuning, and out of these only 16 (34%) reported used hyperparameters (Fig. 4 a)).
Fig. 5 Frequency of algorithms used on datasets with very small (fewer than 300), small (300–700), medium (701–2000), large (2001–20000) and very large (more than 20,000) sample size Table 2 Machine learning and deep learning models used on data collected in one and multiple timepoints, ordered by number of publications Machine Learning (ML) One Timepoint Multiple Timepoints Regression ( n = 61) [ 29 – 39 ] [ 40 – 50 ] [ 40 , 51 – 60 ] [ 61 – 71 ] [ 72 – 75 ] [ 76 ] [ 77 – 86 ] [ 28 , 87 , 88 ] Boosting ( n = 53) [ 31 , 32 , 36 – 40 , 42 , 44 , 89 , 90 ] [ 45 , 48 , 49 , 51 – 53 , 91 – 94 ] [ 40 , 56 , 58 , 60 , 62 – 64 , 95 – 97 ] [ 2 , 33 , 66 , 69 , 72 , 73 , 75 , 98 – 100 ] [ 70 , 74 , 101 ] [ 76 ] [ 102 ] [ 28 , 77 – 79 , 81 – 84 , 103 , 104 ] Random Forest ( n = 50) [ 29 – 33 , 36 , 41 , 42 , 44 , 89 , 105 ] [ 45 – 48 , 50 – 53 , 56 , 92 , 106 ] [ 58 , 60 , 61 , 64 , 65 , 69 , 75 , 95 , 100 , 107 , 108 ] [ 62 , 63 , 71 , 73 , 74 , 97 , 101 ] [ 26 , 28 , 77 , 78 , 80 , 87 , 88 , 104 , 109 , 110 ] Support Vector Machine ( n = 39) [ 29 , 31 , 32 , 34 , 37 , 41 – 43 , 45 , 46 , 48 ] [ 26 , 50 – 52 , 56 , 57 , 92 , 96 , 111 , 112 ] [ 61 , 62 , 64 , 66 , 71 , 74 , 89 , 107 , 108 ] [ 102 ] [ 78 , 81 , 83 , 84 , 87 , 88 , 103 , 104 ] Decision Tree ( n = 24) [ 37 , 39 , 41 , 57 , 92 – 95 , 111 , 112 ] [ 2 , 58 , 63 , 74 , 96 , 98 , 100 , 101 , 108 ] [ 77 , 79 , 84 , 103 , 109 ] K-Nearest-Neighbours ( n = 13) [ 34 , 36 , 37 , 39 , 48 , 56 , 61 , 74 , 113 ] [ 26 , 83 , 87 , 104 ] Na¨ıve Bayes ( n = 12) [ 36 , 44 , 50 , 57 , 70 , 73 , 74 , 94 , 108 ] [ 26 , 104 , 109 ] Voting Classifier ( n = 4) [ 37 , 44 ] [ 28 , 82 ] Discriminant Analysis ( n = 4) [ 34 , 38 , 94 , 108 ] None reported Classification and Regression Tree ( n = 2) [ 34 , 46 ] None reported Super Learner ( n = 2) [ 62 , 69 ] None reported Other ML Methods ( n = 8) Wide and Deep [ 49 ] Stochastic Gradient Descent [ 61 ] Bagging [ 63 ] Bayesian Updating Algorithm [ 66 ] Graphical Gaussian Model [ 67 ] Multivariate Adaptive Regression Spline [ 41 ] Hierarchical Gaussian Process [ 85 ] Autoregressive Integrated Moving Average [ 26 ] Deep Learning (DL) One Timepoint Multiple Timepoints Multilayer Perceptron ( n = 43) [ 29 , 30 , 35 – 37 , 39 – 42 , 45 , 90 ] [ 48 , 52 – 54 , 56 , 92 , 94 , 111 , 112 , 114 ] [ 40 , 57 , 59 , 61 , 64 , 66 , 68 , 96 , 107 , 115 ] [ 72 – 75 ] [ 76 ] [ 26 , 28 , 79 , 83 , 84 , 86 , 104 , 109 , 116 ] Recurrent Neural Network (RNN) ( n = 3) None reported Long-Short Term Memory [ 26 , 27 ] RNN with Gated Recurrent Units [ 28 ] Other DL Methods ( n = 4) Adaptive Neural Network [ 36 ] Stacking Algorithm [ 57 ] Bayesian Network Model [ 117 ] Adaptive Neural Network [ 27 ] In the square brackets we list the number of the cited paper, according to the reference lis Table 3 Methods of addressing temporality in time-series data Level of addressing temporality Method of addressing temporality Description of the method Sample size Frequency Addressed by the model Recurrent Neural Network (RNN) Data transformed to 3D array and fed in the LSTM model [ 26 ] Events encoded by Adaptive Net, pooled by LSTM model [ 27 ] 823 9,500 Weekly Irregular a RNN with GRU, considering each treatment as timestep [ 28 ] 105,129 Irregular a Addressed in the features Measured change Change of measurement from baseline [ 84 ] Change in mean measurements from baseline [ 88 ] Mean daily change from the 24-h baseline period [ 87 ] Change in symptom severity from previous report [ 78 ] 245 31,700 116 34 Every 90 days Daily Twice a day Weekly Binary outcome Variable indicated if a report is followed by exacerbation event [ 80 ] Occurrence of symptom in any day of a time window [ 85 ] 2,374 182,991 Daily Daily (3 days) Dichotomised score one week following the prediction date [ 104 ] 210 Weekly Feature for each timeline Score added as an input feature at every measurement [ 82 ] Created a timeline of best overall responses (BORs) [ 103 ] 83 31 Weekly Weekly Not considered Model for each timeline Treated the 2- and 8-week measures as if assessed at baseline [ 81 ] Three models that used 7, 14, and 21 days as inputs [ 116 ] 1,003 20 3 time-points 3 time-points Selected 1 value for analysis If patient had multiple follow-up events, the first was chosen [ 77 ] The assessment with the highest overall score was used [ 79 ] 494 11,761 Irregular a Bi-weekly Score was updated at each assessment [ 86 ] 212,615 Irregular a a Irregular measurements indicate that the reports were completed at any clinical event that occurred Three papers which used time-series data did not report how the temporality was addressed [ 83 , 109 , 110 ], and are not included in this table
Frequency of algorithms used on datasets with very small (fewer than 300), small (300–700), medium (701–2000), large (2001–20000) and very large (more than 20,000) sample size
Machine learning and deep learning models used on data collected in one and multiple timepoints, ordered by number of publications
[ 29 – 39 ]
[ 40 – 50 ]
[ 40 , 51 – 60 ]
[ 61 – 71 ]
[ 72 – 75 ]
[ 76 ]
[ 77 – 86 ]
[ 28 , 87 , 88 ]
[ 31 , 32 , 36 – 40 , 42 , 44 , 89 , 90 ]
[ 45 , 48 , 49 , 51 – 53 , 91 – 94 ]
[ 40 , 56 , 58 , 60 , 62 – 64 , 95 – 97 ]
[ 2 , 33 , 66 , 69 , 72 , 73 , 75 , 98 – 100 ]
[ 70 , 74 , 101 ]
[ 76 ]
[ 102 ]
[ 29 – 33 , 36 , 41 , 42 , 44 , 89 , 105 ]
[ 45 – 48 , 50 – 53 , 56 , 92 , 106 ]
[ 58 , 60 , 61 , 64 , 65 , 69 , 75 , 95 , 100 , 107 , 108 ]
[ 62 , 63 , 71 , 73 , 74 , 97 , 101 ]
[ 29 , 31 , 32 , 34 , 37 , 41 – 43 , 45 , 46 , 48 ]
[ 26 , 50 – 52 , 56 , 57 , 92 , 96 , 111 , 112 ]
[ 61 , 62 , 64 , 66 , 71 , 74 , 89 , 107 , 108 ]
[ 102 ]
[ 37 , 39 , 41 , 57 , 92 – 95 , 111 , 112 ]
[ 2 , 58 , 63 , 74 , 96 , 98 , 100 , 101 , 108 ]
Wide and Deep [ 49 ]
Stochastic Gradient Descent [ 61 ] Bagging [ 63 ]
Bayesian Updating Algorithm [ 66 ]
Graphical Gaussian Model [ 67 ]
Multivariate Adaptive Regression Spline [ 41 ]
Hierarchical Gaussian Process [ 85 ]
Autoregressive Integrated Moving Average [ 26 ]
[ 29 , 30 , 35 – 37 , 39 – 42 , 45 , 90 ]
[ 48 , 52 – 54 , 56 , 92 , 94 , 111 , 112 , 114 ]
[ 40 , 57 , 59 , 61 , 64 , 66 , 68 , 96 , 107 , 115 ]
[ 72 – 75 ]
[ 76 ]
Long-Short Term Memory [ 26 , 27 ]
RNN with Gated Recurrent Units [ 28 ]
Adaptive Neural Network [ 36 ]
Stacking Algorithm [ 57 ]
Bayesian Network Model [ 117 ]
In the square brackets we list the number of the cited paper, according to the reference lis
Methods of addressing temporality in time-series data
823
9,500
Weekly
Irregular a
Change of measurement from baseline [ 84 ]
Change in mean measurements from baseline [ 88 ]
Mean daily change from the 24-h baseline period [ 87 ]
Change in symptom severity from previous report [ 78 ]
245
31,700
116
34
Every 90 days
Daily
Twice a day
Weekly
2,374
182,991
Daily
Daily (3 days)
83
31
Weekly
Weekly
1,003
20
3 time-points
3 time-points
494
11,761
a Irregular measurements indicate that the reports were completed at any clinical event that occurred
Three papers which used time-series data did not report how the temporality was addressed [ 83 , 109 , 110 ], and are not included in this table
The evaluation metrics varied across the studies. Area under the curve (AUC) was most commonly used ( n = 60, 64%), and 32 (53%) of papers used this value to assess model performance with imbalanced classes. Other frequently used performance metrics were recall, also known as sensitivity ( n = 44, 47%), accuracy ( n = 43, 46%), and specificity ( n = 34, 36%). The majority of the studies used multiple performance metrics ( n = 83, 88%). Variable importance analysis was performed by 64 studies (68%) 61 of which (95%) reported PROMs data being valuable for prediction. Seventy-nine papers (84%) provided information on the best performing model. Regression models were the most frequently selected as best-performing algorithms ( n = 24, 30%), followed by boosting methods ( n = 20, 25%), random forest ( n = 10, 13%) and neural network ( n = 10, 13%).
No studies reported that the developed methods had been applied in the clinical practice. Although we acknowledge, that in such multidisciplinary research clinicians are generally involved in the study design, only 3 papers (3%) explicitly mentioned how clinicians contributed to the model development. They helped selecting input variables [ 30 , 111 ], or creating a testing set [ 100 ]. No papers mentioned patients involvement in the model development or any part of the study design. The majority of papers reported age ( n = 63, 67%) and gender ( n = 60, 64%) of study participants and only 24 (26%) studies reported ethnicity. Papers were classified into 3 different categories, inspired by a previously conducted scoping review [ 118 ]: internal validation (one source of data used for training and validation, including cross-validation or holdout sample for validation on unseen data from the same dataset), external validation (the model developed on one dataset and then tested/validated on a completely new (i.e. external) dataset) or deployment (”integrated into a prototype application, and evaluated for its feasibility in clinical workflows”[ 118 ]). Based on these definitions, 81 papers (86%) were in the internal validation stage, 10 (11%) completed external validation, and 3 papers (3%) were in the deployment stage.
Materials
The methodology of this scoping review was based on the Joanna Briggs Institute (JBI) guidance [ 18 ]. The review protocol is available on Open Science Framework [ 19 ]. The completed PRISMA checklist for scoping reviews [ 20 ] is added in the Supplementary Materials Fig. 1 and 2.
The databases used to search for relevant papers were: Web of Science, IEEE Xplore, ACM Digital Library, Cochrane Central Register of Controlled Trials, Medline and Embase. These databases were selected to include the variety of fields publishing papers on AI in medicine, covering both medical and engineering aspects. The keywords adapted to each database are listed in Supplementary Materials Table 1. Initially, the limited search of Web of Science and Medline was conducted to analyse and approve the keywords. The finalised search of all the records across all the databases was completed on the 7 th November 2023. The reference lists of all relevant reports were also screened. All studies identified through the search strategy were exported to Endnote citation management system.
The inclusion and exclusion criteria followed the Population/Concept/Context (PCC) framework [ 18 ] and are described in Table 1 . The participants in the papers included in the review were patients, whose symptoms and quality of life data were recorded using various PROMs. These can include mobile applications, or surveys completed either online or in a clinic. The type of data can be collected through either validated and widely used PROMs, or any other patient self-reports. Papers reporting the use of PROMs data as both a predictive and predicted variable were included in the study. If PROMs data were only used as a predicted variable, and not included as inputs, the reference was excluded. The concept was the methods of AI used for predictions of the patients’ outcomes. Papers that explicitly mentioned use of AI or Machine Learning methods were included. Any papers using AI models for purposes other than prediction were excluded. Papers that reported prediction of patient outcomes in the healthcare context were included in the analysis. These outcomes should belong to the categories of patient-related outcomes identified by Kersting et al. (2020) [ 21 ], presented in Table 1 . The broad understanding of healthcare context allowed focusing on the AI used for various medical reasons.
Table 1 Inclusion and Exclusion criteria for the study selection Category Inclusion criteria Exclusion criteria Input variables Data reported by patients using standardised PROMs; electronic data collection designed for self-reporting of symptoms or clinical outcomes. These can be used with combination of different non-patient reported data Only not patient-reported data, e.g., data reported through a clinician, recorded, written down during an appointment, data reported on online forums/social media, or data from physiological measurements Output variables Any data describing patient outcomes: sleep behaviour, coping and self-efficacy, healthcare utilisation, body image perception, function, communication skills, reliability of diagnosis and therapy, optimal support, confidence in therapy, satisfaction, cognitive performance, treatment decision, disease control, daily activities, reoperation, mental health, quality of life, mobility, co-morbidities, pain, survival, adverse events, and symptoms Output that does not relate to health or healthcare, or papers which did not aim to predict outcomes Models Machine learning or deep learning predictive models Statistical models (e.g., only regression analysis) Paper type Primary research reported in English language Abstracts only, theses, dissertations, letters to editors, guidelines, commentaries, introductions, papers published not in English, or review papers
Inclusion and Exclusion criteria for the study selection
All duplicates found in the databases were removed automatically in Endnote. Titles and abstracts of the papers were screened by a researcher and re-selected based on inclusion and exclusion criteria, presented in Table 1 . Full texts of articles admitted to the study were assessed against the inclusion and exclusion criteria again. The researcher’s approach was validated through a second reviewer, who repeated scanning through 10% of abstracts, selected full texts and compared their results with the first reviewer. The validation showed high consistency between the reviewers’ decisions, as out of 218 validated papers, 184 (84.4%) were consistently selected or rejected. Therefore, no further validation was performed.
The data was extracted from all papers selected for this review. Extracted and summarised information for each included paper is presented in Supplementary Materials Tables 1 and 2. A second reviewer extracted data from 10% of admitted papers for the purpose of validation, and the extracted information was compared and agreed between the reviewers. The summary of data was reported in tabular form in Excel spreadsheet and presented in a narrative form in this review.
The extracted and analysed data included: Study characteristics (country and year of publication, healthcare domain, input PROMs variables used, types of PROMs, output variable types, and sample sizes) Data pre-processing (missingness in the datasets, missing data imputation techniques, class distribution, techniques for handling class imbalance) Model development (types of AI models used, frequency of AI models used, AI techniques for addressing temporality in data, hyperparameter tuning) Model evaluation (performance metrics used, variable importance, best-performing AI models) Clinical relevance and adoption (patients and clinicians involvement in the study design, validation and deployment stage of research, reporting of sociodemographic information)
Study characteristics (country and year of publication, healthcare domain, input PROMs variables used, types of PROMs, output variable types, and sample sizes)
Data pre-processing (missingness in the datasets, missing data imputation techniques, class distribution, techniques for handling class imbalance)
Model development (types of AI models used, frequency of AI models used, AI techniques for addressing temporality in data, hyperparameter tuning)
Model evaluation (performance metrics used, variable importance, best-performing AI models)
Clinical relevance and adoption (patients and clinicians involvement in the study design, validation and deployment stage of research, reporting of sociodemographic information)