Influence of Preprocessing Methods of Automated Milking Systems Data on the Prediction of Mastitis with Machine Learning Models

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract Missing data and class imbalance represent a hindrance to accurate prediction of rare events such as mastitis (udder inflammation). Various methods are susceptible to handle the problem, however, little is known about their individual and combined effects on the performance of ML models fitted to AMS (automated milking system) data for mastitis prediction. We apply imputation and resampling to improve performance metrics of classifiers (logistic regression, stochastic gradient descent, multilayer perceptron, decision tree and random forest). Three imputation methods: simple imputer (SI), multiple imputer (MICE) and linear interpolation (LI) were compared to complete cases. Three resampling procedures: synthetic minority oversampling technique (SOMTE), Support Vector Machine SMOTE and SMOTE with Edited Nearest Neighbours were compared. We evaluated different techniques by calculating precision, recall, F1 Score and compared models based on kappa score. Both imputation and resampling techniques improved models performance. Complete case analysis suited the Stochastic Gradient Descent (SGD) Classifier better than resampling or imputation (kappa=0.280). The Logistic regression (LR) performed better with SVMSMOTE rand no imputation (kappa= 0.218). The Random Forest (RF), Decision Tree (DT) and Multilayer Perceptron (MLP) performed better than SGD and LR and handled well class imbalance and missing values without preprocessing. We propose careful selection of the technique to handle class imbalance and missing value prior to subjecting data to ML model is crucial to attain best ML model performance.
Full text 165,000 characters · extracted from preprint-html · click to expand
Influence of Preprocessing Methods of Automated Milking Systems Data on the Prediction of Mastitis with Machine Learning Models | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Influence of Preprocessing Methods of Automated Milking Systems Data on the Prediction of Mastitis with Machine Learning Models Kashongwe B.O., Kabelitz T., Amon T., Ammon C, Amon B., Doherr M. This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4629327/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Missing data and class imbalance represent a hindrance to accurate prediction of rare events such as mastitis (udder inflammation). Various methods are susceptible to handle the problem, however, little is known about their individual and combined effects on the performance of ML models fitted to AMS (automated milking system) data for mastitis prediction. We apply imputation and resampling to improve performance metrics of classifiers (logistic regression, stochastic gradient descent, multilayer perceptron, decision tree and random forest). Three imputation methods: simple imputer (SI), multiple imputer (MICE) and linear interpolation (LI) were compared to complete cases. Three resampling procedures: synthetic minority oversampling technique (SOMTE), Support Vector Machine SMOTE and SMOTE with Edited Nearest Neighbours were compared. We evaluated different techniques by calculating precision, recall, F1 Score and compared models based on kappa score. Both imputation and resampling techniques improved models performance. Complete case analysis suited the Stochastic Gradient Descent (SGD) Classifier better than resampling or imputation (kappa=0.280). The Logistic regression (LR) performed better with SVMSMOTE rand no imputation (kappa= 0.218). The Random Forest (RF), Decision Tree (DT) and Multilayer Perceptron (MLP) performed better than SGD and LR and handled well class imbalance and missing values without preprocessing. We propose careful selection of the technique to handle class imbalance and missing value prior to subjecting data to ML model is crucial to attain best ML model performance. Animal Science oversampling undersampling missing value imputation dairy cows machine learning performance analysis Figures Figure 1 Figure 2 Figure 3 1. Introduction Mastitis is the most costly and frequent disease in dairy farming, contributing to considerable (about 125€ per cow and year) and recurring costs incurred through reduction of milk quantity and quality as well as the reproductive performance and longevity of cows [ 1 – 3 ]. The disease can be caused by a wide diversity of pathogens, but is dominated by only few species. The most important mastitis causing bacterial species are Streptococcus agalactiae , Streptococcus dysgalactiae , Streptococcus uberis , Staphylococcus aureus , and Escherichia coli [ 4 ]. The subclinical form of mastitis, without overt symptoms, is by far more prevalent than the clinical form and causes high economic losses [ 2 , 5 ]. Simultaneously, subclinical mastitis is, due to the low level of symptoms, difficult to detect. Therefore, modern data analysis tools like machine leaning could help to identify subclinical mastitis cases more often and accurate. Managing mastitis is even a bigger challenge in large-scale dairy farms despite the use of automated milking systems (AMS) that has steadily increased in Germany and across European countries over the last years [ 6 ]. The AMS uses sensors to collect data such as milk yield and milk components and through management programs, alerts are given in case of variations of either milk production (milk yield and milk flow), electric conductivity, somatic cell counts (SCC), milk temperature or a combination of all these parameters as an indication for mastitis occurrence [ 7 , 8 ]. The data collected at each milking or through monthly milk control programs, is routinely used to help farmers for better decision making on production, reproduction and health. This is specifically the case for mastitis predictions based on milk yield, milk parameters (conductivity, SCC, blood in milk, temperature) and cow characteristics [ 7 – 9 ]. For this prediction purpose, several machine-learning (ML) approaches have been used that aim to improve monitoring of udder health status in general or mastitis specifically, whether subclinical or clinical. [ 10 ] compared eight ML models and achieved prediction accuracy above 75% for all of them in a binary classification with 200,000 cells/ml SCC as threshold for positivity. [ 11 ] trained random forest models to predict mastitis infection patterns in a binary classification where the predictor was contagious vs environmental mastitis or environmental lactation vs environmental dry period. They obtained 98% prediction accuracy. Post et al. [ 12 ] applied ML models to a group of animals with historical records of diseases and achieved higher prediction accuracy than when the model was applied to the whole population. Findings from other studies using similar methodologies show also high prediction accuracy [ 13 , 14 ]. Despite these encouraging results, the application of mastitis prediction ML models in real life conditions remains limited because of a discrepancy between the performance on the training data sets and the actual data. The nature of data recording by sensor systems and low occurrence of the disease appear as major reasons for the difference [ 15 ]. Indeed, farm sensor data present in general two types of challenges that make them difficult to handle by prediction algorithms. First, the data is often noisy with missing values, outliers and skewed values, occurring because of sensor failure during signal transmission [ 16 ]. The missing values or wrongly recorded values fall into the type of either missing values completely at random or missing values at random. Missing values together with a clear definition of positive cases represent a major hindrance for ML algorithms trained on ‘experimental data sets’ to be used under farm conditions [ 17 ]. To handle the problem, it is common in practice to delete missing values completely or at least to apply methods such as list wise deletion but less common is the reporting of the magnitude of missing values or the use of missing data handling methods [ 18 ]. Although working without missing values is convenient, it only produces reliable estimates in limited situations where missing values occur completely at random and only on the dependent variable. In other situations, this result in severely biased estimates, not to mention the potential waste of information in the omitted data and the low practical application of the obtained results [ 19 ]. This is of particular importance in disease prediction where metrics obtained from the training data sets need to be applied in real-life situations [ 17 ]. Indeed, it is very common to have large amounts of missing values in sensor-generated datasets [ 16 ]. Various techniques have been developed to deal with the challenge of missing values in large datasets. The common imputation methods are simple imputation, multiple imputation and linear interpolation. The simple imputation method employs replacement of missing values with mean, median or mode values [ 18 ]. This method is largely proposed for its computational convenience although, in many cases, the results and conclusions are not sensible or generalizable [ 20 ]. The multiple imputation uses the MICE (multiple imputation with chained equations) algorithm, which is a Markov Chain Monte Carlo method that imputes incomplete data in a variable-by-variable way starting with a random draw of the observed data. For instance, A first regression of the first variable with missing values applied to all other variables provided that the rows have observations for the variable of interest. Then missing values in the variable of interest are replaced by simulated draws from its posterior predictive distribution. The process is repeated for all other variables with missing values in turn: this is called a cycle. This process is repeated several times to generate a single imputed dataset and the whole process is repeated 3 to 5 times to obtain stable results [ 21 ]. Although recognised for its robustness, the method suffers the limitation of lack of theoretical rationale [[ 19 , 21 , 22 ]. Linear interpolation estimates the value of the missing data based on the two data points adjacent to the missing one in a one dimensional data sequence [ 23 ]. It is reputed to perform well on time dependent data and on dataset with small to moderate missing values between adjacent points [ 22 ]. The second major hurdle when training models on AMS data to predict mastitis is the class-imbalance between positive and negative cases [ 14 ]. Although frequently observed in dairy farms, mastitis is a rare occurrence when data resolution is increased to either daily basis or animal level or both. This imbalance causes a bias when fitting standard learning classifiers, reflected in their inability to predict correctly the minority class, despite sometimes achieving high prediction accuracy [ 24 ]. Johnson & Khoshgoftaar [ 25 ] noted that the total number of the minority class is more important than the percentage of imbalance. Various methods to handle class imbalance are reported in literature to improve disease prediction [ 14 , 24 ]. Johnson & Khoshgoftaar [ 25 ] categorized them into three groups, data-level methods, algorithm level methods and hybrid methods. Data level methods change the dataset structure by either reducing the majority class (under-sampling) or increasing the minority class (oversampling) or both to achieve a more balanced class distribution [ 24 ]. Among popular resampling techniques, is the Synthetic Minority Oversampling Technique (SMOTE) which produces synthetic samples by interpolating minority samples with their k Nearest Neighbors. The algorithm seems to be improved by taking into consideration the minority class lying along the borderline, hence expanding the minority class area towards the side of the majority class where only few instances of majority class are found [ 26 ]. However, oversampling techniques may lead to overfitting. The Random Under-Sampling is among the first under-sampling techniques developed and works by discarding random samples in the majority class. The technique has been improved with several techniques using nearest neighbours to reduce instances in the majority class. The Edited Nearest Neghbours (ENN) tests every instance with the rest of samples using k-NN and disqualifies incorrectly classified samples. Under-sampling methods have the disadvantage of discarding information that may be useful. Techniques combining over-sampling and under-sampling have been developed to overcome limitations of individual methods. The SMOTE-ENN combines the SMOTE and edited nearest neighbours for under-sampling [ 24 ]. Studies on mastitis prediction with machine learning classifiers use the above mentioned data (pre)processing techniques almost interchangeably, making comparison and evaluation of effectiveness across studies more complex. Hence, we investigated whether resampling or imputation or both techniques were critical to influence the prediction performance of ML classifiers. We used several metrics including accuracy, F1 score, precision, recall and kappa score. 2. Materials and Methods 2.1. Data Collection The data set included records of 232 cows and 75217 milking events, for which daily milk yield, electric conductivity at quarter and cow levels and somatic cell counts were recorded. Data were collected between January 2015 and September 2017 from a dairy farm that uses automated milking systems as well as a dairy herd management program. The Lely Astronaut system (Lely Industries N.V., Maassluis, the Netherlands) equipped with in-line sensors for electric conductivity (EC) and SCC was used to milk cows and monitor their performance. Hence, mastitis cases used in this study referred to the alarm raised by the AMS due to changes detected in milk. Descriptions as abnormal milk (n = 54), mastitis (n = 398), high conductivity (n = 14), watery milk (n = 2) were classified as positive cases, while instances where no alarm was raised were classified as negative cases (n = 74749). Conductivity had an average of 68.44 ± 3.42 µS/cm, SCC had an average of 93.86 ± 188.29 x 10 3 cells/ml. The dataset contained missing values for predictors as presented in Table 1 . Table 1 Description of the dataset. Missing values Total cases Variable 0 75217 Alarm* 36028 39189 EC_FL 36028 39189 EC_FR 36315 38902 EC_BR 36142 39075 EC_BL 37326 37891 EC_ALL 58669 16548 SCC 29 75188 Milk yield *refers to both presence (value = 1) and absence (value = 0) of alarm, EC: electric conductivity, FL: front left, FR: front right, BR: back right, BL: back left, ALL: average all quarters, SCC: Somatic Cell counts. 2.2. Data Preparation 2.2.1. Libraries and packages used for data preparation and modeling All analyses were performed in Python, using numpy, pandas and scikit learn libraries. We used the imblearn and imputer packages to perform the resampling and missing value imputation, respectively. We selected from scikit learn library five common supervised learning models: Stochastic gradient descent (SGD), logistic regression (LR), decision tree (DT), random forest (RF) and multilayer perceptron (MLP). We imported confusion matrix, roc_curve and auc from sklearn metrics to compute precision metrics and plot the ROC curves. Plotting of ROC curves was aided by the pyplot package of matplotlib library [ 20 ]. Finally, we sued statsmodels library to compare the metrics of the models tested and determine for each model type whether resampling, imputation or both influenced the observed performance. We performed five fold cross-validation using k-fold strategy in the Scikit-learn GridsearchCV to obtain optimal parameters search and hyperparameter tuning for each model. 2.2.2. Cross-validation and model tuning The dataset was loaded and checked for inconsistencies before further processing. inconsistencies such as missing dates, missing all values across one observation, outliers that could have resulted from erroneous measurement or recording and the corresponding data were cleaned. The data was then split into training and test set at a ratio of 80:20. The test set was left out of further processing. The training set was further split into training and validation set and subjected to further processing. Two directions were followed for processing. On one hand, all missing values were deleted from the dataset to remain only with complete cases (CC) for which no missing imputation was required. On the other hand, data with missing values were processed with one of three selected imputation techniques: simple imputer (SI), multiple imputer (MICE) and linear interpolation (LI). This concerned both the training and validation set for which a ratio of 60:40 was used for complete case analysis and 80:20 for the other imputation techniques. Model tuning consisted of finding the best parameters for each model resulting from the above mentioned imputation methods. A full description of the hyperparameter tuning and model fitting can be found in the supplementary file 2. In summary, for the stochastic gradient descent models parameters included loss function (hinge, log_loss or modifier huber), penalty criteria (l2, l1, elasticnet) and the value of alpha (0.0001 to 0.1). Additionally 5 fold cross-validation was performed.The best fitted loss functions and penalty criteria were ‘hinge’ and ‘elasticnet’ for CC, MICE, and LI, while it was ‘log_loss’ and ‘L2’ for SI. The best values of alpha were 0.1, 0.01 and 0.001 for CC, SI and MICE and LI, respectively (See Supplementary file 2). The grid search cross-validation for logistic regression included solver (lbfgs, liblinear), penalty criteria (l2, l1, elasticnet), C (0.1 to 100) and 5 fold cross-validation. The lbfgs solver and the L2 penalty were selected for all datasets. The C value of 0.1 selected for all datasets except for the one from LI whose C value was 10. The grid search cross-validation for decision tree models included the criterion (gini, entropy), max depth (None, 2–20), max features (None, sqrt, log2, 0.2 to 0.8), splitter (best or random). The cross-validation was set to 5 folds. The selected criterion (entropy) and max depth (n = 4) were similar for CC and MICE, while it was ‘gini’ for SI and LI. The max depth was n = 6 for SI and n = 5 for LI and n = 4 for MICE (See Supplementary file 2 ). For random forest models, criterion (gini, entropy and log_loss), max depth (None, 2–20), max features (None, sqrt, log2, 0.2–0.8) and 5 folds cross-validation were included in the grid search. The criterion (entropy) was selected only for SI, while entropy was selected for the others. The max depth ranged from None for SI, to 20 for MICE and LI. No max features was selsected for CC and MICE, while SI had ‘sqrt’ and LI had 0.8. The grid search cross-validation parameters for multilayer perceptron models included activation (identity, logistic, tanh, relu), solver (lbfgs, sgd, adam), alpha (0.0001 to 0.01), learning rates (constant, invscaling, adaptive) and cross-validation (cv = 5). The CC and LI had the best activation with identity while the ‘relu’ activation was selected for SI and MICE. The SI and LI datasets had also the learning rate ‘invscaling’ while it was adaptive for CC and constant for MICE. The value of alpha was set at 0.05 for all datasets except for the CC for which the best value was 0.0001 (See supplementary file 2). 2.3. Missing Values Imputation The simple imputer was implemented for the replacement of missing values of individual variables with the mean or median value of the specific variable. The multivariate imputer (implementing the Multiple Imputer Chain Equation (MICE) in Python) was applied using the iterative imputer module of scikit learn that replaced the missing value of individual variable (column) with the value from specific and other variables for the same observation (row). The linear interpolation (LI) was implemented using the interpolate module. We set the method at linear for the replacement of missing values of individual variables based on both previous (limit direction backward) and following (limit direction forward) values to predict the missing value using linear regression. Hence, two datasets with only complete cases and with missing replaced values would form the basis for data processing. They had imbalance ratios of 53.76 and 156.52, respectively (Table 2 ). Table 2 Imbalance ratio of initial datasets. Positive cases Negative cases Imbalance ratio 87 4677 53.76 Complete cases 382 59791 156.52 Data with replaced missing values 2.4. Resampling Methods The two sets of data were submitted to three resampling methods: the Synthetic Minority Oversampling Technique (SMOTE), the SMOTE technique combined with Edited Nearest Neighbours (SMOTEEN) and the SMOTE technique combined with Support Vector Machine (SVM) classifier SVMSMOTE (Table 3 ). The SMOTE technique oversampled the minority class without altering majority class. The SMOTEEN technique not only oversampled the minority class but also under sampled the majority class. The SVMSMOTE technique oversampled the minority class along the borderline and used Support Vector Machine (SVM) classifier to predict new cases. Table 3 Imputed datasets resampled with SMOTE, SMOTEEN, and SVMSMOTE methods. Positive Negative Imbalance ratio Complete cases 87 4677 53.76 No resampling 4677 4677 1.00 SMOTE 4459 4134 0.93 SMOTEEN 4677 4677 1.00 SVMSMOTE Simple Imputer 382 59791 156.52 No resampling 59791 59791 1.00 SMOTE 58100 58387 1.00 SMOTEEN 59791 59791 1.00 SVMSMOTE Multiple Imputer 382 59791 156.52 No resampling 59791 59791 1.00 SMOTE 58674 58519 1.00 SMOTEEN 59791 59791 1.00 SVMSMOTE Linear Interpolation 382 59791 156.52 No resampling 59791 59791 1.00 SMOTE 57820 59521 1.03 SMOTEEN 59791 59791 1.00 SVMSMOTE Paramters used for the application of resampling methods are presented in Table 4 and were similar to those reported in literature [ 27 – 29 ]. Table 4 Parameters settings used for the resampling methods Parameters Methods K-Neighbors = 5 SMOTE K (SMOTE) = 5, K (ENN) = 3 SMOTEEN K-Neighbors = 5, m_Neighbours = 10 SVSMSMOTE 2.5. Processing Method Evaluation and Comparison of Performance metrics After obtaining the resampled and /or imputed datasets, we split the data into training and validation sets using the ratio of 80:20 for the simple imputer, MICE, and linear interpolated data, while we used a ratio of 60:40 for the complete case analysis where the number positive cases were much fewer (Table 3 ). Ebrahimi et al [ 30 ] used a similar approach in their study on prediction of sub-clinical mastitis using machine learning models. The 16 datasets were fitted to five common machine-learning classifiers: Stochastic gradient descent (SGD), logistic regression (LR), decision tree (DT), random forest (RF) and multilayer perceptron (MLP). The performance of the classifiers was evaluated using accuracy, precision, recall, F1 score and kappa metrics. These were obtained from a confusion matrix (Table 5 ), where the correctly classified positive and negative cases are labelled true positive (TP) and true negative (TN). Positive cases incorrectly classified as negative are labelled false negative (FN), while negative cases incorrectly classified as positive are labelled false positive (FP). Table 5 Representation of a Confusion matrix. Predicted negative Predicted positive False negatives FN True positives TP Actual positive True negatives TN False positives FP Actual Negative We used accuracy, the area under the Receiver Operating Characteristic curve (ROC), precision, recall, F1 score and Cohen’s kappa score to evaluate the performance of the models. Accuracy is the most commonly used metric and the starting point to evaluate the performance of classifiers. Accuracy is the proportion of correct predictions (true positive, true negative) among all examined cases [ 7 ]. Accuracy = (TP + TN) / (TP + FP + TN + FN) (1) Sensitivity, recall, true positive rate TPR = TP/ (TP + FN) (2) Positive predictive value, precision PPV = TP/ (TP + FP) (3) False positive rate FPR = FP/ (FP + TN) (4) Specificity, true negative rate TNR = TN/ (TN + FP) (5) Negative predictive value NPV = TN/ (TN + FN) (6) F1 score = 2 * TP / 2 * (TP + FP + FN) = 2 * (precision * recall) / (precision + recall) (7) However, accuracy score for unbalanced problems often provides an overoptimistic estimation of the classifier ability to predict the majority class [ 20 ]. Hence, the use of other performance metrics could strengthen the meaning of models performance obtained. The ROC plots the true positive rate and the false positive rate at various thresholds values. The F1 score is a weighted (harmonic) mean of sensitivity and precision. The Cohen’s Kappa value compares the classifiers performance to the probability that its performance only be based on chance. In general, predictions from models with Kappa values < 0.20 are considered poor, values between 0.21–0.40 are fair and from 0.40–0.60 moderate and above 61 substantial to almost perfect [ 21 ]. Thirty-nine out of 80 models tested for the current study had kappa 50. Thus, we ranked classifiers performance metrics first by Kappa value, then by f1 score and precision/ recall scores regardless of the accuracy score. We also examined the ROC curves of the best models to evaluate their performance at various thresholds. Finally, we assessed the contribution of resampling or imputation techniques or both on the prediction performance of the ML classifiers using AIC score and Residual deviance. An overview of the complete workflow performed in this study is shown in Fig. 1 . 3. Results 3.1. Performance Metrics of ML Models Trained on Data with Different Missing Values Imputation Techniques Results show that imputation techniques improved the prediction performance RF and DT models with SI (kappa = 0.43 and 0.36, respectively) higher than MI (kappa = 0.41 and 0.36, respectively), CC (kappa = 0.40 and 0.31, respectively) and LI (kappa = 0.27 and 0.26, respectively) (Fig. 2 ). The opposite was observed with SGD and LR where CC was higher than MI, SI and LI, respectively. The linear interpolation for MLP performed better (kappa = 0.27) than the other techniques. 3.2. Performance Metrics of ML Models Trained on Data with Different Resampling Techniques Results for resampling methods indicate that ensemble models perform better (kappa scores > 0.30) than MLP (kappa between 0.15 and 0.21) and discriminative classifiers (kappa < 0.12).Random forest classifier had higher kappa score for all resampling methods but not for the data without resampling for which Decision Tree had the highest (kappa = 0.41) (Fig. 3 ). 3.3. Rankings of the Prediction Scores of Fitted Machine Learning Models with Both Resampling and Imputation Methods* The ensemble models had highest performance with SI and MICE resampled by SVMSMOTE, SMOTE, and No resampling. They all produced fair to moderate kappa scores (> 0.35). Discriminative classifiers (SGD, LR) had lower kappa scores (< 0.30). MLP models had fair prediction accuracy (kappa = 0.229–0.332) with LI, CC and MICE with or without resampling methods (Table 6 ). Table 6 Performance metrics of best-ranked ML models. Overall Rank** Kappa Recall Precision F1Score Accuracy Resampling Imputation Stochastic gradient descent (SGD)and Logistic regression (LR) models 27 0.280 0.957 0.979 0.967 0.957 No resampling CC (SGD) 38 0.225 0.940 0.979 0.957 0.940 SVMSMOTE CC (SGD) 39 0.218 0.944 0.978 0.959 0.944 SVMSMOTE CC (LR) Multilayer perceptron (MLP) models 19 0.332 0.995 0.996 0.996 0.991 SMOTE LI 22 0.322 0.996 0.996 0.996 0.992 No resampling LI 29 0.265 0.986 0.997 0.992 0.984 SMOTEEN LI 33 0.261 0.961 0.992 0.976 0.954 No resampling CC 34 0.261 0.961 0.992 0.976 0.954 SMOTE CC 37 0.229 0.984 0.997 0.990 0.981 SVMSMOTE MICE Decision tree (DT) models 1 0.465 0.997 0.997 0.997 0.994 No resampling SI 2 0.455 0.997 0.997 0.997 0.994 No resampling MICE 10 0.416 0.980 0.981 0.981 0.980 No resampling CC 12 0.404 0.993 0.998 0.995 0.991 SVMSMOTE MICE 15 0.367 0.994 0.997 0.995 0.991 SVMSMOTE SI Random forest (RF) models 3 0.453 0.998 0.996 0.997 0.995 SMOTE SI 4 0.453 0.998 0.996 0.997 0.995 SVMSMOTE SI 5 0.446 0.998 0.997 0.997 0.994 SMOTE MICE 6 0.445 0.998 0.997 0.997 0.994 SVMSMOTE MICE 7 0.442 0.998 0.989 0.993 0.987 No resampling CC *top 5 models with kappa > 0.20 were selected for each model category. **full list of ranked model in supplementary file 3.4. Evaluation of Effects of Resampling Technique and Imputation of ML Models Performance Results of method importance to model predictions are presented in Table 6 . All classifiers except the RF performed better when resampling was associated with imputation. AIC and null deviance values are smallest for these models. RF models with resampling only (AIC = 107.2) or a combination of resampling and imputation (Residual deviance = 4609) were the best. Table 6 Effects of Resampling techniques and imputation on ML models performance. Residual deviance Null deviance AIC Models SGD 14.33 327.4 107.2 Resampling + imputation 290.2 327.4 377.1 Imputation 45.78 327.4 132.7 Resampling LR 21.99 294.1 116.9 Resampling + imputation 259.9 294.1 348.8 Imputation 49.21 294.1 138.1 Resampling DT 6.092 35.05 106.7 Resampling + imputation 12.87 35.05 107.5 Imputation 29.03 35.05 123.7 Resampling RF 4.609 39.91 101.7 Resampling + imputation 35.69 39.91 126.8 Imputation 8.945 39.91 100.1 Resampling MLP 152.3 405.4 242.5 Resampling + imputation 380.8 405.4 465 Imputation 175.4 405.4 259.5 Resampling Overall 546.6 1146 966.3 Resampling + imputation 1060 1146 1474 Imputation 611.3 1146 1025 Resampling 4. Discussion The study demonstrated the influence of resampling and imputation techniques on the prediction performance of three types of machine learning models trained to detect mastitis alarms from automated milking system data. Features included quarter and cow level conductivity and in-line somatic cell count from a conventional dairy farm in Germany. Three types of classifiers we evaluated are classical discriminative classifiers (SGD and LR), ensemble classifiers (DT and RF) and a neural network based classifier (MLP). We found that imputation techniques improved the prediction performance for RF and DT models among which SI (kappa = 0.43 and 0.36, respectively) was higher than MI (kappa = 0.41 and 0.36, respectively), CC (kappa = 0.40 and 0.31, respectively) and LI (kappa = 0.27 and 0.26, respectively). This trend was confirmed by the ROC curves for RF and DT that had higher TPR (> 90%) and lower FPR (< 10%) for RF with imputation methods compared to CC (Appendix 1). DT models also had better performance than CC. Poulos & Valle [ 22 ], testing RF models at various levels of missing data, compared the performance of the models with complete case or missing imputation and reported also a higher performance of ensemble models to which missing values were imputed than CC. In the current study, linear interpolation did not perform as good as simple imputation or multiple imputation for random forest and decision tree models. According to [ 15 ], this could be due to no data segregation before applying linear imputation to datasets. They suggest that LI methods estimate the value of the missing data based on two adjacent data points in a one- dimensional sequence. Hence, for datasets where many consecutive data points are missing, such as in AMS data, the performance of LI may not be optimal. Yet, ensemble methods work by segregating data into similar packets small enough to identify their inherent patterns in the terminal nodes. For example, decision trees having two kinds of nodes, determine each leaf node that has a class label with a majority vote of training examples reached by the leaf. Further, they treat each internal node to represent a question on features that will be branching out according to the answers found. Hence they split leaves of a tree until questions are exhausted [ 31 ]. Therefore, it could be that these intrinsic characteristics of the ensemble models and LI led to lower performance than other imputation techniques or complete cases. Following [ 23 ] approach, it could be beneficial to segregate the data prior to submitting it to LI for better results. This may not be useful for ensemble models, reputed robust enough to yield good performance with simple imputation techniques and sometimes without imputing missing values as explained above [ 32 ]. Indeed, two of the top 10 models in this study were RF and DT without missing imputation or resampling (kappa = 0.442 (No7), and 0.416 (No10), respectively). A different trend was observed with SGD and LR where CC had higher scores than MI, SI and LI, respectively. Although the overall Kappa scores for these models were lower than ensemble models. The ROC curves with higher or similar performance for CC than imputation techniques regardless of the resampling techniques confirm this (Appendix 1). Mukaka et al. [ 33 ] also found better results for CC analysis compared to imputation techniques for binary outcomes with LR and recommended complementing the use of imputation techniques with CC analysis. Other authors [ 25 , 26 ] found the opposite and suggested imputation was better than CC analysis. On one hand this can be explained by the fact that imputation techniques, especially for multiple imputation, increase the variability in the outcome values that inflates the standard error of the effect size estimate, probably caused by a random component added to the missing outcome values [ 20 , 21 , 33 ]. On the other hand, the difference can also be attributed to the mechanisms of occurrence of missing values for which [ 33 ] provided an in depth analysis and suggested a thorough examination before deciding on the imputation method to apply. The linear interpolation for MLP had better performance metrics (kappa = 0.27) and ROC characteristics than the other imputation techniques for the same model (Appendix 1). The method relies mostly on time dependent missing value imputation as opposed to inter-attribute correlations employed by other imputation techniques [ 34 ]. For this reason LI is especially efficient for time series and has been reported to improve the performance of neural network based classifier in other studies [ 35 , 36 ]. The performance of ML models from resampled datasets showed a similar trend with the missing imputation. The RF models had highest metrics, followed by DT, MLP, LR and SGD, respectively. Resampling data with SVMSMOTE seemed to result in better performance when subjected to SGD and LR models reported to perform better with more balanced datasets [ 27 ]. The MLP and RF models had better classification performance with SMOTE which is consistent with the findings reported in other studies [ 27 , 37 ]. The evaluation of model fit revealed that both resampling and missing imputation are relevant to explain the performance of most of the tested ML models. The MLP model had fair performance with linear interpolation imputation without resampling. This is in line with reports that the method improves the performance of neural network based models, hence could be applied to these types of ML models without resampling [ 35 , 36 ]. The same behavior was observed for simple and multiple imputated data fitted to decision tree models that resulted in good performance (kappa = 0.465 and 0.455 respectively). The SVMSMOTE resampling method for SGD and LR performed better without imputation than the data where missing values were imputed. This suggests that the improvement in the class imbalance between majority and minority class achieved through borderline classification of SVMSMOTE outweighs the need for replacing missing values prior to analysis for these classifiers [ 27 , 38 , 39 ]. Indeed, studies have applied resampling methods without imputation with satisfactory prediction performance. Random Forest, DT and to some extent MLP had models with fair to good performance (kappa = 0.322 to 0.442) without resampling or missing value imputation [ 32 , 36 , 40 ]. These models are reported to be robust enough to handle imbalance and missing values. They are sometimes used for preprocessing the data and predictions [ 41 – 43 ]. Our findings suggest that depending on the ML models of interest, missing value imputation and resampling techniques need careful consideration. Complete case analysis had higher kappa score than missing imputation techniques for LR and SGD, while RF, DT and MLP had higher performance with imputation techniques. We observed wide variations between models as well as agreement between accuracy, F1 score, precision and recall metric with kappa. For ensemble models resampling with SMOTE or SVMSMOTE improves classification performance with either simple imputation or complete cases. MLP classification is improved by LI and SMOTE resampling, while SGD and LR have better classification performance with complete cases and SVMSMOTE resampling. Hence, careful consideration of the mechanism of occurrence of missing values, class imbalance and the intended ML models to train the data on are needed to generate more reliable mastitis predictions with AMS data. Declarations Supplementary Materials: The following supporting information can be downloaded at: www.mdpi.com/xxx/s1, Figure S1: Receiver Operating curves for the resampling methods; Figure S2: Receiver Operating Curves for the Machine learning model performance with resampling methods; Table S1: Rankings of the prediction scores of the fitted machine learning models with both resampling and imputation methods (kappa >0.20). Author Contributions: Conceptualization: K.O, A.C and K.T; methodology: K.O and C.A; data processing and analysis: KO; validation: A.C, K.T, K.O, A.T, A.B, D.M; original draft preparation: K.O; review and editing: K.T, A.C, A.T, A.B, D.M; supervision: A.T, A.B; project administration: K.T., A.T, A.B; funding acquisition: A.T, A.B. All authors have read and agreed to the published version of the manuscript.” Please turn to the CRediT taxonomy for the term explanation. Authorship must be limited to those who have contributed substantially to the work reported. Funding: This research was funded by the Federal Ministry of Food and Agriculture (BMEL, Germany) based on a resolution of the German Bundestag. The article is funded from the project MEDICow, funding was carried out by the Federal Office of agriculture and food (BLE, Germany) within the framework of the federal program “Livestock Husbandry" (grant number: 28N206601). Data Availability Statement: The data used to produce the results presented in this manuscript will be made available upon reasonable request. The source code used to produce the results presented in this manuscript can be availaed freely at (https://github.com/Okashongwe/resampling_mast.git). Acknowledgments: Authors wish to thank the dairy farm for the support and provision of data. Conflicts of Interest: The authors declare no conflict of interest. Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. References Cheng, W.N.; Han, S.G. Bovine Mastitis: Risk Factors, Therapeutic Strategies, and Alternative Treatments—A Review. Asian-Australasian journal of animal sciences 2020 , 33 , 1699. Aghamohammadi, M.; Haine, D.; Kelton, D.F.; Barkema, H.W.; Hogeveen, H.; Keefe, G.P.; Dufour, S. Herd-Level Mastitis-Associated Costs on Canadian Dairy Farms. Front. Vet. Sci. 2018 , 5 , doi:10.3389/fvets.2018.00100. Hogeveen, H.; Steeneveld, W.; Wolf, C.A. Production Diseases Reduce the Efficiency of Dairy Production: A Review of the Results, Methods, and Approaches Regarding the Economics of Mastitis. Annual Review of Resource Economics 2019 , 11 , 289–312, doi:10.1146/annurev-resource-100518-093954. Baskaran, S.A.; Kazmer, G.; Hinckley, L.; Andrew, S.; Venkitanarayanan, K. Antibacterial Effect of Plant-Derived Antimicrobials on Major Bacterial Mastitis Pathogens in Vitro. Journal of dairy science 2009 , 92 , 1423–1429. Martins, S.A.; Martins, V.C.; Cardoso, F.A.; Germano, J.; Rodrigues, M.; Duarte, C.; Bexiga, R.; Cardoso, S.; Freitas, P.P. Biosensors for On-Farm Diagnosis of Mastitis. Frontiers in bioengineering and biotechnology 2019 , 7 , 186. Bernhardt, H.; Höhendinger, M.; Gräff, A.; Hijazi, O.; Höld, M.; Reger, M.; Stumpenhausen, J. Development of Automatic Milking in Germany. In Proceedings of the 2019 ASABE Annual International Meeting; American Society of Agricultural and Biological Engineers, 2019; p. 1. Steeneveld, W.; Vernooij, J.; Hogeveen, H. Effect of Sensor Systems for Cow Management on Milk Production, Somatic Cell Count, and Reproduction. Journal of dairy science 2015 , 98 , 3896–3905. Sitkowska, B.; Piwczynski, D.; Aerts, J.; Kolenda, M.; ÖZKAYA, S. Detection of High Levels of Somatic Cells in Milk on Farms Equippedwith an Automatic Milking System by Decision Trees Technique. Turkish Journal of Veterinary & Animal Sciences 2017 , 41 , 532–540. Bonestroo, J.; van der Voort, M.; Hogeveen, H.; Emanuelson, U.; Klaas, I.C.; Fall, N. Forecasting Chronic Mastitis Using Automatic Milking System Sensor Data and Gradient-Boosting Classifiers. Computers and Electronics in Agriculture 2022 , 198 , 107002. Bobbo, T.; Biffani, S.; Taccioli, C.; Penasa, M.; Cassandro, M. Comparison of Machine Learning Methods to Predict Udder Health Status Based on Somatic Cell Counts in Dairy Cows. Scientific Reports 2021 , 11 , 13642. Hyde, R.M.; Down, P.M.; Bradley, A.J.; Breen, J.E.; Hudson, C.; Leach, K.A.; Green, M.J. Automated Prediction of Mastitis Infection Patterns in Dairy Herds Using Machine Learning. Scientific reports 2020 , 10 , 4289. Post, C.; Rietz, C.; Büscher, W.; Müller, U. Using Sensor Data to Detect Lameness and Mastitis Treatment Events in Dairy Cows: A Comparison of Classification Models. Sensors 2020 , 20 , 3863. Fadul-Pacheco, L.; Delgado, H.; Cabrera, V.E. Exploring Machine Learning Algorithms for Early Prediction of Clinical Mastitis. International Dairy Journal 2021 , 119 , 105051, doi:10.1016/j.idairyj.2021.105051. Abdul Ghafoor, N.; Sitkowska, B. MasPA: A Machine Learning Application to Predict Risk of Mastitis in Cattle from AMS Sensor Data. AgriEngineering 2021 , 3 , 575–583. Hogeveen, H.; Kamphuis, C.; Steeneveld, W.; Mollenhorst, H. Sensors and Clinical Mastitis—The Quest for the Perfect Alert. Sensors 2010 , 10 , 7991–8009. Li, Z.; Jiang, Y.; Hu, C.; Peng, Z. Recent Progress on Decoupling Diagnosis of Hybrid Failures in Gear Transmission Systems Using Vibration Sensor Signal: A Review. Measurement 2016 , 90 , 4–19. Dominiak, K.N.; Kristensen, A.R. Prioritizing Alarms from Sensor-Based Detection Models in Livestock Production - A Review on Model Performance and Alarm Reducing Methods. Computers and Electronics in Agriculture 2017 , 133 , 46–67, doi:https://doi.org/10.1016/j.compag.2016.12.008. Van Buuren, S. Flexible Imputation of Missing Data ; CRC press, 2018; Madley-Dowd, P.; Hughes, R.; Tilling, K.; Heron, J. The Proportion of Missing Data Should Not Be Used to Guide Decisions on Multiple Imputation. Journal of clinical epidemiology 2019 , 110 , 63–73. Pham, T.M.; Pandis, N.; White, I.R. Missing Data: Issues, Concepts, Methods. Seminars in Orthodontics 2024 , 30 , 37–44, doi:https://doi.org/10.1053/j.sodo.2024.01.007. White, I.R.; Royston, P.; Wood, A.M. Multiple Imputation Using Chained Equations: Issues and Guidance for Practice. Statistics in medicine 2011 , 30 , 377–399. Noor, M.; Al Bakri, A.; Yahaya, A.; Ramli, N.; Fitri, N. Estimation of Missing Values in Environmental Data Set Using Interpolation Technique: Fitting on Lognormal Distribution. Aust. J. Basic Appl. Sci 2013 , 7 , 336–341. Huang, G. Missing Data Filling Method Based on Linear Interpolation and Lightgbm. In Proceedings of the Journal of Physics: Conference Series; IOP Publishing, 2021; Vol. 1754, p. 012187. Khushi, M.; Shaukat, K.; Alam, T.M.; Hameed, I.A.; Uddin, S.; Luo, S.; Yang, X.; Reyes, M.C. A Comparative Performance Analysis of Data Resampling Methods on Imbalance Medical Data. IEEE Access 2021 , 9 , 109960–109975. Johnson, J.M.; Khoshgoftaar, T.M. A Survey on Classifying Big Data with Label Noise. J. Data and Information Quality 2022 , 14 , 23:1-23:43, doi:10.1145/3492546. Nguyen, H.M.; Cooper, E.W.; Kamei, K. Borderline Over-Sampling for Imbalanced Data Classification. International Journal of Knowledge Engineering and Soft Data Paradigms 2011 , 3 , 4–21, doi:10.1504/IJKESDP.2011.039875. Ghorbani, R.; Ghousi, R. Comparing Different Resampling Methods in Predicting Students’ Performance Using Machine Learning Techniques. IEEE Access 2020 , 8 , 67899–67911, doi:10.1109/ACCESS.2020.2986809. Bagui, S.S.; Mink, D.; Bagui, S.C.; Subramaniam, S. Determining Resampling Ratios Using BSMOTE and SVM-SMOTE for Identifying Rare Attacks in Imbalanced Cybersecurity Data. Computers 2023 , 12 , 204, doi:10.3390/computers12100204. Tarimo, C.S.; Bhuyan, S.S.; Li, Q.; Ren, W.; Mahande, M.J.; Wu, J. Combining Resampling Strategies and Ensemble Machine Learning Methods to Enhance Prediction of Neonates with a Low Apgar Score After Induction of Labor in Northern Tanzania. Risk Management and Healthcare Policy 2021 , 14 , 3711–3720, doi:10.2147/RMHP.S331077. Ebrahimi, M.; Mohammadi-Dehcheshmeh, M.; Ebrahimie, E.; Petrovski, K.R. Comprehensive Analysis of Machine Learning Models for Prediction of Sub-Clinical Mastitis: Deep Learning and Gradient-Boosted Trees Outperform Other Models. Computers in Biology and Medicine 2019 , 114 , 103456, doi:10.1016/j.compbiomed.2019.103456. Abidin, N.Z.; Ritahani, A.; A., N. Performance Analysis of Machine Learning Algorithms for Missing Value Imputation. ijacsa 2018 , 9 , doi:10.14569/IJACSA.2018.090660. Shah, A.D.; Bartlett, J.W.; Carpenter, J.; Nicholas, O.; Hemingway, H. Comparison of Random Forest and Parametric Imputation Models for Imputing Missing Data Using MICE: A CALIBER Study. American Journal of Epidemiology 2014 , 179 , 764–774, doi:10.1093/aje/kwt312. Mukaka, M.; White, S.A.; Terlouw, D.J.; Mwapasa, V.; Kalilani-Phiri, L.; Faragher, E.B. Is Using Multiple Imputation Better than Complete Case Analysis for Estimating a Prevalence (Risk) Difference in Randomized Controlled Trials When Binary Outcome Observations Are Missing? Trials 2016 , 17 , 341, doi:10.1186/s13063-016-1473-3. Moritz, S.; Bartz-Beielstein, T. imputeTS: Time Series Missing Value Imputation in R. The R Journal 2017 , 9 , 207, doi:10.32614/RJ-2017-009. Park, I.; Kim, H.S.; Lee, J.; Kim, J.H.; Song, C.H.; Kim, H.K. Temperature Prediction Using the Missing Data Refinement Model Based on a Long Short-Term Memory Neural Network. Atmosphere 2019 , 10 , 718, doi:10.3390/atmos10110718. Moon, T.; Hong, S.; Choi, H.Y.; Jung, D.H.; Chang, S.H.; Son, J.E. Interpolation of Greenhouse Environment Data Using Multilayer Perceptron. Computers and Electronics in Agriculture 2019 , 166 , 105023, doi:10.1016/j.compag.2019.105023. Buabeng, A.; Simons, A.; Frempong, N.K.; Ziggah, Y.Y. A Novel Hybrid Predictive Maintenance Model Based on Clustering, Smote and Multi-Layer Perceptron Neural Network Optimised with Grey Wolf Algorithm. SN Appl. Sci. 2021 , 3 , 593, doi:10.1007/s42452-021-04598-1. Wongvorachan, T.; He, S.; Bulut, O. A Comparison of Undersampling, Oversampling, and SMOTE Methods for Dealing with Imbalanced Classification in Educational Data Mining. Information 2023 , 14 , 54, doi:10.3390/info14010054. Jian, C.; Gao, J.; Ao, Y. A New Sampling Method for Classifying Imbalanced Data Based on Support Vector Machine Ensemble. Neurocomputing 2016 , 193 , 115–122, doi:https://doi.org/10.1016/j.neucom.2016.02.006. Poulos, J.; Valle, R. Missing Data Imputation for Supervised Learning. Applied Artificial Intelligence 2018 , 32 , 186–196. Upadhyay, A.; Singh, M.; Yadav, V.K. Improvised Number Identification Using SVM and Random Forest Classifiers. Journal of Information and Optimization Sciences 2020 , 41 , 387–394, doi:10.1080/02522667.2020.1723934. Phiri, D.; Morgenroth, J.; Xu, C.; Hermosilla, T. Effects of Pre-Processing Methods on Landsat OLI-8 Land Cover Classification Using OBIA and Random Forests Classifier. International Journal of Applied Earth Observation and Geoinformation 2018 , 73 , 170–178, doi:10.1016/j.jag.2018.06.014. Iliou, T.; Anagnostopoulos, C.-N.; Stephanakis, I.M.; Anastassopoulos, G. A Novel Data Preprocessing Method for Boosting Neural Network Performance: A Case Study in Osteoporosis Prediction. Information Sciences 2017 , 380 , 92–100, doi:10.1016/j.ins.2015.10.026. Additional Declarations The authors declare no competing interests. Supplementary Files Appendix.docx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4629327","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":318247626,"identity":"6416de24-3ce0-400d-ac6f-446878fd593d","order_by":0,"name":"Kashongwe B.O.","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABEklEQVRIiWNgGAWjYDACHgYGZgiL+QADYwOEKQHEMDY+LWwJDAdBytiI18JjQJwWfp7Djz8XVNyx23C85+PnjzsY8vjl2x/e+JnDINuPQ4tkb5uZ9Iwzz5I3nDm7WeLgGYZiyTYeY8vebQzGM3FYY3CewYyZt+1wstmN3A0SB9v+J244xsMmwbuNIXHDAVxa2D9/Bmu5/+bxj4NtQJXH2J9J/gVq2Y9Ly9keA2mgFjuzG0DDIVoYzKTBtuDyS8+ZMmmeM4cT7M+kmVmcBfslx9hadpuE8QwctvDzpG/+zFNx2F6y/fDjG5WgEGM+/vDm2202sv04vA8DiTD5BCgtgV89ENjDGAl4FI2CUTAKRsEIBQDlaWSG34RiwwAAAABJRU5ErkJggg==","orcid":"","institution":"Leibniz Institut für Agrartechnik und Bioökonomie, e.V. (ATB)","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Kashongwe","middleName":"","lastName":"B.O.","suffix":""},{"id":318247627,"identity":"87f7b1af-3310-4708-a571-1eaa211607f9","order_by":1,"name":"Kabelitz T.","email":"","orcid":"","institution":"Leibniz Institut für Agrartechnik und Bioökonomie, e.V. (ATB)","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Kabelitz","middleName":"","lastName":"T.","suffix":""},{"id":318247628,"identity":"89384ed4-e966-4060-9d7e-94980509351f","order_by":2,"name":"Amon T.","email":"","orcid":"","institution":"Leibniz Institut für Agrartechnik und Bioökonomie, e.V. (ATB)","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Amon","middleName":"","lastName":"T.","suffix":""},{"id":318247629,"identity":"84524805-5ff8-429b-8639-d8ae00f7cfa9","order_by":3,"name":"Ammon C","email":"","orcid":"","institution":"Leibniz Institut für Agrartechnik und Bioökonomie, e.V. (ATB)","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Ammon","middleName":"","lastName":"C","suffix":""},{"id":318247630,"identity":"69494e9d-7730-49d1-aaa6-1d0dd5d56f83","order_by":4,"name":"Amon B.","email":"","orcid":"","institution":"Leibniz Institut für Agrartechnik und Bioökonomie, e.V. (ATB)","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Amon","middleName":"","lastName":"B.","suffix":""},{"id":318247631,"identity":"cd710763-da73-4f46-8fb0-b67711f2fa73","order_by":5,"name":"Doherr M.","email":"","orcid":"","institution":"Free University Berlin","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Doherr","middleName":"","lastName":"M.","suffix":""}],"badges":[],"createdAt":"2024-06-24 10:09:09","currentVersionCode":1,"declarations":{"humanSubjects":false,"vertebrateSubjects":true,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":false,"humanSubjectConsent":false,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":true,"coiExplicitlySet":false},"doi":"10.21203/rs.3.rs-4629327/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4629327/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":59031569,"identity":"c3c149f9-c263-4204-9699-f86040f65071","added_by":"auto","created_at":"2024-06-25 14:11:19","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":151828,"visible":true,"origin":"","legend":"\u003cp\u003eData processing and analysis workflow.\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-4629327/v1/ccfadedddb17621534f3a78b.png"},{"id":59031508,"identity":"b9584e9b-ff70-4e42-84ec-efa139233a03","added_by":"auto","created_at":"2024-06-25 14:11:17","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":348058,"visible":true,"origin":"","legend":"\u003cp\u003ePerformance metrics of ML models from data without missing value imputation techniques (a), with simple imputer (b) linear interpolation (c) and multiple imputer (d).\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-4629327/v1/18b02729659a5dd352e3cdf4.png"},{"id":59031564,"identity":"1690dd06-12cf-4627-a2e5-acbe503d1e63","added_by":"auto","created_at":"2024-06-25 14:11:19","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":361047,"visible":true,"origin":"","legend":"\u003cp\u003ePerformance metrics of ML models from data without resampling techniques (a), with SMOTE resampling (b), with SMOTEEN (c) and SVMSMOTE (d).\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-4629327/v1/60d17c1b52913ec04bfde20d.png"},{"id":59031606,"identity":"b0ff0d74-92e0-46c2-b2c5-0837d13f04fd","added_by":"auto","created_at":"2024-06-25 14:11:27","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2085972,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4629327/v1/c57f3453-1edb-43f2-acfd-7e347d1272c0.pdf"},{"id":59031567,"identity":"50e4a7d1-c3ec-4b03-aafd-aa71501b1c08","added_by":"auto","created_at":"2024-06-25 14:11:19","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":1016301,"visible":true,"origin":"","legend":"","description":"","filename":"Appendix.docx","url":"https://assets-eu.researchsquare.com/files/rs-4629327/v1/9dfbf40b050095296ec55bb3.docx"}],"financialInterests":"The authors declare no competing interests.","formattedTitle":"\u003cp\u003e\u003cstrong\u003eInfluence of Preprocessing Methods of Automated Milking Systems Data on the Prediction of Mastitis with Machine Learning Models\u003c/strong\u003e\u003c/p\u003e","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003e \u003cdiv class=\"BlockQuote\"\u003e \u003cp\u003eMastitis is the most costly and frequent disease in dairy farming, contributing to considerable (about 125\u0026euro; per cow and year) and recurring costs incurred through reduction of milk quantity and quality as well as the reproductive performance and longevity of cows [\u003cspan additionalcitationids=\"CR2\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. The disease can be caused by a wide diversity of pathogens, but is dominated by only few species. The most important mastitis causing bacterial species are \u003cem\u003eStreptococcus agalactiae\u003c/em\u003e, \u003cem\u003eStreptococcus dysgalactiae\u003c/em\u003e, \u003cem\u003eStreptococcus uberis\u003c/em\u003e, \u003cem\u003eStaphylococcus aureus\u003c/em\u003e, and \u003cem\u003eEscherichia coli\u003c/em\u003e [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. The subclinical form of mastitis, without overt symptoms, is by far more prevalent than the clinical form and causes high economic losses [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. Simultaneously, subclinical mastitis is, due to the low level of symptoms, difficult to detect. Therefore, modern data analysis tools like machine leaning could help to identify subclinical mastitis cases more often and accurate.\u003c/p\u003e \u003cp\u003eManaging mastitis is even a bigger challenge in large-scale dairy farms despite the use of automated milking systems (AMS) that has steadily increased in Germany and across European countries over the last years [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. The AMS uses sensors to collect data such as milk yield and milk components and through management programs, alerts are given in case of variations of either milk production (milk yield and milk flow), electric conductivity, somatic cell counts (SCC), milk temperature or a combination of all these parameters as an indication for mastitis occurrence [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. The data collected at each milking or through monthly milk control programs, is routinely used to help farmers for better decision making on production, reproduction and health. This is specifically the case for mastitis predictions based on milk yield, milk parameters (conductivity, SCC, blood in milk, temperature) and cow characteristics [\u003cspan additionalcitationids=\"CR8\" citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. For this prediction purpose, several machine-learning (ML) approaches have been used that aim to improve monitoring of udder health status in general or mastitis specifically, whether subclinical or clinical. [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e] compared eight ML models and achieved prediction accuracy above 75% for all of them in a binary classification with 200,000 cells/ml SCC as threshold for positivity. [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e] trained random forest models to predict mastitis infection patterns in a binary classification where the predictor was contagious vs environmental mastitis or environmental lactation vs environmental dry period. They obtained 98% prediction accuracy. Post et al. [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e] applied ML models to a group of animals with historical records of diseases and achieved higher prediction accuracy than when the model was applied to the whole population. Findings from other studies using similar methodologies show also high prediction accuracy [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e, \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. Despite these encouraging results, the application of mastitis prediction ML models in real life conditions remains limited because of a discrepancy between the performance on the training data sets and the actual data. The nature of data recording by sensor systems and low occurrence of the disease appear as major reasons for the difference [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eIndeed, farm sensor data present in general two types of challenges that make them difficult to handle by prediction algorithms. First, the data is often noisy with missing values, outliers and skewed values, occurring because of sensor failure during signal transmission [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]. The missing values or wrongly recorded values fall into the type of either missing values completely at random or missing values at random. Missing values together with a clear definition of positive cases represent a major hindrance for ML algorithms trained on \u0026lsquo;experimental data sets\u0026rsquo; to be used under farm conditions [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]. To handle the problem, it is common in practice to delete missing values completely or at least to apply methods such as list wise deletion but less common is the reporting of the magnitude of missing values or the use of missing data handling methods [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. Although working without missing values is convenient, it only produces reliable estimates in limited situations where missing values occur completely at random and only on the dependent variable. In other situations, this result in severely biased estimates, not to mention the potential waste of information in the omitted data and the low practical application of the obtained results [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. This is of particular importance in disease prediction where metrics obtained from the training data sets need to be applied in real-life situations [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]. Indeed, it is very common to have large amounts of missing values in sensor-generated datasets [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eVarious techniques have been developed to deal with the challenge of missing values in large datasets. The common imputation methods are simple imputation, multiple imputation and linear interpolation. The simple imputation method employs replacement of missing values with mean, median or mode values [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. This method is largely proposed for its computational convenience although, in many cases, the results and conclusions are not sensible or generalizable [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. The multiple imputation uses the MICE (multiple imputation with chained equations) algorithm, which is a Markov Chain Monte Carlo method that imputes incomplete data in a variable-by-variable way starting with a random draw of the observed data. For instance, A first regression of the first variable with missing values applied to all other variables provided that the rows have observations for the variable of interest. Then missing values in the variable of interest are replaced by simulated draws from its posterior predictive distribution. The process is repeated for all other variables with missing values in turn: this is called a cycle. This process is repeated several times to generate a single imputed dataset and the whole process is repeated 3 to 5 times to obtain stable results [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]. Although recognised for its robustness, the method suffers the limitation of lack of theoretical rationale [[\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e, \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e, \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e]. Linear interpolation estimates the value of the missing data based on the two data points adjacent to the missing one in a one dimensional data sequence [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e]. It is reputed to perform well on time dependent data and on dataset with small to moderate missing values between adjacent points [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThe second major hurdle when training models on AMS data to predict mastitis is the class-imbalance between positive and negative cases [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. Although frequently observed in dairy farms, mastitis is a rare occurrence when data resolution is increased to either daily basis or animal level or both. This imbalance causes a bias when fitting standard learning classifiers, reflected in their inability to predict correctly the minority class, despite sometimes achieving high prediction accuracy [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e]. Johnson \u0026amp; Khoshgoftaar [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e] noted that the total number of the minority class is more important than the percentage of imbalance. Various methods to handle class imbalance are reported in literature to improve disease prediction [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e, \u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e]. Johnson \u0026amp; Khoshgoftaar [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e] categorized them into three groups, data-level methods, algorithm level methods and hybrid methods. Data level methods change the dataset structure by either reducing the majority class (under-sampling) or increasing the minority class (oversampling) or both to achieve a more balanced class distribution [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e]. Among popular resampling techniques, is the Synthetic Minority Oversampling Technique (SMOTE) which produces synthetic samples by interpolating minority samples with their k Nearest Neighbors. The algorithm seems to be improved by taking into consideration the minority class lying along the borderline, hence expanding the minority class area towards the side of the majority class where only few instances of majority class are found [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e]. However, oversampling techniques may lead to overfitting. The Random Under-Sampling is among the first under-sampling techniques developed and works by discarding random samples in the majority class. The technique has been improved with several techniques using nearest neighbours to reduce instances in the majority class. The Edited Nearest Neghbours (ENN) tests every instance with the rest of samples using k-NN and disqualifies incorrectly classified samples. Under-sampling methods have the disadvantage of discarding information that may be useful. Techniques combining over-sampling and under-sampling have been developed to overcome limitations of individual methods. The SMOTE-ENN combines the SMOTE and edited nearest neighbours for under-sampling [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eStudies on mastitis prediction with machine learning classifiers use the above mentioned data (pre)processing techniques almost interchangeably, making comparison and evaluation of effectiveness across studies more complex. Hence, we investigated whether resampling or imputation or both techniques were critical to influence the prediction performance of ML classifiers. We used several metrics including accuracy, F1 score, precision, recall and kappa score.\u003c/p\u003e \u003c/div\u003e \u003c/p\u003e"},{"header":"2. Materials and Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e2.1. Data Collection\u003c/h2\u003e \u003cp\u003e \u003cdiv class=\"BlockQuote\"\u003e \u003cp\u003eThe data set included records of 232 cows and 75217 milking events, for which daily milk yield, electric conductivity at quarter and cow levels and somatic cell counts were recorded. Data were collected between January 2015 and September 2017 from a dairy farm that uses automated milking systems as well as a dairy herd management program. The Lely Astronaut system (Lely Industries N.V., Maassluis, the Netherlands) equipped with in-line sensors for electric conductivity (EC) and SCC was used to milk cows and monitor their performance. Hence, mastitis cases used in this study referred to the alarm raised by the AMS due to changes detected in milk. Descriptions as abnormal milk (n\u0026thinsp;=\u0026thinsp;54), mastitis (n\u0026thinsp;=\u0026thinsp;398), high conductivity (n\u0026thinsp;=\u0026thinsp;14), watery milk (n\u0026thinsp;=\u0026thinsp;2) were classified as positive cases, while instances where no alarm was raised were classified as negative cases (n\u0026thinsp;=\u0026thinsp;74749). Conductivity had an average of 68.44\u0026thinsp;\u0026plusmn;\u0026thinsp;3.42 \u0026micro;S/cm, SCC had an average of 93.86\u0026thinsp;\u0026plusmn;\u0026thinsp;188.29 x 10\u003csup\u003e3\u003c/sup\u003e cells/ml. The dataset contained missing values for predictors as presented in Table \u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e \u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eDescription of the dataset.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMissing values\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eTotal cases\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eVariable\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e75217\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAlarm*\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e36028\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e39189\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eEC_FL\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e36028\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e39189\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eEC_FR\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e36315\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e38902\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eEC_BR\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e36142\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e39075\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eEC_BL\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e37326\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e37891\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eEC_ALL\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e58669\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e16548\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eSCC\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e29\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e75188\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMilk yield\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"3\" nameend=\"c3\" namest=\"c1\"\u003e \u003cp\u003e*refers to both presence (value\u0026thinsp;=\u0026thinsp;1) and absence (value\u0026thinsp;=\u0026thinsp;0) of alarm, EC: electric conductivity, FL: front left, FR: front right, BR: back right, BL: back left, ALL: average all quarters, SCC: Somatic Cell counts.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e2.2. Data Preparation\u003c/h2\u003e \u003cdiv id=\"Sec5\" class=\"Section3\"\u003e \u003ch2\u003e2.2.1. Libraries and packages used for data preparation and modeling\u003c/h2\u003e \u003cp\u003e \u003cdiv class=\"BlockQuote\"\u003e \u003cp\u003eAll analyses were performed in Python, using \u003cem\u003enumpy, pandas\u003c/em\u003e and \u003cem\u003escikit learn\u003c/em\u003e libraries. We used the \u003cem\u003eimblearn\u003c/em\u003e and \u003cem\u003eimputer\u003c/em\u003e packages to perform the resampling and missing value imputation, respectively. We selected from \u003cem\u003escikit learn\u003c/em\u003e library five common supervised learning models: Stochastic gradient descent (SGD), logistic regression (LR), decision tree (DT), random forest (RF) and multilayer perceptron (MLP). We imported \u003cem\u003econfusion matrix, roc_curve\u003c/em\u003e and \u003cem\u003eauc\u003c/em\u003e from \u003cem\u003esklearn metrics\u003c/em\u003e to compute precision metrics and plot the ROC curves. Plotting of ROC curves was aided by the \u003cem\u003epyplot\u003c/em\u003e package of \u003cem\u003ematplotlib\u003c/em\u003e library [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. Finally, we sued statsmodels library to compare the metrics of the models tested and determine for each model type whether resampling, imputation or both influenced the observed performance. We performed five fold cross-validation using k-fold strategy in the Scikit-learn GridsearchCV to obtain optimal parameters search and hyperparameter tuning for each model.\u003c/p\u003e \u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section3\"\u003e \u003ch2\u003e2.2.2. Cross-validation and model tuning\u003c/h2\u003e \u003cp\u003e \u003cdiv class=\"BlockQuote\"\u003e \u003cp\u003eThe dataset was loaded and checked for inconsistencies before further processing. inconsistencies such as missing dates, missing all values across one observation, outliers that could have resulted from erroneous measurement or recording and the corresponding data were cleaned. The data was then split into training and test set at a ratio of 80:20. The test set was left out of further processing. The training set was further split into training and validation set and subjected to further processing. Two directions were followed for processing. On one hand, all missing values were deleted from the dataset to remain only with complete cases (CC) for which no missing imputation was required. On the other hand, data with missing values were processed with one of three selected imputation techniques: simple imputer (SI), multiple imputer (MICE) and linear interpolation (LI). This concerned both the training and validation set for which a ratio of 60:40 was used for complete case analysis and 80:20 for the other imputation techniques.\u003c/p\u003e \u003cp\u003eModel tuning consisted of finding the best parameters for each model resulting from the above mentioned imputation methods. A full description of the hyperparameter tuning and model fitting can be found in the supplementary file 2. In summary, for the stochastic gradient descent models parameters included loss function (hinge, log_loss or modifier huber), penalty criteria (l2, l1, elasticnet) and the value of alpha (0.0001 to 0.1). Additionally 5 fold cross-validation was performed.The best fitted loss functions and penalty criteria were \u0026lsquo;hinge\u0026rsquo; and \u0026lsquo;elasticnet\u0026rsquo; for CC, MICE, and LI, while it was \u0026lsquo;log_loss\u0026rsquo; and \u0026lsquo;L2\u0026rsquo; for SI. The best values of alpha were 0.1, 0.01 and 0.001 for CC, SI and MICE and LI, respectively (See Supplementary file 2).\u003c/p\u003e \u003cp\u003eThe grid search cross-validation for logistic regression included solver (lbfgs, liblinear), penalty criteria (l2, l1, elasticnet), C (0.1 to 100) and 5 fold cross-validation. The lbfgs solver and the L2 penalty were selected for all datasets. The C value of 0.1 selected for all datasets except for the one from LI whose C value was 10.\u003c/p\u003e \u003cp\u003eThe grid search cross-validation for decision tree models included the criterion (gini, entropy), max depth (None, 2\u0026ndash;20), max features (None, sqrt, log2, 0.2 to 0.8), splitter (best or random). The cross-validation was set to 5 folds. The selected criterion (entropy) and max depth (n\u0026thinsp;=\u0026thinsp;4) were similar for CC and MICE, while it was \u0026lsquo;gini\u0026rsquo; for SI and LI. The max depth was n\u0026thinsp;=\u0026thinsp;6 for SI and n\u0026thinsp;=\u0026thinsp;5 for LI and n\u0026thinsp;=\u0026thinsp;4 for MICE (See Supplementary file 2 ).\u003c/p\u003e \u003cp\u003eFor random forest models, criterion (gini, entropy and log_loss), max depth (None, 2\u0026ndash;20), max features (None, sqrt, log2, 0.2\u0026ndash;0.8) and 5 folds cross-validation were included in the grid search. The criterion (entropy) was selected only for SI, while entropy was selected for the others. The max depth ranged from None for SI, to 20 for MICE and LI. No max features was selsected for CC and MICE, while SI had \u0026lsquo;sqrt\u0026rsquo; and LI had 0.8.\u003c/p\u003e \u003cp\u003eThe grid search cross-validation parameters for multilayer perceptron models included activation (identity, logistic, tanh, relu), solver (lbfgs, sgd, adam), alpha (0.0001 to 0.01), learning rates (constant, invscaling, adaptive) and cross-validation (cv\u0026thinsp;=\u0026thinsp;5). The CC and LI had the best activation with identity while the \u0026lsquo;relu\u0026rsquo; activation was selected for SI and MICE. The SI and LI datasets had also the learning rate \u0026lsquo;invscaling\u0026rsquo; while it was adaptive for CC and constant for MICE. The value of alpha was set at 0.05 for all datasets except for the CC for which the best value was 0.0001 (See supplementary file 2).\u003c/p\u003e \u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003e2.3. Missing Values Imputation\u003c/h2\u003e \u003cp\u003e \u003cdiv class=\"BlockQuote\"\u003e \u003cp\u003eThe simple imputer was implemented for the replacement of missing values of individual variables with the mean or median value of the specific variable. The multivariate imputer (implementing the Multiple Imputer Chain Equation (MICE) in Python) was applied using the iterative imputer module of scikit learn that replaced the missing value of individual variable (column) with the value from specific and other variables for the same observation (row). The linear interpolation (LI) was implemented using the interpolate module. We set the method at linear for the replacement of missing values of individual variables based on both previous (limit direction backward) and following (limit direction forward) values to predict the missing value using linear regression. Hence, two datasets with only complete cases and with missing replaced values would form the basis for data processing. They had imbalance ratios of 53.76 and 156.52, respectively (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eImbalance ratio of initial datasets.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePositive cases\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNegative cases\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eImbalance ratio\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e87\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e4677\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e53.76\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eComplete cases\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e382\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e59791\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e156.52\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eData with replaced missing values\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003e2.4. Resampling Methods\u003c/h2\u003e \u003cp\u003e \u003cdiv class=\"BlockQuote\"\u003e \u003cp\u003eThe two sets of data were submitted to three resampling methods: the Synthetic Minority Oversampling Technique (SMOTE), the SMOTE technique combined with Edited Nearest Neighbours (SMOTEEN) and the SMOTE technique combined with Support Vector Machine (SVM) classifier SVMSMOTE (Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). The SMOTE technique oversampled the minority class without altering majority class. The SMOTEEN technique not only oversampled the minority class but also under sampled the majority class. The SVMSMOTE technique oversampled the minority class along the borderline and used Support Vector Machine (SVM) classifier to predict new cases.\u003c/p\u003e \u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eImputed datasets resampled with SMOTE, SMOTEEN, and SVMSMOTE methods.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePositive\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNegative\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eImbalance ratio\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cem\u003eComplete cases\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e87\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e4677\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e53.76\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eNo resampling\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e4677\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e4677\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSMOTE\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e4459\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e4134\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.93\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSMOTEEN\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e4677\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e4677\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSVMSMOTE\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cem\u003eSimple Imputer\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e382\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e59791\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e156.52\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eNo resampling\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e59791\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e59791\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSMOTE\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e58100\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e58387\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSMOTEEN\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e59791\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e59791\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSVMSMOTE\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cem\u003eMultiple Imputer\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e382\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e59791\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e156.52\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eNo resampling\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e59791\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e59791\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSMOTE\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e58674\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e58519\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSMOTEEN\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e59791\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e59791\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSVMSMOTE\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cem\u003eLinear Interpolation\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e382\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e59791\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e156.52\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eNo resampling\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e59791\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e59791\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSMOTE\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e57820\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e59521\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.03\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSMOTEEN\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e59791\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e59791\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSVMSMOTE\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eParamters used for the application of resampling methods are presented in Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e and were similar to those reported in literature [\u003cspan additionalcitationids=\"CR28\" citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e].\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eParameters settings used for the resampling methods\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eParameters\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMethods\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eK-Neighbors\u0026thinsp;=\u0026thinsp;5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSMOTE\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eK\u003csub\u003e(SMOTE)\u003c/sub\u003e\u0026thinsp;=\u0026thinsp;5, K\u003csub\u003e(ENN)\u003c/sub\u003e\u0026thinsp;=\u0026thinsp;3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSMOTEEN\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eK-Neighbors\u0026thinsp;=\u0026thinsp;5, m_Neighbours\u0026thinsp;=\u0026thinsp;10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSVSMSMOTE\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003e2.5. Processing Method Evaluation and Comparison of Performance metrics\u003c/h2\u003e \u003cp\u003e \u003cdiv class=\"BlockQuote\"\u003e \u003cp\u003eAfter obtaining the resampled and /or imputed datasets, we split the data into training and validation sets using the ratio of 80:20 for the simple imputer, MICE, and linear interpolated data, while we used a ratio of 60:40 for the complete case analysis where the number positive cases were much fewer (Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). Ebrahimi et al [\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e] used a similar approach in their study on prediction of sub-clinical mastitis using machine learning models.\u003c/p\u003e \u003cp\u003eThe 16 datasets were fitted to five common machine-learning classifiers: Stochastic gradient descent (SGD), logistic regression (LR), decision tree (DT), random forest (RF) and multilayer perceptron (MLP). The performance of the classifiers was evaluated using accuracy, precision, recall, F1 score and kappa metrics. These were obtained from a confusion matrix (Table\u0026nbsp;\u003cspan refid=\"Tab5\" class=\"InternalRef\"\u003e5\u003c/span\u003e), where the correctly classified positive and negative cases are labelled true positive (TP) and true negative (TN). Positive cases incorrectly classified as negative are labelled false negative (FN), while negative cases incorrectly classified as positive are labelled false positive (FP).\u003c/p\u003e \u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab5\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eRepresentation of a Confusion matrix.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePredicted negative\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePredicted positive\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFalse negatives FN\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eTrue positives TP\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eActual positive\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTrue negatives TN\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eFalse positives FP\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eActual Negative\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"BlockQuote\"\u003e \u003cp\u003eWe used accuracy, the area under the Receiver Operating Characteristic curve (ROC), precision, recall, F1 score and Cohen\u0026rsquo;s kappa score to evaluate the performance of the models. Accuracy is the most commonly used metric and the starting point to evaluate the performance of classifiers. Accuracy is the proportion of correct predictions (true positive, true negative) among all examined cases [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eAccuracy = (TP\u0026thinsp;+\u0026thinsp;TN) / (TP\u0026thinsp;+\u0026thinsp;FP\u0026thinsp;+\u0026thinsp;TN\u0026thinsp;+\u0026thinsp;FN) (1)\u003c/p\u003e \u003cp\u003eSensitivity, recall, true positive rate TPR\u0026thinsp;=\u0026thinsp;TP/ (TP\u0026thinsp;+\u0026thinsp;FN) (2)\u003c/p\u003e \u003cp\u003ePositive predictive value, precision PPV\u0026thinsp;=\u0026thinsp;TP/ (TP\u0026thinsp;+\u0026thinsp;FP) (3)\u003c/p\u003e \u003cp\u003eFalse positive rate FPR\u0026thinsp;=\u0026thinsp;FP/ (FP\u0026thinsp;+\u0026thinsp;TN) (4)\u003c/p\u003e \u003cp\u003eSpecificity, true negative rate TNR\u0026thinsp;=\u0026thinsp;TN/ (TN\u0026thinsp;+\u0026thinsp;FP) (5)\u003c/p\u003e \u003cp\u003eNegative predictive value NPV\u0026thinsp;=\u0026thinsp;TN/ (TN\u0026thinsp;+\u0026thinsp;FN) (6)\u003c/p\u003e \u003cp\u003eF1 score\u0026thinsp;=\u0026thinsp;2 * TP / 2 * (TP\u0026thinsp;+\u0026thinsp;FP\u0026thinsp;+\u0026thinsp;FN)\u0026thinsp;=\u0026thinsp;2 * (precision * recall) / (precision\u0026thinsp;+\u0026thinsp;recall) (7)\u003c/p\u003e \u003cp\u003eHowever, accuracy score for unbalanced problems often provides an overoptimistic estimation of the classifier ability to predict the majority class [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. Hence, the use of other performance metrics could strengthen the meaning of models performance obtained. The ROC plots the true positive rate and the false positive rate at various thresholds values. The F1 score is a weighted (harmonic) mean of sensitivity and precision. The Cohen\u0026rsquo;s Kappa value compares the classifiers performance to the probability that its performance only be based on chance. In general, predictions from models with Kappa values\u0026thinsp;\u0026lt;\u0026thinsp;0.20 are considered poor, values between 0.21\u0026ndash;0.40 are fair and from 0.40\u0026ndash;0.60 moderate and above 61 substantial to almost perfect [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]. Thirty-nine out of 80 models tested for the current study had kappa\u0026thinsp;\u0026lt;\u0026thinsp;21 and none of the model had kappa value\u0026thinsp;\u0026gt;\u0026thinsp;50. Thus, we ranked classifiers performance metrics first by Kappa value, then by f1 score and precision/ recall scores regardless of the accuracy score. We also examined the ROC curves of the best models to evaluate their performance at various thresholds. Finally, we assessed the contribution of resampling or imputation techniques or both on the prediction performance of the ML classifiers using AIC score and Residual deviance. An overview of the complete workflow performed in this study is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e \u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"3. Results","content":"\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e\n \u003ch2\u003e3.1. Performance Metrics of ML Models Trained on Data with Different Missing Values Imputation Techniques\u003c/h2\u003e\n \u003cdiv class=\"BlockQuote\"\u003e\n \u003cp\u003eResults show that imputation techniques improved the prediction performance RF and DT models with SI (kappa\u0026thinsp;=\u0026thinsp;0.43 and 0.36, respectively) higher than MI (kappa\u0026thinsp;=\u0026thinsp;0.41 and 0.36, respectively), CC (kappa\u0026thinsp;=\u0026thinsp;0.40 and 0.31, respectively) and LI (kappa\u0026thinsp;=\u0026thinsp;0.27 and 0.26, respectively) (Fig. \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e). The opposite was observed with SGD and LR where CC was higher than MI, SI and LI, respectively. The linear interpolation for MLP performed better (kappa\u0026thinsp;=\u0026thinsp;0.27) than the other techniques.\u003c/p\u003e\n \u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec12\" class=\"Section2\"\u003e\n \u003ch2\u003e3.2. Performance Metrics of ML Models Trained on Data with Different Resampling Techniques\u003c/h2\u003e\n \u003cdiv class=\"BlockQuote\"\u003e\n \u003cp\u003eResults for resampling methods indicate that ensemble models perform better (kappa scores\u0026thinsp;\u0026gt;\u0026thinsp;0.30) than MLP (kappa between 0.15 and 0.21) and discriminative classifiers (kappa\u0026thinsp;\u0026lt;\u0026thinsp;0.12).Random forest classifier had higher kappa score for all resampling methods but not for the data without resampling for which Decision Tree had the highest (kappa\u0026thinsp;=\u0026thinsp;0.41) (Fig. \u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e).\u003c/p\u003e\n \u003c/div\u003e\n \u003cp\u003e\u003cem\u003e\u003cstrong\u003e3.3. Rankings of the Prediction Scores of Fitted Machine Learning Models with Both Resampling and Imputation Methods*\u003c/strong\u003e\u003c/em\u003e\u003c/p\u003e\n \u003cdiv class=\"BlockQuote\"\u003e\n \u003cp\u003eThe ensemble models had highest performance with SI and MICE resampled by SVMSMOTE, SMOTE, and No resampling. They all produced fair to moderate kappa scores (\u0026gt;\u0026thinsp;0.35). Discriminative classifiers (SGD, LR) had lower kappa scores (\u0026lt;\u0026thinsp;0.30). MLP models had fair prediction accuracy (kappa\u0026thinsp;=\u0026thinsp;0.229\u0026ndash;0.332) with LI, CC and MICE with or without resampling methods (Table \u003cspan class=\"InternalRef\"\u003e6\u003c/span\u003e).\u003c/p\u003e\n \u003c/div\u003e\n \u003cdiv class=\"gridtable\"\u003e\u0026nbsp;\u003ctable id=\"Tab6\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 6\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003ePerformance metrics of best-ranked ML models.\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003ccolgroup cols=\"8\"\u003e\u003c/colgroup\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eOverall Rank**\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eKappa\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eRecall\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003ePrecision\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eF1Score\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eAccuracy\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eResampling\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eImputation\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\" colspan=\"7\"\u003e\n \u003cp\u003e\u003cem\u003eStochastic gradient descent (SGD)and Logistic regression (LR) models\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e27\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e0.280\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.957\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.979\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.967\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.957\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eNo resampling\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eCC (SGD)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e38\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e0.225\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.940\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.979\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.957\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.940\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eSVMSMOTE\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eCC (SGD)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e39\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e0.218\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.944\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.978\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.959\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.944\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eSVMSMOTE\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eCC (LR)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\" colspan=\"7\"\u003e\n \u003cp\u003e\u003cem\u003eMultilayer perceptron (MLP) models\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e19\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e0.332\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.995\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.996\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.996\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.991\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eSMOTE\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eLI\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e22\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e0.322\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.996\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.996\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.996\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.992\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eNo resampling\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eLI\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e29\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e0.265\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.986\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.997\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.992\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.984\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eSMOTEEN\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eLI\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e33\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e0.261\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.961\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.992\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.976\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.954\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eNo resampling\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eCC\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e34\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e0.261\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.961\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.992\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.976\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.954\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eSMOTE\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eCC\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e37\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e0.229\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.984\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.997\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.990\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.981\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eSVMSMOTE\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eMICE\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\" colspan=\"7\"\u003e\n \u003cp\u003e\u003cem\u003eDecision tree (DT) models\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e0.465\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.997\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.997\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.997\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.994\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eNo resampling\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eSI\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e0.455\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.997\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.997\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.997\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.994\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eNo resampling\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eMICE\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e0.416\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.980\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.981\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.981\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.980\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eNo resampling\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eCC\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e12\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e0.404\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.993\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.998\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.995\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.991\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eSVMSMOTE\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eMICE\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e15\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e0.367\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.994\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.997\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.995\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.991\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eSVMSMOTE\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eSI\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\" colspan=\"7\"\u003e\n \u003cp\u003e\u003cem\u003eRandom forest (RF) models\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e0.453\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.998\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.996\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.997\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.995\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eSMOTE\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eSI\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e0.453\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.998\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.996\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.997\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.995\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eSVMSMOTE\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eSI\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e0.446\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.998\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.997\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.997\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.994\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eSMOTE\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eMICE\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e0.445\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.998\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.997\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.997\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.994\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eSVMSMOTE\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eMICE\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e7\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e0.442\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.998\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.989\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.993\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.987\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eNo resampling\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eCC\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\" colspan=\"7\"\u003e\n \u003cp\u003e*top 5 models with kappa\u0026thinsp;\u0026gt;\u0026thinsp;0.20 were selected for each model category. **full list of ranked model in supplementary file\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec13\" class=\"Section2\"\u003e\n \u003ch2\u003e3.4. Evaluation of Effects of Resampling Technique and Imputation of ML Models Performance\u003c/h2\u003e\n \u003cdiv class=\"BlockQuote\"\u003e\n \u003cp\u003eResults of method importance to model predictions are presented in Table \u003cspan class=\"InternalRef\"\u003e6\u003c/span\u003e. All classifiers except the RF performed better when resampling was associated with imputation. AIC and null deviance values are smallest for these models. RF models with resampling only (AIC\u0026thinsp;=\u0026thinsp;107.2) or a combination of resampling and imputation (Residual deviance\u0026thinsp;=\u0026thinsp;4609) were the best.\u003c/p\u003e\n \u003c/div\u003e\n \u003cdiv class=\"gridtable\"\u003e\u0026nbsp;\u003ctable id=\"Tab7\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 6\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003eEffects of Resampling techniques and imputation on ML models performance.\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003ccolgroup cols=\"4\"\u003e\u003c/colgroup\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eResidual deviance\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eNull deviance\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eAIC\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eModels\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cem\u003eSGD\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e14.33\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e327.4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e107.2\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eResampling\u0026thinsp;+\u0026thinsp;imputation\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e290.2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e327.4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e377.1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eImputation\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e45.78\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e327.4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e132.7\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eResampling\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colspan=\"4\"\u003e\n \u003cp\u003e\u003cem\u003eLR\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e21.99\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e294.1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e116.9\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eResampling\u0026thinsp;+\u0026thinsp;imputation\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e259.9\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e294.1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e348.8\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eImputation\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e49.21\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e294.1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e138.1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eResampling\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cem\u003eDT\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e6.092\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e35.05\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e106.7\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eResampling\u0026thinsp;+\u0026thinsp;imputation\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e12.87\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e35.05\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e107.5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eImputation\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e29.03\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e35.05\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e123.7\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eResampling\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colspan=\"4\"\u003e\n \u003cp\u003e\u003cem\u003eRF\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e4.609\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e39.91\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e101.7\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eResampling\u0026thinsp;+\u0026thinsp;imputation\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e35.69\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e39.91\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e126.8\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eImputation\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e8.945\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e39.91\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e100.1\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eResampling\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colspan=\"4\"\u003e\n \u003cp\u003e\u003cem\u003eMLP\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e152.3\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e405.4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e242.5\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eResampling\u0026thinsp;+\u0026thinsp;imputation\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e380.8\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e405.4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e465\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eImputation\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e175.4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e405.4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e259.5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eResampling\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cem\u003eOverall\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e546.6\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e1146\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e966.3\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eResampling\u0026thinsp;+\u0026thinsp;imputation\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e1060\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e1146\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e1474\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eImputation\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e611.3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e1146\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e1025\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eResampling\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003c/div\u003e\n\u003c/div\u003e"},{"header":"4. Discussion","content":"\u003cp\u003e \u003cdiv class=\"BlockQuote\"\u003e \u003cp\u003eThe study demonstrated the influence of resampling and imputation techniques on the prediction performance of three types of machine learning models trained to detect mastitis alarms from automated milking system data. Features included quarter and cow level conductivity and in-line somatic cell count from a conventional dairy farm in Germany. Three types of classifiers we evaluated are classical discriminative classifiers (SGD and LR), ensemble classifiers (DT and RF) and a neural network based classifier (MLP). We found that imputation techniques improved the prediction performance for RF and DT models among which SI (kappa\u0026thinsp;=\u0026thinsp;0.43 and 0.36, respectively) was higher than MI (kappa\u0026thinsp;=\u0026thinsp;0.41 and 0.36, respectively), CC (kappa\u0026thinsp;=\u0026thinsp;0.40 and 0.31, respectively) and LI (kappa\u0026thinsp;=\u0026thinsp;0.27 and 0.26, respectively). This trend was confirmed by the ROC curves for RF and DT that had higher TPR (\u0026gt;\u0026thinsp;90%) and lower FPR (\u0026lt;\u0026thinsp;10%) for RF with imputation methods compared to CC (Appendix 1). DT models also had better performance than CC. Poulos \u0026amp; Valle [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e], testing RF models at various levels of missing data, compared the performance of the models with complete case or missing imputation and reported also a higher performance of ensemble models to which missing values were imputed than CC. In the current study, linear interpolation did not perform as good as simple imputation or multiple imputation for random forest and decision tree models. According to [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e], this could be due to no data segregation before applying linear imputation to datasets. They suggest that LI methods estimate the value of the missing data based on two adjacent data points in a one- dimensional sequence. Hence, for datasets where many consecutive data points are missing, such as in AMS data, the performance of LI may not be optimal. Yet, ensemble methods work by segregating data into similar packets small enough to identify their inherent patterns in the terminal nodes. For example, decision trees having two kinds of nodes, determine each leaf node that has a class label with a majority vote of training examples reached by the leaf. Further, they treat each internal node to represent a question on features that will be branching out according to the answers found. Hence they split leaves of a tree until questions are exhausted [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]. Therefore, it could be that these intrinsic characteristics of the ensemble models and LI led to lower performance than other imputation techniques or complete cases. Following [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e] approach, it could be beneficial to segregate the data prior to submitting it to LI for better results. This may not be useful for ensemble models, reputed robust enough to yield good performance with simple imputation techniques and sometimes without imputing missing values as explained above [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e]. Indeed, two of the top 10 models in this study were RF and DT without missing imputation or resampling (kappa\u0026thinsp;=\u0026thinsp;0.442 (No7), and 0.416 (No10), respectively).\u003c/p\u003e \u003cp\u003eA different trend was observed with SGD and LR where CC had higher scores than MI, SI and LI, respectively. Although the overall Kappa scores for these models were lower than ensemble models. The ROC curves with higher or similar performance for CC than imputation techniques regardless of the resampling techniques confirm this (Appendix 1). Mukaka et al. [\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e] also found better results for CC analysis compared to imputation techniques for binary outcomes with LR and recommended complementing the use of imputation techniques with CC analysis. Other authors [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e, \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e] found the opposite and suggested imputation was better than CC analysis. On one hand this can be explained by the fact that imputation techniques, especially for multiple imputation, increase the variability in the outcome values that inflates the standard error of the effect size estimate, probably caused by a random component added to the missing outcome values [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e, \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e, \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e]. On the other hand, the difference can also be attributed to the mechanisms of occurrence of missing values for which [\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e] provided an in depth analysis and suggested a thorough examination before deciding on the imputation method to apply.\u003c/p\u003e \u003cp\u003eThe linear interpolation for MLP had better performance metrics (kappa\u0026thinsp;=\u0026thinsp;0.27) and ROC characteristics than the other imputation techniques for the same model (Appendix 1). The method relies mostly on time dependent missing value imputation as opposed to inter-attribute correlations employed by other imputation techniques [\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e]. For this reason LI is especially efficient for time series and has been reported to improve the performance of neural network based classifier in other studies [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e, \u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThe performance of ML models from resampled datasets showed a similar trend with the missing imputation. The RF models had highest metrics, followed by DT, MLP, LR and SGD, respectively. Resampling data with SVMSMOTE seemed to result in better performance when subjected to SGD and LR models reported to perform better with more balanced datasets [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]. The MLP and RF models had better classification performance with SMOTE which is consistent with the findings reported in other studies [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e, \u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThe evaluation of model fit revealed that both resampling and missing imputation are relevant to explain the performance of most of the tested ML models. The MLP model had fair performance with linear interpolation imputation without resampling. This is in line with reports that the method improves the performance of neural network based models, hence could be applied to these types of ML models without resampling [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e, \u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e]. The same behavior was observed for simple and multiple imputated data fitted to decision tree models that resulted in good performance (kappa\u0026thinsp;=\u0026thinsp;0.465 and 0.455 respectively). The SVMSMOTE resampling method for SGD and LR performed better without imputation than the data where missing values were imputed. This suggests that the improvement in the class imbalance between majority and minority class achieved through borderline classification of SVMSMOTE outweighs the need for replacing missing values prior to analysis for these classifiers [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e, \u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e, \u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e]. Indeed, studies have applied resampling methods without imputation with satisfactory prediction performance. Random Forest, DT and to some extent MLP had models with fair to good performance (kappa\u0026thinsp;=\u0026thinsp;0.322 to 0.442) without resampling or missing value imputation [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e, \u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e, \u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e]. These models are reported to be robust enough to handle imbalance and missing values. They are sometimes used for preprocessing the data and predictions [\u003cspan additionalcitationids=\"CR42\" citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eOur findings suggest that depending on the ML models of interest, missing value imputation and resampling techniques need careful consideration. Complete case analysis had higher kappa score than missing imputation techniques for LR and SGD, while RF, DT and MLP had higher performance with imputation techniques. We observed wide variations between models as well as agreement between accuracy, F1 score, precision and recall metric with kappa. For ensemble models resampling with SMOTE or SVMSMOTE improves classification performance with either simple imputation or complete cases. MLP classification is improved by LI and SMOTE resampling, while SGD and LR have better classification performance with complete cases and SVMSMOTE resampling. Hence, careful consideration of the mechanism of occurrence of missing values, class imbalance and the intended ML models to train the data on are needed to generate more reliable mastitis predictions with AMS data.\u003c/p\u003e \u003c/div\u003e \u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eSupplementary Materials:\u0026nbsp;\u003c/strong\u003eThe following supporting information can be downloaded at: www.mdpi.com/xxx/s1, Figure S1: Receiver Operating curves for the resampling methods; Figure S2: Receiver Operating Curves for the Machine learning model performance with resampling methods; Table S1: Rankings of the prediction scores of the fitted machine learning models with both resampling and imputation methods (kappa \u0026gt;0.20).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor Contributions:\u0026nbsp;\u003c/strong\u003eConceptualization: K.O, A.C and K.T; methodology: K.O and C.A; data processing and analysis: KO; validation: A.C, K.T, K.O, A.T, A.B, D.M; original draft preparation: K.O; review and editing: K.T, A.C, A.T, A.B, D.M; supervision: A.T, A.B; project administration: K.T., A.T, A.B; funding acquisition: A.T, A.B. All authors have read and agreed to the published version of the manuscript.\u0026rdquo; Please turn to the CRediT taxonomy for the term explanation. Authorship must be limited to those who have contributed substantially to the work reported.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding:\u003c/strong\u003e This research was funded by the Federal Ministry of Food and Agriculture (BMEL, Germany) based on a resolution of the German Bundestag. The article is funded from the project MEDICow, funding was carried out by the Federal Office of agriculture and food (BLE, Germany) within the framework of the federal program \u0026ldquo;Livestock Husbandry\u0026quot; (grant number: 28N206601).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData Availability Statement:\u0026nbsp;\u003c/strong\u003eThe data used to produce the results presented in this manuscript will be made available upon reasonable request. The source code used to produce the results presented in this manuscript can be availaed freely at (https://github.com/Okashongwe/resampling_mast.git).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgments:\u003c/strong\u003e Authors wish to thank the dairy farm for the support and provision of data.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConflicts of Interest:\u0026nbsp;\u003c/strong\u003eThe authors declare no conflict of interest.\u003c/p\u003e\n\u003cp dir=\"LTR\"\u003e\u003cstrong\u003eDisclaimer/Publisher’s Note:\u003c/strong\u003e The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n \u003cli\u003eCheng, W.N.; Han, S.G. Bovine Mastitis: Risk Factors, Therapeutic Strategies, and Alternative Treatments\u0026mdash;A Review. \u003cem\u003eAsian-Australasian journal of animal sciences\u003c/em\u003e \u003cstrong\u003e2020\u003c/strong\u003e, \u003cem\u003e33\u003c/em\u003e, 1699.\u003c/li\u003e\n \u003cli\u003eAghamohammadi, M.; Haine, D.; Kelton, D.F.; Barkema, H.W.; Hogeveen, H.; Keefe, G.P.; Dufour, S. Herd-Level Mastitis-Associated Costs on Canadian Dairy Farms. \u003cem\u003eFront. Vet. Sci.\u003c/em\u003e \u003cstrong\u003e2018\u003c/strong\u003e, \u003cem\u003e5\u003c/em\u003e, doi:10.3389/fvets.2018.00100.\u003c/li\u003e\n \u003cli\u003eHogeveen, H.; Steeneveld, W.; Wolf, C.A. Production Diseases Reduce the Efficiency of Dairy Production: A Review of the Results, Methods, and Approaches Regarding the Economics of Mastitis. \u003cem\u003eAnnual Review of Resource Economics\u003c/em\u003e \u003cstrong\u003e2019\u003c/strong\u003e, \u003cem\u003e11\u003c/em\u003e, 289\u0026ndash;312, doi:10.1146/annurev-resource-100518-093954.\u003c/li\u003e\n \u003cli\u003eBaskaran, S.A.; Kazmer, G.; Hinckley, L.; Andrew, S.; Venkitanarayanan, K. Antibacterial Effect of Plant-Derived Antimicrobials on Major Bacterial Mastitis Pathogens in Vitro. \u003cem\u003eJournal of dairy science\u003c/em\u003e \u003cstrong\u003e2009\u003c/strong\u003e, \u003cem\u003e92\u003c/em\u003e, 1423\u0026ndash;1429.\u003c/li\u003e\n \u003cli\u003eMartins, S.A.; Martins, V.C.; Cardoso, F.A.; Germano, J.; Rodrigues, M.; Duarte, C.; Bexiga, R.; Cardoso, S.; Freitas, P.P. Biosensors for On-Farm Diagnosis of Mastitis. \u003cem\u003eFrontiers in bioengineering and biotechnology\u003c/em\u003e \u003cstrong\u003e2019\u003c/strong\u003e, \u003cem\u003e7\u003c/em\u003e, 186.\u003c/li\u003e\n \u003cli\u003eBernhardt, H.; H\u0026ouml;hendinger, M.; Gr\u0026auml;ff, A.; Hijazi, O.; H\u0026ouml;ld, M.; Reger, M.; Stumpenhausen, J. Development of Automatic Milking in Germany. In Proceedings of the 2019 ASABE Annual International Meeting; American Society of Agricultural and Biological Engineers, 2019; p. 1.\u003c/li\u003e\n \u003cli\u003eSteeneveld, W.; Vernooij, J.; Hogeveen, H. Effect of Sensor Systems for Cow Management on Milk Production, Somatic Cell Count, and Reproduction. \u003cem\u003eJournal of dairy science\u003c/em\u003e \u003cstrong\u003e2015\u003c/strong\u003e, \u003cem\u003e98\u003c/em\u003e, 3896\u0026ndash;3905.\u003c/li\u003e\n \u003cli\u003eSitkowska, B.; Piwczynski, D.; Aerts, J.; Kolenda, M.; \u0026Ouml;ZKAYA, S. Detection of High Levels of Somatic Cells in Milk on Farms Equippedwith an Automatic Milking System by Decision Trees Technique. \u003cem\u003eTurkish Journal of Veterinary \u0026amp; Animal Sciences\u003c/em\u003e \u003cstrong\u003e2017\u003c/strong\u003e, \u003cem\u003e41\u003c/em\u003e, 532\u0026ndash;540.\u003c/li\u003e\n \u003cli\u003eBonestroo, J.; van der Voort, M.; Hogeveen, H.; Emanuelson, U.; Klaas, I.C.; Fall, N. Forecasting Chronic Mastitis Using Automatic Milking System Sensor Data and Gradient-Boosting Classifiers. \u003cem\u003eComputers and Electronics in Agriculture\u003c/em\u003e \u003cstrong\u003e2022\u003c/strong\u003e, \u003cem\u003e198\u003c/em\u003e, 107002.\u003c/li\u003e\n \u003cli\u003eBobbo, T.; Biffani, S.; Taccioli, C.; Penasa, M.; Cassandro, M. Comparison of Machine Learning Methods to Predict Udder Health Status Based on Somatic Cell Counts in Dairy Cows. \u003cem\u003eScientific Reports\u003c/em\u003e \u003cstrong\u003e2021\u003c/strong\u003e, \u003cem\u003e11\u003c/em\u003e, 13642.\u003c/li\u003e\n \u003cli\u003eHyde, R.M.; Down, P.M.; Bradley, A.J.; Breen, J.E.; Hudson, C.; Leach, K.A.; Green, M.J. Automated Prediction of Mastitis Infection Patterns in Dairy Herds Using Machine Learning. \u003cem\u003eScientific reports\u003c/em\u003e \u003cstrong\u003e2020\u003c/strong\u003e, \u003cem\u003e10\u003c/em\u003e, 4289.\u003c/li\u003e\n \u003cli\u003ePost, C.; Rietz, C.; B\u0026uuml;scher, W.; M\u0026uuml;ller, U. Using Sensor Data to Detect Lameness and Mastitis Treatment Events in Dairy Cows: A Comparison of Classification Models. \u003cem\u003eSensors\u003c/em\u003e \u003cstrong\u003e2020\u003c/strong\u003e, \u003cem\u003e20\u003c/em\u003e, 3863.\u003c/li\u003e\n \u003cli\u003eFadul-Pacheco, L.; Delgado, H.; Cabrera, V.E. Exploring Machine Learning Algorithms for Early Prediction of Clinical Mastitis. \u003cem\u003eInternational Dairy Journal\u003c/em\u003e \u003cstrong\u003e2021\u003c/strong\u003e, \u003cem\u003e119\u003c/em\u003e, 105051, doi:10.1016/j.idairyj.2021.105051.\u003c/li\u003e\n \u003cli\u003eAbdul Ghafoor, N.; Sitkowska, B. MasPA: A Machine Learning Application to Predict Risk of Mastitis in Cattle from AMS Sensor Data. \u003cem\u003eAgriEngineering\u003c/em\u003e \u003cstrong\u003e2021\u003c/strong\u003e, \u003cem\u003e3\u003c/em\u003e, 575\u0026ndash;583.\u003c/li\u003e\n \u003cli\u003eHogeveen, H.; Kamphuis, C.; Steeneveld, W.; Mollenhorst, H. Sensors and Clinical Mastitis\u0026mdash;The Quest for the Perfect Alert. \u003cem\u003eSensors\u003c/em\u003e \u003cstrong\u003e2010\u003c/strong\u003e, \u003cem\u003e10\u003c/em\u003e, 7991\u0026ndash;8009.\u003c/li\u003e\n \u003cli\u003eLi, Z.; Jiang, Y.; Hu, C.; Peng, Z. Recent Progress on Decoupling Diagnosis of Hybrid Failures in Gear Transmission Systems Using Vibration Sensor Signal: A Review. \u003cem\u003eMeasurement\u003c/em\u003e \u003cstrong\u003e2016\u003c/strong\u003e, \u003cem\u003e90\u003c/em\u003e, 4\u0026ndash;19.\u003c/li\u003e\n \u003cli\u003eDominiak, K.N.; Kristensen, A.R. Prioritizing Alarms from Sensor-Based Detection Models in Livestock Production - A Review on Model Performance and Alarm Reducing Methods. \u003cem\u003eComputers and Electronics in Agriculture\u003c/em\u003e \u003cstrong\u003e2017\u003c/strong\u003e, \u003cem\u003e133\u003c/em\u003e, 46\u0026ndash;67, doi:https://doi.org/10.1016/j.compag.2016.12.008.\u003c/li\u003e\n \u003cli\u003eVan Buuren, S. \u003cem\u003eFlexible Imputation of Missing Data\u003c/em\u003e; CRC press, 2018;\u003c/li\u003e\n \u003cli\u003eMadley-Dowd, P.; Hughes, R.; Tilling, K.; Heron, J. The Proportion of Missing Data Should Not Be Used to Guide Decisions on Multiple Imputation. \u003cem\u003eJournal of clinical epidemiology\u003c/em\u003e \u003cstrong\u003e2019\u003c/strong\u003e, \u003cem\u003e110\u003c/em\u003e, 63\u0026ndash;73.\u003c/li\u003e\n \u003cli\u003ePham, T.M.; Pandis, N.; White, I.R. Missing Data: Issues, Concepts, Methods. \u003cem\u003eSeminars in Orthodontics\u003c/em\u003e \u003cstrong\u003e2024\u003c/strong\u003e, \u003cem\u003e30\u003c/em\u003e, 37\u0026ndash;44, doi:https://doi.org/10.1053/j.sodo.2024.01.007.\u003c/li\u003e\n \u003cli\u003eWhite, I.R.; Royston, P.; Wood, A.M. Multiple Imputation Using Chained Equations: Issues and Guidance for Practice. \u003cem\u003eStatistics in medicine\u003c/em\u003e \u003cstrong\u003e2011\u003c/strong\u003e, \u003cem\u003e30\u003c/em\u003e, 377\u0026ndash;399.\u003c/li\u003e\n \u003cli\u003eNoor, M.; Al Bakri, A.; Yahaya, A.; Ramli, N.; Fitri, N. Estimation of Missing Values in Environmental Data Set Using Interpolation Technique: Fitting on Lognormal Distribution. \u003cem\u003eAust. J. Basic Appl. Sci\u003c/em\u003e \u003cstrong\u003e2013\u003c/strong\u003e, \u003cem\u003e7\u003c/em\u003e, 336\u0026ndash;341.\u003c/li\u003e\n \u003cli\u003eHuang, G. Missing Data Filling Method Based on Linear Interpolation and Lightgbm. In Proceedings of the Journal of Physics: Conference Series; IOP Publishing, 2021; Vol. 1754, p. 012187.\u003c/li\u003e\n \u003cli\u003eKhushi, M.; Shaukat, K.; Alam, T.M.; Hameed, I.A.; Uddin, S.; Luo, S.; Yang, X.; Reyes, M.C. A Comparative Performance Analysis of Data Resampling Methods on Imbalance Medical Data. \u003cem\u003eIEEE Access\u003c/em\u003e \u003cstrong\u003e2021\u003c/strong\u003e, \u003cem\u003e9\u003c/em\u003e, 109960\u0026ndash;109975.\u003c/li\u003e\n \u003cli\u003eJohnson, J.M.; Khoshgoftaar, T.M. A Survey on Classifying Big Data with Label Noise. \u003cem\u003eJ. Data and Information Quality\u003c/em\u003e \u003cstrong\u003e2022\u003c/strong\u003e, \u003cem\u003e14\u003c/em\u003e, 23:1-23:43, doi:10.1145/3492546.\u003c/li\u003e\n \u003cli\u003eNguyen, H.M.; Cooper, E.W.; Kamei, K. Borderline Over-Sampling for Imbalanced Data Classification. \u003cem\u003eInternational Journal of Knowledge Engineering and Soft Data Paradigms\u003c/em\u003e \u003cstrong\u003e2011\u003c/strong\u003e, \u003cem\u003e3\u003c/em\u003e, 4\u0026ndash;21, doi:10.1504/IJKESDP.2011.039875.\u003c/li\u003e\n \u003cli\u003eGhorbani, R.; Ghousi, R. Comparing Different Resampling Methods in Predicting Students\u0026rsquo; Performance Using Machine Learning Techniques. \u003cem\u003eIEEE Access\u003c/em\u003e \u003cstrong\u003e2020\u003c/strong\u003e, \u003cem\u003e8\u003c/em\u003e, 67899\u0026ndash;67911, doi:10.1109/ACCESS.2020.2986809.\u003c/li\u003e\n \u003cli\u003eBagui, S.S.; Mink, D.; Bagui, S.C.; Subramaniam, S. Determining Resampling Ratios Using BSMOTE and SVM-SMOTE for Identifying Rare Attacks in Imbalanced Cybersecurity Data. \u003cem\u003eComputers\u003c/em\u003e \u003cstrong\u003e2023\u003c/strong\u003e, \u003cem\u003e12\u003c/em\u003e, 204, doi:10.3390/computers12100204.\u003c/li\u003e\n \u003cli\u003eTarimo, C.S.; Bhuyan, S.S.; Li, Q.; Ren, W.; Mahande, M.J.; Wu, J. Combining Resampling Strategies and Ensemble Machine Learning Methods to Enhance Prediction of Neonates with a Low Apgar Score After Induction of Labor in Northern Tanzania. \u003cem\u003eRisk Management and Healthcare Policy\u003c/em\u003e \u003cstrong\u003e2021\u003c/strong\u003e, \u003cem\u003e14\u003c/em\u003e, 3711\u0026ndash;3720, doi:10.2147/RMHP.S331077.\u003c/li\u003e\n \u003cli\u003eEbrahimi, M.; Mohammadi-Dehcheshmeh, M.; Ebrahimie, E.; Petrovski, K.R. Comprehensive Analysis of Machine Learning Models for Prediction of Sub-Clinical Mastitis: Deep Learning and Gradient-Boosted Trees Outperform Other Models. \u003cem\u003eComputers in Biology and Medicine\u003c/em\u003e \u003cstrong\u003e2019\u003c/strong\u003e, \u003cem\u003e114\u003c/em\u003e, 103456, doi:10.1016/j.compbiomed.2019.103456.\u003c/li\u003e\n \u003cli\u003eAbidin, N.Z.; Ritahani, A.; A., N. Performance Analysis of Machine Learning Algorithms for Missing Value Imputation. \u003cem\u003eijacsa\u003c/em\u003e \u003cstrong\u003e2018\u003c/strong\u003e, \u003cem\u003e9\u003c/em\u003e, doi:10.14569/IJACSA.2018.090660.\u003c/li\u003e\n \u003cli\u003eShah, A.D.; Bartlett, J.W.; Carpenter, J.; Nicholas, O.; Hemingway, H. Comparison of Random Forest and Parametric Imputation Models for Imputing Missing Data Using MICE: A CALIBER Study. \u003cem\u003eAmerican Journal of Epidemiology\u003c/em\u003e \u003cstrong\u003e2014\u003c/strong\u003e, \u003cem\u003e179\u003c/em\u003e, 764\u0026ndash;774, doi:10.1093/aje/kwt312.\u003c/li\u003e\n \u003cli\u003eMukaka, M.; White, S.A.; Terlouw, D.J.; Mwapasa, V.; Kalilani-Phiri, L.; Faragher, E.B. Is Using Multiple Imputation Better than Complete Case Analysis for Estimating a Prevalence (Risk) Difference in Randomized Controlled Trials When Binary Outcome Observations Are Missing? \u003cem\u003eTrials\u003c/em\u003e \u003cstrong\u003e2016\u003c/strong\u003e, \u003cem\u003e17\u003c/em\u003e, 341, doi:10.1186/s13063-016-1473-3.\u003c/li\u003e\n \u003cli\u003eMoritz, S.; Bartz-Beielstein, T. imputeTS: Time Series Missing Value Imputation in R. \u003cem\u003eThe R Journal\u003c/em\u003e \u003cstrong\u003e2017\u003c/strong\u003e, \u003cem\u003e9\u003c/em\u003e, 207, doi:10.32614/RJ-2017-009.\u003c/li\u003e\n \u003cli\u003ePark, I.; Kim, H.S.; Lee, J.; Kim, J.H.; Song, C.H.; Kim, H.K. Temperature Prediction Using the Missing Data Refinement Model Based on a Long Short-Term Memory Neural Network. \u003cem\u003eAtmosphere\u003c/em\u003e \u003cstrong\u003e2019\u003c/strong\u003e, \u003cem\u003e10\u003c/em\u003e, 718, doi:10.3390/atmos10110718.\u003c/li\u003e\n \u003cli\u003eMoon, T.; Hong, S.; Choi, H.Y.; Jung, D.H.; Chang, S.H.; Son, J.E. Interpolation of Greenhouse Environment Data Using Multilayer Perceptron. \u003cem\u003eComputers and Electronics in Agriculture\u003c/em\u003e \u003cstrong\u003e2019\u003c/strong\u003e, \u003cem\u003e166\u003c/em\u003e, 105023, doi:10.1016/j.compag.2019.105023.\u003c/li\u003e\n \u003cli\u003eBuabeng, A.; Simons, A.; Frempong, N.K.; Ziggah, Y.Y. A Novel Hybrid Predictive Maintenance Model Based on Clustering, Smote and Multi-Layer Perceptron Neural Network Optimised with Grey Wolf Algorithm. \u003cem\u003eSN Appl. Sci.\u003c/em\u003e \u003cstrong\u003e2021\u003c/strong\u003e, \u003cem\u003e3\u003c/em\u003e, 593, doi:10.1007/s42452-021-04598-1.\u003c/li\u003e\n \u003cli\u003eWongvorachan, T.; He, S.; Bulut, O. A Comparison of Undersampling, Oversampling, and SMOTE Methods for Dealing with Imbalanced Classification in Educational Data Mining. \u003cem\u003eInformation\u003c/em\u003e \u003cstrong\u003e2023\u003c/strong\u003e, \u003cem\u003e14\u003c/em\u003e, 54, doi:10.3390/info14010054.\u003c/li\u003e\n \u003cli\u003eJian, C.; Gao, J.; Ao, Y. A New Sampling Method for Classifying Imbalanced Data Based on Support Vector Machine Ensemble. \u003cem\u003eNeurocomputing\u003c/em\u003e \u003cstrong\u003e2016\u003c/strong\u003e, \u003cem\u003e193\u003c/em\u003e, 115\u0026ndash;122, doi:https://doi.org/10.1016/j.neucom.2016.02.006.\u003c/li\u003e\n \u003cli\u003ePoulos, J.; Valle, R. Missing Data Imputation for Supervised Learning. \u003cem\u003eApplied Artificial Intelligence\u003c/em\u003e \u003cstrong\u003e2018\u003c/strong\u003e, \u003cem\u003e32\u003c/em\u003e, 186\u0026ndash;196.\u003c/li\u003e\n \u003cli\u003eUpadhyay, A.; Singh, M.; Yadav, V.K. Improvised Number Identification Using SVM and Random Forest Classifiers. \u003cem\u003eJournal of Information and Optimization Sciences\u003c/em\u003e \u003cstrong\u003e2020\u003c/strong\u003e, \u003cem\u003e41\u003c/em\u003e, 387\u0026ndash;394, doi:10.1080/02522667.2020.1723934.\u003c/li\u003e\n \u003cli\u003ePhiri, D.; Morgenroth, J.; Xu, C.; Hermosilla, T. Effects of Pre-Processing Methods on Landsat OLI-8 Land Cover Classification Using OBIA and Random Forests Classifier. \u003cem\u003eInternational Journal of Applied Earth Observation and Geoinformation\u003c/em\u003e \u003cstrong\u003e2018\u003c/strong\u003e, \u003cem\u003e73\u003c/em\u003e, 170\u0026ndash;178, doi:10.1016/j.jag.2018.06.014.\u003c/li\u003e\n \u003cli\u003eIliou, T.; Anagnostopoulos, C.-N.; Stephanakis, I.M.; Anastassopoulos, G. A Novel Data Preprocessing Method for Boosting Neural Network Performance: A Case Study in Osteoporosis Prediction. \u003cem\u003eInformation Sciences\u003c/em\u003e \u003cstrong\u003e2017\u003c/strong\u003e, \u003cem\u003e380\u003c/em\u003e, 92\u0026ndash;100, doi:10.1016/j.ins.2015.10.026.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"Leibniz Institut für Agrartechnk und Bioökonomie","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"oversampling, undersampling, missing value imputation, dairy cows, machine learning performance analysis","lastPublishedDoi":"10.21203/rs.3.rs-4629327/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4629327/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eMissing data and class imbalance represent a hindrance to accurate prediction of rare events such as mastitis (udder inflammation). Various methods are susceptible to handle the problem, however, little is known about their individual and combined effects on the performance of ML models fitted to AMS (automated milking system) data for mastitis prediction. We apply imputation and resampling to improve performance metrics of classifiers (logistic regression, stochastic gradient descent, multilayer perceptron, decision tree and random forest). Three imputation methods: simple imputer (SI), multiple imputer (MICE) and linear interpolation (LI) were compared to complete cases. Three resampling procedures: synthetic minority oversampling technique (SOMTE), Support Vector Machine SMOTE and SMOTE with Edited Nearest Neighbours were compared. We evaluated different techniques by calculating precision, recall, F1 Score and compared models based on kappa score. Both imputation and resampling techniques improved models performance. Complete case analysis suited the Stochastic Gradient Descent (SGD) Classifier better than resampling or imputation (kappa=0.280). The Logistic regression (LR) performed better with SVMSMOTE rand no imputation (kappa= 0.218). The Random Forest (RF), Decision Tree (DT) and Multilayer Perceptron (MLP) performed better than SGD and LR and handled well class imbalance and missing values without preprocessing. We propose careful selection of the technique to handle class imbalance and missing value prior to subjecting data to ML model is crucial to attain best ML model performance.\u003c/p\u003e","manuscriptTitle":"Influence of Preprocessing Methods of Automated Milking Systems Data on the Prediction of Mastitis with Machine Learning Models","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-06-25 14:11:12","doi":"10.21203/rs.3.rs-4629327/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"ef992583-ecfa-49a2-93ea-6223dda6fbc4","owner":[],"postedDate":"June 25th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":33647070,"name":"Animal Science"}],"tags":[],"updatedAt":"2024-06-25T14:11:12+00:00","versionOfRecord":[],"versionCreatedAt":"2024-06-25 14:11:12","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-4629327","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4629327","identity":"rs-4629327","version":["v1"]},"buildId":"omnImTCwR2MFx8CMYfrG7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-08-14T06:25:32.811723+00:00
License: CC-BY-4.0