Improving COVID-19 Mortality Predictions: A Stacking Ensemble Approach with Diverse Classifiers | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Improving COVID-19 Mortality Predictions: A Stacking Ensemble Approach with Diverse Classifiers Farideh Mohtasham, MohamadAmin Pourhoseingholi, Seyed Saeed Hashemi Nazari, and 2 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-5018487/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Ensemble approaches are vital for developing effective machine learning methods by integrating multiple models to enhance performance and reduce bias and variance. This study utilized ensemble techniques to predict COVID-19 mortality using various classifiers. We first mapped the original dataset to a lower-dimensional space to improve training diversity. We then trained multiple base classifiers and ensemble methods, assessing their diversity through pairwise evaluations to create diverse combinations. A Stacking ensemble method was implemented with different meta-learners for improved predictive performance. All models were rigorously evaluated using standard discrimination and calibration metrics, along with statistical tests to identify significant performance differences. Various feature importance methods were applied to clarify the contributors to our model's predictions. The experimental results demonstrated the superiority of our stacking framework, specifically combining Random Forest and Extreme Gradient Boosting (XGBoost) with a Neural Network as the meta-learner on COVID-19 mortality prediction. This model achieved an accuracy of 0.914 (95% CI: 0.898, 0.928), precision of 0.818, F1-score of 0.801, Matthew’s correlation coefficient (MCC) of 0.746, and a ROC AUC of 0.955. These findings indicate that our framework is more effective than individual classifiers and existing ensemble methods, providing valuable insights for medical decision-making. Health sciences/Medical research/Epidemiology Biological sciences/Computational biology and bioinformatics/Machine learning Ensemble learning machine learning algorithms feature selection diversity measurement stacking Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Figure 9 Figure 10 Figure 11 Figure 12 Full Text Additional Declarations No competing interests reported. Tables 1 to 6 are available in the Supplementary Files section Supplementary Files additionalfile.pdf Tables.docx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-5018487","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":361283101,"identity":"7efd695f-ea12-444c-a739-372f3c9d0264","order_by":0,"name":"Farideh Mohtasham","email":"","orcid":"","institution":"Shahid Beheshti University of Medical Sciences","correspondingAuthor":false,"prefix":"","firstName":"Farideh","middleName":"","lastName":"Mohtasham","suffix":""},{"id":361283102,"identity":"788f939e-2f4c-4693-a309-2fdea8670633","order_by":1,"name":"MohamadAmin Pourhoseingholi","email":"","orcid":"","institution":"University of Nottingham","correspondingAuthor":false,"prefix":"","firstName":"MohamadAmin","middleName":"","lastName":"Pourhoseingholi","suffix":""},{"id":361283103,"identity":"f2244514-48c8-4194-8b29-67d7c427ad47","order_by":2,"name":"Seyed Saeed Hashemi Nazari","email":"","orcid":"","institution":"Shahid Beheshti University of Medical Sciences (SBMU)","correspondingAuthor":false,"prefix":"","firstName":"Seyed","middleName":"Saeed Hashemi","lastName":"Nazari","suffix":""},{"id":361283104,"identity":"6cdd0e0d-1723-4362-9ab4-026d7c5e1d2d","order_by":3,"name":"Kaveh Kavousi","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAwklEQVRIiWNgGAWjYBACCR4IzcMPFeAhXotkAzOJWhgMDjAT6TDJnsPPJD621ckY38g/wPCjhkHGvIGAFmneNjPJmW2HecxuJDMw9hxj4JE5QECLHD+DsTFv2wGwFgbeBgYeCUIOk+Nn/wzUUsdjPANoy19itEjz9hg+5m1j5jGQSGZgJsoWyZ4zhQ9nnDvMI3HmscFhmWMShLVInEnfcOBDWZ09f3viw4dvamzsCWpBAQeARpCkYRSMglEwCkYBDgAA9t0yKO3aljwAAAAASUVORK5CYII=","orcid":"","institution":"University of Tehran","correspondingAuthor":true,"prefix":"","firstName":"Kaveh","middleName":"","lastName":"Kavousi","suffix":""},{"id":361283105,"identity":"bf0ff858-cf45-45ed-92de-82679febc31c","order_by":4,"name":"Mohammad Reza Zali","email":"","orcid":"","institution":"Shahid Beheshti University of Medical Sciences","correspondingAuthor":false,"prefix":"","firstName":"Mohammad","middleName":"Reza","lastName":"Zali","suffix":""}],"badges":[],"createdAt":"2024-09-02 12:53:21","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-5018487/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-5018487/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":65895702,"identity":"3cf72556-d012-4061-a8fd-978852b74f30","added_by":"auto","created_at":"2024-10-04 06:29:30","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":58779,"visible":true,"origin":"","legend":"\u003cp\u003eComorbidity Analysis of our trained data\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-5018487/v1/458ae221b4fc29e6792385c6.png"},{"id":65896087,"identity":"15fecbb8-6c3d-4b55-81dc-7fb47fe2f7df","added_by":"auto","created_at":"2024-10-04 06:37:30","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":58729,"visible":true,"origin":"","legend":"\u003cp\u003eCorrelation between the selected features and the outcome class in the training dataset\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-5018487/v1/b367b5aa8a7858a2c91b478f.png"},{"id":65896910,"identity":"d9b1f860-f6a4-415b-9c84-9e28e781433e","added_by":"auto","created_at":"2024-10-04 06:45:30","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":68424,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eComparison of performance of Base, Boosting and Bagging Machine Learning Algorithms using 10 repeated 10-fold Cross Validation\u003c/em\u003e\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-5018487/v1/4b74137f75faae7949983b1e.png"},{"id":65895712,"identity":"01aa323c-4a26-4e0a-86d8-1ad80fa2ed39","added_by":"auto","created_at":"2024-10-04 06:29:30","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":92532,"visible":true,"origin":"","legend":"\u003cp\u003eDisagreement among the classifiers' predictions\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-5018487/v1/46229ab2790ec115b530b6f5.png"},{"id":65895705,"identity":"dfa74ca8-4e21-470a-8c21-356399121d12","added_by":"auto","created_at":"2024-10-04 06:29:30","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":85752,"visible":true,"origin":"","legend":"\u003cp\u003eInterrater agreement among the classifiers' predictions\u003c/p\u003e","description":"","filename":"5.png","url":"https://assets-eu.researchsquare.com/files/rs-5018487/v1/501670bc62b9f51b0c7cb506.png"},{"id":65896089,"identity":"390ea22f-876d-48e4-9ef4-cc27917b6c2d","added_by":"auto","created_at":"2024-10-04 06:37:30","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":29950,"visible":true,"origin":"","legend":"\u003cp\u003eComparison of Diversity Metrics by Ensembles\u003c/p\u003e","description":"","filename":"6.png","url":"https://assets-eu.researchsquare.com/files/rs-5018487/v1/f0b313397523741815a7b4d4.png"},{"id":65895710,"identity":"ef5e646e-f37d-4ac2-943e-49cb0f6f67ea","added_by":"auto","created_at":"2024-10-04 06:29:30","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":77444,"visible":true,"origin":"","legend":"\u003cp\u003eModel Performance Metrics for selected stacked model\u003c/p\u003e","description":"","filename":"7.png","url":"https://assets-eu.researchsquare.com/files/rs-5018487/v1/ac5eeecefcdd2e5496b1f080.png"},{"id":65895714,"identity":"4670a8da-e983-4383-b5d6-faf35abfd709","added_by":"auto","created_at":"2024-10-04 06:29:30","extension":"png","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":39940,"visible":true,"origin":"","legend":"\u003cp\u003eROC curve for best stacked models on the test data\u003c/p\u003e","description":"","filename":"8.png","url":"https://assets-eu.researchsquare.com/files/rs-5018487/v1/27b20226927cb1145f84cc2a.png"},{"id":65897664,"identity":"99147e15-ec35-4303-890a-6c2b28684118","added_by":"auto","created_at":"2024-10-04 06:53:34","extension":"png","order_by":9,"title":"Figure 9","display":"","copyAsset":false,"role":"figure","size":49340,"visible":true,"origin":"","legend":"\u003cp\u003eThe performance of stack.nnet.rf.xgb model using repeated 10 repeated 10-fold Cross Validation dataset\u003c/p\u003e","description":"","filename":"9.png","url":"https://assets-eu.researchsquare.com/files/rs-5018487/v1/f391a86ba680213e9b823cfe.png"},{"id":65895707,"identity":"76759ec6-a537-4b0d-af0b-f3cc1c5c441b","added_by":"auto","created_at":"2024-10-04 06:29:30","extension":"png","order_by":10,"title":"Figure 10","display":"","copyAsset":false,"role":"figure","size":30062,"visible":true,"origin":"","legend":"\u003cp\u003eThe most influential predictors in determining the outcome of 'Death' in the predictions made by the stack.nnet.rf.xgb model\u003c/p\u003e","description":"","filename":"10.png","url":"https://assets-eu.researchsquare.com/files/rs-5018487/v1/9c46fac14836337c8242b4e7.png"},{"id":65896092,"identity":"e44708d0-57a8-48ef-b85c-9f9840c5810d","added_by":"auto","created_at":"2024-10-04 06:37:30","extension":"png","order_by":11,"title":"Figure 11","display":"","copyAsset":false,"role":"figure","size":42266,"visible":true,"origin":"","legend":"\u003cp\u003eInteraction effects between the feature Age and various clinical predictors in relation to the Death outcomes in the predictions made by the stack.nnet.rf.xgb model\u003c/p\u003e","description":"","filename":"11.png","url":"https://assets-eu.researchsquare.com/files/rs-5018487/v1/fc6b9e4a670d0e0ae99dd351.png"},{"id":65896094,"identity":"02248aae-2a68-408c-a0db-8543ff277297","added_by":"auto","created_at":"2024-10-04 06:37:30","extension":"png","order_by":12,"title":"Figure 12","display":"","copyAsset":false,"role":"figure","size":41409,"visible":true,"origin":"","legend":"\u003cp\u003eInterpretation of the stack model according to the SHAP (Shapley Additive Explanations) value\u003c/p\u003e","description":"","filename":"12.png","url":"https://assets-eu.researchsquare.com/files/rs-5018487/v1/292e7716041020ffa397548e.png"},{"id":73731306,"identity":"c5a464a6-db20-4033-9084-17f0e2210391","added_by":"auto","created_at":"2025-01-14 05:41:48","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":789362,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5018487/v1_covered_14af99c5-15ea-4026-8da4-e2ad487c9556.pdf"},{"id":65896911,"identity":"64167ef8-e294-40ac-8dba-82334ed917d4","added_by":"auto","created_at":"2024-10-04 06:45:30","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":418292,"visible":true,"origin":"","legend":"","description":"","filename":"additionalfile.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5018487/v1/a7b781f47cd2fe2ea79d021a.pdf"},{"id":65896090,"identity":"4da0f707-484a-4300-af05-feb4c91e7b33","added_by":"auto","created_at":"2024-10-04 06:37:30","extension":"docx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":31369,"visible":true,"origin":"","legend":"","description":"","filename":"Tables.docx","url":"https://assets-eu.researchsquare.com/files/rs-5018487/v1/606c91d31d0e05c2154d47f8.docx"}],"financialInterests":"\u003cp\u003eNo competing interests reported.\u003c/p\u003e\n\u003cp\u003eTables 1 to 6 are available in the Supplementary Files section\u003c/p\u003e","formattedTitle":"Improving COVID-19 Mortality Predictions: A Stacking Ensemble Approach with Diverse Classifiers","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Ensemble learning, machine learning algorithms, feature selection, diversity measurement, stacking","lastPublishedDoi":"10.21203/rs.3.rs-5018487/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-5018487/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eEnsemble approaches are vital for developing effective machine learning methods by integrating multiple models to enhance performance and reduce bias and variance. This study utilized ensemble techniques to predict COVID-19 mortality using various classifiers. We first mapped the original dataset to a lower-dimensional space to improve training diversity. We then trained multiple base classifiers and ensemble methods, assessing their diversity through pairwise evaluations to create diverse combinations. A Stacking ensemble method was implemented with different meta-learners for improved predictive performance. All models were rigorously evaluated using standard discrimination and calibration metrics, along with statistical tests to identify significant performance differences. Various feature importance methods were applied to clarify the contributors to our model's predictions. The experimental results demonstrated the superiority of our stacking framework, specifically combining Random Forest and Extreme Gradient Boosting (XGBoost) with a Neural Network as the meta-learner on COVID-19 mortality prediction. This model achieved an accuracy of 0.914 (95% CI: 0.898, 0.928), precision of 0.818, F1-score of 0.801, Matthew\u0026rsquo;s correlation coefficient (MCC) of 0.746, and a ROC AUC of 0.955. These findings indicate that our framework is more effective than individual classifiers and existing ensemble methods, providing valuable insights for medical decision-making.\u003c/p\u003e","manuscriptTitle":"Improving COVID-19 Mortality Predictions: A Stacking Ensemble Approach with Diverse Classifiers","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-10-04 06:29:25","doi":"10.21203/rs.3.rs-5018487/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"0df137a8-1a0f-4622-ae4e-d031007f4e8e","owner":[],"postedDate":"October 4th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":38439432,"name":"Health sciences/Medical research/Epidemiology"},{"id":38439433,"name":"Biological sciences/Computational biology and bioinformatics/Machine learning"}],"tags":[],"updatedAt":"2025-01-14T05:24:52+00:00","versionOfRecord":[],"versionCreatedAt":"2024-10-04 06:29:25","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-5018487","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-5018487","identity":"rs-5018487","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.