A Comprehensive Active Learning Model for Predicting Reliable Charges of Health Insurance | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article A Comprehensive Active Learning Model for Predicting Reliable Charges of Health Insurance Moyaj Ibne Akbar, Md. Shahriare Satu, Md. Ziaul Haque, Sheikh M. Shariful Islam, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8938788/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 5 You are reading this latest preprint version Abstract The rapid growth of data analytics has rendered a critical imperative to accurately predict health insurance premiums, enabling the design of personalized, data-driven customer strategies. Despite the progress of machine learning in streamlining insurance, existing methodologies are unable to effectively handle large volumes of unlabeled data alongside limited training data. In this regard, this study proposes an active learning-based framework that strategically leverages scarce labeled data to optimize the prediction of health insurance premiums. Three publicly available datasets were employed, upon which comprehensive data transformation and selection techniques were applied to construct diverse and informative feature subsets. Then, the active learning was applied to both the primary and transformed datasets, in which multiple regression models were iteratively trained and evaluated to identify the most robust premium predictors. Notably, Gradient Boosting (GB) and Random Forest (RF) consistently deliver the most accurate predictions of health insurance premiums throughout all experimental configurations. Furthermore, explainable AI techniques were integrated into the analytical pipeline to identify influential factors contributing to high prediction accuracy. In this work, smoking status, age, body mass index, and the presence of chronic conditions are the most influential determinants to predict health-care premiums. These findings carry significant practical implications for equipping actionable health insights and designing evidence-based premium structures. This experimental evaluation demonstrates that the proposed framework consistently outperforms state-of-the-art models, achieving substantially lower error rates and affirming its generalizability across diverse datasets. Biological sciences/Computational biology and bioinformatics Health sciences/Health care Physical sciences/Mathematics and computing Active learning prediction uncertainty sampling health insurance interpretability Full Text Additional Declarations No competing interests reported. Supplementary Files SupplimentaryTable1.csv SupplimentaryTable2.csv SupplimentaryTable3.csv Cite Share Download PDF Status: Under Review Version 1 posted Reviewers invited by journal 29 Apr, 2026 Editor assigned by journal 28 Mar, 2026 Editor invited by journal 27 Feb, 2026 Submission checks completed at journal 25 Feb, 2026 First submitted to journal 25 Feb, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8938788","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":632010005,"identity":"31ba8782-cc8e-4ecb-964a-a45c334978f2","order_by":0,"name":"Moyaj Ibne Akbar","email":"","orcid":"","institution":"Noakhali Science and Technology University","correspondingAuthor":false,"prefix":"","firstName":"Moyaj","middleName":"Ibne","lastName":"Akbar","suffix":""},{"id":632010007,"identity":"3bc9d7f6-873b-4037-9c8b-d41078a5ffbb","order_by":1,"name":"Md. Shahriare Satu","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABF0lEQVRIiWNgGAWjYDACCQY2BgYDJC4/lMXYwMAD5OLVwgzhSjaAVBPUwgDXAtR+gIAW+dnNzx5XFNjly7f3H5NgzLGwN77d/vwxD4ON7IYDvAdvYNFicOeYueEZg2TLDWcOs0kwbpNI3HbnjGEzD0Oa8YYDfMkW2LRIJJhJNhgwGxhIJIO1JJjdyGEEajmcuOEAjxlWh81I/wbUUm8gP/8xWIu98Yz0h0At/3FqYbiRA7LlsAHDDWawFsYNEgkghx3AqcXgRk4ZUMtxA4MzycYWiUC/zLiRYzhzjkGy8czD2P0CdNg2yYY/1Qby7Qcf3vi4rc6ef0b6gw9vKuxk+473Yg0xVJCAsJ0BHlGjYBSMglEwCkgHANUJXyojF0FMAAAAAElFTkSuQmCC","orcid":"","institution":"Noakhali Science and Technology University","correspondingAuthor":true,"prefix":"","firstName":"Md.","middleName":"Shahriare","lastName":"Satu","suffix":""},{"id":632010009,"identity":"fbe78310-a670-4dd5-9cad-dcbf84a65ec0","order_by":2,"name":"Md. Ziaul Haque","email":"","orcid":"","institution":"Noakhali Science and Technology University","correspondingAuthor":false,"prefix":"","firstName":"Md.","middleName":"Ziaul","lastName":"Haque","suffix":""},{"id":632010013,"identity":"279a7241-0660-4132-bd0e-07c8e01e2869","order_by":3,"name":"Sheikh M. Shariful Islam","email":"","orcid":"","institution":"Texas Tech University","correspondingAuthor":false,"prefix":"","firstName":"Sheikh","middleName":"M. Shariful","lastName":"Islam","suffix":""},{"id":632010015,"identity":"e487119f-b9a4-4e3f-96e7-1415360a0dc1","order_by":4,"name":"Mohammad Ali Moni","email":"","orcid":"","institution":"Charles Sturt University","correspondingAuthor":false,"prefix":"","firstName":"Mohammad","middleName":"Ali","lastName":"Moni","suffix":""}],"badges":[],"createdAt":"2026-02-22 11:38:19","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8938788/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8938788/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":108809749,"identity":"f3e1df5b-df4a-44b2-9891-c3e009d85757","added_by":"auto","created_at":"2026-05-08 15:55:17","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":732048,"visible":true,"origin":"","legend":"","description":"","filename":"ManuscriptSR.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8938788/v1_covered_fbeee781-6928-486d-946c-5a6e6ddad40d.pdf"},{"id":108806504,"identity":"2a3902ab-b38c-4278-9763-00a78ca69a3a","added_by":"auto","created_at":"2026-05-08 15:28:46","extension":"csv","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":1438,"visible":true,"origin":"","legend":"","description":"","filename":"SupplimentaryTable1.csv","url":"https://assets-eu.researchsquare.com/files/rs-8938788/v1/44a60e9a9657f606de9739fb.csv"},{"id":108732870,"identity":"eb8881da-1042-466c-9a9f-3e2cc5816f8b","added_by":"auto","created_at":"2026-05-07 19:20:39","extension":"csv","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":1422,"visible":true,"origin":"","legend":"","description":"","filename":"SupplimentaryTable2.csv","url":"https://assets-eu.researchsquare.com/files/rs-8938788/v1/2ad6f3ad6921843b612fb68e.csv"},{"id":108805931,"identity":"86778418-689d-4102-a81a-f2d009e8c05f","added_by":"auto","created_at":"2026-05-08 15:27:13","extension":"csv","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":1442,"visible":true,"origin":"","legend":"","description":"","filename":"SupplimentaryTable3.csv","url":"https://assets-eu.researchsquare.com/files/rs-8938788/v1/a98ae25fe9705c703b83249f.csv"}],"financialInterests":"No competing interests reported.","formattedTitle":"A Comprehensive Active Learning Model for Predicting Reliable Charges of Health Insurance","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Active learning, prediction, uncertainty sampling, health insurance, interpretability","lastPublishedDoi":"10.21203/rs.3.rs-8938788/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8938788/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"The rapid growth of data analytics has rendered a critical imperative to accurately predict health insurance premiums, enabling the design of personalized, data-driven customer strategies. Despite the progress of machine learning in streamlining insurance, existing methodologies are unable to effectively handle large volumes of unlabeled data alongside limited training data. In this regard, this study proposes an active learning-based framework that strategically leverages scarce labeled data to optimize the prediction of health insurance premiums. Three publicly available datasets were employed, upon which comprehensive data transformation and selection techniques were applied to construct diverse and informative feature subsets. Then, the active learning was applied to both the primary and transformed datasets, in which multiple regression models were iteratively trained and evaluated to identify the most robust premium predictors. Notably, Gradient Boosting (GB) and Random Forest (RF) consistently deliver the most accurate predictions of health insurance premiums throughout all experimental configurations. Furthermore, explainable AI techniques were integrated into the analytical pipeline to identify influential factors contributing to high prediction accuracy. In this work, smoking status, age, body mass index, and the presence of chronic conditions are the most influential determinants to predict health-care premiums. These findings carry significant practical implications for equipping actionable health insights and designing evidence-based premium structures. This experimental evaluation demonstrates that the proposed framework consistently outperforms state-of-the-art models, achieving substantially lower error rates and affirming its generalizability across diverse datasets.","manuscriptTitle":"A Comprehensive Active Learning Model for Predicting Reliable Charges of Health Insurance","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-05-07 19:20:35","doi":"10.21203/rs.3.rs-8938788/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"reviewersInvited","content":"","date":"2026-04-29T11:26:31+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-03-28T14:58:21+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2026-02-27T09:36:27+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-02-25T20:33:17+00:00","index":"","fulltext":""},{"type":"submitted","content":"Scientific Reports","date":"2026-02-25T20:28:12+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"1f9f8a9e-a55b-42dc-930a-31bebca1acb9","owner":[],"postedDate":"May 7th, 2026","published":true,"recentEditorialEvents":[{"type":"reviewersInvited","content":"10","date":"2026-04-29T11:26:31+00:00","index":"","fulltext":""}],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[{"id":67271853,"name":"Biological sciences/Computational biology and bioinformatics"},{"id":67271854,"name":"Health sciences/Health care"},{"id":67271855,"name":"Physical sciences/Mathematics and computing"}],"tags":[],"updatedAt":"2026-05-07T19:20:35+00:00","versionOfRecord":[],"versionCreatedAt":"2026-05-07 19:20:35","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8938788","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8938788","identity":"rs-8938788","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.