Optimized Renal Replacement Therapy Decisions in Intensive Care: A Reinforcement Learning Approach | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Optimized Renal Replacement Therapy Decisions in Intensive Care: A Reinforcement Learning Approach Lorenz Kapral, Razvan Bologheanu, Mohammad Mahdi Azarbeik, Aylin Bilir, and 6 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6243566/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Purpose Acute kidney injury (AKI) is highly prevalent in intensive care units (ICUs) and often requires renal replacement therapy (RRT). However, the optimal timing for initiating RRT remains controversial. The aim of this study was to develop a reinforcement learning (RL) model to support individualized RRT decision-making for critically ill AKI patients. Methods We trained and validated our RL model using ICU data from two cohorts: the publicly available MIMIC-IV database and a dataset from the Medical University of Vienna (MUW). Patients with AKI of stage I or higher were included, and those with chronic kidney disease or prior kidney transplantation were excluded. We extracted 88 features, employing weighted K-means clustering for state definition. A Q-learning–based RL approach was applied, with off-policy evaluation to assess the policy’s performance versus clinician decisions. Results In both the MIMIC and MUW cohorts, the RL model demonstrated a high level of concordance (up to 98.5%) with clinicians but exhibited superior performance on key metrics. Notably, the model identified a reproducible patient subgroup with greater illness severity for whom earlier or more frequent RRT could improve outcomes, suggesting a beneficial role for AI-driven decision support. Conclusions Our RL model provides dynamic, data-driven recommendations for initiating and ceasing RRT, closely aligning with clinical practice and identifying high-risk patients who may benefit from earlier intervention. Acute kidney injury (AKI) Renal replacement therapy (RRT) Reinforcement learning (RL) Intensive care Clinical decision support (CDS) Figures Figure 1 Figure 2 Figure 3 Figure 4 Full Text Supplementary Files RRTfiguressupplement.docx TRIPODAI.pdf Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6243566","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":433025122,"identity":"184a4f42-0f34-4def-9492-caa5c12c3370","order_by":0,"name":"Lorenz Kapral","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA/UlEQVRIie3NsYrCQBCA4ZGFTTOYdkC4vMKGhWCRwkfZbVKKkOYKi1Sx8uoDRZ/BxloQYuMDBM7CcC+wcCAHF47ToIVFEssD92fZhWE/BsBm+4d1r6+o7uNrX7USfkfUntTdoJ3o9BHiTIvPUQnSnU0Loxc09AasMDAu6wnupHxHCOiwk6TXFPsbLgmyhi0U8R4ShJBHcCF6lUBwnjcS5wcFhF4esW89vxDnBPDbvIWhgkDkESedkF4CBtBJGwhmrIcbkv4h432VUSwYxqTfZC1xnbTzhWXoLz5SlptxOPQmk5Uxp5daco2qUyW250u1gZuq8pKHvttsNtsT9QeCiEVF7/PfSQAAAABJRU5ErkJggg==","orcid":"https://orcid.org/0000-0003-3316-9024","institution":"Medical University of Vienna: Medizinische Universitat Wien","correspondingAuthor":true,"prefix":"","firstName":"Lorenz","middleName":"","lastName":"Kapral","suffix":""},{"id":433025123,"identity":"ff3622e7-5fce-49b1-94cb-9b4e9a13db89","order_by":1,"name":"Razvan Bologheanu","email":"","orcid":"","institution":"Medical University of Vienna: Medizinische Universitat Wien","correspondingAuthor":false,"prefix":"","firstName":"Razvan","middleName":"","lastName":"Bologheanu","suffix":""},{"id":433025124,"identity":"87b951d4-2df0-44ae-b768-87ce05e7be7f","order_by":2,"name":"Mohammad Mahdi Azarbeik","email":"","orcid":"","institution":"Technical University of Vienna: Technische Universitat Wien","correspondingAuthor":false,"prefix":"","firstName":"Mohammad","middleName":"Mahdi","lastName":"Azarbeik","suffix":""},{"id":433025125,"identity":"0aecc0fc-3294-4ac0-ba9b-6bbc0af39e42","order_by":3,"name":"Aylin Bilir","email":"","orcid":"","institution":"Medical University of Vienna: Medizinische Universitat Wien","correspondingAuthor":false,"prefix":"","firstName":"Aylin","middleName":"","lastName":"Bilir","suffix":""},{"id":433025126,"identity":"83bdb1cf-92ac-470b-8aa6-9700b7aa450b","order_by":4,"name":"Richard Weiss","email":"","orcid":"","institution":"Technical University of Vienna: Technische Universitat Wien","correspondingAuthor":false,"prefix":"","firstName":"Richard","middleName":"","lastName":"Weiss","suffix":""},{"id":433025127,"identity":"e75cdaad-090a-4226-b44b-000236314773","order_by":5,"name":"Stefan Bartos","email":"","orcid":"","institution":"Medical University of Vienna: Medizinische Universitat Wien","correspondingAuthor":false,"prefix":"","firstName":"Stefan","middleName":"","lastName":"Bartos","suffix":""},{"id":433025128,"identity":"4ab9b3ff-e00a-416e-be22-01916ad8fdd4","order_by":6,"name":"Stefan Schaller","email":"","orcid":"","institution":"Medical University of Vienna: Medizinische Universitat Wien","correspondingAuthor":false,"prefix":"","firstName":"Stefan","middleName":"","lastName":"Schaller","suffix":""},{"id":433025129,"identity":"91cdb285-6eca-446a-992b-21e9e466af22","order_by":7,"name":"Clemens Heitzinger","email":"","orcid":"","institution":"Technical University of Vienna: Technische Universitat Wien","correspondingAuthor":false,"prefix":"","firstName":"Clemens","middleName":"","lastName":"Heitzinger","suffix":""},{"id":433025130,"identity":"bd965773-d6e0-4a09-bf00-7802bd63cb33","order_by":8,"name":"Eva Schaden","email":"","orcid":"","institution":"Medical University of Vienna: Medizinische Universitat Wien","correspondingAuthor":false,"prefix":"","firstName":"Eva","middleName":"","lastName":"Schaden","suffix":""},{"id":433025131,"identity":"a1939263-07d7-4548-95bd-90a465f814ce","order_by":9,"name":"Oliver Kimberger","email":"","orcid":"","institution":"Medical University of Vienna: Medizinische Universitat Wien","correspondingAuthor":false,"prefix":"","firstName":"Oliver","middleName":"","lastName":"Kimberger","suffix":""}],"badges":[],"createdAt":"2025-03-17 10:21:06","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6243566/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6243566/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":79391579,"identity":"390a4135-dade-4057-af1f-4d07014c37b8","added_by":"auto","created_at":"2025-03-27 20:35:23","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":206614,"visible":true,"origin":"","legend":"\u003cp\u003eReinforcement learning algorithm including evaluation and clustering optimization. Vital parameters, drug administration details, and renal replacement therapy (RRT) indications were extracted from the MIMIC dataset, which was split into training and test sets. A random forest model was trained on the training set to rank feature importance. Kullback-Leiber divergence and matrix norms were then used to compare train and test state-transition probabilities and determine the optimal number of features and clusters. Based on these results, a weighted k-means algorithm was applied to cluster all subsequent data. A reinforcement learning (RL) model was trained to optimize RRT timing based on a reward linked to 90-day mortality, and it was evaluated using off-policy evaluation methods (WIS and DICE). Finally, the RL model was validated using both an internal (MIMIC test) and an external (MUW) dataset.\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-6243566/v1/bd96e7aa0007cbdcdf7ba104.png"},{"id":79391578,"identity":"6377e017-017e-425b-9ed0-a0edd6f137b2","added_by":"auto","created_at":"2025-03-27 20:35:23","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":96311,"visible":true,"origin":"","legend":"\u003cp\u003eThis figure illustrates the model’s performance (dark blue) at varying levels of RRT initiation penalty within the validation set. The clinicians’ performance is shown in light blue. The black curve represents the proportion of patients receiving AI-RRT recommendations as the penalty changes, and the black dashed lines show the average RRT treatment rates in the MIMIC and MUW test sets. The red dashed line represents the selected model, which we selected for 2 reasons: Firstly, the treatment rate yielded is analogous to that of the MIMIC dataset, thereby reflecting real-world constraints. Secondly, among the models exhibiting a treatment rate comparable to that of MIMIC, the 22% penalty model demonstrates the highest average WIS.\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-6243566/v1/99187842da6e3928de1e60fa.png"},{"id":79391876,"identity":"fab8f43e-a5f6-42da-9b40-52a13507feb9","added_by":"auto","created_at":"2025-03-27 20:43:23","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":65314,"visible":true,"origin":"","legend":"\u003cp\u003eThis figure demonstrates the proportion of AI-recommended renal replacement therapy (RRT) in the test set compared to actual clinical practice, displayed across all time steps and for all patients, stratified by ICU groups and SOFA score categories. The term \"time steps\" is used to denote the proportion of 12-hour steps in which the model or clinician recommended RRT in the data set. While the overall treatment strategies appear similar, the AI tends to recommend RRT less frequently and demonstrates a more pronounced response to increasing SOFA scores.\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-6243566/v1/afe852556dfe4800ff3a630f.png"},{"id":79391581,"identity":"6fbd91d8-75cd-43b7-8258-23b28f226366","added_by":"auto","created_at":"2025-03-27 20:35:23","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":68145,"visible":true,"origin":"","legend":"\u003cp\u003eThis figure presents the survival probability of four patient groups categorized by treatment type: AI-recommended RRT, Clinician-initiated RRT, Both AI- and Clinician-initiated RRT, and Neither RRT. Additionally, the 90-day mortality rate is shown on the right axis. The legend indicates the number of patients in each group. Notably, the group receiving only AI-recommended RRT exhibits higher mortality, suggesting a potential benefit from optimized treatment strategies.\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-6243566/v1/ed2c75eb7ed3b8e0507bf393.png"},{"id":80121235,"identity":"9ecd225c-9e06-4bca-8bbd-1463a3095a95","added_by":"auto","created_at":"2025-04-08 07:32:18","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":926488,"visible":true,"origin":"","legend":"","description":"","filename":"RRTmanuscript4.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6243566/v1_covered_b52717a6-f07d-4dbe-b9c7-21f1d6b6fd87.pdf"},{"id":79391584,"identity":"f4a06ca1-2a0d-4986-9009-127fcefa9e83","added_by":"auto","created_at":"2025-03-27 20:35:23","extension":"docx","order_by":4,"title":"","display":"","copyAsset":false,"role":"supplement","size":403266,"visible":true,"origin":"","legend":"","description":"","filename":"RRTfiguressupplement.docx","url":"https://assets-eu.researchsquare.com/files/rs-6243566/v1/5641bb4aefabc641da4fdf87.docx"},{"id":79391589,"identity":"4e336873-af4e-42b0-a211-d9f1d3ab6338","added_by":"auto","created_at":"2025-03-27 20:35:23","extension":"pdf","order_by":5,"title":"","display":"","copyAsset":false,"role":"supplement","size":1925125,"visible":true,"origin":"","legend":"","description":"","filename":"TRIPODAI.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6243566/v1/fe5643b31fe360f426e2e3fb.pdf"}],"financialInterests":"","formattedTitle":"Optimized Renal Replacement Therapy Decisions in Intensive Care: A Reinforcement Learning Approach","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Acute kidney injury (AKI), Renal replacement therapy (RRT), Reinforcement learning (RL), Intensive care, Clinical decision support (CDS)","lastPublishedDoi":"10.21203/rs.3.rs-6243566/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6243566/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003ePurpose\u003c/h2\u003e \u003cp\u003eAcute kidney injury (AKI) is highly prevalent in intensive care units (ICUs) and often requires renal replacement therapy (RRT). However, the optimal timing for initiating RRT remains controversial. The aim of this study was to develop a reinforcement learning (RL) model to support individualized RRT decision-making for critically ill AKI patients.\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e \u003cp\u003eWe trained and validated our RL model using ICU data from two cohorts: the publicly available MIMIC-IV database and a dataset from the Medical University of Vienna (MUW). Patients with AKI of stage I or higher were included, and those with chronic kidney disease or prior kidney transplantation were excluded. We extracted 88 features, employing weighted K-means clustering for state definition. A Q-learning\u0026ndash;based RL approach was applied, with off-policy evaluation to assess the policy\u0026rsquo;s performance versus clinician decisions.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eIn both the MIMIC and MUW cohorts, the RL model demonstrated a high level of concordance (up to 98.5%) with clinicians but exhibited superior performance on key metrics. Notably, the model identified a reproducible patient subgroup with greater illness severity for whom earlier or more frequent RRT could improve outcomes, suggesting a beneficial role for AI-driven decision support.\u003c/p\u003e\u003ch2\u003eConclusions\u003c/h2\u003e \u003cp\u003eOur RL model provides dynamic, data-driven recommendations for initiating and ceasing RRT, closely aligning with clinical practice and identifying high-risk patients who may benefit from earlier intervention.\u003c/p\u003e","manuscriptTitle":"Optimized Renal Replacement Therapy Decisions in Intensive Care: A Reinforcement Learning Approach","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-03-27 20:35:18","doi":"10.21203/rs.3.rs-6243566/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"761b97db-b8d6-4218-b3cf-c71246f5f1b2","owner":[],"postedDate":"March 27th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-04-08T07:24:10+00:00","versionOfRecord":[],"versionCreatedAt":"2025-03-27 20:35:18","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6243566","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6243566","identity":"rs-6243566","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.