Risk-Sensitive Joint Inventory-Maintenance Strategy for Bearing Health Management under Prognostic Uncertainty: An Uncertainty-Aware Deep Reinforcement Learning Approach | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Risk-Sensitive Joint Inventory-Maintenance Strategy for Bearing Health Management under Prognostic Uncertainty: An Uncertainty-Aware Deep Reinforcement Learning Approach Zehua Zhang, Ning Shen, Xingfen Wang This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8614588/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 22 You are reading this latest preprint version Abstract Traditional separation of inventory management and Prognostics and Health Management (PHM) often leads to resource misallocation. While Deep Reinforcement Learning (DRL) offers a promising solution for joint decision-making, standard agents typically treat Prognostics and Health Management predictions as deterministic ground truths. However, in real-world scenarios, RUL predictions inherently contain stochastic errors. Ignoring this uncertainty leads to risk-blind policies that fail to buffer against sudden failures when prediction confidence is low. To address this, this paper proposes an Uncertainty-Aware collaborative adaptive inventory strategy. First, we introduce a Bayesian uncertainty quantification mechanism using Monte Carlo Dropout to estimate not only the RUL value but also its prediction variance. Second, to overcome the agent's myopic behavior, a novel Asymmetric Cost-Aware Reward Shaping mechanism is designed. By strategically decoupling the training and evaluation reward functions—specifically by introducing safety stock penalties and attenuating holding costs during training—the agent is guided to establish robust inventory buffers against supply chain uncertainties. Simulation results demonstrate that the proposed Risk-Sensitive PPO strategy significantly outperforms deterministic baselines, reducing total costs by 40.3% under high-noise environments. Physical sciences/Engineering Physical sciences/Mathematics and computing Bearing Health Management Inventory Management Deep Reinforcement Learning Reward Shaping Joint Decision-Making Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Editorial decision: Revision requested 23 Feb, 2026 Reviews received at journal 21 Feb, 2026 Reviews received at journal 16 Feb, 2026 Reviews received at journal 11 Feb, 2026 Reviewers agreed at journal 02 Feb, 2026 Reviews received at journal 01 Feb, 2026 Reviewers agreed at journal 30 Jan, 2026 Reviewers agreed at journal 30 Jan, 2026 Reviewers agreed at journal 30 Jan, 2026 Reviews received at journal 30 Jan, 2026 Reviewers agreed at journal 29 Jan, 2026 Reviewers agreed at journal 28 Jan, 2026 Reviewers agreed at journal 28 Jan, 2026 Reviewers agreed at journal 27 Jan, 2026 Reviewers agreed at journal 27 Jan, 2026 Reviewers agreed at journal 27 Jan, 2026 Reviewers agreed at journal 27 Jan, 2026 Reviewers invited by journal 27 Jan, 2026 Editor assigned by journal 27 Jan, 2026 Editor invited by journal 27 Jan, 2026 Submission checks completed at journal 26 Jan, 2026 First submitted to journal 26 Jan, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8614588","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":581885445,"identity":"eed0c3de-a297-41c4-9c4c-86c7c77edfc5","order_by":0,"name":"Zehua Zhang","email":"","orcid":"","institution":"Beijing Jiaotong University","correspondingAuthor":false,"prefix":"","firstName":"Zehua","middleName":"","lastName":"Zhang","suffix":""},{"id":581885446,"identity":"bdf18f51-fea6-4dab-b3e6-df3cf3f9b50b","order_by":1,"name":"Ning Shen","email":"","orcid":"","institution":"Beijing Information Science \u0026 Technology University","correspondingAuthor":false,"prefix":"","firstName":"Ning","middleName":"","lastName":"Shen","suffix":""},{"id":581885447,"identity":"ea1d85d5-dbc8-4833-b8ac-3e9bfc1f3d05","order_by":2,"name":"Xingfen Wang","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA3klEQVRIiWNgGAWjYDAC5gMg0iaBDcolQgtbAohMA2phJk3L4QQGorUYHOMx/Fzw63weH//5YxIMFdaJDexnD+DVItnGYyw9s+92MZtEMpsEw5n0xAaevAS8WvjlezdI8/bcTmyTYGaTYGw7nNggwWOA3ytsvJt/8/acS2zjPwzU8o8ILfxsvNukeX4cSGxjADqMsYEILZJt/N+seRuSQX4xtkg4lm7cxpODX4vBMbbk2zx/7PLk+w8+vPGhxlq2n/0Mfi1gwNgGZSSAfEdYPQj8IU7ZKBgFo2AUjFAAAC6LPYiDCXKyAAAAAElFTkSuQmCC","orcid":"","institution":"Beijing Jiaotong University","correspondingAuthor":true,"prefix":"","firstName":"Xingfen","middleName":"","lastName":"Wang","suffix":""}],"badges":[],"createdAt":"2026-01-16 02:38:26","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8614588/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8614588/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":101751195,"identity":"355b01d0-5ad9-4885-80c6-4e4d347c8acc","added_by":"auto","created_at":"2026-02-03 10:18:04","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":775362,"visible":true,"origin":"","legend":"","description":"","filename":"AnUncertaintyAwareDeepReinforcementLearningApproach.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8614588/v1_covered_352d49ed-c84f-4e8a-aab7-bb650c578036.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Risk-Sensitive Joint Inventory-Maintenance Strategy for Bearing Health Management under Prognostic Uncertainty: An Uncertainty-Aware Deep Reinforcement Learning Approach","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Bearing Health Management, Inventory Management, Deep Reinforcement Learning, Reward Shaping, Joint Decision-Making","lastPublishedDoi":"10.21203/rs.3.rs-8614588/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8614588/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eTraditional separation of inventory management and Prognostics and Health Management (PHM) often leads to resource misallocation. While Deep Reinforcement Learning (DRL) offers a promising solution for joint decision-making, standard agents typically treat Prognostics and Health Management predictions as deterministic ground truths. However, in real-world scenarios, RUL predictions inherently contain stochastic errors. Ignoring this uncertainty leads to risk-blind policies that fail to buffer against sudden failures when prediction confidence is low. To address this, this paper proposes an Uncertainty-Aware collaborative adaptive inventory strategy. First, we introduce a Bayesian uncertainty quantification mechanism using Monte Carlo Dropout to estimate not only the RUL value but also its prediction variance. Second, to overcome the agent's myopic behavior, a novel Asymmetric Cost-Aware Reward Shaping mechanism is designed. By strategically decoupling the training and evaluation reward functions\u0026mdash;specifically by introducing safety stock penalties and attenuating holding costs during training\u0026mdash;the agent is guided to establish robust inventory buffers against supply chain uncertainties. Simulation results demonstrate that the proposed Risk-Sensitive PPO strategy significantly outperforms deterministic baselines, reducing total costs by 40.3% under high-noise environments.\u003c/p\u003e","manuscriptTitle":"Risk-Sensitive Joint Inventory-Maintenance Strategy for Bearing Health Management under Prognostic Uncertainty: An Uncertainty-Aware Deep Reinforcement Learning Approach","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-01-29 11:49:26","doi":"10.21203/rs.3.rs-8614588/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2026-02-23T18:48:07+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-02-21T12:38:59+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-02-16T05:38:15+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-02-11T15:29:41+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"265857767728056790325649609147258810935","date":"2026-02-02T20:09:52+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-02-01T06:07:53+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"102877620560216110329626141611317209778","date":"2026-01-30T14:35:41+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"248240636300541164571888186919917634623","date":"2026-01-30T09:06:36+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"105102509343415226149428276114836009564","date":"2026-01-30T08:30:58+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-01-30T07:06:28+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"332761731393374894592467540409838026014","date":"2026-01-29T18:14:33+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"207462768664186077546176463688759073070","date":"2026-01-28T09:04:47+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"120638818089709728441583710630060702015","date":"2026-01-28T06:51:19+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"94965402054780981338350830122247667414","date":"2026-01-28T01:54:03+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"22108375897959440398952137617267684047","date":"2026-01-28T01:12:50+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"330644980183910226956622131205105542439","date":"2026-01-28T00:56:51+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"166791469484429386174422709284858308552","date":"2026-01-27T22:07:11+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-01-27T17:57:38+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-01-27T17:55:39+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2026-01-27T17:51:01+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-01-26T05:35:20+00:00","index":"","fulltext":""},{"type":"submitted","content":"Scientific Reports","date":"2026-01-26T05:28:57+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"0c654613-f14c-43d0-bcac-d14435748bf6","owner":[],"postedDate":"January 29th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[{"id":61907696,"name":"Physical sciences/Engineering"},{"id":61907697,"name":"Physical sciences/Mathematics and computing"}],"tags":[],"updatedAt":"2026-04-27T18:08:14+00:00","versionOfRecord":[],"versionCreatedAt":"2026-01-29 11:49:26","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8614588","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8614588","identity":"rs-8614588","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.