{"paper_id":"023ff4d6-8661-404b-90f5-2a704a2acc38","body_text":"Research on the Optimization of Hospital Contract Energy Management Strategy Based on Reinforcement Learning | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Research on the Optimization of Hospital Contract Energy Management Strategy Based on Reinforcement Learning Qinhao Chen, Jiahong Lu, Chuangyin Dang This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6795020/v1 This work is licensed under a CC BY 4.0 License Status: Under Revision Version 1 posted 12 You are reading this latest preprint version Abstract This paper proposes an optimization method for hospital contract energy management strategies based on reinforcement learning, aimed at improving the intelligence of strategies and resource allocation efficiency in long-term energy-saving cooperation for healthcare buildings. Addressing the high volatility of hospital energy consumption behavior, the information incompleteness during contract execution, and the uncertainty of service responses, a master-slave game framework is constructed to model the interaction between hospitals and Energy Service Companies (ESCOs). The improved Deep Deterministic Policy Gradient (Hospital-DDPG) algorithm is introduced for strategy optimization. The proposed method adopts an Actor-Critic structure to implement strategy generation and value evaluation, with a Recurrent Neural Network (RNN) embedded in the critic network to capture the temporal dependencies of energy-saving behavior. Additionally, the Burn-In initialization mechanism is incorporated to enhance the model's ability to fit historical states. Moreover, a Value Decomposition Network (VDN) is introduced to achieve collaborative learning among multiple agents, optimizing global energy-saving performance while ensuring data privacy and decision autonomy. Simulation experiments are conducted based on real or constructed hospital energy consumption data. The results show that, compared to traditional DDPG and MADDPG methods, Hospital-DDPG exhibits significant advantages in strategy convergence speed, energy-saving contract benefits, and market adaptability. It effectively reduces reliance on the main platform, minimizes penalty for non-performance, and enhances local energy-saving response efficiency. This research provides a feasible strategy optimization approach and engineering practice reference for smart hospital contract energy management systems. Physical sciences/Engineering/Mechanical engineering Physical sciences/Mathematics and computing Physical sciences/Mathematics and computing/Computer science Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Revision Version 1 posted Editorial decision: Revision requested 04 Sep, 2025 Reviews received at journal 02 Sep, 2025 Reviews received at journal 28 Aug, 2025 Reviews received at journal 25 Aug, 2025 Reviewers agreed at journal 21 Aug, 2025 Reviewers agreed at journal 19 Aug, 2025 Reviewers agreed at journal 19 Aug, 2025 Reviewers invited by journal 19 Aug, 2025 Editor invited by journal 19 Aug, 2025 Submission checks completed at journal 14 Aug, 2025 Editor assigned by journal 14 Aug, 2025 First submitted to journal 08 Jun, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {\"props\":{\"pageProps\":{\"initialData\":{\"identity\":\"rs-6795020\",\"acceptedTermsAndConditions\":true,\"allowDirectSubmit\":false,\"archivedVersions\":[],\"articleType\":\"Article\",\"associatedPublications\":[],\"authors\":[{\"id\":505481048,\"identity\":\"a47fe74b-393e-4928-84fe-9f6e9b1a8402\",\"order_by\":0,\"name\":\"Qinhao Chen\",\"email\":\"\",\"orcid\":\"\",\"institution\":\"City University of Hong Kong\",\"correspondingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Qinhao\",\"middleName\":\"\",\"lastName\":\"Chen\",\"suffix\":\"\"},{\"id\":505481049,\"identity\":\"b9554247-0481-406f-8fe5-43c92d0e6cb0\",\"order_by\":1,\"name\":\"Jiahong Lu\",\"email\":\"\",\"orcid\":\"\",\"institution\":\"Southwestern University of Finance and Economics\",\"correspondingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Jiahong\",\"middleName\":\"\",\"lastName\":\"Lu\",\"suffix\":\"\"},{\"id\":505481051,\"identity\":\"bcc23a54-4177-497f-aa74-4682dbee7dd4\",\"order_by\":2,\"name\":\"Chuangyin Dang\",\"email\":\"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA3klEQVRIie3RPQrCMBTA8SeF1qHYNQHxDBGhXqdZnBScSocqSiUd1L3HqIsfW6XQLrFzx+YIbo7GxbHNKJj/8CDwfkMSAJ3uF8vA7G1MsB2jt228IFQnIxxHEWl4oU4m5FAyLHZGtxiU3BWJv6IpoiygEjvx3mslmM+n47Qq6UWSml6HgPgjbSWktl0sWEFvyYdwEwhaKJK0pmxJmaFITiycEH5noEQwn/k4qTL5yNsIebywO+8yKPMzPvhr+ZWWeL6CcOTEx3YC0Cdy5N+j3bH+yWrkWCss6nQ63d/2BpleU4g8vqltAAAAAElFTkSuQmCC\",\"orcid\":\"\",\"institution\":\"City University of Hong Kong\",\"correspondingAuthor\":true,\"prefix\":\"\",\"firstName\":\"Chuangyin\",\"middleName\":\"\",\"lastName\":\"Dang\",\"suffix\":\"\"}],\"badges\":[],\"createdAt\":\"2025-06-01 09:53:19\",\"currentVersionCode\":1,\"declarations\":\"\",\"doi\":\"10.21203/rs.3.rs-6795020/v1\",\"doiUrl\":\"https://doi.org/10.21203/rs.3.rs-6795020/v1\",\"draftVersion\":[],\"editorialEvents\":[],\"editorialNote\":\"\",\"failedWorkflow\":false,\"files\":[{\"id\":89985329,\"identity\":\"df258377-744d-46f6-9395-3eccd1c1c4f6\",\"added_by\":\"auto\",\"created_at\":\"2025-08-27 06:43:25\",\"extension\":\"pdf\",\"order_by\":1,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"manuscript-pdf\",\"size\":1179458,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"ResearchontheOptimizationofHospitalContractEnergyManagementStrategyBasedonReinforcementLearning1.pdf\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-6795020/v1_covered_afcc8ab2-4390-4ea7-b66d-fa243640c6dc.pdf\"}],\"financialInterests\":\"No competing interests reported.\",\"formattedTitle\":\"Research on the Optimization of Hospital Contract Energy Management Strategy Based on Reinforcement Learning\",\"fulltext\":[],\"fulltextSource\":\"\",\"fullText\":\"\",\"funders\":[],\"hasAdminPriorityOnWorkflow\":false,\"hasManuscriptDocX\":false,\"hasOptedInToPreprint\":true,\"hasPassedJournalQc\":\"\",\"hasAnyPriority\":false,\"hideJournal\":false,\"highlight\":\"\",\"institution\":\"\",\"isAcceptedByJournal\":false,\"isAuthorSuppliedPdf\":true,\"isDeskRejected\":\"\",\"isHiddenFromSearch\":false,\"isInQc\":false,\"isInWorkflow\":false,\"isPdf\":true,\"isPdfUpToDate\":true,\"isWithdrawnOrRetracted\":false,\"journal\":{\"display\":true,\"email\":\"info@researchsquare.com\",\"identity\":\"scientific-reports\",\"isNatureJournal\":false,\"hasQc\":true,\"allowDirectSubmit\":false,\"externalIdentity\":\"scirep\",\"sideBox\":\"Learn more about [Scientific Reports](http://www.nature.com/srep/)\",\"snPcode\":\"\",\"submissionUrl\":\"\",\"title\":\"Scientific Reports\",\"twitterHandle\":\"\",\"acdcEnabled\":true,\"dfaEnabled\":true,\"editorialSystem\":\"stoa\",\"reportingPortfolio\":\"Scientific Reports\",\"inReviewEnabled\":true,\"inReviewRevisionsEnabled\":true},\"keywords\":\"\",\"lastPublishedDoi\":\"10.21203/rs.3.rs-6795020/v1\",\"lastPublishedDoiUrl\":\"https://doi.org/10.21203/rs.3.rs-6795020/v1\",\"license\":{\"name\":\"CC BY 4.0\",\"url\":\"https://creativecommons.org/licenses/by/4.0/\"},\"manuscriptAbstract\":\"This paper proposes an optimization method for hospital contract energy management strategies based on reinforcement learning, aimed at improving the intelligence of strategies and resource allocation efficiency in long-term energy-saving cooperation for healthcare buildings. Addressing the high volatility of hospital energy consumption behavior, the information incompleteness during contract execution, and the uncertainty of service responses, a master-slave game framework is constructed to model the interaction between hospitals and Energy Service Companies (ESCOs). The improved Deep Deterministic Policy Gradient (Hospital-DDPG) algorithm is introduced for strategy optimization. The proposed method adopts an Actor-Critic structure to implement strategy generation and value evaluation, with a Recurrent Neural Network (RNN) embedded in the critic network to capture the temporal dependencies of energy-saving behavior. Additionally, the Burn-In initialization mechanism is incorporated to enhance the model's ability to fit historical states. Moreover, a Value Decomposition Network (VDN) is introduced to achieve collaborative learning among multiple agents, optimizing global energy-saving performance while ensuring data privacy and decision autonomy. Simulation experiments are conducted based on real or constructed hospital energy consumption data. The results show that, compared to traditional DDPG and MADDPG methods, Hospital-DDPG exhibits significant advantages in strategy convergence speed, energy-saving contract benefits, and market adaptability. It effectively reduces reliance on the main platform, minimizes penalty for non-performance, and enhances local energy-saving response efficiency. This research provides a feasible strategy optimization approach and engineering practice reference for smart hospital contract energy management systems.\",\"manuscriptTitle\":\"Research on the Optimization of Hospital Contract Energy Management Strategy Based on Reinforcement Learning\",\"msid\":\"\",\"msnumber\":\"\",\"nonDraftVersions\":[{\"code\":1,\"date\":\"2025-08-27 06:27:16\",\"doi\":\"10.21203/rs.3.rs-6795020/v1\",\"editorialEvents\":[{\"type\":\"communityComments\",\"content\":0},{\"type\":\"decision\",\"content\":\"Revision requested\",\"date\":\"2025-09-04T20:13:37+00:00\",\"index\":\"\",\"fulltext\":\"\"},{\"type\":\"editorInvitedReview\",\"content\":\"\",\"date\":\"2025-09-02T05:44:26+00:00\",\"index\":\"hide\",\"fulltext\":\"\"},{\"type\":\"editorInvitedReview\",\"content\":\"\",\"date\":\"2025-08-28T15:00:37+00:00\",\"index\":\"hide\",\"fulltext\":\"\"},{\"type\":\"editorInvitedReview\",\"content\":\"\",\"date\":\"2025-08-25T22:56:42+00:00\",\"index\":\"hide\",\"fulltext\":\"\"},{\"type\":\"reviewerAgreed\",\"content\":\"150377892381594018943729130458037792019\",\"date\":\"2025-08-21T05:22:26+00:00\",\"index\":\"hide\",\"fulltext\":\"\"},{\"type\":\"reviewerAgreed\",\"content\":\"199912669754168339797260973564101938397\",\"date\":\"2025-08-19T12:20:33+00:00\",\"index\":\"hide\",\"fulltext\":\"\"},{\"type\":\"reviewerAgreed\",\"content\":\"156227637687609201354107789266469593207\",\"date\":\"2025-08-19T06:19:52+00:00\",\"index\":\"hide\",\"fulltext\":\"\"},{\"type\":\"reviewersInvited\",\"content\":\"\",\"date\":\"2025-08-19T04:45:07+00:00\",\"index\":\"\",\"fulltext\":\"\"},{\"type\":\"editorInvited\",\"content\":\"\",\"date\":\"2025-08-19T04:29:53+00:00\",\"index\":\"\",\"fulltext\":\"\"},{\"type\":\"checksComplete\",\"content\":\"\",\"date\":\"2025-08-14T11:38:36+00:00\",\"index\":\"\",\"fulltext\":\"\"},{\"type\":\"editorAssigned\",\"content\":\"\",\"date\":\"2025-08-14T10:45:06+00:00\",\"index\":\"\",\"fulltext\":\"\"},{\"type\":\"submitted\",\"content\":\"Scientific Reports\",\"date\":\"2025-06-08T13:47:38+00:00\",\"index\":\"\",\"fulltext\":\"\"}],\"status\":\"published\",\"journal\":{\"display\":true,\"email\":\"info@researchsquare.com\",\"identity\":\"scientific-reports\",\"isNatureJournal\":false,\"hasQc\":true,\"allowDirectSubmit\":false,\"externalIdentity\":\"scirep\",\"sideBox\":\"Learn more about [Scientific Reports](http://www.nature.com/srep/)\",\"snPcode\":\"\",\"submissionUrl\":\"\",\"title\":\"Scientific Reports\",\"twitterHandle\":\"\",\"acdcEnabled\":true,\"dfaEnabled\":true,\"editorialSystem\":\"stoa\",\"reportingPortfolio\":\"Scientific Reports\",\"inReviewEnabled\":true,\"inReviewRevisionsEnabled\":true}}],\"origin\":\"\",\"ownerIdentity\":\"d86ad392-44eb-453d-8ab6-2677dff8f33a\",\"owner\":[],\"postedDate\":\"August 27th, 2025\",\"published\":true,\"recentEditorialEvents\":[],\"rejectedJournal\":[],\"revision\":\"\",\"amendment\":\"\",\"status\":\"in-revision\",\"subjectAreas\":[{\"id\":53688850,\"name\":\"Physical sciences/Engineering/Mechanical engineering\"},{\"id\":53688851,\"name\":\"Physical sciences/Mathematics and computing\"},{\"id\":53688852,\"name\":\"Physical sciences/Mathematics and computing/Computer science\"}],\"tags\":[],\"updatedAt\":\"2026-01-19T14:56:30+00:00\",\"versionOfRecord\":[],\"versionCreatedAt\":\"2025-08-27 06:27:16\",\"video\":\"\",\"vorDoi\":\"\",\"vorDoiUrl\":\"\",\"workflowStages\":[]},\"version\":\"v1\",\"identity\":\"rs-6795020\",\"journalConfig\":\"researchsquare\"},\"__N_SSP\":true},\"page\":\"/article/[identity]/[[...version]]\",\"query\":{\"redirect\":\"/article/rs-6795020\",\"identity\":\"rs-6795020\",\"version\":[\"v1\"]},\"buildId\":\"XKTyCvWXoU3ODBz1xrDgd\",\"isFallback\":false,\"isExperimentalCompile\":false,\"dynamicIds\":[84888],\"gssp\":true,\"scriptLoader\":[]}","source_license":"CC-BY-4.0","license_restricted":false}