Reliability Evaluation of Reinforcement Learning Methods for Mechanical Systems with Increasing Complexity | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Reliability Evaluation of Reinforcement Learning Methods for Mechanical Systems with Increasing Complexity Peter Manzl, Oleg Rogov, Johannes Gerstmayr, Aki Mikkola, Grzegorz Orzechowski This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-3066420/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 22 Dec, 2023 Read the published version in Multibody System Dynamics → Version 1 posted 7 You are reading this latest preprint version Abstract Reinforcement Learning (RL) is one of the emerging fields of Artificial Intelligence (AI) intended for designing agents that take actions in the physical environment. RL has many vital applications, including robotics and autonomous vehicles. The key characteristic of RL is its ability to learn from experience without requiring direct programming or supervision. To learn, an agent interacts with an environment by acting and observing the resulting states and rewards. In most practical applications, an environment is implemented as a virtual system due to cost, time, and safety concerns. Simultaneously, Multibody System Dynamics (MSD) is a framework for efficiently and systematically developing virtual systems of arbitrary complexity. MSD is commonly used to create virtual models of robots, vehicles, machinery, and humans. The features of RL and MSD make them perfect companions in building sophisticated, automated, and autonomous mechatronic systems. The research demonstrates the use of RL in controlling multibody systems. While AI methods are used to solve some of the most challenging tasks in engineering, their proper understanding and implementation are demanding. Therefore, we introduce and detail three commonly used RL algorithms to control the inverted N-pendulum on the cart. Single-, double-, and triple-pendulum configurations are investigated, showing the capability of RL methods to handle increasingly complex dynamical systems. We show 2D state space zones where the agent succeeds or fails the stabilization. Despite passing randomized tests during training, "blind spots" may occur where the agent's policy fails. Results confirm that RL is a versatile, although complex, control engineering approach. reinforcement learning reliability analysis inverse pendulum machine learning dynamical systems Full Text Additional Declarations Competing interest reported. Non-financial interests: Johannes Gerstmayr is a member of the Editorial Advisory Board of Multibody System Dynamics. Aki Mikkola is co-editor-in-chief of the journal Multibody System Dynamics. Johannes Gerstmayr and Grzegorz Orzechowski are guest editors of the special issue on "Artificial intelligence within multibody system dynamics" in the journal Multibody System Dynamics. None of the listed editors or editorial board members receive compensation for the listed appointments. Supplementary Files Videos.zip supplementaryCode.zip Cite Share Download PDF Status: Published Journal Publication published 22 Dec, 2023 Read the published version in Multibody System Dynamics → Version 1 posted Editorial decision: Revision requested 07 Nov, 2023 Reviews received at journal 12 Oct, 2023 Reviewers agreed at journal 20 Sep, 2023 Reviewers invited by journal 16 Sep, 2023 Editor assigned by journal 30 Jun, 2023 Submission checks completed at journal 16 Jun, 2023 First submitted to journal 15 Jun, 2023 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-3066420","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":210491296,"identity":"0c09832f-7fa0-400f-afab-54e82511d3cd","order_by":0,"name":"Peter Manzl","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA+0lEQVRIie3PMWrDMBTG8U8I7MXg1YZCriBjSJeUXKXFkClDoUugi0JAUyAXyCG0ZrN40C6BHsAZYgLu4qHdGgi0cig0S+WOHfSHx9Py4yHA5/uHjeV5kR0+30OUopeIknfr1Q5biB9y20s+O6IS4C/keWUOM1SD69iox+h+lyOkPY4fjr8siWdbVNlGGlVFohkimgi2dF1JiiCVqJg2847QCJgCkYsMDuFRgsaamHo4k7gFOzmv8IBZcqefmOKWDJFMwZ1XoiJPpaBCb9kiXYsmD5JG0NXEQUJTv8sZ3egXMm/taZet4qKu29Hv5BtevAM7ZR/w+Xw+n7svG/lSvc/57k8AAAAASUVORK5CYII=","orcid":"","institution":"Universität Innsbruck","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Peter","middleName":"","lastName":"Manzl","suffix":""},{"id":210491297,"identity":"226f5d4c-7bd4-4e2a-b582-b59e6877dc9a","order_by":1,"name":"Oleg Rogov","email":"","orcid":"","institution":"Lappeenranta University of Technology","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Oleg","middleName":"","lastName":"Rogov","suffix":""},{"id":210491298,"identity":"203b1455-6cf4-471d-b123-824504d38a76","order_by":2,"name":"Johannes Gerstmayr","email":"","orcid":"","institution":"Universität Innsbruck","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Johannes","middleName":"","lastName":"Gerstmayr","suffix":""},{"id":210491299,"identity":"9b1f0b4d-ee9b-445c-957e-5b4cd7102d8d","order_by":3,"name":"Aki Mikkola","email":"","orcid":"","institution":"Lappeenranta University of Technology","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Aki","middleName":"","lastName":"Mikkola","suffix":""},{"id":210491300,"identity":"c85bc77b-fa57-4fe8-8f48-88dae2a095f6","order_by":4,"name":"Grzegorz Orzechowski","email":"","orcid":"","institution":"Lappeenranta University of Technology","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Grzegorz","middleName":"","lastName":"Orzechowski","suffix":""}],"badges":[],"createdAt":"2023-06-15 07:44:33","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-3066420/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-3066420/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1007/s11044-023-09960-2","type":"published","date":"2023-12-22T15:00:53+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":48776742,"identity":"51792ef0-dd31-4f46-a51f-d591cd923bdb","added_by":"auto","created_at":"2023-12-25 15:07:41","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":6449797,"visible":true,"origin":"","legend":"","description":"","filename":"ManuscriptReliabilityOfRLMethodsForMechanicalSystems.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3066420/v1_covered_e6a31e44-2897-4806-801d-597de287f67e.pdf"},{"id":38791242,"identity":"4ce764a7-a9c4-40c5-b1ce-668f7eab4182","added_by":"auto","created_at":"2023-06-20 02:43:10","extension":"zip","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":1393277,"visible":true,"origin":"","legend":"","description":"","filename":"Videos.zip","url":"https://assets-eu.researchsquare.com/files/rs-3066420/v1/a28c74891cc2a717b0ff6039.zip"},{"id":38791241,"identity":"562ba074-2fb3-416f-b8f3-8291c02b01e3","added_by":"auto","created_at":"2023-06-20 02:43:10","extension":"zip","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":38638,"visible":true,"origin":"","legend":"","description":"","filename":"supplementaryCode.zip","url":"https://assets-eu.researchsquare.com/files/rs-3066420/v1/a4f7d6d3b1428c2a65bc4882.zip"}],"financialInterests":"Competing interest reported. Non-financial interests: Johannes Gerstmayr is a member of the Editorial Advisory Board of Multibody System Dynamics. \nAki Mikkola is co-editor-in-chief of the journal Multibody System Dynamics. Johannes Gerstmayr and Grzegorz Orzechowski are guest editors of the special issue on \"Artificial intelligence within multibody system dynamics\" in the journal Multibody System Dynamics.\nNone of the listed editors or editorial board members receive compensation for the listed appointments.","formattedTitle":"Reliability Evaluation of Reinforcement Learning Methods for Mechanical Systems with Increasing Complexity","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"multibody-system-dynamics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"mubo","sideBox":"Learn more about [Multibody System Dynamics](http://link.springer.com/journal/11044)","snPcode":"11044","submissionUrl":"https://submission.nature.com/new-submission/11044/3","title":"Multibody System Dynamics","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"reinforcement learning, reliability analysis, inverse pendulum, machine learning, dynamical systems","lastPublishedDoi":"10.21203/rs.3.rs-3066420/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-3066420/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"Reinforcement Learning (RL) is one of the emerging fields of Artificial Intelligence (AI) intended for designing agents that take actions in the physical environment. RL has many vital applications, including robotics and autonomous vehicles. The key characteristic of RL is its ability to learn from experience without requiring direct programming or supervision. To learn, an agent interacts with an environment by acting and observing the resulting states and rewards. In most practical applications, an environment is implemented as a virtual system due to cost, time, and safety concerns. Simultaneously, Multibody System Dynamics (MSD) is a framework for efficiently and systematically developing virtual systems of arbitrary complexity. MSD is commonly used to create virtual models of robots, vehicles, machinery, and humans. The features of RL and MSD make them perfect companions in building sophisticated, automated, and autonomous mechatronic systems. The research demonstrates the use of RL in controlling multibody systems. While AI methods are used to solve some of the most challenging tasks in engineering, their proper understanding and implementation are demanding. Therefore, we introduce and detail three commonly used RL algorithms to control the inverted N-pendulum on the cart. Single-, double-, and triple-pendulum configurations are investigated, showing the capability of RL methods to handle increasingly complex dynamical systems. We show 2D state space zones where the agent succeeds or fails the stabilization. Despite passing randomized tests during training, \"blind spots\" may occur where the agent's policy fails. Results confirm that RL is a versatile, although complex, control engineering approach.","manuscriptTitle":"Reliability Evaluation of Reinforcement Learning Methods for Mechanical Systems with Increasing Complexity","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2023-06-20 02:43:05","doi":"10.21203/rs.3.rs-3066420/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2023-11-07T12:54:06+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2023-10-12T17:02:31+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"461e878f-1339-45e0-8999-fce2ceab17d5","date":"2023-09-20T07:49:11+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2023-09-16T22:47:26+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2023-06-30T17:59:39+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2023-06-16T10:22:33+00:00","index":"","fulltext":""},{"type":"submitted","content":"Multibody System Dynamics","date":"2023-06-15T07:41:21+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"multibody-system-dynamics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"mubo","sideBox":"Learn more about [Multibody System Dynamics](http://link.springer.com/journal/11044)","snPcode":"11044","submissionUrl":"https://submission.nature.com/new-submission/11044/3","title":"Multibody System Dynamics","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"3e264b6b-eb87-4b91-8c92-28a4737fadc3","owner":[],"postedDate":"June 20th, 2023","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2023-12-25T15:02:44+00:00","versionOfRecord":{"articleIdentity":"rs-3066420","link":"https://doi.org/10.1007/s11044-023-09960-2","journal":{"identity":"multibody-system-dynamics","isVorOnly":false,"title":"Multibody System Dynamics"},"publishedOn":"2023-12-22 15:00:53","publishedOnDateReadable":"December 22nd, 2023"},"versionCreatedAt":"2023-06-20 02:43:05","video":"","vorDoi":"10.1007/s11044-023-09960-2","vorDoiUrl":"https://doi.org/10.1007/s11044-023-09960-2","workflowStages":[]},"version":"v1","identity":"rs-3066420","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-3066420","identity":"rs-3066420","version":["v1"]},"buildId":"FbvkV6FR0MCFSLy54lSbu","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.