Deep Reinforcement Learning-Based Optimization of Reconfigurable Intelligent Surfaces (RIS) for 6G Multi-User Connectivity in NLOS Environments | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Deep Reinforcement Learning-Based Optimization of Reconfigurable Intelligent Surfaces (RIS) for 6G Multi-User Connectivity in NLOS Environments Zacheous Aasa This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8668566/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract The transition towards the sixth generation (6G) of wireless networks requires ultra-high data rates and seamless connectivity, even in frequency bands like mmWave and THz, which are highly susceptible to physical blockages. RIS have emerged as a key technology in mitigating NLOS limitations through programmable steering of electromagnetic waves by passive reflection. However, real-time optimization of high-dimensional RIS phase shifts in dynamic multi-user environments remains an NP-hard challenge that is hardly solvable efficiently by traditional methods based on mathematical optimization. This paper presents a new DRL framework, utilizing a TD3 architecture, for the optimization of sum-rate performance of multi-user links. The model incorporates real-time environmental feedback and imperfect CSI for ensuring robust connectivity. Simulation results show that the DRL-RIS framework proposed here achieves remarkable gains in the sum-rate performance and outperforms conventional convex optimization baselines in computational latency and energy efficiency. Keywords: 6G Networks, Reconfigurable Intelligent Surfaces (RIS), Deep Reinforcement Learning (DRL), Non-Line-of-Sight (NLOS), Beamforming 6G Networks Reconfigurable Intelligent Surfaces (RIS) Deep Reinforce-ment Learning (DRL) Non-Line-of-Sight (NLOS) Beamforming Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8668566","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":578682281,"identity":"228fc2a4-a7b8-4035-8909-961383caf33e","order_by":0,"name":"Zacheous Aasa","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA8klEQVRIiWNgGAWjYHACNiR2BRAzMzeQouUMSAsjKVoY28Akfi3y7YefPfhQcdiegf3w4w8/59VG87cDtfyo2IZTC2NPmrnhjDOHExt40swke7cdz51xmLGBsefMbZxamBly2KR52w4nMDAkmDHwbjuW2wDUwszYhlsLG/8bNum/bUCH8T///PHvnGO58wlp4ZEA2sLYBlQmkWMgzdtQk7uBkBYJiWdmkj1n0hPbJN6UScscO5C7EajlID6/yPcnP5P4UWFtz8+fvvnjm5q63HnnDx988KMCtxaEpyDUYTB5gLB6BKgjRfEoGAWjYBSMEAAAvS9W0AOCaloAAAAASUVORK5CYII=","orcid":"","institution":"Ladoke Akintola University of Technology","correspondingAuthor":true,"prefix":"","firstName":"Zacheous","middleName":"","lastName":"Aasa","suffix":""}],"badges":[],"createdAt":"2026-01-22 10:39:43","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8668566/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8668566/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":100940718,"identity":"606a5efa-f03d-4c48-8549-4f7d682c0654","added_by":"auto","created_at":"2026-01-23 04:32:24","extension":"dotx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":1465340,"visible":true,"origin":"","legend":"","description":"","filename":"AASADeepReinforcementLearningBasedOptimization202100574.dotx","url":"https://assets-eu.researchsquare.com/files/rs-8668566/v1/6331869e42cb1b83cee1bf89.dotx"},{"id":100940717,"identity":"f4b77975-2661-4608-9cd0-16c137913231","added_by":"auto","created_at":"2026-01-23 04:32:23","extension":"json","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":3214,"visible":true,"origin":"","legend":"","description":"","filename":"c7c33dceeabb4008aad531a374bd9bd1.json","url":"https://assets-eu.researchsquare.com/files/rs-8668566/v1/de48289fa0e7eaaed5b8d690.json"},{"id":106804704,"identity":"1b5e5054-460c-485b-8349-08e718ebaaf9","added_by":"auto","created_at":"2026-04-13 15:12:14","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":624401,"visible":true,"origin":"","legend":"","description":"","filename":"AASADeepReinforcementLearningBasedOptimization202100574.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8668566/v1_covered_61b345ae-cf51-4c77-87b4-e3a2aafd5f71.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Deep Reinforcement Learning-Based Optimization of Reconfigurable Intelligent Surfaces (RIS) for 6G Multi-User Connectivity in NLOS Environments","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"6G Networks, Reconfigurable Intelligent Surfaces (RIS), Deep Reinforce-ment Learning (DRL), Non-Line-of-Sight (NLOS), Beamforming","lastPublishedDoi":"10.21203/rs.3.rs-8668566/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8668566/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"The transition towards the sixth generation (6G) of wireless networks requires ultra-high data rates and seamless connectivity, even in frequency bands like mmWave and THz, which are highly susceptible to physical blockages. RIS have emerged as a key technology in mitigating NLOS limitations through programmable steering of electromagnetic waves by passive reflection. However, real-time optimization of high-dimensional RIS phase shifts in dynamic multi-user environments remains an NP-hard challenge that is hardly solvable efficiently by traditional methods based on mathematical optimization. This paper presents a new DRL framework, utilizing a TD3 architecture, for the optimization of sum-rate performance of multi-user links. The model incorporates real-time environmental feedback and imperfect CSI for ensuring robust connectivity. Simulation results show that the DRL-RIS framework proposed here achieves remarkable gains in the sum-rate performance and outperforms conventional convex optimization baselines in computational latency and energy efficiency.\nKeywords: 6G Networks, Reconfigurable Intelligent Surfaces (RIS), Deep Reinforcement Learning (DRL), Non-Line-of-Sight (NLOS), Beamforming","manuscriptTitle":"Deep Reinforcement Learning-Based Optimization of Reconfigurable Intelligent Surfaces (RIS) for 6G Multi-User Connectivity in NLOS Environments","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-01-23 04:32:19","doi":"10.21203/rs.3.rs-8668566/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"1fcd830d-9aaf-4d3a-ab05-dc25c4ca0d81","owner":[],"postedDate":"January 23rd, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2026-04-13T19:15:39+00:00","versionOfRecord":[],"versionCreatedAt":"2026-01-23 04:32:19","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8668566","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8668566","identity":"rs-8668566","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.