Automatic GPU Memory Access Optimization for AoSoA-based Application in OP2 Framework

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Portable parallel programming methods are attractive for application developers as hardware architectures are increasingly diverse. OP2 is a domain-specific programming framework for unstructured mesh applications that supports unified programming for multiple hardware platforms. Structure, as a user-defined data type that groups items of possibly different types into a single type, is commonly used in many applications. Current OP2 implementations face limitations in leveraging GPU memory hierarchies when handling complex data organizations, specifically the Array of Structure of Array (AoSoA) patterns, in which each element of the top-level array is a structure of multi-dimensional arrays. To address this issue, we first propose a new SoA (Structure of Array) layout transformation algorithm for AoSoA-based application to optimize the data access locality. Then, we introduce new OP2 primitives to enable CUDA codes to utilize the local memory and the shared memory. These enhancements, integrated into OP2's library and source-to-source translator, enable automatic generation of optimized CUDA code. We evaluate the proposed approaches with a high-order unstructured CFD application on representative GPUs. Compared to the original implementation, the optimized implementation improves the performance for up to 25.68x on NVIDIA V100S, 4.74x on NVIDIA A100 and 3.7x on Hygon Z100 DCU. We also measure a selected set of GPU low-level performance metrics to better explain the results.
Full text 15,668 characters · extracted from preprint-html · click to expand
Automatic GPU Memory Access Optimization for AoSoA-based Application in OP2 Framework | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Automatic GPU Memory Access Optimization for AoSoA-based Application in OP2 Framework Tong Lei, Zongjing Chen, Yonggang Che, Chuanfu Xu, Zhe Dai, Jian Zhang This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6318689/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 22 You are reading this latest preprint version Abstract Portable parallel programming methods are attractive for application developers as hardware architectures are increasingly diverse. OP2 is a domain-specific programming framework for unstructured mesh applications that supports unified programming for multiple hardware platforms. Structure, as a user-defined data type that groups items of possibly different types into a single type, is commonly used in many applications. Current OP2 implementations face limitations in leveraging GPU memory hierarchies when handling complex data organizations, specifically the Array of Structure of Array (AoSoA) patterns, in which each element of the top-level array is a structure of multi-dimensional arrays. To address this issue, we first propose a new SoA (Structure of Array) layout transformation algorithm for AoSoA-based application to optimize the data access locality. Then, we introduce new OP2 primitives to enable CUDA codes to utilize the local memory and the shared memory. These enhancements, integrated into OP2's library and source-to-source translator, enable automatic generation of optimized CUDA code. We evaluate the proposed approaches with a high-order unstructured CFD application on representative GPUs. Compared to the original implementation, the optimized implementation improves the performance for up to 25.68x on NVIDIA V100S, 4.74x on NVIDIA A100 and 3.7x on Hygon Z100 DCU. We also measure a selected set of GPU low-level performance metrics to better explain the results. OP2 framework unstructured mesh applications GPU memory hierarchy code generation automatic performance optimization Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Editorial decision: Revision requested 01 Jul, 2025 Reviews received at journal 30 Jun, 2025 Reviews received at journal 30 Jun, 2025 Reviews received at journal 29 Jun, 2025 Reviewers agreed at journal 23 Jun, 2025 Reviewers agreed at journal 23 Jun, 2025 Reviewers agreed at journal 23 Jun, 2025 Reviewers agreed at journal 18 Jun, 2025 Reviewers agreed at journal 18 Jun, 2025 Reviewers agreed at journal 18 Jun, 2025 Reviewers agreed at journal 17 Jun, 2025 Reviews received at journal 17 Jun, 2025 Reviewers agreed at journal 16 Jun, 2025 Reviewers agreed at journal 15 Jun, 2025 Reviewers agreed at journal 15 Jun, 2025 Reviewers agreed at journal 15 Jun, 2025 Reviewers agreed at journal 13 Jun, 2025 Reviewers agreed at journal 13 Jun, 2025 Reviewers invited by journal 13 Jun, 2025 Editor assigned by journal 31 Mar, 2025 Submission checks completed at journal 31 Mar, 2025 First submitted to journal 27 Mar, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6318689","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":472038199,"identity":"ef8c31aa-3a25-407a-a9c1-2112d4bb30b4","order_by":0,"name":"Tong Lei","email":"","orcid":"","institution":"National University of Defense Technology","correspondingAuthor":false,"prefix":"","firstName":"Tong","middleName":"","lastName":"Lei","suffix":""},{"id":472038200,"identity":"842de42a-f633-44ad-9d07-473acb2cea33","order_by":1,"name":"Zongjing Chen","email":"","orcid":"","institution":"National University of Defense Technology","correspondingAuthor":false,"prefix":"","firstName":"Zongjing","middleName":"","lastName":"Chen","suffix":""},{"id":472038201,"identity":"4fdf53ed-6532-4691-9237-da7317908bef","order_by":2,"name":"Yonggang Che","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAwElEQVRIiWNgGAWjYHACNgaGCiB1AIh5iNdyBkgdI0kLYxspWgzOL3724OO8O4l99xsYH7xtY5A3J6jlxjNzw5nbniXOPMbAbDi3jcFwZwNBLWfYpHm3HU7ccIwByGhjSDA4QIyWv3PAWth/E6flfA+bNGMDxBZmorRI3mAzk+w5dth45rHEZsk55yQMNxDSwnf+8DOJHzWHZfsOHz744U2ZjTxBWxRuJMCYjA1AQoKAeiCQ7ydk6CgYBaNgFIwCAPwcRwubBqVpAAAAAElFTkSuQmCC","orcid":"","institution":"National University of Defense Technology","correspondingAuthor":true,"prefix":"","firstName":"Yonggang","middleName":"","lastName":"Che","suffix":""},{"id":472038202,"identity":"4604cb17-e12b-4937-ad49-89d612cb2b4b","order_by":3,"name":"Chuanfu Xu","email":"","orcid":"","institution":"National University of Defense Technology","correspondingAuthor":false,"prefix":"","firstName":"Chuanfu","middleName":"","lastName":"Xu","suffix":""},{"id":472038203,"identity":"56754c73-77ef-4dde-a115-41e7ad542b4d","order_by":4,"name":"Zhe Dai","email":"","orcid":"","institution":"China Aerodynamics Research and Development Center","correspondingAuthor":false,"prefix":"","firstName":"Zhe","middleName":"","lastName":"Dai","suffix":""},{"id":472038204,"identity":"e96138fe-7594-475a-91b8-1fad6473c0dd","order_by":5,"name":"Jian Zhang","email":"","orcid":"","institution":"China Aerodynamics Research and Development Center","correspondingAuthor":false,"prefix":"","firstName":"Jian","middleName":"","lastName":"Zhang","suffix":""}],"badges":[],"createdAt":"2025-03-27 08:53:28","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6318689/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6318689/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":84766268,"identity":"5bd4c18c-a130-4b52-a03c-87a25adbc9ef","added_by":"auto","created_at":"2025-06-17 07:10:25","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1035589,"visible":true,"origin":"","legend":"","description":"","filename":"AutomaticGPUMemoryAccessOptimizationforAoSoAbasedApplicationinOP2Framework.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6318689/v1_covered_42a3ee54-0328-4219-b59e-4584f09720fb.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Automatic GPU Memory Access Optimization for AoSoA-based Application in OP2 Framework","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"the-journal-of-supercomputing","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"Learn more about [The Journal of Supercomputing](https://www.springer.com/journal/11227)","snPcode":"11227","submissionUrl":"https://submission.nature.com/new-submission/11227/3","title":"The Journal of Supercomputing","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"OP2 framework, unstructured mesh applications, GPU memory hierarchy, code generation, automatic performance optimization","lastPublishedDoi":"10.21203/rs.3.rs-6318689/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6318689/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"Portable parallel programming methods are attractive for application developers as hardware architectures are increasingly diverse. OP2 is a domain-specific programming framework for unstructured mesh applications that supports unified programming for multiple hardware platforms. Structure, as a user-defined data type that groups items of possibly different types into a single type, is commonly used in many applications. Current OP2 implementations face limitations in leveraging GPU memory hierarchies when handling complex data organizations, specifically the Array of Structure of Array (AoSoA) patterns, in which each element of the top-level array is a structure of multi-dimensional arrays. To address this issue, we first propose a new SoA (Structure of Array) layout transformation algorithm for AoSoA-based application to optimize the data access locality. Then, we introduce new OP2 primitives to enable CUDA codes to utilize the local memory and the shared memory. These enhancements, integrated into OP2's library and source-to-source translator, enable automatic generation of optimized CUDA code. We evaluate the proposed approaches with a high-order unstructured CFD application on representative GPUs. Compared to the original implementation, the optimized implementation improves the performance for up to 25.68x on NVIDIA V100S, 4.74x on NVIDIA A100 and 3.7x on Hygon Z100 DCU. We also measure a selected set of GPU low-level performance metrics to better explain the results.","manuscriptTitle":"Automatic GPU Memory Access Optimization for AoSoA-based Application in OP2 Framework","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-06-17 07:02:20","doi":"10.21203/rs.3.rs-6318689/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2025-07-01T15:11:08+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-06-30T23:00:34+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-06-30T20:39:11+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-06-29T14:22:16+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"164716974431958508883216149809391319789","date":"2025-06-23T09:25:52+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"230799913328920975231970326730934222873","date":"2025-06-23T08:37:28+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"39424780614350446288677300879742925418","date":"2025-06-23T07:51:35+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"175244072557074771495683806564435281075","date":"2025-06-18T21:27:42+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"217047140787212686079052832331742055426","date":"2025-06-18T17:55:04+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"159054314157409308628682626943081521042","date":"2025-06-18T15:32:20+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"140675580839311336348570852918356620255","date":"2025-06-17T22:33:40+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-06-17T09:04:58+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"315104186702367549367644367919596059787","date":"2025-06-16T07:19:35+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"293581707860029120125316141466725317112","date":"2025-06-15T18:35:24+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"47855220856640407974093687392066153607","date":"2025-06-15T16:40:28+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"96071713244884135192296042966936105474","date":"2025-06-15T14:18:10+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"166744690164143480521335928225913508933","date":"2025-06-13T21:03:22+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"182858594380475803766644567584294918627","date":"2025-06-13T17:26:40+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-06-13T14:26:13+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-03-31T08:35:07+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-03-31T08:34:00+00:00","index":"","fulltext":""},{"type":"submitted","content":"The Journal of Supercomputing","date":"2025-03-27T08:43:14+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"the-journal-of-supercomputing","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"Learn more about [The Journal of Supercomputing](https://www.springer.com/journal/11227)","snPcode":"11227","submissionUrl":"https://submission.nature.com/new-submission/11227/3","title":"The Journal of Supercomputing","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"231ecfbb-d544-4004-b380-01c495eab90b","owner":[],"postedDate":"June 17th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2025-07-30T23:38:19+00:00","versionOfRecord":[],"versionCreatedAt":"2025-06-17 07:02:20","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6318689","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6318689","identity":"rs-6318689","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00