Towards a GPU-enabled billionare SVD in pyLOM | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Towards a GPU-enabled billionare SVD in pyLOM Arnau Miró, Benet Eiximeno, Lucas Gasparino, Nathan Kutz, Ivette Rodriguez, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7678279/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 17 Jan, 2026 Read the published version in Acta Mechanica → Version 1 posted You are reading this latest preprint version Abstract We develop and implement an accelerated high-performance and open-source computing environment for model order reduction in fluid dynamics called pyLOM. It contains the algorithms of proper orthogonal decomposition, dynamic mode decomposition and spectral proper order decomposition that are based on parallel GPU-accelerated algorithms. The library is profiled in detail under the MareNostrum V supercomputer. The largest case run has been of a billion nodes with a thousand snapshots, computed under 20 seconds with 100 GPUs. While the studied applications are memory bound, a hybrid parallel randomized QR factorization has been found to be able to leverage such large matrices. The largest speedup factor of 83 has been found on the QR factorization, while the matrix--matrix multiplication has shown a speedup factor of about 2. Additionally, two examples of application are provided in the flow around a cylinder at $Re_D=10^4$ and the Windsor body at a Reynolds number of $Re_L = 2.9\times10^6$. The largest dataset of the Windsor body consists of 422 snapshots on a grid of 1.4 billion nodes, and its POD is computed under 3 seconds with 100 GPUs. This showcases the efficiency of GPUs, resulting in a 97% reduction in energy to solution and a reduction of 0.11 kg of $CO_2$ emissions. The scalability and efficiency achieved suggest that this framework can play a key role in reducing the energy demands and environmental impact of large-scale data analysis and model order reduction across a wide range of applications. GPU acceleration fluid dynamics reduced order models single value decomposition Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Published Journal Publication published 17 Jan, 2026 Read the published version in Acta Mechanica → Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7678279","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":525871743,"identity":"c619ccf5-add5-4152-af08-54c37af1a5be","order_by":0,"name":"Arnau Miró","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAxUlEQVRIiWNgGAWjYBADHgYG5gMQ5gHitbAlkKYFpMuAOC38s88e/MzDsE3GvL3nm3RBDYMc340E/FokzuUlS/Mw3OaROXN2m/SMYwzGkoS0MJzhMQBrkZDI3SbN28CQuIGQFvkzPMa/IVpynoG01BPUYnCGxwxqSw4bSEuCASEthkAtlnMMgFp4jhlbzzgmYTjzzAP8WuSADrvxpuK2vQR788PbBTU28nzHCdgCdR6EYgaGIDHKkQAziepHwSgYBaNghAAAMhU6Fr59MqgAAAAASUVORK5CYII=","orcid":"","institution":"Barcelona Supercomputing Center","correspondingAuthor":true,"prefix":"","firstName":"Arnau","middleName":"","lastName":"Miró","suffix":""},{"id":525871744,"identity":"5b38e757-bf9e-4a4a-8743-da6201076b8d","order_by":1,"name":"Benet Eiximeno","email":"","orcid":"","institution":"Barcelona Supercomputing Center","correspondingAuthor":false,"prefix":"","firstName":"Benet","middleName":"","lastName":"Eiximeno","suffix":""},{"id":525871745,"identity":"cb0bd134-057c-4d60-b3a7-3e0ca385922a","order_by":2,"name":"Lucas Gasparino","email":"","orcid":"","institution":"Barcelona Supercomputing Center","correspondingAuthor":false,"prefix":"","firstName":"Lucas","middleName":"","lastName":"Gasparino","suffix":""},{"id":525871747,"identity":"7bcbf313-665c-4ebc-bcf3-ccdd0269af4a","order_by":3,"name":"Nathan Kutz","email":"","orcid":"","institution":"University of Washington","correspondingAuthor":false,"prefix":"","firstName":"Nathan","middleName":"","lastName":"Kutz","suffix":""},{"id":525871750,"identity":"75f710d2-f87d-4478-826a-09c98cce66c0","order_by":4,"name":"Ivette Rodriguez","email":"","orcid":"","institution":"Universitat Politècnica de Catalunya","correspondingAuthor":false,"prefix":"","firstName":"Ivette","middleName":"","lastName":"Rodriguez","suffix":""},{"id":525871752,"identity":"a1d27564-3160-481f-a694-7bc8a195f0e9","order_by":5,"name":"Oriol Lehmkuhl","email":"","orcid":"","institution":"Barcelona Supercomputing Center","correspondingAuthor":false,"prefix":"","firstName":"Oriol","middleName":"","lastName":"Lehmkuhl","suffix":""}],"badges":[],"createdAt":"2025-09-22 14:23:25","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7678279/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7678279/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1007/s00707-025-04621-1","type":"published","date":"2026-01-17T16:31:11+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":93271379,"identity":"94ba22cc-0cd1-4225-ad24-f7ceb88b80f7","added_by":"auto","created_at":"2025-10-10 23:47:22","extension":"json","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":8818,"visible":true,"origin":"","legend":"","description":"","filename":"ff67f0aaec8f4a58baacef420ae40191.json","url":"https://assets-eu.researchsquare.com/files/rs-7678279/v1/42415c08f16629055696d37f.json"},{"id":100616457,"identity":"52211212-9fa6-4dd5-8203-6fa0e2a107a3","added_by":"auto","created_at":"2026-01-19 17:43:03","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1056533,"visible":true,"origin":"","legend":"","description":"","filename":"pyLOMActaMechanica.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7678279/v1_covered_897812f4-ca32-4206-a5fa-5b8b097fb5b7.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Towards a GPU-enabled billionare SVD in pyLOM","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"GPU acceleration, fluid dynamics, reduced order models, single value decomposition","lastPublishedDoi":"10.21203/rs.3.rs-7678279/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7678279/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\nWe develop and implement an accelerated high-performance and open-source computing environment for model order reduction in fluid dynamics called pyLOM. It contains the algorithms of proper orthogonal decomposition, dynamic mode decomposition and spectral proper order decomposition that are based on parallel GPU-accelerated algorithms. The library is profiled in detail under the MareNostrum V supercomputer. The largest case run has been of a billion nodes with a thousand snapshots, computed under 20 seconds with 100 GPUs. While the studied applications are memory bound, a hybrid parallel randomized QR factorization has been found to be able to leverage such large matrices. The largest speedup factor of 83 has been found on the QR factorization, while the matrix--matrix multiplication has shown a speedup factor of about 2. Additionally, two examples of application are provided in the flow around a cylinder at $Re_D=10^4$ and the Windsor body at a Reynolds number of $Re_L = 2.9\\times10^6$. The largest dataset of the Windsor body consists of 422 snapshots on a grid of 1.4 billion nodes, and its POD is computed under 3 seconds with 100 GPUs. This showcases the efficiency of GPUs, resulting in a 97\\% reduction in energy to solution and a reduction of 0.11 kg of $CO_2$ emissions. The scalability and efficiency achieved suggest that this framework can play a key role in reducing the energy demands and environmental impact of large-scale data analysis and model order reduction across a wide range of applications.\n","manuscriptTitle":"Towards a GPU-enabled billionare SVD in pyLOM","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-10-10 23:47:17","doi":"10.21203/rs.3.rs-7678279/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"d5331a98-8544-4200-8317-61ed1f3a95eb","owner":[],"postedDate":"October 10th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2026-01-19T17:08:27+00:00","versionOfRecord":{"articleIdentity":"rs-7678279","link":"https://doi.org/10.1007/s00707-025-04621-1","journal":{"identity":"acta-mechanica","isVorOnly":false,"title":"Acta Mechanica"},"publishedOn":"2026-01-17 16:31:11","publishedOnDateReadable":"January 17th, 2026"},"versionCreatedAt":"2025-10-10 23:47:17","video":"","vorDoi":"10.1007/s00707-025-04621-1","vorDoiUrl":"https://doi.org/10.1007/s00707-025-04621-1","workflowStages":[]},"version":"v1","identity":"rs-7678279","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7678279","identity":"rs-7678279","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.