Floating-point matrix multiplication optimized through the integration of hybrid multiplier techniques | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Floating-point matrix multiplication optimized through the integration of hybrid multiplier techniques Y. R.Annie Bessant, A. Chrispin Jiji, J. Grace Jency Gnanammal This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7629322/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract In order to support high power computing applications, we aimed to design a novel framework for floating point matrix multiplication. In the proposal of matrix multiplication processing elements (PEs), multipliers such as Strassen, Karatsuba and Vedic are utilized to maximize the architecture performance. The Urdhava Tiryabhyam sutra from the Vedic multiplier is one of them that decreases partial products in multiplication while simultaneously speeding up the process. This paper proposes the design of Strassen matrix multiplication algorithm with the combined features of Urdhava Tiryabhyam and Karatsuba multiplier. The Karatsuba algorithm is inefficient at lower orders of multiplication, but it works well at higher orders. Thus, at a lower level, the Urdhava Tiryabhyam Multiplier is applied. A floating-point matrix multiplication architecture was developed using Verilog HDL and implemented on the Virtex-5 FPGA – XC5VSX95T. The design was developed, verified through simulation, and implemented using Xilinx ISE 13.5. Multiplier Processing Element Floating point sub matrices performance Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7629322","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":542616601,"identity":"bfef72b5-5dfd-486c-8782-10ba26dd1a67","order_by":0,"name":"Y. R.Annie Bessant","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAvUlEQVRIiWNgGAWjYJCCAwwVMCYbGBGj5QypWhgY25C0EAT8s3sPHvg4ry6xv/3wAYYPZYcZ+KQb8GuRuHMu4eDMbYeNJc6kJTDOOHeYgU3mAAFrbuQYHObddkDOQILHgJm3DahFIgG/DnmQlr9z6ngMJPg/MP8lRosBSAtjAzPIFgZmRmK0GAK1HOw5BvYLkHEunYegFrkbOcYfftSAQ+zhgx9l1nLyMwhoQQEHgJiHBPWjYBSMglEwCnABAPv+QQvnIkuKAAAAAElFTkSuQmCC","orcid":"","institution":"St.Xavier’s Catholic college of Engineering","correspondingAuthor":true,"prefix":"","firstName":"Y.","middleName":"R.Annie","lastName":"Bessant","suffix":""},{"id":542616602,"identity":"2bdc4c3f-4ab2-4f84-bbb1-f12cd653d096","order_by":1,"name":"A. Chrispin Jiji","email":"","orcid":"","institution":"Cambridge Institute of Technology","correspondingAuthor":false,"prefix":"","firstName":"A.","middleName":"Chrispin","lastName":"Jiji","suffix":""},{"id":542616603,"identity":"5e904917-ede7-4e08-a974-7a4539b308bb","order_by":2,"name":"J. Grace Jency Gnanammal","email":"","orcid":"","institution":"Sri Krishna College of Engineering and Technology","correspondingAuthor":false,"prefix":"","firstName":"J.","middleName":"Grace Jency","lastName":"Gnanammal","suffix":""}],"badges":[],"createdAt":"2025-09-16 10:38:19","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7629322/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7629322/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":95665690,"identity":"b7e63f8b-6272-4e00-a6ee-1b6dc67a8e6e","added_by":"auto","created_at":"2025-11-11 16:47:01","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":1054895,"visible":true,"origin":"","legend":"","description":"","filename":"journalsubmitnew16092025.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7629322/v1/b95c14c873dcbecab3d184d6.pdf"},{"id":95665689,"identity":"78c67b47-347d-4a6b-b215-90a68fc47ffa","added_by":"auto","created_at":"2025-11-11 16:47:01","extension":"json","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":4818,"visible":true,"origin":"","legend":"","description":"","filename":"e91faa2488e5411ba0a1dbe1aa63e785.json","url":"https://assets-eu.researchsquare.com/files/rs-7629322/v1/5410e6deb18ba765f879e534.json"},{"id":96164985,"identity":"7d318254-aa21-489d-be49-c6844c108f27","added_by":"auto","created_at":"2025-11-18 09:24:30","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":704502,"visible":true,"origin":"","legend":"","description":"","filename":"journalsubmitnew16092025.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7629322/v1_covered_6aa998b3-91e6-49e4-9ac2-972e880b158f.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Floating-point matrix multiplication optimized through the integration of hybrid multiplier techniques","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Multiplier, Processing Element, Floating point, sub matrices, performance","lastPublishedDoi":"10.21203/rs.3.rs-7629322/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7629322/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"In order to support high power computing applications, we aimed to design a novel framework for floating point matrix multiplication. In the proposal of matrix multiplication processing elements (PEs), multipliers such as Strassen, Karatsuba and Vedic are utilized to maximize the architecture performance. The Urdhava Tiryabhyam sutra from the Vedic multiplier is one of them that decreases partial products in multiplication while simultaneously speeding up the process. This paper proposes the design of Strassen matrix multiplication algorithm with the combined features of Urdhava Tiryabhyam and Karatsuba multiplier. The Karatsuba algorithm is inefficient at lower orders of multiplication, but it works well at higher orders. Thus, at a lower level, the Urdhava Tiryabhyam Multiplier is applied. A floating-point matrix multiplication architecture was developed using Verilog HDL and implemented on the Virtex-5 FPGA – XC5VSX95T. The design was developed, verified through simulation, and implemented using Xilinx ISE 13.5.","manuscriptTitle":"Floating-point matrix multiplication optimized through the integration of hybrid multiplier techniques","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-11-11 16:46:56","doi":"10.21203/rs.3.rs-7629322/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"f47dfca0-7b16-412b-bfad-74be666d920f","owner":[],"postedDate":"November 11th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-11-18T09:24:04+00:00","versionOfRecord":[],"versionCreatedAt":"2025-11-11 16:46:56","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7629322","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7629322","identity":"rs-7629322","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.