Natural Language-Driven Zero-Shot Generalization of Robotic Grasping Skills | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Natural Language-Driven Zero-Shot Generalization of Robotic Grasping Skills Hanwen Zhang, Yunxi Zhang, Anastasiia Voronina This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9385278/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 5 You are reading this latest preprint version Abstract Robotic grasping in unstructured open environments faces core challenges including highly diverse object categories, absent semantic understanding, and insufficient cross-category generalization. Existing methods predominantly rely on closed-set category training, and the fusion of language instructions with three-dimensional geometric information remains shallow, making zero-shot grasping of unseen objects extremely difficult. To address these limitations, this paper proposes a Natural Language-Driven Zero-Shot Generalization framework for Robotic Grasping Skills (NL-ZSGrasp), which achieves end-to-end six-degree-of-freedom zero-shot grasping of arbitrarily described objects through three collaboratively designed core modules. The Multimodal Semantic Alignment Module (MSAM) leverages a pre-trained CLIP model to map natural language instructions into scene semantic attention heatmaps and filter target point clouds. The Geometric Context-Aware Encoder (GCAE) constructs enhanced geometric representations through hierarchical PointNet++ set abstraction combined with local surface normal and principal curvature estimation. The Cross-Modal Attention Fusion Decoder (CMAF) deeply fuses semantic and geometric features via a multi-head cross-attention mechanism and jointly predicts rotation, translation, gripper width, and grasp quality score through four-branch MLP output heads. Systematic experiments on the GraspNet-1Billion benchmark dataset and a real UR10e robotic platform demonstrate that NL-ZSGrasp achieves a zero-shot grasping success rate of 83.6% on unseen object categories, outperforming the current state-of-the-art by 8.5 percentage points, while maintaining superior robustness under multiple environmental disturbances including occlusion, illumination variation, and background clutter. Ablation studies further validate the independent contribution of each module. This work provides a new methodological framework and practical reference for general-purpose language-guided robotic grasping. Robotic grasping Zero-shot learning Vision-language pre-trained models Natural language instructions Point cloud geometric perception Cross-modal attention fusion Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Reviewers agreed at journal 11 May, 2026 Reviewers invited by journal 05 May, 2026 Editor assigned by journal 25 Apr, 2026 Submission checks completed at journal 22 Apr, 2026 First submitted to journal 22 Apr, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9385278","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":638302841,"identity":"38ebe3e1-0ee5-4f65-90b9-3ed44a64b668","order_by":0,"name":"Hanwen Zhang","email":"","orcid":"","institution":"Stony Brook University","correspondingAuthor":false,"prefix":"","firstName":"Hanwen","middleName":"","lastName":"Zhang","suffix":""},{"id":638302843,"identity":"4bc49107-2f85-4934-b661-3185cc58b732","order_by":1,"name":"Yunxi Zhang","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABFElEQVRIie2RMUvEMBTHXynkluKt6XDfIVJw8sO8IHhLrkvhuMGhU13UWxVBv8K5uBotdBJvkxxxqLg6OEmnYpqrgkOKo0N+wwu85Jd/HgHweP4hTJqC0KJZgrpvBl0zyIcV0ikh++4OKvF2wyqE/kkZ01lFa4jS8eXx86I5uk3ZWsLDpoDJ0qEQmh4wBJrRl8f55qTSGVMI5ayA5MIRQ6jYrREYz5U41EA0XymwCl9Jp8IkAvJro2TQGsU8zCp3bqVLkebyaRUGhVEk9imuWaL3hCGT/EaJMD4901mseF6KJ5qcu1JGYi9uFi2/UtPXj+ZTpzvr8v5NzPcnS8f4W+wfRrZi/yN04PgPo7pXPB6Px/ObLweNaBZ/bvehAAAAAElFTkSuQmCC","orcid":"","institution":"Ludwig-Maximilians-Universität München","correspondingAuthor":true,"prefix":"","firstName":"Yunxi","middleName":"","lastName":"Zhang","suffix":""},{"id":638302844,"identity":"4dd98045-4f2a-477c-89ee-59d800b83af1","order_by":2,"name":"Anastasiia Voronina","email":"","orcid":"","institution":"Tomsk State University","correspondingAuthor":false,"prefix":"","firstName":"Anastasiia","middleName":"","lastName":"Voronina","suffix":""}],"badges":[],"createdAt":"2026-04-11 06:53:26","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9385278/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9385278/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":109222289,"identity":"95e49d30-f779-4daa-9475-3ac7b4e4965c","added_by":"auto","created_at":"2026-05-13 21:06:54","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1641176,"visible":true,"origin":"","legend":"","description":"","filename":"DiscoverAIManuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9385278/v1_covered_8ad70fd4-c274-409b-861b-d20ee9e47148.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Natural Language-Driven Zero-Shot Generalization of Robotic Grasping Skills","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"discover-artificial-intelligence","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"diai","sideBox":"Learn more about [Discover Artificial Intelligence](https://www.springer.com/44163)","snPcode":"","submissionUrl":"","title":"Discover Artificial Intelligence","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Discover Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Robotic grasping, Zero-shot learning, Vision-language pre-trained models, Natural language instructions, Point cloud geometric perception, Cross-modal attention fusion","lastPublishedDoi":"10.21203/rs.3.rs-9385278/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9385278/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eRobotic grasping in unstructured open environments faces core challenges including highly diverse object categories, absent semantic understanding, and insufficient cross-category generalization. Existing methods predominantly rely on closed-set category training, and the fusion of language instructions with three-dimensional geometric information remains shallow, making zero-shot grasping of unseen objects extremely difficult. To address these limitations, this paper proposes a Natural Language-Driven Zero-Shot Generalization framework for Robotic Grasping Skills (NL-ZSGrasp), which achieves end-to-end six-degree-of-freedom zero-shot grasping of arbitrarily described objects through three collaboratively designed core modules. The Multimodal Semantic Alignment Module (MSAM) leverages a pre-trained CLIP model to map natural language instructions into scene semantic attention heatmaps and filter target point clouds. The Geometric Context-Aware Encoder (GCAE) constructs enhanced geometric representations through hierarchical PointNet++ set abstraction combined with local surface normal and principal curvature estimation. The Cross-Modal Attention Fusion Decoder (CMAF) deeply fuses semantic and geometric features via a multi-head cross-attention mechanism and jointly predicts rotation, translation, gripper width, and grasp quality score through four-branch MLP output heads. Systematic experiments on the GraspNet-1Billion benchmark dataset and a real UR10e robotic platform demonstrate that NL-ZSGrasp achieves a zero-shot grasping success rate of 83.6% on unseen object categories, outperforming the current state-of-the-art by 8.5 percentage points, while maintaining superior robustness under multiple environmental disturbances including occlusion, illumination variation, and background clutter. Ablation studies further validate the independent contribution of each module. This work provides a new methodological framework and practical reference for general-purpose language-guided robotic grasping.\u003c/p\u003e","manuscriptTitle":"Natural Language-Driven Zero-Shot Generalization of Robotic Grasping Skills","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-05-13 10:04:36","doi":"10.21203/rs.3.rs-9385278/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"reviewerAgreed","content":"56807945917763730258070686399444447136","date":"2026-05-11T13:23:40+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-05-05T11:17:19+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-04-25T08:00:09+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-04-22T05:21:40+00:00","index":"","fulltext":""},{"type":"submitted","content":"Discover Artificial Intelligence","date":"2026-04-22T05:17:53+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"discover-artificial-intelligence","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"diai","sideBox":"Learn more about [Discover Artificial Intelligence](https://www.springer.com/44163)","snPcode":"","submissionUrl":"","title":"Discover Artificial Intelligence","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Discover Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"e0fc914e-34b8-43cd-8724-f1ffeec19e8d","owner":[],"postedDate":"May 13th, 2026","published":true,"recentEditorialEvents":[{"type":"reviewerAgreed","content":"56807945917763730258070686399444447136","date":"2026-05-11T13:23:40+00:00","index":64,"fulltext":""},{"type":"reviewersInvited","content":"40","date":"2026-05-05T11:17:19+00:00","index":"","fulltext":""}],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-05-13T10:04:37+00:00","versionOfRecord":[],"versionCreatedAt":"2026-05-13 10:04:36","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9385278","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9385278","identity":"rs-9385278","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.