VMSG: a video caption network based on multimodal semantic grouping and semantic attention | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article VMSG: a video caption network based on multimodal semantic grouping and semantic attention Xin Yang, Xiangchen Wang, Xiaohui Ye, Tao Li This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-1542723/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 13 Jun, 2023 Read the published version in Multimedia Systems → Version 1 posted 9 You are reading this latest preprint version Abstract Network video typically contains a variety of information that is used by the video caption model to generate video tags. The process of creating video captions is divided into two steps: video information extraction and natural language generation. Existing models have the problem of redundant information in continuous frames when generating natural language, which affects the accuracy of the caption. As a result, this paper proposes a Multimodal Semantic Grouping and Semantic Attention Video Caption Model (VMSG). VMSG uses a novel semantic grouping method for decoding, which divides the video with the same semantics into a semantic group for decoding and predicting the next word, to reduce the redundant information of continuous video frames, which differs from the decoding mode of grouping by frame. Because the importance of each semantic group varies, we investigate a semantic attention mechanism to add weight to the semantic group and use a single-layer LSTM to simplify the model. Experiments show that VMSG outperforms some state-of-the-art models in terms of caption generation performance and alleviates the problem of redundant information in continuous video frames. Video caption multimodal semantic grouping semantic attention Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Published Journal Publication published 13 Jun, 2023 Read the published version in Multimedia Systems → Version 1 posted Editorial decision: Major revision 16 Apr, 2023 Reviews received at journal 24 Mar, 2023 Reviewers agreed at journal 17 Mar, 2023 Reviewers agreed at journal 03 May, 2022 Reviewers agreed at journal 02 May, 2022 Reviewers invited by journal 02 May, 2022 Editor assigned by journal 26 Apr, 2022 Submission checks completed at journal 12 Apr, 2022 First submitted to journal 10 Apr, 2022 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-1542723","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":98079085,"identity":"9468dd3d-4225-4da3-9567-62d677dc1d11","order_by":0,"name":"Xin Yang","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA20lEQVRIiWNgGAWjYNACAwY5KIuZeC3GpGphYEhsIFqLwfGzh1/zFNxJ3y6R/PADQ4V1YgP72QP4tZzJS7OcYfAsd+eMNGMJhjPpiQ08eQn4tRzIMTP4YHA4d8ONBAMJxrbDiQ0SPAb4tZx/Y2aQYHA43eBG+ucfjP+I0XIjx/gB0JYEIMNMgrGBCC2SN96YMc4wOGy44cybMouEY+nGbTw5+LXwnc8x/szz57C8wfH0zTc+1FjL9rOfwa9F4QADmwSclwDEbHjVA4F8AwPzB0KKRsEoGAWjYIQDAAvvSZDgbzb8AAAAAElFTkSuQmCC","orcid":"","institution":"Nanjing University of Aeronautics and Astronautics","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Xin","middleName":"","lastName":"Yang","suffix":""},{"id":98079086,"identity":"9eac6fa8-40da-4fc8-9ffd-77e8243032d5","order_by":1,"name":"Xiangchen Wang","email":"","orcid":"","institution":"Nanjing University of Aeronautics and Astronautics","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Xiangchen","middleName":"","lastName":"Wang","suffix":""},{"id":98079087,"identity":"625afcc5-1119-44cc-94d2-974cbdd094c3","order_by":2,"name":"Xiaohui Ye","email":"","orcid":"","institution":"Nanjing University of Aeronautics and Astronautics","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Xiaohui","middleName":"","lastName":"Ye","suffix":""},{"id":98079088,"identity":"625f7609-3262-4386-beb2-61d6f5ae560a","order_by":3,"name":"Tao Li","email":"","orcid":"","institution":"Nanjing University of Aeronautics and Astronautics","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Tao","middleName":"","lastName":"Li","suffix":""}],"badges":[],"createdAt":"2022-04-10 11:59:03","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-1542723/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-1542723/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1007/s00530-023-01124-8","type":"published","date":"2023-06-13T21:12:56+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":20354980,"identity":"5f3883d2-bb69-44d7-9ec4-67ef56a9071f","added_by":"auto","created_at":"2022-04-14 16:50:56","extension":"pdf","order_by":2,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":829676,"visible":true,"origin":"","legend":"","description":"","filename":"Manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-1542723/v1_covered.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"VMSG: a video caption network based on multimodal semantic grouping and semantic attention","fulltext":[{"header":"Full Text","content":"This preprint is available for \u003ca href='/article/rs-1542723/latest.pdf' target='_blank'\u003edownload as a PDF\u003c/a\u003e."}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"multimedia-systems","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"mmsj","sideBox":"Learn more about [Multimedia Systems](http://link.springer.com/journal/530)","snPcode":"530","submissionUrl":"https://submission.nature.com/new-submission/530/3","title":"Multimedia Systems","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Video caption, multimodal, semantic grouping, semantic attention","lastPublishedDoi":"10.21203/rs.3.rs-1542723/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-1542723/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eNetwork video typically contains a variety of information that is used by the video caption model to generate video tags. The process of creating video captions is divided into two steps: video information extraction and natural language generation. Existing models have the problem of redundant information in continuous frames when generating natural language, which affects the accuracy of the caption. As a result, this paper proposes a Multimodal Semantic Grouping and Semantic Attention Video Caption Model (VMSG). VMSG uses a novel semantic grouping method for decoding, which divides the video with the same semantics into a semantic group for decoding and predicting the next word, to reduce the redundant information of continuous video frames, which differs from the decoding mode of grouping by frame. Because the importance of each semantic group varies, we investigate a semantic attention mechanism to add weight to the semantic group and use a single-layer LSTM to simplify the model. Experiments show that VMSG outperforms some state-of-the-art models in terms of caption generation performance and alleviates the problem of redundant information in continuous video frames.\u003c/p\u003e","manuscriptTitle":"VMSG: a video caption network based on multimodal semantic grouping and semantic attention","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2022-04-14 16:50:50","doi":"10.21203/rs.3.rs-1542723/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Major revision","date":"2023-04-16T20:17:25+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2023-03-24T07:49:51+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"4339d122-15d4-4034-b6ef-e67b556063d1","date":"2023-03-17T05:55:41+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"62529046-2b93-413a-bf9d-260afe9327b7","date":"2022-05-04T01:27:25+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"f709db11-522b-464e-9688-05147622c376","date":"2022-05-03T03:01:56+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2022-05-02T11:16:11+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2022-04-27T02:00:17+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2022-04-12T11:34:44+00:00","index":"","fulltext":""},{"type":"submitted","content":"Multimedia Systems","date":"2022-04-10T11:44:09+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"multimedia-systems","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"mmsj","sideBox":"Learn more about [Multimedia Systems](http://link.springer.com/journal/530)","snPcode":"530","submissionUrl":"https://submission.nature.com/new-submission/530/3","title":"Multimedia Systems","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"34e4f296-56a2-4dc7-82ba-e036d5bb0304","owner":[],"postedDate":"April 14th, 2022","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2023-10-16T21:39:37+00:00","versionOfRecord":{"articleIdentity":"rs-1542723","link":"https://doi.org/10.1007/s00530-023-01124-8","journal":{"identity":"multimedia-systems","isVorOnly":false,"title":"Multimedia Systems"},"publishedOn":"2023-06-13 21:12:56","publishedOnDateReadable":"June 13th, 2023"},"versionCreatedAt":"2022-04-14 16:50:50","video":"","vorDoi":"10.1007/s00530-023-01124-8","vorDoiUrl":"https://doi.org/10.1007/s00530-023-01124-8","workflowStages":[]},"version":"v1","identity":"rs-1542723","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-1542723","identity":"rs-1542723","version":["v1"]},"buildId":"rHA-KDH7Qsr4HCuvH75dn","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.