Multimodal Attention-Based Feature Fusion for Short Video Classification | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Multimodal Attention-Based Feature Fusion for Short Video Classification Calderon Sinclair, Ambrose Callahan, Sullivan Everly, Jessamine Beckett This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-3936566/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract In recent times, Short videos have emerged as a prevailing trend in social networks, prompting the rise of dedicated Short video social applications and a steady surge in Short video creation. Key attributes of Short videos encompass their concise length, rapid dissemination, straightforward production, broad category spectrum, and active user participation and interaction. In contrast to conventional videos, Short videos face a scarcity of dedicated datasets, and academic exploration into Short video content recognition remains limited. Short videos possess distinctive features necessitating specialized model design. While existing models leverage Short video datasets for training, the underlying model structures often adhere to generic architectures commonly employed in the broader field of behavior recognition. This highlights a substantial research gap in the realm of Short video content recognition, calling for the development of models tailored to the unique attributes of Short videos. The potential for innovation and exploration in this domain is significant, as existing models predominantly focus on specific tasks without fully harnessing the inherent characteristics of Short videos. Therefore, ample research opportunities exist to advance the field of Short video content recognition. This paper focuses on several key research objectives. Firstly, it delves into the exploration of extraction methods for visual, audio, and text features within videos. With the aim of addressing video classification tasks, a novel combinatorial network model is introduced. This model adeptly combines discrete features from each modality into comprehensive features across diverse modalities through the network. Information Retrieval and Management Short video Multimodal Fusion Attention Feature Extraction. Full Text Additional Declarations The authors declare no competing interests. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-3936566","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":271536607,"identity":"842ec053-b3ad-4395-b9e0-ba2210a4d04d","order_by":0,"name":"Calderon Sinclair","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABMklEQVRIie3RsUrEMACA4ZRAbomtY0oVXyFSKBYOfZWEQqdDh5tE0ZNCuxzeqtO9wk2ZHFoKdw59gMItVwXndlBOEDGti3At6iaSH5qWkI+2CQAq1Z8MjsjnXV4rQIDRi5KVfMZ6J9G+ECaJOc48WhP0LQENAYDmzGlmuogRJVd5ddc/3ou2kpKdH5xI4p8+Dw53EIDFQ75JSMYD9+bJH9JU9wibE/c2K+bLXeHJD0O2PdgkFPDQwnHKZxA7gCFC9QXzl6aAkmBktRGjCK03SaZBTd4JBTFzhqa47CZEvgVIMkol4SGh2/fM0SqRdhKSF4E7jn0+S7FN+DWh9SZbmlhgBNv/xZh4Sb6O+3w6yfbL8uWC1kdZvYqzI6MXFI8tpD2Im/Gny+u09W9Wq1Qq1X/vAzWMZGTN/QlTAAAAAElFTkSuQmCC","orcid":"","institution":"University of Wisconsin-Parkside","correspondingAuthor":true,"prefix":"","firstName":"Calderon","middleName":"","lastName":"Sinclair","suffix":""},{"id":271536608,"identity":"1514b65d-2077-4f89-8ad5-46b879caa3bb","order_by":1,"name":"Ambrose Callahan","email":"","orcid":"","institution":"University of Wisconsin-Parkside","correspondingAuthor":false,"prefix":"","firstName":"Ambrose","middleName":"","lastName":"Callahan","suffix":""},{"id":271536609,"identity":"f9dba5a5-cdaa-42b2-ad9f-12bb4394ce04","order_by":2,"name":"Sullivan Everly","email":"","orcid":"","institution":"University of Wisconsin-Parkside","correspondingAuthor":false,"prefix":"","firstName":"Sullivan","middleName":"","lastName":"Everly","suffix":""},{"id":271536610,"identity":"4cc26b66-e31d-4009-8d92-46160719e75d","order_by":3,"name":"Jessamine Beckett","email":"","orcid":"","institution":"University of Wisconsin-Parkside","correspondingAuthor":false,"prefix":"","firstName":"Jessamine","middleName":"","lastName":"Beckett","suffix":""}],"badges":[],"createdAt":"2024-02-07 10:51:39","currentVersionCode":1,"declarations":{"humanSubjects":false,"vertebrateSubjects":false,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":false,"humanSubjectConsent":false,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":false},"doi":"10.21203/rs.3.rs-3936566/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-3936566/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":50828343,"identity":"abe7ccdb-1781-43d2-8c8e-9d4d6ca1eec9","added_by":"auto","created_at":"2024-02-08 01:55:27","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":452467,"visible":true,"origin":"","legend":"","description":"","filename":"MultimodalAttentionBasedFeatureFusionforShortVideoClassification.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3936566/v1_covered_c03f2c2d-85f3-4def-97c3-eeab081833c9.pdf"}],"financialInterests":"The authors declare no competing interests.","formattedTitle":"\u003cp\u003eMultimodal Attention-Based Feature Fusion for Short Video Classification\u003c/p\u003e","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"University of Wisconsin–Parkside","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Short video, Multimodal, Fusion, Attention, Feature Extraction.","lastPublishedDoi":"10.21203/rs.3.rs-3936566/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-3936566/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eIn recent times, Short videos have emerged as a prevailing trend in social networks, prompting the rise of dedicated Short video social applications and a steady surge in Short video creation. Key attributes of Short videos encompass their concise length, rapid dissemination, straightforward production, broad category spectrum, and active user participation and interaction. In contrast to conventional videos, Short videos face a scarcity of dedicated datasets, and academic exploration into Short video content recognition remains limited. Short videos possess distinctive features necessitating specialized model design. While existing models leverage Short video datasets for training, the underlying model structures often adhere to generic architectures commonly employed in the broader field of behavior recognition. This highlights a substantial research gap in the realm of Short video content recognition, calling for the development of models tailored to the unique attributes of Short videos. The potential for innovation and exploration in this domain is significant, as existing models predominantly focus on specific tasks without fully harnessing the inherent characteristics of Short videos. Therefore, ample research opportunities exist to advance the field of Short video content recognition. This paper focuses on several key research objectives. Firstly, it delves into the exploration of extraction methods for visual, audio, and text features within videos. With the aim of addressing video classification tasks, a novel combinatorial network model is introduced. This model adeptly combines discrete features from each modality into comprehensive features across diverse modalities through the network.\u003c/p\u003e","manuscriptTitle":"Multimodal Attention-Based Feature Fusion for Short Video Classification","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-02-08 01:47:17","doi":"10.21203/rs.3.rs-3936566/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"c476afdb-9b94-4c63-9a24-867c1c06ffbf","owner":[],"postedDate":"February 8th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":28626766,"name":"Information Retrieval and Management"}],"tags":[],"updatedAt":"2024-02-08T01:47:17+00:00","versionOfRecord":[],"versionCreatedAt":"2024-02-08 01:47:17","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-3936566","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-3936566","identity":"rs-3936566","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.