REACT: Recognize Every Action Everywhere All At Once | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article REACT: Recognize Every Action Everywhere All At Once Naga VS Raviteja Chappa, Pha Nguyen, Page Daniel Dobbs, Khoa Luu This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4109494/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 01 Jul, 2024 Read the published version in Machine Vision and Applications → Version 1 posted 10 You are reading this latest preprint version Abstract Group Activity Recognition (GAR) is a fundamental problem in computer vision, with diverse applications in sports video analysis, video surveillance, and social scene understanding. Unlike conventional action recognition, GAR aims to classify the actions of a group of individuals as a whole, requiring a deep understanding of their interactions and spatiotemporal relationships. To address the challenges in GAR, we present REACT (Recognize Every Action Everywhere All At Once), a novel architecture inspired by the transformer encoder-decoder model explicitly designed to model complex contextual relationships within videos, including multi-modality and spatio-temporal features.Our architecture features a cutting-edge Vision-Language Encoder block for integrated temporal, spatial, and multi-modal interaction modeling. This component efficiently encodes spatiotemporal interactions, even with sparsely sampled frames, and recovers essential local information. Our Action Decoder Block refines the joint understanding of text and video data, allowing us to precisely retrieve bounding boxes, enhancing the link between semantics and visual reality. At the core, our Actor Fusion Block orchestrates a fusion of actor-specific data and textual features, striking a balance between specificity and context.Our method outperforms state-of-the-art GAR approaches in extensive experiments, demonstrating superior accuracy in recognizing and understanding group activities. Our architecture's potential extends to diverse real-world applications, offering empirical evidence of its performance gains. This work significantly advances the field of group activity recognition, providing a robust framework for nuanced scene comprehension. Group Activity Recognition (GAR) Action Retrieval Vision-Language Modeling Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Published Journal Publication published 01 Jul, 2024 Read the published version in Machine Vision and Applications → Version 1 posted Editorial decision: Revision requested 07 May, 2024 Reviews received at journal 07 May, 2024 Reviews received at journal 23 Apr, 2024 Reviewers agreed at journal 15 Apr, 2024 Reviewers agreed at journal 02 Apr, 2024 Reviewers agreed at journal 18 Mar, 2024 Reviewers invited by journal 18 Mar, 2024 Editor assigned by journal 16 Mar, 2024 Submission checks completed at journal 16 Mar, 2024 First submitted to journal 15 Mar, 2024 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4109494","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":281063424,"identity":"7b7029de-8b8b-4d6a-bdea-8c7b6ab20fd0","order_by":0,"name":"Naga VS Raviteja Chappa","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAuUlEQVRIiWNgGAWjYFACHjApB8TMUEykFmPStSQ2EK3FnP3swc+FO+zSNxw/+9iAocIapBc/sOzJS5aeeSY5d8OZdOMEhjPphLUY3OAxkOZtY87dcCCN+QBj22GitBj/5m2rTzc4/wyo5R9xWsyAthxOMLiRxpzA2ECMljM5ZtYzzxw3nHnjGbNBwrF0Y8Jajp8xvl24o1qe73was8SHGmtZglpAgJkRpiyBGOWoWkbBKBgFo2AUYAMAzsE9LEuN0+IAAAAASUVORK5CYII=","orcid":"","institution":"University of Arkansas at Fayetteville","correspondingAuthor":true,"prefix":"","firstName":"Naga","middleName":"VS Raviteja","lastName":"Chappa","suffix":""},{"id":281063425,"identity":"ad285cbb-4b77-4959-b307-b2bae0435ffb","order_by":1,"name":"Pha Nguyen","email":"","orcid":"","institution":"University of Arkansas at Fayetteville","correspondingAuthor":false,"prefix":"","firstName":"Pha","middleName":"","lastName":"Nguyen","suffix":""},{"id":281063426,"identity":"d681cb19-8570-45be-aa61-228a4b9ca697","order_by":2,"name":"Page Daniel Dobbs","email":"","orcid":"","institution":"University of Arkansas at Fayetteville","correspondingAuthor":false,"prefix":"","firstName":"Page","middleName":"Daniel","lastName":"Dobbs","suffix":""},{"id":281063427,"identity":"d62330b3-a98b-40c2-8826-e8fccf889c7a","order_by":3,"name":"Khoa Luu","email":"","orcid":"","institution":"University of Arkansas at Fayetteville","correspondingAuthor":false,"prefix":"","firstName":"Khoa","middleName":"","lastName":"Luu","suffix":""}],"badges":[],"createdAt":"2024-03-15 17:14:16","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4109494/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4109494/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1007/s00138-024-01561-z","type":"published","date":"2024-07-01T19:40:12+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":61525068,"identity":"7fd7cbed-7bc5-4b82-8378-b548df95d79e","added_by":"auto","created_at":"2024-07-31 19:40:18","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1218453,"visible":true,"origin":"","legend":"","description":"","filename":"ISVCREACT.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4109494/v1_covered_c9f6c547-7b22-4cec-ac19-77599376092e.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"REACT: Recognize Every Action Everywhere All At Once","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"machine-vision-and-applications","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"mvap","sideBox":"Learn more about [Machine Vision and Applications](https://www.springer.com/journal/138)","snPcode":"138","submissionUrl":"https://submission.springernature.com/new-submission/138/3","title":"Machine Vision and Applications","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Group Activity Recognition (GAR), Action Retrieval, Vision-Language Modeling","lastPublishedDoi":"10.21203/rs.3.rs-4109494/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4109494/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"Group Activity Recognition (GAR) is a fundamental problem in computer vision, with diverse applications in sports video analysis, video surveillance, and social scene understanding. Unlike conventional action recognition, GAR aims to classify the actions of a group of individuals as a whole, requiring a deep understanding of their interactions and spatiotemporal relationships. To address the challenges in GAR, we present REACT (Recognize Every Action Everywhere All At Once), a novel architecture inspired by the transformer encoder-decoder model explicitly designed to model complex contextual relationships within videos, including multi-modality and spatio-temporal features.Our architecture features a cutting-edge Vision-Language Encoder block for integrated temporal, spatial, and multi-modal interaction modeling. This component efficiently encodes spatiotemporal interactions, even with sparsely sampled frames, and recovers essential local information. Our Action Decoder Block refines the joint understanding of text and video data, allowing us to precisely retrieve bounding boxes, enhancing the link between semantics and visual reality. At the core, our Actor Fusion Block orchestrates a fusion of actor-specific data and textual features, striking a balance between specificity and context.Our method outperforms state-of-the-art GAR approaches in extensive experiments, demonstrating superior accuracy in recognizing and understanding group activities. Our architecture's potential extends to diverse real-world applications, offering empirical evidence of its performance gains. This work significantly advances the field of group activity recognition, providing a robust framework for nuanced scene comprehension.","manuscriptTitle":"REACT: Recognize Every Action Everywhere All At Once","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-03-19 04:26:17","doi":"10.21203/rs.3.rs-4109494/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2024-05-07T15:10:19+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2024-05-07T11:36:58+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2024-04-23T12:56:35+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"65ee7cac-6a78-484e-ae6f-9638ee05fb7d","date":"2024-04-16T01:35:11+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"bbbed472-758a-423a-be3f-1099a0e9bc08","date":"2024-04-02T15:54:56+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"99543d58-4643-49a1-a6db-ff168cc883f8","date":"2024-03-19T02:17:47+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2024-03-18T16:44:52+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2024-03-16T19:34:27+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2024-03-16T11:01:56+00:00","index":"","fulltext":""},{"type":"submitted","content":"Machine Vision and Applications","date":"2024-03-15T17:01:02+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"machine-vision-and-applications","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"mvap","sideBox":"Learn more about [Machine Vision and Applications](https://www.springer.com/journal/138)","snPcode":"138","submissionUrl":"https://submission.springernature.com/new-submission/138/3","title":"Machine Vision and Applications","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"9037bf3c-302f-4e19-827e-0e0448a2a210","owner":[],"postedDate":"March 19th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2024-07-31T19:40:12+00:00","versionOfRecord":{"articleIdentity":"rs-4109494","link":"https://doi.org/10.1007/s00138-024-01561-z","journal":{"identity":"machine-vision-and-applications","isVorOnly":false,"title":"Machine Vision and Applications"},"publishedOn":"2024-07-01 19:40:12","publishedOnDateReadable":"July 1st, 2024"},"versionCreatedAt":"2024-03-19 04:26:17","video":"","vorDoi":"10.1007/s00138-024-01561-z","vorDoiUrl":"https://doi.org/10.1007/s00138-024-01561-z","workflowStages":[]},"version":"v1","identity":"rs-4109494","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4109494","identity":"rs-4109494","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.