REACT: Recognize Every Action Everywhere All At Once

preprint OA: closed
Full text JSON View at publisher
AI-generated deep summary by claude@2026-07, 2026-07-04 · read from full text

This paper studied Group Activity Recognition (GAR), proposing the REACT (Recognize Every Action Everywhere All At Once) transformer-inspired architecture to classify group actions by modeling actors’ interactions and spatiotemporal context, using a vision-language approach. The method uses a Vision-Language Encoder for integrated temporal, spatial, and multi-modal interaction modeling (including recovery of local information under sparsely sampled frames), an Action Decoder that retrieves bounding boxes while linking semantics to visuals, and an Actor Fusion Block that combines actor-specific data with textual features. Across extensive experiments, REACT reportedly outperformed state-of-the-art GAR methods with improved accuracy in recognizing and understanding group activities, and the work notes it is a preprint and not peer reviewed. This paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Abstract Group Activity Recognition (GAR) is a fundamental problem in computer vision, with diverse applications in sports video analysis, video surveillance, and social scene understanding. Unlike conventional action recognition, GAR aims to classify the actions of a group of individuals as a whole, requiring a deep understanding of their interactions and spatiotemporal relationships. To address the challenges in GAR, we present REACT (Recognize Every Action Everywhere All At Once), a novel architecture inspired by the transformer encoder-decoder model explicitly designed to model complex contextual relationships within videos, including multi-modality and spatio-temporal features.Our architecture features a cutting-edge Vision-Language Encoder block for integrated temporal, spatial, and multi-modal interaction modeling. This component efficiently encodes spatiotemporal interactions, even with sparsely sampled frames, and recovers essential local information. Our Action Decoder Block refines the joint understanding of text and video data, allowing us to precisely retrieve bounding boxes, enhancing the link between semantics and visual reality. At the core, our Actor Fusion Block orchestrates a fusion of actor-specific data and textual features, striking a balance between specificity and context.Our method outperforms state-of-the-art GAR approaches in extensive experiments, demonstrating superior accuracy in recognizing and understanding group activities. Our architecture's potential extends to diverse real-world applications, offering empirical evidence of its performance gains. This work significantly advances the field of group activity recognition, providing a robust framework for nuanced scene comprehension.
Full text 13,840 characters · extracted from preprint-html · click to expand
REACT: Recognize Every Action Everywhere All At Once | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article REACT: Recognize Every Action Everywhere All At Once Naga VS Raviteja Chappa, Pha Nguyen, Page Daniel Dobbs, Khoa Luu This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4109494/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 01 Jul, 2024 Read the published version in Machine Vision and Applications → Version 1 posted 10 You are reading this latest preprint version Abstract Group Activity Recognition (GAR) is a fundamental problem in computer vision, with diverse applications in sports video analysis, video surveillance, and social scene understanding. Unlike conventional action recognition, GAR aims to classify the actions of a group of individuals as a whole, requiring a deep understanding of their interactions and spatiotemporal relationships. To address the challenges in GAR, we present REACT (Recognize Every Action Everywhere All At Once), a novel architecture inspired by the transformer encoder-decoder model explicitly designed to model complex contextual relationships within videos, including multi-modality and spatio-temporal features.Our architecture features a cutting-edge Vision-Language Encoder block for integrated temporal, spatial, and multi-modal interaction modeling. This component efficiently encodes spatiotemporal interactions, even with sparsely sampled frames, and recovers essential local information. Our Action Decoder Block refines the joint understanding of text and video data, allowing us to precisely retrieve bounding boxes, enhancing the link between semantics and visual reality. At the core, our Actor Fusion Block orchestrates a fusion of actor-specific data and textual features, striking a balance between specificity and context.Our method outperforms state-of-the-art GAR approaches in extensive experiments, demonstrating superior accuracy in recognizing and understanding group activities. Our architecture's potential extends to diverse real-world applications, offering empirical evidence of its performance gains. This work significantly advances the field of group activity recognition, providing a robust framework for nuanced scene comprehension. Group Activity Recognition (GAR) Action Retrieval Vision-Language Modeling Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Published Journal Publication published 01 Jul, 2024 Read the published version in Machine Vision and Applications → Version 1 posted Editorial decision: Revision requested 07 May, 2024 Reviews received at journal 07 May, 2024 Reviews received at journal 23 Apr, 2024 Reviewers agreed at journal 15 Apr, 2024 Reviewers agreed at journal 02 Apr, 2024 Reviewers agreed at journal 18 Mar, 2024 Reviewers invited by journal 18 Mar, 2024 Editor assigned by journal 16 Mar, 2024 Submission checks completed at journal 16 Mar, 2024 First submitted to journal 15 Mar, 2024 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4109494","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":281063424,"identity":"7b7029de-8b8b-4d6a-bdea-8c7b6ab20fd0","order_by":0,"name":"Naga VS Raviteja Chappa","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAuUlEQVRIiWNgGAWjYFACHjApB8TMUEykFmPStSQ2EK3FnP3swc+FO+zSNxw/+9iAocIapBc/sOzJS5aeeSY5d8OZdOMEhjPphLUY3OAxkOZtY87dcCCN+QBj22GitBj/5m2rTzc4/wyo5R9xWsyAthxOMLiRxpzA2ECMljM5ZtYzzxw3nHnjGbNBwrF0Y8Jajp8xvl24o1qe73was8SHGmtZglpAgJkRpiyBGOWoWkbBKBgFo2AUYAMAzsE9LEuN0+IAAAAASUVORK5CYII=","orcid":"","institution":"University of Arkansas at Fayetteville","correspondingAuthor":true,"prefix":"","firstName":"Naga","middleName":"VS Raviteja","lastName":"Chappa","suffix":""},{"id":281063425,"identity":"ad285cbb-4b77-4959-b307-b2bae0435ffb","order_by":1,"name":"Pha Nguyen","email":"","orcid":"","institution":"University of Arkansas at Fayetteville","correspondingAuthor":false,"prefix":"","firstName":"Pha","middleName":"","lastName":"Nguyen","suffix":""},{"id":281063426,"identity":"d681cb19-8570-45be-aa61-228a4b9ca697","order_by":2,"name":"Page Daniel Dobbs","email":"","orcid":"","institution":"University of Arkansas at Fayetteville","correspondingAuthor":false,"prefix":"","firstName":"Page","middleName":"Daniel","lastName":"Dobbs","suffix":""},{"id":281063427,"identity":"d62330b3-a98b-40c2-8826-e8fccf889c7a","order_by":3,"name":"Khoa Luu","email":"","orcid":"","institution":"University of Arkansas at Fayetteville","correspondingAuthor":false,"prefix":"","firstName":"Khoa","middleName":"","lastName":"Luu","suffix":""}],"badges":[],"createdAt":"2024-03-15 17:14:16","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4109494/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4109494/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1007/s00138-024-01561-z","type":"published","date":"2024-07-01T19:40:12+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":61525068,"identity":"7fd7cbed-7bc5-4b82-8378-b548df95d79e","added_by":"auto","created_at":"2024-07-31 19:40:18","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1218453,"visible":true,"origin":"","legend":"","description":"","filename":"ISVCREACT.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4109494/v1_covered_c9f6c547-7b22-4cec-ac19-77599376092e.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"REACT: Recognize Every Action Everywhere All At Once","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"machine-vision-and-applications","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"mvap","sideBox":"Learn more about [Machine Vision and Applications](https://www.springer.com/journal/138)","snPcode":"138","submissionUrl":"https://submission.springernature.com/new-submission/138/3","title":"Machine Vision and Applications","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Group Activity Recognition (GAR), Action Retrieval, Vision-Language Modeling","lastPublishedDoi":"10.21203/rs.3.rs-4109494/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4109494/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"Group Activity Recognition (GAR) is a fundamental problem in computer vision, with diverse applications in sports video analysis, video surveillance, and social scene understanding. Unlike conventional action recognition, GAR aims to classify the actions of a group of individuals as a whole, requiring a deep understanding of their interactions and spatiotemporal relationships. To address the challenges in GAR, we present REACT (Recognize Every Action Everywhere All At Once), a novel architecture inspired by the transformer encoder-decoder model explicitly designed to model complex contextual relationships within videos, including multi-modality and spatio-temporal features.Our architecture features a cutting-edge Vision-Language Encoder block for integrated temporal, spatial, and multi-modal interaction modeling. This component efficiently encodes spatiotemporal interactions, even with sparsely sampled frames, and recovers essential local information. Our Action Decoder Block refines the joint understanding of text and video data, allowing us to precisely retrieve bounding boxes, enhancing the link between semantics and visual reality. At the core, our Actor Fusion Block orchestrates a fusion of actor-specific data and textual features, striking a balance between specificity and context.Our method outperforms state-of-the-art GAR approaches in extensive experiments, demonstrating superior accuracy in recognizing and understanding group activities. Our architecture's potential extends to diverse real-world applications, offering empirical evidence of its performance gains. This work significantly advances the field of group activity recognition, providing a robust framework for nuanced scene comprehension.","manuscriptTitle":"REACT: Recognize Every Action Everywhere All At Once","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-03-19 04:26:17","doi":"10.21203/rs.3.rs-4109494/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2024-05-07T15:10:19+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2024-05-07T11:36:58+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2024-04-23T12:56:35+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"65ee7cac-6a78-484e-ae6f-9638ee05fb7d","date":"2024-04-16T01:35:11+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"bbbed472-758a-423a-be3f-1099a0e9bc08","date":"2024-04-02T15:54:56+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"99543d58-4643-49a1-a6db-ff168cc883f8","date":"2024-03-19T02:17:47+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2024-03-18T16:44:52+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2024-03-16T19:34:27+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2024-03-16T11:01:56+00:00","index":"","fulltext":""},{"type":"submitted","content":"Machine Vision and Applications","date":"2024-03-15T17:01:02+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"machine-vision-and-applications","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"mvap","sideBox":"Learn more about [Machine Vision and Applications](https://www.springer.com/journal/138)","snPcode":"138","submissionUrl":"https://submission.springernature.com/new-submission/138/3","title":"Machine Vision and Applications","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"9037bf3c-302f-4e19-827e-0e0448a2a210","owner":[],"postedDate":"March 19th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2024-07-31T19:40:12+00:00","versionOfRecord":{"articleIdentity":"rs-4109494","link":"https://doi.org/10.1007/s00138-024-01561-z","journal":{"identity":"machine-vision-and-applications","isVorOnly":false,"title":"Machine Vision and Applications"},"publishedOn":"2024-07-01 19:40:12","publishedOnDateReadable":"July 1st, 2024"},"versionCreatedAt":"2024-03-19 04:26:17","video":"","vorDoi":"10.1007/s00138-024-01561-z","vorDoiUrl":"https://doi.org/10.1007/s00138-024-01561-z","workflowStages":[]},"version":"v1","identity":"rs-4109494","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4109494","identity":"rs-4109494","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00