MASA-RTNet: A Multimodal Adaptive-Stream-Attention Network for Real-Time Video Suspicious-Behaviour Detection

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Real–time video anomaly detection (VAD) must reconcile three competing goals: (i) state–of–the–art accuracy in one–class settings where true anomalies are unseen during training; (ii) sub–frame latency on resource-constrained edge devices; and (iii) robustness across the appearance, motion, and physical interaction cues that jointly characterise abnormal behaviour. We introduce MASA-RTNet, a Multimodal Adaptive Stream–Attention network that fuses appearance, optical flow, and a lightweight physics-informed graph branch inside a parameter-free attention gate. Two adaptive early-exit classifiers decide on-the-fly whether intermediate features already suffice for a confident verdict, yielding up to 5.6× average FLOP reduction. Trained exclusively on normal data with a curriculum of synthetic outliers, MASA-RTNet attains new state-of-the-art frame-level AUCs of 97.3 % on UCSD-Ped2, 87.9 % on CUHK Avenue, and 74.5 % on ShanghaiTech, while sustaining 37 fps (26.9 ms) on an NVIDIA Jet-son Xavier NX with only 5.8 M trainable parameters. Extensive ablations confirm that every modality and the MASA gate contribute meaningfully, and that INT8 quantisation plus structured pruning incur negligible (< 0.1 pp) accuracy loss. The full code, trained checkpoints, and reproducible Docker environment are released for the community. 1 We propose the EE-OneClass paradigm that marries early-exit inference with one-class energy modelling, enabling accurate VAD on edge GPUs (27 ms per 256 × 256 px frame on Jetson Xavier). 2 We design a tri-modal encoder that fuses appearance, motion and physics in a parameter-efficient manner (5.8 M trainable parameters), and show that each modality makes complementary contributions.
Full text 10,949 characters · extracted from preprint-html · click to expand
MASA-RTNet: A Multimodal Adaptive-Stream-Attention Network for Real-Time Video Suspicious-Behaviour Detection | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF correspondence MASA-RTNet: A Multimodal Adaptive-Stream-Attention Network for Real-Time Video Suspicious-Behaviour Detection Lucky Rajpoot, Rosy Madaan This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7362672/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Real–time video anomaly detection (VAD) must reconcile three competing goals: (i) state–of–the–art accuracy in one–class settings where true anomalies are unseen during training; (ii) sub–frame latency on resource-constrained edge devices; and (iii) robustness across the appearance, motion, and physical interaction cues that jointly characterise abnormal behaviour. We introduce MASA-RTNet, a Multimodal Adaptive Stream–Attention network that fuses appearance, optical flow, and a lightweight physics-informed graph branch inside a parameter-free attention gate. Two adaptive early-exit classifiers decide on-the-fly whether intermediate features already suffice for a confident verdict, yielding up to 5.6× average FLOP reduction. Trained exclusively on normal data with a curriculum of synthetic outliers, MASA-RTNet attains new state-of-the-art frame-level AUCs of 97.3 % on UCSD-Ped2, 87.9 % on CUHK Avenue, and 74.5 % on ShanghaiTech, while sustaining 37 fps (26.9 ms) on an NVIDIA Jet-son Xavier NX with only 5.8 M trainable parameters. Extensive ablations confirm that every modality and the MASA gate contribute meaningfully, and that INT8 quantisation plus structured pruning incur negligible (< 0.1 pp) accuracy loss. The full code, trained checkpoints, and reproducible Docker environment are released for the community. 1 We propose the EE-OneClass paradigm that marries early-exit inference with one-class energy modelling, enabling accurate VAD on edge GPUs (27 ms per 256 × 256 px frame on Jetson Xavier). 2 We design a tri-modal encoder that fuses appearance, motion and physics in a parameter-efficient manner (5.8 M trainable parameters), and show that each modality makes complementary contributions. video suspicious detection early-exit inference multimodal fusion energy-based model edge AI graph neural networks Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7362672","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"correspondence","associatedPublications":[],"authors":[{"id":503447935,"identity":"503a04a5-08b3-4c3d-a3dc-07e77929cb39","order_by":0,"name":"Lucky Rajpoot","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA5klEQVRIiWNgGAWjYBACPgYGZgYGAxsgwXwAyJeQIaiFDaylIg1IsCWAtPAQqeXMISCTxwAkQIQWieTHxrxtB9jNpXs+v7pRY8HDwH746Ab8WtKMk3nb7jBbzjm7zTrnGNBhPGlpN/BryWE+zNv2jNngRu424xw2oBYJHjNitBwGasl5Zpzzj0gtyTxnwFqYH+e2EaOF55mx4RxgIFvOSDNjzu2T4GEj5Bd+9uTHEm8MbJLNgUH3OedbnRw/++FjeLWAABMwLpINQI4E20tIOQgw/mBgsANqYf5AjOpRMApGwSgYeQAAmGdA9e6vohoAAAAASUVORK5CYII=","orcid":"","institution":"Manav Rachna International Institute of Research and Studies","correspondingAuthor":true,"prefix":"","firstName":"Lucky","middleName":"","lastName":"Rajpoot","suffix":""},{"id":503447937,"identity":"4d2b98a2-2f06-4fd5-8c06-7821757c1bc9","order_by":1,"name":"Rosy Madaan","email":"","orcid":"","institution":"Manav Rachna International Institute of Research and Studies","correspondingAuthor":false,"prefix":"","firstName":"Rosy","middleName":"","lastName":"Madaan","suffix":""}],"badges":[],"createdAt":"2025-08-13 08:38:21","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7362672/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7362672/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":107716437,"identity":"af1943a8-fe80-4833-821c-89ea1de599b2","added_by":"auto","created_at":"2026-04-24 10:06:36","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":3113501,"visible":true,"origin":"","legend":"","description":"","filename":"MASARTNet.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7362672/v1_covered_21dc86df-9969-433e-a3cf-8d075b9e8905.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"MASA-RTNet: A Multimodal Adaptive-Stream-Attention Network for Real-Time Video Suspicious-Behaviour Detection","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"video suspicious detection, early-exit inference, multimodal fusion, energy-based model, edge AI, graph neural networks","lastPublishedDoi":"10.21203/rs.3.rs-7362672/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7362672/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"Real–time video anomaly detection (VAD) must reconcile three competing goals: (i) state–of–the–art accuracy in one–class settings where true anomalies are unseen during training; (ii) sub–frame latency on resource-constrained edge devices; and (iii) robustness across the appearance, motion, and physical interaction cues that jointly characterise abnormal behaviour. We introduce MASA-RTNet, a Multimodal Adaptive Stream–Attention network that fuses appearance, optical flow, and a lightweight physics-informed graph branch inside a parameter-free attention gate. Two adaptive early-exit classifiers decide on-the-fly whether intermediate features already suffice for a confident verdict, yielding up to 5.6× average FLOP reduction. Trained exclusively on normal data with a curriculum of synthetic outliers, MASA-RTNet attains new state-of-the-art frame-level AUCs of 97.3 % on UCSD-Ped2, 87.9 % on CUHK Avenue, and 74.5 % on ShanghaiTech, while sustaining 37 fps (26.9 ms) on an NVIDIA Jet-son Xavier NX with only 5.8 M trainable parameters. Extensive ablations confirm that every modality and the MASA gate contribute meaningfully, and that INT8 quantisation plus structured pruning incur negligible (\u003c 0.1 pp) accuracy loss. The full code, trained checkpoints, and reproducible Docker environment are released for the community. 1 We propose the EE-OneClass paradigm that marries early-exit inference with one-class energy modelling, enabling accurate VAD on edge GPUs (27 ms per 256 × 256 px frame on Jetson Xavier). 2 We design a tri-modal encoder that fuses appearance, motion and physics in a parameter-efficient manner (5.8 M trainable parameters), and show that each modality makes complementary contributions.","manuscriptTitle":"MASA-RTNet: A Multimodal Adaptive-Stream-Attention Network for Real-Time Video Suspicious-Behaviour Detection","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-08-27 15:45:32","doi":"10.21203/rs.3.rs-7362672/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"ce145cc7-7ad8-456e-a62f-9cfb50bf4803","owner":[],"postedDate":"August 27th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2026-04-24T10:03:52+00:00","versionOfRecord":[],"versionCreatedAt":"2025-08-27 15:45:32","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7362672","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7362672","identity":"rs-7362672","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00