Human Action Recognition Using YOLOv11 Ultralytics: A Comprehensive Study for Real-Time Applications

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract Human action recognition (HAR) is a pivotal task in computer vision, with applications in surveillance, healthcare, robotics, and human-computer interaction. This study presents a novel framework for HAR using the YOLOv11 model by Ultralyt ics, a state-of-the-art object detection architecture optimized for real-time performance. We trained and evaluated the model on a custom dataset comprising 18 distinct human actions, captured in indoor environments using fisheye cameras. The actions range from everyday activities (e.g., walking, sitting) to specialized tasks (e.g., patient on stretcher, patient on wheelchair). Our results show that YOLOv11 achieves a mean Average Precision ([email protected]) of 0.401, with exceptional performance on actions like ”cleaning” ([email protected]: 0.760), ”searching” ([email protected]: 0.695), and ”patient on wheelchair” ([email protected]: 0.995). We provide an in-depth analysis of the model’s training metrics, bounding box distributions, precision-recall curves, F1-confidence curves, recall-confidence curves, and confusion matrices. Additionally, we present extensive qualitative results to demonstrate the model’s robustness in real-world scenarios. A comparison with existing methods, such as two-stream CNNs and Transformer-based models, highlights YOLOv11’s superior balance of accuracy and speed, making it a promising solution for real-time HAR applications. This study also discusses the model’s limitations and outlines directions for future research, paving the way for enhanced action recognition systems.
Full text 13,120 characters · extracted from preprint-html · click to expand
Human Action Recognition Using YOLOv11 Ultralytics: A Comprehensive Study for Real-Time Applications | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Human Action Recognition Using YOLOv11 Ultralytics: A Comprehensive Study for Real-Time Applications Seung Jin Kim This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6393539/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Human action recognition (HAR) is a pivotal task in computer vision, with applications in surveillance, healthcare, robotics, and human-computer interaction. This study presents a novel framework for HAR using the YOLOv11 model by Ultralyt ics, a state-of-the-art object detection architecture optimized for real-time performance. We trained and evaluated the model on a custom dataset comprising 18 distinct human actions, captured in indoor environments using fisheye cameras. The actions range from everyday activities (e.g., walking, sitting) to specialized tasks (e.g., patient on stretcher, patient on wheelchair). Our results show that YOLOv11 achieves a mean Average Precision ( [email protected] ) of 0.401, with exceptional performance on actions like ”cleaning” ( [email protected] : 0.760), ”searching” ( [email protected] : 0.695), and ”patient on wheelchair” ( [email protected] : 0.995). We provide an in-depth analysis of the model’s training metrics, bounding box distributions, precision-recall curves, F1-confidence curves, recall-confidence curves, and confusion matrices. Additionally, we present extensive qualitative results to demonstrate the model’s robustness in real-world scenarios. A comparison with existing methods, such as two-stream CNNs and Transformer-based models, highlights YOLOv11’s superior balance of accuracy and speed, making it a promising solution for real-time HAR applications. This study also discusses the model’s limitations and outlines directions for future research, paving the way for enhanced action recognition systems. Artificial Intelligence and Machine Learning Human Action Recognition YOLOv11 Ultralytics Deep Learning Computer Vision Real-time Detection Surveillance Healthcare Full Text Additional Declarations The authors declare no competing interests. Ethics Statement: This study was approved by the Institutional Review Board (IRB) of Assist University, Seoul, Korea (Approval Number: AU-IRB-2023-015). The IRB reviewed the study protocol, including the data collection process and usage of video footage for human action recognition, and determined that it adhered to ethical standards for research involving human subjects. Consent Statement: All participants (or their legal guardians, where applicable) provided informed consent for the collection and use of video footage in this study. The consent process included a clear explanation of the study's purpose, the nature of the data collection, and the intended use of the data for research and publication. For cases involving vulnerable populations (e.g., patients in healthcare settings), the need for individual consent was waived by the Assist University IRB, as the data was anonymized and posed minimal risk to participants. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6393539","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":439476693,"identity":"b5522b4f-841f-491d-a40d-8856a10b57b9","order_by":0,"name":"Seung Jin Kim","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABKklEQVRIie3RMUvDQBjG8SccpIMXsl4IjV/hSsBS9MNcENIlQ90UtAQC6SSuCn6IiiCOlUCzRFwLLoUOLg7NlkIEE1OzJFZHh/svL8nlB+8RQCb7jzHwahIoy2IqfvkkqkPyGyG8+PSb8D8QQGU1wQ7SN4Mp0scLGJM992yTRV2/E6/WyzyHPpkR+7RJBrfzkXKTxDCJ9vBKRWT71DtgTsjBEkGcpEn4wuNEC+ewSgIROT48FY5fLLYAefJ/IB9bcpKVRH9brUXOsb+LKOH512KgJWGCM6Hy4gjEaSGDa3ekXIYzagTavUndoR2y9/IuNu0lTtBrIX12PMUmHFvs5fkuzY4Ou1f6cJVmuWVZcRQZbYsBnQyIaP1G3U6K+h81SNm47Uwmk8lkVZ+xWGNE4AbvqQAAAABJRU5ErkJggg==","orcid":"","institution":"AI Convergence Engineering Assist University Seoul, Korea","correspondingAuthor":true,"prefix":"","firstName":"Seung","middleName":"Jin","lastName":"Kim","suffix":""}],"badges":[],"createdAt":"2025-04-07 11:21:25","currentVersionCode":1,"declarations":{"humanSubjects":false,"vertebrateSubjects":true,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":false,"humanSubjectConsent":false,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":true},"doi":"10.21203/rs.3.rs-6393539/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6393539/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":80525630,"identity":"076fb0d2-f092-4e53-a9d4-aa4d83eae138","added_by":"auto","created_at":"2025-04-14 09:53:23","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":3617675,"visible":true,"origin":"","legend":"","description":"","filename":"f77f22a27df841c389a05d4e21decc8b.revised.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6393539/v1_covered_de3bb5e6-0d92-4328-aa4c-b55d1cf383f1.pdf"}],"financialInterests":"\u003cp\u003eThe authors declare no competing interests.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;Ethics Statement:\u003c/p\u003e\n\u003cp\u003eThis study was approved by the Institutional Review Board (IRB) of Assist University, Seoul, Korea (Approval Number: AU-IRB-2023-015). The IRB reviewed the study protocol, including the data collection process and usage of video footage for human action recognition, and determined that it adhered to ethical standards for research involving human subjects.\u003c/p\u003e\n\u003cp\u003eConsent Statement:\u003c/p\u003e\n\u003cp\u003eAll participants (or their legal guardians, where applicable) provided informed consent for the collection and use of video footage in this study. The consent process included a clear explanation of the study's purpose, the nature of the data collection, and the intended use of the data for research and publication. For cases involving vulnerable populations (e.g., patients in healthcare settings), the need for individual consent was waived by the Assist University IRB, as the data was anonymized and posed minimal risk to participants.\u003c/p\u003e","formattedTitle":"\u003cp\u003eHuman Action Recognition Using YOLOv11 Ultralytics: A Comprehensive Study for Real-Time Applications\u003c/p\u003e","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"assist university","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Human Action Recognition, YOLOv11, Ultralytics, Deep Learning, Computer Vision, Real-time Detection, Surveillance, Healthcare","lastPublishedDoi":"10.21203/rs.3.rs-6393539/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6393539/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eHuman action recognition (HAR) is a pivotal task \u0026nbsp;in computer vision, with applications in surveillance, healthcare, \u0026nbsp;robotics, and human-computer interaction. This study presents a \u0026nbsp;novel framework for HAR using the YOLOv11 model by Ultralyt ics, a state-of-the-art object detection architecture optimized for \u0026nbsp;real-time performance. We trained and evaluated the model on a \u0026nbsp;custom dataset comprising 18 distinct human actions, captured \u0026nbsp;in indoor environments using fisheye cameras. The actions range \u0026nbsp;from everyday activities (e.g., walking, sitting) to specialized \u0026nbsp;tasks (e.g., patient on stretcher, patient on wheelchair). Our \u0026nbsp;results show that YOLOv11 achieves a mean Average Precision \u0026nbsp;([email protected]) of 0.401, with exceptional performance on actions like \u0026nbsp;”cleaning” ([email protected]: 0.760), ”searching” ([email protected]: 0.695), \u0026nbsp;and ”patient on wheelchair” ([email protected]: 0.995). We provide \u0026nbsp;an in-depth analysis of the model’s training metrics, bounding \u0026nbsp;box distributions, precision-recall curves, F1-confidence curves, \u0026nbsp;recall-confidence curves, and confusion matrices. Additionally, we \u0026nbsp;present extensive qualitative results to demonstrate the model’s \u0026nbsp;robustness in real-world scenarios. A comparison with existing \u0026nbsp;methods, such as two-stream CNNs and Transformer-based \u0026nbsp;models, highlights YOLOv11’s superior balance of accuracy \u0026nbsp;and speed, making it a promising solution for real-time HAR \u0026nbsp;applications. This study also discusses the model’s limitations \u0026nbsp;and outlines directions for future research, paving the way for \u0026nbsp;enhanced action recognition systems.\u003c/p\u003e","manuscriptTitle":"Human Action Recognition Using YOLOv11 Ultralytics: A Comprehensive Study for Real-Time Applications","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-04-14 09:45:14","doi":"10.21203/rs.3.rs-6393539/v1","editorialEvents":[{"type":"communityComments","content":14}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"ae92ab0b-0ad3-4344-8eae-4c5e0f9e18e6","owner":[],"postedDate":"April 14th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":46785651,"name":"Artificial Intelligence and Machine Learning"}],"tags":[],"updatedAt":"2025-06-05T03:23:22+00:00","versionOfRecord":[],"versionCreatedAt":"2025-04-14 09:45:14","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6393539","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6393539","identity":"rs-6393539","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-24T02:00:01.246996+00:00
License: CC-BY-4.0